$ 60 entries match — showing 1–24 · sort: Free credits ↓
388 routers & providers · 813 models tracked · $3.8k free credits
- Official free tierAPI keydown4322msupdated 23h agoconsole.mistral.aisource · mistral.ai/pricingoverview
Mistral's La Plateforme is a first-party EU-based API with per-1M-token pricing, a 50% batch discount and up to 90% cached-input savings. The free plan includes $10/month of API credits plus Vibe and Studio access; Pro ($14.99/mo) raises that to $30/month. Open-weight models are available for self-hosting (Apache 2.0 for research/individual; commercial use requires a Mistral license).
free quotaCommunity reported approximate monthly value of the Experiment free tier for developers.
free modelsMistral Small 4Mistral Medium 3.5Mixtral 8x22B InstructMistral Large 2407Ministral 3B/8BPixtral+5see console.mistral.ai’s free quota / risk review → - Official free tierAPI keyup966msupdated 23h agotoken.sensenova.cnsource · sensenova.ai/token-planoverview
token.sensenova.cn is the official Token Plan for SenseTime's SenseNova multimodal AI platform. During its public beta, it offers generous free quotas for models like SenseNova 6.8 Flash Lite and U1 Fast, suitable for complex office workflows.
free quotaFree public beta (extended to end of July): 1,500 calls/5h per model (verified: 60k points/5h on token-plan page); ~7,200 calls/day ~ $200/mo equivalent at ~1K tokens/call. Phone signup, up to 20 API keys.
free modelsSenseNova 6.7 Flash-LiteSenseNova U1 FastDeepSeek V4 FlashGLM 5.2see token.sensenova.cn’s free quota / risk review → - Official free tierAPI keyup303msupdated 23h agoconsole.groq.comsource · console.groq.com/docs/rate-limitsoverview
Groq runs LPU inference hardware tuned for raw tokens-per-second. The API offers a documented free developer tier — per-model rate limits such as 30 RPM / 1,000 RPD / 8K TPM on gpt-oss-120b — with full x-ratelimit-* header observability, and paid Developer plans that add Batch and Flex processing plus higher ceilings.
free quotaCommunity reported approximate monthly value of free tier rate limits for individual developers.
free modelsQwen familyWhisper Large v3Llama-3.1-8B-InstructLlama 3.3 70B Instructsee console.groq.com’s free quota / risk review → - Official free tierAPI keyunknownupdated 23d agopioneer.aisource · pioneer.ai/pricingoverview
Pioneer AI, by Fastino Labs, provides an inference API for specialized tasks like data extraction and classification. It offers seat-based subscriptions with included platform credits and priority support.
free quota$75 free usage credits — no credit card required
free models— none trackedsee pioneer.ai’s free quota / risk review → - Official free tierAPI keyup72msupdated 23h agoapi.cerebras.aisource · cerebras.ai/pricingoverview
Cerebras Inference offers ultra-fast AI inference powered by its Wafer-Scale Engine (WSE-3) technology, achieving record-breaking speeds up to 3000 tokens/s. The official API is OpenAI-compatible and features a generous free tier alongside professional pay-as-you-go access.
free quotaFree tier ≈1M tokens/day (resets UTC 00:00, no card, no waitlist) → ≈$30/mo blended at $1/1M; 8K context cap; RPM reported 5–30 by different sources. Paid dev tier from $10.
free modelsLlama 3.1/3.3 familyQwen 3 family (per official model list)Llama 4 Scoutgpt-oss-120bsee api.cerebras.ai’s free quota / risk review → - Official free tierAPI keyupupdated 13d agomodal.comsource · modal.com/pricingoverview
Modal is serverless compute for AI workloads rather than a token-billed model API: you pay per second for CPU, memory and GPU time (H100 SXM5 $0.001097/sec, B200 $0.001736/sec, CPU $0.0000131 per core per second) and never for idle resources. The Starter plan is $0 plus compute and includes $30 per month of free compute with 3 seats, 100 containers and 10 GPU concurrency; Team is $250 per month with $100 per month of free compute. Volumes include 1 TiB per month free, then $0.09/GiB per month.
free quotaStarter plan includes $30 per month of free compute, stated on the Modal pricing page (Team includes $100 per month on the $250 plan); academics can apply separately for up to $10,000 in credits.
free modelsNo fixed free-model list (self-deploy any supported model, billed by compute)see modal.com’s free quota / risk review → - Official free tierAPI keyup1934msupdated 23h agoapi.z.aisource · z.ai/pricingoverview
Z.ai is Zhipu AI's international developer platform, offering access to the GLM (General Language Model) family. It provides a high-performance alternative to Western models, with the GLM-Flash series offered permanently for free to developers to encourage global adoption of its flagship reasoning models.
free quotaGLM Flash models (4.7/4.5 text, 4.6V vision) listed $0 on official pricing — ongoing free models, not trial credits; ~1 QPS and ~1K calls/day (third-party, unofficial) → ≈$30/mo blended. Flagship GLM-5.x paid.
free modelsGLM-4.5-FlashGLM-4.6V-FlashGLM-4.7-Flashsee api.z.ai’s free quota / risk review → - Official free tierAPI keyup159msupdated 23h agoapi.x.aisource · x.ai/apioverview
xAI (founded by Elon Musk) offers the official API for the Grok series of LLMs. Known for leading benchmarks in coding and reasoning, the Grok-4.6 flagship model is available via an OpenAI-compatible API, featuring low-latency inference and high agentic tool-calling capabilities.
free quotaNo permanent free tier; ~$25 promo credit on signup (30-day expiry, varies by region/promotion). $150/mo data-sharing credits still listed by 2026 guides despite mid-2025 end reports — rely on what the console shows.
free modelsGrok 4.20see api.x.ai’s free quota / risk review → - Official free tierAPI keyup288msupdated 23h agoai.google.devsource · ai.google.dev/gemini-api/docs/pricingoverview
Google AI Studio gives free API access to Gemini models — free input and output tokens under per-model rate limits, not a depleting credit. Current-generation Flash models (including Gemini 3.8 Flash) are free-tier eligible. Paid tiers unlock higher limits, context caching, the 50%-off Batch API, and the most advanced models on the same API; free-tier data is used by Google to improve its products, paid-tier data is not.
free quotaEstimated based on high free-tier rate limits (RPD) vs paid input/output rates for equivalent usage.
free modelsGemma 4 26B A4B Gemma 4 31BGemini 2.5 Flash LiteGemini 3.1 Flash Lite PreviewGemma 2 27BGemini 3.8 Flash+2see ai.google.dev’s free quota / risk review → - Official free tierAPI keyup80msupdated 23h agolightning.aisource · lightning.aioverview
Lightning AI (LitAI) provides a unified gateway to frontier models through an OpenAI-compatible API. New users receive 40 million free tokens (approx. $15 in credits) to test models like Claude, GPT, and Gemini. Usage beyond the initial credit is billed as pay-as-you-go.
free quota~30M free tokens per user per month (~15 credits = $15 value), no card required; credits refresh monthly and expire if unused; plus 1 free 4-CPU Studio and 10GB Drive.
free modelsgpt-5 / gpt-5.5 familygemini-3.5-flash / gemini-2.5-pronemotron-3-ultra-550bclaude-opus-4-8claude-fable-5deepseek-v4-pro+1see lightning.ai’s free quota / risk review → - Official free tierAPI keyupupdated 13d agonlpcloud.comsource · nlpcloud.com/pricing.htmloverview
NLP Cloud is an API platform offering pre-trained and custom models for summarization, classification, embeddings, ASR, speech synthesis and large language models such as GPT-OSS 120B, LLaMA 3.1 405B and ChatDolphin. Its pay-as-you-go plan has no fixed monthly cost and automatically grants a $15 free credit, then bills per request at $0.003 on CPU or $0.005 on GPU with token surcharges for the largest models. Published prepaid plans run from Starter at $29 per month (10 parallel requests) to the Large Language Models plan at $2,499 per month (50 parallel requests), with dedicated fine-tuning plans from $399 per month.
free quotaOfficial pricing page states that the pay-as-you-go plan automatically gives a $15 FREE credit with no fixed monthly cost (nlpcloud.com/pricing.html, checked 2026-09-16).
free modelsLlama 3.1 405Bgpt-oss-120bLlama 3.3 70B Instructsee nlpcloud.com’s free quota / risk review → - Official free tierAPI keyup85msupdated 23h agoapi.ai21.comsource · docs.ai21.com/docs/usage-costoverview
AI21 Labs sells API access to its Jamba foundation models through AI21 Studio on usage-based pricing, with a self-serve free trial and a custom plan for volume. The free trial is published as $10 in credits for 7 days with no credit card required. Published pay-as-you-go rates are Jamba Mini at $0.2 per 1M input tokens and $0.4 per 1M output tokens, and Jamba Large at $2 per 1M input tokens and $8 per 1M output tokens. AI21 states that on Studio models an average token equals about 1 word or 6 characters of English text.
free quotaOfficial pricing page states a Free Trial of $10 credits for 7 days, no credit card needed (ai21.com/pricing, checked 2026-09-16).
free modelsJamba Large (1.7)Jamba Mini (1.7)Jamba 1.6 Largesee api.ai21.com’s free quota / risk review → - Official free tierAPI keyupupdated 13d agoconsole.upstage.aisource · upstage.ai/blog/en/guide-1-upstage-console-apioverview
Upstage Console provides an OpenAI-compatible API for their high-performance Solar models. It features a developer-friendly $10 signup credit and commitment-based tiers for businesses requiring higher rate limits and support.
free quotaOne-time $10 credit valid for 3 months for new signups.
free modelsSolar ProSolar MiniSolar Pro 3see console.upstage.ai’s free quota / risk review → - Official free tierAPI keyup133msupdated 23h agoapi.inceptionlabs.aisource · inceptionlabs.aioverview
Inception Labs offers a cloud inference platform for its in-house diffusion-based LLMs, such as Mercury 2 and Mercury Coder. Designed for extreme efficiency, it provides an OpenAI-compatible API that significantly undercuts GPU-based inference costs while maintaining high reasoning quality.
free quotaFree plan ≈10M tokens, officially described as ~$7–10 of credit (low end counted), no card; self-serve key at platform.inceptionlabs.ai; then PAYG (Mercury 2 ~$0.25/1M in, $0.75–1.00/1M out).
free modelsmercury (general chat diffusion model)mercury-coder (code, with fim/completions)mercury-2see api.inceptionlabs.ai’s free quota / risk review → - Official free tierAPI keyupupdated 13d agocerebras.aisource · cerebras.ai/pricingoverview
Cerebras Inference serves open-weight models on wafer-scale CS-4 hardware — gpt-oss-120b at ~3,000 tokens/s for $0.35/$0.75 per 1M in/out, Qwen 3.8 27B at $0.99/$1.49. The self-serve developer tier is pay-as-you-go with $5 of free credit to start; the enterprise tier (production capacity, priority limits, fine-tuning) is quote-based. Also available via AWS Marketplace, OpenRouter, Hugging Face and Vercel.
free quotaSelf-serve developer tier includes $5 free credit to start (pricing page, checked 2026-09-16). Replaces an older $30 figure that no longer matches the official page.
free modelsgpt-oss-120bQwen3 32BLlama 4 Scoutzai-glm-4.7see cerebras.ai’s free quota / risk review → - Official free tierAPI keyup1082msupdated 23h agochat.intern-ai.org.cnsource · internlm.intern-ai.org.cn/api/documentoverview
InternLM (Shanghai AI Laboratory) provides an open platform with an OpenAI-compatible API. It offers a monthly free token allowance for its InternVL and long-thinking reasoning models. Details were checked against the provider's own page at internlm.intern-ai.org.cn.
free quotaCommunity users: ~1M input + 3M output tokens free per month (~10 RPM, keys valid 6 months) ≈ $5/mo at $0.5 in / $1.5 out per 1M; monthly usage visible in console under API Usage.
free modelsintern-latestintern-s1intern-s1-miniintern-s1-prointernvl3.5-latestsee chat.intern-ai.org.cn’s free quota / risk review → - Official free tierAPI keyup281msupdated 23h agoinference.api.nscale.comsource · nscale.com/services/ai-servicesoverview
Nscale's Serverless Inference is an OpenAI-compatible endpoint service for open models - the catalogue pages list Meta Llama 4 Scout, Alibaba Qwen, DeepSeek, OpenAI GPT OSS 20B/120B, Mistral 8x22B Instruct, FLUX.1 [schnell] and Stable Diffusion XL Base 1.0 - with token-based pay-as-you-go billing. Per Nscale's own launch materials, pricing is based on input and output tokens for text and vision models, while image generation is priced on output image dimensions. Per-token rates are not published on a static official page: the docs direct users to the console's AI Services to Models list or the /v1/models API endpoint for pricing and context length. Every new user can claim $5 of free credit, and Nscale states the platform does not log or train on request or response content.
free quotaEvery new user can claim $5 of free credit, stated in Nscale's Serverless Inference launch post (2 April 2025) and in its announcement that pricing is based on input and output tokens for text and vision models.
free modelsLlama familyDeepSeek familygpt-oss-120bsee inference.api.nscale.com’s free quota / risk review → - Official free tierAPI keyup95msupdated 23h agovercel.comsource · vercel.com/docs/ai-gateway/pricingoverview
Vercel AI Gateway is a unified API gateway that allows routing to hundreds of models from various providers through a single OpenAI-compatible endpoint. It provides observability, budgets, and zero-data-retention (ZDR) options with no token markup over upstream list prices.
free quotaEach team account gets $5 of AI Gateway credit every 30 days to try any model; after your first payment you become a paid customer and stop receiving the free credit. Official figure.
free modelsAccess to OpenAI/Anthropic/Google/Meta/DeepSeek and more via the gateway (at upstream list prices)see vercel.com’s free quota / risk review → - Official free tierAPI keyunknownupdated 23d agoinference.dahl.globalsource · inference.dahl.globaloverview
Dahl Inference is a high-performance platform offering decentralized access to open models like DeepSeek and MiniMax. It features an OpenAI-compatible API and a generous signup promotion of 100 million free tokens.
free quotaValue based on 100M free tokens at an estimated rate of $0.03/1M.
free modelsLlama-3.1-8B-InstructQwen2.5 72B Instructsee inference.dahl.global’s free quota / risk review → - Official free tierAPI keyup1579msupdated 23h agos.qiniu.comsource · qiniu.com/ai/chatoverview
Qiniu Cloud AI LLM Inference is a managed MaaS platform that provides a unified, compliant API for over 50 mainstream LLMs. It is compatible with both OpenAI and Anthropic protocols, offering new users substantial free tokens and flexible resource packages for production scaling.
free quotaFirst API key activates a shared 3M-token free resource pack (~$3 at $1/1M; auto-granted ~1h after first billable usage), then PAYG; several open-source models are permanently fully free (bonus, unquantified).
free modelsDeepSeek-V4 seriesQwenGLM (Zhipu)Longcat (Meituan)ArceeNemotron (NVIDIA)see s.qiniu.com’s free quota / risk review → - Official free tierAPI keyup1273msupdated 23h agoapi.siliconflow.cnsource · siliconflow.cn/pricingoverview
SiliconFlow (SiliconCloud) is a premier Chinese AI model platform offering high-performance inference for a wide range of open-source models (Qwen, Llama, DeepSeek). It is notable for providing free, permanent API access to models with parameters under 9B, alongside competitive pay-as-you-go rates for larger flagship
free quotaSignup grants ~CNY 14 (≈$2.08) credit, plus ~CNY 16 (≈$2.38) voucher after real-name verification; many <=9B open-source models (Qwen2.5-7B, Llama3.1-8B, etc.) are permanently free; direct access within mainland China. Since 2026-05-15, unverified users are capped at 10 RPM and delinquent/unverified accounts are cut off.
free modelsQwen2.5-7B-InstructLlama-3.1-8B-Instructsee api.siliconflow.cn’s free quota / risk review → - Official free tierAPI keyup2280msupdated 23h agoplatform.moonshot.cnsource · platform.moonshot.cn/pricingoverview
platform.moonshot.cn is the China-domain entry point for Moonshot AI's Kimi platform, and it serves the same documentation tree as platform.kimi.com - the fetched pricing page points its documentation index at platform.kimi.com/docs/llms.txt, so model IDs and parameters are identical across the two domains. Prices are quoted in CNY on the China-facing site, which lists K3 at 20.00 CNY input / 2.00 CNY cached input / 100.00 CNY output per 1M tokens, K2.7 Code at 6.50 / 1.30 / 27.00 and K2.6 at 6.50 / 1.10 / 27.00, while the international list quotes K3 at $3.00 / $0.30 / $15.00 per 1M tokens. As on the parent platform, billing is prepaid token usage and the only documented free grant is the 15 CNY verification voucher, which excludes K3.
free quotaSame account system as platform.kimi.com: a 15 CNY voucher is granted after identity verification (platform.kimi.com/docs/guide/account-and-payments), converted at roughly 7.1 CNY/USD, and cannot be used on Kimi K3.
free modelsmoonshot-v1-8kmoonshot-v1-32kmoonshot-v1-128kkimi-k2see platform.moonshot.cn’s free quota / risk review → - Official free tierAPI keyup84msupdated 23h agoapi.cohere.aisource · cohere.com/pricingoverview
Cohere's API splits keys into a free Trial tier (auto-created at sign-up: 1,000 API calls/month, 20 RPM on chat models, no commercial use) and an application-gated Production tier (500 RPM chat, pay-as-you-go, invoiced monthly or at $250 outstanding). Legacy Command models price from $0.30–$2.50 per 1M in; dedicated Model Vault instances run $4–$10/hr.
free quotaTrial key (auto-created at sign-up): 1,000 calls/month + 20 RPM chat (docs.cohere.com rate-limits, verified 2026-09-16); no production/commercial use. Dollar value is an estimate.
free modelsCommand R (08-2024)Command R+ (08-2024)Command R+Command R7B (and all Cohere models — trial and production have the same model access)Command Asee api.cohere.ai’s free quota / risk review → - Official free tierAPI keyunknownupdated 23d agobytez.comsource · bytez.com/pricingoverview
Bytez is an AI model hosting platform that offers an OpenAI-compatible API for a wide range of models. It provides $1 in free credits that refresh every 4 weeks and uses a low-cost pay-as-you-go pricing model.
free quota$1 free credits, refreshes every 4 weeks
free modelsLlama-3.1-8B-InstructAnthropic: Claude 3.5 SonnetGPT-4osee bytez.com’s free quota / risk review →