Free AI APIs for Developers (2026): Real Rate Limits

Free AI APIs for Developers (2026): Real Rate Limits — illustration

Every AI company offers a free tier. Most developer guides list them without mentioning the actual limits. Here’s the honest breakdown — what you get for free, when you’ll hit the wall, and what the upgrade costs.

If you want the short answer: start on Gemini 3.8 Flash. It’s on Google’s free tier with input listed as “Free of charge”, though the actual limits now only show up in your own AI Studio console. Use Groq’s Free plan with openai/gpt-oss-120b as the fallback, at 30 RPM, 1,000 requests a day and 200K tokens a day. Anthropic gives new users only “a small amount” of credit with no figure attached, OpenAI’s GPT-4.1 mini isn’t on its Free tier and needs $5 paid first, and Together AI doesn’t offer free trials at all.

Last verified: 26 September 2026.

What changed on 26 September: Google now limits Gemini 2.5 to past users, so the Gemini section is rebuilt around 3.8 Flash. Groq decommissioned its two compound models on 21 September. And xAI now points developers to Grok 4.7 instead of 4.6. Older changes are in the log at the bottom.

The Complete Free Tier Comparison

ProviderFree accessRate limitVolume capBest model availableSource (checked 2026-09-26)
Google GeminiFree tier, input “Free of charge”No longer publishedNo longer publishedGemini 3.8 Flash (2.5 Flash now restricted to past users)rate limits, changelog
GroqFree plan, $030 RPM, 8K tokens/min1,000 req/day, 200K tokens/daygpt-oss-120brate limits
Anthropic Claude”A small amount” of credit, no figure givenNo published free-tier RPMStart tier: $500/month spend capClaude Fable 5.1pricing
OpenAIFree tier exists, but not for GPT-4.1 mini500 RPM at Tier 1, after $5 paid; 200K tokens/min10,000 req/dayGPT-4.1 minigpt-4.1-mini
Together AINone. No free trialsn/a$5 minimum purchase to use at allLlama 3.3 70Bbilling docs
HuggingFace$0.10/month creditn/a$0.10/monthMostly CPU-scale modelspricing
ReplicateSelect models free to run, then billing required6 req/min (granted credit, no card on file); 600 req/min otherwisePay per predictionStable Diffusion, Fluxrate limits
xAI / SpaceXAI (Grok)Signup credit not confirmable150 RPS (T0, Grok 4.7)50M tokens/min (T0)Grok 4.7rate limits
MistralFree mode, limits shown only in your consoleNo longer publishedNo longer publishedMistral Small 4API pricing
CohereFree trial key20 req/min (Chat API)1,000 API calls/monthCommand A+rate limits

Tier 1: Best Free Tiers (Actually Usable)

Google Gemini API

Still the most generous free tier I’ve used. It’s also the one that got harder to write about honestly, and this month it moved out from under the code sample that used to sit here.

Correction, 26 September: don’t start a new project on Gemini 2.5 Flash. Google’s changelog entry for 18 September says “we are limiting access to the 2.5 models to users who have actively used them in the past.” It’s careful to add that they “are not deprecated and will continue to be served until further notice through the API,” and the deprecations page lists Gemini 2.5 Flash as “Access restricted to past users” with no shutdown date. So if 2.5 Flash is already running in your app, it keeps running. If you’re new, the model this page used to call the sensible default, and the gemini-2.5-flash string the old sample used, may simply not be open to you. The changelog doesn’t say what “limiting access” looks like from a new account’s side, whether that’s an error or just worse service, and I’m not going to guess. Google’s instruction for everyone else is short: “For any new projects, use our latest models: 3.5 Flash-Lite or 3.8 Flash.”

Google no longer publishes free-tier RPM, TPM or requests-per-day in a per-model table. The rate limits page now says your limits “depend on a variety of factors (such as your usage tier) and can be viewed in Google AI Studio,” and the link to view them requires a Google login. The “15 RPM, 1M tokens/day” figure you’ll still find quoted around the web is no longer confirmable from a primary source. Check your own console. That’s the only number that applies to your account anyway.

Gemini 3.8 Flash went GA on 2 September, and Google describes it as “our most intelligent Flash model.” It’s also the model Google’s own quickstart now uses, through a different SDK pattern from the genai.configure() / GenerativeModel() one this page used to show. Here’s the quickstart’s sample exactly as Google prints it, prompt included:

from google import genai

client = genai.Client()

interaction = client.interactions.create(
    model="gemini-3.8-flash",
    input="Explain how AI works in a few words"
)
print(interaction.output_text)

I haven’t edited a character of it. Google calls this the Interactions API, and the quickstart walks through the rest, so read that page before you wire it into anything.

Free limit: input is listed as “Free of charge” on the free tier. The request ceiling is console-only now. Paid tier (3.8 Flash, standard): $0.75 per 1M input and $3.75 per 1M output through 31 December 2026, then $1.50/$7.50 from 1 January 2027 Best for: Prototyping, high-volume batch processing, long-context tasks

Google’s pricing page lists that exact price, and the same 1 January jump, for 3.7 Flash and 3.6 Flash too. Three generations at one price is odd enough that I checked the page twice. Same answer both times. I’d still look at the rendered page yourself before you budget a year out on it. Correction (15 August): this page once printed $1.50/$7.50 as today’s Flash price. That’s the 2027 price, twice what you’d pay right now.

If you’re choosing between the Flash models, 3.7 Flash (GA 13 August) is already labelled “previous-generation” on Google’s models page, six weeks after launch. Take 3.8.

Groq — Fastest Inference

Groq runs open-source models on custom LPU hardware. Its console lists two separate plans, a Free plan at $0 and a paid, pay-per-token Developer plan. The numbers below are the Free plan’s, and that Free plan is still the one entry in this table that’s both genuinely free and genuinely usable. The Developer plan is what you upgrade to for higher limits plus Batch and Flex processing.

from groq import Groq

client = Groq(api_key="YOUR_API_KEY")
response = client.chat.completions.create(
    model="openai/gpt-oss-120b",
    messages=[{"role": "user", "content": "Write a Python function to validate email"}]
)
print(response.choices[0].message.content)

Free limit (Free plan, openai/gpt-oss-120b): 30 RPM, 8,000 tokens/min, 1,000 requests/day, 200,000 tokens/day Paid tier: Developer plan, pay per token Best for: Real-time applications, chatbots, any use case where latency matters

The daily ceilings are the ones nobody notices until they run a batch job, and then they notice hard. 200K tokens/day is maybe eighty decent-sized prompts with long answers. Budget accordingly.

Correction, 21 August: llama-3.3-70b-versatile is gone. Groq shut it down on 16 August 2026, along with llama-3.1-8b-instant, the day after my 15 August check found it listed as active production, and its deprecations page names openai/gpt-oss-120b as the replacement (and openai/gpt-oss-20b for 8B Instant). The model’s own docs page still shows it with an “Enterprise” badge and no deprecation banner, but the deprecations list and the Free plan table both say it’s gone for free-plan users, and I’ve dropped the old limits this page printed for it rather than compare against numbers I can’t re-check, so re-tune any retry logic against the four above.

Correction, 26 September: groq/compound and groq/compound-mini are gone. This section used to warn that they had no tokens-per-day figure but a hard 250 requests/day; Groq announced their deprecation on 24 August and decommissioned both on 21 September, after which requests to those model IDs return errors. The same page shows qwen/qwen3.6-27b, the Preview model Groq had also offered as a Llama 3.3 70B replacement, shut down on 14 September; its successor qwen/qwen3.8-27b is still a Preview model, on the same 30 RPM / 1,000 requests a day / 8,000 tokens a minute / 200,000 tokens a day limits as gpt-oss-120b.

Tier 2: Credits That Aren’t What You Think

Anthropic Claude API

Claude is still what I reach for when the code has to be right. The free-tier story changed shape completely, though.

import anthropic

client = anthropic.Anthropic(api_key="YOUR_API_KEY")
message = client.messages.create(
    model="claude-sonnet-5",
    max_tokens=1024,
    messages=[{"role": "user", "content": "Write a React hook for debounced search"}]
)
print(message.content[0].text)

The model ID matters. This page previously used claude-sonnet-4-20250514, which is now listed as retired except on Bedrock and Google Cloud — it will error against the first-party API. And the table said “Claude 3.5 Sonnet,” which no longer appears on Anthropic’s current or legacy model tables — it survives only in the deprecation history, retired 28 October 2025. The correct current string is claude-sonnet-5. If you copied the old snippet, that’s your bug.

Price — and the expiry date on it is gone: Claude Sonnet 5 is $2 per MTok input and $10 per MTok output, and as of August 2026 that’s no longer introductory. Anthropic’s pricing page now says it plainly: “The $2/$10 per million input/output token pricing for Claude Sonnet 5, announced at launch as introductory pricing through August 31, 2026, is now the standard price. The previously scheduled increase to $3/$15 per million input/output tokens on September 1, 2026 will not occur.” If you budgeted for the September jump this page previously warned about, you can unbudget it — re-checked against that page on 26 September 2026.

Free credit: Anthropic’s pricing FAQ now says only that “new users receive a small amount of free credits to test the API.” No dollar figure. The old $5 number is still repeated everywhere, but it’s not on Anthropic’s page anymore, so I’m not printing it as fact.

Rate limits: there is no numbered free tier. Anthropic publishes usage tiers keyed to a monthly spend cap. The lowest documented one (Start) gives Sonnet 5 a 1,000 RPM ceiling, 2,000,000 input tokens/min and 400,000 output tokens/min against a $500/month cap. New organisations may land in an “Evaluation tier” with lower limits that Anthropic doesn’t publish.

If you want Claude for daily coding, the API isn’t the play — Claude Code is included in Claude Pro, which is $17/month billed annually or $20/month billed monthly, and that is far more cost-effective than metering an agentic session. I’ve written up whether the Claude API is free in more detail.

OpenAI API

GPT-4.1 mini has no Free-tier access. None. OpenAI’s model page starts its rate-limit table at Tier 1, and Tier 1 requires $5 paid, not $5 given. Once you’ve paid it, you get 500 RPM, 10,000 requests/day and 200,000 tokens/min.

There is a genuine no-payment Free tier at OpenAI, restricted to allowed geographies with a $100/month usage limit. GPT-4.1 mini just isn’t on it.

from openai import OpenAI

client = OpenAI(api_key="YOUR_API_KEY")
response = client.chat.completions.create(
    model="gpt-4.1-mini",
    messages=[{"role": "user", "content": "Generate a SQL schema for a blog"}]
)
print(response.choices[0].message.content)

Cost to start: $5, paid Paid tier: GPT-4.1 mini at $0.40/$1.60 per 1M tokens, cached input $0.10 Best for: General-purpose tasks, function calling, structured outputs

That $0.40/$1.60 was the single number in the big August sweep that verified as still exactly correct. Everything else moved.

One housekeeping note: OpenAI’s docs moved from platform.openai.com to developers.openai.com. Old links 301 to the new host, but if you’ve got docs URLs in a runbook, update them.

Together AI

Together is still widely listed as giving new accounts “$5 free credit.” It doesn’t. Together’s own billing docs say: “Together AI does not currently offer free trials.”

There is no credit. There’s a $5 minimum purchase required before you can use the platform at all. It’s still a good service. Llama 3.3 70B serverless runs $1.04 per 1M tokens in both directions, and you get dozens of open models behind one API without touching a GPU. But it does not belong in a free-tier article as anything other than a correction, so I’ve pulled it from the free stack below.

xAI / SpaceXAI (Grok API)

The widely-repeated claim is a $25 signup credit plus $150/month for opting into data sharing. I could not confirm either figure. I checked xAI’s billing FAQ, its console billing docs, the rate-limits page, the models page, the release notes and the docs home, most recently on 26 September. None of them mentions a signup credit or a data-sharing credit, and none mentions $25 in any context. The only purchase figure xAI documents is on the console billing page, which sets the top-up amount at “(minimum $5)”. This page previously said the top-up minimum was $25, and I’d guessed that was where the rumour started. That $25 isn’t on the page now, so the guess goes with it. The promo may well exist. But xAI doesn’t document it, and I’m not stating it as fact because affiliate sites copied each other.

What xAI does publish is rate limits in requests per second. The entry tier (T0) on Grok 4.7 is 150 RPS and 50M tokens/min. xAI derives that per-second figure from your per-minute budget — its docs state the per-second limit is RPM divided by 60, so you can’t spend a whole minute’s requests in one burst.

A naming note: xAI was absorbed into SpaceX and rebranded SpaceXAI on 6 July 2026, and the docs are switching over unevenly. As of 26 September the release notes describe Grok 4.7 as “SpaceXAI’s frontier model for coding, agentic tasks, and knowledge work” while still saying “xAI API” and “xAI console” elsewhere on the same page. The billing page, the billing FAQ and the rate-limits page still say xAI and nothing else. The models page carries neither name, and the domain is still docs.x.ai. I use both names here so this page turns up whichever one you searched.

The lineup keeps moving. “Grok 3 mini” is gone, and the flagship has changed twice since August. Grok 4.6 took over between my 10 and 15 August checks; as of 26 September xAI’s models page points past it to Grok 4.7: “For everything else, including code, use Grok 4.7. It is the most capable model we’ve built.” Grok 4.7 costs the same as Grok 4.6, $2.00/$6.00 per 1M tokens under 200K context and $4.00/$12.00 at or over it. Grok 4.6 carries no deprecation notice; it has simply stopped being the one xAI recommends.

The rest of the lineup is Grok 4.6, Grok 4.5, Grok 4.3, the Grok 4.20 family, and Grok Build 0.1. Grok Build 0.1 is the cheapest at $1.00/$2.00 per 1M tokens under 200K context.

Tier 3: Specialized Free Tiers

HuggingFace Inference

Restructured, and not in your favour. The old rate-limited-but-free serverless Inference API has been folded into “Inference Providers,” and free accounts now get $0.10 per month in credit — a figure HuggingFace’s own table marks “subject to change”. PRO accounts get $2.00/month, Team and Enterprise get $2.00 per seat per month.

Ten cents a month is not a free tier for large models. It’s a taste. This page previously showed a code sample hitting meta-llama/Llama-3.3-70B-Instruct through the old serverless endpoint, and I’ve removed it rather than leave a snippet that will burn your entire monthly allowance in a handful of calls. HuggingFace’s own docs note that since July 2025 the hf-inference provider focuses mostly on CPU inference: embeddings, text ranking, text classification, smaller LLMs. Use it for what it’s now scoped to, and read the pricing page before you wire anything up.

Replicate

Pay-per-prediction pricing. Correction (15 August): this page previously said Replicate offers “trial credits” — its pricing page publishes no credit figure of any kind, and what the billing docs actually say is “You can run select models on Replicate for free, but after a bit you’ll be asked to set up billing.” So what you can count on is free runs of selected models. Replicate does grant credit to some accounts — its rate-limits page refers to accounts that “have been granted credit and don’t have a payment method on file” — but it never says how much, or who gets it.

Billing is per-second for hosted models and per-token for the LLMs it fronts.

Correction (29 August): this section used to warn you off Replicate’s Claude listing, and that warning is now wrong. Until recently the model behind it was anthropic/claude-3.7-sonnet at $3.00/$15.00 per million tokens — an old model at a markup. That listing is gone (its page 404s, and claude no longer appears anywhere on Replicate’s pricing page), and Replicate now fronts nine Anthropic models, including Claude Sonnet 5, Fable 5, Sonnet 4.6 and Opus 4.7. Claude Sonnet 5 there is $2 per million input tokens and $10 per million output — exactly what Anthropic charges first-party. The markup is gone along with the stale model, so routing Claude through Replicate to keep one billing relationship no longer costs you anything on the token price. Opus 4.7’s Replicate page publishes no price at all, so check that one in your own console before you rely on it.

Best for: Image generation, audio models, niche open-source models

Mistral

Mistral still has a free API tier, but it no longer publishes the numbers for it. Its help centre says “Free mode (default) has the lowest limits, intended for evaluation and prototyping”, and points you to the Limits page of your Admin panel for the actual figures. Neither mistral.ai/pricing/api nor the models docs carries an RPM or tokens-per-day figure for it. So the “1 RPM, 500K tokens/day” figure still circulating for Mistral’s free tier is unconfirmable from anything Mistral publishes.

The current small model is specifically Mistral Small 4 (v26.03), a hybrid instruct/reasoning/coding model. Mistral’s docs don’t designate a flagship: Medium 3.5 (v26.04) is described only as “our frontier-class multimodal model optimized for agentic and coding use cases”, and the same page also lists Mistral Large 3 (v25.12), “a state-of-the-art, open-weight, general-purpose multimodal model”. Either way, Small is not the model Mistral points you at first.

Cohere

Two numbers get conflated constantly here, and they’re separate things: the trial-key rate limit is 20 requests/minute on the Chat API, and the volume cap is 1,000 API calls per month. A rate and a quota, not one figure.

Cohere splits its trial limits across two tables. The Chat API table lists nine models and puts every one of them at 20 req/min. The other-endpoints table has eight rows, and four of them are far tighter than the Chat number anyone quotes:

Trial-key endpointLimit
Audio Transcriptions5 req/min
Embed (Images)5 inputs/min
EmbedJob5 req/min
Rerank10 req/min
Tokenize100 req/min
Embed2,000 inputs/min
Parse500 req/min
Default (anything not covered above)500 req/min

Cohere doesn’t publish a token ceiling for trial keys separately from the 1,000-call cap — the page states request-rate and input-count limits only, with no tokens-per-minute or tokens-per-month figure for trial keys anywhere on it.

Command R+ is no longer where Cohere’s lineup ends. Cohere’s models doc describes Command A+ (command-a-plus-05-2026) as “the last model in the Command A family, while being Cohere’s first Mixture of Experts model”. Note that Cohere does not call it the flagship; the strongest ranking language on that page is “our most performant model to date”, and it applies that to Command A, not Command A+. It’s on a trial key — the rate-limit table’s first row is Command A+ at 20 req/min — and Cohere’s changelog says it “is now available for all Cohere users through our standard API endpoints”. Here’s the odd part: Cohere’s public pricing page doesn’t list a price for Command A+ at all. Its price table stops at Command R+ 08-2024 at $2.50/$10.00 per 1M. So Cohere ships a model it won’t quote you a rate for, which usually means enterprise sales.

The Free Stack: Running AI for $0/Month

The honest version of this stack is shorter than it was in April.

LayerProviderWhy
Primary LLMGemini 3.8 Flash (free tier)Input free of charge, check your console for limits; 2.5 Flash is restricted to past users
Fast inferenceGroq (free tier)30 RPM, but watch the 1,000 req/day cap
Fallback LLMCohere trial key20 req/min, 1,000 calls/month
EmbeddingsHuggingFace hf-inferenceNow CPU-scoped, and only $0.10/month of credit
Image generationReplicate (select models free)Until you’re asked to set up billing

Together AI is out of this table this month for the reason above. HuggingFace stays only because embeddings are one of the few workloads that still fit inside ten cents, and even then I’d treat it as a stopgap. If you want more genuinely free infrastructure around this, free developer tool credits are still holding up better than the API tiers are.

Code Example: Smart Fallback Chain

async def generate_response(prompt: str) -> str:
    providers = [
        ("gemini", call_gemini),   # free tier, limits console-only
        ("groq", call_groq),       # 30 RPM, 1,000 req/day
        ("cohere", call_cohere),   # 20 req/min, 1,000 calls/month
    ]

    for name, call_fn in providers:
        try:
            return await call_fn(prompt)
        except RateLimitError:
            logger.warning(f"{name} rate limited, falling back")
            continue

    raise Exception("All providers exhausted")

This pattern is how production AI apps work behind the scenes. Primary provider handles most requests. Fallbacks catch the rest. Log which provider served each request, because when a vendor quietly cuts a quota you want to see it in your own metrics before you read about it.

When to Pay

The free tiers hit their limits when you need:

For side projects and low-traffic apps, Gemini and Groq will genuinely carry you. For anything with real users, budget $50-200/month and stop counting requests.

Every row in the comparison table at the top has moved at least once since April, so if you’re reading a free-tier comparison that hasn’t been touched in six months, assume most of it is fiction.

Change log

I re-check this page against each vendor’s own docs, not against other people’s blog posts. Where a vendor no longer publishes a number, I say so instead of quoting the old one.


Related: Free Developer Tool Credits | Is the Claude API Free? | Best AI Tools Developers Actually Use

Frequently Asked Questions

What are the rate limits for free AI API tiers in 2026?

Groq's Free plan gives openai/gpt-oss-120b 30 requests/minute, 8,000 tokens/minute, 1,000 requests/day and 200,000 tokens/day, and Cohere's trial key is 20 requests/minute on the Chat API with 1,000 API calls per month. Google and Mistral no longer publish their free-tier rate limits on a public docs page.

Which AI API providers offer free tiers with no daily token limits?

For a general-purpose LLM you can build on, no. The last exceptions were Groq's groq/compound and groq/compound-mini, which showed no tokens-per-day figure (they were held to 250 requests/day instead); Groq decommissioned both on 21 September 2026 and they are gone from its Free plan table. Every language-model row left on that table carries a tokens-per-day cap. Together AI, which used to be the answer here, states in its own billing docs that it does not currently offer free trials and requires a $5 minimum purchase.

How can I build a production-ready AI backend at zero cost?

Google's Gemini free tier and Groq's Free plan are the two that still cost $0 and handle real traffic, so run Gemini as primary and Groq as fallback. Start new Gemini projects on Gemini 3.8 Flash: since 18 September 2026 Google limits the 2.5 models to users who have actively used them in the past. HuggingFace barely qualifies anymore: its free Inference Providers allowance is now a $0.10/month credit pool ('subject to change', per its own pricing table), enough for light embeddings work but not LLM traffic.

Written by Hirak Banerjee

Indie dev and maker. I build AI-powered apps and write about the tools I actually use. Follow on X · GitHub

Get told when these numbers change

Every figure here is re-checked against the vendor's own docs. Leave your email and we'll tell you when one moves — only when something actually changes.

Join builders who ship faster. No spam.

Comments