- Space Bunny Alpha listed free Sep 23
- Baidu: ERNIE 4.5 VL 424B A47B expires Oct 8
- MiniMax: MiniMax M2.1 expires Oct 8
- Qwen: Qwen Plus 0728 expires Oct 9
- GitHub Models retired on july 30, 2026
Free models,
actually tested.
Prices for 436 AI models, free tiers decoded, and reliability measured on real agent work.
- Models priced
- 436 refreshed hourly
- Routes ending
- 24 next: Oct 8
- Tokens measured
- 250.8M signed cheapoS reports
| Compare | Model | Price / M | Context | Coding | Reliability measured | Avg time |
|---|---|---|---|---|---|---|
| NVIDIA: Nemotron 3 Super (free)nvidia/nemotron-3-super-120b-a12b:freetools · reasoning | Free | 262K | 37.7 | 93.9%likely 89–97% · 164 outcomes | 13.2 s | |
| Dots Studio: Dots3-Note Preview (free)dots-studio/dots-3-note-preview:freetools · vision · reasoning | Free | 512K | — | 92.3%likely 83–97% · 65 outcomes | 17.0 s | |
| NVIDIA: Nemotron 3 Ultra (free)nvidia/nemotron-3-ultra-550b-a55b:freetools · reasoning | Free | 1M | 49.3 | 89.2%likely 81–94% · 93 outcomes | 26.8 s | |
| Poolside: Laguna XS 2.1 (free)poolside/laguna-xs-2.1:freetools · reasoning | Free | 262K | — | 68.4%likely 53–81% · 38 outcomes | 16.3 s | |
| Poolside: Laguna S 2.1 (free)poolside/laguna-s-2.1:freetools · reasoning | Free | 262K | — | 72.7%likely 52–87% · 22 outcomes | 40.1 s | |
| inclusionAI: Ling 3.0 Flash Sante (free)inclusionai/ling-3.0-flash-sante:freetools · reasoning | Free | 262K | — | 58.3%likely 39–76% · 24 outcomes | 11.2 s | |
| NVIDIA: Nemotron 3 Nano Omni (free)nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:freetools · vision · reasoning | Free | 256K | 13.8 | 53.8%likely 36–71% · 26 outcomes | 19.6 s | |
| Google: Gemma 4 31B (free)google/gemma-4-31b-it:freetools · vision · reasoning | Free | 262K | 43.4 | 5%likely 1–24% · 20 outcomes | 6.8 s | |
| Google: Gemma 4 26B A4B (free)google/gemma-4-26b-a4b-it:freetools · vision · reasoning | Free | 262K | 39.3 | 0%likely 0–13% · 26 outcomes | 7.9 s | |
| Cohere: North Mini Code (free)cohere/north-mini-code:freetools · reasoning | Free | 256K | 36.5 | Too few13 of 17 answered · rates start at 20 | 11.2 s | |
| Free Models Routeropenrouter/freetools · vision · reasoning | Free | 200K | — | No reports | — | |
| LiquidAI: LFM2.5-2.6B (free)liquid/lfm-2.5-2.6b:freetools · reasoning | Free | 66K | — | Too few9 of 13 answered · rates start at 20 | 12.8 s | |
| NVIDIA: Nemotron 3.5 Lightning (free)nvidia/nemotron-3.5-lightning:freetools · reasoning | Free | 1M | 26.8 | Too few7 of 14 answered · rates start at 20 | 52.4 s | |
| Qwen: Qwen3.8 27B (free)qwen/qwen3.8-27b:freetools · vision · reasoning | Free | 262K | 68.1 | Too few0 of 8 answered · rates start at 20 | 9.1 s | |
| Space Bunny Alphastealth/space-bunny-alphatools · vision · reasoning | Free | 1M | — | No reports | — | |
| Thinking Machines: Inkling (free)thinkingmachines/inkling:freetools · vision · reasoning | Free | 1.05M | 52.1 | Too few0 of 7 answered · rates start at 20 | 7.6 s | |
| Thinking Machines: Inkling Small (free)thinkingmachines/inkling-small:freetools · vision · reasoning | Free | 1.05M | 52.9 | Too few0 of 10 answered · rates start at 20 | 8.7 s |
Prices are OpenRouter’s published text-token rates per million tokens; $0 isn’t unlimited. Reliability comes from signed cheapoS reports (1 member shares model details): a rate appears after 20 known outcomes, with its likely range. How we measure
How we measure
Published
Prices, context and capabilities come from the OpenRouter catalog, refreshed hourly (last Sep 29, 2026, 2:11 PM UTC). Prices are USD per million text tokens. A listing isn't a live availability check, and $0 comes with account rules and limits.
Benchmark
Coding, agentic and intelligence indices from Artificial Analysis, as distributed by OpenRouter. They measure test suites, not reliability on real work. Missing scores stay missing.
Measured
Accepted, signed request reports from cheapoS members who share model details. A rate appears only after 20 known outcomes, always with its likely range (95% Wilson interval); averages after 5 timed reports. “Answered” means the provider responded, not that the task succeeded. The sample is early and can be one person.
Before you pick a model
Are free models really free?
They list zero input and output token prices. You still need a provider account and API key, and usage limits apply. Optional tools or features may cost extra. See what free means at each provider.
Which free model is best for coding?
Filter for tool calling and sort by the coding benchmark as a starting point. Then check each model page for Club reports and their sample size, and try a small task in your own project. Benchmarks, request reliability and finished work answer different questions.
Why does a model have no Club reports?
It may be new, unused by members who share model details, or recorded under a route we can't match exactly. It stays listed because published availability is useful on its own. We never fill the gap with a guessed rate.
Can I help make the directory more useful?
Use cheapoS on real projects and opt in to Club sharing, including model details if you choose. Only accepted, signed reports count. Join the Club or read what gets shared.
Community compute notebookExact totals, role breakdown and shareable badge
Free intelligence, put to work.
Total: 1 public reporting members (up to 100). Breakdowns: all 2 sharing members. Models include shared names only.
Report fetched 2026-09-29 15:00:50 UTC. Updates when new reports arrive.
Illustrative reference: $3 USD per 1M reported tokens.
Free intelligence, real scale. This compares reported public-free, included-access and local tokens with a flat reference rate. Paid and unknown-access tokens are excluded.
An illustration, not actual charges or verified savings. Subscription, hardware and electricity costs are not deducted.
Input vs output
Input is context sent to models, including repeated context. Output is what models generate. Token volume measures usage, not work quality.
- Input
- 246,815,36698.4%
- Output
- 3,935,5701.6%
7,422 reported requests across 1 reporting public profile. Public-free, included-access and local usage only.
Cached input & reasoning output
- Cached input
- 27,254,813 tokens · 1,355 of 7,422 requests reported
- Reasoning output
- 580,329 tokens · 1,321 of 7,422 requests reported
These reported subsets are already included in input or output. Missing details stay unknown; they are not added to the total.
- Input
- 176,783,951
- Output
- 1,692,469
- Input
- 49,216,215
- Output
- 1,642,827
- Input
- 11,860,762
- Output
- 364,708
- Input
- 107,308
- Output
- 3,875
Get cheapoS Free
Autonomous coding with local test execution and zero subscription tax. Choose your preferred environment:
git clone https://github.com/cheapos/CheapoS.git && cd CheapoS && python3 run.pyOpens the app in your browser at http://127.0.0.1:5173/. On macOS, you can also double-click Start CheapOS.command after cloning.
Grab 2 Free API Keys
cheapos harnesses generous free quotas. Get a free key from Google AI Studio (Gemini 2.5 Flash) and GroqCloud (Llama 3.3 70B). Zero credit card required.
Enter Any Task Prompt
Type what you want to build (CLI tool, micro-app, test suite). The fast worker drafts code while the reviewer critiques and fixes mistakes autonomously.
⚡ 0 bill shock · 100% local executionDeterministic Test & Ship
cheapos never stops until local unit tests pass (e.g. 28/28 tests passing). Review the diff and 1-click share your finished product to the Workbench.
Explore 10 showcase builds →