⚡ New — Kimi K3 is live: bring your own Moonshot key →

Which model should you use?

The cheapest model, the snappiest model and the fastest-streaming model are usuallythree different models — and the winner flips with your workload. Pick your use case; we rank the catalog by what that workload actually feels:₹ per task, time to first token andsustained throughput, measured through the production gateway (sweep of 2026-07-02), not read off datasheets.

Start from your application

Agent loops make dozens of short, tool-calling turns per task — time to first token dominates how fast the agent feels, throughput matters for big diffs, and cost adds up across the loop. Reasoning support is required.

#ModelBest route₹/Mtok (blended)First tokentok/s
1gpt-oss-120bopen weights🇮🇳 India route🔧 toolsfireworks₹391.2s469
2qwen3-32bopen weights🔧 toolsgroqBYOK₹222.1s382
3gpt-oss-20bopen weights🇮🇳 India route🔧 toolskrutrim₹244.7s606
4glm-4.7open weights🔧 toolsopenrouterBYOK₹1361.9s68
5gemini-2.5-flash🔧 toolsopenrouter₹1872.2s132
6glm-4.7-flashopen weights🔧 toolsprice-rankedzhipuBYOKfree——
7glm-4.5-flashopen weights🔧 toolsprice-rankedzhipuBYOKfree——
8nemotron-3.5-lightning-30b-a3bopen weights🔧 toolsprice-rankednvidiaBYOKfree——

Blended ₹/Mtok = cheapest route at a 1:3 input:output token mix (generation-dominant). First token and tok/s are the best measured route per model. Rankings are a weighted percentile score per lens — details in the methodology below.

The full picture — every chat model, three lenses

Click a metric column to sort by that lens. The same model often wins one and loses another.

ModelRoutes₹/Mtok ↕First token ↕tok/s ↕
glm-4.7-flashreasoning🔧 toolszhipuopenrouterfree——
glm-4.5-flashreasoning🔧 toolszhipufree——
nemotron-3.5-lightning-30b-a3breasoning🔧 toolsnvidiafree——
nemotron-nano-omni-30breasoningdgxspark2 🇮🇳₹1.8——
qwen2.5-coder-7bbharatrouter 🇮🇳₹3.5——
qwen2.5-7b-instruct🔧 toolsbharatrouter 🇮🇳₹3.5——
qwen3-8breasoning🔧 toolsbharatrouter 🇮🇳₹3.5——
nemotron-super-120breasoning🔧 toolsdgxspark 🇮🇳₹3.5——
qwen2.5-vl-7b-instructbharatrouter 🇮🇳₹5.3——
gemma-4-e4b-it🔧 toolskrutrim 🇮🇳₹6.8——
qwen3.5-9breasoning🔧 toolskrutrim 🇮🇳₹6.8——
llama-3.1-8b-instruct🔧 toolsopenroutergroqfireworks₹7.0428ms groq548
glm-4-32b-0414-128k🔧 toolszhipu₹10——
command-r7b🔧 toolscohere₹12——
qwen3-32breasoning🔧 toolsgroqopenrouterfireworks₹222.1s groq382
gemma-4-26b-a4b-it🔧 toolskrutrim 🇮🇳₹23——
qwen3.6-35b-a3breasoning🔧 toolskrutrim 🇮🇳₹23——
gpt-oss-20breasoning🔧 toolskrutrim 🇮🇳groqfireworks₹244.7s krutrim606
deepseek-v4-flashreasoning🔧 toolsdeepseek₹24——
devstral-small🔧 toolsmistral₹24——
llama-3.3-70b🔧 toolsgroqopenrouterfireworks₹262.0s openrouter218
gemma-4-31b-it🔧 toolskrutrim 🇮🇳₹27——
llama-4-scout🔧 toolsgroqfireworks₹28——
glm-4.6v-flashxreasoning🔧 toolszhipu₹30——
glm-4.7-flashxreasoning🔧 toolszhipu₹30——
gemini-2.5-flash-litereasoning🔧 toolsgemini₹31——
gpt-oss-120breasoning🔧 toolskrutrim 🇮🇳groqbasetenfireworks₹391.2s fireworks469
glm-5.3-flashreasoning🔧 toolskrutrim 🇮🇳openrouter₹40——
mistral-small🔧 toolsmistral₹47——
command-r🔧 toolscohere₹47——
gpt-4o-mini🔧 toolsopenai₹47——
gpt-4o-mini-search-preview🔧 toolsopenai₹47——
nemotron-superreasoning🔧 toolsbaseten₹61——
glm-4.5-airreasoning🔧 toolszhipuopenrouterfireworks₹65——
glm-4.6vreasoning🔧 toolszhipuopenrouter₹72——
codestral🔧 toolsmistral₹72——
deepseek-v4-proreasoning🔧 toolsdeepseekbasetenfireworks₹74——
deepseek-v3🔧 toolsopenrouter₹8116.3s openrouter38
gpt-5.6-lunareasoning🔧 toolsopenai₹91——
sonar🔧 toolsperplexity₹96——
grok-code-fast-1reasoning🔧 toolsxai₹113——
gemini-3.1-flash-litereasoning🔧 toolsgemini₹114——
mistral-largereasoning🔧 toolsmistral₹120——
glm-4.7reasoning🔧 toolsbasetenzhipuopenrouter₹1361.9s openrouter68
gpt-5-minireasoning🔧 toolsopenai₹15010.2s openai63
glm-5reasoning🔧 toolsbasetenzhipuopenrouter₹15312.2s zhipu53
glm-4.6reasoning🔧 toolszhipuopenrouter₹156——
grok-build-0.1reasoning🔧 toolsxaiopenrouter₹168——
glm-4.5reasoning🔧 toolszhipuopenrouter₹173——
kimi-k2.5reasoning🔧 toolsbasetenopenrouter₹173——
kimi-k2🔧 toolsgroqopenrouter₹180688ms groq178
nemotron-ultrareasoning🔧 toolsbaseten₹187——
gemini-2.5-flashreasoning🔧 toolsgeminiopenrouter₹1872.2s openrouter132
gemini-3.5-flash-litereasoning🔧 toolsgemini₹187——
grok-4.3reasoning🔧 toolsxaiopenrouter₹210——
grok-4.20reasoning🔧 toolsxaiopenrouter₹210——
glm-5.2reasoning🔧 toolsbasetenzhipuopenrouterfireworks₹2385.7s zhipu68
glm-5.1reasoning🔧 toolsbasetenzhipuopenrouter₹242——
kimi-k2.7-codereasoning🔧 toolsfireworksmoonshotbasetenopenrouter₹27019.6s openrouter39
gemini-3.7-flashreasoning🔧 toolsgemini₹288——
kimi-k2.6reasoning🔧 toolsfireworksmoonshotbasetenopenrouter₹311——
glm-5-turboreasoning🔧 toolszhipu₹317——
glm-5v-turboreasoning🔧 toolszhipuopenrouter₹317——
glm-5.3reasoning🔧 toolsfireworkszhipu₹350——
glm-4.5-airxreasoning🔧 toolszhipu₹351——
claude-haiku-4.5reasoning🔧 toolsanthropicopenrouter₹38416.8s openrouter102
qwen3.8-maxreasoning🔧 toolsalibabaopenrouter₹480——
pixtral-large🔧 toolsmistral₹480——
grok-4.6reasoning🔧 toolsxai₹480——
grok-4.5reasoning🔧 toolsxai₹480——
gemini-3.6-flashreasoning🔧 toolsgemini₹576——
mistral-medium-3.5🔧 toolsmistral₹576——
kimi-k2.7-code-highspeedreasoning🔧 toolsmoonshot₹622——
sonar-reasoning-proreasoning🔧 toolsperplexity₹624——
sonar-deep-researchreasoning🔧 toolsperplexity₹624——
gemini-3.5-flashreasoning🔧 toolsgemini₹684——
glm-4.5-xreasoning🔧 toolszhipu₹693——
gemini-2.5-proreasoning🔧 toolsgeminiopenrouter₹750——
gpt-5reasoning🔧 toolsopenai₹750——
claude-sonnet-5reasoning🔧 toolsanthropicopenrouter₹768——
command-a🔧 toolscohere₹780——
gpt-4o-search-preview🔧 toolsopenai₹780——
gpt-5.6-terrareasoning🔧 toolsopenai₹912——
kimi-k3reasoning🔧 toolsfireworksmoonshotopenrouterbaseten₹1152——
sonar-pro🔧 toolsperplexity₹1152——
claude-opus-4.8reasoning🔧 toolsanthropicopenrouter₹1920——
claude-opus-5reasoning🔧 toolsanthropic₹1920——
gpt-5.6-solreasoning🔧 toolsopenai₹2280——
gpt-5.6reasoning🔧 toolsopenai₹2280——
claude-fable-5reasoning🔧 toolsanthropicopenrouter₹3840——
gpt-6-astrareasoning🔧 toolsopenai₹3840——

In the open — how these numbers are made

Perf numbers are medians from a multi-run streamed sweep through the production gateway on 2026-07-02: every (model × provider) route gets the same ~300-token prompt, rounds interleaved across hosts so no provider owns a time-of-day advantage. First token counts reasoning tokens (it's what you see). tok/s is the post-first-token decode rate. Models not yet swept show “—” and rank on price with a neutral perf score. Routes we could not measure are listed openly in theAPI response(9 skipped this sweep), never silently dropped. Numbers refresh with each sweep; live per-route health is on /models.

Agents get this same chooser as JSON:GET /v1/compare/models — rankings, per-route pricing, measured perf and live failure rates, no auth required.

Get a keyBrowse the catalog