⚡ New — Kimi K3 is live: bring your own Moonshot key →

← Blog · Coding · Models

GLM-5.3-Flash for coding agents

GLM-5.3-Flash for coding agents on BharatRouter — 1M context, measured latency, and the prefix cache that carries agent loops.

A fast, India-resident model with a 1 million-token context, verified tool-calling, and a prefix cache that keeps long agent loops fast. Here's where it sits next to the Claude/Codex tiers you know, the latency we actually measured, and how to point your coding agent at it in one line.

What it is

glm-5.3-flash is served on Krutrim Cloud and reachable on the BharatRouter platform key — no BYOK. It's MIT-licensed, runs fp8 in an India datacenter, and carries a 1,048,576-token context window (131,072 max output). We confirmed it emits proper OpenAI tool_calls live, so agentic harnesses (brcode, Codex, Claude Code, Hermes) drive it the same way they drive a frontier model.

In Claude/Codex parlance

Capability tiers are approximate and task-dependent — benchmarks aren't the same as how a model feels in your repo — but a useful frame:

If you're used to…Reach for…Why
Claude Haiku (fast, cheap, high-volume)glm-5.3-flashSame speed/cost role — but a 1M window (vs ~200K) and real reasoning + tools. Great for whole-repo prompts and long files.
Claude Sonnet / Codex default (day-to-day agentic coding)gpt-oss-120bThe heavier agentic-coding workhorse; the model Codex CLI ships open-weight support for. Pick this for long tool-loop tasks.
Claude Opus (frontier)—No exact platform-served match yet; the honest gap. GLM-5 full or a larger model on dedicated hardware is the closest.

Short version: Haiku-class speed and cost, Sonnet-ish context and reasoning.For the heaviest step-by-step agentic coding, gpt-oss-120b still leads; for fast iteration and large context, glm-5.3-flash is the better pick.

How big are they — and why the flash models are quick

Both of ours are open-weight, so their sizes are public. Both are Mixture-of-Experts: a large total parameter count, but only a small activeslice runs per token — which is exactly why they stream fast for their capability.

ModelParametersWeights
glm-5.3-flash320B total · 18B active (MoE)Open (MIT)
gpt-oss-120b117B total · 5.1B active (MoE, 128 experts)Open (Apache-2.0)
Claude Haiku 4.5 / Sonnet 5Not disclosedProprietary
GPT-5 / GPT-5-miniNot disclosedProprietary

Sizes: gpt-oss-120b from its OpenAI model card(116.8B total, 5.1B active); glm-5.3-flash per Zhipu's model card (320B total, 18B active). Anthropic and OpenAI don't publish parameter counts for their hosted models — one honest edge of the open-weight option is that you can see, run and audit exactly what you're using.

Which one should I use?

Especially if you're coming from Codex or Claude — a practical guide so you're not guessing:

Your taskReach for
Day-to-day coding, agent loops, tool-heavy workgpt-oss-120b
Fast iteration, whole-repo / long-file prompts, quick Q&Aglm-5.3-flash (1M context)
The absolute hardest reasoning (Opus-tier)Neither is Opus-class yet — the one real gap

Coming from…

Rule of thumb: not sure? Start with gpt-oss-120b for coding; switch toglm-5.3-flash if it feels slow or you're feeding a big codebase. Both do real tool-calling, so your agent works either way — and you can switch anytime (brcode model <id>, or just change the model field).

Reasoning & effort

Both models are reasoning models — they can think before answering. The knob is the standard OpenAI reasoning_effort:

Why it matters: with a reasoning model, max_tokens is shared between the hidden reasoning and the answer. Too small a budget and the model can run out mid-thought — so either lower reasoning_effort or raise max_tokens.

How they compare — measured

Streamed, single-request, warm, from Bengaluru — the same three coding prompts (a small function, an LRU cache, a concept + example) across every model, medians, max_tokens 1500. GLM-5.3-Flash runs at its shipped default. Every model answered all three correctly, so this isn't a capability benchmark — it's a responsiveness comparison, which is what you actually feel in an editor.

Time to first token

How long until text starts appearing. Under ~1s feels instant; several seconds feels stalled.

Off the chart above: gpt-5-mini ~7.9s · gpt-5 ~8.3s — 8–16× slower to first token. OpenAI's gpt-5 family reasons fully before emitting anything, so you wait ~8s for the first character. Our two models and Claude all start in under ~1.1s.

Full completion time

Total wall-clock for the whole answer, single request.

An honest caveat: part of this gap is geography — Krutrim serves from India, Claude and OpenAI answer from the US, so some of it is round-trip distance, not the model. For a team in India that's a real everyday advantage — but it's distance, plus our models' tighter output, not only raw speed.

Under load

At 20 concurrent streams against production, both our models heldzero failures and zero empty responses. GLM-5.3-Flash stayed fast (p50 ~7s) — its 1M context and prefix cache make it comfortable for high-concurrency, whole-repo work. gpt-oss-120b, a heavy reasoning model, is more sensitive to load (p50 ~16s at 20 concurrent), so for the busiest team-wide agentic workloads, GLM-5.3-Flash is the better high-concurrency pick.

Reproduce the first-token number yourself:

# TTFT + total, streamed, from your machine — swap in your br- key
curl -sN -o /dev/null -w 'ttfb=%{time_starttransfer}s total=%{time_total}s\n' \
  https://api.bharatrouter.com/v1/chat/completions \
  -H "Authorization: Bearer br-..." -H 'content-type: application/json' \
  -d '{"model":"glm-5.3-flash","stream":true,"max_tokens":120,
       "messages":[{"role":"user","content":"Count to ten."}]}'

Quality & multi-turn — measured, and honest

Speed only counts if the code is right. So we measured correctness ourselves — and we're straight about what it does and doesn't tell you.

HumanEval pass@1 — measured through BharatRouter

We ran HumanEval (first 80 problems) throughapi.bharatrouter.com, executing each completion against the hidden unit tests in a sandboxed subprocess. pass@1 = solved on the first try.

Read this the right way: HumanEval is saturated — every current coding model sits in the high-90s, so ours land at the ceiling alongside the frontier, and a one-problem gap is noise. It's a correctness floor (does the model write correct Python?), not a measure of agentic skill. (claude-sonnet-5 isn't shown — our test key's route to it errored during the run; its HumanEval would be at ceiling too.)

Multi-turn & agentic — where the real difference is

Coding agents don't write one function and stop — they read a repo, run tools, hit failing tests, and iterate. The community standard for that isSWE-bench Verified (real GitHub issues; the patch must make the project's own tests pass). Our HumanEval number says nothing about this axis — here's where things stand on published scores:

Honest summary: on standard correctness our models are at ceiling with the frontier; on the hardest multi-turn agentic benchmark, frontier Claude/GPT still lead, and gpt-oss-120b is the open model that closes most of that gap. Reach for glm-5.3-flash for speed and 1M context, gpt-oss-120b for heavy agentic loops.

Method: HumanEval pass@1 measured by us on 80 problems via the BharatRouter API, executed in a sandboxed subprocess — a single-turn correctness check, not an agentic evaluation. SWE-bench figures are third-party published scores; vendor-reported numbers use vendor harnesses and aren't strictly head-to-head.

The cache you'll feel

Agentic coding resends the same system prompt + repo context on every tool-loop turn. A prefix cache reuses the already-computed attention state for that shared prefix instead of recomputing it — so the second, third and hundredth turn of a session start faster than the first. This is why a "flash" model comfortably carries long agent sessions: the heavy compute on your big context is done once, so latency stays low turn after turn.

It's also why the tier table above understates glm-5.3-flash for agent work — the 1M window plus prefix reuse means you can keep a large codebase in context across a whole session and each turn still responds quickly. (Today that reuse is a latency win; billed token-level cache discounts are on the way.)

How to use it

One br- key; set the model to glm-5.3-flash. Pick your harness.

🤖 Setting up with an AI agent? Point it at the machine-readable version — glm-flash-coding-agent.md — with curl / Python / Node / brcode / Codex snippets it can run directly.

Any OpenAI-compatible client (curl / Python / Hermes)

curl https://api.bharatrouter.com/v1/chat/completions \
  -H "Authorization: Bearer br-..." \
  -H "Content-Type: application/json" \
  -d '{
    "model": "glm-5.3-flash",
    "messages": [{"role": "user", "content": "Refactor this function and explain why."}]
  }'
# pip install openai
from openai import OpenAI

client = OpenAI(base_url="https://api.bharatrouter.com/v1", api_key="br-...")

resp = client.chat.completions.create(
    model="glm-5.3-flash",
    messages=[{"role": "user", "content": "Write a FastAPI healthcheck endpoint."}],
    stream=True,
)
for chunk in resp:
    print(chunk.choices[0].delta.content or "", end="", flush=True)

Hermes, or anything that takes a base URL + key + model, is the same three settings:

# Hermes (or any OpenAI-compatible client): point it at BharatRouter.
BASE_URL = https://api.bharatrouter.com/v1
API_KEY  = br-...
MODEL    = glm-5.3-flash

brcode (Codex-/Claude-Code-style agent)

brew install bharatrouter/tap/brcode      # or: curl -fsSL https://bharatrouter.com/install/brcode.sh | bash
brcode login && brcode init                # zero-paste browser login + your org profile
brcode model glm-5.3-flash                 # (or one-off: BR_MODEL=glm-5.3-flash brcode)
brcode                                      # start coding — same AGENTS.md, Codex-shaped TUI

AGENTS.md travels with your repo unchanged. Full guide:the brcode reference · migrating fromCodex CLI orClaude Code.

Keep the real Codex CLI

BharatRouter speaks the Responses API, so the codex binary can route through it:

# Keep the real Codex CLI — route it through BharatRouter (Responses API)
# ~/.codex/config.toml
model = "glm-5.3-flash"
model_provider = "bharatrouter"

[model_providers.bharatrouter]
name = "BharatRouter"
base_url = "https://api.bharatrouter.com/v1"
wire_api = "responses"
env_key = "BR_API_KEY"     # export BR_API_KEY=br-...

Metering & budgets

Every call is metered and audited on your key. Admins set ₹/month or token budgetsper org, team, member or key — and the caps enforce (a request over budget gets a clean 429), so usage stays governed. Set them on the console's Spend & Budgetspanel.

"Claude", "Codex" and "Sonnet/Haiku/Opus" are products of Anthropic and OpenAI; tier comparisons are approximate positioning, not head-to-head benchmarks. Latency numbers are measured from Bengaluru and will vary with your location and load.