OpenRouter Rankings July 2026
Who's Actually Winning the AI Model Race?

If you are still picking models based on impressions from a few months ago, OpenRouter traffic through July 25, 2026 tells a different story: Xiaomi Mimo V2.5 leads at 1.4 trillion tokens per day, Chinese labs collectively hold roughly 46% of share (up from under 2% a year ago), and the US big three have slid from ~70% to 30%–36%. This article is for developers and engineering leads building multi-model routing — we break down the Top 12 model table, provider share, why usage does not equal quality, the hidden app layer, price benchmarks, August outlook, and a six-step tiered routing playbook.

01

Why "Number One on the Leaderboard" Stopped Being Enough in July 2026

The OpenRouter rankings sort by real token volume — what developers actually pay to run, not who scores highest on a benchmark. The July board moves fast (Mimo V2.5 reshuffled the Top 10 between July 24 and 25 alone). If you are still using an old mental model for model selection, you will keep hitting the same traps:

  1. 01

    Treating MMLU as a production metric: Benchmarks measure ceiling capability; OpenRouter measures billing habits. Cheap, fast models attached to high-traffic apps can climb to #1 even when complex reasoning is mediocre.

  2. 02

    Ignoring the usage-vs-quality split: Volume leaders like cheap open-weight models barely register on classification and complex-reasoning spend share — Claude Sonnet 4.6 and Claude Opus 4.7 each hold 13.5% there.

  3. 03

    Watching models but not apps: openrouter.ai/apps shows Hermes Agent alone at roughly 45% of application tokens; roleplay apps collectively drive a large share of open-source traffic that enterprise coverage rarely mentions.

  4. 04

    Treating a daily #1 as a permanent lead: DeepSeek holds the steadiest provider share (roughly 16%–18%), but the monthly model crown has rotated from MiniMax M2.5 to MiMo-V2-Pro to Mimo V2.5 — competition inside the Chinese camp is just as fierce.

  5. 05

    Missing the 35x price gap: DeepSeek V4 Flash inputs run around $0.05–0.14/M while GPT-5.5 sits near $5/M — when open models are "good enough," rational developers vote with their wallets.

  6. 06

    Leaving safety off the scorecard: This week brought OpenAI sandbox-escape headlines and US congressional progress on an "AI kill switch" bill — vendor safety track records will weigh more heavily in enterprise Agent deployments.

"Capability and popularity are diverging — the smarter bet is tiered routing, not obsessing over who holds #1 today."

02

July Rankings: Xiaomi Takes #1, Chinese Models Hit ~46% Share

Data as of July 25, 2026 (verify live numbers at openrouter.ai/rankings before you ship). The top three by daily token volume: Xiaomi Mimo V2.5 (1.4T/day), DeepSeek V4 Flash (943.9B/day), Tencent Hy3 (590B/day). Seven of the Top 12 models come from Chinese labs.

Model Token Volume Top 12 (Daily)

RankModelProviderDaily Tokens30-Day Total
1Mimo V2.5Xiaomi1.4T31.2T
2DeepSeek V4 FlashDeepSeek943.9B23.6T
3Hy3Tencent590B23.4T
4Nemotron 3 Ultra 550B (free)NVIDIA428.6B9T
5DeepSeek V4 ProDeepSeek413.7B11.6T
6GLM 5.2Z.ai316.7B13.3T
7MiniMax M3MiniMax262.5B15.1T
8Step 3.7 FlashStepFun204.8B5.9T
9Kimi K3Moonshot157.6B1.6T (new entry)
10Ling 3.0 FlashAnt InclusionAI128.3B417.3B
11Gemini 3 Flash PreviewGoogle106.3B4T
12Claude Sonnet 5Anthropic99.5B3.6T

Provider Share (Multi-Source 7-Day Window, Approximate)

ProviderRegionToken Share (Approx.)
DeepSeekChina16%–18% (most sources rank #1)
XiaomiChina8%–18% (Mimo V2.5 surge; highest volatility)
AnthropicUS10%–15%
TencentChina8%–13%
GoogleUS8%–13%
Z.aiChina4%–7%
OpenAIUS6%–8%

Chinese providers collectively hold roughly 46%; US providers sit around 30%–36% (down from ~70% a year ago). Compare month-over-month shifts in our June 2026 OpenRouter analysis.

info

Citable hard numbers: (1) Mimo V2.5 at 1.4T/day on 7/25. (2) DeepSeek provider share at 16%–18%, the most stable #1. (3) US vs China share flipped by roughly 40 percentage points over twelve months.

03

The Other Side of the Board: High Volume Does Not Mean High Quality

OpenRouter's spend share by task type (dollar-weighted, not token-weighted) paints a different picture: general chat 35.7%, Agent workflows 30.4%, code 26.5%, data processing 7.5%. But zoom into classification and complex reasoning:

DimensionVolume Tier (Cheap Open-Weight Dominant)Hard-Task Tier (Closed Frontier Dominant)
Typical WorkloadsCasual chat, creative writing, roleplay, light coding assistComplex reasoning, enterprise Agent planning, high-consistency classification
Spend Share LeadersMimo V2.5, V4 Flash, Hy3, and peersClaude Sonnet 4.6 / Opus 4.7 at 13.5% each; GPT-5.5 at 11.6%
Pricing LogicUp to 35x cheaper; absorbs massive fault-tolerant trafficOpus 5 holds at $5/$25; FrontierBench v0.1 at 43.3%

On July 24, Anthropic shipped Claude Opus 5: 43.3% on the new FrontierBench v0.1 benchmark (GPT-5.6 Sol at 37.5%), pricing unchanged at $5/$25 — half of Fable 5's input tier. The message is clear: "I cost more because I'm worth it on the hardest tasks." See our Claude Opus 5 launch and Kimi K3 controversy coverage for the full breakdown.

warning

Selection takeaway: Closed frontier models handle the hardest 5%–10% of tasks; Chinese open-weight models absorb the remaining 90%+ of throughput. That is the clearest signal in July's data — not a binary "who is stronger" debate.

04

The App Layer: Coding Agents Rule, Roleplay Is the Invisible Half

Model rankings show which "brain" developers prefer; the application leaderboard shows what those brains actually do in production:

RankApplicationCategoryShare (Approx.)
1Hermes AgentPersonal / CLI Agent~45%
2Kilo CodeCoding Agent~13%
3OpenClawGeneral Agent~9%
4Claude CodeCoding Agent~6%
5DescriptContent production~4.5%
7–9Lemonade / ISEKAI ZERO / Janitor AICompanion / roleplay (new entries)~2% each
10ClineCoding Agent (IDE extension)~1.7%

Cline, Roo Code, and Kilo Code share the same codebase lineage across three forks — the "grandchild" Kilo Code has already overtaken ancestor Cline. OpenRouter and a16z's State of AI report finds creative roleplay alone accounts for more than half of total open-source model usage. If you only read enterprise AI coverage, you will miss this entire half of the market.

05

August Outlook, Price Benchmarks, and Citable Parameters

Based on July trends and industry signals, here is our August read — plus a price positioning table you can pair with our OpenRouter integration guide for hands-on setup:

ModelInput $/MOutput $/MPositioning
DeepSeek V4 Flash~$0.05–0.14~$0.24–0.28Best value; top pick for Agentic coding
MiniMax M3$0.10$1.21Long-context / multimodal on a budget
GLM 5.2$0.45$3.31Near-Opus planning quality among open models
Kimi K3~$3~$151.4TB open weights, closed-tier capability
Claude Opus 5$5 (fast tier $10)$25 (fast tier $50)Closed frontier; strongest July benchmark scores
  • Share forecast: Combined Chinese open-weight share likely keeps climbing — could cross 50% this year unless US providers cut prices significantly.
  • Crown rotation: The monthly #1 model will keep changing hands; watch each lab's August release cadence.
  • Kimi K3 quantization: Community quant versions of the 1.4TB weights should land within 2–4 weeks — the point when smaller teams can actually self-host.
  • Security and compliance: Vendor safety reputation will carry more weight in enterprise procurement; autonomous Agent deployments need tighter permission scoping.
06

Six-Step Tiered Routing Playbook and Role-Specific Guidance

Independent developers: OpenRouter is a fine sandbox for model evaluation. For coding tasks, start with DeepSeek V4 Flash and GLM 5.2; reserve Opus 5 / GPT-5.6 for steps that stall. Expect ~180–250ms latency from some regions — evaluate compliant domestic routing before production.

Engineering leads: Do not pick models from the volume leaderboard alone. Route by task type and add vendor safety track records to your selection scorecard.

Six-Step Implementation Checklist

  1. 01

    Review openrouter.ai/rankings weekly: Log Top 12 and provider share shifts with the data cutoff date — the board moves daily.

  2. 02

    Write gateway rules by scenario: Chat/creative to V4 Flash / Mimo V2.5; complex reasoning to Claude Opus 5; high-risk classification to Sonnet 4.6 / Opus 4.7.

  3. 03

    Split token bills from dollar bills: If tokens cluster on Flash but dollars cluster on Claude, you are already tiered — codify the routing table explicitly.

  4. 04

    Track the app layer, not just models: Coding Agents (Kilo Code, Claude Code) and roleplay traffic have very different model profiles — match models to your product's actual workload.

  5. 05

    Keep a regression suite ready for August releases: When a new #1 appears, run the same eval within 48 hours and update routing — do not rewrite the application.

  6. 06

    Pin down your Agent execution layer: Long-session CLI Agents and sensitive prefill belong on dedicated SSH nodes; let OpenRouter handle elastic peaks. See rental pricing for specs.

A laptop that sleeps or a cheap Linux VPS cannot sustain 12-hour Agent loops or run macOS-only tooling like xcodebuild and notarytool. Pairing weekly leaderboard reviews with a fixed execution environment is more sustainable than chasing a single "best model" every week.

For iOS CI/CD and AI Agent automation teams that need stable SSH sessions, Keychain isolation, and predictable bandwidth, the winning pattern is clear: define OpenRouter multi-model routing in your gateway and place heavy workloads on a dedicated cloud Mac rather than routing every token through public APIs. NodeMini cloud Mac Mini rental works as the Agent execution layer — when you rotate API keys or swap model endpoints, your SSH node and CI labels stay put. Setup details live in the Help Center.

FAQ

Frequently Asked Questions

As of July 25, 2026, Xiaomi Mimo V2.5 led daily token volume at roughly 1.4T/day, followed by DeepSeek V4 Flash (943.9B/day) and Tencent Hy3 (590B/day). Rankings shift daily — check openrouter.ai/rankings for live data.

No — different measurement windows (seven-day provider share vs overall token volume including free and long-tail models) and whether first-party app traffic is included explain the gap. The provider table here uses a multi-source seven-day approximation. The trend that matters: from under 2% to ~46% in twelve months. See our June analysis for the monthly series.

Use OpenRouter for elastic multi-model routing and billing visibility; run long-session CLI Agents and local prefill on a dedicated cloud Mac with fixed monthly cost. Setup and key configuration: Help Center. Hardware specs and pricing: rental rates.

(1) Whether combined Chinese open-weight share breaks 50%. (2) Whether Anthropic ships a lower-priced tier to chase volume. (3) Community Kimi K3 quantization landing. (4) US AI safety legislation and White House pre-release review frameworks affecting enterprise model selection.