If you are still picking models based on impressions from a few months ago, OpenRouter traffic through July 25, 2026 tells a different story: Xiaomi Mimo V2.5 leads at 1.4 trillion tokens per day, Chinese labs collectively hold roughly 46% of share (up from under 2% a year ago), and the US big three have slid from ~70% to 30%–36%. This article is for developers and engineering leads building multi-model routing — we break down the Top 12 model table, provider share, why usage does not equal quality, the hidden app layer, price benchmarks, August outlook, and a six-step tiered routing playbook.
The OpenRouter rankings sort by real token volume — what developers actually pay to run, not who scores highest on a benchmark. The July board moves fast (Mimo V2.5 reshuffled the Top 10 between July 24 and 25 alone). If you are still using an old mental model for model selection, you will keep hitting the same traps:
Treating MMLU as a production metric: Benchmarks measure ceiling capability; OpenRouter measures billing habits. Cheap, fast models attached to high-traffic apps can climb to #1 even when complex reasoning is mediocre.
Ignoring the usage-vs-quality split: Volume leaders like cheap open-weight models barely register on classification and complex-reasoning spend share — Claude Sonnet 4.6 and Claude Opus 4.7 each hold 13.5% there.
Watching models but not apps: openrouter.ai/apps shows Hermes Agent alone at roughly 45% of application tokens; roleplay apps collectively drive a large share of open-source traffic that enterprise coverage rarely mentions.
Treating a daily #1 as a permanent lead: DeepSeek holds the steadiest provider share (roughly 16%–18%), but the monthly model crown has rotated from MiniMax M2.5 to MiMo-V2-Pro to Mimo V2.5 — competition inside the Chinese camp is just as fierce.
Missing the 35x price gap: DeepSeek V4 Flash inputs run around $0.05–0.14/M while GPT-5.5 sits near $5/M — when open models are "good enough," rational developers vote with their wallets.
Leaving safety off the scorecard: This week brought OpenAI sandbox-escape headlines and US congressional progress on an "AI kill switch" bill — vendor safety track records will weigh more heavily in enterprise Agent deployments.
"Capability and popularity are diverging — the smarter bet is tiered routing, not obsessing over who holds #1 today."
Data as of July 25, 2026 (verify live numbers at openrouter.ai/rankings before you ship). The top three by daily token volume: Xiaomi Mimo V2.5 (1.4T/day), DeepSeek V4 Flash (943.9B/day), Tencent Hy3 (590B/day). Seven of the Top 12 models come from Chinese labs.
| Rank | Model | Provider | Daily Tokens | 30-Day Total |
|---|---|---|---|---|
| 1 | Mimo V2.5 | Xiaomi | 1.4T | 31.2T |
| 2 | DeepSeek V4 Flash | DeepSeek | 943.9B | 23.6T |
| 3 | Hy3 | Tencent | 590B | 23.4T |
| 4 | Nemotron 3 Ultra 550B (free) | NVIDIA | 428.6B | 9T |
| 5 | DeepSeek V4 Pro | DeepSeek | 413.7B | 11.6T |
| 6 | GLM 5.2 | Z.ai | 316.7B | 13.3T |
| 7 | MiniMax M3 | MiniMax | 262.5B | 15.1T |
| 8 | Step 3.7 Flash | StepFun | 204.8B | 5.9T |
| 9 | Kimi K3 | Moonshot | 157.6B | 1.6T (new entry) |
| 10 | Ling 3.0 Flash | Ant InclusionAI | 128.3B | 417.3B |
| 11 | Gemini 3 Flash Preview | 106.3B | 4T | |
| 12 | Claude Sonnet 5 | Anthropic | 99.5B | 3.6T |
| Provider | Region | Token Share (Approx.) |
|---|---|---|
| DeepSeek | China | 16%–18% (most sources rank #1) |
| Xiaomi | China | 8%–18% (Mimo V2.5 surge; highest volatility) |
| Anthropic | US | 10%–15% |
| Tencent | China | 8%–13% |
| US | 8%–13% | |
| Z.ai | China | 4%–7% |
| OpenAI | US | 6%–8% |
Chinese providers collectively hold roughly 46%; US providers sit around 30%–36% (down from ~70% a year ago). Compare month-over-month shifts in our June 2026 OpenRouter analysis.
Citable hard numbers: (1) Mimo V2.5 at 1.4T/day on 7/25. (2) DeepSeek provider share at 16%–18%, the most stable #1. (3) US vs China share flipped by roughly 40 percentage points over twelve months.
OpenRouter's spend share by task type (dollar-weighted, not token-weighted) paints a different picture: general chat 35.7%, Agent workflows 30.4%, code 26.5%, data processing 7.5%. But zoom into classification and complex reasoning:
| Dimension | Volume Tier (Cheap Open-Weight Dominant) | Hard-Task Tier (Closed Frontier Dominant) |
|---|---|---|
| Typical Workloads | Casual chat, creative writing, roleplay, light coding assist | Complex reasoning, enterprise Agent planning, high-consistency classification |
| Spend Share Leaders | Mimo V2.5, V4 Flash, Hy3, and peers | Claude Sonnet 4.6 / Opus 4.7 at 13.5% each; GPT-5.5 at 11.6% |
| Pricing Logic | Up to 35x cheaper; absorbs massive fault-tolerant traffic | Opus 5 holds at $5/$25; FrontierBench v0.1 at 43.3% |
On July 24, Anthropic shipped Claude Opus 5: 43.3% on the new FrontierBench v0.1 benchmark (GPT-5.6 Sol at 37.5%), pricing unchanged at $5/$25 — half of Fable 5's input tier. The message is clear: "I cost more because I'm worth it on the hardest tasks." See our Claude Opus 5 launch and Kimi K3 controversy coverage for the full breakdown.
Selection takeaway: Closed frontier models handle the hardest 5%–10% of tasks; Chinese open-weight models absorb the remaining 90%+ of throughput. That is the clearest signal in July's data — not a binary "who is stronger" debate.
Model rankings show which "brain" developers prefer; the application leaderboard shows what those brains actually do in production:
| Rank | Application | Category | Share (Approx.) |
|---|---|---|---|
| 1 | Hermes Agent | Personal / CLI Agent | ~45% |
| 2 | Kilo Code | Coding Agent | ~13% |
| 3 | OpenClaw | General Agent | ~9% |
| 4 | Claude Code | Coding Agent | ~6% |
| 5 | Descript | Content production | ~4.5% |
| 7–9 | Lemonade / ISEKAI ZERO / Janitor AI | Companion / roleplay (new entries) | ~2% each |
| 10 | Cline | Coding Agent (IDE extension) | ~1.7% |
Cline, Roo Code, and Kilo Code share the same codebase lineage across three forks — the "grandchild" Kilo Code has already overtaken ancestor Cline. OpenRouter and a16z's State of AI report finds creative roleplay alone accounts for more than half of total open-source model usage. If you only read enterprise AI coverage, you will miss this entire half of the market.
Based on July trends and industry signals, here is our August read — plus a price positioning table you can pair with our OpenRouter integration guide for hands-on setup:
| Model | Input $/M | Output $/M | Positioning |
|---|---|---|---|
| DeepSeek V4 Flash | ~$0.05–0.14 | ~$0.24–0.28 | Best value; top pick for Agentic coding |
| MiniMax M3 | $0.10 | $1.21 | Long-context / multimodal on a budget |
| GLM 5.2 | $0.45 | $3.31 | Near-Opus planning quality among open models |
| Kimi K3 | ~$3 | ~$15 | 1.4TB open weights, closed-tier capability |
| Claude Opus 5 | $5 (fast tier $10) | $25 (fast tier $50) | Closed frontier; strongest July benchmark scores |
Independent developers: OpenRouter is a fine sandbox for model evaluation. For coding tasks, start with DeepSeek V4 Flash and GLM 5.2; reserve Opus 5 / GPT-5.6 for steps that stall. Expect ~180–250ms latency from some regions — evaluate compliant domestic routing before production.
Engineering leads: Do not pick models from the volume leaderboard alone. Route by task type and add vendor safety track records to your selection scorecard.
Review openrouter.ai/rankings weekly: Log Top 12 and provider share shifts with the data cutoff date — the board moves daily.
Write gateway rules by scenario: Chat/creative to V4 Flash / Mimo V2.5; complex reasoning to Claude Opus 5; high-risk classification to Sonnet 4.6 / Opus 4.7.
Split token bills from dollar bills: If tokens cluster on Flash but dollars cluster on Claude, you are already tiered — codify the routing table explicitly.
Track the app layer, not just models: Coding Agents (Kilo Code, Claude Code) and roleplay traffic have very different model profiles — match models to your product's actual workload.
Keep a regression suite ready for August releases: When a new #1 appears, run the same eval within 48 hours and update routing — do not rewrite the application.
Pin down your Agent execution layer: Long-session CLI Agents and sensitive prefill belong on dedicated SSH nodes; let OpenRouter handle elastic peaks. See rental pricing for specs.
A laptop that sleeps or a cheap Linux VPS cannot sustain 12-hour Agent loops or run macOS-only tooling like xcodebuild and notarytool. Pairing weekly leaderboard reviews with a fixed execution environment is more sustainable than chasing a single "best model" every week.
For iOS CI/CD and AI Agent automation teams that need stable SSH sessions, Keychain isolation, and predictable bandwidth, the winning pattern is clear: define OpenRouter multi-model routing in your gateway and place heavy workloads on a dedicated cloud Mac rather than routing every token through public APIs. NodeMini cloud Mac Mini rental works as the Agent execution layer — when you rotate API keys or swap model endpoints, your SSH node and CI labels stay put. Setup details live in the Help Center.
As of July 25, 2026, Xiaomi Mimo V2.5 led daily token volume at roughly 1.4T/day, followed by DeepSeek V4 Flash (943.9B/day) and Tencent Hy3 (590B/day). Rankings shift daily — check openrouter.ai/rankings for live data.
No — different measurement windows (seven-day provider share vs overall token volume including free and long-tail models) and whether first-party app traffic is included explain the gap. The provider table here uses a multi-source seven-day approximation. The trend that matters: from under 2% to ~46% in twelve months. See our June analysis for the monthly series.
Use OpenRouter for elastic multi-model routing and billing visibility; run long-session CLI Agents and local prefill on a dedicated cloud Mac with fixed monthly cost. Setup and key configuration: Help Center. Hardware specs and pricing: rental rates.
(1) Whether combined Chinese open-weight share breaks 50%. (2) Whether Anthropic ships a lower-priced tier to chase volume. (3) Community Kimi K3 quantization landing. (4) US AI safety legislation and White House pre-release review frameworks affecting enterprise model selection.