OpenRouter Rankings July 2026: Who's Actually Winning the AI Model Race

If you still pick models in OpenRouter, Cursor, or a custom agent stack based on impressions from a few months ago, you are likely behind. This article anchors on OpenRouter's real paid-token data as of July 25, 2026: Xiaomi's surge to #1, Chinese providers at ~46% share, the usage-vs-quality barbell, the app-layer agent and roleplay hidden market, August trend forecasts, and a 5-step tiered routing Runbook.

Abstract data flow and neural network visualization representing OpenRouter global developer model traffic statistics

Table of Contents

1. Three Selection Pain Points: Rankings, Quality, and Real Workloads Are Out of Sync

  1. Treating the token volume leaderboard as a quality leaderboard. OpenRouter ranks real developer spend on routed tokens, not benchmark scores. Cheap, fast models behind high-traffic apps can dominate the chart while barely appearing on hard reasoning workloads.
  2. Ignoring the app layer that drives model traffic. Hermes Agent alone accounts for ~45% of app-layer tokens; the Cline → Roo Code → Kilo Code fork chain shows open-source coding agents iterate fast. Model rankings alone do not tell you what those brains are actually doing.
  3. Hard-coding a single model cannot keep up with monthly champion rotation. February's MiniMax M2.5, April's MiMo-V2-Pro, July's Mimo V2.5 — the monthly #1 seat keeps changing, while DeepSeek remains the most stable provider at ~16%–18% share.

2. July Rankings: Xiaomi Takes #1, Chinese Models Hold ~46% Share

Source: OpenRouter official rankings (openrouter.ai/rankings), data as of July 25, 2026. Rankings update daily — verify live data before production decisions.

Top 12 Models by Daily Token Volume

RankModelProviderDaily Tokens30-Day Total
1Mimo V2.5Xiaomi1.4T31.2T
2DeepSeek V4 FlashDeepSeek943.9B23.6T
3Hy3Tencent590B23.4T
4Nemotron 3 Ultra 550B (free)NVIDIA428.6B9T
5DeepSeek V4 ProDeepSeek413.7B11.6T
6GLM 5.2Z.ai316.7B13.3T
7MiniMax M3MiniMax262.5B15.1T
8Step 3.7 FlashStepFun204.8B5.9T
9Kimi K3Moonshot157.6B1.6T (new entry)
10Ling 3.0 FlashAnt InclusionAI128.3B417.3B
11Gemini 3 Flash PreviewGoogle106.3B4T
12Claude Sonnet 5Anthropic99.5B3.6T

Seven of the top ten models are from Chinese providers; Nemotron 3 Ultra, Claude, and Gemini hold positions for the US camp. On July 24, Claude Opus 4.8 ranked #10; by the next day Ling 3.0 Flash had pushed it out of the top 12 — the board moves fast.

Provider Share (Blended 7-Day Window, Approximate)

ProviderRegionToken Share (Approx.)
DeepSeek🇨🇳 China16%–18%
Xiaomi🇨🇳 China8%–18% (Mimo V2.5 surge — highest volatility)
Anthropic🇺🇸 US10%–15%
Tencent🇨🇳 China8%–13%
Google🇺🇸 US8%–13%
Z.ai🇨🇳 China4%–7%
OpenAI🇺🇸 US6%–8%
NVIDIA🇺🇸 US5%

Chinese providers combined: ~46% (under 2% a year ago); US providers combined: ~30%–36% (about 70% a year ago). Price is the main driver: DeepSeek V4 Flash input runs ~$0.05–0.14/M vs GPT-5.5 at ~$5/M — a ~35× spread.

3. The Other Side of the Leaderboard: High Usage ≠ High Quality

By spend share (not token volume) across task types: general chat 35.7%, agent workflows 30.4%, code 26.5%, data processing 7.5%. On classification and complex reasoning, the picture flips — Claude Sonnet 4.6 and Claude Opus 4.7 each hold ~13.5% spend share, tied for #1; GPT-5.5 follows at ~11.6%. Cheap open-source volume leaders barely register in this dimension.

The market is barbelling: inexpensive Chinese open models absorb massive, low-barrier, fault-tolerant throughput (chat, creative, roleplay, simple coding); closed-source flagships retain pricing power on high-value, low-tolerance hard tasks. Anthropic's July 24 Claude Opus 5 release scored 43.3% on FrontierBench v0.1 (GPT-5.6 Sol 37.5%), priced at the Opus tier $5/$25 per M — "expensive, but worth it."

4. App-Layer Secret: Coding Agents Rule, Roleplay Is the Invisible Half

Model rankings show which brain is popular; app rankings (openrouter.ai/apps) show what that brain is used for:

RankAppTypeShare (Approx.)
1Hermes AgentPersonal agent / CLI agent~45%
2Kilo CodeCoding agent~13%
3OpenClawGeneral agent~9%
4Claude CodeCoding agent~6%
5DescriptContent production~4.5%
6piAgent~3.3%
7–9Lemonade / ISEKAI ZERO / Janitor AICompanion / roleplay~2% each
10ClineCoding agent (IDE plugin)~1.7%

Cline → Roo Code → Kilo Code share the same code lineage across three forks — the "grandchild" Kilo Code has already overtaken its ancestor. Roleplay and companion apps (Janitor AI, ISEKAI ZERO, SillyTavern, HammerAI) together drive a substantial share of open-model traffic; OpenRouter × a16z State of AI reports creative roleplay accounts for more than half of total open-model usage — a market enterprise AI coverage rarely mentions.

5. Pricing and Positioning Comparison

ModelInput / MOutput / MContextPositioning
DeepSeek V4 Flash~$0.05–0.14~$0.24–0.281MBest value; top pick for agentic coding
Nemotron 3 Ultra$0.42 (free tier available)$2.61US open-source, NVIDIA ecosystem
MiniMax M3$0.10$1.21Long contextBudget pick for multimodal / image input
GLM 5.2$0.45$3.31Opus-class planning among open models
Kimi K3~$3~$151MLargest open weights (1.4TB), closed-tier capability
Claude Opus 5$5 (fast tier $10)$25 (fast tier $50)1MClosed flagship, strongest July benchmarks

6. August Trend Forecast

  1. Combined Chinese open-model share likely keeps climbing, potentially crossing 50% this year, unless US providers cut prices significantly — no sign of that yet.
  2. The "monthly champion model" seat will keep rotating: Xiaomi, DeepSeek, Tencent, Z.ai, MiniMax, and Moonshot are not slowing price wars or release cadence.
  3. Anthropic may launch a cheaper tier (Haiku-class) to capture volume; Opus 5 is the fourth flagship in two months (after Mythos 5, Fable 5, Sonnet 5) — a multi-tier full-coverage strategy.
  4. Kimi K3's 1.4TB weights may get community quantizations within 2–4 weeks, enabling real adoption by smaller teams; today the main beneficiaries are large labs and inference clouds (e.g., Fireworks AI).
  5. Security and compliance become new selection variables: the OpenAI unreleased model sandbox escape on Hugging Face, the US Congress "AI kill switch" bill, and the White House frontier-model pre-release review framework expected before August will raise "provider security reputation" weight in enterprise procurement.

7. Role-Based Takeaways

Indie Developers / Small Teams

Enterprise Technical Leaders

AI Agent / Coding Tool Builders

8. 5-Step Tiered Routing Runbook

Step 1 — Split tiers by task risk and fault tolerance

Throughput / fault-tolerant → DeepSeek V4 Flash, Mimo V2.5; hard reasoning / high-risk agents → Claude Opus 5.

Step 2 — Configure primary model and fallbacks on OpenRouter

{ "agents": { "defaults": { "model": { "primary": "openrouter/deepseek/deepseek-v4-flash", "fallbacks": [ "openrouter/anthropic/claude-opus-5", "openrouter/z-ai/glm-5.2" ] } } } }

Step 3 — Calculate monthly bills against the 35× price spread

At ~10M input tokens/day: DeepSeek V4 Flash roughly $5–14/day vs GPT-5.5 ~$50/day — tiered routing ROI is obvious.

Step 4 — Move Gateway to a 24/7 Mac cloud node

Run OpenClaw under launchd with API keys in environment variables. See the OpenRouter API tutorial and June 2026 rankings analysis in this series.

Step 5 — Monthly review of rankings and your own eval set

openclaw doctor && openclaw channels status --probe # Cross-check openrouter.ai/rankings and update routes

9. Citable Technical Facts

10. Conclusion: Capability and Popularity Are Diverging

The one sentence worth remembering from July's leaderboard: capability and popularity are diverging. Chinese open models win traffic on price; US closed flagships guard pricing power and security reputation on hard tasks. Rather than obsessing over who is #1, invest in evaluation frameworks and tiered routing strategy.

Running a multi-model Gateway on a laptop or plain Linux VPS has real limits: sleep disconnects, no native Apple toolchain, harder ops. If you need OpenClaw or Cursor agents routing DeepSeek, Opus, and GLM 24/7, renting a VPSMAC M4 Mac cloud node is the more reliable production path — swap models as rankings shift, keep the runtime stable.