· 5 min read

China's Open Models Now Out-Download America's 2 to 1. Here's Which One a Solo Stack Should Actually Route To.

AI researcher Nathan Lambert published a data-driven case this week that Chinese open-weight models have decisively pulled ahead of their American counterparts, and the numbers are specific enough to change how I'd advise a solo builder to route open-model traffic today. Chinese models hold about 3.2 billion Hugging Face downloads against 1.6 billion for American open models. On the Artificial Analysis Index, the leading Chinese models score 42 to 45 against 23 to 26 for the best American open options. OpenRouter usage data shows more than 80% of open-model traffic already going to Chinese-origin models.

The actual numbers

On the Chinese side: Alibaba's Qwen family, Z.ai's GLM-5.2 and 5.3, Moonshot AI's Kimi K3, and DeepSeek's models. On the American side: Meta's Llama, OpenAI's gpt-oss, Google's Gemma 4, Nvidia's Nemotron, and Thinking Machines' Inkling. Lambert's piece also notes Chinese models now appear in roughly 40% of ML papers against 30% for American open models, which is a slower-moving but real signal that the research community's default reference point has shifted too.

The part I'd flag before treating any of this as settled: Lambert estimates that distillation from American frontier labs accounts for only one to two months of the performance gap, meaning the case he's making is for genuine technical progress at Chinese labs, not primarily a story about Chinese teams training on outputs from GPT or Claude and calling it their own. That's a real distinction if you're deciding whether this trend is durable or a temporary artifact of a training-data shortcut. I haven't independently reproduced his methodology, and a single researcher's benchmark analysis, however well-sourced, is one data point rather than a consensus. But the download and OpenRouter usage numbers are close to directly observable facts, not model performance claims subject to benchmark-gaming concerns, and those alone already tell you where the actual traffic is going.

Why this is a stack decision, not a geopolitics story

I don't cover this to weigh in on US-China AI competition as a policy question. I cover it because if you're a solo builder currently defaulting to Llama for a self-hosted feature, or routing OpenRouter traffic to whatever the "safe American choice" is out of habit, the benchmark and cost case for that default no longer holds up on the numbers. If GLM or Qwen scores meaningfully higher on the tasks you actually care about, at a lower inference cost, sticking with Llama because it's the familiar name is leaving real performance and real margin on the table.

That said, this isn't a costless switch. Licensing terms differ meaningfully between these model families, and you need to actually read them rather than assume open-weight means unrestricted-use. Data residency matters if you're processing anything remotely sensitive through a hosted Chinese-origin model API versus a self-hosted weights file you control directly, those are very different risk profiles even when the underlying model is identical. And there's a real, separate conversation about running weights from a lab you can't audit versus one under different regulatory jurisdiction, which is a business risk question independent of raw benchmark performance.

What I'd actually do

If you're self-hosting weights, the licensing question is the one to resolve first: check whether the specific Qwen, GLM, Kimi, or DeepSeek release you're evaluating carries usage restrictions that matter for your product, some do, some are genuinely permissive, and it varies by model and by version. If you're routing through OpenRouter or a similar aggregator, the calculation is more straightforward: run your own benchmark against your own real prompts, not the Artificial Analysis Index's general tasks, and compare actual output quality and cost per request across two or three candidate models before defaulting to whichever one is topping a leaderboard this month. Leaderboards move fast in this space, and the model that's ahead in September isn't guaranteed to hold that position by year end.

The honest counter-take: OpenRouter's 80%+ usage figure could partly reflect price sensitivity among OpenRouter's specific user base rather than a universal verdict on model quality, developers routing through an aggregator skew toward cost-optimizing, which is exactly the population most likely to pick the cheaper Chinese-origin option regardless of the raw benchmark gap. Usage share and quality leadership are correlated here, but they're not the same claim, and it's worth keeping them separate when you're the one making the actual routing decision.

Author

Sources

Stay in the Loop

Get new posts delivered to your inbox. No spam, unsubscribe anytime.

Newsletter coming soon. Set PUBLIC_CONVERTKIT_FORM_ID in .env to activate.

Related Posts