Mozilla Report: China-U.S. AI Model Capability Gap Narrows to 4.4 Months
Mozilla published the second edition of its State of Open Source AI report this month, and buried in the charts is a number that should worry anyone betting the farm on a closed-model moat: 4.4 months. That’s roughly how far behind the best open source AI models now trail the best closed ones, based on Mozilla’s own fit against METR’s time-horizon data. A year ago, the gap looked closer to a full year. It’s closing fast enough that “open source is always a generation behind” is starting to sound like something people say out of habit rather than observation.
The practical version of that stat: today’s closed frontier models can reliably handle tasks that would take a skilled human 8 to 12 hours to finish. Open models get there too, just about four months later. On raw capability scores the gap is even thinner. The best open model trails the closed leader by roughly three points on the Artificial Analysis Intelligence Index, while costing about 60 percent less to run.
The OpenRouter Tipping Point
If you want proof this isn’t purely academic, look at where developers actually send their traffic. Google held the top request-volume spot on OpenRouter for 51 straight weeks. Then in August, DeepSeek took over, the first time an open model has led the platform’s rankings outright. The same shift shows up on other infrastructure: Vercel’s AI Gateway data found DeepSeek overtaking Google for the number-two spot by token volume in July, running more than twice Google’s traffic on that platform. By Mozilla’s count, eight of the top ten models by token volume on OpenRouter that month were open weights, and seven of those eight were built in China. Increasingly, the search for the best open source AI is a search through Chinese labs’ release notes, not Silicon Valley’s.
That’s not an accident. Beijing has made open-weight proliferation explicit industrial policy, backing it with training slots pledged to the Global South and a state-linked coalition, WAICO, that has grown to 37 member countries and, tellingly, no major Western democracy.
The Fable 5 Blackout
Mozilla’s report doubles as a case study in why “closed and hosted” carries risks that have nothing to do with model quality. On June 9, Anthropic shipped Fable 5 and Mythos 5. Three days later, the U.S. Commerce Department barred foreign-national access to comply with export controls, and Anthropic pulled both models offline entirely rather than build a system that sorted users by nationality mid-session. They stayed dark for nineteen days, restored on July 1 once the restriction lifted.
Nobody outside Anthropic and the Commerce Department controlled that outage. Enterprises that had built products on Fable 5 got a hard lesson in what renting a model actually means: access can disappear with essentially no notice, for reasons that have nothing to do with a contract. Open weights don’t fail that way. As Mozilla frames it, you can switch off a model, but you can’t switch off a copy already running on hardware someone else owns.
Sovereignty, Not Just Savings
That distinction is doing most of the work in Mozilla’s adoption numbers. The report’s developer survey puts open-model usage at 79 percent, ahead of closed models at 71 percent (plenty of teams run both, and about half use them side by side). The reasons developers give for sticking with open weights track less with ideology and more with control: dodging cloud egress fees that can run $90,000 to $120,000 per petabyte, avoiding the metered-pricing shocks that have blown up AI budgets at companies like Uber, and not wanting a single vendor or government holding a kill switch over production infrastructure.
None of this means closed models are finished. Fable 5 still beats Moonshot’s Kimi K3 by a wide margin on expert-level knowledge work, and frontier labs retain an edge in long-context reliability that open models haven’t matched yet. But the case for treating open weights as the backup plan rather than the default gets weaker every quarter. Mozilla’s framing reads more like a warning than a celebration: the capability gap resets every release cycle, but the sovereignty gap, once it opens in the wrong direction, doesn’t close on its own.
