Paying for Frontier AI Buys 4-Month Head Start at 5x Cost
Mozilla's State of Open Source AI report finds frontier AI models buy a 4.4-month head start at five times the per-task cost.
Frontier AI models command a premium that buys roughly four months of capability, according to a new analysis of the widening contest between closed American systems and open-weight rivals built largely in China. The price of that head start runs about five times higher per completed task.
Those numbers come from the latest State of Open Source AI report from Mozilla, published September 15. The findings land at an awkward moment for anyone writing checks to the biggest US labs, because the performance gap they are paying to maintain keeps shrinking.
The Gap Is Now Measured in Months
Mozilla puts the distance between frontier AI models from US tech companies and the best open-weights models from Chinese companies at just 4.4 months. That figure captures something companies are already acting on. Moonshot AI's Kimi K3, a leading open model, lands a composite score on the Artificial Analysis Intelligence Index that trails Anthropic's Fable 5 closed model by only three points. It costs 30 percent of what Fable 5 costs.
Raffi Krikorian is Mozilla's chief technology officer. He frames the buying decision as a matter of workload, not organizational philosophy, and when you sit with that distinction long enough you start to see why he keeps pulling the conversation away from abstract debates about open versus closed and back toward the far more practical question of what a given team is actually asking the model to do. A closed model earns its premium. It's narrow. But the places where it pays off are specific: expert professional work, high-intensity retrieval, and long context. And that's it.
The mechanics behind that gap involve how long a job a model can reliably finish. The research nonprofit METR defines a model's time horizon as the length of tasks, measured by how long human experts need to complete them, that a system can handle with a reliable 50 percent success rate. That horizon has been doubling on an accelerating cadence.
Seven Hours Versus Twelve
Right now the best closed model can handle a job 1.7 times as long as the longest one the best open model can finish reliably. Krikorian offers a concrete illustration. If the open frontier can handle a seven-hour job, the closed frontier handles a 12-hour one. Four months later, the open model takes the 12-hour job and the closed one reaches something around 20.
That creates a surprisingly narrow band where the premium makes sense. Tasks needing between eight and 12 hours are the ones a closed frontier model can do and open models cannot yet. Nothing on the market reliably handles work beyond 12 hours. Anything under eight hours can go to either type, which means it can go to the cheaper option.
"Pay when that head start is worth it, something like a deadline that lands before the open frontier catches up would be here. Routine work you'll still be doing next quarter is not, because you'll be able to do it for a fifth of the cost soon, and the model won't be the bottleneck anyway."
The comparison carries caveats. Closed frontier models often come with their own software layer, the part that helps a model reach tools and memory so it can act with some autonomy. Labs build these harnesses specifically for their own systems, and they can lift performance in ways third-party harnesses do not.
Benchmarking company Vals AI has tried to cancel out that advantage by running every model through its own neutral software layer. On the Terminal-Bench 2.1 evaluation, the open-weights model GLM 5.2 from Chinese company Z.ai scored within a point of Anthropic's Claude Opus 4.7 and 4.8 while costing about five times less per completed task.
Why Companies Still Write the Check
Organizations keep paying for closed frontier AI models for reasons that have little to do with raw scores. The systems work out of the box and arrive bundled with compliance packaging, support, and accountability. Many organizations simply lack the staff to run open-weights models well.

Open-weights models let anyone download the main model components and run them locally, but developers typically withhold training data, the data pipeline, and training code. US companies like Anthropic and OpenAI keep everything proprietary and charge accordingly.
DoorDash has already split the difference. The delivery company uses Kimi for routine work while reserving Fable for harder tasks that would take human experts longer to complete.
OpenRouter shows the adoption. It's the AI gateway and marketplace that gives developers access to hundreds of models. And eight of the top 10 models ranked by token volumes in August 2026 provide open weights, which is a striking detail when you consider how these rankings get compiled and what they say about where the market is actually heading.
Revenue tells the opposite story.
The Concentration Problem Nobody Wants to Name
The uncomfortable reality sits with geography. Most open models the world runs on are Chinese. Krikorian describes the dynamic as the same playbook American companies ran with Android: give the software away, own the ecosystem around it.
His concern is not Chineseness so much as concentration. The best open models cluster in China while the best closed frontier AI models come from US companies. He wants US and European labs competing in the same open lane so no single country sets the world's defaults.
Open-weight AI's plural ecosystem is largely funded by Chinese capital right now. Krikorian calls that a plurality and a concentration at once. He points to the history of open source infrastructure like Linux, funded by a coalition of organizations that each needed the commodity layer to exist, and it's a history that shows how shared dependence can hold a system together even when the money behind it comes from a single, powerful source. But that's the tension. They're not the same thing.
Any new coalition, he argues, would likely consist of institutions with a mission rather than a market. His sketch includes public compute programs funding fully open reference models, with Switzerland's national compute producing Apertus as one example. Foundations would hold neutral ground for models, harnesses, and agentic standards. Companies that benefit from commodity models would participate. Philanthropy would cover evaluation and audit infrastructure.
- Frontier closed models: roughly 4.4 months ahead, about 5x the per-task cost
- The sweet spot for paying: tasks taking eight to 12 hours
- Routine work under eight hours: hand it to open models
- Revenue split: 4 percent open, 96 percent closed, per older Linux Foundation data
Krikorian pushes the argument past open weights entirely. He contends it is hard to fully trust a model with decisions when you cannot tell how it was trained or what it was evaluated against.
For buyers, the practical read is blunt. If a deadline lands before the open frontier catches up, the premium is worth it. If the work will still be waiting next quarter, the math flips fast.
Frequently Asked Questions
How large is the capability gap between closed frontier AI models and the best open-weight models, and what does that premium cost?
Mozilla puts the distance between frontier AI models from US tech companies and the best open-weights models from Chinese companies at just 4.4 months. The price of that head start runs about five times higher per completed task.
According to Raffi Krikorian, which specific tasks justify paying the premium for a closed frontier model?
Krikorian frames the buying decision as a matter of workload, and a closed model earns its premium only in narrow areas: expert professional work, high-intensity retrieval, and long context. Tasks needing between eight and 12 hours are the ones a closed frontier model can do and open models cannot yet.
How does METR define a model's time horizon, and how does the closed-versus-open comparison play out in practice?
METR defines a model's time horizon as the length of tasks, measured by how long human experts need to complete them, that a system can handle with a reliable 50 percent success rate. Krikorian illustrates that if the open frontier can handle a seven-hour job, the closed frontier handles a 12-hour one, and four months later the open model takes the 12-hour job while the closed one reaches something around 20.
Why do organizations keep paying for closed frontier AI models even when open models score comparably?
Organizations keep paying for reasons that have little to do with raw scores: the systems work out of the box and arrive bundled with compliance packaging, support, and accountability. Many organizations simply lack the staff to run open-weights models well.
What is the concentration concern Krikorian raises about the open-weight ecosystem, and what coalition does he sketch as a response?
Krikorian's concern is not Chineseness so much as concentration, since the best open models cluster in China while the best closed frontier AI models come from US companies. His sketch includes public compute programs funding fully open reference models, foundations holding neutral ground for models, harnesses, and agentic standards, companies that benefit from commodity models participating, and philanthropy covering evaluation and audit infrastructure.
๐ฌ Comments (0)
No comments yet. Be the first!













