The Supercomputer Era of AI Is Over

Okay, so the supercomputer era of AI is over. Let me explain what I'm on about here.
For the last two years, we've been repeatedly impressed by the capabilities coming out of frontier labs like OpenAI, Anthropic, and, in many cases, Google, especially on creative outputs like video and images. But we've been thinking about their capabilities the wrong way, similar to how we viewed supercomputers in the 1990s and early 2000s. We keep throwing massive projects at frontier models, and we express envy when increasingly sophisticated ones come out that can accomplish them. That framing is already starting to change, and the frontier labs are in for some trouble.
The Frontier Labs Are Feeling It
First, a bit of data. These companies are beginning to feel the headwinds. OpenAI felt the pressure first, six months ago, when Claude surpassed them in growth. Now Anthropic is experiencing the same struggle, especially with large enterprises limiting engineers' open access to AI agents, and with the semiconductor market's stagnation upon the launch of Anthropic's Fable model. There are also repeated reports from insiders that spending growth for both companies is dramatically slowing down.
On the other hand, open-weight models, particularly those coming out of China, are drastically outperforming on platforms like OpenRouter and other AI model routing and inference providers. Why is that? I believe it's because the way people use AI is fundamentally changing.
AI Is Having Its Data Warehouse Moment
A good way to conceptualize this is to compare it to a data warehouse product. As companies like Snowflake and Databricks grew, they initially operated largely as analytics query engines. Most transactions involved users inputting data, and the data ingressed and egressed, with complex computations performed by individuals at keyboards. That's how people interact with AI today: an assistant-based method where users either code or talk to chat support bots.
However, we believe at Cotera, and are beginning to see, that most people will be running smaller workflows that don't require as much context, in parallel, across various types of agents. Cost becomes crucial as that becomes popular. Right now, we estimate about 90% of use cases involve someone chatting with a bot or actions driven by human intervention. In the future, that flips: from only 10% happening in the background, triggered by some external event, to 90%.
When that happens, cost is paramount, because suddenly there's time to think about a system rather than just reacting ad hoc. And we have found that open-weight models perform quite well here.
What Our Benchmarks Actually Show
Mistral's Large model series (most recently Mistral 3.5 Large, which we haven't run through our official benchmarks yet) is an impressive product. In Cotera's testing, Mistral Large 3 currently sits at the top of our leaderboard for what we describe as real-world use cases: background agent tasks like enriching contacts, conducting deep-dive intelligence summaries on people, writing creative briefs, using tools to build presentations, and generating tickets from calls and follow-up materials.
Tasks passed, all 24 models
One square per task: Sales, Marketing, CX, Coding, Web Scraping. ● marks open-weight models.
NVIDIA's Nemotron 3 Super also excels under the same conditions, passing four of five tasks for about 60 cents total. The interesting exception is the Chinese open-weight models: Qwen went zero for five, and even the stronger ones (the GLM series, Kimi, DeepSeek, MiniMax) topped out at three of five, mostly by stalling out mid-run before writing a final answer.
The Cost Gap Is Not Subtle
When it comes to cost, Mistral and NVIDIA come in at one-tenth to one-twentieth of their frontier counterparts while doing the same job. A full 5/5 sweep of our benchmark suite cost Mistral Large 3 $0.27. Claude Sonnet 4.6 spent $3.83 to do the same work. Opus 4.8 spent $4.58. That's a 17x difference for identical output, graded by the same rubric.
What a clean 5/5 sweep costs
Total spend across all five agent tasks, models that passed everything. Green bars are open-weight models.
I've written before about how price per token isn't the number that matters, price per correct answer is. This is that argument playing out at the model-selection level instead of the prompt level. If you're curious how deep the cost obsession goes here, we also switched our agent payloads from JSON to TOON and saved 30% on tokens.
AI Is a Commodity, Not Software
We at Cotera believe that AI is not software, contrary to what much of the venture capital world asserts. AI increasingly resembles a commodity, similar to telecom, where switching costs are relatively low and privacy and security are the real concerns. That seems to be the trajectory the world is taking.
In our own operations, we recently moved off Claude's Sonnet and Opus models in favor of Mistral Large 3 to onboard most of our customers and explore their use cases. Our customer support teams use it to QA tickets. Sales teams use it to create follow-up presentations. Product teams use it to analyze Sentry errors and Amplitude events and create Linear tickets to identify areas for improvement.
A lot of the work that follows is still done by individuals using Claude, but we don't expect that to last forever. The era of the supercomputer in AI is coming to an end, and anyone still maximizing tokens or relying solely on frontier models is going to face a real shock as the world moves on without them.
Try These Agents
The background tasks I mentioned above are all things you can run on Cotera today:
- Ticket QA Grader — grade support tickets against your QA rubric automatically
- Lead Enrichment Agent — enrich new contacts with firmographic and intent data in the background
- Auto-Generate Sales Deck from CRM — turn CRM data into follow-up presentations without opening Slides
- Support Ticket → Linear Bug — convert recurring support issues into Linear tickets your product team can act on