# Claude Opus 4.8 on the Cotera agent benchmark

> 5/5 at $4.58. The most expensive perfect score, and the only one we actively don't recommend.

Source: https://cotera.co/benchmarks/models/claude-opus-4-8

---

Anthropic's Opus 4.8 — the top of the Claude 4.X family, frontier reasoning at frontier prices. Marketed for agent work; we ran it on agent work and it got every answer right.

**Score:** 5/5 benchmarks passed · **Cost across the matrix:** $4.58

## Strengths
- 5/5 pass with the deepest verification chains of any model. Opus made 17 tool calls on the Crunchbase brief — the most thorough Sales run in the matrix.
- Best multi-source corroboration. Opus consistently cross-checked Crunchbase data against the company's own marketing page before returning a number.
- Zero JSON drama. If anything, Opus's explanation fields are the most carefully written prose in the matrix.

## Watch-outs
- Crunchbase was $1.22 for one run. Reddit was $1.74. CX was $1.00. Per-task costs that would be a wash on a customer migration are unsustainable as the agent-loop default model.
- Total spend was $4.58 — 17x Mistral, 12x GPT-5 Mini. We saw no quality lift on this rubric that would justify the multiple.
- Opus's verification instinct is great for high-stakes singletons (one-off research, regulated work). Don't put it on a job queue.

