Market context for this story
Loading quotes…
Informational only — not investment advice. Full markets →

Claude Opus 5 Tops AI Intelligence Leaderboard: What Investors Should Watch
💡 Watch for enterprise customers that adopt Claude Opus 5 for high-intelligence tasks – premium pricing could boost Anthropic's revenue. Track speed and latency leaders like Gemini 3.5 Flash-Lite and Mercury 2 for real-time applications; companies offering the lowest latency often win in customer-facing AI. Keep an eye on open-weight models from Meta and others – commercial use restrictions can limit enterprise adoption, creating opportunities for fully open alternatives.
Anthropic's Claude Opus 5 has claimed the top intelligence slot on the Artificial Analysis Intelligence Index, while new speed and cost benchmarks reveal shifting competitive dynamics. Investors can track which players lead in intelligence, speed, latency, and price to identify potential winners and disruptors in the AI arms race.
The latest Artificial Analysis Intelligence Index v4.1 ranks Claude Opus 5 (max) and Claude Opus 5 (xhigh) as the highest intelligence models, followed closely by Claude Fable 5 (with fallback) and GPT-5.6 Sol (max). The index evaluates 586 models using nine benchmarks including GDPval-AA v2, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, and GPQA Diamond. Intelligence alone does not dictate market value, but leadership here often translates into premium pricing power and enterprise adoption.
Speed leaders tell a different story. Mercury 2 leads output speed at 902 tokens per second, with Gemini 3.5 Flash-Lite at 435 t/s, while HyperNova 60B 2605 and Granite 4.0 H Small follow. Gemini 2.5 Flash-Lite also takes the lowest-latency crown at 0.34 seconds, closely followed by Command A+ at 0.42 seconds. These metrics matter for real-time applications like chatbots, code assistants, and agentic workflows.
On the cost side, Devstral 2 and North Mini Code both price at $0.00 per million tokens, with Gemma 3 variants rounding out the cheapest options. Meanwhile, context window size is a differentiator for long-document processing: Llama 4 Scout supports 10 million tokens, and Grok 4.20 0309 handles 2 million tokens. Larger context windows enable more complex reasoning tasks and could unlock new enterprise use cases.
The index also incorporates agentic and coding-specific evaluations, including 𝜏³-Banking, AA-Briefcase, and Harvey LAB-AA. These metrics assess real-world tool use, legal work, and business operations automation. Companies building AI-powered automation tools may benefit most from models that excel in agentic benchmarks, not just raw intelligence.
For investors, the leaderboard creates a clear signal: intelligence leadership alone doesn't guarantee market dominance. The most attractive quadrant in the cost-per-task vs. intelligence graph includes players like Google, Anthropic, OpenAI, NVIDIA, Alibaba, and DeepSeek. Speed, latency, and cost metrics will determine which models get deployed at scale.
Based on reporting from hn-frontpage.
Read the full story
Original reporting and related coverage — attribution links only, not paid recommendations.
Broker buttons use invite / refer-a-friend links (rewards may be capped). Other partner links may pay OppHub America a commission at no extra cost to you.
OppSHOP
Full OppSHOP →Curated tools and reads — shopping here helps keep OppHub America free.
Playbook
New stories get a playbook when they publish. Older articles may not have one yet.
No stored playbook for this article. Going forward, playbooks are generated once at publish and kept on the story.