Four frontier models shipped in five days — Anthropic's Claude Fable 5.1, Google's Gemini 3.8 Flash, Meta's Muse Spark 1.3, and OpenAI's GPT-6 Astra. Each one arrived with a chart showing it was the best. By Friday, Astra's headline ARC AGI score had two different values — 99.9% using OpenAI's harness, 62.7% using the standard one everyone else runs. Fable 5.1 cut cash prices 75%, but the per-task bill went up 20% because it thinks harder. Meta's Muse Spark topped the coding leaderboard but produced the weakest game when a reviewer gave all four models the same brief. The benchmark shortcut just broke three ways at once.
The same week, Apple named John Ternus — a hardware engineer — as its new CEO for the AI era, while data shows 88% of incoming S&P 500 CEOs were promoted from inside, and only 6% of new Fortune 500 directors have ever run tech in the C-suite. Two instruments for choosing broke at once. Replace them with your own tests and your own bench: cost per completed task on your own work, and a succession plan dated against your AI roadmap.