xAI's Grok Sweeps Benchmarks, Lands Federal Contract, and Ships a Voice Agent — All in One Week
Grok 4.20 claimed the top spot across multiple reasoning benchmarks while xAI simultaneously secured USDA adoption and deployed a voice AI agent handling live Starlink customer support calls.
xAI is having the kind of week that forces competitors to update their strategy decks. Grok 4.20 Reasoning has taken the number-one position on BridgeBench, a reasoning benchmark, beating GPT-5.4, Claude Opus 4.6, and Google Gemini, as documented by @cb_doge. The same account noted Grok leads across multiple global benchmarks including AA Omniscience and IFBench, with particular strength in speed, reasoning accuracy, and hallucination control.
Unlock the full briefing
Get every story in today's briefing, the full archive, and the daily AI intelligence brief.
All stories today
Full archive
Daily brief
Cancel anytime. Payments powered by Stripe.