Grok Claims Top Spot in Legal and Government Benchmarks on Chatbot Arena

xAI's Grok-4.20 reportedly ranked first in the Legal & Government category on Chatbot Arena, outperforming Anthropic's Opus 4.6 and Google's Gemini 3.1 Pro — though the claim comes amid executive upheaval at the company.

Grok-4.20 has taken the top position in the Legal & Government category on Chatbot Arena, according to @XFreeze, outperforming both Anthropic's Opus 4.6 and Google's Gemini 3.1 Pro. The result is notable because legal reasoning has historically been one of the hardest domains for language models — it requires precise citation, nuanced interpretation, and low hallucination rates.

Unlock the full briefing

Get every story in today's briefing, the full archive, and the daily AI intelligence brief.

All stories today

Full archive

Daily brief

Cancel anytime. Payments powered by Stripe.