DeepMind Made Four Frontier Models Play Civilization VI. They Couldn't Remember What They Were Doing Ten Minutes Ago.

A study forcing Claude, GPT, Gemini, and Kimi to complete a full game of Civilization VI exposed two structural failures — situational blindness and what one observer called 'organizational ADHD' — that raw model intelligence did nothing to fix.

The most instructive AI research of the week didn't come from a benchmark leaderboard. It came from a video game. Researchers forced Claude, GPT, Gemini, and Kimi to play Civilization VI to completion, and the results, as @thesupermanmx documented in a thread that went viral, read less like a capabilities showcase and more like a diagnosis of everything agentic systems still get wrong.

The first failure the researchers catalogued was total situational blindness. The models could reason brilliantly about any single decision presented in isolation, but they repeatedly lost track of the actual board state — where their units were, what their opponents were doing, what threats were forming three tiles away. Civilization is a game of persistent, evolving context. The models treated each turn as if it had walked into the room cold.

Get our free daily newsletter

Get this article free — plus the lead story every day — delivered to your inbox.

Want every article and the full archive? Upgrade anytime.

No spam. Unsubscribe anytime.