Princeton Gave Claude Opus 4.8 Six Days to Do Research. It Scored 2 Out of 6.

A Princeton evaluation handed a frontier model an open-ended research brief and graded the output like a paper. The result punctures the tidy story of recursive self-improvement — while a 27B specialist quietly outperformed the giants on replication.

The most sobering AI result of the weekend didn't come from a benchmark leaderboard or a launch event. It came from a grading rubric. As @Markestolle reported, a Princeton team gave Claude Opus 4.8 six days to work on open-ended research tasks, then graded the output the way you would grade a graduate student. The model scored two out of six. The conclusion was blunt: full automation of open-ended research is not on the horizon right now.

That sentence deserves to be read slowly, because it cuts against the dominant narrative of the past year. The pitch for frontier models has increasingly been framed around recursive self-improvement — the idea that models will soon conduct their own research, design their own successors, and compress the innovation cycle into something超-human. Princeton's exercise suggests the gap is not in raw capability but in judgment. The models can execute. What they lack is the taste to know which questions are worth asking, which dead ends to abandon, and when a result is actually interesting rather than merely correct.

Get our free daily newsletter

Get this article free — plus the lead story every day — delivered to your inbox.

Want every article and the full archive? Upgrade anytime.

No spam. Unsubscribe anytime.