DeepSeek V4 Flash Runs Agentic Workloads Locally at 31 Tokens/Sec on an M3 Ultra

The open-model camp keeps closing the gap: DeepSeek's V4 Flash is drawing attention for single-box agentic performance that runs entirely on consumer-grade hardware.

While the frontier labs chase million-token budgets, the open-model camp is winning a different race — local, private, cheap. @marcopapa99 reports that DeepSeek V4 Flash "looks nasty on single-box agent workloads," running locally on an M3 Ultra at 31 tokens per second. That last figure is the one that matters.

Unlock the full briefing

Get every story in today's briefing, the full archive, and the daily AI intelligence brief.

All stories today

Full archive

Daily brief

Cancel anytime. Payments powered by Stripe.