DeepSeek V4 Flash Runs Agentic Workloads Locally at 31 Tokens/Sec on an M3 Ultra
The open-model camp keeps closing the gap: DeepSeek's V4 Flash is drawing attention for single-box agentic performance that runs entirely on consumer-grade hardware.
While the frontier labs chase million-token budgets, the open-model camp is winning a different race — local, private, cheap. @marcopapa99 reports that DeepSeek V4 Flash "looks nasty on single-box agent workloads," running locally on an M3 Ultra at 31 tokens per second. That last figure is the one that matters.
Unlock the full briefing
Get every story in today's briefing, the full archive, and the daily AI intelligence brief.
All stories today
Full archive
Daily brief
Cancel anytime. Payments powered by Stripe.