DFlash 2 Hits 70 Tokens/Second on a MacBook, 4.6x Standard Decoding
A new speculative-decoding technique reportedly reaches 70 tokens per second on consumer Apple hardware, roughly 4.6 times the speed of standard autoregressive generation.
Local inference keeps getting faster. @TraffAlex reported that DFlash 2 hits 70 tokens per second on a MacBook — about 4.6 times the speed of standard autoregressive decoding. For anyone who has watched a local model dribble out text token by token, a near-fivefold speedup is the difference between a novelty and a usable tool.
Unlock the full briefing
Get every story in today's briefing, the full archive, and the daily AI intelligence brief.
All stories today
Full archive
Daily brief
Cancel anytime. Payments powered by Stripe.