GLM-5.3 Flash Goes Open-Weight: 320B Multimodal Model, 18B Active, 1M Context, Runs Locally

Chinese labs continued their open-weight push this week, with GLM-5.3 Flash and a 770B Tencent Hunyuan model narrowing the gap between local and closed-API capability.

The open-weight frontier moved again on Thursday. GLM-5.3 went open-weight as a 320-billion-parameter, natively multimodal model with 18 billion active parameters and a 1-million-token context window, according to @EricZhou866. The "Flash" designation, as @marcopapa99 noted, signals optimization for local inference and speed — the sparse mixture-of-experts design keeps active parameters low enough to run on accessible hardware while total capacity stays high.

Unlock the full briefing

Get every story in today's briefing, the full archive, and the daily AI intelligence brief.

All stories today

Full archive

Daily brief

Cancel anytime. Payments powered by Stripe.