This article was written with the assistance of AI.
News Factory APP - agentic news to boost your SEO & AEO.
Alibaba Unveils Qwen3.8-Flash-Next, a 125‑Billion‑Parameter Model with 6‑Billion‑Parameter Active Compute
Key Points
- Alibaba released Qwen3.8-Flash-Next, a 125 B parameter model that activates only 6 B parameters per token.
- Active compute drops from 17 B in Qwen3.7-Plus to roughly one‑third in the new model.
- Four architectural changes: micro‑block sparse attention, gated residuals, no batch‑size warm‑up, and a 51 B embedding indexed by short character fragments.
- Design targets memory‑constrained accelerators, a workaround for export‑control limits on Chinese hardware.
- Weights are available on Hugging Face under a qwen‑community licence that may charge large commercial users.
- EU AI Act’s open‑source exemption could be denied because the licence includes monetized components.
- Thomson Reuters has already built a model on Qwen, indicating immediate industry uptake.