Alibaba’s artificial‑intelligence research unit announced the preview model Qwen3.8-Flash-Next on Thursday, offering a glimpse of the architecture it plans to scale for the forthcoming Qwen4. The model packs 125 billion parameters but fires only 6 billion for each token it generates, a deliberate design choice aimed at curbing inference costs as large‑context, agentic workloads become commonplace.

Compared with the team’s earlier flagship, Qwen3.7-Plus, which holds 397 billion parameters and activates 17 billion per token, the new system runs on roughly a third of the active compute. The efficiency gain stems from three conventional tweaks and one more unconventional addition.

The first three changes are familiar to the LLM community. Qwen3.8‑Flash‑Next employs a new sparse‑attention mechanism that operates on micro‑blocks instead of selecting individual tokens, reducing the number of attention calculations. A gated residual pathway controls the flow of information between layers, and the training recipe eliminates batch‑size warm‑up, streamlining the optimization process.

The fourth modification departs from typical expert‑layer designs. Rather than adding more specialized experts, the model attaches a 51 billion‑parameter embedding matrix indexed by two‑ and three‑character fragments. According to the Qwen team, this approach lowers computational overhead and eases offloading to accelerators with limited memory, a practical response to export‑control restrictions that have limited Chinese labs’ access to top‑tier chips.

Alibaba’s own benchmarks back the performance claims, though external verification remains pending. The company previously billed Qwen3.8 as the world’s second‑best model, a statement reported by The Next Web without independent proof. The model card also notes that its Humanity’s Last Exam score was graded by GPT‑4o rather than the benchmark’s official grader, a transparency move that signals candid reporting but does not replace third‑party evaluation.

Licensing details could shape the model’s adoption in Europe. Qwen3.8‑Flash‑Next’s weights are hosted on Hugging Face under a “qwen‑community” licence that permits commercial use for large customers in exchange for fees. The European AI Act distinguishes between truly free and open‑source software and versions that monetize components; Article 53(2) and Recital 103 suggest that the fee‑based model may not qualify for the exemption that relieves developers of certain documentation duties.

European firms appear undeterred. Thomson Reuters, for example, has already built a proprietary model on top of Qwen, meaning any future licensing shift will affect downstream users. The industry is watching closely as regulators decide whether the Qwen licence satisfies the AI Act’s open‑source criteria.

Dieser Artikel wurde mit Unterstützung von KI verfasst.
News Factory APP - agentische News für besseres SEO & AEO.