Loading prices...
All news
Flat vector illustration of a small glowing multifaceted crystal emitting a powerful beam of light toward a much larger dim geometric sphere on a dark navy background, symbolizing a compact AI model outperforming much larger rival systems

Alibaba's Qwen3.8-Flash-Next beats bigger models on a fraction of cost

23:40 · 26.08.2026
Source: The Decoder
0

Alibaba's Qwen team released Qwen3.8-Flash-Next, a multimodal mixture-of-experts model built as an architecture preview of the coming Qwen4, The Decoder reported. The model carries 125 billion total parameters but activates only 6 billion per token, the kind of ratio that lets it chase performance closer to much larger, much more expensive systems.

Part of that efficiency comes from a novel N-gram embedding layer, one of the architecture changes slated for Qwen4 itself. The layer stores common word groups as standalone entries in something like a phrase dictionary, and it can run in regular system RAM instead of GPU memory at relatively low added cost. It accounts for 51 billion of the model's parameters on its own.

Qwen3.8-Flash-Next natively handles a 262,144-token context window and can stretch to one million tokens using YaRN. The technical report sits on GitHub, and the weights are available on Hugging Face and ModelScope. The production version ships as Qwen3.8-Flash through QwenCloud, priced at $0.16 per million input tokens and $0.47 per million output tokens, with the API expected to go live shortly.

The efficiency story extends to training itself. Qwen's team says Flash-Next beats Qwen3.7-Plus, a 397-billion-parameter model that activates 17 billion parameters per token, nearly three times Flash-Next's active count, while costing roughly one-ninth as much to train. The biggest gains land in coding and office tasks.

Alibaba's published benchmarks put Flash-Next against DeepSeek-V4-Flash (284 billion parameters, 13 billion activated) and Anthropic's Claude Opus 4.6 Max. Both rivals are larger, more expensive, or both, yet Flash-Next leads on most of the tested tasks.

Agentic coding benchmarks, which require a model to independently find and fix bugs in real software projects, show the clearest edge. Flash-Next scored 58.7 on DeepSWE and 62.5 on SWE-bench Pro, ahead of both DeepSeek-V4-Flash and Claude Opus 4.6. The gap widens further on office and productivity work: Flash-Next hit 73.9 on CoWorkBench against DeepSeek-V4-Flash's 45.1, and scored 55.7 on JobBench, nearly double Qwen3.7-Plus's 27.6. Scientific reasoning and competitive programming stay close across the board, with GPQA Diamond at 91.7 and LiveCodeBench v6 at 91.9.

  • Parameters: 125B total, 6B activated per token (Flash-Next); 51B in the N-gram embedding layer alone
  • Context window: 262,144 tokens natively, up to 1 million via YaRN
  • DeepSWE: 58.7 (Flash-Next) vs Claude Opus 4.6 and DeepSeek-V4-Flash, both lower
  • CoWorkBench: 73.9 (Flash-Next) vs 45.1 (DeepSeek-V4-Flash)
  • JobBench: 55.7 (Flash-Next) vs 27.6 (Qwen3.7-Plus)
  • QwenCloud pricing: $0.16 per million input tokens, $0.47 per million output tokens

Claude Opus 4.6 pulls ahead on one benchmark, Humanity's Last Exam, a test built around extremely hard multidisciplinary problems, but it's also an older Anthropic release from February 2026.

Flash-Next sits below Alibaba's own current flagship, Qwen3.8-Max, which launched in early August to compete with Claude Opus 4.8, Gemini 3.1 Pro, and GPT-5.6 Sol. Intokened covered how close that flagship came to Claude Opus 4.8 on release, though Kimi K3 still edged past both on raw benchmark scores. Flash-Next performs slightly below Qwen3.8-Max itself while costing about one-twelfth as much: the flagship runs $2.00 per million input tokens and $6.00 per million output tokens, against Flash-Next's $0.16 and $0.47.

That pricing keeps squeezing OpenAI and Anthropic from below. Qwen3.8-27B has already built a following for running locally at strong performance and minimal cost, assuming a user has the hardware for it. OpenAI answered with steep discounts across its new GPT-5.6 line, a move that helps users but works against the rapid revenue growth AI labs need to keep their investment case intact.

Nothing here should be taken as financial advice — just information to consider.

Published: 23:40 · 26.08.2026
Maks

Author

Maks

Trading man

I've been interested in the cryptocurrency market for a long time, am a trader, and write articles and news about my experience and crypto in simple terms.

Comments (0)

No comments yet — be the first!