Qwen3.8 Flash Next
Qwen3.8-Flash-Next is Alibaba Qwen's open-weight multimodal MoE preview of the architecture planned for Qwen4. Its language model has 125B parameters with 6B activated per token, plus 51B n-gram embeddings and 4B MTP parameters. It combines Gated DeltaNet, Qwen Sparse Attention, gated residuals, vision input, tool use, flexible reasoning, a native 262K context window extensible to 1M tokens, and official serving support through vLLM, SGLang, KTransformers, and TokenSpeed.
Deployment and license note
Qwen3.8 Flash Next is compute-efficient, not small: only 6B parameters activate per token, but the 125B main model, 51B n-gram embeddings, and auxiliary MTP weights still need to be stored. A community Q4 build is roughly a 128GB system-memory class deployment, while official FP8/BF16 serving needs substantially more memory. The weights use the Qwen Community 1.0 license, so review its terms before commercial deployment, and verify that your runtime supports this new architecture and multimodal path.
Task Fit
Tool use, repo work, terminal workflows, and coding benchmarks.
Code generation, debugging, refactoring, and benchmark signal.
General writing, Q&A, and assistant use.
Document QA benefits from long context and instruction following.
Image or visual understanding, not necessarily image generation.
Not marked for image generation in the current library.
Not marked for video generation in the current library.
Not marked for voice in the current library.
Source Confidence
Variants and Quant Artifacts
Choose the artifact first; hardware fit follows from RAM, VRAM, format, and runtime.
| Quant | Format | Quality | Min RAM | Reco RAM | Runtime | Action |
|---|---|---|---|---|---|---|
| Q4_K_M | gguf | balanced | 128GB | 192GB | llama.cpp, lm-studio | Plan with this |
| FP8 | safetensors | balanced | 256GB | 384GB | transformers, vllm, sglang, ktransformers, tokenspeed | Plan with this |
| BF16 | safetensors | high | 384GB | 512GB | transformers, vllm, sglang, ktransformers, tokenspeed | Plan with this |
Recommended Hardware
Lowest estimated 5-year cost that can run this model.
Enough effective VRAM with a balanced 5-year cost.
Highest local performance signal among compatible hardware.
Benchmarks
Source and Review
Execution evidence
Run Qwen3.8 Flash Next with a documented recipe
Recipes connect hardware, a model artifact, tools, settings, verification, and a reportable result.
No verified recipe is linked to this record yet.
Compatibility estimates remain available in the planner. A recipe appears here only after its exact stack and verification protocol are documented.
Similar Models
Qwen3.8-27B is Alibaba Qwen's Apache 2.0 open-weight dense vision-language model for coding, professional work, research, and long-horizon agents. It has 27B parameters, a native 262K context window extensible to 1M tokens, flexible reasoning effort, tool use, official BF16 and FP8 weights, and a broad community quantization ecosystem for consumer GPUs and Apple Silicon.
Alibaba's open-weight Qwen3.6 27B. Strong coding-agent, reasoning, long-context, and vision-language model with 262K native context.
Alibaba's flagship Qwen3. Competitive with GPT-4 class models.