Model Detail Human reviewed Custom license terms

Qwen3.8 Flash Next

Qwen3.8-Flash-Next is Alibaba Qwen's open-weight multimodal MoE preview of the architecture planned for Qwen4. Its language model has 125B parameters with 6B activated per token, plus 51B n-gram embeddings and 4B MTP parameters. It combines Gated DeltaNet, Qwen Sparse Attention, gated residuals, vision input, tool use, flexible reasoning, a native 262K context window extensible to 1M tokens, and official serving support through vLLM, SGLang, KTransformers, and TokenSpeed.

AlibabaQWEN3qwen-community-1.02026-08-26

Deployment and license note

Qwen3.8 Flash Next is compute-efficient, not small: only 6B parameters activate per token, but the 125B main model, 51B n-gram embeddings, and auxiliary MTP weights still need to be stored. A community Q4 build is roughly a 128GB system-memory class deployment, while official FP8/BF16 serving needs substantially more memory. The weights use the Qwen Community 1.0 license, so review its terms before commercial deployment, and verify that your runtime supports this new architecture and multimodal path.

Parameters
176B
6B active
Context window
1M
Long context
Architecture
gated deltanet qwen sparse attention moe
moe
Quality score
94
Planner signal

Task Fit

Code AgentSupported

Tool use, repo work, terminal workflows, and coding benchmarks.

CodeSupported

Code generation, debugging, refactoring, and benchmark signal.

ChatSupported

General writing, Q&A, and assistant use.

RAGSupported

Document QA benefits from long context and instruction following.

VisionSupported

Image or visual understanding, not necessarily image generation.

Image GenerationNot a fit

Not marked for image generation in the current library.

Video GenerationNot a fit

Not marked for video generation in the current library.

VoiceNot a fit

Not marked for voice in the current library.

Source Confidence

Overallhigh · 100/100
ParametersReviewed / seeded
Task fitReviewed / seeded
MemorySeeded artifact
LicenseSource / seed
BenchmarksAvailable
Hardware fitCalculated

Variants and Quant Artifacts

Choose the artifact first; hardware fit follows from RAM, VRAM, format, and runtime.

3 artifacts
QuantFormatQualityMin RAMReco RAMRuntimeAction
Q4_K_Mggufbalanced128GB192GBllama.cpp, lm-studio Plan with this
FP8safetensorsbalanced256GB384GBtransformers, vllm, sglang, ktransformers, tokenspeed Plan with this
BF16safetensorshigh384GB512GBtransformers, vllm, sglang, ktransformers, tokenspeed Plan with this

Benchmarks

Livecodebench V691.9%
GPQA Diamond91.7%
Android World84.5%
SWE Bench Multilingual81%
Toolathlon Verified73.5%
SWE Bench Pro62.5%
Deep SWE 1.158.7%

Source and Review

Hugging FaceQwen/Qwen3.8-Flash-Next
OllamaNot mapped
VerificationHuman reviewed
Artifact sourcecommunity-qwen3.8-flash-next-gguf
Default variantQwen3.8 Flash Next Coder
Tool callingSupported

Execution evidence

Run Qwen3.8 Flash Next with a documented recipe

Recipes connect hardware, a model artifact, tools, settings, verification, and a reportable result.

Browse all recipes →

No verified recipe is linked to this record yet.

Compatibility estimates remain available in the planner. A recipe appears here only after its exact stack and verification protocol are documented.

Similar Models

Continue planning your local AI setup