Local AI model library
Browse open and open-weight models by task fit, memory requirement, quantization, runtime, and source confidence.
DeepSeek-R1-Distill 70B
human-review-recommendedmitDeepSeek-R1 distilled to Llama 70B base. Strong general reasoning.
DeepSeek-R1-Distill 32B
human-review-recommendedmitDeepSeek-R1 reasoning model distilled to 32B. Excellent at math & logic.
DeepSeek V4.1 Flash
Human reviewedmitDeepSeek V4.1 Flash is DeepSeek's MIT-licensed open-weight multimodal MoE for reasoning, coding, agents, and native image understanding. The complete checkpoint is approximately 763B parameters, including a 552B backbone; its Causal Encoder-Decoder architecture activates 8B parameters during input prefill and 16B during output decoding. CSA2 and FP4 KV caching reduce long-context cache cost. The model supports a 1M-token context window and is available through official FP8 weights, the DeepSeek API, and Ollama Cloud. It remains a data-center-class deployment despite the low active-parameter count.
Qwen3 235B-A22B
human-review-recommendedapache-2.0Alibaba's flagship Qwen3. Competitive with GPT-4 class models.
Phi-4 14B
human-review-recommendedmitMicrosoft's compact 14B. Punches way above its weight, especially on math.
GLM 5.2 FP8
Needs reviewmitZ.ai's GLM 5.2 FP8. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.
GLM 5.1 FP8
Needs reviewmitZ.ai's GLM 5.1 FP8. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.
GLM 4.7 Flash
Needs reviewmitZ.ai's GLM 4.7 Flash. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.
Kimi K2.7 Code
Needs reviewmodified-mitMoonshot AI's Kimi K2.7 Code. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.
Kimi K2.6
Needs reviewmodified-mitMoonshot AI's Kimi K2.6. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.
Kimi K2 Thinking
Needs reviewmodified-mitMoonshot AI's Kimi K2 Thinking. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.
MiniMax M3
Needs reviewminimax-communityMiniMax's MiniMax M3. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.
MiniMax M3 MXFP8
Needs reviewminimax-communityMiniMax's MiniMax M3 MXFP8. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.
MiniMax M2.7
Needs reviewminimax-m2.7-non-commercialMiniMax's MiniMax M2.7. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.
Gemma 4 31B It QAT W4a16 Ct
Needs reviewapache-2.0Google's Gemma 4 31B It QAT W4a16 Ct. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.
Gemma 4 12B It QAT W4a16 Ct
Needs reviewapache-2.0Google's Gemma 4 12B It QAT W4a16 Ct. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.
Gemma 4 E4B It QAT Q4 0 GGUF
Needs reviewapache-2.0Google's Gemma 4 E4B It QAT Q4 0 GGUF. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.
FLUX.2 Dev
Needs reviewflux-non-commercial-licenseBlack Forest Labs's FLUX.2 Dev. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.
FLUX.2 Klein 9B
Needs reviewflux-non-commercial-licenseBlack Forest Labs's FLUX.2 Klein 9B. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.
FLUX.2 Klein 4B
Needs reviewapache-2.0Black Forest Labs's FLUX.2 Klein 4B. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.
Wan2.2 T2V A14B
Needs reviewapache-2.0Alibaba's Wan2.2 T2V A14B. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.
Wan2.2 TI2V 5B
Needs reviewapache-2.0Alibaba's Wan2.2 TI2V 5B. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.
Wan2.2 I2V A14B
Needs reviewapache-2.0Alibaba's Wan2.2 I2V A14B. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.
LTX 2.3
Needs reviewltx-2-community-license-agreementLightricks's LTX 2.3. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.
LTX 2.3 FP8
Needs reviewltx-2-community-license-agreementLightricks's LTX 2.3 FP8. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.
GLM 5.2
Needs reviewmitZ.ai's GLM 5.2. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.
Kimi K3
Human reviewedkimi-k3Moonshot AI Kimi K3 is an open-weight 2.8T-parameter multimodal MoE model with 104B active parameters, native MXFP4 weights, always-on reasoning, tool use, and a 1M-token context window. The official release targets large multi-accelerator deployments through vLLM, SGLang, or TokenSpeed rather than consumer local hardware.
Qwen3.8 2.4T-A95B
Human reviewedqwen3.8-maxQwen3.8-2.4T-A95B is Alibaba Qwen's open-weight, text-only flagship MoE model with 2.4T total parameters and 95B activated parameters. It uses mandatory thinking, supports tool use, has a native 262K context window extensible to about 1.01M tokens, and ships in official BF16 and FP8 repositories. Its multi-terabyte weights target distributed data-center inference with vLLM, SGLang, or TokenSpeed rather than consumer local hardware.
DeepSeek V4 Flash 0731
Human reviewedmitDeepSeek V4 Flash 0731 is the official open-weight release that superseded the preview model and added the DSpark speculative decoding module. Its 304B-parameter, approximately 13B-active MoE checkpoint uses mixed FP4/FP8 weights, supports a 1M-token context window, three reasoning-effort levels, and agentic tool use. The downloadable weights remain available, but DeepSeek retired the hosted V4 Flash API alias in September 2026 in favor of DeepSeek V4.1 Flash.
Muse Glimmer 30B
Human reviewedapache-2.0Meta Muse Glimmer 30B is an Apache 2.0 open-weight, dense multimodal agentic model designed for always-on local workflows. It combines text and image understanding, long-context reasoning, coding, tool calling, and computer-use capabilities. Meta publishes BF16 weights, official GGUF quantizations targeting 24GB and 32GB VRAM, and ExecuTorch packages for Metal and CUDA.
Qwen3.8 27B
Human reviewedapache-2.0Qwen3.8-27B is Alibaba Qwen's Apache 2.0 open-weight dense vision-language model for coding, professional work, research, and long-horizon agents. It has 27B parameters, a native 262K context window extensible to 1M tokens, flexible reasoning effort, tool use, official BF16 and FP8 weights, and a broad community quantization ecosystem for consumer GPUs and Apple Silicon.
Qwen3.8 Flash Next
Human reviewedqwen-community-1.0Qwen3.8-Flash-Next is Alibaba Qwen's open-weight multimodal MoE preview of the architecture planned for Qwen4. Its language model has 125B parameters with 6B activated per token, plus 51B n-gram embeddings and 4B MTP parameters. It combines Gated DeltaNet, Qwen Sparse Attention, gated residuals, vision input, tool use, flexible reasoning, a native 262K context window extensible to 1M tokens, and official serving support through vLLM, SGLang, KTransformers, and TokenSpeed.
GLM 5.3 Flash
Human reviewedmitGLM-5.3-Flash is Z.ai's MIT-licensed, open-weight multimodal MoE for coding, long-horizon agents, computer use, and professional work. It has 320B total parameters with 18B activated per token, hybrid linear and sparse attention, native image input, a 1M-token context window, and official support for SGLang, vLLM, and TokenSpeed. Its total weight footprint makes it a server-class local model despite its efficient active-parameter count.
Qwen3.6 27B
Human reviewedapache-2.0Alibaba's open-weight Qwen3.6 27B. Strong coding-agent, reasoning, long-context, and vision-language model with 262K native context.
Qwen3-Coder 30B-A3B
human-review-recommendedapache-2.0Alibaba's Qwen3-Coder 30B with MoE architecture. Top coding model in <30B class.
Qwen2.5-Coder 32B
human-review-recommendedapache-2.0Qwen 2.5 Coder 32B. Predecessor to Qwen3-Coder, still very capable.
Kimi K2.5
Needs reviewmodified-mitMoonshot AI's Kimi K2.5. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.
MiniMax M2.5
Needs reviewmodified-mitMiniMax's MiniMax M2.5. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.
MiniMax M2
Needs reviewmodified-mitMiniMax's MiniMax M2. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.
DeepSeek-Coder-V2 Lite 16B
human-review-recommendeddeepseekDeepSeek's efficient MoE coder. 16B total / 2.4B active.
Llama 3.3 70B
human-review-recommendedllamaMeta's Llama 3.3 70B. Strong all-rounder, 128K context.
Kimi K2 Instruct 0905
Needs reviewmodified-mitMoonshot AI's Kimi K2 Instruct 0905. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.
Diffusiongemma 26B A4B It
Needs reviewapache-2.0Google's Diffusiongemma 26B A4B It. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.
FLUX.1 Kontext Dev
Needs reviewflux-1-dev-non-commercial-licenseBlack Forest Labs's FLUX.1 Kontext Dev. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.
Stable Diffusion 3.5 Medium
Needs reviewstabilityai-ai-communityStability AI's Stable Diffusion 3.5 Medium. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.
Devstral 24B
human-review-recommendedapache-2.0Mistral's coding-focused model. Strong agent capabilities.
Kimi VL A3B Thinking 2506
Needs reviewmitMoonshot AI's Kimi VL A3B Thinking 2506. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.
Stable Diffusion 3.5 Large
Needs reviewstabilityai-ai-communityStability AI's Stable Diffusion 3.5 Large. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.
Wan2.2 Animate 14B
Needs reviewapache-2.0Alibaba's Wan2.2 Animate 14B. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.
Qwen3 30B-A3B
human-review-recommendedapache-2.0Qwen3 30B with MoE. Same family as Coder, general-purpose.
Wan2.2 S2V 14B
Needs reviewapache-2.0Alibaba's Wan2.2 S2V 14B. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.
LTX 2.3 Nvfp4
Needs reviewltx-2-community-license-agreementLightricks's LTX 2.3 Nvfp4. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.
Mistral Small 3.1 24B
human-review-recommendedapache-2.0Mistral Small 3.1 with vision. Strong multilingual support.
Gemma 4 E2B It QAT Mobile Transformers
Needs reviewapache-2.0Google's Gemma 4 E2B It QAT Mobile Transformers. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.
DeepSeek-VL2
Human revieweddeepseekDeepSeek-VL2 is the full 27.5B-parameter MoE vision-language model, activating about 4.5B parameters per token. It provides the strongest capability in the VL2 family for visual question answering, OCR, document, table and chart understanding, and visual grounding.
DeepSeek-OCR 2
Human reviewedapache-2.0DeepSeek-OCR 2 is DeepSeek's 3B-parameter document understanding model based on Visual Causal Flow. It converts document images into structured text or Markdown and supports dynamic-resolution OCR, layout grounding, and local serving through Transformers, vLLM, and SGLang.
DeepSeek-VL2 Small
Human revieweddeepseekDeepSeek-VL2 Small is the middle 16.1B-parameter MoE member of DeepSeek's vision-language family, activating about 2.8B parameters per token. It is designed for visual question answering, OCR, document, table and chart understanding, and visual grounding.
Janus-Pro 7B
Human reviewedmitJanus-Pro 7B is DeepSeek's larger unified multimodal model for image understanding and text-to-image generation. It improves instruction following and generation stability while retaining separate visual encoding paths for understanding and generation.
Gemma 3 27B
human-review-recommendedgemmaGoogle's open multimodal 27B. Vision + language in one model.
Qwen3 8B
human-review-recommendedapache-2.0Qwen3 8B. Compact general-purpose, runs on any laptop.
DeepSeek-VL2 Tiny
Human revieweddeepseekDeepSeek-VL2 Tiny is the smallest official DeepSeek-VL2 vision-language model. Its 3B-parameter MoE activates about 1B parameters per token and targets visual question answering, OCR, document and chart understanding, and visual grounding with a 4K context window.
Qwen2.5-VL 32B
human-review-recommendedapache-2.0Qwen 2.5 VL (Vision-Language). Top open vision model.
LLaVA-1.6 34B
human-review-recommendedapache-2.0LLaVA 1.6 multimodal model. Image understanding + chat.
Wan2.1 14B
human-review-recommendedapache-2.0Alibaba's Wan 2.1 text-to-video. Open weights, runs on consumer GPUs.
CogVideoX 5B
human-review-recommendedapache-2.0Zhipu AI's CogVideoX 5B. Compact open video model.
BGE-M3 Embedding
human-review-recommendedmitBAAI's BGE-M3. Multilingual embedding for RAG.
Nomic Embed Text v1.5
human-review-recommendedapache-2.0Nomic AI's text embedding. Long context, fully open.
Whisper Large v3
human-review-recommendedmitOpenAI's Whisper Large v3. Top open ASR model.
F5-TTS
human-review-recommendedmitF5-TTS. Open-source TTS with voice cloning.
CosyVoice 300M
human-review-recommendedapache-2.0Alibaba's CosyVoice. Multilingual TTS with emotion control.
SmolLM2 1.7B
human-review-recommendedapache-2.0Hugging Face's SmolLM2 1.7B. Tiny but capable, fits anywhere.
Janus-Pro 1B
Human reviewedmitJanus-Pro 1B is DeepSeek's compact unified multimodal model for image understanding and text-to-image generation. Its decoupled visual encoding paths separate understanding from generation while sharing one transformer architecture.
FLUX.1-dev
human-review-recommendedapache-2.0Black Forest Labs' state-of-the-art text-to-image model. Top of GenEval.
Stable Diffusion XL
human-review-recommendedopenrail++Stability AI's classic SDXL. Best bang/buck for image gen.
HunyuanVideo 13B
human-review-recommendedtencent-hunyuan-communityTencent's open HunyuanVideo. 13B params, high quality.
Llama 3.2 3B
human-review-recommendedllamaMeta's tiny Llama 3.2 3B. Runs on phones & weak hardware.
Gemma 3n 4B
human-review-recommendedgemmaGoogle's Gemma 3n. Multimodal (text/vision/audio) at 4B size.
Stable Diffusion 1.5
Human reviewedcreativeml-openrail-mStable Diffusion 1.5 is a mature text-to-image baseline with broad ComfyUI support and a relatively modest local memory requirement.