Model Library

Local AI model library

Browse open and open-weight models by task fit, memory requirement, quantization, runtime, and source confidence.

Models
78
Families
34
Artifacts
197
Human reviewed
16

DeepSeek-R1-Distill 70B

human-review-recommendedmit

DeepSeek-R1 distilled to Llama 70B base. Strong general reasoning.

DeepSeek70B66K contextQ4_K_M80GB reco
Task fit
ChatCodeRAG

DeepSeek-R1-Distill 32B

human-review-recommendedmit

DeepSeek-R1 reasoning model distilled to 32B. Excellent at math & logic.

DeepSeek32B66K contextQ4_K_M48GB reco
Task fit
ChatCodeRAG

DeepSeek V4.1 Flash

Human reviewedmit

DeepSeek V4.1 Flash is DeepSeek's MIT-licensed open-weight multimodal MoE for reasoning, coding, agents, and native image understanding. The complete checkpoint is approximately 763B parameters, including a 552B backbone; its Causal Encoder-Decoder architecture activates 8B parameters during input prefill and 16B during output decoding. CSA2 and FP4 KV caching reduce long-context cache cost. The model supports a 1M-token context window and is available through official FP8 weights, the DeepSeek API, and Ollama Cloud. It remains a data-center-class deployment despite the low active-parameter count.

DeepSeek763B / 16B active1M contextFP81TB reco
Task fit
VisionChatRAGCodeAgent

Qwen3 235B-A22B

human-review-recommendedapache-2.0

Alibaba's flagship Qwen3. Competitive with GPT-4 class models.

Alibaba235B / 22B active131K contextQ2_K192GB reco
Task fit
ChatCodeRAGToolsAgent

Phi-4 14B

human-review-recommendedmit

Microsoft's compact 14B. Punches way above its weight, especially on math.

Microsoft14B16K contextQ4_K_M24GB reco
Task fit
ChatCodeRAG

GLM 5.2 FP8

Needs reviewmit

Z.ai's GLM 5.2 FP8. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.

Z.ai355B / 32B active131K contextFP896GB reco
Task fit
ChatRAGAgentToolsCode

GLM 5.1 FP8

Needs reviewmit

Z.ai's GLM 5.1 FP8. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.

Z.ai355B / 32B active131K contextFP896GB reco
Task fit
ChatRAGAgentToolsCode

GLM 4.7 Flash

Needs reviewmit

Z.ai's GLM 4.7 Flash. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.

Z.ai32B131K contextQ4_K_M32GB reco
Task fit
ChatRAGAgentToolsCode

Kimi K2.7 Code

Needs reviewmodified-mit

Moonshot AI's Kimi K2.7 Code. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.

Moonshot AI1000B / 32B active262K contextQ4_K_M48GB reco
Task fit
VisionChatRAGCodeAgent

Kimi K2.6

Needs reviewmodified-mit

Moonshot AI's Kimi K2.6. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.

Moonshot AI1000B / 32B active262K contextQ4_K_M48GB reco
Task fit
VisionChatRAGCodeAgent

Kimi K2 Thinking

Needs reviewmodified-mit

Moonshot AI's Kimi K2 Thinking. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.

Moonshot AI1000B / 32B active262K contextQ4_K_M48GB reco
Task fit
ChatRAGCodeAgentTools

MiniMax M3

Needs reviewminimax-community

MiniMax's MiniMax M3. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.

MiniMax456B / 46B active262K contextQ4_K_M64GB reco
Task fit
VisionChatRAGCodeAgent

MiniMax M3 MXFP8

Needs reviewminimax-community

MiniMax's MiniMax M3 MXFP8. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.

MiniMax456B / 46B active262K contextMXFP8128GB reco
Task fit
VisionChatRAGCodeAgent

MiniMax M2.7

Needs reviewminimax-m2.7-non-commercial

MiniMax's MiniMax M2.7. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.

MiniMax230B / 10B active262K contextFP832GB reco
Task fit
ChatRAGCodeAgent

Gemma 4 31B It QAT W4a16 Ct

Needs reviewapache-2.0

Google's Gemma 4 31B It QAT W4a16 Ct. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.

Google31B131K contextQ4_032GB reco
Task fit
VisionChatRAG

Gemma 4 12B It QAT W4a16 Ct

Needs reviewapache-2.0

Google's Gemma 4 12B It QAT W4a16 Ct. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.

Google12B131K contextQ4_016GB reco
Task fit
VisionChatRAG

Gemma 4 E4B It QAT Q4 0 GGUF

Needs reviewapache-2.0

Google's Gemma 4 E4B It QAT Q4 0 GGUF. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.

Google4B131K contextQ4_012GB reco
Task fit
VisionChatRAG

FLUX.2 Dev

Needs reviewflux-non-commercial-license

Black Forest Labs's FLUX.2 Dev. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.

Black Forest Labs32B128K contextFP16128GB reco
Task fit
Image

FLUX.2 Klein 9B

Needs reviewflux-non-commercial-license

Black Forest Labs's FLUX.2 Klein 9B. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.

Black Forest Labs9B128K contextFP1648GB reco
Task fit
Image

FLUX.2 Klein 4B

Needs reviewapache-2.0

Black Forest Labs's FLUX.2 Klein 4B. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.

Black Forest Labs4B128K contextFP1640GB reco
Task fit
Image

Wan2.2 T2V A14B

Needs reviewapache-2.0

Alibaba's Wan2.2 T2V A14B. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.

Alibaba14B128K contextFP1664GB reco
Task fit
Video

Wan2.2 TI2V 5B

Needs reviewapache-2.0

Alibaba's Wan2.2 TI2V 5B. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.

Alibaba5B128K contextFP1640GB reco
Task fit
Video

Wan2.2 I2V A14B

Needs reviewapache-2.0

Alibaba's Wan2.2 I2V A14B. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.

Alibaba14B128K contextFP1664GB reco
Task fit
Video

LTX 2.3

Needs reviewltx-2-community-license-agreement

Lightricks's LTX 2.3. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.

Lightricks13B128K contextFP1664GB reco
Task fit
VideoVision

LTX 2.3 FP8

Needs reviewltx-2-community-license-agreement

Lightricks's LTX 2.3 FP8. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.

Lightricks13B128K contextFP840GB reco
Task fit
VideoVision

GLM 5.2

Needs reviewmit

Z.ai's GLM 5.2. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.

Z.ai355B / 32B active131K contextQ4_K_M48GB reco
Task fit
ChatRAGAgentToolsCode

Kimi K3

Human reviewedkimi-k3

Moonshot AI Kimi K3 is an open-weight 2.8T-parameter multimodal MoE model with 104B active parameters, native MXFP4 weights, always-on reasoning, tool use, and a 1M-token context window. The official release targets large multi-accelerator deployments through vLLM, SGLang, or TokenSpeed rather than consumer local hardware.

Moonshot AI2800B / 104B active1M contextMXFP42TB reco
Task fit
VisionChatRAGCodeAgent

Qwen3.8 2.4T-A95B

Human reviewedqwen3.8-max

Qwen3.8-2.4T-A95B is Alibaba Qwen's open-weight, text-only flagship MoE model with 2.4T total parameters and 95B activated parameters. It uses mandatory thinking, supports tool use, has a native 262K context window extensible to about 1.01M tokens, and ships in official BF16 and FP8 repositories. Its multi-terabyte weights target distributed data-center inference with vLLM, SGLang, or TokenSpeed rather than consumer local hardware.

Alibaba2400B / 95B active1M contextFP84TB reco
Task fit
ChatRAGCodeAgentTools

DeepSeek V4 Flash 0731

Human reviewedmit

DeepSeek V4 Flash 0731 is the official open-weight release that superseded the preview model and added the DSpark speculative decoding module. Its 304B-parameter, approximately 13B-active MoE checkpoint uses mixed FP4/FP8 weights, supports a 1M-token context window, three reasoning-effort levels, and agentic tool use. The downloadable weights remain available, but DeepSeek retired the hosted V4 Flash API alias in September 2026 in favor of DeepSeek V4.1 Flash.

DeepSeek304B / 13B active1M contextMIXED_FP4_FP8256GB reco
Task fit
ChatRAGCodeAgentTools

Muse Glimmer 30B

Human reviewedapache-2.0

Meta Muse Glimmer 30B is an Apache 2.0 open-weight, dense multimodal agentic model designed for always-on local workflows. It combines text and image understanding, long-context reasoning, coding, tool calling, and computer-use capabilities. Meta publishes BF16 weights, official GGUF quantizations targeting 24GB and 32GB VRAM, and ExecuTorch packages for Metal and CUDA.

Meta30B131K contextQ4_K_M32GB reco
Task fit
VisionChatRAGCodeAgent

Qwen3.8 27B

Human reviewedapache-2.0

Qwen3.8-27B is Alibaba Qwen's Apache 2.0 open-weight dense vision-language model for coding, professional work, research, and long-horizon agents. It has 27B parameters, a native 262K context window extensible to 1M tokens, flexible reasoning effort, tool use, official BF16 and FP8 weights, and a broad community quantization ecosystem for consumer GPUs and Apple Silicon.

Alibaba27B1M contextQ4_K_M32GB reco
Task fit
VisionChatRAGCodeAgent

Qwen3.8 Flash Next

Human reviewedqwen-community-1.0

Qwen3.8-Flash-Next is Alibaba Qwen's open-weight multimodal MoE preview of the architecture planned for Qwen4. Its language model has 125B parameters with 6B activated per token, plus 51B n-gram embeddings and 4B MTP parameters. It combines Gated DeltaNet, Qwen Sparse Attention, gated residuals, vision input, tool use, flexible reasoning, a native 262K context window extensible to 1M tokens, and official serving support through vLLM, SGLang, KTransformers, and TokenSpeed.

Alibaba176B / 6B active1M contextQ4_K_M192GB reco
Task fit
VisionChatRAGCodeAgent

GLM 5.3 Flash

Human reviewedmit

GLM-5.3-Flash is Z.ai's MIT-licensed, open-weight multimodal MoE for coding, long-horizon agents, computer use, and professional work. It has 320B total parameters with 18B activated per token, hybrid linear and sparse attention, native image input, a 1M-token context window, and official support for SGLang, vLLM, and TokenSpeed. Its total weight footprint makes it a server-class local model despite its efficient active-parameter count.

Z.ai320B / 18B active1M contextFP8512GB reco
Task fit
VisionChatRAGCodeAgent

Qwen3.6 27B

Human reviewedapache-2.0

Alibaba's open-weight Qwen3.6 27B. Strong coding-agent, reasoning, long-context, and vision-language model with 262K native context.

Alibaba27B262K contextQ4_K_M32GB reco
Task fit
ChatCodeAgentRAGVision

Qwen3-Coder 30B-A3B

human-review-recommendedapache-2.0

Alibaba's Qwen3-Coder 30B with MoE architecture. Top coding model in <30B class.

Alibaba30.5B / 3.3B active262K contextQ4_K_M48GB reco
Task fit
CodeAgentTools

Qwen2.5-Coder 32B

human-review-recommendedapache-2.0

Qwen 2.5 Coder 32B. Predecessor to Qwen3-Coder, still very capable.

Alibaba32B33K contextQ4_K_M48GB reco
Task fit
CodeAgentTools

Kimi K2.5

Needs reviewmodified-mit

Moonshot AI's Kimi K2.5. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.

Moonshot AI1000B / 32B active262K contextQ4_K_M48GB reco
Task fit
VisionChatRAGCodeAgent

MiniMax M2.5

Needs reviewmodified-mit

MiniMax's MiniMax M2.5. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.

MiniMax230B / 10B active262K contextFP832GB reco
Task fit
ChatRAGCodeAgent

MiniMax M2

Needs reviewmodified-mit

MiniMax's MiniMax M2. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.

MiniMax230B / 10B active262K contextFP832GB reco
Task fit
ChatRAGCodeAgent

DeepSeek-Coder-V2 Lite 16B

human-review-recommendeddeepseek

DeepSeek's efficient MoE coder. 16B total / 2.4B active.

DeepSeek16B / 2.4B active128K contextQ4_K_M24GB reco
Task fit
CodeAgentTools

Llama 3.3 70B

human-review-recommendedllama

Meta's Llama 3.3 70B. Strong all-rounder, 128K context.

70B131K contextQ4_K_M80GB reco
Task fit
ChatCodeRAGTools

Kimi K2 Instruct 0905

Needs reviewmodified-mit

Moonshot AI's Kimi K2 Instruct 0905. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.

Moonshot AI1000B / 32B active262K contextFP896GB reco
Task fit
ChatRAGCodeAgentTools

Diffusiongemma 26B A4B It

Needs reviewapache-2.0

Google's Diffusiongemma 26B A4B It. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.

Google26B / 4B active128K contextFP16128GB reco
Task fit
ImageVision

FLUX.1 Kontext Dev

Needs reviewflux-1-dev-non-commercial-license

Black Forest Labs's FLUX.1 Kontext Dev. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.

Black Forest Labs8B128K contextFP1648GB reco
Task fit
Image

Stable Diffusion 3.5 Medium

Needs reviewstabilityai-ai-community

Stability AI's Stable Diffusion 3.5 Medium. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.

Stability AI2.5B128K contextFP1624GB reco
Task fit
Image

Devstral 24B

human-review-recommendedapache-2.0

Mistral's coding-focused model. Strong agent capabilities.

Mistral AI24B128K contextQ4_K_M32GB reco
Task fit
CodeAgent

Kimi VL A3B Thinking 2506

Needs reviewmit

Moonshot AI's Kimi VL A3B Thinking 2506. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.

Moonshot AI3B262K contextQ4_K_M8GB reco
Task fit
VisionChatRAGCodeAgent

Stable Diffusion 3.5 Large

Needs reviewstabilityai-ai-community

Stability AI's Stable Diffusion 3.5 Large. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.

Stability AI8B128K contextFP1648GB reco
Task fit
Image

Wan2.2 Animate 14B

Needs reviewapache-2.0

Alibaba's Wan2.2 Animate 14B. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.

Alibaba14B128K contextFP1664GB reco
Task fit
Video

Qwen3 30B-A3B

human-review-recommendedapache-2.0

Qwen3 30B with MoE. Same family as Coder, general-purpose.

Alibaba30.5B / 3.3B active131K contextQ4_K_M48GB reco
Task fit
ChatCodeRAGTools

Wan2.2 S2V 14B

Needs reviewapache-2.0

Alibaba's Wan2.2 S2V 14B. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.

Alibaba14B128K contextFP1664GB reco
Task fit
Video

LTX 2.3 Nvfp4

Needs reviewltx-2-community-license-agreement

Lightricks's LTX 2.3 Nvfp4. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.

Lightricks13B128K contextNVFP424GB reco
Task fit
VideoVision

Mistral Small 3.1 24B

human-review-recommendedapache-2.0

Mistral Small 3.1 with vision. Strong multilingual support.

Mistral AI24B128K contextQ4_K_M40GB reco
Task fit
ChatVisionToolsCode

Gemma 4 E2B It QAT Mobile Transformers

Needs reviewapache-2.0

Google's Gemma 4 E2B It QAT Mobile Transformers. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.

Google2B131K contextQ4_K_M8GB reco
Task fit
VisionChatRAG

DeepSeek-VL2

Human revieweddeepseek

DeepSeek-VL2 is the full 27.5B-parameter MoE vision-language model, activating about 4.5B parameters per token. It provides the strongest capability in the VL2 family for visual question answering, OCR, document, table and chart understanding, and visual grounding.

DeepSeek27.5B / 4.5B active4K contextBF1680GB reco
Task fit
VisionChatRAG

DeepSeek-OCR 2

Human reviewedapache-2.0

DeepSeek-OCR 2 is DeepSeek's 3B-parameter document understanding model based on Visual Causal Flow. It converts document images into structured text or Markdown and supports dynamic-resolution OCR, layout grounding, and local serving through Transformers, vLLM, and SGLang.

DeepSeek3B8K contextBF1616GB reco
Task fit
VisionRAG

DeepSeek-VL2 Small

Human revieweddeepseek

DeepSeek-VL2 Small is the middle 16.1B-parameter MoE member of DeepSeek's vision-language family, activating about 2.8B parameters per token. It is designed for visual question answering, OCR, document, table and chart understanding, and visual grounding.

DeepSeek16.1B / 2.8B active4K contextBF1648GB reco
Task fit
VisionChatRAG

Janus-Pro 7B

Human reviewedmit

Janus-Pro 7B is DeepSeek's larger unified multimodal model for image understanding and text-to-image generation. It improves instruction following and generation stability while retaining separate visual encoding paths for understanding and generation.

DeepSeek7B4K contextBF1632GB reco
Task fit
VisionImageChat

Gemma 3 27B

human-review-recommendedgemma

Google's open multimodal 27B. Vision + language in one model.

Google27B128K contextQ4_K_M40GB reco
Task fit
ChatVisionCode

Qwen3 8B

human-review-recommendedapache-2.0

Qwen3 8B. Compact general-purpose, runs on any laptop.

Alibaba8.2B131K contextQ4_K_M16GB reco
Task fit
ChatCodeRAGTools

DeepSeek-VL2 Tiny

Human revieweddeepseek

DeepSeek-VL2 Tiny is the smallest official DeepSeek-VL2 vision-language model. Its 3B-parameter MoE activates about 1B parameters per token and targets visual question answering, OCR, document and chart understanding, and visual grounding with a 4K context window.

DeepSeek3B / 1B active4K contextBF1616GB reco
Task fit
VisionChatRAG

Qwen2.5-VL 32B

human-review-recommendedapache-2.0

Qwen 2.5 VL (Vision-Language). Top open vision model.

Alibaba32B33K contextQ4_K_M48GB reco
Task fit
VisionChatCode

LLaVA-1.6 34B

human-review-recommendedapache-2.0

LLaVA 1.6 multimodal model. Image understanding + chat.

34B4K contextQ4_K_M48GB reco
Task fit
VisionChat

Wan2.1 14B

human-review-recommendedapache-2.0

Alibaba's Wan 2.1 text-to-video. Open weights, runs on consumer GPUs.

14BContext unknownFP1648GB reco
Task fit
Video

CogVideoX 5B

human-review-recommendedapache-2.0

Zhipu AI's CogVideoX 5B. Compact open video model.

5BContext unknownFP1624GB reco
Task fit
Video

BGE-M3 Embedding

human-review-recommendedmit

BAAI's BGE-M3. Multilingual embedding for RAG.

0.6B8K contextFP168GB reco
Task fit
RAG

Nomic Embed Text v1.5

human-review-recommendedapache-2.0

Nomic AI's text embedding. Long context, fully open.

0.3B8K contextFP164GB reco
Task fit
RAG

Whisper Large v3

human-review-recommendedmit

OpenAI's Whisper Large v3. Top open ASR model.

OpenAI1.5BContext unknownFP168GB reco
Task fit
Voice

F5-TTS

human-review-recommendedmit

F5-TTS. Open-source TTS with voice cloning.

0.3BContext unknownFP168GB reco
Task fit
Voice

CosyVoice 300M

human-review-recommendedapache-2.0

Alibaba's CosyVoice. Multilingual TTS with emotion control.

0.3BContext unknownFP168GB reco
Task fit
Voice

SmolLM2 1.7B

human-review-recommendedapache-2.0

Hugging Face's SmolLM2 1.7B. Tiny but capable, fits anywhere.

1.7B8K contextQ4_K_M4GB reco
Task fit
ChatRAG

Janus-Pro 1B

Human reviewedmit

Janus-Pro 1B is DeepSeek's compact unified multimodal model for image understanding and text-to-image generation. Its decoupled visual encoding paths separate understanding from generation while sharing one transformer architecture.

DeepSeek1B4K contextBF1612GB reco
Task fit
VisionImageChat

FLUX.1-dev

human-review-recommendedapache-2.0

Black Forest Labs' state-of-the-art text-to-image model. Top of GenEval.

Black Forest Labs12BContext unknownFP1648GB reco
Task fit
Image

Stable Diffusion XL

human-review-recommendedopenrail++

Stability AI's classic SDXL. Best bang/buck for image gen.

3.5BContext unknownFP1616GB reco
Task fit
Image

HunyuanVideo 13B

human-review-recommendedtencent-hunyuan-community

Tencent's open HunyuanVideo. 13B params, high quality.

13BContext unknownFP1648GB reco
Task fit
Video

Llama 3.2 3B

human-review-recommendedllama

Meta's tiny Llama 3.2 3B. Runs on phones & weak hardware.

3.2B131K contextQ4_K_M8GB reco
Task fit
ChatRAG

Gemma 3n 4B

human-review-recommendedgemma

Google's Gemma 3n. Multimodal (text/vision/audio) at 4B size.

4B33K contextQ4_K_M8GB reco
Task fit
ChatVisionVoice

Stable Diffusion 1.5

Human reviewedcreativeml-openrail-m

Stable Diffusion 1.5 is a mature text-to-image baseline with broad ComfyUI support and a relatively modest local memory requirement.

Stability AI0.86BContext unknownFP168GB reco
Task fit
Image