Find models that run on your hardware
Start with a Mac, GPU, AI PC, or manual hardware profile. Then filter by task, context length, runtime, and memory headroom.
Recommended models
Cards are deduplicated. The full list below keeps every compatible model visible.
Best combined task, quality, memory, runtime, and freshness score.
Largest estimated model load that fits the selected hardware profile.
Highest task-aware quality signal in the compatible set.
Model list
Click a row to select it. Click again to clear the selection.
| Model | Fit | Quant | Memory | KV extra | Context | Runtime | Action |
|---|---|---|---|---|---|---|---|
Gemma 4 E2B It QAT Mobile Transformers Google · 2B · Latest gen | Good Comfortable fast-memory fit | Q4_K_M | 8.3GB | 0.3GB | 128K | llama.cpp | Plan setup |
Kimi VL A3B Thinking 2506 Moonshot AI · 3B | Good Comfortable fast-memory fit | Q4_K_M | 8.5GB | 0.5GB | 256K | llama.cpp | Plan setup |
Llama 3.2 3B · 3.2B · Previous gen | Good Comfortable fast-memory fit | Q4_K_M | 8.5GB | 0.5GB | 128K | ollama | Plan setup |
Gemma 3n 4B · 4B · Previous gen | Good Comfortable fast-memory fit | Q4_K_M | 8.7GB | 0.7GB | 32K | ollama | Plan setup |
Estimate notes
Uses artifact load RAM/VRAM when available, then adds context and concurrency headroom.
Architecture metadata is incomplete, so first version uses a conservative parameter/context heuristic.
Ollama and llama.cpp can spill model layers into system RAM. This may improve correctness-first choices, but can be much slower.