Hardware First

Find models that run on your hardware

Start with a Mac, GPU, AI PC, or manual hardware profile. Then filter by task, context length, runtime, and memory headroom.

Runnable
4
35 checked
Good+
4
Excellent or good fit
Largest Fit
8.7GB
Gemma 3n 4B
Offload
0
Runs with CPU/RAM offload

Recommended models

Cards are deduplicated. The full list below keeps every compatible model visible.

General Chat & Reasoning

Model list

Click a row to select it. Click again to clear the selection.

ModelFitQuantMemoryKV extraContextRuntimeAction
Gemma 4 E2B It QAT Mobile Transformers
Google · 2B · Latest gen
Good
Comfortable fast-memory fit
Q4_K_M
8.3GB
0.3GB128Kllama.cpp Plan setup
Kimi VL A3B Thinking 2506
Moonshot AI · 3B
Good
Comfortable fast-memory fit
Q4_K_M
8.5GB
0.5GB256Kllama.cpp Plan setup
Llama 3.2 3B
· 3.2B · Previous gen
Good
Comfortable fast-memory fit
Q4_K_M
8.5GB
0.5GB128Kollama Plan setup
Gemma 3n 4B
· 4B · Previous gen
Good
Comfortable fast-memory fit
Q4_K_M
8.7GB
0.7GB32Kollama Plan setup

Estimate notes

Memory

Uses artifact load RAM/VRAM when available, then adds context and concurrency headroom.

KV cache

Architecture metadata is incomplete, so first version uses a conservative parameter/context heuristic.

Offload

Ollama and llama.cpp can spill model layers into system RAM. This may improve correctness-first choices, but can be much slower.