Model Runtimes
Download and run local models on a workstation, laptop, or single server.
Jan
RuntimesJan HQ · Apache-2.0
An open-source desktop chat application and local API for running models privately on a personal computer.
Open-source desktop local chat
llama.cpp
Runtimesggml-org · MIT
A portable C/C++ inference engine that underpins much of the GGUF local-model ecosystem.
Maximum portability and runtime control
LM Studio
RuntimesLM Studio · Proprietary
A polished desktop application for discovering, running, and serving local models.
Exploring local models from a desktop UI
LocalAI
ServingLocalAI Project · MIT
A self-hosted OpenAI-compatible API that runs language, image, audio, and multimodal models through modular local backends.
One private API across several AI modalities
MLX LM
RuntimesApple MLX · MIT
Apple Silicon-native tooling for generating, quantizing, fine-tuning, and serving language models with MLX.
Efficient LLM work on Apple Silicon
Ollama
RuntimesOllama · MIT
A straightforward local model runtime with a CLI, model library, and local API.
Running local models with minimal setup
GPT4All
RuntimesNomic AI · MIT
A private desktop chat application and Python SDK for running GGUF language models on everyday computers.
Offline desktop chat on consumer computers
KoboldCpp
RuntimesKoboldCpp Project · AGPL-3.0
A compact GGUF runtime and web interface focused on local text generation, roleplay, and story workflows.
Local creative writing and roleplay