Local and self-hosted AI tools
Compare model runtimes, serving engines, GPU infrastructure, chat interfaces, knowledge systems, agents, and image workflows by what they actually operate.
All tools
29 matching tools
AnythingLLM
Knowledge & RAGMintplex Labs · MIT
A desktop and self-hosted workspace for document chat, RAG, and AI agents.
Private document question answering
AUTOMATIC1111
ImageAUTOMATIC1111 · AGPL-3.0
A widely used local Stable Diffusion web interface with extensive controls, extensions, and model support.
Detailed Stable Diffusion controls
ComfyUI
ImageComfy Org · GPL-3.0
A node-based workflow engine and interface for local generative image and media models.
Reproducible image-generation workflows
DeepSeek Harness
AgentsDeepSeek AI · MIT
An open-source, plugin-first agent harness with local Web UI, tools, sessions, approvals, and extensible model adapters.
Building highly extensible agent workflows
Dify
AgentsLangGenius · Dify Open Source License
A visual platform for building and operating AI workflows, agents, and RAG applications.
Building AI workflows without coding every integration
Flowise
AgentsFlowiseAI · Apache-2.0 community; commercial enterprise components
A visual low-code platform for composing, evaluating, and deploying AI agents and LLM workflows.
Visual agent and RAG prototyping
GPUStack
GPU InfraGPUStack · Apache-2.0
A GPU cluster manager and model-serving control plane for heterogeneous infrastructure.
Pooling GPUs across multiple servers
InvokeAI
ImageInvoke AI · Apache-2.0
A professional local creative application for generative images, canvas editing, workflows, and asset management.
Canvas-based professional image workflows
Jan
RuntimesJan HQ · Apache-2.0
An open-source desktop chat application and local API for running models privately on a personal computer.
Open-source desktop local chat
Khoj
Knowledge & RAGKhoj AI · AGPL-3.0
A self-hostable personal AI and knowledge assistant with document search, agents, automations, and research workflows.
A private personal knowledge assistant
Langflow
AgentsLangflow · MIT
An open-source visual builder for creating AI agents and workflows that can run as APIs or MCP servers.
Visual agent workflows with Python extensibility
Letta
AgentsLetta · Apache-2.0
An open-source platform for building stateful agents with persistent, editable, and external memory.
Stateful agents with inspectable memory
LibreChat
Chat UIsLibreChat · MIT
A self-hosted multi-provider chat application with agents, files, search, MCP, and model routing.
Self-hosted multi-provider chat
llama.cpp
Runtimesggml-org · MIT
A portable C/C++ inference engine that underpins much of the GGUF local-model ecosystem.
Maximum portability and runtime control
LM Studio
RuntimesLM Studio · Proprietary
A polished desktop application for discovering, running, and serving local models.
Exploring local models from a desktop UI
LocalAI
ServingLocalAI Project · MIT
A self-hosted OpenAI-compatible API that runs language, image, audio, and multimodal models through modular local backends.
One private API across several AI modalities
MLX LM
RuntimesApple MLX · MIT
Apple Silicon-native tooling for generating, quantizing, fine-tuning, and serving language models with MLX.
Efficient LLM work on Apple Silicon
n8n
Agentsn8n · n8n Sustainable Use License
A workflow automation platform that connects AI agents and model calls to hundreds of business applications.
Connecting AI to business systems
Ollama
RuntimesOllama · MIT
A straightforward local model runtime with a CLI, model library, and local API.
Running local models with minimal setup
Open WebUI
Chat UIsOpen WebUI · Open WebUI License
A self-hosted chat and knowledge interface for local and cloud model providers.
Adding a shared UI to Ollama or an API server
TensorRT-LLM
ServingNVIDIA · Apache-2.0
An NVIDIA inference toolkit for building and serving highly optimized language-model engines on NVIDIA GPUs.
Maximum serving performance on NVIDIA GPUs
vLLM
ServingvLLM Project · Apache-2.0
A high-throughput inference and serving engine for production language-model APIs.
High-throughput model APIs
Xinference
ServingXorbits · Apache-2.0
A model-serving platform for deploying language, embedding, reranking, image, and audio models.
Serving several model types from one control plane
GPT4All
RuntimesNomic AI · MIT
A private desktop chat application and Python SDK for running GGUF language models on everyday computers.
Offline desktop chat on consumer computers
KoboldCpp
RuntimesKoboldCpp Project · AGPL-3.0
A compact GGUF runtime and web interface focused on local text generation, roleplay, and story workflows.
Local creative writing and roleplay
KServe
GPU InfraKServe Project · Apache-2.0
A Kubernetes-native inference platform for standardizing scalable predictive and generative model services.
Standardized model services on Kubernetes
LobeChat
Chat UIsLobeHub · LobeHub Community License
A polished chat and agent workspace with provider routing, knowledge bases, plugins, and optional managed cloud.
Polished personal or team AI workspace
RAGFlow
Knowledge & RAGInfiniFlow · Apache-2.0
A self-hosted RAG engine focused on document understanding, retrieval, and traceable answers.
Document-heavy knowledge systems
SGLang
ServingSGLang Project · Apache-2.0
A high-performance serving framework for language and multimodal model workloads.
High-performance language and multimodal serving