Local and self-hosted AI tools
Compare model runtimes, serving engines, GPU infrastructure, chat interfaces, knowledge systems, agents, and image workflows by what they actually operate.
All tools
32 matching tools
AnyJev
DecisionsNokia Applied Research · Apache-2.0
An open-source library that turns supported language models into Jev-style typed decision scorers without task-specific training.
Evaluating an existing open model as a typed decision engine
AnythingLLM
Knowledge & RAGMintplex Labs · MIT
A desktop and self-hosted workspace for document chat, RAG, and AI agents.
Private document question answering
AUTOMATIC1111
ImageAUTOMATIC1111 · AGPL-3.0
A widely used local Stable Diffusion web interface with extensive controls, extensions, and model support.
Detailed Stable Diffusion controls
ComfyUI
ImageComfy Org · GPL-3.0
A node-based workflow engine and interface for local generative image and media models.
Reproducible image-generation workflows
DeepSeek Harness
AgentsDeepSeek AI · MIT
An open-source, plugin-first agent harness with local Web UI, tools, sessions, approvals, and extensible model adapters.
Building highly extensible agent workflows
Dify
AgentsLangGenius · Dify Open Source License
A visual platform for building and operating AI workflows, agents, and RAG applications.
Building AI workflows without coding every integration
Flowise
AgentsFlowiseAI · Apache-2.0 community; commercial enterprise components
A visual low-code platform for composing, evaluating, and deploying AI agents and LLM workflows.
Visual agent and RAG prototyping
GPUStack
GPU InfraGPUStack · Apache-2.0
A GPU cluster manager and model-serving control plane for heterogeneous infrastructure.
Pooling GPUs across multiple servers
InvokeAI
ImageInvoke AI · Apache-2.0
A professional local creative application for generative images, canvas editing, workflows, and asset management.
Canvas-based professional image workflows
Jan
RuntimesJan HQ · Apache-2.0
An open-source desktop chat application and local API for running models privately on a personal computer.
Open-source desktop local chat
Jev
DecisionsTypeSafe AI · Proprietary hosted service
A hosted System One decision model that returns typed choices, scores, and yes/no probabilities instead of generated prose.
Adding typed decisions without operating a local model
JevBench
DecisionsJevBench contributors · Open source; verify repository and benchmark data licenses
An open benchmark suite for comparing Jev-class typed decision models across correctness, latency, reliability, and cost.
Comparing decision engines under one frozen protocol
Khoj
Knowledge & RAGKhoj AI · AGPL-3.0
A self-hostable personal AI and knowledge assistant with document search, agents, automations, and research workflows.
A private personal knowledge assistant
Langflow
AgentsLangflow · MIT
An open-source visual builder for creating AI agents and workflows that can run as APIs or MCP servers.
Visual agent workflows with Python extensibility
Letta
AgentsLetta · Apache-2.0
An open-source platform for building stateful agents with persistent, editable, and external memory.
Stateful agents with inspectable memory
LibreChat
Chat UIsLibreChat · MIT
A self-hosted multi-provider chat application with agents, files, search, MCP, and model routing.
Self-hosted multi-provider chat
llama.cpp
Runtimesggml-org · MIT
A portable C/C++ inference engine that underpins much of the GGUF local-model ecosystem.
Maximum portability and runtime control
LM Studio
RuntimesLM Studio · Proprietary
A polished desktop application for discovering, running, and serving local models.
Exploring local models from a desktop UI
LocalAI
ServingLocalAI Project · MIT
A self-hosted OpenAI-compatible API that runs language, image, audio, and multimodal models through modular local backends.
One private API across several AI modalities
MLX LM
RuntimesApple MLX · MIT
Apple Silicon-native tooling for generating, quantizing, fine-tuning, and serving language models with MLX.
Efficient LLM work on Apple Silicon
n8n
Agentsn8n · n8n Sustainable Use License
A workflow automation platform that connects AI agents and model calls to hundreds of business applications.
Connecting AI to business systems
Ollama
RuntimesOllama · MIT
A straightforward local model runtime with a CLI, model library, and local API.
Running local models with minimal setup
Open WebUI
Chat UIsOpen WebUI · Open WebUI License
A self-hosted chat and knowledge interface for local and cloud model providers.
Adding a shared UI to Ollama or an API server
TensorRT-LLM
ServingNVIDIA · Apache-2.0
An NVIDIA inference toolkit for building and serving highly optimized language-model engines on NVIDIA GPUs.
Maximum serving performance on NVIDIA GPUs
vLLM
ServingvLLM Project · Apache-2.0
A high-throughput inference and serving engine for production language-model APIs.
High-throughput model APIs
Xinference
ServingXorbits · Apache-2.0
A model-serving platform for deploying language, embedding, reranking, image, and audio models.
Serving several model types from one control plane
GPT4All
RuntimesNomic AI · MIT
A private desktop chat application and Python SDK for running GGUF language models on everyday computers.
Offline desktop chat on consumer computers
KoboldCpp
RuntimesKoboldCpp Project · AGPL-3.0
A compact GGUF runtime and web interface focused on local text generation, roleplay, and story workflows.
Local creative writing and roleplay
KServe
GPU InfraKServe Project · Apache-2.0
A Kubernetes-native inference platform for standardizing scalable predictive and generative model services.
Standardized model services on Kubernetes
LobeChat
Chat UIsLobeHub · LobeHub Community License
A polished chat and agent workspace with provider routing, knowledge bases, plugins, and optional managed cloud.
Polished personal or team AI workspace
RAGFlow
Knowledge & RAGInfiniFlow · Apache-2.0
A self-hosted RAG engine focused on document understanding, retrieval, and traceable answers.
Document-heavy knowledge systems
SGLang
ServingSGLang Project · Apache-2.0
A high-performance serving framework for language and multimodal model workloads.
High-performance language and multimodal serving