Model Serving
Expose models through APIs and optimize inference for shared or production workloads.
GPUStack
GPU InfraGPUStack · Apache-2.0
A GPU cluster manager and model-serving control plane for heterogeneous infrastructure.
Pooling GPUs across multiple servers
llama.cpp
Runtimesggml-org · MIT
A portable C/C++ inference engine that underpins much of the GGUF local-model ecosystem.
Maximum portability and runtime control
LocalAI
ServingLocalAI Project · MIT
A self-hosted OpenAI-compatible API that runs language, image, audio, and multimodal models through modular local backends.
One private API across several AI modalities
MLX LM
RuntimesApple MLX · MIT
Apple Silicon-native tooling for generating, quantizing, fine-tuning, and serving language models with MLX.
Efficient LLM work on Apple Silicon
Ollama
RuntimesOllama · MIT
A straightforward local model runtime with a CLI, model library, and local API.
Running local models with minimal setup
TensorRT-LLM
ServingNVIDIA · Apache-2.0
An NVIDIA inference toolkit for building and serving highly optimized language-model engines on NVIDIA GPUs.
Maximum serving performance on NVIDIA GPUs
vLLM
ServingvLLM Project · Apache-2.0
A high-throughput inference and serving engine for production language-model APIs.
High-throughput model APIs
Xinference
ServingXorbits · Apache-2.0
A model-serving platform for deploying language, embedding, reranking, image, and audio models.
Serving several model types from one control plane
KServe
GPU InfraKServe Project · Apache-2.0
A Kubernetes-native inference platform for standardizing scalable predictive and generative model services.
Standardized model services on Kubernetes
SGLang
ServingSGLang Project · Apache-2.0
A high-performance serving framework for language and multimodal model workloads.
High-performance language and multimodal serving