Tool Category

Model Serving

Expose models through APIs and optimize inference for shared or production workloads.

Open source

GPUStack

GPU Infra

GPUStack · Apache-2.0

A GPU cluster manager and model-serving control plane for heterogeneous infrastructure.

Self-hostedClusterLinuxDocker
Best for

Pooling GPUs across multiple servers

Open source

llama.cpp

Runtimes

ggml-org · MIT

A portable C/C++ inference engine that underpins much of the GGUF local-model ecosystem.

Desktop / localSelf-hostedmacOSWindows
Best for

Maximum portability and runtime control

Open source

LocalAI

Serving

LocalAI Project · MIT

A self-hosted OpenAI-compatible API that runs language, image, audio, and multimodal models through modular local backends.

Desktop / localSelf-hostedmacOSLinux
Best for

One private API across several AI modalities

Open source

MLX LM

Runtimes

Apple MLX · MIT

Apple Silicon-native tooling for generating, quantizing, fine-tuning, and serving language models with MLX.

Desktop / localSelf-hostedmacOS
Best for

Efficient LLM work on Apple Silicon

Open source

Ollama

Runtimes

Ollama · MIT

A straightforward local model runtime with a CLI, model library, and local API.

Desktop / localSelf-hostedmacOSWindows
Best for

Running local models with minimal setup

Open source

TensorRT-LLM

Serving

NVIDIA · Apache-2.0

An NVIDIA inference toolkit for building and serving highly optimized language-model engines on NVIDIA GPUs.

Self-hostedClusterLinux
Best for

Maximum serving performance on NVIDIA GPUs

Open source

vLLM

Serving

vLLM Project · Apache-2.0

A high-throughput inference and serving engine for production language-model APIs.

Self-hostedClusterLinux
Best for

High-throughput model APIs

Open source

Xinference

Serving

Xorbits · Apache-2.0

A model-serving platform for deploying language, embedding, reranking, image, and audio models.

Desktop / localSelf-hostedmacOSWindows
Best for

Serving several model types from one control plane

Open source

KServe

GPU Infra

KServe Project · Apache-2.0

A Kubernetes-native inference platform for standardizing scalable predictive and generative model services.

ClusterLinuxKubernetes
Best for

Standardized model services on Kubernetes

Open source

SGLang

Serving

SGLang Project · Apache-2.0

A high-performance serving framework for language and multimodal model workloads.

Self-hostedClusterLinux
Best for

High-performance language and multimodal serving

Explore other tool categories