Reference — Serving the models

Model Runtimes & Gateways

What actually runs an open-weight model from the registry, and what sits in front of it. These tools are routinely compared as if they were alternatives; mostly they are not. They occupy distinct layers, and the common expensive mistake — putting a developer wrapper behind a production API — is a layer error, not a benchmark error.

The registry computes a weight footprint for every self-hostable model — parameters multiplied by bytes per weight, at Q4 and FP8, with the hardware class that implies. It is derived arithmetic rather than a vendor claim, and deliberately a floor rather than a requirement.

What it takes to run a model

The registry's weight footprint is a floor, not a requirement. Four things decide the rest, and all four belong to this layer rather than to the model.

Total parameters set the memory; active parameters only set the compute
A mixture-of-experts model must hold every expert in memory even though a fraction fires per token. DeepSeek V4 Pro is 49B active and still needs roughly 880 GB at Q4, because all 1.6T weights have to be resident. Reading the active count as the requirement is the most common and most expensive mistake in this area, which is why the registry prints both figures side by side.
KV cache is not in that number, and at long context it rivals the weights
Cache grows with context length and again with every concurrent request, so a model that fits a card at 8K context may not fit it at 128K, and almost certainly will not serve ten users there. Computing it needs layer counts, KV-head counts and head dimensions that most closed vendors never publish, so the registry does not pretend to. Budget headroom on top of the printed figure, and treat a 1M-context claim as a statement about the architecture, not about your GPU.
The engine decides what happens when it does not fit
llama.cpp will offload layers to system RAM and keep going slowly; vLLM effectively wants residency and will fail rather than degrade; ExLlamaV3 exists to squeeze a model onto a card it should not fit; KTransformers offloads MoE experts to CPU on purpose. The same model on the same hardware is runnable or not depending on which of these you pick — which is the argument for choosing the engine before the model.
Quantisation is a quality decision, not just a size one
Q4 roughly halves FP8, and the loss is small on many tasks and not small on some — long-context reasoning and code generation degrade earlier than chat. The registry shows both so the trade is visible; measure on your own evaluation before assuming the cheap tier is free.

llama.cpp

🌐 —
ggml-org (open source) · MIT
EngineOpen sourceSelf-hostable
Best at
Portability, edge and CPU or hybrid offload
Concurrency
Parallel slots (--parallel), single-node
Hardware
Effectively everything: x86 AVX2/AVX-512, ARM NEON, CUDA, Metal, Vulkan, ROCm
Formats
GGUF
Registry models that name it
Ceiling
Manual operations, single-node thinking, no fleet-serving story. It is a library and a server binary, not a platform.
The substrate of the local-model world. It created GGUF, and most consumer-hardware tools either wrap it or exist in reaction to it. Go direct when you need embedded deployment, unusual hardware, aggressive CPU/GPU layer splitting, or one less dependency between you and the weights.

vLLM

🌐 —
vLLM project · Apache-2.0
EngineOpen sourceSelf-hostable
Best at
Serving many concurrent users from shared GPUs
Concurrency
PagedAttention with continuous batching
Hardware
NVIDIA, AMD, some TPU and Gaudi
Formats
SafetensorsAWQGPTQFP8
Ceiling
Linux-centric with a heavy footprint, and GGUF is second-class. Not what you want on a laptop.
The default answer for production self-hosting, and the one the industry converged on — Hugging Face's own TGI now points here. Requests are packed into shared compute passes rather than run end to end, which is why throughput climbs with concurrency instead of flattening. Speaks OpenAI, so a gateway fronts it identically to a hosted vendor.

SGLang

🌐 —
SGLang project · Apache-2.0
EngineOpen sourceSelf-hostable
Best at
Agentic and prefix-heavy workloads
Concurrency
RadixAttention — a radix tree of cached prefixes
Hardware
Mostly NVIDIA
Formats
SafetensorsAWQGPTQFP8
Ceiling
Smaller ecosystem and narrower hardware coverage than vLLM.
vLLM's sibling, optimised for reuse rather than raw batching. Requests sharing a system prompt or conversation history skip recomputation, which is the exact shape of agent loops, tool-calling chains and multi-turn chat behind a fat system prompt. Choose it over vLLM when prefix reuse is high and time-to-first-token matters.

TensorRT-LLM

🇺🇸 USA
NVIDIA · Apache-2.0
EngineOpen sourceSelf-hostable
Best at
Peak throughput on an NVIDIA estate
Concurrency
In-flight batching over compiled engines
Hardware
NVIDIA only
Formats
Compiled TensorRT enginesSafetensors (converted)
Registry models that name it
Ceiling
Highest performance ceiling, highest operational cost: models are compiled per configuration, and it leaves the API surface to Triton, NIM or Dynamo.
The choice when the hardware is already bought and the last 30% of utilisation is worth an engineering budget. Rarely the right first step, and rarely the wrong last one.

MLX / mlx-lm

🇺🇸 USA
Apple · MIT
EngineOpen sourceSelf-hostable
Best at
Apple Silicon, particularly prompt processing
Concurrency
Single-node; batching depends on the serving layer above it
Hardware
Apple Silicon only
Formats
MLXSafetensors (converted)
Registry models that name it
Ceiling
One vendor's hardware, and no answer at all for a rack.
Relevant precisely when the fleet is MacBooks — an underrated case for organisations that issue Macs and want inference to stay on the endpoint. Reached most often through a wrapper rather than directly.

LMDeploy

🇨🇳 China
InternLM / OpenMMLab · Apache-2.0
EngineOpen sourceSelf-hostable
Best at
Throughput via its TurboMind runtime
Concurrency
Persistent batching
Hardware
NVIDIA, some Ascend
Formats
SafetensorsAWQ
Ceiling
Smaller community and thinner documentation in English than vLLM or SGLang.
A credible third option in the serving-engine tier, strongest on the models its authors ship. Worth benchmarking against vLLM for a specific model rather than adopting on reputation.

ExLlamaV3

🌐 —
turboderp (open source) · MIT
EngineOpen sourceSelf-hostable
Best at
Fitting large models onto consumer VRAM
Concurrency
Limited — built for single-user quality per gigabyte
Hardware
NVIDIA consumer GPUs
Formats
EXL3Safetensors (quantised)
Ceiling
A quantisation and single-user story, not a serving platform.
The tool for running a model a card should not fit, at a quality per gigabyte the general engines do not match. Niche by design.

Text Generation Inference (TGI)

🇺🇸 USA
Hugging Face · Apache-2.0
EngineOpen sourcemaintenanceSelf-hostable
Best at
Nothing new — migrate off it
Concurrency
Continuous batching
Hardware
NVIDIA, AMD, Gaudi
Formats
SafetensorsAWQGPTQ
Ceiling
In maintenance mode: minor bug fixes, documentation and lightweight maintenance only. The repository was archived on 21 March 2026.
For a period the default self-hosted server, and now the clearest example of why this page records status. Hugging Face's own README points new users to vLLM, SGLang, llama.cpp and MLX. If you are running it, the migration is not optional so much as already overdue.

Ollama

🇺🇸 USA
Ollama · MIT
WrapperOpen sourceSelf-hostable
Best at
The single-developer loop: pull by name and go
Concurrency
Limited, via OLLAMA_NUM_PARALLEL
Hardware
Anything the underlying engine supports
Formats
GGUF
Ceiling
Not a production API server. Throughput flattens almost immediately under concurrent load while a real serving engine climbs — the single most common expensive mistake in this space is putting it behind a shared endpoint.
An OpenAI-compatible endpoint on :11434 and a model library, which is why it appears in so many agent integrations: its own CLI documents launch profiles for Claude Code, Codex, OpenCode, VS Code and Droid. That is also why it shows up on the agents page as the runtime behind several local-model claims. Right for dev loops, CI fixtures and single-user desktop work.

LM Studio

🇺🇸 USA
Element Labs · Proprietary desktop application
WrapperProprietarySelf-hostable
Best at
GUI-first model discovery and evaluation
Concurrency
Parallel slots with continuous batching since 0.4.0
Hardware
macOS, Windows, Linux desktops
Formats
GGUFMLX
Registry models that name it
Ceiling
Proprietary, and its API gives less control over parallel and multi-model routing than Ollama's environment variables.
Version 0.4.0 (28 January 2026) split the engine from the GUI: llmster runs headless on a server or in CI, with continuous batching, an n_parallel setting and a stateful REST API. The real advantage remains evaluation ergonomics — browsing Hugging Face in-app, quantisation recommendations against your actual RAM and VRAM, and visible tokens per second while you compare.

Jan

🌐 —
Menlo Research · AGPL-3.0
WrapperOpen sourceSelf-hostable
Best at
An open-source desktop assistant that stays offline
Concurrency
Single-user desktop
Hardware
macOS, Windows, Linux desktops
Formats
GGUF
Ceiling
Desktop scope, and a smaller ecosystem than Ollama or LM Studio.
The open alternative in the GUI tier, for people who want LM Studio's shape without a proprietary binary. Same engine underneath as most of this layer.

KoboldCpp

🌐 —
LostRuins (open source) · AGPL-3.0
WrapperOpen sourceSelf-hostable
Best at
Single-binary local deployment with heavy sampling control
Concurrency
Limited
Hardware
Anything llama.cpp supports
Formats
GGUF
Ceiling
A hobbyist and single-user tool by design; no fleet story.
A one-file llama.cpp distribution with an unusually deep set of sampling and context controls. Included because it is what a lot of local deployment actually is, rather than what conference talks assume it is.

OpenRouter

🇺🇸 USA
OpenRouter · Hosted service
GatewayProprietaryHosted only
Best at
Breadth of catalogue with zero setup
Ceiling
You are renting someone else's routing decisions. Fine for evaluation, awkward as a production path for anyone with a residency or assessment obligation.
Governance
A single request may be fulfilled by any of several sub-providers, whose retention policies differ. The model name is not the control surface — the provider allowlist and data-policy settings are. Treat it as a catalogue and price-discovery instrument; if it goes near production, put it behind your own gateway with a hard allowlist rather than beside it.
One key, one OpenAI-shaped API, hundreds of hosted models, unified billing — the fastest way to compare models or price-shop. Teams typically leave not for a bigger catalogue but because the problem turns into governance, observability and residency.

LiteLLM

🇺🇸 USA
BerriAI · MIT, with an enterprise tier
GatewayOpen sourceSelf-hostable
Best at
A self-hosted control plane over every provider
Ceiling
You own the uptime. A gateway is another production dependency, and it fails in front of everything.
Governance
Keys, budgets, rate limits, model allowlists and request logs stay on your infrastructure, which is what makes it the usual answer where OpenRouter is not. Also the point at which self-hosted engines and hosted vendors become interchangeable behind one contract.
The pragmatic centre of most serious stacks: one OpenAI-shaped internal contract, with vLLM, Azure, Bedrock or anything else swappable behind it. Its price file is also one of the sources this site refreshes registry pricing against.

Portkey

🇺🇸 USA
Portkey · Open-source gateway with a hosted control plane
GatewayOpen sourceSelf-hostable
Best at
Guardrails and governance in front of the model call
Ceiling
The richest features sit in the hosted control plane, so the self-hosted story is thinner than the marketing implies.
Governance
Positions policy — guardrails, PII handling, request filtering — as gateway concerns rather than application ones, which is the right layer for them if you are enforcing across many teams.
Chosen where the requirement is governance rather than routing, and where a per-request policy engine belongs in the path.

Helicone

🇺🇸 USA
Helicone · Apache-2.0, hosted option
GatewayOpen sourceSelf-hostable
Best at
Observability over LLM traffic
Ceiling
Observability-first, so it is a weaker policy engine than a gateway built for control.
Governance
Logging, tracing and cost attribution per request. The reason to care: without this layer nobody can answer what a prompt cost, what it contained, or which provider served it.
Often deployed alongside rather than instead of a routing gateway.

Cloudflare AI Gateway

🇺🇸 USA
Cloudflare · Hosted service
GatewayProprietaryHosted only
Best at
Caching, rate limiting and analytics at the edge
Ceiling
Hosted only — the gateway itself is another vendor in the path, which is a strange shape if the reason for a gateway was vendor control.
Governance
Caching and rate limiting sit at the edge, and requests transit Cloudflare. Straightforward for cost control; a question to answer explicitly where the point of the exercise was data residency.
Cheap to adopt if you are already a Cloudflare tenant, and the analytics arrive without instrumentation.

NVIDIA Triton Inference Server

🇺🇸 USA
NVIDIA · BSD-3-Clause
OrchestratorOpen sourceSelf-hostable
Best at
Serving many models and frameworks behind one endpoint
Concurrency
Dynamic batching and multi-model scheduling
Hardware
NVIDIA primarily
Formats
TensorRTONNXPyTorchOthers via backends
Ceiling
Substantial operational surface, and it is not itself an LLM engine — it schedules whatever backend is doing the inference.
The layer above the engine: model packing, versioning and multi-framework serving. Usually paired with TensorRT-LLM rather than replacing it.

NVIDIA Dynamo

🇺🇸 USA
NVIDIA · Apache-2.0
OrchestratorOpen sourceSelf-hostable
Best at
Distributed serving with prefill and decode disaggregated
Concurrency
Cross-node scheduling and KV-cache-aware routing
Hardware
NVIDIA clusters
Formats
Engine-dependent — vLLM, SGLang, TensorRT-LLM
Ceiling
Only earns its complexity at multi-node scale.
Splits the compute-bound prefill phase from the memory-bound decode phase onto different resources, which is the current frontier of serving efficiency and irrelevant below a certain size.

Ray Serve

🇺🇸 USA
Anyscale / Ray · Apache-2.0
OrchestratorOpen sourceSelf-hostable
Best at
Autoscaling and composing models with application code
Concurrency
Replica autoscaling and request batching
Hardware
Any cluster Ray runs on
Formats
Engine-dependent
Ceiling
You inherit Ray, which is a distributed-systems commitment rather than a serving library.
The pick when inference is one step in a larger pipeline rather than the whole product, and when scaling policy needs to live in code.

KServe

🌐 —
CNCF · Apache-2.0
OrchestratorOpen sourceSelf-hostable
Best at
Kubernetes-native model serving with standard CRDs
Concurrency
Autoscaling, including scale-to-zero
Hardware
Any Kubernetes cluster
Formats
Engine-dependent — vLLM and others via runtimes
Ceiling
Only makes sense if Kubernetes is already the platform, and adds a control plane if it is not.
The vendor-neutral option in the orchestration tier, and the one that fits an existing platform-engineering practice rather than replacing it.

Together AI

🇺🇸 USA
Together · Hosted service
Hosted providerProprietaryHosted only
Best at
Open-weight models as an API, without owning GPUs
Registry models that name it
Ceiling
Someone else owns the engine, the hardware and the retention policy. Open weights served by a third party is not the same privacy posture as open weights on your own metal.
Governance
The provider's own terms govern retention and residency. This layer is where an open-weight model quietly stops being a self-hosting story.
Representative of the hosted-provider tier alongside Fireworks, Groq, DeepInfra and Novita. Useful when the requirement is an open model rather than a private deployment — a distinction the registry's on-prem flag does not by itself make.

Groq

🇺🇸 USA
Groq · Hosted service
Hosted providerProprietaryHosted only
Best at
Very low latency on custom inference hardware
Hardware
Groq LPU — theirs, not yours
Ceiling
A narrower model catalogue than general providers, and no self-hosted path at all.
Governance
Hosted only. The speed comes from hardware you cannot buy, which is also why it cannot be part of an on-premise plan.
Included because latency-first serving is a genuinely different product from throughput-first serving, and it is the clearest example of the trade.