What actually runs an open-weight model from the registry, and what sits in front of it. These tools are routinely compared as if they were alternatives; mostly they are not. They occupy distinct layers, and the common expensive mistake — putting a developer wrapper behind a production API — is a layer error, not a benchmark error.
The registry computes a weight footprint for every self-hostable model — parameters multiplied by bytes per weight, at Q4 and FP8, with the hardware class that implies. It is derived arithmetic rather than a vendor claim, and deliberately a floor rather than a requirement.
What it takes to run a model
The registry's weight footprint is a floor, not a requirement. Four things decide the rest, and all four belong to this layer rather than to the model.
Total parameters set the memory; active parameters only set the compute
A mixture-of-experts model must hold every expert in memory even though a fraction fires per token. DeepSeek V4 Pro is 49B active and still needs roughly 880 GB at Q4, because all 1.6T weights have to be resident. Reading the active count as the requirement is the most common and most expensive mistake in this area, which is why the registry prints both figures side by side.
KV cache is not in that number, and at long context it rivals the weights
Cache grows with context length and again with every concurrent request, so a model that fits a card at 8K context may not fit it at 128K, and almost certainly will not serve ten users there. Computing it needs layer counts, KV-head counts and head dimensions that most closed vendors never publish, so the registry does not pretend to. Budget headroom on top of the printed figure, and treat a 1M-context claim as a statement about the architecture, not about your GPU.
The engine decides what happens when it does not fit
llama.cpp will offload layers to system RAM and keep going slowly; vLLM effectively wants residency and will fail rather than degrade; ExLlamaV3 exists to squeeze a model onto a card it should not fit; KTransformers offloads MoE experts to CPU on purpose. The same model on the same hardware is runnable or not depending on which of these you pick — which is the argument for choosing the engine before the model.
Quantisation is a quality decision, not just a size one
Q4 roughly halves FP8, and the loss is small on many tasks and not small on some — long-context reasoning and code generation degrade earlier than chat. The registry shows both so the trade is visible; measure on your own evaluation before assuming the cheap tier is free.
❯
Add
llama.cpp
🌐 —
ggml-org (open source) · MIT
EngineOpen sourceSelf-hostable
Best at
Portability, edge and CPU or hybrid offload
Concurrency
Parallel slots (--parallel), single-node
Hardware
Effectively everything: x86 AVX2/AVX-512, ARM NEON, CUDA, Metal, Vulkan, ROCm
Manual operations, single-node thinking, no fleet-serving story. It is a library and a server binary, not a platform.
The substrate of the local-model world. It created GGUF, and most consumer-hardware tools either wrap it or exist in reaction to it. Go direct when you need embedded deployment, unusual hardware, aggressive CPU/GPU layer splitting, or one less dependency between you and the weights.
Linux-centric with a heavy footprint, and GGUF is second-class. Not what you want on a laptop.
The default answer for production self-hosting, and the one the industry converged on — Hugging Face's own TGI now points here. Requests are packed into shared compute passes rather than run end to end, which is why throughput climbs with concurrency instead of flattening. Speaks OpenAI, so a gateway fronts it identically to a hosted vendor.
Smaller ecosystem and narrower hardware coverage than vLLM.
vLLM's sibling, optimised for reuse rather than raw batching. Requests sharing a system prompt or conversation history skip recomputation, which is the exact shape of agent loops, tool-calling chains and multi-turn chat behind a fat system prompt. Choose it over vLLM when prefix reuse is high and time-to-first-token matters.
Highest performance ceiling, highest operational cost: models are compiled per configuration, and it leaves the API surface to Triton, NIM or Dynamo.
The choice when the hardware is already bought and the last 30% of utilisation is worth an engineering budget. Rarely the right first step, and rarely the wrong last one.
One vendor's hardware, and no answer at all for a rack.
Relevant precisely when the fleet is MacBooks — an underrated case for organisations that issue Macs and want inference to stay on the endpoint. Reached most often through a wrapper rather than directly.
Smaller community and thinner documentation in English than vLLM or SGLang.
A credible third option in the serving-engine tier, strongest on the models its authors ship. Worth benchmarking against vLLM for a specific model rather than adopting on reputation.
In maintenance mode: minor bug fixes, documentation and lightweight maintenance only. The repository was archived on 21 March 2026.
For a period the default self-hosted server, and now the clearest example of why this page records status. Hugging Face's own README points new users to vLLM, SGLang, llama.cpp and MLX. If you are running it, the migration is not optional so much as already overdue.
Not a production API server. Throughput flattens almost immediately under concurrent load while a real serving engine climbs — the single most common expensive mistake in this space is putting it behind a shared endpoint.
An OpenAI-compatible endpoint on :11434 and a model library, which is why it appears in so many agent integrations: its own CLI documents launch profiles for Claude Code, Codex, OpenCode, VS Code and Droid. That is also why it shows up on the agents page as the runtime behind several local-model claims. Right for dev loops, CI fixtures and single-user desktop work.
Proprietary, and its API gives less control over parallel and multi-model routing than Ollama's environment variables.
Version 0.4.0 (28 January 2026) split the engine from the GUI: llmster runs headless on a server or in CI, with continuous batching, an n_parallel setting and a stateful REST API. The real advantage remains evaluation ergonomics — browsing Hugging Face in-app, quantisation recommendations against your actual RAM and VRAM, and visible tokens per second while you compare.
An open-source desktop assistant that stays offline
Concurrency
Single-user desktop
Hardware
macOS, Windows, Linux desktops
Formats
GGUF
Ceiling
Desktop scope, and a smaller ecosystem than Ollama or LM Studio.
The open alternative in the GUI tier, for people who want LM Studio's shape without a proprietary binary. Same engine underneath as most of this layer.
Single-binary local deployment with heavy sampling control
Concurrency
Limited
Hardware
Anything llama.cpp supports
Formats
GGUF
Ceiling
A hobbyist and single-user tool by design; no fleet story.
A one-file llama.cpp distribution with an unusually deep set of sampling and context controls. Included because it is what a lot of local deployment actually is, rather than what conference talks assume it is.
You are renting someone else's routing decisions. Fine for evaluation, awkward as a production path for anyone with a residency or assessment obligation.
Governance
A single request may be fulfilled by any of several sub-providers, whose retention policies differ. The model name is not the control surface — the provider allowlist and data-policy settings are. Treat it as a catalogue and price-discovery instrument; if it goes near production, put it behind your own gateway with a hard allowlist rather than beside it.
One key, one OpenAI-shaped API, hundreds of hosted models, unified billing — the fastest way to compare models or price-shop. Teams typically leave not for a bigger catalogue but because the problem turns into governance, observability and residency.
You own the uptime. A gateway is another production dependency, and it fails in front of everything.
Governance
Keys, budgets, rate limits, model allowlists and request logs stay on your infrastructure, which is what makes it the usual answer where OpenRouter is not. Also the point at which self-hosted engines and hosted vendors become interchangeable behind one contract.
The pragmatic centre of most serious stacks: one OpenAI-shaped internal contract, with vLLM, Azure, Bedrock or anything else swappable behind it. Its price file is also one of the sources this site refreshes registry pricing against.
Portkey · Open-source gateway with a hosted control plane
GatewayOpen sourceSelf-hostable
Best at
Guardrails and governance in front of the model call
Ceiling
The richest features sit in the hosted control plane, so the self-hosted story is thinner than the marketing implies.
Governance
Positions policy — guardrails, PII handling, request filtering — as gateway concerns rather than application ones, which is the right layer for them if you are enforcing across many teams.
Chosen where the requirement is governance rather than routing, and where a per-request policy engine belongs in the path.
Observability-first, so it is a weaker policy engine than a gateway built for control.
Governance
Logging, tracing and cost attribution per request. The reason to care: without this layer nobody can answer what a prompt cost, what it contained, or which provider served it.
Often deployed alongside rather than instead of a routing gateway.
Hosted only — the gateway itself is another vendor in the path, which is a strange shape if the reason for a gateway was vendor control.
Governance
Caching and rate limiting sit at the edge, and requests transit Cloudflare. Straightforward for cost control; a question to answer explicitly where the point of the exercise was data residency.
Cheap to adopt if you are already a Cloudflare tenant, and the analytics arrive without instrumentation.
Distributed serving with prefill and decode disaggregated
Concurrency
Cross-node scheduling and KV-cache-aware routing
Hardware
NVIDIA clusters
Formats
Engine-dependent — vLLM, SGLang, TensorRT-LLM
Ceiling
Only earns its complexity at multi-node scale.
Splits the compute-bound prefill phase from the memory-bound decode phase onto different resources, which is the current frontier of serving efficiency and irrelevant below a certain size.
Someone else owns the engine, the hardware and the retention policy. Open weights served by a third party is not the same privacy posture as open weights on your own metal.
Governance
The provider's own terms govern retention and residency. This layer is where an open-weight model quietly stops being a self-hosting story.
Representative of the hosted-provider tier alongside Fireworks, Groq, DeepInfra and Novita. Useful when the requirement is an open model rather than a private deployment — a distinction the registry's on-prem flag does not by itself make.