Local Inference Runtimes
A local deployment has at least four separable layers:
Changing any layer can change quality, memory, latency, or tool behavior. An HTTP 200 only proves that a request returned.
Runtime Map
SGLang, Transformers, MLX, TensorRT-LLM, ROCm-specific, and vendor runtimes may be better for a measured need. The table is not an exhaustive rank.
Compatibility Gate
Before evaluation, freeze:
- exact model repository, revision, filename, and quantization;
- tokenizer, chat template, reasoning format/parser, and stop tokens;
- tool-call format and structured-output behavior;
- native/extended context, KV-cache type, batch/concurrency, and GPU offload;
- runtime and driver versions, checked against the owner's hardware requirements for Ollama, LM Studio, or vLLM.
OpenAI-compatible APIs standardize a useful subset of request shape. They do not guarantee identical reasoning fields, token accounting, tool IDs, streaming chunks, logprobs, cancellation, or errors.
Worked Choice
For one 16 GB consumer GPU and one interactive user:
- start with a supported quantized artifact and native context;
- use llama.cpp when exact GGUF/offload control matters, or Ollama/LM Studio for the quickest manual comparison;
- measure time to first token and decode speed at concurrency 1;
- verify chat and tool formatting before judging intelligence;
- move to vLLM only if its supported artifact and batching/serving benefits solve a real workload.
For a 24 GB always-on multi-client service, throughput, queueing, cancellation, metrics, and isolation may outweigh desktop convenience. This is a workload distinction, not a claim that one engine always produces better tokens.
Evaluation handoff
This page owns runtime compatibility, serving choices, and operational risk. The shared Local Model Evaluation protocol owns run cards, controlled factor changes, latency and throughput metrics, memory measurements, accepted-result rate, and repeated task results.
Security and Operations
For locality-sensitive work, distinguish the local server from where its selected model executes. Ollama cloud models can run through localhost:11434. When local-only execution is required, use a downloaded local model and disable cloud features with OLLAMA_NO_CLOUD=1, then restart Ollama. This disables Ollama cloud models and web search, not external tools in another harness; it is not a network sandbox.
Bind to loopback by default, add authentication before any network exposure, isolate model-serving credentials, review telemetry, cap request/context/output size, and test cancellation. Treat model files and templates as supply-chain inputs. Do not let a compatible endpoint become an unaudited internal network proxy.
Bias Register
This note favors Linux-like developer workflows and popular open ecosystems. Apple Silicon, AMD, Windows, mobile NPUs, energy use, accessibility, and production fleet operations need separate measurements. Official runtime documentation establishes supported features; performance claims require the actual hardware and workload.