Single-Device Deployment: Weights, Memory and Context
This page owns the capacity question: what must fit on a device and what changes as context and concurrency grow. The dated candidates below illustrate those constraints. For model behavior use coding-model selection; for numeric formats and measured speed continue with quantization and inference performance.
“Runs on one GPU” can mean that the weights barely load, that a short request finishes with CPU offload, or that a useful context runs at interactive speed. To decide whether a configuration is practical, check the memory, context length, and speed your task needs. An interactive coding session and an overnight batch can tolerate very different delays.
Start with the memory budget
Raw weight memory is approximately
A 27B model at four bits is about 12.6 GiB before quantization metadata, tensors kept at higher precision, runtime workspaces, and the KV cache. Artifact size is a better starting point than parameter arithmetic when the exact quantization already exists.
For conventional MHA or GQA models, one sequence's KV cache is approximately
where is layer count, is retained sequence length, is cached key/value width per layer, and is bytes per element. MLA and other attention designs use different cache layouts. Batch size and parallel sequences multiply the requirement.
Leave room for output tokens, temporary buffers, the display driver, and ordinary runtime variation. A configuration that survives one prompt with almost no free VRAM is not yet a useful capacity plan.
Reasonable first experiments
These are starting points, not compatibility guarantees. Architecture, quantization format, backend, cache precision, and GPU support can move the boundary substantially.
Current examples worth testing
- Qwen3.6-27B: Apache-2.0, 27B parameters, a vision encoder, 262K native context, and a documented extension path to roughly 1M tokens. The linked card recommends vLLM ≥0.19.0 or SGLang ≥0.5.10 and shows eight-GPU serving examples, not a tested 16 GB quantized configuration or a minimum GPU count. Qwen advises at least 128K context to preserve thinking capability; reducing context to fit is a quality trade-off to evaluate, not a free capacity gain.
- gpt-oss-20b: Apache-2.0, 21B total and 3.6B active parameters. Its model card describes an MXFP4 artifact designed to run within 16 GB memory and requires the Harmony response format.
- Mistral Small 3.1 24B: the model card distinguishes a quantized consumer-GPU path from much larger BF16 serving requirements. It is a useful reminder that model name alone does not state memory use.
A published context limit does not mean the full window fits on one GPU, runs quickly, or retains enough quality for the task.
Offload is a trade, not a failure
CPU or system-memory offload can make a larger model run, but it moves data across a much slower path and may reduce decode speed sharply. The size of that penalty depends on the model, offloaded layers, memory bandwidth, interconnect, backend, prompt length, and batch size. There is no universal tokens-per-second cliff.
Full GPU residency is desirable for latency-sensitive work, not a rule for every workload. An overnight batch may tolerate offload that would make an interactive coding loop unpleasant.
Fit procedure
- Define the task, concurrency, input length, output reserve, and acceptable latency.
- Choose an exact artifact, template, and runtime.
- Start with a short context and one sequence.
- Record peak GPU/RAM use, offload, time to first token, decode speed, and task result.
- Increase context separately from concurrency so the source of a failure is visible.
- Test the longest information dependency the real workload needs, not only a prompt full of filler.
- Keep operational headroom and record the configuration that was actually tested.
A smaller resident model can beat a larger offloaded one in a tool loop. A larger context can also make the system slower without making it better at finding the right evidence. Use the local evaluation protocol to decide with tasks rather than parameter count.
The Accelerate model memory estimator guide shows how to estimate parameter memory for different numeric types without loading all model weights. Compare that estimate with the weight arithmetic above, then measure the complete runtime separately: cache, activations, and temporary buffers still need their own budget.