Inference Performance: Loading, First Tokens, Decoding, and Throughput
Break requests into measurable stages and explain how concurrency, caching, and batching affect waiting time.
Break requests into measurable stages and explain how concurrency, caching, and batching affect waiting time.
Choose a local runtime by artifact compatibility, hardware path, latency, throughput, observability, and security.
Follow a token through MoE routing to distinguish total parameters, active parameters, memory, speed, load balancing, and communication.