How to Use, Optimize, and Serve an LLM in Your Production System

A plain-English map of production LLM serving: making the model itself faster, building a serving system that holds P95/P99, and choosing an inference engine. Written as the overview to a three-part technical series.

Tags

Published June 20, 2026.

You shipped the demo. The model works. Stakeholders are happy. Then real traffic hits, and suddenly:

If any of these hit close to home, this article is the map. I also wrote a three-part technical series that goes an order of magnitude deeper on each section, from GPU kernels to cluster scaling. This is the territory; the series is the map you walk through it with.

Why this problem is harder than it looks

Deploying an LLM isn't like deploying a REST API. A Flask endpoint is stateless, horizontally scalable, and usually CPU-bound. An LLM inference server is none of those things.

**1. Two completely different compute phases, back to back.** Every LLM request has two phases. The *prefill* phase reads your prompt — it's compute-intensive, runs fast, and can process many tokens in parallel. The *decode* phase generates the response one token at a time — it's memory-bandwidth-bound and inherently sequential. The bottleneck for each phase is a different hardware resource, so optimizing one without the other gets you half the gains.

**2. The GPU's memory is both the product and the bottleneck.** The model weights live in GPU memory, and so does the KV cache — the stored attention state for every active request. These two fight each other. A 7B model in FP16 alone eats 14 GB of VRAM, so on a 40 GB A100 that leaves 26 GB for batching real users. How you manage that 26 GB decides whether you serve 4 concurrent users or 40.

**3. Latency and throughput are opposites.** The things that make throughput high — large batches, long waits for requests to fill the batch — make latency worse. The things that make latency low waste GPU capacity. Production systems have to serve both interactive users who need sub-500ms responses and background jobs that need to process millions of documents cheaply. You can't use one configuration for both.

**4. The ecosystem moves fast and breaks things.** vLLM, SGLang, TensorRT-LLM, and TGI have all had major changes in 2025 and 2026. What was best practice 12 months ago may be deprecated today — TGI, for example, was fully archived in March 2026, and if you're still building on it you need to migrate.

The three-layer mental model

Every production LLM problem lives in one of three layers. This is the organizing principle of the full series:

┌─────────────────────────────────────────────────────┐
│  Layer 3: ENGINE CHOICE                             │
│  Which inference runtime? vLLM, SGLang, TRT-LLM?   │
├─────────────────────────────────────────────────────┤
│  Layer 2: SERVING SYSTEM                            │
│  Queuing, traffic routing, autoscaling, SLAs        │
├─────────────────────────────────────────────────────┤
│  Layer 1: THE MODEL ITSELF                          │
│  Quantization, attention, GPU kernels, KV cache     │
└─────────────────────────────────────────────────────┘

Most engineers start at Layer 3 — picking an engine — and then discover their problems are actually in Layer 1 or 2. The right order is bottom-up: make the model efficient, then make the serving layer stable, then pick the engine that fits your workload.

Layer 1 — Making the model itself faster

The core insight: decode is starving your GPU

Modern GPUs have hundreds of TFLOPS of tensor core capacity. But when generating tokens one by one, the GPU is mostly *reading weights from memory and doing almost no math on them*. Arithmetic intensity during decode is roughly 1 FLOP per byte of memory read; saturating tensor cores needs over 200 FLOP/byte. For most of a request's lifetime your expensive GPU sits idle at maybe 5–15% compute utilization, bottlenecked entirely on memory bandwidth.

This is why quantization and caching work so well — they don't make the math faster, they reduce how much data the GPU has to haul across the memory bus per step.

The five-level optimization ladder

**Level 1 — BF16 + Flash Attention (free wins, zero quality loss).** Flash Attention rewrites the attention operation to avoid materializing the full N×N attention matrix in GPU memory. Instead of writing intermediate results back to slow HBM and reading them again, it fuses the operation into a single kernel that stays in fast on-chip SRAM. The result is 30–50% faster prefill on prompts longer than ~2K tokens with zero change to model quality. It's already on by default in vLLM and SGLang; if you're running anything else, check that it is.

**Level 2 — INT8 quantization (1.5–2× batch size, negligible quality loss).** Storing weights as 8-bit integers instead of 16-bit floats halves memory use. Perplexity degradation is under 0.1% on standard benchmarks, which is essentially invisible, and you can fit twice the model — or twice as many concurrent users — into the same VRAM. Libraries like `bitsandbytes` make this close to a one-line change.

**Level 3 — INT4 quantization / AWQ / GPTQ (3–4× memory, 2–3× throughput).** Going to 4-bit gives a 4× memory reduction; the challenge is doing it without quality collapse. Two methods are production-grade:

Both are well-supported in vLLM and SGLang.

**Level 4 — Paged KV cache + KV cache quantization (2–3× concurrency).** Without memory management the KV cache fragments — you might have 40% of your VRAM technically allocated but unusable because it sits in non-contiguous chunks. vLLM's PagedAttention solves this the way an OS solves memory fragmentation: by managing the cache in fixed-size pages of 16–64 tokens. GPU memory utilization typically improves from ~45% to ~85%.

On top of that you can quantize the KV cache itself to FP8. The cost is ~5–10% throughput; the benefit is that the same VRAM holds twice as many active requests. At scale that trade is usually worth it.

**Level 5 — Speculative decoding + tensor parallelism.** Speculative decoding uses a small draft model to predict several tokens ahead, then verifies them all at once with the main model. When predictions are mostly right — common phrases and patterns often are — you effectively generate multiple tokens per main-model forward pass, cutting TTFT and per-token latency by 2–3× on the right workloads. Tensor parallelism splits the model across GPUs, the standard approach for 70B+ models that don't fit on one card.

Most teams stop at Level 3 and see TTFT drop from ~800ms to ~250ms with throughput improving from ~30 to ~80 tokens/second. Levels 4 and 5 are for teams with serious production scale.

Layer 2 — Building a serving system that doesn't fall apart under load

The counterintuitive truth: your latency problem is probably queueing, not compute

Once a model is reasonably optimized, P95 and P99 latency are almost never caused by the GPU being slow. They're caused by *requests waiting*. A request that takes 50ms of actual GPU work can easily spend 800ms in queue because the batch scheduler waited for one more request, or because a 4K-token document summary monopolized the prefill slot while 12 chat queries waited.

The traffic lanes pattern

The single most impactful serving change most teams can make is also the simplest: **run interactive traffic and batch traffic on separate replicas.**

Interactive users — chat UI, streaming responses — need low time-to-first-token and can tolerate lower throughput. Batch jobs like document summarization and embedding pipelines need high throughput and can tolerate higher latency. These two workloads want opposite things from the scheduler, and forcing them to share a GPU means one will always suffer. vLLM lets you tune this per instance:

# Interactive lane: small batches, fast dispatch
--max-num-seqs 4

# Batch lane: large batches, maximum throughput
--max-num-seqs 16

Add short-prompt vs. long-prompt splitting and you've removed the most common source of P99 spikes in production.

Measure the right things

Most teams log "total request time," which is nearly useless for diagnosis. Log these instead, on every request:

Per-lane P99 matters especially. If you mix 50-token chat queries with 4K-token RAG summaries, a global P99 will always look terrible and will always be caused by the wrong thing.

Backpressure, admission control, graceful degradation

Production systems need a way to say "no" before they fall over:

Cold start and autoscaling

LLM replicas are expensive to start. Weight loading, KV cache allocation, and warmup can add 30–90 seconds. During an autoscale event the new replica can't help for over a minute, and during that minute the remaining replicas absorb the full load spike. Two mitigations:

  1. Keep `minReplicaCount: 2` at all times. Never scale to zero for production models.
  2. Scale on `vllm:num_requests_waiting` (queue depth), not CPU/GPU utilization. By the time utilization-based autoscaling triggers, users are already waiting.

Layer 3 — Choosing the right inference engine

The decision shortcuts

The full picture — why all three layers matter together

Here's how the layers interact in a real scenario. Suppose you have a 7B model serving a chat application with 500 concurrent users:

  1. **Layer 1 (model optimization):** you apply AWQ INT4 quantization. The 14 GB FP16 model shrinks to ~4 GB, VRAM is freed for KV cache, and TTFT drops from 800ms to ~250ms.
  2. **Layer 2 (serving system):** you split interactive and batch traffic into separate replicas, add admission control at depth > 20 requests, and configure KEDA to scale on `vllm:num_requests_waiting` with `minReplicaCount: 2`. P99 drops from 3.5 seconds to ~600ms.
  3. **Layer 3 (engine choice):** you're using SGLang because your chat product has a long shared system prompt, so RadixAttention caches it automatically. Every request after the first skips re-computing 500 tokens of prefill, and throughput increases by ~29%.

None of these three layers alone gets you there. Model optimization means nothing if the queueing system serializes every request. Serving-system tuning means nothing if the model is eating all the VRAM. Engine choice means nothing if you haven't quantized.

TL;DR — the one-page cheat sheet

If this was useful, the three-part series goes an order of magnitude deeper on each section.