← Lab notes

How to measure decode bandwidth utilization without hardware counters

A vendor-neutral way to compute memory-bandwidth utilization (MBU) for LLM decode from model geometry and client-side timing — including MoE expert routing and MLA KV caches — and how we validated it against a simulated roofline engine.

Autoregressive decode is usually limited by memory bandwidth: each step streams the active weights and the KV cache from HBM to produce one token per sequence. Decode bandwidth utilization (MBU) asks how much of the rated HBM bandwidth that traffic actually uses. Low MBU at a latency-bound operating point means the software stack is leaving throughput on the table.

Hardware counters can measure bytes moved, but they have two problems for a vendor-neutral benchmark. They are exposed differently on every vendor and are often locked down inside cloud VMs and containers. And they count all traffic, including redundant reads from unfused kernels, so a less efficient stack can look more bandwidth-hungry. We want the opposite: the bytes the model must read, divided by the time it took.

The definition

TokenWatt Bench derives MBU from three things it already has: the model's config.json, the spec, and client-side timing.

T   = output tokens/s for the node                 (measured)
τ   = token-weighted mean inter-token latency      (measured)
R   = data-parallel engine replicas on the node    (spec)
s   = 1/τ          decode steps per second per replica
B   = T·τ / R      sequences decoded per step per replica

W(B) = dense weight bytes per step
     + MoE layers × E[unique experts hit by B tokens] × bytes per expert
KV/s = T × mean context per generated token × KV bytes per token

MBU  = (R · s · W(B) + KV/s) ÷ (N_GPU × rated bandwidth per GPU)

Every term is either measured on the client or read from public model metadata. That makes the number comparable across vendors, engines and parallelism layouts. The weights themselves are never needed.

A few details matter:

  • Mean context per generated token. Token i of a request with input length L reads a KV cache of length L + i. Summed over a request with n output tokens, that's n·L + n(n−1)/2. Bench sums this over every completed request in the window and divides by total output tokens.
  • Inter-token latency includes everything. With continuous batching and chunked prefill, decode steps are interleaved with prefill work. Using the measured τ means MBU reflects the effective decode rate at the SLO operating point, not an idealized kernel benchmark.
  • The LM head is read every step. It's usually kept in BF16 even when the rest of the model is FP8, so Bench counts it at 2 bytes per parameter.

MoE: which experts get read?

For a mixture-of-experts layer with E experts and top-k routing, a batch of B tokens touches a set of distinct experts, and only those are read. Under uniform routing, the expected number of distinct experts is:

E · (1 − (1 − k/E)^B)

For Qwen3-30B-A3B (48 MoE layers, 128 experts, top-8, about 4.7 M parameters per expert), a batch of 1 reads 8 experts per layer. By a batch of 64, nearly all 128 are touched every step. That is why MoE decode cost climbs steeply at small batch and then flattens.

Real routing is skewed, which makes collisions more likely and touches fewer distinct experts. So the uniform assumption is an upper bound on weight traffic, and MBU reads somewhat high at small batch. At the batch sizes typical of SLO-bound serving the bias vanishes, and because every configuration in a comparison uses the same assumption, comparisons hold.

KV cache geometry, including MLA

For standard grouped-query attention, KV elements per token are 2 × layers × kv_heads × head_dim. Qwen3-30B-A3B has 48 layers, 4 KV heads and head dimension 128, so that's 96 KiB per token in BF16.

DeepSeek-style multi-head latent attention caches a compressed latent instead: layers × (kv_lora_rank + qk_rope_head_dim) elements per token. For DeepSeek-V3 that's 61 × (512 + 64). Bench reads these fields from config.json and handles both automatically. Any field can be overridden in the spec for architectures it doesn't know yet.

We check the geometry parser against published parameter counts. Llama 3 8B comes out at 8.0 B, Mixtral 8×7B at 46.7 B, and DeepSeek-V3 at roughly 671 B.

Validating it against a roofline simulator

To test the derivation end to end without GPUs, Bench ships a simulated OpenAI-compatible engine. It does continuous batching with a bandwidth roofline. Each step takes:

(weight bytes + Σ context × KV bytes per token) ÷ (rated bandwidth × efficiency)
  + prefill time of newly admitted requests + fixed overhead

The efficiency knob plays the role of the software stack. Bench treats the simulator as a black box, drives it with Poisson load, finds the SLO boundary and derives MBU from timing alone. The derived MBU should then land close to the efficiency the simulator was given.

It does:

  • With prefill and per-step overhead set to zero: the derived MBU tracks the configured efficiency within a few points. The only gap is asyncio scheduling jitter, and our end-to-end test asserts this.
  • With realistic prefill load: MBU comes out lower than the configured efficiency, by about the fraction of each step spent on prefill and overhead. For example, a simulated stack at 30% efficiency measured about 21% MBU at its SLO boundary. That's the point of the metric: it reports bandwidth delivered to decode at the operating point you'd actually run.

What MBU is and isn't

MBU is a diagnostic. It tells you how much headroom the software leaves, and it's the number that moves when someone fuses a dequantization kernel or fixes an expert-dispatch layout.

It is not the headline. At large batch, decode can become compute-bound or interconnect-bound, and MBU will plateau for reasons that have nothing to do with kernel quality. That's why every TokenWatt report leads with SLO goodput and tokens per joule, and shows MBU alongside them.

The full definitions are in the methodology.