← Lab notes

How many GPUs does an internal ChatGPT need? Fewer than you think — and that's the cost problem

Sizing private LLM inference from a measured H100 benchmark. One H100 can carry the peak chat traffic of roughly 12,000 employees, but at 1,000 employees it sits 99.6% idle and a million tokens costs $67.70 instead of $0.30. Utilization, not GPU count, decides the economics. Includes a calculator.

Companies that can't send their data to an outside API keep asking the same first question: how many GPUs do we need to run our own ChatGPT for our employees?

We can answer it from a measurement rather than a spec sheet. In lab notes #2, one H100 served Qwen3-30B-A3B at 3,996 output tokens per second (9.5 requests per second) while holding p99 time-to-first-token under 2 seconds and p99 time-per-output-token under 50 ms. That's the throughput users actually experience as responsive, not a peak number.

The answer turns out to be "very few GPUs". The harder question is what each token costs once you have them.

The assumptions

Assumption Value Why
Daily active users 50% of employees typical of internal assistant rollouts after the novelty wears off
Requests per active user 20 per workday chat questions, rewrites, summaries
Busiest hour 20% of the day's requests traffic bunches up mid-morning
Request shape our chat-mix-v1 trace: input median 1,024 tokens, output median 256 (mean ≈ 421) the workload we measured
Planning headroom size GPUs to 70% of the measured SLO capacity, plus one spare don't run at the edge; survive a failed GPU
GPU price $4.29 per H100-hour, rented around the clock Lambda on-demand list price on 2026-10-07

Change any of these in the calculator below. The only number that comes from our lab is the per-GPU capacity.

The result

Employees Requests/day Peak req/s H100s (+1 spare) Average utilization Cost/month Cost per 1M output tokens
1,000 10,000 0.56 1 + 1 0.4% $6,263 $67.70
5,000 50,000 2.78 1 + 1 2.2% $6,263 $13.54
10,000 100,000 5.56 1 + 1 4.4% $6,263 $6.77
25,000 250,000 13.9 3 + 1 5.5% $12,527 $5.42
50,000 500,000 27.8 5 + 1 7.3% $18,790 $4.06
100,000 1,000,000 55.6 9 + 1 8.8% $31,317 $3.38

At 70% of its SLO capacity, one H100 can carry the peak chat traffic of about 12,000 employees under these assumptions. Most companies need two GPUs, and one of them is the spare.

The same table shows the catch. At full utilization this setup costs $0.30 per million output tokens. At 1,000 employees it costs $67.70, because the GPUs are idle 99.6% of the time. Even at 100,000 employees, average utilization stays under 10%. Chat traffic follows the workday, and the GPUs are paid for around the clock.

What this means

  1. GPU count is the easy part. For chat on a model of this size, the hardware is small. Don't plan a cluster before you've measured your own traffic.
  2. Utilization decides the economics. The cost per token falls almost exactly as utilization rises. Ways to raise it:
    • Share the GPUs across workloads. Put chat, document summarization, coding assistants and RAG on the same pool instead of buying per project.
    • Fill the nights. Batch jobs (classification, extraction, embedding refreshes, evaluation runs) can take the 16 idle hours.
    • Rent elastically if you can. On-demand or reserved-plus-burst capacity follows the workday. Owned hardware has to earn its keep 24/7.
    • Right-size the model. This is a mixture-of-experts model with about 3 B active parameters. A dense 70 B model needs several times more GPU per request; use it only where the quality difference matters.
  3. Privacy is the reason to self-host; utilization is how you make it pay. If the data can't leave, compare your effective cost per token, not the full-utilization number, against what the business would accept.

Calculator

GPUs—
Peak load—
Average utilization—
Cost per month—
Cost per 1M output tokens—

Capacity per GPU: 9.5 requests/s and 3,996 output tokens/s at p99 TTFT ≤ 2 s and TPOT ≤ 50 ms — measured on one H100 SXM with Qwen3-30B-A3B-FP8 and vLLM 0.31.0 on our chat-mix trace. Includes one spare GPU. 22 workdays and 730 billed hours per month.

What this doesn't cover

  • Other models. The capacity comes from one MoE model with about 3 B active parameters. Dense and larger models serve far fewer requests per GPU, and we haven't measured them yet.
  • RAG and agents. Long retrieved contexts and multi-step agent loops change the request shape a lot. Agents can multiply requests per task by ten or more. Measure those traces before sizing.
  • Owned hardware. We price rented GPUs. Owned hardware swaps the hourly rate for depreciation, power and operations, but the utilization argument is the same.
  • Peaks beyond the busiest hour. Launch days and all-hands demos can spike well above it. The spare GPU and the 70% headroom absorb some of that, not all.

If you're sizing private inference for your own workload, we can measure it on your traffic, your model and your latency targets. Get in touch.