← Lab notes

Lab notes #3: Nine providers, one Llama 3.3 70B — same accuracy, 18× different speed

We audited every OpenRouter provider of Llama 3.3 70B Instruct with the same GSM8K suite and latency probe. Accuracy was statistically indistinguishable (94.0–98.0%), decode speed ranged from 19 to 347 tokens/s, and output prices varied 7×. A naive first pass wrongly scored two providers at 0% and 44%.

When the same open model is served by many providers, buyers usually compare price and maybe speed. They rarely check quality. A provider can quantize harder, change kernels or cap generation, and the model name stays the same.

So we pinned every request to one provider at a time and measured them all the same way. The model was Llama 3.3 70B Instruct, chosen because it has many providers on OpenRouter and they declare different precisions (FP8, BF16, FP16, or nothing).

Short version:

  • Accuracy: no meaningful difference. All nine measurable providers scored between 94.0% and 98.0% on 200 GSM8K questions. That spread is within sampling noise for a 200-item test. Providers declaring FP8 were not worse than those declaring FP16 or BF16.
  • Speed: 18× apart. Median decode speed ranged from 19 to 347 tokens per second. Median time to first token ranged from 330 to 766 ms.
  • Price: 7× apart, and not tied to speed. Output prices ran from $0.32 to $2.25 per million tokens. The cheapest provider was also the most accurate, and the slowest.
  • A naive audit gets this badly wrong. Our first pass scored Cloudflare at 0% and Novita at 44%. The cause was API incompatibilities and rate limits, not model quality.

The whole audit cost $1.13 in API usage. The raw data is at /data/audits/llama-3.3-70b-2026-10-07.json.

Method

Model meta-llama/llama-3.3-70b-instruct via OpenRouter, 2026-10-06/07
Routing each request pinned with provider.only = [<endpoint>], allow_fallbacks = false; a routing check confirmed the serving provider before each audit
Accuracy GSM8K, the same 5-shot prompts, first 200 test questions, temperature 0, stop at the next "Question:", max 384 tokens, scored on the final number
Latency 20 streamed requests per provider, sent one at a time (about 1,000 input words, up to 256 output tokens) — we don't load-test other people's APIs
Failures rate limits and server errors retried up to 4 times with backoff; remaining failures counted separately

Results

Provider Declared precision GSM8K (200) ±95% CI Failed TTFT p50 Decode p50 $ in / out per 1M
DeepInfra (turbo) FP8 98.0% ±1.9 0 514 ms 19 tok/s $0.10 / $0.32
AkashML FP8 96.5% ±2.5 0 732 ms 26 tok/s $0.20 / $0.52
SambaNova (turbo) not declared 96.5% ±2.5 0 428 ms 287 tok/s $0.45 / $0.90
Parasail FP8 95.5% ±2.9 0 624 ms 67 tok/s $0.22 / $0.50
Groq not declared 95.5% ±2.9 0 330 ms 347 tok/s $0.59 / $0.79
Together not declared 95.5% ±2.9 0 643 ms 68 tok/s $1.04 / $1.04
CoreWeave FP16 95.0% ±3.0 0 382 ms 82 tok/s $0.71 / $0.71
Novita BF16 94.5% ±3.2 1 766 ms 52 tok/s $0.14 / $0.40
Cloudflare FP8 94.0% ±3.3 3 743 ms 64 tok/s $0.29 / $2.25

Google Vertex's two endpoints were listed but couldn't be routed to: every pinned request returned "No endpoints found". They aren't in the table.

Is a 4-point spread real? The best and worst providers differ by 4.0 points; on 200 questions each, that gives z ≈ 2.0 before correcting for comparing nine providers at once. After correction it isn't significant. On this test, every provider serves Llama 3.3 70B at essentially the same quality. A harder suite or more samples could still separate them.

Speed and price are separate decisions

  • If you need interactive speed, Groq (347 tok/s) and SambaNova (287 tok/s) are in a different class. The rest decode at 19–82 tok/s from a single request.
  • If you need cheap bulk tokens, DeepInfra's turbo endpoint costs $0.32 per million output tokens and scored highest. It's also the slowest, which is fine for batch jobs and frustrating for chat.
  • Price doesn't buy speed. Cloudflare's output price is the highest on the list at $2.25, with mid-pack speed. Together charges $1.04 for 68 tok/s.

These latencies come from one client location in the US, one request at a time, through OpenRouter, on one evening. Under load, and from other regions, they will differ. Treat them as a snapshot, not a ranking.

The audit that almost went wrong

Our first pass sent the same requests to every provider with a fixed seed and eight requests in parallel. It produced this:

Provider First pass After the fix What went wrong
Cloudflare 0.0% 94.0% rejects the seed parameter; all 200 requests returned HTTP 400
Novita 44.0% 94.5% rate limits (HTTP 429) on parallel requests
Groq 85.0% 95.5% intermittent failures
AkashML 87.5% 96.5% intermittent failures

Had we counted every failed request as a wrong answer and published, we'd have accused two providers of serving a broken model. Bench now retries rate limits and server errors with backoff, drops seed when a provider rejects it (at temperature 0 it doesn't change results), and reports failures separately from wrong answers.

If you compare providers yourself: check failure rates before you trust a score.

Not covered

  • Harder or longer tasks: reasoning, long context, tool use, code. GSM8K is a quick sanity check, not a full quality evaluation.
  • Throughput under load, and latency from other regions or times of day.
  • Variation over time: providers change hardware, kernels and quantization without notice. That's why audits like this need repeating.

Reproduce

cd bench
python scripts/make_gsm8k_suite.py --out suites/gsm8k-200.jsonl
python scripts/audit_openrouter.py --model meta-llama/llama-3.3-70b-instruct --concurrency 4

Next: the same audit for Qwen and DeepSeek models, where providers declare FP4 and FP8 side by side.