Benchmarking LLM Inference Needs a Real Load Client

NVIDIA's AIPerf replaces GenAI-Perf with a multiprocess load client that stops being the bottleneck. Netics on why LLM inference numbers are only worth as much as the client generating them.

Netics feature card for the NVIDIA AIPerf LLM inference benchmarking article with the official NVIDIA identity
Netics editorial feature card using the official NVIDIA identity

TL;DR

  • NVIDIA shipped AIPerf, the designated successor to GenAI-Perf, on the Developer Blog (September 18, 2026).
  • It is a multiprocess load client — workers generate traffic, record-processors handle results, coordinated over ZMQ — so the client stops being the bottleneck.
  • 15+ endpoint types, ShareGPT and trace replay (Mooncake, Baseten, WEKA AgentX), and constant/Poisson/gamma arrival patterns with reproducibility flags.
  • Core metrics: TTFT, ITL, request latency, output token throughput, with p25-p99 breakdowns and optional GPU telemetry.
  • Netics' take: the tool change is real, but the discipline is the same — percentile reading, load shaping, and honest client capacity are what make inference numbers trustworthy.
Official NVIDIA Technical Blog screenshot: the AIPerf metrics summary table at the end of a benchmark run, with effective, active and summary statistics and percentile columns
Official NVIDIA Technical Blog screenshot: AIPerf end-of-run metrics summary (effective, active, summary statistics with percentile breakdowns); source: developer.nvidia.com — Benchmarking LLM Inference at Scale with AIPerf

The client was the silent variable

Every LLM inference benchmark you have ever quoted has a hidden actor: the load client. If the client is a single-process script — curl loops, a hand-rolled asyncio harness, or GenAI-Perf itself on top of Perf Analyzer — Python's GIL caps concurrency, and the numbers you get describe the client's ceiling as much as the server's. NVIDIA's answer is a ground-up rewrite: AIPerf splits load generation (worker processes) from result handling (record-processor services) and coordinates them over ZMQ, so a single benchmark client can saturate a real server without becoming the bottleneck itself.

That is the architectural difference that matters. GenAI-Perf ran on top of Perf Analyzer, which made it a wrapper around an older stack; AIPerf is a clean break, and the migration guide exists precisely because present workflows need porting. For anyone who has fought a benchmark where raising concurrency changed nothing because the client saturated first, this is the fix in the right layer.

Original Netics diagram: one-off curl and asyncio scripts versus AIPerf's multiprocess architecture — worker processes generating load, record-processors handling results, coordinated over ZMQ
Original Netics diagram: single-process load scripts (GIL-capped) versus AIPerf's multiprocess client — workers generate load, record-processors collect results, ZMQ coordinates; source: NVIDIA AIPerf post

Percentiles are the skill, not the tool

AIPerf reports the core four — TTFT (time to first token), ITL (inter-token latency), request latency, and output token throughput — in percentile breakdowns from p25 to p99 with min/max/average and standard deviation. The percentile work is where the real engineering lives: a server with a healthy mean TTFT and an outlier p99 looks fine in aggregate and fails in production. AIPerf's load-shaping knobs feed those percentiles honestly — constant, Poisson, and gamma arrival patterns with tunable burstiness, gradual ramping for concurrency and request rate, and synthetic distributions for variable input/output sequence lengths.

Three flags deserve special attention because they are reproducibility, not decoration. --streaming is required to measure TTFT and ITL at all: without it the server batches the full response, and there are no first-token or decode-token events. min_tokens and ignore_eos pin the output length; without them the model stops whenever it naturally finishes, output counts drift, and your throughput numbers are not reproducible across runs. --random-seed 42 makes the Poisson timing and synthetic length draws reproducible, so the same command produces the same request sequence — the difference between a headline number and a test you can re-run after a config change.

Official NVIDIA Technical Blog screenshot: AIPerf metrics summary from the Poisson arrival pattern run, showing wider percentiles under contended load
Official NVIDIA Technical Blog screenshot: AIPerf summary statistics from the Poisson arrival-pattern run, with visibly wider distributions than the static baseline; source: developer.nvidia.com

Replaying real traffic beats inventing your own

The quiet capability is trace replay. AIPerf supports public datasets like ShareGPT and replay formats from Mooncake, Baseten and WEKA AgentX, which means you can feed a server the shape of traffic it will actually see rather than a synthetic average. Combined with the arrival-pattern controls, that turns a benchmark from a lab exercise into a capacity-planning input: a database vendor's p50 chart tells you nothing about your users' request-rate jitter, but a Poisson run at your expected concurrency tells you the p99 your SREs will actually feel.

GPU telemetry lands in the same run output when DCGM or pynvml is available — power draw, utilization, memory. Correlating a latency spike with a memory-pressure event no longer requires a separate profiling session. For the infrastructure teams this blog serves, that is the difference between a vague complaint about a slow LLM stack and a specific diagnosis: the decode phase is memory-bound at 90% utilization.

Original Netics diagram: the four core metrics AIPerf reports — TTFT, ITL, request latency, output token throughput — with percentile breakdowns
Original Netics diagram: the four core AIPerf metrics — TTFT, ITL, request latency, output token throughput — and the percentile breakdowns the skill depends on; source: NVIDIA AIPerf post

What this changes for an inference rollout

Start by re-running your existing GenAI-Perf scenarios under AIPerf with --streaming and pinned output lengths; the migration guide covers the flag deltas, and the comparison will show where the old client was the limit. Second, adopt one real load shape — Poisson at your expected request rate — and record p50, p90, p99 instead of the mean; that is the number that survives contact with production. Third, if your traffic has recognizable patterns (bursty batches, long-context users), replay a captured trace rather than a synthetic distribution and keep the repeat command line AIPerf prints as your regression test. The tool replaced GenAI-Perf, but the underlying habit — verifying that the measurement client is not part of the bottleneck — is the perimeter to keep.

The broader point connects to everything Netics argues about AI infrastructure: routing decides where inference runs, and benchmarking decides whether you can trust the result after routing. A load client you can saturate-and-trust is the missing third leg for teams buying GPUs on vendor benchmarks or capacity-planning their own vLLM fleets. That, more than the tool swap, is the reason to read the migration guide this quarter. If you want a practical pass over your inference measurement setup, the Netics homepage is where we start.

Sources

Source: "Benchmarking LLM Inference at Scale with AIPerf" — developer.nvidia.com, September 18, 2026.