NVIDIA Topograph Treats GPU Placement as a Live Topology Problem

NVIDIA's open-source Topograph discovers cluster topology from cloud APIs or on-prem fabrics and publishes it as Kubernetes labels, Slurm config, or Slinky ConfigMaps. Netics on why topology

Netics feature card for the NVIDIA Topograph topology-aware scheduling article with the official NVIDIA identity
Netics editorial feature card using the official NVIDIA identity

TL;DR

  • NVIDIA released Topograph (open source, dsx-ai-factory/topograph) on September 22, 2026: it discovers cluster topology from cloud APIs or on-premises fabric systems and normalizes it into a common model for schedulers.
  • It publishes that topology as Kubernetes node labels, Slurm configuration, or Slinky ConfigMaps — the formats workload managers already read.
  • Providers cover Google Cloud, Lambda, Nebius, Nscale, OCI, and Crusoe; on-premises support covers InfiniBand and NetQ-managed Spectrum-X or Multi-Node NVLink domains.
  • Topograph regenerates the view on request and on watched cluster changes, so schedulers never act on a stale drawing.
  • Netics' take: Topograph is a data pipeline wearing a scheduling hat — the win is treating topology as live infrastructure, not as a one-time cluster diagram.
  • For operators: start on Kubernetes, verify the labels appear, simulate first — and treat the topology view as monitored infrastructure.
Official NVIDIA Technical Blog hero image for the Topograph post, showing GPU cluster topology discovery
Official NVIDIA Technical Blog hero for "Topology-Aware Workload Scheduling with NVIDIA Topograph" (developer.nvidia.com, September 22, 2026).

The placement problem is a data problem

GPU workload placement is where AI-factory economics are won or lost. NVLink provides 1.8 TB/s of bidirectional bandwidth per GPU in Blackwell and 3.6 TB/s in Vera Rubin; Quantum InfiniBand ports reach 800 Gb/s. A workload spread across distant topology domains crosses shared links, contends, fragments, and — the part vendors rarely say loudly — leaves GPUs consuming provisioned power while waiting on data. In a power-limited AI factory, that is not an efficiency footnote; it is the difference between a node earning its power draw and a node burning it.

The blocker is not that schedulers cannot do topology-aware placement. Slurm and Kubernetes both support it. The blocker is that a scheduler can only act on the topology it observes — and the current topology is exactly what most operators do not have. Manual cluster drawings go stale the moment a rack changes, a switch fails, or a maintenance window re-cables a domain. Topograph's core move is to treat that data as a live pipeline: discover, normalize, publish, and regenerate when the cluster changes.

Official NVIDIA Technical Blog diagram (Figure 1): Topograph architecture — a generation request selects a cloud topology API or network fabric provider path, the API Server validates and dispatches, the Provider discovers and normalizes topology, and the Engine publishes Slurm config, Kubernetes resources, or instance-oriented topology JSON
Official NVIDIA Technical Blog Figure 1: Topograph architecture — generation request through provider selection, API Server dispatch, discovery and normalization, and engine publication to Slurm, Kubernetes, or instance-oriented topology JSON (developer.nvidia.com, September 22, 2026).

Providers and engines: the pipeline, named

Topograph's two concepts are worth understanding as architecture, because they map to real operational roles. Providers discover topology from cloud APIs or on-premises systems and normalize it into a canonical model; engines translate that model into the format each workload manager expects. One discovery layer, many outputs — Kubernetes node labels, Slurm topology.conf, Slinky ConfigMaps, NFD resources, or instance-oriented JSON.

The provider list reads like the current AI-cloud map: Google Cloud, Lambda, Nebius, Nscale, OCI, Crusoe on the managed side; InfiniBand through ibnetdiscover and NetQ for Spectrum-X or Multi-Node NVLink domains on-premises. The provider interface is open, which matters for private clusters — the operators who need this most are exactly the ones with a proprietary fabric vendor no public provider covers.

Original Netics diagram: what Topograph publishes — Kubernetes labels, Slurm topology.conf, Slinky ConfigMaps, NFD resources, and instance-oriented JSON
Original Netics diagram: the five publication formats Topograph writes — Kubernetes node labels, Slurm topology.conf (tree/block), Slinky ConfigMaps, NFD resources (alpha gate), and instance-oriented topology JSON; source: NVIDIA Topograph post.

Stale topology is the silent tax

The operational claim worth testing is regeneration. Topograph's five components — API Server, Node Observer, Node Data Broker, Provider, Engine — keep the view current: the Node Observer watches configured Kubernetes node or pod changes and requests regeneration with retries; duplicate requests are deduplicated behind a typical 15-second aggregation delay; the API exposes /v1/generate, /v1/topology, /v1/lookup, /healthz, and /metrics. In practice that means a failed switch or a node drain updates the labels schedulers read, instead of waiting for someone to remember to redraw the diagram.

The honest caveat is scope. The matrix in the post shows exactly which engine works for which provider — InfiniBand on bare metal or VMs, for example, publishes to Slinky and NFD but not to Kubernetes node labels; MNNVL NVLink partitions only reach the Slinky engine as DRA block topology. This is a young toolkit with real sharp edges, and the "topology-aware" promise is only as good as the provider-engine pairing you actually deploy.

Original Netics diagram: the stale manual cluster view versus Topograph's live regenerated view — on request and on watched cluster changes
Original Netics diagram: the stale-view versus live-view contrast — a manual cluster drawing that decays versus Topograph's regenerated topology published to the scheduler on request and on watched changes; source: NVIDIA Topograph post.

What an operator should do first

Start on Kubernetes, because that is where the feedback loop is easiest to verify. Deploy via Helm, pick a provider that matches where your GPUs live, and confirm the labels appear: fabric.topograph.run/tier-0..N and accelerator.topograph.run/domain. Those labels are immediately consumable by topology-aware schedulers like KAI Scheduler or Kueue's TAS for gang scheduling — and they are the same labels your observability stack can use to explain why a job landed where it did.

Then use the simulation utilities before touching production hardware: kwok-nodes and the Kind/KWOK helpers model node and switch hierarchies with no GPUs attached, which is the right way to validate provider behavior and engine output without a pilot cluster. When you do go live, treat the topology view as monitored infrastructure: hook /metrics into Prometheus, alert on stalled regeneration, and make "scheduler sees current topology" a documented production check rather than an assumption. Topology-aware placement only helps teams that keep the topology honest — and that is precisely what Topograph is for. If you already route local inference with NVIDIA PAIR, Topograph is the placement layer that decides which nodes PAIR has to choose from in the first place. For a practical pass over AI-factory scheduling, the Netics homepage is where we start.

Sources

Source: "Topology-Aware Workload Scheduling with NVIDIA Topograph" — developer.nvidia.com, September 22, 2026.