NVIDIA Topograph Treats GPU Placement as a Live Topology Problem
NVIDIA's open-source Topograph discovers cluster topology from cloud APIs or on-prem fabrics and publishes it as Kubernetes labels, Slurm config, or Slinky ConfigMaps. Netics on why topology
TL;DR
- NVIDIA released Topograph (open source, dsx-ai-factory/topograph) on September 22, 2026: it discovers cluster topology from cloud APIs or on-premises fabric systems and normalizes it into a common model for schedulers.
- It publishes that topology as Kubernetes node labels, Slurm configuration, or Slinky ConfigMaps — the formats workload managers already read.
- Providers cover Google Cloud, Lambda, Nebius, Nscale, OCI, and Crusoe; on-premises support covers InfiniBand and NetQ-managed Spectrum-X or Multi-Node NVLink domains.
- Topograph regenerates the view on request and on watched cluster changes, so schedulers never act on a stale drawing.
- Netics' take: Topograph is a data pipeline wearing a scheduling hat — the win is treating topology as live infrastructure, not as a one-time cluster diagram.
- For operators: start on Kubernetes, verify the labels appear, simulate first — and treat the topology view as monitored infrastructure.

The placement problem is a data problem
GPU workload placement is where AI-factory economics are won or lost. NVLink provides 1.8 TB/s of bidirectional bandwidth per GPU in Blackwell and 3.6 TB/s in Vera Rubin; Quantum InfiniBand ports reach 800 Gb/s. A workload spread across distant topology domains crosses shared links, contends, fragments, and — the part vendors rarely say loudly — leaves GPUs consuming provisioned power while waiting on data. In a power-limited AI factory, that is not an efficiency footnote; it is the difference between a node earning its power draw and a node burning it.
The blocker is not that schedulers cannot do topology-aware placement. Slurm and Kubernetes both support it. The blocker is that a scheduler can only act on the topology it observes — and the current topology is exactly what most operators do not have. Manual cluster drawings go stale the moment a rack changes, a switch fails, or a maintenance window re-cables a domain. Topograph's core move is to treat that data as a live pipeline: discover, normalize, publish, and regenerate when the cluster changes.

Providers and engines: the pipeline, named
Topograph's two concepts are worth understanding as architecture, because they map to real operational roles. Providers discover topology from cloud APIs or on-premises systems and normalize it into a canonical model; engines translate that model into the format each workload manager expects. One discovery layer, many outputs — Kubernetes node labels, Slurm topology.conf, Slinky ConfigMaps, NFD resources, or instance-oriented JSON.
The provider list reads like the current AI-cloud map: Google Cloud, Lambda, Nebius, Nscale, OCI, Crusoe on the managed side; InfiniBand through ibnetdiscover and NetQ for Spectrum-X or Multi-Node NVLink domains on-premises. The provider interface is open, which matters for private clusters — the operators who need this most are exactly the ones with a proprietary fabric vendor no public provider covers.

Stale topology is the silent tax
The operational claim worth testing is regeneration. Topograph's five components — API Server, Node Observer, Node Data Broker, Provider, Engine — keep the view current: the Node Observer watches configured Kubernetes node or pod changes and requests regeneration with retries; duplicate requests are deduplicated behind a typical 15-second aggregation delay; the API exposes /v1/generate, /v1/topology, /v1/lookup, /healthz, and /metrics. In practice that means a failed switch or a node drain updates the labels schedulers read, instead of waiting for someone to remember to redraw the diagram.
The honest caveat is scope. The matrix in the post shows exactly which engine works for which provider — InfiniBand on bare metal or VMs, for example, publishes to Slinky and NFD but not to Kubernetes node labels; MNNVL NVLink partitions only reach the Slinky engine as DRA block topology. This is a young toolkit with real sharp edges, and the "topology-aware" promise is only as good as the provider-engine pairing you actually deploy.

What an operator should do first
Start on Kubernetes, because that is where the feedback loop is easiest to verify. Deploy via Helm, pick a provider that matches where your GPUs live, and confirm the labels appear: fabric.topograph.run/tier-0..N and accelerator.topograph.run/domain. Those labels are immediately consumable by topology-aware schedulers like KAI Scheduler or Kueue's TAS for gang scheduling — and they are the same labels your observability stack can use to explain why a job landed where it did.
Then use the simulation utilities before touching production hardware: kwok-nodes and the Kind/KWOK helpers model node and switch hierarchies with no GPUs attached, which is the right way to validate provider behavior and engine output without a pilot cluster. When you do go live, treat the topology view as monitored infrastructure: hook /metrics into Prometheus, alert on stalled regeneration, and make "scheduler sees current topology" a documented production check rather than an assumption. Topology-aware placement only helps teams that keep the topology honest — and that is precisely what Topograph is for. If you already route local inference with NVIDIA PAIR, Topograph is the placement layer that decides which nodes PAIR has to choose from in the first place. For a practical pass over AI-factory scheduling, the Netics homepage is where we start.
Sources
- Topology-Aware Workload Scheduling with NVIDIA Topograph — NVIDIA Technical Blog, September 22, 2026.
- Topograph — dsx-ai-factory on GitHub, accessed September 23, 2026.
Source: "Topology-Aware Workload Scheduling with NVIDIA Topograph" — developer.nvidia.com, September 22, 2026.