The Controller That Runs Your AI Data Center Had Hard-Coded Credentials
NVIDIA's September 22 bulletin disclosed a CVSS 9.8 hard-coded-credential flaw in NICo, the Kubernetes-based controller that provisions bare-metal AI infrastructure. The tool that manages ev
TL;DR
- On September 22, 2026, NVIDIA disclosed CVE-2026-65113, a hard-coded-credentials flaw (CWE-798) in the NVIDIA Infrastructure Controller (NICo) for Linux, rated 9.8 out of 10 on the CVSS v3.1 scale and present in every version from 0 through 1.9.
- NICo is not an edge utility: it is the Kubernetes-based control plane that automates bare-metal lifecycle management for AI infrastructure — hardware discovery, firmware validation, DPU provisioning, OS deployment, network isolation, and tenant sanitization, across rack-scale systems including GB200 and GB300.
- The advisory is a 14-CVE batch: one critical (the hard-coded credentials), five high (including SQL injection at 8.8, missing authentication at 8.3, improper authentication at 8.2, OS command injection at 8.0), and eight medium. The pattern across all fourteen is a management plane that accumulated authentication and trust-validation weaknesses; no in-the-wild exploitation was confirmed at disclosure, and the exploit-prediction score was low.
- The fix is NICo version 2.0 from the official repository. The operational response goes beyond upgrading: isolate the management plane from untrusted networks and audit any hard-coded or default credentials before you trust the controller again.
Why this CVE matters more than its CVSS number
A hard-coded credential in a random web app is a bad day. A hard-coded credential in the controller that provisions every server, DPU, and tenant network in an AI data center is a different class of problem, because of what the controller is allowed to do.
NICo's own documentation describes it as delivering "zero-touch lifecycle automation for bare-metal systems that secures datacenter infrastructure at its foundation." The architecture is a Kubernetes-based control plane with a REST API, site agents, and Temporal workflows, deployed on the bare-metal cluster it manages. It handles hardware discovery, firmware validation, DPU provisioning through the DOCA Platform Framework, OS deployment, network isolation, and tenant sanitization. It is the single authority that prepares a bare-metal machine to become part of a multi-tenant AI cloud.

That centrality is exactly what makes CVE-2026-65113 dangerous. The flaw allows an unauthenticated remote attacker to bypass authentication using the static credentials; NVIDIA's advisory describes a successful exploit as potentially leading to escalation of privileges, data tampering, denial of service, and information disclosure. Applied to NICo, that is not one server: it is the plane from which an attacker can reach every tenant's bare metal, every DPU, and the network isolation the controller is supposed to enforce. NVIDIA's own security guidance in the same disclosure cycle is blunt about the class: restrict access to administrative interfaces and isolate management-plane services from untrusted networks.
Fourteen problems, one pattern
The hard-coded credentials are the headline, but they are not the whole story. CVE-2026-65128 is a SQL injection rated 8.8 that low-privilege attackers can use to reach code execution. CVE-2026-65114 is missing authentication on a critical function, rated 8.3. CVE-2026-65121 is improper authentication enabling privilege escalation, at 8.2. CVE-2026-65130 is OS command injection at 8.0. Add improper certificate validation (7.5 and 6.7), external file/path control, XML injection, workflow-enforcement failures, a separate hard-coded password, uncontrolled resource consumption, and uncleared debug information, and the picture is consistent: a management plane that accumulated authentication and trust-validation weaknesses across its release history, disclosed together because they were only reviewed together.

The pattern is worth naming: as AI infrastructure consolidated into GPU clouds and large on-prem clusters, the management plane became the crown jewel, and its security did not scale with its privileges. Cloud partners and GPU-as-a-service providers run NICo precisely because it automates provisioning of DPU-equipped racks; that automation is valuable, and it is valuable to an attacker in exactly the same proportion. NVIDIA's documentation lists Vault and external-secrets as part of the reference deployment, which makes the hard-coded credential all the more conspicuous: the surrounding stack attempts to manage secrets properly while the controller itself ships with one baked in.
The operational response
The direct fix is mechanical: upgrade to NICo version 2.0, or clone the latest release from the project's official GitHub repository. That is the precondition, not the whole answer.

First, treat the management plane as what it is — the most privileged network in the estate. NICo, like other infrastructure controllers, should be reachable only from a dedicated management VLAN or jump-host path for operators, never from tenant networks or the internet. The DPUs it provisions are not controllable by tenant host operating systems, which means the controller is the critical point of failure for the hardware-level security model: if the controller goes, so does the isolation story. Second, audit for credentials before re-trusting the upgrade: scan for hard-coded passwords and default accounts across the controller configuration, and rotate anything that looks inherited. Third, if you are a multi-tenant operator, financial or otherwise — this advisory is aimed squarely at AI-cloud and GPU-as-a-service providers — write the management-plane isolation into your change plan, not just the version bump.

This is the same lesson the BlueField-4 scale-in analysis drew from a different direction: when control leaves the host and moves into infrastructure hardware, the control plane becomes the trust boundary, and every operator needs to know exactly what that plane can do. Here the plane can provision every tenant's hardware, and until this advisory it answered to a hard-coded credential. Upgrade, isolate, audit — in that order, and before the next batch of GPU capacity goes online.
Sources
Source: NVIDIA Security Bulletin 5879 — NVIDIA Infrastructure Controller, September 2026 (nvidia.custhelp.com a_id 5879, initial release 2026-09-22, listed on nvidia.com/en-us/product-security). Source: NICo repository — github.com/dsx-ai-factory/infra-controller (moved from the NVIDIA organization on 2026-09-04). Source: NVIDIA Infra Controller documentation — docs.nvidia.com/infra-controller. Source: NVIDIA's AI Data Center Controller Ships with Hard-Coded Credentials — forkast.news, 2026-09-24. Source: NVIDIA Infrastructure Controller Hit by 14 Flaws — gbhackers.com, 2026-09-23. Internal linkage: BlueField-4 scale-in control plane. More infrastructure-security analysis on neticslabs.com.
Source: NVIDIA Security Bulletin 5879 — 2026-09-22, with NICo repository and documentation pages captured 2026-09-30.