DGX Spark at 64GB: What $4,999 of Local Compute Buys a Small Team
NVIDIA's 64GB DGX Spark configuration lands on 23 October from $4,999, runs models up to 100 billion parameters on the desk, and clusters two units into 128GB of pooled memory. The interesti
TL;DR
- NVIDIA has put a 64GB configuration of DGX Spark on sale from Acer, ASUS, Dell, Gigabyte, HP and MSI from 23 October, starting at $4,999, with the same GB10 Grace Blackwell superchip, DGX OS and AI software stack as the 128GB model.
- One unit runs models up to 100 billion parameters entirely on device; two units connect over a QSFP cable into a 200 GbE fabric that pools 128GB and extends support to 200 billion parameters.
- The vendor's own measurement sets the expectation: in a Qwen 3.8 27B test, two clustered 64GB systems delivered up to 1.7x the performance of one, and the software environment stays identical when a second node appears.
- For a small team, the decision is where the box fits in a working setup: workloads steady enough to justify fixed hardware go local, and bursty ones stay cheaper to rent.
- Before the order goes in, check the constraints nobody benchmarks: memory ceiling per model, one physical site, and somebody who owns driver, model and firmware updates.
What the 64GB configuration ships with

NVIDIA's announcement is specific about what changes and what stays the same. The new configuration keeps the GB10 Grace Blackwell superchip, keeps DGX OS and the full AI software stack, and takes the memory down to 64GB of unified memory for a $4,999 starting price. The vendor positions it for developers, researchers and enthusiasts who want capable agents running on the device, privately and without a cloud dependency, and states support for models up to 100 billion parameters fully on device.
The software list is the part that decides whether the box gets used. The platform ships with the NVIDIA Agent Toolkit, CUDA-X AI libraries and Nemotron open models, and supports the runtimes people already run: Ollama, vLLM and PyTorch among them. Blender is named as an early creator-application supporter with an installer coming. That mix tells you what the box is for. It is a development and always-on inference machine, sized for the models that fit in 64GB of unified memory, rather than a substitute for a rack.
Two units, 128GB and a 200 GbE link
The cluster story is the more interesting half of the announcement, because it changes which models are in reach. Every unit carries a ConnectX-7 network interface, and two units connect directly with a QSFP cable rather than through a switch. That connection pools memory to 128GB, doubles the memory bandwidth and extends model support to 200 billion parameters. NVIDIA's own test on a Qwen 3.8 27B workload measured up to 1.7x the performance of a single system, and the company describes the scaling as seamless on the software side: the NVIDIA Sync Cluster Assistant detects the connected units, validates the configuration and configures the ConnectX-7 network, so both nodes run the same stack.

Read that as a budget decision rather than a performance claim. The ramp from one unit to two is what lets a team start at $4,999 and buy the second unit when a workload outgrows the first, with the memory ceiling as the trigger. It is also the point where the desk-side box stops being a personal device and starts being infrastructure: two units on one site are still one site.
The cost question a small team should actually ask
Local inference is usually sold as a saving, and the honest version of that argument is narrower. Fixed hardware beats rented tokens when a workload runs steadily enough that the hardware stays busy, on data that benefits from never leaving the building, at a model size that fits in the memory you bought. It loses to renting when the workload is spiky, when the model needs to be the largest one available next quarter, or when the team has nobody to own the machine.
Take a concrete case. Imagine a twelve-person engineering firm in Lyon that extracts structured data from client contracts and technical specifications. Its peak is forty documents a day, the data is contractually sensitive, and the pipeline is stable enough that the same model has run unchanged for months. For that team, a device that answers locally and keeps the documents on the desk is easier to put in front of a client than an API call, and the cost of the hardware is a line item rather than a monthly surprise. The same box in a research team that swaps models weekly is a shelf warmer.
The pattern we build for clients who want inference on hardware they control looks at the memory ceiling first and the benchmark headline last, because the ceiling decides which models can be served at all. The other half of that work is unglamorous and it does not ship in a box: a model registry, an update cadence for firmware and runtimes, and a monitoring story that tells you when a node is answering slowly.

Where the box fits in a working setup
Three uses show up in the announcement and each maps to a real team situation. The first is the always-on agent: a coding or research agent that reviews code, reads documents or runs multi-step tasks around the clock, with a cluster when several agents need to run at once. The second is a model server for a small office, where the box answers requests over the local network and the laptops stay free. The third is the development tier, where a team validates an inference setup on hardware that costs less than a month of a rented accelerator, then keeps the same container images for the production deployment.
The end-of-month NVIDIA Sync Model Launcher points at the same idea from the vendor side: download and start Qwen3.8 27B on one unit or a cluster, and have the tooling configure OpenCode to use the model. Convenience features like that are what turn a workstation into shared infrastructure for a small team, and they are also the features to read carefully, because a default that starts a model is a default that starts a process somebody has to own.
What to check before the order goes in
Five checks cover most of the disappointment. Confirm the largest model you actually need fits the ceiling you are buying, and remember that fine-tuning sits in a different size class from inference. Confirm the workload runs at least a few hours a day, because idle hardware is the most expensive token in the building. Confirm where the data has to live, since a box on a desk answers a residency question that a managed service cannot. Confirm who owns the updates, because a local stack means your team tracks drivers, runtimes and model versions. Then confirm the second unit's cost and the cable that joins them, since the cluster path is the reason the 64GB configuration is interesting in the first place.

For teams whose cloud bill is dominated by steady inference on sensitive data, the sovereign infrastructure and cloud exit work is where this decision lands in practice, and our earlier analysis of local inference routing covers the layer that comes next: once two machines answer requests, something has to decide which one takes each job.
Sources
Source: NVIDIA DGX Spark 64GB Gives Developers More Ways to Build and Scale Local AI — blogs.nvidia.com/blog/local-ai-dgx-spark-64gb-sync/, published 2026-10-02, retrieved 2026-10-04 (64GB configuration available from Acer, ASUS, Dell, Gigabyte, HP and MSI from Friday 23 October; starting at $4,999; GB10 Grace Blackwell superchip, DGX OS and full NVIDIA AI software stack retained; support for models up to 100 billion parameters on device; two units connected over a QSFP cable pooling memory to 128GB and extending support to 200 billion parameters with twice the memory bandwidth; up to 1.7x performance for two clustered 64GB systems in NVIDIA's Qwen 3.8 27B test; NVIDIA Sync Cluster Assistant detecting units and configuring the ConnectX-7 network; NVIDIA Agent Toolkit, CUDA-X AI libraries, Nemotron open models and support for Ollama, vLLM and PyTorch; NVIDIA Sync Model Launcher due at the end of the month with Qwen3.8 27B and OpenCode configuration; Blender installer coming; three developer workflow examples; agentic AI playbooks on build.nvidia.com). Source: DGX Spark product page — nvidia.com/en-us/products/workstations/dgx-spark/, retrieved 2026-10-04 (GB10 Grace Blackwell superchip; up to 1 petaFLOP of AI performance at FP4 precision; 64GB or 128GB unified system memory configurations; ConnectX networking enabling up to four DGX Spark systems to work together; 273 GB/s memory bandwidth in the published specification table). Internal linkage: Local AI Needs a Router Before It Needs a Bigger Cluster. More on Netics' work at sovereign infrastructure and cloud exit.
Source: NVIDIA blog — blogs.nvidia.com, published 2026-10-02, retrieved 2026-10-04; NVIDIA DGX Spark product page — nvidia.com, retrieved 2026-10-04. Figures: official NVIDIA announcement and product images, retrieved 2026-10-04.