Every figure this video puts on screen, chased to a primary source. Where a claim in the narration could not be re-found at a primary source, the shot does not render the figure, and the claim is listed at the bottom.
NVIDIA’s own hardware specification for DGX Spark gives the whole of the teardown chapter.
Source: NVIDIA, DGX Spark User Guide, Hardware Overview. https://docs.nvidia.com/dgx/dgx-spark/hardware.html
Up to 1 petaFLOP of FP4 AI compute, and the three model size claims, come from NVIDIA’s own product page:
The four node configuration quadruples available memory to 512 GB.
Source: NVIDIA, DGX Spark product page. https://www.nvidia.com/en-us/products/workstations/dgx-spark/
Two systems connected over a single 200 Gb/s QSFP link reach models up to 405 billion parameters; three systems can be cabled directly, four or more need a switch.
Source: NVIDIA Sync User Guide, Cluster Assistant for a multi node DGX Spark cluster. https://docs.nvidia.com/sync/latest/cluster-assistant.html
The software named on screen, DGX OS, CUDA, NIM, NemoClaw, OpenShell, playbooks and an NVIDIA AI Enterprise trial, is the list NVIDIA gives on the product page above.
Decode throughput, Qwen3 family, same harness and same machine. This is the report the dense against mixture of experts comparison in chapter three is drawn from.
| Model | 512 context | 2048 context |
|---|---|---|
| Qwen3 1.7B | 161.4 tok/s | 146.1 tok/s |
| Qwen3 30B A3B (MoE) | 89.3 tok/s | 83.8 tok/s |
| Qwen3 32B (dense) | 10.7 tok/s | 10.5 tok/s |
The report’s own explanation of the gap: the 30B MoE activates roughly 2.4B parameters per token and so avoids the memory bandwidth bottleneck that limits the dense 32B.
Source: DandinPower, llama.cpp_bench, DGX Spark report. https://github.com/DandinPower/llama.cpp_bench/blob/main/dgx_spark/report.md
Context: the wider llama.cpp DGX Spark performance discussion. https://github.com/ggml-org/llama.cpp/discussions/16578
Throughput in tokens per second at 4k context, on one GB10.
| Model and recipe | tok/s at 4k |
|---|---|
| nvidia/qwen3.6-35b-a3b NVFP4 MTP | 86.3 |
| nvidia/qwen3-30b-a3b NVFP4 | 74.2 |
| ornith-ai/ornith-1.5-35b-a3b NVFP4 MTP | 70.8 |
| saricles/qwen3-coder-next NVFP4 | 58.9 |
These are different recipes, engines and model formats, so they are not directly comparable with each other, which is what the shot that separates them into lanes shows.
Source: SparkBench. https://sparkbench.dev/
Generated tokens per second. This is the “what runs against what feels good” chapter.
| Model | Gen tok/s |
|---|---|
| Llama 3.1 8B | 42.86 |
| Gemma 3 27B | 11.71 |
| Qwen2.5 Coder 32B | 10.36 |
| Qwen3 32B | 9.88 |
| CodeLlama 70B | 5.73 |
| Nemotron 70B | 4.77 |
| Llama 3.1 70B | 4.76 |
| Qwen 2.5 72B | 4.40 |
| Mistral Large 123B | 2.28 |
Source: G3nadh/dgx-spark-benchmarks dataset. https://huggingface.co/datasets/G3nadh/dgx-spark-benchmarks
The model is real and its architecture is confirmed: 176B total, made of a 125B main body plus a 51B n gram embedding table, activating 6B parameters per token, MoE, with a 262,144 token native context.
Source: vLLM recipes, Qwen/Qwen3.8-Flash-Next. https://recipes.vllm.ai/Qwen/Qwen3.8-Flash-Next
The caveats the narration lists are all confirmed by the published DGX Spark recipes: it needs a patched vLLM, the checkpoint is about 122 GiB on disk and needs the n gram table served from NVMe by mmap to fit in 128 GB, and cold start is 8 to 13 minutes.
Source: blazux/qwen3.8-Flash-DGX. https://github.com/blazux/qwen3.8-Flash-DGX Source: madeye, Qwen3.8-Flash-Next on a single NVIDIA DGX Spark. https://madeye.github.io/qwen38-flash-next-on-dgx-spark/
The specific throughput pair the narration gives could not be re-found. Published single stream figures for this model on one DGX Spark range widely by recipe: 21.6 tok/s hybrid quantisation, about 26 to 31 tok/s NVFP4 with MTP, 24.8 tok/s on one vLLM setup and 42.7 tok/s on another measured with the same script. So the shots for those beats draw the architecture, the warm state and the concurrency scaling, and print no throughput figure.
The model and the two node setup are confirmed: 284B total, 13B active, MoE, run across two GB10 nodes with tensor parallel size 2 because a single 128 GB pool is below the checkpoint footprint.
Source: vLLM recipes, deepseek-ai/DeepSeek-V4-Flash. https://recipes.vllm.ai/deepseek-ai/DeepSeek-V4-Flash
That aggregate throughput rises sharply with concurrency is confirmed: 210.8 tok/s combined at six concurrent users with vLLM, and 261 tok/s aggregate at 16 concurrent on a DSpark speculative decoding setup.
Source: DevelopersIO, “Tried running DeepSeek V4 Flash-0731 at 284B on two DGX Spark units”. https://dev.classmethod.jp/en/articles/dgx-spark-2node-deepseek-v4-flash-0731/
The single stream figure the narration gives could not be re-found, and the published vLLM measurements for this configuration are higher, at 61.4 to 76 tok/s. The shot draws the two node split and the concurrency scaling and prints no single stream figure.
GeForce RTX 5090: 32 GB GDDR7 on a 512 bit interface. Source: NVIDIA. https://www.nvidia.com/en-us/geforce/graphics-cards/50-series/rtx-5090/
Mac Studio: configurable to 512 GB unified memory with 1.2 TB/s of memory bandwidth on the M5 Ultra, and 128 GB at 614 GB/s on the M5 Max. Both are well above DGX Spark’s 273 GB/s, which is what the narration says. Source: Apple. https://www.apple.com/mac-studio/specs/
AMD Ryzen AI Halo developer platform, Ryzen AI Max+ 395: up to 128 GB unified system memory, enough headroom to run up to 200 billion parameter models locally, Windows and Linux versions, an NPU, and a published tokens per dollar comparison against DGX Spark at retail prices of $3,999 for the AMD platform against $4,699 for DGX Spark. Source: AMD, “AMD Powers Next-Generation Agent Computers with New Ryzen AI Halo Developer Platform”. https://www.amd.com/en/blogs/2026/amd-powers-next-generation-agent-computers-with-new-ryzen-ai-hal.html