Sources for every figure this video states. Checked 21 September 2026.
| Fact | Value | Source |
|---|---|---|
| Unified system memory | 128 GB LPDDR5x, 256 bit interface at 4266 MHz | NVIDIA DGX Spark User Guide, Hardware Overview |
| Memory bandwidth | 273 GB/s | NVIDIA DGX Spark User Guide, Hardware Overview |
| CPU | 20 core Arm (10 Cortex X925 plus 10 Cortex A725) | NVIDIA DGX Spark User Guide, Hardware Overview |
| GB10 thermal design power | 140 W, with a 240 W external supply for the system | NVIDIA DGX Spark User Guide, Hardware Overview |
| Size | 150 mm x 150 mm x 50.5 mm, about 5.9 by 5.9 by 2.0 inches, 1.2 kg | NVIDIA DGX Spark User Guide, Hardware Overview |
| US list price | $4,699 Founders Edition, raised from $3,999 on 27 February 2026 | VideoCardz, NVIDIA officially raises DGX Spark Founders Edition MSRP to $4,699 |
| Reason for the rise | Global memory supply constraints; the configuration did not change | VideoCardz, Overclock3D |
The memory is unified and shared between CPU, GPU and operating system rather than dedicated graphics memory. NVLink C2C carries the CPU to GPU link.
| Fact | Value | Source |
|---|---|---|
| Dedicated memory | 32 GB GDDR7, 512 bit interface | NVIDIA GeForce RTX 5090 product page |
| Memory bandwidth | 1,792 GB/s | NVIDIA GeForce RTX 5090 product page |
| Total graphics power | 575 W | NVIDIA GeForce RTX 5090 product page |
| Launch price | $1,999 for the card, January 2025 | NVIDIA GeForce RTX 5090 product page |
| Founders Edition size | 304 mm long, 137 mm wide, 48 mm high, two slot | Tom’s Hardware, RTX 5090 Founders Edition review |
LMSYS ran both machines on the same model, GPT OSS 20B in its MXFP4 build, through Ollama.
| Machine | Prefill, input tokens/s | Decode, output tokens/s |
|---|---|---|
| RTX 5090 | 8,519 | 205 |
| DGX Spark | 2,053 | 49.7 |
Source: LMSYS, NVIDIA DGX Spark In-Depth Review: A New Standard for Local AI Inference, Jerry Zhou and Richard Chen, 13 October 2025.
A thousand output tokens at 205 tokens per second is about 4.9 seconds; at 49.7 it is about 20.1 seconds. Both exclude prompt processing.
The llama.cpp project keeps a running DGX Spark benchmark thread, and the numbers move with the build. The February 2026 results are the ones this video uses.
| Model | Size | Generation at depth 0 | Generation at depth 32,768 |
|---|---|---|---|
| GPT OSS 20B MXFP4 | 11.27 GiB | 83.43 tokens/s | — |
| GPT OSS 120B MXFP4 | 59.02 GiB | 58.72 tokens/s | 42.76 tokens/s |
Source: ggml-org/llama.cpp Discussion 16578, Performance of llama.cpp on NVIDIA DGX Spark. Earlier runs in the same thread are substantially lower for the same hardware: the 14 October 2025 entry has GPT OSS 120B at 35.31 tokens/s at depth 0, and the 31 October 2025 entry has 52.87. That spread is software, not hardware.
The generation tests are tg32, which is a 32 token sample. They are not a measurement of
a sustained coding session.
59.02 GiB is 63.4 GB, which is larger than the RTX 5090’s entire 32 GB of VRAM before any conversation state is added.
| Fact | Value | Source |
|---|---|---|
| Parameters | 27 billion, dense | Qwen/Qwen3.8-27B model card |
| Modality | Native vision language model, understands images and video | Qwen/Qwen3.8-27B model card |
| Context | 262,144 native, extensible toward 1M with YaRN | Qwen/Qwen3.8-27B model card |
| Licence | Apache 2.0 | Qwen/Qwen3.8-27B model card |
| SWE bench Pro | 61.7, against 53.5 for Qwen3.6 27B | Qwen/Qwen3.8-27B model card |
| Four bit download | Q4_K_M is 19.0 GB | ggml-org/Qwen3.8-27B-GGUF |
61.7 against 53.5 is a gain of 8.2 points at the same headline parameter count. Both are Alibaba’s own model card numbers rather than independent replications.
The ggml-org four bit pack also ships the separate multi token prediction head used for speculative decoding, which is why it is larger than some other four bit packs of the same model.
| Model | Parameters | SWE bench Verified | Licence |
|---|---|---|---|
| Devstral Small 2 | 24 billion | 68.0% | Apache 2.0 |
| Devstral 2 | 123 billion | 72.2% | Modified MIT |
Source: Mistral AI, Introducing: Devstral 2 and Mistral Vibe CLI, 9 December 2025, and mistralai/Devstral-Small-2-24B-Instruct-2512.
Both carry a 256K context window. Mistral’s own figures.
SWE bench Verified and SWE bench Pro are different benchmarks. 68.0 on Verified and 61.7 on Pro are not comparable numbers, which is why the script says comparing them directly would be nonsense.
| Fact | Value | Source |
|---|---|---|
| Total parameters | 117 billion | OpenAI, Introducing gpt-oss |
| Active parameters per token | 5.1 billion | gpt-oss model card, arXiv 2508.10925 |
| Architecture | Mixture of experts | OpenAI, Introducing gpt-oss |
| Licence | Apache 2.0 | OpenAI, Introducing gpt-oss |
A mixture of experts routes each token to a subset of the experts, so the active count is far below the total. The inactive experts still occupy memory.
| Machine | Memory | Bandwidth | Source |
|---|---|---|---|
| AMD Ryzen AI Max+ 395, Strix Halo | up to 128 GB LPDDR5x 8000, 256 bit | 256 GB/s | AMD developer article |
| Apple M3 Ultra Mac Studio | 96 GB to 512 GB unified | 819 GB/s | Apple Newsroom, Apple reveals M3 Ultra |
The AMD part runs Linux or Windows. Measured sustained bandwidth on Strix Halo is reported at roughly 215 GB/s, about 84% of the 256 GB/s specification.
Apple silicon runs MLX and Metal rather than CUDA, so model support and optimised implementations differ from the NVIDIA path.
DGX Spark runs Linux on an Arm CPU, and NVIDIA publishes a porting guide because an ordinary x86 desktop software package is not automatically compatible. Its supported containers are the documented route to running a separate AI service.
The KV cache stores the attention keys and values for the conversation so far, so its cost grows with the conversation rather than with the model file. Its size depends on the model architecture, the context length, the precision the cache is held at, and the runner. An advertised context window is a property of the model, not a promise that the cache for it fits in a particular GPU’s memory.
These are stated in the narration and could not be chased to a primary source.