Coding Horizon

128GB vs 192GB: The Local AI Upgrade You Might Regret

Checked 11 September 2026. Prices are US dollars before tax. File sizes are the decimal gigabytes the hosting service reports for the download, not the memory a running model needs; every fit shown on screen leaves that overhead as a separate, labelled block.

The machines

On screen Source
DGX Spark: GB10 Grace Blackwell, 128 GB unified LPDDR5x, 4 TB self encrypting NVMe, 273 GB/s Nvidia, DGX Spark specifications
Spark US list price raised from $3,999 to $4,699 with no hardware change; Nvidia cites memory supply constraints Nvidia developer forum, price change announcement, 23 Feb 2026
Dell Pro Max with GB10, built on the same chip Jeff Geerling, ai-benchmarks lists the Dell Pro Max with GB10 beside the Framework and Apple results
Ryzen AI Max+ 395: 16 Zen 5 cores, 32 threads, Radeon 8060S with 40 RDNA 3.5 compute units, XDNA 2 NPU AMD, Ryzen AI Max+ 395 blog; cpu-monkey spec listing
Framework Desktop, 128 GB Ryzen AI Max+ 395: $3,449 before storage and other options; memory listed as non upgradeable; storage priced separately from +$135 Framework configurator
Framework 395 theoretical bandwidth 256 GB/s LPDDR5X-8000 on a 256 bit bus: 8000 MT/s × 256 bits ÷ 8 = 256 GB/s. Framework Desktop specifications
Beelink GTR9 Pro: Ryzen AI Max+ 395, Radeon 8060S, 128 GB LPDDR5X-8000 Beelink, GTR9 Pro
Minisforum MS S1 Max: Ryzen AI Max+ 395, up to 128 GB unified memory Minisforum, MS-S1 Max
BOSGAME M5: Ryzen AI Max+ 395, 96 GB and 128 GB variants BOSGAME, M5
GMKtec EVO X2: $2,199.99 sale price applies to the 64 GB / 1 TB configuration; 128 GB configurations exist without a distinct price on the page GMKtec, EVO X2 product page
Mac Studio M2 Ultra: up to 192 GB unified memory, 800 GB/s, 60 or 76 GPU cores, 24 CPU cores, announced 5 June 2023 Apple Newsroom, M2 Ultra
Reseller listing: Mac Studio M2 Ultra, 76 core GPU, 192 GB, 1 TB, $5,095.02, marked currently unavailable iPowerResale listing. One indexed listing, not an in stock offer or a market price
AMD Variable Graphics Memory allows up to 96 GB for graphics on a 128 GB system AMD, Ryzen AI Max+ 395 blog; AMD, VGM FAQ. A driver setting, not a hardware ceiling; Linux runtimes draw the line differently
Framework 192 GB desktop, coming: Ryzen AI Max+ PRO 495, 40 CU Radeon 8065S, 192 GB LPDDR5X, 273 GB/s, no price, no measured model speed Framework, 192 GB coming soon. 192 ÷ 128 = 1.5; 273 ÷ 256 = 1.066
H100: 80 GB OpenAI, gpt-oss-120b model card names “a single 80GB GPU (like NVIDIA H100 or AMD MI300X)”

The models

On screen Source
Qwen3 Coder 30B A3B Instruct, Q4_K_M: 18.56 GB Unsloth, Qwen3-Coder-30B-A3B-Instruct-GGUF, file size from the repository tree
Llama 70B class at four bit: about 40 to 45 GB Ollama, llama3.1:70b lists the Q4 tag at about 43 GB; drawn as a range
GPT OSS 120B: 117B total, 5.1B active, MoE weights post trained in MXFP4, runs on a single 80 GB GPU; text only OpenAI, gpt-oss-120b model card; OpenAI developer docs
Qwen3.5 122B A10B: 122B total, 10B active, vision language model (image, text, video input) Qwen, Qwen3.5-122B-A10B model card
Qwen3.5 122B A10B, bartowski Q4_K_M: 77.62 GB (shown as 78); Q8_0: 132.60 GB (shown as 133) bartowski, Qwen_Qwen3.5-122B-A10B-GGUF, shard sizes summed from the repository tree
Qwen3 235B A22B, official Q4_K_M: 142.15 GB (shown as 142) Qwen, Qwen3-235B-A22B-GGUF, five shards summed
DeepSeek V4 Flash: 284B total, 13B active per token DeepSeek, DeepSeek-V4-Flash model card
Unsloth DeepSeek V4 Flash GGUF: UD-IQ3_XXS 103.0 GB, UD-Q4_K_XL 155.1 GB Unsloth, DeepSeek-V4-Flash-GGUF, folder sizes summed from the repository tree

Speed and bandwidth

On screen Source
Llama 3.1 70B generation: Framework Desktop mainboard (395+) 4.97 tokens/s (GPU/CPU); Dell Pro Max GB10 4.71 tokens/s (GPU); M1 Ultra 64 GPU core 128 GB 9.84 tokens/s (GPU) Jeff Geerling, ai-benchmarks README. Collected submissions with different software versions and settings, not one controlled run; the M2 Ultra is not in the table
1,000 tokens at 5 tokens/s is 200 s; at 10 tokens/s is 100 s Arithmetic. Excludes loading, prompt processing and any reasoning tokens
Bandwidth: M2 Ultra 800 GB/s, DGX Spark 273 GB/s, Framework 395 256 GB/s theoretical, Framework 495 273 GB/s Apple, Nvidia and Framework pages above
800 ÷ 273 ≈ 2.9, shown only to be struck out Arithmetic; bandwidth ratios are not measured speedups

Benchmarks

On screen Source
SWE-bench Verified: Qwen3.5 122B A10B 72.0, Qwen3.5 27B 72.4. Terminal Bench 2: 49.4 against 41.6 Qwen, Qwen3.5-122B-A10B model card, benchmark table. Developer reported evaluations of the full precision models

Illustrative, not measured

These frames draw an idea rather than a published number, and print no figure for it:

Not checked