Coding Horizon

9 Local AI models ranked: the biggest wastes your VRAM

Every figure, size, score and capability this video states, chased to a primary source. Sizes are the ones the distributor’s own page shows rather than calculated from a parameter count, because the download is what is being compared.

Two kinds of number appear in this video and they are not sourced the same way. A published figure is somebody’s own number and is cited here. A planning range is the script’s own estimate of working memory, and the script says so itself; those are not measurements and are not presented as any.

The models, in ranking order

Llama 3.1 8B instruct — B tier

Gemma 4 26B A4B instruction tuned — A tier

The script says “roughly 26 billion in total, but about four billion active”. The published figures are 25.2B and 3.8B.

DeepSeek R1 distill qwen 7B — C tier

Ministral 3 8B instruct 2512 — A tier

Llama 3.3 70B instruct — D tier

Phi 4 reasoning vision 15B — B tier

Mistral 7B instruct version 0.3 — F tier

Devstral small 2 24B instruct 2512 — A tier

Qwen 3.8 27B — S tier

The local runner

The worked example

The broken checkout is the script’s own worked example and is nobody’s product. Its arithmetic holds both ways round: a $100 cart with a $10 coupon and 10% tax is $99 when the coupon is applied before tax (90 x 1.10) and $100 when it is applied after (110 less 10).

Not chased to a source