Skip to main content

What can it run

Popular open-weight models, and the cheapest box we track that has room for each one. Memory capacity decides whether a model loads at all, so this is the first question worth answering. Prices last checked .

The requirement is the weights at that quantization plus a 20 percent allowance for the KV cache, activations, and the runtime. It is arithmetic on the format, not a measurement. Long contexts and large batches push the real number up, sometimes a lot. Every parameter count links to the model card it came from.

Search

Family

Quantization

Hardware budget

25 of 25 models

Qwen3.5-0.8B

0.8B parameters

Model card lists 0.8B parameters and a native 262,144-token context window.

Quantization ?NeedsCheapest that fitsOptions
Q4 (4-bit)0.5 GBTenstorrent Blackhole p150a ($1,399)19
Q5 (5-bit)0.6 GBTenstorrent Blackhole p150a ($1,399)19
Q8 (8-bit)1.0 GBTenstorrent Blackhole p150a ($1,399)19
FP161.9 GBTenstorrent Blackhole p150a ($1,399)19

Parameter count from Qwen3.5-0.8B model card, fetched

LFM2.5-1.2B-Thinking

1.17B parameters

Model card lists 1.17B parameters and a 32,768-token context window.

Quantization ?NeedsCheapest that fitsOptions
Q4 (4-bit)0.7 GBTenstorrent Blackhole p150a ($1,399)19
Q5 (5-bit)0.9 GBTenstorrent Blackhole p150a ($1,399)19
Q8 (8-bit)1.4 GBTenstorrent Blackhole p150a ($1,399)19
FP162.8 GBTenstorrent Blackhole p150a ($1,399)19

Parameter count from LFM2.5-1.2B-Thinking model card, fetched

Qwen3.5-2B

2B parameters

Model card lists 2B parameters and a native 262,144-token context window.

Quantization ?NeedsCheapest that fitsOptions
Q4 (4-bit)1.2 GBTenstorrent Blackhole p150a ($1,399)19
Q5 (5-bit)1.5 GBTenstorrent Blackhole p150a ($1,399)19
Q8 (8-bit)2.4 GBTenstorrent Blackhole p150a ($1,399)19
FP164.8 GBTenstorrent Blackhole p150a ($1,399)19

Parameter count from Qwen3.5-2B model card, fetched

SmolLM3-3B

3B parameters

Model card calls it a 3B parameter model; the current config supports 65,536 tokens and the card documents YaRN extensions for 128K and 256K.

Quantization ?NeedsCheapest that fitsOptions
Q4 (4-bit)1.8 GBTenstorrent Blackhole p150a ($1,399)19
Q5 (5-bit)2.3 GBTenstorrent Blackhole p150a ($1,399)19
Q8 (8-bit)3.6 GBTenstorrent Blackhole p150a ($1,399)19
FP167.2 GBTenstorrent Blackhole p150a ($1,399)19

Parameter count from SmolLM3-3B model card, fetched

Qwen3.5-4B

4B parameters

Model card lists 4B parameters and a native 262,144-token context window.

Quantization ?NeedsCheapest that fitsOptions
Q4 (4-bit)2.4 GBTenstorrent Blackhole p150a ($1,399)19
Q5 (5-bit)3.0 GBTenstorrent Blackhole p150a ($1,399)19
Q8 (8-bit)4.8 GBTenstorrent Blackhole p150a ($1,399)19
FP169.6 GBTenstorrent Blackhole p150a ($1,399)19

Parameter count from Qwen3.5-4B model card, fetched

Gemma 4 E4B

8B parameters

Google lists 4.5B effective parameters and 8B total with embeddings. Sizing uses all 8B resident weights, not the effective count.

Quantization ?NeedsCheapest that fitsOptions
Q4 (4-bit)4.8 GBTenstorrent Blackhole p150a ($1,399)19
Q5 (5-bit)6.0 GBTenstorrent Blackhole p150a ($1,399)19
Q8 (8-bit)9.6 GBTenstorrent Blackhole p150a ($1,399)19
FP1619.2 GBTenstorrent Blackhole p150a ($1,399)19

Parameter count from Gemma 4 E4B IT model card, fetched

Qwen3-8B

8.2B parameters

Model card: "Number of Parameters: 8.2B". Context 32,768 natively, 131,072 with YaRN.

Quantization ?NeedsCheapest that fitsOptions
Q4 (4-bit)4.9 GBTenstorrent Blackhole p150a ($1,399)19
Q5 (5-bit)6.1 GBTenstorrent Blackhole p150a ($1,399)19
Q8 (8-bit)9.8 GBTenstorrent Blackhole p150a ($1,399)19
FP1619.7 GBTenstorrent Blackhole p150a ($1,399)19

Parameter count from Qwen3-8B model card, fetched

Qwen3.5-9B

9B parameters

Model card lists 9B parameters and a native 262,144-token context window.

Quantization ?NeedsCheapest that fitsOptions
Q4 (4-bit)5.4 GBTenstorrent Blackhole p150a ($1,399)19
Q5 (5-bit)6.8 GBTenstorrent Blackhole p150a ($1,399)19
Q8 (8-bit)10.8 GBTenstorrent Blackhole p150a ($1,399)19
FP1621.6 GBTenstorrent Blackhole p150a ($1,399)19

Parameter count from Qwen3.5-9B model card, fetched

Gemma 4 12B

11.95B parameters

Google lists 11.95B total parameters and a 256K context window for the encoder-free multimodal model.

Quantization ?NeedsCheapest that fitsOptions
Q4 (4-bit)7.2 GBTenstorrent Blackhole p150a ($1,399)19
Q5 (5-bit)9.0 GBTenstorrent Blackhole p150a ($1,399)19
Q8 (8-bit)14.3 GBTenstorrent Blackhole p150a ($1,399)19
FP1628.7 GBTenstorrent Blackhole p150a ($1,399)17

Parameter count from Gemma 4 12B IT model card, fetched

gpt-oss-20b

21B parameters · 3.6B active per token

OpenAI: "21B parameters with 3.6B active parameters"; card says it can run within 16GB of memory.

Quantization ?NeedsCheapest that fitsOptions
Q4 (4-bit)12.6 GBTenstorrent Blackhole p150a ($1,399)19
Q5 (5-bit)15.8 GBTenstorrent Blackhole p150a ($1,399)19
Q8 (8-bit)25.2 GBTenstorrent Blackhole p150a ($1,399)17
FP1650.4 GBApple Mac mini (M4 Pro, 64GB) ($2,800)14

Parameter count from gpt-oss-120b model card, fetched

Mistral Small 3.2 24B

24B parameters

Model card lists 24B params. Mistral does not state a context window on the card.

Quantization ?NeedsCheapest that fitsOptions
Q4 (4-bit)14.4 GBTenstorrent Blackhole p150a ($1,399)19
Q5 (5-bit)18.0 GBTenstorrent Blackhole p150a ($1,399)19
Q8 (8-bit)28.8 GBTenstorrent Blackhole p150a ($1,399)17
FP1657.6 GBApple Mac mini (M4 Pro, 64GB) ($2,800)14

Parameter count from Mistral Small 3.2 24B Instruct model card, fetched

Gemma 4 26B-A4B

25.2B parameters · 3.8B active per token

Mixture of experts: Google lists 25.2B total parameters and 3.8B active, with a 256K context window.

Quantization ?NeedsCheapest that fitsOptions
Q4 (4-bit)15.1 GBTenstorrent Blackhole p150a ($1,399)19
Q5 (5-bit)18.9 GBTenstorrent Blackhole p150a ($1,399)19
Q8 (8-bit)30.2 GBTenstorrent Blackhole p150a ($1,399)17
FP1660.5 GBApple Mac mini (M4 Pro, 64GB) ($2,800)14

Parameter count from Gemma 4 26B-A4B IT model card, fetched

Gemma 3 27B

27B parameters

Model card lists 27B params, with a 128K input context window.

Quantization ?NeedsCheapest that fitsOptions
Q4 (4-bit)16.2 GBTenstorrent Blackhole p150a ($1,399)19
Q5 (5-bit)20.3 GBTenstorrent Blackhole p150a ($1,399)19
Q8 (8-bit)32.4 GBApple Mac mini (M4 Pro, 64GB) ($2,800)15
FP1664.8 GBFramework Desktop (Ryzen AI Max+ 395, 128GB) ($3,449)11

Parameter count from Gemma 3 27B IT model card, fetched

Qwen3.5-27B

27B parameters

Model card lists 27B parameters and a native 262,144-token context window.

Quantization ?NeedsCheapest that fitsOptions
Q4 (4-bit)16.2 GBTenstorrent Blackhole p150a ($1,399)19
Q5 (5-bit)20.3 GBTenstorrent Blackhole p150a ($1,399)19
Q8 (8-bit)32.4 GBApple Mac mini (M4 Pro, 64GB) ($2,800)15
FP1664.8 GBFramework Desktop (Ryzen AI Max+ 395, 128GB) ($3,449)11

Parameter count from Qwen3.5-27B model card, fetched

Qwen3.8-27B

27B parameters

Model card lists 27B parameters and a native 262,144-token context window, extensible to 1,000,000 tokens.

Quantization ?NeedsCheapest that fitsOptions
Q4 (4-bit)16.2 GBTenstorrent Blackhole p150a ($1,399)19
Q5 (5-bit)20.3 GBTenstorrent Blackhole p150a ($1,399)19
Q8 (8-bit)32.4 GBApple Mac mini (M4 Pro, 64GB) ($2,800)15
FP1664.8 GBFramework Desktop (Ryzen AI Max+ 395, 128GB) ($3,449)11

Parameter count from Qwen3.8-27B model card, fetched

Gemma 4 31B

30.7B parameters

Google lists 30.7B total parameters and a 256K context window for the dense model.

Quantization ?NeedsCheapest that fitsOptions
Q4 (4-bit)18.4 GBTenstorrent Blackhole p150a ($1,399)19
Q5 (5-bit)23.0 GBTenstorrent Blackhole p150a ($1,399)19
Q8 (8-bit)36.8 GBApple Mac mini (M4 Pro, 64GB) ($2,800)15
FP1673.7 GBFramework Desktop (Ryzen AI Max+ 395, 128GB) ($3,449)11

Parameter count from Gemma 4 31B IT model card, fetched

Qwen3.5-35B-A3B

35B parameters · 3B active per token

Mixture of experts: model card lists 35B total parameters and 3B active, with a native 262,144-token context window.

Quantization ?NeedsCheapest that fitsOptions
Q4 (4-bit)21.0 GBTenstorrent Blackhole p150a ($1,399)19
Q5 (5-bit)26.3 GBTenstorrent Blackhole p150a ($1,399)17
Q8 (8-bit)42.0 GBApple Mac mini (M4 Pro, 64GB) ($2,800)15
FP1684.0 GBFramework Desktop (Ryzen AI Max+ 395, 128GB) ($3,449)9

Parameter count from Qwen3.5-35B-A3B model card, fetched

Llama 3.3 70B Instruct

70B parameters

Card table lists 70B; the computed safetensors count reads 71B. We use the nominal 70B.

Parameter count from Llama-3.3-70B-Instruct model card, fetched

gpt-oss-120b

117B parameters · 5.1B active per token

OpenAI: "117B parameters with 5.1B active parameters"; designed to fit a single 80GB GPU using MXFP4.

Parameter count from gpt-oss-120b model card, fetched

Qwen3.5-122B-A10B

122B parameters · 10B active per token

Mixture of experts: model card lists 122B total parameters and 10B active, with a native 262,144-token context window.

Parameter count from Qwen3.5-122B-A10B model card, fetched

Qwen3-235B-A22B

235B parameters · 22B active per token

Mixture of experts: "235B in total and 22B activated". All weights must still be resident.

Quantization ?NeedsCheapest that fitsOptions
Q4 (4-bit)141.0 GBApple Mac Studio (M3 Ultra, 512GB) ($17,999)2
Q5 (5-bit)176.3 GBApple Mac Studio (M3 Ultra, 512GB) ($17,999)2
Q8 (8-bit)282.0 GBApple Mac Studio (M3 Ultra, 512GB) ($17,999)2
FP16564.0 GBNothing we track fits it0

Parameter count from Qwen3-235B-A22B model card, fetched

DeepSeek-V4-Flash-0731

284B parameters · 13B active per token

Mixture of experts: the DeepSeek-V4 report lists 284B total parameters and 13B active, with a one-million-token context window.

Quantization ?NeedsCheapest that fitsOptions
Q4 (4-bit)170.4 GBApple Mac Studio (M3 Ultra, 512GB) ($17,999)2
Q5 (5-bit)213.0 GBApple Mac Studio (M3 Ultra, 512GB) ($17,999)2
Q8 (8-bit)340.8 GBApple Mac Studio (M3 Ultra, 512GB) ($17,999)2
FP16681.6 GBNothing we track fits it0

Parameter count from DeepSeek-V4-Flash-0731 model card, fetched , DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence, fetched

DeepSeek-V3

671B parameters · 37B active per token

"671B total parameters" with "37B activated for each token". The HF upload totals 685B including the MTP module.

Quantization ?NeedsCheapest that fitsOptions
Q4 (4-bit)402.6 GBApple Mac Studio (M3 Ultra, 512GB) ($17,999)2
Q5 (5-bit)503.3 GBApple Mac Studio (M3 Ultra, 512GB) ($17,999)2
Q8 (8-bit)805.2 GBNothing we track fits it0
FP161610.4 GBNothing we track fits it0

Parameter count from DeepSeek-V3 model card, fetched

GLM-5.1

744B parameters · 40B active per token

Mixture of experts: the official GLM-5 repository lists GLM-5.1 as 744B total parameters with 40B active.

Quantization ?NeedsCheapest that fitsOptions
Q4 (4-bit)446.4 GBApple Mac Studio (M3 Ultra, 512GB) ($17,999)2
Q5 (5-bit)558.0 GBNothing we track fits it0
Q8 (8-bit)892.8 GBNothing we track fits it0
FP161785.6 GBNothing we track fits it0

Parameter count from GLM-5.1 model card, fetched , GLM-5 official repository, fetched

Kimi K3

2800B parameters · 104B active per token

Mixture of experts: card table reads Total Parameters 2.8T, Activated Parameters 104B, 896 experts with 16 selected per token, 1,048,576 context. All weights must be resident, so nothing we track fits it at any quantization.

Quantization ?NeedsCheapest that fitsOptions
Q4 (4-bit)1680.0 GBNothing we track fits it0
Q5 (5-bit)2100.0 GBNothing we track fits it0
Q8 (8-bit)3360.0 GBNothing we track fits it0
FP166720.0 GBNothing we track fits it0

Parameter count from Kimi-K3 model card, fetched

Mixture-of-experts models activate only a fraction of their weights per token, but every weight still has to be resident in memory, so sizing uses the total. See what can you actually run at home for why capacity and bandwidth answer different questions, and our methodology for how these figures are sourced.