What can it run
Popular open-weight models, and the cheapest box we track that has room for each one. Memory capacity decides whether a model loads at all, so this is the first question worth answering. Prices last checked .
The requirement is the weights at that quantization plus a 20 percent allowance for the KV cache, activations, and the runtime. It is arithmetic on the format, not a measurement. Long contexts and large batches push the real number up, sometimes a lot. Every parameter count links to the model card it came from.
Search
Family
Quantization
Hardware budget
25 of 25 models
Qwen3.5-0.8B
0.8B parameters
Model card lists 0.8B parameters and a native 262,144-token context window.
| Quantization ? | Needs | Cheapest that fits | Options |
|---|---|---|---|
| Q4 (4-bit) | 0.5 GB | Tenstorrent Blackhole p150a ($1,399) | 19 |
| Q5 (5-bit) | 0.6 GB | Tenstorrent Blackhole p150a ($1,399) | 19 |
| Q8 (8-bit) | 1.0 GB | Tenstorrent Blackhole p150a ($1,399) | 19 |
| FP16 | 1.9 GB | Tenstorrent Blackhole p150a ($1,399) | 19 |
Parameter count from Qwen3.5-0.8B model card, fetched
LFM2.5-1.2B-Thinking
1.17B parameters
Model card lists 1.17B parameters and a 32,768-token context window.
| Quantization ? | Needs | Cheapest that fits | Options |
|---|---|---|---|
| Q4 (4-bit) | 0.7 GB | Tenstorrent Blackhole p150a ($1,399) | 19 |
| Q5 (5-bit) | 0.9 GB | Tenstorrent Blackhole p150a ($1,399) | 19 |
| Q8 (8-bit) | 1.4 GB | Tenstorrent Blackhole p150a ($1,399) | 19 |
| FP16 | 2.8 GB | Tenstorrent Blackhole p150a ($1,399) | 19 |
Parameter count from LFM2.5-1.2B-Thinking model card, fetched
Qwen3.5-2B
2B parameters
Model card lists 2B parameters and a native 262,144-token context window.
| Quantization ? | Needs | Cheapest that fits | Options |
|---|---|---|---|
| Q4 (4-bit) | 1.2 GB | Tenstorrent Blackhole p150a ($1,399) | 19 |
| Q5 (5-bit) | 1.5 GB | Tenstorrent Blackhole p150a ($1,399) | 19 |
| Q8 (8-bit) | 2.4 GB | Tenstorrent Blackhole p150a ($1,399) | 19 |
| FP16 | 4.8 GB | Tenstorrent Blackhole p150a ($1,399) | 19 |
Parameter count from Qwen3.5-2B model card, fetched
SmolLM3-3B
3B parameters
Model card calls it a 3B parameter model; the current config supports 65,536 tokens and the card documents YaRN extensions for 128K and 256K.
| Quantization ? | Needs | Cheapest that fits | Options |
|---|---|---|---|
| Q4 (4-bit) | 1.8 GB | Tenstorrent Blackhole p150a ($1,399) | 19 |
| Q5 (5-bit) | 2.3 GB | Tenstorrent Blackhole p150a ($1,399) | 19 |
| Q8 (8-bit) | 3.6 GB | Tenstorrent Blackhole p150a ($1,399) | 19 |
| FP16 | 7.2 GB | Tenstorrent Blackhole p150a ($1,399) | 19 |
Parameter count from SmolLM3-3B model card, fetched
Qwen3.5-4B
4B parameters
Model card lists 4B parameters and a native 262,144-token context window.
| Quantization ? | Needs | Cheapest that fits | Options |
|---|---|---|---|
| Q4 (4-bit) | 2.4 GB | Tenstorrent Blackhole p150a ($1,399) | 19 |
| Q5 (5-bit) | 3.0 GB | Tenstorrent Blackhole p150a ($1,399) | 19 |
| Q8 (8-bit) | 4.8 GB | Tenstorrent Blackhole p150a ($1,399) | 19 |
| FP16 | 9.6 GB | Tenstorrent Blackhole p150a ($1,399) | 19 |
Parameter count from Qwen3.5-4B model card, fetched
Gemma 4 E4B
8B parameters
Google lists 4.5B effective parameters and 8B total with embeddings. Sizing uses all 8B resident weights, not the effective count.
| Quantization ? | Needs | Cheapest that fits | Options |
|---|---|---|---|
| Q4 (4-bit) | 4.8 GB | Tenstorrent Blackhole p150a ($1,399) | 19 |
| Q5 (5-bit) | 6.0 GB | Tenstorrent Blackhole p150a ($1,399) | 19 |
| Q8 (8-bit) | 9.6 GB | Tenstorrent Blackhole p150a ($1,399) | 19 |
| FP16 | 19.2 GB | Tenstorrent Blackhole p150a ($1,399) | 19 |
Parameter count from Gemma 4 E4B IT model card, fetched
Qwen3-8B
8.2B parameters
Model card: "Number of Parameters: 8.2B". Context 32,768 natively, 131,072 with YaRN.
| Quantization ? | Needs | Cheapest that fits | Options |
|---|---|---|---|
| Q4 (4-bit) | 4.9 GB | Tenstorrent Blackhole p150a ($1,399) | 19 |
| Q5 (5-bit) | 6.1 GB | Tenstorrent Blackhole p150a ($1,399) | 19 |
| Q8 (8-bit) | 9.8 GB | Tenstorrent Blackhole p150a ($1,399) | 19 |
| FP16 | 19.7 GB | Tenstorrent Blackhole p150a ($1,399) | 19 |
Parameter count from Qwen3-8B model card, fetched
Qwen3.5-9B
9B parameters
Model card lists 9B parameters and a native 262,144-token context window.
| Quantization ? | Needs | Cheapest that fits | Options |
|---|---|---|---|
| Q4 (4-bit) | 5.4 GB | Tenstorrent Blackhole p150a ($1,399) | 19 |
| Q5 (5-bit) | 6.8 GB | Tenstorrent Blackhole p150a ($1,399) | 19 |
| Q8 (8-bit) | 10.8 GB | Tenstorrent Blackhole p150a ($1,399) | 19 |
| FP16 | 21.6 GB | Tenstorrent Blackhole p150a ($1,399) | 19 |
Parameter count from Qwen3.5-9B model card, fetched
Gemma 4 12B
11.95B parameters
Google lists 11.95B total parameters and a 256K context window for the encoder-free multimodal model.
| Quantization ? | Needs | Cheapest that fits | Options |
|---|---|---|---|
| Q4 (4-bit) | 7.2 GB | Tenstorrent Blackhole p150a ($1,399) | 19 |
| Q5 (5-bit) | 9.0 GB | Tenstorrent Blackhole p150a ($1,399) | 19 |
| Q8 (8-bit) | 14.3 GB | Tenstorrent Blackhole p150a ($1,399) | 19 |
| FP16 | 28.7 GB | Tenstorrent Blackhole p150a ($1,399) | 17 |
Parameter count from Gemma 4 12B IT model card, fetched
gpt-oss-20b
21B parameters · 3.6B active per token
OpenAI: "21B parameters with 3.6B active parameters"; card says it can run within 16GB of memory.
| Quantization ? | Needs | Cheapest that fits | Options |
|---|---|---|---|
| Q4 (4-bit) | 12.6 GB | Tenstorrent Blackhole p150a ($1,399) | 19 |
| Q5 (5-bit) | 15.8 GB | Tenstorrent Blackhole p150a ($1,399) | 19 |
| Q8 (8-bit) | 25.2 GB | Tenstorrent Blackhole p150a ($1,399) | 17 |
| FP16 | 50.4 GB | Apple Mac mini (M4 Pro, 64GB) ($2,800) | 14 |
Parameter count from gpt-oss-120b model card, fetched
Mistral Small 3.2 24B
24B parameters
Model card lists 24B params. Mistral does not state a context window on the card.
| Quantization ? | Needs | Cheapest that fits | Options |
|---|---|---|---|
| Q4 (4-bit) | 14.4 GB | Tenstorrent Blackhole p150a ($1,399) | 19 |
| Q5 (5-bit) | 18.0 GB | Tenstorrent Blackhole p150a ($1,399) | 19 |
| Q8 (8-bit) | 28.8 GB | Tenstorrent Blackhole p150a ($1,399) | 17 |
| FP16 | 57.6 GB | Apple Mac mini (M4 Pro, 64GB) ($2,800) | 14 |
Parameter count from Mistral Small 3.2 24B Instruct model card, fetched
Gemma 4 26B-A4B
25.2B parameters · 3.8B active per token
Mixture of experts: Google lists 25.2B total parameters and 3.8B active, with a 256K context window.
| Quantization ? | Needs | Cheapest that fits | Options |
|---|---|---|---|
| Q4 (4-bit) | 15.1 GB | Tenstorrent Blackhole p150a ($1,399) | 19 |
| Q5 (5-bit) | 18.9 GB | Tenstorrent Blackhole p150a ($1,399) | 19 |
| Q8 (8-bit) | 30.2 GB | Tenstorrent Blackhole p150a ($1,399) | 17 |
| FP16 | 60.5 GB | Apple Mac mini (M4 Pro, 64GB) ($2,800) | 14 |
Parameter count from Gemma 4 26B-A4B IT model card, fetched
Gemma 3 27B
27B parameters
Model card lists 27B params, with a 128K input context window.
| Quantization ? | Needs | Cheapest that fits | Options |
|---|---|---|---|
| Q4 (4-bit) | 16.2 GB | Tenstorrent Blackhole p150a ($1,399) | 19 |
| Q5 (5-bit) | 20.3 GB | Tenstorrent Blackhole p150a ($1,399) | 19 |
| Q8 (8-bit) | 32.4 GB | Apple Mac mini (M4 Pro, 64GB) ($2,800) | 15 |
| FP16 | 64.8 GB | Framework Desktop (Ryzen AI Max+ 395, 128GB) ($3,449) | 11 |
Parameter count from Gemma 3 27B IT model card, fetched
Qwen3.5-27B
27B parameters
Model card lists 27B parameters and a native 262,144-token context window.
| Quantization ? | Needs | Cheapest that fits | Options |
|---|---|---|---|
| Q4 (4-bit) | 16.2 GB | Tenstorrent Blackhole p150a ($1,399) | 19 |
| Q5 (5-bit) | 20.3 GB | Tenstorrent Blackhole p150a ($1,399) | 19 |
| Q8 (8-bit) | 32.4 GB | Apple Mac mini (M4 Pro, 64GB) ($2,800) | 15 |
| FP16 | 64.8 GB | Framework Desktop (Ryzen AI Max+ 395, 128GB) ($3,449) | 11 |
Parameter count from Qwen3.5-27B model card, fetched
Qwen3.8-27B
27B parameters
Model card lists 27B parameters and a native 262,144-token context window, extensible to 1,000,000 tokens.
| Quantization ? | Needs | Cheapest that fits | Options |
|---|---|---|---|
| Q4 (4-bit) | 16.2 GB | Tenstorrent Blackhole p150a ($1,399) | 19 |
| Q5 (5-bit) | 20.3 GB | Tenstorrent Blackhole p150a ($1,399) | 19 |
| Q8 (8-bit) | 32.4 GB | Apple Mac mini (M4 Pro, 64GB) ($2,800) | 15 |
| FP16 | 64.8 GB | Framework Desktop (Ryzen AI Max+ 395, 128GB) ($3,449) | 11 |
Parameter count from Qwen3.8-27B model card, fetched
Gemma 4 31B
30.7B parameters
Google lists 30.7B total parameters and a 256K context window for the dense model.
| Quantization ? | Needs | Cheapest that fits | Options |
|---|---|---|---|
| Q4 (4-bit) | 18.4 GB | Tenstorrent Blackhole p150a ($1,399) | 19 |
| Q5 (5-bit) | 23.0 GB | Tenstorrent Blackhole p150a ($1,399) | 19 |
| Q8 (8-bit) | 36.8 GB | Apple Mac mini (M4 Pro, 64GB) ($2,800) | 15 |
| FP16 | 73.7 GB | Framework Desktop (Ryzen AI Max+ 395, 128GB) ($3,449) | 11 |
Parameter count from Gemma 4 31B IT model card, fetched
Qwen3.5-35B-A3B
35B parameters · 3B active per token
Mixture of experts: model card lists 35B total parameters and 3B active, with a native 262,144-token context window.
| Quantization ? | Needs | Cheapest that fits | Options |
|---|---|---|---|
| Q4 (4-bit) | 21.0 GB | Tenstorrent Blackhole p150a ($1,399) | 19 |
| Q5 (5-bit) | 26.3 GB | Tenstorrent Blackhole p150a ($1,399) | 17 |
| Q8 (8-bit) | 42.0 GB | Apple Mac mini (M4 Pro, 64GB) ($2,800) | 15 |
| FP16 | 84.0 GB | Framework Desktop (Ryzen AI Max+ 395, 128GB) ($3,449) | 9 |
Parameter count from Qwen3.5-35B-A3B model card, fetched
Llama 3.3 70B Instruct
70B parameters
Card table lists 70B; the computed safetensors count reads 71B. We use the nominal 70B.
| Quantization ? | Needs | Cheapest that fits | Options |
|---|---|---|---|
| Q4 (4-bit) | 42.0 GB | Apple Mac mini (M4 Pro, 64GB) ($2,800) | 15 |
| Q5 (5-bit) | 52.5 GB | Apple Mac mini (M4 Pro, 64GB) ($2,800) | 14 |
| Q8 (8-bit) | 84.0 GB | Framework Desktop (Ryzen AI Max+ 395, 128GB) ($3,449) | 9 |
| FP16 | 168.0 GB | Apple Mac Studio (M3 Ultra, 512GB) ($17,999) | 2 |
Parameter count from Llama-3.3-70B-Instruct model card, fetched
gpt-oss-120b
117B parameters · 5.1B active per token
OpenAI: "117B parameters with 5.1B active parameters"; designed to fit a single 80GB GPU using MXFP4.
| Quantization ? | Needs | Cheapest that fits | Options |
|---|---|---|---|
| Q4 (4-bit) | 70.2 GB | Framework Desktop (Ryzen AI Max+ 395, 128GB) ($3,449) | 11 |
| Q5 (5-bit) | 87.8 GB | Framework Desktop (Ryzen AI Max+ 395, 128GB) ($3,449) | 8 |
| Q8 (8-bit) | 140.4 GB | Apple Mac Studio (M3 Ultra, 512GB) ($17,999) | 2 |
| FP16 | 280.8 GB | Apple Mac Studio (M3 Ultra, 512GB) ($17,999) | 2 |
Parameter count from gpt-oss-120b model card, fetched
Qwen3.5-122B-A10B
122B parameters · 10B active per token
Mixture of experts: model card lists 122B total parameters and 10B active, with a native 262,144-token context window.
| Quantization ? | Needs | Cheapest that fits | Options |
|---|---|---|---|
| Q4 (4-bit) | 73.2 GB | Framework Desktop (Ryzen AI Max+ 395, 128GB) ($3,449) | 11 |
| Q5 (5-bit) | 91.5 GB | Framework Desktop (Ryzen AI Max+ 395, 128GB) ($3,449) | 8 |
| Q8 (8-bit) | 146.4 GB | Apple Mac Studio (M3 Ultra, 512GB) ($17,999) | 2 |
| FP16 | 292.8 GB | Apple Mac Studio (M3 Ultra, 512GB) ($17,999) | 2 |
Parameter count from Qwen3.5-122B-A10B model card, fetched
Qwen3-235B-A22B
235B parameters · 22B active per token
Mixture of experts: "235B in total and 22B activated". All weights must still be resident.
| Quantization ? | Needs | Cheapest that fits | Options |
|---|---|---|---|
| Q4 (4-bit) | 141.0 GB | Apple Mac Studio (M3 Ultra, 512GB) ($17,999) | 2 |
| Q5 (5-bit) | 176.3 GB | Apple Mac Studio (M3 Ultra, 512GB) ($17,999) | 2 |
| Q8 (8-bit) | 282.0 GB | Apple Mac Studio (M3 Ultra, 512GB) ($17,999) | 2 |
| FP16 | 564.0 GB | Nothing we track fits it | 0 |
Parameter count from Qwen3-235B-A22B model card, fetched
DeepSeek-V4-Flash-0731
284B parameters · 13B active per token
Mixture of experts: the DeepSeek-V4 report lists 284B total parameters and 13B active, with a one-million-token context window.
| Quantization ? | Needs | Cheapest that fits | Options |
|---|---|---|---|
| Q4 (4-bit) | 170.4 GB | Apple Mac Studio (M3 Ultra, 512GB) ($17,999) | 2 |
| Q5 (5-bit) | 213.0 GB | Apple Mac Studio (M3 Ultra, 512GB) ($17,999) | 2 |
| Q8 (8-bit) | 340.8 GB | Apple Mac Studio (M3 Ultra, 512GB) ($17,999) | 2 |
| FP16 | 681.6 GB | Nothing we track fits it | 0 |
Parameter count from DeepSeek-V4-Flash-0731 model card, fetched , DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence, fetched
DeepSeek-V3
671B parameters · 37B active per token
"671B total parameters" with "37B activated for each token". The HF upload totals 685B including the MTP module.
| Quantization ? | Needs | Cheapest that fits | Options |
|---|---|---|---|
| Q4 (4-bit) | 402.6 GB | Apple Mac Studio (M3 Ultra, 512GB) ($17,999) | 2 |
| Q5 (5-bit) | 503.3 GB | Apple Mac Studio (M3 Ultra, 512GB) ($17,999) | 2 |
| Q8 (8-bit) | 805.2 GB | Nothing we track fits it | 0 |
| FP16 | 1610.4 GB | Nothing we track fits it | 0 |
Parameter count from DeepSeek-V3 model card, fetched
GLM-5.1
744B parameters · 40B active per token
Mixture of experts: the official GLM-5 repository lists GLM-5.1 as 744B total parameters with 40B active.
| Quantization ? | Needs | Cheapest that fits | Options |
|---|---|---|---|
| Q4 (4-bit) | 446.4 GB | Apple Mac Studio (M3 Ultra, 512GB) ($17,999) | 2 |
| Q5 (5-bit) | 558.0 GB | Nothing we track fits it | 0 |
| Q8 (8-bit) | 892.8 GB | Nothing we track fits it | 0 |
| FP16 | 1785.6 GB | Nothing we track fits it | 0 |
Parameter count from GLM-5.1 model card, fetched , GLM-5 official repository, fetched
Kimi K3
2800B parameters · 104B active per token
Mixture of experts: card table reads Total Parameters 2.8T, Activated Parameters 104B, 896 experts with 16 selected per token, 1,048,576 context. All weights must be resident, so nothing we track fits it at any quantization.
| Quantization ? | Needs | Cheapest that fits | Options |
|---|---|---|---|
| Q4 (4-bit) | 1680.0 GB | Nothing we track fits it | 0 |
| Q5 (5-bit) | 2100.0 GB | Nothing we track fits it | 0 |
| Q8 (8-bit) | 3360.0 GB | Nothing we track fits it | 0 |
| FP16 | 6720.0 GB | Nothing we track fits it | 0 |
Parameter count from Kimi-K3 model card, fetched
Mixture-of-experts models activate only a fraction of their weights per token, but every weight still has to be resident in memory, so sizing uses the total. See what can you actually run at home for why capacity and bandwidth answer different questions, and our methodology for how these figures are sourced.