Skip to main content
AI Gear Watch

Rig builds by budget

Updated 2026-08-15

Every buying guide you have read picks winners from benchmark charts. This one cannot, because we do not publish throughput numbers we did not measure. What we can do is better anyway: pick on the two specs that decide what runs at all, and be explicit about the tradeoff at each price point.

Read what can you actually run at home first if you have not. The short version: capacity decides whether a model loads, bandwidth decides how fast it answers.

A note on the prices below

Only some vendors in our catalog publish a list price, and our daily price checks have not yet filled in street prices. Where we quote a number, it is the vendor’s own published figure and we say whose. Where we do not, check the hardware list for what the daily run has confirmed. Do not treat any figure here as today’s street price.

Roughly $1,000: one card, 24 GB

At this level you are buying a single consumer GPU and reusing a machine you already own.

The RX 7900 XTX gives you 24 GB of GDDR6 at 960 GB/s for an AMD suggested price of $999, and it is the only entry in our catalog with a vendor-published price at this tier. The RTX 4090 is the other obvious candidate: also 24 GB, at 1008 GB/s, but it has no current list price and lives on the used market, which the buying used guide covers.

24 GB is the number that matters. It comfortably holds a 30B model at 4-bit and a 13B at 8-bit. It does not hold a 70B at any quantization worth running. If a 70B is your target, skip this tier entirely rather than buying twice.

The catch nobody mentions until the parts arrive: the 7900 XTX draws 355 W and the 4090 draws 450 W, both card-only figures that say nothing about the rest of your machine. Read power, cooling, and noise before you assume your existing power supply copes.

The AMD versus NVIDIA question here is really a software question. CUDA has the deepest support across every runtime; ROCm works and has closed much of the gap. Our setup guide notes where that distinction bites, most sharply with vLLM.

Roughly $2,000 to $3,500: capacity over speed

This is where the interesting decision lives, and it is not “a faster card.” It is whether to buy bandwidth or capacity.

Unified-memory machines buy capacity. The Framework Desktop is $3,449 from Framework with 128 GB of LPDDR5X, and the GMKtec EVO-X2 is the same Strix Halo silicon with the same 128 GB, listed on GMKtec’s own store at $1,999.99 on sale against a $2,199.99 regular price. Both are complete computers, not cards. The GMKtec draws 120 W steadily with peaks to 140 W, which is a whole-system figure and a fraction of what a single high-end GPU pulls.

128 GB changes what is possible. A 70B model at 8-bit fits. A 120B at 4-bit fits. Neither is remotely possible on a 24 GB card.

The honest catch: neither vendor publishes a memory bandwidth figure for these machines, so our catalog records null rather than a guess, and you should assume generation speed is well below a discrete GPU’s. You are buying the ability to run large models slowly instead of small models quickly.

Apple sits in the same category with better-documented bandwidth. The Mac mini M4 Pro offers 64 GB at 273 GB/s, and the Mac Studio M4 Max offers 128 GB at 546 GB/s. That 546 figure is genuinely competitive for a unified-memory machine, roughly half a 4090’s bandwidth with more than five times the capacity.

If you would rather have speed, the RTX 5090 is 32 GB at 1792 GB/s, the highest bandwidth of any consumer card we track. Eight more gigabytes than a 4090 and nearly double the bandwidth. It is also a 575 W card that NVIDIA’s own page describes with two different system power requirements, 1000 W in the spec table and 850 W minimum in the setup section. Budget for the power supply.

Compare the two philosophies directly.

Roughly $5,000: buying real memory

Here the options diverge sharply, and which one is right depends entirely on whether you need to run one enormous model or serve many requests.

The RTX PRO 6000 Blackwell is the standout on paper: 96 GB of GDDR7 at 1792 GB/s. That is 5090 bandwidth with three times the memory, in one card, with a normal fan. It draws 600 W. Nothing else in our catalog combines that much capacity with that much bandwidth in a single desktop-installable card.

The RTX 6000 Ada is the previous generation: 48 GB at 960 GB/s, 300 W. Half the memory, half the bandwidth, half the power. Worth considering if the price gap is large.

Used datacenter cards are the other path. The H100 80GB PCIe delivers 80 GB of HBM2e at 2000 GB/s, and the A100 80GB delivers the same capacity at 1935 GB/s. Both beat everything else here on bandwidth.

Both are also passively cooled, and NVIDIA says so plainly: the H100 PCIe product brief lists Thermal solution: Passive and states it “requires system airflow to operate the card properly within its thermal limits.” In a desktop tower that airflow does not exist. Budget for a shroud and fan kit, and expect noise. This is the single most common expensive mistake at this tier, and both the power guide and the used-buying guide exist largely because of it.

The Mac Studio M3 Ultra belongs in this conversation from the opposite direction: up to 512 GB of unified memory at 819 GB/s, drawing a measured 270 W for the entire machine and idling at 9 W. Less than half the H100’s bandwidth, more than six times its memory, and quiet enough for a desk. For running a very large model at moderate speed, nothing else in the catalog is close on capacity.

No limit

At the top the question stops being which box and becomes how many, which makes multi-GPU splitting the guide you actually need. Read it before buying a second card, because the answer to “will two cards be twice as fast” is usually no, and the reason is your interconnect.

The DGX Spark is NVIDIA’s own take on the desktop AI machine: 128 GB of unified LPDDR5X at 273 GB/s in a 240 W box. Its appeal is the software stack rather than the raw numbers, which are close to what the Strix Halo machines offer for less.

Tenstorrent’s Blackhole p150a is the genuine wildcard: 32 GB of GDDR6 at 512 GB/s for a published $1,399, with an open software stack. We track the actively cooled p150a rather than the identical passive p150b for the reasons above. Software maturity is the question here, not hardware.

How to actually decide

Work in this order and the choice usually makes itself.

First, name the model you want to run. Not “large models”, an actual model at an actual quantization. Then look up what it needs on what can it run. Capacity is a hard wall; everything else is a preference.

Second, decide whether slow-and-large or fast-and-small serves you better. If you are chatting with one model interactively, capacity usually wins, because running the model you want slowly beats not running it. If you are serving many requests or iterating fast, bandwidth wins.

Third, check the power and cooling before you buy, not after. Passive card, power supply cables, where the heat goes. All three are answerable from documentation, and all three have ended builds.

Fourth, prefer one bigger box to two smaller ones when you can afford it. No split to configure, no interconnect bottleneck, no experimental flags.

What is not in this guide

No full parts lists. CPU, motherboard, PSU, and case choices depend on what you already own, and a list assembled without measuring your constraints is decoration.

No performance rankings or tokens-per-second comparisons. Every ranking above is on published specs, and where a vendor publishes nothing we say so instead of estimating. See our methodology.

No claim about what any of this costs today. Prices move, our daily checks are what track them, and the hardware list is the current answer.

All guides