One Strix Halo box for AI and your NAS? Test both together
If you want one box for local AI, file storage, and household services, do not buy it from an AI demo alone. Ask for a test with those jobs running together. I would keep essential services off a machine bought to run a large model near its memory limit, unless that shared setup has been tested.
This is not just a hypothetical parts-list problem. A recent r/homelab buyer wants a NAS, local models, media services, and occasional game servers in one build. Another owner’s Strix Halo NAS write-up prompted questions about slowdowns and reports of host memory pressure. Those are owner reports, not tests we reproduced. They are good reasons to ask what happens outside the chat window.
There is a timely software reason to ask, too. Halogen’s changelog documents an opt-in weight-locking setting added in version 0.13.2 for memory-pressure problems. Its current instructions still say “Give it a machine of its own.” That is a buying constraint, not a footnote to a speed chart.
What the fast demo leaves out
Halogen is a specialized runtime for Qwen3.8-Flash-Next on AMD Strix Halo, not a general promise about every model on every AMD machine. Its maintainer says the weights stay resident and the context-cache pool is reserved up front. The same documentation warns that other large processes competing for the remaining memory can produce long stalls.
That distinction matters in this week’s AMD-box discussion, where readers recommend specialized runtimes and debate buying a different platform. A result from one of those runtimes is a reason to test that exact setup. It is not evidence that a NAS, photo library, and coding assistant will all stay responsive beside it.
Even the memory display needs context. In its
shared-host notes,
Halogen’s maintainer warns that free and MemAvailable can overstate
usable room in the described configuration. The server prints a separate
host-memory estimate. Do not treat a reassuring dashboard as proof that
another large job will fit.
The weight-locking option is not extra capacity. The changelog reports that it keeps registered weights from being reclaimed, but also describes the server being killed when another process exhausts the remaining memory. The sharing instructions offer smaller cache pools and other settings, with tradeoffs in resident conversations, context, or speed. Do not copy a tuning flag and assume the shared-host problem is solved.
These are Halogen’s documented limits and maintainer-reported results. They do not establish that every Strix Halo system has the same problem, or that Ollama behaves the same way. We have not benchmarked this setup. Also read Halogen’s license: the engine ships as a compiled binary under its own agreement. A public GitHub repository is not a promise that you can maintain the engine yourself.
The test I would ask for before buying
Write down what must keep working while AI runs. File access? DNS? Media playback? Then test on a backed-up system during a maintenance window, not by pushing the family’s only server to failure.
- Use the exact software you plan to keep. Record the model file, quantization, runtime version, context settings, and container or VM setup. A bare-metal demo does not test your proposed virtualized build.
- Start with your normal services already running. Load the model, send a realistic long prompt, and check those services throughout the run. Repeat with the simultaneous AI requests you actually need.
- Test a cold start and a follow-up. Record the wait for the first useful answer, not just the rate once text starts arriving. Then restart the AI service and check that it can load again without disrupting the rest.
- Check memory from both sides. Read the runtime’s startup messages as well as host metrics. Stop the test if services become unresponsive; do not count “the model eventually answered” as a pass.
- Test giving the memory back. Stop the model and confirm that normal service behavior returns. If AI is occasional, make unloading part of the plan rather than leaving every model resident.
For Ollama users, the
official FAQ
documents ollama stop and the API’s keep_alive: 0 for unloading. It also
explains that parallel requests increase context-memory allocation. These
are useful controls to test, not guarantees that your whole server fits.
Use our model-fit tool to narrow the hardware list. It is an arithmetic estimate, not a test of your other services or Halogen’s cache settings. The home inference guide and runtime setup guide help with the next questions. If your shortlist includes a Framework Desktop and a DGX Spark, use the comparison page for catalog details, not as a benchmark of this shared-server setup.
My buying rule: a shared box makes sense only if it passes the shared-work test. If the model needs nearly the whole machine, price the build as a dedicated AI server. Keeping an existing NAS separate may be the less disruptive choice, even when one new box looks neater.