Skip to main content

GGUF in Transformers: check the memory path before buying

Do not buy a smaller-memory machine just because Transformers can load a small GGUF file. It may unpack those weights into a larger model when it loads them. The useful part of Hugging Face’s September 22 announcement is narrower: supported models can now keep their weights compressed while running through Transformers on Apple Silicon.

That gives people who use Python and PyTorch another reason to try the Mac they already own. It does not establish a new memory budget for every Mac, NVIDIA card, or model.

The buying question keeps coming up. A recent r/LocalLLM thread asks what practical limits separate Apple/Metal from CUDA and ROCm. Another r/LocalLLaMA discussion weighs a small GPU build against a Mac mini, with prompt-processing speed part of the decision. Those are hardware questions, but the software path needs to be settled before the shopping list.

One file, two very different memory paths

The current Transformers docs describe two ways to load a GGUF:

Installing kernels is part of the setup, not proof that your model will take the packed path. The announcement’s limitations say it is MPS-only for now. Architecture support covers Qwen3.5 dense and mixture-of-experts models, including compatible Qwen3.8 checkpoints. The docs say other architectures use the legacy loader, which always dequantizes. “Supports GGUF” does not mean “keeps every GGUF compressed on every device.”

There is a separate trap for people buying a machine to fine-tune. The training example in the announcement explicitly uses GgufConfig(dequantize=True). That is not a demonstration of training while keeping the weights packed. Do not use the inference example as a training-memory requirement.

What this changes for a buyer

Already have an Apple Silicon Mac and need Transformers? Try it before upgrading. Hugging Face describes uses such as inspecting activations, evaluating quantized checkpoints, and changing generation behavior with Python. This is the appeal: working with a GGUF inside familiar development tools, not a promise that you need a newer computer.

Just want local chat? There is no need to switch for this announcement alone. Hugging Face still recommends llama.cpp when efficient local inference is the priority. Our runtime setup guide covers the existing options.

Buying an NVIDIA or AMD system? Do not count this as a new packed-GGUF path for that purchase. The announced path is MPS-only. That says nothing about what a different runtime can do on your card. Ask for a test of the runtime you will actually use.

Also resist turning the announcement’s chart into a hardware ranking. Its benchmark notes describe testing on an M2 Max MacBook Pro. The Transformers measurement includes prompt processing and takes the best warmed run; the llama.cpp measurement excludes prompt processing and averages repetitions. It is not a matched test of a Mac against a GPU workstation. We have not run our own comparison.

The test to ask for before paying

Start with the exact model file, not just the model family. Write down its quantization, your intended prompt length, and whether you need one conversation, several requests, or training.

Then use a separate environment and follow Hugging Face’s current setup instructions. At publication, those instructions require Transformers from main until the next release, plus compatible PyTorch and kernel builds. Do not replace a working environment just to try this.

Check which loading path was selected, the device, any warnings, and observed memory use. If the result is unclear, do not assume it stayed packed. Try your real long prompt and record both the wait for the first token and the time to finish. If you need simultaneous requests, test them: Hugging Face says the initial target is a single interactive conversation and that padding and batching still need work.

The Reddit announcement thread already has questions about other applications and training. Treat those as things to verify, not features your purchase can rely on.

Use our model-fit tool and quantization guide to narrow the choices, not to certify this loader. For shopping context, the tracked Mac Studio and RTX 5090 have a side-by-side catalog view. Those pages are not benchmarks of this integration.

My rule: choose the model and software path first. Buy more memory only after a test shows what is missing. A small download is not a memory guarantee.

All news