hardware · research
Best mini PC for a useful local LLM? Start with memory
A practical reading of the 2024 M4 Mac mini for local models, including memory tiers, model-weight arithmetic and the evidence still missing.
· 2 min read
Evidence: 2024 Apple specifications reviewed 7 August 2026; this machine has not been benchmarked by Out of Browser
Review summary
- Verdict
- The 32GB M4 Mac mini is the sensible starting configuration; choose 48GB or more only for a measured larger-model need.
- Best for
- A quiet, compact local-inference machine with unified memory
- Not for
- Maximum generation speed, upgradeable RAM or a universal best-PC claim
For local language models, memory capacity removes candidates before processor branding decides anything. The machine must hold model weights, context cache, runtime overhead and the rest of the system. Buying the fastest chip with too little memory produces an expensive machine that cannot run the model you bought it for.
This is a specification-led recommendation, not a comparative hardware benchmark. The product is Apple’s 2024 M4 Mac mini, whose published specifications we checked on 7 August 2026. Out of Browser has not measured its tokens per second, time to first token or wall power against rival mini PCs, so “best” here means the clearest fit for a stated constraint—not a universal winner.
The verdict
For a new local-LLM box, 32GB of unified memory is the sensible M4 starting point. It leaves room for useful quantised models, context and the operating system without paying immediately for the M4 Pro tier. If your planned evaluation proves that you need a model around the 30B class or large contexts, move to 48GB or 64GB. Do not buy the base 16GB configuration for an unspecified future local-AI workload.
Apple lists the M4 model with 120GB/s memory bandwidth and configurations up to 32GB. The M4 Pro version lists 273GB/s and configurations up to 64GB. Those numbers make the product interesting; they do not tell you application speed by themselves.
Work out what must fit
A rough first pass for 4-bit weights is half a byte per parameter: about 4GB for 8B, 7GB for 14B, and 16GB for 32B before metadata, context cache and runtime overhead. Real files are larger and the system shares the same memory, so leave headroom. The fuller memory arithmetic explains why model weights alone are not the answer.
Write down your target model, quantisation and context length before selecting memory. If you cannot name them, buy for the smallest credible workload and keep the unspent money until the evidence changes.
What to benchmark before calling it best
Run the same model file and prompts on every shortlisted machine. Record prompt-processing speed, generation speed, time to first token, peak memory, watts at the wall and total joules per completed task. Also record noise and whether performance changes over a sustained 30-minute run. A fast first minute is not a server.
The Mac mini loses points for memory that cannot be upgraded later. Its compactness only wins if the fixed configuration matches the workload for the machine’s useful life.
Sources: Apple’s Mac mini (2024) technical specifications and Ollama.
Product links
Check the current product, price, plan and regional availability on the seller's page.
- See Mac mini — Official Apple link
- Visit Ollama — Official Ollama link
Take it further
- Local or nothing — Run open-weight models on hardware you own, and know what they cost you in watts and seconds. (5 lessons, 6 min, 2 free)
- Where the energy actually goes — Joules, not vibes. Measure what your own work costs before joining an argument about data centres. (3 lessons, 65 min, 3 free)