I leave it running at night. No danger of burning my token subscriptions and it has hours and hours to run slowly with a manager like: github.com/kunchenguid/gnhf
If I had to pick a product, I'd say an affordable 32GB mac would be the sweet spot for running local models that function well like Qwen 3.8.
It's true, most people don't run models, but being the default platform for running open weights seems like it has plenty of advantages right now. Just like sales benefited from developers defaulting to MacOS for most open source languages like Ruby, Go, Rust, and TypeScript.
32GB of fast unified memory is enough for Qwen 3.8 27B.
- 16GB for the weights at Q4
- 9GB for the full 256K context at Q8
- 7GB spare for overhead and system.
The problem is that these Macs have 32GB of slow unified memory.
Edit: I'm thinking of a headless Mac mini, if you meant running it on the same machine you're using of course you'll need more memory, but LLMs are best served from a headless server so that's what I'd recommend.
Your agent(s) need to work on something, ie. running TypeScript, your app, your tests, Docker, Redis and/or database you need to run harness and user side apps ie. VScode, browser etc. it all adds up quickly.
Single user conversation spawns multiple parallel backend conversations, you need extra room for it as well, not just single context.
This plus usual apps like Mail, Spotify, iTerm2, SourceTree etc. also fill in memory.
Also 4 bit quantization is already quite aggressive compromise (measurable but sometimes acceptable loss, compared to ie. 8 bits which are often practically lossless) – for weights it's ok'ish, sometimes (especially if model was trained as 4 bit quants aware), but activations need to stay at higher bits taking more memory, otherwise quality degrades a lot.
For a dedicated headless setup, I’d probably use something like NVIDIA DGX Spark rather than a Mac (to be more precise NVIDIA GB10 Grace Blackwell from other suppliers than directly NVidia, they are much cheaper and have same insides). Linux is much better for running headless server, you also get standard NVIDIA/CUDA ecosystem instead of being tied to Metal/macOS.
For people who are interested in buying IMHO I'd wait a bit – next generation of Spark and/or Macs that are going to come out next year will be much better / will cross the line of being actually useful, not just a toy with goldfish LLM.
Keep in mind that if you want MTP it adds a few gigs. If you use sub-agents it turns already slow generation into even slower generation. Won't be doing any compling (so rust, c and probably go are not avaiable) becase those add memory pressure during compiling.
32gb of unified memory is enough enough for system to be used for anything other than LLM generation.
You can, you just need a beefier PC, and it's more annoying in terms of noise and heat vs throwing something on your server closet. Plus you don't need to worry about other software stealing resources and whatnot.
> Edit: I'm thinking of a headless Mac mini, if you meant running it on the same machine you're using of course you'll need more memory, but LLMs are best served from a headless server so that's what I'd recommend.
What? LLMs are best served from a massive PD disaggregated cluster of B300s connected via NVLink.
If you're running LLMs on a Mac Mini, it's because you want to run local, not because it's the best setup.
>massive PD disaggregated cluster of B300s connected via NVLink.
So a headless server.
Macs were mentioned because that's what the post is about. It could be a PC (I use a 2x3090 PC). The point is that it's a better experience to have a box dedicated to the LLM than running it in your system. Obviously in your home, so local.
> If I had to pick a product, I'd say an affordable 32GB mac would be the sweet spot for running local models that function well like Qwen 3.8.
32GB is not enough RAM. I don't even own a device with less than 36GB at this point, and that device I only have because my employer is being cheap. 64GB is a reasonable starting point for running local LLMs + normal tasks. 128GB let's you really run most smaller models like Qwen 27B and 35BA3B with good context. Even Qwen3.8-Flash-Next runs in 128GB with a 4-bit quant.
32GB would be limited to running models like Gemma4 12B and smaller dense Qwen versions like 9B unless you were using very small quants which damages quality of response.
You are mistaken. I'm running Qwen 3.7 28B 4bit (MLX) with a 200k context window and everything total is 32GB RSS.
Is this the best? No. That's why I said the sweet spot. Getting from 16GB macs to 32GB is perhaps possible. Jumping to 64GB or 128GB as the default is simply unreasonable right now.
I assume you mean Qwen 3.8-27B? Yes, you can run this in 32GB of RAM, but it's very context limited. With KV cache compression and other techniques, it's better now than in the past, but I'd still want more RAM, personally.
EDIT to add that you need to reserve 8GB for the system if you don't want to cause problems on macOS, which means 32GB RAM = 24GB max for model + context. It takes 18-19GB to load a 4-bit quant of Qwen3.8-27B, so I'd be really surprised if you can actually get a 200k context window. You need to fit within a 24GB WSS (which is generally a more constrained RSS) to get stable performance on 32GB RAM.
I run Qwen 3.8 27B just fine on my Mac mini M4 24GB. I use Unsloth's Q3 XXS with 128k context. It successfully completes long horizon tasks with OpenCode.
I have a similar machine, and briefly poked at running a local LLM, but got discouraged after a couple days. The quality, responsiveness, and impact on the rest of the system didn’t seem worth it to me.
What sorts of things are you doing with the local LLM? Anything interactive? Should I take another look?
Yes, 15-30 t/sec is pretty slow for local models so I recommend running local LLM tasks overnight where (vs paid plans) there isn't a risk of chewing through your token budget from a rogue loop or sub-agent. Even if it takes hours, you're sleeping anyway so no concern. herdr + pi works great for this but there are lots of harnesses.
Htmx 4.0 is over 100kb / 2,000 lines of code. Isn't there a way to slim this down some? I mean, isn't the basic idea just a couple lines of code? I remember AJAX 2.0 or whatever loaders that were only 1-5kb wrappers around XHR back in the day. Now we have fetch() which is even more basic:
I don't mean to imply this is close to total number of Htmx features, but fetching islands of content is the core idea and I feel like we're way past feature bloat at this point for something that is supposed to be simpler than React (Preact is only 10kB)
Yeah but when you're reading the actual source code you're not thinking in terms of bytes, you're looking and files and line counts. Htmx is deliberately maintained as a single, no-dependency file of about 5000loc. The maintainers are of the opinion that it reduces complexity, and I tend to agree.
Well, I think the cybertruck was announced at least a year or two before they shipped. A month wait is one thing, but giving people literal years to think about and find flaws with your product is a really bad move.
> First, generic methods are now supported
> Generic functions can now be used without explicit type arguments
Great! This was an ergonomic code issue I hit when trying to create a universal handler/controller generic that could hydrate/populate function arguments (from a request body) without having an actual copy of the arguments: https://github.com/xeoncross/mid/blob/main/handler.go#L12
Thanks for the link, looks like a very pleasant framework to use! I was interested to see "Mid-fasthttp" on the slower end in the benchmarks at the bottom, do you know why that is?
Btw, the Gin and Echo examples reference an "input" variable but it doesn't seem to be defined there? Maybe it was intentional, since the examples are just to give a general idea of how the handler looks in each library, but thought I would let you know just in case it wasn't.
Due to the fact that fasthttp is not net/http compatible it is due to the conversion being required from a https://pkg.go.dev/net/http#Handler. It was included for information purposes as I though someone would be curious.
Watching time compress as you age is indeed illuminating.
Another thing is arguments you have as a child with an adult. It's not that kids are always "technically" wrong when they argue their cases or reasons for things - It's that you realize how impactful actions/words are and often technical details don't matter as much as the grand scheme of things.
reply