Skip to content
AI

Testing the 9x Smaller Local LLM Claim on a GPU-less VPS

Testing the 9x Smaller Local LLM Claim on a GPU-less VPS

Last month I did something no sane person should try on a Tuesday night: I downloaded a custom fork of llama.cpp from a company I’d never heard of, just to run a 27 billion parameter model that supposedly fits in under 6 gigabytes. I’d read about it in Simon Willison’s comment on the release, a ternary-quantized model called Bonsai 2 27B, pitched as “near-lossless compression in a 9x smaller footprint.” Near-lossless is a big claim for something throwing away most of the bits in every weight. I wanted to know whether that phrase meant what I hoped, or whether it was the kind of number that only survives contact with a benchmark table. So instead of running it on a beefy Mac like Willison did, I pointed it at the cheapest box I own: the small Hetzner VPS that runs this very blog’s publishing pipeline. Here’s what actually happened.

What “ternary” buys you and what it costs

Most quantized models you download are 4-bit or 8-bit: each weight still gets a handful of possible values. Ternary quantization is far more aggressive. Every weight collapses to one of three values, roughly minus one, zero, or plus one, and the model leans on scaling factors to recover something close to the original behavior. Prism’s Ternary-Bonsai-2-27B-gguf takes a 27 billion parameter model and compresses it down to a single GGUF file just under 6 gigabytes, using a PTQ1_0 format that most mainstream llama.cpp builds can’t even read yet. That’s why you need Prism’s llama.cpp fork instead of the stock binary: the 1-bit quantization support isn’t merged upstream.

The tradeoff is obvious once you say it out loud. You’re asking a model to represent everything it knows using something closer to a light switch than a dial. Whether that “near-lossless” label holds up depends entirely on what you’re asking it to do, and I was not about to take a marketing line at face value without running it myself.

Getting it running on a box with no GPU

Willison’s setup used the macOS Apple Silicon build and Metal acceleration, and reported 20 to 44 tokens per second depending on the run. I don’t have an M-series Mac sitting around for this kind of experiment, and I was more curious about the worst case anyway: what happens on a machine with no GPU at all. Prism ships CPU-only Ubuntu binaries alongside the CUDA and Metal ones, so I grabbed that instead.

curl -fL https://github.com/PrismML-Eng/llama.cpp/releases/download/prism-b10685-7dffb15/llama-prism-b10685-7dffb15-bin-ubuntu-x64.tar.gz -o bonsai-runtime.tar.gz
tar -xzf bonsai-runtime.tar.gz

curl -fL https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf/resolve/main/Ternary-Bonsai-2-27B-PTQ1_0.gguf -o bonsai.gguf

./llama-prism-b10685-7dffb15/llama-server -m bonsai.gguf --port 8331 -c 8192

The download itself was the fast part, five and a half minutes on the VPS’s connection for a model that would normally arrive as forty-plus gigabytes uncompressed. The server started fine and the built-in web UI came up at localhost, same as Willison described. Where it stopped being fun was throughput: without a GPU to lean on, I was getting something in the neighborhood of two to three tokens a second on simple prompts. That’s the kind of speed where you start a request, go refill your coffee, and come back to find it’s maybe a third done. Usable for a background batch job. Not usable for anything where a person is sitting there waiting.

I also hit a smaller version of the same oddity Willison mentioned: restarting the server sometimes changed the throughput noticeably, once by close to 40 percent, with no configuration change on my end. I don’t have a clean explanation for that, and neither did the release notes.

Memory turned out to be the easier problem to reason about. The GGUF itself sits under 6 gigabytes, but once you add the context window and the server’s own overhead, plan on closer to 8 or 9 gigabytes of RAM before you touch a single prompt. My VPS has 16, so there was headroom, but I’ve seen enough OOM-killed processes on smaller boxes to say: check free -h before you start the server, not after it silently dies mid-download.

So is it actually near lossless

Here’s where I have to be honest about the limits of a one-evening test. I did not run a proper benchmark suite against the full-precision version of this model. What I did was throw a handful of the kinds of questions I actually ask models day to day: summarizing a paragraph of documentation, writing a short regex, explaining a stack trace, renaming a batch of variables to something less embarrassing. On those, the quantized model held up better than I expected for something described in bits rather than bytes. It didn’t fall apart or produce garbage. But “didn’t fall apart on five casual prompts” and “near-lossless” are different claims, and I’m not going to pretend I verified the second one just because the first one felt true.

What convinces me the underlying technique deserves attention, rather than dismissal, is the pace of the surrounding field this year. Willison’s own year-in-review post on 2026’s LLM developments makes the case that most of the meaningful progress hasn’t been raw capability, it’s been models crossing a threshold from “often makes mistakes” to “reliable enough to use daily” at a given size and cost. A 9x compression ratio that gets even 80 percent of the way to lossless changes which threshold a given piece of hardware can clear. That’s a more interesting question than whether any single benchmark number is exactly right.

When self-hosting something like this actually makes sense

I already wrote about what this blog’s own infrastructure costs to run, and the honest answer is that a hosted API call is almost always cheaper than the electricity and hardware amortization of running your own inference, unless you’re doing enough volume to change that math or you have a hard requirement that the data never leaves your machine. Ternary quantization doesn’t flip that equation for most people. What it does is lower the floor: a model that used to need a $2,000 GPU to run at a usable speed might now limp along, slowly, on a $5 VPS, or run at a genuinely usable speed on a laptop that couldn’t fit the full-precision version in memory at all.

That’s a narrow use case. If you’re prototyping something that has to work entirely offline, or you’re testing whether a smaller footprint model is “good enough” before committing to a paid deployment, this is worth twenty minutes of your time. If you just want a model that answers questions quickly, pay for the API and skip the custom fork.

There’s also a version of this that has nothing to do with speed. If you’re building something where the data genuinely can’t leave a customer’s premises, a compression scheme that turns a 27 billion parameter model into a file smaller than a typical Blu-ray disc is a real answer to “can we even fit this on the hardware they gave us,” even before you ask how fast it runs. I’ve had exactly one client conversation this year where that constraint was non-negotiable, and it wasn’t fast, but it worked, which mattered more than tokens per second in that specific case. Most of the client infrastructure work I take on through my consulting practice still ends up on a hosted API for exactly this reason.

What I’d actually try this week

If you want to reproduce this without the CPU-only pain, grab whichever prebuilt binary matches your own hardware from Prism’s release page, not necessarily the Ubuntu CPU one I used, and time a handful of real prompts against whatever you’re currently paying for. Write down the tokens per second and the answer quality side by side. That fifteen-minute comparison will tell you more about whether “near-lossless” and “9x smaller” apply to your workload than any number in a release announcement, mine included.