The largest open models need about 240 GB of GPU memory even at 4-bit precision. Almost nobody has that in one machine. A lot of people have one gaming GPU with 16 or 24 GB, and it sits idle most of the day.
So I wanted to know: can a group of ordinary computers run one very large model together, each holding a small piece? Not in a datacenter with fast private links, but the way home computers actually connect, behind routers that do not accept incoming connections.
I have been building an open-source project called Sangama to find out. This post is the result of two days of testing it on Qwen3.5-397B, one of the largest open models. The short version: it works, the GPUs are almost idle, and nearly everything interesting turned out to be about the network.
The Setup
A language model is a stack of layers, and a token has to pass through them in order. That makes it easy to cut: give the first three layers to one machine, the next three to another, and pass the intermediate numbers along. This is called pipeline parallelism.
- Model: Qwen3.5-397B-A17B at 4-bit precision, 60 layers, cut into 20 stages of 3 layers.
- Machines: 20 rented GPUs (mostly RTX 4090 and 5090), each capped at 16 GB. No machine held more than 3 of the 60 layers, and each downloaded only its own 12 GB.
- Engine: llama.cpp, patched so a process can load and run just a range of layers.
- Network: an invitation-only, encrypted peer-to-peer mesh. I forced every connection through one relay, because that is what machines behind home routers have to do.
client -> relay -> machine 1 -> relay -> machine 2 -> ... -> relay -> machine 20
|
client <---------------------- next token -------------------------------+
The First Run: 0.6 Tokens per Second
It worked on the first full run, and it was slow: 0.6 tokens per second, with over seven seconds before the first token.
The surprise was where the time went. All 20 GPUs together spent 21 milliseconds computing each token. The model only activates 17 billion of its 397 billion parameters per token, so the arithmetic is cheap. The other 1.5 seconds or so, 98.6% of the total, was the network. Each GPU drew about 63 watts and showed close to 0% utilisation.
That reframes the whole problem. A 397B model on 20 small machines is not limited by computing power or by bandwidth (each machine sent about 0.1 Mbit/s). It is limited by how long a message takes to cross 21 hops.
Six Rounds to 5.5 Tokens per Second
Over the first day I made the same fleet nine times faster without touching the model.
| Change | What it did |
|---|---|
| Pass the result forward | Each machine used to wait for the rest of the chain before replying, so every token crossed every hop twice. Now each one hands off and replies at once: 2.2×. |
| Half-size numbers | Send the intermediate values in 16-bit instead of 32-bit format. Another 1.4 to 2×. |
| Prompt in pieces | Send a long prompt in 64-token pieces, so machines work on different pieces at the same time. First token 2.7× sooner on a long prompt. |
| Closer machines | Measure each machine's distance to the relay and replace the farthest. Decode speed more than doubled. |
The end of day one: 5.0 to 5.5 tokens per second, first token in 1.3 to 6.2 seconds. Each token now took about 190 ms, still almost all of it network.
The bugs were more instructive than the optimisations. The relay library's default rate limits silently disconnected most of my machines within minutes, because they all appeared to come from one address. And every machine was asking every other machine what it held, through the relay: 170 connections where the route needed 39, which made each request take about a second.
Day Two: The Fleet Disappeared
I came back the next morning to a dead relay. The provider I had rented the fastest machines from lists them for a few hours at a time, and when a listing ends the machine stops and cannot be restarted. Half the fleet went within two hours.
I rebuilt it on machines with weeks left on their listings. Two lessons from that, neither about AI:
- Listed locations can be wrong. Two machines listed in California added 220 ms and over 400 ms per hop. From one of them, connecting to the relay looked instant, because a proxy on that host answered. Timing a real reply from the relay showed the true delay.
- Reliable machines were further apart. The new fleet ran one request at 1.25 tokens per second instead of 5.4. Everything below compares settings on this slower fleet.
Many People at Once
If each GPU is busy for one millisecond per token and idle for the other 800, it can serve other requests in between. So I taught each machine to keep up to 96 separate sessions in memory and let requests take turns.
| Requests at once | Total tokens per second | Each request |
|---|---|---|
| 1 | 1.3 | 1.25 |
| 16 | 30 | 1.88 |
| 32 | 60 | 1.87 |
| 51 | 90.9 | 1.79 |
| 96 | 67.6 | 0.70 |
Up to about 50 requests, each one kept its full speed, so the total grew with the number of users: about 70 times the throughput of a single request, on the same hardware. All 96 requests in the largest run produced exactly the same tokens as when run alone. Past 50, each request slowed and the total stopped rising. Processors and GPUs were still mostly idle, so I have not found that limit yet.
Getting there meant removing limits I had set for one user. My own network layer allowed 4 requests in flight per link, then 16 calls per machine, then 64 sessions. The model runtime kept memory for all 60 layers in every session, about 306 MB each, although each machine runs only 3; fixing that brought it to under 20 MB.
Guessing Ahead
One user is still stuck at one token per trip through 20 machines. The way around that is to guess several tokens ahead and check all the guesses in a single trip.
My first attempt guessed by copying earlier text. The model agreed with 5 to 26% of those guesses, and undoing the wrong ones on 20 machines cost more than the right ones saved. It was correct and slower.
Qwen3.5 ships with a small extra layer trained to predict the token after next. The published 4-bit file leaves it out, so I converted it from the original release and loaded it on the last machine. After each token, that machine guesses the next few and sends them back with the result.
- The guesses are good: 41 of 42 accepted with two guesses per step on a coding prompt.
- Up to 5.3 tokens per trip instead of one.
- Up to 2.2× faster for one request (1.25 to 2.69 tokens per second on code; 1.2 to 1.4× on prose).
The gap between 5.3 tokens per trip and 2.2× is the cost of checking. A trip that checks guesses took 1.5 to 2 seconds against 0.8, because every machine saves its state first and rewinds when a guess is rejected. That is the next thing to make cheap.
With many users and guesses together, 48 requests gave 83 tokens per second against 51 without. But the gain was not consistent between runs, and I would want more repeats before quoting it.
The Thing I Got Wrong
After day one I wrote that every configuration produced exactly the same tokens. On day two, runs with guesses were fluent but often differed from plain runs at close word choices: "iteratively." in one, "using an iterative approach." in the other.
Guessing was not the cause. Plain decoding gave different tokens at the same place when I only changed how many prompt tokens were sent at a time. Processing several positions together changes the arithmetic very slightly, and in this model that can tip a near-tie. A run repeated with the same settings always gave the same tokens. So the honest claim is narrower: the output is repeatable for a given setting, and concurrent requests match the same request run alone. I corrected the earlier write-up.
What I Took From It
- The idea holds. A model no single machine can load ran on 20 machines with 16 GB each, through a relay, and gave coherent answers.
- It is a throughput machine, not a fast one. One user gets a few tokens per second. Many users share the same machines at almost no cost to each other.
- Geography is the speed limit. The same software ran at 5.4 tokens per second on machines near the relay and 1.25 on machines spread over three states.
- Every limit was found by pushing until something refused. None of the concurrency limits showed up in a design review. They showed up at 5, 17 and 49 requests.
- Rented hardware is its own test. Hosts vanished, lied about where they were, or could not reach the internet. A network for home computers has to expect all of that.
The whole fleet cost about $13 an hour to rent, and I tore it down after the last run. The code is open source under the MIT licence, and the full results, with raw data, are on the project site: the first day and the second.
Next is making rejected guesses cheap, finding what stops the total at about 90 tokens per second, and placing machines by measured delay automatically. If you have a spare GPU and want to try being one of the twenty, the project is on GitHub. And if you liked the small-machines angle, my three-node home lab on refurbished ThinkCentres is an earlier experiment in the same spirit.
~ Comments & Discussion ~
Have thoughts on this post? Join the discussion below! Comments are powered by Disqus.