LLM inference for a fleet of machines. One node per machine serves every model on it behind one OpenAI-compatible API. Nodes find each other and form a mesh; any node is an entry point, and a request goes to a free slot wherever one exists.
I had 50 Radeon VII cards sitting in servers in my mother-in-law's garage (who doesn't?) and wanted to do something useful with them. viiwork turns a pile of aging-but-capable GPUs into a practical inference cluster.
That fleet is still the reference deployment, but viiwork runs on anything from one gaming GPU to racks of cards — AMD, NVIDIA or an Apple Silicon Mac, all in one mesh. It doesn't do inference itself: it drives existing engines — llama.cpp, vLLM and FreeToken — and turns them into one fleet. Each node's GPU resources are available from any endpoint.
| One API per machine | Every model on the box on port 8086 → configuration |
| Three engines | llama.cpp, vLLM, FreeToken side by side → models |
| Self-assembling mesh | No peer lists; routes to whoever has a free slot → mesh |
| Aliases | Stable names like stable-coder, switched once for the fleet → aliases |
| Dashboards | Models, jobs, backends, prompts, power → dashboards |
viiwork top |
The mesh live in a terminal → operations |
| Power and cost | Wattage, spot price, per-model kWh → power and energy |
| Pipelines | Several LLM steps as one model name → pipelines |
| MCP server | The cluster as tools for assistants → MCP |
| Autodiscovery | Coding clients find models and context on their own → autodiscovery |
You need GGUF model files and one of:
| Hardware | What the wizard does |
|---|---|
| Radeon VII / MI50 / MI60 (gfx906), Linux + Docker | Writes the config; you build and start the image (below) |
| NVIDIA, Linux + Docker + NVIDIA Container Toolkit | Writes the config and starts the node |
| Apple Silicon Mac | Writes the config and starts the node natively |
Full requirements: setup.
v=vX.Y.Z # newest from https://github.com/janit/viiwork/releases
os=linux_amd64 # or linux_arm64, darwin_arm64
base=https://github.com/janit/viiwork/releases/download/$v
curl -fLO "$base/viiwork_${v}_${os}.tar.gz" -fLO "$base/SHA256SUMS"
sha256sum --check --ignore-missing SHA256SUMS # Mac: shasum -a 256 --check ...
tar xzf "viiwork_${v}_${os}.tar.gz" && cd "viiwork_${v}_${os}"
sudo ./viiwork init # Mac: ./viiwork initThe wizard finds your GPUs and models, proposes a layout, and shows every file before writing anything. → setup · verifying signatures
On gfx906, build the image from the same release and start it:
git clone https://github.com/janit/viiwork && cd viiwork && git checkout "$v"
make docker
sudo cp configs/docker-compose.v2.example.yaml /etc/viiwork/docker-compose.yaml
# set the models mount in that file, then:
sudo docker compose -f /etc/viiwork/docker-compose.yaml -p viiwork up -d→ when it writes the config only
Add the next machine: print a join code on any node, run the wizard on the new one, paste the code. The code is the mesh secret.
sudo sh -c 'set -a; . /etc/viiwork/mesh.env; /usr/local/bin/viiwork join-code' # Linux
~/.local/bin/viiwork join-code # MacTest it:
curl http://localhost:8086/v1/models
curl http://localhost:8086/v1/chat/completions -H "Content-Type: application/json" \
-d '{"model":"<name from /v1/models>","messages":[{"role":"user","content":"Hello"}]}'Day to day: viiwork top, viiwork update (updating),
viiwork stop / start, viiwork uninstall (keeps your models).
Skip the wizard: copy viiwork.yaml.example to /etc/viiwork/viiwork.yaml,
put a VIIWORK_MESH_SECRET in /etc/viiwork/mesh.env (or set mesh.open: true),
then build and start the image as above.
→ configuration
One file per machine. Each model has a name, its GPUs and its slots:
models:
- name: Qwen3.8-27B
engine: llamacpp
path: /models/Qwen3.8-27B-UD-Q4_K_XL.gguf
gpus: [0, 1]
gpus_per_backend: 2 # one backend across both cards
context: 49152 # tokens PER SLOT
parallel: 2- A GPU belongs to one model
SIGHUPreloads: added models start, removed drain, changed restart- Validated at startup and reload, errors name the field
- Ports: 8086 tcp (API, dashboards), 7946 tcp+udp (gossip)
- Discovery: Tailscale (default) or LAN mDNS; optional seeds
- Secured or open: set
VIIWORK_MESH_SECRET, or declaremesh.open: true - Routing: local free slot → member with most free slots → queue (20 s)
| Page | What it is |
|---|---|
/ |
This node |
/mesh |
The whole cluster, from any node |
/chat |
Chat UI (?model=&host=) |
/prompt |
One request's prompt and output |
| Engine | Runs | Shape |
|---|---|---|
llamacpp |
llama-server (reference) |
CPU or GPU |
vllm |
vllm serve |
tensor-parallel across its cards |
freetoken |
ft serve |
one card per process, MoE experts in host RAM |
All three can run on one node. Adding an engine is one package plus one line. → docs/adding-an-engine.md
- gfx906 rule: a dense model that doesn't fit one card costs ~3× throughput
- FreeToken exception: big sparse MoE models on a single card, at the cost of load time
What the Radeon VII hosts serve today:
| Model | Hosts | Context per slot | Role |
|---|---|---|---|
Qwen3.8-27B |
2 | 49152 | coder and prose |
granite-4.2-8b |
2 | 16384 | fast utility |
translategemma-27b-it |
3 | 4096 | translation |
gemma-4-31B-it |
3 | 6144 | prose |
Ornith-1.5-35B-A3B |
1 | 262144 | long context |
Throughput, validated configs, failed bring-ups, tuning rules → docs/models.md
viiwork authenticates nothing. Reachability is the authorization; run it on a tailnet. A CORS allowlist (default: your own tailnet) stops browser pages driving the fleet. → docs/security.md
| Endpoint | What |
|---|---|
/v1/chat/completions, /v1/completions, /v1/embeddings |
Inference, routed by model (?host= pins a node) |
/v1/models |
Every model in the mesh, with context length |
/health |
Node health |
/v1/capacity, /v1/fleet/capacity |
Slots and load, this node or the whole mesh |
/v1/status, /v1/cluster |
Node and member state |
/api.json, /v1/model/info |
Discovery for OpenCode and Roo Code |
| Image | Engine | Make target |
|---|---|---|
viiwork |
llamacpp on ROCm / gfx906 |
make docker |
viiwork-vllm |
vllm |
make docker-vllm |
viiwork-freetoken |
freetoken |
make docker-freetoken |
Releases also ship signed static binaries for linux/amd64, linux/arm64 and darwin/arm64. → BUILDS.md · docs/releases.md
| Document | About |
|---|---|
| setup.md | First-run wizard, join codes, updating, uninstalling |
| configuration.md | Every config key |
| mesh.md | Discovery, security modes, routing, aliases |
| models.md | Measured catalogue and gfx906 tuning |
| dashboards.md | The web pages |
| api-integration.md | Full API reference |
| autodiscovery.md | Clients discovering the fleet |
| consuming-fleet-capacity.md | Clients reacting to capacity |
| operations.md | Scripts, acceptance checks, viiwork top, MCP |
| power-and-energy.md | Power, cost, energy store, IPMI |
| security.md | Trust model and CORS |
| releases.md | Verifying, updating, publishing |
| macos.md | A Mac as a node |
| thinking-models.md | Controlling reasoning output |
| adding-an-engine.md | The engine contract |
| energy-store-format.md | Energy file format |
| tensor-split-design.md | Multi-GPU backends |
| migrating-to-v2.md | v1 → v2 |
| BUILDS.md | Images and pins |
make build # binary
make test # unit tests
make docker # ROCm image
go test -tags=integration ./mesh/... ./internal/proxy/ ./internal/alias/ ./internal/node/ # multi-node, no GPU
go test -bench=. -benchmem ./internal/proxy ./internal/route # compare allocs, not timeGo 1.27.1. Dependencies: yaml.v3, memberlist, mdns, x/term.
