A safe haven for AI inference
Self-hostable Docker stack: local AI inference through Ollama, SSH workspace, web IDE, and up to ten coding assistants wired to your models.
Built by InferHaven, a one-person research studio for privacy and local-first AI.
Quick Start • Configuration • Bulkhead CLI • Full Guide • Harnesses & Models • Bulkhead Cloud ↗
Note
This project was previously called InferHaven Core; its hosted edition was InferHaven Cloud, now Bulkhead Cloud. The command is now bulkhead (short form: bh). If you have scripts using haven, they still work: haven prints a one-line notice and forwards to bulkhead.
Bulkhead is a self-hostable Docker stack that turns hardware you control into your own AI coding server. It runs Ollama for local models and gives you a pre-configured workspace over SSH and a web IDE, with your own models and up to ten coding assistants already wired to them. No per-token meter. Free to self-host.
It is not another coding assistant competing for your editor. It is the box the assistants run in.
- Local inference: Ollama with an OpenAI-compatible API and any open-weight model. NVIDIA and AMD GPU support out of the box.
- Cloud models: use a popular provider's models instead of, or alongside, your local ones.
- Privacy: with local models, your code and the model weights stay on your own hardware; with a cloud model provider, your prompts go to that provider.
- Terminal-first workspace: SSH and mosh, tmux sessions that save themselves, zsh with Starship, neovim, ripgrep, fzf, supercronic, lazygit, git-delta, direnv, zoxide, eza, mise, atuin and tmate.
- Web IDE: VS Code in the browser through code-server.
- Coding assistants: Claude Code, OpenCode, Aider, Qwen Code, Amp, Gemini CLI, Goose, Continue CLI, Pi and Avante, each installable and pre-configured from
.env. Seven of them (opencode,aider,qwencode,pi,goose,continue,avante) re-render their config on every model pull. - Security: SSH is key-only; the Ollama and code-server ports are not exposed by default, and all traffic routes through the Caddy reverse proxy, which provides HTTPS.
- Multi-user: provision extra users with their own SSH keys via
.env. - Devcontainer-ready: VS Code Dev Containers, GitHub Codespaces, DevPod, JetBrains Gateway and the headless
@devcontainers/cli. Two flavours: a light Codespaces flavour for CPU-only quick starts, and a full-stack flavour with the web IDE, Caddy and optional GPU passthrough. Nested devcontainers throughbulkhead devcontainer. The Codespaces flavour boots a small model (qwen3:4b) withopencodeandaiderpreinstalled. - Backups:
bulkhead backup configuresets up an rclone remote;bulkhead backup push <remote:path>snapshots your home directory and assistant configs. - Fast rebuilds: warm rebuilds take under 30 seconds with BuildKit cache mounts.
You can, and if you do, you've built the first couple of layers of what Bulkhead ships whole. The DIY path is Ollama, plus a web UI, plus each assistant's config, plus a reverse proxy, HTTPS, SSH, and backups: a weekend to assemble and a maintenance tab that never closes, since every assistant's config drifts the moment you pull a new model.
Bulkhead is those same parts, assembled and kept in tune. docker compose up -d brings the whole stack up, and seven assistants (opencode, aider, qwencode, pi, goose, continue, avante) re-render their config automatically on every model pull. It's still just Docker. The exit door is the same size as the front door.
Just want to try it first? No install needed: click Open in GitHub Codespaces (badge above). It boots the CPU-only flavor with a small model (
qwen3:4b) andopencode+aiderpreinstalled. For real use (GPU, web IDE, your own models), self-host below.
Requirements: Linux, Docker, Docker Compose v2.
See docs/gpu-setup.md for full GPU configuration.
git clone https://github.com/InferHaven/bulkhead.git
cd bulkhead
cp .env.example .env
chmod 600 .env # contains API keys — keep it owner-only
# Edit .env — set CODE_SERVER_PASSWORD, AUTHORIZED_KEYS, and any API keys.
# Edit docker-compose.yml - enable required GPU settings, Nvidia + AMD (vulkan or RocM) supported
docker compose up -dOnce running:
| Access | Method |
|---|---|
| SSH | ssh -p 2222 haven@localhost |
| Web IDE | http://localhost |
| Ollama / OpenAI API | http://localhost / http://localhost/v1/ |
Ollama and code-server exposed ports are commented out in
docker-compose.ymlby default. All traffic routes through Caddy. To expose them directly, uncomment theports:blocks for theollamaandcode-serverservices (routes around Caddy security).
For a step-by-step walkthrough see docs/quickstart.md.
All configuration lives in .env (copy from .env.example).
| Variable | Default | Description |
|---|---|---|
DEFAULT_MODEL |
qwen3.5:9b |
Model pulled on first startup |
CODE_SERVER_PASSWORD |
inferhaven |
Web IDE password (change this) |
AUTHORIZED_KEYS |
(empty) | SSH public key(s) |
DOMAIN |
localhost |
Domain for auto-HTTPS via Caddy |
SSH_PORT |
2222 |
SSH port |
OLLAMA_PORT |
11434 |
Ollama API port |
INSTALL_ASSISTANTS |
(empty) | Harnesses to auto-install on first boot |
HAVEN_CTX |
32768 |
Context window target for auto-tune on pull. Use 16384 on memory-constrained hardware |
HAVEN_AUTO_TUNE |
1 |
Auto-run bulkhead tune after every pull / pullback and for DEFAULT_MODEL on boot. Set 0 to disable |
HAVEN_FORCE_FAMILY |
(empty) | Bypass family detection in bulkhead tune. Values: qwen3 qwen25 llama3 deepseek mistral phi4 codellama gemma. Use for custom finetunes you know are template-compatible |
GOOSE_CTX_LIMIT |
32768 |
Maximum context passed to Goose sessions. Caps the KV-cache budget regardless of model tuning; reduce to 16384 if you see 30 s stream stalls on constrained hardware |
ANTHROPIC_API_KEY |
(empty) | For Claude Code, Aider (claude backend), Amp |
OPENAI_API_KEY |
(empty) | For OpenCode, Aider (openai backend) |
GEMINI_API_KEY |
(empty) | For Gemini CLI |
OPENROUTER_API_KEY |
(empty) | For OpenRouter-compatible tools |
GITHUB_TOKEN |
(empty) | GitHub CLI auth |
CLAUDE_CODE_DISABLE_TELEMETRY |
true |
Disable Claude Code telemetry |
INSTALL_STARSHIP |
1 |
Starship prompt (0 = keep Oh My Zsh robbyrussell) |
MOSH_PORTS |
60000-60010 |
Host UDP range for mosh. Set empty to skip host mapping (mosh still works internally) |
HAVEN_EXTRA_USERS |
(empty) | Comma-separated extra users provisioned alongside haven (e.g. alice,bob) |
HAVEN_EXTRA_USERS_SUDO |
(empty) | Subset of HAVEN_EXTRA_USERS granted passwordless sudo |
AUTHORIZED_KEYS_<USER> |
(empty) | Per-extra-user SSH key. USER is uppercase (e.g. AUTHORIZED_KEYS_ALICE) |
DOTFILES_REPO |
(empty) | Git URL cloned to ~/.dotfiles on first boot; runs install.sh once |
Set INSTALL_ASSISTANTS and any API keys in .env before the first start. Harnesses are installed in the background. SSH is available immediately and they are ready within a minute or two.
INSTALL_ASSISTANTS=claudecode,opencode,aider
ANTHROPIC_API_KEY=sk-ant-...Supported Harnesses: claudecode, opencode, aider, qwencode, amp, gemini, pi, goose, continue, avante
When opencode, aider, qwencode, pi, goose, continue, or avante is included, local Ollama models are auto-configured and kept in sync: every bulkhead pull, bulkhead tune, and bulkhead remove updates all harness configs immediately. Most harnesses use an internal sentinel so Bulkhead never touches user-customised configs; continue syncs whenever cn is installed (opt out: touch ~/.continue/.no-autosync).
See docs/harnesses.md for opt-out, per-project override instructions, and for per-harness setup details and recommended models.
Single key (no quotes needed):
AUTHORIZED_KEYS=ssh-ed25519 AAAA... user@hostMultiple keys, wrap in double quotes with a real newline:
AUTHORIZED_KEYS="ssh-ed25519 AAAA...key1 user@host
ssh-ed25519 AAAA...key2 user2@host"Bulkhead's CLI, bulkhead, works in two contexts:
From the host (repo directory) manages Docker services:
./scripts/bulkhead up # start all services
./scripts/bulkhead down # stop all services
./scripts/bulkhead restart # restart all services
./scripts/bulkhead logs # stream logs (all services)
./scripts/bulkhead logs ollama # stream logs for a specific service
./scripts/bulkhead update # pull latest images and restart
./scripts/bulkhead reset # remove all data (careful)
./scripts/bulkhead status # service status
./scripts/bulkhead doctor # diagnose the host environment
./scripts/bulkhead ssh-key "<pubkey>" # add an SSH public key
./scripts/bulkhead ssh # show SSH connection command
./scripts/bulkhead ide # show web IDE URLInside the workspace (after SSH-ing in): full feature set:
# Models
bulkhead models # list downloaded models
bulkhead pull <model> # download a model (foreground, with live progress)
bulkhead pullback <model> # download a model in the background — keep working
bulkhead pullback status # show all background download progress
bulkhead pullback cancel <model> # cancel a background download
bulkhead remove <model> # remove a model (updates harness configs)
bulkhead show <model> # model details: params, template, system prompt
bulkhead show <model> --modelfile # print raw Modelfile
bulkhead ps # models currently loaded in GPU/RAM
bulkhead unload <model> # force-unload from GPU/RAM
bulkhead cp <src> <dest> # copy / rename a model
bulkhead chat [model] # interactive chat (defaults to DEFAULT_MODEL)
bulkhead run <model> # same as chat — TTY interactive session
bulkhead run <model> "your prompt" # one-shot: print response and exit (scriptable)
echo "prompt" | bulkhead run <model> # pipe stdin into model
bulkhead bench [model] # benchmark tokens/sec (--tokens N --prompt ".." --runs K --json)
# ollama.com account
bulkhead push <model> # push a model to ollama.com
bulkhead signin # authenticate with ollama.com
bulkhead signout # sign out of ollama.com
# Model parameters (instant — no re-download)
bulkhead params <model> # show current parameters
bulkhead params <model> set num_ctx 32768 # context window size
bulkhead params <model> set temperature 0.3 # creativity (0.0–2.0)
bulkhead params <model> set num_predict 4096 # max tokens (-1 = unlimited)
bulkhead params <model> set top_p 0.9 # nucleus sampling threshold
bulkhead params <model> set top_k 40 # top-k candidates per step
bulkhead params <model> set repeat_penalty 1.1 # penalise repeated tokens
bulkhead params <model> reset # reset all params to defaults
# Model tuning (no re-download — sets num_ctx, stop tokens, template per family)
bulkhead tune <model> # optimise for harness use
# Families: qwen2.5 · qwen3 · llama3 · deepseek · mistral · phi4 · codellama
# Harnesses
bulkhead harness # show installed harnesses + OpenCode config summary
bulkhead claude # launch Claude Code with a local Ollama model
bulkhead aider # launch Aider with a local Ollama model
bulkhead goose # launch Goose with a local Ollama model
bulkhead qwen # launch Qwen Code with a local Ollama model
# Status & diagnostics
bulkhead status # service status + model count
bulkhead logs [service] # stream service logs
bulkhead doctor # diagnose the container environment (incl. P1/P2 binaries, swap, cgroup)
bulkhead service <name> <action> # docker compose wrapper: status / restart / stop / start / logs
bulkhead limits # show container cgroup limits vs host capacity (memory/CPU/swap)
bulkhead gpu-info # canonical GPU readout from the metrics-server (driver, util, VRAM)
# Pair-programming + backup (P2)
bulkhead tmate # start a backgrounded tmate session — prints SSH/web URLs
bulkhead tmate status # print current tmate URLs + uptime
bulkhead tmate fg # attach to the active tmate session
bulkhead tmate kill # tear down the tmate session
bulkhead backup configure # interactive rclone remote setup wizard
bulkhead backup status # show local backup paths + configured rclone remotes
bulkhead backup status <remote:path> # also show size + top-level contents of that remote
bulkhead backup push <remote:path> # snapshot ~/.haven + ~/.config + ~/.continue to an rclone remote
bulkhead backup pull <remote:path> # restore from an rclone remote
# Tool-config sync (re-render coding-assistant configs from the live model list)
bulkhead sync # re-sync all 7 supported tools in parallel
bulkhead sync <tool> # opencode | aider | qwencode | pi | goose | continue | avante
bulkhead sync list # list supported tools
# Tmux workspace (sessions auto-save every 15 min, fully restored after restarts)
bulkhead tmux # attach to the always-running 'Haven' session
bulkhead tmux attach [name] # attach to a session (default: Haven)
bulkhead tmux ls # list all active sessions
bulkhead tmux new <name> # create and attach to a new named session
bulkhead tmux kill <name> # kill a session
bulkhead tmux save # manually save sessions to disk
bulkhead tmux restore # manually restore from last save
bulkhead tmux plugin list # list installed plugins
bulkhead tmux plugin install # install plugins from ~/.tmux.conf
bulkhead tmux plugin update # update all plugins
bulkhead tmux plugin bootstrap # reinstall all plugins from scratch
bulkhead tmux help # full subcommand reference
# Packages (persist across container restarts)
bulkhead apt install <pkg...> # install and track apt packages
bulkhead apt remove <pkg...> # stop tracking a package
bulkhead apt list # show tracked packages
bulkhead apt update # refresh package lists
bulkhead apt upgrade # upgrade all tracked packages
# SSH / IDE
bulkhead ssh-key "<pubkey>" # add an SSH public key
bulkhead ssh # show SSH connection command
bulkhead ide # show web IDE URL
bulkhead help # show all commandsFor workspace-specific features (model tuning, background downloads, Starship prompt, persistent packages, and status bar alerts), see docs/workspace.md.
Four Docker services in a bridge network:
| Service | Image | Purpose |
|---|---|---|
ollama |
ollama/ollama |
AI inference, OpenAI-compatible API on :11434 |
workspace |
Custom build | SSH terminal (:2222) + bulkhead CLI + harnesses |
code-server |
linuxserver/code-server |
VS Code in Browser |
caddy |
caddy:2-alpine |
Reverse proxy, auto-HTTPS |
┌──────────────────────────────────────────────┐
│ Bulkhead │
│ │
│ ┌───────────┐ ┌──────────┐ ┌───────────┐ │
│ │ Workspace │ │ Ollama │ │Code Server│ │
│ │ (SSH/tmux)│ │ (AI) │ │ (IDE) │ │
│ └────┬──────┘ └────┬─────┘ └─────┬─────┘ │
│ └──────────────┴──────────────┘ │
│ │ │
│ ┌─────────┴──────────┐ │
│ │ Caddy │ │
│ │ (Proxy + auto-TLS) │ │
│ └────────────────────┘ │
└──────────────────────────────────────────────┘
:2222 SSH :80/:443 HTTP/HTTPS :11434 Ollama APICaddy routes: /status → status page, /ide* → code-server, /api/* and /v1/* → Ollama, default → code-server.
Contributions are welcome. Please read CONTRIBUTING.md before submitting a pull request.
Found a vulnerability? Don't open a public issue. Report it privately per our Security Policy (email lookout@inferhaven.com). For sensitive reports, encrypt to our OpenPGP key (inferhaven_pub.asc, also at https://inferhaven.com/pgpkey.asc):
OpenPGP fingerprint:
4992 80D5 D75E 3A4F 837C 6A68 85D8 E097 0D05 CEC0
For deployment hardening (access control, network exposure, TLS, secrets, and the stack's intended trust boundaries), see docs/security.md.
AI assistants are part of how Bulkhead is built. We use them to accelerate the work: drafting code, refactoring, generating tests, and writing documentation.
What doesn't change: every change is reviewed, understood, and manually tested by a human before it merges. Bulkhead is owned and maintained by its human author(s). AI is a tool we use, not the author.
Bulkhead is licensed under the Functional Source License 1.1 with Apache 2.0 Future License (FSL-1.1-Apache-2.0).
What this means in practice:
- ✅ You can use, modify, and self-host Bulkhead for any purpose: personal, commercial, internal, or research.
- ✅ Enterprises can deploy it on their own infrastructure, integrate it with internal tools, and modify it as needed.
- ✅ Consultants and integrators can offer professional services around it.
- ❌ You cannot offer a commercial managed-hosting service that competes with Bulkhead Cloud (until each version's two-year window expires).
- 🔄 Two years after each version's release, that version automatically converts to the Apache License 2.0, a fully permissive open-source license with no restrictions.
See the LICENSE file for full terms, docs/licensing.md for a plain-language explainer, and fsl.software for background on the license.
Bulkhead · A safe haven for AI inference · Built by InferHaven
Website •
Discord •
Twitter

