modelserver is a self-hosted server for local GGUF and ONNX models. One Go binary
provides the HTTP API, an embedded web dashboard, a SQLite model registry, and model
lifecycle management.
- GGUF models run through local
llama-serverchild processes. - ONNX models run in-process through ONNX Runtime.
- Multiple models can be registered and loaded at the same time, subject to available RAM/VRAM and per-model port configuration.
The build requires Go 1.25+. Node 20+ is required when rebuilding the dashboard. Build the dashboard before building the Go binary so the UI is embedded:
git clone https://github.com/donvito/open-models-server.git
cd open-models-server
bash build.shStart the API and dashboard together:
./bin/modelserver serveOn Windows PowerShell or Command Prompt, build with ./build.cmd instead. Git Bash
can use bash build.sh as shown above. Both produce bin/modelserver.exe on Windows;
start it with ./bin/modelserver.exe serve. Stop a running server before rebuilding.
The dashboard and API are available at
http://127.0.0.1:9090. Use --headless for an API-only server.
If a runtime is installed outside the normal search paths, pass its executable or shared library explicitly:
./bin/modelserver serve \
--llama-binary "C:/path/to/llama-server.exe" \
--onnx-library "C:/path/to/onnxruntime.dll"The two runtime options are independent; omit either one when that runtime is not needed.
Register and load a GGUF model:
./bin/modelserver models add ./models/gemma-2b-it.gguf --name gemma --loadRegister and load an ONNX text classifier:
./bin/modelserver models add ./models/sst2-onnx \
--name sentiment --runtime onnx --task classification --loadUse the dashboard, or manage models from the CLI:
./bin/modelserver models list
./bin/modelserver models status gemma
./bin/modelserver models unload gemmaGGUF models expose OpenAI-compatible chat and completion endpoints. ONNX models use the generic prediction endpoint:
curl http://127.0.0.1:9090/v1/models/sentiment/predict \
-H 'Content-Type: application/json' \
-d '{"input":"I loved this movie"}'Configuration precedence is:
CLI flags > MODELSERVER_* environment variables > modelserver.yaml > defaults
Copy modelserver.example.yaml when you need persistent settings. The main settings are the server address and port, database path, model directory, runtime paths, internal llama.cpp port range, and API keys.
- Setup and startup — build, run, Windows paths, embedded UI, and development mode.
- GGUF and llama.cpp — install
llama-server, load GGUF models, vision models, per-model settings, and troubleshooting. - ONNX Runtime — install the native runtime, use ONNX model layouts, supported tasks, and a working classifier example.
- Configuration — configuration file, environment variables, ports, authentication, and network access.
- HTTP API — management, prediction, and OpenAI-compatible endpoints.
- Development — tests, linting, frontend development, and repository layout.
MIT