Sitelet https://github.com/plotarmordev/thegrill
Skip to content
plotarmordevPublic

About

Experimental Rust CLI for language-model evaluation evidence and bounded serving-performance comparisons.

Resources

Contributing

Security policy

Stars

4 stars

Watchers

0 watching

Forks

Repository files navigation

The Grill: an open charcoal barbecue

The Grill

Measure how fast a serving setup runs an LLM, then compare setups.

A serving setup is the engine plus its settings, such as vLLM with a quantized model at a given context length. Run the same test on each setup, then compare the saved results.

I want to... Use What you get
Check whether a serving setup is faster Speed benchmarks
grill-perf Ready to try
Time per request group, combined tokens per second, comparison of saved runs
Check whether a model answers tasks correctly Quality evaluation
grill Work in progress
Graded answers, with wrong, refused, cut-off, and missing results kept separate

Both tools run on your machine and connect to a server you already run. No project account or results upload is required. This is experimental software.

Works on

Platform How to install What works
Linux x86_64 and aarch64 (Ubuntu 24.04) Release archive, no Rust needed Everything in grill-perf (opt-in Python producers need their own runtime)
Apple Silicon macOS (11+) Release archive, no Rust needed; unsigned, not notarized Serving speed and deployment comparison; macOS memory observation. Linux resource and external-program collectors are not available

grill-perf talks to any OpenAI-compatible Chat Completions server that reports streaming token usage. It has been run against vLLM, TensorFold (on NVIDIA and on Apple Silicon) and oMLX. It never starts, configures or restarts your server.

The latest pre-release is v0.6.0. Earlier releases stay available; an older binary does not contain newer features.

Quick start

Download and verify the release archive for your platform as described in Install and first capture, then set GRILL_PERF to its bin/grill-perf. On a Mac, download with curl as shown there; a file saved by a browser is blocked by Gatekeeper until you run xattr -d com.apple.quarantine on it.

Then describe your server once and capture a baseline next to it:

mkdir -p results
jq -n --arg m "your/model@revision" --arg r "engine and version" \
  --arg h "device" --arg s "serving flags" \
  '{model_revision:$m, runtime:$r, hardware:$h, settings:$s}' > serving.json
"$GRILL_PERF" baseline --workload portable-v1 --client-placement same-host \
  --endpoint http://127.0.0.1:8000/v1/chat/completions --local-http \
  --model your-model --deployment serving.json --out results/before

Change your server yourself, then check and compare the saved runs. The full walk-through, including an unchanged control run, is in Install and first capture.

Example: the same model on a Mac and on a DGX Spark

Qwen3.8-Flash-Next, the same 4-bit MLX checkpoint (Vontra/...-MLX-4bit-MTP@dadefa80), served by TensorFold 0.6.0 on each machine, measured with the same grill-perf source build, client on the serving host (A/B/A2 deployment comparison):

DGX Spark (GB10) Mac Studio (M5 Ultra)
portable-v1, one request at a time 118 tok/s 246 tok/s
Verdict MEASURED FASTER, +108.7% (95% range +107.9% to +109.6%)
Repeat of the Spark run (drift control) -0.2%
sparkDash decode, 4 requests at once (descriptive) 350 tok/s total 354 tok/s total
sparkDash prefill, 4K-32K (descriptive) 2,340-2,450 tok/s 2,870-2,960 tok/s

This compares whole deployments, not chips: the Mac's TensorFold build rejected two settings the Spark used (int8 KV cache and an MTP confidence threshold), and the counting prompt favours speculative decoding. Verdicts cover the built-in single-request workload; selected workloads such as the sparkDash copies are descriptive side-by-side results (the Mac's sparkDash runs completed 3 of 8 rounds).

Speed benchmarks

Use grill-perf against a server you already run. check reuses the baseline's endpoint, model, workload and credential-variable name. Deployment details are what you declare, not something the tool verifies.

Workloads. The default is a short structured single-request (C1) check; portable-v1 is the same idea for servers that do not support exact-length controls, such as MLX servers. Selected workloads, all descriptive, include:

  • sparkdash-decode-portable-v2 and sparkdash-prefill-portable-v1: the sparkDash tests, matching sparkDash 1.8.7;
  • prefill-prose-portable-v1: the same prefill sizes with varied text instead of one repeated word;
  • realistic coding and edit prompts, a C1-C8 concurrency ladder, long-context decode and prefill, and conversation history checks.

See the selection list and recipe embedding. Contributors making performance claims start from the shared claim map and report template.

Results.

Display label JSON result code Meaning
MEASURED FASTER IMPROVED Higher measured throughput between these capture periods
MEASURED SLOWER REGRESSED Lower measured throughput between these periods
COMPLETE - DESCRIPTIVE ONLY DESCRIPTIVE A selected comparison completed; no faster/slower verdict
INCONCLUSIVE INCONCLUSIVE No direction established, or not enough evidence; not the same as "equal"
VERDICT PENDING PENDING A deployment candidate is captured; the verdict needs the unchanged reference run
INVALID INVALID Response, identity or evidence checks failed; the report says why

A direction is not a cause: sequential runs cannot separate your change from time, load or cache effects, and there is no guaranteed precision. Runs have a fixed time and request budget with no automatic retries; reports and raw evidence stay on your machine, and compare re-checks saved runs without network calls.

Full performance guide and measurement limits.

Quality evaluation (WIP)

Use grill to collect model answers and check them against task rules.

This part is a work in progress. The included questions are synthetic examples, not a validated intelligence test. It currently supports text answers, not agents or code execution.

Build this separate source-only CLI on Linux with Rust/Cargo 1.98, a C/C++ toolchain and CMake; it is not shipped in the performance archive:

git clone https://github.com/plotarmordev/thegrill.git
cd thegrill
cargo build -p grill --release --locked
mkdir -p results

Start your server separately. Replace the endpoint and model selector below. For remote use, select HTTPS; local HTTP requires a literal loopback address. If authentication is required, supply the credential independently in your environment and add --auth-env NAME to run, never the key itself.

target/release/grill run examples/synthetic-pack.json \
  --endpoint http://127.0.0.1:8000/v1/chat/completions \
  --local-http --model your-model --token-cap 4096 --stream \
  --out results/answers

target/release/grill inspect results/answers --json
target/release/grill regrade results/answers --out results/answers-regraded

Inspection and regrading use saved files. They do not call the model again. Wrong answers, refusals, cut-off responses, and missing results stay separate.

Task formats and quality evaluation guide

Keep your results safe

Use a new output directory for each run. Results contain prompts and model responses, so review them before sharing.

Contributing · Code organization · Security · Apache 2.0 license

Apache License 2.0 covers this project's code and documentation; releases up to v0.5.0 were published under MIT. Benchmark data and model weights keep their own licenses.

About

Experimental Rust CLI for language-model evaluation evidence and bounded serving-performance comparisons.

Resources

Contributing

Security policy

Stars

4 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages