Sitelet https://github.com/computesdk/benchmarks/pull/253
Skip to content

Split @benchsdk/client (REST) and @benchsdk/runner (framework); rework benchmark authoring to config + task - #253

Merged
HeyGarrison merged 6 commits into
masterfrom
devin/1785369687-benchsdk-client-runner-split
Jul 30, 2026
Merged

HeyGarrison merged 6 commits into
masterfrom
devin/1785369687-benchsdk-client-runner-split

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Jul 30, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Splits the two SDK packages around a clean boundary and reworks benchmark authoring into a fully declarative model. No published users, so this is a breaking rework rather than a compat shim; scale stays on its own orchestration path.

  • @benchsdk/client = REST transport + worker engine. Removed the authoring factories defineStep/defineTask/defineWorker/defineBench and the runBenchmarkWorker free function. client.runWorker({ task }) takes a raw TaskFunction whose context now exposes three channels: step() (imperative named steps), measure(data), and log(msg, meta?).
  • @benchsdk/runner = the authoring framework (renamed from @benchsdk/cli). A *.bench.ts exports exactly two things — config and task — and never calls the runner itself; the new bench run <file> CLI owns the entrypoint.

Authoring model (the whole surface)

export const config = defineBenchmarkConfig({
  benchmarkSlug, benchmarkName, iterations, concurrency,
  participants,                       // orchestration data lives in config
  onComplete: (outcome) => writeLegacyResults(outcome.participants),  // run-level aggregate hook
});

export const task = defineTask(async ({ participant, step, measure, log }) => {
  const sandbox = await step('create', () => participant.createCompute().sandbox.create()); // return threads to next step
  try {
    const t0 = performance.now();
    await step('exec', () => sandbox.runCommand('node -v'));
    measure({ ttiMs: performance.now() - t0 });   // metrics → platform (step data / record data)
  } finally {
    await step('destroy', () => sandbox.destroy());
  }
});
bench run benchmarks/sandbox/sequential.bench.ts --iterations 100 --concurrency 20 --provider e2b,modal

Key semantics:

  • step() returns are control-flow values, never auto-recorded — you thread live objects (a sandbox, a client) between steps. measure(data) is the explicit metric channel: inside a step() it merges into that step's data; at task top level it merges into the record. log() is human-readable timeline narration.
  • A task with no explicit step() calls is recorded as one implicit 'task' step spanning the task's wall-clock, carrying any task-level measurements. Measurements are preserved even when a task throws.
  • onResult is gone from the author surface (the runner still prints a default per-record line); per-iteration concerns live in the task, aggregate concerns in config.onComplete.

bench run entrypoint

cli.ts imports the module, validates config/task, and calls the now-internal runBenchmark(config, task, argv); NoAvailableParticipantsError maps to a clean exit. The bin re-imports the package by name (marked external in tsup) so the bin and the dynamically-imported *.bench.ts share one @benchsdk/runner instance — otherwise instanceof TaskError / NoAvailableParticipantsError would break across separately-bundled entries.

Migrations & housekeeping

  • All 9 *.bench.ts migrated to config + task; each file's legacy aggregate writer moved into config.onComplete; removed runBenchmark tails, logTti/logAiGateway/logDax, and the bench-exit util. AI Gateway keeps its cold/warm phase parsing and zero-phase skip.
  • Root package.json bench:* scripts now run tsx …/dist/bin.js run <file>.
  • GitHub Actions benchmark workflows repointed to the runner CLI. Because the migrated *.bench.ts no longer self-run, the 7 workflows that invoked npx tsx <file>.bench.ts (sandbox-tti, sandbox-dax, storage, snapshot-fork, browser, browser-throughput, ai-gateway) would have silently run zero benchmarks. They now call npx tsx packages/benchsdk-runner/dist/bin.js run <file> … (dist is produced by the runner's prepare: tsup on pnpm install), and each bench step's load-vault-secrets allowlist gains the org-scoped BENCHMARKS_PLATFORM_API_KEY the runner requires (prod platform URL is the default, so no URL var needed). The other sandbox-* micro-benchmarks use the separate benchmarks/src/run.ts harness and are unaffected.
  • create-bench scaffolds a declarative config+task project on @benchsdk/runner.
  • READMEs, the local-e2e skill, and changesets updated; CLI unit tests + fixtures added.

Testing

  • pnpm -r --filter "./packages/**" build — builds (CJS/ESM/DTS)
  • pnpm typecheck — clean
  • pnpm --filter @benchsdk/client test — 118 passed, 1 skipped
  • pnpm --filter @benchsdk/runner test — 53 passed (incl. new CLI tests)
  • pnpm --filter create-bench test — 1 passed
  • Local platform E2E: bench run smoke against a local benchmarks-platform (Postgres/MinIO/ClickHouse) created a run and reported success (exit 0).
  • CI invocation path verified locally: npx tsx packages/benchsdk-runner/dist/bin.js run <file> for sequential/storage/browser/ai-gateway all exit 0 via the "no participants" path with no creds; pnpm install --frozen-lockfile rebuilds dist/bin.js via prepare.

Tests were updated to match the intentionally-changed public API, not to paper over failures.

Note: the runner requires benchmarks-platform at origin/main (org-scoped API-key auth + organizationSlug returned from createRun); confirm prod is at that revision before dispatching the workflows.

Link to Devin session: https://app.devin.ai/sessions/3dbf4532c9e94ab6811ff85b49f1a203
Requested by: @HeyGarrison


Open in Devin Review

…k benchmark authoring to config + task

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@HeyGarrison HeyGarrison self-assigned this Jul 30, 2026
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

🤖 Devin AI Engineer

I'll be helping with this pull request! Here's what you should know:

✅ I will automatically:

  • Address comments on this PR. Add '(aside)' to your comment to have me ignore it.
  • Look at CI failures and help fix them

Note: I can only respond to comments from users who have write access to this repository.

⚙️ Control Options:

  • Disable automatic comment, CI, and merge conflict monitoring

@open-cla

open-cla Bot commented Jul 30, 2026

Copy link
Copy Markdown

Contributor License Agreement

All contributors are covered by a CLA.

@devin-ai-integration devin-ai-integration Bot left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Devin Review: No Issues Found

Devin Review analyzed this PR and found no potential bugs to report.

View in Devin Review to see 1 additional finding.

Open in Devin Review

devin-ai-integration Bot and others added 3 commits July 30, 2026 01:21
…and e2e skill

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…participants+onComplete

- @benchsdk/runner: add `bench run <file>` CLI (cli.ts + bin.ts); bin imports the package by name (external) so it and the loaded *.bench.ts share one module instance, keeping instanceof checks valid.
- Task context gains measure()/log(); a task with no explicit steps is recorded as one implicit 'task' step; onComplete + participants live in config; onResult removed from the author surface.
- Migrate all 9 *.bench.ts to config+task (incl. ai-gateway phases via onComplete); drop runBenchmark tails, logTti/logAiGateway/logDax, and bench-exit util.
- create-bench template scaffolds declarative config+task; READMEs, e2e skill, and changesets updated.
- Add CLI unit tests + fixtures.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

@superagent-security superagent-security Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Superagent found 1 security concern(s).

Comment thread packages/benchsdk-runner/src/cli.ts
HeyGarrison and others added 2 commits July 30, 2026 16:00
The declarative rework made *.bench.ts files export only config+task with no
imperative runBenchmark tail, so invoking them directly (npx tsx <file>.bench.ts)
is now a no-op. Repoint the 7 benchmark workflows at the runner bin
(packages/benchsdk-runner/dist/bin.js run <file>) and add the org-scoped
BENCHMARKS_PLATFORM_API_KEY the runner requires to each bench step's vault
allowlist.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@HeyGarrison
HeyGarrison merged commit 9ec0632 into master Jul 30, 2026
55 checks passed
@github-actions

Copy link
Copy Markdown
Contributor

Snapshot/Fork Benchmark Results

small dataset

# Provider Score Snapshot create Fork (snapshot) Fork (live) First read Status
1 Tigris 99.5 0.18s 0.37s 0.39s 0.31s 1/1
2 Azure-blob 98.7 0.92s 0.83s 0.88s 0.02s 1/1
3 Cloudflare R2 95.6 3.09s 2.56s 3.23s 0.10s 1/1
4 AWS S3 0.0 0.00s 0.00s 0.00s 0.00s 0/1

View full run

@github-actions

Copy link
Copy Markdown
Contributor

Storage Benchmark Results

1MB Files

# Provider Score Download Throughput Upload Status
1 Azure Blob Storage 96.3 0.03s 320.5 Mbps 0.07s 10/10
2 Tigris 94.9 0.04s 204.6 Mbps 0.53s 10/10
3 AWS S3 94.8 0.10s 83.3 Mbps 0.20s 10/10
4 Cloudflare R2 94.4 0.14s 63.7 Mbps 0.39s 10/10
5 Vercel Blob 94.4 0.22s 37.4 Mbps 0.19s 10/10
6 Google Cloud Storage 0.0 0.00s 0.0 Mbps 0.00s 0/10

View full run · SVGs available as build artifacts

@github-actions

Copy link
Copy Markdown
Contributor

Browser Benchmark Results

# Provider Score Create Connect Navigate Release Total Status
1 Kernel 97.6 0.04s 0.08s 0.13s 0.05s 0.30s 10/10
2 Tilion 96.9 0.02s 0.05s 0.11s 0.01s 0.20s 10/10
3 Browseruse 95.1 0.21s 0.17s 0.07s 0.04s 0.53s 10/10
4 Browserbase 91.5 0.23s 0.13s 0.10s 0.15s 0.62s 10/10
5 Hyperbrowser 86.5 0.32s 0.56s 0.70s 0.10s 1.68s 10/10
6 Steel 81.6 0.67s 0.41s 0.13s 0.53s 1.83s 10/10
7 Notte 57.6 0.62s 2.04s 0.50s 0.42s 4.16s 10/10

View full run · SVG available as build artifact

@github-actions

Copy link
Copy Markdown
Contributor

Browser Throughput Benchmark Results

# Provider Score APS (med) Task (med) Task (p95) Screenshot Status
1 Kernel 76.4 5.03/s 1.99s 2.93s 286ms 3/3
2 Browseruse 68.4 3.33/s 3.00s 3.35s 306ms 3/3
3 Browserbase 66.0 2.95/s 3.39s 4.28s 283ms 3/3
4 Tilion 64.0 2.74/s 3.65s 5.55s 354ms 3/3
5 Steel 51.3 1.16/s 8.65s 8.70s 616ms 3/3
6 Hyperbrowser 49.7 1.10/s 9.09s 9.90s 1016ms 3/3
7 Notte 18.0 0.38/s 26.25s 45.74s 3351ms 3/3

View full run · SVG available as build artifact

@github-actions

Copy link
Copy Markdown
Contributor

Sandbox Benchmark Results

Sequential

# Provider Score Median TTI P95 P99 Status
1 declaw 99.4 0.05s 0.09s 0.09s 10/10
2 northflank 98.8 0.10s 0.16s 0.16s 10/10
3 createos 98.1 0.08s 0.35s 0.35s 10/10
4 archil 97.8 0.14s 0.33s 0.33s 10/10
5 lightning 97.2 0.19s 0.41s 0.41s 10/10
6 tensorlake 96.6 0.32s 0.37s 0.37s 10/10
7 blaxel 96.3 0.24s 0.55s 0.55s 10/10
8 superserve 96.1 0.36s 0.44s 0.44s 10/10
9 upstash 95.9 0.36s 0.49s 0.49s 10/10
10 isorun 95.7 0.02s 1.03s 1.03s 10/10
11 beam 95.3 0.10s 1.02s 1.02s 10/10
12 runloop 94.5 0.45s 0.70s 0.70s 10/10
13 modal 94.3 0.54s 0.62s 0.62s 10/10
14 vercel 93.6 0.37s 1.04s 1.04s 10/10
15 cloud-run 93.3 0.60s 0.76s 0.76s 10/10
16 e2b 89.0 0.83s 1.50s 1.50s 10/10
17 tenki 88.2 0.83s 1.70s 1.70s 10/10
18 opencomputer 76.3 1.81s 3.20s 3.20s 10/10
19 sandbox0 64.3 2.71s 4.85s 4.85s 10/10
20 cloudflare 48.8 3.80s 7.10s 7.10s 10/10
21 daytona 48.7 1.88s 11.13s 11.13s 10/10
22 codesandbox 44.6 2.57s 20.79s 20.79s 10/10
23 hopx 0.0 0.00s 0.00s 0.00s 0/10

Staggered

# Provider Score Median TTI P95 P99 Status
1 declaw 99.3 0.05s 0.11s 0.11s 10/10
2 createos 99.0 0.08s 0.13s 0.13s 10/10
3 archil 98.4 0.13s 0.21s 0.21s 10/10
4 northflank 98.2 0.17s 0.19s 0.19s 10/10
5 blaxel 97.2 0.25s 0.34s 0.34s 10/10
6 superserve 96.9 0.29s 0.33s 0.33s 10/10
7 tensorlake 96.6 0.30s 0.39s 0.39s 10/10
8 lightning 96.6 0.23s 0.51s 0.51s 10/10
9 vercel 95.7 0.39s 0.48s 0.48s 10/10
10 beam 95.4 0.15s 0.93s 0.93s 10/10
11 upstash 95.0 0.47s 0.54s 0.54s 10/10
12 modal 92.7 0.57s 0.95s 0.95s 10/10
13 runloop 90.6 0.56s 1.51s 1.51s 10/10
14 e2b 90.2 0.85s 1.18s 1.18s 10/10
15 tenki 89.0 0.85s 1.48s 1.48s 10/10
16 daytona 84.8 1.28s 1.87s 1.87s 10/10
17 opencomputer 79.0 2.03s 2.21s 2.21s 10/10
18 codesandbox 74.9 2.31s 2.80s 2.80s 10/10
19 cloud-run 67.3 2.82s 3.95s 3.95s 10/10
20 cloudflare 66.3 2.92s 4.04s 4.04s 10/10
21 isorun 59.8 0.03s 51.74s 51.74s 10/10
22 sandbox0 16.8 7.95s 8.89s 8.89s 10/10
23 hopx 0.0 0.00s 0.00s 0.00s 0/10

Burst

# Provider Score Median TTI P95 P99 Status
1 declaw 99.0 0.10s 0.11s 0.11s 10/10
2 createos 98.5 0.15s 0.15s 0.15s 10/10
3 lightning 97.8 0.21s 0.24s 0.24s 10/10
4 archil 97.7 0.22s 0.25s 0.25s 10/10
5 northflank 97.4 0.26s 0.27s 0.27s 10/10
6 blaxel 96.8 0.31s 0.33s 0.33s 10/10
7 superserve 96.3 0.37s 0.38s 0.38s 10/10
8 tensorlake 95.7 0.42s 0.46s 0.46s 10/10
9 vercel 94.9 0.43s 0.63s 0.63s 10/10
10 upstash 94.5 0.51s 0.62s 0.62s 10/10
11 modal 92.7 0.61s 0.93s 0.93s 10/10
12 e2b 90.4 0.87s 1.09s 1.09s 10/10
13 beam 90.2 0.98s 1.00s 1.00s 10/10
14 daytona 89.4 0.95s 1.23s 1.23s 10/10
15 tenki 86.4 1.34s 1.39s 1.39s 10/10
16 cloud-run 79.8 0.62s 4.13s 4.13s 10/10
17 opencomputer 77.0 2.24s 2.40s 2.40s 10/10
18 runloop 71.6 2.24s 3.74s 3.74s 10/10
19 sandbox0 67.9 3.15s 3.31s 3.31s 10/10
20 codesandbox 67.7 2.91s 3.70s 3.70s 10/10
21 isorun 59.7 0.06s 30.12s 30.12s 10/10
22 cloudflare 54.8 3.30s 6.36s 6.36s 10/10
23 hopx 0.0 0.00s 0.00s 0.00s 0/10

View full run · SVGs available as build artifacts

@github-actions

Copy link
Copy Markdown
Contributor

Sandbox Dax Benchmark Results

Provider Phases Total Prepare Clone Install Typecheck Status
archil 0/7 0.00s 0.00s 0.00s 0.00s 0.00s 0/1 OK
beam 0/7 0.00s 0.00s 0.00s 0.00s 0.00s 0/1 OK
blaxel 6/7 47.20s 2.14s 2.56s 12.18s 0.00s 0/1 OK
cloud-run 2/7 0.93s 0.19s 0.00s 0.00s 0.00s 0/1 OK
cloudflare 0/7 0.00s 0.00s 0.00s 0.00s 0.00s 0/1 OK
codesandbox 0/7 0.00s 0.00s 0.00s 0.00s 0.00s 0/1 OK
createos 6/7 48.73s 3.07s 1.96s 12.73s 0.00s 0/1 OK
daytona 7/7 69.63s 2.93s 2.50s 13.45s 35.70s 1/1 OK
declaw 7/7 165.49s 53151.25s 4.82s 13.48s 127.83s 1/1 OK
e2b 7/7 79.17s 9.31s 4.11s 16.07s 41.09s 1/1 OK
hopx -- 0.00s 0.00s 0.00s 0.00s 0.00s 0/1 OK
isorun -- 0.00s 0.00s 0.00s 0.00s 0.00s 0/1 OK
lightning 7/7 43.33s 4.91s 2.68s 12.52s 17.54s 1/1 OK
modal 7/7 98.74s 3.98s 2.52s 14.87s 69.13s 1/1 OK
namespace 7/7 45.67s 5.25s 1.22s 8.13s 24.47s 1/1 OK
northflank -- 0.00s 0.00s 0.00s 0.00s 0.00s 0/1 OK
opencomputer 6/7 26.79s 4.24s 1.89s 16.84s 0.00s 0/1 OK
runloop 2/7 83.44s 0.00s 0.00s 18.52s 35.96s 1/1 OK
sandbox0 7/7 180.06s 4.15s 7.62s 73.29s 75.32s 1/1 OK
superserve 7/7 108.56s 6.47s 4.83s 24.98s 68.02s 1/1 OK
tenki 7/7 70.33s 4.83s 4.16s 15.31s 39.03s 1/1 OK
tensorlake 7/7 56.42s 11.93s 1.53s 10.39s 28.14s 1/1 OK
upstash 7/7 52.40s 4.20s 1.60s 10.91s 24.92s 1/1 OK
vercel 7/7 76.58s 17.01s 1.79s 12.81s 37.32s 1/1 OK

View full run

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant