Split @benchsdk/client (REST) and @benchsdk/runner (framework); rework benchmark authoring to config + task - #253
Merged
HeyGarrison merged 6 commits intoJul 30, 2026
Conversation
…k benchmark authoring to config + task Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Contributor
Author
🤖 Devin AI EngineerI'll be helping with this pull request! Here's what you should know: ✅ I will automatically:
Note: I can only respond to comments from users who have write access to this repository. ⚙️ Control Options:
|
Contributor License AgreementAll contributors are covered by a CLA. |
…nchsdk-client-runner-split
…and e2e skill Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…participants+onComplete - @benchsdk/runner: add `bench run <file>` CLI (cli.ts + bin.ts); bin imports the package by name (external) so it and the loaded *.bench.ts share one module instance, keeping instanceof checks valid. - Task context gains measure()/log(); a task with no explicit steps is recorded as one implicit 'task' step; onComplete + participants live in config; onResult removed from the author surface. - Migrate all 9 *.bench.ts to config+task (incl. ai-gateway phases via onComplete); drop runBenchmark tails, logTti/logAiGateway/logDax, and bench-exit util. - create-bench template scaffolds declarative config+task; READMEs, e2e skill, and changesets updated. - Add CLI unit tests + fixtures. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
The declarative rework made *.bench.ts files export only config+task with no imperative runBenchmark tail, so invoking them directly (npx tsx <file>.bench.ts) is now a no-op. Repoint the 7 benchmark workflows at the runner bin (packages/benchsdk-runner/dist/bin.js run <file>) and add the org-scoped BENCHMARKS_PLATFORM_API_KEY the runner requires to each bench step's vault allowlist. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Contributor
Snapshot/Fork Benchmark Resultssmall dataset
|
Contributor
Storage Benchmark Results1MB Files
View full run · SVGs available as build artifacts |
Contributor
Browser Benchmark Results
View full run · SVG available as build artifact |
Contributor
Browser Throughput Benchmark Results
View full run · SVG available as build artifact |
Contributor
Sandbox Benchmark ResultsSequential
Staggered
Burst
View full run · SVGs available as build artifacts |
Contributor
Sandbox Dax Benchmark Results
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Splits the two SDK packages around a clean boundary and reworks benchmark authoring into a fully declarative model. No published users, so this is a breaking rework rather than a compat shim; scale stays on its own orchestration path.
@benchsdk/client= REST transport + worker engine. Removed the authoring factoriesdefineStep/defineTask/defineWorker/defineBenchand therunBenchmarkWorkerfree function.client.runWorker({ task })takes a rawTaskFunctionwhose context now exposes three channels:step()(imperative named steps),measure(data), andlog(msg, meta?).@benchsdk/runner= the authoring framework (renamed from@benchsdk/cli). A*.bench.tsexports exactly two things —configandtask— and never calls the runner itself; the newbench run <file>CLI owns the entrypoint.Authoring model (the whole surface)
Key semantics:
step()returns are control-flow values, never auto-recorded — you thread live objects (a sandbox, a client) between steps.measure(data)is the explicit metric channel: inside astep()it merges into that step'sdata; at task top level it merges into the record.log()is human-readable timeline narration.step()calls is recorded as one implicit'task'step spanning the task's wall-clock, carrying any task-level measurements. Measurements are preserved even when a task throws.onResultis gone from the author surface (the runner still prints a default per-record line); per-iteration concerns live in the task, aggregate concerns inconfig.onComplete.bench runentrypointcli.tsimports the module, validatesconfig/task, and calls the now-internalrunBenchmark(config, task, argv);NoAvailableParticipantsErrormaps to a clean exit. Thebinre-imports the package by name (markedexternalin tsup) so the bin and the dynamically-imported*.bench.tsshare one@benchsdk/runnerinstance — otherwiseinstanceof TaskError/NoAvailableParticipantsErrorwould break across separately-bundled entries.Migrations & housekeeping
*.bench.tsmigrated toconfig+task; each file's legacy aggregate writer moved intoconfig.onComplete; removedrunBenchmarktails,logTti/logAiGateway/logDax, and thebench-exitutil. AI Gateway keeps its cold/warm phase parsing and zero-phase skip.package.jsonbench:*scripts now runtsx …/dist/bin.js run <file>.*.bench.tsno longer self-run, the 7 workflows that invokednpx tsx <file>.bench.ts(sandbox-tti, sandbox-dax, storage, snapshot-fork, browser, browser-throughput, ai-gateway) would have silently run zero benchmarks. They now callnpx tsx packages/benchsdk-runner/dist/bin.js run <file> …(dist is produced by the runner'sprepare: tsuponpnpm install), and each bench step'sload-vault-secretsallowlist gains the org-scopedBENCHMARKS_PLATFORM_API_KEYthe runner requires (prod platform URL is the default, so no URL var needed). The othersandbox-*micro-benchmarks use the separatebenchmarks/src/run.tsharness and are unaffected.create-benchscaffolds a declarativeconfig+taskproject on@benchsdk/runner.Testing
pnpm -r --filter "./packages/**" build— builds (CJS/ESM/DTS)pnpm typecheck— cleanpnpm --filter @benchsdk/client test— 118 passed, 1 skippedpnpm --filter @benchsdk/runner test— 53 passed (incl. new CLI tests)pnpm --filter create-bench test— 1 passedbench runsmoke against a local benchmarks-platform (Postgres/MinIO/ClickHouse) created a run and reported success (exit 0).npx tsx packages/benchsdk-runner/dist/bin.js run <file>for sequential/storage/browser/ai-gateway all exit 0 via the "no participants" path with no creds;pnpm install --frozen-lockfilerebuildsdist/bin.jsviaprepare.Tests were updated to match the intentionally-changed public API, not to paper over failures.
Note: the runner requires benchmarks-platform at
origin/main(org-scoped API-key auth +organizationSlugreturned fromcreateRun); confirm prod is at that revision before dispatching the workflows.Link to Devin session: https://app.devin.ai/sessions/3dbf4532c9e94ab6811ff85b49f1a203
Requested by: @HeyGarrison