Sitelet https://github.com/NodeOps-app/benchmarks/commit/9ec0632d05af58a89115789d157ed7f0b953fc9e
Skip to content

Commit 9ec0632

Browse files
Split @benchsdk/client (REST) and @benchsdk/runner (framework); rework benchmark authoring to config + task (computesdk#253)
* Split @benchsdk/client (REST) and @benchsdk/runner (framework); rework benchmark authoring to config + task Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * Merge master: adopt BENCHMARKS_PLATFORM_API_KEY auth in runner tests and e2e skill Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * Finish declarative authoring: bench run CLI, ctx.measure/log, config participants+onComplete - @benchsdk/runner: add `bench run <file>` CLI (cli.ts + bin.ts); bin imports the package by name (external) so it and the loaded *.bench.ts share one module instance, keeping instanceof checks valid. - Task context gains measure()/log(); a task with no explicit steps is recorded as one implicit 'task' step; onComplete + participants live in config; onResult removed from the author surface. - Migrate all 9 *.bench.ts to config+task (incl. ai-gateway phases via onComplete); drop runBenchmark tails, logTti/logAiGateway/logDax, and bench-exit util. - create-bench template scaffolds declarative config+task; READMEs, e2e skill, and changesets updated. - Add CLI unit tests + fixtures. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * ci: run benchmark workflows through the runner CLI (bench run) The declarative rework made *.bench.ts files export only config+task with no imperative runBenchmark tail, so invoking them directly (npx tsx <file>.bench.ts) is now a no-op. Repoint the 7 benchmark workflows at the runner bin (packages/benchsdk-runner/dist/bin.js run <file>) and add the org-scoped BENCHMARKS_PLATFORM_API_KEY the runner requires to each bench step's vault allowlist. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: garrison <garrison@computesdk.com> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
1 parent 5b4de68 commit 9ec0632

72 files changed

Lines changed: 1233 additions & 1994 deletions

File tree

Some content is hidden

Large Commits have some content hidden by default. Use the searchbox below for content that may be hidden.

‎.agents/skills/local-platform-e2e/SKILL.md‎

Lines changed: 16 additions & 15 deletions
Original file line numberDiff line numberDiff line change
@@ -1,9 +1,9 @@
11
---
22
name: local-platform-e2e
3-
description: Stand up benchmarks-platform locally (Postgres + MinIO + ClickHouse in docker) and run a real @benchsdk/cli benchmark against it, with no cloud or provider credentials. Use when testing @benchsdk/client / @benchsdk/cli against the platform end to end, or when debugging benchmark reporting, worker planning, artifacts, or dashboard results locally.
3+
description: Stand up benchmarks-platform locally (Postgres + MinIO + ClickHouse in docker) and run a real @benchsdk/runner benchmark against it, with no cloud or provider credentials. Use when testing @benchsdk/client / @benchsdk/runner against the platform end to end, or when debugging benchmark reporting, worker planning, artifacts, or dashboard results locally.
44
---
55

6-
# Local end-to-end: @benchsdk/cli ↔ benchmarks-platform
6+
# Local end-to-end: @benchsdk/runner ↔ benchmarks-platform
77

88
Goal: exercise upsert benchmark → create run → planWorkers → claim → heartbeat →
99
task_results → artifact upload → complete → dashboard, with zero external
@@ -130,29 +130,30 @@ Check `imported`/`failed`/`failureSamples` in the response.
130130
Build first (`packages/*/dist` is not committed): `pnpm install && pnpm -r --filter "./packages/**" build`.
131131

132132
Write a throwaway bench inside the repo (untracked, e.g. `e2e-local/local.bench.ts`)
133-
so pnpm workspace resolution finds `@benchsdk/cli`, with a fake participant:
133+
so pnpm workspace resolution finds `@benchsdk/runner`, with a fake participant:
134134

135135
```ts
136-
import { defineBenchmark, runBenchmark } from '@benchsdk/cli';
137-
const config = defineBenchmark({
136+
import { defineBenchmarkConfig, defineTask } from '@benchsdk/runner';
137+
export const config = defineBenchmarkConfig({
138138
benchmarkSlug: 'e2e-local', benchmarkName: 'E2E', iterations: 4, concurrency: 1,
139-
task: async (ctx) => { await ctx.step('create', () => new Promise(r => setTimeout(r, 50))); },
139+
participants: [{ name: 'local', requiredEnvVars: [] }],
140+
});
141+
export const task = defineTask(async (ctx) => {
142+
await ctx.step('create', () => new Promise((r) => setTimeout(r, 50)));
143+
ctx.measure({ ok: true });
140144
});
141-
runBenchmark(config, [{ name: 'local', requiredEnvVars: [] } as any], process.argv.slice(2));
142145
```
143146

144-
Run it:
147+
Run it via the `bench run` CLI (under tsx so the `.bench.ts` module loads without a build):
145148
```bash
146-
BENCHMARKS_PLATFORM_URL=http://localhost:3000 COMPUTESDK_ADMIN_API_KEY=local-admin-key \
147-
npx tsx e2e-local/local.bench.ts --iterations 4 --concurrency 2
149+
BENCHMARKS_PLATFORM_URL=http://localhost:3000 BENCHMARKS_PLATFORM_API_KEY=<org bp_ key> \
150+
npx tsx packages/benchsdk-runner/dist/bin.js run e2e-local/local.bench.ts --iterations 4 --concurrency 2
148151
```
149152
`BENCHMARKS_PLATFORM_URL` is the **root** URL (the runner appends `/api/v1`).
150153

151-
The runner reads `COMPUTESDK_ADMIN_API_KEY ?? COMPUTESDK_API_KEY`, so to prove the master key
152-
is not needed, pass an org `bp_` key as `COMPUTESDK_API_KEY` and strip the admin vars from the
153-
child env (`env -u COMPUTESDK_ADMIN_API_KEY -u ADMIN_API_KEY …`). Note the "View at:" URL the
154-
CLI prints uses `BENCHMARKS_PLATFORM_ORG_SLUG` (default `computesdk`), **not** the org that owns
155-
the key, so with an org key the printed link may 404 — set that env var to the key's org slug.
154+
The runner authenticates with `BENCHMARKS_PLATFORM_API_KEY` (an org-scoped `bp_` key — mint one
155+
locally per the section below) and pulls the owning org slug from the server, so the "View at:"
156+
URL it prints always points at the run's real org (no `BENCHMARKS_PLATFORM_ORG_SLUG` needed).
156157

157158
## 6b. Probing tenant isolation
158159

‎.changeset/benchsdk-cli-initial.md‎

Lines changed: 0 additions & 5 deletions
This file was deleted.
Lines changed: 5 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,5 @@
1+
---
2+
"@benchsdk/client": minor
3+
---
4+
5+
Make `@benchsdk/client` a pure REST + worker-engine package. The benchmark authoring factories `defineStep`, `defineTask`, `defineWorker`, `defineBench`, and the `runBenchmarkWorker` free function have been removed — that authoring model now lives in `@benchsdk/runner`. `client.runWorker({ task })` now accepts a raw `TaskFunction` whose context exposes `step(...)` (imperative named steps), `measure(data)` (explicit metrics — merged into the active step's data, or the task record outside a step; a task with no explicit steps is recorded as one implicit `'task'` step carrying its measurements, and measurements are preserved when a task throws), and `log(message, meta?)` (buffered per worker and uploaded once as a `worker.log` artifact). `createBenchmarkClient`, the REST methods, `BenchmarkReporter`, and the system-metrics collector are unchanged.
Lines changed: 5 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,5 @@
1+
---
2+
"@benchsdk/runner": patch
3+
---
4+
5+
Initial publication of `@benchsdk/runner`, the benchmark authoring framework (renamed from `@benchsdk/cli`). A `*.bench.ts` file exports exactly two things: a **config** (`defineBenchmarkConfig({ benchmarkSlug, iterations, concurrency, participants, onComplete, ... })` — orchestration knobs, the `participants` to run against, and an optional run-level `onComplete` hook) and a **task** (`defineTask(fn)`, the workload for one iteration, with named steps via `ctx.step` supporting closures and `try/finally`). The task context also exposes `ctx.measure(data)` (explicit metric channel — merges into the active step's data, or the task record outside a step; a task with no explicit steps is recorded as one implicit `'task'` step) and `ctx.log(message, meta?)` (timeline narration). The `bench run <file>` CLI owns the entrypoint: it imports the module, reads `config`/`task`, applies CLI overrides (`--iterations`, `--concurrency`, `--stagger-delay-ms`, `--group-by`, `--provider`), and drives the run against `@benchsdk/client`; benchmark files no longer call the runner themselves. `NoAvailableParticipantsError` (every participant env-gated out) exits cleanly. Also exports `TaskError`.

.changeset/benchsdk-cli-no-available-participants.md renamed to .changeset/benchsdk-runner-no-available-participants.md

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -1,5 +1,5 @@
11
---
2-
"@benchsdk/cli": patch
2+
"@benchsdk/runner": patch
33
---
44

55
`runBenchmark()` now rejects with the exported `NoAvailableParticipantsError` (carrying the `skipped` participants and their missing env vars) instead of a plain `Error` when every participant is env-gated out, so callers can treat an unprovisioned provider as a skip rather than a failure.

‎.github/workflows/ai-gateway-benchmarks.yml‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -83,7 +83,7 @@ jobs:
8383
if [ -n "${{ github.event.inputs.provider }}" ]; then
8484
PROVIDER_FLAG="--provider ${{ github.event.inputs.provider }}"
8585
fi
86-
npx tsx benchmarks/ai-gateway/ai-gateway.bench.ts $PROVIDER_FLAG \
86+
npx tsx packages/benchsdk-runner/dist/bin.js run benchmarks/ai-gateway/ai-gateway.bench.ts $PROVIDER_FLAG \
8787
--iterations ${{ (github.event_name == 'push' && '10') || github.event.inputs.iterations || '10' }}
8888
- run: pnpm run generate-ai-gateway-svg
8989
- name: Upload results and SVG as artifacts

‎.github/workflows/browser-benchmarks.yml‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -67,7 +67,7 @@ jobs:
6767
run: |
6868
. benchmarks/scripts/load-vault-secrets.sh '^(BROWSERBASE_API_KEY|BROWSERBASE_PROJECT_ID|BROWSER_USE_API_KEY|HYPERBROWSER_API_KEY|KERNEL_API_KEY|NOTTE_API_KEY|STEEL_API_KEY|TILION_API_KEY|TILION_BASE_URL|COMPUTESDK_ADMIN_API_KEY|BENCHMARKS_PLATFORM_API_KEY)'
6969
70-
npx tsx benchmarks/browser/browser.bench.ts \
70+
npx tsx packages/benchsdk-runner/dist/bin.js run benchmarks/browser/browser.bench.ts \
7171
--provider ${{ matrix.provider }} \
7272
--iterations ${{ github.event_name == 'push' && '10' || github.event.inputs.iterations || '100' }}
7373
- name: Upload results

‎.github/workflows/browser-throughput-benchmarks.yml‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -120,7 +120,7 @@ jobs:
120120
run: |
121121
. benchmarks/scripts/load-vault-secrets.sh '^(BROWSERBASE_API_KEY|BROWSERBASE_PROJECT_ID|BROWSER_USE_API_KEY|HYPERBROWSER_API_KEY|KERNEL_API_KEY|NOTTE_API_KEY|STEEL_API_KEY|TILION_API_KEY|TILION_BASE_URL|COMPUTESDK_ADMIN_API_KEY|BENCHMARKS_PLATFORM_API_KEY)'
122122
123-
npx tsx benchmarks/browser/browser-throughput.bench.ts \
123+
npx tsx packages/benchsdk-runner/dist/bin.js run benchmarks/browser/browser-throughput.bench.ts \
124124
--provider ${{ matrix.provider }} \
125125
--iterations ${{ github.event_name == 'push' && '3' || github.event.inputs.iterations || '100' }}
126126
- name: Upload results

‎.github/workflows/ci.yml‎

Lines changed: 4 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -32,11 +32,11 @@ jobs:
3232

3333
- name: Build packages
3434
# The root typecheck consumes the built dist output, so build
35-
# @benchsdk/client, @benchsdk/cli, and create-bench before running
35+
# @benchsdk/client, @benchsdk/runner, and create-bench before running
3636
# typecheck.
3737
run: |
3838
pnpm --filter @benchsdk/client run build
39-
pnpm --filter @benchsdk/cli run build
39+
pnpm --filter @benchsdk/runner run build
4040
pnpm --filter create-bench run build
4141
4242
- name: Typecheck
@@ -45,8 +45,8 @@ jobs:
4545
- name: Test
4646
run: pnpm --filter @benchsdk/client run test
4747

48-
- name: Test @benchsdk/cli
49-
run: pnpm --filter @benchsdk/cli run test
48+
- name: Test @benchsdk/runner
49+
run: pnpm --filter @benchsdk/runner run test
5050

5151
- name: Test create-bench
5252
run: pnpm --filter create-bench run test

‎.github/workflows/release.yml‎

Lines changed: 2 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -46,13 +46,13 @@ jobs:
4646
- name: Build packages
4747
run: |
4848
pnpm --filter @benchsdk/client run build
49-
pnpm --filter @benchsdk/cli run build
49+
pnpm --filter @benchsdk/runner run build
5050
pnpm --filter create-bench run build
5151
5252
- name: Test
5353
run: |
5454
pnpm --filter @benchsdk/client run test
55-
pnpm --filter @benchsdk/cli run test
55+
pnpm --filter @benchsdk/runner run test
5656
pnpm --filter create-bench run test
5757
5858
- name: Create Release Pull Request or Publish to npm

0 commit comments

Comments
 (0)