feat(racetrack): run Gherkin scenarios against real editors and return JSON evidence - #3333
christianhg wants to merge 2 commits into
Conversation
Tools built on the test harness had no way to observe an editor from the start: `createTestEditor` and `createTestEditors` hand the editor back only after it has rendered and become editable, so a subscription made then misses `ready` and anything emitted during startup. `onTestEditorCreated` registers a listener that receives each editor (named `A`, or `B` for the second editor of `createTestEditors`) from a plugin mounted as the first child of the provider. Its effect runs before `EditorProvider` starts the editor, so an `editor.on` subscription made in the listener sees the startup events. With no listener registered the plugin does nothing. The hook is internal to the harness and not exported from any package entry.
…n JSON evidence Racetrack becomes a harness for agents: `pnpm racetrack run <file.feature>` runs a scenario against real editors in a Playwright browser, through a throwaway vitest browser test built from the shared editor step definitions, and writes an evidence bundle. The bundle holds checkpoints (textspec, value, model and DOM selection per editor) taken at `Then capture the state` and at a failing step, a passive log of every emitted event from editor creation on, console errors, and the outcome. `passed` requires a completed run with at least one assertion, so a scenario that only captures state never passes. `pnpm racetrack steps` prints the step catalogue generated from the definitions. Checkpoints settle on two animation frames and never wait for mutation flushes, so capture does not change batching. Each scenario gets its own step context, a timed-out step marks the remaining scenarios as not run instead of letting them share a page with it, and failure capture is best-effort so the original failure and its evidence survive. The teaser app is removed.
|
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
Bundle Stats✅ No significant changes. All scenario measurements (7)🗺️
Significant means at least 1.0 KB and 1% gzip, or at least 5 ms and 10% import time. |
Agents proving something about the editor either write a test file or claim. Racetrack gives them a third route: run a Gherkin scenario against real editors and get back what the editors did.
A run executes the scenario in a Playwright browser through a throwaway vitest test built from the shared editor step definitions, and writes an evidence bundle: checkpoints at
Then capture the stateand at a failing step (textspec, value, model and DOM selection per editor), every emitted event from editor creation on, console errors, and the outcome. A run only passes with at least one assertion, so capturing state proves nothing by itself.apps/racetrack/README.mdcovers the bundle, the limits inherited from the step definitions, and the manual path from a passing scenario intogherkin-spec/.Capture never waits for mutation flushes, so it does not change batching. To record from creation, the test harness gains an internal
onTestEditorCreatedhook. It does nothing without a listener, and the full editor chromium suite passes with it (193 files).Racetrack's own unit and e2e tests pin scenario isolation, timeout isolation, failure capture, and the pass rules. Not covered yet: an end-to-end unhandled page error (only the aggregation is unit-tested) and the hook under React strict mode.