Sitelet https://github.com/NodeOps-app/benchmarks/commit/f493beaae24bfbd77440862735e638f0b38f8307
Skip to content

Commit f493bea

Browse files
committed
Professionalize benchmarks for sponsorship program
- Overhaul README with product-focused structure, transparency section, and roadmap - Add METHODOLOGY.md with detailed, auditable benchmark documentation - Add SPONSORSHIP.md with pricing, benefits, and editorial independence policy - Add CONTRIBUTING.md for providers and community contributors - Add MIT LICENSE - Improve JSON output with rounded numbers and environment metadata - Remove scratch notes file (benchmarking.md) - Rename 'gateway' to 'orchestrator' throughout
1 parent b1a23b4 commit f493bea

10 files changed

Lines changed: 630 additions & 46 deletions

‎CONTRIBUTING.md‎

Lines changed: 112 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,112 @@
1+
# Contributing
2+
3+
ComputeSDK Benchmarks is open source. We welcome contributions that improve measurement accuracy, add providers, or enhance the project.
4+
5+
## For Sandbox Providers
6+
7+
Want your provider included in the benchmark?
8+
9+
**Option 1: Become a Sponsor**
10+
11+
Sponsorship includes full integration and maintenance. See [SPONSORSHIP.md](./SPONSORSHIP.md).
12+
13+
**Option 2: Submit a PR**
14+
15+
We accept community contributions to add providers. Requirements:
16+
17+
1. **Public SDK**: Your provider must have a publicly available SDK (npm package preferred)
18+
2. **Standard Interface**: Must support basic sandbox operations (create, run command, destroy)
19+
3. **Free Tier or Credits**: We need ongoing API access for daily benchmarks
20+
4. **Documentation**: Clear setup instructions for API keys/credentials
21+
22+
### Adding a Provider (Direct Mode)
23+
24+
1. Create a new SDK wrapper in `src/direct-providers.ts`:
25+
26+
```typescript
27+
export const yourProvider: DirectBenchmarkConfig = {
28+
name: 'your-provider',
29+
requiredEnvVars: ['YOUR_API_KEY'],
30+
createCompute: () => new YourSDK({
31+
apiKey: process.env.YOUR_API_KEY!,
32+
}),
33+
};
34+
```
35+
36+
2. Add to the providers array in `src/direct-run.ts`
37+
38+
3. Update `env.example` with required environment variables
39+
40+
4. Submit a PR with:
41+
- The code changes
42+
- Documentation for obtaining API credentials
43+
- Confirmation you can provide ongoing API access
44+
45+
### Adding a Provider (Magic Mode)
46+
47+
Magic Mode requires integration with the ComputeSDK orchestrator. Contact us at benchmarks@computesdk.com to discuss.
48+
49+
## For General Contributors
50+
51+
### Bug Fixes
52+
53+
Found a bug? Please:
54+
55+
1. Check existing issues first
56+
2. Open an issue describing the bug
57+
3. Submit a PR with the fix (reference the issue)
58+
59+
### Methodology Improvements
60+
61+
We're open to improving how we measure performance. Before making changes:
62+
63+
1. Open an issue describing the proposed change
64+
2. Explain why it improves accuracy or fairness
65+
3. Wait for maintainer feedback before implementing
66+
67+
Methodology changes require careful consideration since they affect historical comparability.
68+
69+
### Documentation
70+
71+
Documentation improvements are always welcome. No issue required for typos, clarifications, or formatting fixes.
72+
73+
## Development Setup
74+
75+
```bash
76+
git clone https://github.com/computesdk/benchmarks.git
77+
cd benchmarks
78+
npm install
79+
cp env.example .env
80+
```
81+
82+
### Running Tests Locally
83+
84+
```bash
85+
# Run direct mode benchmarks (requires API keys in .env)
86+
npm run bench:direct
87+
88+
# Run single provider
89+
npm run bench:direct:e2b
90+
91+
# Run with custom iterations
92+
npm run bench:direct -- --iterations 5
93+
```
94+
95+
### Code Style
96+
97+
- TypeScript with strict mode
98+
- ES modules (`import`/`export`)
99+
- Prettier for formatting (run `npm run format` if available)
100+
101+
## Code of Conduct
102+
103+
- Be respectful and constructive
104+
- Focus on technical merit
105+
- No promotional content in issues/PRs
106+
- Disclose any conflicts of interest (e.g., if you work for a benchmarked provider)
107+
108+
## Questions
109+
110+
- **General questions**: Open a GitHub issue
111+
- **Sponsorship**: sponsors@computesdk.com
112+
- **Security issues**: security@computesdk.com (do not open public issues)

‎LICENSE‎

Lines changed: 21 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,21 @@
1+
MIT License
2+
3+
Copyright (c) 2026 ComputeSDK
4+
5+
Permission is hereby granted, free of charge, to any person obtaining a copy
6+
of this software and associated documentation files (the "Software"), to deal
7+
in the Software without restriction, including without limitation the rights
8+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9+
copies of the Software, and to permit persons to whom the Software is
10+
furnished to do so, subject to the following conditions:
11+
12+
The above copyright notice and this permission notice shall be included in all
13+
copies or substantial portions of the Software.
14+
15+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21+
SOFTWARE.

‎METHODOLOGY.md‎

Lines changed: 259 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,259 @@
1+
# Benchmark Methodology
2+
3+
This document describes how ComputeSDK Benchmarks measures sandbox provider performance. Our goal is transparent, reproducible, and fair measurement.
4+
5+
## What We Measure
6+
7+
### Time to Interactive (TTI)
8+
9+
**Definition**: The wall-clock time from initiating a sandbox creation request to successfully executing the first command.
10+
11+
TTI captures the complete developer experience:
12+
13+
```
14+
┌─────────────────────────────────────────────────────────────────────────┐
15+
│ Time to Interactive (TTI) │
16+
├─────────────┬─────────────────┬──────────────┬─────────────┬───────────┤
17+
│ API Latency │ Provisioning │ Boot Time │ Health Check│ Command │
18+
│ │ │ │ Polling │ Execution │
19+
└─────────────┴─────────────────┴──────────────┴─────────────┴───────────┘
20+
```
21+
22+
This metric matters because it's what developers actually experience—the time spent waiting before they can use the sandbox.
23+
24+
### What's Included in TTI
25+
26+
- Network round-trip to provider API
27+
- Queue time (if provider has provisioning queues)
28+
- Infrastructure allocation (VM, container, or serverless spin-up)
29+
- Operating system and runtime boot
30+
- Provider daemon/agent initialization
31+
- Health check and readiness polling
32+
- First command network round-trip
33+
- Command execution time (trivial for our test command)
34+
35+
### What's NOT Included
36+
37+
- Sandbox teardown/destruction time
38+
- Subsequent command execution times
39+
- File system operations
40+
- Network transfer speeds within the sandbox
41+
42+
## Test Procedure
43+
44+
Each benchmark iteration executes the following steps:
45+
46+
```typescript
47+
// 1. Start timer
48+
const start = performance.now();
49+
50+
// 2. Create sandbox and wait until ready
51+
const sandbox = await compute.sandbox.create();
52+
53+
// 3. Execute a trivial command to confirm interactivity
54+
await sandbox.runCommand('echo "benchmark"');
55+
56+
// 4. Stop timer
57+
const ttiMs = performance.now() - start;
58+
59+
// 5. Cleanup (not timed)
60+
await sandbox.destroy();
61+
```
62+
63+
### Why `echo "benchmark"`?
64+
65+
We use a minimal command to isolate sandbox startup time from command complexity. The command:
66+
- Has negligible execution time
67+
- Requires no file system access
68+
- Produces deterministic output
69+
- Validates the full request/response cycle
70+
71+
## Test Configuration
72+
73+
### Daily Automated Runs
74+
75+
| Parameter | Value |
76+
|-----------|-------|
77+
| Iterations per provider | 10 |
78+
| Timeout per iteration | 120 seconds |
79+
| Run frequency | Daily at 00:00 UTC |
80+
| Runner environment | GitHub Actions (ubuntu-latest) |
81+
| Node.js version | 20.x |
82+
83+
### Provider Execution Order
84+
85+
Providers are tested **sequentially** to:
86+
- Avoid resource contention on the test runner
87+
- Prevent rate limiting issues
88+
- Ensure consistent network conditions per provider
89+
90+
The order is randomized each run to prevent systematic bias from time-of-day effects.
91+
92+
## Statistical Reporting
93+
94+
For each provider, we report:
95+
96+
| Metric | Description |
97+
|--------|-------------|
98+
| **Min** | Fastest iteration (best case) |
99+
| **Max** | Slowest iteration (worst case) |
100+
| **Median** | Middle value (typical case) |
101+
| **Average** | Arithmetic mean |
102+
| **Success Rate** | Iterations completed without error |
103+
104+
We emphasize **median** as the primary metric because it's robust to outliers and represents the typical developer experience.
105+
106+
## Benchmark Modes
107+
108+
### Direct Mode
109+
110+
Tests each provider's native SDK without any abstraction layer.
111+
112+
```typescript
113+
import { E2B } from '@computesdk/e2b';
114+
115+
const compute = new E2B({ apiKey: process.env.E2B_API_KEY });
116+
const sandbox = await compute.sandbox.create();
117+
```
118+
119+
**Purpose**: Measure raw provider performance.
120+
121+
### Magic Mode
122+
123+
Tests providers through the ComputeSDK orchestrator.
124+
125+
```typescript
126+
import { compute } from 'computesdk';
127+
128+
compute.setConfig({ provider: 'e2b', ... });
129+
const sandbox = await compute.sandbox.create();
130+
```
131+
132+
**Purpose**: Measure the experience when using ComputeSDK's abstraction layer.
133+
134+
**Note**: Magic Mode includes additional latency from the ComputeSDK orchestrator (Tributary routing + Daemon protocol). This is intentional—it measures the real-world experience for ComputeSDK users.
135+
136+
## Environment & Infrastructure
137+
138+
### Test Runner
139+
140+
All benchmarks run on GitHub Actions hosted runners:
141+
142+
- **OS**: Ubuntu (latest LTS)
143+
- **CPU**: 2-core x86_64
144+
- **Memory**: 7 GB
145+
- **Network**: GitHub's shared datacenter network
146+
- **Location**: Azure US regions (GitHub's infrastructure)
147+
148+
### Network Considerations
149+
150+
Network latency between the GitHub runner and each provider's API endpoints varies. This is **intentional**—it reflects real-world conditions where developers call these APIs from various locations.
151+
152+
We do not:
153+
- Run from provider-specific regions to artificially reduce latency
154+
- Use dedicated/reserved network capacity
155+
- Retry failed requests (failures count against success rate)
156+
157+
## Fairness & Limitations
158+
159+
### What This Benchmark Shows
160+
161+
- Relative performance between providers under consistent conditions
162+
- Typical cold-start times for on-demand sandbox creation
163+
- Provider reliability (success rate over time)
164+
165+
### What This Benchmark Does NOT Show
166+
167+
- Performance with pre-warmed pools or snapshots
168+
- Performance under high concurrency (coming Q2 2026)
169+
- Geographic variation (coming Q3 2026)
170+
- Cost efficiency
171+
- Feature differences between providers
172+
173+
### Provider-Specific Notes
174+
175+
Some providers offer optimizations that aren't captured in our default test:
176+
177+
| Provider | Available Optimization | Benchmark Status |
178+
|----------|----------------------|------------------|
179+
| E2B | Snapshots | Not tested (yet) |
180+
| Daytona | Templates | Not tested (yet) |
181+
| Modal | Warm containers | Not tested (yet) |
182+
| Namespace | Liquid pools | Not tested (yet) |
183+
184+
We plan to add warm-start benchmarks in Q3 2026.
185+
186+
## Data & Reproducibility
187+
188+
### Raw Data
189+
190+
All benchmark results are committed to this repository:
191+
192+
```
193+
results/
194+
├── 2026-02-19T00-30-31-832Z.json # Magic mode results
195+
├── direct-2026-02-19T00-30-31-832Z.json # Direct mode results
196+
└── ...
197+
```
198+
199+
### JSON Schema
200+
201+
```json
202+
{
203+
"timestamp": "ISO 8601 timestamp",
204+
"results": [
205+
{
206+
"provider": "provider-name",
207+
"iterations": [
208+
{ "ttiMs": 123.45 },
209+
{ "ttiMs": 0, "error": "error message" }
210+
],
211+
"summary": {
212+
"ttiMs": {
213+
"min": 100.0,
214+
"max": 150.0,
215+
"median": 125.0,
216+
"avg": 124.5
217+
}
218+
},
219+
"skipped": false,
220+
"skipReason": null
221+
}
222+
]
223+
}
224+
```
225+
226+
### Running Locally
227+
228+
Reproduce our results:
229+
230+
```bash
231+
git clone https://github.com/computesdk/benchmarks.git
232+
cd benchmarks
233+
npm install
234+
cp env.example .env # Add your API keys
235+
236+
# Run with same settings as CI
237+
npm run bench:direct -- --iterations 10
238+
```
239+
240+
**Note**: Your results will differ based on your network location and conditions.
241+
242+
## Changelog
243+
244+
| Date | Change |
245+
|------|--------|
246+
| 2026-02-19 | Initial methodology documentation |
247+
| 2026-02-01 | Increased default iterations from 3 to 10 |
248+
| 2026-01-15 | Added Direct Mode benchmarks |
249+
250+
## Questions & Disputes
251+
252+
Providers or users who have questions about methodology or wish to dispute results should open a GitHub issue. We commit to:
253+
254+
- Responding within 5 business days
255+
- Investigating any reproducible discrepancies
256+
- Updating methodology if we identify unfairness
257+
- Publishing corrections if errors are found
258+
259+
Contact: benchmarks@computesdk.com

0 commit comments

Comments
 (0)