|
| 1 | +# Benchmark Methodology |
| 2 | + |
| 3 | +This document describes how ComputeSDK Benchmarks measures sandbox provider performance. Our goal is transparent, reproducible, and fair measurement. |
| 4 | + |
| 5 | +## What We Measure |
| 6 | + |
| 7 | +### Time to Interactive (TTI) |
| 8 | + |
| 9 | +**Definition**: The wall-clock time from initiating a sandbox creation request to successfully executing the first command. |
| 10 | + |
| 11 | +TTI captures the complete developer experience: |
| 12 | + |
| 13 | +``` |
| 14 | +┌─────────────────────────────────────────────────────────────────────────┐ |
| 15 | +│ Time to Interactive (TTI) │ |
| 16 | +├─────────────┬─────────────────┬──────────────┬─────────────┬───────────┤ |
| 17 | +│ API Latency │ Provisioning │ Boot Time │ Health Check│ Command │ |
| 18 | +│ │ │ │ Polling │ Execution │ |
| 19 | +└─────────────┴─────────────────┴──────────────┴─────────────┴───────────┘ |
| 20 | +``` |
| 21 | + |
| 22 | +This metric matters because it's what developers actually experience—the time spent waiting before they can use the sandbox. |
| 23 | + |
| 24 | +### What's Included in TTI |
| 25 | + |
| 26 | +- Network round-trip to provider API |
| 27 | +- Queue time (if provider has provisioning queues) |
| 28 | +- Infrastructure allocation (VM, container, or serverless spin-up) |
| 29 | +- Operating system and runtime boot |
| 30 | +- Provider daemon/agent initialization |
| 31 | +- Health check and readiness polling |
| 32 | +- First command network round-trip |
| 33 | +- Command execution time (trivial for our test command) |
| 34 | + |
| 35 | +### What's NOT Included |
| 36 | + |
| 37 | +- Sandbox teardown/destruction time |
| 38 | +- Subsequent command execution times |
| 39 | +- File system operations |
| 40 | +- Network transfer speeds within the sandbox |
| 41 | + |
| 42 | +## Test Procedure |
| 43 | + |
| 44 | +Each benchmark iteration executes the following steps: |
| 45 | + |
| 46 | +```typescript |
| 47 | +// 1. Start timer |
| 48 | +const start = performance.now(); |
| 49 | + |
| 50 | +// 2. Create sandbox and wait until ready |
| 51 | +const sandbox = await compute.sandbox.create(); |
| 52 | + |
| 53 | +// 3. Execute a trivial command to confirm interactivity |
| 54 | +await sandbox.runCommand('echo "benchmark"'); |
| 55 | + |
| 56 | +// 4. Stop timer |
| 57 | +const ttiMs = performance.now() - start; |
| 58 | + |
| 59 | +// 5. Cleanup (not timed) |
| 60 | +await sandbox.destroy(); |
| 61 | +``` |
| 62 | + |
| 63 | +### Why `echo "benchmark"`? |
| 64 | + |
| 65 | +We use a minimal command to isolate sandbox startup time from command complexity. The command: |
| 66 | +- Has negligible execution time |
| 67 | +- Requires no file system access |
| 68 | +- Produces deterministic output |
| 69 | +- Validates the full request/response cycle |
| 70 | + |
| 71 | +## Test Configuration |
| 72 | + |
| 73 | +### Daily Automated Runs |
| 74 | + |
| 75 | +| Parameter | Value | |
| 76 | +|-----------|-------| |
| 77 | +| Iterations per provider | 10 | |
| 78 | +| Timeout per iteration | 120 seconds | |
| 79 | +| Run frequency | Daily at 00:00 UTC | |
| 80 | +| Runner environment | GitHub Actions (ubuntu-latest) | |
| 81 | +| Node.js version | 20.x | |
| 82 | + |
| 83 | +### Provider Execution Order |
| 84 | + |
| 85 | +Providers are tested **sequentially** to: |
| 86 | +- Avoid resource contention on the test runner |
| 87 | +- Prevent rate limiting issues |
| 88 | +- Ensure consistent network conditions per provider |
| 89 | + |
| 90 | +The order is randomized each run to prevent systematic bias from time-of-day effects. |
| 91 | + |
| 92 | +## Statistical Reporting |
| 93 | + |
| 94 | +For each provider, we report: |
| 95 | + |
| 96 | +| Metric | Description | |
| 97 | +|--------|-------------| |
| 98 | +| **Min** | Fastest iteration (best case) | |
| 99 | +| **Max** | Slowest iteration (worst case) | |
| 100 | +| **Median** | Middle value (typical case) | |
| 101 | +| **Average** | Arithmetic mean | |
| 102 | +| **Success Rate** | Iterations completed without error | |
| 103 | + |
| 104 | +We emphasize **median** as the primary metric because it's robust to outliers and represents the typical developer experience. |
| 105 | + |
| 106 | +## Benchmark Modes |
| 107 | + |
| 108 | +### Direct Mode |
| 109 | + |
| 110 | +Tests each provider's native SDK without any abstraction layer. |
| 111 | + |
| 112 | +```typescript |
| 113 | +import { E2B } from '@computesdk/e2b'; |
| 114 | + |
| 115 | +const compute = new E2B({ apiKey: process.env.E2B_API_KEY }); |
| 116 | +const sandbox = await compute.sandbox.create(); |
| 117 | +``` |
| 118 | + |
| 119 | +**Purpose**: Measure raw provider performance. |
| 120 | + |
| 121 | +### Magic Mode |
| 122 | + |
| 123 | +Tests providers through the ComputeSDK orchestrator. |
| 124 | + |
| 125 | +```typescript |
| 126 | +import { compute } from 'computesdk'; |
| 127 | + |
| 128 | +compute.setConfig({ provider: 'e2b', ... }); |
| 129 | +const sandbox = await compute.sandbox.create(); |
| 130 | +``` |
| 131 | + |
| 132 | +**Purpose**: Measure the experience when using ComputeSDK's abstraction layer. |
| 133 | + |
| 134 | +**Note**: Magic Mode includes additional latency from the ComputeSDK orchestrator (Tributary routing + Daemon protocol). This is intentional—it measures the real-world experience for ComputeSDK users. |
| 135 | + |
| 136 | +## Environment & Infrastructure |
| 137 | + |
| 138 | +### Test Runner |
| 139 | + |
| 140 | +All benchmarks run on GitHub Actions hosted runners: |
| 141 | + |
| 142 | +- **OS**: Ubuntu (latest LTS) |
| 143 | +- **CPU**: 2-core x86_64 |
| 144 | +- **Memory**: 7 GB |
| 145 | +- **Network**: GitHub's shared datacenter network |
| 146 | +- **Location**: Azure US regions (GitHub's infrastructure) |
| 147 | + |
| 148 | +### Network Considerations |
| 149 | + |
| 150 | +Network latency between the GitHub runner and each provider's API endpoints varies. This is **intentional**—it reflects real-world conditions where developers call these APIs from various locations. |
| 151 | + |
| 152 | +We do not: |
| 153 | +- Run from provider-specific regions to artificially reduce latency |
| 154 | +- Use dedicated/reserved network capacity |
| 155 | +- Retry failed requests (failures count against success rate) |
| 156 | + |
| 157 | +## Fairness & Limitations |
| 158 | + |
| 159 | +### What This Benchmark Shows |
| 160 | + |
| 161 | +- Relative performance between providers under consistent conditions |
| 162 | +- Typical cold-start times for on-demand sandbox creation |
| 163 | +- Provider reliability (success rate over time) |
| 164 | + |
| 165 | +### What This Benchmark Does NOT Show |
| 166 | + |
| 167 | +- Performance with pre-warmed pools or snapshots |
| 168 | +- Performance under high concurrency (coming Q2 2026) |
| 169 | +- Geographic variation (coming Q3 2026) |
| 170 | +- Cost efficiency |
| 171 | +- Feature differences between providers |
| 172 | + |
| 173 | +### Provider-Specific Notes |
| 174 | + |
| 175 | +Some providers offer optimizations that aren't captured in our default test: |
| 176 | + |
| 177 | +| Provider | Available Optimization | Benchmark Status | |
| 178 | +|----------|----------------------|------------------| |
| 179 | +| E2B | Snapshots | Not tested (yet) | |
| 180 | +| Daytona | Templates | Not tested (yet) | |
| 181 | +| Modal | Warm containers | Not tested (yet) | |
| 182 | +| Namespace | Liquid pools | Not tested (yet) | |
| 183 | + |
| 184 | +We plan to add warm-start benchmarks in Q3 2026. |
| 185 | + |
| 186 | +## Data & Reproducibility |
| 187 | + |
| 188 | +### Raw Data |
| 189 | + |
| 190 | +All benchmark results are committed to this repository: |
| 191 | + |
| 192 | +``` |
| 193 | +results/ |
| 194 | +├── 2026-02-19T00-30-31-832Z.json # Magic mode results |
| 195 | +├── direct-2026-02-19T00-30-31-832Z.json # Direct mode results |
| 196 | +└── ... |
| 197 | +``` |
| 198 | + |
| 199 | +### JSON Schema |
| 200 | + |
| 201 | +```json |
| 202 | +{ |
| 203 | + "timestamp": "ISO 8601 timestamp", |
| 204 | + "results": [ |
| 205 | + { |
| 206 | + "provider": "provider-name", |
| 207 | + "iterations": [ |
| 208 | + { "ttiMs": 123.45 }, |
| 209 | + { "ttiMs": 0, "error": "error message" } |
| 210 | + ], |
| 211 | + "summary": { |
| 212 | + "ttiMs": { |
| 213 | + "min": 100.0, |
| 214 | + "max": 150.0, |
| 215 | + "median": 125.0, |
| 216 | + "avg": 124.5 |
| 217 | + } |
| 218 | + }, |
| 219 | + "skipped": false, |
| 220 | + "skipReason": null |
| 221 | + } |
| 222 | + ] |
| 223 | +} |
| 224 | +``` |
| 225 | + |
| 226 | +### Running Locally |
| 227 | + |
| 228 | +Reproduce our results: |
| 229 | + |
| 230 | +```bash |
| 231 | +git clone https://github.com/computesdk/benchmarks.git |
| 232 | +cd benchmarks |
| 233 | +npm install |
| 234 | +cp env.example .env # Add your API keys |
| 235 | + |
| 236 | +# Run with same settings as CI |
| 237 | +npm run bench:direct -- --iterations 10 |
| 238 | +``` |
| 239 | + |
| 240 | +**Note**: Your results will differ based on your network location and conditions. |
| 241 | + |
| 242 | +## Changelog |
| 243 | + |
| 244 | +| Date | Change | |
| 245 | +|------|--------| |
| 246 | +| 2026-02-19 | Initial methodology documentation | |
| 247 | +| 2026-02-01 | Increased default iterations from 3 to 10 | |
| 248 | +| 2026-01-15 | Added Direct Mode benchmarks | |
| 249 | + |
| 250 | +## Questions & Disputes |
| 251 | + |
| 252 | +Providers or users who have questions about methodology or wish to dispute results should open a GitHub issue. We commit to: |
| 253 | + |
| 254 | +- Responding within 5 business days |
| 255 | +- Investigating any reproducible discrepancies |
| 256 | +- Updating methodology if we identify unfairness |
| 257 | +- Publishing corrections if errors are found |
| 258 | + |
| 259 | +Contact: benchmarks@computesdk.com |
0 commit comments