# Benchable Store and visualize arbitrary benchmarks — latency, build time, bundle size — with statistically tested verdicts, traces and budgets. Self-hostable. Currently v0.1.0. Post any benchmark from CI — `go test`, `hyperfine`, `k6`, Lighthouse — or let your coding agent record it, and get a statistically tested verdict, the span that moved and the commits that did it. Self-hostable, on [Postgres](https://www.postgresql.org), by [linesofcode](https://x.com/linesofcode). ## CI fails on a header Gate merges without writing a parser. `POST /api/v1/runs` answers with a verdict per metric and `X-Benchable-Regressions`. See [sending a run](/reference#sending-a-run). ```sh curl -X POST http://localhost:3000/api/v1/runs \ -H "Authorization: Bearer $BENCHABLE_KEY" \ -H "Content-Type: application/json" \ -H "Idempotency-Key: $GITHUB_RUN_ID" \ -d '{ "branch": "main", "commitSha": "9f3c1ab", "environment": "ci-linux-x64", "metadata": { "runner": "github-actions", "node": "24" }, "metrics": { "build.time_ms": 4210, "api.latency_ms": { "value": 118.4, "p50": 110, "p95": 180, "p99": 260, "samples": 500 } }, "spans": [ { "id": "root", "name": "POST /checkout", "startMs": 0, "durationMs": 118.4 }, { "id": "db", "parentId": "root", "name": "SELECT orders", "startMs": 12, "durationMs": 40 } ] }' ``` ## Typed client, one file Autocomplete on the payload the server validates. The `benchable` SDK has no dependencies — vendor it and call `.run()`. See [SDK](/reference#sdk). ```ts import { Benchable } from "benchable"; const benchable = new Benchable({ apiKey: process.env.BENCHABLE_KEY! }); const result = await benchable.run({ branch: "main", commitSha: process.env.GITHUB_SHA, metrics: { "api.latency_ms": { value: 118.4, p95: 180, p99: 260 } }, }); if (result.regressions > 0) process.exit(1); ``` ## Only real changes fail Fewer false alarms, so nobody mutes the channel. Send `stddev` or raw `values` and the verdict comes from Welch's t-test or Mann–Whitney U, with false-discovery control across the suite. See [is the change real?](/reference#is-the-change-real). - metrics tested: 312 - flagged by raw p-value: 15 - flagged after control: 2 ```json "fdr": { "applied": true, "tested": 312, "flaggedBefore": 15, "flaggedAfter": 2, "q": 0.05 } ``` ```json { "key": "api.latency_ms", "value": 106, "baseline": 100, "deltaPct": 6, "verdict": "neutral", "reason": "not significant (p = 0.287)", "significance": { "test": "welch", "p": 0.287, "effectSize": 0.52, "interval": { "low": -5.9, "high": 17.9, "level": 0.95 }, "significant": false } } ``` ## Noise bands measure themselves Thresholds that match your hardware, not a guess. Each metric learns its own `stability profile`; switch the band to **Measured** to gate on it. See [noise bands](/reference#noise-bands-the-metric-measures-for-itself). - api.latency_ms measured noise: ±9.4% - runs of history: 46 - cold_start — too unstable to gate: ±10.5% ``` GET /api/v1/metrics -H 'Accept: text/plain' key name unit better noise band stability -------------------- ---------------- ---- ------ ---------- --------------------------------------- api.latency_ms Api Latency ms lower ±9.4% auto noisy, run-to-run noise ±9.4% over 46 runs bundle.main_kb Bundle Main KB lower ±5% stable, run-to-run noise ±0.0% over 40 runs cold_start.p99_ms Cold Start P99 ms lower ±5% flaky — too unstable to gate on, ±10.5% ``` ## Laptops never race CI A hardware change stops looking like a regression. Baselines are searched narrowest-first, and `baselineScope` says which one matched. See [baselines](/reference#baselines-that-compare-like-with-like). | `baselineScope` | Meaning | | --- | --- | | `branch+environment` | Same branch, same environment — a like-for-like comparison | | `branch` | Same branch, another environment | | `default-branch+environment` | The default branch, same environment | | `default-branch` | The default branch, another environment | ## Post the file you have Adopt it this afternoon. No translation shim. `POST /api/v1/import?format=auto` detects the format from the content. See [importing](/reference#importing-your-tools-output). Formats: `benchable`, `go-bench`, `hyperfine`, `pytest-benchmark`, `google-benchmark`, `criterion`, `vitest-bench`, `k6`, `lighthouse`, `jmh`, `prometheus`, `csv`, `otlp-trace`, `jaeger`, `zipkin`, `chrome-trace` ```sh go test -bench=. -benchmem ./... > bench.txt curl -X POST "$BENCHABLE_URL/api/v1/import?branch=main&commitSha=$(git rev-parse HEAD)" \ -H "Authorization: Bearer $BENCHABLE_KEY" \ -H "Accept: text/plain" \ --data-binary @bench.txt ``` ``` run WJ6S2WQYUr8NCdyWUSvLH on main (go-bench) metric value baseline change -------------------------------- ------------ -------- --------- go.BenchmarkEncode.ns_per_op 1.05µs — first run go.BenchmarkEncode.B_per_op 512 B — first run go.BenchmarkStream.MB_per_s 452 MB/s — first run No regressions. ``` ## One regression, one alert Your channel fires once, not forty times. Later detections bump `occurrences` silently; a run back at the old level resolves it. See [regressions](/reference#regressions-are-states-not-events). ``` GET /api/v1/regressions -H 'Accept: text/plain' id metric status kind branch now was change seen since -------------------- ------------------ ------ ---------- ------ ----- ----- ------- ---- ---------- kf3ZcUxhw3I3EKIcBKJI api.latency_ms open regression main 155ms 100ms +55.0% 3x 2026-05-02 llT2cPh0mA26KPed4cQz bundle.main_kb open budget main 310KB 240KB +29.2% 7x 2026-04-28 ``` ## Find the query that moved From "checkout is slow" to the exact query. Attach `spans`, or POST an OTLP, Jaeger, Zipkin or Chrome trace — the run page renders every trace as a filterable waterfall, and `GET /api/v1/compare` ranks both sides by each span's own time. See [which span moved](/reference#which-span-moved). ``` trace 118.0ms → 196.0ms span self total change state ------------------ ------- ------- ------- ------- SELECT orders +78.0ms 118.0ms +195.0% changed INSERT audit_log — 12.0ms — added POST /checkout −12.0ms 196.0ms +66.1% changed Ranked by change in the span's own work, which is what attributes a regression. ``` ## Hold the absolute line The bundle cannot grow 1% a day forever. `budgetMax` and `budgetMin` fail a run on its value, whatever the baseline says. See [budgets](/reference#performance-budgets). ``` metric value baseline change budget -------------------- ------ -------- --------- ---------------- budget.bundle_kb 251 KB 240 KB · 4.58% over by 0% api.latency_ms 118ms 120ms · −1.67% 32% left ``` ## Catch the slow creep Hand the range to git log. The cause is inside. `GET /api/v1/metrics/{key}/changepoints` finds each level shift and its commit range. See [change points](/reference#change-points). ``` search.throughput_ops on main — 40 runs analysed when level change from to commit range p ---------- --------- ------ ------------ ------------ ---------------- ------ 2026-09-10 regressed -6.30% 23.67k ops/s 22.18k ops/s aaaaaaa..aaa0aaa <0.001 2026-09-14 regressed -5.63% 22.32k ops/s 21.07k ops/s aaaf1da..a2aafab 0.002 The cause of each change is in the commits inside its range. ``` ## Your agent reads it too Let Claude or Cursor triage the regression queue. An MCP server at `/api/mcp` answers in text, not JSON. See [agents](/reference#agents). ```json { "mcpServers": { "benchable": { "url": "https://benchable.example.com/api/mcp", "headers": { "Authorization": "Bearer bmk_..." } } } } ``` ## Agents benchmark, you get charts Ask whether it got slower. Get a tested answer and a link. `npx skills add TimMikeladze/benchable` teaches Claude Code, Codex and OpenCode to `benchable record` before and after a change, online or offline. See [agent skill](/reference#agent-skill). ```text metric value baseline change verdict test ---------------------- ----- -------- -------- -------- -------------------- hyperfine.parse.mean_ms 9.0ms 14.6ms ▼ -38.4% improved mann-whitney p<0.001 view: https://benchable.sh/p/web/runs/r_8c1… ``` ## Full documentation - Reference (the README in full): https://benchable.sh/reference - Repository: https://github.com/TimMikeladze/benchable - This landing page as Markdown: https://benchable.sh/index.md ## API Every endpoint below takes a project API key as a bearer token: Authorization: Bearer bmk_... The key scopes every request to one project. Create one in the project's Settings tab. ## MCP There is an MCP server at https://benchable.sh/api/mcp (Streamable HTTP, stateless). Use the same API key as the bearer token. Tools: project_summary, project_report, list_metrics, get_metric_series, list_runs, get_run, compare_runs, list_branches, list_artifacts, submit_run, import_run, explain_regression, list_comments, post_comment, resolve_comment. The team discusses results in comment threads. Read list_comments before post_comment so you add to the discussion rather than repeat a teammate, and resolve_comment when it is settled. Client config: { "benchable": { "url": "https://benchable.sh/api/mcp", "headers": { "Authorization": "Bearer bmk_..." } } } ## Plain text Every read endpoint honours `Accept: text/plain` or `?format=text` and answers with a compact digest instead of JSON. Prefer it — it is a fraction of the tokens and needs no parsing. ## Endpoints POST https://benchable.sh/api/v1/runs Record a run. Body: { branch?, commitSha?, environment?, label?, startedAt?, metadata?, metrics: { "": number | { value, unit?, direction?, samples?, min?, max?, mean?, p50?, p95?, p99?, stddev?, labels? } }, spans?: [{ id, parentId?, name, startMs, durationMs, status?, attributes? }] } Send an Idempotency-Key header (or idempotencyKey in the body) so a retried CI step returns the original run instead of creating a duplicate. Returns 201 with each metric's baseline, deltaPct and verdict (improved|regressed|neutral), plus a regressions count — enough to fail a build without a second request. Also url (the run page) and shareUrl (only when the project is already shared); hand these to the person. POST https://benchable.sh/api/v1/import?format=auto Send a benchmark tool's own output as the raw body. Query: format, branch, commitSha, environment, label, idempotencyKey. GET https://benchable.sh/api/v1/import POST https://benchable.sh/api/v1/connect body { client: "claude-code"|"codex"|"opencode"|"cli"|"other" } POST https://benchable.sh/api/v1/connect/token body { deviceCode } `benchable login` as a device-code flow, no key needed to start. Show the person verificationUrl and userCode; poll token every 2s. It answers pending|denied|expired, then approved with a project API key exactly once. Lists the supported formats and how to produce each one. GET https://benchable.sh/api/v1/summary Every metric with its latest value and change, the newest run, and the branches reporting. The cheapest way to see the state of a project. GET https://benchable.sh/api/v1/report?period=daily|weekly Where the project stands over a window: health score and components, runs that landed, open/new/resolved regressions, and the metrics that moved or sit over budget. Ask this when the question is "how are the benchmarks doing" rather than "what are the numbers". GET https://benchable.sh/api/v1/metrics GET https://benchable.sh/api/v1/metrics/{key}/series?branch=&limit=&format=json|text|csv GET https://benchable.sh/api/v1/metrics/{key}/changepoints?branch=&environment=&minShiftPct= Where the metric changed level, each with the commit range that contains the cause. Use this, not two-run comparison, to answer "why is this slower than it used to be": it finds slow drift and steps buried in noise, which per-run verdicts cannot see. Without environment, the environment of the newest run on the branch is used. GET https://benchable.sh/api/v1/runs?branch=&environment=&limit=&cursor= GET https://benchable.sh/api/v1/runs/{runId} GET https://benchable.sh/api/v1/runs/{runId}/report?format=markdown|text&marker=&full= The run as a Markdown comment for a pull request: headline verdict, what moved, budgets, and the project's open regressions. The marker parameter embeds an HTML comment so a CI job can find and update its own comment instead of stacking a new one on every push. GET https://benchable.sh/api/v1/compare?base={runId}&head={runId} Metric-by-metric difference, plus a span-level trace diff when both runs carried a trace. The trace diff turns "checkout got 80ms slower" into "SELECT orders got 78ms slower"; spans are matched by path and ranked by change in their own work. POST https://benchable.sh/api/v1/runs/{runId}/analysis AI diagnosis of one run. 503 with code "ai_unavailable" when the server has no AI key. POST https://benchable.sh/api/v1/runs/{runId}/artifacts?name={filename}&kind={kind} Attach a file to a run: a flamegraph, a profile, a Lighthouse report, a screenshot. The body is the raw file, so `curl --data-binary @flame.svg` is the whole call; a multipart/form-data body with a `file` field works too. kind is optional and inferred from the name when omitted: image, flamegraph, profile, trace, log, report, file. Re-uploading a name replaces that file rather than adding a second one. GET https://benchable.sh/api/v1/runs/{runId}/artifacts What is attached to a run: name, kind, size, sha256 and a URL each. The sha256 is the one recorded at upload — verify a download against it. GET https://benchable.sh/api/v1/artifacts/{artifactId} The bytes. Requires the same bearer token — artifacts are never public. DELETE https://benchable.sh/api/v1/artifacts/{artifactId} GET https://benchable.sh/api/v1/regressions?status=&branch= Regression records: one per problem, not one per detection. status defaults to the live ones (open, acknowledged); pass resolved, wontfix, flaky or a comma-separated list. POST https://benchable.sh/api/v1/regressions/{id} Body: { status?, note? }. wontfix and flaky suppress new records for that metric until a person reopens it, so do not use them to quieten something you have not understood. note: null clears the note. Reopening a record while a newer one tracks the same metric answers 409 conflict with details.liveId; triage that record instead. GET https://benchable.sh/api/v1/comments?runId=&metricKey=&open=true|false&limit= Discussion threads, newest first, up to limit (default 50, max 200). open=true keeps the unresolved ones, open=false the resolved ones; it combines with runId and metricKey. POST https://benchable.sh/api/v1/comments Leave a comment. Body: { body, runId?, metricKey?, parentId?, author? }. Anchor it to a run, a benchmark, one measurement (both), or the project (neither). Comments posted with an API key are shown as automated and attributed to the key's name, never to a person. POST https://benchable.sh/api/v1/comments/{id}/resolve Body: { resolved: boolean }. GET https://benchable.sh/api/health ## Import formats - benchable: The native payload: a metrics object, with optional spans and metadata. produce: echo '{"metrics":{"build.time_ms":4210}}' - otlp-trace: OTLP/JSON ResourceSpans; becomes trace waterfalls plus summary timings. produce: Export OTLP/JSON from your collector, or POST the payload your app already sends. - jaeger: Jaeger JSON export; becomes trace waterfalls plus summary timings. produce: Jaeger UI → Share → Download JSON, or `GET /api/traces` from your Jaeger query service. - zipkin: Zipkin v2 JSON; becomes trace waterfalls plus summary timings. produce: `GET /api/v2/traces` from your Zipkin server, or POST /api/v2/spans output. - chrome-trace: Trace-events JSON (DevTools Performance, Perfetto); becomes a trace waterfall. produce: DevTools → Performance → record → Export JSON, or Perfetto's trace-events JSON. - lighthouse: Category scores and core web vitals from a Lighthouse JSON report. produce: lighthouse https://example.com --output=json --output-path=lhr.json - google-benchmark: C++ microbenchmarks, folding mean/median/stddev aggregates into one metric. produce: ./bench --benchmark_format=json --benchmark_repetitions=5 > bench.json - pytest-benchmark: Python benchmark stats from pytest-benchmark, with median, IQR and ops/s. produce: pytest --benchmark-json=bench.json - hyperfine: Command-line benchmark timings from hyperfine, with mean, stddev, min and max. produce: hyperfine --export-json bench.json './build.sh' - jmh: Java microbenchmarks from JMH, with score error and percentiles. produce: java -jar benchmarks.jar -rf json -rff bench.json - k6: Load-test summary from k6, with request percentiles, rates and check counts. produce: k6 run --summary-export=summary.json script.js - vitest-bench: JavaScript microbenchmarks from `vitest bench --outputJson`, or raw tinybench. produce: vitest bench --outputJson=bench.json - criterion: Rust microbenchmarks from Criterion's NDJSON output, using typical and median. produce: cargo criterion --message-format=json > bench.ndjson - go-bench: Text output from `go test -bench`, including -benchmem counters. produce: go test -bench=. -benchmem ./... | tee bench.txt - prometheus: A /metrics scrape in text exposition format; quantiles fold into percentiles. produce: curl -s http://localhost:9090/metrics > metrics.txt - csv: Two to four columns: name, value, optional unit, optional lower|higher. produce: printf "build.time_ms,4210,ms\napi.rps,1840,ops/s,higher\n" > bench.csv ## Errors Every error is { "error": { "code", "message", "details"? }, "requestId" }. Codes: unauthorized, forbidden, not_found, invalid_payload, invalid_query, payload_too_large, format_unrecognized, rate_limited, ai_unavailable, internal_error. Rate limits are per API key and reported in X-RateLimit-Limit / -Remaining / -Reset. A 429 carries Retry-After. ## Verdicts A metric's direction decides whether a rise is good. A change smaller than the metric's noise band (5% by default, editable per metric) is "neutral". A metric's baseline is the newest earlier run on the same branch that reported that metric, falling back to the project's default branch. It is resolved per metric, so interleaving two benchmark suites on one branch does not break comparisons.