# Benchable

Store and visualize arbitrary benchmarks — latency, build time, bundle size — with statistically tested verdicts, traces and budgets. Self-hostable.

> Catch the slowdown before it merges — currently v0.1.0.

Post any benchmark from CI — `go test`, `hyperfine`, `k6`, Lighthouse — or let your coding agent record it, and get a statistically tested verdict, the span that moved and the commits that did it. Self-hostable, on [Postgres](https://www.postgresql.org), by [linesofcode](https://x.com/linesofcode).

## CI fails on a header

Gate merges without writing a parser. `POST /api/v1/runs` answers with a verdict per metric and `X-Benchable-Regressions`. See [sending a run](/reference#sending-a-run).

```sh
curl -X POST http://localhost:3000/api/v1/runs \
  -H "Authorization: Bearer $BENCHABLE_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: $GITHUB_RUN_ID" \
  -d '{
    "branch": "main",
    "commitSha": "9f3c1ab",
    "environment": "ci-linux-x64",
    "metadata": { "runner": "github-actions", "node": "24" },
    "metrics": {
      "build.time_ms": 4210,
      "api.latency_ms": { "value": 118.4, "p50": 110, "p95": 180, "p99": 260, "samples": 500 }
    },
    "spans": [
      { "id": "root", "name": "POST /checkout", "startMs": 0, "durationMs": 118.4 },
      { "id": "db", "parentId": "root", "name": "SELECT orders", "startMs": 12, "durationMs": 40 }
    ]
  }'
```

## Typed client, one file

Autocomplete on the payload the server validates. The `benchable` SDK has no dependencies — vendor it and call `.run()`. See [SDK](/reference#sdk).

```ts
import { Benchable } from "benchable";

const benchable = new Benchable({ apiKey: process.env.BENCHABLE_KEY! });

const result = await benchable.run({
  branch: "main",
  commitSha: process.env.GITHUB_SHA,
  metrics: { "api.latency_ms": { value: 118.4, p95: 180, p99: 260 } },
});

if (result.regressions > 0) process.exit(1);
```

## Only real changes fail

Fewer false alarms, so nobody mutes the channel. Send `stddev` or raw `values` and the verdict comes from Welch's t-test or Mann–Whitney U, with false-discovery control across the suite. See [is the change real?](/reference#is-the-change-real).

- metrics tested: 312
- flagged by raw p-value: 15
- flagged after control: 2

```json
"fdr": { "applied": true, "tested": 312, "flaggedBefore": 15, "flaggedAfter": 2, "q": 0.05 }
```

```json
{ "key": "api.latency_ms", "value": 106, "baseline": 100, "deltaPct": 6,
  "verdict": "neutral", "reason": "not significant (p = 0.287)",
  "significance": { "test": "welch", "p": 0.287, "effectSize": 0.52,
                    "interval": { "low": -5.9, "high": 17.9, "level": 0.95 },
                    "significant": false } }
```

## Noise bands measure themselves

Thresholds that match your hardware, not a guess. Each metric learns its own `stability profile`; switch the band to **Measured** to gate on it. See [noise bands](/reference#noise-bands-the-metric-measures-for-itself).

- api.latency_ms measured noise: ±9.4%
- runs of history: 46
- cold_start — too unstable to gate: ±10.5%

```
GET /api/v1/metrics -H 'Accept: text/plain'

key                    name              unit  better  noise band  stability
--------------------  ----------------  ----  ------  ----------  ---------------------------------------
api.latency_ms        Api Latency       ms    lower   ±9.4% auto  noisy, run-to-run noise ±9.4% over 46 runs
bundle.main_kb        Bundle Main       KB    lower   ±5%         stable, run-to-run noise ±0.0% over 40 runs
cold_start.p99_ms     Cold Start P99    ms    lower   ±5%         flaky — too unstable to gate on, ±10.5%
```

## Laptops never race CI

A hardware change stops looking like a regression. Baselines are searched narrowest-first, and `baselineScope` says which one matched. See [baselines](/reference#baselines-that-compare-like-with-like).

| `baselineScope` | Meaning |
| --- | --- |
| `branch+environment` | Same branch, same environment — a like-for-like comparison |
| `branch` | Same branch, another environment |
| `default-branch+environment` | The default branch, same environment |
| `default-branch` | The default branch, another environment |

## Post the file you have

Adopt it this afternoon. No translation shim. `POST /api/v1/import?format=auto` detects the format from the content. See [importing](/reference#importing-your-tools-output).

Formats: `benchable`, `go-bench`, `hyperfine`, `pytest-benchmark`, `google-benchmark`, `criterion`, `vitest-bench`, `k6`, `lighthouse`, `jmh`, `prometheus`, `csv`, `otlp-trace`, `jaeger`, `zipkin`, `chrome-trace`

```sh
go test -bench=. -benchmem ./... > bench.txt

curl -X POST "$BENCHABLE_URL/api/v1/import?branch=main&commitSha=$(git rev-parse HEAD)" \
  -H "Authorization: Bearer $BENCHABLE_KEY" \
  -H "Accept: text/plain" \
  --data-binary @bench.txt
```

```
run WJ6S2WQYUr8NCdyWUSvLH on main (go-bench)

metric                                   value  baseline     change
--------------------------------  ------------  --------  ---------
go.BenchmarkEncode.ns_per_op            1.05µs         —  first run
go.BenchmarkEncode.B_per_op              512 B         —  first run
go.BenchmarkStream.MB_per_s           452 MB/s         —  first run

No regressions.
```

## One regression, one alert

Your channel fires once, not forty times. Later detections bump `occurrences` silently; a run back at the old level resolves it. See [regressions](/reference#regressions-are-states-not-events).

```
GET /api/v1/regressions -H 'Accept: text/plain'

id                     metric              status  kind        branch    now    was   change  seen   since
--------------------  ------------------  ------  ----------  ------  -----  -----  -------  ----  ----------
kf3ZcUxhw3I3EKIcBKJI  api.latency_ms      open    regression  main    155ms  100ms   +55.0%    3x  2026-05-02
llT2cPh0mA26KPed4cQz  bundle.main_kb      open    budget      main    310KB  240KB   +29.2%    7x  2026-04-28
```

## Find the query that moved

From "checkout is slow" to the exact query. Attach `spans`, or POST an OTLP, Jaeger, Zipkin or Chrome trace — the run page renders every trace as a filterable waterfall, and `GET /api/v1/compare` ranks both sides by each span's own time. See [which span moved](/reference#which-span-moved).

```
trace 118.0ms → 196.0ms

span                   self    total   change    state
------------------  -------  -------  -------  -------
  SELECT orders     +78.0ms  118.0ms  +195.0%  changed
  INSERT audit_log        —   12.0ms        —    added
POST /checkout      −12.0ms  196.0ms   +66.1%  changed

Ranked by change in the span's own work, which is what attributes a regression.
```

## Hold the absolute line

The bundle cannot grow 1% a day forever. `budgetMax` and `budgetMin` fail a run on its value, whatever the baseline says. See [budgets](/reference#performance-budgets).

```
metric                 value  baseline     change            budget
--------------------  ------  --------  ---------  ----------------
budget.bundle_kb      251 KB    240 KB     · 4.58%       over by 0%
api.latency_ms         118ms     120ms     · −1.67%      32% left
```

## Catch the slow creep

Hand the range to git log. The cause is inside. `GET /api/v1/metrics/{key}/changepoints` finds each level shift and its commit range. See [change points](/reference#change-points).

```
search.throughput_ops on main — 40 runs analysed

when            level  change          from            to      commit range       p
----------  ---------  ------  ------------  ------------  ----------------  ------
2026-09-10  regressed  -6.30%  23.67k ops/s  22.18k ops/s  aaaaaaa..aaa0aaa  <0.001
2026-09-14  regressed  -5.63%  22.32k ops/s  21.07k ops/s  aaaf1da..a2aafab   0.002

The cause of each change is in the commits inside its range.
```

## Your agent reads it too

Let Claude or Cursor triage the regression queue. An MCP server at `/api/mcp` answers in text, not JSON. See [agents](/reference#agents).

```json
{
  "mcpServers": {
    "benchable": {
      "url": "https://benchable.example.com/api/mcp",
      "headers": { "Authorization": "Bearer bmk_..." }
    }
  }
}
```

## Agents benchmark, you get charts

Ask whether it got slower. Get a tested answer and a link. `npx skills add TimMikeladze/benchable` teaches Claude Code, Codex and OpenCode to `benchable record` before and after a change, online or offline. See [agent skill](/reference#agent-skill).

```text
metric                  value  baseline    change   verdict  test
----------------------  -----  --------  --------  --------  --------------------
hyperfine.parse.mean_ms  9.0ms    14.6ms  ▼ -38.4%  improved  mann-whitney p<0.001

view: https://benchable.sh/p/web/runs/r_8c1…
```

## What holds, what is a guess

Holds:

- Evidence only ever downgrades a verdict — what gated CI yesterday still gates today.
- Ingest is idempotent: a retried step gets the original run back.
- API keys are stored as SHA-256 hashes. Artifacts are never public.

Judgements:

- The ±5% default band is a guess — switch the metric to Measured.
- False-discovery control runs at q = 0.05, the field default.
- Stability needs six runs of history before it is reported.

Not here yet:

- AI diagnosis needs `AI_GATEWAY_API_KEY`; without it you get `503 ai_unavailable`.
- No scheduler — CI owns the cadence, Benchable ingests.
- Cross-environment baselines are labelled, not corrected.
- The `benchable` package is not on npm yet — vendor the SDK file.

## Priced by how big it gets

The free plan is the entire product — every ingest format, statistical verdicts, noise bands, budgets, traces, regressions, AI diagnosis, artifacts, webhooks and the MCP server. What it caps is volume. Pro removes every cap, per member. Enterprise runs inside your own perimeter.

- **Free** — $0 forever. The whole product, sized for one repository. every ingest format, verdict and chart; aI diagnosis, artifacts, webhooks and share links; no card, no trial clock.
- **Pro** — $20 per member, per month. Everything, uncapped, billed by the people using it. unlimited projects, runs, members and history; retention you choose, not one we impose; seats follow your team — add someone, we prorate.
- **Enterprise** — Let's talk annual, invoiced. Run it inside your own perimeter. self-hosted, in your cloud account; vPC peering and private networking; sSO, SCIM and an audit log; custom retention, invoicing and support terms.

## Get your first verdict

Sign up, create a project, copy its key from **Settings**, and post a run. The first `POST` creates every metric in it.

1. Create a project
2. Copy the API key
3. Post a run from CI

Sign up: https://benchable.sh/sign-up. Rather run it yourself? One Postgres URL and Bun.

```sh
bun install
cp .env.example .env     # set DATABASE_URL and BETTER_AUTH_SECRET
bun run db:migrate
bun run db:seed          # optional: a demo workspace with 54 runs across 3 branches
bun run dev              # picks the first free port at or after 3000
```

```sh
export BENCHABLE_URL=https://benchable.example.com
export BENCHABLE_KEY=bmk_...

bunx benchable import  --file bench.txt --branch main --commit "$(git rev-parse HEAD)"
bunx benchable submit  --metrics metrics.json --branch main
bunx benchable summary
bunx benchable report --run last --marker pr-42   # Markdown for a PR comment
bunx benchable regressions --status live          # the regression queue
bunx benchable upload --run last --file flame.svg --kind flamegraph
bunx benchable artifacts --run last
bunx benchable download --run last --name flame.svg
bunx benchable formats
```

## Links

- Full reference: https://benchable.sh/reference
- Repository: https://github.com/TimMikeladze/benchable
- Pricing: https://benchable.sh/#pricing

Benchable is built by [linesofcode](https://linesofcode.dev). Every example on this page is read out of the README at build time — when the behaviour changes, the page breaks. © 2026 linesofcode
