Measuring the plan, without watching you

The plan is in the issues, starting with the master plan, #1. It has success metrics in it. This page says how each one gets measured, and — because the acceptance criterion is that no metric is published on the site without data behind it — it says which ones currently have data and which ones are still empty.

It does not add telemetry. There is no metrics client in the CLI, no counters, no analytics. Every number below comes from a stopwatch held by a person, from a log that already exists for another reason, or from a benchmark that runs in our own CI. If a number cannot be produced without the tool phoning home, it is not a metric we collect; it is a metric we ask for.

The rule

One rule decides whether a metric may be published, and it is not a judgement call:

A metric may be published whenIt may not be published when
A collection method above has produced at least one recorded readingThe method exists but has never been run
The reading carries what it was measured on — version, corpus, hardwareThe number is a budget, a target or an estimate
Someone can re-run the method and get a comparable numberOnly our own CI produced it and nobody else could check it

A budget is not a measurement. docs/adr/0013-rust-workspace.md pins the check budget at cold 1 s and warm p95 100 ms; that is a ceiling the CI job enforces, not evidence that anyone hit it. The distinction is the same one CALIBRATION.md makes between a confidence that has a report behind it and one that is only an example.

The four metrics

MetricWhat countsHow it is measuredData today
Time to first shipWall-clock from the first command to the first successful deploy, per onboarding sessionStopwatch, written down by the person being onboardedNone — no sessions run yet
Repos shippingDistinct repos with a ship.ks that completed a deployCounted by the hosted services where a repo uses them; otherwise self-reportedNone
check speed and local/CI parityWall-clock of keepshipping check; whether the same file behaves the same on a laptop and in CIcargo bench in our own CI; parity read off run logs a person comparesSpeed: yes. Parity: none
Product telemetryNothing, until someone asks for itOpt-in and anonymous (#156)Does not exist

Three of the four are empty. That is the honest state of the project, and it is recorded here rather than left implied by silence.

Time to first ship

A stopwatch, not a log. The measurement is taken by the person being onboarded, written down by hand at the end of the session, and reported as one number with the two timestamps: when they started (first command run, before they had installed anything) and when the first deploy succeeded. Sessions that never reached a deploy are recorded as did not finish with the step they stopped at, because a time to first ship that only counts the successes hides exactly the number worth knowing.

The report is one line per session. There is no session identifier, no account, and nothing sent to us unless the person sends it: the person being onboarded is the measurement instrument, and an instrument that reports to the vendor is a different instrument.

There is no data, and the reason is worth stating plainly: there is no onboarding path yet. CLI.md says that two of the three subcommands are implemented end to end and run is not, so there is no documented path from nothing to a first deploy to put a stopwatch around. The metric is defined and ready; the measurement starts when a session can be run. Publishing a number before then would mean publishing ours, not a user's.

Repos shipping through Keep Shipping

A repo counts when a ship.ks in it completes a deploy — not when it is created, not when it is checked, not when it is imported. Two ways to count it, and they are not interchangeable:

The two are reported as separate counts and are never added into one headline number. A total across them would imply a coverage we do not have: the self-reported count is a floor, not a sample, and the gap between the two is itself the honest answer to "how many repos ship through this".

There is no data. The hosted surface in this repository is the early-access waitlist, which counts people who signed up, not repos that deployed. There is no repo-counting endpoint, and adding one purely to produce this metric would be the surveillance this page rules out — so it stays self-reported until a control plane that already knows the answer exists.

check speed and local/CI parity

Speed is the one metric with data behind it, and it is measured by a benchmark that already gates the build.

$ cargo bench -p ks-cli --bench check
cold:  6.008ms (budget 1s)
warm p95 of 50: 6.172ms (budget 100ms)
script cold:  89.070ms (budget 500ms)
script warm p95 of 50: 4.928ms (budget 50ms)

Measured 2026-10-08 on 3 vCPU AMD EPYC, Linux x86_64, against the CLI test fixtures (crates/cli/tests/fixtures/site). The budgets are the ones in crates/cli/benches/check.rs and the ones .github/workflows/ci.yml fails the build on, so this number and the gate move together: when the budget is blown, CI says so and the number here is stale. Re-run the command above to refresh it. One caveat on that gate: the bench job runs on ubuntu-latest only, so it does not cover the macOS leg of the gates matrix.

Parsing is measured the same way, and reported here with its caveat:

$ cargo bench -p ks-lang
parsing a 522-line, 11 KiB source
median over 100 runs: 206 us

This one has a ceiling as well, enforced in crates/lang/benches/parse.rs: the run fails when the median exceeds 5 ms, and the nightly job — .github/workflows/nightly.yml, bench (500-line parse under 5 ms) — fails the build when that happens. 206 us against a 5 ms budget is about 24x of headroom. The two gates are not the same shape, though. The check bench runs on every push; the parse bench runs nightly. A parse regression is therefore caught a day late rather than immediately, which is a real difference between the two numbers above and worth keeping in mind before quoting either as "fast".

Local/CI parity has no data, and the gap is a wiring gap rather than a measurement gap. Parity here means what CI.md means by it — a run on Buildkite and the same run on a laptop are the same run, per docs/adr/0005-local-first-core-with-optional-hosted-services.md. The run-context modules for GitHub Actions and GitLab CI are built and tested, but the command dispatch still has no run context to hand them (crates/cli/src/main.rs), so the two have not been compared on a real deploy. When they are, the reading is a pair of logs from the same ship.ks, compared by a person — which is the same self-reported shape as the repo count, and is counted the same way.

Product telemetry

There is no product telemetry, and this section exists to say what adding one would owe.

If it is ever added (#156), it is opt-in and anonymous: off by default, carrying no identifier, no repo name, no file contents and no machine identity, and documented here before it ships rather than after. A metric that arrives in this file after the code that collects it has landed has already been collected without anyone's consent, which makes the consent meaningless.

The test a proposed metric has to pass is the one this page has been applying all along: could the same number be produced without the tool watching the person? If yes, that is how it is produced. If no, it is a privacy decision for the person who owns the tool, made in the open, and the default stays off.

What this does not measure

Where to read more

ForRead
The plan and the metrics it commits to#1
What check and the CLI actually doCLI.md
What parity means, and what --ci reportsCI.md
Why a budget is not a measurementCALIBRATION.md
Why the local path exists at alladr/0005
Moving to Keep Shipping without a migration projectCOEXISTENCE.md