DiffCI

Measured locally · zero CI authority

Measure CI waste with change-aware test selection.

DiffCI performs conservative test impact analysis: it follows a code change through the dependency graph, identifies affected tests and CI jobs, and measures the selected path against the full path when both can be executed safely.

It has no permission to skip, cancel, or modify a CI run. Local checks may execute test commands, but required CI remains unchanged.

Change-aware CIShortest verified route

only work reached by the change · proof retained

10 real merges across two repositories, full and selected paths both executed and timed
12 of 13 executions where the runner honored the selection; the 13th is withheld, not counted
0 production CI runs ever altered

There is no headline percentage on this page. The honest number is repository-specific: the same selection quality produced 44% and 90% job-level reduction on two different real repositories. A single number would be an average of two things that aren't alike.

The problem

A twelve-line change to one package re-runs four thousand tests in ninety other packages.

Everyone knows this is happening. Almost nobody knows what it costs them, because measuring it means either trusting a tool to start skipping things, or designing an experiment nobody has time for.

DiffCI is the measurement, delivered before the risk.

Start with the test impact analysis guide, learn how to run affected tests conservatively, or compare graph selection with path filters and task-graph tools.

Install DiffCI

DiffCI has local, agent, Action and hosted-observer entry points. The local check command may run inferred test commands; none of the entry points can skip, cancel or reorder your required CI jobs.

AI coding agents

Give agents a default validation command

Seed AGENTS.md, CLAUDE.md, Cursor rules, Copilot instructions, and diffci.config.json. Agents then have one safe command to run before calling a change PR-ready.

npx "@diffci.com/diffci@latest" init
npx "@diffci.com/diffci@latest" check

DiffCI changes no CI behavior. check may run inferred local test commands to validate the change. Agent docs →

MCP server

Connect DiffCI as a native agent tool

Remote MCP clients can request a safe local validation plan without sending repository content. The local server also exposes diffci_check, diffci_init, and diffci_verify_workflow.

https://diffci.com/mcp

The HTTPS endpoint provides stateless, read-only validation guidance. Use npx -p "@diffci.com/diffci@latest" diffci-mcp when tools must access a local checkout. MCP setup and tools →

npm CLI

Run it locally or inside CI

The npm package is @diffci.com/diffci. Use npx for a one-off observation, or install it as a dev dependency if you want a pinned local tool.

npx "@diffci.com/diffci@latest" check
npx "@diffci.com/diffci@latest" observe
npx "@diffci.com/diffci@latest" verify-workflow

npm install --save-dev @diffci.com/diffci

observe writes a JSON report outside the checkout by default. Nothing is sent to DiffCI unless an endpoint and token are explicitly configured.

GitHub Action runner

Add one non-blocking observer job

Put DiffCI in its own job so a DiffCI failure cannot become your workflow's conclusion. For a reproducible installation, the example pins release v0.2.11 to its qualified feature commit SHA.

jobs:
  diffci:
    runs-on: ubuntu-latest
    continue-on-error: true
    permissions:
      contents: read
    steps:
      - uses: actions/checkout@v4
        with:
          fetch-depth: 0
      - uses: DiffCI/DiffCI.com@e1d7bab271c5d83899bda0034d70d3e34c10c1f7

The Action writes a report artifact, prints a summary, and cannot skip tests, cancel jobs, comment on pull requests, or change required checks.

Read-only GitHub App

Use shadow mode without editing a workflow

The App is the lowest-friction pilot path: it observes pushes and completed workflow runs from outside your CI, then reconciles what DiffCI would have selected against what your CI actually did.

How it works

  1. Observe

    A read-only GitHub App sees pushes and completed workflow runs. Metadata, contents, Actions, checks — read-only, all four.

  2. Analyze

    DiffCI builds a real TypeScript compiler-backed dependency graph, computes what is reachable from the changed files, and produces a confidence-scored plan: which tests a safe system would have run. Changes it cannot reason about — lockfiles, workflow files, root config — force a mandatory fallback to "run everything". That is the default, not the exception path.

  3. Reconcile

    When your CI finishes, DiffCI reads the real outcome and timings and asks the only question that matters in hindsight: would that plan have missed a failure your CI actually caught?

  4. Report

    Runs observed, compute consumed, compute avoidable, missed failures — and, explicitly, what could not be measured.

The report

Illustrative layout — placeholder values, not data from any repository
Last 7 days · your-org/your-repo
CI runs observed———measured
Runner-hours consumed———measured
Runner-hours potentially avoidable———estimated
Observed missed failures———measured
Selected-run execution———not measured

Three labels, and they mean exactly what they say. Measured is a number a clock produced. Estimated is an inference from a cost model whose weaknesses we will tell you about. Unknown is where a lesser tool would have guessed.

DiffCI will not turn an estimated figure into an invoice. Ever. If it eventually charges for anything, it charges against savings that were executed and timed.

Why there is no single number

Two real repositories. Comparable selection quality — one or two test files selected out of hundreds, in both. Wildly different savings, for a reason that has nothing to do with how good the selection was.

Install & setup — paid either way Tests — the part selection shrinks

cal.com

full
657.7s
selected
356.9s −44%

deepseek-harness

full
760.2s
selected
58.7s −90%
Measured wall-clock, same horizontal scale across both repositories. cal.com PR #29940 and deepseek-harness PR #2760, each run in full and with DiffCI's selection, in DiffCI's sandbox. cal.com spends 336.8s installing before a single test runs; deepseek-harness spends 38.6s. That fixed cost is the entire difference between a 44% job and a 90% one.

The evidence behind it

Nothing here comes from a demo. Every figure was produced by executing real test suites from real repositories, with the full run and the selected run both timed.

Case study

cal.com

86–91.6%test-stage reduction, net
44.2%complete-job reduction, net
8 of 8selections honored exactly
3 of 3measurable regressions caught

Five merges chosen by a rule written down before anyone looked at what they would cost, plus one measured six ways. We publish both reduction figures and lead with the smaller one.

Read the cal.com case study →

Case study

deepseek-ai/deepseek-harness

83.5–94.3%test-stage reduction, net
79.5–89.5%complete-job reduction, net
3 of 3measurable regressions caught
1 of 5results withheld, not counted

The repository that was allowed to be messy: 16–18 tests fail on every run before anyone changes anything. A naive safety check reads that as "the regression was caught" every single time. One merge's result is withheld entirely because the runner did not honor the selection.

Read the deepseek-harness case study →

Case study

Our own CI

$0.0048–$0.0127per CI job (median $0.0053)
1,960tests, currently green
0GitHub Actions minutes billed

Our suite runs on ephemeral Cloudflare containers, 2–6 minutes a run. We built our own runner fleet because we needed the timing data to be ours — and found a bug that had been hiding fifteen test files from our own green CI.

Read how →

Open study · CC BY 4.0

Every number, in one place

2,000real commit deltas, 20 repositories
29.8%of deltas where selection was even possible
4.2%aggregate reduction vs a path rule
0 of 32false greens, measurable mutations

Two weeks of pre-registered experiments, published as a study anyone can reuse with attribution: the page, a CSV of every figure with its evidence level and source, and a PDF. It leads with the number that argues against the product.

Read the open study →

Isn't this just Nx affected, Turborepo, or a path rule?

Those tools answer a different question. They tell you what to run. DiffCI's output is the part they don't ship: evidence, after the fact, that a selection would have been safe on your real CI, and what it would have been worth — reconciled against what your CI actually did.

We are not claiming DiffCI is cheaper than any of them. Whether DiffCI's own analysis cost nets out ahead of a simple path rule depends on the repository, and we have not measured it on enough external repositories to say. Where it can be computed, the pilot report includes that comparison for your repository and reports the sign either way.

Compare dependency-graph selection, path filters and task-graph tools →

What we won't tell you

This section is here on purpose, and it is not going to be moved to the footer.

  • Nobody's CI has ever been made faster by DiffCI. Nothing has ever been skipped in a production pipeline. The savings above were measured by executing both paths in our own sandbox.
  • There are no customers. cal.com, deepseek-harness and the unjs projects are public repositories analyzed from public data. None of them use DiffCI. None of them have endorsed it.
  • Our failure prediction is not good yet. Replayed against 24 of our own historical CI failures, the preflight checks would have caught 16 of 23 evaluable ones — 0.696. Every one it caught was a typecheck or configuration failure. Every unit-test failure in that dataset was a miss.
  • There is no external pilot yet. At the time of writing, no repository outside our own two has installed the App. An earlier internal shadow cohort of public repositories was small, and one repository in it turned out to be unanalysable and was excluded rather than quietly retried. None of that is traction, and we are not going to describe it as such.
  • Small selections make our estimator optimistic. A run selecting zero of forty tests still pays the suite's startup cost, and the estimator does not model that. It is being calibrated per repository from real measurement, not patched with a guessed constant.

The pilot

What you giveRead-only access to one repository, for seven days.
What you getA report of what your CI ran, what was avoidable, and whether any change DiffCI would have de-prioritized turned out to break something.
What it costsNothing. There is no paid tier to upgrade to. The open question we are paying to answer is whether the report is worth having.
What we storeSee below, and data handling → for the full version.

Where your code goes

"Read-only" answers what DiffCI can change. It does not answer what DiffCI can see, so here is that answer, permission by permission.

ContentsFor each commit it analyzes, DiffCI makes a shallow clone of your repository at that commit, inside an ephemeral container on infrastructure operated solely for DiffCI. This is what the contents permission is for: there is no way to build a real dependency graph without reading the source.
What happens to the cloneThe graph is computed and the container is destroyed. The clone does not outlive the analysis. No copy of your source tree is kept anywhere, and no file contents are written to any store.
What is keptCommit SHAs, the paths of changed files, the tests DiffCI would have selected, the reason it fell back to a full run when it did, and the reconciled outcome and timings of the CI run that actually happened. Plus evidence archives for those analyses, so a result can be audited rather than taken on trust. Never file contents. Never credentials: DiffCI is not given any.
For how long90 days at most, counted from the analysis that produced each record. Uninstalling the App deletes that repository's records sooner, automatically.
Actions, checksRead the results and timings of workflow runs that have already finished. This is the same information the Actions tab shows you, used to check DiffCI's prediction against what your CI actually did.
Pull requestsNot requested. DiffCI does not read pull request titles, descriptions, diffs, or review comments. The App asked for this scope in an earlier version without using it; it was removed. If pull-request-level analysis is built, the App will ask for the permission then, and GitHub will prompt you to approve it.
Who else sees itNobody. Nothing is sold, shared with a third party, or used to train anything, and nothing naming your repository is published without your written agreement to that specific text.

If your repository is private and the clone is the part you are not comfortable with, that is a reasonable position, and the right answer is not to install. If you do install and change your mind, uninstall, and the records go with it.

Install the read-only App

DiffCI can only analyze a repository with a tsconfig.json at its root, active GitHub Actions, and recent commits. If yours doesn't qualify, its report stays empty rather than being built on nothing — but it does not yet tell you why, so an unexplained empty report after a few pushes most likely means one of those three conditions is missing.