DiffCI

Executed evidence

CI test-selection case studies with the failures left in.

These studies report full and selected execution, analysis overhead, selection-honoring checks, mutation recall, repository scope, and results that were withheld when the experiment did not support a claim.

Compare the measured cases

cal.com · six merges

Setup cost changes the answer

44.2%job-equivalent net reduction
8 of 8selections honored exactly

Large test-stage reductions became a smaller job-level result after install and generation were included.

Read the case study and source commits →

deepseek-harness · five merges

A dirty baseline changes recall

79.5–89.5%eligible job-level reductions
1 of 5results withheld

Sixteen to eighteen pre-existing failures forced baseline-relative recall, and one broadened execution was excluded.

Read the case study and source commits →

DiffCI's own repository

Green CI was not running every test

15test files hidden by a shell glob
0.696historical preflight recall

The self-study documents infrastructure defects, a misleading green suite, and an initially overstated replay score.

Read the self-study and repository evidence →

Use the evidence ledger, not only the summaries

The DiffCI Open Evidence Study consolidates the benchmark method, source reports, measured and inferred labels, limitations, downloadable CSV, and citable PDF. The case studies explain individual repositories; the ledger is the cross-repository record.