Which model should we standardise on — proven on our code?
The harness is MIT and yours to run. When you would rather have the answer than the instrument, this is the engagement that produces it.
Not a subscription
Priced against the decision, not the compute.
An annual AI tooling budget is a six-figure line item chosen, almost everywhere, on marketing plus a feeling. This engagement replaces the feeling with a measurement taken on the only corpus that predicts your outcomes: your own repository.
It is deliberately a consulting engagement rather than a seat licence. One engagement at a time, run end to end, with the harness you can keep running yourself afterwards — because the harness is MIT and the method is published.
Later, if the pull is there: scheduled re-runs as models ship. “Did the new release actually get better on our code?” is the same question, asked on a cadence.
How it runs
Four stages. Your code never leaves.
- 01
Scope
Which repositories, which candidate configs, and — the part that matters — which decision the result has to settle. A model standardisation, a spend tier, an agent harness, a reasoning-effort default.
A written scope and a candidate matrix.
- 02
Mine & gate
Guignet runs on your infrastructure. Your history becomes candidate tasks; the validity gate replays each one and discards everything that doesn’t reproduce. You see the soundness rate before a single agent runs.
An admitted suite, with per-candidate discard reasons.
- 03
Run & score
N attempts per task per config, in disposable worktrees, with cost parsed from each harness’s own transcript. Contamination controls applied and reported, never quietly.
Verdicts, costs, confidence intervals, flag rates.
- 04
Readout
The leaderboard, the recommendation, the tuned configs — and an explicit statement of what would change the answer. A number you can’t act on isn’t a deliverable.
The report, the configs, and a live walkthrough.
Specimen — what you receive
One self-contained file, and its JSON twin.
Renders offline. No framework, no CDN. Regenerates from stored runs without re-executing a single attempt.
- 01Leaderboard
- Every candidate config ranked on your admitted suite, with 95% Wilson confidence intervals rendered on every solve rate.
- 02$ per solved task
- The executive number. Cost parsed from harness transcripts, never self-reported, with partial coverage marked as a lower bound.
- 03Taxonomy breakdowns
- By type (bugfix / feature / refactor), by size bucket, and by path-based area tag — “wins on backend bugfixes, loses on UI features” is the actionable shape.
- 04Cutoff-split columns
- Scores split by each model’s training cutoff, framed by your repository’s visibility, with regurgitation flag rates per config.
- 05Tuned configs
- The run configs that produced the result, ready to keep — plus the prompt and effort settings that moved the numbers.
- 06The JSON twin
- Every number in the HTML, also as JSON. Auditable, and yours to re-run whenever you like.
Every figure in the report is explained in the published methodology, including the limitations. If a number isn’t explained there, treat it as unexplained.
Start
Tell me the decision you’re trying to settle.
A short note about the repositories, the configs you are weighing, and the deadline is enough to know whether this is worth either of our time. If your history won’t support a sound suite, I will say so.