Skip to content
The record
Report3 October 20265 min read

Seven reviewers and a refuting verifier for every finding: 55 of 57 confirmed, five of them major

An adversarial review of AI-implemented data contracts, with seven reviewers on separate dimensions and a separate verifier trying to refute each finding, confirmed 55 of 57 de-duplicated findings (0 blocker, 5 major, 50 minor) and led to 53 fixes, though as one review of one codebase by agents from the same model family it gives no measure of the defects it missed.

methodsystems

What was done

After the Phase 2 contracts were implemented, seven reviewer agents each examined the code along one dimension, and every finding was then handed to a separate verifier agent whose task was to refute it. Of 57 findings left after de-duplication, 2 were refuted and 55 confirmed: 0 blocker, 5 major and 50 minor. 53 were fixed, each with a regression test where practical, and 2 were kept as documented design decisions. The lead then re-ran the full gate itself rather than accepting the agents' own reports.

 7 reviewers, one dimension each
              |
 57 findings after de-duplication
              |
 1 refuting verifier per finding
       |                  |
   refuted: 2       confirmed: 55
                    0 blocker / 5 major / 50 minor
                          |
                  fixed: 53
                  kept as documented decisions: 2
              |
 lead re-runs the whole gate: 1,119 pass, 0 fail

Every finding had to survive someone trying to knock it down before it changed the code.


Status

Completed 3 October 2026, within the Phase 2 work committed that evening.

The outcome

On this phase's code, a review in which each finding had to survive a separate refutation attempt confirmed 55 of 57 findings, including five major findings, all of which received fixes before the phase was committed.

The method

StepWhat happened
ReviewSeven reviewers, one per review dimension, each working on the implemented code
De-duplicationOverlapping findings merged, leaving 57
RefutationOne separate verifier per finding, tasked with showing it was wrong
Fix53 confirmed findings fixed, with a regression test where practical
Decision2 confirmed findings kept, documented as deliberate design choices
Re-verificationThe lead re-ran the full gate itself

The refutation step exists to stop a plausible but wrong finding from causing an unnecessary change; the regression tests exist to stop a fixed defect from returning.

The owner, Stephen Ukaegbu, set the specification and the decisions this work answers to. The reviewers, verifiers and implementers were AI agents: Anthropic's Claude Code worked as lead engineer, orchestrated the review, and carried out the re-run of the gate itself, rather than relying on the agents' reports.

The count

MeasureCount
Findings after de-duplication57
Refuted2
Confirmed55
Blocker / major / minor0 / 5 / 50
Fixed53
Kept as documented design decisions2

The repository records the total of five majors and a fix for each. Which individual findings carried the "major" label is recorded only in the working-session log. According to the working-session log, the five majors were:

  • A stable schema could be edited and re-locked without detection. The registry test now compares each locked entry with the previous registry, not only with itself.
  • An interpretation could cite a seller's claim as its evidence, contrary to the layering rules. The evidence an interpretation may cite is now restricted so it cannot point at seller claims.
  • A stale copy on the USB stick could reopen a visit that had already been closed. The agent now withholds every copy of a session once any copy of it is closed.
  • An integrity check ignored telemetry streams that satisfy a requirement. It now covers them.
  • Linux CI exercised only the weak-identity path. A CI step now runs the real-hardware identity and scan tests as root on Linux, and treats an unidentified machine as a failure. This fix could not be demonstrated at the gate, because CI had never run.

After the fixes, the lead's own re-run passed 1,119 tests with none failing, plus 12,551 fixture checks.

Limitations

  • One review of one codebase. These counts describe a single phase's code under a single review. They are not a rate that can be carried to other code.
  • No measure of what was missed. The record contains no measure of the defects the reviewers did not find, so the number of real defects missed is unknown. Defects that neither the tests nor the reviewers anticipated can remain.
  • Every agent involved shares a model family. The reviewers, the refuting verifiers, the implementers and the lead that re-ran the gate are all instances of the same underlying system, so their blind spots may be correlated and the independence between them is partial. The lead's re-run replaces the agents' own reports with fresh test results; it does not remove that correlation. Real hardware, CI and the owner's decisions are the checks that sit outside the model.
  • Refutation was also done by an agent. Two refuted findings show the step can reject a finding; they do not show how often it rejects a correct one.
  • Severity labels are partly unverified. The count of five majors is recorded in the repository; the assignment of the label to each finding is not.

Open questions

How many real defects the review missed is not known, and how often the refutation step rejects a correct finding has not been measured.


The code

The fix for the stale-copy finding (a major, per the working-session log): when any copy of a session is no longer open, every copy of that session is withheld, so an open copy beside a closed one cannot be continued.

public static IReadOnlyList<SessionSummary> OfferedCopies(IReadOnlyList<SessionSummary> sessions, List<string> log)
{
    var withheld = new HashSet<SessionSummary>(ReferenceEqualityComparer.Instance);
    foreach (var copies in sessions.GroupBy(s => s.Id).Where(g => g.Count() > 1))
    {
        if (copies.FirstOrDefault(c => !c.IsOpen) is { } ended)
        {
            withheld.UnionWith(copies);
            // ... log the open copies as a stale branch of a finished visit ...
            continue;
        }
        // ... otherwise offer one copy, preferring a proven descendant ...
    }
    return sessions.Where(s => !withheld.Contains(s)).ToList();
}
From platforms/windows-agent/src/PcDiag.WindowsAgent.Core/ContinueWorkflow.cs — an excerpt, trimmed to the technique it illustrates.

Not shown. The fixes for the other majors touch the schema's evidence-reference rules, an integrity rule and the identity tests, which are internal design and are withheld. The review findings themselves, their numbering and the reviewers' dimensions are not published, and nothing is shown of how a session is matched to a machine.

A published copy. Commit references and internal identifiers have been removed and the operator is not named; the engineering, the counts and the stated limits are unchanged.