A diagnostic core that says what it cannot do, and a Linux test that assumed Windows
Phase 3 built the architecture of the Rust diagnostic core, from isolated provider calls and honest capability reporting to a test registry and pure safety logic, passed a lead-run gate of 358 Rust tests plus 30 in release mode, 95 C# and 979 TypeScript, and reached green CI on Linux and Windows after the first run exposed one test that had hard-coded the answer for a Windows host.
What was done
Built the architecture layer of the Rust diagnostic core: a provider model in which every kind of hardware source reports its capabilities and availability, isolation of every provider call behind a deadline, a registry and resolver for test definitions, a confirmation gate for tests that write, and pure safety logic that judges telemetry samples against stop conditions. The design came from three competing proposals, three judges, a synthesis and two critique rounds; eight implementation agents built it, and an adversarial review confirmed 40 findings, of which 39 were fixed with regression tests and 1 only partly. The first CI run on the commit failed one Linux test that had hard-coded the answer expected on a Windows host; once the test derived its expectation instead, all four CI jobs passed on Linux and Windows.
planning one mock test definition
it supports: pre-boot, network boot, Windows
on a Windows runner on a Linux runner
environment unsupported
mock dataset mock dataset
draft dataset draft dataset
no implementation no implementation
no safety telemetry no safety telemetry
the test expected the left column on every host
fix: derive the first refusal from the definition
and the host environment the plan reportsThe resolver was right on both hosts; the test had assumed which host it would run on.
Status
Completed 4 October 2026. The Phase 3 commit, a1c0dca, was made at 12:36 BST. CI run 37199342984 on it passed three of four jobs and failed on Linux. The correction, a9b0246, was committed at 16:12 BST, and CI run 37212131115 on it ran from 16:12 to 16:22 BST with all four jobs passing. Green CI is a condition of the gate, so the gate report records the phase as passed on a9b0246; the Phase 3 commit on its own does not pass.
This is the third of the specification's 23 phases.
The problem
At the end of Phase 2 the core could read hardware and manage sessions, but it called each provider directly, with no deadline and no protection against a crashing provider, so one hung query could stall a whole scan and one crash could end it. The core had no registry or resolver to load, validate and refuse test definitions before a run, and no logic for deciding when telemetry means a test must stop.
How it was built
| Step | What happened |
|---|---|
| Design | Three proposals, each led by a different priority: safety and honesty; runtime and concurrency; extensibility. Three judges all chose the safety proposal as the base |
| Synthesis | Interrupted by a session usage limit and resumed from cached results |
| Critique | Two adversarial rounds before any code was written |
| Lead review | Kept the threaded runtime that will execute tests as a specification only, because nothing in this build calls it; the stable session contract stays at 0.2.0 |
| Implementation | Eight agents, in dependency order |
| Review | Eight reviewers, one per dimension, then a separate verifier per finding whose task was to refute it |
| Fixes | Four parallel agents and a gate agent; the first attempt was interrupted by a usage limit and restarted from its partial work |
| Verification | The lead re-ran the full gate and the release-binary checks itself |
| Toolchain | A toolchain file now pins Rust 1.99.0, so a commit determines the compiler it was tested with |
The owner, Stephen Ukaegbu, set the specification and the decisions this phase answers to. The engineering was AI-assisted: Anthropic's Claude Code worked as lead engineer, orchestrating the designer, judge, critic, implementer, reviewer and verifier agents. The lead and the agents share a model family, so the lead's re-run replaces the agents' own reports with fresh results but is not an independent check. Real hardware, CI and the owner's decisions are the checks that sit outside the model.
What the core now does
- Every kind of source is accounted for. Eleven kinds are reported. On the development laptop six providers are real and probed, and the other five kinds (sensor, telemetry, error, test and benchmark) are reported as having no provider in this build, each with its reason, never silently absent.
- Unproven is not available. A capability whose usefulness a probe cannot prove is reported as unverified, never as available or partial.
- A failing source cannot take the scan with it. Each provider call is isolated with a deadline, so a source that hangs or crashes is reported as unavailable, with its cause, instead of stopping the scan. CI checks the release binary as well as the test builds.
- An identity that cannot be trusted stops the command. Nothing is written, rather than a weak fingerprint that would split one visit into unrelated sessions.
- Test definitions are pinned, not guessed. Reference datasets are hash-pinned and validated offline, there is never an implicit "latest" version, a bad safety value is refused rather than clamped, and every resolved parameter records where its value came from.
- Unsafe work is refused before it starts. A definition that needs safety telemetry the build cannot supply is refused up front. Tests that write to a disk stay impossible in this build, because the confirmation gate cannot yet tie them to an identified disk.
- Refusals are data. New read-only commands,
diag providersanddiag tests list,describeandplan, report every refusal with its reason and exit normally.
How it was proved
The lead's re-run on the development laptop:
| Suite | Passed | Notes |
|---|---|---|
| Rust | 358 | 0 failed, 0 ignored; includes real-hardware scan, identity and provider tests run as a standard user |
| Rust, release mode | 30 | The isolation, command-line and build-hygiene tests against an optimised build |
| C# | 95 | 0 skipped; integration tests drive the real Rust core |
| TypeScript and database | 979 | Run with one worker, for the reason given below |
The fixture verifier checked 143 fixture files, including 326 negative cases. The release binary passed its build checks and contained no test double.
Real output. On the laptop, diag providers listed all eleven kinds, and diag tests refused every test, each with an explicit reason: only mock and draft datasets exist, no implementation is built, and safety telemetry is unavailable. Session output matched Phase 2 field for field, apart from the version strings and one added dataset entry.
The review
| Measure | Count |
|---|---|
| Findings raised | 56 |
| After de-duplication | 49 |
| Refuted | 9 |
| Confirmed | 40 |
| Graded blocker / major / minor | 0 / 3 / 37 |
| Fixed, each with a regression test | 39 |
| Partly fixed | 1 |
The lead raised one minor finding to major because it concerned when a safety stop is triggered; it was fixed with a regression test. Every repaired test that previously could not fail was shown to fail under a deliberate mutation before it was trusted.
The partly fixed finding concerns one build check that does not yet cover every class of write. Extending it needs a design decision that had not been made at the time of this record.
The first CI run on the commit
Run 37199342984 started at 12:36 BST. The Windows Rust job passed, including the test that had failed in the project's first CI run because the runner keeps the checkout and the temporary directory on different drives. The C# and TypeScript jobs passed.
The Linux Rust job failed one test, which plans a mock stress-test definition on the measured host. The test had hard-coded the list of refusals produced on Windows. On a Linux runner the resolver correctly refuses the environment first, because the mock definition supports only pre-boot, network boot and Windows, so the list had one more entry. The product was right and the test's premise was wrong, so it was classified as a test defect.
The coverage this cost. Cargo stopped at the first failing test binary, so only three Linux test binaries ran in that job. The isolation, provider-model, real-hardware, registry and safety tests did not run, and the later steps, among them the release-mode tests and the real-hardware tests as root, were skipped.
The correction. The test now derives its expectation from the definition and from the host environment the plan reports, never from the operating system it was built on. CI now runs the Rust tests without fail-fast, so one failing binary cannot hide the others.
The green run
Run 37212131115 on a9b0246:
| Job | Duration | Detail |
|---|---|---|
| Rust, Ubuntu | 10 min 27 s | Formatting; Clippy for both targets; a check that the test-double crate never reaches the shipped binary; all tests without fail-fast; release-mode isolation, command-line and build-hygiene tests; checks on the release binary, including that it has no test double; real-hardware tests as root |
| Rust, Windows | 7 min 51 s | The same steps; the root step applies only to Linux and was skipped |
| C# Windows agent | 3 min 37 s | 95 passed, 0 failed, 0 skipped |
| TypeScript contracts and database | 2 min 34 s | 979 tests across 18 files |
Totals parsed from the Rust job logs: Linux 406 passed, 0 failed, 0 ignored, across 39 test-result lines; Windows 388 passed, 0 failed, 0 ignored, across 36. Windows ran the same 358 workspace tests as the lead's local gate plus the 30 release-mode tests. Linux ran 356 workspace tests, because some tests are compiled for one operating system only, then the 30 release-mode tests and 20 real-hardware tests as root. The totals count test executions, so a test that runs in more than one step is counted each time.
This was the first run in which the whole Phase 3 Rust suite executed on Linux, including the isolation, threading, probe and real-hardware code; the earlier run had reached only the three test binaries of the command-line crate, which is how the defect was found. GitHub-hosted runners are virtual machines, so the Linux real-hardware tests ran against virtualised hardware.
Problems on the development PC
- Memory. With default parallelism, the TypeScript suite's in-process PostgreSQL (PGlite) tests exhausted the PC's memory commit, so the local run used one worker. CI is unaffected.
- Disk space. The C: drive fell to about 1.3 GB free during the work (working-session log). 12 GB of scratch build copies were deleted during the gate; that they were the lead's own copies comes from the working-session log.
- Two usage-limit interruptions, in the design synthesis and in the fix workflow. Each resumed from saved or partial work.
Limitations
- No test runs yet. There is no real test or benchmark provider, and the threaded runtime that would execute a test is not built, so every test is refused. The safety logic is exercised by tables of samples, not by live telemetry.
- The safety values are provisional. The ceilings, the staleness limit and the hold before a breach can clear have not had a laboratory review.
- Memory stress is refused by design for now. No sensor channel covers a memory module, so every memory definition that is not read-only is refused.
- One partly fixed finding, as described above.
- Linux evidence comes from virtual machines. All physical real-hardware verification is still on one Lenovo laptop (Intel Core i5-10210U, 16 GB, Windows 11 Pro, UEFI with Secure Boot disabled).
- Shared model family. The reviewers, the refuting verifiers, the implementers and the lead are instances of the same underlying system, so their blind spots may be correlated.
Open questions
Whether the provisional safety values hold under real load has not been shown, because nothing in this build runs a load. The gap in that build check was still open at the time of this record.
The code
How the core accounts for every kind of source: a kind with no provider in the build is reported with a reason, never left out of the report.
pub fn kind_report(&self) -> Vec<KindReport> {
let descriptors: Vec<&'static ProviderDescriptor> = self.descriptors().collect();
ProviderKind::ALL
.into_iter()
.map(|kind| {
let providers: Vec<&'static str> =
descriptors.iter().filter(|d| d.kind == kind).map(|d| d.id).collect();
let status = if providers.is_empty() {
KindStatus::NoneInBuild { reason: none_in_build(kind) }
} else {
KindStatus::Provided { providers }
};
KindReport { kind, status }
})
.collect()
}core/diagnostic-core/src/provider/platform.rs, as it stands at a9b0246 (unchanged since a1c0dca) — the whole method, without its doc comment and re-indented.
Not shown. The isolation wrapper, the safety evaluator and the confirmation gate's checks are withheld, because they are safeguards still in use. The per-kind reason texts, the test-definition schema and the dataset format are internal design and are not published, and individual review findings are described only at summary level.
A published copy. Commit references and internal identifiers have been removed and the operator is not named; the engineering, the counts and the stated limits are unchanged.