Skip to content
All work
Software DevelopmentDiagnosticsResearch

PC Diagnostic Platform

An engineering and research project to build a single PC diagnostic system in which every hardware value records whether it was measured, where it came from, and why it is missing if it is, so that any conclusion can be traced back to recorded evidence. It is built in gated phases from an owner-written specification, with a Rust core, a C# Windows agent and a TypeScript and PostgreSQL contract layer. Three of the specification's 23 phases have passed their gates; the work continues.

Role
Project owner — specification, direction and decisions; AI-assisted engineering (Claude Code as lead engineer)
Timeline
Oct 2026 — Present
Status
in progress
RustC# / .NET 8TypeScriptJSON Schema (2020-12)PostgreSQL (PGlite)ajvVitestxUnitGitHub ActionsWMICPUIDClaude Code
Specification phases through their gate
3 of 23Phases 1 and 2, committed on 3 October 2026, and Phase 3, completed on 4 October 2026 once CI passed on Linux and Windows; each has a written gate report.
Tests passed at the Phase 2 gate
1,119Rust 79, C# 92, TypeScript and database 948; 0 failed, 1 skipped (a test that needs the commit to exist, which passed afterwards). Re-run by the lead, rather than taken from the agents' reports, before the gate was recorded, alongside 12,551 fixture checks including 326 negative cases.
Confirmed review findings fixed
53 of 55Phase 2 adversarial review: seven reviewers, with a separate verifier trying to refute each finding. 57 findings after de-duplication, 2 refuted, 55 confirmed (none blocking, 5 major); the other 2 were kept as documented design decisions.
Raw serial numbers in test output
06 raw serials from the test laptop checked across 9 real-hardware test runs at the Phase 2 gate; none appeared in any output.

The problem

Whether a PC is worth upgrading, keeping or replacing depends on facts that are hard to see: what is actually installed, what condition it is in, and what it can run. Those facts are easy to state and hard to check: an advertised component or figure is a claim until it is detected and tested. The specification does not accept a conclusion such as 'CPU: good' without the measurements behind it, and a value that could not be read should say so rather than be quietly filled in. The project sets out to build one diagnostic system that records what a PC contains, what condition it is in and what it can do, with evidence attached to every value, so that anything concluded later can be traced back to what was actually measured.

Research

The specification states requirements, not answers, and several of them are open technical questions, so the project tracks them as research questions and records evidence for and against each approach as phases complete. None is answered yet. How can evidence be preserved so that every conclusion traces to recorded, reproducible observations and missing data is never silently filled in? How can stress testing run within hardware protections, and refuse to run when safety telemetry is missing? Can a pre-boot environment work with Secure Boot left on, through a signed boot chain, and share one diagnostic system with Legacy boot? Can a component be recognised across boot environments and visits without storing its raw serial numbers? How should measured performance be kept apart from estimates, so that an estimate can never pass as verified? And how can AI-assisted, multi-agent engineering be verified to the standard a diagnostic tool needs? Technology choices were investigated the same way, with the rejected options written down: WinPE was ruled out for the pre-boot environment because redistribution needs separate Microsoft licensing and modern kits have no Legacy BIOS path; a bare UEFI application has no drivers for USB networking, NVMe health or GPUs; and an unsigned image that asks users to disable Secure Boot was rejected outright.

Planning

I wrote a master engineering specification before any code existed, on 3 October 2026. It fixes the stack, sets a 23-phase build order and puts a gate at the end of every phase: run the tests, build what changed, check for regressions, document what works and what fails, commit, update the architecture documents, and do not proceed on an unstable foundation. Each gate produces a written report of files changed, tests run, passed and failed, limitations and the next step. In Phase 2, after the design was settled, a lead review drew a cut line: machinery that no real caller could exercise yet was specified but deferred to the phase that will first use it, so that raw disk identity collection, for example, waits until it can be tested against real hardware. The aim is that nothing is built that the gate cannot test. Alongside the code, the project keeps a dated, append-only record of every stage, decision, problem and fix.

Design

The design takes the specification's evidential rules and turns each one into something the data model enforces rather than a convention. Every hardware value records whether it was measured, unavailable, not applicable or mock, together with its source and unit; an unavailable value must carry the error behind it, and mock data is labelled at every level, so a guess cannot pass as a reading. Raw measurement, derived metric, finding, interpretation and recommendation are separate kinds of record: a finding must cite evidence and the rule and rule version that produced it, an interpretation cannot cite a seller's claim as its evidence, and an estimate carries a range and can never be marked verified. History is append-only. Each save records its sequence and a hash of the save before it, and saves are ordered by that lineage rather than by timestamp, because a pre-boot environment and Windows keep different clocks and a correct save can appear to go backwards in time. Stable schema files are hash-locked, so changing one means publishing a new version. The Phase 2 design came from three independent proposals written to different priorities, scored by three judges, combined in one synthesis and put through two adversarial critique rounds that closed 46 gaps before any of it was implemented.

Architecture

A polyglot monorepo, with the language for each part fixed by the specification. A Rust core reads hardware directly — CPUID on any platform; WMI, the registry and the Win32 firmware interface on Windows; sysfs and efivarfs on Linux — and writes a diagnostic session through a command-line tool. A C# (.NET 8) Windows agent finds open sessions, on the USB data volume first and then in the local store, and asks the core to continue them; it never modifies another machine's session or a mock one. A TypeScript layer holds the contract types, strict schema validation, the schema registry and the integrity and lifecycle checks, and a PostgreSQL ingest library stores every revision append-only, tested against a real PostgreSQL engine running in-process. One JSON Schema (draft 2020-12) is the contract between all of them: the Rust, C# and TypeScript types are written by hand, and each language's tests validate fixtures and real output against the same schema, rather than generating code from it. One diagnostic visit is one session; it can be continued in another environment only on the same machine, and a weakly identified machine is never merged with another. The pre-boot environment is designed as a minimal Linux image behind a signed shim, GRUB and kernel chain, so that one USB can serve UEFI with Secure Boot enabled, Legacy BIOS and network boot; no bootable image has been built yet. GitHub Actions runs four jobs: Rust on Windows, Rust on Linux (including a step that runs the real-hardware tests as root), the C# agent with integration tests that drive the real Rust core, and the TypeScript, schema, fixture and database suite.

Challenges

  1. 01

    A CPU clock the operating system got wrong

    On the test laptop, WMI reported a base clock of 2101 MHz for a processor rated at 1.6 GHz. The CPUID leaf that gives the base frequency directly was masked by the Windows hypervisor, because the machine runs virtualisation-based security. No automated test caught it; checking real output against known facts about the machine did. The fix is a documented order of sources — the CPUID leaf first, then the rated frequency in the processor's brand string, then WMI, labelled as unverified — with a unit test. The same check found a second defect before the Phase 1 commit: the Windows agent crashed on a relative path to the core, now resolved to an absolute path, with launch failures reported instead of thrown.

  2. 02

    A locked schema that could be edited unnoticed

    Stable schema files are meant to be immutable, but the adversarial review showed that one could be edited and its hash lock updated to match without anything noticing. The registry test now compares stable entries against the previous commit's registry, and after the Phase 2 commit an in-place edit with an updated lock was shown to be rejected. The same check showed that the schema-check script on its own would still accept such an edit, so this protection depends on the TypeScript registry test running in CI — which, at that point, had never run.

  3. 03

    A stale USB copy that could reopen a closed visit

    A session can live on the USB stick and in the local store at the same time. The review found that an out-of-date copy on the stick could reopen a visit that had already been closed. The Windows agent now withholds every copy of a session that has been closed. Across the review, fixed findings were given regression tests where practical, so that a fixed defect cannot quietly return.

  4. 04

    Linux code that had only ever been type-checked

    Development happens on Windows, so the Linux hardware providers were compiled and linted locally but never run, and the review found that CI as written would only have exercised the weak-identity path. A CI step was added that runs the real-hardware identity and scan tests as root on Linux, with an unidentified machine counted as a failure. In the first CI run those 13 tests ran and passed — the first execution of the Linux code — though on a GitHub-hosted virtual machine rather than physical hardware.

  5. 05

    A first CI run that failed on Windows

    The first CI run, on 4 October 2026 (UK time), failed one Windows-only test. The test built a relative path from the checkout to the temporary directory, and on GitHub's Windows runners those sit on different drives, where no relative path can exist. Because the premise of the test was wrong rather than the code it tested, it was classified as a test defect, and the test now creates its temporary directory under the current directory. Cargo stops at the first failing test binary, so several Windows test binaries did not run in that job. The corrected test was committed with Phase 3 and passed on GitHub's Windows runner on 4 October, and CI now runs the Rust tests without fail-fast, so one failing test binary can no longer hide the others.

The solution

After three phases the platform is a thin, verified path through every layer rather than a feature set. On the test laptop the core created a session from a real hardware scan, continued it and closed it across three saves, each with a verified parent hash; the Windows agent continued, closed and mirrored the same session; the TypeScript layer read every revision with no integrity problems; and the database ingested them as linked revisions of one strongly identified PC. The method around the code is as much a part of the solution as the code. The engineering was AI-assisted: Anthropic's Claude Code (model Claude Opus 5.5) acted as lead engineer under my specification and decisions, orchestrating multi-agent workflows of independent designers, judges, critics, implementers working in dependency order, adversarial reviewers and refuting verifiers. The agents' own reports are not accepted as evidence. Every gate is re-run by the lead itself before it is recorded; real-hardware output is checked against known facts about the machine; and decisions with a material effect on architecture, security, data integrity, publication or disclosure come to me.

Result

Phase 1, committed on 3 October 2026, passed its gate with 45 of 45 tests (Rust 21, C# 8, TypeScript 16). Phase 2, committed the same day, defined 29 entities across 22 schema files, 15 of them stable and hash-locked, with 133 deterministic mock fixtures; its gate, re-run by the lead, passed 1,119 tests with none failing and one skipped, plus 12,551 fixture checks including 326 negative cases. Shared test vectors gave identical results across languages, including seven migration vectors that were deep-equal in Rust, in TypeScript and in the fixture generator's reference implementation. A scan of the working tree for the test laptop's real serial numbers and system UUID found none. The repository was then published privately, and the first CI run started at 00:15 BST on 4 October (23:15 UTC on 3 October): three of four jobs passed, including the TypeScript suite with 949 tests and the C# agent with 92. Phase 3, completed on 4 October 2026, built the architecture of the Rust diagnostic core: isolated provider calls with deadlines, a report of every kind of hardware source including those the build does not have, a registry and resolver for test definitions, and pure safety logic. Its gate, re-run by the lead, passed 358 Rust tests plus 30 in release mode, 95 C# tests and 979 TypeScript tests, and an adversarial review confirmed 40 of 49 de-duplicated findings, of which 39 were fixed with regression tests and one was partly fixed. The first CI run on the Phase 3 commit failed one Linux test that had hard-coded the result expected on Windows; once the test was corrected, all four CI jobs passed, with 406 Rust test executions on Linux and 388 on Windows (including release-mode and root re-runs) and none failing. The limits are stated as plainly as the results. All real-hardware verification so far is on one Lenovo laptop (Intel Core i5-10210U, 16 GB, Windows 11 Pro, UEFI with Secure Boot disabled); the Linux runs were on a virtual machine; no bootable image exists yet; no real test or benchmark runs yet, so the safety logic has been exercised only by tables of samples; and the contracts for tests, telemetry and findings have so far been exercised only by fixtures.

What I took from it

  • Real output catches what tests do not. Both Phase 1 defects were found by checking what the core reported against known facts about the machine, not by the test suite that existed at the time.

  • A protection that only runs in CI is only as good as CI running. The schema lock is enforced only by a test, and that test protects later changes only when it runs in CI, which needed a remote repository. Until the first CI run, the protection existed on paper.

  • Defer what you cannot exercise. Specifying machinery early is cheap; building it before anything can call it produces code the gate cannot test.

  • A finding deserves a challenge before a fix. In the Phase 2 review, a verifier whose job was to refute each finding rejected 2 of 57 before they prompted any change; whether both rejections were right was not independently checked.

  • Ordering by timestamp fails when two environments keep different clocks. Lineage — sequence and parent hash — orders saves correctly where wall-clock time can run backwards.

  • AI-assisted engineering needs evidence from outside the model. The reviewing and implementing agents share a model family, so their blind spots may be correlated. The checks that do not share them are real hardware, CI and the owner's decisions; the lead's re-runs of each gate replace the agents' own reports with fresh test results, but they are made by the same model family.