Telemetry without a kernel driver: recording what a PC does over time, and what it cannot report
Phase 5 added a recorder that samples CPU, GPU, memory and storage readings over time from sources that need no kernel driver and wake no device, writes them to a crash-safe evidence log, records every metric it cannot read as unavailable with its reason, and was held to a hard limit on its own cost, measured at 0.536 % of one logical processor in CPU cycles over ten minutes.
What was done
Built the platform's telemetry layer: a recorder in the Rust core that samples CPU, GPU, memory and storage readings at fixed rates for as long as a recording runs, and writes them to an evidence log that survives a crash or a pulled USB stick. Every source it reads works without a kernel driver and, on Windows, without elevation, under the rule that no read may wake a device; a metric the machine cannot report is recorded as unavailable, with the reason, in the same log. The sources were measured on the development PC before any reader was written, against rules fixed in advance, and the recorder's own cost is a hard limit of the gate, measured in CPU cycles. The local gate on the development PC passed on 11 October 2026 and CI then passed on the Phase 5 commit; the elevated run on that PC, a real sleep during a recording and every Linux sensor read on physical hardware did not take place and are stated limitations.
asked of the development PC what the log holds
----------------------------- --------------------------
CPU load, clock, power readings
GPU engine activity, memory readings
disk activity, NVMe health readings
the recorder's own cost readings
CPU temperature (Windows) unavailable: needs a
kernel driver; none used
GPU temperature and fan unavailable: the driver
reports none
board sensors, fan headers unavailable: no source
without a driver
81 channels of readings, 24 recorded as unavailableA metric the machine cannot report is recorded as unavailable with its reason, never left out and never filled in.
Status
The design was committed on 9 October 2026 (2ff6920). The implementation was committed for CI on 10 October (d4ba30b, on a work-in-progress branch), before the adversarial review, which is recorded separately. On 10 October the owner committed the tree part-way through the fix round (b6b4df0) and moved the repository to another drive, and the work continued in the moved copy. A gate agent ran the full Phase 5 local gate on the development PC on 11 October, on the uncommitted working tree, and every line passed. The work was then committed as the Phase 5 commit (53010e8, 11 October, 06:09 BST); no code file changed between the gate run and that commit.
CI run 38113880489 on the Phase 5 commit then passed all four jobs the same morning, including the Linux steps as a standard user and as root, which run on a GitHub-hosted virtual machine. That completed the phase.
This is the fifth of the specification's 23 phases.
The problem
Until this phase the platform recorded what a PC is: its parts, their identity and, for storage, one health snapshot. A diagnosis also needs to know how the machine behaves over time: how loaded it is, how fast it runs, how much power it draws, how hot it gets and whether it is being held back. One reading cannot show that, so the specification asks for measurements recorded over time, each tied to its source.
Most sensor programs for Windows get those readings by loading a kernel driver that reads processor registers and motherboard chips directly. The platform never loads a kernel driver, and without one some readings do not exist on Windows at all, the CPU package temperature among them. A second difficulty is that a read which is safe once may not be safe several times a second for hours. Several sources that look like plain reads do more: a driver-served GPU value can resume a suspended GPU, a disk temperature read can spin a disk up, and a Windows performance-counter query can load third-party code into the process. A third is that a recorder changes what it measures, and a recording that is lost when the machine crashes is no evidence.
How it was built
| Step | What happened |
|---|---|
| Design | Three proposals, each led by a different priority: device safety, honest readings and the recorder's own effect; low-level correctness across operating systems; the runtime, the evidence log and scope discipline. Three judges each ranked the first highest; the synthesis started from it and took the evidence-log writer from the third |
| Critique | Two review rounds before any code: 34 gaps, then 24, none a blocker; every major gap was fixed in the design |
| Lead review | Cut the readers that could not run on this PC or in CI (Linux GPU sensors, and firmware thermal zones on both operating systems) and two database projections that nothing reads yet: all stay designed, and their metrics are recorded as unavailable, not in this build. No gate step was made to depend on the owner |
| Measurement | Five sets of measurements on the development PC before any reader was written, each with its deciding rule written down first |
| Implementation | Eleven agents in dependency order: the measurements, the core's evidence log, the TypeScript side and the fixture generator, the server, the documents, the sampling engine, the Windows and Linux readers, the command line and agent, verification and integration |
| Review and fixes | An adversarial review, a fix round, a verification of every fix and a regression review, recorded separately |
| Verification | A gate agent ran the full local gate, including a ten-minute cost measurement and end-to-end runs of the release binary, on the development PC |
The owner, Stephen Ukaegbu, wrote the specification and made the decisions this phase answers to. The engineering was AI-assisted: Anthropic's Claude Code worked as lead engineer, orchestrating the designer, judge, critic, implementer, reviewer, verifier and gate agents. The lead and the agents share a model family, so their results are not independent checks of each other. Real hardware, CI and the owner's decisions are the checks that sit outside the model.
The rules
- No kernel driver, and no data that depends on one. The core loads no driver, reads no processor register or motherboard chip directly, loads no GPU vendor library, and does not read values published by third-party monitoring programs that load a driver.
- Nothing is woken or commanded. Sources come from a closed list compiled into the core, and no function accepts a path, counter or query from its caller. Each source has a compiled maximum rate. On Windows, GPU queries that could wake a powered-down GPU are built but disabled, because that could not be checked on this PC. The NVMe health log, the one command sent to a storage device, is read at a low rate and only after a check that the controller is awake. A machine-wide lock per disk keeps every running copy of the tool to one command at a time.
- Raw readings only. A recorded value is what the source states, an exact unit conversion of it, or a rate over two consecutive reads of one counter. Nothing is smoothed, clamped or combined across sources.
- A metric means one thing. A table compiled into the core decides which source may carry which metric, and a conformance test refuses anything else. A firmware thermal zone is never recorded as a CPU temperature, a share of a power limit is never converted to watts, and a processor power domain is never relabelled as GPU power.
- Absence is evidence. For each kind of reading the specification asks for, every component present has either a channel of readings or a channel whose every frame says unavailable, with the reason.
- The recorder states its own effect. Its lateness and read time are recorded with every frame, and its processor time, memory and writes once a second.
What the platform now records
On Windows, as a standard user:
- CPU: utilisation in total and per logical processor, the clock frequency the operating system measures, its performance-limit indicators, and package power from the processor's own energy counters as the operating system exposes them.
- GPU: how busy each engine of the GPU is and how much memory each of its memory segments holds, from the graphics kernel's statistics.
- Memory: available physical memory, and the processor-reported power of the memory domain.
- Storage: read and write rates, operation rates and queue depth for each disk; for NVMe disks the temperature, spare capacity, wear and error counters of the health log, at a low rate.
- The recorder itself: lateness and read duration with every frame; processor time, memory, write rate and the state of its writer once a second.
On Linux the core reads processor load, disk activity and available memory from the kernel's statistics files and, where the kernel exposes them, CPU temperatures, clock, thermal-throttle counts and, as root, package power. These readers are tested against synthetic file trees and on CI's virtual machines; none has run on physical hardware.
On the development PC at the gate, a recording at the standard profile holds 81 channels of readings in six streams, at rates from four a second for memory and disk activity down to a much lower rate for the NVMe health log, and a seventh stream for the 24 channels recorded as unavailable. Those 24 are the CPU temperatures, the throttle count and the requested clock, which Windows does not state without a driver; the GPU temperature, fan and engine clock, which this display driver reports as unsupported; GPU hotspot and memory temperature and GPU power; motherboard temperature, voltage and fans; firmware thermal zones, whose reader is not in this build; for the PC's ATA disk, temperature and health, which this build does not read; and the idle-time metric of both disks, which is not offered.
Five new commands list the channels a machine offers, record, take a short sample, check a recording without writing anything, and recover an interrupted one. The session contract did not change: it stays at version 0.3.0, the telemetry formats defined in Phase 2 were implemented as they stood, and no byte of a stable schema or of an existing fixture changed.
CPU temperature on Windows: unavailable by decision
The CPU package temperature is a processor register that Windows exposes only to kernel-mode code. What Windows offers a standard user is firmware thermal zones, and on the development PC one such zone reads about 72 °C and changes. Nothing states what a zone measures: firmware may track a board sensor, a copy of a CPU reading or a skin temperature. The decision, recorded before any code, has three parts. On Windows the CPU package temperature is unavailable at every privilege level, and the log says why. A thermal zone is a different metric and is never relabelled. The design decision of 9 October assigns CPU thermal tests to a pre-boot Linux environment, where the kernel states the package temperature; none exists.
The decision has a cost that was accepted with it. The platform refuses to start a load test without its safety reading, so its planning command reports the safety requirement of a CPU stress test as unmet on Windows; the rule was not weakened to avoid that. In this phase no zone reader was built at all, because reading a zone may make firmware query the embedded controller, and checking that needs an elevated trace that was not taken.
Measuring first, with the rules fixed in advance
Before any reader was written, throw-away programs measured each candidate source on the development PC, as a standard user, with read-only queries. The rule each measurement would decide was written into the design first, so that a result could not be argued into the answer wanted.
- Processor energy counters. Rule: the fastest allowed rate is the highest at which, under load, no read shows time advancing while energy stands still. At 10 and at 4 reads a second such reads occurred; at 2 and at 1 a second none did in 60 seconds each. The cap is 2 a second, and the standard profile's rate for that group fell from 4 to 2. The counter's unit was confirmed against the operating system's own power figure to within 0.05 %.
- Disk activity. The query is documented as switching the disk's activity counting on. Rule: it ships only if counting was already on before the first call. It was: the first call returned totals accumulated since boot.
- Disk idle time. Rule: offered only if the counter is seen to advance over a 60-second window with no disk activity. The only disk was the busy system disk and no such window occurred, so the metric is not offered, although the counter did advance in all 27 shorter idle intervals observed.
- NVMe health log, one read a second for 60 seconds. 60 reads, no errors, the controller awake before each, and no two commands closer than 945.8 ms.
- Flushing to disk. The slowest 1 % of flushes took about 213 ms, well inside the ten seconds of readings the writer's queue holds.
The measurements were taken while parallel build and test jobs held the processor at 93 to 100 %, so their costs and latencies are upper bounds, and the design says so. Some were not made: the two that need an elevated trace, the behaviour of the clocks across a real sleep, and flushing to a USB stick, because none was attached.
Recordings that survive a crash
A recording is an append-only log written in chunks, one set per stream. A chunk is listed in the session only after its first readings and a checkpoint are on disk, and a further checkpoint is written about every five seconds, or with every frame on a slower stream, so a crash or a pulled stick loses only the end of a recording, and the loss is recorded. The next command that saves the session recovers interrupted streams first: it keeps a copy of the torn file, replaces it with the part that verifies and marks the stream as recovered after an interruption. A lock stops two writers working on one session. If the PC sleeps during a recording the stream continues with a marked gap, and the first counter-based value after waking is marked invalid.
The log has three readers that must agree: the Rust core, the fixture generator's reference verifier and the TypeScript contract layer. The writer's output is compared byte for byte with the generator's, and shared test vectors for recovery give the same result in Rust and TypeScript.
It was tested by killing it. A real-hardware test kills the recorder five times, between 1.3 and about 17 seconds into a recording; each time the check command writes nothing, recovery completes in one save and both other readers verify the result. The end-to-end run with the release binary killed two more recordings, and both were recovered and closed.
The recorder's own cost
The cost is a gate limit with no allowance. At the standard profile on the development PC, over ten minutes with the release build:
| Measure | Result | Limit |
|---|---|---|
| Processor time, by CPU cycles | 0.536 % of one logical processor | 1 % |
| Peak memory | 26.1 MiB | 48 MiB |
| Written, including session saves | 2.51 KiB a second | 8 KiB a second |
| Flushes to disk | 1.08 a second | 2 a second |
| Lateness of a read, 99th percentile | 15.9 ms | 20 ms |
The processor figure is 6,791,459,161 cycles of the recorder's process at a counter rate of 2.1120 GHz, which the test calibrates itself: 3,215.6 ms of processor time in 600.0 seconds. A run of the same tree earlier that day gave 0.507 %. The first version of this gate used the operating system's tick-based accounting and reported 0.06 %; the review showed that such accounting misses most of the work of a program that wakes on timers, and that figure was withdrawn. The machine was not quiet during the gate run: another program used 23 to 34 % of one processor throughout, so the lateness figures are indicative.
Exceeding a limit in normal use produces a warning and never an automatic change of rate, because silently changing rates would distort the data.
How it was proved
The local gate on the development PC, run as a standard user:
| Check | Result |
|---|---|
| Rust workspace | 1,178 passed, 0 failed, 5 ignored: three helper processes of other tests, the ten-minute cost measurement, which was run separately, and a 30-minute power comparison, which was not run |
| Rust, release mode | Isolation 16, command line 10, build hygiene 20, vector files 17 |
| Real hardware | Discovery 22, identity and leak 14, scan 10, providers 8, telemetry 22 |
| TypeScript and database | 1,879 passed in 23 files, under two versions of Node |
| C# Windows agent | 103 passed, 0 skipped; build with 0 warnings |
| Fixtures | 733 generated files identical on regeneration; 748 files, 386 negative cases and 53,918 checks verified |
| Formatting and lints | Clean on the Windows and Linux targets |
End to end with the release binary. A 60-second recording at the standard profile closed all seven streams as complete, with no warning, and verified. A 30-second recording was added to a session that a scan had created, then verified and closed, and the C# agent continued and closed a session that a recording had left open. The TypeScript layer read all five outputs with no integrity violation, and the in-process PostgreSQL ingest stored 21 revisions of 5 sessions with none quarantined.
Unchanged contract and identity. Against the commit that closed Phase 4, the schema registry is byte-identical and the only schema file changed is the draft describing command-line output. A scan by the Phase 5 release build gave the same machine identity and the same component keys as the Phase 4 release build.
Leak check. 111 search strings of 30 kinds, taken from the development PC's own identifiers (among them serial numbers, the system UUID, MAC addresses, interface GUIDs and the computer name), were searched for across 619 files: every file changed or added since Phase 4, every end-to-end output and the gate logs. None was found, and a positive control found all 111. The repository's own leak test also searches what the telemetry commands print and save.
Source scan. A build check fails if the source names a forbidden device access or telemetry source, or if a disabled reader's switch is anything but off; controls prove that each scan can fail.
CI
- Run 38026119970, before the review (commit
d4ba30b, 10 October, from 06:02 BST): the TypeScript and C# jobs passed. Both Rust jobs failed, on three tests whose assumptions GitHub's virtual machines do not meet. The leak test searched for the account name, which in the Linux root step is a generic name that occurs in harmless output. On the Windows runner's current image every disk is of NVMe type with no controller the scan can locate, which had already failed the same test on the main branch (run 37986651545, on the commit that closed Phase 4), so it was not a Phase 5 regression. And the graphics statistics offered a channel for a virtual adapter that the scan could not match to a GPU. All three became review findings. - Run 38071634315, on the owner's mid-fix commit (
b6b4df0, 10 October, from 18:24 BST): the TypeScript and C# jobs passed. Both Rust jobs failed the formatting check and two tests that expected a text an earlier fix had changed; the steps that run the hardware tests on the runners' virtual machines, including the Linux root step, passed. Each of those failures was fixed in the working tree by the end of the fix round. - Run 38113880489, on the Phase 5 commit (
53010e8, 11 October, 06:09 to 06:50 BST): all four jobs passed. Rust ran 1,250 test executions on Linux and 1,285 on Windows, with none failing; the C# job passed 103 tests and the TypeScript job 1,879, with the same 53,918 fixture checks as the local gate. On the Linux virtual machine the recorder was killed five times as a standard user and five times as root, and every stream was recovered each time; the root step also traced the files the sampler threads open and ran a recording on a file system that fills up. The Windows runner is elevated, but in this run it presented no NVMe controller, so no NVMe test ran there, and neither check that needs an elevated and a standard-user run on one machine was covered. The ten-minute cost measurement does not run in CI.
Problems during the work
- An interruption. The application restarted during the fix round. Partial work was kept and finished, and the continuation checked every earlier result against the code.
- A move of the repository. On 10 October, at about 18:24 BST, the owner committed the tree as it stood and copied the repository to another drive. The work continued in the copy. The build directory there was built afresh, which matters for one test recorded in the review: it compares a result that an elevated run leaves in that directory, and no elevated run has been made.
- A second disk. The development PC gained a disk the same day. An ATA disk now enumerates first and the NVMe system disk second, which broke three real-hardware tests that had assumed the NVMe disk came first, and a fourth carried the same assumption; all four now find a disk by its transport or by the test's own write. The new disk also put on real hardware, for the first time, paths that had run only on synthetic devices or in CI: a disk whose identity is recorded as none, because the Windows identity query for ATA disks stays disabled, and whose temperature and health are recorded as unavailable.
- A busy machine. The measurements before the build ran at 93 to 100 % processor load, and the gate's cost measurement ran beside another program.
Limitations
- One physical machine. All physical verification is on the development PC (a Lenovo laptop) with an integrated GPU and 8 logical processors. Linux evidence comes from synthetic file trees and GitHub's virtual machines; no Linux sensor reader has run on physical hardware.
- Readings this PC cannot exercise. Its display driver reports no GPU temperature, fan, clock or power, so the code that records those values has run only against synthetic tables. The disabled GPU queries have never met a discrete or hybrid GPU.
- No CPU temperature on Windows, by decision, and no motherboard sensors or fan speeds on either operating system in this build.
- No elevated run on real hardware, and the test in which an elevated and a standard-user process share the per-disk lock was not run.
- No real sleep. Suspend handling is proved with simulated clocks only; no sleep or Modern Standby period was observed during a recording.
- Cost at scale is extrapolated. The cost of per-processor readings on larger machines rests on one measurement on 8 logical processors, taken under full load. On Linux machines with about 200 or more logical processors the default plan at the standard profile is still refused, with an error that names the setting to change.
- Designed but not built, or not offered. Linux GPU sensors, firmware thermal zones, Linux NVMe and ATA temperature and Windows ATA health are designed but not built; the idle-time metric for disks is not offered.
- Not measured. Whether the disk-activity query reaches the device was not traced; it is taken from the documented behaviour of the operating system component that answers it. The comparison of package power under the two sampling profiles was run only on the build from before the fixes (7.954 W against 8.223 W, means of three five-minute windows each, with the recorder measuring its own effect) and was not repeated on the fixed build.
- A gate on a machine that was not quiet. The cost figure is per process and little affected; the lateness figures may be worse or better on a quiet machine.
Open questions
Whether the disabled and unverified readers behave as designed on discrete GPUs, on SATA disks and on physical Linux machines was not shown at the time of this record. Neither was the recorder's cost on a quiet machine or on one with many more processors.
The code
How an absence is compiled into the core: the decision that no CPU temperature is recorded on Windows, with the metrics it covers, its cause and the fixed reason that every frame of those channels carries.
PolicyDecision {
name: "windows-no-driver-free-cpu-temperature",
metrics: &[
"cpu.package.temperature",
"cpu.core.temperature",
"cpu.control.temperature",
"cpu.die.temperature",
"cpu.ccd.temperature",
],
cause: PolicyCause::SourceAbsent,
// ...
reason: "no driver-free source of the CPU package temperature on Windows: the package digital thermal sensor is an MSR, which needs a kernel driver, and kernel drivers are never used; ACPI thermal zones are another metric, acpi.thermal_zone.temperature (ADR-0026)",
},core/diagnostic-core/src/telemetry/bindings.rs, as it stands at the Phase 5 commit 53010e8 — one entry of the table of policy decisions; a two-line comment is elided.
Not shown. The list of allowed sources, the table that binds sources to metrics, the evidence-log line format, the recovery procedure, the session and per-disk locks, the conditions under which a GPU is queried, the rates of the NVMe health-log read and the budget and cost values the planner uses are withheld, because they are internal design or safeguards still in use.
A published copy. Commit references and internal identifiers have been removed and the operator is not named; the engineering, the counts and the stated limits are unchanged.