Skip to content
The record
Report11 October 202615 min read

Reviewing the telemetry recorder: misfiled readings, a trusted path and a cost gate that could not see the cost

An adversarial review of the Phase 5 telemetry recorder confirmed 103 of 109 de-duplicated findings, 2 of them blockers: readings written under another stream's header, and a recovery step that trusted a stored path; separate verifiers found 92 fixes complete, 10 partial and 1 worse, a remediation pass closed those 11, and of 10 regressions then found, 9 were fixed and 1 waits for an elevated run.

methodsystemssecurity

What was done

Ran an adversarial review of the Phase 5 telemetry implementation before its gate: thirteen reviewers raised 120 findings, and a separate verifier per finding tried to refute each one, leaving 103 confirmed, 2 of them blockers. The lead settled the findings whose fix had options in 13 written decisions, eight fix owners worked on their own parts of the code, and a second set of verifiers then checked every fix against the code: 92 were fixed, 10 partly fixed and 1 made worse, and a remediation pass then closed those 11. A review of the fix diff raised 10 regression claims; all 10 held, 9 were fixed in the same pass, whose own fixes were not verified a second time, and the tenth needs an elevated run that only the owner can make. The round was interrupted part-way by an application restart and by a move of the repository to another drive, and was finished by a continuation that checked every finding against the code again.

 10 reviewers, one dimension each    104 raised
  3 fresh-eyes reviewers             +16 raised
 de-duplicated                       109
 a verifier per finding tries to       6 refuted
 refute it (38 by reproduction too)  103 confirmed
   blocker / major / minor             2 / 34 / 67
 8 fix owners, one area each         the fix round
 53 verifiers check each first fix   92 fixed, 10 partly,
                                      1 made worse
 remediation                         those 11 closed
 review of the fix diff              10 regressions
 each claim checked, then fixed      10 held, 9 fixed

Every finding was challenged before it was fixed, and every first fix was checked by an agent that had not written it.


Status

The review ran on 10 October 2026 on the implementation as committed for CI in d4ba30b. The fix round started the same day and was interrupted; the owner committed the tree as it stood (b6b4df0, 18:24 BST) and moved the repository to another drive, and a continuation finished the fixes, their verification, the regression review and the remediation overnight. A gate agent ran the full local gate on the development PC on 11 October, and every line passed; the tree was then committed as the Phase 5 commit (53010e8, 06:09 BST).

CI run 38113880489 on the Phase 5 commit then passed all four jobs, including the Linux steps as a standard user and as root, which run on a GitHub-hosted virtual machine.

The problem

The implementation had passed its own local gate, including a ten-minute measurement of the recorder's cost and a test that kills the recorder and recovers the recording. A recorder that runs for hours beside real devices, and whose output is the evidence that conclusions rest on, needs more than its authors' tests. The review asked one question per dimension: device safety; whether readings say exactly what their source says; the Windows readers; the Linux readers; the sampling engine and its concurrency; the evidence-log writer, recovery and the contract; the command line and sessions; the safety feed; tests and CI; and privacy and whether the documents were true. Three fresh-eyes reviewers were given no dimension: one looked across the parts, one at failure paths, one at the specification and safety.

The two blockers:

  • Readings filed under another stream's header. Each sampler was told which stream to write to when the plan was built. If a source earlier in the plan timed out while opening, the streams that did start were renumbered but the samplers were not, so every later sampler wrote into its neighbour's log. In the reproduction a CPU utilisation value was sealed under the header of a disk write rate, every frame of another stream was rejected, and the summary still reported a stream as closed with 15 frames and no file.
  • Recovery trusted a path stored in the session file. Recovery of an interrupted recording opened each log file by the path the session file gave for it, without checking that path. A session file travels on removable media, so an edited one could name a file outside the session's own folder, and the next command that saved the session would copy that file onto the stick. The verifier reproduced it with a planted file.

Among the 34 majors:

  • A cost gate that could not see the cost. The recorder's processor time, the limit on it and the gate's independent check all read the operating system's process-time accounting, which charges whole clock ticks to whichever thread is running when a tick fires. The recorder's threads wake on timers and run for far less than a tick, so most of their work was never charged. The gate's first figure, 0.06 %, was that artefact; the design's own estimate, made beforehand from cycle counts, had been 0.4 to 0.5 %.
  • Default plans refused on many-core PCs. The planner refuses a plan whose estimated cost exceeds a fixed budget, and per-processor readings cost time for every logical processor. With the readers' real costs the default plan was refused from about 16 to 32 logical processors, depending on the profile, so a default recording on a common 32-thread desktop would have ended with an error and recorded nothing. The tests said to cover 128 and 256 processors had used a token cost per channel in place of the readers' real one, so they never reached the budget.
  • Readings that said more than their source. On Linux a throttle-time counter was read as a share of each interval, although the kernel adds a whole episode's duration in one step when the episode ends. On AMD processors a power domain that covers one core was recorded as the power of all cores. And a disk's readings followed the kernel's name for it, so a different disk that took the same name would have been recorded as the first.
  • Wrong after a pause. A late wake-up made every sampler read twice back to back; the recorder's own blocking work was declared a system-wide stall; and after a suspend on Linux a frame could be given a time early by the whole period asleep.
  • Failures reported as success. A recording that failed saved its surviving streams as complete, the check command reported no problem while an open stream's file was missing, and recovery recorded an unreadable last chunk as a stream in which nothing had been recorded.
  • Tests that could not fail, or never ran. The second verifier of real recordings, in TypeScript, had never run anywhere, because Node could not load the package, and CI did not require it. A test of the writer's write order could not fail. Three tests made assumptions that GitHub's virtual machines do not meet, which the first CI run had already shown.

The lead's decisions

The lead recorded 13 decisions where a fix had options or changed the design. For the other findings the verifier's proposed fix was applied, or the departure from it and the reason were written into the design.

How each class was fixed

ClassWhat changed
Misfiled readingsA sampler is given its stream only once the list of streams that actually started is final. A test with a hung source followed by working ones asserts that every frame lands in its own stream and that the hung source's metrics are recorded as unavailable
Trusted pathsNo path stored in a session file is trusted: it is validated before any file is touched, and a path that fails is refused. Tests cover crafted paths of several kinds. The validation rule is withheld
The cost gateThe hard limit is measured by the test in CPU cycles of the recorder's process, converted with a counter rate that the test calibrates itself before and after the run. The tick-based figure and the recorder's own report are printed beside it as information, and the recorded channel's definition states its granularity. The limit was not relaxed
Many-core plansWhen the operator gave no rate, the planner picks the highest of a fixed ladder of rates at which the per-processor readings fit the budget, records the chosen rate in the stream, reports it and warns. An explicit rate that does not fit is still refused, with an error that names the group and a rate that would fit. The plan tests use the readers' real costs
Readings that overstatedThe throttle-time counter is no longer read, and its two metrics were removed before any release. The per-core power domain is recorded as all-cores power on Intel processors only. On Linux a source's identity is established again at every read, so a reused name ends the channel instead of continuing it
Time and pausesA sampler chooses its slot before and after it waits, so a late wake-up skips the missed slot. Only lateness the recorder cannot account for is a stall, one interruption gives one marked gap, and the first frame after a suspend carries its own fresh time reference
Failure reportingA failed recording completes no stream: the surviving ones are left open for recovery. The check command reports a missing or unreadable file as a failure, and a lost last chunk is recorded as lost
Tests and CIThe contract package was changed so that Node can load it, and CI requires the TypeScript verification in both Linux steps that run the hardware tests. The leak test no longer searches for a generic account name as a bare word. On a CI runner a virtual NVMe disk with no controller is asserted as a well-formed unavailable observation and reported as not exercised

No stable schema or existing fixture byte changed in the round; new test vectors live in new directories.

How the fixes were verified

Several fix owners reported reverting their own fixes to show the new tests failing. Those reports were not taken as the evidence. In the continuation, 53 verifiers went through all 103 findings against the code as it then stood. They ran the tests again and, in most cases, undid the fix in a scratch copy to see its test fail. Their verdicts: 92 fixed, 10 partly fixed, 1 made worse.

  • The one made worse. The fix for the AMD power domain compared the processor's vendor with a spelling that the scan does not record, so the comparison was false on every processor and Intel machines would have lost the reading as well. One shared definition of the vendor names replaced it, and three tests fail against the old comparison.
  • Partly fixed, for want of proof. The cycle-based cost gate had been written, but its ten-minute run had not been made, so the figure the decision asked for did not exist. It was then run: 0.507 %, and 0.536 % in the gate run, against the limit of 1 %.
  • Partly fixed, for want of a decision. One fix made the first read of the NVMe health log keep its distance from the scan's last command to the disk. With that in place the first frames of a one-a-second recording ran late, beyond the bound the design sets, and a real-hardware test failed on almost every run (it passed once). A slot that would be read too late is now skipped and counted, which costs at most one frame at the start of such a recording.

A remediation pass closed all eleven. Nine have a test that fails without the fix; for two of the nine that follows from what the test asserts and was not run against the old code. The other two were a measurement that had to be made and a correction to the documents. No verifier ran after this pass: its fixes, and the nine regression fixes described below, were checked by its own tests and by the full local gate, not by a second agent.

The review of the fixes

Two reviewers read the fix diff for regressions, one for safety and honest readings and one for correctness and tests, and raised five claims each: 3 major and 7 minor. Each was checked against the code before anything changed, and all 10 held.

  • The many-core fix overreached (two claims, one defect). The new ladder also made room when the operator had explicitly raised another group's rate: instead of refusing, it lowered the per-processor readings, on this 8-processor PC too, and its warning blamed the processor count. The ladder's step is now chosen from the plan without the operator's increases, so an explicit rate that does not fit is refused.
  • A real-hardware test that failed on almost every run. This is the delayed first read described above; the regression review raised it independently of the verifiers.
  • A comparison that passed without comparing. A test compares component keys between a standard-user and an elevated scan, using a result each run leaves in the build directory. With no elevated result there, it passed without comparing anything, and a moved or rebuilt checkout has none. No elevated run has been made on this checkout, so the comparison is still not exercised. It can now be made mandatory, so that an unexercised comparison fails instead of passing. This is the one claim not closed.

The other claims were smaller: a cause recorded too broadly for disks whose health is not read, a per-disk lock that two code paths could have taken differently, a hardware adapter listed as a software one, a refusal message that claimed more than was true, a leftover duplicate of a mapping function, and a document that claimed more than the code does on very large Linux machines.

The interruption

The fix round did not run in one piece. On 10 October the application restarted during the stage for findings that span several owners, after the eight fix owners had reported. At about 18:24 BST the owner committed the tree as it stood (b6b4df0), pushed it and copied the repository to another drive. The continuation started in the copy. It took the first run's reports from the workflow journal and its digests, among them 82 requests that fix owners had addressed to one another, and checked them against the code instead of accepting them. At its start the tree did not pass the formatting check and five tests failed.

Three of those five failed because the development PC had gained a second disk that day. An ATA disk now enumerates first and the NVMe system disk second, and three real-hardware tests had assumed the NVMe disk came first; a fourth carried the same assumption. All four now find a disk by its transport or by the test's own write. CI on the mid-fix commit (run 38071634315) failed both Rust jobs on the formatting check and on the other two tests, which expected a text an earlier fix had changed; all of it was fixed in the tree by the end of the round.

Results

MeasureCount
Findings raised120
After de-duplication109
Refuted6
Confirmed103
Graded blocker / major / minor2 / 34 / 67
Lead decisions13
Fixed at the first verification92
Partly fixed, then completed10
Made worse, then fixed1
Regression claims on the fixes10, all confirmed; 9 fixed, 1 waiting for an elevated run

The local gate after the fixes passed with 1,178 Rust tests, 103 C# and 1,879 TypeScript, the release-mode suites and the real-hardware suites; at the first local gate, before the review, the same suites had 945, 102 and 1,806. The recorder's cost measured 0.536 % of one logical processor by CPU cycles over ten minutes.

Limitations

  • Shared model family. The reviewers, the verifiers, the fix owners, the gate agent and the lead are instances of the same underlying system, so their blind spots may be correlated. CI and real hardware are the checks outside the model: the three failing CI tests and the second disk each exposed assumptions that no agent had questioned.
  • Refutations were not re-checked. Whether the six refuted findings were rightly refuted was not independently checked.
  • The under-reporting was shown by reproduction, not by the gate. The reviewer and the verifier showed with small test programs that tick-based accounting misses short timer-driven bursts. In the gate run itself, on a machine that was not quiet, the tick-based figure (0.547 %) and the cycle figure (0.536 %) were close.
  • Many fixes are proved only on synthetic devices. The many-core ladder was tested with modelled machines of up to 256 processors, using a per-processor cost extrapolated from one measurement on 8; the largest real machine is this PC. Suspend and stall handling ran against simulated clocks. The Linux fixes ran on synthetic file trees and in CI's virtual machines.
  • Limits the round states and does not fix. Among them: on Linux machines with about 200 or more logical processors the default plan at the standard profile is still refused; the final save of a session on a medium that stops responding can still block; and a console window's close event is tested by calling the handler directly, not by closing a real window.
  • The remediation was not independently verified. The separate verifiers judged the first fixes. The completions of the eleven findings and the nine regression fixes were written in one pass and checked by its own tests and by the gate; for some, the new test was not run against the old code.
  • One regression claim is open, as described above, until an elevated run is made.

Open questions

Whether the recorder's cost stays within its limit on a quiet machine and on machines with many more processors, and whether component keys stay equal between a standard and an elevated scan of the same machine, were not shown at the time of this record.


The code

The instrument the cost gate now uses: the CPU cycles the operating system has charged to the recorder's process, and a counter rate that the test measures itself, so that cycles become time on a processor without any figure from the product.

    pub fn cycle_time(process: HANDLE) -> Option<u64> {
        use windows_sys::Win32::System::WindowsProgramming::QueryProcessCycleTime;
        let mut cycles = 0u64;
        // SAFETY: `process` is a process handle with query rights; `cycles` is a valid out pointer.
        (unsafe { QueryProcessCycleTime(process, &mut cycles) } != 0).then_some(cycles)
    }
 
    // ...
    #[cfg(target_arch = "x86_64")]
    pub fn tsc_hz(over: std::time::Duration) -> Option<f64> {
        let (t0, c0) = tsc_now();
        std::thread::sleep(over);
        let (t1, c1) = tsc_now();
        let seconds = t1.duration_since(t0).as_secs_f64();
        (c1 > c0 && seconds > 0.0).then(|| (c1 - c0) as f64 / seconds)
    }
From core/diag-cli/tests/support/independent.rs, as it stands at the Phase 5 commit 53010e8 — two whole functions of the test support code, without their doc comments; the lines between them are elided.

Not shown. The path validation, the stream-assignment code, the planner's budget and ladder values and the recovery procedure are withheld, because they are internal design or safeguards still in use. Apart from the two blockers, which are described as problems and outcomes, individual findings are described only at summary level, and the lead's decisions are not published.

A published copy. Commit references and internal identifiers have been removed and the operator is not named; the engineering, the counts and the stated limits are unchanged.