SECOND MEASUREMENTS · REVIEW, OCTOBER 2026

What six second measurements found, and what happened next

Markovian Protocol2 October 2026, second editionCovers SM-001 to SM-006 and seven null checks

Six claims about public logs and research datasets were rechecked in September, and four turned out wrong in a way that mattered. A month on, two of the four are fixed or confirmed and two are still wrong, and one of them sits under at least 39 published studies.

Disclosure: Markovian Protocol holds no financial position in any organisation named here, was paid by no one for this work, and showed it to no one before publication except the organisation measured. How we work.

Details

A second measurement takes a number someone has published about their own system or dataset and gets it again by a different route, with code anyone can rerun.1 Since 15 September we have published six papers and seven checks that found nothing. This review puts them side by side, reports what each looks like when rerun today, and follows one finding, about the AIDev dataset of AI coding-agent pull requests, into the studies built on it.

Since this review, 2 October

Three more papers and two more claims that held, all after this edition closed. SM-009: Firefox’s revocation list misses a third of certificates revoked on their first day. SM-010: Hugging Face’s malware scanner reads nothing over 2 GB, and the badge says safe. SM-011: that same Firefox list is months behind the largest certificate logs, and one in eleven top sites, mozilla.org included, gets no revocation check. Held: Debian’s reproducibility gate and every Certificate Transparency log’s merge-delay promise, on the index. The next edition will fold them in.

Every check so far

CheckThe claimHeld?Since then
SM-001Google's Pixel log covers every factory image from Pixel 6 on2No: the January 2026 release (29 images) was missingGoogle confirmed and added them on 22 Sep3
SM-002AIDev captures each agent's pull requests, Claude Code's included4No: its search finds about 1 in 11, and larger ones than usualIndependently confirmed by a 180-million-repository census5; no reply from the authors yet
SM-003Wild SBOMs are written by many practitioners6No: 47.5% come from one automated runOne author replied; no correction yet
SM-004whisper.online's ledger proves nothing was changed7Yes, every proof checkedOperator confirmed, fixed a duplicate-update bug we measured
SM-005Meta republishes its pulled-software list every 3 hours and ships server software weekly8No: every 6 hours since January; new software in 37 of 70 weeksStill 6 hours on 2 Oct; whitepaper unchanged
SM-006136.7 million x402 payments on Base in 280 days, most of them manufactured9Not tested. Our one-day count reproduced exactly; payments then fell by three-quarters, and two wallet groups still sent about 80%Rechecked monthly from now on
NullArmored Witness firmware log names its source1027 of 28—
NullCloudflare's Plexi audits every WhatsApp key-transparency epoch11150 of 150 sampled—
NullAn MSR 2026 dataset describes real-world MCP servers12Yes, though it keeps only repos with 50+ stars, unstatedAuthors told
Null17% of 2025 PyPI uploads carry an attestation1317.9% in our sample—
Null4.2% of MCP registry servers redirected their endpoint144.02%—
NullGitHub leaves reporter credit out of the CVEs it assigns150 of 150 carry it—
NullHomebrew attests every bottle its CI builds16599 of 600One bottle reported to Homebrew17

"Held?" refers to the claim as its owner stated it. Null checks are listed on the measurements index with their scripts.

Big organisations measuring their own systems mostly held up: PyPI, Homebrew, GitHub and Cloudflare all did. The two lasting misses came from somewhere else: a vendor whitepaper that stopped matching its own log (SM-005), and a research dataset built on a single search (SM-002).

Rechecked today

Three of the checks now rerun on a schedule, and every result is kept on a rechecks page.

When published2 October 2026
SM-001 Pixel images missing from the log33 (15 Sep)0
SM-001 Pixel entries a phone can't look up103128
SM-005 Meta pulled-software list, median gap6.0 h6.0 h (whitepaper still says 3)
SM-005 newest server-software entry29 Sep29 Sep
SM-006 x402 payments in one day234,490 (5 Aug)63,003 (1 Oct)

Pixel figures from Google's image page against the log's image_info.txt; Meta figures from Cloudflare Plexi, namespace prod.pc.revocation_list over the last 30 days.

The Pixel fix held, but the other Pixel problem grew. Entries filed under a label phones don't report, so a phone can't find its own, went from 103 to 128.2

One dataset, 39 studies, and the Claude Code it doesn’t see

AIDev finds Claude Code’s pull requests by searching for a line Claude Code writes in commits, not the line it writes in pull-request descriptions. Over the same dates that search finds 16,065; searching for both finds 173,066. The ones it finds are also bigger than typical: a median of 736 changed lines against 447.1 A census of 180 million repositories reached the same conclusion from the commit side: AIDev holds 5,137 Claude Code pull requests against 850,157 Claude Code commits that census found, and a pull-request census like AIDev misses 79% of the projects where they found Claude Code commits.5 A third team, building a tool that tells agents apart, found that “every one of the 458 Claude Code PRs includes a ‘Co-Authored-By: Claude’ line”, so naming the agent from that text “would be circular”.32

AIDev is the field’s shared dataset for studying coding agents: all 62 papers in this year’s MSR Mining Challenge used it.22 Of the 50 we could read, 24 report a result about Claude Code, and 15 papers elsewhere do too. That makes AIDev’s Claude Code slice, about 1 in 11 of the real thing, the basis for most of what has been published about how Claude Code behaves on GitHub. The full map lists all 47 papers with each finding and the sample behind it.

For most of them the slice matters little: many authors already note their Claude Code samples are small, and our own test found no significant difference in merge rate between AIDev’s slice and typical Claude Code PRs (82.7% against 88.3%, p = 0.22). It matters most where a finding depends on change size, or compares Claude Code on a handful of cases against thousands for other agents. Four examples:

StudyIts Claude Code findingBehind itWhere the slice bears on it
Agent contributions31“Claude (84.3%) and Codex (73.5%) achieve the highest PRs merge probabilities”, and Claude’s PRs “modify significantly more LOC”219 PRsThe PRs were picked by the “Co-Authored-By: Claude” line, which favours larger changes, so the size difference is partly built in.
Review wait20“Claude Code PRs wait longest for a first human review (median 12.6 hours)”, which the authors tie to PR size459 PRsAdds to the authors’ own size explanation: AIDev’s Claude Code slice runs larger than typical Claude Code PRs.
CI/CD reliability28“Claude had the lowest reliability at 64.86%”37 workflow runsThe comparison rests on 37 runs against tens of thousands for other agents.
Co-authorship21A positive Claude Code effect (+33.8 points) once Claude’s own address is excluded47 PRsThe PRs were found by the co-author line the study measures; the authors flag this themselves.

Quotes and counts are each paper’s own, checked against its text on 2 October 2026.

None of this makes these findings wrong. It is a reason to read Claude Code results built on AIDev as results about a particular, larger-than-usual set of Claude Code pull requests.

What we can't be sure of

Run it yourself

Every script is on GitHub, one folder per paper, with the null checks in nulls/ and the scheduled rechecks in rechecks/: github.com/MarkovianProtocol/second-measurements.

References

  1. Markovian Protocol. A research dataset finds only 1 in 11 of Claude Code's pull requests. SM-002, 2026. markovianprotocol.com/measurements/sm-002.html.
  2. Markovian Protocol. Google's Pixel software ledger was missing a whole month. SM-001, 2026. DOI 10.5281/zenodo.23070510.
  3. Google. Pixel Binary Transparency log, developers.google.com/android/binary_transparency/image_info.txt, tree size 1,163, last modified 22 September 2026.
  4. Hao Li et al. The Rise of AI Teammates in Software Engineering 3.0. arXiv:2507.15003, 2025.
  5. Detecting AI Coding Agents in Open Source: A Validated Multi-Method Census of 180 Million Repositories. arXiv:2606.24429, 2026. Table 6.
  6. Markovian Protocol. Half of a well-known collection of software ingredient lists came from one robot. SM-003, 2026. markovianprotocol.com/measurements/sm-003.html.
  7. Markovian Protocol. We checked whisper.online's public ledger with our own code, and it holds up. SM-004, 2026. DOI 10.5281/zenodo.23071390.
  8. Meta. Private Processing technical whitepaper, V2, 16 March 2026, pp. 16–17; Markovian Protocol, SM-005, markovianprotocol.com/measurements/sm-005.html.
  9. Shengchen Ling, Yajin Zhou, Lei Wu, Cong Wang. How Agentic Is Agentic Commerce? arXiv:2607.12575, 2026; Markovian Protocol, SM-006, markovianprotocol.com/measurements/sm-006.html.
  10. transparency.dev. Armored Witness production firmware log. Check script.
  11. Cloudflare. Plexi key-transparency auditor, namespace whatsapp.key-transparency.v2. Check script.
  12. Benny Toeppe, Amine Barrak, Emna Ksontini. A Large-Scale Dataset of MCP Implementations on GitHub. MSR 2026, arXiv:2607.10123. Check script.
  13. PyPI. PyPI in 2025: A Year in Review. blog.pypi.org, 31 December 2025. Check script.
  14. Same Name, Different Server: A Security Census of Silent Drift in the Model Context Protocol Ecosystem. arXiv:2609.14119, 2026. Check script.
  15. Never Emitted: Reporter Attribution in GitHub's Machine-Readable Vulnerability Records. arXiv:2609.33099, 2026. Check script.
  16. William Woodruff, Hayden Blauzvern. Homebrew's Sigstore-powered provenance is in beta. blog.sigstore.dev, 14 May 2024. Check script.
  17. Homebrew/homebrew-core issue 314985, 2 October 2026.
  18. Comparing AI Coding Agents: A Task-Stratified Analysis of Pull Request Acceptance. arXiv:2602.08915, 2026. Tables 1 and 3.
  19. A Task-Level Evaluation of AI Agents in Open-Source Projects. arXiv:2602.02345, 2026. Table 1.
  20. Not All Agents Are Equal: Code Quality and Post-Merge Maintenance Across Five Autonomous Coding Agents in the Wild. arXiv:2609.17598, 2026. Section 4.
  21. Beyond Simpson's Paradox: A Cascade of Confounders in AI Agent Pull-Request Co-Authorship. arXiv:2606.22711, 2026. Measurement limitations.
  22. MSR 2026 Mining Challenge, accepted papers. 2026.msrconf.org.
  23. AI builds, We Analyze: An Empirical Study of AI-Generated Build Code Quality. MSR 2026 Mining Challenge, arXiv:2601.16839.
  24. Quality and Security Signals in AI-Generated Python Refactoring Pull Requests. arXiv:2605.21453.
  25. Beyond Bug Fixes: An Empirical Investigation of Post-Merge Code Quality. MSR 2026 Mining Challenge, arXiv:2601.20109. Tables 1–2.
  26. How do Agents Refactor: An Empirical Study. MSR 2026 Mining Challenge, arXiv:2601.20160. Table 1.
  27. The Quiet Contributions: Insights into AI-Generated Silent Pull Requests. MSR 2026 Mining Challenge, arXiv:2601.21102. Figure 1.
  28. Reliability of AI Bots Footprints in GitHub Actions CI/CD Workflows. MSR 2026 Mining Challenge, arXiv:2604.18334. Table 1.
  29. Merged, Not Measured: An Empirical Study of Performance Issues Fixed by Coding Agents. arXiv:2609.37985.
  30. Security in the Age of AI Teammates: An Empirical Study of Agentic Pull Requests. arXiv:2601.00477.
  31. How Do AI Coding Agents Contribute to Software Development? arXiv:2607.21832. Pages 6, 7, 10 and 27.
  32. AgenTag: Attribution of AI Coding Agents from Behavioral Fingerprints. arXiv:2608.00966.

Cite as

@misc{markovian-review-2026-10,
  author = {{Markovian Protocol}},
  title  = {What six second measurements found, and what happened next},
  year   = {2026},
  month  = oct,
  url    = {https://markovianprotocol.com/measurements/review-2026-10.html}
}