SECOND MEASUREMENTS · REVIEW, OCTOBER 2026
What six second measurements found, and what happened next
Six claims about public logs and research datasets were rechecked in September, and four turned out wrong in a way that mattered. A month on, two of the four are fixed or confirmed and two are still wrong, and one of them sits under at least 39 published studies.
Disclosure: Markovian Protocol holds no financial position in any organisation named here, was paid by no one for this work, and showed it to no one before publication except the organisation measured. How we work.
A second measurement takes a number someone has published about their own system or dataset and gets it again by a different route, with code anyone can rerun.1 Since 15 September we have published six papers and seven checks that found nothing. This review puts them side by side, reports what each looks like when rerun today, and follows one finding, about the AIDev dataset of AI coding-agent pull requests, into the studies built on it.
Three more papers and two more claims that held, all after this edition closed. SM-009: Firefox’s revocation list misses a third of certificates revoked on their first day. SM-010: Hugging Face’s malware scanner reads nothing over 2 GB, and the badge says safe. SM-011: that same Firefox list is months behind the largest certificate logs, and one in eleven top sites, mozilla.org included, gets no revocation check. Held: Debian’s reproducibility gate and every Certificate Transparency log’s merge-delay promise, on the index. The next edition will fold them in.
Every check so far
| Check | The claim | Held? | Since then |
|---|---|---|---|
| SM-001 | Google's Pixel log covers every factory image from Pixel 6 on2 | No: the January 2026 release (29 images) was missing | Google confirmed and added them on 22 Sep3 |
| SM-002 | AIDev captures each agent's pull requests, Claude Code's included4 | No: its search finds about 1 in 11, and larger ones than usual | Independently confirmed by a 180-million-repository census5; no reply from the authors yet |
| SM-003 | Wild SBOMs are written by many practitioners6 | No: 47.5% come from one automated run | One author replied; no correction yet |
| SM-004 | whisper.online's ledger proves nothing was changed7 | Yes, every proof checked | Operator confirmed, fixed a duplicate-update bug we measured |
| SM-005 | Meta republishes its pulled-software list every 3 hours and ships server software weekly8 | No: every 6 hours since January; new software in 37 of 70 weeks | Still 6 hours on 2 Oct; whitepaper unchanged |
| SM-006 | 136.7 million x402 payments on Base in 280 days, most of them manufactured9 | Not tested. Our one-day count reproduced exactly; payments then fell by three-quarters, and two wallet groups still sent about 80% | Rechecked monthly from now on |
| Null | Armored Witness firmware log names its source10 | 27 of 28 | — |
| Null | Cloudflare's Plexi audits every WhatsApp key-transparency epoch11 | 150 of 150 sampled | — |
| Null | An MSR 2026 dataset describes real-world MCP servers12 | Yes, though it keeps only repos with 50+ stars, unstated | Authors told |
| Null | 17% of 2025 PyPI uploads carry an attestation13 | 17.9% in our sample | — |
| Null | 4.2% of MCP registry servers redirected their endpoint14 | 4.02% | — |
| Null | GitHub leaves reporter credit out of the CVEs it assigns15 | 0 of 150 carry it | — |
| Null | Homebrew attests every bottle its CI builds16 | 599 of 600 | One bottle reported to Homebrew17 |
"Held?" refers to the claim as its owner stated it. Null checks are listed on the measurements index with their scripts.
Big organisations measuring their own systems mostly held up: PyPI, Homebrew, GitHub and Cloudflare all did. The two lasting misses came from somewhere else: a vendor whitepaper that stopped matching its own log (SM-005), and a research dataset built on a single search (SM-002).
Rechecked today
Three of the checks now rerun on a schedule, and every result is kept on a rechecks page.
| When published | 2 October 2026 | |
|---|---|---|
| SM-001 Pixel images missing from the log | 33 (15 Sep) | 0 |
| SM-001 Pixel entries a phone can't look up | 103 | 128 |
| SM-005 Meta pulled-software list, median gap | 6.0 h | 6.0 h (whitepaper still says 3) |
| SM-005 newest server-software entry | 29 Sep | 29 Sep |
| SM-006 x402 payments in one day | 234,490 (5 Aug) | 63,003 (1 Oct) |
Pixel figures from Google's image page against the log's image_info.txt; Meta figures from Cloudflare Plexi, namespace prod.pc.revocation_list over the last 30 days.
The Pixel fix held, but the other Pixel problem grew. Entries filed under a label phones don't report, so a phone can't find its own, went from 103 to 128.2
One dataset, 39 studies, and the Claude Code it doesn’t see
AIDev finds Claude Code’s pull requests by searching for a line Claude Code writes in commits, not the line it writes in pull-request descriptions. Over the same dates that search finds 16,065; searching for both finds 173,066. The ones it finds are also bigger than typical: a median of 736 changed lines against 447.1 A census of 180 million repositories reached the same conclusion from the commit side: AIDev holds 5,137 Claude Code pull requests against 850,157 Claude Code commits that census found, and a pull-request census like AIDev misses 79% of the projects where they found Claude Code commits.5 A third team, building a tool that tells agents apart, found that “every one of the 458 Claude Code PRs includes a ‘Co-Authored-By: Claude’ line”, so naming the agent from that text “would be circular”.32
AIDev is the field’s shared dataset for studying coding agents: all 62 papers in this year’s MSR Mining Challenge used it.22 Of the 50 we could read, 24 report a result about Claude Code, and 15 papers elsewhere do too. That makes AIDev’s Claude Code slice, about 1 in 11 of the real thing, the basis for most of what has been published about how Claude Code behaves on GitHub. The full map lists all 47 papers with each finding and the sample behind it.
For most of them the slice matters little: many authors already note their Claude Code samples are small, and our own test found no significant difference in merge rate between AIDev’s slice and typical Claude Code PRs (82.7% against 88.3%, p = 0.22). It matters most where a finding depends on change size, or compares Claude Code on a handful of cases against thousands for other agents. Four examples:
| Study | Its Claude Code finding | Behind it | Where the slice bears on it |
|---|---|---|---|
| Agent contributions31 | “Claude (84.3%) and Codex (73.5%) achieve the highest PRs merge probabilities”, and Claude’s PRs “modify significantly more LOC” | 219 PRs | The PRs were picked by the “Co-Authored-By: Claude” line, which favours larger changes, so the size difference is partly built in. |
| Review wait20 | “Claude Code PRs wait longest for a first human review (median 12.6 hours)”, which the authors tie to PR size | 459 PRs | Adds to the authors’ own size explanation: AIDev’s Claude Code slice runs larger than typical Claude Code PRs. |
| CI/CD reliability28 | “Claude had the lowest reliability at 64.86%” | 37 workflow runs | The comparison rests on 37 runs against tens of thousands for other agents. |
| Co-authorship21 | A positive Claude Code effect (+33.8 points) once Claude’s own address is excluded | 47 PRs | The PRs were found by the co-author line the study measures; the authors flag this themselves. |
Quotes and counts are each paper’s own, checked against its text on 2 October 2026.
None of this makes these findings wrong. It is a reason to read Claude Code results built on AIDev as results about a particular, larger-than-usual set of Claude Code pull requests.
What we can't be sure of
- 12 of the 62 Mining Challenge papers have no public preprint and could not be read, so 25 is a lower bound.
- The map covers arXiv and the Mining Challenge. Papers in other venues are likely missing.
- Our test of AIDev's slice used 120 and 150 pull requests. It found a size difference, not a merge-rate one; a bigger sample could find either.
- Recheck numbers are single runs on the dates shown.
Run it yourself
Every script is on GitHub, one folder per paper, with the null checks in nulls/ and the scheduled rechecks in rechecks/: github.com/MarkovianProtocol/second-measurements.
References
- Markovian Protocol. A research dataset finds only 1 in 11 of Claude Code's pull requests. SM-002, 2026. markovianprotocol.com/measurements/sm-002.html.
- Markovian Protocol. Google's Pixel software ledger was missing a whole month. SM-001, 2026. DOI 10.5281/zenodo.23070510.
- Google. Pixel Binary Transparency log,
developers.google.com/android/binary_transparency/image_info.txt, tree size 1,163, last modified 22 September 2026. - Hao Li et al. The Rise of AI Teammates in Software Engineering 3.0. arXiv:2507.15003, 2025.
- Detecting AI Coding Agents in Open Source: A Validated Multi-Method Census of 180 Million Repositories. arXiv:2606.24429, 2026. Table 6.
- Markovian Protocol. Half of a well-known collection of software ingredient lists came from one robot. SM-003, 2026. markovianprotocol.com/measurements/sm-003.html.
- Markovian Protocol. We checked whisper.online's public ledger with our own code, and it holds up. SM-004, 2026. DOI 10.5281/zenodo.23071390.
- Meta. Private Processing technical whitepaper, V2, 16 March 2026, pp. 16–17; Markovian Protocol, SM-005, markovianprotocol.com/measurements/sm-005.html.
- Shengchen Ling, Yajin Zhou, Lei Wu, Cong Wang. How Agentic Is Agentic Commerce? arXiv:2607.12575, 2026; Markovian Protocol, SM-006, markovianprotocol.com/measurements/sm-006.html.
- transparency.dev. Armored Witness production firmware log. Check script.
- Cloudflare. Plexi key-transparency auditor, namespace
whatsapp.key-transparency.v2. Check script. - Benny Toeppe, Amine Barrak, Emna Ksontini. A Large-Scale Dataset of MCP Implementations on GitHub. MSR 2026, arXiv:2607.10123. Check script.
- PyPI. PyPI in 2025: A Year in Review. blog.pypi.org, 31 December 2025. Check script.
- Same Name, Different Server: A Security Census of Silent Drift in the Model Context Protocol Ecosystem. arXiv:2609.14119, 2026. Check script.
- Never Emitted: Reporter Attribution in GitHub's Machine-Readable Vulnerability Records. arXiv:2609.33099, 2026. Check script.
- William Woodruff, Hayden Blauzvern. Homebrew's Sigstore-powered provenance is in beta. blog.sigstore.dev, 14 May 2024. Check script.
- Homebrew/homebrew-core issue 314985, 2 October 2026.
- Comparing AI Coding Agents: A Task-Stratified Analysis of Pull Request Acceptance. arXiv:2602.08915, 2026. Tables 1 and 3.
- A Task-Level Evaluation of AI Agents in Open-Source Projects. arXiv:2602.02345, 2026. Table 1.
- Not All Agents Are Equal: Code Quality and Post-Merge Maintenance Across Five Autonomous Coding Agents in the Wild. arXiv:2609.17598, 2026. Section 4.
- Beyond Simpson's Paradox: A Cascade of Confounders in AI Agent Pull-Request Co-Authorship. arXiv:2606.22711, 2026. Measurement limitations.
- MSR 2026 Mining Challenge, accepted papers. 2026.msrconf.org.
- AI builds, We Analyze: An Empirical Study of AI-Generated Build Code Quality. MSR 2026 Mining Challenge, arXiv:2601.16839.
- Quality and Security Signals in AI-Generated Python Refactoring Pull Requests. arXiv:2605.21453.
- Beyond Bug Fixes: An Empirical Investigation of Post-Merge Code Quality. MSR 2026 Mining Challenge, arXiv:2601.20109. Tables 1–2.
- How do Agents Refactor: An Empirical Study. MSR 2026 Mining Challenge, arXiv:2601.20160. Table 1.
- The Quiet Contributions: Insights into AI-Generated Silent Pull Requests. MSR 2026 Mining Challenge, arXiv:2601.21102. Figure 1.
- Reliability of AI Bots Footprints in GitHub Actions CI/CD Workflows. MSR 2026 Mining Challenge, arXiv:2604.18334. Table 1.
- Merged, Not Measured: An Empirical Study of Performance Issues Fixed by Coding Agents. arXiv:2609.37985.
- Security in the Age of AI Teammates: An Empirical Study of Agentic Pull Requests. arXiv:2601.00477.
- How Do AI Coding Agents Contribute to Software Development? arXiv:2607.21832. Pages 6, 7, 10 and 27.
- AgenTag: Attribution of AI Coding Agents from Behavioral Fingerprints. arXiv:2608.00966.
Cite as
@misc{markovian-review-2026-10,
author = {{Markovian Protocol}},
title = {What six second measurements found, and what happened next},
year = {2026},
month = oct,
url = {https://markovianprotocol.com/measurements/review-2026-10.html}
}