SECOND MEASUREMENT SM-002

The dataset behind 39 studies of Claude Code can see 1 in 11 of its pull requests

Markovian ProtocolMeasured 2026-09-30Status: sent to the authors 2026-09-30, response pendingDOI: 10.5281/zenodo.23123115

AIDev is the field's go-to dataset for studying AI coding agents, and 39 published studies use it to say how Claude Code behaves. To find Claude Code's pull requests it searches for a line Claude Code writes in commits, and skips the line it writes by default in pull-request descriptions. Over the same dates that search finds 16,065; searching for both finds 173,066. The ones it does catch run bigger than typical: a median of 736 changed lines against 447.

The short version

Disclosure: Markovian Protocol holds no financial position in any organisation named here, was paid by no one for this work, and showed it to no one before publication except the organisation measured. How we work.

Details

AIDev is a public dataset of pull requests (proposed code changes) opened by AI coding tools on GitHub. All 62 papers accepted to this year's MSR Mining Challenge, a research competition, used it. AIDev finds Claude Code's pull requests by searching their descriptions for "Co-Authored-By: Claude". Claude Code writes that line into commit messages, not descriptions. In descriptions it writes "Generated with Claude Code".

We ran both searches on GitHub over AIDev's own dates. AIDev's search returns 16,065; searching for either line returns 173,066, almost 11 times as many. To check this isn't just GitHub search changing over time, we reran AIDev's searches for the other AI tools too. Every one comes back at 85–88% of what AIDev recorded, same as Claude Code.

We also took a random sample from the bigger set and compared it with AIDev's. AIDev's pull requests are bigger: a typical one has 3 commits and 736 changed lines, against 2 and 447, and both differences are statistically significant. Merge rates were too close to call (82.7% against 88.3%).

What they said, what we found

They saidWe found
Claude Code pull requests are found by searching for “Co-Authored-By: Claude”, which Claude Code adds by default.1That line goes in commit messages. In pull request descriptions, where the search looks, Claude Code writes “🤖 Generated with Claude Code”.
Claude Code made 18,232 of AIDev's 2.7 million agent pull requests: 0.66%, last of six tools.2Searching for both lines finds 173,066 over the same dates, so AIDev caught about 1 in 11. That would put Claude Code fourth of six, not last.
The dataset's Claude Code pull requests are what researchers study as Claude Code's work.3The ones AIDev caught are bigger than typical: 736 lines changed against 447, and 3 commits against 2.

Why the search misses most of them

Four of AIDev's tools are found by something they stamp on every pull request: a branch name like codex/ or cursor/, or a bot account like Devin's. Claude Code is found by a line of text in the description:

ToolAIDev's GitHub searchFrom
OpenAI Codexis:pr head:codex/2025-05-16
Devinis:pr author:devin-ai-integration[bot]2024-12-24
GitHub Copilotis:pr head:copilot/ + login check2025-01-01
Cursoris:pr head:cursor/2025-01-01
Claude Codeis:pr "Co-Authored-By: Claude"2025-02-24

AIDev's searches, from the dataset paper's Table 2.

The line AIDev looks for, Co-Authored-By: Claude, is what Claude Code writes into commit messages. It only lands in a pull request description when the description happens to be built from commit text. The line Claude Code writes into descriptions is 🤖 Generated with [Claude Code](https://claude.ai/code) (later versions link to claude.com/claude-code).

Counting them

We ran each of AIDev's searches on GitHub today, over AIDev's own dates: each tool's start date to 24 October 2025, the last date in the dataset. Pull requests get deleted over time, so today's counts come out a bit lower than AIDev's. To measure that, we reran the searches for the other tools as a control. Every one comes back at 85–88% of what AIDev recorded, and Claude Code's comes back at 88%, right in line. So today's search is a fair stand-in.

SearchAIDev v4GitHub todayRatio
OpenAI Codex2,069,5951,752,8780.847
Devin43,29838,0720.879
Cursor212,544183,6670.864
Claude Code, AIDev's search18,23216,0650.881
Claude Code, the description line—171,794—
Claude Code, either line—173,066—

Same searches, same dates, run 30 September 2026. Copilot is left out because its login check can't be done in search.

AIDev's search catches 16,065 of 173,066, or 9.3%. Allow for the same 0.881 drift and the full set was about 196,000 when AIDev was collected. That would put Claude Code fourth of six tools, behind Codex, Copilot and Cursor. The other tools' searches have blind spots of their own (head:cursor/ only sees pull requests from a cursor/ branch, for example), and we didn't test those.

The Mining Challenge used an earlier cut of the dataset, up to 1 August 2025, with 5,137 Claude Code pull requests.4 Over those dates the gap is wider: AIDev's search finds 4,320 today, and either line finds 64,017, 14.8 times as many.

Is the description line really Claude Code?

We read 120 pull requests carrying it, 15 from each month. 113 had the exact line. 4 had a variant: the line without the link, or a sentence saying the pull request was made with Claude Code. 3 we couldn't tie to Claude Code from what's there now, and one of those was a Dependabot pull request. So the line is right 94–97.5% of the time. These 120 came from search result pages, so they aren't a perfectly random draw.

Are AIDev's ones typical?

We drew 120 Claude Code pull requests at random from AIDev (105 still exist) and 150 at random from the full set. GitHub search only returns 1,000 results per query, so to sample the full set fairly we picked random one-hour windows and kept each with a chance proportional to how many pull requests it had. We drew 792 windows. 2 had more than 150 pull requests, so they're slightly under-weighted.

AIDev's Claude CodeAll Claude Codep
Still exist105 of 120150 of 150
Typical commits320.001
Single-commit pull requests30.5%45.3%0.017
Typical lines changed7364470.012
Merged, of those closed82.7% (81 of 98)88.3% (128 of 145)0.22
Description has the Claude Code line95150
Description has the commit line967

Random samples, both from 24 February to 24 October 2025. “Typical” is the median. The full set includes AIDev's own (about 9%), so the real gap is a little bigger than shown. Only 7 of the 150 carry the exact commit line, because search matches phrases more loosely than an exact string.

AIDev's pull requests are bigger, by commits and by lines. We didn't work out why a description ends up carrying the commit line. Merge rates differ by 5.6 points, but these samples are too small to tell whether that's real; it would take about 600 closed pull requests on each side.

Who this affects

Any study that reports a Claude Code count, share or growth curve from AIDev is looking at about one pull request in eleven. Any study comparing Claude Code with other tools on size or commit structure is comparing a slice that runs bigger than usual. For example, a task-stratified study reports “Claude Code leads in documentation (92.3%)” from 139 Claude Code pull requests in AIDev's popular-repository subset.5 Its authors flag the small sample. Those 139 also came through a filter that picks larger pull requests. Acceptance itself didn't differ detectably in our data.

The evidence

Fetched 2026-10-02T18:11:35Z. Copies and SHA-256 hashes: SHA256SUMS.

Exhibit 1 · AIDev’s search, Table 2 of its paper

Agent           GitHub Search Query                     Start Date
Claude Code     is:pr “Co-Authored-By: Claude”          2025-02-24

arXiv:2507.15003 · saved copy aidev-arxiv-2507.15003.pdf

Exhibit 2 · AIDev on what Claude Code writes

“While Claude Code includes a default ‘Co-Authored-By: Claude’ message (which can be disabled), OpenAI Codex provides no attribution at all.”

arXiv:2507.15003, same copy

Exhibit 3 · GitHub’s own count, 2026-10-02

is:pr "Co-Authored-By: Claude" created:2025-02-24..2025-10-24
  total_count: 16,063

is:pr "Generated with Claude Code" OR "Co-Authored-By: Claude" created:2025-02-24..2025-10-24
  total_count: 173,101

Same dates. Search counts drift a little day to day as pull requests are deleted.

api.github.com/search/issues · raw github-search-counts.json

Disclosure timeline

30 SepSent to AIDev’s authors.
1 OctPublished; follow-up sent.
2 OctWe find two other teams had independently reached the same result (see the review).

Waiting on AIDev’s authors since 30 Sep.

What we can't be sure of

Since publication

2 October 2026. At least 39 published studies draw conclusions about Claude Code from AIDev’s slice, 24 of them in the MSR 2026 Mining Challenge; the October review lists them with the sample behind each finding. Two other teams reached the undercount independently: a census of 180 million repositories finds AIDev misses 79% of projects with Claude Code commits (arXiv:2606.24429), and a team building an agent classifier notes that every Claude Code PR in AIDev carries the line it was found by (arXiv:2608.00966).

Run it yourself

The scripts are also on GitHub: github.com/MarkovianProtocol/second-measurements/sm-002.

Scripts need Python 3, the GitHub CLI logged in, and for the AIDev sample pyarrow and huggingface_hub.

python3 counts.py         # the counts
python3 aidev_sample.py   # 120 AIDev v4 Claude Code PRs, seed 20260930
python3 union_sample.py   # 150 union PRs by hour-window rejection sampling, seed 20260930
python3 compare.py        # the comparison
be7d2e6f3faea12a9a83e80d2a1c50f0c476ec8037eada50bbdbcdfbf05ebede  counts.py
6a84658973658ceb7c4e8540e98cc7790c36e106a9ecf93cdcab6e64354bbb8e  aidev_sample.py
0c9e0aa3bca14719dfac3e83460d81672dff3e1a744f9c71bb846daa3b64fca6  union_sample.py
a20e9a0f5d82024d83dd55ee092202f879b5e3ce48c2ec6aa0c127e5ff00b8d8  compare.py

Search counts move as PRs are deleted, and the union sample depends on search results at run time. A rerun will land near these figures but not on them.

Reproduced from the public repository on 3 October 2026 on a clean machine: the counts land within 0.1% (16,060 and 173,135 against 16,065 and 173,066; GitHub search moves daily). The union sample takes about 50 minutes of search calls. Full table on the track record.

What you can do with this

References

  1. H. Li, H. Zhang, A. E. Hassan. The Rise of AI Teammates in Software Engineering (SE) 3.0: How Autonomous Coding Agents Are Reshaping Software Engineering. arXiv:2507.15003, 2025. Table 2 and §3.1. arxiv.org/abs/2507.15003.
  2. H. Li et al. AIDev dataset, v4 (AIDev-2.7M). huggingface.co/datasets/hao-li/AIDev, tag v4. Accessed 2026-09-30.
  3. MSR 2026. Mining Challenge. 2026.msrconf.org/track/msr-2026-mining-challenge.
  4. H. Li, H. Zhang, A. E. Hassan. AIDev: Studying AI Coding Agents on GitHub. Mining Challenge paper, dataset cutoff 2025-08-01. github.com/SAILResearch/AI_Teammates_in_SE3.
  5. G. Pinna, J. Gong, D. Williams, F. Sarro. Comparing AI Coding Agents: A Task-Stratified Analysis of Pull Request Acceptance. MSR 2026 Mining Challenge. arXiv:2602.08915. arxiv.org/abs/2602.08915.

Cite as

@misc{markovian-sm002,
  author = {{Markovian Protocol}},
  title  = {How AIDev identifies Claude Code pull requests,
            and what the identification misses},
  number = {SM-002},
  doi    = {10.5281/zenodo.23123115},
  year   = {2026},
  month  = sep,
  url    = {https://markovianprotocol.com/measurements/sm-002.html}
}