SECOND MEASUREMENT SM-003

Wild SBOMs was billed as the work of many developers. Nearly half came from one robot

Markovian ProtocolMeasured 2026-09-30Status: sent to the authors 2026-09-30, response pendingDOI: 10.5281/zenodo.23123301

An SBOM is an ingredients list for software, and Wild SBOMs is a research collection of 78,612 of them, described as written by practitioners in the wild. We traced where they came from: 37,306, almost half, were made by a single automated run in 2022. That robot then sent a 2026 study down the wrong path. The “GitHub tool” it blamed for missing dependency links was the robot's own tool, from another company; GitHub's real exporter records those links in 78 of 78 projects we checked.

The short version

Disclosure: Markovian Protocol holds no financial position in any organisation named here, was paid by no one for this work, and showed it to no one before publication except the organisation measured. How we work.

Details

Wild SBOMs is a dataset of 78,612 SBOMs collected from public code through Software Heritage, an archive of public source code. Its authors describe the files as committed by many developers under many different conditions. A 2026 study used it to report that 52.9% of SBOMs don't record which component depends on which, and that GitHub's SBOM exporter never records it.

The dataset comes with a table of where each file was found. Matching the two shows that 37,306 files (47.5%) came from one Bitbucket repository, and none of them shows up anywhere else. That repository was a test setup that generated SBOMs of GitHub projects in bulk between April and August 2022. The tool it used, labelled "Github Extractor", is Fortress Information Security's Github Extractor 1.15.1, not GitHub's.

GitHub's own exporter appears in only 17 files, and 16 of them do record dependencies. We ran it today on 78 projects that have dependencies, and it recorded them for all 78. Using the 2026 study's own definitions, 55.2% of the dataset has no dependency links by our count, and 46.2% once the test setup's files are removed.

What they said, what we found

They saidWe found
The files were committed by practitioners, giving “more diversity in the conditions of generation”.147.5% came from one repository, made by one script, over five months in 2022.
“GitHub's dependency-graph exporter never emits edges (100%).”2That “GitHub Extractor” belongs to another company, Fortress Information Security. GitHub's real exporter recorded dependencies for 78 of the 78 projects we tried today.
52.9% of SBOMs in the wild have no dependency links.255.2% by our count, and 46.2% once you take the robot out.

Meet the robot

The repository is bitbucket.org/dlambert_fis/invoke_github_extractor.3 Its own description: “Python class to multiprocess github_extractor and validate resulting SBoMs.” Between 18 April and 25 August 2022, it generated SBOMs for other people's GitHub projects and saved each one as owner_repo_sbom.json: TheCherno_Hazel_sbom.json, faif_python-patterns_sbom.json, and 37,304 more.

Software Heritage has seen the dataset's files at 94 million other places on the internet. None of the robot's 37,306 shows up at any of them. They're one tool's test output, sitting in one repository, and they make up nearly half the dataset.

The GitHub tool that wasn't GitHub's

36,547 files name their tool as “Fortress Information Security Github Extractor 1.15.1”, and every one came from the robot. 4,630 of them name no other tool, and not one of those records dependencies. So the study's 100% is right, but it belongs to a different company.

GitHub's real exporter signs its files “GitHub.com-Dependency-Graph”.4 It appears in just 17 files in the dataset (June 2023 to May 2024), and 16 of them record dependencies. We also ran it live on 90 popular repositories. 84 returned an SBOM, and all 78 that had any dependencies listed them.

One catch: GitHub's older files only say “the project depends on each of these packages”, a flat list. Today, 41 of the 78 also record which package depends on which.

How much of the study's tool table is the robot

ToolFiles in the studyMade by the robot
Node.js module10,3209,901
cdx-php-composer6,3945,993
cyclonedx-gomod6,4145,557
“GitHub Extractor”4,5044,630
cdxgen17,4779,742

All five of the study's biggest tool rows are mostly robot output, cdxgen only just, all from tool versions pinned in 2022. (The robot's “GitHub Extractor” count is higher than the study's because the study doesn't say how it credits a file that lists several tools; we credited the first one listed.)

Take the robot out

No dependency linksConnected
Whole dataset (our count)55.2%32.7%
The robot alone64.9%27.0%
Everything except the robot46.2%38.2%

Our count for the whole dataset differs from the study's (55.2% against 52.9%). The study doesn't say exactly which relationship types it counts as dependencies, and that choice moves files between groups. So the fair comparison is within our own count, where removing the robot drops “no links” by 9 points.

What still holds

The study's main point survives. Whether an SBOM records dependencies depends on the tool that made it, and that's true even inside the robot's own run. The same script produced Node.js files mostly with no links and Go and PHP files fully linked. What changes is the phrase “in the wild”: nearly half of that wild is one test run from 2022. Anyone using Wild SBOMs should report their numbers with and without this repository.

How we checked

The dataset ships a table of every place Software Heritage saw each file: 123 million rows. We went through it twice, first to pull every file from the robot's repository, then to see whether any of those files shows up anywhere else. We scanned the 5 GB archive for every SBOM that mentions GitHub, recording its tools, date and dependency links. We sorted every file into the study's three groups (no links, mostly unconnected, connected) using the study's written definitions. And we asked GitHub's real exporter for the SBOMs of 10 random repositories with more than 500 stars in each of 9 languages.

The evidence

Fetched 2026-10-02T18:12:45Z. Copies and SHA-256 hashes: SHA256SUMS.

Exhibit 1 · How the dataset describes itself

“We introduce such a dataset, consisting of over 78 thousand unique SBOM files, deduplicated from those found in over 94 million repositories.”

zenodo.org/records/14250103 · saved record wild-sboms-zenodo-record.json

Exhibit 2 · Where 37,306 of them came from

swhid,origin_url
swh:1:cnt:00015f2ed2427040de2b35ae33e5c2cbfbae438e,https://bitbucket.org/dlambert_fis/invoke_github_extractor.git
swh:1:cnt:0003673596babe6957305248078a25e8ebe2053a,https://bitbucket.org/dlambert_fis/invoke_github_extractor.git
…
37,306 rows, 1 distinct origin

Software Heritage origins table in the dataset’s replication package (sboms-03-origins) · extract wild-sboms-origins-harness.csv

Exhibit 3 · The 2026 study’s conclusion

“GitHub’s dependency-graph exporter never emits edges (100%)”

The files behind it were tagged “GitHub Extractor”, a Fortress tool run by the robot, not GitHub’s exporter.

arXiv:2607.22140 · saved copy study-arxiv-2607.22140.pdf

Disclosure timeline

30 SepSent to the Wild SBOMs authors and the 2026 study’s author.
1 OctA co-author’s note reaches us; we send a clearer explanation and the specific ask.
1 OctPublished.

Waiting on the Wild SBOMs authors since 30 Sep.

What we can't be sure of

Run it yourself

The scripts are also on GitHub: github.com/MarkovianProtocol/second-measurements/sm-003.

./origins.sh               # the robot's files (streams the 12 GB origins table twice)
python3 scan_github.py < corpus.tar       # tool labels
python3 classify_all.py < corpus.tar      # the three groups
python3 github_live.py && python3 github_live_depth.py   # GitHub's exporter, live
python3 analyze.py         # every table

The corpus tar is sbom-files.tar.ztsd from Zenodo record 14250103, piped through zstd -dc. analyze.py also needs sboms-01.csv from the same record.

627a78afebcd01556f7e364066f43bd758ea53942fdad3184e1f79ea0d176928  scan_github.py
17679685a9818dfe68e7fd0d7704f70df42365d8bc74bfa744cd908b0859c703  classify_all.py
21ad65c43b629e091ea5a5bed1e31f8d1d5bf62f289e1c9b0fa1d2a538e3413a  origins.sh
721298a97ea94a9e8aad38efed7221e5618c28b43d09453fb851dc89bf705dc1  github_live.py
fa2708a8516cd5b1a5b7ff120f8d0df0e84c57a7312b17062d70ebe2215a42d0  github_live_depth.py
47ce16cfdd6a34f2bd6987d47eb46f125b885fa38bb991482a07c33a87805418  analyze.py

Reproduced from the public repository on 3 October 2026 on a clean machine: the live GitHub check holds on a fresh repository set (80 of 86 exported SBOMs carry links). The corpus steps need the 12 GB Zenodo table and were not run in the window. Full table on the track record.

What you can do with this

References

  1. L. Soeiro, T. Robert, S. Zacchiroli. Wild SBOMs: a Large-scale Dataset of Software Bills of Materials from Public Code. MSR 2025 Data and Tool Showcase. arXiv:2503.15021. Dataset: doi.org/10.5281/zenodo.14250103.
  2. A. Zięba-Kozarzewski. No Edges, No Verdict: A Large-Scale Empirical Study of Declared Dependency Graphs in 78K SBOMs in the Wild. arXiv:2607.22140v2, 2026. arxiv.org/abs/2607.22140.
  3. dlambert_fis. invoke_github_extractor. bitbucket.org/dlambert_fis/invoke_github_extractor. Accessed 2026-09-30.
  4. GitHub. Export a software bill of materials for your repository (REST: dependency-graph/sbom). docs.github.com/en/rest/dependency-graph/sboms.

Cite as

@misc{markovian-sm003,
  author = {{Markovian Protocol}},
  title  = {Half of the Wild SBOMs corpus is one test harness},
  number = {SM-003},
  doi    = {10.5281/zenodo.23123301},
  year   = {2026},
  month  = sep,
  url    = {https://markovianprotocol.com/measurements/sm-003.html}
}