SECOND MEASUREMENT SM-010
Hugging Face says every file goes through its malware scanner. Nothing over 2 GB does, and the badge says safe anyway
Hugging Face’s documentation says every file in every repository goes through a malware scanner at each commit, within minutes. We read the per-file scan status Hugging Face itself publishes on 3,038 files across 450 repositories, including the 150 most-downloaded models. Under 2 gigabytes the claim holds: 2,539 of 2,551 files scanned. At 2 gigabytes and above it stops: 0 of 487. Those are the model weights, 89% of the bytes in the sample and 92% in the top models, and 476 of the 494 files the scanner never read wear a green “safe” badge.
- Hugging Face: “We run every file of your repositories through a malware scanner. Scanning is triggered at each commit.”
- Its own API reports each file’s result. We read it for 3,038 files in 450 repositories, twice, by two endpoints.
- Files under 2 GiB: 2,539 of 2,551 scanned. Files of 2 GiB or more: 0 of 487. The largest scanned file is 1.994 GiB; the smallest skipped one over a gigabyte is 2.026 GiB.
- That is 89% of the bytes in the sample, and 55 of the 149 most-downloaded models have at least one skipped file.
- 476 of the 494 skipped files show the “safe” badge, because other checks passed. Seven are PyTorch pickle files, the format that can run code on load; the separate pickle scanner cleared them.
- No size limit is documented. Most of the skipped weight files are safetensors or GGUF, formats that do not execute code, so the gap is in the promise more than in the risk.
Disclosure: Markovian Protocol holds no financial position in any organisation named here, was paid by no one for this work, and showed it to no one before publication except the organisation measured. How we work.
Every file on the Hugging Face Hub carries a security status that the website shows as a badge and the API returns as a field, securityFileStatus, with one entry per scanner: Hugging Face’s own malware scanner (avScan, ClamAV according to the documentation), its pickle-import scanner, and partner scanners from Protect AI, JFrog and VirusTotal. The overall badge is a roll-up. We drew 450 repositories: the 150 most-downloaded models, the 150 most-downloaded datasets, and 150 models picked at random from recently updated ones. For each we listed the root directory with the API’s expand option, which returns the status fields, and kept files committed at least a day earlier. 54 files in three gated repositories (two meta-llama models and one dataset) can’t be read without accepting terms and were left out. Every file the listing called unscanned or left blank was looked up again through a second endpoint, paths-info; the two agreed on 479 of 483 unscanned verdicts; the other 4 were in gated repositories, and most blanks (105 of 121 readable) turned out to be listing gaps; 15 came back unscanned and are counted, so the counts here use the second answer.
What they said, what we found
| Hugging Face said | We found |
|---|---|
| Hugging Face, pull request huggingface_hub#2397, July 2024: “ClamAV scans files of up to 3.9GB. By lowering the shard size, we can ensure files are scanned.” A user asked on the Hub forum in August 2023 how files over 2 GB were scanned at all.4 | The cut we measured is 2 GiB, not 3.9 GB: 0 of 487 files at 2 GiB or more carry a scan result. The limit was known inside Hugging Face in 2024; the badge did not change. |
| “We run every file of your repositories through a malware scanner. Scanning is triggered at each commit.”1 | Under 2 GiB: 2,539 of 2,551 files scanned. 2 GiB and over: 0 of 487, including every weight file in the most-downloaded models. |
| “It can take up to a few minutes to be scanned.”1 | The skipped files were committed between 15 days and 6.8 years before the run. The oldest dates from December 2019. |
| “If at least one file has a been scanned as unsafe, a message will warn the users” (sic)1 | 476 of the 494 skipped files show the “safe” badge; 12 show “unscanned”; 6 show “queued”. |
| JFrog: “Model files are scanned by the JFrog scanner and we expose the scanning results on the Hub interface.”2 | 49 of the 1,132 files older than a year that carry a JFrog status are still “queued”, including five files in GPT-2 committed in 2019, 2020 and 2021. |
The cut-off
The malware scanner stops at 2 GiB, and the edge is sharp. Nothing above it is scanned; almost nothing below it is missed.
| File size | Files | Scanned | Not scanned |
|---|---|---|---|
| under 1 MB | 1,699 | 1,693 | 6 |
| 1 MB to 100 MB | 348 | 344 | 0 |
| 100 MB to 1 GB | 342 | 340 | 1 |
| 1 GB to 2 GB | 145 | 145 | 0 |
| 2 GB to 3 GB | 78 | 17 | 61 |
| 3 GB to 5 GB | 252 | 0 | 252 |
| over 5 GB | 174 | 0 | 174 |
Decimal gigabytes in the buckets; the cut sits at 2 GiB (2,147,483,648 bytes), which falls inside the 2 to 3 GB row. The 17 scanned files in that row are all below 2 GiB. Four files marked “suspicious” and one “error” are counted as neither.
Where the bytes are, by group:
| Group | Files | Not scanned | Bytes not scanned | Repositories with a skipped file |
|---|---|---|---|---|
| 150 most-downloaded models | 1,955 | 388 | 92% | 55 of 149 |
| 150 most-downloaded datasets | 856 | 102 | 80% | 6 of 122 |
| 150 recently updated models, random | 227 | 4 | 75% | 1 of 23 |
Repository counts are those with at least one readable file older than a day. Random recent models are small; most have no file near the limit.
The ten most-downloaded models by skipped volume:
| Repository | Files not scanned | Volume | Badge shown |
|---|---|---|---|
| unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF | 26 of 31 | 445 GB | safe |
| unsloth/Qwen3.8-27B-GGUF | 24 of 30 | 414 GB | safe |
| zai-org/GLM-5.3-Flash | 43 of 49 | 231 GB | safe |
| ornith-ai/Ornith-1.5-35B-A3B-GGUF | 5 of 8 | 185 GB | queued |
| deepseek-ai/DeepSeek-V3.2 | 40 of 47 | 173 GB | safe |
| deepseek-ai/DeepSeek-V4-Flash-0731 | 42 of 48 | 150 GB | safe |
| farbodtavakkoli/OTel-2.0-LLM-31B-IT | 28 of 39 | 126 GB | safe |
| dphn/dolphin-2.9.1-yi-1.5-34b | 14 of 24 | 68 GB | safe |
| prism-ml/Ternary-Bonsai-2-27B-gguf | 3 of 10 | 67 GB | safe |
| Qwen/Qwen3-32B | 17 of 27 | 66 GB | safe |
Most of that is safetensors and GGUF, formats that hold numbers and cannot run code when loaded. The format that can is pickle, which PyTorch’s .bin files use. Seven of those in the sample are over the limit. The malware scanner never read them; Hugging Face’s separate pickle-import scanner did, and found only ordinary imports.
| File | Size | Pickle scanner | Badge |
|---|---|---|---|
| BAAI/bge-m3/pytorch_model.bin | 2.27 GB | safe | safe |
| intfloat/multilingual-e5-large/pytorch_model.bin | 2.24 GB | safe | safe |
| openai/whisper-large-v3/pytorch_model.bin | 3.09 GB | safe | safe |
| openai/whisper-large-v3/pytorch_model.fp32-00001-of-00002.bin | 4.99 GB | safe | safe |
| E-MIMIC/inclusively-reformulation-it5/pytorch_model.bin | 3.13 GB | safe | safe |
| facebook/esm2_t33_650M_UR50D/pytorch_model.bin | 2.61 GB | safe | safe |
| FacebookAI/xlm-roberta-large/pytorch_model.bin | 2.24 GB | safe | safe |
How this ends
It ends when the Hub scans files of 2 GiB and above, or when the documentation and the badge say that it doesn’t. The recheck re-reads the same 487 files; the day their status changes from “unscanned”, or the page says “not scanned” next to the badge, is the day this closes.
The evidence
Every quote below was copied from the source and checked against a saved copy, fetched 2026-10-02. The copies, the API responses, and the full sample with both readings per file are published next to this page (SHA256SUMS).
Exhibit 1 · The claim
“We run every file of your repositories through a malware scanner. Scanning is triggered at each commit. … If your file has neither an ok nor infected badge, it could mean that it is either currently being scanned, waiting to be scanned, or that there was an error during the scan. It can take up to a few minutes to be scanned.”
huggingface.co/docs/hub/security-malware · saved copy hf_security_malware.html
Exhibit 2 · openai/whisper-large-v3, the model’s own answer
POST /api/models/openai/whisper-large-v3/paths-info/main {"paths": [...], "expand": true}
pytorch_model.bin 3,087,394,553 bytes status: "safe" avScan: "unscanned"
model.safetensors 3,087,130,976 bytes status: "safe" avScan: "unscanned"
config.json 1,272 bytes status: "safe" avScan: "safe"raw response whisper-large-v3-paths-info.json · the same for four of Qwen3-8B’s five weight shards; the fifth, 1.24 GB, is scanned: qwen3-8b-tree.json
Exhibit 3 · GPT-2, queued since 2019
openai-community/gpt2 64-fp16.tflite committed 2019-12-12 status: "queued" avScan: "safe" protectAiScan: "safe" virusTotalScan: "safe" jFrogScan: "queued"
raw response gpt2-paths-info.json; the same three calls returned the same answer
Exhibit 4 · Two endpoints, one answer
80 files re-read through paths-info: 40 marked "unscanned" by the listing -> 40 "unscanned" (overall badge differed on 2) 20 with no status in the listing -> 12 "safe", 1 "unscanned", 7 gated (401) 20 marked "safe" -> 18 "safe", 2 gated Then every non-safe file (658) re-read the same way: 479 of 483 "unscanned" confirmed.
verify_paths_info.json · full sample with both readings: hf_scan_sample.json
Our questions to Hugging Face
Sent to security@huggingface.co and filed as huggingface/hub-docs issue 2850 on 2 October 2026. Answers will be printed here as written.
- Is 2 GiB a deliberate limit on the malware scanner? If so, can the documentation say so?
- Why does the overall badge show “safe” for a file the malware scanner has not read?
- Are large pickle-format files, such as the seven PyTorch
.binfiles over 2 GiB here, covered by any scanner other than the pickle-import check? - JFrog shows “queued” on files committed in 2019. Is that queue still being worked, or is it abandoned?
- Do Enterprise Hub customers get a different size limit?
Their reply, scored
Disclosure timeline
| 2 Oct | Published; sent to Hugging Face with the five questions above, as huggingface/hub-docs issue 2850 and by email to security@huggingface.co. |
Waiting on Hugging Face since 2 Oct.
What we can’t be sure of
- We read Hugging Face’s reported status, not the scanner itself. A file could have been scanned and its result lost; the field says “unscanned”, and we take it at its word, as users do.
- Only root directories were listed. Files in subdirectories, common in large multi-part models, are not in the sample. Their sizes suggest they would push the share of skipped bytes up, not down.
- The sample over-represents popular repositories. That is where the big files are; a random walk of the Hub would find far fewer files over 2 GiB.
- A first version of this check, in September, died on our own bug: the listing API’s recursive option drops the status field, which we had read as missing scans. This version uses direct listings and a second endpoint for every non-safe file.
- Safetensors and GGUF cannot execute code on load, so a malware scan of them catches little that matters. The gap is between the documentation and the practice. The pickle files are the exception, and the pickle-import scanner did read them.
Run it yourself
Python 3 standard library. About half an hour against the public API; no token needed except for gated repositories, which this skips.
python3 hf_scan_sample.py # 450 repositories, root listings with expand=true python3 hf_requery.py # paths-info for every non-safe file, then the tables
Reproduced from the public repository on 3 October 2026 on a clean machine, on a fresh sample: holds (0 of 688 files of 2 GiB or more scanned; 2,502 of 2,801 smaller files safe; 519 unscanned files behind a safe badge). Full table on the track record.
What you can do with this
- If you download a model file of 2 GiB or more from the Hub, the green badge means nothing about that file; scan it yourself.
- If you publish large weights, say in the model card that the Hub's scanner skipped them.
References
- Hugging Face. Malware Scanning. huggingface.co/docs/hub/security-malware.
- Hugging Face. Third-party scanner: JFrog. huggingface.co/docs/hub/security-jfrog; Third-party scanner: Protect AI; Pickle Scanning.
- Hugging Face Hub API:
GET /api/{models|datasets}/{repo}/tree/main?expand=true,POST /api/{models|datasets}/{repo}/paths-info/main. - huggingface/huggingface_hub, pull request #2397, “Reduce max shard size”, July 2024; Hugging Face forum, ClamAV - Scanning Files Larger than 2GB, August 2023. Saved copies in the exhibits.
Cite as
@misc{markovian-sm010,
author = {{Markovian Protocol}},
title = {Hugging Face malware scanning by file size, October 2026},
number = {SM-010},
doi = {10.5281/zenodo.23123005},
year = {2026},
month = oct,
url = {https://markovianprotocol.com/measurements/sm-010.html}
}