One-Third of New arXiv Papers Now Read as Machine-Written
A calibrated detector finds ~32% of recent arXiv papers flag as AI-generated, with computer science at 65% and mathematics near zero — but the limitations matter.
There's a new entry in the 'N% of X is now AI' genre, and this one actually did its homework. The team behind Unslop ran a detector over 12,750 arXiv papers, anchoring their false-positive rate to pre-ChatGPT text so that only 0.4% of genuine human writing from 2021–2022 triggers a flag. The result: about a third of new papers submitted in the most recent quarter read as machine-written, with computer science leading at 65% and mathematics trailing at 0.7%.
The method, briefly
The detector is calibrated for academic prose. By setting the threshold so that pre-LLM papers flag at just 0.4%, the authors built in a control: if the rise were an artifact, 2021 and 2022 would show the same spike as 2026. They don't. The flagged share lifts off within months of ChatGPT's release and climbs in two waves, hitting ~39% in early 2026 before settling to ~32% over the most recent quarter.
Each paper was scored on its full body text, not just the abstract, because abstracts understate the signal — the same paper can score under 20% on its abstract and over 70% on its body. The sample covers ten field groups, roughly 25 papers per field per month from January 2023 to July 2026, plus eight control months across 2021 and 2022.
Field-by-field breakdown
- Computer science: 65% (control: 0.2%)
- Quantitative biology: 56.3% (control: 3.5%)
- Electrical engineering & systems: 51.3% (control: 1.7%)
- Economics & finance: 47% (control: 2.5%)
- Applied physics: 34% (control: 1.3%)
- Statistics: 31.3% (control: 1.8%)
- Condensed matter: 24% (control: 0%)
- High-energy physics: 14% (control: 0.5%)
- Astrophysics: 10.7% (control: 0%)
- Mathematics: 0.7% (control: 0%)
The fields that rise most are not the ones with the highest pre-LLM control level, so an elevated starting point doesn't explain the trend.
Where the measurement breaks
The authors are refreshingly honest about limitations. The per-field control samples are small (200 papers each), so the per-field false-positive rates are approximate. More importantly, a low score can mean either low adoption or a detector blind spot. Mathematics is the clearest example: papers dominated by notation and theorem-proof structure leave sparse prose that is out of distribution for the detector. A mathematics paper drafted with heavy model assistance may score low simply because its prose doesn't look like the training data. The authors note that the reported prevalence is a lower bound — incomplete detector coverage means the true share is at least what they measured.
Finally, a flag is not authorship. The detector estimates whether text reads as machine-written at a calibrated probability with a known error rate. It cannot separate a lightly-edited document from a wholly-generated one, and a single score is never grounds to accuse a specific person. The study reports the prevalence of machine-like writing, which includes heavy AI-assisted editing.
Take it for a spin
The detector is free and available for any arXiv paper or your own text. Worth a look if you're curious how much of the latest research in your field might have had AI assistance.
Source: unslop
Discussion
0 Comments
Be the first to start the discussion.