🔍 Read the full analysis: The AI Content Surge Is Outpacing Our Ability To Check It on ThorstenMeyerAI.com
Get office and shipping supplies delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
A report from ThorstenMeyerAI.com argues that AI is increasing the supply of mathematical, software and professional work faster than people can check it. The cited figures suggest review delays, missed reviews and continuing dependence on human judgment, though several software metrics come from vendors that sell review tools.
A report published this week by ThorstenMeyerAI.com says AI systems are producing mathematical manuscripts, code changes and contract work faster than people can verify them, creating a growing review-capacity gap. The report cites OpenAI’s publication of 722 mathematics manuscripts and software-industry data showing longer review waits and, in some cases, changes merged without human review; the figures point to a challenge for organisations seeking to use AI output safely and reliably.
In mathematics, OpenAI posed about 4,000 problems to a model and published 722 manuscripts grouped into 372 families, according to the source. The average result reportedly required about three hours of computing. Some work was formally checked using Lean, a proof assistant, while OpenAI cautioned that unformalized results could have issues. The source contrasts that output with the careful verification by five leading mathematicians of an earlier result from the programme: a counterexample to an old Erdős conjecture.
The report also points to software data as evidence of pressure on review. Faros AI found teams merged 98% more pull requests during periods of high AI adoption, while review time rose 91%. LinearB, analysing 8.1 million pull requests from 4,800 organisations, reported that AI-generated changes waited 4.6 times longer for review to begin. It found those changes were accepted 32.7% of the time, compared with 84.4% for human-written changes. A peer-reviewed 2026 study cited in the report found 61% of AI-agent pull requests received no human review before they were merged or closed.
Those measurements have limitations. Faros AI and LinearB sell software related to engineering workflows or code review, so their figures should be read with that commercial interest in mind. The report says Faros also found merges with zero review rose 31.3% during high-adoption periods. The numbers come from different studies and measures; they do not establish that AI alone caused every change in review practice.
In professional services, the source describes OpenAI’s partnership with contract-software company Ironclad, which involved training GPT-6 Astra on real contracting workflows. On 11 tasks, Astra met an average of 55% of the evaluation criteria, the source says. That result indicates progress in the tested tasks, but leaves human reviewers to identify errors or missed requirements before a draft can be used. The source does not provide further details about the evaluation method.
The referee shortage: AI made doing cheap and checking expensive
OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.
Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.
Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.
Real progress. Someone still has to find the other 45% before the work can be used.
Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.
A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.
Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.
of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).
of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.
OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.
Budget review hours next to model spend.
Provers, types, tests, policy engines.
Experts only where consequences are high.
Who profits from generation pays for checking.
Keep some production human for learners.
The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.
Reviewers Become the Constraint
The gap matters because a high volume of output does not automatically translate into useful, safe work. Someone must decide whether a result answers the right question, whether a code change fits the system, or whether a contract draft meets the relevant requirements. When that work is delayed or skipped, organisations face a choice between accepting greater risk and limiting how much AI output they can put into use.
The report describes three possible responses already visible in its cited examples: rubber-stamping work that receives no review, delaying machine-generated changes out of caution, and relying on the producer to decide which outputs merit attention. The last approach can be useful, but it leaves the producer’s selection process doing work that independent review would otherwise perform.
It also raises a workforce question. Senior reviewers typically develop judgment through years of doing and checking the work themselves. If AI takes over much of that early-career work, employers may eventually have fewer people with the experience required to assess complex outputs. The report argues that organisations will need to preserve ways for junior staff to build expertise while adopting AI.
As an affiliate, we earn on qualifying purchases.
Three Fields, Similar Pressure
The report’s examples span mathematics, software and contracting, but the verification problem differs in each. Formal proof tools can establish that a proof follows from its stated assumptions; they do not determine whether the theorem is the one researchers needed to address or whether it is important. In software, tests can check specified behaviour without showing that the tests cover every requirement. Contract review also involves obligations and institutional rules that may not be captured by a model’s evaluation score.
Across these fields, the source distinguishes checking an output against a defined test from judging whether the test itself captures the real-world need. It also points to accountability: mathematicians, engineers and lawyers can be questioned or held responsible for work in ways that a model cannot. Those limits help explain why faster generation does not automatically produce faster, dependable adoption.
The source frames the economic effect as a shift in the value of work: when production becomes cheaper but trusted review remains scarce, people able to verify and take responsibility for results may become a bottleneck. That is an interpretation of the cited examples, not a measured forecast of wages or employment.
Limits of the Available Evidence
The cited figures do not show how representative the results are across all organisations or fields. The source gives no common method for comparing the mathematics, software and contract examples, and the commercial interests of some software-data providers warrant caution. The 55% contract-task score is not accompanied here by details of the evaluation criteria, sample or comparison model.
It is also unclear how much of the review backlog is caused by AI-generated volume rather than staffing, workflow design or other changes. The report’s examples indicate pressure and possible failure points, but do not establish a single cause or show the overall rate of errors that reach users. The long-term effect on training junior professionals is likewise a concern raised by the report, not a measured outcome in the supplied evidence.
How Review Practices Respond
The next test is whether organisations can match faster production with adequate review. That may involve prioritising high-risk work, documenting which outputs have been checked, and maintaining clear responsibility for final decisions. The source material does not identify a new policy, deadline or follow-up study, so no specific industry response is confirmed.
For now, the reported evidence leaves a practical question for employers and professional fields: can they expand review capacity while still giving junior workers enough direct experience to develop judgment? Further data on review time, error rates and outcomes across comparable workflows would help show whether the gap is narrowing or widening.
Key Questions
What is the main development described in the report?
The report says AI is producing work faster than people can verify it, citing examples in mathematics, software and contracting.
Did OpenAI formally verify all 722 mathematics manuscripts?
No. The source says some results were formally checked in Lean and quotes OpenAI cautioning that unformalized results could have issues.
What did the software data show?
The cited studies reported more pull requests and longer review waits during periods of AI adoption, along with substantial shares of AI-agent changes receiving no human review in one study. The measures differ, and some data comes from companies that sell review-related tools.
Does the report prove AI caused review backlogs?
No. The figures show patterns reported by the cited sources, but the material does not establish that AI alone caused longer waits, missed reviews or lower acceptance rates.
Why could AI adoption affect future reviewer supply?
The report argues that professionals often learn judgment by doing the work they later review. If AI reduces those early-career opportunities, employers may need to find other ways to build that experience.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
