🔍 Read the full analysis: Cheap AI Creation Is Putting Pressure On Reviewers on ThorstenMeyerAI.com
TL;DR
AI is making it faster and cheaper to produce mathematical manuscripts, software changes and contract work, while review remains dependent on limited human expertise. Figures cited by OpenAI, software analytics firms and a peer-reviewed study suggest review delays and gaps are growing, though some sources sell review tools and their data should be read with care.
AI is producing work faster and more cheaply in mathematics, software and contract workflows, while the people responsible for checking that work remain a limited resource. The mismatch is visible in OpenAI’s publication of 722 mathematical manuscripts and in software-industry data showing more code changes alongside longer or missing reviews; the full scale of the trend remains uncertain because the figures come from different sources and methods.
OpenAI said its model was given about 4,000 mathematical problems and produced 722 manuscripts grouped into 372 families. The source material reports an average of roughly three hours of compute per result. Some manuscripts have been formally checked using Lean, a proof-assistant system; OpenAI cautioned that unformalized results could have issues. Producing a manuscript is not the same as establishing that its claims are sound or important.
Software data points to a similar tension. Faros AI reported that teams merged 98% more pull requests during high-AI-adoption periods, while review time rose 91%. LinearB, in an analysis of 8.1 million pull requests across 4,800 organizations, reported that AI-generated changes waited 4.6 times longer for review to begin and were accepted 32.7% of the time, compared with 84.4% for human-written changes. These are findings from the named companies, not a single standardized industry-wide measurement.
A peer-reviewed 2026 study cited in the source material found that 61% of AI-agent pull requests received no human review before they were merged or closed. Faros separately reported a 31.3% rise in merges with zero review during high-adoption periods. The source also describes OpenAI’s partnership with contract-software company Ironclad: GPT-6 Astra met 55% of evaluation criteria on average across 11 tasks. That result indicates progress against the evaluation, but also leaves criteria unmet; the source does not specify how each shortfall would affect real contracts.
The referee shortage: AI made doing cheap and checking expensive
OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.
Some Lean-checked; OpenAI warns unformalized ones „could have issues.“ Verification abundance, adjudication scarcity.
Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.
Real progress. Someone still has to find the other 45% before the work can be used.
Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.
A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.
Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.
of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).
of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.
OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.
Budget review hours next to model spend.
Provers, types, tests, policy engines.
Experts only where consequences are high.
Who profits from generation pays for checking.
Keep some production human for learners.
The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.
Review Capacity Sets the Pace
These examples matter because generating work and validating it are different tasks. When production speeds up but review does not, organizations may face queues, accept work with less scrutiny or divert senior staff from other responsibilities. More output does not automatically translate into more usable output.
The consequences vary by field. In software, insufficient review can leave defects or security problems undetected. In mathematics, a proof assistant can check whether a proof follows from a stated theorem, but experts still need to judge whether the theorem is meaningful and whether the formalized statement matches the intended claim. In legal work, a missed contractual condition can have practical consequences, and a person or organization still needs to take responsibility for the final document.
The pressure may also affect how people gain expertise. Experienced reviewers typically develop judgment through years of doing the underlying work. If AI takes over too much drafting or coding from junior staff, workplaces could reduce the practice that helps create future reviewers. That is a concern raised by the source material, not an established outcome; how training changes will depend on how organizations deploy the tools.
As an affiliate, we earn on qualifying purchases.
Three Fields, Similar Bottleneck
The reported examples come from mathematics, software and contracting, but they do not measure the same thing. OpenAI’s manuscript count describes model output. Faros and LinearB track software work using their own datasets and definitions. The Ironclad evaluation concerns performance on 11 contract-related tasks. Together, they illustrate a shared issue, but they cannot be combined into one measure of AI quality or review capacity.
The software figures need particular care: Faros AI and LinearB sell products related to software development and review, so their findings should be read with that commercial connection in mind. Their data may still offer useful signals, but the source material gives no independent, common benchmark that resolves differences in sample, time period or methodology. The peer-reviewed study provides another data point, but its stated 61% finding applies to the study’s sample, not necessarily every organization.
The mathematical example also shows why formal verification does not settle every question. A tool such as Lean can establish that a proof meets formal rules for a specified statement. That check does not, by itself, determine whether the statement addresses the intended problem or whether the result advances the field. The source characterizes this distinction as “verification abundance, adjudication scarcity”: automated checks can expand, while expert interpretation remains limited.
As an affiliate, we earn on qualifying purchases.
How Wide Is the Review Gap?
The available figures do not establish one economy-wide rate of review failure. The source material does not provide the full methodologies, time windows or definitions behind every company metric, and the sources cover different populations. In particular, longer review waits do not alone show that a change is incorrect, while a lack of recorded human review does not reveal whether other checks took place.
It is also unclear how much of the pressure can be addressed by automated review tools. AI can assist with testing, proof checking and document comparison, but the supplied material does not quantify how often those systems catch consequential errors or how their performance compares with human review. Nor does the Ironclad evaluation explain the practical importance of the 45% of criteria Astra did not meet.
The reported concern about junior training is prospective. The source argues that reduced hands-on drafting and coding could weaken the future supply of experienced reviewers, but it gives no longitudinal data showing that this has already happened. The pace and effects will depend on how employers distribute work and training.
mathematical proof verification software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Watch Review Practices Change
The next useful evidence will be independent, comparable measurements showing how AI changes review time, error rates and outcomes across organizations. For software, that means distinguishing between changes that were reviewed, tested or merged without review, and tracking whether defects emerge later. For mathematics and contract work, evaluation should clarify both what automated checks verify and what still requires expert judgment.
Organizations adopting AI will also need to decide who is accountable for approving its output and how junior staff can build the expertise required for that role. The source material does not identify a common policy or next milestone across the fields discussed. For now, the reported figures point to a developing operational challenge: AI can increase the supply of drafts and results, but the capacity to judge and take responsibility for them may not grow at the same pace.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is the main development in this report?
Examples from mathematics, software and contract work suggest that AI is increasing the volume of generated output faster than human review capacity is expanding. The evidence comes from separate sources and does not establish a single overall rate.
Did OpenAI say all 722 mathematical manuscripts were verified?
No. The source says some results were formally checked in Lean, while OpenAI warned that some unformalized results could have issues. The manuscript count should not be read as a count of independently confirmed discoveries.
What did the software data report?
Faros AI reported more merged pull requests alongside longer review time during high-AI-adoption periods. LinearB reported longer waits for AI-generated changes to receive review and lower acceptance rates than for human-written changes. The figures use different datasets and should not be treated as directly comparable.
Can AI systems review AI-generated work?
Automated systems can check specific properties, such as whether code passes tests or a formal proof follows stated rules. Those checks do not necessarily confirm that the tests or theorem reflect the real requirement, or that a contract is appropriate for a particular situation.
What remains unknown?
The source material does not establish how representative the reported figures are across industries, how often automated review catches serious mistakes, or whether reduced junior-level work will lead to fewer experienced reviewers over time.
Source: ThorstenMeyerAI.com