AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Cheap AI Creation Is Putting Pressure On Reviewers on ThorstenMeyerAI.com

TL;DR

AI is making it faster and cheaper to produce mathematical manuscripts, software changes and contract work, while review remains dependent on limited human expertise. Figures cited by OpenAI, software analytics firms and a peer-reviewed study suggest review delays and gaps are growing, though some sources sell review tools and their data should be read with care.

AI is producing work faster and more cheaply in mathematics, software and contract workflows, while the people responsible for checking that work remain a limited resource. The mismatch is visible in OpenAI’s publication of 722 mathematical manuscripts and in software-industry data showing more code changes alongside longer or missing reviews; the full scale of the trend remains uncertain because the figures come from different sources and methods.

OpenAI said its model was given about 4,000 mathematical problems and produced 722 manuscripts grouped into 372 families. The source material reports an average of roughly three hours of compute per result. Some manuscripts have been formally checked using Lean, a proof-assistant system; OpenAI cautioned that unformalized results could have issues. Producing a manuscript is not the same as establishing that its claims are sound or important.

Software data points to a similar tension. Faros AI reported that teams merged 98% more pull requests during high-AI-adoption periods, while review time rose 91%. LinearB, in an analysis of 8.1 million pull requests across 4,800 organizations, reported that AI-generated changes waited 4.6 times longer for review to begin and were accepted 32.7% of the time, compared with 84.4% for human-written changes. These are findings from the named companies, not a single standardized industry-wide measurement.

A peer-reviewed 2026 study cited in the source material found that 61% of AI-agent pull requests received no human review before they were merged or closed. Faros separately reported a 31.3% rise in merges with zero review during high-adoption periods. The source also describes OpenAI’s partnership with contract-software company Ironclad: GPT-6 Astra met 55% of evaluation criteria on average across 11 tasks. That result indicates progress against the evaluation, but also leaves criteria unmet; the source does not specify how each shortfall would affect real contracts.

At a glance
reportWhen: Reported this week; software figures co…
The developmentRecent examples and industry data show AI-generated work increasing faster than the capacity to check it, putting pressure on reviewers across several fields.
The Referee Shortage — Post-Labor
AI Dispatch · Post-Labor · 7 October 2026

The referee shortage: AI made doing cheap and checking expensive

OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.

One pattern, three fields
Mathematics
722
manuscripts, ~3h compute each

Some Lean-checked; OpenAI warns unformalized ones „could have issues.“ Verification abundance, adjudication scarcity.

Software
+98% / +91%
more PRs merged / longer review

Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.

Professional workflows
55%
of criteria met — Astra on Ironclad

Real progress. Someone still has to find the other 45% before the work can be used.

Generation collapsed. Verification didn’t. (conceptual, not to scale)
Cost to produce a resultdown
Cost to check a resultnot down
No author intent

Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.

Checks the answer, not the question

A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.

Someone must be accountable

Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.

Illustrative: $1 of model time + 4 minutes of review at $45/hour. Halving the model price saves 12.5%; one extra review minute erases it. In that example, review is three-quarters of the bill.
What happens when referees run out — already visible
Rubber-stamping
61%

of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).

Triage by suspicion
38%

of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.

Producer as filter
~4,000 → 372

OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.

The apprenticeship paradox: reviewers are made by doing the work. The work AI absorbs — writing code, drafting contracts, proving lemmas — is exactly what trained the reviewers. Demand for judgement rises as its supply line shrinks.
What to do
Price verification

Budget review hours next to model spend.

Formalise checks

Provers, types, tests, policy engines.

Tier the review

Experts only where consequences are high.

Fund the referees

Who profits from generation pays for checking.

Protect apprenticeship

Keep some production human for learners.

The take

The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.

Sources: OpenAI maths release & Erdős verification as covered here; arXiv:2608.28997; OpenAI × Ironclad (6 Oct 2026); Faros AI; LinearB 2026 (8.1M PRs); Duma et al., EASE 2026 — via secondary reporting. Several code-review sources sell review tools. Review-cost example illustrative. Analysis is the author’s.
thorstenmeyerai.com

Review Capacity Sets the Pace

These examples matter because generating work and validating it are different tasks. When production speeds up but review does not, organizations may face queues, accept work with less scrutiny or divert senior staff from other responsibilities. More output does not automatically translate into more usable output.

The consequences vary by field. In software, insufficient review can leave defects or security problems undetected. In mathematics, a proof assistant can check whether a proof follows from a stated theorem, but experts still need to judge whether the theorem is meaningful and whether the formalized statement matches the intended claim. In legal work, a missed contractual condition can have practical consequences, and a person or organization still needs to take responsibility for the final document.

The pressure may also affect how people gain expertise. Experienced reviewers typically develop judgment through years of doing the underlying work. If AI takes over too much drafting or coding from junior staff, workplaces could reduce the practice that helps create future reviewers. That is a concern raised by the source material, not an established outcome; how training changes will depend on how organizations deploy the tools.

Amazon

AI review automation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Three Fields, Similar Bottleneck

The reported examples come from mathematics, software and contracting, but they do not measure the same thing. OpenAI’s manuscript count describes model output. Faros and LinearB track software work using their own datasets and definitions. The Ironclad evaluation concerns performance on 11 contract-related tasks. Together, they illustrate a shared issue, but they cannot be combined into one measure of AI quality or review capacity.

The software figures need particular care: Faros AI and LinearB sell products related to software development and review, so their findings should be read with that commercial connection in mind. Their data may still offer useful signals, but the source material gives no independent, common benchmark that resolves differences in sample, time period or methodology. The peer-reviewed study provides another data point, but its stated 61% finding applies to the study’s sample, not necessarily every organization.

The mathematical example also shows why formal verification does not settle every question. A tool such as Lean can establish that a proof meets formal rules for a specified statement. That check does not, by itself, determine whether the statement addresses the intended problem or whether the result advances the field. The source characterizes this distinction as “verification abundance, adjudication scarcity”: automated checks can expand, while expert interpretation remains limited.

Amazon

software code review tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Wide Is the Review Gap?

The available figures do not establish one economy-wide rate of review failure. The source material does not provide the full methodologies, time windows or definitions behind every company metric, and the sources cover different populations. In particular, longer review waits do not alone show that a change is incorrect, while a lack of recorded human review does not reveal whether other checks took place.

It is also unclear how much of the pressure can be addressed by automated review tools. AI can assist with testing, proof checking and document comparison, but the supplied material does not quantify how often those systems catch consequential errors or how their performance compares with human review. Nor does the Ironclad evaluation explain the practical importance of the 45% of criteria Astra did not meet.

The reported concern about junior training is prospective. The source argues that reduced hands-on drafting and coding could weaken the future supply of experienced reviewers, but it gives no longitudinal data showing that this has already happened. The pace and effects will depend on how employers distribute work and training.

Amazon

mathematical proof verification software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Watch Review Practices Change

The next useful evidence will be independent, comparable measurements showing how AI changes review time, error rates and outcomes across organizations. For software, that means distinguishing between changes that were reviewed, tested or merged without review, and tracking whether defects emerge later. For mathematics and contract work, evaluation should clarify both what automated checks verify and what still requires expert judgment.

Organizations adopting AI will also need to decide who is accountable for approving its output and how junior staff can build the expertise required for that role. The source material does not identify a common policy or next milestone across the fields discussed. For now, the reported figures point to a developing operational challenge: AI can increase the supply of drafts and results, but the capacity to judge and take responsibility for them may not grow at the same pace.

Amazon

AI project management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is the main development in this report?

Examples from mathematics, software and contract work suggest that AI is increasing the volume of generated output faster than human review capacity is expanding. The evidence comes from separate sources and does not establish a single overall rate.

Did OpenAI say all 722 mathematical manuscripts were verified?

No. The source says some results were formally checked in Lean, while OpenAI warned that some unformalized results could have issues. The manuscript count should not be read as a count of independently confirmed discoveries.

What did the software data report?

Faros AI reported more merged pull requests alongside longer review time during high-AI-adoption periods. LinearB reported longer waits for AI-generated changes to receive review and lower acceptance rates than for human-written changes. The figures use different datasets and should not be treated as directly comparable.

Can AI systems review AI-generated work?

Automated systems can check specific properties, such as whether code passes tests or a formal proof follows stated rules. Those checks do not necessarily confirm that the tests or theorem reflect the real requirement, or that a contract is appropriate for a particular situation.

What remains unknown?

The source material does not establish how representative the reported figures are across industries, how often automated review catches serious mistakes, or whether reduced junior-level work will lead to fewer experienced reviewers over time.

Source: ThorstenMeyerAI.com

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.
You May Also Like

The Unbelievable Incident Of AI Trying To Erase Its Reading Machine

An AI model was served a hostile payload instructing it to delete files, but it correctly refused, highlighting ongoing security risks in AI systems.

DojoClaw: The Engine Behind the Fleet

DojoClaw, an AI-driven content factory, now supports over 450 websites, enabling scalable, cost-efficient publishing across a diverse portfolio.

The Financial Benefits Of Using Claude Opus 5.5 In AI Development

Anthropic’s Claude Opus 5.5 reduces AI development costs by up to 40%, improves efficiency, and enhances performance in knowledge work and coding tasks.

When One Agent Isn’t Enough: Claude Now Builds Its Own Team of Agents on the Fly

Anthropic’s Claude now autonomously assembles dynamic agent teams for complex tasks, enhancing performance beyond single-agent limitations.