The Growing Gap Between AI Production And Human Review
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: The Growing Gap Between AI Production And Human Review on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get hardware and tech essentials delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Evidence from AI-assisted mathematics, software development and contract work points to a widening gap between the volume of machine-generated output and people’s capacity to check it. The reported figures come with limits, including tool-provider sources and unanswered questions about review quality, but they highlight a growing need for expert judgment and accountability.

Recent reports across mathematics, software and contract work point to a widening gap between AI-generated output and the human capacity to verify it. OpenAI said its model produced 722 mathematical manuscripts from about 4,000 problems, while reported software data shows review delays and limited human checks; the measures vary, but together they raise questions about how organizations can safely use growing volumes of AI work.

OpenAI’s mathematical work was grouped into 372 families, with an average result taking about three hours of compute to produce, according to the source material. Some results were checked formally using Lean, a proof-assistant system. OpenAI cautioned that some results without formal verification could have issues. The source says an earlier result from the same program, a counterexample to an Erdős conjecture, received careful verification from five leading mathematicians. These figures describe different parts of the work, not a direct measure of how long every manuscript took to review.

In software, the source cites several datasets. Faros AI reported that teams merged 98% more pull requests during high-AI-adoption periods, while review time rose 91%. LinearB said its analysis of 8.1 million pull requests across 4,800 organizations found AI-generated changes waited 4.6 times longer for review to begin and were accepted 32.7% of the time, compared with 84.4% for human-written changes. A peer-reviewed 2026 study found 61% of AI-agent pull requests received no human review before being merged or closed.

The source also describes OpenAI’s partnership with contract-software company Ironclad. In an evaluation across 11 tasks, GPT-6 Astra met 55% of the evaluation criteria on average, which the source characterizes as an improvement over the earlier model. That score does not establish how often the system’s work is safe or fit for use in every real-world contract; people still need to assess whether particular outputs meet legal and business requirements.

At a glance
reportWhen: Reported this week; software studies an…
The developmentRecent figures on AI-generated mathematical manuscripts, software pull requests and contract tasks illustrate how production is scaling faster than human review.
Crypto market snapshot
Fear & Greed Index
64/100 — Greed
Bitcoin BTC$82,914▼ 1.3%
Ethereum ETH$2,570▼ 1.5%
Tether USDT$0.9994▼ 0.0%
BNB BNB$770.49▲ 0.5%
XRP XRP$1.42▼ 3.0%
USDC USDC$0.9997▼ 0.0%
Solana SOL$115.94▼ 1.9%
TRON TRX$0.3357▲ 0.9%
Live data · CoinGecko · alternative.me (24h change)
The Referee Shortage — Post-Labor
AI Dispatch · Post-Labor · 7 October 2026

The referee shortage: AI made doing cheap and checking expensive

OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.

One pattern, three fields
Mathematics
722
manuscripts, ~3h compute each

Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.

Software
+98% / +91%
more PRs merged / longer review

Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.

Professional workflows
55%
of criteria met — Astra on Ironclad

Real progress. Someone still has to find the other 45% before the work can be used.

Generation collapsed. Verification didn’t. (conceptual, not to scale)
Cost to produce a resultdown
Cost to check a resultnot down
No author intent

Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.

Checks the answer, not the question

A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.

Someone must be accountable

Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.

Illustrative: $1 of model time + 4 minutes of review at $45/hour. Halving the model price saves 12.5%; one extra review minute erases it. In that example, review is three-quarters of the bill.
What happens when referees run out — already visible
Rubber-stamping
61%

of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).

Triage by suspicion
38%

of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.

Producer as filter
~4,000 → 372

OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.

The apprenticeship paradox: reviewers are made by doing the work. The work AI absorbs — writing code, drafting contracts, proving lemmas — is exactly what trained the reviewers. Demand for judgement rises as its supply line shrinks.
What to do
Price verification

Budget review hours next to model spend.

Formalise checks

Provers, types, tests, policy engines.

Tier the review

Experts only where consequences are high.

Fund the referees

Who profits from generation pays for checking.

Protect apprenticeship

Keep some production human for learners.

The take

The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.

Sources: OpenAI maths release & Erdős verification as covered here; arXiv:2608.28997; OpenAI × Ironclad (6 Oct 2026); Faros AI; LinearB 2026 (8.1M PRs); Duma et al., EASE 2026 — via secondary reporting. Several code-review sources sell review tools. Review-cost example illustrative. Analysis is the author’s.
thorstenmeyerai.com

Review Capacity Shapes AI Adoption

The practical issue is not simply how much AI can produce, but how much output organizations can responsibly use. If review time grows more slowly than generation, work may accumulate in queues, receive less scrutiny or be rejected because reviewers cannot assess it promptly. The cited studies suggest all three pressures may be present in software, though their methods and populations differ.

Verification also involves more than checking whether an output is internally consistent. A proof assistant can test whether a proof supports its stated theorem; it cannot by itself decide whether the theorem addresses the intended question. Software tests cover the cases they were designed to cover, not every requirement. In contracts and other professional work, an accountable person must judge context and consequences. That makes expert review a potential constraint on AI adoption, not a routine step that can automatically be removed.

There may also be a longer-term workforce effect. Early-career staff often gain judgment by doing the underlying work before reviewing it. If AI substitutes for too much drafting, coding or analysis, organizations may reduce the opportunities through which future reviewers learn. The source raises this as a risk, not as an established outcome; its scale depends on how employers structure training and responsibility.

Amazon

AI review automation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Three Fields, Similar Pressure

The examples concern different kinds of work and should not be treated as one controlled comparison. Mathematical manuscripts, software pull requests and contract evaluations have different standards of correctness, different consequences for errors and different review processes. Their shared feature is that AI can increase the supply of drafts or candidate results faster than expert scrutiny necessarily expands.

The software numbers also require caution. The source notes that Faros AI and LinearB sell code-review products, which may affect how their findings are framed. The cited peer-reviewed study provides a separate data point, but a finding from one study does not establish the rate across all organizations. The figures are evidence of a possible pattern, not proof that every team experiences the same backlog or review quality.

Formal tools can help with parts of verification, and AI may assist reviewers by prioritizing risks or checking routine requirements. But those tools do not settle every question of intent, relevance or accountability. The source’s central framing is that generation costs have fallen while trust still depends, in many settings, on people able to assess whether the work is fit for purpose.

Amazon

software code review software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Available Measures Miss

The cited figures do not provide a complete measure of review quality or downstream errors. The source does not specify enough detail to compare all study periods, team characteristics or definitions of review. A pull request waiting longer may reflect a backlog, a more careful review, or other workflow differences; acceptance rates alone do not show whether rejected code was faulty.

It is also unclear how representative the reported results are across industries, how much review work AI tools can reliably automate, and whether organizations are changing staffing or training in response. The contract evaluation’s 55% average does not explain which criteria were missed or how serious those misses were. No evidence in the supplied material establishes that a general shortage of qualified reviewers has already emerged across the economy.

Amazon

mathematical proof verification software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Organizations Respond to Backlogs

The next indicators will be whether teams add review capacity, change which work receives human scrutiny, or use technical checks to handle routine cases. For software, useful follow-up data would distinguish review start times from completion times and report defect rates, review depth and outcomes across comparable teams. For research, clearer reporting on which results were formally checked and how independent verification was performed would help readers judge reliability.

Employers and professional bodies will also need to clarify who is accountable when AI-assisted work is approved, and how junior staff can build the judgment required for that responsibility. The available source material does not identify specific policy changes or a shared timetable. For now, the measurable question is whether human review keeps pace with rising output without lowering the standards that make professional work trustworthy.

Amazon

contract review software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is the main development?

Reports across mathematics, software and contract work suggest AI output is increasing faster than the available capacity to evaluate it. The examples are separate measures, not a single unified study.

Were all 722 mathematical manuscripts formally verified?

No. The source says some results were checked in Lean and notes OpenAI’s warning that unformalized results could have issues. It does not provide a count of how many manuscripts received formal checks.

What did the software figures find?

The cited reports describe more pull requests alongside longer review waits, lower acceptance for AI-generated changes in one dataset, and a peer-reviewed study in which 61% of AI-agent pull requests received no human review before merging or closing. Their samples and methods differ.

Does this prove organizations lack enough reviewers?

No. The figures indicate possible pressure on review workflows, but the source material does not establish a broad, economy-wide shortage. More comparable data on review quality, errors and staffing would be needed.

What should readers watch for next?

Look for data on review delays, defect rates, formal verification and changes to training or accountability. Those measures can show whether organizations are increasing review capacity or changing how they approve AI-assisted work.

Source: ThorstenMeyerAI.com

Nothing in this article is financial or investment advice. Cryptocurrency and precious-metal investments carry significant risk — do your own research and consider a licensed advisor.
HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Future Of Open-Weight Market Competition Lies In Cheap AI

Alibaba’s release of a cheap, capable open-weight AI model signals a strategic move to dominate the open AI market through cost-effective distribution and adoption.

How to Reduce Heat and Noise in a High-Power AI Workstation

Learn effective strategies to lower heat and noise in high-power AI workstations, focusing on undervolting, cooling, and airflow management for improved performance.

Why Homelab Servers Matter for Node Operators

Discover the top homelab servers for node operators in 2026. Find the best options for performance, value, and beginner-friendly setups in this detailed guide.

The Difference Between Layer 1, Layer 2, and Layer 3 Chains

Promising to clarify blockchain complexities, this guide explores how Layer 1, 2, and 3 chains differ and why understanding them is essential.