📊 Full opportunity report: Engineering Is Automated. Research Is the Residual. on ThorstenMeyerAI.com — validation score, market gap, and execution plan.
TL;DR
AI systems have achieved near-saturation in automating core engineering tasks, such as reproducing research and optimizing kernels. Research remains less automated but is rapidly advancing. This shift could accelerate AI development and reshape research practices.
Recent AI benchmark results confirm that AI systems have achieved near-complete automation of core engineering tasks, such as research reproduction and kernel optimization, shifting the residual challenge to AI research itself. This development has significant implications for the pace of AI innovation and the future of research practices.
Multiple independent benchmarks—CORE-Bench, MLE-Bench, and kernel design studies—show AI systems are approaching or have achieved saturation in automating core engineering skills. For example, CORE-Bench, which assesses the reproduction of research papers, reached 95.5% reliability in December 2025, with the benchmark’s author stating it is ’solved.‘ Similarly, MLE-Bench, evaluating performance on Kaggle competitions, hit 64.4% in February 2026, approaching mid-tier human performance. Meanwhile, advances in kernel design, including automated GPU kernel creation and optimization, are documented through numerous research papers demonstrating production-grade capabilities. These developments suggest that the bottleneck in AI R&D is shifting from engineering execution to research innovation, which remains less automated.
Engineering is automated.
Research is the residual.
Six skill benchmarks. Edison’s framing. The question Clark leaves open is whether research is just engineering at scale.
Jack Clark’s Import AI #455 catalogs six benchmarks measuring AI capability on AI R&D tasks and concludes „AI can today automate vast swatches, perhaps the entirety, of AI engineering.“ The residual question is research. The structural read on the residual: it may not be a permanent moat.
Six skills. One trajectory.
Clark catalogs six benchmarks measuring AI capability on AI R&D-relevant tasks. Each individual benchmark could be noise. Six benchmarks moving together is a curve. The pattern is the cascade observed across the broader Clark series — visible here in the specific R&D-skill domain.

AI Workflow Automation for Bloggers: Build a Simple Content System to Research, Write, Optimize, and Repurpose Posts Faster with AI and No-Code Tools (AI Toolkit for Bloggers 2026 Book 8)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Three data points. Mixed signal.
Clark provides three data points on the creative-spark question. Yes-evidence: Erdős-1051, centaur math discovery, sporadic Move-37-style moments. No-evidence: low yield, framing dependence, absence of acceleration. The mixed signal is the honest read.
The data supports two readings. Pessimistic: rare moments suggest creative insight is qualitatively distinct from engineering work. Optimistic: rare moments are an artifact of low-volume exploration; more shots on goal yields more discoveries. Both readings are consistent with Clark’s „vast swatches, perhaps the entirety“ claim. They differ on the residual.

GPU-Accelerated Computing with Python 3 and CUDA: From low-level kernels to real-world applications in scientific computing and machine learning
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Five dimensions Clark gestures at but leaves underdeveloped.
Clark’s section is rigorous on the empirical evidence. Five strategic dimensions matter for the institutional response that the Clark series synthesis argues is structurally inadequate.
![WavePad Audio Editing Software - Professional Audio and Music Editor for Anyone [Download]](https://m.media-amazon.com/images/I/B1fcLEGCs6S._SL500_.png)
WavePad Audio Editing Software – Professional Audio and Music Editor for Anyone [Download]
Full-featured professional audio and music editor that lets you record and edit music, voice and other audio recordings
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Two readings. Different equilibria.
The structural question Clark leaves open: is research a permanent moat that bounds automated AI R&D, or is it engineering at scale that dissolves with more shots on goal? Both readings are consistent with the current data. They differ by orders of magnitude in consequences.
Productivity multiplier years
Recursive loop operational
automated AI benchmarking tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Five audiences. Asymmetric cost of being wrong.
The institutional response should not bet on inspiration being a permanent moat. If the distinction holds, capacity built is still useful. If it closes, capacity is necessary. Asymmetric cost-of-being-wrong points toward building now.
IN INDUSTRY
IN ACADEMIA
POLICYMAKERS
INVESTORS
EVERYONE ELSE
Engineering is automated. The residual is the question. The institutional response should not bet on inspiration being a permanent moat.
Implications for AI Development Speed and Research Practices
The near-complete automation of engineering tasks means that the primary remaining challenge in AI progress is now research itself. This could significantly accelerate the development cycle, reduce costs, and alter the role of human researchers. It also raises questions about the future of research workflows, intellectual property, and the potential for AI to generate novel scientific insights independently.
Progress in AI Benchmarks and Engineering Capabilities
Over the past year, multiple benchmarks have tracked AI progress in core research skills. CORE-Bench, assessing research reproduction, improved from 21.5% in September 2024 to 95.5% in December 2025. MLE-Bench, measuring performance on Kaggle competitions, advanced from 16.9% in October 2024 to 64.4% in February 2026. Additionally, research into kernel design—such as automated GPU kernel creation and optimization—has produced multiple papers demonstrating production-ready tools. These parallel advancements indicate a saturation point in engineering capabilities, with progress now primarily driven by research innovation.
„Research reproduction, which has been an ongoing crisis in academic ML, is now a solved engineering problem.“
— Thorsten Meyer
Remaining Uncertainties About Research Automation
While engineering tasks are nearing full automation, it remains unclear how much of AI research—such as hypothesis generation, experimental design, and creative problem-solving—can be automated. The structural question about whether research is itself a form of engineering at scale is still open, and the pace of progress in automating research remains uncertain.
Next Steps in AI Automation and Research Development
Expect continued rapid advancements in AI capabilities for research tasks, with benchmarks evolving to measure more complex research activities. Researchers and institutions should prepare for increased automation in research workflows, potential shifts in intellectual property regimes, and new opportunities for AI-driven scientific discovery. Monitoring benchmark progress and understanding limitations will be critical in the coming months.
Key Questions
What does automation of engineering mean for AI development?
It indicates that many routine and complex engineering tasks can now be handled by AI, potentially accelerating development timelines and reducing costs.
Can AI fully automate scientific research now?
Not yet. While engineering tasks are approaching full automation, research activities involving hypothesis creation and creative insight remain less automated and are the current residual challenge.
What are the risks of increased automation in research?
Potential risks include reliance on AI for critical scientific insights, challenges to intellectual property rights, and ethical concerns about autonomous research processes.
How might this shift impact human researchers?
Human researchers may shift focus from routine tasks to higher-level analysis, interpretation, and strategic research planning, potentially changing the research workforce dynamics.
What should institutions do to prepare for these changes?
Invest in understanding AI capabilities, update research workflows, and develop policies for AI-assisted research and intellectual property management.
Source: ThorstenMeyerAI.com