skip to content
$empowered.guru

AI & Machine Learning

Shadow Evaluations: AI Agents Aced the Engineering, Flunked the Science

Princeton's shadow evaluation gave frontier agents real unpublished NeurIPS questions, six days, and thousands in compute. The papers came back 2/6 and 1/6. Engineering works. Science doesn't.

September 16, 20265 min read
B

Brian Marvin

Published September 16, 2026

Shadow Evaluations: AI Agents Aced the Engineering, Flunked the Science

Shadow Evaluations: AI Agents Aced the Engineering, Flunked the Science

A Princeton-led team gave frontier agents the central research question from two unpublished NeurIPS 2026 papers, six days, and thousands of dollars of compute. The code worked. The reviewers rejected both papers anyway. Here is what that actually proves.

A claim crossed my feed this week that should have been bigger news than it was. AI agents, handed real unpublished machine learning research questions, produced two finished papers from scratch. The catch is what happened when the original authors graded the work.

The study is Can AI agents conduct open-ended AI research? Early evidence from two case studies, out in late July from a 24-person team led by Peter Kirgis and Sayash Kapoor at Princeton, with researchers from Stanford, the UK AI Security Institute, and Johns Hopkins. They call their method a shadow evaluation, and it is the most honest test I have seen of the claim everyone keeps making about agents doing research.

How a shadow evaluation works

The setup is what makes it interesting. The team took two high-quality, not yet public NeurIPS 2026 submissions and handed each agent the central research question, with the original paper kept hidden. The agent gets roughly six days, internet access, thousands of dollars of compute, and a virtual machine. No hints, no human in the loop. When it is done, the paper's own authors grade the output with the real conference rubric.

Why this beats the alternatives is worth sitting with. Standard ML evals test narrow, verifiable tasks, which is not what open-ended research is. Blind peer review of AI-generated papers is stretched thin, stochastic, and reviewers are overloaded. A shadow evaluation puts the person who actually did the research in the grader's chair, and the question cannot be retrieved from training data because the paper is not public yet. No memorization bonus available.

What happened

The agents finished all of the engineering without human help. Experiments ran, code executed, papers got written. Then the original authors graded the submissions. One paper earned a 2 out of 6. The other earned a 1 out of 6. Both were unambiguously rejected.

The reviewers' complaint was not that the code did not work. It was that the scientific logic was aimless and added zero novel value to the field. The machine could grind. It could not figure out what would matter.

Five ways the agents lost the thread

The authors catalogued five recurring failure modes, and a second model on a different scaffold reproduced all of them, so this is a pattern rather than one vendor having a bad week:

  • Poor judgment about what the bar for publishable research actually is
  • Uncreative responses when the research design was not working
  • Ineffective backtracking from dead ends
  • Poor awareness of the resources being burned
  • Instruction drift as the run went long

Read that list like a fractional CTO and it starts to look familiar. It is roughly the failure list of a junior engineer given a six-day budget and no supervision, except the juniors eventually learn taste.

Why this should change how you buy AI

Two sides of the AI argument should both be uncomfortable after this study.

For the vendors selling you autonomous research agents and self-improving systems, the honest read is that models cannot rewrite themselves. A deployed model is a frozen file of weights. The "recursive self-improvement" demos are multigenerational training loops with humans wiring each generation together. That is automated engineering assistance, which is real and useful, but it is not a machine bootstrapping its own intelligence. Nobody is close to that.

For the panic side, the same evidence cuts the other way. A system that cannot notice it has walked into a dead end is not about to organize a breakout. The systems are powerful pattern matchers executing paths humans already built. The realistic danger has always been people wiring them into production with guardrails lowered, which is an engineering and governance problem, not a sci-fi problem.

For founders, the actionable version is simple. The agents I deploy are enormously good at the engineering layer: code, tests, configs, scraping, pipelines, the parts where success is verifiable. That matches what I wrote in my determinism post: keep the tasks bounded, the outputs checkable, and the human holding the hypothesis. What the shadow evaluation killed is the idea that you can point an agent at an open problem, walk away, and come back to a breakthrough. That product does not exist, and this paper is the evidence.

What I would do with this

If you are running AI in a business, do three things. First, split your workload the way the study splits cleanly: automation that ships, iterates, and tests belongs to agents. Research questions and product bets belong to people. Second, when a vendor claims self-improving agents or autonomous research, ask which human is in the loop between generations, and move on if the answer involves a shrug. Third, budget for agent-assisted research instead: humans pick the questions, agents grind the engineering, humans judge whether the result argues for itself.

The paper is worth reading in full, including the released expert reviews and agent logs. The link is https://arxiv.org/abs/2607.27191.

About the Author

I'm Brian Marvin, an AI-native Fractional CTO with 30 years in technical leadership. At empowered.guru, I help startups build MVPs, shape roadmaps, and make AI-powered technology decisions that scale.

Filed under

AI ResearchShadow EvaluationsNeurIPSAI AgentsFractional CTO
$empowered.guru --book-session

Keep exploring

Turn the next insight into a shipped product.

Bring us the product, architecture, or delivery problem you are working through. We will help you find the clearest path forward.