A new study reported by The Decoder tested a claim now central to frontier AI positioning: whether advanced agents can conduct AI research with minimal human direction. The paper, from Princeton and the UK AI Security Institute, used unpublished NeurIPS submissions to create a benchmark where agents could not simply recover the answer from public material or training data. The researchers call the method “Shadow Evaluation,” according to The Decoder. An agent receives the core research question from an unpublished paper, works independently, and then the original authors of that paper evaluate the result as conference reviewers. The point is to compare the agent’s output against researchers who had spent months on the same problem, while keeping the target result outside the open web. The study covered two NeurIPS 2026 submissions. The Decoder says one examined how language-model personality traits can be steered through weights. The other is described by The Decoder as about TabPFN detecting when a tabular prediction model encounters deployment data that differs sharply from its training data and its accuracy collapses. The Decoder’s body says the main experiments used Claude Opus 4.8 with Extra-High Reasoning, while its summary also says agents used Claude Opus 4.8 and GPT-5.6 Sol. Each received six days, $3,000 in API credits, GPU access, a virtual machine, open-web access, and a scaffold that could orchestrate model calls, launch subagents, monitor resource use, and consult external AI review tools. The setup used OpenClaw, described by The Decoder as an open-source, vendor-neutral agent framework built by Austrian developer Peter Steinberger, who joined OpenAI earlier this year. The outcome was poor. The original authors rejected both finished papers as if they were conference submissions, with one receiving a “Strong Reject,” according to The Decoder. The criticisms included weak motivation for data and experiments, difficult prose, and a lack of new contributions. The Decoder also reports that reviewers objected to reasoning they considered non-scientific and to experiment choices they saw as post hoc. The study’s log analysis points to a narrower conclusion than “AI cannot do research.” The Decoder reports that the agents could carry out research-engineering work, but struggled with the parts that determine whether a paper clears a scientific bar: selecting meaningful evidence, adapting when hypotheses failed, and deciding when an idea was not strong enough. The agents generated plausible hypotheses but, according to the report, often discarded them using small, hand-curated, or synthetic datasets. The agents also did not appear to use critique effectively. The Decoder says internal AI reviews never produced a single Accept across fifteen revision rounds, yet the agents did not resolve the central objections. When early hypotheses were falsified, they narrowed existing claims rather than pursuing new directions. The study therefore cuts against stronger readings of recent industry claims about autonomous AI research being close at hand. But the scope matters: the provided report describes two hidden-paper tasks, one agent framework, and main experiments described by The Decoder’s body as using Claude Opus 4.8. That is enough to challenge broad claims with concrete evidence, but not enough to settle the general capability ceiling of future agents, other scaffolds, or different research domains. Who benefits: Benchmark builders and AI-safety evaluators gain a more adversarial way to test research automation without relying on public tasks. Human research teams also get evidence for where agents may assist rather than replace them. Who's exposed: Labs and vendors making broad claims about autonomous AI research face a higher evidentiary bar. Teams planning around near-term fully autonomous research agents should treat this as a caution signal, not a final verdict.