[{"data":1,"prerenderedAt":134},["ShallowReactive",2],{"notes-why-crux-2-differs-from-ai-scientist":3},{"id":4,"title":5,"artifacts":6,"authors":7,"body":12,"citation":6,"date":122,"description":123,"devOnly":124,"extension":125,"image":6,"lede":6,"meta":126,"navigation":127,"path":128,"pdf":6,"rawbody":129,"seo":130,"slug":131,"stem":132,"__hash__":133},"notes\u002Fresearch-notes\u002Fwhy-crux-2-differs-from-ai-scientist.md","Why do our findings differ from Sakana’s AI Scientist?",null,[8,9,10,11],"Stephan Rabanser","Peter Kirgis","Sayash Kapoor","Arvind Narayanan",{"type":13,"value":14,"toc":116},"minimark",[15,27,36,45,50,59,62,65,68,72,75,104,107],[16,17,18,19,26],"p",{},"In the ",[20,21,25],"a",{"href":22,"rel":23},"https:\u002F\u002Fcruxevals.com\u002Fcrux\u002Fcan-ai-agents-conduct-research\u002F",[24],"nofollow","latest iteration"," of CRUX, we investigate whether today's frontier agents can conduct open-ended AI research autonomously. We identified recurring failures of research judgment, backtracking, resource awareness, and instruction following. The full CRUX 2 write-up describes the setup, findings, and limitations in detail.",[16,28,29,30,35],{},"CRUX 2 has drawn reactions from across the spectrum of views on recursive self-improvement: from people who strongly doubt it is achievable anytime soon to people who believe it may be just around the corner, with many positions in between. We were grateful that our paper was ",[20,31,34],{"href":32,"rel":33},"https:\u002F\u002Fjack-clark.net\u002F2026\u002F08\u002F03\u002Fimport-ai-467-self-sustaining-ai-viruses-pacing-ai-progress-confusion-about-ai-and-creativity\u002F",[24],"helpful"," to people thinking carefully about the impact of AI agents conducting open-ended AI research.",[16,37,38,39,44],{},"In this research note, we address one point of criticism. In a recent ",[20,40,43],{"href":41,"rel":42},"https:\u002F\u002Fwww.nature.com\u002Farticles\u002Fd41586-026-02494-5",[24],"Nature article covering CRUX 2",", Cong Lu, a co-creator of The AI Scientist from Sakana AI, argued that some of the problems we observed “would have been trivially fixable with a more constrained harness.” He suggested that a harness (which we call scaffold) closer to the design of The AI Scientist might produce a different and, presumably, stronger result.",[46,47,49],"h2",{"id":48},"why-we-dont-think-its-the-scaffold","Why we don’t think it’s the scaffold",[16,51,52,53,58],{},"The CRUX 2 scaffold already incorporated several components of ",[20,54,57],{"href":55,"rel":56},"https:\u002F\u002Farxiv.org\u002Fhtml\u002F2504.08066",[24],"the AI Scientist",". It supported literature review through Semantic Scholar, instructed the agent to explore multiple distinct approaches in parallel, and asked it to collect examples from the literature that could serve as templates for the structure of the final paper. The agent also maintained an append-only research log, a revisable planning document, and a daily list of lessons learned. While we did not reproduce the AI Scientist's specific tree-search implementation, our approach was still heavily inspired by it. An orchestrator agent managed the research lifecycle and delegated phases of work to subagents. That lifecycle covered literature review, open-ended exploration, hypothesis generation, experimentation, and drafting. The last three phases were connected in both directions so that negative reviews from subagents or external AI review tools could lead the agent to revise its hypotheses and experiments. We also used adversarial critique loops to reduce path dependence and under-exploration.",[16,60,61],{},"We believe several parts of the CRUX 2 scaffold also improved on the scaffold described in the AI Scientist. It better supported large-scale, compute-intensive experiments that ran for days (whereas the AI Scientist capped each experimental node at one hour with its full paper-generation process only lasted several hours), used external AI review tools and subagents, allowed the agent to surface broken tools or other environmental problems to a human, and specified budgets and targets for key resources that the agent could monitor in real time.",[16,63,64],{},"While we think that these adaptations are already meaningful, we agree that parts of the CRUX 2 harness can and should be improved. For example, we are working on re-calibrating reviewer feedback to focus more on high-level issues, improving the delegation to sub-agents, coming up with a more persistent task context, transitioning to more adaptive resource planning, and adding a final presentation pass to the produced paper to make the contributions and results more accessible. In our latest experiments on new papers, we also give the agents a summary of the preliminary CRUX 2 findings to help them avoid the failure modes we observed. Finally, we are also investigating whether increasing the API and compute budget will lead to improved results. We are actively working on these improvements and hope to share our findings soon.",[16,66,67],{},"At the same time, we are skeptical that trivial or minor changes to scaffolds or prompts would result in significant improvements in the produced artifacts. The reason is that the main failures were not simply a lack of reminders, review steps, or opportunities to revise. In fact, across 15 revision rounds, the agents received many of the same criticisms that were later raised by the human experts. But instead of rethinking the research plan, the agent usually responded with caveats or local fixes. While we are optimistic that a stricter scaffold can force another review, ablation, or restart, we are less convinced that it can, by itself, make the agent identify the decisive experiment, recognize that a direction has failed, or find a promising alternative. These choices require significant scientific judgment and an agent that would exercise this degree of creativity would be a substantive advance beyond a minor prompt or guardrail change.",[46,69,71],{"id":70},"our-setup-is-different-from-the-ai-scientists","Our setup is different from the AI scientist’s",[16,73,74],{},"In our view, the actual reason for the differences in our findings is that the experimental pipelines differ between the AI Scientist and CRUX 2. We consider four aspects as especially relevant in comparison to our workflow:",[76,77,78,86,92,98],"ol",{},[79,80,81,85],"li",{},[82,83,84],"strong",{},"Human filtering and selection."," The AI scientist authors manually filtered outputs at the idea, experiment, and manuscript stages before choosing three papers to submit. They argue that this filtering only reduced cost and did not change the selected papers' scientific content. However, they report no comparison without filtering. Therefore, the results provide little indication of how many irrelevant or poor-quality papers would have been produced alongside the successful papers. They also do not discuss the cost and reviewer effort that was required to identify and reject these false positives.",[79,87,88,91],{},[82,89,90],{},"Workshop bar and acceptance rate",". The three papers authored by the AI scientist were submitted to a workshop hosted at the ICLR ML conference. One of the three submissions met the workshop's acceptance threshold. That was a notable result, but the workshop accepted 70% of submissions (which is significantly higher compared to the 32% for ICLR's main conference that year). The authors' own assessment was that none of the three papers met the main-conference bar.",[79,93,94,97],{},[82,95,96],{},"Fit to the workshop theme."," The successful paper reported a negative result, making it particularly well suited to a workshop devoted to interesting negative findings. Interestingly, this particular fit of a negative results workshop aligns well with the behavioral patterns we observed from the agent in CRUX 2: the agents were all too keen to narrow the scope of the papers in response to negative AI reviews rather than go back to the original research question and try out a new approach.",[79,99,100,103],{},[82,101,102],{},"Review methodology and expertise",". The AI Scientist papers went through blind peer review, whereas in CRUX we ask the original paper authors who have spent months on the same research question to evaluate the work. While blind review is the standard practice of evaluating submitted scientific work at publishing venues in many fields, reviewing is often overstretched, noisy, and often poorly matched to exact expertise—especially at ML venues. In contrast, asking the original authors to review offers much deeper subject-matter knowledge, while introducing potential biases from non-blind evaluation.",[16,105,106],{},"In summary, CRUX 2 did not reproduce the AI Scientist's exact scaffold, but it incorporated many of the same elements and extended some of them to enable longer-running, more resource-intensive research. The AI Scientist shows that a highly structured pipeline can produce promising workshop-level work under modest human supervision. However, it does, in our opinion, not show that straightforward harness changes would overcome the judgment failures observed on our held-out research questions\u002Fpapers.",[16,108,109,110,115],{},"To make sure that we are setting up our agents to have the potential to succeed, we remain open to adversarial collaborations with researchers who disagree with our design choices or conclusions. If you believe that a different harness would change the results, we would be glad to test that together. We think our initial choices in CRUX 2 were reasonable, but also recognize that other designs may perform better. If an improved, domain-general harness succeeds while preserving the shadow-evaluation requirement of minimal task-specific guidance and interventions, we would update our view accordingly. We are continuing to tune the harness and expanding the number and variety of papers for the full CRUX 2 release. If you have a strong unpublished paper with an open-ended research question that a frontier agent could attempt, ",[20,111,114],{"href":112,"rel":113},"https:\u002F\u002Fdocs.google.com\u002Fforms\u002Fd\u002Fe\u002F1FAIpQLScbbO-tI6Igd7lqBKVGXRgoLkRWNiQiVz8y8qxIpnYpD5o2tA\u002Fviewform",[24],"submit it through our collaborator form",".",{"title":117,"searchDepth":118,"depth":118,"links":119},"",2,[120,121],{"id":48,"depth":118,"text":49},{"id":70,"depth":118,"text":71},"2026-08-19","In the latest iteration of CRUX, we investigate whether today's frontier agents can conduct open-ended AI research autonomously. We identified recurring failures of research judgment, backtracking, resource awareness, and instruction following. The full CRUX 2 write-up describes the setup, findings, and limitations in detail.",false,"md",{},true,"\u002Fresearch-notes\u002Fwhy-crux-2-differs-from-ai-scientist","---\ntitle: \"Why do our findings differ from Sakana’s AI Scientist?\"\nauthors:\n  - Stephan Rabanser\n  - Peter Kirgis\n  - Sayash Kapoor\n  - Arvind Narayanan\ndate: 2026-08-19\nslug: why-crux-2-differs-from-ai-scientist\n---\n\nIn the [latest iteration](https:\u002F\u002Fcruxevals.com\u002Fcrux\u002Fcan-ai-agents-conduct-research\u002F) of CRUX, we investigate whether today's frontier agents can conduct open-ended AI research autonomously. We identified recurring failures of research judgment, backtracking, resource awareness, and instruction following. The full CRUX 2 write-up describes the setup, findings, and limitations in detail.\n\nCRUX 2 has drawn reactions from across the spectrum of views on recursive self-improvement: from people who strongly doubt it is achievable anytime soon to people who believe it may be just around the corner, with many positions in between. We were grateful that our paper was [helpful](https:\u002F\u002Fjack-clark.net\u002F2026\u002F08\u002F03\u002Fimport-ai-467-self-sustaining-ai-viruses-pacing-ai-progress-confusion-about-ai-and-creativity\u002F) to people thinking carefully about the impact of AI agents conducting open-ended AI research.\n\nIn this research note, we address one point of criticism. In a recent [Nature article covering CRUX 2](https:\u002F\u002Fwww.nature.com\u002Farticles\u002Fd41586-026-02494-5), Cong Lu, a co-creator of The AI Scientist from Sakana AI, argued that some of the problems we observed “would have been trivially fixable with a more constrained harness.” He suggested that a harness (which we call scaffold) closer to the design of The AI Scientist might produce a different and, presumably, stronger result.\n\n## Why we don’t think it’s the scaffold\n\nThe CRUX 2 scaffold already incorporated several components of [the AI Scientist](https:\u002F\u002Farxiv.org\u002Fhtml\u002F2504.08066). It supported literature review through Semantic Scholar, instructed the agent to explore multiple distinct approaches in parallel, and asked it to collect examples from the literature that could serve as templates for the structure of the final paper. The agent also maintained an append-only research log, a revisable planning document, and a daily list of lessons learned. While we did not reproduce the AI Scientist's specific tree-search implementation, our approach was still heavily inspired by it. An orchestrator agent managed the research lifecycle and delegated phases of work to subagents. That lifecycle covered literature review, open-ended exploration, hypothesis generation, experimentation, and drafting. The last three phases were connected in both directions so that negative reviews from subagents or external AI review tools could lead the agent to revise its hypotheses and experiments. We also used adversarial critique loops to reduce path dependence and under-exploration.\n\nWe believe several parts of the CRUX 2 scaffold also improved on the scaffold described in the AI Scientist. It better supported large-scale, compute-intensive experiments that ran for days (whereas the AI Scientist capped each experimental node at one hour with its full paper-generation process only lasted several hours), used external AI review tools and subagents, allowed the agent to surface broken tools or other environmental problems to a human, and specified budgets and targets for key resources that the agent could monitor in real time.\n\nWhile we think that these adaptations are already meaningful, we agree that parts of the CRUX 2 harness can and should be improved. For example, we are working on re-calibrating reviewer feedback to focus more on high-level issues, improving the delegation to sub-agents, coming up with a more persistent task context, transitioning to more adaptive resource planning, and adding a final presentation pass to the produced paper to make the contributions and results more accessible. In our latest experiments on new papers, we also give the agents a summary of the preliminary CRUX 2 findings to help them avoid the failure modes we observed. Finally, we are also investigating whether increasing the API and compute budget will lead to improved results. We are actively working on these improvements and hope to share our findings soon.\n\nAt the same time, we are skeptical that trivial or minor changes to scaffolds or prompts would result in significant improvements in the produced artifacts. The reason is that the main failures were not simply a lack of reminders, review steps, or opportunities to revise. In fact, across 15 revision rounds, the agents received many of the same criticisms that were later raised by the human experts. But instead of rethinking the research plan, the agent usually responded with caveats or local fixes. While we are optimistic that a stricter scaffold can force another review, ablation, or restart, we are less convinced that it can, by itself, make the agent identify the decisive experiment, recognize that a direction has failed, or find a promising alternative. These choices require significant scientific judgment and an agent that would exercise this degree of creativity would be a substantive advance beyond a minor prompt or guardrail change.\n\n## Our setup is different from the AI scientist’s\n\nIn our view, the actual reason for the differences in our findings is that the experimental pipelines differ between the AI Scientist and CRUX 2. We consider four aspects as especially relevant in comparison to our workflow:\n\n1. **Human filtering and selection.** The AI scientist authors manually filtered outputs at the idea, experiment, and manuscript stages before choosing three papers to submit. They argue that this filtering only reduced cost and did not change the selected papers' scientific content. However, they report no comparison without filtering. Therefore, the results provide little indication of how many irrelevant or poor-quality papers would have been produced alongside the successful papers. They also do not discuss the cost and reviewer effort that was required to identify and reject these false positives.\n\n2. **Workshop bar and acceptance rate**. The three papers authored by the AI scientist were submitted to a workshop hosted at the ICLR ML conference. One of the three submissions met the workshop's acceptance threshold. That was a notable result, but the workshop accepted 70% of submissions (which is significantly higher compared to the 32% for ICLR's main conference that year). The authors' own assessment was that none of the three papers met the main-conference bar.\n\n3. **Fit to the workshop theme.** The successful paper reported a negative result, making it particularly well suited to a workshop devoted to interesting negative findings. Interestingly, this particular fit of a negative results workshop aligns well with the behavioral patterns we observed from the agent in CRUX 2: the agents were all too keen to narrow the scope of the papers in response to negative AI reviews rather than go back to the original research question and try out a new approach.\n\n4. **Review methodology and expertise**. The AI Scientist papers went through blind peer review, whereas in CRUX we ask the original paper authors who have spent months on the same research question to evaluate the work. While blind review is the standard practice of evaluating submitted scientific work at publishing venues in many fields, reviewing is often overstretched, noisy, and often poorly matched to exact expertise—especially at ML venues. In contrast, asking the original authors to review offers much deeper subject-matter knowledge, while introducing potential biases from non-blind evaluation.\n\nIn summary, CRUX 2 did not reproduce the AI Scientist's exact scaffold, but it incorporated many of the same elements and extended some of them to enable longer-running, more resource-intensive research. The AI Scientist shows that a highly structured pipeline can produce promising workshop-level work under modest human supervision. However, it does, in our opinion, not show that straightforward harness changes would overcome the judgment failures observed on our held-out research questions\u002Fpapers.\n\nTo make sure that we are setting up our agents to have the potential to succeed, we remain open to adversarial collaborations with researchers who disagree with our design choices or conclusions. If you believe that a different harness would change the results, we would be glad to test that together. We think our initial choices in CRUX 2 were reasonable, but also recognize that other designs may perform better. If an improved, domain-general harness succeeds while preserving the shadow-evaluation requirement of minimal task-specific guidance and interventions, we would update our view accordingly. We are continuing to tune the harness and expanding the number and variety of papers for the full CRUX 2 release. If you have a strong unpublished paper with an open-ended research question that a frontier agent could attempt, [submit it through our collaborator form](https:\u002F\u002Fdocs.google.com\u002Fforms\u002Fd\u002Fe\u002F1FAIpQLScbbO-tI6Igd7lqBKVGXRgoLkRWNiQiVz8y8qxIpnYpD5o2tA\u002Fviewform).\n",{"title":5,"description":123},"why-crux-2-differs-from-ai-scientist","research-notes\u002Fwhy-crux-2-differs-from-ai-scientist","O4iOMItzscEboljXLIZuJypl3rCKB5oRWUdkxZqI2DE",1787167692873]