AI agents can't yet do open-ended AI research, shadow evaluation finds
Original titleAI agents can't yet do open-ended AI research
AISummary
A shadow evaluation found that frontier AI agents, given six days and thousands of dollars in credits, produced two research papers that the original authors unambiguously rejected. The authors' log analysis cited poor judgment, underused budgets, weak responses to feedback, and failure to backtrack or follow instructions as main causes.
Source: AI Snake Oil · normaltech.aiPublished · added here