AI Agents3 mins read

AI Agents Take On More Model-Development Work, but Humans Still Make the Calls

A research team’s analysis of AI-assisted model development shows agents are doing more task-level work, while people continue to set goals, choose methods, and resolve key judgment calls.

What the Research Team Measured

Eight bar charts comparing Atria Dawn Preview with several AI models across multiple benchmarks.
Image credits:Atria Team

A team involving researchers from China’s Fudan University analyzed 769 task logs from its own AI model development work. The project centered on Atria Dawn Preview, an agentic language model built on a mixture-of-experts architecture with 744 billion parameters and designed for research and engineering tasks.

The study tracked how humans and AI agents contributed across tasks, including proposals, decisions, revisions, and problem-solving. Its core finding is specific: agents are doing more of the work, but that does not automatically mean they are becoming more autonomous.

AI Expanded What Researchers Could Attempt

Line chart showing the daily median of agent actions per human prompt rising from 11.0 to 28.5 over four weeks.
Image credits:Atria Team

AI was used in 96.5 percent of the reviewed tasks, and participants handed off more activity to agents over time. The median ratio of agent actions to human inputs rose from 11 to 28.5 over four weeks.

The authors caution that this increase should not be read as rising autonomy. In many cases, one human decision simply triggered a longer chain of agent steps. Of 455 completed AI-assisted tasks, 151 were rated infeasible without AI, meaning AI did not just speed up work — it made some work possible at the stated scope and quality.

The Dominant Pattern: AI Proposes, Humans Choose

Stacked bar chart showing who proposes and who decides on goals, methods, and acceptance criteria.
Image credits:Atria Team

For methods and parameters, the most common workflow was “AI proposes, human selects,” at 55.4 percent. Overall, humans made 85.5 percent of decisions about methods and parameters, while AI made 9.2 percent.

Humans also made final decisions on goals and scope in 93.4 percent of cases. Even in tasks participants said could not have been completed without AI, humans chose the goal 95.4 percent of the time. The takeaway: AI may broaden the menu of options, but humans still control the final direction.

Oversight Is Becoming a Judgment Problem

Horizontal bar chart showing how tasks progressed after a problem, including added context, human diagnosis, and AI self-recovery.
Image credits:Atria Team

When tasks ran into difficulty, human intervention moved work forward in 76 percent of recorded cases, while agents solved the problem on their own in 23 percent. Human help usually meant adding context, clarifying requirements, diagnosing issues, or switching methods — not manually taking over the work.

The study warns that oversight becomes harder when each human decision triggers a long chain of agent work that may be difficult to fully review. The risk is that humans become rubber-stamp reviewers rather than meaningful decision-makers. For teams using agents, the practical lesson is to define where human authority sits before convenience turns autonomous modes into default practice.

Discover More