Microsoft researchers claim that AI models are still not capable of solving complex problems.

Microsoft researchers claim that AI models are still not capable of solving complex problems.

64 hardware

Microsoft researchers identified limitations of modern large language models (LLMs) in performing complex multi-step tasks

In a series of experiments, Microsoft Research found that even the most advanced AI models make serious mistakes when tasked with carrying out long and multi-stage operations. Among the systems tested—Gemini 3.1 Pro, Claude 4.6 Opus, and GPT 5.4—each lost on average about 25 % of document content after autonomous processing.

What was actually tested
* Benchmark DELEGATE‑52

Created by Philip Laban, Tobias Schnabel, and Jennifer Neville. It models workflows in 52 professional domains: from programming and musical notation to crystallography.

* Evaluation criterion

Models were evaluated on how well they preserve document integrity after 20 processing cycles. The “readiness” threshold was set at 98 %—if the result fell below that, the task was considered a failure.

Key findings
Domain Best results Programming Strongest performance Natural language Least reliable models
* Quality drop

In more than 80 % of document combinations, quality dropped to 80 % or lower. The best model (Gemini 3.1 Pro) met readiness criteria in only 11 out of 52 domains.

* Jump‑like errors

Losses did not occur gradually but “in jumps”: within a single interaction cycle the model could lose between 10 and 30 % accuracy. More advanced models (Gemini 3.1 Pro, Claude 4.6, GPT 5.4) tried to avoid small mistakes by deferring their handling to later stages and reducing the number of interactions.

* Agentic control

When using tools in agentic mode, results did not improve: by the end of the cycle quality fell by about 6 %.

What this means for users
Scientists emphasize that users still need to carefully monitor AI systems when delegating authority. Current models can operate autonomously only in narrow areas.

Nevertheless, the benchmark authors note significant progress: the OpenAI family improved performance metrics from 14.7 % to 71.5 % over 16 months. This confirms that LLMs continue to evolve but remain limited in complex multi‑step tasks.

Comments (0)

Share your thoughts — please be polite and stay on topic.

No comments yet. Leave a comment — share your opinion!

To leave a comment, please log in.

Log in to comment