When AIs take control of a company: a fiasco beyond expectations

découvrez comment l'intégration des intelligences artificielles dans une entreprise a conduit à un échec spectaculaire, dépassant toutes les anticipations.

The integration of artificial intelligence within companies raises numerous questions about its capabilities and limitations. A recent project led by researchers at Carnegie Mellon University revealed that when AI agents were tasked with managing a company, the results could be disappointing. The agents, representing various technologies like Claude from Anthropic and Gemini from Google, failed to complete most of the tasks assigned to them. This article examines this fiasco and what it means for the future of AI in the professional world.

The objectives of the experiment

The project carried out by these researchers aimed to simulate a company by assigning specific roles to artificial intelligence agents, such as financial analyst or project manager. In parallel, the group used a separate platform to simulate human colleagues with whom these agents were supposed to interact. The goal was clear: to assess the autonomy and ability of these AIs to manage complex tasks equivalent to those of a human team.

The disappointing results of the AI agents

The results of this experiment were revealing. The AI agents failed to complete more than three-quarters of the tasks assigned to them. Despite relatively simple specifications, such as analyzing databases or organizing visits to new premises, the performance was alarming. For example, Claude 3.5 Sonnet, one of the best agents, managed to complete only 24% of the tasks, with an aggregated score of 34.4% considering the tasks partially completed.

The challenges encountered when using AI

The difficulties faced by these agents highlight the current limitations of artificial intelligence. The researchers noted that often these systems did not understand implicit instructions. For example, when tasked with saving results in a “.docx” format file, they did not comprehend that it referred to a Microsoft Word format. Furthermore, navigating the web posed a significant challenge, especially when it came to bypassing pop-ups or redirects.

The inadequate social skills of the agents

In addition to technical difficulties, the lack of social skills of the AIs was also identified as a major issue. When an agent found itself struggling, it often attempted to circumvent complex steps rather than proactively solve problems, leading to errors in task execution.

The question of operating costs

Another interesting aspect of this experiment was the analysis of the operating costs of the different agents. Claude 3.5 Sonnet, although more efficient, cost $6.34 per task, while Gemini 2.0 Flash represented a more economical option at $0.79. However, these figures raise another question: why invest in an expensive AI if its performance remains largely inadequate?

The implications for the future of businesses

The fiasco observed in this experiment raises major questions about the integration of artificial intelligence in business management. While some companies, like Meta from Mark Zuckerberg, aspire to see AI play a key role in directing operations, the reality of this simulation serves as a reminder that AI is still far from being able to operate autonomously and effectively. These results demonstrate that it is crucial to remain vigilant regarding the exploitation of these technologies to avoid situations where excessive trust in AI capabilities could lead to significant failures.

Scroll to Top