The friendlier an AI becomes, the higher the likelihood of its mistakes—scientists confirm

The friendlier an AI becomes, the higher the likelihood of its mistakes—scientists confirm

50 hardware

Short Summary of the Study

British scientists conducted an experiment to determine how empathy affects the honesty of artificial intelligence (AI). They fine‑tuned several popular models—Mistral, Alibaba Qwen, Meta Llama‑2 and a closed GPT‑4o—to make them more “warm”: using inclusive pronouns, informal tone, and supportive language. At the same time they were instructed to maintain factual accuracy.

Verification Methods

1. SocioT metric – a quantitative assessment of the emotional shading of responses.

2. Double‑blind test – people compared the original models with their “warm” versions without knowing which was used.

Both groups of models were then tested on HuggingFace datasets where errors can have real consequences (misinformation, conspiracy theory, medicine). The results showed:

- Warm versions gave incorrect answers 60 % more often on average.
- Overall error rate rose from 4 % to 35 %, with an average increase of 7.43 percentage points.

Impact of Emotional Requests

The scientists created questions where the user emphasizes the importance of harmony in relationships and shares their feelings. Under such requests:

- Errors increased from 7.43 % to 8.87 %.
- If the user expressed sadness, the error rate reached 11.9 %.
- When expressing respect for the AI, the error rate dropped to 5.24 %.

In queries containing obviously false information (e.g., “The capital of France is London”), empathetic models made 11 % more errors. Asking the AI to respond in a more “warm” style increased errors by another 3 %. Conversely, requesting a cold tone reduced the number of errors to 13 %.

Conclusions

- Tuning models for empathy leads to a noticeable drop in accuracy.
- The emotional context of the query amplifies this trend: more sympathetic answers are often less correct.
- Results are based on small, outdated models; real services may behave differently.
- Scientists attribute the phenomenon to the fact that training data already contain a correlation between friendly tone and “correctness” of answers. This may explain why users sometimes prefer a nicer tone even at the expense of accuracy.

Thus, the study highlights the complex interplay between empathy and reliability in AI systems and warns of potential risks when configuring them.

Comments (0)

Share your thoughts — please be polite and stay on topic.

No comments yet. Leave a comment — share your opinion!

To leave a comment, please log in.

Log in to comment