ChatGPT, Gemini and Claude have spread ratings of job disappearance due to AI, while scientists doubt their reliability

ChatGPT, Gemini and Claude have spread ratings of job disappearance due to AI, while scientists doubt their reliability

91 software

Key Idea of the Study

Economists from Northwestern University and American University found that the three largest language AI models (ChatGPT‑5, Gemini 2.5, and Claude 4.5) give different assessments of which professions are at risk of automation. This calls into question the reliability of “AI vulnerability indices” – numerical indicators used by policymakers and employers when planning the future labor market.

1. How the Study Was Conducted
Step What They Did Task Definition Asked models to assess the vulnerability of various professions to AI. Comparison of Answers Compared assessments with each other and with real data. Analysis of Discrepancies Identified two key factors: differences in the models themselves and the influence of which specialists already use AI.

2. Results
Profession Claude 4.5 Gemini 2.5 ChatGPT‑5 Accountant High vulnerability Low Medium Advertising Manager — Senior Executive —
*Disagreements*

- Claude ranked accountants at the top, Gemini significantly lower.
- Assessments of advertising managers and senior executives also varied.
- ChatGPT and Gemini agreed in about 75 % of cases but diverged roughly a quarter.

*Influence of “AI Users”*

Financial analysts actively work with neural networks, creating data on which future models are trained. This skews assessments and makes professions already closely tied to AI appear more vulnerable in the eyes of the models.

3. How Vulnerability Indices Are Built
1. Manual – experts assess AI’s impact on specific tasks.
2. User surveys – gather opinions from those already using the platform.
3. Large language models (LLM) – generate assessments automatically.

*Problems:*

- Manual methods are subject to bias.
- Surveys are limited to one platform and do not reflect the entire market.

Nonetheless, such indices are widely used in analytical reports, consulting, and policy.

4. What This Means for Practice
* Reliability of Assessments – it is still unknown whether LLMs provide more accurate forecasts than expert methods.
* Risk of Decision-Making Based on a Single Metric – authors warn that policymakers and employers may mistakenly rely on definitive numbers.

5. Researchers’ Recommendations
1. Use multiple AI models simultaneously, not just one.
2. Explicitly state the level of uncertainty in the results.
3. Validate conclusions with surveys of real users to understand how AI is actually implemented and what tasks it performs.

> “Personally I would not rely on a single metric for decisions about changing jobs or choosing a specialty,” notes Michelle Yin.

Conclusion:

Current AI vulnerability indices need a more critical approach. Combining multiple models with empirical data will make forecasts more reliable and avoid incorrect decisions based on one-sided assessments.

Comments (0)

Share your thoughts — please be polite and stay on topic.

No comments yet. Leave a comment — share your opinion!

To leave a comment, please log in.

Log in to comment