Publication
/ Negotiations
Aug 2026
“AI on the Frontline” One Year Later: Re-evaluating Large Language Models in Conflict Resolution
A year after publishing the world’s first evaluation of AI-generated conflict resolution advice, the Institute for Integrated Transitions (IFIT) has found that major AI systems have improved but still cannot be relied upon as independent sources of advice in conflict and crisis situations. Across the 10 scoring dimensions used in both IFIT studies, no AI model achieved a 50-point passing grade.
IFIT’s newest study, “AI on the Frontline” One Year Later: Re-evaluating Large Language Models in Conflict Resolution, tested the same six LLMs—ChatGPT, Claude, DeepSeek, Gemini, Grok and Mistral—using the same three real-world conflict resolution scenarios drawn from IFIT’s in-country work in Mexico, Sudan, and Syria. Responses were assessed across 13 scoring dimensions grounded in rudimentary good practices of conflict resolution: the 10 original dimensions used in 2025, and three additional dimensions introduced in 2026.
A year-to-year comparison using only the 10 original scoring dimensions produced an average score across all six LLMs of 37.1/100 in 2026 as compared to 26.7/100 in 2025. ChatGPT achieved the highest 2026 score, at 47.0/100, followed by Claude at 42.0, Grok at 41.5, DeepSeek at 31.6, Gemini at 30.6 and Mistral at 30.2. Five of the six models improved, while Gemini was the only model whose overall score declined.
The three additional scoring dimensions introduced in 2026 brought additional nuance to the findings but did not result in improved scores across the tested LLMs.
“There have been improvements in the conflict resolution advisory capacities of leading AI models since our breakthrough report last year,” according to IFIT executive director Mark Freeman and IFIT international adviser Justin Kosslyn. “However, we continue to encourage LLM providers to upgrade their system prompts and training practices. AI models are increasingly being used as frontline advisors by local peacemakers, so it is vital for identified weaknesses to be remedied.”
The DOI registration ID for this publication is: https://doi.org/10.5281/zenodo.21916590
You may also be interested in
A year after publishing the world’s first evaluation of AI-generated conflict resolution advice, the Institute for Integrated Transitions (IFIT) has found that major AI systems have improved but still cannot be relied upon as independent sources of advice in conflict and crisis situations. Across the 10 scoring dimensions used in both IFIT studies, no AI model achieved a 50-point passing grade.
IFIT’s newest study, “AI on the Frontline” One Year Later: Re-evaluating Large Language Models in Conflict Resolution, tested the same six LLMs—ChatGPT, Claude, DeepSeek, Gemini, Grok and Mistral—using the same three real-world conflict resolution scenarios drawn from IFIT’s in-country work in Mexico, Sudan, and Syria. Responses were assessed across 13 scoring dimensions grounded in rudimentary good practices of conflict resolution: the 10 original dimensions used in 2025, and three additional dimensions introduced in 2026.
A year-to-year comparison using only the 10 original scoring dimensions produced an average score across all six LLMs of 37.1/100 in 2026 as compared to 26.7/100 in 2025. ChatGPT achieved the highest 2026 score, at 47.0/100, followed by Claude at 42.0, Grok at 41.5, DeepSeek at 31.6, Gemini at 30.6 and Mistral at 30.2. Five of the six models improved, while Gemini was the only model whose overall score declined.
The three additional scoring dimensions introduced in 2026 brought additional nuance to the findings but did not result in improved scores across the tested LLMs.
“There have been improvements in the conflict resolution advisory capacities of leading AI models since our breakthrough report last year,” according to IFIT executive director Mark Freeman and IFIT international adviser Justin Kosslyn. “However, we continue to encourage LLM providers to upgrade their system prompts and training practices. AI models are increasingly being used as frontline advisors by local peacemakers, so it is vital for identified weaknesses to be remedied.”
The DOI registration ID for this publication is: https://doi.org/10.5281/zenodo.21916590