In plain words: A multilingual test suite checks chatbots for made-up answers, social bias, and harmful output, then pinpoints exactly where each one fails instead of just scoring them. Testing 17 leading models found repeated weaknesses: agreeing too readily, answers shifting with wording, and repeating stereotypes.
Abstract · Phare: A Safety Probe for Large Language Models
Ensuring the safety of large language models (LLMs) is critical for responsible deployment, yet existing evaluations often prioritize performance over identifying failure modes. We introduce Phare, a multilingual diagnostic framework to probe and evaluate LLM behavior across three critical dimensions: hallucination and reliability, social biases, and harmful content generation. Our evaluation of 17 state-of-the-art LLMs reveals patterns of systematic vulnerabilities across all safety dimensions, including sycophancy, prompt sensitivity, and stereotype reproduction. By highlighting these specific failure modes rather than simply ranking models, Phare provides researchers and practitioners with actionable insights to build more robust, aligned, and trustworthy language systems.
Pierre Le Jeune, Benoît Malézieux, Weixuan Xiao, Matteo Dora
arXiv:2505.11365 · cs.CY, cs.AI, cs.CL, cs.CR · submitted May 16, 2025 · updated May 26, 2025
abstract · pdf · html