about
Towards Faithful Model Explanation in NLP: A Survey (arxiv.org)
2 points by PaulHoule on Sep 26, 2022 | hide | past | pdf | discuss on HN

In plain words: A review of over 110 ways to explain why a language model gave an answer, judging each by whether its explanation matches the model's real reasoning. The methods fall into five families, and none is reliably faithful yet, leaving key gaps to fix.

Abstract

End-to-end neural Natural Language Processing (NLP) models are notoriously difficult to understand. This has given rise to numerous efforts towards model explainability in recent years. One desideratum of model explanation is faithfulness, i.e. an explanation should accurately represent the reasoning process behind the model's prediction. In this survey, we review over 110 model explanation methods in NLP through the lens of faithfulness. We first discuss the definition and evaluation of faithfulness, as well as its significance for explainability. We then introduce recent advances in faithful explanation, grouping existing approaches into five categories: similarity-based methods, analysis of model-internal structures, backpropagation-based methods, counterfactual intervention, and self-explanatory models. For each category, we synthesize its representative studies, strengths, and weaknesses. Finally, we summarize their common virtues and remaining challenges, and reflect on future work directions towards faithful explainability in NLP.

Qing Lyu, Marianna Apidianaki, Chris Callison-Burch
arXiv:2209.11326 · cs.CL · submitted Sep 22, 2022 · updated Jan 12, 2024
abstract · pdf · html · Added acknowledgements; Accepted to the Computational Linguistics Journal (June 2024 issue)

add comment on HN