about
Regulation of LLMs with Interpretability Will Likely Result in Performance Loss (arxiv.org)
1 point by PaulHoule on Jan 17, 2025 | hide | past | pdf | discuss on HN

In plain words: A language model was built to base its answers only on human-chosen, visible features, so regulators can see and check what drives its decisions. This cut classification accuracy by 7.34% compared with a normal free-to-use-anything model, yet people working with it finished tasks faster and judged their confidence more appropriately than working without AI help.

Abstract · Regulation of Language Models With Interpretability Will Likely Result In A Performance Trade-Off

Regulation is increasingly cited as the most important and pressing concern in machine learning. However, it is currently unknown how to implement this, and perhaps more importantly, how it would effect model performance alongside human collaboration if actually realized. In this paper, we attempt to answer these questions by building a regulatable large-language model (LLM), and then quantifying how the additional constraints involved affect (1) model performance, alongside (2) human collaboration. Our empirical results reveal that it is possible to force an LLM to use human-defined features in a transparent way, but a "regulation performance trade-off" previously not considered reveals itself in the form of a 7.34% classification performance drop. Surprisingly however, we show that despite this, such systems actually improve human task performance speed and appropriate confidence in a realistic deployment setting compared to no AI assistance, thus paving a way for fair, regulatable AI, which benefits users.

Eoin M. Kenny, Julie A. Shah
arXiv:2412.12169 · cs.LG, cs.AI, cs.CY · submitted Dec 12, 2024
abstract · pdf · html

add comment on HN