about
Running cognitive evaluations on large language models: The do's and the don'ts (arxiv.org)
2 points by belter on Dec 5, 2023 | hide | past | pdf | discuss on HN

In plain words: Fourteen practical guidelines lay out how to design studies of language models' thinking abilities so results are solid, apply more widely, and are read correctly. The advice covers experiment design and result interpretation instead of reporting a new test or model.

Abstract · Toward best research practices in AI Psychology

Language models have become an essential part of the burgeoning field of AI Psychology. I discuss 14 methodological considerations that can help design more robust, generalizable studies evaluating the cognitive abilities of language-based AI systems, as well as to accurately interpret the results of these studies.

Anna A. Ivanova
arXiv:2312.01276 · cs.AI, cs.CL · submitted Dec 3, 2023 · updated Oct 29, 2024
abstract · pdf · html

add comment on HN