ML News
new
|
past
|
best
|
rss
|
submit
about
Stories from November 11, 2025 (UTC)
Go back a
day
,
month
, or
year
. Go forward a
day
.
1.
Measuring What Matters: Construct Validity in Large Language Model Benchmarks
(
arxiv.org
)
1 point
by
Cynddl
327 days ago
|
hide
|
past
|
pdf
|
discuss
2.
Too Good to Be Bad: On the Failure of LLMs to Role-Play Villains [pdf]
(
arxiv.org
)
1 point
by
SerCe
327 days ago
|
hide
|
past
|
pdf
|
discuss
About
|
RSS
|
RSS (all)
|
HN arXiv