| 2911. |
Revealing Political Bias in LLMs Through Structured Multi-Agent Debate (arxiv.org) |
|
1 point by rntn on Jun 16, 2025 | hide | past | pdf | discuss
|
| 2912. |
Appraisal-Based Chain-of-Emotion Improves AI Persona Accuracy (arxiv.org) |
|
1 point by virtual_rf on Jun 16, 2025 | hide | past | pdf | discuss
|
| 2913. |
The Illusion of the Illusion of Thinking – A Comment on Shojaee et al. (2025) (arxiv.org) |
|
16 points by gfortaine on Jun 16, 2025 | hide | past | pdf | 14 comments
|
| 2914. |
LiveCodeBench Pro: How Olympiad Medalists Judge LLMs in Competitive Programming? (arxiv.org) |
|
2 points by EvgeniyZh on Jun 16, 2025 | hide | past | pdf | discuss
|
| 2915. |
Towards Understanding Sycophancy in Language Models (arxiv.org) |
|
9 points by fzliu on Jun 16, 2025 | hide | past | pdf | 2 comments
|
| 2916. |
Vision Transformers Don't Need Trained Registers (arxiv.org) |
|
4 points by avd4292 on Jun 16, 2025 | hide | past | pdf | discuss
|
| 2917. |
A Word Is Worth 4-Bit: Log Parsing with Binary Coded Decimal Recognition (arxiv.org) |
|
7 points by PaulHoule on Jun 15, 2025 | hide | past | pdf | discuss
|
| 2918. |
Large language models often know when they are being evaluated (arxiv.org) |
|
89 points by jonbaer on Jun 15, 2025 | hide | past | pdf | 130 comments
|
| 2919. |
Clinical knowledge in LLMs does not translate to human interactions (arxiv.org) |
|
102 points by insistent on Jun 14, 2025 | hide | past | pdf | 39 comments
|
| 2920. |
Exploring the Best Input Representation for Electrocardiogram-Language Models (arxiv.org) |
|
2 points by PaulHoule on Jun 14, 2025 | hide | past | pdf | discuss
|
| 2921. |
Eliciting Fine-Tuned Transformer Capabilities via Inference-Time Techniques (arxiv.org) |
|
1 point by codelion on Jun 14, 2025 | hide | past | pdf | discuss
|
| 2922. |
Relic: Evaluating Compositional Instruction Following via Language Recognition (arxiv.org) |
|
1 point by demirbey05 on Jun 14, 2025 | hide | past | pdf | discuss
|
| 2923. |
Resa: Transparent Reasoning Models via SAEs (arxiv.org) |
|
1 point by Bogdanp on Jun 14, 2025 | hide | past | pdf | discuss
|
| 2924. |
Unsupervised Elicitation of Language Models (arxiv.org) |
|
135 points by kordlessagain on Jun 14, 2025 | hide | past | pdf | 24 comments
|
| 2925. |
CRMArena-Pro: LLM Agents Assessed Across Diverse Business Scenarios (arxiv.org) |
|
1 point by felineflock on Jun 14, 2025 | hide | past | pdf | discuss
|
| 2926. |
Memoir: Lifelong Model Editing with Minimal Overwrite Informed Retention for LLM (arxiv.org) |
|
1 point by dataminer on Jun 14, 2025 | hide | past | pdf | discuss
|
| 2927. |
Comment on the Illusion of Thinking (arxiv.org) |
|
4 points by esafak on Jun 14, 2025 | hide | past | pdf | 1 comment
|
| 2928. |
Rethinking Losses for Diffusion Bridge Samplers (arxiv.org) |
|
10 points by badmonster on Jun 14, 2025 | hide | past | pdf | 1 comment
|
| 2929. |
Unsupervised Elicitation of Language Models (arxiv.org) |
|
7 points by xianshou on Jun 13, 2025 | hide | past | pdf | discuss
|
| 2930. |
Self-Adapting Language Models (arxiv.org) |
|
246 points by archon1410 on Jun 13, 2025 | hide | past | pdf | 73 comments
|
| 2931. |
How Do Large Language Monkeys Get Their Power (Laws)? (arxiv.org) |
|
2 points by RSchaeffer on Jun 13, 2025 | hide | past | pdf | 1 comment
|
| 2932. |
Holistic Assessment of LLM Agents Across Diverse Scenarios and Interactions (arxiv.org) |
|
2 points by prisenco on Jun 13, 2025 | hide | past | pdf | discuss
|
| 2933. |
TabM: Advancing Tabular Deep Learning with Parameter-Efficient Ensembling (arxiv.org) |
|
3 points by felineflock on Jun 13, 2025 | hide | past | pdf | discuss
|
| 2934. |
Jailbreak attacks against DNA language models with pathogenicity guidance (arxiv.org) |
|
2 points by PaulHoule on Jun 12, 2025 | hide | past | pdf | 1 comment
|
| 2935. |
TabM: Advancing Tabular Deep Learning with Parameter-Efficient Ensembling (arxiv.org) |
|
2 points by m0nhawk on Jun 12, 2025 | hide | past | pdf | discuss
|
| 2936. |
Learning Semantically Faithful EEG-to-Text Generation (arxiv.org) |
|
1 point by PaulHoule on Jun 12, 2025 | hide | past | pdf | discuss
|
| 2937. |
Autonomous Behavior and Whole-Brain Dynamics Emerge in Embodied Zebrafish Agents (arxiv.org) |
|
1 point by leokoz8 on Jun 12, 2025 | hide | past | pdf | discuss
|
| 2938. |
Reasoning Language Models: A Blueprint (arxiv.org) |
|
2 points by Anon84 on Jun 12, 2025 | hide | past | pdf | 1 comment
|
| 2939. |
Text-to-LoRA: Instant Transformer Adaption (arxiv.org) |
|
5 points by hardmaru on Jun 12, 2025 | hide | past | pdf | discuss
|
| 2940. |
Can Theoretical Physics Research Benefit from Language Agents? (arxiv.org) |
|
1 point by in_a_hole on Jun 12, 2025 | hide | past | pdf | discuss
|
| More |