about
2911. Revealing Political Bias in LLMs Through Structured Multi-Agent Debate (arxiv.org)
1 point by rntn on Jun 16, 2025 | hide | past | pdf | discuss
2912. Appraisal-Based Chain-of-Emotion Improves AI Persona Accuracy (arxiv.org)
1 point by virtual_rf on Jun 16, 2025 | hide | past | pdf | discuss
2913. The Illusion of the Illusion of Thinking – A Comment on Shojaee et al. (2025) (arxiv.org)
16 points by gfortaine on Jun 16, 2025 | hide | past | pdf | 14 comments
2914. LiveCodeBench Pro: How Olympiad Medalists Judge LLMs in Competitive Programming? (arxiv.org)
2 points by EvgeniyZh on Jun 16, 2025 | hide | past | pdf | discuss
2915. Towards Understanding Sycophancy in Language Models (arxiv.org)
9 points by fzliu on Jun 16, 2025 | hide | past | pdf | 2 comments
2916. Vision Transformers Don't Need Trained Registers (arxiv.org)
4 points by avd4292 on Jun 16, 2025 | hide | past | pdf | discuss
2917. A Word Is Worth 4-Bit: Log Parsing with Binary Coded Decimal Recognition (arxiv.org)
7 points by PaulHoule on Jun 15, 2025 | hide | past | pdf | discuss
2918. Large language models often know when they are being evaluated (arxiv.org)
89 points by jonbaer on Jun 15, 2025 | hide | past | pdf | 130 comments
2919. Clinical knowledge in LLMs does not translate to human interactions (arxiv.org)
102 points by insistent on Jun 14, 2025 | hide | past | pdf | 39 comments
2920. Exploring the Best Input Representation for Electrocardiogram-Language Models (arxiv.org)
2 points by PaulHoule on Jun 14, 2025 | hide | past | pdf | discuss
2921. Eliciting Fine-Tuned Transformer Capabilities via Inference-Time Techniques (arxiv.org)
1 point by codelion on Jun 14, 2025 | hide | past | pdf | discuss
2922. Relic: Evaluating Compositional Instruction Following via Language Recognition (arxiv.org)
1 point by demirbey05 on Jun 14, 2025 | hide | past | pdf | discuss
2923. Resa: Transparent Reasoning Models via SAEs (arxiv.org)
1 point by Bogdanp on Jun 14, 2025 | hide | past | pdf | discuss
2924. Unsupervised Elicitation of Language Models (arxiv.org)
135 points by kordlessagain on Jun 14, 2025 | hide | past | pdf | 24 comments
2925. CRMArena-Pro: LLM Agents Assessed Across Diverse Business Scenarios (arxiv.org)
1 point by felineflock on Jun 14, 2025 | hide | past | pdf | discuss
2926. Memoir: Lifelong Model Editing with Minimal Overwrite Informed Retention for LLM (arxiv.org)
1 point by dataminer on Jun 14, 2025 | hide | past | pdf | discuss
2927. Comment on the Illusion of Thinking (arxiv.org)
4 points by esafak on Jun 14, 2025 | hide | past | pdf | 1 comment
2928. Rethinking Losses for Diffusion Bridge Samplers (arxiv.org)
10 points by badmonster on Jun 14, 2025 | hide | past | pdf | 1 comment
2929. Unsupervised Elicitation of Language Models (arxiv.org)
7 points by xianshou on Jun 13, 2025 | hide | past | pdf | discuss
2930. Self-Adapting Language Models (arxiv.org)
246 points by archon1410 on Jun 13, 2025 | hide | past | pdf | 73 comments
2931. How Do Large Language Monkeys Get Their Power (Laws)? (arxiv.org)
2 points by RSchaeffer on Jun 13, 2025 | hide | past | pdf | 1 comment
2932. Holistic Assessment of LLM Agents Across Diverse Scenarios and Interactions (arxiv.org)
2 points by prisenco on Jun 13, 2025 | hide | past | pdf | discuss
2933. TabM: Advancing Tabular Deep Learning with Parameter-Efficient Ensembling (arxiv.org)
3 points by felineflock on Jun 13, 2025 | hide | past | pdf | discuss
2934. Jailbreak attacks against DNA language models with pathogenicity guidance (arxiv.org)
2 points by PaulHoule on Jun 12, 2025 | hide | past | pdf | 1 comment
2935. TabM: Advancing Tabular Deep Learning with Parameter-Efficient Ensembling (arxiv.org)
2 points by m0nhawk on Jun 12, 2025 | hide | past | pdf | discuss
2936. Learning Semantically Faithful EEG-to-Text Generation (arxiv.org)
1 point by PaulHoule on Jun 12, 2025 | hide | past | pdf | discuss
2937. Autonomous Behavior and Whole-Brain Dynamics Emerge in Embodied Zebrafish Agents (arxiv.org)
1 point by leokoz8 on Jun 12, 2025 | hide | past | pdf | discuss
2938. Reasoning Language Models: A Blueprint (arxiv.org)
2 points by Anon84 on Jun 12, 2025 | hide | past | pdf | 1 comment
2939. Text-to-LoRA: Instant Transformer Adaption (arxiv.org)
5 points by hardmaru on Jun 12, 2025 | hide | past | pdf | discuss
2940. Can Theoretical Physics Research Benefit from Language Agents? (arxiv.org)
1 point by in_a_hole on Jun 12, 2025 | hide | past | pdf | discuss