about
3691. A Survey on Large Language Models (2025) (arxiv.org)
1 point by OutOfHere on Feb 12, 2025 | hide | past | pdf | discuss
3692. Competitive Programming with Large Reasoning Models (arxiv.org)
16 points by t55 on Feb 12, 2025 | hide | past | pdf | 1 comment
3693. Reducing the Transformer Architecture to a Minimum [pdf] (arxiv.org)
2 points by DoctorOetker on Feb 12, 2025 | hide | past | pdf | discuss
3694. The Curse of Depth in Large Language Models [pdf] (arxiv.org)
2 points by mkaic on Feb 11, 2025 | hide | past | pdf | discuss
3695. Turning Up the Heat: Min-P Sampling for Creative and Coherent LLM Outputs (arxiv.org)
1 point by Der_Einzige on Feb 11, 2025 | hide | past | pdf | discuss
3696. LLMs can teach themselves to better predict the future (arxiv.org)
176 points by bturtel on Feb 11, 2025 | hide | past | pdf | 86 comments
3697. Time to act on the risk of efficient personalized text generation (arxiv.org)
57 points by Jimmc414 on Feb 11, 2025 | hide | past | pdf | 34 comments
3698. Resurrecting saturated LLM benchmarks with adversarial encoding (arxiv.org)
1 point by bearseascape on Feb 11, 2025 | hide | past | pdf | discuss
3699. Training LLMs to Reason Efficiently (arxiv.org)
2 points by omarsar on Feb 11, 2025 | hide | past | pdf | discuss
3700. Frontier AI systems have surpassed the self-replicating red line (arxiv.org)
3 points by rahton on Feb 11, 2025 | hide | past | pdf | 4 comments
3701. Deep Networks Always Grok and Here Is Why (arxiv.org)
1 point by ziofill on Feb 11, 2025 | hide | past | pdf | discuss
3702. Matryoshka Quantization (arxiv.org)
2 points by fzliu on Feb 11, 2025 | hide | past | pdf | discuss
3703. Frontier AI systems have surpassed the self-replicating red line (arxiv.org)
10 points by ryan_j_naughton on Feb 10, 2025 | hide | past | pdf | 4 comments
3704. s1: Simple Test-Time Scaling (arxiv.org)
2 points by btilly on Feb 10, 2025 | hide | past | pdf | discuss
3705. Scaling up test-time compute with latent reasoning: A recurrent depth approach (arxiv.org)
149 points by timbilt on Feb 10, 2025 | hide | past | pdf | 44 comments
3706. LLM Failure Modes in Medical QA Arising from Inflexible Reasoning (arxiv.org)
3 points by docere on Feb 10, 2025 | hide | past | pdf | discuss
3707. Self-Backtracking for Boosting Reasoning of LLMs (arxiv.org)
1 point by omarsar on Feb 10, 2025 | hide | past | pdf | discuss
3708. ZebraLogic: On The Scaling Limits Of LLMs For Logical Reasoning (arxiv.org)
1 point by optimalsolver on Feb 10, 2025 | hide | past | pdf | 1 comment
3709. Adaptive Computation Time for Recurrent Neural Networks (2016) (arxiv.org)
2 points by tosh on Feb 10, 2025 | hide | past | pdf | discuss
3710. Scaling Up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach (arxiv.org)
1 point by tosh on Feb 10, 2025 | hide | past | pdf | discuss
3711. The Impact of Prompt Programming on Function-Level Code Generation (arxiv.org)
3 points by bobrenjc93 on Feb 10, 2025 | hide | past | pdf | discuss
3712. Extractive Schema Linking for Text-to-SQL (arxiv.org)
1 point by PaulHoule on Feb 9, 2025 | hide | past | pdf | discuss
3713. The Differences Between Direct Alignment Algorithms Are a Blur (arxiv.org)
8 points by t55 on Feb 9, 2025 | hide | past | pdf | discuss
3714. PhD Knowledge Not Required: A Reasoning Challenge for Large Language Models (arxiv.org)
174 points by enum on Feb 9, 2025 | hide | past | pdf | 80 comments
3715. LIMO: Less Is More for Reasoning (arxiv.org)
389 points by trott on Feb 9, 2025 | hide | past | pdf | 128 comments
3716. Frontier AI systems have surpassed the self-replicating red line (arxiv.org)
24 points by LLcolD on Feb 9, 2025 | hide | past | pdf | 5 comments
3717. Demystifying Long Chain-of-Thought Reasoning in LLMs (arxiv.org)
11 points by Anon84 on Feb 8, 2025 | hide | past | pdf | discuss
3718. STP: Self-Play LLM Theorem Provers with Iterative Conjecturing and Proving (arxiv.org)
3 points by heydenberk on Feb 8, 2025 | hide | past | pdf | discuss
3719. Bolt: Bootstrap long chain-of-thought in LLMs without distillation [pdf] (arxiv.org)
15 points by TaurenHunter on Feb 8, 2025 | hide | past | pdf | 5 comments
3720. Value-Based Deep RL Scales Predictably (arxiv.org)
68 points by bearseascape on Feb 8, 2025 | hide | past | pdf | 3 comments