about
2071. Deep sequence models tend to memorize geometrically; it is unclear why (arxiv.org)
4 points by amichail 333 days ago | hide | past | pdf | discuss
2072. Accumulating Context Changes the Beliefs of Language Models (arxiv.org)
2 points by Anon84 333 days ago | hide | past | pdf | discuss
2073. Continuous Autoregressive Language Models (arxiv.org)
115 points by Anon84 334 days ago | hide | past | pdf | 10 comments
2074. Evaluating Probabilistic Reasoning in LLMs Through Language-Only Decision Tasks (arxiv.org)
1 point by PaulHoule 334 days ago | hide | past | pdf | 1 comment
2075. Best-Of-∞: Asymptotic Performance of Test-Time Compute (Infinite Compute Budget) (arxiv.org)
2 points by SweetSoftPillow 334 days ago | hide | past | pdf | discuss
2076. Pre-training under infinite compute (arxiv.org)
20 points by SweetSoftPillow 334 days ago | hide | past | pdf | discuss
2077. Kosmos: An AI Scientist for Autonomous Discovery (arxiv.org)
60 points by belter 334 days ago | hide | past | pdf | 20 comments
2078. AutoCode: LLMs as Problem Setters for Competitive Programming (arxiv.org)
1 point by PaulHoule 334 days ago | hide | past | pdf | discuss
2079. Benchmarking multilingual long-context language models (arxiv.org)
3 points by sysoleg 334 days ago | hide | past | pdf | discuss
2080. Can LLMs Subtract Numbers? (arxiv.org)
2 points by belter 334 days ago | hide | past | pdf | discuss
2081. Towards General Auditory Intelligence (arxiv.org)
2 points by selimonder 335 days ago | hide | past | pdf | 1 comment
2082. Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity (arxiv.org)
1 point by JnBrymn 335 days ago | hide | past | pdf | discuss
2083. Continuous Autoregressive Language Models (arxiv.org)
3 points by guybedo 335 days ago | hide | past | pdf | 1 comment
2084. Cache-to-Cache: Direct Semantic Communication Between Large Language Models (arxiv.org)
14 points by jonbaer 335 days ago | hide | past | pdf | discuss
2085. LLM's Report Subjective Experience Under Self-Referential Processing (arxiv.org)
1 point by gradus_ad 335 days ago | hide | past | pdf | discuss
2086. LLMZip: Lossless Text Compression Using Large Language Models (arxiv.org)
2 points by jfantl 336 days ago | hide | past | pdf | 4 comments
2087. Deep sequence models tend to memorize geometrically (arxiv.org)
3 points by jonbaer 336 days ago | hide | past | pdf | discuss
2088. Streaming DiLoCo: Towards a Distributed Free Lunch (arxiv.org)
2 points by nathan-barry 336 days ago | hide | past | pdf | discuss
2089. Kimi Linear: An Expressive, Efficient Attention Architecture (arxiv.org)
6 points by birriel 337 days ago | hide | past | pdf | discuss
2090. Defeating the Training-Inference Mismatch via FP16 (arxiv.org)
1 point by billyzs 337 days ago | hide | past | pdf | discuss
2091. R2T: Rule-Encoded Loss Functions for Low-Resource Sequence Tagging (arxiv.org)
4 points by PaulHoule 337 days ago | hide | past | pdf | discuss
2092. Humains-Junior: A 3.8B Language Model Achieving GPT-4o-Level Factual Accuracy (arxiv.org)
2 points by gidellav 338 days ago | hide | past | pdf | discuss
2093. Watermarking for Generative AI (arxiv.org)
17 points by gidellav 338 days ago | hide | past | pdf | discuss
2094. Agentic AI Home Energy Management System: Residential Load Scheduling (arxiv.org)
2 points by simonpure 338 days ago | hide | past | pdf | 2 comments
2095. TempoPFN: Synthetic Pre-Training of Linear RNNs for Zero-Shot Timeseries Forecas (arxiv.org)
1 point by jul8234 338 days ago | hide | past | pdf | discuss
2096. Remote Labor Index: Measuring AI Automation of Remote Work (arxiv.org)
2 points by belter 338 days ago | hide | past | pdf | discuss
2097. LLMs Report Subjective Experience Under Self-Referential Processing (arxiv.org)
3 points by j_crick 339 days ago | hide | past | pdf | 1 comment
2098. Defeating the Training-Inference Mismatch via FP16 (arxiv.org)
2 points by matt_d 339 days ago | hide | past | pdf | discuss
2099. Benchmarking On-Device Machine Learning on Apple Silicon with MLX (arxiv.org)
2 points by PaulHoule 339 days ago | hide | past | pdf | discuss
2100. Chain-of-Thought Hijacking (arxiv.org)
2 points by belter 339 days ago | hide | past | pdf | discuss