about
5101. SWE-Agent: Agent-Computer Interfaces Enable Automated Software Engineering (arxiv.org)
1 point by broyojo on May 28, 2024 | hide | past | pdf | discuss
5102. Transformers Can Do Arithmetic with the Right Embeddings (arxiv.org)
207 points by byt3h3ad on May 28, 2024 | hide | past | pdf | 211 comments
5103. Grokked Transformers Are Implicit Reasoners (arxiv.org)
239 points by jasondavies on May 27, 2024 | hide | past | pdf | 61 comments
5104. Extracting Prompts by Inverting LLM Outputs (arxiv.org)
2 points by xcccube on May 27, 2024 | hide | past | pdf | 1 comment
5105. AstroPT: Scaling Large Observation Models for Astronomy (arxiv.org)
2 points by Smith42 on May 27, 2024 | hide | past | pdf | 1 comment
5106. Neuromorphic dreaming: A pathway to efficient learning in artificial agents (arxiv.org)
2 points by belter on May 27, 2024 | hide | past | pdf | discuss
5107. Kolmogorov-Arnold Networks (KANs) for Time Series Analysis (arxiv.org)
3 points by jasondavies on May 27, 2024 | hide | past | pdf | discuss
5108. Principled Instructions Are All You Need for Questioning LLaMA-1/2, GPT-3.5/4 (arxiv.org)
2 points by belter on May 27, 2024 | hide | past | pdf | discuss
5109. Models That Prove Their Own Correctness (arxiv.org)
4 points by laybak on May 27, 2024 | hide | past | pdf | discuss
5110. Look Once to Hear: Target Speech Hearing with Noisy Examples (arxiv.org)
1 point by jasondavies on May 26, 2024 | hide | past | pdf | discuss
5111. How Far Are We from AGI (arxiv.org)
5 points by RafelMri on May 26, 2024 | hide | past | pdf | 4 comments
5112. Exploring Autonomous Agents Through the Lens of Large Language Models (arxiv.org)
2 points by sabrina_ramonov on May 26, 2024 | hide | past | pdf | discuss
5113. Feather: Data Reordering for On-Chip Dataflow Switching (arxiv.org)
1 point by breck on May 26, 2024 | hide | past | pdf | discuss
5114. Attention as an RNN (arxiv.org)
2 points by RafelMri on May 26, 2024 | hide | past | pdf | discuss
5115. Some models are useful, but for how long?: When to refit prediction models (arxiv.org)
1 point by nequo on May 26, 2024 | hide | past | pdf | discuss
5116. MoRA: High-Rank Updating for Parameter-Efficient Fine-Tuning (arxiv.org)
3 points by jasondavies on May 25, 2024 | hide | past | pdf | discuss
5117. A Prefrontal Cortex-Inspired Architecture for Planning in Large Language Models (arxiv.org)
2 points by rntn on May 25, 2024 | hide | past | pdf | 1 comment
5118. Evaluation of the Programming Skills of Large Language Models (arxiv.org)
3 points by belter on May 25, 2024 | hide | past | pdf | discuss
5119. Not All Language Model Features Are Linear (arxiv.org)
1 point by belter on May 25, 2024 | hide | past | pdf | discuss
5120. A Case Study in CUDA Kernel Fusion (arxiv.org)
1 point by veryluckyxyz on May 25, 2024 | hide | past | pdf | discuss
5121. Lessons from the trenches on reproducible evaluation of language models (arxiv.org)
42 points by veryluckyxyz on May 25, 2024 | hide | past | pdf | 3 comments
5122. A ConvNet for the 2020s (arxiv.org)
18 points by laybak on May 25, 2024 | hide | past | pdf | discuss
5123. Bird's-Eye View to Street-View: A Survey (arxiv.org)
1 point by PaulHoule on May 24, 2024 | hide | past | pdf | discuss
5124. "Yes I Would Recommend Calling the Police":Norm Inconsistency in LLM Decisions (arxiv.org)
1 point by belter on May 24, 2024 | hide | past | pdf | discuss
5125. Lessons from the Trenches on Reproducible Evaluation of Language Models (arxiv.org)
1 point by tosh on May 24, 2024 | hide | past | pdf | discuss
5126. YOLOv10: Real-Time End-to-End Object Detection (arxiv.org)
2 points by zerojames on May 24, 2024 | hide | past | pdf | discuss
5127. Thermodynamic Natural Gradient Descent (arxiv.org)
200 points by jasondavies on May 24, 2024 | hide | past | pdf | 32 comments
5128. A Declarative System for Optimizing AI Workloads (arxiv.org)
2 points by zerojames on May 24, 2024 | hide | past | pdf | discuss
5129. Representation noising effectively prevents harmful fine-tuning on LLMs (arxiv.org)
1 point by darosati on May 24, 2024 | hide | past | pdf | discuss
5130. You Only Cache Once: Decoder-Decoder Architectures for Language Models (arxiv.org)
3 points by belter on May 24, 2024 | hide | past | pdf | discuss