| 2821. |
Robustly improving LLM fairness in realistic settings via interpretability (arxiv.org) |
|
1 point by like_any_other on Jul 1, 2025 | hide | past | pdf | discuss
|
| 2822. |
Small language models are the future of agentic AI (arxiv.org) |
|
113 points by favoboa on Jul 1, 2025 | hide | past | pdf | 45 comments
|
| 2823. |
Survey on Evaluation of LLM-Based Agents (arxiv.org) |
|
2 points by simonpure on Jul 1, 2025 | hide | past | pdf | discuss
|
| 2824. |
BIG-Bench Extra Hard (arxiv.org) |
|
1 point by optimalsolver on Jun 30, 2025 | hide | past | pdf | discuss
|
| 2825. |
Transformers Are Graph Neural Networks (arxiv.org) |
|
33 points by Anon84 on Jun 30, 2025 | hide | past | pdf | 2 comments
|
| 2826. |
Embodied AI Agents: Modeling the World (arxiv.org) |
|
3 points by lucaspauker on Jun 30, 2025 | hide | past | pdf | discuss
|
| 2827. |
Large Language Model-Powered Agent for C to Rust Code Translation (arxiv.org) |
|
4 points by elashri on Jun 30, 2025 | hide | past | pdf | discuss
|
| 2828. |
Mechanistic Interpretability of Emotion Inference in Large Language Models (arxiv.org) |
|
3 points by cainxinth on Jun 30, 2025 | hide | past | pdf | discuss
|
| 2829. |
Alice's Adventures in a Differentiable Wonderland (arxiv.org) |
|
157 points by henning on Jun 30, 2025 | hide | past | pdf | 26 comments
|
| 2830. |
Execution Outcomes of LLM-Generated versus Human Research Ideas (arxiv.org) |
|
1 point by eamag on Jun 30, 2025 | hide | past | pdf | discuss
|
| 2831. |
Sequential Diagnosis with Language Models (arxiv.org) |
|
2 points by FlyingLawnmower on Jun 30, 2025 | hide | past | pdf | 1 comment
|
| 2832. |
Pretrained Transformers as Universal Computation Engines (arxiv.org) |
|
2 points by bilsbie on Jun 30, 2025 | hide | past | pdf | discuss
|
| 2833. |
Distillation Robustifies Unlearning (arxiv.org) |
|
3 points by PaulHoule on Jun 30, 2025 | hide | past | pdf | discuss
|
| 2834. |
Microsoft releases foundation model of quantum wavefunctions (arxiv.org) |
|
3 points by ae-foster on Jun 30, 2025 | hide | past | pdf | discuss
|
| 2835. |
WorldVLA: Towards Autoregressive Action World Model (arxiv.org) |
|
25 points by chrsw on Jun 29, 2025 | hide | past | pdf | 2 comments
|
| 2836. |
Attention Is All You Need (arxiv.org) |
|
6 points by Bluestein on Jun 29, 2025 | hide | past | pdf | 2 comments
|
| 2837. |
Storm – Help LLMs to write very long articles (arxiv.org) |
|
2 points by mococa on Jun 29, 2025 | hide | past | pdf | discuss
|
| 2838. |
Evaluating World Models with LLM for Decision Making (arxiv.org) |
|
4 points by Bluestein on Jun 29, 2025 | hide | past | pdf | discuss
|
| 2839. |
Enhancing LLM Reasoning with Reward-Guided Tree Search (arxiv.org) |
|
6 points by Bluestein on Jun 29, 2025 | hide | past | pdf | discuss
|
| 2840. |
LLMs Capable of Metacognitive Monitoring Control of Their Internal Activations (arxiv.org) |
|
6 points by Bluestein on Jun 29, 2025 | hide | past | pdf | discuss
|
| 2841. |
Mercury: Ultra-Fast Language Models Based on Diffusion (arxiv.org) |
|
10 points by simonpure on Jun 29, 2025 | hide | past | pdf | 2 comments
|
| 2842. |
Universal pre-training by iterated random computation (arxiv.org) |
|
37 points by liamdgray on Jun 29, 2025 | hide | past | pdf | 6 comments
|
| 2843. |
Hype, Sustainability, and the Price of the Bigger-Is-Better Paradigm in AI (arxiv.org) |
|
3 points by weird_trousers on Jun 28, 2025 | hide | past | pdf | discuss
|
| 2844. |
Potemkin Understanding in Large Language Models (arxiv.org) |
|
5 points by nsagent on Jun 28, 2025 | hide | past | pdf | discuss
|
| 2845. |
FineWeb2: Adapting Pre-Training Data Processing to Every Language (arxiv.org) |
|
7 points by hynky on Jun 27, 2025 | hide | past | pdf | discuss
|
| 2846. |
Theoretical Analysis of Positional Encodings in Transformer Models (arxiv.org) |
|
37 points by PaulHoule on Jun 27, 2025 | hide | past | pdf | 4 comments
|
| 2847. |
A Guide to Failure in Machine Learning (arxiv.org) |
|
4 points by belter on Jun 27, 2025 | hide | past | pdf | discuss
|
| 2848. |
Code Researcher: Deep Research Agent for Large Systems Code and Commit History (arxiv.org) |
|
2 points by PaulHoule on Jun 27, 2025 | hide | past | pdf | discuss
|
| 2849. |
Exploiting Local KV Cache Asymmetry for Long-Context LLMs (arxiv.org) |
|
6 points by PaulHoule on Jun 27, 2025 | hide | past | pdf | discuss
|
| 2850. |
Natural Language Database Interaction in the Internet of Battlefield Things (arxiv.org) |
|
2 points by PaulHoule on Jun 27, 2025 | hide | past | pdf | discuss
|
| More |