about
2821. Robustly improving LLM fairness in realistic settings via interpretability (arxiv.org)
1 point by like_any_other on Jul 1, 2025 | hide | past | pdf | discuss
2822. Small language models are the future of agentic AI (arxiv.org)
113 points by favoboa on Jul 1, 2025 | hide | past | pdf | 45 comments
2823. Survey on Evaluation of LLM-Based Agents (arxiv.org)
2 points by simonpure on Jul 1, 2025 | hide | past | pdf | discuss
2824. BIG-Bench Extra Hard (arxiv.org)
1 point by optimalsolver on Jun 30, 2025 | hide | past | pdf | discuss
2825. Transformers Are Graph Neural Networks (arxiv.org)
33 points by Anon84 on Jun 30, 2025 | hide | past | pdf | 2 comments
2826. Embodied AI Agents: Modeling the World (arxiv.org)
3 points by lucaspauker on Jun 30, 2025 | hide | past | pdf | discuss
2827. Large Language Model-Powered Agent for C to Rust Code Translation (arxiv.org)
4 points by elashri on Jun 30, 2025 | hide | past | pdf | discuss
2828. Mechanistic Interpretability of Emotion Inference in Large Language Models (arxiv.org)
3 points by cainxinth on Jun 30, 2025 | hide | past | pdf | discuss
2829. Alice's Adventures in a Differentiable Wonderland (arxiv.org)
157 points by henning on Jun 30, 2025 | hide | past | pdf | 26 comments
2830. Execution Outcomes of LLM-Generated versus Human Research Ideas (arxiv.org)
1 point by eamag on Jun 30, 2025 | hide | past | pdf | discuss
2831. Sequential Diagnosis with Language Models (arxiv.org)
2 points by FlyingLawnmower on Jun 30, 2025 | hide | past | pdf | 1 comment
2832. Pretrained Transformers as Universal Computation Engines (arxiv.org)
2 points by bilsbie on Jun 30, 2025 | hide | past | pdf | discuss
2833. Distillation Robustifies Unlearning (arxiv.org)
3 points by PaulHoule on Jun 30, 2025 | hide | past | pdf | discuss
2834. Microsoft releases foundation model of quantum wavefunctions (arxiv.org)
3 points by ae-foster on Jun 30, 2025 | hide | past | pdf | discuss
2835. WorldVLA: Towards Autoregressive Action World Model (arxiv.org)
25 points by chrsw on Jun 29, 2025 | hide | past | pdf | 2 comments
2836. Attention Is All You Need (arxiv.org)
6 points by Bluestein on Jun 29, 2025 | hide | past | pdf | 2 comments
2837. Storm – Help LLMs to write very long articles (arxiv.org)
2 points by mococa on Jun 29, 2025 | hide | past | pdf | discuss
2838. Evaluating World Models with LLM for Decision Making (arxiv.org)
4 points by Bluestein on Jun 29, 2025 | hide | past | pdf | discuss
2839. Enhancing LLM Reasoning with Reward-Guided Tree Search (arxiv.org)
6 points by Bluestein on Jun 29, 2025 | hide | past | pdf | discuss
2840. LLMs Capable of Metacognitive Monitoring Control of Their Internal Activations (arxiv.org)
6 points by Bluestein on Jun 29, 2025 | hide | past | pdf | discuss
2841. Mercury: Ultra-Fast Language Models Based on Diffusion (arxiv.org)
10 points by simonpure on Jun 29, 2025 | hide | past | pdf | 2 comments
2842. Universal pre-training by iterated random computation (arxiv.org)
37 points by liamdgray on Jun 29, 2025 | hide | past | pdf | 6 comments
2843. Hype, Sustainability, and the Price of the Bigger-Is-Better Paradigm in AI (arxiv.org)
3 points by weird_trousers on Jun 28, 2025 | hide | past | pdf | discuss
2844. Potemkin Understanding in Large Language Models (arxiv.org)
5 points by nsagent on Jun 28, 2025 | hide | past | pdf | discuss
2845. FineWeb2: Adapting Pre-Training Data Processing to Every Language (arxiv.org)
7 points by hynky on Jun 27, 2025 | hide | past | pdf | discuss
2846. Theoretical Analysis of Positional Encodings in Transformer Models (arxiv.org)
37 points by PaulHoule on Jun 27, 2025 | hide | past | pdf | 4 comments
2847. A Guide to Failure in Machine Learning (arxiv.org)
4 points by belter on Jun 27, 2025 | hide | past | pdf | discuss
2848. Code Researcher: Deep Research Agent for Large Systems Code and Commit History (arxiv.org)
2 points by PaulHoule on Jun 27, 2025 | hide | past | pdf | discuss
2849. Exploiting Local KV Cache Asymmetry for Long-Context LLMs (arxiv.org)
6 points by PaulHoule on Jun 27, 2025 | hide | past | pdf | discuss
2850. Natural Language Database Interaction in the Internet of Battlefield Things (arxiv.org)
2 points by PaulHoule on Jun 27, 2025 | hide | past | pdf | discuss