about
1681. Towards Execution-Grounded Automated AI Research (arxiv.org)
2 points by abracos 255 days ago | hide | past | pdf | discuss
1682. The unreasonable effectiveness of pattern matching (arxiv.org)
3 points by chbint 255 days ago | hide | past | pdf | 2 comments
1683. Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in CLIs (arxiv.org)
2 points by matt_d 256 days ago | hide | past | pdf | discuss
1684. Us-vs-Them Bias in Large Language Models (arxiv.org)
1 point by geox 256 days ago | hide | past | pdf | discuss
1685. Reasoning Models Generate Societies of Thought (arxiv.org)
3 points by Anon84 256 days ago | hide | past | pdf | discuss
1686. Convolutional-neural-operator-based transfer learning for solving PDEs (arxiv.org)
2 points by PaulHoule 256 days ago | hide | past | pdf | 1 comment
1687. SlimEdge: Lightweight Distributed DNN Deployment on Constrained Hardware (arxiv.org)
1 point by PaulHoule 256 days ago | hide | past | pdf | discuss
1688. DiffRatio: A SOTA one-step Diffusion model with 50% less GPU memory (arxiv.org)
1 point by LoMoGan 256 days ago | hide | past | pdf | discuss
1689. Human-Like Working Memory from Artificial Intrinsic Plasticity Neurons (arxiv.org)
1 point by PaulHoule 257 days ago | hide | past | pdf | discuss
1690. Delegated Authorization Constraining Agents to Semantic Task-to-Scope Matching (arxiv.org)
1 point by mooreds 257 days ago | hide | past | pdf | discuss
1691. DiffRatio – A One-Step Diffusion Model with SOTA quality and 50% less memory (arxiv.org)
4 points by LoMoGan 257 days ago | hide | past | pdf | 1 comment
1692. Provably unmasking malicious behavior through execution traces (arxiv.org)
46 points by PaulHoule 258 days ago | hide | past | pdf | 5 comments
1693. The Legal Embedding Benchmark (MLEB) (arxiv.org)
1 point by fzliu 258 days ago | hide | past | pdf | discuss
1694. WildCAT3D: Appearance-Aware Multi-View Diffusion in the Wild (arxiv.org)
3 points by PaulHoule 258 days ago | hide | past | pdf | discuss
1695. Repeating your prompt twice before sending it to an LLM improves accuracy (arxiv.org)
2 points by daviducolo 258 days ago | hide | past | pdf | discuss
1696. Perturb Your Data: Paraphrase-Guided Training Data Watermarking (arxiv.org)
1 point by PaulHoule 258 days ago | hide | past | pdf | discuss
1697. Fair is Better than Sensational:Man is to Doctor as Woman is to Doctor (2019) (arxiv.org)
1 point by bhickey 258 days ago | hide | past | pdf | discuss
1698. DiffusionBlocks: Block-Wise Neural Network Training (arxiv.org)
3 points by E-Reverance 259 days ago | hide | past | pdf | 5 comments
1699. The Assistant Axis: Situating/Stabilizing the Default Persona of Language Models (arxiv.org)
1 point by lawrenceyan 259 days ago | hide | past | pdf | discuss
1700. The unreasonable effectiveness of pattern matching (arxiv.org)
7 points by headalgorithm 259 days ago | hide | past | pdf | 1 comment
1701. Prompt Repetition Improves Non-Reasoning LLMs (arxiv.org)
3 points by Tomte 259 days ago | hide | past | pdf | discuss
1702. Do You Trust Me? Cognitive-Affective Signatures of Trustworthiness in LLMs (arxiv.org)
4 points by 7777777phil 259 days ago | hide | past | pdf | discuss
1703. Too Helpful to Be Safe: User-Mediated Attacks on Planning and Web-Use Agents (arxiv.org)
4 points by 7777777phil 259 days ago | hide | past | pdf | discuss
1704. From Code Foundation Models to Agents and Applications (arxiv.org)
2 points by tamnd 260 days ago | hide | past | pdf | discuss
1705. Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity (arxiv.org)
1 point by ycombiredd 260 days ago | hide | past | pdf | 2 comments
1706. VaultGemma: A Differentially Private LLM (arxiv.org)
3 points by todsacerdoti 260 days ago | hide | past | pdf | discuss
1707. Prompt Repetition Improves Non-Reasoning LLMs (arxiv.org)
2 points by UntitledNo4 261 days ago | hide | past | pdf | discuss
1708. Private LLM Inference on Consumer Blackwell GPUs (arxiv.org)
3 points by Teever 261 days ago | hide | past | pdf | discuss
1709. Categorizing Variants of Goodhart's Law (arxiv.org)
1 point by foster_nyman 261 days ago | hide | past | pdf | discuss
1710. Future-as-Label: Scalable Supervision from Real-World Outcomes (arxiv.org)
17 points by bturtel 262 days ago | hide | past | pdf | discuss