about
3211. Don't be lazy: CompleteP enables compute-efficient deep transformers (arxiv.org)
4 points by nsdey on May 6, 2025 | hide | past | pdf | discuss
3212. Cost-of-Pass: An Economic Framework for Evaluating Language Models (arxiv.org)
1 point by JumpCrisscross on May 6, 2025 | hide | past | pdf | discuss
3213. Unveiling the Hidden: Movie Genre and User Bias in Spoiler Detection (arxiv.org)
3 points by PaulHoule on May 6, 2025 | hide | past | pdf | discuss
3214. HalluMix Benchmark: Detecting Hallucinations in Real-World Scenarios (arxiv.org)
2 points by rancar2 on May 6, 2025 | hide | past | pdf | discuss
3215. (How) Do reasoning models reason? (arxiv.org)
3 points by YeGoblynQueenne on May 6, 2025 | hide | past | pdf | discuss
3216. Evaluating Frontier Models for Stealth and Situational Awareness (arxiv.org)
2 points by badmonster on May 6, 2025 | hide | past | pdf | discuss
3217. Analyzing Modern Nvidia GPU Cores (arxiv.org)
178 points by mfiguiere on May 5, 2025 | hide | past | pdf | 37 comments
3218. ReadMe.LLM: LLM-Oriented Documentation (arxiv.org)
2 points by dewley452 on May 5, 2025 | hide | past | pdf | 1 comment
3219. Multimodal Doctor-in-the-Loop: A Clinically-Guided Explainable Framework (arxiv.org)
2 points by badmonster on May 5, 2025 | hide | past | pdf | discuss
3220. CrashFixer: A crash resolution agent for the Linux kernel (arxiv.org)
2 points by chrsw on May 5, 2025 | hide | past | pdf | discuss
3221. Learning Large-Scale Competitive Team Behaviors with Mean-Field Interactions (arxiv.org)
4 points by PaulHoule on May 5, 2025 | hide | past | pdf | discuss
3222. JPEC: A Novel GNN for Competitor Retrieval in Financial Knowledge Graphs (arxiv.org)
2 points by sonabinu on May 5, 2025 | hide | past | pdf | discuss
3223. Evaluating Frontier Models for Stealth and Situational Awareness (arxiv.org)
1 point by nkko on May 5, 2025 | hide | past | pdf | discuss
3224. Llama-Nemotron: Efficient Reasoning Models (arxiv.org)
7 points by nkko on May 5, 2025 | hide | past | pdf | discuss
3225. TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks (arxiv.org)
2 points by EvgeniyZh on May 5, 2025 | hide | past | pdf | discuss
3226. Matrix-vector multiplication implemented in off-the-shelf DRAM for Low-Bit LLMs (arxiv.org)
230 points by cpldcpu on May 4, 2025 | hide | past | pdf | 53 comments
3227. Your ViT Is Secretly an Image Segmentation Model (arxiv.org)
10 points by lamename on May 4, 2025 | hide | past | pdf | discuss
3228. Chain-of-Draft surpasses CoT with only 7.6% of tokens (arxiv.org)
1 point by felineflock on May 4, 2025 | hide | past | pdf | discuss
3229. Robotic Visual Instruction (arxiv.org)
3 points by badmonster on May 4, 2025 | hide | past | pdf | discuss
3230. CMU TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks (arxiv.org)
7 points by walterbell on May 4, 2025 | hide | past | pdf | discuss
3231. Can Transformers Reason Logically? A Study in SAT Solving (arxiv.org)
1 point by sonabinu on May 4, 2025 | hide | past | pdf | discuss
3232. A Survey of AI Agent Protocols (arxiv.org)
91 points by distalx on May 4, 2025 | hide | past | pdf | 63 comments
3233. Reinforcement Learning for Reasoning in LLMs with One Training Example (arxiv.org)
1 point by mococa on May 3, 2025 | hide | past | pdf | 1 comment
3234. Understanding Is Compression (arxiv.org)
3 points by liamdgray on May 3, 2025 | hide | past | pdf | 1 comment
3235. Stop treating `AGI' as the north-star goal of AI research (arxiv.org)
46 points by todsacerdoti on May 3, 2025 | hide | past | pdf | 32 comments
3236. AI-LieDar: Examine the Trade-Off Between Utility and Truthfulness in LLM Agents (arxiv.org)
1 point by msvana on May 2, 2025 | hide | past | pdf | discuss
3237. AstroAgents: MultiAgent AI for Hypothesis Generation from Mass Spectrometry Data (arxiv.org)
2 points by rntn on May 2, 2025 | hide | past | pdf | discuss
3238. Step-by-step reasoning verifiers that think (arxiv.org)
6 points by mkhalifa on May 2, 2025 | hide | past | pdf | 1 comment
3239. Who Gets the Callback? Generative AI and Gender Bias (arxiv.org)
1 point by rntn on May 2, 2025 | hide | past | pdf | discuss
3240. Exploring the Impact of Personality Traits on Conversational Recommender Systems (arxiv.org)
1 point by PaulHoule on May 2, 2025 | hide | past | pdf | discuss