about
601. CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs(2024) (arxiv.org)
3 points by hamburgererror 80 days ago | hide | past | pdf | discuss
602. The Refusal Residue: When Probes Catch Alignment Faking and When They Don't (arxiv.org)
5 points by sbulaev 80 days ago | hide | past | pdf | discuss
603. Fleet: Hierarchical Task-Based Abstraction for Megakernels on Multi-Die GPUs (arxiv.org)
10 points by matt_d 80 days ago | hide | past | pdf | 1 comment
604. Can LLMs Perform Deep Technical Comprehension of Computer Architecture Papers (arxiv.org)
83 points by Jimmc414 80 days ago | hide | past | pdf | 26 comments
605. SynapticOS: An Inference-First Runtime Architecture for Neural Processing Units (arxiv.org)
5 points by Jimmc414 80 days ago | hide | past | pdf | discuss
606. Rethinking the Evaluation of Harness Evolution for Agents (arxiv.org)
4 points by Anon84 80 days ago | hide | past | pdf | discuss
607. Bridging the Gap Between Latent and Explicit Reasoning with Looped Transformers (arxiv.org)
3 points by simonpure 80 days ago | hide | past | pdf | discuss
608. LLM-as-a-Verifier: A General-Purpose Verification Framework (arxiv.org)
5 points by root-parent 81 days ago | hide | past | pdf | discuss
609. Detecting Prompt Injection Attacks on Purpose-Specific LLM Agents (arxiv.org)
3 points by zhinit 81 days ago | hide | past | pdf | discuss
610. Remember When It Matters: Proactive Memory Agent for Long-Horizon Agents (arxiv.org)
2 points by gmays 81 days ago | hide | past | pdf | discuss
611. GPU-Tile-SIM: Tile-Centric GPU Simulation for LLM Hardware-Software Co-Design (arxiv.org)
3 points by rbanffy 81 days ago | hide | past | pdf | discuss
612. Neural Texture Compression Using Hypernetworks (arxiv.org)
5 points by ibobev 81 days ago | hide | past | pdf | discuss
613. Unsupervised Representation Learning with Deep Convolutional GANs (arxiv.org)
3 points by ronfriedhaber 81 days ago | hide | past | pdf | discuss
614. Skills That Don't Exist: A Large-Scale Study of Hallucinated Skill (arxiv.org)
4 points by sbulaev 81 days ago | hide | past | pdf | discuss
615. Trust but Verify? Uncovering the Security Debt of Autonomous Coding Agents (arxiv.org)
4 points by Timofeibu 81 days ago | hide | past | pdf | discuss
616. Mako: A Self-Evolving Agentic Operating System for Autonomous Web Exploitation (arxiv.org)
2 points by praneethn 81 days ago | hide | past | pdf | discuss
617. LLM-as-a-Verifier: A General-Purpose Verification Framework (arxiv.org)
3 points by gmays 81 days ago | hide | past | pdf | discuss
618. Gemma 4 Technical Report (arxiv.org)
1 point by gmays 82 days ago | hide | past | pdf | discuss
619. The Ramanujan Challenge for AI (arxiv.org)
2 points by root-parent 82 days ago | hide | past | pdf | discuss
620. Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning (arxiv.org)
2 points by gmays 82 days ago | hide | past | pdf | discuss
621. Auditing the Risk Claims of Distributional Reinforcement Learning (arxiv.org)
2 points by sbulaev 82 days ago | hide | past | pdf | discuss
622. Coding agents think ahead of time (arxiv.org)
96 points by andre15silva 82 days ago | hide | past | pdf | 78 comments
623. The Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge (arxiv.org)
1 point by sbulaev 82 days ago | hide | past | pdf | discuss
624. ModelDNA: Verifying the lineage of open-weight LLMs from weight fingerprints (arxiv.org)
2 points by saadaamir14 82 days ago | hide | past | pdf | discuss
625. GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning (arxiv.org)
4 points by handfuloflight 82 days ago | hide | past | pdf | discuss
626. CTA-Pipelining: A Latency-Oriented Spatial Scaling Method for Multi-GPU Systems (arxiv.org)
2 points by matt_d 82 days ago | hide | past | pdf | discuss
627. Deceptive Grounding: Entity Attribution Failure in Clinical RAG (arxiv.org)
2 points by sbulaev 82 days ago | hide | past | pdf | discuss
628. Measuring Agents in Production – ICML (arxiv.org)
2 points by haritha1313 82 days ago | hide | past | pdf | discuss
629. Who Needs DRAM? We Have Fiber (arxiv.org)
12 points by alechammond 82 days ago | hide | past | pdf | 2 comments
630. Seeing Is Free, Speaking Is Not: Uncovering the True Energy Bottleneck in Edge (arxiv.org)
2 points by sbulaev 82 days ago | hide | past | pdf | discuss