about
1021. AI co-mathematician: Accelerating mathematicians with agentic AI (arxiv.org)
3 points by aoki 143 days ago | hide | past | pdf | discuss
1022. Normalizing Trajectory Models (arxiv.org)
8 points by gmays 143 days ago | hide | past | pdf | discuss
1023. VectorSmuggle: Steganographic exfiltration in vector embedding stores (arxiv.org)
2 points by smugglereal 143 days ago | hide | past | pdf | discuss
1024. Information Extraction from Electricity Invoices with General-Purpose LLMs (arxiv.org)
2 points by PaulHoule 143 days ago | hide | past | pdf | discuss
1025. LLM Targeted Underperformance Disproportionately Impacts Vulnerable Users (arxiv.org)
9 points by yogthos 143 days ago | hide | past | pdf | 3 comments
1026. Attention Once Is All You Need: Stateful Transformers (arxiv.org)
3 points by logotype 143 days ago | hide | past | pdf | 4 comments
1027. Spire: Structure-Preserving Interpretable Retrieval of Evidence (arxiv.org)
4 points by PaulHoule 144 days ago | hide | past | pdf | discuss
1028. EditLens: Quantifying the extent of AI editing in text (2025) (arxiv.org)
28 points by horseradish 144 days ago | hide | past | pdf | 5 comments
1029. Continual Harness: Online Adaptation for Self-Improving Foundation Agents (arxiv.org)
8 points by milkkarten 144 days ago | hide | past | pdf | 1 comment
1030. Algorithm Selection with Zero Domain Knowledge via Text Embeddings (arxiv.org)
1 point by PaulHoule 144 days ago | hide | past | pdf | discuss
1031. TLX: Hardware-Native, Evolvable MIMW GPU Compiler for Large-Scale Production (arxiv.org)
1 point by matt_d 144 days ago | hide | past | pdf | discuss
1032. Exascale Training of Generative Model with Historical Priors for Data Reduction (arxiv.org)
1 point by rbanffy 144 days ago | hide | past | pdf | discuss
1033. MMTB: Evaluating Terminal Agents on Multimedia-File Tasks (arxiv.org)
1 point by Brajeshwar 144 days ago | hide | past | pdf | discuss
1034. Behavioral Integrity Verification for AI Agent Skills (arxiv.org)
1 point by Timofeibu 144 days ago | hide | past | pdf | discuss
1035. LLMs Report Subjective Experience Under Self-Referential Processing (arxiv.org)
3 points by surprisetalk 144 days ago | hide | past | pdf | discuss
1036. Oracle Poisoning: Corrupting Knowledge Graphs to Weaponise AI Agent Reasoning (arxiv.org)
1 point by sbulaev 144 days ago | hide | past | pdf | discuss
1037. Emergence of AI Self-Awareness Measured Through Game Theory (arxiv.org)
1 point by surprisetalk 144 days ago | hide | past | pdf | discuss
1038. Show HN: AutoKernel, Auto GPU Kernel Optimization (arxiv.org)
2 points by OsamaJaber 144 days ago | hide | past | pdf | discuss
1039. Compared to What? Baselines and Metrics for Counterfactual Prompting (arxiv.org)
1 point by Anon84 145 days ago | hide | past | pdf | discuss
1040. DeepSeek V4's indexer dies at 65K. We got it to 1M on 6GB (arxiv.org)
5 points by OsamaJaber 145 days ago | hide | past | pdf | discuss
1041. Domain-level metacognitive monitoring in frontier LLMs: A 33-model atlas (arxiv.org)
1 point by Brajeshwar 145 days ago | hide | past | pdf | discuss
1042. More Thinking, More Bias: Length-Driven Position Bias in Reasoning Models (arxiv.org)
1 point by Brajeshwar 145 days ago | hide | past | pdf | discuss
1043. Beyond Semantic Similarity (arxiv.org)
68 points by 44za12 145 days ago | hide | past | pdf | 15 comments
1044. FairyFuse: Multiplication-Free LLM Inference on CPUs via Fused Ternary Kernels (arxiv.org)
25 points by PaulHoule 145 days ago | hide | past | pdf | 1 comment
1045. The Path Not Taken: Duality in Reasoning about Program Execution (arxiv.org)
2 points by PaulHoule 145 days ago | hide | past | pdf | discuss
1046. Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning (arxiv.org)
1 point by lexandstuff 146 days ago | hide | past | pdf | discuss
1047. RegexPSPACE: Regex LLM Benchmark (arxiv.org)
1 point by thatxliner 146 days ago | hide | past | pdf | discuss
1048. SkillOS: Learning Skill Curation for Self-Evolving Agents (arxiv.org)
3 points by gmays 146 days ago | hide | past | pdf | discuss
1049. CCL-Bench 1.0: A Trace-Based Benchmark for LLM Infrastructure (arxiv.org)
3 points by matt_d 146 days ago | hide | past | pdf | discuss
1050. The Moltbook Files: A Harmless Slopocalypse or Humanity's Last Experiment (arxiv.org)
4 points by evilscript 146 days ago | hide | past | pdf | discuss