about
2161. Parse: LLM Driven Schema Optimization for Reliable Entity Extraction (arxiv.org)
2 points by PaulHoule 347 days ago | hide | past | pdf | discuss
2162. Antislop: A framework for eliminating repetitive patterns in language models (arxiv.org)
120 points by Der_Einzige 347 days ago | hide | past | pdf | 110 comments
2163. Enhancing Transformer-Based Rerankers with Synthetic Data and LLM Supervision (arxiv.org)
1 point by PaulHoule 347 days ago | hide | past | pdf | discuss
2164. CausalRAG: Integrating Causal Graphs into RAG (arxiv.org)
2 points by page_index 348 days ago | hide | past | pdf | 1 comment
2165. OpenEstimate Evaluating LLMs on Reasoning Under Uncertainty with Real-World Data (arxiv.org)
1 point by matt_d 348 days ago | hide | past | pdf | discuss
2166. The Free Transformer (arxiv.org)
6 points by aaraujo002 348 days ago | hide | past | pdf | discuss
2167. A Novel Spinor-Based Embedding Model for Transformers (arxiv.org)
4 points by haxiomic 348 days ago | hide | past | pdf | discuss
2168. Poisoning Attacks on LLMs Require a Near-Constant Number of Poison Samples (arxiv.org)
2 points by ievans 348 days ago | hide | past | pdf | discuss
2169. Can We Trust Functionally Correct Patches Generated by Code Agents? (arxiv.org)
1 point by bikenaga 348 days ago | hide | past | pdf | discuss
2170. Scaling Reinforcement Learning for Trillion-Scale Thinking Model (arxiv.org)
4 points by omarsar 348 days ago | hide | past | pdf | discuss
2171. The Dragon Hatchling: The missing link between the transformer and brain models (arxiv.org)
134 points by thatxliner 348 days ago | hide | past | pdf | 99 comments
2172. Grasp Any Region: Precise, Contextual Pixel Understanding for Multimodal LLMs (arxiv.org)
1 point by badmonster 349 days ago | hide | past | pdf | discuss
2173. A tiny spectral PMF estimator for large discrete supports (arxiv.org)
2 points by alexshtf 349 days ago | hide | past | pdf | discuss
2174. Query Decomposition for RAG (arxiv.org)
1 point by salkahfi 349 days ago | hide | past | pdf | discuss
2175. Thinking Sparks: Emergent Attention Heads in Reasoning Models (arxiv.org)
1 point by diwank 349 days ago | hide | past | pdf | discuss
2176. Large Language Models Inference Engines Based on Spiking Neural Networks (arxiv.org)
3 points by PaulHoule 349 days ago | hide | past | pdf | discuss
2177. Knowledge Transfer from High-Resource to Low-Resource Languages for Code LLMs (2023) (arxiv.org)
1 point by peatmoss 349 days ago | hide | past | pdf | discuss
2178. OptPipe: Memory- and Scheduling-Optimized Pipeline Parallelism for LLM Training (arxiv.org)
11 points by PaulHoule 349 days ago | hide | past | pdf | discuss
2179. Why can't transformers learn multiplication? (arxiv.org)
161 points by PaulHoule 349 days ago | hide | past | pdf | 107 comments
2180. Prompt Baking (arxiv.org)
8 points by jxmorris12 349 days ago | hide | past | pdf | 1 comment
2181. Binary Retrieval-Augmented Reward Mitigates Hallucinations (arxiv.org)
44 points by MarlonPro 349 days ago | hide | past | pdf | 3 comments
2182. Explaining Why Hallucinate Large Language Models (arxiv.org)
1 point by gagan30 349 days ago | hide | past | pdf | 1 comment
2183. Verifiable ML Without Determinism: Tolerance-Aware Optimistic Verification (arxiv.org)
1 point by cnyaojz 349 days ago | hide | past | pdf | 1 comment
2184. Tensor Logic: The Language of AI (arxiv.org)
5 points by fofoz 350 days ago | hide | past | pdf | discuss
2185. Reasoning with Sampling: Your Base Model Is Smarter Than You Think (arxiv.org)
3 points by Anon84 350 days ago | hide | past | pdf | discuss
2186. Evaluating Agentic Cybersecurity in Attack/Defense CTFs: Offensive Is Not Better (arxiv.org)
2 points by vmayoral 350 days ago | hide | past | pdf | 1 comment
2187. Qwen Language Confusion Gate (arxiv.org)
2 points by CollinZ 350 days ago | hide | past | pdf | discuss
2188. Reasoning with Sampling: Your Base Model Is Smarter Than You Think (arxiv.org)
2 points by kesor 350 days ago | hide | past | pdf | discuss
2189. Modeling Others' Minds as Code (arxiv.org)
66 points by PaulHoule 350 days ago | hide | past | pdf | 42 comments
2190. A Survey of Vibe Coding with Large Language Models (arxiv.org)
2 points by Anon84 351 days ago | hide | past | pdf | discuss