about
421. A Watermark for Large Language Models (arxiv.org)
4 points by matheusmoreira 53 days ago | hide | past | pdf | discuss
422. Towards Expert-Level Medical AI for Real-Time Video Consultations (arxiv.org)
3 points by tcp_handshaker 53 days ago | hide | past | pdf | discuss
423. Emergent Introspective Awareness in Large Language Models (arxiv.org)
72 points by doener 53 days ago | hide | past | pdf | 33 comments
424. Previous-Token Prediction Based LLM Near-Exact Prompt Reconstruction (arxiv.org)
2 points by doener 53 days ago | hide | past | pdf | discuss
425. Stealing Reasoning Traces from Proprietary LLM APIs (arxiv.org)
9 points by samvher 53 days ago | hide | past | pdf | discuss
426. ArchAgent v2: A Case Study with the Data Prefetching Championship (arxiv.org)
3 points by root-parent 54 days ago | hide | past | pdf | discuss
427. Stealing Reasoning Traces from Proprietary LLM APIs (arxiv.org)
5 points by sbulaev 54 days ago | hide | past | pdf | discuss
428. Ouroboros: A coding agent that evolves its own harness and tops agent benchmarks (arxiv.org)
5 points by ndreeew 54 days ago | hide | past | pdf | 1 comment
429. Attention-Only Transformers (arxiv.org)
3 points by MediaSquirrel 54 days ago | hide | past | pdf | 1 comment
430. A Watermark for Large Language Models (arxiv.org)
3 points by m-chrzan 54 days ago | hide | past | pdf | discuss
431. EvoHarnessRL: Learning Self-Evolving Runtime Harness for Long-Horizon LLM Agents (arxiv.org)
3 points by matt_d 54 days ago | hide | past | pdf | discuss
432. Mind the Gaps: Mixture-of-Minds for Human Simulation (arxiv.org)
2 points by jtewright 54 days ago | hide | past | pdf | 1 comment
433. HLSmith: An Expert-Guided Agentic Framework for C/C++-to-HLS Translation (arxiv.org)
2 points by rbanffy 54 days ago | hide | past | pdf | discuss
434. Composer 2 Technical Report (arxiv.org)
1 point by ronfriedhaber 55 days ago | hide | past | pdf | discuss
435. Diffusion LLMs as Targets and Adversaries: Mechanistic Safety Exploits (arxiv.org)
2 points by sbulaev 55 days ago | hide | past | pdf | discuss
436. HCCL: Collective Communication for Meta Training and Inference Accelerators (arxiv.org)
1 point by matt_d 55 days ago | hide | past | pdf | discuss
437. MatrAIx: Simulating the World with 8.3B Persona Agents (arxiv.org)
4 points by birriel 55 days ago | hide | past | pdf | 2 comments
438. LLMs Get Lost in Evolving User Intent (arxiv.org)
2 points by Anon84 55 days ago | hide | past | pdf | discuss
439. ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidence (arxiv.org)
2 points by simonpure 55 days ago | hide | past | pdf | discuss
440. Answer First, Reason Later: Commitment Order in Diffusion LLMs (arxiv.org)
3 points by sbulaev 56 days ago | hide | past | pdf | discuss
441. Benchmarking LLMs on File System Design and Implementation (arxiv.org)
4 points by matt_d 56 days ago | hide | past | pdf | discuss
442. Towards a Risk Assessment of Malicious Skill Files in Coding Agents (arxiv.org)
1 point by sbulaev 57 days ago | hide | past | pdf | discuss
443. Robust AI Security and Alignment: A Sisyphean Endeavor? (arxiv.org)
3 points by joshcsimmons 57 days ago | hide | past | pdf | discuss
444. TensorLift: Auto Extraction of ISA Semantics from Accelerator RTL via MLIR (arxiv.org)
1 point by matt_d 57 days ago | hide | past | pdf | discuss
445. HarnessOpt-Bench: Evaluating LLMs at Harness Optimization (arxiv.org)
3 points by wslh 57 days ago | hide | past | pdf | discuss
446. AI Meet CAD: Beat Opus/Mythos 5, GPT 5.6 Sol on BenchCAD (arxiv.org)
2 points by SUPERustam 58 days ago | hide | past | pdf | 1 comment
447. The Illusion of Visual Tool-Use: A Causal Audit of Thinking with Images (arxiv.org)
1 point by sbulaev 58 days ago | hide | past | pdf | discuss
448. INT2 KV-cache quantization is getting surprisingly good (arxiv.org)
3 points by yunuyean 58 days ago | hide | past | pdf | discuss
449. SparSEEty: Extracting Tokens from Sparsity-Exploiting LLM Serving Systems (arxiv.org)
1 point by sbulaev 58 days ago | hide | past | pdf | discuss
450. Benchmarking the Residual: What Long-Horizon Evals Add Beyond Short-Task Perf (arxiv.org)
1 point by matt_d 58 days ago | hide | past | pdf | discuss