| 3571. |
Flash Interpretability: Decoding Specialised Feature Neurons in LLM (arxiv.org) |
|
1 point by mococa on Mar 3, 2025 | hide | past | pdf | discuss
|
| 3572. |
Mixtera: A Data Plane for Foundation Model Training (arxiv.org) |
|
1 point by mboether on Mar 3, 2025 | hide | past | pdf | discuss
|
| 3573. |
HW-Aligned Sparse Attention Architecture for Efficient Long-Context Modeling (arxiv.org) |
|
2 points by PaulHoule on Mar 2, 2025 | hide | past | pdf | discuss
|
| 3574. |
Infinite Retrieval: Attention enhanced LLMs in long-context processing (arxiv.org) |
|
37 points by TaurenHunter on Mar 1, 2025 | hide | past | pdf | 7 comments
|
| 3575. |
NeoBERT: A Next-Generation Bert (arxiv.org) |
|
2 points by amrrs on Mar 1, 2025 | hide | past | pdf | discuss
|
| 3576. |
Why Are Web AI Agents More Vulnerable Than Standalone LLMs? (arxiv.org) |
|
1 point by saikatsg on Mar 1, 2025 | hide | past | pdf | discuss
|
| 3577. |
Chain of Draft: Thinking Faster by Writing Less (arxiv.org) |
|
1 point by amichail on Mar 1, 2025 | hide | past | pdf | discuss
|
| 3578. |
A Comprehensive Survey on Concept Erasure in Text-to-Image Diffusion Models (arxiv.org) |
|
2 points by PaulHoule on Feb 28, 2025 | hide | past | pdf | discuss
|
| 3579. |
Enhancing Frame Detection with Retrieval Augmented Generation (arxiv.org) |
|
37 points by PaulHoule on Feb 28, 2025 | hide | past | pdf | 12 comments
|
| 3580. |
GneissWeb: Preparing High Quality Data for LLMs at Scale (arxiv.org) |
|
1 point by PaulHoule on Feb 28, 2025 | hide | past | pdf | discuss
|
| 3581. |
PhantomWiki: On-Demand Datasets for Reasoning and Retrieval Evaluation (arxiv.org) |
|
1 point by belter on Feb 28, 2025 | hide | past | pdf | discuss
|
| 3582. |
Towards an AI Co-Scientist (arxiv.org) |
|
47 points by Anon84 on Feb 28, 2025 | hide | past | pdf | 17 comments
|
| 3583. |
Increasing Transformer Context Length with Sparse Graph Processing Techniques (arxiv.org) |
|
1 point by PaulHoule on Feb 28, 2025 | hide | past | pdf | discuss
|
| 3584. |
Strassen Multisystolic Array Hardware Architectures (arxiv.org) |
|
1 point by emacs28 on Feb 28, 2025 | hide | past | pdf | discuss
|
| 3585. |
System prompts in LLMs do not reliably override user prompts (arxiv.org) |
|
1 point by macleginn on Feb 28, 2025 | hide | past | pdf | discuss
|
| 3586. |
Belief State Transformer (arxiv.org) |
|
1 point by jonbaer on Feb 28, 2025 | hide | past | pdf | discuss
|
| 3587. |
Chain of Draft: Thinking Faster by Writing Less (arxiv.org) |
|
2 points by oleg_tarasov on Feb 28, 2025 | hide | past | pdf | discuss
|
| 3588. |
Querying Databases with Function Calling (arxiv.org) |
|
2 points by PaulHoule on Feb 27, 2025 | hide | past | pdf | discuss
|
| 3589. |
Implicit Language Models Are RNNs: Balancing Parallelization and Expressivity (arxiv.org) |
|
1 point by PaulHoule on Feb 27, 2025 | hide | past | pdf | discuss
|
| 3590. |
Kitsune: Enabling Dataflow Execution on GPUs (arxiv.org) |
|
2 points by matt_d on Feb 27, 2025 | hide | past | pdf | discuss
|
| 3591. |
Contrastive Learning for Cold Start Recommendation with Adaptive Feature Fusion (arxiv.org) |
|
1 point by PaulHoule on Feb 27, 2025 | hide | past | pdf | discuss
|
| 3592. |
Bitnet.cpp: Efficient Inference for 1.58bit LLMs (arxiv.org) |
|
1 point by galeos on Feb 27, 2025 | hide | past | pdf | discuss
|
| 3593. |
(Mis)Fitting: A Survey of Scaling Laws (arxiv.org) |
|
2 points by bearseascape on Feb 27, 2025 | hide | past | pdf | discuss
|
| 3594. |
Diffusion LLM Has Arrived (arxiv.org) |
|
2 points by ironbound on Feb 27, 2025 | hide | past | pdf | discuss
|
| 3595. |
Prompt-to-Leaderboard (arxiv.org) |
|
1 point by CrypticShift on Feb 26, 2025 | hide | past | pdf | discuss
|
| 3596. |
Dataflow Execution in GPU (arxiv.org) |
|
3 points by ban-lan-gen on Feb 26, 2025 | hide | past | pdf | discuss
|
| 3597. |
Improving Consistency in Large Language Models Through Chain of Guidance (arxiv.org) |
|
11 points by Anon84 on Feb 26, 2025 | hide | past | pdf | discuss
|
| 3598. |
Fractal Generative Models (arxiv.org) |
|
2 points by lnyan on Feb 26, 2025 | hide | past | pdf | discuss
|
| 3599. |
MITRE's Offensive Security Evaluation Framework for LLMs (arxiv.org) |
|
5 points by fourierslide on Feb 26, 2025 | hide | past | pdf | 1 comment
|
| 3600. |
Twilight: Adaptive Attention Sparsity with Hierarchical Top-$P$ Pruning (arxiv.org) |
|
2 points by PaulHoule on Feb 26, 2025 | hide | past | pdf | discuss
|
| More |