| 5431. |
Griffin: RNN for Efficient Language Models (arxiv.org) |
|
2 points by milliondreams on Apr 10, 2024 | hide | past | pdf | discuss
|
| 5432. |
Small Changes and Jailbreaks Affect Large Language Model Performance (arxiv.org) |
|
2 points by belter on Apr 10, 2024 | hide | past | pdf | discuss
|
| 5433. |
Language Models Learn Rare Phenomena from Less Rare Phenomena (arxiv.org) |
|
2 points by PaulHoule on Apr 10, 2024 | hide | past | pdf | discuss
|
| 5434. |
Does Transformer Interpretability Transfer to RNNs? (arxiv.org) |
|
3 points by veryluckyxyz on Apr 10, 2024 | hide | past | pdf | discuss
|
| 5435. |
InternLM-XComposer2-4KHD: A Pioneering LVLM Handling Resolutions from 336 to 4K (arxiv.org) |
|
2 points by yhzan on Apr 10, 2024 | hide | past | pdf | discuss
|
| 5436. |
MiniCPM: Potential of Small Language Models W Scalable Training Strategies (arxiv.org) |
|
2 points by veryluckyxyz on Apr 10, 2024 | hide | past | pdf | discuss
|
| 5437. |
Eagle and Finch: RWKV with Matrix-Valued States and Dynamic Recurrence (arxiv.org) |
|
1 point by tosh on Apr 10, 2024 | hide | past | pdf | discuss
|
| 5438. |
Text-to-SQL that asks the LLM to predict the result set (arxiv.org) |
|
2 points by aazo11 on Apr 10, 2024 | hide | past | pdf | 1 comment
|
| 5439. |
DE-Cop: Detecting Copyrighted Content in Language Models Training Data (arxiv.org) |
|
1 point by kristianp on Apr 10, 2024 | hide | past | pdf | discuss
|
| 5440. |
Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient LMs (arxiv.org) |
|
9 points by kristianp on Apr 10, 2024 | hide | past | pdf | 1 comment
|
| 5441. |
A Survey on Red Teaming for Generative Models (arxiv.org) |
|
16 points by sonabinu on Apr 10, 2024 | hide | past | pdf | discuss
|
| 5442. |
Eagle and Finch: RWKV with Matrix-Valued States and Dynamic Recurrence (arxiv.org) |
|
1 point by Multiset on Apr 10, 2024 | hide | past | pdf | discuss
|
| 5443. |
Categorical Deep Learning: An Algebraic Theory of Architectures (arxiv.org) |
|
3 points by milliondreams on Apr 10, 2024 | hide | past | pdf | discuss
|
| 5444. |
A Study on Scaling Up Multilingual News Framing Analysis (arxiv.org) |
|
1 point by PaulHoule on Apr 9, 2024 | hide | past | pdf | discuss
|
| 5445. |
Training LLMs over Neurally Compressed Text (arxiv.org) |
|
1 point by milliondreams on Apr 9, 2024 | hide | past | pdf | discuss
|
| 5446. |
Evaluating faithfulness and content selection of LLMs in book-length summaries (arxiv.org) |
|
71 points by passwordoops on Apr 9, 2024 | hide | past | pdf | 7 comments
|
| 5447. |
Griffin: Mixing Gated Linear Recurrences with Local Attention (arxiv.org) |
|
6 points by tosh on Apr 9, 2024 | hide | past | pdf | discuss
|
| 5448. |
No "Zero-Shot" Without Exponential Data (arxiv.org) |
|
2 points by kmdupree on Apr 9, 2024 | hide | past | pdf | discuss
|
| 5449. |
Quantum Circuit Optimization with AlphaTensor (arxiv.org) |
|
1 point by lairv on Apr 9, 2024 | hide | past | pdf | discuss
|
| 5450. |
Chops: Chat with CustOmer Profile Systems for Customer Service with LLMs (arxiv.org) |
|
1 point by PaulHoule on Apr 9, 2024 | hide | past | pdf | discuss
|
| 5451. |
Aragog: Advanced RAG Output Grading (arxiv.org) |
|
2 points by bbzjk7 on Apr 9, 2024 | hide | past | pdf | discuss
|
| 5452. |
Diffusion-RWKV: Scaling RWKV-Like Architectures for Diffusion Models (arxiv.org) |
|
2 points by tosh on Apr 9, 2024 | hide | past | pdf | discuss
|
| 5453. |
Apple Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs (arxiv.org) |
|
53 points by tosh on Apr 9, 2024 | hide | past | pdf | 7 comments
|
| 5454. |
Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws (arxiv.org) |
|
5 points by Jimmc414 on Apr 9, 2024 | hide | past | pdf | discuss
|
| 5455. |
Enhancing Efficiency in Sparse Models with Sparser Selection (arxiv.org) |
|
2 points by PaulHoule on Apr 8, 2024 | hide | past | pdf | discuss
|
| 5456. |
World Model on Million-Length Video and Language with Ring Attention (arxiv.org) |
|
2 points by shinryudbz on Apr 8, 2024 | hide | past | pdf | discuss
|
| 5457. |
TableLlama: Towards Open Large Generalist Models for Tables (arxiv.org) |
|
2 points by belter on Apr 8, 2024 | hide | past | pdf | discuss
|
| 5458. |
Foundation Models for Time Series Analysis: A Tutorial and Survey (arxiv.org) |
|
2 points by Anon84 on Apr 8, 2024 | hide | past | pdf | discuss
|
| 5459. |
The Efficiency of Convolutional Neural Networks (arxiv.org) |
|
1 point by boldi on Apr 8, 2024 | hide | past | pdf | discuss
|
| 5460. |
Direct Nash Optimization: Teaching language models to self-improve (arxiv.org) |
|
52 points by tosh on Apr 8, 2024 | hide | past | pdf | 11 comments
|
| More |