about
5431. Griffin: RNN for Efficient Language Models (arxiv.org)
2 points by milliondreams on Apr 10, 2024 | hide | past | pdf | discuss
5432. Small Changes and Jailbreaks Affect Large Language Model Performance (arxiv.org)
2 points by belter on Apr 10, 2024 | hide | past | pdf | discuss
5433. Language Models Learn Rare Phenomena from Less Rare Phenomena (arxiv.org)
2 points by PaulHoule on Apr 10, 2024 | hide | past | pdf | discuss
5434. Does Transformer Interpretability Transfer to RNNs? (arxiv.org)
3 points by veryluckyxyz on Apr 10, 2024 | hide | past | pdf | discuss
5435. InternLM-XComposer2-4KHD: A Pioneering LVLM Handling Resolutions from 336 to 4K (arxiv.org)
2 points by yhzan on Apr 10, 2024 | hide | past | pdf | discuss
5436. MiniCPM: Potential of Small Language Models W Scalable Training Strategies (arxiv.org)
2 points by veryluckyxyz on Apr 10, 2024 | hide | past | pdf | discuss
5437. Eagle and Finch: RWKV with Matrix-Valued States and Dynamic Recurrence (arxiv.org)
1 point by tosh on Apr 10, 2024 | hide | past | pdf | discuss
5438. Text-to-SQL that asks the LLM to predict the result set (arxiv.org)
2 points by aazo11 on Apr 10, 2024 | hide | past | pdf | 1 comment
5439. DE-Cop: Detecting Copyrighted Content in Language Models Training Data (arxiv.org)
1 point by kristianp on Apr 10, 2024 | hide | past | pdf | discuss
5440. Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient LMs (arxiv.org)
9 points by kristianp on Apr 10, 2024 | hide | past | pdf | 1 comment
5441. A Survey on Red Teaming for Generative Models (arxiv.org)
16 points by sonabinu on Apr 10, 2024 | hide | past | pdf | discuss
5442. Eagle and Finch: RWKV with Matrix-Valued States and Dynamic Recurrence (arxiv.org)
1 point by Multiset on Apr 10, 2024 | hide | past | pdf | discuss
5443. Categorical Deep Learning: An Algebraic Theory of Architectures (arxiv.org)
3 points by milliondreams on Apr 10, 2024 | hide | past | pdf | discuss
5444. A Study on Scaling Up Multilingual News Framing Analysis (arxiv.org)
1 point by PaulHoule on Apr 9, 2024 | hide | past | pdf | discuss
5445. Training LLMs over Neurally Compressed Text (arxiv.org)
1 point by milliondreams on Apr 9, 2024 | hide | past | pdf | discuss
5446. Evaluating faithfulness and content selection of LLMs in book-length summaries (arxiv.org)
71 points by passwordoops on Apr 9, 2024 | hide | past | pdf | 7 comments
5447. Griffin: Mixing Gated Linear Recurrences with Local Attention (arxiv.org)
6 points by tosh on Apr 9, 2024 | hide | past | pdf | discuss
5448. No "Zero-Shot" Without Exponential Data (arxiv.org)
2 points by kmdupree on Apr 9, 2024 | hide | past | pdf | discuss
5449. Quantum Circuit Optimization with AlphaTensor (arxiv.org)
1 point by lairv on Apr 9, 2024 | hide | past | pdf | discuss
5450. Chops: Chat with CustOmer Profile Systems for Customer Service with LLMs (arxiv.org)
1 point by PaulHoule on Apr 9, 2024 | hide | past | pdf | discuss
5451. Aragog: Advanced RAG Output Grading (arxiv.org)
2 points by bbzjk7 on Apr 9, 2024 | hide | past | pdf | discuss
5452. Diffusion-RWKV: Scaling RWKV-Like Architectures for Diffusion Models (arxiv.org)
2 points by tosh on Apr 9, 2024 | hide | past | pdf | discuss
5453. Apple Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs (arxiv.org)
53 points by tosh on Apr 9, 2024 | hide | past | pdf | 7 comments
5454. Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws (arxiv.org)
5 points by Jimmc414 on Apr 9, 2024 | hide | past | pdf | discuss
5455. Enhancing Efficiency in Sparse Models with Sparser Selection (arxiv.org)
2 points by PaulHoule on Apr 8, 2024 | hide | past | pdf | discuss
5456. World Model on Million-Length Video and Language with Ring Attention (arxiv.org)
2 points by shinryudbz on Apr 8, 2024 | hide | past | pdf | discuss
5457. TableLlama: Towards Open Large Generalist Models for Tables (arxiv.org)
2 points by belter on Apr 8, 2024 | hide | past | pdf | discuss
5458. Foundation Models for Time Series Analysis: A Tutorial and Survey (arxiv.org)
2 points by Anon84 on Apr 8, 2024 | hide | past | pdf | discuss
5459. The Efficiency of Convolutional Neural Networks (arxiv.org)
1 point by boldi on Apr 8, 2024 | hide | past | pdf | discuss
5460. Direct Nash Optimization: Teaching language models to self-improve (arxiv.org)
52 points by tosh on Apr 8, 2024 | hide | past | pdf | 11 comments