| 5611. |
Logits of API-Protected LLMs Leak Proprietary Information (arxiv.org) |
|
1 point by renonce on Mar 17, 2024 | hide | past | pdf | 1 comment
|
| 5612. |
Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference (arxiv.org) |
|
3 points by veryluckyxyz on Mar 17, 2024 | hide | past | pdf | discuss
|
| 5613. |
Mosaic: A Modular System for Assistive and Interactive Cooking (arxiv.org) |
|
2 points by rntn on Mar 16, 2024 | hide | past | pdf | discuss
|
| 5614. |
Can Large Language Models Reason and Plan? (arxiv.org) |
|
4 points by YeGoblynQueenne on Mar 16, 2024 | hide | past | pdf | 1 comment
|
| 5615. |
Open-source, language-agnostic dev environment for Neural Machine Translation (arxiv.org) |
|
3 points by PaulHoule on Mar 16, 2024 | hide | past | pdf | discuss
|
| 5616. |
Apple MM1: Methods, Analysis and Insights from Multimodal LLM Pre-Training (arxiv.org) |
|
70 points by Anon84 on Mar 16, 2024 | hide | past | pdf | 2 comments
|
| 5617. |
AutoDev: Automated AI-driven development by Microsoft (arxiv.org) |
|
163 points by saran945 on Mar 16, 2024 | hide | past | pdf | 213 comments
|
| 5618. |
MM1: Methods, Analysis and Insights from Multimodal LLM Pre-training (arxiv.org) |
|
179 points by lord_sudo on Mar 16, 2024 | hide | past | pdf | 60 comments
|
| 5619. |
Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient LMs (arxiv.org) |
|
1 point by panabee on Mar 15, 2024 | hide | past | pdf | discuss
|
| 5620. |
Wukong: Towards a Scaling Law for Large-Scale Recommendation (arxiv.org) |
|
1 point by PaulHoule on Mar 15, 2024 | hide | past | pdf | discuss
|
| 5621. |
Simple and Scalable Strategies to Continually Pre-Train Large Language Models (arxiv.org) |
|
1 point by Anon84 on Mar 15, 2024 | hide | past | pdf | discuss
|
| 5622. |
Apple MM1: Methods, Analysis and Insights from Multimodal LLM Pre-Training (arxiv.org) |
|
1 point by CharlesW on Mar 15, 2024 | hide | past | pdf | 1 comment
|
| 5623. |
Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking (arxiv.org) |
|
280 points by hackerlight on Mar 15, 2024 | hide | past | pdf | 264 comments
|
| 5624. |
MM1: Methods, Analysis and Insights from Multimodal LLM Pre-Training (arxiv.org) |
|
3 points by kmdupree on Mar 15, 2024 | hide | past | pdf | discuss
|
| 5625. |
Search-Based Optimization of LLM Learning Shots for Story Point Estimation (arxiv.org) |
|
2 points by zkirby on Mar 14, 2024 | hide | past | pdf | 1 comment
|
| 5626. |
Why Are Sensitive Functions Hard for Transformers? (arxiv.org) |
|
1 point by nyrikki on Mar 14, 2024 | hide | past | pdf | 1 comment
|
| 5627. |
Chronos: Learning the Language of Time Series (arxiv.org) |
|
6 points by ctrk on Mar 14, 2024 | hide | past | pdf | 1 comment
|
| 5628. |
Can Large Language Models Do Analytical Reasoning? (arxiv.org) |
|
2 points by PaulHoule on Mar 13, 2024 | hide | past | pdf | discuss
|
| 5629. |
Neural Exec: Learning Execution Triggers for Prompt Injection Attacks (arxiv.org) |
|
1 point by InfiniteWhisp on Mar 13, 2024 | hide | past | pdf | discuss
|
| 5630. |
Show HN: Data Interpreter: An LLM Agent for Data Science (arxiv.org) |
|
2 points by metagpt on Mar 13, 2024 | hide | past | pdf | discuss
|
| 5631. |
Algorithmic Progress in Language Models (arxiv.org) |
|
2 points by Anon84 on Mar 13, 2024 | hide | past | pdf | discuss
|
| 5632. |
Adding NVMe SSDs to Enable and Accelerate 100B Model Fine-Tuning on a Single GPU (arxiv.org) |
|
3 points by dataminer on Mar 13, 2024 | hide | past | pdf | 1 comment
|
| 5633. |
Human vs. Machine: Language Models and Wargames (arxiv.org) |
|
2 points by jonbaer on Mar 13, 2024 | hide | past | pdf | discuss
|
| 5634. |
COA-GPT: Gen AI for Course of Action Development in Military Operations (arxiv.org) |
|
3 points by thenaturalist on Mar 13, 2024 | hide | past | pdf | discuss
|
| 5635. |
Design2Code: How Far Are We from Automating Front-End Engineering? (arxiv.org) |
|
2 points by taubek on Mar 13, 2024 | hide | past | pdf | discuss
|
| 5636. |
AdaPT: Adaptive Point Cloud Transformer (arxiv.org) |
|
6 points by teleforce on Mar 13, 2024 | hide | past | pdf | discuss
|
| 5637. |
GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection (arxiv.org) |
|
2 points by mau on Mar 13, 2024 | hide | past | pdf | discuss
|
| 5638. |
Google's Gemini 1.5 Pro whitepaper (arxiv.org) |
|
2 points by MrCocotoso on Mar 13, 2024 | hide | past | pdf | 1 comment
|
| 5639. |
Random Networks are not Random Functions (arxiv.org) |
|
3 points by fzliu on Mar 12, 2024 | hide | past | pdf | discuss
|
| 5640. |
Here Comes the AI Worm (arxiv.org) |
|
5 points by jonbaer on Mar 12, 2024 | hide | past | pdf | discuss
|
| More |