| 4471. |
To CoT or not? Chain-of-thought helps mainly on math and symbolic reasoning (arxiv.org) |
|
1 point by amichail on Oct 3, 2024 | hide | past | pdf | discuss
|
| 4472. |
Serving 70B-scale LLMs efficiently on low-resource edge devices [pdf] (arxiv.org) |
|
248 points by simonpure on Oct 3, 2024 | hide | past | pdf | 58 comments
|
| 4473. |
Emergent Abilities of Large Language Models (2022) (arxiv.org) |
|
1 point by squircle on Oct 3, 2024 | hide | past | pdf | discuss
|
| 4474. |
Embers of Autoregression in OpenAI O1 (arxiv.org) |
|
1 point by hdvr on Oct 3, 2024 | hide | past | pdf | discuss
|
| 4475. |
Do Large Language Models Need a Content Delivery Network? (arxiv.org) |
|
2 points by PaulHoule on Oct 3, 2024 | hide | past | pdf | discuss
|
| 4476. |
Paged KV-Cache Compression with Variable Compression Rates per Attention Head (arxiv.org) |
|
2 points by jgrahamc on Oct 2, 2024 | hide | past | pdf | discuss
|
| 4477. |
MM1.5: Methods, Analysis and Insights from Multimodal LLM Fine-Tuning (arxiv.org) |
|
5 points by dailcooper on Oct 2, 2024 | hide | past | pdf | discuss
|
| 4478. |
TPI-LLM: Serving 70B-Scale LLMs Efficiently on Low-Resource Edge Devices (arxiv.org) |
|
2 points by CrypticShift on Oct 2, 2024 | hide | past | pdf | discuss
|
| 4479. |
L-Mul: A New Algorithm for 95% AI Energy Reduction Without Performance Sacrifice (arxiv.org) |
|
3 points by kevin8704 on Oct 2, 2024 | hide | past | pdf | 1 comment
|
| 4480. |
Nvidia – NVLM: Open Frontier-Class Multimodal LLMs (arxiv.org) |
|
2 points by belter on Oct 2, 2024 | hide | past | pdf | discuss
|
| 4481. |
Vera: Validation and Enhancement for Retrieval Augmented Systems (arxiv.org) |
|
1 point by PaulHoule on Oct 1, 2024 | hide | past | pdf | discuss
|
| 4482. |
ForecastBench: A Dynamic Benchmark of AI Forecasting Capabilities (arxiv.org) |
|
2 points by dash2 on Oct 1, 2024 | hide | past | pdf | discuss
|
| 4483. |
Inducing anxiety in large language models increases exploration and bias (arxiv.org) |
|
2 points by robg on Oct 1, 2024 | hide | past | pdf | discuss
|
| 4484. |
A Comprehensive Analysis of Package Hallucinations by Code Generating LLMs (arxiv.org) |
|
31 points by rntn on Oct 1, 2024 | hide | past | pdf | 9 comments
|
| 4485. |
Results of the Big ANN: NeurIPS'23 Competition (arxiv.org) |
|
2 points by fzliu on Sep 30, 2024 | hide | past | pdf | discuss
|
| 4486. |
Efficient Streaming Inference of Multimodal Large Language Models on 1 GPU (arxiv.org) |
|
1 point by PaulHoule on Sep 29, 2024 | hide | past | pdf | discuss
|
| 4487. |
Context-Dependent Interactable UI Element Detection for Spatial Computing (arxiv.org) |
|
1 point by PaulHoule on Sep 29, 2024 | hide | past | pdf | discuss
|
| 4488. |
Explaining Data in Words: Statistical Models with Natural Language Parameters (arxiv.org) |
|
2 points by PaulHoule on Sep 27, 2024 | hide | past | pdf | discuss
|
| 4489. |
LlamaF: An Efficient Llama2 Architecture Accelerator on Embedded FPGAs (arxiv.org) |
|
124 points by PaulHoule on Sep 27, 2024 | hide | past | pdf | 29 comments
|
| 4490. |
Small Language Models: Survey, Measurements, and Insights (arxiv.org) |
|
1 point by alanzhuly on Sep 26, 2024 | hide | past | pdf | discuss
|
| 4491. |
ZeusAI: Learning to Play 7 Wonders Duel Without Human Supervision (arxiv.org) |
|
1 point by Narann on Sep 26, 2024 | hide | past | pdf | discuss
|
| 4492. |
Reconsidering the energy efficiency of spiking neural networks (arxiv.org) |
|
1 point by PaulHoule on Sep 26, 2024 | hide | past | pdf | discuss
|
| 4493. |
Chain-of-thought helps mainly on math and symbolic reasoning (arxiv.org) |
|
1 point by Anon84 on Sep 26, 2024 | hide | past | pdf | discuss
|
| 4494. |
Small Language Models: Survey, Measurements, and Insights (arxiv.org) |
|
1 point by sebg on Sep 26, 2024 | hide | past | pdf | discuss
|
| 4495. |
Zero-shot forecasting of chaotic systems (arxiv.org) |
|
3 points by sebg on Sep 26, 2024 | hide | past | pdf | discuss
|
| 4496. |
The WMDP Benchmark: Measuring and Reducing Malicious Use with Unlearning (arxiv.org) |
|
1 point by falcor84 on Sep 26, 2024 | hide | past | pdf | discuss
|
| 4497. |
Extracting Memorized Training Data via Decomposition (arxiv.org) |
|
2 points by PaulHoule on Sep 26, 2024 | hide | past | pdf | discuss
|
| 4498. |
Towards Empathetic Conversational Recommender Systems (arxiv.org) |
|
1 point by PaulHoule on Sep 25, 2024 | hide | past | pdf | discuss
|
| 4499. |
MIMO: Controllable Character Video Synthesis with Spatial Decomposed Modeling (arxiv.org) |
|
1 point by justhw on Sep 25, 2024 | hide | past | pdf | 2 comments
|
| 4500. |
From Text to Treatment Effects: A Meta-Learning Approach (arxiv.org) |
|
1 point by hdvr on Sep 25, 2024 | hide | past | pdf | discuss
|
| More |