| 4651. |
LLM Pruning and Distillation in Practice: The Minitron Approach (arxiv.org) |
|
2 points by kmdupree on Aug 22, 2024 | hide | past | pdf | 1 comment
|
| 4652. |
The Vizier Gaussian Process Bandit Algorithm (arxiv.org) |
|
2 points by alphabetting on Aug 22, 2024 | hide | past | pdf | discuss
|
| 4653. |
First open source Legal AI retrieval benchmark for RAG finally released (arxiv.org) |
|
9 points by ghita_ on Aug 22, 2024 | hide | past | pdf | discuss
|
| 4654. |
To Code, or Not to Code? Exploring Impact of Code in Pre-Training (arxiv.org) |
|
1 point by quxinxin on Aug 22, 2024 | hide | past | pdf | discuss
|
| 4655. |
Human-Like Episodic Memory for Infinite Context LLMs (arxiv.org) |
|
3 points by geuds on Aug 22, 2024 | hide | past | pdf | discuss
|
| 4656. |
Min P Sampling: Balancing Creativity and Coherence at High Temperature (arxiv.org) |
|
1 point by fzliu on Aug 22, 2024 | hide | past | pdf | discuss
|
| 4657. |
From pixels to planning: scale-free active inference (arxiv.org) |
|
2 points by birriel on Aug 22, 2024 | hide | past | pdf | discuss
|
| 4658. |
Predict the Next Token and Diffuse Images with One Multi-Modal Model (arxiv.org) |
|
1 point by fzliu on Aug 21, 2024 | hide | past | pdf | discuss
|
| 4659. |
To Code, or Not to Code? Exploring Impact of Code in Pre-Training (arxiv.org) |
|
1 point by tosh on Aug 21, 2024 | hide | past | pdf | discuss
|
| 4660. |
Can Large Language Models Reason? A Characterization via 3-SAT (arxiv.org) |
|
1 point by YeGoblynQueenne on Aug 21, 2024 | hide | past | pdf | discuss
|
| 4661. |
Exploring Impact of Code in Pre-Training (arxiv.org) |
|
5 points by ijk on Aug 21, 2024 | hide | past | pdf | 2 comments
|
| 4662. |
A Comparison of LLM and Human Performance on Random Number Generation Tasks[pdf] (arxiv.org) |
|
1 point by bikenaga on Aug 21, 2024 | hide | past | pdf | discuss
|
| 4663. |
Information-Theoretic Measures Reveal Grokking Is an Emergent Phase Transition (arxiv.org) |
|
2 points by puttycat on Aug 20, 2024 | hide | past | pdf | discuss
|
| 4664. |
Unlocking the Power of LSTM for Long Term Time Series Forecasting (arxiv.org) |
|
2 points by tosh on Aug 20, 2024 | hide | past | pdf | discuss
|
| 4665. |
Luna: High Accuracy Low Cost Evaluation Foundation Model to Catch Hallucinations (arxiv.org) |
|
1 point by teleforce on Aug 20, 2024 | hide | past | pdf | discuss
|
| 4666. |
Performance Law of Large Language Models (arxiv.org) |
|
1 point by minhuw on Aug 20, 2024 | hide | past | pdf | discuss
|
| 4667. |
A Robust Deep Learning Enabled Semantic Communication System for Text (2022) (arxiv.org) |
|
2 points by squircle on Aug 19, 2024 | hide | past | pdf | discuss
|
| 4668. |
JPEG-LM: LLMs as Image Generators with Canonical Codec Representations (arxiv.org) |
|
1 point by mkaic on Aug 19, 2024 | hide | past | pdf | discuss
|
| 4669. |
A Foundation Model Based on Recordings of People's Emotions and Physiology (arxiv.org) |
|
1 point by PaulHoule on Aug 19, 2024 | hide | past | pdf | discuss
|
| 4670. |
Treating the Intent Detection Problem as Dynamics in a Low-Dimensional Space (arxiv.org) |
|
1 point by PaulHoule on Aug 19, 2024 | hide | past | pdf | discuss
|
| 4671. |
Automated Design of Agentic Systems (arxiv.org) |
|
4 points by hardmaru on Aug 19, 2024 | hide | past | pdf | discuss
|
| 4672. |
MINT-1T: Open-Source Multimodal Dataset with One Trillion Tokens (arxiv.org) |
|
3 points by teleforce on Aug 19, 2024 | hide | past | pdf | discuss
|
| 4673. |
MiniCTX: Neural Theorem Proving with (Long-)Contexts (arxiv.org) |
|
3 points by PaulHoule on Aug 19, 2024 | hide | past | pdf | discuss
|
| 4674. |
LLMs consistently generate high-quality content for election disinformation (arxiv.org) |
|
3 points by edh649 on Aug 19, 2024 | hide | past | pdf | discuss
|
| 4675. |
JPEG-LM: LLMs as Image Generators with Canonical Codec Representations (arxiv.org) |
|
5 points by hardmaru on Aug 19, 2024 | hide | past | pdf | 1 comment
|
| 4676. |
Assessing the Learning Limits of LLMs with Synthetic Impossible Languages (arxiv.org) |
|
1 point by tampueroc on Aug 18, 2024 | hide | past | pdf | discuss
|
| 4677. |
The Impact of Positional Encoding on Length Generalization in Transformers (arxiv.org) |
|
2 points by cscurmudgeon on Aug 18, 2024 | hide | past | pdf | 1 comment
|
| 4678. |
Logistic Regression makes small LLMs strong "tens-of-shot" classifiers (arxiv.org) |
|
1 point by PaulHoule on Aug 18, 2024 | hide | past | pdf | discuss
|
| 4679. |
Inductive or Deductive? Rethinking the Fundamental Reasoning Abilities of LLMs (arxiv.org) |
|
1 point by PaulHoule on Aug 17, 2024 | hide | past | pdf | discuss
|
| 4680. |
Liquid Time-Constant Networks (arxiv.org) |
|
1 point by rbanffy on Aug 17, 2024 | hide | past | pdf | discuss
|
| More |