about
4651. LLM Pruning and Distillation in Practice: The Minitron Approach (arxiv.org)
2 points by kmdupree on Aug 22, 2024 | hide | past | pdf | 1 comment
4652. The Vizier Gaussian Process Bandit Algorithm (arxiv.org)
2 points by alphabetting on Aug 22, 2024 | hide | past | pdf | discuss
4653. First open source Legal AI retrieval benchmark for RAG finally released (arxiv.org)
9 points by ghita_ on Aug 22, 2024 | hide | past | pdf | discuss
4654. To Code, or Not to Code? Exploring Impact of Code in Pre-Training (arxiv.org)
1 point by quxinxin on Aug 22, 2024 | hide | past | pdf | discuss
4655. Human-Like Episodic Memory for Infinite Context LLMs (arxiv.org)
3 points by geuds on Aug 22, 2024 | hide | past | pdf | discuss
4656. Min P Sampling: Balancing Creativity and Coherence at High Temperature (arxiv.org)
1 point by fzliu on Aug 22, 2024 | hide | past | pdf | discuss
4657. From pixels to planning: scale-free active inference (arxiv.org)
2 points by birriel on Aug 22, 2024 | hide | past | pdf | discuss
4658. Predict the Next Token and Diffuse Images with One Multi-Modal Model (arxiv.org)
1 point by fzliu on Aug 21, 2024 | hide | past | pdf | discuss
4659. To Code, or Not to Code? Exploring Impact of Code in Pre-Training (arxiv.org)
1 point by tosh on Aug 21, 2024 | hide | past | pdf | discuss
4660. Can Large Language Models Reason? A Characterization via 3-SAT (arxiv.org)
1 point by YeGoblynQueenne on Aug 21, 2024 | hide | past | pdf | discuss
4661. Exploring Impact of Code in Pre-Training (arxiv.org)
5 points by ijk on Aug 21, 2024 | hide | past | pdf | 2 comments
4662. A Comparison of LLM and Human Performance on Random Number Generation Tasks[pdf] (arxiv.org)
1 point by bikenaga on Aug 21, 2024 | hide | past | pdf | discuss
4663. Information-Theoretic Measures Reveal Grokking Is an Emergent Phase Transition (arxiv.org)
2 points by puttycat on Aug 20, 2024 | hide | past | pdf | discuss
4664. Unlocking the Power of LSTM for Long Term Time Series Forecasting (arxiv.org)
2 points by tosh on Aug 20, 2024 | hide | past | pdf | discuss
4665. Luna: High Accuracy Low Cost Evaluation Foundation Model to Catch Hallucinations (arxiv.org)
1 point by teleforce on Aug 20, 2024 | hide | past | pdf | discuss
4666. Performance Law of Large Language Models (arxiv.org)
1 point by minhuw on Aug 20, 2024 | hide | past | pdf | discuss
4667. A Robust Deep Learning Enabled Semantic Communication System for Text (2022) (arxiv.org)
2 points by squircle on Aug 19, 2024 | hide | past | pdf | discuss
4668. JPEG-LM: LLMs as Image Generators with Canonical Codec Representations (arxiv.org)
1 point by mkaic on Aug 19, 2024 | hide | past | pdf | discuss
4669. A Foundation Model Based on Recordings of People's Emotions and Physiology (arxiv.org)
1 point by PaulHoule on Aug 19, 2024 | hide | past | pdf | discuss
4670. Treating the Intent Detection Problem as Dynamics in a Low-Dimensional Space (arxiv.org)
1 point by PaulHoule on Aug 19, 2024 | hide | past | pdf | discuss
4671. Automated Design of Agentic Systems (arxiv.org)
4 points by hardmaru on Aug 19, 2024 | hide | past | pdf | discuss
4672. MINT-1T: Open-Source Multimodal Dataset with One Trillion Tokens (arxiv.org)
3 points by teleforce on Aug 19, 2024 | hide | past | pdf | discuss
4673. MiniCTX: Neural Theorem Proving with (Long-)Contexts (arxiv.org)
3 points by PaulHoule on Aug 19, 2024 | hide | past | pdf | discuss
4674. LLMs consistently generate high-quality content for election disinformation (arxiv.org)
3 points by edh649 on Aug 19, 2024 | hide | past | pdf | discuss
4675. JPEG-LM: LLMs as Image Generators with Canonical Codec Representations (arxiv.org)
5 points by hardmaru on Aug 19, 2024 | hide | past | pdf | 1 comment
4676. Assessing the Learning Limits of LLMs with Synthetic Impossible Languages (arxiv.org)
1 point by tampueroc on Aug 18, 2024 | hide | past | pdf | discuss
4677. The Impact of Positional Encoding on Length Generalization in Transformers (arxiv.org)
2 points by cscurmudgeon on Aug 18, 2024 | hide | past | pdf | 1 comment
4678. Logistic Regression makes small LLMs strong "tens-of-shot" classifiers (arxiv.org)
1 point by PaulHoule on Aug 18, 2024 | hide | past | pdf | discuss
4679. Inductive or Deductive? Rethinking the Fundamental Reasoning Abilities of LLMs (arxiv.org)
1 point by PaulHoule on Aug 17, 2024 | hide | past | pdf | discuss
4680. Liquid Time-Constant Networks (arxiv.org)
1 point by rbanffy on Aug 17, 2024 | hide | past | pdf | discuss