about
Stories from October 10, 2023 (UTC)
Go back a day, month, or year. Go forward a day.
1. HyperAttention: Long-Context Attention in Near-Linear Time (arxiv.org)
73 points by kelseyfrog on Oct 10, 2023 | hide | past | pdf | 13 comments
2. Outlier Weighed Layerwise Sparsity: A Missing Secret Sauce for Pruning LLMs (arxiv.org)
5 points by amilios on Oct 10, 2023 | hide | past | pdf | 1 comment
3. Fast Neural Rendering with Multi-Input Multi-Output Neural Radiance Fields (arxiv.org)
3 points by PaulHoule on Oct 10, 2023 | hide | past | pdf | discuss
4. Finetuning 3-Bit LLMs on Consumer GPUs by Integrating with Modular Quantizers (arxiv.org)
2 points by volodia on Oct 10, 2023 | hide | past | pdf | discuss
5. Just Ask for Calibration: Eliciting Calibrated Confidence Scores from LLMs (arxiv.org)
2 points by famouswaffles on Oct 10, 2023 | hide | past | pdf | discuss
6. Grokking as Compression: A Nonlinear Complexity Perspective (arxiv.org)
2 points by amilios on Oct 10, 2023 | hide | past | pdf | 1 comment
7. Why do we need weight decay in modern deep learning? (arxiv.org)
2 points by max-andr on Oct 10, 2023 | hide | past | pdf | 1 comment
8. Self-Taught Optimizer (Stop): Recursively Self-Improving Code Generation (arxiv.org)
1 point by haltist on Oct 10, 2023 | hide | past | pdf | discuss
9. Enabling Level-4 Autonomous Driving on a Single $1k Off-the-Shelf Card (arxiv.org)
1 point by robertlagrant on Oct 10, 2023 | hide | past | pdf | discuss
10. PB-LLM: Partially Binarized Large Language Models (arxiv.org)
1 point by tosh on Oct 10, 2023 | hide | past | pdf | discuss