about
3991. Tokenisation Is NP-Complete (arxiv.org)
117 points by belter on Dec 22, 2024 | hide | past | pdf | 24 comments
3992. Specification-Driven Code Translation Powered by LLMs: How Far Are We? (arxiv.org)
4 points by PaulHoule on Dec 22, 2024 | hide | past | pdf | discuss
3993. Syzygy: Dual Code-Test C to Rust Translation Using LLMs and Dynamic Analysis (arxiv.org)
4 points by slimshetty on Dec 22, 2024 | hide | past | pdf | discuss
3994. Chatting with Logs: An Exploratory Study on Finetuning LLMs for LogQL (arxiv.org)
1 point by PaulHoule on Dec 21, 2024 | hide | past | pdf | discuss
3995. Relational inductive biases, deep learning, and graph networks (arxiv.org)
2 points by sonabinu on Dec 21, 2024 | hide | past | pdf | discuss
3996. On the Measure of Intelligence (arxiv.org)
2 points by liuliu on Dec 21, 2024 | hide | past | pdf | discuss
3997. Posterior Mean Matching: Generative Modeling Through Online Bayesian Inference (arxiv.org)
3 points by gradientgoose on Dec 20, 2024 | hide | past | pdf | discuss
3998. Apollo: SGD-like Memory, AdamW-level Performance (arxiv.org)
2 points by PaulHoule on Dec 20, 2024 | hide | past | pdf | discuss
3999. SpikeFI: A Fault Injection Framework for Spiking Neural Networks (arxiv.org)
1 point by PaulHoule on Dec 20, 2024 | hide | past | pdf | discuss
4000. Video Representation Learning with Joint-Embedding Predictive Architectures (arxiv.org)
2 points by fofoz on Dec 20, 2024 | hide | past | pdf | discuss
4001. Monolith: Real Time Recommendation System with Collisionless Embedding Table (arxiv.org)
2 points by mpweiher on Dec 20, 2024 | hide | past | pdf | discuss
4002. Glider: Small model beats GPT on eval tasks (arxiv.org)
2 points by rebeccatqian on Dec 19, 2024 | hide | past | pdf | discuss
4003. Etalumis: Bringing Probabilistic Programming to Scientific Simulators at Scale (arxiv.org)
3 points by Anon84 on Dec 19, 2024 | hide | past | pdf | discuss
4004. ModernBERT (arxiv.org)
3 points by fzliu on Dec 19, 2024 | hide | past | pdf | 1 comment
4005. Lightweight Safety Classification Using Pruned Language Models (arxiv.org)
19 points by sandijean90 on Dec 19, 2024 | hide | past | pdf | 3 comments
4006. Study: Better Large Language Models Process Text More Like Human Brains Do (arxiv.org)
2 points by nopinsight on Dec 19, 2024 | hide | past | pdf | discuss
4007. Design choices made by LLM-based test generators prevent them from finding bugs (arxiv.org)
1 point by ingve on Dec 19, 2024 | hide | past | pdf | discuss
4008. Pattern Matching in AI Compilers and Its Formalization (Extended Version) (arxiv.org)
2 points by matt_d on Dec 19, 2024 | hide | past | pdf | discuss
4009. Rethinking the Combination of Graph Neural Network and Large Language Model (arxiv.org)
2 points by PaulHoule on Dec 18, 2024 | hide | past | pdf | discuss
4010. Selfish Evolution: Making Discoveries in Extreme Label Noise (arxiv.org)
2 points by PaulHoule on Dec 18, 2024 | hide | past | pdf | discuss
4011. Cultural Evolution of Cooperation Among LLM Agents (arxiv.org)
246 points by Anon84 on Dec 18, 2024 | hide | past | pdf | 131 comments
4012. Meshtron: High-fidelity 3D mesh generation from point clouds (arxiv.org)
3 points by werediver on Dec 18, 2024 | hide | past | pdf | 1 comment
4013. No More Adam: Learning Rate Scaling at Initialization Is All You Need (arxiv.org)
91 points by jinqueeny on Dec 18, 2024 | hide | past | pdf | 28 comments
4014. Rethinking Emotion Annotations in the Era of Large Language Models (arxiv.org)
1 point by PaulHoule on Dec 18, 2024 | hide | past | pdf | discuss
4015. GuardSplat: Efficient and Robust Watermarking for 3D Gaussian Splatting (arxiv.org)
1 point by PaulHoule on Dec 17, 2024 | hide | past | pdf | discuss
4016. Superhuman performance of an LLM on the reasoning tasks of a physician (arxiv.org)
4 points by wumeow on Dec 17, 2024 | hide | past | pdf | discuss
4017. NewsEdits 2.0: Learning the Intentions Behind Updating News (arxiv.org)
1 point by PaulHoule on Dec 17, 2024 | hide | past | pdf | discuss
4018. Undecidability of Underfitting in Learning Algorithms (arxiv.org)
2 points by pizza on Dec 17, 2024 | hide | past | pdf | discuss
4019. Meta's new Video Understanding Multimodal Model used Qwen model for training (arxiv.org)
7 points by BUFU on Dec 16, 2024 | hide | past | pdf | 1 comment
4020. DeepSeek-VL2 (arxiv.org)
1 point by omarsar on Dec 16, 2024 | hide | past | pdf | discuss