about
5791. BlackJAX: Composable Bayesian Inference in Jax (arxiv.org)
3 points by sebg on Feb 20, 2024 | hide | past | pdf | discuss
5792. Fit: Flexible Vision Transformer for Diffusion Model (arxiv.org)
1 point by upmind on Feb 20, 2024 | hide | past | pdf | discuss
5793. RecallM: Architecture for Temporal Context Understanding and Question Answering (arxiv.org)
1 point by chuckhend on Feb 20, 2024 | hide | past | pdf | discuss
5794. Speculative Streaming: Fast LLM Inference Without Auxiliary Models (arxiv.org)
2 points by gok on Feb 20, 2024 | hide | past | pdf | 1 comment
5795. Pushing the limits of mathematical reasoning in open language models (arxiv.org)
1 point by dr_dshiv on Feb 20, 2024 | hide | past | pdf | discuss
5796. Fit: Flexible Vision Transformer for Diffusion Model (arxiv.org)
3 points by dihuang on Feb 20, 2024 | hide | past | pdf | 2 comments
5797. Personalized Language Modeling from Personalized Human Feedback (arxiv.org)
1 point by PaulHoule on Feb 19, 2024 | hide | past | pdf | discuss
5798. A Survey on Transformer Compression (arxiv.org)
1 point by PaulHoule on Feb 19, 2024 | hide | past | pdf | discuss
5799. Hydragen: High-Throughput LLM Inference with Shared Prefixes (arxiv.org)
1 point by PaulHoule on Feb 19, 2024 | hide | past | pdf | discuss
5800. The boundary of neural network trainability is fractal (arxiv.org)
200 points by RafelMri on Feb 19, 2024 | hide | past | pdf | 65 comments
5801. More Agents Is All You Need (arxiv.org)
3 points by PaulHoule on Feb 18, 2024 | hide | past | pdf | discuss
5802. ARB: Advanced Reasoning Benchmark For Large Language Models (2023) (arxiv.org)
3 points by optimalsolver on Feb 18, 2024 | hide | past | pdf | 1 comment
5803. Chain-of-Thought Reasoning Without Prompting (arxiv.org)
94 points by famouswaffles on Feb 17, 2024 | hide | past | pdf | 25 comments
5804. Bridging the Empirical-Theoretical Gap in Formal Language Learning Using MDL (arxiv.org)
1 point by bottencat on Feb 17, 2024 | hide | past | pdf | discuss
5805. A claim for Linearly scalable Attention ( Breakthrough if true) (arxiv.org)
2 points by guywithabowtie on Feb 17, 2024 | hide | past | pdf | 1 comment
5806. Counterfactual Tasks to Evaluate the Generality of Analogical Reasoning in LLMs (arxiv.org)
2 points by MAXPOOL on Feb 17, 2024 | hide | past | pdf | discuss
5807. LLM agents can autonomously hack websites (arxiv.org)
85 points by pella on Feb 16, 2024 | hide | past | pdf | 21 comments
5808. BitDelta: Your Fine-Tune May Only Be Worth One Bit (arxiv.org)
2 points by convexstrictly on Feb 16, 2024 | hide | past | pdf | 2 comments
5809. A Comprehensive Survey of 400 Activation Functions (arxiv.org)
5 points by jonbaer on Feb 16, 2024 | hide | past | pdf | discuss
5810. Training LLMs to generate text with citations via fine-grained rewards (arxiv.org)
170 points by PaulHoule on Feb 16, 2024 | hide | past | pdf | 34 comments
5811. Lumiere: A Space-Time Diffusion Model for Video Generation (arxiv.org)
5 points by fzliu on Feb 15, 2024 | hide | past | pdf | discuss
5812. Large-Scale Generative AI Models Lack Visual Number Sense (arxiv.org)
1 point by PaulHoule on Feb 15, 2024 | hide | past | pdf | 1 comment
5813. Understanding LLMs: A Comprehensive Overview from Training to Inference (arxiv.org)
2 points by pella on Feb 15, 2024 | hide | past | pdf | discuss
5814. Evaluating the Generality of Analogical Reasoning in LLMs (arxiv.org)
1 point by max_ on Feb 15, 2024 | hide | past | pdf | discuss
5815. LLM Agents can Autonomously Hack Websites (arxiv.org)
2 points by nopinsight on Feb 15, 2024 | hide | past | pdf | discuss
5816. Neural networks for abstraction and reasoning: Towards broad generalization (arxiv.org)
3 points by PaulHoule on Feb 15, 2024 | hide | past | pdf | discuss
5817. Study Reveals Gender Bias in ChatGPT Translations (arxiv.org)
2 points by thatha7777 on Feb 14, 2024 | hide | past | pdf | discuss
5818. Towards a Certified Proof Checker for Deep Neural Network Verification (arxiv.org)
1 point by bluish29 on Feb 14, 2024 | hide | past | pdf | discuss
5819. BASE TTS: Lessons from building a billion-parameter Text-to-Speech model (arxiv.org)
3 points by jcuenod on Feb 14, 2024 | hide | past | pdf | discuss
5820. Disinformation Capabilities of Large Language Models (arxiv.org)
5 points by empiko on Feb 14, 2024 | hide | past | pdf | 1 comment