about
4501. Hype, Sustainability, and the Price of the Bigger-Is-Better Paradigm in AI (arxiv.org)
6 points by lamename on Sep 25, 2024 | hide | past | pdf | discuss
4502. Ian Soboroff: Don't Use LLMs to Make Relevance Judgments (arxiv.org)
1 point by danielsgriffin on Sep 24, 2024 | hide | past | pdf | 2 comments
4503. Efficient Acquisition of Robot Cooking Skills (arxiv.org)
1 point by hdvr on Sep 24, 2024 | hide | past | pdf | discuss
4504. A Preliminary Study of O1 in Medicine: Are We Closer to an AI Doctor (arxiv.org)
2 points by hdvr on Sep 24, 2024 | hide | past | pdf | discuss
4505. The Impact of Element Ordering on LM Agent Performance (arxiv.org)
43 points by PaulHoule on Sep 24, 2024 | hide | past | pdf | 2 comments
4506. Marca: Mamba Accelerator with ReConfigurable Architecture (arxiv.org)
2 points by PaulHoule on Sep 23, 2024 | hide | past | pdf | discuss
4507. LLM-Powered Text Simulation Attack Against ID-Free Recommender Systems (arxiv.org)
1 point by PaulHoule on Sep 23, 2024 | hide | past | pdf | discuss
4508. Real-World ML Systems: A Data-Oriented Architecture Perspective Survey (2023) (arxiv.org)
2 points by rntn on Sep 23, 2024 | hide | past | pdf | discuss
4509. LLMs Still Can't Plan; Can LRMs? A Preliminary Evaluation of O1 on PlanBench (arxiv.org)
2 points by nijaar on Sep 23, 2024 | hide | past | pdf | 1 comment
4510. Quantized neural network for complex hologram generation (arxiv.org)
1 point by PaulHoule on Sep 23, 2024 | hide | past | pdf | discuss
4511. Nudge: Lightweight Non-Parametric Fine-Tuning of Embeddings for Retrieval (arxiv.org)
2 points by PaulHoule on Sep 22, 2024 | hide | past | pdf | discuss
4512. Towards Understanding Human Emotional Fluctuations with Sparse Check-In Data (arxiv.org)
1 point by PaulHoule on Sep 22, 2024 | hide | past | pdf | discuss
4513. Kan or MLP: A Fairer Comparison (arxiv.org)
1 point by lnyan on Sep 22, 2024 | hide | past | pdf | discuss
4514. Facial Recognition Technology Detects Entrepreneurs, Outperforming Human Experts (arxiv.org)
3 points by PaulHoule on Sep 22, 2024 | hide | past | pdf | 1 comment
4515. WaveletGPT: Wavelets Meet Large Language Models (arxiv.org)
4 points by hdvr on Sep 22, 2024 | hide | past | pdf | discuss
4516. Iteration of Thought: Leveraging Inner Dialogue for Autonomous LLM Reasoning (arxiv.org)
1 point by ringer007 on Sep 22, 2024 | hide | past | pdf | discuss
4517. Long Context Evaluations Beyond Haystacks via Latent Structure Queries (arxiv.org)
2 points by Shipped02 on Sep 22, 2024 | hide | past | pdf | discuss
4518. Dissociating language and thought in large language models (arxiv.org)
42 points by rntn on Sep 21, 2024 | hide | past | pdf | 4 comments
4519. A Multimodal User Embedding Provides Personalized Explanations (arxiv.org)
1 point by PaulHoule on Sep 21, 2024 | hide | past | pdf | discuss
4520. SwiGLU activation function causes instability in FP8 LLM training (arxiv.org)
10 points by LarsDu88 on Sep 21, 2024 | hide | past | pdf | 2 comments
4521. The consistent reasoning paradox of intelligence and optimal trust in AI (arxiv.org)
1 point by zhamisen on Sep 20, 2024 | hide | past | pdf | 1 comment
4522. Ranking of popular image generation AI models (incl. Flux) from 2M votes (arxiv.org)
1 point by maalber on Sep 20, 2024 | hide | past | pdf | 1 comment
4523. Training Language Models to Self-Correct via Reinforcement Learning (arxiv.org)
230 points by weirdcat on Sep 20, 2024 | hide | past | pdf | 92 comments
4524. To CoT or not to CoT? Chain-of-thought helps on math and symbolic reasoning (arxiv.org)
1 point by healthypunk on Sep 20, 2024 | hide | past | pdf | discuss
4525. Kolmogorov-Arnold Transformer (arxiv.org)
2 points by GaggiX on Sep 20, 2024 | hide | past | pdf | discuss
4526. Late Chunking: Contextual Chunk Embeddings Using Long-Context Embedding Models (arxiv.org)
20 points by mfiguiere on Sep 20, 2024 | hide | past | pdf | 3 comments
4527. Breaking ReCAPTCHAv2 (arxiv.org)
5 points by GaggiX on Sep 20, 2024 | hide | past | pdf | discuss
4528. A Primer on the Inner Workings of Transformer-Based Language Models (arxiv.org)
4 points by Anon84 on Sep 19, 2024 | hide | past | pdf | discuss
4529. A Survey on Statistical Theory of Deep Learning (arxiv.org)
1 point by sebg on Sep 19, 2024 | hide | past | pdf | discuss
4530. Photorealistic Synthetic Data for Object Detection in Urban Streetscapes (arxiv.org)
1 point by PaulHoule on Sep 18, 2024 | hide | past | pdf | discuss