about
4531. Core-Bench: Computational Reproducibility Agent Benchmark (arxiv.org)
1 point by randomwalker on Sep 18, 2024 | hide | past | pdf | discuss
4532. LLM as Interpreter for Natural Language Programming (arxiv.org)
2 points by hdvr on Sep 18, 2024 | hide | past | pdf | discuss
4533. The real Q-Star paper: Critical Planning Step Learning (arxiv.org)
3 points by vegax87 on Sep 17, 2024 | hide | past | pdf | discuss
4534. Diagram of Thought (arxiv.org)
2 points by taikon on Sep 17, 2024 | hide | past | pdf | discuss
4535. Breaking ReCAPTCHAv2 (arxiv.org)
3 points by ryzvonusef on Sep 17, 2024 | hide | past | pdf | 2 comments
4536. Schrodinger's Memory: Large Language Models (arxiv.org)
2 points by Anon84 on Sep 17, 2024 | hide | past | pdf | discuss
4537. Facial Wrinkle Segmentation for Cosmetic Dermatology (arxiv.org)
2 points by surprisetalk on Sep 17, 2024 | hide | past | pdf | discuss
4538. ODYSSEE: Oyster Detection Yielded by Sensor Systems on Edge Electronics (arxiv.org)
2 points by surprisetalk on Sep 17, 2024 | hide | past | pdf | discuss
4539. Can Generative Multi-Agents Spontaneously Form a Society? (arxiv.org)
48 points by geuds on Sep 17, 2024 | hide | past | pdf | 5 comments
4540. Planning in Natural Language Improves LLM Search for Code Generation (arxiv.org)
3 points by hughzhang on Sep 17, 2024 | hide | past | pdf | discuss
4541. Trustworthiness in Retrieval-Augmented Generation Systems: A Survey (arxiv.org)
1 point by mfiguiere on Sep 17, 2024 | hide | past | pdf | discuss
4542. RAG Based Question-Answering for Contextual Response Prediction System (arxiv.org)
1 point by kungfudoi on Sep 17, 2024 | hide | past | pdf | discuss
4543. Single prompt achieves competitive results with o1-preview (arxiv.org)
3 points by saran945 on Sep 17, 2024 | hide | past | pdf | 1 comment
4544. Chain of Thought empowers transformers to solve inherently serial problems (arxiv.org)
261 points by krackers on Sep 17, 2024 | hide | past | pdf | 184 comments
4545. Are Pre-trained Convolutions Better than Pre-trained Transformers? (2021) (arxiv.org)
3 points by fzliu on Sep 16, 2024 | hide | past | pdf | discuss
4546. Towards Measuring and Modeling "Culture" in LLMs: A Survey (arxiv.org)
1 point by rntn on Sep 16, 2024 | hide | past | pdf | discuss
4547. Deep Knowledge-Infusion for Explainable Depression Detection (arxiv.org)
9 points by PaulHoule on Sep 15, 2024 | hide | past | pdf | discuss
4548. Double Descent Demystified (arxiv.org)
1 point by veryluckyxyz on Sep 15, 2024 | hide | past | pdf | discuss
4549. The Influence of Faulty Labels in Data Sets on Human Pose Estimation (arxiv.org)
1 point by PaulHoule on Sep 15, 2024 | hide | past | pdf | discuss
4550. Skeleton-of-Thought: Prompting LLMs for Efficient Parallel Generation (arxiv.org)
1 point by wseqyrku on Sep 15, 2024 | hide | past | pdf | discuss
4551. LLMs Will Always Hallucinate, and We Need to Live with This (arxiv.org)
291 points by Anon84 on Sep 14, 2024 | hide | past | pdf | 261 comments
4552. Neurosymbolic Methods for Dynamic Knowledge Graphs (arxiv.org)
2 points by PaulHoule on Sep 14, 2024 | hide | past | pdf | discuss
4553. Teaching Models to Express Their Uncertainty in Words (2022) (arxiv.org)
7 points by Bluestein on Sep 14, 2024 | hide | past | pdf | discuss
4554. Synthetic Continued Pretraining (arxiv.org)
3 points by veryluckyxyz on Sep 14, 2024 | hide | past | pdf | discuss
4555. Can We Count on LLMs? The Fixed-Effect Fallacy and Claims of GPT-4 Capabilities (arxiv.org)
1 point by belter on Sep 13, 2024 | hide | past | pdf | discuss
4556. Scaling LLM Test-Time can be More Effective than Scaling Parameters (arxiv.org)
1 point by krackers on Sep 13, 2024 | hide | past | pdf | discuss
4557. EMP: Enhance Memory in Data Pruning (arxiv.org)
7 points by PaulHoule on Sep 13, 2024 | hide | past | pdf | discuss
4558. Localization of Synthetic Manipulations in Western Blot Images (arxiv.org)
1 point by hentrep on Sep 13, 2024 | hide | past | pdf | discuss
4559. Let's Verify Step by Step (arxiv.org)
3 points by gmays on Sep 13, 2024 | hide | past | pdf | discuss
4560. Can Large Language Models Unlock Novel Scientific Research Ideas? [pdf] (arxiv.org)
1 point by SerCe on Sep 13, 2024 | hide | past | pdf | discuss