about
Stories from February 22, 2025 (UTC)
Go back a day, month, or year. Go forward a day.
1. BaxBench: Can LLMs Generate Correct and Secure Back Ends? (arxiv.org)
3 points by elashri on Feb 22, 2025 | hide | past | pdf | discuss
2. Evaluating LLMs Capabilities Towards Understanding Social Dynamics (arxiv.org)
1 point by rntn on Feb 22, 2025 | hide | past | pdf | discuss
3. Experiments in News Bias Detection with Pre-Trained Neural Transformers (arxiv.org)
1 point by rntn on Feb 22, 2025 | hide | past | pdf | discuss
4. NoLiMa: Long-Context Evaluation Beyond Literal Matching (arxiv.org)
1 point by hexhowells on Feb 22, 2025 | hide | past | pdf | discuss