about
3631. Fire-Flyer AI-HPC: Cost-Effective Software-Hardware Co-Design for Deep Learning (arxiv.org)
2 points by doener on Feb 21, 2025 | hide | past | pdf | discuss
3632. Performance of Zero-Shot Time Series Foundation Models on Cloud Data (arxiv.org)
4 points by wanderingmind on Feb 21, 2025 | hide | past | pdf | discuss
3633. Presumed Cultural Identity: How Names Shape LLM Responses (arxiv.org)
1 point by SerCe on Feb 21, 2025 | hide | past | pdf | discuss
3634. Idiosyncrasies in Large Language Models (arxiv.org)
2 points by mfiguiere on Feb 20, 2025 | hide | past | pdf | discuss
3635. AI Alignment at Your Discretion (arxiv.org)
3 points by maartenbuyl on Feb 20, 2025 | hide | past | pdf | discuss
3636. ReLU Strikes Back: Exploiting Activation Sparsity in Large Language Models(2023) (arxiv.org)
2 points by martinloretz on Feb 20, 2025 | hide | past | pdf | 1 comment
3637. WonderHuman: 3D avatars from single-view video (arxiv.org)
38 points by jinqueeny on Feb 20, 2025 | hide | past | pdf | 6 comments
3638. Aide: AI-Driven Exploration in the Space of Code (Arxiv) (arxiv.org)
1 point by WecoAI on Feb 19, 2025 | hide | past | pdf | 1 comment
3639. Native Sparse Attention: Hardware-Aligned and Natively Trainable (arxiv.org)
2 points by teepo on Feb 19, 2025 | hide | past | pdf | discuss
3640. STT: Stateful Tracking with Transformers for Autonomous Driving [pdf] (arxiv.org)
1 point by lawrenceyan on Feb 19, 2025 | hide | past | pdf | discuss
3641. Karatsuba Matrix Multiplication and Its Efficient Hardware Implementations (arxiv.org)
2 points by emacs28 on Feb 19, 2025 | hide | past | pdf | discuss
3642. NSA: Hardware-Aligned and Natively Trainable Sparse Attention (arxiv.org)
4 points by unignorant on Feb 19, 2025 | hide | past | pdf | 2 comments
3643. Deep Lake: A Lakehouse for Deep Learning (arxiv.org)
2 points by teleforce on Feb 19, 2025 | hide | past | pdf | discuss
3644. Tails Tell Tales: Chapter-Wide Manga Transcriptions with Character Names (arxiv.org)
2 points by famouswaffles on Feb 18, 2025 | hide | past | pdf | 1 comment
3645. Tensor evolution: A framework for fast tensor computations using recurrences (arxiv.org)
53 points by matt_d on Feb 18, 2025 | hide | past | pdf | 18 comments
3646. DeepSeek Native Sparse Attention (arxiv.org)
16 points by bandwitch on Feb 18, 2025 | hide | past | pdf | 1 comment
3647. The Curse of Depth in Large Language Models (arxiv.org)
2 points by jonbaer on Feb 18, 2025 | hide | past | pdf | discuss
3648. Mamba-Shedder: Post-Transformer Compression for Efficient SSMs (arxiv.org)
1 point by fovc on Feb 18, 2025 | hide | past | pdf | discuss
3649. Native Sparse Attention: Hardware-Aligned, Natively Trainable Sparse Attention (arxiv.org)
15 points by mfiguiere on Feb 18, 2025 | hide | past | pdf | 2 comments
3650. SWE-Lancer: a benchmark of freelance software engineering tasks from Upwork (arxiv.org)
111 points by zone411 on Feb 18, 2025 | hide | past | pdf | 74 comments
3651. WonderHuman: Hallucinating Unseen Parts in Dynamic 3D Human Reconstruction (arxiv.org)
1 point by jinqueeny on Feb 18, 2025 | hide | past | pdf | discuss
3652. Zep: A Temporal Knowledge Graph Architecture for Agent Memory (arxiv.org)
1 point by TaurenHunter on Feb 18, 2025 | hide | past | pdf | discuss
3653. Pretraining on the Test Set Is All You Need (arxiv.org)
2 points by Poalopat on Feb 17, 2025 | hide | past | pdf | discuss
3654. EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges (arxiv.org)
17 points by apsec112 on Feb 17, 2025 | hide | past | pdf | discuss
3655. Matryoshka Quantization (arxiv.org)
1 point by wseqyrku on Feb 17, 2025 | hide | past | pdf | discuss
3656. Step-Video-T2V: The Practice, Challenges, and Future of Video Foundation Model (arxiv.org)
41 points by limoce on Feb 17, 2025 | hide | past | pdf | 5 comments
3657. ZeroBench: An Impossible Visual Benchmark for Contemporary LMMs (arxiv.org)
9 points by taesiri on Feb 17, 2025 | hide | past | pdf | 3 comments
3658. Large Language Diffusion Models (arxiv.org)
2 points by Philpax on Feb 17, 2025 | hide | past | pdf | discuss
3659. OpenAI: Competitive Programming with Large Reasoning Models (arxiv.org)
2 points by Anon84 on Feb 16, 2025 | hide | past | pdf | 1 comment
3660. Matryoshka Quantization (arxiv.org)
2 points by galeos on Feb 16, 2025 | hide | past | pdf | discuss