about
4201. Efficient Machine Translation with a BiLSTM-Attention Approach (arxiv.org)
1 point by PaulHoule on Nov 13, 2024 | hide | past | pdf | discuss
4202. Sparse experts with automatic token-level specialization to model time series (arxiv.org)
1 point by TaurenHunter on Nov 13, 2024 | hide | past | pdf | discuss
4203. In Search of Forgotten Domain Generalization (arxiv.org)
1 point by fzliu on Nov 13, 2024 | hide | past | pdf | discuss
4204. Language models are few-shot learners (2020) (arxiv.org)
1 point by gone35 on Nov 12, 2024 | hide | past | pdf | discuss
4205. Helmet: How to Evaluate Long-Context Language Models Effectively and Thoroughly (arxiv.org)
2 points by nopinsight on Nov 12, 2024 | hide | past | pdf | discuss
4206. Hello SME! Generating Fast Matrix Multiplication Kernels Using the SME (arxiv.org)
1 point by matt_d on Nov 12, 2024 | hide | past | pdf | discuss
4207. A Survey of Explainable AI in Financial Forecasting (arxiv.org)
1 point by nickpsecurity on Nov 12, 2024 | hide | past | pdf | 1 comment
4208. Convolutional Differentiable Logic Gate Networks (arxiv.org)
3 points by simonpure on Nov 12, 2024 | hide | past | pdf | discuss
4209. The Super Weight in Large Language Models (arxiv.org)
6 points by georgehill on Nov 12, 2024 | hide | past | pdf | discuss
4210. Tiny Transformers Excel at Sentence Compression (arxiv.org)
2 points by PaulHoule on Nov 11, 2024 | hide | past | pdf | discuss
4211. Hunyuan-Large: An Open-Source Moe Model with 52B Activated Parameters (arxiv.org)
7 points by tu7001 on Nov 11, 2024 | hide | past | pdf | 1 comment
4212. Qwen2.5-Coder Technical Report (arxiv.org)
4 points by timbilt on Nov 11, 2024 | hide | past | pdf | discuss
4213. Combining Induction and Transduction for Abstract Reasoning (arxiv.org)
2 points by lnyan on Nov 11, 2024 | hide | past | pdf | discuss
4214. LLM outputs explained using Game Theory (arxiv.org)
1 point by behnamoh on Nov 10, 2024 | hide | past | pdf | discuss
4215. The Multiple Dimensions of Spuriousness in Machine Learning (arxiv.org)
5 points by sjb326 on Nov 10, 2024 | hide | past | pdf | 1 comment
4216. Building, Reusing, Generalizing Abstract Representations from Concrete Sequences (arxiv.org)
2 points by PaulHoule on Nov 9, 2024 | hide | past | pdf | discuss
4217. Rethinking Code Refinement: Learning to Judge Code Efficiency (arxiv.org)
2 points by PaulHoule on Nov 9, 2024 | hide | past | pdf | 1 comment
4218. Personalization of Large Language Models: A Survey (arxiv.org)
1 point by PaulHoule on Nov 9, 2024 | hide | past | pdf | discuss
4219. An Efficient and Effective Approach for Repairing Programming Assignments (arxiv.org)
1 point by PaulHoule on Nov 9, 2024 | hide | past | pdf | discuss
4220. GPT-4o reads the mind in the eyes (arxiv.org)
2 points by PaulHoule on Nov 9, 2024 | hide | past | pdf | discuss
4221. Neural spell-checker: Beyond words with synthetic data generation (arxiv.org)
1 point by PaulHoule on Nov 9, 2024 | hide | past | pdf | discuss
4222. Less Is More: DocString Compression in Code Generation (arxiv.org)
1 point by PaulHoule on Nov 9, 2024 | hide | past | pdf | discuss
4223. Smaller Large Language Models Can Do Moral Self-Correction (arxiv.org)
2 points by PaulHoule on Nov 9, 2024 | hide | past | pdf | discuss
4224. BitNet a4.8: 4-bit Activations for 1-bit LLMs (arxiv.org)
3 points by ivanfioravanti on Nov 9, 2024 | hide | past | pdf | discuss
4225. Adopt: Modified Adam Can Converge with Any $β_$2 with the Optimal Rate (arxiv.org)
2 points by tosh on Nov 8, 2024 | hide | past | pdf | discuss
4226. Distinguishing Ignorance from Error in LLM Hallucinations (arxiv.org)
2 points by Anon84 on Nov 8, 2024 | hide | past | pdf | discuss
4227. LoRA vs. Full Fine-Tuning: An Illusion of Equivalence (arxiv.org)
236 points by timbilt on Nov 8, 2024 | hide | past | pdf | 53 comments
4228. AI Knowledge and Reasoning: Emulating Expert Creativity in Scientific Research (arxiv.org)
1 point by sinsentidos on Nov 8, 2024 | hide | past | pdf | 2 comments
4229. Small Language Models: Techniques, Enhancements, Applications, Trustworthiness (arxiv.org)
3 points by ignoramous on Nov 8, 2024 | hide | past | pdf | discuss
4230. Adaptive Length Image Tokenization via Recurrent Allocation (arxiv.org)
1 point by fzliu on Nov 8, 2024 | hide | past | pdf | discuss