about
5251. A Primer on the Inner Workings of Transformer-Based Language Models (arxiv.org)
4 points by jonbaer on May 4, 2024 | hide | past | pdf | discuss
5252. StructLM: Towards Building Generalist Models for Structured Knowledge Grounding (arxiv.org)
78 points by PaulHoule on May 3, 2024 | hide | past | pdf | discuss
5253. LoRA Land: 310 Fine-Tuned LLMs That Rival GPT-4, a Technical Report (arxiv.org)
3 points by milliondreams on May 3, 2024 | hide | past | pdf | discuss
5254. Kolmogorov-Arnold Networks, an alternative to Multi-Layer Perceptrons (arxiv.org)
2 points by thatxliner on May 3, 2024 | hide | past | pdf | discuss
5255. The Matrix: A Bayesian learning model for LLMs (arxiv.org)
3 points by smaddox on May 3, 2024 | hide | past | pdf | discuss
5256. Network reconstruction via the minimum description length principle (arxiv.org)
2 points by Anon84 on May 3, 2024 | hide | past | pdf | discuss
5257. Kolmogorov–Arnold Networks: Alternative to Multilayer Perceptrons. (arxiv.org)
2 points by georgehill on May 3, 2024 | hide | past | pdf | 3 comments
5258. Examination of Large Language Model Performance on Grade School Arithmetic (arxiv.org)
2 points by s-macke on May 3, 2024 | hide | past | pdf | discuss
5259. Evoke: Emotion Enabled Virtual Avatar Mapping Optimized Knowledge Distillation (arxiv.org)
1 point by sandwichukulele on May 3, 2024 | hide | past | pdf | 1 comment
5260. An Examination of Large Language Model Performance on Grade School Arithmetic (arxiv.org)
8 points by kmdupree on May 2, 2024 | hide | past | pdf | 1 comment
5261. AdvPrompter: Fast Adaptive Adversarial Prompting for LLMs (arxiv.org)
2 points by zerojames on May 2, 2024 | hide | past | pdf | discuss
5262. Scalable Bayesian Inference in the Era of Deep Learning (arxiv.org)
1 point by georgehill on May 2, 2024 | hide | past | pdf | discuss
5263. Capabilities of Gemini Models in Medicine (arxiv.org)
1 point by tosh on May 2, 2024 | hide | past | pdf | discuss
5264. Alice's Adventures in a Differentiable Wonderland (arxiv.org)
1 point by Schiphol on May 2, 2024 | hide | past | pdf | 2 comments
5265. The Feasibility of Implementing Large-Scale Transformers on Multi-FPGA Platforms (arxiv.org)
3 points by PaulHoule on May 2, 2024 | hide | past | pdf | discuss
5266. Shape of Money Laundering: Subgraph Representation Learning on the Blockchain (arxiv.org)
2 points by Anon84 on May 2, 2024 | hide | past | pdf | discuss
5267. Is Model Collapse Inevitable? (arxiv.org)
5 points by tosh on May 2, 2024 | hide | past | pdf | 1 comment
5268. Kan: Kolmogorov-Arnold Networks (arxiv.org)
4 points by hardmaru on May 2, 2024 | hide | past | pdf | discuss
5269. A Careful Examination of LLM Performance on Grade School Arithmetic (arxiv.org)
2 points by GaggiX on May 2, 2024 | hide | past | pdf | discuss
5270. SoccerNet Game State Reconstruction: End-to-End Athlete Tracking (arxiv.org)
1 point by jonbaer on May 1, 2024 | hide | past | pdf | discuss
5271. When can transformers reason with abstract symbols? (arxiv.org)
3 points by zerojames on May 1, 2024 | hide | past | pdf | discuss
5272. Geo-Entity Linking for Noisy Multilingual User Input (arxiv.org)
1 point by PaulHoule on May 1, 2024 | hide | past | pdf | discuss
5273. Kan: Kolmogorov-Arnold Networks (arxiv.org)
3 points by ototot on May 1, 2024 | hide | past | pdf | 1 comment
5274. Iterative Reasoning Preference Optimization (arxiv.org)
2 points by kmdupree on May 1, 2024 | hide | past | pdf | discuss
5275. Extending Llama-3's Context Ten-Fold Overnight (arxiv.org)
2 points by Jimmc414 on May 1, 2024 | hide | past | pdf | 1 comment
5276. Machine Unlearning in Large Language Models (arxiv.org)
2 points by PaulHoule on May 1, 2024 | hide | past | pdf | discuss
5277. Predicting SSH Keys in Open SSH Memory Dumps (arxiv.org)
2 points by PaulHoule on May 1, 2024 | hide | past | pdf | discuss
5278. BlenderAlchemy: Editing 3D Graphics with Vision-Language Models (arxiv.org)
2 points by PaulHoule on May 1, 2024 | hide | past | pdf | discuss
5279. Predicting Question Quality on StackOverflow with Neural Networks (arxiv.org)
1 point by PaulHoule on May 1, 2024 | hide | past | pdf | discuss
5280. Better and Faster Large Language Models via Multi-Token Prediction (arxiv.org)
302 points by jasondavies on May 1, 2024 | hide | past | pdf | 128 comments