| 5251. |
A Primer on the Inner Workings of Transformer-Based Language Models (arxiv.org) |
|
4 points by jonbaer on May 4, 2024 | hide | past | pdf | discuss
|
| 5252. |
StructLM: Towards Building Generalist Models for Structured Knowledge Grounding (arxiv.org) |
|
78 points by PaulHoule on May 3, 2024 | hide | past | pdf | discuss
|
| 5253. |
LoRA Land: 310 Fine-Tuned LLMs That Rival GPT-4, a Technical Report (arxiv.org) |
|
3 points by milliondreams on May 3, 2024 | hide | past | pdf | discuss
|
| 5254. |
Kolmogorov-Arnold Networks, an alternative to Multi-Layer Perceptrons (arxiv.org) |
|
2 points by thatxliner on May 3, 2024 | hide | past | pdf | discuss
|
| 5255. |
The Matrix: A Bayesian learning model for LLMs (arxiv.org) |
|
3 points by smaddox on May 3, 2024 | hide | past | pdf | discuss
|
| 5256. |
Network reconstruction via the minimum description length principle (arxiv.org) |
|
2 points by Anon84 on May 3, 2024 | hide | past | pdf | discuss
|
| 5257. |
Kolmogorov–Arnold Networks: Alternative to Multilayer Perceptrons. (arxiv.org) |
|
2 points by georgehill on May 3, 2024 | hide | past | pdf | 3 comments
|
| 5258. |
Examination of Large Language Model Performance on Grade School Arithmetic (arxiv.org) |
|
2 points by s-macke on May 3, 2024 | hide | past | pdf | discuss
|
| 5259. |
Evoke: Emotion Enabled Virtual Avatar Mapping Optimized Knowledge Distillation (arxiv.org) |
|
1 point by sandwichukulele on May 3, 2024 | hide | past | pdf | 1 comment
|
| 5260. |
An Examination of Large Language Model Performance on Grade School Arithmetic (arxiv.org) |
|
8 points by kmdupree on May 2, 2024 | hide | past | pdf | 1 comment
|
| 5261. |
AdvPrompter: Fast Adaptive Adversarial Prompting for LLMs (arxiv.org) |
|
2 points by zerojames on May 2, 2024 | hide | past | pdf | discuss
|
| 5262. |
Scalable Bayesian Inference in the Era of Deep Learning (arxiv.org) |
|
1 point by georgehill on May 2, 2024 | hide | past | pdf | discuss
|
| 5263. |
Capabilities of Gemini Models in Medicine (arxiv.org) |
|
1 point by tosh on May 2, 2024 | hide | past | pdf | discuss
|
| 5264. |
Alice's Adventures in a Differentiable Wonderland (arxiv.org) |
|
1 point by Schiphol on May 2, 2024 | hide | past | pdf | 2 comments
|
| 5265. |
The Feasibility of Implementing Large-Scale Transformers on Multi-FPGA Platforms (arxiv.org) |
|
3 points by PaulHoule on May 2, 2024 | hide | past | pdf | discuss
|
| 5266. |
Shape of Money Laundering: Subgraph Representation Learning on the Blockchain (arxiv.org) |
|
2 points by Anon84 on May 2, 2024 | hide | past | pdf | discuss
|
| 5267. |
Is Model Collapse Inevitable? (arxiv.org) |
|
5 points by tosh on May 2, 2024 | hide | past | pdf | 1 comment
|
| 5268. |
Kan: Kolmogorov-Arnold Networks (arxiv.org) |
|
4 points by hardmaru on May 2, 2024 | hide | past | pdf | discuss
|
| 5269. |
A Careful Examination of LLM Performance on Grade School Arithmetic (arxiv.org) |
|
2 points by GaggiX on May 2, 2024 | hide | past | pdf | discuss
|
| 5270. |
SoccerNet Game State Reconstruction: End-to-End Athlete Tracking (arxiv.org) |
|
1 point by jonbaer on May 1, 2024 | hide | past | pdf | discuss
|
| 5271. |
When can transformers reason with abstract symbols? (arxiv.org) |
|
3 points by zerojames on May 1, 2024 | hide | past | pdf | discuss
|
| 5272. |
Geo-Entity Linking for Noisy Multilingual User Input (arxiv.org) |
|
1 point by PaulHoule on May 1, 2024 | hide | past | pdf | discuss
|
| 5273. |
Kan: Kolmogorov-Arnold Networks (arxiv.org) |
|
3 points by ototot on May 1, 2024 | hide | past | pdf | 1 comment
|
| 5274. |
Iterative Reasoning Preference Optimization (arxiv.org) |
|
2 points by kmdupree on May 1, 2024 | hide | past | pdf | discuss
|
| 5275. |
Extending Llama-3's Context Ten-Fold Overnight (arxiv.org) |
|
2 points by Jimmc414 on May 1, 2024 | hide | past | pdf | 1 comment
|
| 5276. |
Machine Unlearning in Large Language Models (arxiv.org) |
|
2 points by PaulHoule on May 1, 2024 | hide | past | pdf | discuss
|
| 5277. |
Predicting SSH Keys in Open SSH Memory Dumps (arxiv.org) |
|
2 points by PaulHoule on May 1, 2024 | hide | past | pdf | discuss
|
| 5278. |
BlenderAlchemy: Editing 3D Graphics with Vision-Language Models (arxiv.org) |
|
2 points by PaulHoule on May 1, 2024 | hide | past | pdf | discuss
|
| 5279. |
Predicting Question Quality on StackOverflow with Neural Networks (arxiv.org) |
|
1 point by PaulHoule on May 1, 2024 | hide | past | pdf | discuss
|
| 5280. |
Better and Faster Large Language Models via Multi-Token Prediction (arxiv.org) |
|
302 points by jasondavies on May 1, 2024 | hide | past | pdf | 128 comments
|
| More |