about
5161. Chameleon: Mixed-Modal Early-Fusion Foundation Models (arxiv.org)
1 point by rando_person_1 on May 20, 2024 | hide | past | pdf | discuss
5162. Minimum Description Length Recurrent Neural Networks (2022) (arxiv.org)
1 point by puttycat on May 19, 2024 | hide | past | pdf | discuss
5163. Exploring the Adversarial Robustness of Multimodal Large Language Models (arxiv.org)
1 point by belter on May 18, 2024 | hide | past | pdf | 1 comment
5164. Mitigating Exaggerated Safety in Large Language Models (arxiv.org)
1 point by PaulHoule on May 18, 2024 | hide | past | pdf | discuss
5165. Large Language Models Show Human-Like Social Desirability Biases (arxiv.org)
1 point by PaulHoule on May 18, 2024 | hide | past | pdf | discuss
5166. MarkLLM: An Open-Source Toolkit for LLM Watermarking (arxiv.org)
2 points by belter on May 18, 2024 | hide | past | pdf | discuss
5167. Learning quantum Hamiltonians at any temperature in polynomial time (arxiv.org)
1 point by westurner on May 18, 2024 | hide | past | pdf | 1 comment
5168. ReAct: Synergizing Reasoning and Acting in Language Models (2022) (arxiv.org)
1 point by tosh on May 17, 2024 | hide | past | pdf | discuss
5169. Arctic-Embed: Scalable, Efficient, and Accurate Text Embedding Models (arxiv.org)
3 points by PaulHoule on May 17, 2024 | hide | past | pdf | discuss
5170. Unveiling Covert Harms and Social Threats in LLM Generated Conversations (arxiv.org)
2 points by PaulHoule on May 17, 2024 | hide | past | pdf | 1 comment
5171. Sakuga-42M Dataset: Scaling Up Cartoon Research (arxiv.org)
85 points by snats on May 17, 2024 | hide | past | pdf | 31 comments
5172. How Far Are We from AGI (arxiv.org)
14 points by belter on May 17, 2024 | hide | past | pdf | 7 comments
5173. LoRA Learns Less and Forgets Less (arxiv.org)
177 points by wolecki on May 17, 2024 | hide | past | pdf | 60 comments
5174. HMT: Hierarchical Memory Transformer for Long Context Language Processing (arxiv.org)
87 points by jasondavies on May 17, 2024 | hide | past | pdf | 6 comments
5175. Chameleon: Mixed-Modal Early-Fusion Foundation Models (arxiv.org)
6 points by jasondavies on May 17, 2024 | hide | past | pdf | discuss
5176. Harnessing Deep Metric Learning to Circumvent Video Streaming Encryption (arxiv.org)
3 points by Hard_Space on May 17, 2024 | hide | past | pdf | discuss
5177. Moment: A Family of Open Time-Series Foundation Models (arxiv.org)
55 points by sarusso on May 17, 2024 | hide | past | pdf | 5 comments
5178. Chameleon: Mixed-Modal Early-Fusion Foundation Models (arxiv.org)
4 points by ericflo on May 17, 2024 | hide | past | pdf | discuss
5179. Thinking Tokens for Language Modeling (arxiv.org)
6 points by Jimmc414 on May 17, 2024 | hide | past | pdf | 1 comment
5180. Special Characters Attack: Toward Scalable Training Data Extraction from LLMs (arxiv.org)
10 points by PaulHoule on May 16, 2024 | hide | past | pdf | discuss
5181. Scaling the AI Memory Wall with Dataflow and Composition of Experts (arxiv.org)
3 points by belter on May 16, 2024 | hide | past | pdf | discuss
5182. LLM4ED: Large Language Models for Automatic Equation Discovery (arxiv.org)
5 points by cscurmudgeon on May 16, 2024 | hide | past | pdf | discuss
5183. Zero-Shot Tokenizer Transfer (arxiv.org)
2 points by veryluckyxyz on May 15, 2024 | hide | past | pdf | discuss
5184. Deep learning guided Android malware and anomaly detection (arxiv.org)
1 point by nikolamilosevic on May 15, 2024 | hide | past | pdf | discuss
5185. More accurate biomedical info from GPT/Llama with optimized token usage (arxiv.org)
1 point by karthiksoman on May 15, 2024 | hide | past | pdf | discuss
5186. Mutlimodal neural networks converge to a shared statistical model of reality (arxiv.org)
34 points by ignoramous on May 15, 2024 | hide | past | pdf | 6 comments
5187. New SOTA 4-bit weight and KV quantization with optimized GPU performance (arxiv.org)
3 points by Hugsun on May 14, 2024 | hide | past | pdf | discuss
5188. Memory Mosaics (arxiv.org)
3 points by beefman on May 14, 2024 | hide | past | pdf | discuss
5189. Gemini 1.5: Multimodal understanding across tokens of context (arxiv.org)
5 points by divbzero on May 14, 2024 | hide | past | pdf | discuss
5190. An Empirical Model of Large-Batch Training (arxiv.org)
2 points by veryluckyxyz on May 14, 2024 | hide | past | pdf | discuss