In plain words: It turns images, video, audio, and motion-sensor readings into text-like signals so a strong text-only language model can reason over them and reply in words. After training on hand-collected multimodal instructions, it beat the best prior systems on a range of multimodal tasks.
Abstract
We present Any-Modality Augmented Language Model (AnyMAL), a unified model that reasons over diverse input modality signals (i.e. text, image, video, audio, IMU motion sensor), and generates textual responses. AnyMAL inherits the powerful text-based reasoning abilities of the state-of-the-art LLMs including LLaMA-2 (70B), and converts modality-specific signals to the joint textual space through a pre-trained aligner module. To further strengthen the multimodal LLM's capabilities, we fine-tune the model with a multimodal instruction set manually collected to cover diverse topics and tasks beyond simple QAs. We conduct comprehensive empirical analysis comprising both human and automatic evaluations, and demonstrate state-of-the-art performance on various multimodal tasks.
Seungwhan Moon, Andrea Madotto, Zhaojiang Lin, Tushar Nagarajan, Matt Smith, Shashank Jain, Chun-Fu Yeh, Prakash Murugesan, Peyman Heidari, Yue Liu, Kavya Srinet, Babak Damavandi, et al.
arXiv:2309.16058 · cs.LG, cs.CL, cs.CV · submitted Sep 27, 2023
abstract · pdf · html
https://www.anybotics.com/robotics/anymal/