The Revolution of Multimodal Large Language Models: A Survey (2402.12451v2)

Published 19 Feb 2024 in cs.CV, cs.AI, cs.CL, and cs.MM

Abstract: Connecting text and visual modalities plays an essential role in generative intelligence. For this reason, inspired by the success of LLMs, significant research efforts are being devoted to the development of Multimodal LLMs (MLLMs). These models can seamlessly integrate visual and textual modalities, while providing a dialogue-based interface and instruction-following capabilities. In this paper, we provide a comprehensive review of recent visual-based MLLMs, analyzing their architectural choices, multimodal alignment strategies, and training techniques. We also conduct a detailed analysis of these models across a wide range of tasks, including visual grounding, image generation and editing, visual understanding, and domain-specific applications. Additionally, we compile and describe training datasets and evaluation benchmarks, conducting comparisons among existing models in terms of performance and computational requirements. Overall, this survey offers a comprehensive overview of the current state of the art, laying the groundwork for future MLLMs.

References (219)

Citations (13)

View on Semantic Scholar

Summary

We haven't generated a summary for this paper yet.

Summarize Now

Tweets

https://twitter.com/_reachsumit/status/1760202566473068721

https://twitter.com/morris_phd/status/1760683694241812767

https://twitter.com/gastronomy/status/1760169660236886339

https://twitter.com/Umberto_Senpai/status/1760388714969092540

YouTube

Show All Videos

The Revolution of Multimodal Large Language Models: A Survey (2402.12451v2)

Summary

Related Papers

Tweets

YouTube