A curated list of awesome Multimodal studies.
Contribution
If you have published a high-quality paper or come across one that you think is valuable, feel free to contribute! To submit a paper, please open an issue and include the following information in the specified format:
Submission Format
{
"title": paper title,
"url": paper URL,
"venue": the venue where the paper was published, such as ICML 2025, CVPR 2025 or arXiv,
"category": one or more relevant categories from our directory, or feel free to propose a new, more suitable category,
"code": [Optional] code URL,
"project_page": [Optional] project page URL,
"dataset": [Optional] HuggingFace Dataset URL,
"collections": [Optional] HuggingFace Collections URL
}
- Awesome-Multimodal-Papers
- Foundation Model (Textual and Multimodal)
- Visual Understanding
- Omni Understanding
- Unified Understanding and Generation
- Coding / GUI
- Diffusion MLLM
- Multimodal Embedding/Retrieval
- Image Understanding Benchmark
- Video Understanding Benchmark
- Audio
- Multimodal Dialogue
- Multimodal Learning
- Image Generation
- Video Generation
- Multimodal Dataset
- Multimodal Survey
| Title | Date | Code | Supplement |
|---|---|---|---|
| [DeepSeek-V4] DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence | 2026-04 | - | |
| [GLM-5] GLM-5: from Vibe Coding to Agentic Engineering | 2026-02 | ||
| [Kimi K2.5] Kimi K2.5: Visual Agentic Intelligence | 2026-02 | ||
| [MiMo-V2-Flash] MiMo-V2-Flash Technical Report | 2026-01 | ||
| [Qwen3-VL] Qwen3-VL Technical Report | 2025-11 | ||
| [GLM-4.5] GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models | 2025-08 | ||
| [Kimi K2] Kimi K2: Open Agentic Intelligence | 2025-07 | ||
| [InternVL3] InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models | 2025-04 | ||
| [Qwen2.5-VL] Qwen2.5-VL Technical Report | 2025-01 | ||
| [InternVL2.5] Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling | 2024-12 | ||
| [Qwen2-VL] Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution | 2024-09 |
Vibe Coding / GUI / UI Design
| Title | Venue | Date | Code | Supplement |
|---|---|---|---|---|
| Muddit: Liberating Generation Beyond Text-to-Image with a Unified Discrete Diffusion Model | arXiv | 2025-05-29 | - | |
| FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities | arXiv | 2025-05-26 | - | |
| Dimple: Discrete Diffusion Multimodal Large Language Model with Parallel Decoding (NUS) | arXiv | 2025-05-22 | - | |
| LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning (Gaoling) | arXiv | 2025-05-22 | ||
| LaViDa: A Large Diffusion Language Model for Multimodal Understanding (UCLA, Panasonic AI, Salesforce, Adobe) | arXiv | 2025-05-22 | ||
| MMaDA: Multimodal Large Diffusion Language Models (ByteDance Seed) | arXiv | 2025-05-21 |
| Title | Venue | Date | Code | Supplement |
|---|---|---|---|---|
| SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities | EMNLP 2023 (Findings) | 2023-05-18 |