MIT

Awesome-RL-based-Reasoning-MLLMs

This repository provides valuable reference for researchers in the field of multimodality, please start your exploratory travel in RL-based Reasoning MLLMs!

S

Sun-Haoyuan23

Dernière activité 2 août 2026
Sun-Haoyuan23/Awesome-RL-based-Reasoning-MLLMs

1,4 k

étoiles

72

forks

11

issues ouvertes

Ce README est souvent en anglais.

Awesome RL-based Reasoning MLLMs

License: MIT Awesome

Recent advancements in leveraging reinforcement learning to enhance LLM reasoning capabilities have yielded remarkably promising results, exemplified by DeepSeek-R1, Kimi k1.5, OpenAI o3-mini, Grok 3. These exhilarating achievements herald ascendance of Large Reasoning Models, making us advance further along the thorny path towards Artificial General Intelligence (AGI). Study of LLM reasoning has garnered significant attention within the community, and researchers have concurrently summarized Awesome RL-based LLM Reasoning. Recently, researchers have also compiled a collection of some projects with detailed configurations about Large Reasoning Models in Awesome RL Reasoning Recipes ("Triple R"). Meanwhile, we have observed that remarkably awesome work has already been done in the domain of RL-based Reasoning Multimodal Large Language Models (MLLMs). We aim to provide the community with a comprehensive and timely synthesis of this fascinating and promising field, as well as some insights into it.

"The senses are the organs by which man perceives the world, and the soul acts through them as through tools."
— Leonardo da Vinci

This repository provides valuable reference for researchers in the field of multimodality, please start your exploratory travel in RL-based Reasoning MLLMs!

News

🔥🔥🔥[2025-5-24] We write the position paper Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models that summarizes recent advancements on the topic of RFT for MLLMs. We focus on answering the following three questions: 1. What background should researchers interested in this field know? 2. What has the community done? 3. What could the community do next? We hope that this position paper will provide valuable insights to the community at this pivotal stage in the advancement toward AGI.

📧📧📧[2025-4-10] Based on existing work in the community, we provide some insights into this field, which you can find in the PowerPoint presentation file.

image

Figure 1: An overview of the works done on reinforcement fine-tuning (RFT) for multimodal large language models (MLLMs). Works are sorted by release time and are collected up to May 15, 2025.

Papers (Sort by Time of Release)📄

Vision (Image)👀

Vision (Video)📹

Medical Vision🏥

Embodied Vision🤖

Multimodal Reward Model 💯

Audio👂

Omni☺️

GUI Agent📲

Web Agent🌏

Autonomous Driving🚙

3D & Metaverse🌠

Benchmarks and Datasets📊

Open-Source Projects🌐

Training Framework 🗼

  • EasyR1 💻 EasyR1 (An Efficient, Scalable, Multi-Modality RL Training Framework)

  • VeRL-Omni 💻 VeRL-Omni (Easy, fast, and stable RL training for diffusion and omni-modality models) [Docs 🌐]

  • Agent-R1 💻 Agent-R1 (A flexible RL training framework that supports training Agents with Multimodal LLM backbones.)

Vision (Image) 👀

Vision (Video)📹

Agent 👥

Contribution and Acknowledgment❤️

This is an active repository and your contributions are always welcome! If you have any question about this opinionated list, do not hesitate to contact me sun-hy23@mails.tsinghua.edu.cn.

I extend my sincere gratitude to all community members who provided valuable supplementary support.

Citation📑

If you find this repository useful for your research and applications, please star us ⭐ and consider citing:

@misc{sun2025reinforcementfinetuningpowersreasoning,
      title={Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models}, 
      author={Haoyuan Sun and Jiaqi Wu and Bo Xia and Yifu Luo and Yifei Zhao and Kai Qin and Xufei Lv and Tiantian Zhang and Yongzhe Chang and Xueqian Wang},
      year={2025},
      eprint={2505.18536},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2505.18536}, 
}

and

@misc{sun2025RL-Reasoning-MLLMs,
  title={Awesome RL-based Reasoning MLLMs},
  author={Haoyuan Sun, Xueqian Wang},
  year={2025},
  howpublished={\url{https://github.com/Sun-Haoyuan23/Awesome-RL-based-Reasoning-MLLMs}},
  note={Github Repository},
}

Star Chart⭐

Star History Chart

Projets similaires

A comprehensive list of papers using large language/multi-modal models for Robotics/RL, including papers, codes, and related websites

GGT-RIPL
4,5 k étoiles342

Latest Advances on Multimodal Large Language Models

chain-of-thoughtin-context-learninginstruction-following
BBradyFU
18 k étoiles1,1 k

A Survey of Reinforcement Learning for Large Reasoning Models

TeXawesome-listdeepseek-r1llm
TTsinghuaC3I
2,5 k étoiles134