Multimodal Chain-of-Thought Reasoning in Language Models

Z Zhang, A Zhang, M Li, H Zhao, G Karypis, A Smola
[Shanghai Jiao Tong University & Amazon Web Services]


  1. 多模态思维链推理问题的正式研究;

  2. 提出了一种两阶段框架,用于微调语言模型以纳入视觉信号;

  3. 在ScienceQA基准上取得新的最先进性能。






Large language models (LLMs) have shown impressive performance on complex reasoning by leveraging chain-of-thought (CoT) prompting to generate intermediate reasoning chains as the rationale to infer the answer. However, existing CoT studies are mostly isolated in the language modality with LLMs, where LLMs are hard to deploy. To elicit CoT reasoning in multimodality, a possible solution is to fine-tune small language models by fusing the vision and language features to perform CoT reasoning. The key challenge is that those language models tend to generate hallucinated reasoning chains that mislead the answer inference. To mitigate the effect of such mistakes, we propose Multimodal-CoT that incorporates vision features in a decoupled training framework. The framework separates the rationale generation and answer inference into two stages. By incorporating the vision features in both stages, the model is able to generate effective rationales that contribute to answer inference. With Multimodal-CoT, our model under 1 billion parameters outperforms the previous state-of-the-art LLM (GPT-3.5) by 16% (75.17%->91.68%) on the ScienceQA benchmark and even surpasses human performance.
