Visual-Text Multimodal Large Language Models | Views : 0 下载量: 1754 CSCD: 0
  • Export

  • Share

  • Collection

  • Album

    • Multimodal large model-based method for generating visual Q&A data for electronic document images

    • Vol. 30, Issue 9, Pages: 3083-3096(2025)   

      Received:16 October 2024

      Revised:2025-02-16

      Accepted:25 February 2025

      Online First:26 February 2025

      Published:16 September 2025

    • DOI: 10.11834/jig.240610     

    移动端阅览

  • Li Yuzhe, Fu Ling, Zhu Linghao, Luo Qidi, Tu Lai. 2025. Multimodal large model-based method for generating visual Q&A data for electronic document images. Journal of Image and Graphics, 30(9):3083-3096 DOI: 10.11834/jig.240610.
  •  
  •  
Alert me when the article has been cited
提交

相关作者

Yang Yi 浙江大学脑机智能全国重点实验室;浙江大学人工智能学院
Wang Jun OPPO研究院
Wang Wenguan 浙江大学脑机智能全国重点实验室;浙江大学人工智能学院
Liu Rui 浙江大学脑机智能全国重点实验室;浙江大学人工智能学院
Yin Wenti 华中科技大学人工智能与自动化学院
Wu Meiqi 中国科学院自动化研究所
Wang Xiang 华中科技大学人工智能与自动化学院
Tan Chuangchuang 北京交通大学计算机 科学与技术学院

相关机构

The State Key Lab of Brain-Machine Intelligence, Zhejiang University
College of Artificial Intelligence, Zhejiang University
OPPO Research Institute
Institute of Automation, Chinese Academy of Sciences
School of Computer Science and Technology, Beijing Jiaotong University
0