Visual-Text Multimodal Large Language Models | Views : 0 下载量: 1625 CSCD: 0
  • Export

  • Share

  • Collection

  • Album

    • Multimodal large model-based method for generating visual Q&A data for electronic document images

    • Vol. 30, Issue 9, Pages: 3083-3096(2025)   

      Received:16 October 2024

      Revised:2025-02-16

      Accepted:25 February 2025

      Online First:26 February 2025

      Published:16 September 2025

    • DOI: 10.11834/jig.240610     

    移动端阅览

  • Li Yuzhe, Fu Ling, Zhu Linghao, Luo Qidi, Tu Lai. 2025. Multimodal large model-based method for generating visual Q&A data for electronic document images. Journal of Image and Graphics, 30(9):3083-3096 DOI: 10.11834/jig.240610.
  •  
  •  
Alert me when the article has been cited
提交

相关作者

Yin Wenti 华中科技大学人工智能与自动化学院
Wu Meiqi 中国科学院自动化研究所
Wang Xiang 华中科技大学人工智能与自动化学院
Tan Chuangchuang 北京交通大学计算机 科学与技术学院
Kao Yueying 北京市科学技术研究院信息与人工智能技术研究所
Gao Changxin 华中科技大学人工智能与自动化学院
Zhao Yao 北京交通大学计算机 科学与技术学院
Huang Kaiqi 中国科学院自动化研究所

相关机构

Institute of Automation, Chinese Academy of Sciences
School of Computer Science and Technology, Beijing Jiaotong University
Institute of Information and Artificial Intelligence Technology, Beijing Academy of Science and Technology
School of Computer Science, Northwestern Polytechnical University
School of Electronic and Information Engineering, South China University of Technology
0