最新刊期

    31 7 2026

      Review

    • Frontier trends and top 10 advances in 3D vision in 2025 AI导读

      Liu Yebin, Mu Yao, Ye Qi, Gao Lin, Han Xiaoguang, Chen Anpei, Duan Yueqi, Peng Sida, Shao Tianjia, Zhang Hongwen, Zhang Li, Liao Yiyi, Xu Lan, Liu Xihui, Yao Yao, Hu Ruizhen, Yi Li, Guo Yulan, Lian Zhouhui, Liu Ziwei, Chen Baoquan
      Vol. 31, Issue 7, Pages: 2279-2316(2026) DOI: 10.11834/jig.260114
      Frontier trends and top 10 advances in 3D vision in 2025
      摘要:As an interdisciplinary field that spans computer vision, graphics, artificial intelligence (AI), and optical imaging, 3D vision serves as the cornerstone for constructing embodied general intelligence and the metaverse. The “scaling law” paradigm, upon which AI development relies, faces significantly diminishing marginal returns and encounters bottlenecks. Thus, the focus of academia and industry is pivoting ever more clearly toward foundational subjects that are closely related to 3D vision, such as “world models”, “spatial intelligence”, and “embodied intelligence”, granting 3D vision with unprecedented strategic attention and developmental opportunities. In 2025, the primary frontier trends in the field of 3D vision can be summarized as follows. 1) Feed-forward 3D reconstruction that supports spatiotemporal multi-image input: with breakthroughs in feed-forward 3D reconstruction technologies, such as visual geometry grounded Transformer(VGGT), obtaining scene structure and motion information through spatiotemporal multi-image feed-forward methods has become increasingly simple, exerting two profound effects. First, it provides a solid foundation to 3D scene understanding for spatial intelligence, allowing many traditional 2D vision problems to be solved more fundamentally in 3D space. Second, combined with efficient rendering technologies, such as 3D Gaussian splatting (3DGS), the threshold for high-quality 3D content production has been significantly lowered, paving the way for large-scale applications, such as digital twins and the metaverse. 2) Gradual fusion of 3D generation and 3D reconstruction: 3D AI generated content technologies, such as SAM3D support compositional and instance-level object generation under single-image input, with generation quality gradually reaching industrial-grade scanning standards while simultaneously integrating feed-forward reconstruction methods to gradually achieve the generation of authentic 3D structures and textures that are consistent with the input images. This condition will support the feed-forward multi-instance reconstruction of dynamic complex scenes, significantly improving real-time, multimodal perception and understanding capabilities in complex scenarios. 3) Integration of video generation and world models into embodied intelligence: Video generation technology is rapidly incorporating explicit or implicit 3D representations and evolving toward multi-view consistency, long sequences, and physical plausibility, directly driving the development of integrated “perception-generation-interaction” world model technologies. These types of world models, combined with feed-forward 3D reconstruction technology, will form a complete “multimodal perception-3D modeling-4D generation-real-time interaction” 4D world model. Simultaneously, world model methods have begun to serve embodied intelligence, and a unified framework of “understanding-generation-execution” has begun to emerge. World models are widely regarded by the academic community as the key path to achieving generalizable embodied intelligence and ultimately leading to artificial general intelligence. 4) Human behavior and video data becoming the core fuel that drives breakthroughs: Human operational spaces and interaction videos constitute a “data goldmine” for training embodied intelligence. The vast amount of human behavior videos on the internet, along with first-person perspective data collected through simple devices, contain physical common sense, causal reasoning, and interaction preferences that serve as the natural fuel for breaking through the current data bottlenecks of embodied intelligence. By performing explicit 3D perceptual reconstruction or latent space action alignment and learning on these data, a “data pyramid” base can be constructed to drive the scaling of embodied intelligence. 5) Evolution of the embodied training paradigm from imitation learning to interaction-driven reinforcement learning: The technical evolution of embodied intelligence vision-language-action (VLA) models is shifting from a supervised fine-tuning paradigm that relies on expert demonstrations to a composite training architecture that integrates online reinforcement learning. This shift effectively breaks the dependence on scarce high-quality data, enabling policies driven by sparse rewards to obtain generalization and exploration capabilities that surpass those of imitation learning, and solving the challenges of exploration and stable updates in continuous action spaces. Simultaneously, the development of high-performance training systems and action-conditioned world models provides infrastructure support for large-scale interaction data generation and efficient policy evolution, marking a new “post-training” stage for embodied intelligence that is centered on “interaction-driven” approaches. The top 10 selected research advancements of 2025 in the field of 3D vision include the following: 1) feed-forward 3D reconstruction that builds the foundation models for 3D vision (spatial intelligence); 2) the convergence of the reconstruction and generation technical routes (video generation/3D generation), moving away from mutual assistance toward preliminary integration; 3) 3DGS/4DGS continuously improving representation efficiency, triggering a surge in scene modeling and volumetric video applications; 4) 3D generation: a shift from single-object visual realism to structuralized components/scenes and physical interactivity; 5) from video generation to world models: oriented toward spatiotemporal consistency, physical plausibility, and interactivity; 6) unified multimodal large models for understanding and generation, serving spatial intelligence perception; 7) frontier shifts in digital humans: from appearance modeling to multimodal interaction; 8) human data becoming the essential fuel for breaking through the scaling law of embodied intelligence; 9) embodied intelligence foundation models that are evolving toward unified models of integrated “understanding-imagination-execution”; and 10) “post-training” moment of embodied intelligence: the paradigm shift of VLA models from imitation learning to online reinforcement learning. Collectively, these breakthroughs have established the prototype of an integrated intelligent architecture characterized by “multimodal perception-3D modeling-4D generation-real-time interaction”, providing critical technical support for the substantive advancement of spatial and embodied intelligence. To promote academic discourse, this study extensively analyzes frontier trends in 3D vision and curates the top 10 annual research advances, offering valuable reference perspectives for academia and industry.  
      关键词:3D vision;Embodied AI;World model;reconstruction and generation;spatial intelligence   
      416
      |
      353
      |
      0
      <HTML>
      <L-PDF><WORD><Meta-XML>
      <引用本文> <批量引用> 155139672 false
      更新时间:2026-07-20
    • Progress in audio-driven digital human technologies AI导读

      Li Weibin, Gao Jiafeng, Xu Bing, Hou Biao, Jiao Licheng
      Vol. 31, Issue 7, Pages: 2317-2338(2026) DOI: 10.11834/jig.260103
      Progress in audio-driven digital human technologies
      摘要:With the exponential growth of the metaverse, immersive virtual reality, and advanced human-computer interaction paradigms, audio-driven digital human generation has rapidly emerged as a paramount research frontier in computer vision and computer graphics. This transformative technology aims to establish a robust and seamless cross-modal mapping from 1D speech signals to 3D visual video streams. In contrast with traditional motion capture-based approaches, which are constrained by expensive hardware and rigid physical setups, audio-driven methodologies offer a highly scalable, cost-effective, and interactive solution. From a cognitive psychology perspective, human speech perception is inherently a multimodal process. The classic McGurk effect demonstrates that visual cues, such as precise lip movements, play a decisive role in deciphering auditory information, especially in noisy environments. Consequently, synthesizing a lifelike digital human requires not only ultrahigh-definition visual rendering but also millimeter-level, time-synchronized alignment between auditory signals and visual lip dynamics to prevent the notorious “uncanny valley” effect, where subtle unnaturalness provokes discomfort in human observers. Despite the proliferation of extensive research and rapid algorithmic advancements, generating photorealistic and real-time interactive digital avatars continues to encounter three formidable challenges. First, bridging the immense semantic gap between distinct modalities—specifically, translating semantic and acoustic features into complex spatial facial muscle deformations—remains fundamentally difficult. Second, the trade-off between high-fidelity rendering and real-time inference latency poses a critical bottleneck. Achieving pore-level skin details and realistic hair rendering typically demands immense computational resources, conflicting with the strict low-latency requirements (often under 100 ms) necessary for seamless live streaming or real-time conversational agents. Third, zero-shot cross-identity generalization and the expansion of driving scope—from isolated talking heads to full-body expressive agents that encompass co-speech gestures—demand robust structural priors that current isolated models struggle to provide. To deconstruct this complex domain systematically, this comprehensive review first proposes a unified “encoding-mapping-rendering” technical framework that encapsulates the majority of audio-driven digital human architectures. The framework initiates with an audio feature extractor, utilizing state-of-the-art self-supervised foundational models (e.g., WavLM and Whisper) to capture rich phonetic, semantic, and emotional representations from raw audio waveforms. These acoustic features are then fed into visual geometric representations, transitioning from sparse 2D landmarks into sophisticated 3D morphable models, which provide explicit facial motion space. Finally, a cross-modal mapping and rendering network translates these spatial configurations into final visual video frames. Following this structural foundation, this study meticulously reviews the evolutionary trajectory of audio-driven rendering technologies, categorizing them into four dominant paradigms. First, early 2D image-based generation methods, predominantly driven by generative adversarial networks, are examined. Although these models excel in achieving rapid inference speed and fundamental lip synchronization, they inherently lack topological constraints, suffering from severe structural degradation and texture stretching when handling large-pose head rotations. Second, the integration of neural radiance fields (NeRFs) introduces a paradigm shift by leveraging implicit volumetric representations. NeRF-based approaches successfully resolve multi-view consistency issues; however, the heavy computational burden of volumetric ray marching restricts their real-time application, while their subject-specific optimization nature limits broad generalization. Third, the recent breakthrough of 3D Gaussian splatting (3DGS) has revolutionized the field by combining explicit geometric primitives with highly optimized rasterization pipelines. 3DGS-based models break real-time rendering constraints, easily surpassing 100 frames per second, although they still struggle with dynamic consistency during extreme facial deformations due to their discrete point cloud nature. Lastly, diffusion models represent the current pinnacle of visual fidelity and generative diversity. By iteratively denoising latent spaces, these models synthesize unprecedented skin textures, hair dynamics, and micro-expressions, although their inherent Markov chain sampling mechanisms introduce significant inference latency and temporal jitter challenges into long video generation. To facilitate an objective assessment of these diverse methodologies, this study synthesizes mainstream datasets and rigorous evaluation metrics. We detail the utilization of large-scale in-the-wild datasets, such as VoxCeleb, alongside high-resolution benchmarks, like high-definition talking face, and emotional datasets, such as the multi-view emotional audio-visual dataset. The evaluation criteria are multidimensional, encompassing lip-sync accuracy through SyncNet metrics (lip sync error distance and lip sync error confidence), visual quality via Fréchet inception distance and learned perceptual image patch similarity, temporal video consistency that utilizes Fréchet video distance, and identity preservation measured using cosine similarity. Finally, this study provides an in-depth analysis of open challenges and outlines prospective future research trajectories. We hypothesize that the integration of multimodal large language models will fundamentally redefine cross-identity generalization, serving as powerful world models for decoding audio into structured motion commands. Furthermore, transitioning from purely correlation-based learning to causal modeling will be essential for decoupling authentic emotional expressions from entangled acoustic noise. To address the theoretical bottlenecks of long-sequence temporal consistency, we propose exploring state space models, such as Mamba architectures, and flow matching trajectories. Concurrently, given that these generative capabilities approach indistinguishable realism, ethical implications, particularly those that concern deepfake misuse, necessitate the parallel development of robust forensic detection and defense mechanisms. Ultimately, this review aims to serve as a definitive guide for researchers, navigating the transition of digital humans from mere visual replicas into fully embodied, emotionally intelligent interactive agents in the digital era.  
      关键词:digital human;Audio-Driven;Talking Head Generation;neural radiance field (NeRF);3D Gaussian splatting(3DGS);diffusion model   
      75
      |
      49
      |
      0
      <HTML>
      <L-PDF><WORD><Meta-XML>
      <引用本文> <批量引用> 157321213 false
      更新时间:2026-07-20
    • Bao Hong, Liang Tianjiao, Zheng Ying
      Vol. 31, Issue 7, Pages: 2339-2357(2026) DOI: 10.11834/jig.260112
      The new-generation artificial neural networks in the 21st century — embodied cognitive neural networks: a review
      摘要:For the 21st century, this study is dedicated to exploring the fundamental theories, models, and architectures of next-generation artificial neural networks (ANNs), with the objective of constructing high-performance ANNs that featuring adaptive topology, interpretability, strong generalization capability, and high energy efficiency. Since the initial proposition of ANNs in the 1940s, the field has witnessed over 80 years of development. Extending to the mid-21st century, ANNs can be categorized into five generations based on five core dimensions, and such generational evolution constitutes the core trajectory of neural network research and advancement. This study defines five generations of ANNs (abbreviated as xG-ANNs) from five perspectives: neuronal unit, information coding, network structure, learning mechanism, and Turing test. 1G-ANNs: Threshold logic networks, represented by the McCulloch-Pitts model and the perceptron; 2G-ANNs: Continuous activation networks (e.g., sigmoid, tanh), typified by the classical back-propagation network; 3G-ANNs: Represented by spiking neural networks 4G-ANNs: Deep neural networks, represented by AlexNet, ResNet, Transformer architectures, and the attention mechanism, which have passed the conversational Turing test for disembodied intelligence; 5G-ANNs: Cognitive neural networks, with five defining characteristics: 1) cognitive units that integrate memory, reasoning, and human-like attention; 2) hybrid coding of semantic symbols and distributed representations; 3) modular cognitive architecture with dynamic reconfigurable topology; 4) meta-learning, causal reasoning, and lifelong learning; and 5) the embodied Turing test that remains as a breakthrough. A core academic consensus has been formed regarding 5G-ANNs. That is, such networks integrate neural computing, symbolic reasoning, and cognitive architecture with low-power consumption, support dynamic topology, memory cognition, neuro-symbolic fusion, and embodied/world models. They exhibit intrinsic advantages, including adaptive structure, few-shot generalization, interpretability, low energy consumption, and embodied cognition. At present, the field is in the era of 4G-ANNs, which are characterized by data-driven fitting, deep learning, the attention mechanism, and Transformer frameworks. Represented by large language model-based ChatGPT, 4G-ANNs have passed the conversational Turing test, but such validation is a “black box” assessment that is restricted to the emergence of disembodied intelligence. The root causes lie in the inherent asymmetry of large language models built on the scaling law of parameter expansion, their lack of comprehension of physical laws that govern the real world, and critical drawbacks, including jagged multimodal and multiform intelligent output and inferior energy efficiency. By contrast, neural networks that are deployed in embodied intelligent robots lack autonomous intelligence and can only execute predefined actions per programmed instructions, leading to a huge gap between disembodied intelligence and embodied intelligence in 4G-ANN systems. Addressing these critical limitations of 4G-ANNs calls for the support of novel theories, models, and architectural designs. At present, extensive debates and divergence persist over the developmental orientation and technical routes of next-generation ANNs. This study analyzes and summarizes the developmental progress of mainstream theories, models, and architectures of the first four generations of ANNs. It focuses on the characteristics of several typical 4G-ANN models and their enhanced variants, and reviews representative architectures for 5G-ANNs, including world models, joint embedding predictive architecture, cognitive spiral models, and intelligence-trace cellular network frameworks. Finally, grounded in theory of cognitive physics and the practical technologies of driving brain cognition, this work proposes a lightweight architecture, called embodied cognitive physics neural network (E-CoPNN), which incorporates the core characteristics of 5G-ANNs. In particular, E-CoPNN features dynamic topology (three-layer nested structure and dual systems), memory cognition (three categories of memory), neuro-symbolic fusion (integration of Euclidean and non-Euclidean data, combination of lexemes and concepts), and embodied cognition/world models (fusion of disembodied and embodied intelligence, integration of cognitive space and physical space). The proposed architecture embodies the typical properties of 5G-ANNs: an adaptive structure, few-shot generalization, interpretability, low power consumption, and embodied cognition. Targeting brain-like general intelligence, 5G-ANNs regard dynamic topology, autonomous memory, cognitive reasoning, and structural evolution as core attributes, shifting the ANN research paradigm from data fitting to structural reconstruction, and thus, representing a pivotal direction for achieving machine autonomous intelligence and continuous learning.Conclusion and Significance: At present, the development of next-generation ANNs for the 21st century will drive a series of transformative revolutions: philosophically, a paradigm shift from mind-body dualism to embodied cognitive science based on embodied perception monism; theoretically, a disciplinary expansion from 20th century biophysics to 21st century cognitive physics; model-wise, a fundamental transition of ANN research from data fitting into structural reconstruction; in application, bridging the long-standing gap between disembodied intelligence and embodied intelligence in ANN development; generationally, an evolutionary leap from 4G-ANNs to 5G-ANNs characterized by brain-like cognitive capabilities and adaptive topological structures. This work lays a solid foundation for the widespread application of embodied cognitive robots with learning, self-development, self-correction, and human-robot interaction capabilities. Furthermore, it supports the convergent development of nano-bio-info-cogno, or the four cutting-edge technologies led by cognitive integration. Finally, this work enhances human intellectual competence and ushers in a new round of cognitive revolution.  
      关键词:Artificial Neural Network (ANN);fifth generation ANN embodied intelligence;disembodied intelligence;embodied cognition;cognitive physics(cognitophysics);Embodied Cognitive Physics Neural Network (E-CoPNN);unified field theory;intelligence field;human attention mechanism;selective mechanism;driving brain cognition;review   
      270
      |
      249
      |
      0
      <HTML>
      <L-PDF><WORD><Meta-XML>
      <引用本文> <批量引用> 153982493 false
      更新时间:2026-07-20
    • Zhang Hui, Tian Yonglin, Wang Yutong, Gou Chao, Li Xuan, Wang Feiyue
      Vol. 31, Issue 7, Pages: 2358-2380(2026) DOI: 10.11834/jig.250231
      Parallel images: the theoretical foundation and multimodal vision applications
      摘要:Driven by the rapid advancement of deep learning and large-scale computing, visual perception systems have achieved remarkable progress in a wide range of domains, including autonomous driving, intelligent transportation, security surveillance, medical diagnostics, industrial inspection, and human-robot interaction. Modern vision algorithms, empowered by massive datasets and increasingly sophisticated neural architectures, are now capable of performing object detection, semantic segmentation, scene understanding, and even causal reasoning with unprecedented accuracy. Despite this rapid growth, the development of visual intelligence still faces several critical bottlenecks. Chief among these challenges are the highly imbalanced data distributions that characterize real-world environments, the scarcity of long-tail or rare-event samples essential for robustness, and the substantial human and financial cost associated with large-scale manual annotation. These factors significantly hinder the performance, safety, and generalizability of deep perception systems, especially in complex, dynamic, or safety-critical scenarios. Parallel images technology, emerging as a novel image generation and modeling methodology grounded in parallel systems theory and the artificial systems, computational experiments, and parallel execution (ACP) framework, offers a promising pathway to address these limitations. The core idea behind parallel images is to construct controllable, high-fidelity artificial scene systems that reflect the structure, behavior, physics, and semantics of their real-world counterparts. Within these artificial systems, computational experiments can be conducted at scale, allowing for the controlled generation of diverse visual data that capture variations in illumination, geometry, environmental conditions, sensor characteristics, and task-specific factors. Through the interaction and iterative feedback between virtual and real environments, parallel images establish a closed-loop mechanism of “modeling-training-feedback-optimization”, enabling perception models to evolve continuously, validate hypotheses, and improve performance under systematically generated variations. This closed-loop mechanism differentiates parallel images from traditional synthetic data generation in several important aspects. First, instead of passively producing static rendered images, parallel images emphasize dynamic parallelism, where virtual agents, environments, and tasks evolve in sync with real-world processes. Second, the approach integrates multimodal feedback, bridging visual, geometric, physical, and semantic modalities to ensure consistency and translatability across domains. Third, the framework supports scalable modeling of rare, dangerous, or expensive scenarios that are difficult or impossible to capture in real life, such as near-crash events in autonomous driving, rare diseases in medical imaging, or hazardous industrial operations. These capabilities make parallel images a powerful tool for enhancing the robustness, safety, and domain generalization of modern perception systems. This paper provides a comprehensive, systematic review of the theoretical foundations, methodological innovations, and developmental trajectory of parallel images technology. This paper begins by revisiting its roots in parallel intelligence and the ACP paradigm and detailing how artificial systems serve as controlled experimental platforms that complement real-world data collection. Then, this paper examines recent technical advances encompassing three major research directions aligned with the ACP framework: 1) multimodal data-driven virtual scene generation, which employs generative adversarial networks, diffusion models, neural radiance fields, and 3D Gaussian splatting to overcome data scarcity and annotation bottlenecks, enabling the creation of controllable, editable, and semantically consistent synthetic environments; 2) multiview feature fusion and virtual-real model transfer, aimed at addressing feature discrepancies and semantic misalignment across heterogeneous visual modalities through cross-modal alignment, multigranularity adaptive transfer, and domain-bridging strategies that enhance generalization and adaptability in hybrid virtual-real environments; 3) parallel reasoning through heterogeneous data and knowledge fusion, which integrates structured information extraction, external knowledge guidance, scene graphs, temporal logic, and large language models to advance perceptual understanding toward semantic-level reasoning and decision-making, thereby supporting continuous optimization and closed-loop evolution in complex scenes. Beyond summarizing technological developments, this paper also situates parallel images within the broader context of emerging trends in generative artificial intelligence and foundation models. With the rise of diffusion models, neural radiance fields, and large-scale multimodal models, parallel images are poised to integrate more deeply with generative simulation pipelines. This paper discusses how these innovations can strengthen the fidelity, controllability, and adaptability of artificial visual data and potentially enable new capabilities such as task-conditioned scene synthesis, human-AI cosimulation, interactive data generation, and closed-loop autonomous scenario exploration, and provides key capabilities for building general visual systems with continuous learning and feedback optimization. This paper provides crucial support for building general visual systems with continuous learning and feedback-driven optimization. Finally, this paper identifies several open challenges and future research directions that are essential for advancing the development of parallel images systems. These challenges include achieving high-quality expansion of virtual data, bridging the semantic gap between virtual and real domains, and enabling real-time, tightly coupled virtual-real interaction. Addressing these issues will require advances in intelligent generation models, self-supervised quality evaluation, unified data standards, causality-aware cross-domain alignment, and low-latency virtual-real collaboration supported by next-generation communication and sensing technologies. Solving these challenges will be critical for pushing forward the frontier of synthetic visual intelligence and unlocking the full potential of parallel images in real-world applications.  
      关键词:Parallel Images;Generative Artificial Intelligence;Virtual-Real Integration;scene generation;Digital Twin   
      85
      |
      209
      |
      0
      <HTML>
      <L-PDF><WORD><Meta-XML>
      <引用本文> <批量引用> 153576772 false
      更新时间:2026-07-20
    • Advances in open vocabulary perception for remote sensing images AI导读

      Li Kaiyu, Cao Xiangyong, Jiang Zixuan, Meng Deyu
      Vol. 31, Issue 7, Pages: 2381-2407(2026) DOI: 10.11834/jig.260163
      Advances in open vocabulary perception for remote sensing images
      摘要:Remote sensing technology serves as the core mechanism for the observation of the Earth and the understanding of surface environments. It plays an irreplaceable role in critical fields, such as natural disaster monitoring, urban planning, resource exploration, and ecological protection. Driven by the rapid advancement of deep learning over the past decade, the intelligent interpretation of remote sensing images has achieved breakthrough progress in fundamental vision tasks. However, the traditional deep learning paradigm is intrinsically built upon a close-set assumption, i.e., that models can only recognize a predefined and human-annotated set of fixed categories during the inference stage. When confronted with highly complex surface environments in real-world Earth observation scenarios, dynamic object morphology, and rare ground objects with long-tail distributions, this traditional paradigm not only incurs prohibitive costs for the construction of massive pixel-level annotated datasets but also easily falls into the trap of domain-specific overfitting. Consequently, the generalization and response capabilities of this paradigm are severely challenged by unseen categories or sudden events, making this paradigm inadequate for meeting the highly dynamic interpretation demands of the open world. In recent years, the rapid development of vision-language models has catalyzed a paradigm shift in artificial intelligence from task-specific models into general-purpose perception models. By mapping visual representations and natural language into a unified feature space through contrastive learning on massive image-text pairs, these models have broken the constraints of discrete labels, enabling a direct response to arbitrary natural language prompts. This capability is known as open vocabulary perception. Although this technology has demonstrated remarkable zero-shot generalization and cross-modal reasoning capabilities in the natural image domain, the direct application of these general vision-language models to the remote sensing domain encounters a severe domain gap. The uniqueness of remote sensing data poses multiple challenges to the adaptability of existing models. First, the distinct overhead imaging perspective causes drastic variations in object scale and complex background textures. Second, Earth observation tasks rely on multisource heterogeneous data from synthetic-aperture radar, multispectral or hyperspectral imaging, and thermal infrared sensors. The underlying physical mechanisms of these sensors exceed the inherent inductive biases of models that are pretrained solely on natural RGB images. Third, remote sensing objects often exhibit strong geospatial attributes and complex topological associations. To address these critical challenges, this study provides a comprehensive and systematic review of recent advancements in open vocabulary perception for remote sensing images. We first delve into the foundational aspect of this field: vision-language pretraining for remote sensing. We extensively review the evolution of construction strategies for large-scale datasets. We highlight the transition from limited, human-annotated image-text pairs into massive datasets generated via heuristic rules, the integration of geographic metadata, and advanced multimodal large language models, including innovative approaches that leverage OpenStreetMap and geographical coordinates to produce fine-grained, physics-aware descriptions across multiple modalities. Concurrently, we systematically summarize the progression of pretraining methodologies. Although early approaches have primarily focused on simple domain adaptation through continuous pretraining, recent state-of-the-art frameworks emphasize physics-aware encoding, fine-grained multilevel consistency learning, and geography-enhanced architectures. These frameworks better capture the intricate spatial relationships and modality diversities that are inherent in Earth observation data. Subsequently, this review conducts an in-depth analysis of the adaptation and optimization of open vocabulary perception techniques across a wide spectrum of crucial downstream tasks. For zero-shot scene classification and cross-modal retrieval, we discuss advanced strategies that are designed to mitigate the high intra-class similarity and complex interclass variances typical in remote sensing. We emphasize the shift toward fine-grained local-global alignment, hard negative mining, dynamic soft labeling, and prompt engineering. In the realm of open vocabulary image segmentation, we categorize the existing literature into training-based methods and training-free or annotation-free paradigms. Training-based methods leverage base categories to adapt models while preventing catastrophic forgetting through pseudo-label distillation and knowledge retention mechanisms. Training-free paradigms synergize foundational models, such as CLIP and the segment anything model, to extract structural masks and align semantics without the updating of network weights. For open vocabulary object detection and remote sensing visual grounding, we explore the approaches of researchers to deal with extreme scale variations, arbitrary orientations, and dense object distributions. These approaches include innovative frameworks for pseudo-label generation, multi-scale feature alignment, cross-modality context modeling, and interactive grounding mechanisms. Furthermore, we examine open vocabulary change detection, wherein recent studies employ either combinations of pretrained vision-language models or generative models to generate large-scale data. These approaches aim to identify arbitrary, text-specified surface transitions and simulate complex spatiotemporal changes without reliance on massive and costly bitemporal pixel-level annotations. We also briefly discuss emerging open vocabulary applications in 3D urban point clouds and cross-domain archeological remote sensing, illustrating the expanding horizon of this technology. Despite remarkable progress, the field of open vocabulary perception for remote sensing remains in a crucial developmental stage and faces several critical bottlenecks. This study critically identifies the limitations of current research, including the severe scarcity of high-quality and geographically balanced training data. This scarcity leads to geographic biases and performance degradation in data-poor regions. In addition, genuinely fine-grained and long-tailed open vocabulary evaluation benchmarks that can accurately reflect the performance of a model are prominently absent in extreme or unknown real-world scenarios. The inadequate physical understanding of heterogeneous modalities and the inherent black box unreliability of current large models in high-stake decision-making scenarios further constrain practical deployments. To chart the course for future research, we outline several promising and essential trajectories. First, we anticipate a paradigm shift toward generative perception that is driven by multimodal large language models. This shift unifies various spatial localization tasks into the direct generation of coordinate sequences or geometric property tokens to utilize fully the logical reasoning capabilities of foundational models. Second, we strongly advocate for the construction of rigorous, real-world, and fine-grained evaluation systems that incorporate complex spatiotemporal logic, diverse geographic conditions, and comprehensive evaluation metrics. Third, the development of omni-modal foundation models that explicitly integrate physical priors and deep learning is considered crucial for the achievement of all-weather and all-spectrum Earth observations, moving beyond pure data-driven approaches. Furthermore, we highlight the necessity to extend perception from static spatial analysis to dynamic spatiotemporal causal reasoning to decode the evolutionary processes of the Earth. Finally, addressing the severe conflict between the massive parameter scale of foundation models and the limited computing power of aerospace edge devices requires focused research into efficient, trustworthy, and safe edge-cloud collaborative computing architectures. By systematically synthesizing these advancements and challenges, this comprehensive review aims to serve as a foundational road map for researchers and practitioners. It accelerates the transition of the intelligent interpretation of remote sensing from isolated, close-set recognition toward artificial general intelligence that is capable of highly reliable, dynamic, and open world perception.  
      关键词:remote sensing image;open vocabulary perception;vision-language model(VLM);zero-shot learning;intelligent interpretation   
      282
      |
      290
      |
      0
      <HTML>
      <L-PDF><WORD><Meta-XML>
      <引用本文> <批量引用> 155139984 false
      更新时间:2026-07-20
    • Xue Xuqian, Wen Jie, Liu Xinwang, Zhang Junping
      Vol. 31, Issue 7, Pages: 2408-2433(2026) DOI: 10.11834/jig.260125
      Prior-driven multi-view clustering: a perspective from geometry and semantics to LMM cognition
      摘要:With the rapid emergence of large multimodal models (LMMs), massive heterogeneous data are proliferating across various industrial and scientific domains. In this context, multi-view clustering (MVC) serves as a cornerstone technology for unsupervised knowledge discovery and latent correlation mining. At present, MVC is undergoing a profound and historical paradigm shift. Traditional surveys predominantly focus on the horizontal categorization of algorithmic network structures. However, this approach often fails to reveal the intrinsic evolutionary logic across different technological eras. Departing from these conventions, this study proposes a pioneering, prior-driven theoretical perspective to reconstruct systematically the developmental trajectory of MVC over the past two decades. This goal is achieved through a trans-paradigm analytical framework of geometry-semantics-cognition. During the initial stage of shallow structural mining, the research paradigm focused on explicit mathematical constraints within original or kernel-induced feature spaces. Euclidean space methods, such as multi-view K-Means and nonnegative matrix factorization, identify global prototypes by minimizing squared error or Frobenius norm reconstruction loss. By contrast, affine space methods leverage self-representation properties to model data as a union of low-dimensional subspaces, while manifold space techniques utilize spectral graph theory to transform clustering into optimal graph cut problems by preserving local topological correlations. As the field transitioned into deep spatial modeling based on semantic collaborative priors, researchers utilized the powerful nonlinear mapping capabilities of deep neural networks to project heterogeneous data into high-order semantic spaces. This evolution encompasses several distinct research paradigms. 1) Embedding space research focuses on deep subspace clustering, employing autoencoders to learn discriminative features while maintaining cross-view consistency. 2) Latent space methods utilize probabilistic generative models, such as variational autoencoders and generative adversarial networks, to align latent distributions and infer missing view information through adversarial games. 3) Augmented space paradigms introduce contrastive learning to maximize mutual information between views via InfoNCE-like losses, enhancing representation robustness. 4) Topological space studies leverage graph neural networks to mine intra-view geometric structures and inter-view complementary semantics synergistically. Moving into the current era of LMMs, this study progressively explores deep alignment based on cognitive priors, where LMMs are viewed not merely as data sources but as knowledge bases that contain human-level common sense and logical reasoning capability. We systematically elucidate the infrastructural role of MVC in empowering massive data governance. Specifically, MVC facilitates semantic deduplication to enhance data quality and employs token-level clustering to optimize mixture-of-experts(MoE) routing for expert specialization. Furthermore, MVC enables hierarchical semantic chunking, which is critical for precise document retrieval within retrieval-augmented generation(RAG) frameworks. Beyond these applications, we analyze the potential for LMM logical reasoning to back-propagate into clustering tasks. This synergy elevates MVC from pure statistical feature alignment to a new dimension of knowledge-driven cognitive logic consistency. Beyond theoretical frameworks, the review highlights the transformative effect of MVC in diverse real-world scenarios, ranging from multi-omics fusion in smart healthcare for cancer subtyping and spatial-spectral fusion in remote sensing for urban functional zone identification to cross-perspective pedestrian recognition in public security and unsupervised anomaly detection in industrial internet of things networks. Despite these advancements, achieving a transition from laboratory benchmarks to robust industrial infrastructure requires addressing several core challenges. First, the scaling bottleneck where in the O(N2) or O(N3) computational complexity of traditional methods must be reduced to linear levels to handle million-scale web data. Second, the extreme robustness required in open environments to handle not only severely incomplete views but also the pervasive long-tailed category imbalance where minority abnormal samples are frequently overwhelmed by dominant patterns. Third, the urgent need for evaluation system reconstruction to move away from legacy small-scale datasets with shallow features toward native heterogeneous multimodal benchmarks that reflect real-world weak alignment and complex noise distributions. Ultimately, this survey aims to provide a novel research road map for theoretical innovation and engineering practice, advocating for a transition toward intelligent decision systems characterized by high-level semantic decoupling and cognitive logic alignment in the era of LMMs.  
      关键词:multi-view clustering (MVC);Prior-driven learning;large multimodal models (LMMs);Geometric structure;Semantic collaboration;Cognitive alignment;review   
      112
      |
      223
      |
      0
      <HTML>
      <L-PDF><WORD><Meta-XML>
      <引用本文> <批量引用> 155139902 false
      更新时间:2026-07-20
    • Bi Shifan, Ye Liang, Wang Zhixiang, Zhang Ziyang, Hong Hanyu, Sang Nong
      Vol. 31, Issue 7, Pages: 2434-2470(2026) DOI: 10.11834/jig.250459
      Progress of research on object detection and tracking methods for nighttime UAV aerial imagery
      摘要:Nighttime perception remains a critical bottleneck for the autonomous operation of unmanned aerial vehicles (UAVs). Under low-light conditions, weak and uneven illumination amplifies noise and glare, background interference becomes increasingly complex, and targets often appear small, low in contrast, or partially occluded. Meanwhile, UAV platforms impose strict constraints on onboard computation and power consumption, requiring detection and tracking methods to achieve balance between robustness and efficiency. This paper presents a comprehensive review and empirical study of the application of deep learning in UAV nighttime target detection and tracking by synthesizing methodological trends, experimental findings, and future research directions. We characterize the unique challenges of UAV nighttime vision in contrast to daytime scenarios and highlight four interrelated issues: 1) limited perceptual signals and variable illumination, leading to degraded pixel-level cues; 2) feature degradation and background interference caused by artificial light sources and shadows; 3) restricted onboard resources that make the deployment of large-scale models difficult; and 4) complex imaging conditions, where target scale and appearance fluctuate frequently. These constraints underscore the need for specialized algorithmic paradigms rather than direct transfer of daytime methods. From the perspective of UAV nighttime target detection, this study systematically reviews four major research directions. First, nighttime image enhancement methods employ dedicated restoration networks to improve brightness and contrast before detection. Although these approaches can effectively enhance visual quality, the two-stage pipeline may introduce artifacts and does not always optimally align with detector objectives. Second, domain adaptation methods leverage adversarial learning, style transfer, or feature alignment to bridge the day-night distribution gap. Such methods reduce the dependence on annotated nighttime data but remain sensitive to interdomain diversity and often require carefully designed pseudo-labeling or regularization. Third, multimodal perception fusion, especially RGB-infrared (IR) fusion, utilizes complementary thermal cues to mitigate illumination deficiencies. State-of-the-art fusion networks (e.g., M2D-LIF) have demonstrated substantial mean average precision (mAP) improvements on UAV nighttime benchmarks, with some models achieving over 80% mAP@0.5 on the DroneVehicle-Night subset. Fourth, lightweight models (e.g., CSPDarkNet-based backbones) focus on balancing accuracy and efficiency under limited computational budgets. Experiments indicate that backbone choice is critical for enabling real-time nighttime detection without a considerable accuracy loss. For nighttime target tracking, this study categorizes four representative paradigms. 1) Enhancement-before-tracking: low-light enhancers are used as a preprocessing stage (sometimes co-optimized) to improve input quality and thereby downstream tracker performance. 2) Domain adaptation: extending static feature alignment to temporal modeling, sometimes combined with synthetic nighttime sequences or foundation models, to enhance robustness. 3) Prompt learning: reframing day-night transfer as a fast adaptation task and leveraging small-scale prompts for efficient transfer with minimal parameter updates. 4) Curriculum learning: ordering training samples from easy to difficult (e.g., progressively increasing darkness or motion blur) to stabilize training and improve performance under extreme degradation. 5) Multimodal fusion: we further analyze the interplay among these paradigms and the tension between restoration- and task-oriented objectives. Subsequently, commonly used evaluation metrics and benchmark datasets for nighttime and all-weather UAV detection and tracking are introduced, and a nighttime UAV vehicle detection dataset called DroneVehicle-Night is constructed. Two complementary empirical studies are conducted to support this survey. The first study, which is based on DroneVehicle-Night (derived from DroneVehicle), contains tens of thousands of paired RGB-IR annotated images covering urban roads, residential areas, and parking lots, with five classes of vehicles annotated. Experiments reveal that under identical training conditions, IR consistently outperforms RGB (e.g., CSPDarkNet-based detectors achieve high mAP@0.5 in IR), and multimodal fusion methods, such as M2D-LIF, achieve over 83.9% mAP@0.5, underscoring the value of modality complementarity in nighttime detection. The second study evaluates the transferability of generic UAV detectors to the nighttime scenario. From VisDrone2019, we construct a paired subset (521 nighttime/513 daytime images). Although generic detectors achieve 30%-45% mAP@0.5 on daytime data, their performance declines by 8%-12% at night, with small and low-contrast objects exhibiting the most severe degradation. These findings confirm that specialized nighttime and modality-aware methods are effective, and simple day-night transfer is insufficient. To further reflect the state of literature on nighttime tracking, this study summarizes the performance of 20 representative trackers on UAVDark135 and NAT2021 (results are cited from the original publications). Among them, Mamba-based methods (e.g., MambaNUT-Small and MambaTrack) and several domain adaptation methods (e.g., DARTer and DCPT) consistently rank among the top, achieving robust AUC-accuracy trade-offs while maintaining reasonable frame rates. These results highlight the importance of efficient sequence modeling and domain-aware adaptation in nighttime tracking. On the basis of the survey and empirical evidence, several key conclusions are derived. First, sensor modality selection is critical. IR and RGB-IR fusion substantially enhance robustness to low-light and complex backgrounds, whereas RGB-only pipelines remain fragile. Second, backbone design strongly influences the accuracy-efficiency balance. Compared with heavy backbones, such as ResNet50, CSPDarkNet variants offer more attractive real-time trade-offs on embedded platforms. Third, aside from frame-by-frame strategies, explicit temporal modeling and sequence-oriented architectures are indispensable for robust nighttime tracking. Fourth, day-night transfer remains a major challenge, as evidenced by the VisDrone experiments that highlight the urgent need for specialized nighttime designs and datasets. Finally, four promising research directions are identified: 1) cross-modal vision-language fusion that leverages large-scale language/vision models to enhance multimodal alignment and interactive tracking; 2) construction of rich multimodal nighttime datasets that extend beyond urban scenes to cover wild, mountainous, and aquatic environments; 3) unsupervised and weakly supervised learning to reduce annotation costs and improve cross-domain robustness; and 4) lightweight architecture design. Aside from accuracy, UAV tasks demand real-time and efficient models. Future work may explore long-sequence modeling with Mamba in combination with neural architecture search, dynamic inference, and efficient operators to balance accuracy, latency, and energy. In conclusion, this survey and empirical study provide a clear roadmap toward robust UAV nighttime perception and practical deployment. All algorithms and datasets used in this work are publicly available at https://github.com/bsfsf/DroneVehicle-Night and https://doi.org/10.57760/sciencedb.32435, to facilitate future research.  
      关键词:deep learning;unmanned aerial vehicles aerial image;nighttime UAV;target tracking;object detection   
      424
      |
      389
      |
      0
      <HTML>
      <L-PDF><WORD><Meta-XML>
      <引用本文> <批量引用> 144740091 false
      更新时间:2026-07-20
    • Jiang Jianxiang, Wu Yiquan
      Vol. 31, Issue 7, Pages: 2471-2504(2026) DOI: 10.11834/jig.250435
      Progress of research on noncooperative space object recognition and pose estimation on the basis of vision and deep learning
      摘要:With the recent increase in space missions and the proliferation of space debris, the perception of noncooperative objects has become a critical issue for on-orbit servicing and space situational awareness. Noncooperative objects are objects in space that lack any cooperative markers, communication channels, or navigation aids, such as defunct satellites, orbital debris, small debris, or unknown spacecraft. Identifying such objects and estimating their six-degree-of-freedom pose is crucial for various space activities, including debris removal, rendezvous and docking, satellite servicing, and space security. Reliable identification provides prior semantic and structural information for subsequent pose estimation, and accurate pose estimation supports autonomous navigation, capture, and manipulation. Therefore, progress in this field is not only of scientific value but also of practical importance for ensuring the safety and sustainability of space operations. This study comprehensively reviews the progress of the latest research on noncooperative space object recognition and pose estimation on the basis of vision and deep learning. Unlike traditional computer vision methods that rely heavily on handcrafted features and geometric modeling, deep learning methods demonstrate high robustness, adaptability, and generalization advantages in complex space environments, such as those with weak textures, drastic illumination variations, rapid pose changes, and partial occlusion. We categorize existing research into two major areas, namely, object recognition and pose estimation, and systematically analyze representative methods along with their strengths, limitations, and applicability. For object recognition, three main research directions are identified. The first is multimodal visual fusion, which integrates data from multiple sensors, such as RGB cameras, infrared imaging, depth maps, or radar. By leveraging complementary features across modalities, these methods can overcome the limitations of single-sensor systems, particularly in low visibility or with incomplete imagery. However, multimodal fusion often incurs additional computational overhead and sensor dependency. The second research direction is object detection and segmentation methods. Powered by deep models, such as you only look once (YOLO), single-shot multibox detector, faster region-based convolutional neural network, and mask region-based convolutional neural network, these methods have substantially improved spatial object localization and fine-grained analysis capabilities. However, these methods typically require large-scale annotated datasets and often struggle to identify very small or distant objects. The third direction is transfer learning and few-shot learning methods, which mitigate data scarcity by adopting pretrained models or meta-learning frameworks. These methods have produced promising results in identifying satellites or debris with limited sample sizes, although domain transfer between synthetic and real images remains a challenge. For pose estimation, this study reviews four major categories of methods. First, direct regression methods regard pose estimation as an end-to-end mapping problem and are highly effective when dealing with symmetric or highly variable objects. Second, keypoint detection methods combine landmark localization with geometric constraints to achieve robust performance. Third, unsupervised and domain adaptation methods reduce the reliance on expensive labeled data and address the gap between synthetic and real domains through reconstruction losses, contrastive learning, or adversarial adaptation. Finally, multimodal fusion methods combine visual data with LiDAR, radar, or infrared sensing to improve robustness in extreme environments, but the complexity of the integration limits their widespread application. In addition to methodological analysis, this study summarizes public datasets, such as SPEED, SPEED+, URSO, and BUAA-SID, and self-constructed datasets. Evaluation metrics, such as position error, pose error, mean average precision, intersection over union, and inference speed, are summarized to provide strong data support for researchers. For spatial object recognition, the traditional principal component analysis feature extraction method combined with random forest classification achieves an accuracy of approximately 89.13%. The deep convolutional neural network (DCNN) approach substantially improves recognition performance, with its accuracy increments ranging from 4.31% to 15.03%. By further employing data augmentation strategies, DCNN’s recognition accuracy is as high as 99.90%, demonstrating exceptional generalization and robustness. In the pose estimation task, with the SPEED dataset as a benchmark, the traditional perspective-n-point algorithm achieves a pose angle error of 8.74°, with its performance degrading by over 40% in low-resolution imagery. By contrast, the deep learning-based DSOAE-Net model reduces the pose angle error to 7.53° and maintains stable accuracy in low-resolution scenarios, demonstrating high environmental adaptability and robustness. Comparative analysis shows that deep learning-based methods outperform traditional methods in terms of robustness to illumination variations, occlusions, and structural ambiguity. Multimodal fusion and domain adaptation strategies are particularly effective in improving cross-domain generalization capabilities. In sum, this study summarizes the main challenges that face object recognition and pose estimation methods and proposes future development directions to provide a reference for technical research and engineering applications in this field. The methods, datasets, and evaluation metrics mentioned are linked at https://github.com/viskyll/openResource/tree/main.  
      关键词:non-cooperative space target;target recognition;pose estimation;vision;deep learning;multimodal visual fusion;direct regression;key point detection   
      552
      |
      457
      |
      0
      <HTML>
      <L-PDF><WORD><Meta-XML>
      <引用本文> <批量引用> 139329322 false
      更新时间:2026-07-20

      Foundation Models for Medical Image Analysis

    • Wen Xinhao, Liu Wei, Yue Xiaodong, Chen Yufei
      Vol. 31, Issue 7, Pages: 2505-2525(2026) DOI: 10.11834/jig.250475
      Survey of trustworthy medical image analysis methods on the basis of evidential deep learning
      摘要:The integration of deep learning into medical diagnostics has heralded a new era of data-driven healthcare, demonstrating profound potential in tasks ranging from radiological image interpretation and pathological slide analysis to genomic sequencing and clinical outcome prediction. These advanced computational models have achieved, and in some cases even surpassed, human-level performance on various well-defined tasks, promising to enhance diagnostic accuracy, streamline clinical workflows, and facilitate personalized medicine. However, the remarkable success of these models is fundamentally predicated on an implicit assumption: the availability of large-scale, high-quality, comprehensively annotated datasets. This assumption is in stark contrast with the reality of clinical data, which are intrinsically complex and challenging. Real-world medical data are typically characterized by scarcity, heterogeneity, and imbalance. The deployment of conventional deep learning models, which often behave as black boxes in such complex data environments, poses critical risks. These models are notoriously prone to producing overconfident yet erroneous predictions and lack a mechanism to express their uncertainty or to reliably signal when they are operating outside their domain of expertise. In the high-stake context of clinical decision-making, where a single incorrect prediction can have dire effects on patient safety, this lack of trustworthy uncertainty awareness is a critical barrier to the responsible adoption of artificial intelligence (AI). This review is motivated by the urgent need to address the challenge of real-world complex data in medical AI. We systematically structure the landscape of data complexities into three principal categories, which form the organizational framework of our analysis. The first one is data scarcity, which is common in rare diseases, underrepresented populations, and emerging pathologies; it leads to overfitting and poor generalization. The second one is data heterogeneity, which includes label noise, feature noise, multimodal fusion conflicts, and out-of-distribution (OOD) samples. The third one is data imbalance, where the dominance of normal over abnormal cases biases models toward the majority class, undermining their sensitivity for rare pathologies. A new paradigm—one that goes beyond mere prediction to incorporate principled uncertainty quantification—is required to address these challenges. This review centers on evidential deep learning (EDL), an emergent and powerful framework designed for this purpose. Unlike traditional models that output a single point-estimate probability via a softmax function, EDL reframes classification by placing a high-order probability distribution over the categorical distribution of classes. The core contribution of this review is that it provides a systematic, in-depth analysis of how the EDL framework and its associated uncertainty metrics can be specifically utilized to mitigate the problems arising from real-world complex medical data. We meticulously map EDL-based solutions to our three-pronged taxonomy of data challenges. For data scarcity, EDL models naturally express high epistemic uncertainty when making predictions for classes with insufficient training evidence, thereby providing a quantitative and actionable signal that the model’s output should be treated with caution. For data heterogeneity, EDL offers robust solutions. In cases of data noise and multimodal conflicts, the conflicting or low-quality evidence is reflected in the parameters of the learned Dirichlet distribution, leading to high overall uncertainty and preventing an overconfident but incorrect conclusion. EDL is particularly effective for OOD detection; an OOD sample, by definition, lacks supporting evidence in the training data, causing a well-trained EDL model to produce a prediction with high epistemic uncertainty and a“vacuous”evidence vector. This situation enables the model to effectively“abstain”from making a high-risk prediction and instead refer the case to a human expert. In the context of data imbalance, although not a direct remedy, EDL provides a more nuanced view than standard outputs. It can reveal when a model is uncertain about its predictions for minority classes, even if the predicted probability is low, thus helping identify instances that require further scrutiny. This study, therefore, provides a comprehensive and structured survey of the current state of the art in applying EDL to medical image analysis under real-world complex data conditions. We synthesize the disparate literature into a coherent conceptual framework, elucidating the underlying principles, core methodologies, and practical applications of EDL in this critical domain. In conclusion, we posit that evidential deep learning represents a pivotal step toward building safe, reliable, and trustworthy AI systems for medical applications. However, the field is still nascent and faces several challenges that open up avenues for future research. These challenges include the need for robust theoretical foundations for evidence modeling, improving the computational scalability of EDL methods, and developing techniques for effective uncertainty calibration to ensure that the quantified uncertainty meaningfully corresponds to the true likelihood of error. The clinical translatability of uncertainty needs to be explored, with focus on how to present uncertainty information to clinicians in an intuitive and actionable manner. Moreover, large-scale, prospective clinical validation studies are imperative to move beyond retrospective dataset benchmarks and rigorously assess the real-world effect and safety of EDL-based diagnostic systems. This review aims to serve as a valuable resource and roadmap, inspiring further research to bridge the gap between the theoretical promise of EDL and its practical deployment in future clinical workflows. The codes link: https://github.com/oscarab/Medical-EDL.  
      关键词:evidential deep learning(EDL);medical image analysis;uncertainty quantification;subjective logic;evidence theory   
      389
      |
      474
      |
      0
      <HTML>
      <L-PDF><WORD><Meta-XML>
      <引用本文> <批量引用> 143378785 false
      更新时间:2026-07-20
    • Brain structure-function coupling hypergraph neural network AI导读

      Lei Mengqi, Han Xiangmin, Li Siqi, Gao Yue
      Vol. 31, Issue 7, Pages: 2526-2537(2026) DOI: 10.11834/jig.250535
      Brain structure-function coupling hypergraph neural network
      摘要:ObjectiveIn recent years, brain networks have become an indispensable cornerstone in brain disorder research. At present, functional magnetic resonance imaging (fMRI) is commonly used to construct functional networks, while diffusion tensor imaging (DTI) is utilized to build structural networks. However, two limitations remain prominent. First, many approaches still operate on a single modality or adopt shallow multimodal fusion, leaving the potential dependency between structural connectivity (SC) and functional connectivity (FC) under-explored. Second, although structure-function coupling (SFC) has been demonstrated to differ across brain regions between healthy controls and patients, its relationship to classification has not been systematically integrated into representation learning to hold biomarker potential. This study aims to operationalize SFC as an explicit prior that guides multimodal representation learning, such that high-order cross-modal associations can be identified and translated into robust diagnostic performance.MethodThis study presents SFC-hypergraph neural network (HGNN), an SFC-guided foundational framework that combines a dual-stream encoder, cross-modal reconstruction pretraining, and hypergraph computation to model high-order relationships inside and between SC and FC. The core design is to use an SFC matrix as a bridge that explicitly guides two HGNN branches—one per modality—such that each branch learns modality-appropriate group relations while remaining informed by the other modality. Concretely, the SFC matrix is first constructed to capture the consistency between FC and SC patterns across brain regions by using a rank-based correlation measure. This matrix is fused with each modality through shallow feature mapping to form initial node-level features for the two streams. On the functional stream, hyperedges are formed via the sparse representation method, providing data-driven, high-order groupings that are appropriate for FC. On the structural stream, hyperedges are built with the k-nearest neighbor strategy to reflect global structural associations. Then, each branch employs HGNN layers to realize high-order messages that are passing on its respective hypergraph, encoding latent dependency between SC and FC under SFC guidance. To fully utilize cross-modal dependency, a cross-modal reconstruction pretraining task is introduced. During pretraining, the functional stream is trained to reconstruct the SC matrix while the structural stream reconstructs the FC matrix via decoders. The decoders are optimized with a reconstruction objective combined with a constraint that emphasizes the appropriate symmetry and sparsity of the reconstructed matrices. This pretraining forces the encoder to internalize modality-bridging information under SFC guidance. In the subsequent downstream tuning phase, the encoder parameters are frozen. Then, latent node representations from the two branches are flattened and concatenated into a global feature vector, which is fed into a lightweight multilayer perceptron classifier that is optimized with cross-entropy. Freezing the encoder keeps downstream training simple and stable while preserving the cross-modal dependency captured during pretraining.ResultExperiments are conducted on two public multimodal brain imaging datasets. On the Alzheimer’s disease neuroimaging initiative (ADNI), we use 332 subjects, specifically, 64 Alzheimer’s Disease (AD), 129 mild cognitive impairment (MCI), and 139 normal control (NC). On the autism brain imaging data exchange (ABIDE), we include 86 subjects with paired fMRI and DTI for autism spectrum disorder (ASD)/NC. We adopt the automated anatomical labeling atlas for both datasets. fMRI data are preprocessed with the data processing assistant for resting-state fMRI and DTI data with the pipeline for analyzing brain diffusion images. The model is implemented in PyTorch and trained on an NVIDIA RTX 4090 GPU. We follow a two-stage training protocol (cross-modal pretraining followed by frozen-encoder tuning) and evaluate with fivefold cross-validation, reporting accuracy (ACC), area under the curve (AUC), F1, and specificity. We compare SFC-HGNN against representative single-modality baselines (BrainNetCNN, GAT, HGNN+, BrainGNN) and multi-modality methods (MME-GCN, Cross-GNN). Across diagnostic tasks, SFC-HGNN achieves state-of-the-art performance, consistently improving ACC, AUC, and F1. Although certain baselines occasionally yield higher specificity, these cases are frequently accompanied by markedly lower F1, suggesting reduced stability. By contrast, SFC-HGNN maintains better overall balance among metrics and demonstrates stronger robustness. Ablation studies further underscore the contributions of the proposed components: 1) introducing SFC as a cross-modal bridge (without pretraining) improves accuracy by 1.6% and 2.3% on AD versus NC and ASD versus NC, respectively, and 2) adding cross-modal reconstruction pretraining brings additional gains of 2.1% and 1.9%. These results indicate that SFC guidance, along with cross-modal reconstruction, effectively encourages the encoder to capture latent SC-FC dependency that transfers to downstream classification. Finally, we conduct significant hyperedge analysis to assess interpretability by using group-level t-tests and visualize discriminative functional and structural hyperedges for AD versus NC and ASD versus NC. The identified regions and connections align with known patterns of network alteration, supporting the neurobiological plausibility of the learned high-order representations.ConclusionBy utilizing SFC as an explicit bridge and pairing dual HGNN encoders with a cross-modal reconstruction pretraining paradigm, this study introduces a principled multimodal framework for brain network analysis and brain disorder diagnosis. The approach encourages each modality to be informed by the other and encode high-order associations that are useful for downstream classification, while a frozen encoder tuning strategy keeps optimization stable and lightweight. On the ADNI and ABIDE datasets, SFC-HGNN consistently surpasses single-modal and multimodal baselines, with gains reflected not only in ACC and AUC but also in better balance between F1 and specificity, highlighting robustness. The significant hyperedge findings further provide neurobiologically plausible insights that complement the quantitative improvements. Overall, SFC-HGNN advances multimodal brain disorder diagnosis by unifying SFC-guided hypergraph encoding, cross-reconstruction pretraining, and simple downstream tuning into a coherent pipeline that achieves superior and reliable performance across datasets.  
      关键词:brain disorder diagnosis;multimodal brain network;structure-function coupling(SFC);hypergraph neural network(HGNN);cross-modal reconstruction;brain network foundation models   
      254
      |
      310
      |
      0
      <HTML>
      <L-PDF><WORD><Meta-XML>
      <引用本文> <批量引用> 154040441 false
      更新时间:2026-07-20
    • Deep unfolding network for brain image registration AI导读

      Sun Jiaying, Yang Shuangyan, Ying Shihui
      Vol. 31, Issue 7, Pages: 2538-2552(2026) DOI: 10.11834/jig.250283
      Deep unfolding network for brain image registration
      摘要:ObjectiveMedical image registration plays an important role in clinical applications such as surgical navigation, disease progression monitoring, and multimodal diagnosis. The accuracy of image registration is critical to ensuring the reliability of subsequent image analysis. Traditional methods formulate the registration process as an optimization problem by maximizing the similarity between the fixed and warped moving images and employ iterative optimization strategies to update the deformation field gradually. While such approaches provide clear theoretical interpretability, their high computational costs fail to meet real-time clinical demands. By contrast, deep learning-based methods leverage large-scale datasets to learn the deformation field directly. Although these data-driven approaches have demonstrated remarkable improvements in registration speed and accuracy, they operate as black-box models with limited interpretability. Moreover, most current approaches continue to depend on manually designed regularization terms, which may lack adaptability to diverse anatomical structures. To address these challenges, we propose an unfolding-based registration network that combines the theoretical guarantees of traditional optimization with the high efficiency and adaptive learning capabilities of deep neural networks.MethodSpecifically, by introducing variable splitting, we first decouple the registration problem into two subproblems: a similarity-constraint subproblem and a deformation-regularization subproblem. This decomposition reduces optimization complexity and allows each subproblem to be addressed separately. Inspired by deep unfolding networks, the method maps the alternating update process between the two subproblems to a multistage cascaded architecture. Each stage corresponds to one complete iteration of the optimizer, providing clear physical and mathematical foundations and enhancing interpretability compared with traditional end-to-end models. Within this framework, two specialized modules are designed to address each subproblem. The fidelity module (FM) is employed to solve the similarity-constraint subproblem. Convolutional neural networks (CNN) are widely used in deep learning-based registration. However, the limited receptive field of the convolution kernel leads the network to focus on local feature information, while the Transformer uses the self-attention mechanism to capture global dependencies effectively. By integrating the strengths of these two architectures, the FM is designed on the basis of the TransUnet framework. This hybrid design allows the FM to combine local features with global context. The denoising module (DM) is designed to solve the deformation-regularization subproblem. Instead of using handcrafted regularization terms, we approximate the regularization term using a trainable neural network. Concretely, the DM adopts the denoising convolutional neural network architecture and incorporates residual connections, allowing the model to learn priors adaptively in a data-driven manner and thus improving generalization across different anatomical structures. These two modules jointly complete a single alternating iteration and are cascaded to form the deep unfolding network. The overall network is trained with a loss function that balances image similarity, deformation consistency between the outputs of the FM and DM, and smoothness of the deformation field. The proposed network is implemented using the PyTorch framework and trained on an NVIDIA Tesla V100 GPU with 32 GB of memory. The model is optimized using the Adam optimizer with a learning rate of 1E-4, a batch size of 1, and is trained for 500 epochs. The Transformer component in the FM was configured with four attention heads, eight encoder layers, and an MLP dimension of 96. The unfolding network was constructed with three cascaded stages.ResultExperiments were conducted on two publicly available brain magnetic resonance imaging (MRI) datasets: LPBA40 and OASIS. We compared our method with a range of state-of-the-art models, including traditional registration methods (i.e., large deformation diffeomorphic metric mapping and symmetric normalization) and learning-based methods (e.g., VoxelMorph, TransMorph, UTSRMorph and H-SGANet). The quantitative evaluation metrics contained a dice similarity coefficient (Dice), the percentage of nonpositive Jacobian determinants, and inference time. The proposed method achieved an average Dice of 70.1% on LPBA40 and 79.9% on OASIS, while the percentage of nonpositive Jacobian determinants remained at a low level of 0.01% on both datasets, and the inference time was 0.55 s on LPBA40 and 0.37 s on OASIS. Although slightly higher than some deep learning-based models due to the cascaded unfolding architecture and Transformer modules, the time remained clinically feasible. Visualized results also indicated that the deformation fields produced by the proposed method was smoother and more consistent with anatomical priors. Additionally, we conducted a set of ablation studies and reported corresponding quantitative results. These ablation studies explicitly demonstrated the effectiveness of individual modules in our proposed model, including the variable splitting strategy, the TransUnet-based fidelity module, and the residual connection-based denoising module and assessed the influence of key hyperparameters.ConclusionIn this study, we proposed a deep unfolding network for medical image registration that effectively combines the interpretability of traditional iterative methods with the efficiency and adaptability of deep learning. By unrolling a decoupled optimization model into a trainable network with fidelity and denoising modules, the proposed approach strictly corresponds to a specific step in iterative algorithms. Experiments demonstrate that the proposed method remarkably enhances interpretability while maintaining high efficiency and accuracy.  
      关键词:medical image registration;unfolding network;model decoupling;regularization learning;brain image   
      166
      |
      530
      |
      0
      <HTML>
      <L-PDF><WORD><Meta-XML>
      <引用本文> <批量引用> 147851942 false
      更新时间:2026-07-20
    • Chen Yu, Zhu Dongxiao, Su Lei, Li Ke, Pecht Michael
      Vol. 31, Issue 7, Pages: 2553-2568(2026) DOI: 10.11834/jig.250334
      Multitask learning model for fetal heart standard plane recognition with high-recall classification
      摘要:ObjectiveFetal cardiac ultrasound examination plays a pivotal role in the prenatal diagnosis of congenital heart disease, which is one of the most prevalent and serious birth defects worldwide. However, clinical acquisition of fetal heart standard planes is highly challenging because of several physiological and technical constraints. The fetal heart is small in size, has a high beating rate, and its position in the maternal abdomen is easily influenced by fetal posture and movement. As a result, capturing the five clinically required standard planes, such as the four-chamber and outflow tract views, demands high temporal precision and operator experience. Moreover, the acquisition process suffers from poor repeatability, further hindering consistent and efficient screening outcomes. These challenges limit the overall effectiveness of fetal cardiac ultrasound screening in routine clinical workflows. To address these limitations and promote the automation of fetal cardiac ultrasound analysis, this study proposes FHSP-Net, a novel multitask deep learning model designed for accurate and comprehensive recognition of fetal heart standard planes. The model adopts a detection-then-classification strategy, which explicitly separates the anatomical structure detection task from the subsequent plane classification task, enabling effective feature extraction and decision-making for each subtask. This design reflects the clinical reasoning process in which a sonographer locates key anatomical structures and interprets them in the context of a specific standard plane.MethodIn the detection task, FHSP-Net incorporates several carefully designed components to handle the inherent complexities of fetal ultrasound imaging. A path-weaved network (PWN) is developed to improve the multiscale representation of features, particularly enhancing the detection of small and morphologically variable cardiac structures. In addition, a target aware enhancer (TAE) module is introduced to strengthen object-level context modeling, making the detection branch highly sensitive to anatomical structures that are often obscured by noise, artifacts, or low contrast. Equalization loss function version 2 (EQLv2) is adopted to further mitigate the effect of data imbalance, which is common in clinical datasets where some structures or views are underrepresented. This loss formulation ensures that the model maintains stable learning dynamics across different categories and structure sizes, yielding robust and generalizable detection results. While the detection task ensures a strong foundation by locating relevant anatomical regions, the plane recognition task focuses on accurately classifying ultrasound views on the basis of the detected structures. Traditional algorithms that rely on whether certain anatomical structures are present in an image (i.e., anatomical inclusion conditions) tend to have high precision but suffer from low recall. In other words, although they make only a few mistakes in their predictions, they often fail to identify all valid instances of each standard plane, resulting in missed diagnoses. To overcome this issue, we propose a candidate view scoring algorithm that evaluates and ranks multiple candidate views generated during the detection stage. This algorithm integrates structural cues and contextual features to refine classification decisions, substantially improving the model’s ability to capture valid plane instances across varied imaging conditions.ResultExperiments are conducted on a self-built fetal cardiac ultrasound dataset to evaluate the effectiveness of FHSP-Net. The dataset comprises a wide range of clinical scenarios and includes comprehensive annotations for anatomical structures and standard plane labels. Quantitative results demonstrate the superior performance of FHSP-Net across both subtasks. In the detection task, the model achieves a mean average precision of 0.962, representing a 0.027 improvement over a strong baseline model. This finding indicates that the proposed enhancements in feature fusion, object sensitivity, and loss balancing substantially benefit structure localization performance. In the plane recognition task, the candidate view scoring algorithm boosts classification accuracy by 0.10 compared with conventional structure-inclusion-based approaches. Ultimately, FHSP-Net attains an overall standard plane recognition accuracy of 0.959, representing a 0.194 improvement over the baseline model and the anatomical-inclusion-condition-based classification method. These results strongly validate the effectiveness and advantage of the proposed multitask design and its components.ConclusionThe FHSP-Net multi-task model demonstrates strong capabilities in anatomical structure detection and fetal heart standard plane recognition. The integration of PWN, TAE, and EQLv2 enhances robustness and precision at the detection stage even under challenging imaging conditions. The candidate view scoring algorithm effectively addresses the recall bottleneck of previous classification methods, ensuring that the model captures a large portion of valid standard plane instances without sacrificing accuracy. Together, these innovations make FHSP-Net a reliable and efficient solution for intelligent ultrasound analysis. FHSP-Net has potential for advancing computer-aided diagnosis in prenatal cardiology.  
      关键词:fetal cardiac ultrasound detection;standard plane recognition;key anatomical structure recognition;deep learning;object detection   
      208
      |
      460
      |
      0
      <HTML>
      <L-PDF><WORD><Meta-XML>
      <引用本文> <批量引用> 143378957 false
      更新时间:2026-07-20
    • Wu Qiqi, Deng Xing, Shao Haijian, Wang Fei
      Vol. 31, Issue 7, Pages: 2569-2586(2026) DOI: 10.11834/jig.250488
      ViBound-Net: a hybrid attention network for colorectal polyp segmentation integrating visual and boundary-focused mechanisms
      摘要:ObjectiveAccurate delineation of colorectal polyps in colonoscopy images is a key prerequisite for a computer-aided diagnosis (CAD) in colorectal cancer (CRC) screening, as adenomatous polyps are common precursors of CRC, and their timely detection and removal can substantially reduce incidence and mortality. However, in routine endoscopy, automated polyp segmentation remains challenging for three main reasons. First, polyp boundaries are frequently indistinct: subtle intensity transitions, mucus coverage, and folds of surrounding mucosa can obscure lesion contours, making boundary localization error-prone. Second, polyp appearance markedly varies across size, shape, texture, and attachment morphology. This diversity complicates stable feature representation and can trigger over-segmentation (including mucosal folds) or under-segmentation (missing flat or small lesions). Third, colonoscopy imaging often exhibits non-uniform illumination, specular highlights, and lens-related artifacts. These factors distort local contrast and degrade model robustness. Meanwhile, practical deployment introduces additional constraints: endoscopy systems typically process high-resolution streams at high frame rates and must respond with low latency on resource-limited hardware. Accordingly, the objective of this study is to develop a lightweight yet accurate polyp segmentation model that 1) balances global semantic understanding with fine-grained boundary detail, 2) achieves efficiency suitable for routine clinical workflows under deployment constraints, and 3) generalizes robustly across heterogeneous datasets collected with different devices and acquisition conditions.MethodWe propose ViBound-Net, a lightweight hybrid-attention segmentation network designed to jointly emphasize semantic context and boundary fidelity while maintaining a compact computational footprint. The network adopts EfficientNetV2S as a lightweight encoder to provide strong feature extraction with favorable parameter-compute efficiency. We introduce a local-global contextual attention fusion (LGCAF) module that integrates fine structural cues (effectively captured by local receptive fields) with global contextual dependencies (captured via attention) to address the local-global trade-off that commonly limits purely convolutional or purely attention-based approaches. This fusion aims to improve semantic consistency in challenging backgrounds while preserving lesion-specific structural patterns that are essential for small or camouflaged polyps. We further design a boundary-guided scale-aware feature enhancement (BSFE) module to explicitly strengthen boundary modeling. BSFE leverages multi-scale representations and boundary-aware guidance to refine edge-related features, targeting failure modes, such as contour breaks, boundary blurring, and partial omission around low-contrast margins. In decoding, we utilized a multi-scale residual attention fusion block to aggregate cross-scale information while stabilizing optimization through residual pathways. This design encourages coherent reconstruction of polyp regions across scales and reduces sensitivity to size variability. In addition, compact channel-spatial attention is incorporated to promote feature consistency and adaptively emphasize salient lesion cues without introducing heavy computation. Deep supervision is used to guide intermediate predictions and improve convergence stability, particularly for boundary-sensitive learning. Model training adopts a compound Dice + focal loss to mitigate foreground-background imbalance and enhance learning stability under diverse lesion sizes and appearance variations. Images are resized to 352 × 352, normalized, and augmented using random scaling, flipping, and rotation to improve robustness to viewpoint and scale changes commonly observed in endoscopy. Optimization follows a unified protocol using AdamW (learning rate 1 × 10⁻⁴, batch size 4) to ensure a fair comparison across methods. Experiments are conducted on three widely used public colonoscopy benchmarks——Kvasir-SEG, CVC-ClinicDB, and CVC-300——which collectively cover diverse resolutions, polyp sizes and shapes, and background complexity, for evaluation. We report segmentation accuracy using Dice similarity coefficient (Dice) and mean intersection over union (mIoU). Moreover, we report the mean absolute error (MAE) to effectively reflect boundary sensitivity and contour adherence. Given that deployability is an explicit goal, we assess the model efficiency through parameter count, FLOPs, and per-image inference time under the same experimental protocol and input scale. Finally, we provide qualitative visual comparisons to analyze contour adherence, over-/under-segmentation behaviors, and robustness under low contrast, specular highlights, and lesions that resemble surrounding mucosa.ResultUnder the unified evaluation protocol, ViBound-Net achieves strong performance across all three benchmarks. Specifically, this model attains mDice/mIoU of 0.908 1/0.852 7 on Kvasir-SEG, 0.936 8/0.887 8 on CVC-ClinicDB, and 0.871 0/0.798 9 on CVC-300. In boundary-sensitive evaluation on the ETIS-LaribPolypDB dataset, ViBound-Net reports an MAE of 0.011 7, indicating improved contour sharpness and faithful delineation. These results suggest that the proposed local-global fusion and boundary-guided enhancement effectively reduce common segmentation artifacts, such as contour leakage into mucosal folds, fragmented edges around flat polyps, and missing regions under low contrast. ViBound-Net delivers competitive or superior accuracy on most metrics while maintaining a compact design compared with strong convolutional and transformer-based baselines under the same protocol. For example, on Kvasir-SEG, this model improves over M2SNet in mDice and mIoU by 2.6 and 3.5 percentage points, respectively, and it reduces boundary error relative to DUCK-Net, reflecting a reliable edge localization. Across datasets, qualitative comparisons show that ViBound-Net produces smooth, better-adhered contours with few over-/under-segmentation artifacts, particularly for small lesions, low-contrast polyps, and cases affected by specular highlights. From an efficiency perspective, ViBound-Net consistently demonstrates lower computational burden than many transformer-heavy counterparts (as reflected by fewer parameters and reduced FLOPs), supporting deployment-oriented constraints where latency and memory are critical. Overall, these results indicate that the proposed architecture provides a favorable accuracy-efficiency trade-off, aligning with the practical requirements of routine endoscopy systems.ConclusionThis study presents ViBound-Net, a lightweight hybrid-attention network for colorectal polyp segmentation that explicitly targets the simultaneous demands of semantic reliability, boundary precision, and deployment efficiency. ViBound-Net achieves a high segmentation accuracy and a low boundary error on three public benchmarks while remaining computationally efficient by combining a compact EfficientNetV2S encoder with a LGCAF module and a boundary-guided scale-aware feature enhancement module and reinforcing multi-scale decoding via residual attention fusion and compact channel-spatial attention. These findings support the feasibility of integrating ViBound-Net into CAD pipelines for CRC screening, where stable boundary delineation can assist downstream clinical tasks, such as polyp characterization, measurement, and resection guidance. Future work will extend the method in several directions: 1) incorporating temporal cues to exploit video continuity and minimize frame-to-frame inconsistency in real endoscopy streams; 2) improving cross-device and cross-center robustness through explicit domain adaptation strategies; 3) enhancing interpretability to effectively communicate model decisions to clinicians; and 4) conducting prospective clinical validation to quantify the workflow-level effect, including time savings, detection consistency, and potential reduction of missed lesions in routine practice.  
      关键词:medical image segmentation;colorectal polyp segmentation;hybrid attention network;boundary guidance;multi-scale feature enhancement;lightweight model   
      117
      |
      390
      |
      0
      <HTML>
      <L-PDF><WORD><Meta-XML>
      <引用本文> <批量引用> 146770240 false
      更新时间:2026-07-20
    • Edge-guided multiperception decoder for medical image segmentation AI导读

      Huang Guanhui, Liu Xiang, Shi Yunyu, Ji Yu, Wang Shuohong
      Vol. 31, Issue 7, Pages: 2587-2600(2026) DOI: 10.11834/jig.250251
      Edge-guided multiperception decoder for medical image segmentation
      摘要:ObjectiveMedical image segmentation refers to the process of accurately separating the regions of interest in medical images from the background to extract key information, such as organs and lesion areas, and provide support for subsequent diagnosis, analysis, evaluation, and treatment. Benefiting from the widespread application of deep learning, especially convolutional neural networks and Transformer frameworks, it has evolved from using traditional manual methods to automated and universal intelligent segmentation. Many AI-assisted diagnosis and treatment systems are used in practice. Compared with simple pixel computing, existing network structures can recognize various biological information well by extracting and classifying image features, and automated segmentation has achieved substantial progress. However, challenging issues, such as how to solve the difficulty of edge acquisition, how to fully utilize the obtained features, and how to determine which features among multiple ones play a more important role in segmentation judgment, remain. With the deepening of the network layer, the loss of high-frequency detail information leads to blurred and diffused image edges. Reasonable utilization and accurate fusion of multilevel and multiscale feature information directly affect the accuracy of image segmentation. To address these issues, we introduce three modules and propose a novel edge-guided multiperception decoder for medical image segmentation (EMPD). The three modules are edge information perception module, dual-scale cascade gating, and three-dimensional(3D) attention fusion module.MethodUsing EMPD as a decoder after some high-performance encoders can effectively improve the model’s segmentation accuracy. The input image first enters the three-dimensional(3D) attention fusion module, which utilizes the encoder input features to interactively fuse positional and channel features. By rearranging the attention mechanism at the pixel location, features from the position, channel, and pixel dimensions are interactively combined to generate a three-dimensional(3D) attention map. This attention map is then applied to the input to extract features at multiple scales for feature enrichment. These features are interactively combined with encoder information from the previous layer through two-scale cascade gating. In this part, we simultaneously extract and interactively merge information from different scales and layers to highlight primary information and suppress secondary information. An edge-aware module processes the original image and performs corresponding downsampling and edge extraction operations at different scales. During the progressive upsampling process, multiple operators extract edge information from different perspectives. This combination of multiple edge information, coupled with the introduction of learnable variables, makes our operator more informative, more learnable, and less expensive than others. The final learned result is inputted into the next layer of three-dimensional(3D) attention fusion module, and the multilayer predicted image is combined with the MUTATION loss to obtain the final segmentation result.ResultThis study conducted experimental evaluations on seven datasets representing three medical image segmentation tasks. The datasets are ClincDB, ColonDB, CVC-300, Kvasir, and EITS datasets for single-class polyp segmentation; the ISIC2017 dataset for single-class skin lesion segmentation; and the Synapse dataset for multiorgan segmentation. The proposed method achieved a 1.58% improvement in segmentation performance on the ClinicDB dataset compared with state-of-the-art methods. On the five Polyp and ISIC2017 datasets, the Dice score improved by an average of 0.77% compared with the best-performing network. On the Synapse dataset, the Dice score and mean intersection over union improved by 2.16% and 3.81%, respectively, compared with PVT Cascade. In the segmentation of the left and right kidneys, the Dice score improved by 2.92% and 2.50%, respectively. In the large stomach region, the proposed method achieved a 3.28% improvement compared with PVT Cascade. Ablation experiments conducted on this architecture demonstrated that on multiple datasets, the Dice coefficient improved by more than 3 percentage points compared with a simple encoder cascade architecture. Introducing edge information and combining it with an effective fusion strategy substantially improved the problem of poor edge fitting and enhanced the overall segmentation accuracy, proving the effectiveness of the proposed module. However, having abundant edge information is not always good. The number and distribution of edge-aware modules need to be controlled. If the image depth is too large, a small number of learnable parameters may be insufficient to compensate for the conflict between semantic and edge information, potentially leading to a decrease in accuracy.ConclusionEMPD, combined with a high-performance encoder, achieved optimal performance on the five datasets, and its results approached the current best results on the two remaining datasets. Its performance surpassed that of mainstream image segmentation methods, and it demonstrated good generalization properties. This work provides new tools and model frameworks for medical image segmentation research, thus helping achieve effective medical-assisted diagnosis.  
      关键词:edge perception;dual-scale cascaded gating(DSCG);three-dimensional attention;feature fusion;medical image segmentation   
      312
      |
      417
      |
      0
      <HTML>
      <L-PDF><WORD><Meta-XML>
      <引用本文> <批量引用> 139328000 false
      更新时间:2026-07-20
    • Wang Xuan, Deng Na, Xie Feng, Xing Wenyu
      Vol. 31, Issue 7, Pages: 2601-2613(2026) DOI: 10.11834/jig.250510
      DMA-UNet:dual-domain multiscale attention analysis network for hemorrhage regions in abdominal ultrasound images
      摘要:ObjectiveAbdominal hemorrhage is a critical medical condition characterized by a dangerous clinical course, making rapid and accurate diagnosis imperative. This study addresses the challenges in ultrasound image segmentation for abdominal hemorrhage, including low contrast, high noise, and blurred boundaries, which often hinder the precise delineation of bleeding regions. To overcome these limitations, we propose a novel dual-domain and multiscale attention U-Net (DMA-UNet) that integrates spatial-and frequency-domain features for enhanced segmentation of abdominal hemorrhage in ultrasound images. This life-threatening condition typically results from trauma, ruptured aneurysms, or postoperative complications. Its abrupt onset and rapid progression can swiftly lead to hemorrhagic shock and multiple organ failure, rendering time a decisive factor in determining patient outcomes. In such critical scenarios, ultrasonography has become an indispensable first-line diagnostic tool because of its portability, absence of radiation exposure, and capability for rapid bedside assessment. Accurately identifying and quantifying the extent of hemorrhage on ultrasound scans directly influence clinical decision-making, such as choosing between conservative management or urgent surgical intervention, and provides crucial prognostic information. However, achieving precise automated segmentation of hemorrhage regions presents multiple inherent challenges. Ultrasound images are inherently characterized by low signal-to-noise ratios, pervasive speckle noise, poor tissue contrast, and various artifact interferences. The heterogeneity of pathological manifestations further compounds this difficulty: hemorrhage areas may present as anechoic, hyperechoic, or mixed echogenicity patterns, often with irregular and poorly defined boundaries that blend with surrounding parenchymal or adipose tissues. This combination of technical limitations and pathological variability makes traditional manual segmentation methods time consuming and operator dependent and subject to substantial interobserver variability. Moreover, complex factors in the clinical examination environment pose additional challenges. Ultrasound examinations are frequently conducted under suboptimal conditions, including patient movement, limited acoustic windows, uneven probe pressure, and intestinal gas interference, all of which considerably degrade image quality and consistency. In emergency and intraoperative settings, the inherent tension between diagnostic speed requirements and accuracy demands creates substantial pressure on medical personnel. Although deep learning techniques, particularly convolutional neural network architectures represented by U-Net and its variants, have achieved remarkable progress in medical image analysis for modalities, such as CT and MRI, their direct application to abdominal ultrasound image segmentation continues to face fundamental challenges. These models are typically designed for images with high contrast and standardized spatial characteristics. When processing ultrasound data, they often exhibit limited receptive fields, insufficient global contextual relationship modeling, and inefficient multiscale feature fusion. This phenomenon results in suboptimal performance when segmenting diffuse, low-contrast hemorrhage regions and frequently leads to under-segmentation of scattered bleeding areas or over-segmentation of anatomically similar soft tissues.MethodTo bridge this technological gap, we propose a sophisticated deep learning architecture called dual-domain multiscale attention network (DMA-Net). This framework is meticulously designed around three synergistic and innovative core modules, each specifically optimized to address the distinct limitations of conventional segmentation models when processing complex ultrasound data. The core architecture incorporates the following key innovative modules. 1) Multiscale atrous attention module in the encoder: This module employs a parallel multibranch architecture that systematically expands the receptive field through atrous convolutional layers with varying dilation rates while adaptively calibrating feature weights via dual channel-spatial attention mechanisms. Each branch integrates contextual information at different scales, thereby enhancing the sensitivity to subtle hemorrhage regions while preserving feature map resolution. A residual attention pathway is specifically designed to retain original texture details. The module dynamically suppresses speckle noise interference through learnable spatial masks, and its hierarchical dilation design captures local vascular structures and global bleeding distribution patterns. This multiscale fusion strategy substantially improves the detection accuracy of early-stage and scattered hemorrhage foci under challenging acoustic conditions. 2) Frequency-domain global modeling module in the bottleneck layer: This module implements a frequency-spatial dual-stream architecture that transforms feature maps into the spectral domain via fast Fourier transform, enabling holistic analysis of hemorrhage distribution patterns. In the frequency pathway, learnable spectral filters suppress ultrasound-specific artifacts while amplifying hemorrhage-related frequency components. The spatial pathway maintains convolutional operations to preserve anatomical details. Through residual connections and frequency-domain attention gates, the module selectively enhances cross-region dependencies that are critical for identifying discontinuous bleeding patterns. This approach effectively bridges the gap between localized feature extraction and global contextual understanding and is particularly valuable for segmenting diffuse hemorrhage regions with irregular boundaries. The integrated design achieves computational efficiency while remarkably improving the model’s ability to interpret complex acoustic signatures in abdominal ultrasound imaging. 3) Cross-level joint attention fusion module: a dynamic feature selection mechanism is introduced in the encoder-decoder skip connections. Through multiscale feature cross-computation, it generates adaptive attention weights that intelligently filter semantic features across different levels, suppress noise transmission, and promote the optimized fusion of deep semantic information with shallow spatial details for accurate boundary reconstruction.ResultThe experimental framework was established using a clinically heterogeneous dataset comprising 363 intra-abdominal hemorrhage ultrasound images sourced from two independent medical centers to ensure anatomical and acquisition variability. A comprehensive data augmentation pipeline was implemented to enhance model robustness and generalizability; this pipeline incorporates spatial transformations (rotation, scaling, and elastic deformation), intensity variations, and speckle noise simulation. Quantitative evaluation on hold-out test sets from both institutions demonstrated the model’s strong segmentation performance and cross-center adaptability. For Center A, the model achieved a Dice similarity coefficient of 0.879 7 and an intersection over onion (IoU) of 0.796 1. For Center B, the Dice similarity coefficient and IoU reached 0.933 9 and 0.876 2, respectively, reflecting consistent reliability across different imaging protocols and patient populations.ConclusionDMA-UNet introduces an ultrasound-adapted segmentation framework that effectively integrates spatial- and frequency-domain representations through coordinated attention mechanisms. By jointly enhancing global context modeling, multiscale feature fusion, and imaging artifact resilience, the network offers a clinically viable solution for rapid and accurate hemorrhage assessment in critical care. Its strong performance across independent datasets underscores its translational potential for use in bedside decision support systems.  
      关键词:Abdominal hemorrhage;fast Fourier convolution(FFC);Multi-scale representation;joint attention structure(JAS);Ultrasound image segmentation   
      212
      |
      331
      |
      0
      <HTML>
      <L-PDF><WORD><Meta-XML>
      <引用本文> <批量引用> 146822182 false
      更新时间:2026-07-20

      Image Analysis and Recognition

    • Liu Dian, Yin Na, Zhang Zening, Lyu Hongyan, Luo Xubin, Wang Haoxiang
      Vol. 31, Issue 7, Pages: 2614-2630(2026) DOI: 10.11834/jig.250354
      Highway parking violation detection by integrating dual-attention and spatiotemporal features
      摘要:ObjectiveAutomatic detection of illegal parking on highways through surveillance video is a crucial task for modern traffic management systems because it directly affects traffic flow, public safety, and operational efficiency. However, detecting parking violations in highway environments presents substantial challenges because of complex traffic dynamics, varying lighting conditions, and broad surveillance coverage. These challenges include frequent vehicle occlusions in dense traffic that lead to identity confusion and incorrect tracking, substantial perspective distortion that causes slow-moving vehicles to be mistakenly identified as parked, and degraded detection performance under low-illumination conditions, where vehicle features become blurred and are prone to false detection because of environmental lighting effects.MethodTo address these problems, this study proposes a novel highway illegal parking detection algorithm that integrates dual-attention feature fusion and spatiotemporal behavior analysis. The approach introduces a dual-attention fusion module to enhance multiscale feature representations extracted from YOLOv8, which are used as appearance descriptors in a vehicle tracking framework on the basis of BotSORT-ReID. This framework directly leverages YOLOv8’s multiscale feature outputs and enhances them with attention mechanisms in spatial and channel dimensions. The spatial attention module emphasizes critical regions related to vehicle targets, suppressing background noise, whereas the channel attention mechanism selectively emphasizes discriminative feature channels relevant to object identity. These enhancements substantially improve the robustness and accuracy of vehicle tracking, especially in scenarios involving occlusion, background clutter, or varying viewing angles. The second component of the algorithm focuses on violation judgment through spatiotemporal feature accumulation. Specifically, it uses an adaptive interframe displacement estimation method on the basis of exponentially weighted moving averages to identify low-speed or stationary vehicles over a sequence of frames. A normalized displacement model is employed to correct perspective distortions across different regions of the image, ensuring that movement estimates are reliable even when vehicle scale varies with the distance from the camera. This method smooths short-term fluctuations in object position and helps avoid misclassification caused by detection box jitter or gradual deceleration. A secondary validation mechanism is proposed to address the complexity of nighttime detection, especially in environments with poor visibility or interference from artificial lights. This process involves analyzing the temporal characteristics of brightness changes to identify flashing hazard lights, which are a strong indicator of stationary vehicles in emergency situations. Time-series analysis is conducted on grayscale intensity values extracted from target bounding boxes across frames. Preprocessing steps, including Savitzky-Golay filtering, detrending, and temporal gradient computation, are applied to isolate periodic brightness variations. Subsequently, Welch’s method is used for power spectral density estimation, and dominant frequencies are extracted and compared with known hazard light flashing frequency ranges. Only when these frequencies match expected patterns is the object confirmed to be illegally parked, thereby reducing false positives from nonvehicle light sources or reflections.ResultGiven the absence of publicly available datasets for highway parking violation detection, three dedicated highway datasets are constructed for object detection, vehicle tracking, and parking violation recognition. The datasets consist of 8 532 labeled images for detection, 34 600 frames for tracking, and 77 annotated video sequences (average over 1 min, covering 98 illegal parking vehicles) for violation recognition. Moreover, they cover diverse conditions, including daytime, nighttime, tunnels, and adverse weather (e.g., rain and fog), thereby ensuring comprehensive evaluation across complex highway scenarios. Experimental results on the constructed datasets show that the proposed algorithm consistently outperforms several benchmark methods in terms of precision, recall, and F1 score. Notably, dual-attention feature extraction improves identity consistency in vehicle tracking, thereby reducing ID switches and increasing tracking accuracy. The adaptive displacement method demonstrates strong resistance to jitter and perspective artifacts, and the hazard light detection module substantially reduces nighttime false positives and improves robustness under difficult lighting. Ablation studies further confirm the contribution of each component. For instance, removing the dual-attention module results in reduced tracking quality and an increase in identity confusion. Without the adaptive displacement method, the system becomes highly vulnerable to misclassification because of minor motion or camera instability. Meanwhile, excluding the temporal signal-based hazard light detection leads to a rise in false alarms during nighttime evaluations. The full system achieves a balanced trade-off between high precision and high recall, demonstrating its effectiveness in real-world applications.ConclusionOverall, the proposed algorithm presents a unified, efficient, accurate solution for illegal parking detection on highways. By combining deep learning-based object detection, attention-guided feature fusion, motion pattern recognition, and signal-based validation, this algorithm addresses the key limitations of current methods and achieves robust performance across varying highway surveillance conditions. The integration of lightweight attention modules into the YOLOv8 backbone and the design of spatiotemporal behavioral analysis allow for real-time processing while maintaining high accuracy. The system is thus well-suited for deployment in intelligent transportation systems and smart city infrastructure where reliability and adaptability are paramount. The methodology also provides a scalable framework for extending similar detection strategies to other traffic violations or abnormal behaviors in complex road environments.  
      关键词:highway;spatiotemporal feature;attention feature;parking violation detection;dual attention   
      132
      |
      245
      |
      0
      <HTML>
      <L-PDF><WORD><Meta-XML>
      <引用本文> <批量引用> 144739999 false
      更新时间:2026-07-20
    • Yang Yuqing, Chen Yi, Lyu Cheng
      Vol. 31, Issue 7, Pages: 2631-2644(2026) DOI: 10.11834/jig.250398
      FixMatch++: a method for expanding limited image label data on the basis of semi-supervised learning
      摘要:ObjectiveDeep visual recognition models have achieved remarkable success in diverse domains ranging from everyday object classification to specialized tasks, such as medical image analysis. However, their performance remains heavily dependent on large-scale labeled datasets, the construction of which is costly, time consuming, and often impractical in real-world scenarios. Semi-supervised learning (SSL), which leverages abundant unlabeled data to compensate for limited labeled samples, has therefore emerged as a promising paradigm. Among existing SSL methods, FixMatch is a widely adopted baseline because of its simple yet powerful combination of consistency regularization and pseudo-labeling, driven by weak-strong augmentation. However, FixMatch still faces notable limitations: feature extraction instability during early training, entanglement of content and style caused by acquisition or imaging variations, and noise accumulation from single-view pseudo-labels. These issues reduce the robustness and generalizability of FixMatch in realistic deployment scenarios.MethodTo address these challenges, we propose FixMatch++, an enhanced framework that introduces three tightly integrated improvements. The first improvement is learnable-shift batch normalization (LS-BN) and dual-scale parallel convolution (DSPC). Conventional batch normalization assumes stable channel statistics, but the aggressive augmentations employed in FixMatch frequently distort these distributions, introducing bias. We design LS-BN that adaptively corrects augmentation-induced deviations, and DSPC incorporates parallel convolutional branches with different receptive fields, thus enabling rich multiscale representations and improving robustness to variations in object size and texture. The second improvement is content-style dual-branch representation with dynamic residual gating. Natural images often embed intertwined structural (content) and appearance (style) factors. A single-stream backbone entangles these signals, reducing discriminative power. We disentangle features into two dedicated branches: one focusing on content (shape and semantics) and the other on style (texture, illumination, and acquisition artifacts). A dynamic residual gating mechanism adaptively fuses the two streams and selectively emphasizes informative features under varying conditions, thereby stabilizing representation learning and enhancing generalizability. The third improvement is multilevel pseudo-label fusion. Unlike FixMatch that generates pseudo-labels from a single augmented view, FixMatch++ aggregates predictions from weak, medium, and strong augmentations. Their class probabilities are fused through weighted averaging, ensuring that noisy predictions from any single view are mitigated by corroborating evidence from others. In addition, we introduce a class-wise confidence thresholding strategy, which applies tailored thresholds for different categories to prevent dominant classes from overwhelming minority ones and to ensure balanced pseudo-label adoption. Together, these mechanisms substantially increase the quantity and quality of pseudo-labels, which are critical for effective semi-supervised training.ResultWe conduct extensive experiments on three widely used SSL benchmarks: CIFAR-10, CIFAR-100, and SVHN. We benchmark FixMatch++ against seven state-of-the-art baselines, namely, MeanTeacher, UDA, ReMixMatch, FixMatch, FlexMatch, SoftMatch, and UES. Across all datasets, FixMatch++ consistently achieves the lowest classification error rates among the approaches. Gains are particularly pronounced in low-label regimes, which best reflect real-world annotation scarcity. For instance, under the challenging CIFAR-10 setting with only 250 labeled samples, FixMatch++ reduces the error rate to 4.56%, representing an absolute improvement of 0.35%—27.76% compared with the seven baselines. On CIFAR-100, which involves 100 categories and is highly susceptible to interclass confusion, FixMatch++ demonstrates clear adaptability, outperforming FlexMatch and ReMixMatch. On SVHN, which features digit recognition in complex street-view conditions with substantial imaging variability, our framework surpasses UDA and SoftMatch, validating its robustness to acquisition artifacts. We conduct ablation studies to further validate the contribution of each proposed component. Results show that removing learnable-shift batch normalization considerably increases the error rates, confirming the necessity of correcting augmentation-induced biases. Excluding the content-style disentanglement module leads to unstable training and increased sensitivity to domain shifts, revealing the importance of separating and adaptively fusing features. Eliminating multilevel pseudo-label fusion degrades label quality and amplifies noise, particularly at the late training stages, thereby verifying the effectiveness of the fusion strategy. We also perform visualization analyses, including feature map inspection, class activation mapping, and pseudo-label confidence distribution plots. These visualizations provide intuitive evidence of how FixMatch++ stabilizes feature extraction, disentangles factors of variation, and generates reliable pseudo-labels. Together, the analyses confirm not only the empirical improvements but also the interpretability of the model’s internal processes.ConclusionIn summary, FixMatch++ substantially advances semi-supervised image classification by introducing improvements at the architectural and algorithmic levels. The learnable-shift batch normalization and dual-scale convolution modules enhance the robustness of feature extraction. The content-style dual-branch representation with dynamic residual gating disentangles and adaptively fuses complementary features, and the multilevel pseudo-label fusion mechanism ensures high-quality supervision from unlabeled data. Comprehensive experiments and analyses demonstrate that these improvements yield consistent and notable performance gains over strong SSL baselines across multiple datasets and labeling regimes. Owing to the modular design of FixMatch++, the proposed components can be seamlessly integrated into other SSL frameworks, extending the applicability and influence of FixMatch++. FixMatch++ enlarges the effective training set under limited-label conditions and strengthens the reliability and interpretability of SSL. It offers a practical solution for real-world scenarios where annotation resources are scarce yet robust performance is essential.  
      关键词:image classification;semi-supervised learning(SSL);weakly labeled data;data augmentation;pseudo-label fusion;cross-domain adaptation   
      143
      |
      354
      |
      0
      <HTML>
      <L-PDF><WORD><Meta-XML>
      <引用本文> <批量引用> 143378753 false
      更新时间:2026-07-20
    • Co-salient object detection with automatic semantic discovery AI导读

      Chen Huachao, Chen Yingzhou, Wang Bin, Wang Feng, Zhao Jia
      Vol. 31, Issue 7, Pages: 2645-2659(2026) DOI: 10.11834/jig.250377
      Co-salient object detection with automatic semantic discovery
      摘要:ObjectiveCo-salient object detection (Co-SOD) aims to identify target regions from a set of related images that are not only visually salient but also semantically consistent across the group. The core challenge of Co-SOD lies in detecting visually prominent objects within individual images and determining which objects exhibit cross-image consistency among multiple images. Traditional Co-SOD methods mainly rely on low-level visual features to construct inter-image relationships. These approaches extract discriminative representations on the basis of handcrafted features, such as color, shape, and texture, to mine common salient objects. Although such methods perform reasonably well under simple backgrounds, they often fail in complex real-world scenarios with strong background clutter, severe occlusions, or multiple coexisting objects. In these cases, non-co-salient but visually salient objects can easily interfere with detection, leading to high false detection rates, poor robustness, and limited generalization. In recent years, with the rapid progress of deep neural networks, Co-SOD methods that are based on convolutional neural networks or vision Transformer have been proposed. By learning high-level semantic features across images, these methods substantially improve detection performance. However, they typically still rely on modeling image groups as the basis of detection without sufficiently incorporating explicit semantic priors. As a result, they lack the capability for proactive understanding of image categories or user-specific targets and therefore struggle in open-domain, multiobject scenarios or when fine-grained semantic distinctions are required. Moreover, when an image contains multiple visually salient but semantically irrelevant objects, these models often fail to distinguish them correctly, mistakenly including non-co-salient objects in the detection results, thus reducing the overall performance. For instance, in multiclass image datasets, such as CoCA, a single image often contains multiple visually prominent objects, but only a subset of them is semantically related and constitute co-salient targets. Other salient regions, although visually striking, act as semantic distractors. Without explicit semantic filtering mechanisms, traditional methods are often unable to distinguish which objects are the true co-salient targets, thus decreasing the accuracy of Co-SOD. To address these limitations, this study proposes a semantic-guided open-vocabulary Co-SOD framework. The proposed method leverages language prompts to guide the model’s attention toward target regions that are semantically aligned with image categories, thereby enabling semantic filtering of co-salient objects in complex multiobject scenarios and improving detection accuracy and semantic generalization ability.MethodThe proposed semantic-guided open-vocabulary Co-SOD framework integrates three key components, namely, a CLIP-based group-consistent category generation method, the open-vocabulary object detection model OWLv2, and the universal segmentation model SAM, forming a complete pipeline that begins from semantic prompt generation to object detection and pixel-level segmentation. First, the system introduces a CLIP-based group-consistent category generation method to automatically derive the most representative group-consistent category for a given image set. This procedure eliminates the need for manually specifying class names, reduces human effort, and provides accurate and stable semantic prompts for subsequent detection. Specifically, the method computes cross-modal similarities between the visual features of each image in the group and candidate category text features and automatically generates the most representative category name that is consistent across the group to serve as a reliable input for the detection model. Second, the generated group-consistent category name is used as a natural language prompt and fed into the open-vocabulary object detection model OWLv2. OWLv2, which is pretrained on large-scale image-text pairs, possesses strong cross-modal semantic alignment and matching capabilities. By mapping image region and text prompt features into a unified semantic space, OWLv2 can identify target regions that are highly relevant to the natural language prompts. To further improve the robustness of semantic understanding, this work constructs multiple semantically related natural language phrases for each category name and adopts a multiprompt parallel input mechanism to feed them into OWLv2 simultaneously. This process enhances the model’s capacity to capture diverse semantic descriptions, improves stability and robustness in semantic matching, and effectively addresses uncertainties caused by language ambiguity. During the candidate region generation stage, similarities between region feature and text semantic vectors are calculated to identify candidate regions that are highly related to the input prompts, with the model outputting corresponding bounding boxes and confidence scores. A dynamic thresholding strategy is introduced to further improve detection flexibility and adaptability. This mechanism automatically adjusts detection thresholds on the basis of the confidence levels of semantic matches, thereby increasing the recall of low-confidence targets while maintaining precision. As a result, the framework can effectively detect occluded or weakly salient targets and avoid missed detections caused by overly strict thresholds. At the saliency segmentation stage, the bounding box regions filtered during the semantic detection stage are passed to segment anything model (SAM). SAM then performs pixel-level mask generation for each region of interest. Leveraging its powerful structural perception and boundary modeling capabilities, SAM produces high-quality segmentation results. The final system outputs salient masks that are structurally complete and have clear boundaries. This design effectively decouples the two subtasks of semantic perception and visual detail modeling, allowing each module to fully leverage its strengths, thereby enhancing the overall semantic guidance ability and object delineation quality of the system.ResultExperiments are conducted on three publicly available datasets: CoCA, CoSal2015, and CoSOD3k. The proposed method is compared with 15 state-of-the-art Co-SOD methods. On the CoCA dataset, our method achieves a 25.7% improvement in F measure and a 5.2% reduction in the mean absolute error (MAE) compared with the second-best model. On the CoSOD3k dataset, our method improves the F measure by 2.5%. On the CoSal2015 dataset, it improves the F measure by 2.2% and reduces MAE by 1.2%. These results comprehensively validate the strong ability of the proposed framework to recognize semantically relevant co-salient objects in multiobject open scenarios. Ablation experiments conducted on the CoSOD3k dataset further confirm the effectiveness of the proposed approach in enhancing Co-SOD performance.ConclusionThis study proposes a semantic-guided open-vocabulary Co-SOD framework. The framework leverages a CLIP-based group-consistent category generation method to automatically generate semantic prompts for image groups and integrates OWLv2 for open-vocabulary object detection with SAM for universal segmentation. This process enables explicit semantic filtering of co-salient targets and fine-grained pixel-level segmentation. The proposed framework substantially enhances semantic generalization and target recognition robustness, enabling the effective localization of semantically consistent salient objects in multiobject and complex background scenarios. By unifying three critical components, namely, semantic prompt generation, open-vocabulary detection, and saliency segmentation, this framework provides an innovative and practical solution for advancing Co-SOD in open environments.  
      关键词:co-salient object detection(Co-SOD);semantic guidance;open vocabulary;co-salient object detection framework;Cross-modal Similarity;dynamic threshold adaptation strategy   
      351
      |
      646
      |
      0
      <HTML>
      <L-PDF><WORD><Meta-XML>
      <引用本文> <批量引用> 139330173 false
      更新时间:2026-07-20

      Image Understanding and Computer Vision

    • Wang Jiacheng, Zhou Shubo, Pan Feng, Jiang Xueqin, Huang Rong, Xie Yinghua
      Vol. 31, Issue 7, Pages: 2660-2673(2026) DOI: 10.11834/jig.250353
      Two-stage mask-guided network for car surface specular highlight removal by using the shifted window attention mechanism
      摘要:ObjectiveSpecular highlight detection and removal, with the primary objective of accurately identifying highlight regions exhibiting mirror-like reflection characteristics and precisely recovering the underlying pixel information, constitute a fundamental research topic in computer vision and image processing. This technology demonstrates considerable practical value across various visual applications, including object detection, medical image analysis, industrial defect identification, and 3D imaging. From an optical perspective, the specular highlight phenomenon originates from the directional reflection behavior of incident light on smooth non-Lambertian surfaces, and its physical mechanism is theoretically describable through Fresnel equations. Traditional highlight removal methods typically rely on strict physical assumptions, which primarily encompass three classical approaches: 1) chromaticity space analysis-based techniques, 2) illumination estimation-based algorithms, and 3) polarization information-utilizing solutions. However, these methods exhibit limited adaptability to complex and variable highlight removal requirements in real-world scenarios because of their constraints regarding specific scenes or object types. By contrast, deep learning-based highlight detection and removal techniques adopt a data-driven approach to automatically learn the complex nonlinear mapping relationships between highlight reflections and scene content, thereby effectively circumventing the dependency on prior assumptions inherent in traditional methods. The shifted window attention mechanism demonstrates unique advantages in highlight removal tasks and effectively addresses issues, such as texture information loss and color distortion, caused by specular highlights. This mechanism integrates local window partitioning with cross-window interaction, thus reducing computational complexity while maintaining global modeling capability. Building upon the exceptional modeling capability demonstrated by the shifted window attention mechanism during the feature extraction stage, this study proposes a two-stage mask-guided highlight removal network architecture. The proposed framework employs the shifted window attention mechanism to extract multiscale cross-channel features and utilizes highlight masks as prior knowledge to guide subsequent pixel recovery in highlight regions, thereby achieving highlight localization and removal. Notably, existing deep learning methods face challenges regarding insufficient training data. Current publicly available highlight datasets predominantly employ synthetic generation or laboratory-controlled lighting acquisition methods, failing to meet the requirements of real-world applications. To address this limitation, this study additionally constructs an in-situ automotive highlight dataset captured under multiple illumination conditions. This dataset has the following characteristics: 1) coverage of diverse natural illumination conditions, 2) inclusion of complex background interference, and 3) precise pixel-level annotations, thereby providing a data foundation for enhancing model generalization performance in real-world scenarios.MethodIn this study, we construct a real highlight dataset with cars as the main object, which is based on common cars in daily life with different highlight regions and intensities. Each pair of highlight images has its corresponding real highlight mask image. For highlight detection and removal, we propose a mask-guided two-stage highlight removal network (MG-TransUNet). This network explicitly divides the highlight detection and removal task into two phases. In the first phase, we train with a lightweight U-Net architecture to extract the highlight regions in the highlight images and generate the highlight mask map. In the second phase, a shifted window attention Transformer network that is based on the U-Net architecture is designed to recover the diffuse images of highlight regions by using the predicted highlight mask map as an a priori guide. Our experimental dataset is selected from three major public datasets (i.e., paired specular-diffuse image dataset(PSD), specular highlight image quadruples(SHIQ) and synthetic specular highlight removal dataset(SSHR)) with our proposed automotive dataset. Among the four datasets, only SHIQ and automobile datasets provide highlight masks. To make our method trainable on the PSD and SSHR datasets, we subtract and threshold the highlight images and the corresponding specular no-highlight images from the two datasets to obtain the highlight mask for each pair of images. The highlight detection network is trained with the highlight images and masks at the first stage, and the obtained highlight masks are used as a priori information to guide the highlight removal task at the second stage. Our mask-guided highlight removal network is implemented by the Pytorch toolbox. The input images and ground truth maps are resized to 256 × 256 for training, the momentum parameter is set to 0.99, and the initial learning rate is set to 2 × 10-4, which is reduced using a decay coefficient of 0.8 every 5 cycles until it reaches 1 × 10-6. The stochastic gradient descent learning process is accelerated using an NVIDIA GeForce GTX3090 GPU device, and 100 iterations requires about half a day.ResultExperiments are conducted on the proposed automotive dataset and three publicly available datasets to compare eight state-of-the-art methods. The quantitative evaluation metrics include the structural similarity index measure and peak signal-to-noise ratio (PSNR). Experimental results show that the PSNR value ranks second on the PSD dataset. It is improved by 1.11 compared with that in the model with second-best performance on the SHIQ dataset and by 0.38 compared with that in the model with the second-best performance on the SSHR dataset. Comparative experiments are also performed on automotive datasets to validate the effectiveness of the proposed dataset.ConclusionWe propose a car highlight dataset for real scenarios and a mask-guided network incorporating a window-shift attention mechanism to perform the highlight removal task. The experimental results show that our highlight removal network can accurately recognize highlight regions and remove highlight images. The code link: https://github.com/chenWULUQI/MG-Trans_Unet.  
      关键词:highlight detection;highlight removal;highlight dataset;U-Net model;shifted window attention mechanism   
      71
      |
      170
      |
      0
      <HTML>
      <L-PDF><WORD><Meta-XML>
      <引用本文> <批量引用> 151835511 false
      更新时间:2026-07-20

      Virtual Reality and Augmented Reality

    • Meng Qifan, Zhang Yiran, Zhang Hongwen, Hu Xiaoyan, Luo Yanhong
      Vol. 31, Issue 7, Pages: 2674-2686(2026) DOI: 10.11834/jig.250425
      Virtual embodied interaction technology for virtual-reality physics experiments
      摘要:ObjectiveThe continuous iteration of virtual reality (VR) technology has driven the digital transformation of education, and its application in physics experiment teaching has gradually become a research hotspot. However, existing consumer-grade VR devices are limited by their tracking nodes; they fail to realize natural and smooth full-body motion reconstruction, leading to insufficient user immersion, discontinuous interaction experiences, and even visual discomfort (e.g., motion sickness caused by conflicts between visual and vestibular perceptions). These issues severely restrict the application of VR systems in high-demand immersive physics experiment scenarios that require precise interactions and strong sense of presence. This study aims to 1) address the technical bottlenecks of poor full-body motion reconstruction and low interaction fidelity in current VR-based physics experiments; 2) optimize the biomechanical rationality, motion naturalness, and interaction accuracy of virtual embodiments while achieving dynamic proportion calibration for users with different body types and consistency adjustment for cross-platform models; 3) explore the mechanism by which virtual embodied interaction influences user experience, realizing the synergistic optimization of immersive experience enhancement and physiological comfort (i.e., reducing motion sickness); and 4) provide a theoretical basis and practical guidance for the future implementation of multiuser collaborative virtual embodied interaction in the field of immersive physics education.MethodTo achieve these objectives, this study proposes a virtual embodied motion-solving framework that integrates three core technologies: bone retargeting, inverse kinematics (IK), and data-driven motion matching. First, a hierarchical 3D human skeletal structure model is constructed. This model divides the human body into two biomechanically distinct subsystems, namely, spinal composite chains (pelvis-vertebrae-neck-head) and limb chains (upper limbs: shoulder-upper arm-lower arm-hand; lower limbs: pelvis-thigh-shank-foot), and adopts vertex skinning technology with blend shapes to avoid unnatural joint deformation caused by traditional linear blend skinning. For bone retargeting, a dual-calibration mechanism is designed. Global proportion calibration uses position data from head-mounted displays and hand controllers to calculate vertical and horizontal scaling factors; it adjusts the root joint position to match the user’s overall proportion while maintaining foot-ground contact. Local skeletal chain calibration adjusts the length of individual segments in five key chains (spinal chain, two arm chains, and two leg chains) while preserving internal length ratios to adapt to local body differences. For IK solving, a segmented strategy is adopted. An improved Gauss-Seidel algorithm (with joint angle constraints and multithreaded acceleration) is used for the head-spine composite chain to ensure stable convergence. A geometric analytical method with physiological limits (e.g., elbow flexion 0°-150°) is applied to limb chains to avoid nonphysical phenomena, such as limb penetration. Meanwhile, a hybrid hand control scheme (integrating repositioning and motion correction) on the basis of the OpenXR skeleton system is implemented to ensure fine interaction accuracy. For data-driven motion matching, an animation database is built using an Xsens motion capture system; this database covers diverse gaits, speeds, and turning motions with key features, such as joint position and foot contact state. Real-time frame matching (via hierarchical indexing) and motion smoothing (via inertial interpolation) are realized to compensate for insufficient lower-limb tracking data. For evaluation, 64 healthy college students from Beijing Normal University (31 males and 33 females aged 18-40, with no prior virtual embodiment experience) are stratified into an experimental group (using the improved virtual embodied interaction model) and a control group (using meta movement-based virtual embodiment) on the basis of scores from the visually induced motion sickness susceptibility questionnaire (VIMSSQ). The experiment is conducted on a hardware platform that has an Intel i7-12700F processor, NVIDIA RTX 3070 graphics card, and Meta Quest 3 HMD, and scenes are developed in Unity 2021.3.32f1c1 (frame rate > 90 frame/s to meet real-time requirements). Participants are asked to complete three basic tasks (i.e., sitting on a chair, picking up a cup, button interaction with left/right hands) and a momentum-conservation physics experiment. Objective data (task completion time and completion rate) and subjective data (obtained via the embodiment questionnaire (EQ), igroup presence questionnaire (IPQ), and simulator sickness questionnaire (SSQ)) are collected and analyzed using MATLAB R2022b, Origin 2024b, and SPSS 27 (including independent samples t-test, Mann-Whitney U test, Pearson correlation, and multiple linear regression).ResultComprehensive experimental results verify the effectiveness of the proposed framework. In terms of quantitative task performance, Task 1 (sitting) has both groups achieve 100% completion rate, with similar average times (experimental group: 32.25 s; control group: 32.85 s). In Task 2 (cup picking), the experimental group has a shorter average time compared with the control group (30.78 s vs. 32.00 s) and a stable completion rate (90% vs. 95%). Task 3 (button interaction) reveals the experimental group’s advantage, especially in nondominant left-hand operations (average time: 36.45 s vs. 38.36 s; completion rate: 70% vs. 55%). In terms of subjective experience, the EQ results show that the experimental group has significantly high scores in agency (p = 0.003, Cohen’s d = 1.028) and self-location (p = 0.037, Cohen’s d = 1.272) but no significant difference in ownership. The IPQ results indicate that the experimental group scores significantly high in general presence (p = 0.032), involvement (p = 0.012), and overall immersion (p = 0.045). Multiple linear regression (R² = 0.543, p < 0.001) shows that the experimental group has low SSQ scores (β = -0.495, p < 0.001); the IPQ scores negatively predict motion sickness (β = -0.493, p = 0.017). The interaction terms (VIMSSQ × Group, β = -0.365, p = 0.024; IPQ × Group, β = -0.326, p = 0.026) confirm that the framework weakens the effect of motion sickness susceptibility and strengthens immersion’s mitigating effect on discomfort.ConclusionThe proposed virtual embodied motion solving framework effectively addresses the technical limitations of current VR-based immersive physics experiments. By integrating bone retargeting, IK, and data-driven methods, it substantially enhances the biomechanical rationality (avoiding nonphysical phenomena, such as limb penetration), motion naturalness (via hierarchical skeletal modeling and inertial interpolation), and interaction accuracy (through hand optimization and local chain calibration) of virtual embodiments. The framework improves users’ embodiment sense (especially agency and self-location), immersion, and spatial motion perception and reduces visual discomfort via a dual-path mechanism (direct experience enhancement through improved presence; compensatory mitigation of motion sickness via strengthened immersion). Future research may focus on three directions: exploring high-precision and personalized virtual embodiment generation (e.g., physics experiment-specific motion libraries driven by large models), integrating physical laws and environmental factors to enhance interaction authenticity, and optimizing real-time performance for multiuser collaborative scenarios. This technology has broad application prospects in advancing immersive physics education and promoting the popularization of high-quality VR-based experimental teaching.  
      关键词:virtual reality (VR);virtual embodied generation;inverse kinematics (IK);virtual physics experiment;user experience   
      170
      |
      419
      |
      0
      <HTML>
      <L-PDF><WORD><Meta-XML>
      <引用本文> <批量引用> 143378877 false
      更新时间:2026-07-20
    0