Welcome to Journal of Beijing Institute of Technology
Peng Lu, Xinfu Liu, Yuting Zhou. A Distributed Cross-Modal Representation Structure for Zero-Shot Medical Image ClassificationJ. JOURNAL OF BEIJING INSTITUTE OF TECHNOLOGY, 2026, 35(4): 411-426. DOI: 10.15918/j.jbit1004-0579.2025.062
Citation: Peng Lu, Xinfu Liu, Yuting Zhou. A Distributed Cross-Modal Representation Structure for Zero-Shot Medical Image ClassificationJ. JOURNAL OF BEIJING INSTITUTE OF TECHNOLOGY, 2026, 35(4): 411-426. DOI: 10.15918/j.jbit1004-0579.2025.062

A Distributed Cross-Modal Representation Structure for Zero-Shot Medical Image Classification

  • Recent advances in distributed computing enable privacy-preserving aggregation and artificial intelligence (AI)-driven analysis of multimodal medical data, empowering real-time distributed healthcare applications and large-scale disease detection systems. Inspired by its remarkable performance, we proposed to discover medical knowledge embedded in big data with a large multimodal model, offering intelligent disease analytics via this promising AI route. In this paper, we proposed a distributed multimodal representation learning structure for zero-shot medical image classification. Specifically, we simultaneously adopted the implicit knowledge extracted from a large multimodal model built on images, and the explicit knowledge extracted from the medical knowledge graph built on textual records. Facing the inconsistent alignment in latent space constructed by multimodal data, a cross-modal alignment strategy was proposed to adjust intra- and inter-modal representations for convinced learning. Experiments on several public datasets proved that the proposed framework could improve the accuracy of zero-shot medical image classification, achieving robust and accurate disease analytical results.
  • loading

Catalog

    /

    DownLoad:  Full-Size Img  PowerPoint
    Return
    Return