SGG-Mvar
Cross-Modal Retrieval With Scene Graph Generation and Multiview Attribute Relationship Guidance
Dados Bibliográficos
| ID | 22106930 |
|---|---|
| Autores | Suping Wang (0000-0002-1476-5595, Zhejiang Meteorological Bureau), Fei Zhou (0000-0001-6207-6236, Guangxi Zhuang Autonomous Region Health and Family Planning), Ming Yang (0000-0002-6975-2548, Zhejiang Meteorological Bureau), Lei Shi (0000-0003-1203-9984, Communication University of China), Chaohong Tan (0000-0003-3337-3447, Guangxi Zhuang Autonomous Region Health and Family Planning) |
| Ano | 2025 |
| Volume | 12 |
| Fascículo | 5 |
| Páginas | 3671-3683 |
| Data de publicação | 2025-10-01 |
| Peer Reviewed | Sim |
| Open Access | Sim |
| Tipo | ARTICLE |
| Periódico | IEEE Transactions on Computational Social Systems (JOURNAL) |
| Identificadores do periódico | ISSN: 2329-924X • E-ISSN: 2373-7476 |
| Editora | Institute of Electrical and Electronics Engineers (IEEE) (PUBLISHER) |
| DOI | 10.1109/tcss.2024.3524297 |
| OpenAlex | W4406321963 |
| Idioma | EN |
| Citações recebidas | 1 |
| Referências citadas | 34 |
Cross-modal retrieval is crucial for achieving accurate and efficient information retrieval by establishing semantic correlations between heterogeneous images and text. However, traditional image-text training sets suffer from information asymmetry, which includes short lengths and limited sentence structures. This phenomenon often results in insufficient representations of essential visual information. We introduce RichDataset, which offers extensive semantic information. It includes diverse real-life image-text pairs and AI-generated content across domains such as news, entertainment, education, and posters. Compared with classic benchmarks such as Flickr30k and MS-COCO, RichDataset exhibits a novel and balanced distribution. Existing cross-modal retrieval models face challenges in extracting distinct features from the emerging data, leading to low retrieval accuracy. We propose SGG-MVAR, a comprehensive retrieval model guided by multiview scene information and semantic relationships. Leveraging a scene knowledge database, our model parses scene graphs and identifies differences in attributes and relationships. We conduct extensive experiments to evaluate our proposed dataset and model. All experimental results consistently demonstrate a significant improvement in recall for cross-modal retrieval
Computer vision · Graph · Modal · Advanced Image and Video Retrieval Techniques · Computer Science · Image Retrieval and Classification Techniques · Materials Science · Multimodal Machine Learning Applications · Artificial Intelligence · Theoretical Computer Science
| Obras citantes distintas | 1 |
|---|---|
| Citações por ano | 1 |
| Intervalo de citações | 2026 - 2026 (1) |
| Velocidade de citação | current |
| Altamente citado | Não |
| Tipos de citação | Neutras: 1 |