Pular para o conteúdo principal

ETHNOS_APP

Início • Busca • Periódicos • Lista 0

Extracting Geoscientific Dataset Names from the Literature Based on the Hierarchical Temporal Memory Model

Dados Bibliográficos

ID22033554
AutoresKai Wu (0000-0003-4684-9758, Henan University of Science and Technology), Zugang Chen (0000-0002-7782-9488, Chinese Academy of Sciences, autor correspondente), Xinqian Wu (Henan University of Science and Technology), Guoqing Li (0000-0001-6799-372X, Chinese Academy of Sciences), Jing Li (0000-0001-7792-4322, Chinese Academy of Sciences), Shaohua Wang (0000-0003-4451-8883, Aerospace Information Research Institute, Chinese Academy of Sciences, Beijing 100094, China), Shao‐Hua Wang (0000-0002-4347-8245, Chinese Academy of Sciences), Haodong Wang (0000-0001-7459-5893, Zhengzhou University), Hang Feng (0000-0002-8771-9573, Zhengzhou University)
Ano2024
Volume13
Fascículo7
Páginas260
Data de publicação2024-07-21
Peer ReviewedSim
Open AccessSim
TipoARTICLE
PeriódicoISPRS International Journal of Geo-Information (JOURNAL)
Identificadores do periódicoISSN: 2220-9964 • E-ISSN: 2220-9964
EditoraMDPI AG (PUBLISHER • IT)
DOI10.3390/ijgi13070260
OpenAlexW4400878070
IdiomaEN
Referências citadas52

Extracting geoscientific dataset names from the literature is crucial for building a literature–data association network, which can help readers access the data quickly through the Internet. However, the existing named-entity extraction methods have low accuracy in extracting geoscientific dataset names from unstructured text because geoscientific dataset names are a complex combination of multiple elements, such as geospatial coverage, temporal coverage, scale or resolution, theme content, and version. This paper proposes a new method based on the hierarchical temporal memory (HTM) model, a brain-inspired neural network with superior performance in high-level cognitive tasks, to accurately extract geoscientific dataset names from unstructured text. First, a word-encoding method based on the Unicode values of characters for the HTM model was proposed. Then, over 12,000 dataset names were collected from geoscience data-sharing websites and encoded into binary vectors to train the HTM model. We conceived a new classifier scheme for the HTM model that decodes the predictive vector for the encoder of the next word so that the similarity of the encoders of the predictive next word and the real next word can be computed. If the similarity is greater than a specified threshold, the real next word can be regarded as part of the name, and a successive word set forms the full geoscientific dataset name. We used the trained HTM model to extract geoscientific dataset names from 100 papers. Our method achieved an F1-score of 0.727, outperforming the GPT-4- and Claude-3-based few-shot learning (FSL) method, with F1-scores of 0.698 and 0.72, respectively

Data mining · Information retrieval · Natural language processing · Advanced Graph Neural Networks · Biomedical Text Mining and Ontologies · Computer Science · Topic Modeling · Artificial Intelligence

  • SciBert

    Open Access•Iz Beltagy, Kyle Lo et al.•Proceedings of the 2019…•2019

  • Statistical tests, P values, confidence intervals, and power

    Open Access•Sander Greenland, Stephen Senn et al.•European Journal of Epidemiology•2016

  • Geographic Named Entity Recognition by Employing Natural Language Processing and an Improved Bert Model

    Open Access•Liufeng Tao, Zhong Xie et al.•ISPRS International Journal of…•2022

Velocidade de citaçãohistorical
Altamente citadoNão
Ethnos_APP • Projeto Open Source • Licença MIT • Frontend v2.0.0 • Privacidade e Cookies • Documentação da API: api.ethnos.app/docs • Código da API: GitHub • DOI: 10.5281/zenodo.17049435 • Código do Frontend: GitHub • DOI: 10.5281/zenodo.17050053 • cruz.rio.br • Expectantes Misericordiae