Pre-Training With Whole Word Masking for Chinese Bert
Bibliographic Data
| ID | 23319672 |
|---|---|
| Authors | Yiming Cui (0000-0002-2452-375X, Harbin Institute of Technology), Wanxiang Che (0000-0002-3907-0335, Harbin Institute of Technology), Ting Liu (0000-0001-5479-7228, Harbin Institute of Technology), Bing Qin (0000-0003-2842-9637, Harbin Institute of Technology), Ziqing Yang (0009-0002-1144-3408, Institute of Animal Sciences) |
| Year | 2021 |
| Volume | 29 |
| Pages | 3504-3514 |
| Publication date | 2021-01-01 |
| Peer Reviewed | Yes |
| Open Access | Yes |
| Type | ARTICLE |
| Venue | IEEE/ACM Transactions on Audio, Speech, and Language Processing (JOURNAL) |
| Journal identifiers | ISSN: 2329-9290 • E-ISSN: 2329-9304 |
| Publisher | Institute of Electrical and Electronics Engineers (IEEE) (PUBLISHER) |
| DOI | 10.1109/taslp.2021.3124365 |
| OpenAlex | W2952370363 |
| Language | EN |
| Citations received | 51 |
| References cited | 16 |
Bidirectional Encoder Representations from Transformers (BERT) has shown marvelous improvements across various NLP tasks, and its consecutive variants have been proposed to further improve the performance of the pre-trained language models. In this paper, we aim to first introduce the whole word masking (wwm) strategy for Chinese BERT, along with a series of Chinese pre-trained language models. Then we also propose a simple but effective model called MacBERT, which improves upon RoBERTa in several ways. Especially, we propose a new masking strategy called MLM as correction (Mac). To demonstrate the effectiveness of these models, we create a series of Chinese pre-trained language models as our baselines, including BERT, RoBERTa, ELECTRA, RBT, etc. We carried out extensive experiments on ten Chinese NLP tasks to evaluate the created Chinese pre-trained language models as well as the proposed MacBERT. Experimental results show that MacBERT could achieve state-of-the-art performances on many NLP tasks, and we also ablate details with several findings that may help future research. We open-source our pre-trained language models for further facilitating our research community.
Chinese language · Encoder · Language model · Masking (illustration) · Transformer · Word (group theory) · Generative Adversarial Networks and Image Synthesis · Multimodal Machine Learning Applications · Topic Modeling
Assessing multidimensional changes in human settlements across urban regeneration paradigms
A Brief Overview of ChatGPT
Identification and Impact Analysis of Family History of Psychiatric Disorder in Mood Disorder Patients With Pretrained Language Model
Bridging Cultures in the Era of Big Data
An interactive mobile visual search model for narrative murals enhanced by the large multimodal model
Examining the escalation of hostility in social media
Polarization of public opinions on feminism in China
Sinophobia was popular in Chinese language communities on Twitter during the early Covid-19 pandemic
The risk effects of corporate digitalization
Empowering college students to select ideal advisors
Tibetan-Bert-wwm
Metaphors as Semantic Anchors
Sail
PLP-Ssaf
Two-Stage Construction Method of Event Knowledge Graph for Emergency Disposal of Gas Accidents
Evidence Mining for Interpretable Charge Prediction via Prompt Learning
Improved Target-Specific Stance Detection on Social Media Platforms by Delving Into Conversation Threads
Research on the identification and evolution of health industry policy instruments in China
Enhancing Electric Vehicle Charging Infrastructure Planning with Pre-Trained Language Models and Spatial Analysis
A Novel Address-Matching Framework Based on Region Proposal
Critique
Large language model-assisted public opinion analysis of emergency events of major transportation infrastructure through social media
Research on online public opinion in the investigation of the “7–20” extraordinary rainstorm and flooding disaster in Zhengzhou, China
How Digital Technology Affects the New Quality Productivity Forces of Enterprises
Cultural differences in information exchange
Public attitudes toward the Israeli-Palestinian conflict in China
Exploring an effective automated grading model with reliability detection for large‐scale online peer assessment
Investigating patients' adoption of online medical advice
Residential sentiments and spatial environments in urban villages
What matters in patent claims
Exploring the climate change discourse on Chinese social media and the role of social bots
The impact of tax reform on firms' digitalization in China
Attention-aware semantic relevance predicting Chinese sentence reading
Decoding urban policies
A novel domain knowledge augmented large language model based medical conversation system for sustainable smart city development
Exploring temporal and spatial patterns and nonlinear driving mechanism of park perceptions
Examining heat risk inequality from the perspective of urban resilience
Differentiation and unity
Unveiling the global hijab discourse on Instagram
Social media insights into spatio-temporal emotional responses to Covid-19 crisis
Greenspace exposure is conducive to the resilience of public sentiment during the Covid-19 pandemic
Governing the Smart City in the Age of AI
Evaluating community resilience through social media during China’s first post-Covid-19 reopening
The viral gender antagonism on social media
Attention and attitudes of Chinese social media users towards autonomous vehicles
From trade protectionism to legal protectionism
Natural language processing for social science research
How does three-dimensional landscape pattern affect urban residents' sentiments
Electoral systems and geographically targeted oversight
Beyond five senses
Fine-grained extraction of geospatial and temporal information from Chinese historical newspapers
| Unique citing works | 51 |
|---|---|
| Citations per year | 12,75 |
| Citation span | 2022 - 2026 (5) |
| Citation velocity | current |
| Highly cited | No |
| Citation types | Neutral: 50 |