Large Language Model Embedding for Cold-Start Item Recommendation via Data Augmentation and Regularization
Dados Bibliográficos
| ID | 22106830 |
|---|---|
| Autores | Kangyi Xie (0009-0000-7621-5743, Fujian Normal University), Ayong Ye (0000-0002-2606-5406, Fujian Normal University), Chuang Huang (0000-0001-7873-8978, Fujian Normal University), Jing Chen (0000-0002-7243-0806, Fujian Normal University), Jingxu Chen (0000-0003-2395-6120, Fujian Normal University), Fu Liao (0009-0003-5861-964X, Fujian Normal University) |
| Ano | 2026 |
| Páginas | 1-10 |
| Data de publicação | 2026-01-01 |
| Peer Reviewed | Sim |
| Open Access | Sim |
| Tipo | ARTICLE |
| Periódico | IEEE Transactions on Computational Social Systems (JOURNAL) |
| Identificadores do periódico | ISSN: 2329-924X • E-ISSN: 2373-7476 |
| Editora | Institute of Electrical and Electronics Engineers (IEEE) (PUBLISHER) |
| DOI | 10.1109/tcss.2026.3654086 |
| OpenAlex | W7130426290 |
| Idioma | EN |
The cold-start problem constitutes a persistent challenge in recommender systems (RecSys). Recent advances in large language model (LLM) for natural language processing have inspired researchers to utilize their semantic understanding and knowledge extraction capabilities in recommendation tasks. LLM shows clear advantages in addressing cold-start items. However, current methods often rely on direct online LLM inference. This approach causes high computational costs and latency. To overcome these limitations, we propose large language model embedding for cold-start item recommendation via data augmentation and regularization (LLM-DAR). Our method transfers LLM capabilities to RecSys through semantic embeddings. This reduces computational costs and latency in online recommendations, making the system more suitable for high-concurrency real-time recommendation scenarios. Specifically, we first use LLM to extract semantic embeddings from item descriptions. Based on these embeddings, we implement data augmentation by associating cold-start items with similar popular items from users’ historical interactions, generating synthetic interaction samples to bridge knowledge gaps. To ensure effective learning of semantic relationships, we design a regularization constraint. This constraint pulls semantically similar items closer in the vector space
Data modeling · Deep learning · Embedding · Language model · Natural language · Recommender system · Semantic data model · Word embedding · Recommender Systems and Techniques · Sentiment Analysis and Opinion Mining · Topic Modeling
| Velocidade de citação | historical |
|---|---|
| Altamente citado | Não |