An Embedded Unlabeled Data Partitioning PU-Learning Method Based on Genetic Programming
Dados Bibliográficos
| ID | 22106950 |
|---|---|
| Autores | Yu Zhou (0009-0001-5712-6902, Shenzhen University), Nanjian Yang (0009-0005-1312-1957, Shenzhen University), Ran Wang (0000-0002-1581-4651, Shenzhen University) |
| Ano | 2025 |
| Volume | 12 |
| Fascículo | 4 |
| Páginas | 1858-1868 |
| Data de publicação | 2025-08-01 |
| Peer Reviewed | Sim |
| Open Access | Sim |
| Tipo | ARTICLE |
| Periódico | IEEE Transactions on Computational Social Systems (JOURNAL) |
| Identificadores do periódico | ISSN: 2329-924X • E-ISSN: 2373-7476 |
| Editora | Institute of Electrical and Electronics Engineers (IEEE) (PUBLISHER) |
| DOI | 10.1109/tcss.2024.3406377 |
| OpenAlex | W4399727987 |
| Idioma | EN |
| Referências citadas | 30 |
In traditional binary classification tasks, learning algorithms conventionally distinguish positive and negative samples by leveraging fully labeled training data. However, in practical applications, there frequently arises a scenario where only a small number of positive samples and a large volume of unlabeled data exist, with the latter potentially containing a mix of positive and negative instances. This situation is known as positive-unlabeled learning (PUL) and has drawn considerable interest across various domains, such as text categorization. Despite numerous studies addressing PUL issues, few have focused on improving the two-step methodology in the data partitioning process, and even fewer have explored the application of genetic programming (GP) to PUL challenges. This article introduces a GP-based embedded unlabeled data partitioning method (EPGP) tailored for the PUL problem, particularly in the context of few-shot learning scenarios. The approach adopts a multiphase data partitioning strategy, integrating the partitioning process into the evolutionary cycle of the GP population, progressively isolating positive samples from the unlabeled data to yield a PU dataset closer to the real-world distribution. To achieve smoother data partitioning, a dynamically adjusted partitioning threshold strategy is incorporated. Finally, an ensemble method is devised, capitalizing on the high confidence associated with the originally labeled positive samples to generate the final classification outcome via weighted voting. Experimental evaluations of EPGP against other state-of-the-art PUL methods on 22 datasets show significant advantages in terms of balanced accuracy and macro-F1 scores on 17 datasets. Moreover, a case study on text data classification tasks demonstrates that EPGP consistently delivers substantial performance enhancements under varying proportions of labeled positive samples
Genetic algorithm · Genetic programming · Machine learning · Advanced Algorithms and Applications · Computer Science · Evolutionary Algorithms and Applications · Industrial Vision Systems and Defect Detection · Artificial Intelligence
| Velocidade de citação | historical |
|---|---|
| Altamente citado | Não |