Suzan Verberne
Biographic Data
| ID | 6536244 |
|---|---|
| NAME | Suzan Verberne |
| GIVEN NAMES | Suzan |
| FAMILY NAME | Verberne |
| SIGNATURE | VERBERNE S |
| AFFILIATIONS | Leiden University |
| ORCID | 0000-0002-9609-9505 |
| VERIFIED | Yes |
| TOTAL WORKS | 6 |
| TOTAL CITATIONS | 2 |
| AUTHOR COUNT | 6 |
| EDITOR COUNT | 0 |
| FIRST PUBLICATION YEAR | 2010 |
| LATEST PUBLICATION YEAR | 2025 |
| H-INDEX | 1 |
Which topics are best represented by science maps? An analysis of clustering effectiveness for citation and text similarity networks
A science map of topics is a visualization that shows topics identified algorithmically based on the bibliographic metadata of scientific publications. In practice not all topics are well represented in a science map. We analyzed how effectively different topics are represented in science maps created by clustering biomedical publications. To achieve this, we investigated which topic categories, obtained from MeSH terms, are better represented in…
Academic information retrieval using citation clusters: In-depth evaluation based on systematic reviews
The field of science mapping has shown the power of citation-based clusters for literature analysis, yet this technique has barely been used for information retrieval tasks. This work evaluates the performance of citation-based clusters for information retrieval tasks. We simulated a search process with a tree hierarchy of clusters and a cluster selection algorithm. We evaluated the task of finding the relevant documents for 25 systematic reviews…
Pretrained Transformers for Text Ranking: Bert and Beyond
User Requirement Solicitation for an Information Retrieval System Applied to Dutch Grey Literature in the Archaeology Domain
In this paper, we present the results of user requirement solicitation for a search system of grey literature in archaeology, specifically Dutch excavation reports. This search system uses Named Entity Recognition and Information Retrieval techniques to create an effective and effortless search experience. Specifically, we used Conditional Random Fields to identify entities, with an average accuracy of 56%. This is a baseline result, and we ident…
Text Representations for Patent Classification
With the increasing rate of patent application filings, automated patent classification is of rising economic importance. This article investigates how patent classification can be improved by using different representations of the patent documents. Using the Linguistic Classification System (LCS), we compare the impact of adding statistical phrases (in the form of bigrams) and linguistic phrases (in two different dependency formats) to the stand…
What Is Not in the Bag of Words for Why -QA
While developing an approach to why-QA, we extended a passage retrieval system that uses off-the-shelf retrieval technology with a re-ranking step incorporating structural information. We get significantly higher scores in terms of MRR@150 (from 0.25 to 0.34) and success@10. The 23% improvement that we reach in terms of MRR is comparable to the improvement reached on different QA tasks by other researchers in the field, although our re-ranking ap…
Text Representations for Patent Classification
With the increasing rate of patent application filings, automated patent classification is of rising economic importance. This article investigates how patent classification can be improved by using different representations of the patent documents. Using the Linguistic Classification System (LCS), we compare the impact of adding statistical phrases (in the form of bigrams) and linguistic phrases (in two different dependency formats) to the stand…
What Is Not in the Bag of Words for Why -QA
While developing an approach to why-QA, we extended a passage retrieval system that uses off-the-shelf retrieval technology with a re-ranking step incorporating structural information. We get significantly higher scores in terms of MRR@150 (from 0.25 to 0.34) and success@10. The 23% improvement that we reach in terms of MRR is comparable to the improvement reached on different QA tasks by other researchers in the field, although our re-ranking ap…
Text Representations for Patent Classification
With the increasing rate of patent application filings, automated patent classification is of rising economic importance. This article investigates how patent classification can be improved by using different representations of the patent documents. Using the Linguistic Classification System (LCS), we compare the impact of adding statistical phrases (in the form of bigrams) and linguistic phrases (in two different dependency formats) to the stand…
User Requirement Solicitation for an Information Retrieval System Applied to Dutch Grey Literature in the Archaeology Domain
In this paper, we present the results of user requirement solicitation for a search system of grey literature in archaeology, specifically Dutch excavation reports. This search system uses Named Entity Recognition and Information Retrieval techniques to create an effective and effortless search experience. Specifically, we used Conditional Random Fields to identify entities, with an average accuracy of 56%. This is a baseline result, and we ident…
Pretrained Transformers for Text Ranking: Bert and Beyond
Academic information retrieval using citation clusters: In-depth evaluation based on systematic reviews
The field of science mapping has shown the power of citation-based clusters for literature analysis, yet this technique has barely been used for information retrieval tasks. This work evaluates the performance of citation-based clusters for information retrieval tasks. We simulated a search process with a tree hierarchy of clusters and a cluster selection algorithm. We evaluated the task of finding the relevant documents for 25 systematic reviews…
Which topics are best represented by science maps? An analysis of clustering effectiveness for citation and text similarity networks
A science map of topics is a visualization that shows topics identified algorithmically based on the bibliographic metadata of scientific publications. In practice not all topics are well represented in a science map. We analyzed how effectively different topics are represented in science maps created by clustering biomedical publications. To achieve this, we investigated which topic categories, obtained from MeSH terms, are better represented in…
Computer Science (6 works) · Information retrieval (6 works) · Artificial Intelligence (5 works) · Advanced Text Analysis Techniques (3 works) · Mathematics (3 works) · Natural language processing (3 works) · Biomedical Text Mining and Ontologies (2 works) · Citation (2 works) · Data mining (2 works) · Data science (2 works)