Antoni Oliver
Biographic Data
| ID | 6640550 |
|---|---|
| NAME | Antoni Oliver |
| GIVEN NAMES | Antoni |
| FAMILY NAME | Oliver |
| SIGNATURE | OLIVER A |
| AFFILIATIONS | Universitat Oberta de Catalunya |
| ORCID | 0000-0001-8399-3770 |
| VERIFIED | Yes |
| TOTAL WORKS | 8 |
| TOTAL CITATIONS | 6 |
| AUTHOR COUNT | 8 |
| EDITOR COUNT | 0 |
| FIRST PUBLICATION YEAR | 2006 |
| LATEST PUBLICATION YEAR | 2026 |
| H-INDEX | 1 |
Terminological consistency in legal EU documents translated into Catalan using machine translation
Minority languages in Europe have a relevant position to strengthen cultural and linguistic communities and preserve societal cohesion. Nowadays, machine translation (MT) technologies have achieved relevant advances using neural networks, which are able to increase language use over communities. Although this is a promising scenario, the latest technological advances in MT technologies are not in-depth tailored and tested for some minority langua…
Assessing MT with measures of PE effort
Recent improvements in quality obtained by neural machine translation (NMT) have boosted its presence in the translation industry. In many domains and language combinations, translators post-edit raw MT output: they edit and correct the pre-translated text to produce the final translation. However, this process can only produce the expected results if the quality of the raw MT can be assured. MT is usually assessed with automatic metrics, as they…
Improving term candidates selection using terminological tokens
The identification of reliable terms from domain-specific corpora using computational methods is a task that has to be validated manually by specialists, which is a highly time-consuming activity. To reduce this effort and improve term candidate selection, we implemented the Token Slot Recognition method, a filtering method based on terminological tokens which is used to rank extracted term candidates from domain-specific corpora. This paper pres…
Cadlaws – An English–French Parallel Corpus of Legally Equivalent Documents
This article presents Cadlaws, a new English–French corpus built from Canadian legal documents, and describes the corpus construction process and preliminary statistics obtained from it. The corpus contains over 16 million words in each language and includes unique features since it is composed of documents that are legally equivalent in both languages but not the result of a translation. The corpus is built upon enactments co-drafted by two juri…
Metaphors of mental illness
In this paper we describe the building, manual annotation and analysis of a balanced corpus to assess conceptual metaphors on mental illness as used in Spanish blogger writing by patients and mental health professionals. The corpus was structured as eight subgroups: four patient subgroups (composed of persons who declared having been diagnosed with major depression, schizophrenia, bipolar disorder, or obsessive-compulsive disorder) and four menta…
Language industry views on the profile of the post-editor
The more language service companies (LSCs) include machine translation post-editing (MTPE) in their workflows, the more important it is to know how the PE task is performed, who the post-editors are, and what skills they should have. This research is designed to address such questions. It aims to deepen our knowledge of current practices to later create new training content and adapt existing training methodologies to different types of audiences…
Using open data to create the Catalan Iate e-dictionary
Linguistic resources available in the form of open data are an essential source of information for creating e-dictionaries, but access to these linguistic resources is still limited. This paper presents a method for maximising use of open access linguistic resources and integrating them into specialised e-dictionaries. The method combines automatic compilation of terminology data with the creation of specialised linguistic corpora to produce a Ca…
Bilingual Newsgroups in Catalonia
This paper presents a linguistic analysis of a corpus of messages written in Catalan and Spanish, which come from several informal newsgroups on the Universitat Oberta de Catalunya (Open University of Catalonia; henceforth, UOC) Virtual Campus. The surrounding environment is one of extensive bilingualism and contact between Spanish and Catalan. The study was carried out as part of the INTERLINGUA project conducted by the UOC's Internet Interdisci…
Bilingual Newsgroups in Catalonia
This paper presents a linguistic analysis of a corpus of messages written in Catalan and Spanish, which come from several informal newsgroups on the Universitat Oberta de Catalunya (Open University of Catalonia; henceforth, UOC) Virtual Campus. The surrounding environment is one of extensive bilingualism and contact between Spanish and Catalan. The study was carried out as part of the INTERLINGUA project conducted by the UOC's Internet Interdisci…
Bilingual Newsgroups in Catalonia
This paper presents a linguistic analysis of a corpus of messages written in Catalan and Spanish, which come from several informal newsgroups on the Universitat Oberta de Catalunya (Open University of Catalonia; henceforth, UOC) Virtual Campus. The surrounding environment is one of extensive bilingualism and contact between Spanish and Catalan. The study was carried out as part of the INTERLINGUA project conducted by the UOC's Internet Interdisci…
Using open data to create the Catalan Iate e-dictionary
Linguistic resources available in the form of open data are an essential source of information for creating e-dictionaries, but access to these linguistic resources is still limited. This paper presents a method for maximising use of open access linguistic resources and integrating them into specialised e-dictionaries. The method combines automatic compilation of terminology data with the creation of specialised linguistic corpora to produce a Ca…
Language industry views on the profile of the post-editor
The more language service companies (LSCs) include machine translation post-editing (MTPE) in their workflows, the more important it is to know how the PE task is performed, who the post-editors are, and what skills they should have. This research is designed to address such questions. It aims to deepen our knowledge of current practices to later create new training content and adapt existing training methodologies to different types of audiences…
Cadlaws – An English–French Parallel Corpus of Legally Equivalent Documents
This article presents Cadlaws, a new English–French corpus built from Canadian legal documents, and describes the corpus construction process and preliminary statistics obtained from it. The corpus contains over 16 million words in each language and includes unique features since it is composed of documents that are legally equivalent in both languages but not the result of a translation. The corpus is built upon enactments co-drafted by two juri…
Metaphors of mental illness
In this paper we describe the building, manual annotation and analysis of a balanced corpus to assess conceptual metaphors on mental illness as used in Spanish blogger writing by patients and mental health professionals. The corpus was structured as eight subgroups: four patient subgroups (composed of persons who declared having been diagnosed with major depression, schizophrenia, bipolar disorder, or obsessive-compulsive disorder) and four menta…
Improving term candidates selection using terminological tokens
The identification of reliable terms from domain-specific corpora using computational methods is a task that has to be validated manually by specialists, which is a highly time-consuming activity. To reduce this effort and improve term candidate selection, we implemented the Token Slot Recognition method, a filtering method based on terminological tokens which is used to rank extracted term candidates from domain-specific corpora. This paper pres…
Assessing MT with measures of PE effort
Recent improvements in quality obtained by neural machine translation (NMT) have boosted its presence in the translation industry. In many domains and language combinations, translators post-edit raw MT output: they edit and correct the pre-translated text to produce the final translation. However, this process can only produce the expected results if the quality of the raw MT can be assured. MT is usually assessed with automatic metrics, as they…
Terminological consistency in legal EU documents translated into Catalan using machine translation
Minority languages in Europe have a relevant position to strengthen cultural and linguistic communities and preserve societal cohesion. Nowadays, machine translation (MT) technologies have achieved relevant advances using neural networks, which are able to increase language use over communities. Although this is a promising scenario, the latest technological advances in MT technologies are not in-depth tailored and tested for some minority langua…
Artificial Intelligence (6 works) · Computer Science (6 works) · Natural Language Processing Techniques (6 works) · Machine translation (5 works) · Natural language processing (5 works) · Linguistics (4 works) · Catalan (3 works) · linguistics and terminology studies (3 works) · Translation Studies and Practices (3 works) · Business (2 works)