arXiv:2608.00713v1 Announce Type: new Abstract: This paper describes Observatorio L'azaro, a language resource that monitors unassimilated lexical borrowings (predominantly English lexical borrowings or anglicisms) in the Spanish digital press. Since April 2020 the system has automatically processed the daily output of a collection of news outlets, detected borrowings with a neural sequence-labeling model, and made the results available through a public web interface and API. The result is a continuously updated diachronic database which, at the time of writing, records more than two million borrowings across 1.88 million articles and 993 million running tokens of text (2020-2026). The paper documents the resource: we describe the end-to-end pipeline (acquisition, detection, post-processing, storage and access), the data model and the terms of availability; we evaluate the resource through the detector's held-out performance (span-level F1=0.86 for the borrowing class), inter-annotator agreement on the training corpus (Cohen's kappa=0.91) and a manual precision audit of 1,000 spans from the deployed data; and we situate it with respect to Spanish borrowing lexicography, annotated borrowing corpora and neology-monitoring observatories. The data shows that unassimilated anglicisms are used in the Spanish press at a frequency of approximately two anglicisms per thousand tokens, and that this rate remains stable. Our statistical analysis over six years reveals that the anglicism vocabulary in Spanish behaves as an open and growing class, with 58.7% of its types attested only once (53.6% after correcting for detection precision), and that its density is highest in the fashion, technology and lifestyle sections and lowest in political and institutional news. The resource is intended to complement static borrowing dictionaries and one-off annotated corpora by providing a continuously updated record of borrowing in the Spanish press.
Observatorio Lazaro: A self-populating database of anglicism usage in the Spanish press
Researchers at Observatorio Lázaro have built a self-populating database that has recorded more than two million English lexical borrowings (anglicisms) in the Spanish digital press, processing 1.88 million articles and 993 million running tokens from 2020 to 2026. The system, described in a paper on arXiv (2608.00713v1), uses a neural sequence-labeling model to detect borrowings with a span-level F1 of 0.86 and inter-annotator agreement of Cohen's kappa 0.91, finding a stable frequency of about two anglicisms per thousand tokens and that 58.7% of anglicism types appear only once. The resource, available via a public web interface and API, aims to complement static dictionaries and annotated corpora by providing a continuously updated record of borrowing in Spanish news.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.