{"slug": "rethinking-gender-annotation-for-bias-evaluation-in-machine-translation-can-llms", "title": "Rethinking Gender Annotation for Bias Evaluation in Machine Translation: Can LLMs Improve Reliability?", "summary": "A study by Chiara Manna, Argentina Anna Rescigno, and Eva Vanmassenhove compared the automated WinoMT annotation pipeline with the instruction-tuned LLM Qwen3-8B for annotating grammatical gender in English–Italian machine translation, finding both reach similar agreement with a human-annotated gold standard but fail in distinct ways. The WinoMT pipeline is sensitive to alignment shifts and morphological tagging limits, often producing indeterminate gender labels, while Qwen3-8B tends to favor binary gender labels even for morphologically gender-invariant noun phrases. The authors report that qualitative analysis of model outputs and reasoning traces suggests a lack of stable generalization of grammatical gender rules even with in-context exemplars, highlighting methodological limits of using LLMs as objective annotators for gender evaluation, in the paper published in the Proceedings of the 4th Workshop on Gender-Inclusive Translation Technologies (GITT 2026), pages 16–30.", "body_md": "##### Abstract\n\nThe assessment of gender bias in Machine Translation critically depends on reliable methods for identifying grammatical gender in system outputs. In this paper, we compare the automated WinoMT annotation pipeline, based on word alignment and morphological tagging, with an instruction-tuned LLM (Qwen3-8B) used to annotate grammatical gender in English–Italian translations. Both approaches achieve similar levels of agreement with a human-annotated gold standard, but exhibit distinct systematic weaknesses. The WinoMT pipeline is sensitive to alignment shifts and morphological tagging limitations, often resulting in indeterminate gender labels. The LLM, in contrast, tends to favour binary gender labels even when noun phrases are morphologically gender-invariant. A qualitative analysis of model outputs and reasoning traces further suggests a lack of stable generalization of grammatical gender rules, even when illustrative exemplars are provided for in-context learning. This highlights important methodological limitations in using LLMs as objective annotators for gender evaluation tasks.\n- Anthology ID:\n- 2026.gitt-1.3\n- Volume:\n- [Proceedings of the 4th Workshop on Gender-Inclusive Translation Technologies (GITT 2026)](https://aclanthology.org/volumes/2026.gitt-1/)\n- Month:\n- June\n- Year:\n- 2026\n- Address:\n- Tilburg, the Netherlands\n- Editors:\n- [Manuel Lardelli](https://aclanthology.org/people/manuel-lardelli/) ,[Beatrice Savoldi](https://aclanthology.org/people/beatrice-savoldi/) ,[Janiça Hackenbuchner](https://aclanthology.org/people/janica-hackenbuchner/) ,[Luisa Bentivogli](https://aclanthology.org/people/luisa-bentivogli/) ,[Eleni Gkovedarou](https://aclanthology.org/people/eleni-gkovedarou/) ,[Joke Daems](https://aclanthology.org/people/joke-daems/)\n- Venues:\n- [GITT](https://aclanthology.org/venues/gitt/) |[WS](https://aclanthology.org/venues/ws/)\n- SIG:\n- Publisher:\n- European Association for Machine Translation\n- Note:\n- Pages:\n- 16–30\n- Language:\n- URL:\n- [https://aclanthology.org/2026.gitt-1.3/](https://aclanthology.org/2026.gitt-1.3/)\n- DOI:\n- Cite (ACL):\n- Chiara Manna, Argentina Anna Rescigno, and Eva Vanmassenhove. 2026. [Rethinking Gender Annotation for Bias Evaluation in Machine Translation: Can LLMs Improve Reliability?](https://aclanthology.org/2026.gitt-1.3/) . In*Proceedings of the 4th Workshop on Gender-Inclusive Translation Technologies (GITT 2026)* , pages 16–30, Tilburg, the Netherlands. European Association for Machine Translation.\n- Cite (Informal):\n- [Rethinking Gender Annotation for Bias Evaluation in Machine Translation: Can LLMs Improve Reliability?](https://aclanthology.org/2026.gitt-1.3/) (Manna et al., GITT 2026)\n- PDF:\n- [https://aclanthology.org/2026.gitt-1.3.pdf](https://aclanthology.org/2026.gitt-1.3.pdf)", "url": "https://wpnews.pro/news/rethinking-gender-annotation-for-bias-evaluation-in-machine-translation-can-llms", "canonical_source": "https://aclanthology.org/2026.gitt-1.3/", "published_at": "2026-09-17 00:00:00+00:00", "updated_at": "2026-09-22 18:24:44.523027+00:00", "lang": "en", "topics": ["machine-learning", "natural-language-processing", "large-language-models", "ai-ethics"], "entities": ["Chiara Manna", "Argentina Anna Rescigno", "Eva Vanmassenhove", "Qwen3-8B", "WinoMT", "European Association for Machine Translation", "GITT 2026"], "alternates": {"html": "https://wpnews.pro/news/rethinking-gender-annotation-for-bias-evaluation-in-machine-translation-can-llms", "markdown": "https://wpnews.pro/news/rethinking-gender-annotation-for-bias-evaluation-in-machine-translation-can-llms.md", "text": "https://wpnews.pro/news/rethinking-gender-annotation-for-bias-evaluation-in-machine-translation-can-llms.txt", "jsonld": "https://wpnews.pro/news/rethinking-gender-annotation-for-bias-evaluation-in-machine-translation-can-llms.jsonld"}}