cd /news/artificial-intelligence/lora-for-gender-inclusive-rewriting-… · home topics artificial-intelligence article
[ARTICLE · art-76367] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

LoRA for Gender-Inclusive Rewriting and Activation Steering for Counter-Narrative Generation

The IHLC system achieved an official score of 80.00% for gender-inclusive rewriting using parameter-efficient Low-Rank Adaptation (LoRA) fine-tuning, and 78.12% for counter-narrative generation via a compute-efficient inference-time representation engineering approach that injects a principal steering direction into Gemma-3-4B-it's intermediate representations. The system's manual analysis identified key failure modes including semantic drift, residual bias leakage, layer sensitivity, over-steering, and text degeneration, highlighting both the potential and limitations of activation steering for socially aligned language generation.

read1 min views1 publishedJul 28, 2026

arXiv:2607.23083v1 Announce Type: new Abstract: Gender-inclusive language generation seeks to transform biased text into inclusive alternatives while preserving semantic meaning and contextual coherence. This paper presents the IHLC system for the LT-EDI 2026 Shared Task, addressing both gender-inclusive rewriting and counter-narrative generation. For gender-inclusive rewriting, we employ parameter-efficient Low-Rank Adaptation (LoRA) fine-tuning, achieving an official score of 80.00%. Our primary contribution is a compute-efficient inference-time representation engineering approach for counter-narrative generation. We derive a principal steering direction from contrastive hidden-state activations using principal component analysis (PCA) and inject it into the intermediate representations of Gemma-3-4B-it during inference, enabling behavioral steering toward inclusive responses without modifying model weights. Combined with constrained prompting, this approach produces polite and contextually appropriate counter-narratives, achieving an official score of 78.12%. We further present a manual analysis of steering behavior, identifying key failure modes including semantic drift, residual bias leakage, layer sensitivity, over-steering, and text degeneration. Our findings highlight both the practical potential and current limitations of activation steering as a lightweight alternative to parameter updates for controllable and socially aligned language generation.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @ihlc 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/lora-for-gender-incl…] indexed:0 read:1min 2026-07-28 ·