cd /news/natural-language-processing/writerslogic-at-the-clef-2026-simple… · home › topics › natural-language-processing › article
[ARTICLE · art-145195] src=machinebrief.com ↗ pub= topic=natural-language-processing verified=true sentiment=↑ positive

Writerslogic at the CLEF 2026 SimpleText Track: Multi-Candidate LLM Simplification and Stacked Complexity Spotting

The Writerslogic team's Claude Sonnet 4 submission ranked first among sentence-level systems in Task 1.1 of the CLEF 2026 SimpleText shared task with a SARI of 47.43 and BLEU of 14.21, placing third on the combined Task 1 leaderboard behind two document-level submissions. In Task 2.1, the team's fine-tuned DeBERTa-v3-large NLI model, trained on 350K labeled (source, sentence) pairs, reached 0.8081 document-level macro F1 (0.8085 in its best ensemble), the top-ranked entry within the identification track and second among teams overall behind AIIR Lab's 0.8197. The team's best Task 2.2 submission reached 0.804 multiclass accuracy, ranking second among unique teams behind AIIR Lab's 0.827, with both tasks evaluated on English and multilingual biomedical text from Cochrane systematic reviews.

by read1 min views1 publishedOct 5, 2026

arXiv:2610.03567v1 Announce Type: new Abstract: We describe the Writerslogic team's participation in the CLEF 2026 SimpleText shared task, addressing Task 1 (text simplification) and Task 2 (complexity spotting). For Task 1, we develop a multi-candidate generation pipeline using GPT-4o-mini that produces five simplification candidates per sentence at varying temperatures, then selects the best candidate using a reference-free scoring heuristic that rewards compression, source word retention, Cochrane Plain Language Summary vocabulary usage, and lexical simplicity. On Task 1.1 (sentence-level simplification), our Claude Sonnet 4 submission achieves SARI 47.43 and BLEU 14.21, the top-ranked sentence-level system (3rd on the combined Task 1 leaderboard, behind two document-level submissions). For Task 2, we fine-tune a DeBERTa-v3-large NLI model on 350K labeled (source, sentence) pairs, framing hallucination detection as natural language inference. The model reads the most relevant source sentence as premise and the candidate as hypothesis, directly learning to distinguish grounded from hallucinated content. On Task 2.1 (binary overgeneration identification), our fine-tuned DeBERTa system achieves 0.8081 document-level macro F1 (0.8085 in our best ensemble), the top-ranked entry within the identification track and 2nd among teams overall, behind AIIR Lab (0.8197). On Task 2.2 (multi-class error classification), our best submission reaches 0.804 multiclass accuracy, ranking 2nd among unique teams behind AIIR Lab (0.827). We evaluate both tasks on English and multilingual biomedical text from Cochrane systematic reviews.

── more in #natural-language-processing 4 stories · sorted by recency
── more on @writerslogic 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/writerslogic-at-the-…] indexed:0 read:1min 2026-10-05 · —