Writerslogic at the CLEF 2026 SimpleText Track: Multi-Candidate LLM Simplification and Stacked Complexity Spotting The Writerslogic team's Claude Sonnet 4 submission ranked first among sentence-level systems in Task 1.1 of the CLEF 2026 SimpleText shared task with a SARI of 47.43 and BLEU of 14.21, placing third on the combined Task 1 leaderboard behind two document-level submissions. In Task 2.1, the team's fine-tuned DeBERTa-v3-large NLI model, trained on 350K labeled (source, sentence) pairs, reached 0.8081 document-level macro F1 (0.8085 in its best ensemble), the top-ranked entry within the identification track and second among teams overall behind AIIR Lab's 0.8197. The team's best Task 2.2 submission reached 0.804 multiclass accuracy, ranking second among unique teams behind AIIR Lab's 0.827, with both tasks evaluated on English and multilingual biomedical text from Cochrane systematic reviews. arXiv:2610.03567v1 Announce Type: new Abstract: We describe the Writerslogic team's participation in the CLEF 2026 SimpleText shared task, addressing Task 1 text simplification and Task 2 complexity spotting . For Task 1, we develop a multi-candidate generation pipeline using GPT-4o-mini that produces five simplification candidates per sentence at varying temperatures, then selects the best candidate using a reference-free scoring heuristic that rewards compression, source word retention, Cochrane Plain Language Summary vocabulary usage, and lexical simplicity. On Task 1.1 sentence-level simplification , our Claude Sonnet 4 submission achieves SARI 47.43 and BLEU 14.21, the top-ranked sentence-level system 3rd on the combined Task 1 leaderboard, behind two document-level submissions . For Task 2, we fine-tune a DeBERTa-v3-large NLI model on 350K labeled source, sentence pairs, framing hallucination detection as natural language inference. The model reads the most relevant source sentence as premise and the candidate as hypothesis, directly learning to distinguish grounded from hallucinated content. On Task 2.1 binary overgeneration identification , our fine-tuned DeBERTa system achieves 0.8081 document-level macro F1 0.8085 in our best ensemble , the top-ranked entry within the identification track and 2nd among teams overall, behind AIIR Lab 0.8197 . On Task 2.2 multi-class error classification , our best submission reaches 0.804 multiclass accuracy, ranking 2nd among unique teams behind AIIR Lab 0.827 . We evaluate both tasks on English and multilingual biomedical text from Cochrane systematic reviews.