cd /news/large-language-models/every-time-i-hire-a-linguist-inferen… · home topics large-language-models article
[ARTICLE · art-78065] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Every Time I Hire a Linguist, Inference Costs Go Down: On Linguistic Rules as Effective Prompt Compressors

A new study from arXiv:2607.25335 shows that linguistic rules alone can serve as effective prompt compressors for large language models, reducing inference costs without requiring LM-based scoring at compression time. The evolved compressors, which use only CPU-side processing, achieve performance similar to advanced prompt-compression strategies across short passages, multi-document reasoning, and dialogue-memory QA datasets, with strongest results under light-to-moderate compression.

read1 min views1 publishedJul 29, 2026

arXiv:2607.25335v1 Announce Type: new Abstract: Prompt compression shortens LLM input to reduce inference cost, yet existing methods score token importance through LM forward passes. It remains questionable whether such nuanced, costly token selection is necessary. Compression requires identifying informative content, a problem that linguistic research has long addressed through cues that can be operationalized as deterministic rules. We therefore ask: can \textbf{linguistic rules alone} serve as effective prompt compressors, without LM-based scoring at compression time? To address this, we conduct offline evolutionary search over lexical, syntactic, semantic, and discourse seeds to find competitive rule combinations. The resulting linguistic compressor requires no LM forward pass at deployment and uses only CPU-side processing for compression. We evaluate it with a dual-path protocol to balance compression quality and reconstruction fidelity. Across short passages, multi-document reasoning, and dialogue-memory QA datasets, evolved compressors achieve performance similar to that of recent advanced prompt-compression strategies. Performance is strongest under light-to-moderate compression and degrades as compression becomes more aggressive, while the Direct and Reconstruction paths exhibit distinct patterns. Evolutionary analysis reveals that effective compression fuses signals across linguistic levels and, as the compression ratio increases, rules shift from token pruning to sentence extraction.

── more in #large-language-models 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/every-time-i-hire-a-…] indexed:0 read:1min 2026-07-29 ·