cd /news/large-language-models/is-human-readable-text-necessary-for… · home › topics › large-language-models › article
[ARTICLE · art-142210] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=↑ positive

Is Human-Readable Text Necessary for Effective LLM Fine-Tuning?

A new arXiv paper (2609.35868v1) proposes Desired-Update-Aligned Synthetic Data (DASA), a method that fine-tunes large language models using continuous synthetic input embeddings guided by activation-gradient feedback from a frozen reference model, rather than human-readable text. Across six Llama and Qwen models from 1B to 32B parameters and six benchmarks covering knowledge, mathematical reasoning, code generation, and commonsense reasoning, DASA matched the source natural-language data under matched LoRA settings and surpassed it in multiple configurations, while outperforming GRADMM in most comparisons. The authors report DASA delivers a 3.6–4.9x speedup over GRADMM at comparable peak GPU memory.

by read1 min views1 publishedSep 30, 2026

arXiv:2609.35868v1 Announce Type: new Abstract: Is human readability necessary for effective fine-tuning of large language models? We investigate whether model-conditioned training representations can preserve or improve adaptation utility without requiring a human-readable textual form. We propose Desired-Update-Aligned Synthetic Data (DASA), which uses activation-gradient feedback from a frozen reference model to guide the optimization of continuous synthetic input embeddings. Inspired by the role of activation gradients in local risk reduction, DASA targets useful adaptation updates rather than source-text reconstruction or linguistic fluency. The resulting embeddings are used directly for downstream fine-tuning; discrete token projections are employed only for qualitative inspection. Experiments on six models from the Llama and Qwen families, ranging from 1B to 32B parameters, cover six benchmarks spanning knowledge, mathematical reasoning, code generation, and commonsense reasoning. Under matched LoRA adaptation settings, DASA achieves performance comparable to the source natural-language data and surpasses it in multiple configurations, while outperforming GRADMM in most comparisons. Further experiments cover general-domain and task-specialized source data. Under the evaluated synthesis settings, DASA provides a $3.6$--$4.9\times$ speedup over GRADMM with comparable peak GPU memory.

── more in #large-language-models 4 stories · sorted by recency
── more on @desired-update-aligned synthetic data 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/is-human-readable-te…] indexed:0 read:1min 2026-09-30 · —