cd /news/large-language-models/ppl-factory-task-aware-and-budget-aw… · home topics large-language-models article
[ARTICLE · art-66449] src=machinebrief.com ↗ pub= topic=large-language-models verified=true sentiment=↑ positive

PPL-Factory: Task-Aware and Budget-Aware Data Selection from Language Modeling to Reasoning

Researchers propose PPL-Factory, a task-aware and budget-aware data selection framework for large language model fine-tuning that uses perplexity-based scores to select informative training samples. Experiments on GSM8K show PPL-Factory outperforms other state-of-the-art methods using only 1% of the training set, and with 10% of the data it exceeds full-data fine-tuning accuracy by 0.9 on GSM8K and 4.8 on MATH.

read1 min views3 publishedJul 21, 2026

arXiv:2607.18199v1 Announce Type: new Abstract: Not all training samples contribute equally to large language model fine-tuning. Selecting informative training samples can reduce the computational cost while preserving downstream performance. Many existing data selection methods rely on indirect heuristics, such as data quality, diversity or reasoning trace length. However, the effectiveness of these fixed criteria is task-dependent and difficult to generalize across diverse downstream tasks. Perplexity-based data selection provides a simple and model-aware solution to estimate the sample difficulty, but existing approaches typically score the entire training sequence and ignore the difference in learning objectives of language modeling and reasoning tasks. In this paper, we propose PPL-Factory, a simple and interpretable data selection framework that combines task-aware perplexity-based scores and data budget-aware selection criteria. Experiments on GSM8K demonstrate that PPL-Factory outperforms other state-of-the-art data selection methods using only $1%$ of the training set. With $10%$ of the data, PPL-Factory exceeds full-data fine-tuning accuracy by 0.9 on GSM8K and 4.8 on MATH. Overall, our results demonstrate that task-aware and budget-aware perplexity-based selection provides an effective and applicable approach for efficient fine-tuning.

── more in #large-language-models 4 stories · sorted by recency
── more on @ppl-factory 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/ppl-factory-task-awa…] indexed:0 read:1min 2026-07-21 ·