cd /news/machine-learning/evaluating-federated-pre-training-on… · home topics machine-learning article
[ARTICLE · art-84186] src=arxiv.org ↗ pub= topic=machine-learning verified=true sentiment=· neutral

Evaluating Federated Pre-Training: On the Reliability of Downstream Fine-Tuning and Intrinsic Evaluation

A new study from arXiv (arXiv:2607.28658v1) finds that downstream fine-tuning on GLUE does not reliably preserve the reference ranking of federated pre-trained models, whereas direct next-token prediction on GLUE text strongly corresponds with pre-training test perplexity. The researchers trained a 16M parameter transformer model on identical client data under centralized and federated settings, and compared evaluation protocols including full, head-only, and reduced-data fine-tuning against intrinsic next-token prediction. The findings suggest that relying solely on downstream fine-tuning can mislead comparisons of federated pre-trained models, and evaluation signals closer to the pre-training objective should be prioritized.

read1 min views1 publishedAug 3, 2026

arXiv:2607.28658v1 Announce Type: new Abstract: Federated pre-training offers a way to train foundation models on private or distributed data without centralizing the underlying datasets. However, evaluating federated pre-training remains challenging because differences in client participation and local data availability can make directly comparable evaluation difficult. Moreover, pre-training test perplexity is tied to the pre-training distribution, while downstream benchmarks introduce task-specific adaptation that may not faithfully reflect the test perplexity established during pre-training. In this work, we study which evaluation protocol more reliably reflects federated pre-training quality. Using a controlled set of centralized and federated-trained models of a 16M parameter transformer model trained on identical client data, we assess evaluation protocols by whether they preserve a reference ranking established on the same pre-training testset. We compare downstream fine-tuning on GLUE, including full, head-only, and reduced-data variants, with next-token prediction on GLUE text as an intrinsic evaluation signal. Our results show that downstream fine-tuning does not reliably preserve the pre-training ranking, whereas direct next-token prediction exhibits a strong correspondence with the pre-training test perplexity. These findings suggest that downstream fine-tuning alone can be misleading when comparing federated pre-trained models, and that evaluation signals closer to the original pre-training objective deserve greater attention.

── more in #machine-learning 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/evaluating-federated…] indexed:0 read:1min 2026-08-03 ·