{"slug": "evaluating-federated-pre-training-on-the-reliability-of-downstream-fine-tuning", "title": "Evaluating Federated Pre-Training: On the Reliability of Downstream Fine-Tuning and Intrinsic Evaluation", "summary": "A new study from arXiv (arXiv:2607.28658v1) finds that downstream fine-tuning on GLUE does not reliably preserve the reference ranking of federated pre-trained models, whereas direct next-token prediction on GLUE text strongly corresponds with pre-training test perplexity. The researchers trained a 16M parameter transformer model on identical client data under centralized and federated settings, and compared evaluation protocols including full, head-only, and reduced-data fine-tuning against intrinsic next-token prediction. The findings suggest that relying solely on downstream fine-tuning can mislead comparisons of federated pre-trained models, and evaluation signals closer to the pre-training objective should be prioritized.", "body_md": "arXiv:2607.28658v1 Announce Type: new\nAbstract: Federated pre-training offers a way to train foundation models on private or distributed data without centralizing the underlying datasets. However, evaluating federated pre-training remains challenging because differences in client participation and local data availability can make directly comparable evaluation difficult. Moreover, pre-training test perplexity is tied to the pre-training distribution, while downstream benchmarks introduce task-specific adaptation that may not faithfully reflect the test perplexity established during pre-training. In this work, we study which evaluation protocol more reliably reflects federated pre-training quality. Using a controlled set of centralized and federated-trained models of a 16M parameter transformer model trained on identical client data, we assess evaluation protocols by whether they preserve a reference ranking established on the same pre-training testset. We compare downstream fine-tuning on GLUE, including full, head-only, and reduced-data variants, with next-token prediction on GLUE text as an intrinsic evaluation signal. Our results show that downstream fine-tuning does not reliably preserve the pre-training ranking, whereas direct next-token prediction exhibits a strong correspondence with the pre-training test perplexity. These findings suggest that downstream fine-tuning alone can be misleading when comparing federated pre-trained models, and that evaluation signals closer to the original pre-training objective deserve greater attention.", "url": "https://wpnews.pro/news/evaluating-federated-pre-training-on-the-reliability-of-downstream-fine-tuning", "canonical_source": "https://arxiv.org/abs/2607.28658", "published_at": "2026-08-03 04:00:00+00:00", "updated_at": "2026-08-03 04:04:13.092031+00:00", "lang": "en", "topics": ["machine-learning", "large-language-models", "ai-research"], "entities": ["arXiv", "GLUE"], "alternates": {"html": "https://wpnews.pro/news/evaluating-federated-pre-training-on-the-reliability-of-downstream-fine-tuning", "markdown": "https://wpnews.pro/news/evaluating-federated-pre-training-on-the-reliability-of-downstream-fine-tuning.md", "text": "https://wpnews.pro/news/evaluating-federated-pre-training-on-the-reliability-of-downstream-fine-tuning.txt", "jsonld": "https://wpnews.pro/news/evaluating-federated-pre-training-on-the-reliability-of-downstream-fine-tuning.jsonld"}}