cd /news/artificial-intelligence/agentic-coding-without-the-cloud-eva… · home topics artificial-intelligence article
[ARTICLE · art-74955] src=machinebrief.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Agentic coding without the cloud: evaluating open-weight large language models on longitudinal data preparation tasks

A new open-source framework from University College London (UCL) evaluates open-weight large language models on longitudinal data preparation tasks without sending data to the cloud. Benchmarking models up to 31-35B parameters achieved up to 87.9% average task completion on 20 tasks creating 102 variables, showing promise for AI-assisted data preparation in governance-restricted research settings.

read1 min views1 publishedJul 27, 2026

arXiv:2607.21482v2 Announce Type: replace-cross Abstract: Large language models (LLMs) and agents are now widely used tools in code development, with data typically sent to third-party cloud-based models. Their adoption in research using personal data is constrained by governance requirements that typically prohibit data transmission to external services. Locally deployable open-weight models offer an alternative since sensitive data never leave the local environment. We introduce an open-source framework for evaluating the efficacy of AI agents powered by open-weight LLMs on one of the most persistent bottlenecks in research on longitudinal population studies: data preparation. The framework comprises: a curated ground-truth dataset (cleaning scripts preparing six sweeps of data from a British cohort study), task definitions encompassing tasks such as category harmonization and multi-wave merging, and automated routines for evaluating the LLM-produced R code and outputted data. We benchmark LLMs across the (consumer grade) deployment spectrum to assess their efficacy in 20 data preparation tasks (creation of 102 variables). Current state-of-the-art, 31-35B parameter models almost saturated our benchmark ('average task completion' up to 87.9%). The performance of open-weight LLMs running on consumer-grade hardware shows promise of a viable path toward AI-assisted data preparation in governance-restricted research settings. Our framework is publicly available at: https://github.com/UCL-ARC/RRBench.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @university college london 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/agentic-coding-witho…] indexed:0 read:1min 2026-07-27 ·