cd /news/artificial-intelligence/crossaudit-a-git-native-cross-vendor… · home topics artificial-intelligence article
[ARTICLE · art-117364] src=machinebrief.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

CrossAudit: A Git-Native, Cross-Vendor Audit Loop for Agentic Science

A new protocol called CrossAudit, described in an arXiv preprint (arXiv:2608.28631v1), proposes that AI research pipelines be audited by agents from different vendors to avoid self-grading bias. The protocol, implemented with GitHub Actions and Python, was tested in a seeded-defect trial with 30 increments and 43 seeded defects, showing that two vendors read the same rulebook differently. The authors report a live deployment in a computational-chemistry pipeline and have adopted the findings of a cross-vendor audit of their own repository.

read2 min views1 publishedSep 1, 2026

arXiv:2608.28631v1 Announce Type: new Abstract: An AI scientist should not grade its own homework. Yet in the systems we examined, the agent that reviews the work usually comes from the same model family as the agent that produced it, or at least from the same vendor. Model evaluators are known to favour their own generations. Whether models trained alike also share blind spots is a conjecture, not a settled finding, but if they do, the reviewer inherits the author's. The record of what was flagged and what was waved through often sits in platform logs that nobody outside can replay. We present CrossAudit, a protocol for supervising autonomous research pipelines. It rests on three commitments. Each increment of work is audited by an agent from a different vendor against a rulebook a human wrote and versioned. Reports, verdicts, disputes and rulings are git commits, so the supervision history can be re-read and cited; raw model exchanges are not yet part of that record. Scripted checks run before any model does. Advisory judgement never gates the pipeline: a model blocks only by citing a rule, and no model may waive a deterministic failure. Blockers that survive a bounded number of revision rounds go to a person. We state the protocol as eight invariants. We describe a reference implementation built from GitHub Actions and a few hundred lines of Python, and report a live deployment of a closely related variant in a computational-chemistry pipeline. We also ran a seeded-defect trial (30 increments, 43 seeded defects, one run per configuration). A cross-vendor audit of our own repository then voided its blinding. We adopt that audit's findings and report the corrected results. The trial shows that two vendors read the same rulebook differently. It does not show that either is better. The strongest evidence here is the committed, uncontrolled record of cross-vendor audits of this paper itself.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @crossaudit 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/crossaudit-a-git-nat…] indexed:0 read:2min 2026-09-01 ·