{"slug": "crossaudit-a-git-native-cross-vendor-audit-loop-for-agentic-science", "title": "CrossAudit: A Git-Native, Cross-Vendor Audit Loop for Agentic Science", "summary": "A new protocol called CrossAudit, described in an arXiv preprint (arXiv:2608.28631v1), proposes that AI research pipelines be audited by agents from different vendors to avoid self-grading bias. The protocol, implemented with GitHub Actions and Python, was tested in a seeded-defect trial with 30 increments and 43 seeded defects, showing that two vendors read the same rulebook differently. The authors report a live deployment in a computational-chemistry pipeline and have adopted the findings of a cross-vendor audit of their own repository.", "body_md": "arXiv:2608.28631v1 Announce Type: new\nAbstract: An AI scientist should not grade its own homework. Yet in the systems we examined, the agent that reviews the work usually comes from the same model family as the agent that produced it, or at least from the same vendor. Model evaluators are known to favour their own generations. Whether models trained alike also share blind spots is a conjecture, not a settled finding, but if they do, the reviewer inherits the author's. The record of what was flagged and what was waved through often sits in platform logs that nobody outside can replay.\nWe present CrossAudit, a protocol for supervising autonomous research pipelines. It rests on three commitments. Each increment of work is audited by an agent from a different vendor against a rulebook a human wrote and versioned. Reports, verdicts, disputes and rulings are git commits, so the supervision history can be re-read and cited; raw model exchanges are not yet part of that record. Scripted checks run before any model does. Advisory judgement never gates the pipeline: a model blocks only by citing a rule, and no model may waive a deterministic failure. Blockers that survive a bounded number of revision rounds go to a person.\nWe state the protocol as eight invariants. We describe a reference implementation built from GitHub Actions and a few hundred lines of Python, and report a live deployment of a closely related variant in a computational-chemistry pipeline. We also ran a seeded-defect trial (30 increments, 43 seeded defects, one run per configuration). A cross-vendor audit of our own repository then voided its blinding. We adopt that audit's findings and report the corrected results. The trial shows that two vendors read the same rulebook differently. It does not show that either is better. The strongest evidence here is the committed, uncontrolled record of cross-vendor audits of this paper itself.", "url": "https://wpnews.pro/news/crossaudit-a-git-native-cross-vendor-audit-loop-for-agentic-science", "canonical_source": "https://www.machinebrief.com/news/crossaudit-a-git-native-cross-vendor-audit-loop-for-agentic-ai7p", "published_at": "2026-09-01 04:00:00+00:00", "updated_at": "2026-09-01 04:27:13.969235+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-research", "ai-safety", "ai-agents"], "entities": ["CrossAudit", "arXiv", "GitHub Actions"], "alternates": {"html": "https://wpnews.pro/news/crossaudit-a-git-native-cross-vendor-audit-loop-for-agentic-science", "markdown": "https://wpnews.pro/news/crossaudit-a-git-native-cross-vendor-audit-loop-for-agentic-science.md", "text": "https://wpnews.pro/news/crossaudit-a-git-native-cross-vendor-audit-loop-for-agentic-science.txt", "jsonld": "https://wpnews.pro/news/crossaudit-a-git-native-cross-vendor-audit-loop-for-agentic-science.jsonld"}}