{"slug": "openai-pitches-codex-for-tax-prep-after-a-7000-return-pilot", "title": "OpenAI pitches Codex for tax prep after a 7,000-return pilot", "summary": "OpenAI is pitching its Codex agent harness for tax preparation after a 2025 tax-year pilot processed 7,000 returns and cut accountants' preparation time by 31%, according to Current (formerly Crete Professionals Alliance). OpenAI president Greg Brockman highlighted the results on August 20th, positioning Codex as infrastructure for vertical agents beyond software development. The pilot, built with Thrive Holdings and Current, also increased throughput by about 50% and achieved up to 97% accuracy on drafts, with practitioners retaining review responsibility.", "body_md": "# OpenAI pitches Codex for tax prep after a 7,000-return pilot\n\n**The 2025 tax-year pilot saved accountants 31% of preparation time, according to Current, while keeping practitioners responsible for review.**\n\nBy [Ryan Merket](/author/ryan-merket)\n· Published\n\nPrimary source: [X](https://x.com/gdb/status/2090246288478814281)\n\n## Why it matters\n\nOpenAI is positioning Codex as infrastructure for vertical agents, with its Thrive stake providing the real workflows and expert feedback that generic model access cannot supply.\n\n## Video version\n\n[Greg Brockman (@gdb)](https://x.com/gdb/status/2090246288478814281), OpenAI's president and co-founder, pitched Codex on August 20th as infrastructure for products outside software development, pointing to a tax-preparation system that processed 7,000 returns and reduced accountants' preparation time by about a third.\n\nThe figures describe a pilot from the 2025 tax year, rather than a new deployment. [OpenAI first detailed the project on May 27th](https://openai.com/index/building-self-improving-tax-agents-with-codex/), after six months of work with Thrive Holdings and Crete Professionals Alliance, the accounting network that [rebranded as Current on June 2nd](https://www.current.co/news/crete-professionals-alliance-rebrands-as-current). Brockman, who was Stripe's chief technology officer before helping found OpenAI, resurfaced the results to argue that developers can use Codex's open-source harness as the operating layer for specialized agents.\n\nThat is a broader product pitch than the coding assistant OpenAI originally put in terminals and development environments. The [Codex repository](https://github.com/openai/codex) is published under the Apache 2.0 license, and its App Server exposes the agent harness through a bidirectional interface that developers can embed in other products. The open-source code handles agent threads, tool execution, configuration and approvals; access to OpenAI's models still requires a ChatGPT account or API setup.\n\n### What the tax system handled\n\nTax AI was built for Current's network of accounting firms. Participating accountants uploaded source documents and client notes, and the system extracted information and prepared submissions for tax-engine review. The pilot covered 1040 individual returns and 1041 returns for estates and trusts.\n\nOpenAI said data entry for medium- and high-complexity filings can consume as much as eight hours per return. The work includes pulling information from prior-year filings, spreadsheets and other inconsistent client documents, then mapping it to the correct tax fields.\n\nCurrent reported an average 31% reduction in preparation time across the 7,000 returns. OpenAI said the system increased throughput by about 50% and produced drafts with up to 97% accuracy. Current later described accuracy as high as 98%. Those figures are reported by the organizations that built and deployed Tax AI, and the top-line accuracy number does not describe how results varied by return complexity.\n\nOpenAI supplied a more useful measure of improvement over time. When Tax AI launched, one-quarter of evaluated returns reached at least 75% correct field completion. Six weeks later, 86% reached that threshold, even as Tax AI moved from relatively straightforward W-2 and 1099 inputs into K-1 forms, rental-property schedules and other complicated filings.\n\nPractitioners remained responsible for reviewing the work and approving final returns. OpenAI limited Codex's automated engineering work to the extraction and mapping layer, while engineers retained control over architecture, product decisions and production releases.\n\n### Corrections became engineering tasks\n\nThe central mechanism was a feedback loop built around accountants' corrections. Tax AI recorded what the system proposed, what a practitioner changed and what ultimately entered the filed return. Repeated errors could then be grouped into a finding, converted into a targeted evaluation and assigned to Codex as a bounded engineering task.\n\nCodex received the relevant production trace, source documents, expected tax-engine output, code and test commands. It could inspect a failure, propose a change and run targeted and regression evaluations. Ambiguous cases went back to engineers rather than being turned automatically into code changes.\n\nThat distinction matters in tax preparation, where a mismatch can reflect an extraction error, an accountant's judgment, a value carried over from a previous return or a change introduced elsewhere in the filing process. A raw correction is not necessarily evidence that the agent was wrong.\n\nOpenAI said rental-property support took roughly six weeks and substantial engineering oversight to reach 90% precision and recall. The resulting evaluation and review patterns were then reused for other schedules.\n\n### OpenAI has a stake in the rollout\n\nThe pilot also reflects OpenAI's strategy for moving agents into established service businesses. [OpenAI took an ownership stake in Thrive Holdings in December 2025](https://openai.com/index/thrive-holdings/), with accounting and IT services named as the partnership's first targets. OpenAI agreed to embed research, product and engineering staff inside Thrive-owned operations.\n\nThat structure gives OpenAI direct access to production workflows, expert corrections and a distribution channel across acquired businesses. Thrive gets purpose-built automation for labor-intensive services, while OpenAI gets a controlled proving ground for agent systems that need domain feedback and repeated evaluation.\n\nCurrent says its network has since grown to 48 firms with more than 2,000 employees across 39 states. The accounting group plans to expand Tax AI across additional firms, while Thrive and OpenAI are applying the same design to bookkeeping, audit and IT help-desk workflows.\n\nBrockman's post turns the pilot into a developer pitch: Codex can supply the agent loop beneath a vertical product, while domain experts, production traces and evaluations determine whether that product becomes reliable. The harness is available to copy. The difficult part remains securing the workflow, review process and proprietary feedback that made the tax pilot improve.", "url": "https://wpnews.pro/news/openai-pitches-codex-for-tax-prep-after-a-7000-return-pilot", "canonical_source": "https://runtimewire.com/article/openai-codex-tax-prep-7000-return-pilot", "published_at": "2026-08-20 01:28:31+00:00", "updated_at": "2026-08-20 01:43:48.239316+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-agents", "ai-products", "ai-infrastructure"], "entities": ["OpenAI", "Greg Brockman", "Codex", "Thrive Holdings", "Current", "Crete Professionals Alliance", "Tax AI"], "alternates": {"html": "https://wpnews.pro/news/openai-pitches-codex-for-tax-prep-after-a-7000-return-pilot", "markdown": "https://wpnews.pro/news/openai-pitches-codex-for-tax-prep-after-a-7000-return-pilot.md", "text": "https://wpnews.pro/news/openai-pitches-codex-for-tax-prep-after-a-7000-return-pilot.txt", "jsonld": "https://wpnews.pro/news/openai-pitches-codex-for-tax-prep-after-a-7000-return-pilot.jsonld"}}