{"slug": "what-counts-as-valuable-company-data-is-changing", "title": "What counts as valuable company data is changing", "summary": "Google agreed to pay $10 million for the internal business data of bankrupt Spirit Airlines, including emails, Teams messages, spreadsheets, and operating records, which it plans to use for product development and AI. Spirit flight attendants objected to the sale and sought additional protections for employee data, and a bankruptcy judge has delayed approval while the dispute is resolved. A developer argues that such workflow history — not just final documents — is becoming a valuable asset for training AI agents, citing research that compiled recurring agent traces into reusable workflows and cut a task from 34 API calls to 11.", "body_md": "# **When Work History Becomes an AI Asset**\n\n[Google recently agreed](https://edition.cnn.com/2026/08/18/business/google-spirit-airlines-data?utm_source=gradientflow&utm_medium=newsletter) to pay $10 million for the internal business data of Spirit Airlines. The airline is bankrupt, but its emails, Teams messages, software, spreadsheets, and operating records apparently still have value. Google plans to use the material for product development and AI. Spirit’s flight attendants objected to the sale and sought additional protections for employee data, and a bankruptcy judge has delayed approval while the dispute gets sorted out.\n\nI find the deal interesting for a different reason. I [recently noted](https://gradientflow.substack.com/p/i-think-ai-teams-are-defending-the?r=ks4p&utm_campaign=post&utm_medium=web&showWelcomeOnShare=true) that as access to capable models gets easier, more of the durable value moves into what a company learns by operating its systems. The Spirit deal made me want to look more closely at one part of that argument: the record of how people actually do the work. That made me wonder whether this kind of operating history is more valuable than companies have realized.\n\n#### **From Outputs to Process**\n\nMost of what companies keep is a record of results. You have the finished financial analysis, the resolved support ticket, the code that shipped, or the recommendation that went to a customer.\n\nAn agent trying to perform the work may need more than the final result. It can benefit from seeing which information someone looked for, what they ignored, which tools they used, where they got stuck, how they corrected course, and how they decided the job was finished. **Final documents can tell an AI what a company knows. Workflow history can teach it how the company operates.**\n\nThis is starting to move beyond theory. [Researchers recently used](https://arxiv.org/abs/2607.04948) software-development event logs to infer the roles people perform and generate corresponding agent specifications. Companies are also [building training environments](https://www.mercor.com/blog/mercor-to-acquire-deeptune/) where agents can practice complete tasks and get scored on whether they actually accomplished them.\n\nThe underlying idea should be familiar. People get better at a job by accumulating experience, not just by reading the finished work of people who came before them. Agents may need something similar.\n\n#### **The Archive Is Not the Asset**\n\nThat does not mean your Slack archive is suddenly a competitive moat. Most workplace exhaust is probably just exhaust. I think there are three useful layers here. At the bottom are messages, meetings, tickets, commits, clicks, and logs. Above that are **trajectories**, where those fragments are connected into a sequence from goal to decision to action to outcome. The most valuable layer is what you can build from those trajectories: good examples, failure cases, evaluations, corrections, and reusable workflows.\n\nOne recent system illustrates the progression nicely. [It looks across messy traces left by agents](https://arxiv.org/abs/2608.02680), finds procedures that recur, and turns some of them into executable workflows. In [one example](https://arxiv.org/abs/2608.02680), a task that had previously involved 34 API calls was reduced to 11 when the recurring procedure was compiled. There was extra work involved in doing the compilation, so this is not evidence that every workflow suddenly gets cheaper. What matters is that experience that used to disappear into logs can sometimes be turned into something reusable. **The raw archive is not the asset. The asset is the ability to [turn operating experience into something a machine can learn from and reuse](https://gradientflow.substack.com/p/the-missing-layer-in-todays-agent).**\n\nThis also creates a preservation problem. Companies may worry about valuable workflow data leaking when the more immediate issue is that they never captured the valuable parts. A decision gets separated from its outcome. An unusual exception disappears into chat. Someone fixes a problem but nobody records why the fix worked. Companies may keep plenty of activity and still lose much of the experience.\n\n#### **Same Model, Different Operating History**\n\nImagine two competitors with access to essentially the same capable model and comparable internal documents. The first connects the model to those documents. The second also has hundreds or thousands of examples showing how experienced employees investigate problems, choose tools, deal with unusual cases, recover from mistakes, and check their work. It has connected those examples to outcomes and built evaluations that tell it whether its agent is actually improving.\n\nThe second company does not just have more data. It has built a learning system around its own operating experience. Companies increasingly have [credible choices at the model layer](https://gradientflow.substack.com/p/the-big-ai-labs-are-suddenly-competing), even if open models are not always the cheapest or best option for a particular task.\n\nSo I would not frame this as open versus closed AI. A frontier API might be right for one workload, a hosted open model for another, and something running inside your own infrastructure for a third. **Once a workflow contains valuable company know-how, the question is no longer just which model works best. It is also where you are comfortable having that know-how processed and stored.**\n\n#### **Start With the Work Worth Teaching**\n\nIf I were running an enterprise AI program, I would start by identifying work that is valuable, comes up often, requires real expertise, and produces outcomes I can evaluate. A simple test is whether you would invest meaningful time teaching a new employee to do the task well. If you would, ask whether some of that accumulated experience could be made useful to an AI system too.\n\nBut I would not start recording everyone all day. The useful material is not the volume of activity. It is the decisions, exceptions, failed approaches, corrections, and feedback that explain what good performance looks like. The goal is not a bigger archive. It is a better record of how the work succeeds.\n\nI would also settle the rights question early. Can the company use those records for AI, and what employee or customer information is mixed in? Teams should also decide what can be transferred after an acquisition, what an outside model provider may retain, and which workflows are distinctive enough that tighter control is worth the additional operational burden.\n\nThe [Spirit dispute is a good preview](https://arstechnica.com/tech-policy/2026/08/flight-attendants-freaked-out-that-google-to-buy-tons-of-spirit-employee-data/?utm_source=gradientflow&utm_medium=newsletter) of what happens when those questions wait until after somebody decides the data is valuable. Companies have spent years preserving the outputs of work. Agents may make the process behind those outputs just as important. The opportunity is not to save everything. It is to recognize which parts of that experience are worth keeping, and [turn them into something](https://gradientflow.substack.com/p/the-missing-layer-in-todays-agent) an AI system [can actually learn from.](https://gradientflow.substack.com/p/the-big-ai-labs-are-suddenly-competing)\n\n# **[Water, Power, Noise: The Local Impact of Data Centers](https://gradientflow.com/ai-data-centers-water-noise-power/)**\n\n# AI’s Knowledge-Work Ratchet\n\n**[Ben Lorica](https://gradientflow.com/disclosure/)** edits [Ethics.dev](https://ethics.dev/) and the [Gradient Flow newsletter](https://gradientflow.substack.com/), and he hosts the **[Data Exchange podcast](https://thedataexchange.media/)**. He helps organize the **[AI Conference](https://aiconference.com/?utm_source=gradientflow&utm_medium=newsletter)** and the **[Agent Conference](https://agentconference.com/?utm_source=gradientflow&utm_medium=newsletter)**. You can follow him on [Linkedin](https://www.linkedin.com/in/benlorica/), [X](https://x.com/bigdata), [Mastodon](https://indieweb.social/@bigdata), [Reddit](https://www.reddit.com/r/GradientFlow/), [Bluesky](https://bsky.app/profile/gradientflow.com), [YouTube](https://www.youtube.com/c/GradientFlow), or [TikTok](https://www.tiktok.com/@gradientflow). This newsletter is produced by [Gradient Flow](https://gradientflow.com/blog/).", "url": "https://wpnews.pro/news/what-counts-as-valuable-company-data-is-changing", "canonical_source": "https://gradientflow.substack.com/p/what-counts-as-valuable-company-data", "published_at": "2026-09-15 13:00:54+00:00", "updated_at": "2026-09-15 13:14:55.726684+00:00", "lang": "en", "topics": ["ai-agents", "artificial-intelligence", "ai-products", "ai-research"], "entities": ["Google", "Spirit Airlines", "Mercor", "DeepTune"], "alternates": {"html": "https://wpnews.pro/news/what-counts-as-valuable-company-data-is-changing", "markdown": "https://wpnews.pro/news/what-counts-as-valuable-company-data-is-changing.md", "text": "https://wpnews.pro/news/what-counts-as-valuable-company-data-is-changing.txt", "jsonld": "https://wpnews.pro/news/what-counts-as-valuable-company-data-is-changing.jsonld"}}