The null result in OpenAI's enterprise AI paper OpenAI, in a working paper with Columbia and Wharton based on ChatGPT Enterprise telemetry through March 2026, found that among adopting firms, larger headcount predicts lower usage per employee, but the fourth column shows no significant difference in messages per weekly active user, indicating the usage gap is due to penetration, not engagement. The paper also found that SG&A stock per employee (a proxy for organizational capability) was the strongest predictor of adoption, while physical capital intensity was negatively associated. Hype-o-meter The null result in OpenAI's enterprise AI paper OpenAI published a working paper this week with Columbia and Wharton, built on ChatGPT Enterprise telemetry through March 2026. The findings everyone will quote are in the abstract. The one worth your time is a coefficient in Table 2 that the paper reports and then moves past. It's a good paper. It's also a vendor publishing data about its own customers, with two co-authors on its payroll, so read it the way you'd read anything else in that category. The headline findings are what you'd expect. Usage grew sevenfold. Adopters skew large. Task use is broad rather than concentrated. All plausible, none of it especially surprising. What I keep coming back to is elsewhere. Among firms that have already adopted, larger headcount predicts lower usage per employee. Fewer messages per head, fewer weekly active users per head, fewer tokens per head, all statistically significant. This is presented, reasonably, as a scaling effect. Big companies dilute. But there are four columns in that table, and the fourth is messages per weekly active user. Nothing. Flat. A standard error four times the size of the estimate. A 200,000-person company's engaged AI users are as engaged as a 500-person company's engaged AI users. One honest caveat before I lean on this. That fourth regression has an R² of 0.100, against 0.36 to 0.48 for the other three. It's a noisier specification, so the null is weaker evidence than a null in a tightly fitted model would be. But the point estimate is essentially zero, not merely imprecise, and the sign flips nothing. Which means the large-enterprise usage gap has nothing to do with how people use the tool once they're using it. It's entirely a question of how many people are using it at all. If your problem is engagement, you buy training. If your problem is penetration, none of that touches the constraint. I don't think this is a small distinction, because the two readings send you to completely different places. If your problem is engagement, you're buying better prompts, better training, better internal evangelism, better use case libraries. That's the entire enablement industry right now. If your problem is penetration, you should be looking at who never logged in a second time and why. It's easy to measure the first thing and assume it explains the second. This paper is decent evidence that it doesn't. What the data can't tell you is why the non-returners don't return, and that's the question the null result makes urgent. The dataset has no denominator for who was offered a seat and declined it, no record of the second session that never happened. Whatever is happening there is happening before any of the enablement machinery gets a chance to work, which is a fairly awkward finding for an industry that has organised itself almost entirely around the post-onboarding phase. The second thing worth sitting with is what predicts adoption in the first place. The researchers tested three kinds of accumulated intangible capital, all measured in fiscal 2021, well before any of these firms bought anything: R&D stock per employee, capitalised software per employee, and SG&A stock per employee. SG&A is the crude proxy here for organisational and managerial capability. It came out strongest and most robust. Physical capital intensity ran the other direction, negatively associated with adoption once you control for scale. So the firms that moved first weren't the ones with the best engineering or the deepest software estate. They were the ones that had already spent years building the overhead function that redesigns how work happens. The unglamorous layer. The layer that gets cut first in a downturn and that no one puts in an earnings call. I find this uncomfortable and probably correct. It's consistent with everything we know about general purpose technologies, and it's consistent with the pattern where the technically strongest organisation in a given industry is frequently not the one that adapts fastest. Capability to absorb is a separate asset from capability to build, and we're not very good at measuring or funding it. It also suggests a fairly grim near-term dynamic. If the firms best positioned to absorb this are also the largest and most valuable ones, and the models themselves are broadly available at similar prices to everyone, then general availability doesn't level anything. It widens the gap. The paper says as much, briefly, and then leaves it alone. The third piece is the one that changes what I'd want to look at internally. Six months after adoption, the intensity gradient by seniority runs the wrong way relative to how most governance is designed. Early-career workers and trainees send roughly eight to nine more messages per week than the average active user at the same firm. Executives, founders and partners send fewer. Analysts and marketing and communications staff also run hot. And the task mix splits along the same line. Across the whole sample, the dominant activities are documentation and technical writing — more than half of all active users — technical digital work, and drafting messages. Production. Executives cluster somewhere else: topic overviews, facts and figures, legal and regulatory questions, financial and tax queries. Consumption and synthesis. None of that is scandalous on its own. But put it together and the operating picture is that your least experienced people are producing a high volume of written material with model assistance, your most senior people are reading summaries, and the review layer between them was designed for a world where the junior output was slower and more obviously effortful. There's a related finding on task risk. Legal, regulatory, financial and tax tasks show up as widespread by user reach but small by message share. A lot of people touch them occasionally. If you're monitoring by volume, which most organisations are because volume is what the admin console gives you, you will sample those categories least and they are the ones where a bad output is most expensive. That's a straightforwardly fixable measurement error and I'd guess it's near-universal. Some caveats, because they matter more than usual here. Everything above is correlational. The paper is careful about this and I'll be careful too. It measures usage only inside ChatGPT Enterprise: not the API, not internal tooling, not other vendors, not personal accounts. Firms classified as non-adopters may be running heavy AI programmes that this dataset simply cannot see, which attenuates every adopter-versus-non-adopter comparison in the paper. Job titles are point-in-time and incomplete. The financial analysis covers US public companies only, so it tells you nothing about mid-market or private firms. And critically, none of this measures output. Not quality, not revenue, not headcount. Usage is not productivity and the authors never claim it is. The framing I'd push back on hardest is the implicit one, which is that more tokens is better. Most of the paper's intensity measures are volume measures. It's entirely possible for a firm to be in the bottom quartile on messages per employee and extracting more value than someone in the top, and nothing in this dataset would distinguish them. What the paper does establish, and what I think holds up, is in its conclusion: adoption is only the beginning of deployment. Firms aren't deciding whether to use this anymore. They're working out where it belongs, and that process is slow, organisational, and mostly invisible from outside. The gap that opens over the next two years won't come from model access. Everyone has that. It'll come from what happens in the eighteen months after the contract is signed, and almost nobody is measuring that period well enough to know whether they're winning it. Chatterji, A., Holtz, D., Rakholia, N., Tambe, P., & Weeratunga, G. 2026 . How Organizations Use AI: Evidence from ChatGPT. Working paper, 11 August 2026. PDF https://cdn.openai.com/pdf/how-organizations-use-chatgpt.pdf One of these a week. No hype. The Working Model reads the research, the filings and the release notes so you don't have to, and tells you what actually changed. Free.