cd /news/artificial-intelligence/pew-research-finds-ai-authorship-acr… · home topics artificial-intelligence article
[ARTICLE · art-107212] src=letsdatascience.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Pew Research Finds AI Authorship Across New Webpages

Pew Research Center's August 20 analysis found that 35% of English-language webpages published after ChatGPT's November 2022 release showed significant signs of AI authorship or editing, based on a study of nearly 500,000 archived pages using the Open Pangram detection tool. The study also found that 10% of pages in a mixed-age July 2026 sample showed AI-authorship signals, with .com domains showing the highest share at about 10%, compared to 4.6% for .org and roughly 1% for .edu and .gov domains. Pew cautioned that AI detectors can misclassify text, making the findings a population-level pattern rather than definitive authorship determinations.

read3 min views1 publishedAug 22, 2026
Pew Research Finds AI Authorship Across New Webpages
Image: Letsdatascience (auto-discovered)

Pew Research Center's August 20 analysis found that 35% of English-language webpages published after ChatGPT's November 2022 release showed significant signs of AI authorship or editing. The study analyzed nearly 500,000 pages from the previous five years using the Open Pangram detection tool. Pew cautioned that AI detectors can misclassify both human and AI-written text.

Pew Research Center found that 35% of English-language webpages published after ChatGPT's November 2022 public release showed significant signs of AI authorship or substantial AI editing. The August 20 study analyzed nearly half a million archived webpages spanning roughly five years, then used the Open Pangram AI-detection tool to identify language patterns associated with machine-generated text.

In a random July 2026 sample of 10,000 webpages that included both old and new material, Pew found significant AI-authorship signals in 10% of pages. The higher 35% figure applies only after excluding pages published before ChatGPT, which Pew noted could not have been produced with modern generative AI tools.

What the study measured

Pew describes its result as evidence of likely AI writing or substantial editing, rather than a count of webpages generated entirely by models. Open Pangram evaluates linguistic patterns, including words, phrases, and other stylistic features that occur more often in AI-generated than human-written text.

That distinction matters. A published page can include AI-produced passages alongside human reporting, editing, research, or revision. TechCrunch reported that Pew's methodology was designed to capture both direct AI authorship and substantial AI-assisted editing.

Pew also explicitly cautioned that automated detection is imperfect: such systems can classify human text as AI-generated and AI-generated text as human-written. The study uses a large corpus to measure a population-level pattern, not to make definitive authorship determinations for individual pages.

Distribution differs by domain

The study found that AI-authorship signals are unevenly distributed across top-level domains. CNET's account of the Pew findings reported that .com pages had the highest observed share, at about one in 10 pages, compared with 4.6% for .org pages and roughly 1% for .edu and .gov pages.

Pew reported that these domain-level differences were less apparent around ChatGPT's launch and widened afterward. The source material does not establish why the rates differ across commercial, nonprofit, educational, and government domains.

Implications for data and content systems

For ML teams building web corpora, retrieval systems, search products, or evaluation sets, the findings reinforce a broader data-provenance problem. Comparable shifts in web publishing can change the composition of newly crawled text faster than historical datasets reveal, while detector outputs remain probabilistic rather than ground truth. That makes source metadata, publication dates, deduplication, quality filtering, and human review important complements to AI-content detection.

The research arrives amid wider discussion of automated content production and consumption online. TechCrunch noted separately reported Cloudflare data indicating bot traffic had exceeded human traffic, but Pew's analysis concerns the authorship characteristics of webpages, not the identity of their visitors.

Key Points #

  • 1Pew found AI-authorship signals in 35% of pages published after ChatGPT, versus 10% across its mixed-age July sample.
  • 2Commercial .com domains showed higher detected AI-content rates than .org, .edu, and .gov domains, according to coverage of Pew's findings.
  • 3Web-corpus builders increasingly need provenance and quality controls because detector results are probabilistic and AI-assisted text can be partially human-authored.

Scoring Rationale #

The study provides a large-scale, recent measurement of AI-authorship signals in web text, a core input to training-data, retrieval, and search workflows. Its detection-based methodology limits page-level certainty, but the reported scale is notable for practitioners managing web-derived datasets.

Sources #

Primary source and supporting public references used for this report.

Practice interview problems based on real data

1,625 SQL & Python problems across 15 industry datasets — the exact type of data you work with.

Try 250 free problems

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @pew research center 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/pew-research-finds-a…] indexed:0 read:3min 2026-08-22 ·