A Pew Research Center analysis has found a sharp rise in web pages showing signs of AI-authored text since ChatGPT’s launch in late 2022, according to The Decoder. The study examined nearly half a million English-language web pages and used the Common Crawl web archive as its source material. The Decoder reports that Pew checked the pages with Open Pangram, an AI detection tool. In a July 2026 sample, about 10% of all pages examined showed clear signs of AI authorship. When Pew narrowed the sample to pages published after ChatGPT’s release, the share was much higher: more than a third of those newer pages showed signs of machine authorship. The domain split was uneven. According to The Decoder’s summary of the Pew analysis, about one in ten .com pages showed signs of AI authorship. The reported figure was 4.6% for .org domains, and about 1% each for .edu and .gov pages. On that basis, commercial websites were roughly ten times more likely than school or government pages to contain AI-written text. Pew’s analysis also looked at language patterns that have become more common on the web since 2023. The Decoder reports that em dashes appeared about twice as often as they did in 2023, while Oxford comma use rose 63%. Several words described as AI-favored — including “delve,” “interplay,” “pivotal,” “landscape,” “meticulous,” and “vibrant” — more than doubled in frequency. The study also tracked a pattern built around negative parallelism, such as the “it’s not just X, it’s Y” construction. The Decoder reports that this pattern nearly tripled, though it remains rare in absolute terms. A separate study of corporate public-relations documents found that the phrase quadrupled since 2022, according to the same report. The Pew findings are broadly consistent with another April 2026 study cited by The Decoder from Imperial College London, the Internet Archive, and Stanford University. That work reportedly found that roughly 35% of newly published websites were fully or partly AI-generated. It also found 33% higher semantic similarity between AI texts and a more positive tone overall, while cautioning that public perceptions of harm often run beyond what the data can establish. The central caveat is measurement. The Decoder notes that there is no settled definition of “AI text”: a page could be fully generated by a model, drafted by a person and polished with AI, or mostly human-written with model-generated passages inserted. Detection tools can flag likely machine authorship, but they cannot reliably determine how much of a text came from AI or at what stage of writing the tool was used. That distinction matters for policy and interpretation. The reported Pew numbers indicate that more than a third of newer pages in the sample showed signs of AI authorship, with higher detected rates on commercial domains. They do not, on their own, show how much of that material was generated end-to-end by models, how much was edited by humans, or whether the resulting text is better or worse than what it replaced. Who benefits: Too early to tell. The source material reports detected prevalence of AI-like text, but does not establish specific beneficiaries. Who's exposed: Publishers and organizations that need clear authorship standards are exposed to ambiguity, because detectors can flag likely AI text but not reliably describe how AI was used. In the reported domain breakdown, .edu and .gov pages had lower detected rates than .com pages.