cd/entity/Common Crawl· home entities Common Crawl
grep -l @common crawl /news/*.json | wc -l → 37

Common Crawl

mentions 37 type Person page 2/2 feed RSS

// recent coverage 37 mentions

18:34
2026-07-29
dev.to
artificial-intelligence

Block AI Crawlers: The 15 Bots That Matter

A developer at techpotions has compiled a verified registry of 15 AI crawlers across four categories, correcting the common misconception that Google-Extended and Applebot-Extended are crawlers rather…

03:12
2026-07-29
promptcube3.com
generative-ai

Generative AI Training Data: The Pirated Book Controversy

A dataset containing roughly 200,000 pirated books from shadow libraries has been used to train generative AI models, raising legal and ethical concerns about copyright infringement. The unlicensed da…

09:51
2026-07-28
promptcube3.com
artificial-intelligence

Why AI companies are digitizing rare books at the cost of

AI companies are digitizing rare books to train large language models, but the aggressive scanning process is damaging the original physical artifacts, according to a report. The shift toward 'dark da…

23:26
2026-07-22
github.com
ai-tools

Local agent first AI search optimization tooling

Canonry has released an open-source, self-hosted AI engine optimization (AEO) operating platform that tracks citations across Gemini, ChatGPT, Claude, Perplexity, and local LLMs, and includes tools fo…

17:20
2026-07-22
github.com
artificial-intelligence

GigaToken: ~1000x faster Language model tokenization

GigaToken, a new open-source tokenizer, claims to be up to ~1000x faster than HuggingFace's tokenizers and tiktoken for language model tokenization, achieving speeds of over 24 GB/s on an AMD EPYC 956…

09:58
2026-07-04
dev.to
large-language-models

The Training Data Effect: Why Some Brands Dominate AI Responses

Large language models exhibit brand bias because training data distribution determines which companies appear as defaults in AI responses. Brands that left deep textual footprints across high-quality …

14:03
2026-06-28
dev.to
artificial-intelligence

What AI Crawlers Actually Do to a Small Blog: 9 Days of Logs

A small Home Assistant blog received 18,209 AI crawler requests in nine days, accounting for 5.2% of total traffic. The majority came from ChatGPT-User (6,687 requests), which performs live fetches fo…

02:21
2026-06-22
theatlantic.com
artificial-intelligence

AI Watchdog

The Atlantic's investigation reveals that tech companies have used at least 15 million videos and millions of songs to train AI models, often without permission. The report highlights the industry's r…

08:24
2026-06-17
blog.mozilla.org
machine-learning

Firefox suggests tab groups with local AI (2025)

Mozilla launched an AI tab grouping feature in Firefox in early 2025 that suggests group titles and tabs to add, running entirely locally on the user's device using a small T5-based model fine-tuned o…

← prev page 2 / 2
// co-occurs with top 8 entities
// topics top 6 topics