

Oops, realized I didn’t answer your question about actually crawling, dig into common crawl documentation, they provide a bunch of technical data and stats that show you the scale…2-4billion pages per month
And note CC just does a sample of the pages it finds. So the more monthly dumps don’t contain all of the data afaik
And the number above are for one of the monthly dumps
https://commoncrawl.github.io/cc-crawl-statistics/plots/crawlermetrics








So this would work for you?
“She floats, so she’s a witch, and all witches are guilty so she’s guilty”