I run a calculators-and-tools site, solo, five months old. About 5,800 pages across four language sections. Google has 642 of them in its index. Not great. The rest sit in two buckets: 1,153 "crawled, currently not indexed" and 4,161 "discovered, currently not indexed."
Like everyone in that situation, I had a theory. Somewhere in those 5,800 pages I had generated 1,432 timezone pages — every URL a pair of cities, /date-time/time-difference/london-and-buenos-aires, that sort of thing. One calculation, different arguments, spread across a URL space. It looks exactly like what Google's spam policy calls scaled content abuse, and I assumed it was dragging the whole host down. The plan was to collapse them into one hub page and 301 the rest away.
In Search Console: Indexing → Pages, then the link "View data about indexed pages." At the bottom of the examples table, change rows-per-page to 500. The examples table caps around 1,000 URLs, so if your indexed count is under that, you get the complete list, not a sample. Mine was 642, and I pulled 638 unique URLs out of it.
Then classify them. I did it with a throwaway script that buckets by URL prefix:
Do the same for the "crawled, not indexed" and "discovered, not indexed" drilldowns and you get the whole picture in three exports.
One trap if you automate the reading: the console reuses its DOM between drilldowns. My first pass at the "crawled, not indexed" list came back looking almost identical to the indexed list, and it took a set intersection to notice that 638 of those 660 rows were the indexed list, still sitting in the page. Reload the drilldown URL directly before reading it.
| Family | Total | Indexed | Rate | Crawled, rejected | Never crawled | |---|---|---|---|---|---| | Generated timezone pairs | 1,432 | 428 | 29.9% | 647 | 24 | | Generated finance scenarios | 180 | 74 | 41.1% | 65 | 4 | | Cooking guides | 530 | 61 | 11.5% | 36 | 95 | | Tools and hubs | 1,377 | 65 | 4.7% | 45 | 407 | | Market-specific pages | 593 | 8 | 1.3% | 0 | 138 | | Handwritten recipes | 2,290 | 0 | 0% | 0 | 195 |
The two generated families I was about to delete are 502 of the 638 indexed pages. Seventy-nine percent of my index. They also have the two highest acceptance rates on the site.
The 2,290 recipes — the ones with photographs, per-serving nutrition, ingredient scaling rules, written natively in two languages rather than machine-translated — have zero pages in the index. Not rejected. Zero of them appear in the "crawled, not indexed" bucket either. Google has never fetched a single one.
(The counts in the last two columns are from samples — those drilldowns cap out around a thousand example URLs, and my buckets are larger than that. The indexed column is complete.)
The obvious reading is "Google likes generated pages better," which is nonsense. Look at the last two columns instead.
For the timezone family, Google is finished. It crawled essentially all 1,432, kept 428, threw out 647, and has almost nothing left in the queue. That is a completed judgement, and 30% is the verdict.
For the recipes, Google hasn't started. They are all sitting in "discovered, currently not indexed," which does not mean rejected. It means Google knows the URLs exist and has not spent a request on them.
Those are two completely different failure modes and the summary number blends them into one scary total. "Crawled, not indexed" is a quality verdict. "Discovered, not indexed" is a scheduling decision — Google deciding your host isn't worth more requests right now. My Links report says external links: 0. That's the whole story: crawl allowance got spent on whatever was linked from the hubs first, and the rest of the site never got its turn.
So the site does not have a content-quality problem in the family I suspected. It has a "nobody links to this domain" problem, which no amount of deleting pages will fix.
Day one, the old URLs are still in the index. Someone searches for a city pair, finds the old URL, clicks, gets a 301 to the hub with the pair pre-filled. Fine — the reader gets their answer.
Week four, Google re-crawls, sees the redirects, drops the old URLs and shows the hub. Now one generic hub page has to rank for every city-pair query on its own. A hub with two dropdowns competes badly against a page whose title is literally the query.
Net effect: hand back 67% of the index in exchange for rankings the hub would have to earn from scratch. To fix a quality problem that the data says is not there.
