Search and SEO
What is Index bloat?
Index bloat is having large numbers of low-value URLs indexed — filter combinations, internal search results, thin tag archives, printer-friendly duplicates. It dilutes the perceived quality of a site and spends crawl attention on pages that will never earn a visitor.
Also called: bloated index
How to detect it
Compare the number of URLs you deliberately publish against the number reported as indexed. A site with 80 real pages and 4,000 indexed URLs has a problem that no amount of content work will fix.
The usual sources
- Faceted navigation and sort parameters
- Internal site search results, which are infinite
- Tag and archive pages generated automatically by a CMS
- Paginated series with no canonical strategy
- Auto-generated location or keyword pages with no unique content
Fixing it without making things worse
Noindex the low-value pages and allow them to be crawled so the noindex is read. Only once they have dropped out should you consider a robots.txt disallow to save crawl attention.
Doing it in the other order strands the URLs in the index with no description and no route out.
Where this is covered in depth
A definition can only go so far. The e-commerce SEO checklist, in the order the checks actually matter covers this properly — 3 minutes, free, no email required.
Who wrote this
Anas Bin Masud builds e-commerce sites and does technical SEO for businesses in the UK, Canada and Pakistan. These definitions come from client work rather than from a content brief — where an entry describes a mistake, it is usually one found on a real site. More about how I work.