Ranking Recovery · Part 6
What is a sitemap?
A sitemap is a file on your website that lists the addresses of your pages for search engines. It is not a page for visitors. It is a plain list, usually at an address like yourdomain.com/sitemap.xml, written in a simple format called XML. Each entry is one page address, often with the date that page last changed.
Its purpose is discovery. A search engine finds pages by following links. A sitemap hands it the list directly, so nothing depends on a link being found. For a small, well-linked site, a sitemap is a convenience. After a hack, it becomes one of the few direct lines you have to Google — and it was very likely one of the tools used against you.
What is a sitemap XML file made of?
Open one in a browser and you will see something like a list of blocks, each holding two things.
- loc — the full address of a page.
- lastmod — the date the page was last modified.
Larger sites use a sitemap index: one file that lists several sitemaps, each holding a group of pages. WordPress produces one automatically at /wp-sitemap.xml, split into posts, pages and other types.
Two rules matter more than the format. A sitemap should list only pages you want indexed. And every address in it should answer with a success code, not a redirect or an error. A sitemap is a statement of what your site consists of. It should be true.
How attackers use sitemaps
The same feature that helps you get pages found helps an intruder get spam found. A sitemap is an invitation to crawl, and it does not need your permission to be issued.
From the case file. The attackers in this case never touched the Sitemaps report in Search Console. They did not need to. They wrote a robots.txt file on a quiet site in the same hosting account that declared twenty sitemaps, and added server rules that sent every sitemap request to a script generating them on demand. A sitemap named in robots.txt gets crawled without anyone submitting anything. The owner’s Search Console showed a clean submitted-sitemaps list the whole time.
So after a hack, check three places, not one.
- Search Console → Sitemaps. Anything submitted that you do not recognize.
- Your robots.txt file. Any
Sitemap:line you did not write. - The sitemap addresses themselves. Open each one and read what it lists.
The sitemap you need after the cleanup
Once the site is clean, the sitemap’s job is to tell Google what the real site is. It should be boring and exact.
- Only real pages. No spam addresses, no deleted pages, no test pages.
- Only pages that answer 200. If a listed address redirects or errors, take it out or fix it.
- Honest dates. The lastmod date is how Google learns that a page changed. If your pages were restored or rewritten after the hack, the dates should say so.
- No clutter. WordPress adds archive listings for authors, categories and tags. On a small business site these are rarely pages you want found, and a user sitemap can reveal login names.
Then submit it in Search Console, and submit it again after significant changes. Resubmitting does not force a crawl, but it prompts Google to read the list afresh.
Google Search Console sitemap status: what to watch
The Sitemaps report shows each sitemap you submitted, when it was last read, and how many pages were discovered in it.
Last read is the line to watch. If Google read your sitemap yesterday, your list of real pages is current in its hands. If the date is weeks old, resubmit.
Status should say Success. “Couldn’t fetch” after a hack often means a leftover rule is interfering with the address — sometimes the attackers’ own rule that hijacked every request ending in .xml.
Discovered pages should match the number of pages you actually have. A number far larger than your site means something is listing addresses that are not yours.
The cache problem nobody mentions
If your site sits behind a service like Cloudflare, the file Google receives may not be the file on your server. It may be a copy the service saved hours ago.
For robots.txt in particular, this matters. You remove a bad rule, check the file on the server, see the correction — and Google goes on receiving the old version from the cache. The fix exists and has not been delivered.
After changing robots.txt or a sitemap, fetch it from outside and check what actually comes back. If the service is holding an old copy, clear it, and consider excluding those two files from caching altogether. The hardening series covers this setting.
Using links to get dead addresses re-fetched
Here is the difficulty from earlier in this series. Thousands of spam addresses answer “gone,” but Google only learns that when it requests one. Many were discovered through links and never requested. They sit in the records, waiting for a visit that is not scheduled.
The attackers got those addresses discovered by putting links to them where Google would find them. The same mechanism works in reverse.
From the case file. The owner built plain pages on his own site consisting of nothing but links to the dead spam addresses, about a thousand to a page. He then used Search Console’s Test Live URL on each of those pages, which had Google fetch the page and see its links. Google followed them, requested the addresses, and received 410 for each. He repeated it in batches. He is candid that he cannot prove how much time it saved, since there is no second copy of the site to compare against. But it gave Google a reason to ask, and every address it asked for was told the truth.
A temporary sitemap listing the dead addresses is a variation on the same idea. Either way the principle holds: you cannot make Google forget an address, but you can give it cause to check one.
Two cautions. Do this only when every listed address reliably answers 410, verified from outside. And take the pages down when the job is done. They are scaffolding.
A sitemap is not a ranking tool
A sitemap gets pages discovered. It does not make them rank, and listing a page does not oblige Google to index it. If a page is in your sitemap and still excluded, the sitemap has done its job and the page has not passed the next test. That is a question about the page itself.
With the real pages listed and being read, the remaining question is how Google understands them. The Free Visibility Snapshot shows what Google sees when it reads your site today. If the sitemap and structure need rebuilding and you would rather do it yourself, the DIY SEO Audit sets out those fixes in order.
See what Google seesFree · the Free Visibility Snapshot™ · a narrow look at four signals, not an audit
Back: Google Search Console Request Indexing: Getting Your Real Pages Back
Next: Google Search Console Not Updating: Why Recovery Data Lags
Hub: Ranking Recovery Blog Series
ProVAE builds and recovers websites in Douglas, Georgia, serving South Georgia.
