We own 15 registrable domains. Some are our own projects, some are client work, and several are tools only our team should ever see. Nobody had checked all of them for the basics a search engine reads first, so we built a small script and ran it against every one.
The script visits each site logged out, the way a crawler would. It reads the robots file, the sitemap, any instruction that says "don't index this," the canonical address, the structured data, and the links to the business's social profiles. We ran it before and after our changes so we had a real comparison.
What was broken
Mostly small, boring things.
- Several sites returned a 404 for their robots file and their sitemap. Those now return the right files.
- A few returned the site's app page where a sitemap should be. A crawler that gets a web page when it expects a sitemap learns nothing.
- One site has no sitemap at all. The app belongs to another team, so we flagged it to them.
None of this shows up when you look at a site in a browser. The pages look fine. The files search engines rely on were missing or wrong.
The public record we built for CLR
The Council on Local Relations (CLR) is a UMS client, and we built its public record. Its sitemap listed 77 of its 207 pages. We added a pages sitemap, and it now lists 211 pages, all returning 200. We also added links from the site to CLR's Facebook Page and newsletter profile, so search engines can tie the site to those accounts.
Then we found a second problem. The per-town sitemap files had been copied by hand when the site was stood up and never regenerated. Some listed pages for the wrong town, and one listed a page that returns a 404. In all, crawlers were being handed 103 extra URLs next to the 211 correct ones. We changed the sitemap index so it stops naming those old files. It now points only to the place sitemap and the pages sitemap.
An older CLR domain had been answering every address with the same single line of text. It now redirects to the main site, keeping the path and query, so anyone arriving there lands in the right place.
Keeping the static sites current
Several of our sites are static, and each deploy overwrites the site's folder. If we put the robots and sitemap files in that folder, the next deploy would wipe them. So we keep them outside the folder and serve them from there. A timer regenerates them every hour, so they follow the site as it changes.
One rule worth copying: a page marked noindex is left out of the sitemap, but it is not blocked in the robots file. A crawler has to be able to read the page to see the noindex instruction. Block it in robots and the instruction is never seen.
Closing off the private tools
We run staging copies of our sites, internal dashboards and admin tools. Search engines have no business seeing them. On one server we added a noindex instruction to 14 blocks of configuration, and we added the same on another server for our internal tools. Every one of these now sends a header telling search engines not to index it, and we confirmed that from outside. The public sites do not send that header.
We left two kinds of site indexable on purpose: our video site and our podcast pages. Those are meant to be found.
What it took
We read the live sites, edited the web server configuration, wrote a generator for the sitemaps, and checked the result from outside. We backed up every config file before changing it. One change, the CLR sitemap cleanup, was written but not deployed at first because the deploy was refused by a permission check. It went live once it was approved.
What is still open
A few things remain. Several of our sites still lack the links to our social profiles. One older domain redirects to an http address when it should go to https, and that is a setting at the host. One app's sitemap belongs to the team that owns it.
The larger work is off the website. We ranked it by value:
- Verify and complete the Google Business Profiles for Unrivaled Interiors and UMS, and put the site address on each. Local-pack placement drives calls directly, so this comes first.
- Fill in the Website field on every Facebook Page, so the link runs both ways.
- Import the same businesses into Bing Places and Apple Business Connect.
- Get listed with local chambers of commerce for a high-trust local link.
What you can take from this
Open your own robots file and your sitemap in a browser, logged out. Type your address followed by /robots.txt, then /sitemap.xml.
You want a plain text file for the first and a list of your real pages for the second. If you see a 404, a copy of your home page, or addresses for pages that no longer exist, search engines are seeing the same thing. It costs nothing to check, and it is usually a quick fix.
If your sitemap was set up once by hand and never touched again, assume it has drifted. Ours had.