Recomma docs

Site

Your own site as the engines meet it — what was crawled, what the markup says, and who is allowed to read it.

Every other page in Recomma is downstream of an answer. This one is not: it is the state of your own site whether or not a single chat has been sampled, and whether or not anybody has decided to act on it.

The Site page: crawl counts across the top, the crawler list beside the failing markup rules, and the page inventory below.
What was crawled, which AI crawlers the site restricts, and every rule it currently fails.

It reads from two things the product was already collecting and neither of which had a home: the crawl of your sitemap, and the audit findings behind Site health tickets. Findings could previously only be reached through a ticket, which meant the state of your site was legible only where somebody had already decided to do something about it.

Crawl

Four counts, from the sitemap inward:

CountWhat it counts
FoundPages discovered in the sitemap
ReadPages fetched successfully
FailedPages that answered with an error
WaitingDiscovered, not yet fetched

Pages are found from your sitemap rather than from your citations, because the pages with problems are usually the ones nothing has quoted yet. It is a backlog worked on the schedule, so Waiting above zero is normal on a large site and is the reason an audit can be a true statement about part of a site.

Crawlability

What your robots.txt says about the crawlers that fetch pages in order to answer questions — GPTBot, ClaudeBot, PerplexityBot and the rest.

Only findings are listed. Most sites carry housekeeping disallows that every crawler on the web falls under, and reporting those would say every site blocks every agent.

  • Blocked — the agent is shut out of the site.
  • Partly blocked — shut out of some paths, which are listed beneath.
  • named — somebody wrote this rule for this agent specifically, rather than it falling under the wildcard with every other bot.

Blocking a training crawler is a decision, not a fault

Only the crawlers that fetch a page in order to answer a question matter here. Keeping CCBot or Bytespider out costs you nothing in the engines this product measures — they are listed because a reader asking “who can read my site” wants the whole answer, not because they are a finding.

If robots.txt could not be fetched, the panel says so rather than reporting the site as open. Unreachable is not permissive.

Markup

Every audit rule the site currently fails, with how many pages fail it. Open a row and it gives the rule in a sentence, the pages it fires on, and the markup on each one.

The State column is the link back to work:

  • ticketed — an action is open against this rule on Opportunities.
  • untriaged — no ticket. Which is not the same as no problem: a ticket is raised when a fault clears an impact bar, and a rule failing on four pages of a large site may never clear it. Somebody reading their own site should see it either way.

That distinction is why this page is read-only. A ticket is work somebody has decided to do; this is the site’s condition whether anybody has decided anything, and the two answer different questions. Run the audit, take tickets and mark them done on Opportunities.

Pages

The crawled inventory, worst first: pages that failed, then pages nothing has read yet, then the oldest readings. Each row carries the URL and the status the crawler got back.

The table pages through the whole inventory rather than showing a sample of it.