Public AI readability report

blog.dogtorcito.com

Scanned Fri, 07 Aug 2026 11:09:02 GMT · https://blog.dogtorcito.com/

77
out of 100

Critical access or response-content failures found

Your server or CDN blocks 4 AI crawlers at the network layer

Access52/100
Can crawlers fetch it
Readability100/100
Is content in the HTML
Structure100/100
Machine-readable meaning
Answerability60/100
Shaped to be quoted

Send this to whoever can deploy the fix

We’ll email the full findings plus a prioritised fix checklist for blog.dogtorcito.com. Access failures come first, followed by response-content and structure fixes. No account needed.

Initial response and browser render

Word counts of readable text before and after browser execution. The gap is a rendering dependency, not a claim about every vendor’s crawler behavior.

Initial HTML response
625
words of readable text
Browser render after JavaScript
625
words of readable text

Crawler access

Live requests were sent with each vendor’s published user-agent string. This tests how the site handles the claimed identity; a user-agent alone does not authenticate the vendor. A robots.txt pass with a failed request points to another access control.

CrawlerVendorRolerobots.txtLive request
GPTBotOpenAITrainingBlockedHTTP 403
OAI-SearchBotOpenAISearch indexretrievalAllowedHTTP 403
ChatGPT-UserOpenAILive fetchretrievalAllowednot probed
ClaudeBotAnthropicTrainingBlockedHTTP 403
Claude-SearchBotAnthropicSearch indexretrievalAllowednot probed
Claude-UserAnthropicLive fetchretrievalAllowednot probed
PerplexityBotPerplexitySearch indexretrievalAllowedHTTP 403
Perplexity-UserPerplexityLive fetchretrievalAllowednot probed
Google-ExtendedGoogleTrainingBlockednot probed
Applebot-ExtendedAppleTrainingBlockednot probed
meta-externalagentMetaTrainingBlockednot probed
AmazonbotAmazonSearch indexBlockednot probed
BytespiderByteDanceTrainingBlockednot probed
CCBotCommon CrawlTrainingBlockednot probed

Findings (3)

1 critical access or response issue found in this scan.

critical

Your server or CDN blocks 4 AI crawlers at the network layer

Evidence
OAI-SearchBot: HTTP 403; PerplexityBot: HTTP 403; ClaudeBot: HTTP 403; GPTBot: HTTP 403
Why it matters
robots.txt permits OAI-SearchBot, PerplexityBot, but the live test using those published identities was refused before it reached the content. A robots.txt-only check would miss this access layer.
Fix
Review the WAF or bot-management event that handled the request. Allow the crawler identities you chose. For a security bypass, verify vendor IP ranges or reverse DNS instead of trusting a user-agent string alone.
medium

No question-shaped content

Evidence
No question headings and no FAQPage schema found.
Why it matters
Assistants answer questions. Pages organised as explicit question/answer pairs are markedly easier to retrieve and quote than undifferentiated marketing prose.
Fix
Add a server-rendered FAQ section using the real questions prospects ask, marked up with FAQPage schema.
low

robots.txt blocks 8 AI training crawlers

Evidence
GPTBot (OpenAI), ClaudeBot (Anthropic), Google-Extended (Google), Applebot-Extended (Apple), meta-externalagent (Meta), Amazonbot (Amazon), Bytespider (ByteDance), CCBot (Common Crawl)
Why it matters
This can be a deliberate publisher choice. These controls apply to training use, while vendors document separate agents for search and user-triggered retrieval.
Fix
No action needed if this is intentional. Verify retrieval crawlers remain allowed.

Get these fixed and keep them fixed

The diagnosis stays free. Paid plans generate reviewable robots.txt, JSON-LD and llms.txt artifacts, then re-scan on schedule and email you when a check regresses.

See plansScan another site
AI readability report for blog.dogtorcito.com — 77/100 · Crawlable