# robots.txt

Published articles for robots.txt.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## Cloudflare update distinguishes search crawlers from AI training bots

DevFeed: [Cloudflare update distinguishes search crawlers from AI training bots](<https://devfeed.tech/articles/cloudflare-40899.md>)

Original publisher: [Read original article](<https://habr.com/ru/news/1083142/>)

Author: neuromamontov

Published: 2026-09-16T21:09:58Z

Content type: news

Language: ru

Sources: [Tagir Valeev](<https://devfeed.tech/sources/tagir-valeev.md>)

Topics: [Cloudflare](<https://devfeed.tech/topics/cloudflare.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>)

Tags: [2025](<https://devfeed.tech/tags/2025.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-5ebc2dd0d80c](<https://devfeed.tech/tags/ai-5ebc2dd0d80c.md>), [cloudflare](<https://devfeed.tech/tags/cloudflare.md>), [geo](<https://devfeed.tech/tags/geo.md>), [llms-txt](<https://devfeed.tech/tags/llms-txt.md>), [robots-txt](<https://devfeed.tech/tags/robots-txt.md>), [tag-3ec72a25b0de](<https://devfeed.tech/tags/tag-3ec72a25b0de.md>), [tag-6dbce23d3386](<https://devfeed.tech/tags/tag-6dbce23d3386.md>), [tag-ba47f5ee56ab](<https://devfeed.tech/tags/tag-ba47f5ee56ab.md>), [tag-c94f063f9c70](<https://devfeed.tech/tags/tag-c94f063f9c70.md>), [tag-d5d895f883db](<https://devfeed.tech/tags/tag-d5d895f883db.md>)

### AI overview

The article discusses Cloudflare's update for distinguishing search crawlers from AI training bots. It explains why website owners may want to allow search indexing while restricting AI training access, and notes that the earlier llms.txt approach was not widely adopted.

### Source excerpt

Буду первым, кто не просто перепостит новость про обновление Cloudflare, а ещё и по существу прокомментирует, почему тема разделения ботов на поисковых и обучающих так важна для владельцев сайтов. Что показательно, тема ещё и совсем не новая: эти разговоры я слышал ещё года полтора назад -- тогда все обсуждали формат llms.txt, который, к слову, так толком и не взлетел. Но ключевая идея была очевидна уже тогда: желающих закрыть свой сайт от обучающих ботов в разы больше чем желающих закрыть сайт от поисковых ботов -- 17% против менее чем 1% (оценка самого Cloudflare, 2025). Читать далее

## Cloudflare Adds Setting to Block AI Training Crawlers While Allowing Search Crawlers

DevFeed: [Cloudflare Adds Setting to Block AI Training Crawlers While Allowing Search Crawlers](<https://devfeed.tech/articles/cloudflare-just-gave-ai-training-bots-the-middle-finger-31387.md>)

Original publisher: [Read original article](<https://webdesignerdepot.com/cloudflare-just-gave-ai-training-bots-the-middle-finger/>)

Author: Alex Harper

Published: 2026-09-16T17:18:57Z

Content type: news

Language: en

Sources: [Web Designer Depot](<https://devfeed.tech/sources/web-designer-depot.md>)

Topics: [Cloudflare](<https://devfeed.tech/topics/cloudflare.md>), [AI Bots](<https://devfeed.tech/topics/ai-bots.md>), [Crawler](<https://devfeed.tech/topics/crawler.md>), [Google Search](<https://devfeed.tech/topics/google-search.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-crawlers](<https://devfeed.tech/tags/ai-crawlers.md>), [ai-tech](<https://devfeed.tech/tags/ai-tech.md>), [ai-training](<https://devfeed.tech/tags/ai-training.md>), [anthropic](<https://devfeed.tech/tags/anthropic.md>), [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [cloudflare](<https://devfeed.tech/tags/cloudflare.md>), [content-protection](<https://devfeed.tech/tags/content-protection.md>), [future-of-the-web](<https://devfeed.tech/tags/future-of-the-web.md>), [google-extended](<https://devfeed.tech/tags/google-extended.md>), [google-search](<https://devfeed.tech/tags/google-search.md>), [googlebot](<https://devfeed.tech/tags/googlebot.md>), [openai](<https://devfeed.tech/tags/openai.md>), [publishers](<https://devfeed.tech/tags/publishers.md>), [robots-txt](<https://devfeed.tech/tags/robots-txt.md>), [search](<https://devfeed.tech/tags/search.md>), [search-engines](<https://devfeed.tech/tags/search-engines.md>), [web](<https://devfeed.tech/tags/web.md>), [web-design](<https://devfeed.tech/tags/web-design.md>), [web-development](<https://devfeed.tech/tags/web-development.md>), [web-publishing](<https://devfeed.tech/tags/web-publishing.md>), [web-scraping](<https://devfeed.tech/tags/web-scraping.md>), [website-traffic](<https://devfeed.tech/tags/website-traffic.md>)

### AI overview

Cloudflare launched a Disallow AI Training setting that lets website owners allow traditional search crawlers while blocking training-only crawlers from companies including Amazon, Anthropic, Meta, and OpenAI. The article notes that robots.txt depends on crawler compliance and that blocking Google-Extended does not remove content from Google Search features such as AI Overviews or AI Mode.

### Source excerpt

Cloudflare just gave website owners a new weapon against AI crawlers: keep the search traffic, block the AI training. After years of watching bots consume the web's content, publishers finally have an easier way to tell AI companies where to go.

## Your robots.txt Says Yes. Your Firewall Says 403.

DevFeed: [Your robots.txt Says Yes. Your Firewall Says 403.](<https://devfeed.tech/articles/your-robots-txt-says-yes-your-firewall-says-403-30874.md>)

Original publisher: [Read original article](<https://brent.leekley.me/blog/robots-vs-firewall/>)

Author: Brent Leekley

Published: 2026-06-11T00:00:00Z

Content type: article

Language: en

Sources: [brent.leekley.me blog](<https://devfeed.tech/sources/brent-leekley-me-blog.md>)

Topics: [Cloudflare](<https://devfeed.tech/topics/cloudflare.md>), [AI Bots](<https://devfeed.tech/topics/ai-bots.md>), [AI search](<https://devfeed.tech/topics/ai-search.md>), [Crawler](<https://devfeed.tech/topics/crawler.md>), [Firewall](<https://devfeed.tech/topics/firewall.md>)

Tags: [403](<https://devfeed.tech/tags/403.md>), [aeo](<https://devfeed.tech/tags/aeo.md>), [ai-bots](<https://devfeed.tech/tags/ai-bots.md>), [ai-crawl-control](<https://devfeed.tech/tags/ai-crawl-control.md>), [ai-crawlers](<https://devfeed.tech/tags/ai-crawlers.md>), [ai-search](<https://devfeed.tech/tags/ai-search.md>), [ai-visibility](<https://devfeed.tech/tags/ai-visibility.md>), [blocking](<https://devfeed.tech/tags/blocking.md>), [bot-management](<https://devfeed.tech/tags/bot-management.md>), [chatgpt-user](<https://devfeed.tech/tags/chatgpt-user.md>), [cloudflare](<https://devfeed.tech/tags/cloudflare.md>), [firewall](<https://devfeed.tech/tags/firewall.md>), [gptbot](<https://devfeed.tech/tags/gptbot.md>), [oai-searchbot](<https://devfeed.tech/tags/oai-searchbot.md>), [perplexity](<https://devfeed.tech/tags/perplexity.md>), [robots](<https://devfeed.tech/tags/robots.md>), [robots-txt](<https://devfeed.tech/tags/robots-txt.md>)

### AI overview

A field note explains how a Cloudflare AI-bot blocking setting returned 403 responses to all AI agents even though the site's robots.txt allowed AI search and blocked training crawlers. It distinguishes training crawlers, search indexers, and user-triggered fetchers, and recommends auditing enforcement at the firewall layer.

### Source excerpt

A client's robots.txt welcomed AI search and blocked training crawlers, but Cloudflare's blunt AI-bot toggle was returning 403 to every AI agent at the edge. How the block was found, the AI Crawl Control fix, and why you should audit enforcement, not intent.

## Three Things Made My Blog Agent-Ready. Five I Skipped on Purpose.

DevFeed: [Three Things Made My Blog Agent-Ready. Five I Skipped on Purpose.](<https://devfeed.tech/articles/three-things-made-my-blog-agent-ready-five-i-skipped-on-purpose-40137.md>)

Original publisher: [Read original article](<https://korbonits.com/blog/2026-04-30-three-things-made-my-blog-agent-ready/>)

Published: 2026-04-30T00:00:00Z

Content type: opinion

Language: en

Sources: [Alex Korbonits](<https://devfeed.tech/sources/alex-korbonits.md>)

Topics: [Web](<https://devfeed.tech/topics/web.md>), [Netlify](<https://devfeed.tech/topics/netlify.md>), [Cloudflare](<https://devfeed.tech/topics/cloudflare.md>), [RSS Feed](<https://devfeed.tech/topics/rss-feed.md>)

Tags: [ai-agents](<https://devfeed.tech/tags/ai-agents.md>), [blog](<https://devfeed.tech/tags/blog.md>), [cloudflare](<https://devfeed.tech/tags/cloudflare.md>), [ietf](<https://devfeed.tech/tags/ietf.md>), [markdown](<https://devfeed.tech/tags/markdown.md>), [mcp-oauth](<https://devfeed.tech/tags/mcp-oauth.md>), [netlify](<https://devfeed.tech/tags/netlify.md>), [robots-txt](<https://devfeed.tech/tags/robots-txt.md>), [rss](<https://devfeed.tech/tags/rss.md>), [web](<https://devfeed.tech/tags/web.md>)

### AI overview

The author describes improving a personal blog's score on Cloudflare's agent-readiness checker from 25 to 50 by adding content signals in robots.txt and Link headers through Netlify. The article argues that personal blogs should implement applicable agent-facing features while skipping API, OAuth, and MCP requirements that do not fit their purpose.

### Source excerpt

Cloudflare's isitagentready.com scored korbonits.com at 25 / Level 1 'Basic Web Presence.' Two hours and three small additions later it scored 50 / Level 4 'Agent-Integrated.' The other half of the points came from checks that don't apply to a personal blog -- and shipping fake compliance for them would be theatre, not value.

## Preventing Google from Indexing Staging Sites

DevFeed: [Preventing Google from Indexing Staging Sites](<https://devfeed.tech/articles/preventing-google-from-indexing-staging-sites-31289.md>)

Original publisher: [Read original article](<https://nystudio107.com/blog/prevent-google-from-indexing-staging-sites>)

Author: andrew@nystudio107.com (Andrew Welch)

Published: 2017-02-02T09:00:00Z

Content type: tutorial

Language: en

Sources: [nystudio107 | Articles on modern web development.](<https://devfeed.tech/sources/nystudio107-articles-on-modern-web-development.md>)

Topics: [Search engine optimization (SEO)](<https://devfeed.tech/topics/seo.md>), [Crawler](<https://devfeed.tech/topics/crawler.md>), [Content Management System](<https://devfeed.tech/topics/cms.md>), [Web Development](<https://devfeed.tech/topics/web-development.md>)

Tags: [config](<https://devfeed.tech/tags/config.md>), [diluting](<https://devfeed.tech/tags/diluting.md>), [environments](<https://devfeed.tech/tags/environments.md>), [google](<https://devfeed.tech/tags/google.md>), [indexing](<https://devfeed.tech/tags/indexing.md>), [insights](<https://devfeed.tech/tags/insights.md>), [multi-environment](<https://devfeed.tech/tags/multi-environment.md>), [prevent](<https://devfeed.tech/tags/prevent.md>), [robots-txt](<https://devfeed.tech/tags/robots-txt.md>), [seo](<https://devfeed.tech/tags/seo.md>), [seomatic](<https://devfeed.tech/tags/seomatic.md>), [sites](<https://devfeed.tech/tags/sites.md>), [staging](<https://devfeed.tech/tags/staging.md>), [value](<https://devfeed.tech/tags/value.md>)

### AI overview

This tutorial explains how to prevent Google and other search engines from indexing staging sites. It presents robots.txt and SEOmatic as an alternative to password protection, allowing external performance and SEO testing while avoiding duplicate-content concerns.

### Source excerpt

SEOmatic and a multi-environment config can prevent Google from indexing your staging sites, and diluting your SEO value

## Commenting powered by Discourse

DevFeed: [Commenting powered by Discourse](<https://devfeed.tech/articles/commenting-powered-by-discourse-41332.md>)

Original publisher: [Read original article](<https://samsaffron.com/archive/2013/12/04/commenting-powered-by-discourse>)

Author: Sam Saffron

Published: 2013-12-04T03:12:48Z

Content type: opinion

Language: en

Sources: [Sam Saffron](<https://devfeed.tech/sources/sam-saffron.md>)

Topics: [web applications](<https://devfeed.tech/topics/web-applications.md>), [Software](<https://devfeed.tech/topics/software.md>), [Markdown](<https://devfeed.tech/topics/markdown.md>), [notifications](<https://devfeed.tech/topics/notifications.md>), [email](<https://devfeed.tech/topics/email.md>), [Crawler](<https://devfeed.tech/topics/crawler.md>)

Tags: [blog](<https://devfeed.tech/tags/blog.md>), [email](<https://devfeed.tech/tags/email.md>), [markdown](<https://devfeed.tech/tags/markdown.md>), [notifications](<https://devfeed.tech/tags/notifications.md>), [robots-txt](<https://devfeed.tech/tags/robots-txt.md>), [spam](<https://devfeed.tech/tags/spam.md>)

### AI overview

The author explains why Discourse powers comments on the blog. Although using it adds friction because readers must log in on another site, the author values its support for thoughtful discussions, email follow-up, Markdown editing, moderation, and spam prevention. The author reports receiving no spam over the previous 55 days.

### Source excerpt

Leaving comments on this blog requires a certain amount of commitment. You have to jump to another site to log in. ac12507b13e4685b.png765x428 13.4 KB Compare this to the "least amount of friction possible". A lot of work. When I made the move to Discourse I thought of this state-of-affairs as a temporary situation. I would add the "traditional" comment box at the bottom and make it super easy to add comments, I would transparently create accounts and all that jazz. A couple of months in, I am not so sure. When I think about comments on my blog these are my priorities. Give users room to type in interesting and insightful comments. Provide great support for followup (reply by email, email notifications) Rich markdown editor with edit and preview support. Comment format must be markdown, anything else and I am risking complex conversion later on. Zero spam Comments / emails / content are unconditionally hosted on my server and under my control. Not hosted by some third party under their rules with their advertising injected and my readers tracked. Trivial for me to moderate. You may notice that, "Make it super easy for anybody on the Internet to contribute a random unfiltered opinion" is surprisingly missing from this list. What does Discourse score in my 7 point dream list? A solid 7 out of 7. The extra friction completely eradicated spam and enables me to have rich conversations with my readers. I have their emails, I can communicate with them. ###About Spam twitter.com Kelly Sommers @kellabyte Trying out disqus cuz I just got 600 more spam comments on my blog since earlier. Gotta end this madness. 12:19 AM - 1 Dec 2013 1 I eliminated the vast majority of spam on this blog a while back. I blogged about it 2 years ago. Critics claimed that this approach was doomed to fail if it ever got popular. Discourse is popular. Yet I get zero spam. By zero I mean that in the last 55 days I got no spam on this blog, nor did I have to delete any spam from this blog. http://discu