# AI crawlers

Published articles for AI crawlers.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## Cloudflare Adds Setting to Block AI Training Crawlers While Allowing Search Crawlers

DevFeed: [Cloudflare Adds Setting to Block AI Training Crawlers While Allowing Search Crawlers](<https://devfeed.tech/articles/cloudflare-just-gave-ai-training-bots-the-middle-finger-31387.md>)

Original publisher: [Read original article](<https://webdesignerdepot.com/cloudflare-just-gave-ai-training-bots-the-middle-finger/>)

Author: Alex Harper

Published: 2026-09-16T17:18:57Z

Content type: news

Language: en

Sources: [Web Designer Depot](<https://devfeed.tech/sources/web-designer-depot.md>)

Topics: [Cloudflare](<https://devfeed.tech/topics/cloudflare.md>), [AI Bots](<https://devfeed.tech/topics/ai-bots.md>), [Crawler](<https://devfeed.tech/topics/crawler.md>), [Google Search](<https://devfeed.tech/topics/google-search.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-crawlers](<https://devfeed.tech/tags/ai-crawlers.md>), [ai-tech](<https://devfeed.tech/tags/ai-tech.md>), [ai-training](<https://devfeed.tech/tags/ai-training.md>), [anthropic](<https://devfeed.tech/tags/anthropic.md>), [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [cloudflare](<https://devfeed.tech/tags/cloudflare.md>), [content-protection](<https://devfeed.tech/tags/content-protection.md>), [future-of-the-web](<https://devfeed.tech/tags/future-of-the-web.md>), [google-extended](<https://devfeed.tech/tags/google-extended.md>), [google-search](<https://devfeed.tech/tags/google-search.md>), [googlebot](<https://devfeed.tech/tags/googlebot.md>), [openai](<https://devfeed.tech/tags/openai.md>), [publishers](<https://devfeed.tech/tags/publishers.md>), [robots-txt](<https://devfeed.tech/tags/robots-txt.md>), [search](<https://devfeed.tech/tags/search.md>), [search-engines](<https://devfeed.tech/tags/search-engines.md>), [web](<https://devfeed.tech/tags/web.md>), [web-design](<https://devfeed.tech/tags/web-design.md>), [web-development](<https://devfeed.tech/tags/web-development.md>), [web-publishing](<https://devfeed.tech/tags/web-publishing.md>), [web-scraping](<https://devfeed.tech/tags/web-scraping.md>), [website-traffic](<https://devfeed.tech/tags/website-traffic.md>)

### AI overview

Cloudflare launched a Disallow AI Training setting that lets website owners allow traditional search crawlers while blocking training-only crawlers from companies including Amazon, Anthropic, Meta, and OpenAI. The article notes that robots.txt depends on crawler compliance and that blocking Google-Extended does not remove content from Google Search features such as AI Overviews or AI Mode.

### Source excerpt

Cloudflare just gave website owners a new weapon against AI crawlers: keep the search traffic, block the AI training. After years of watching bots consume the web's content, publishers finally have an easier way to tell AI companies where to go.

## Keeping AI crawlers off my Forgejo server

DevFeed: [Keeping AI crawlers off my Forgejo server](<https://devfeed.tech/articles/keeping-ai-crawlers-off-my-forgejo-server-38540.md>)

Original publisher: [Read original article](<https://msfjarvis.dev/posts/keeping-ai-crawlers-off-my-forgejo-server/>)

Author: Harsh Shandilya

Published: 2026-06-29T04:53:18Z

Content type: article

Language: en

Sources: [Posts on Harsh Shandilya](<https://devfeed.tech/sources/posts-on-harsh-shandilya.md>)

Topics: [forgejo](<https://devfeed.tech/topics/forgejo.md>), [AI Bots](<https://devfeed.tech/topics/ai-bots.md>), [Cloudflare](<https://devfeed.tech/topics/cloudflare.md>), [Security](<https://devfeed.tech/topics/security.md>), [Fail2ban](<https://devfeed.tech/topics/fail2ban.md>), [Server](<https://devfeed.tech/topics/server.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-crawlers](<https://devfeed.tech/tags/ai-crawlers.md>), [cloudflare](<https://devfeed.tech/tags/cloudflare.md>), [crawlers](<https://devfeed.tech/tags/crawlers.md>), [fail2ban](<https://devfeed.tech/tags/fail2ban.md>), [forgejo](<https://devfeed.tech/tags/forgejo.md>), [security](<https://devfeed.tech/tags/security.md>), [server](<https://devfeed.tech/tags/server.md>)

### AI overview

A developer describes mitigating AI crawler traffic against a Forgejo instance using fail2ban, Caddy configuration, and Cloudflare rules. The measures reduced some traffic but also exposed client-IP handling and availability issues, while IP bans approached Cloudflare's access-rule limit.

### Source excerpt

The short and bumbling journey to finally giving my tiny VPS some respite

## Your robots.txt Says Yes. Your Firewall Says 403.

DevFeed: [Your robots.txt Says Yes. Your Firewall Says 403.](<https://devfeed.tech/articles/your-robots-txt-says-yes-your-firewall-says-403-30874.md>)

Original publisher: [Read original article](<https://brent.leekley.me/blog/robots-vs-firewall/>)

Author: Brent Leekley

Published: 2026-06-11T00:00:00Z

Content type: article

Language: en

Sources: [brent.leekley.me blog](<https://devfeed.tech/sources/brent-leekley-me-blog.md>)

Topics: [Cloudflare](<https://devfeed.tech/topics/cloudflare.md>), [AI Bots](<https://devfeed.tech/topics/ai-bots.md>), [AI search](<https://devfeed.tech/topics/ai-search.md>), [Crawler](<https://devfeed.tech/topics/crawler.md>), [Firewall](<https://devfeed.tech/topics/firewall.md>)

Tags: [403](<https://devfeed.tech/tags/403.md>), [aeo](<https://devfeed.tech/tags/aeo.md>), [ai-bots](<https://devfeed.tech/tags/ai-bots.md>), [ai-crawl-control](<https://devfeed.tech/tags/ai-crawl-control.md>), [ai-crawlers](<https://devfeed.tech/tags/ai-crawlers.md>), [ai-search](<https://devfeed.tech/tags/ai-search.md>), [ai-visibility](<https://devfeed.tech/tags/ai-visibility.md>), [blocking](<https://devfeed.tech/tags/blocking.md>), [bot-management](<https://devfeed.tech/tags/bot-management.md>), [chatgpt-user](<https://devfeed.tech/tags/chatgpt-user.md>), [cloudflare](<https://devfeed.tech/tags/cloudflare.md>), [firewall](<https://devfeed.tech/tags/firewall.md>), [gptbot](<https://devfeed.tech/tags/gptbot.md>), [oai-searchbot](<https://devfeed.tech/tags/oai-searchbot.md>), [perplexity](<https://devfeed.tech/tags/perplexity.md>), [robots](<https://devfeed.tech/tags/robots.md>), [robots-txt](<https://devfeed.tech/tags/robots-txt.md>)

### AI overview

A field note explains how a Cloudflare AI-bot blocking setting returned 403 responses to all AI agents even though the site's robots.txt allowed AI search and blocked training crawlers. It distinguishes training crawlers, search indexers, and user-triggered fetchers, and recommends auditing enforcement at the firewall layer.

### Source excerpt

A client's robots.txt welcomed AI search and blocked training crawlers, but Cloudflare's blunt AI-bot toggle was returning 403 to every AI agent at the edge. How the block was found, the AI Crawl Control fix, and why you should audit enforcement, not intent.

## How to add llms.txt to a Hugo Blog

DevFeed: [How to add llms.txt to a Hugo Blog](<https://devfeed.tech/articles/how-to-add-llms-txt-to-a-hugo-blog-27728.md>)

Original publisher: [Read original article](<https://gagor.pro/2025/10/how-to-add-llms.txt-to-a-hugo-blog/>)

Author: Tom

Published: 2025-10-16T00:00:00Z

Content type: tutorial

Language: en

Sources: [Tomasz Gągor](<https://devfeed.tech/sources/tomasz-gagor.md>)

Topics: [Hugo](<https://devfeed.tech/topics/hugo.md>), [AI Bots](<https://devfeed.tech/topics/ai-bots.md>), [Search engine optimization (SEO)](<https://devfeed.tech/topics/seo.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-content-visibility](<https://devfeed.tech/tags/ai-content-visibility.md>), [ai-crawlers](<https://devfeed.tech/tags/ai-crawlers.md>), [blog](<https://devfeed.tech/tags/blog.md>), [generative-engine-optimization](<https://devfeed.tech/tags/generative-engine-optimization.md>), [geo](<https://devfeed.tech/tags/geo.md>), [how-to](<https://devfeed.tech/tags/how-to.md>), [hugo](<https://devfeed.tech/tags/hugo.md>), [hugo-llms-txt](<https://devfeed.tech/tags/hugo-llms-txt.md>), [hugo-seo](<https://devfeed.tech/tags/hugo-seo.md>), [llm-metadata](<https://devfeed.tech/tags/llm-metadata.md>), [markdown](<https://devfeed.tech/tags/markdown.md>), [seo](<https://devfeed.tech/tags/seo.md>)

### AI overview

A tutorial showing how to add an llms.txt Markdown file to a Hugo blog using a custom media type, output format, and template. It presents the file as a way to help AI systems understand and discover site content, while noting that its effect on organic traffic is uncertain.

### Source excerpt

Learn how to add an llms.txt file to your Hugo blog to make it more visible to AI agents and improve Generative Engine Optimization (GEO).

## What is llms.txt? Breaking down the skepticism

DevFeed: [What is llms.txt? Breaking down the skepticism](<https://devfeed.tech/articles/what-is-llms-txt-breaking-down-the-skepticism-31107.md>)

Original publisher: [Read original article](<https://www.mintlify.com/blog/what-is-llms-txt>)

Author: Tiffany Chen

Published: 2025-04-02T00:00:00Z

Content type: article

Language: en

Sources: [Mintlify Blog](<https://devfeed.tech/sources/mintlify-blog.md>)

Topics: [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Documentation](<https://devfeed.tech/topics/documentation.md>), [AI Bots](<https://devfeed.tech/topics/ai-bots.md>), [Search engine optimization (SEO)](<https://devfeed.tech/topics/seo.md>), [Web](<https://devfeed.tech/topics/web.md>)

Tags: [ai-crawlers](<https://devfeed.tech/tags/ai-crawlers.md>), [ai-trends](<https://devfeed.tech/tags/ai-trends.md>), [article](<https://devfeed.tech/tags/article.md>), [documentation](<https://devfeed.tech/tags/documentation.md>), [javascript](<https://devfeed.tech/tags/javascript.md>), [llms](<https://devfeed.tech/tags/llms.md>), [seo](<https://devfeed.tech/tags/seo.md>)

### AI overview

This article explains llms.txt, a proposed Markdown standard that provides large language models and AI crawlers with a structured summary and reading guide for website documentation. It describes early adoption, ongoing skepticism, and the distinction between improving AI context and improving search-engine traffic.

### Source excerpt

You might be seeing llms.txt pop up more lately. It's a new standard that's quickly gaining traction that makes web content easier for large language models (LLMs) to understand.