# crawlers

Published articles for crawlers.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## Keeping AI crawlers off my Forgejo server

DevFeed: [Keeping AI crawlers off my Forgejo server](<https://devfeed.tech/articles/keeping-ai-crawlers-off-my-forgejo-server-38540.md>)

Original publisher: [Read original article](<https://msfjarvis.dev/posts/keeping-ai-crawlers-off-my-forgejo-server/>)

Author: Harsh Shandilya

Published: 2026-06-29T04:53:18Z

Content type: article

Language: en

Sources: [Posts on Harsh Shandilya](<https://devfeed.tech/sources/posts-on-harsh-shandilya.md>)

Topics: [forgejo](<https://devfeed.tech/topics/forgejo.md>), [AI Bots](<https://devfeed.tech/topics/ai-bots.md>), [Cloudflare](<https://devfeed.tech/topics/cloudflare.md>), [Security](<https://devfeed.tech/topics/security.md>), [Fail2ban](<https://devfeed.tech/topics/fail2ban.md>), [Server](<https://devfeed.tech/topics/server.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-crawlers](<https://devfeed.tech/tags/ai-crawlers.md>), [cloudflare](<https://devfeed.tech/tags/cloudflare.md>), [crawlers](<https://devfeed.tech/tags/crawlers.md>), [fail2ban](<https://devfeed.tech/tags/fail2ban.md>), [forgejo](<https://devfeed.tech/tags/forgejo.md>), [security](<https://devfeed.tech/tags/security.md>), [server](<https://devfeed.tech/tags/server.md>)

### AI overview

A developer describes mitigating AI crawler traffic against a Forgejo instance using fail2ban, Caddy configuration, and Cloudflare rules. The measures reduced some traffic but also exposed client-IP handling and availability issues, while IP bans approached Cloudflare's access-rule limit.

### Source excerpt

The short and bumbling journey to finally giving my tiny VPS some respite

## RFC 9969: IAB AI-CONTROL Workshop Report

DevFeed: [RFC 9969: IAB AI-CONTROL Workshop Report](<https://devfeed.tech/articles/rfc-9969-iab-ai-control-workshop-report-41814.md>)

Original publisher: [Read original article](<https://www.bortzmeyer.org/9969.html>)

Published: 2026-05-20T00:00:00Z

Content type: article

Language: fr

Sources: [Blog de Stéphane Bortzmeyer](<https://devfeed.tech/sources/blog-de-stephane-bortzmeyer.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Internet Engineering Task Force (IETF)](<https://devfeed.tech/topics/ietf.md>), [Web](<https://devfeed.tech/topics/web.md>), [Bot](<https://devfeed.tech/topics/bot.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [blog](<https://devfeed.tech/tags/blog.md>), [bots](<https://devfeed.tech/tags/bots.md>), [crawlers](<https://devfeed.tech/tags/crawlers.md>), [ietf](<https://devfeed.tech/tags/ietf.md>), [llm](<https://devfeed.tech/tags/llm.md>), [perplexity](<https://devfeed.tech/tags/perplexity.md>), [report](<https://devfeed.tech/tags/report.md>), [web](<https://devfeed.tech/tags/web.md>)

### AI overview

This French article reports on RFC 9969, an IAB workshop report about controlling the use of Web content for large-language-model training. It discusses data legitimacy, server load, crawler-based collection, and possible controls for webmasters, while distinguishing the RFC's account from the author's opinions.

### Source excerpt

Ah, l'IA... Vaste sujet, et d'actualité. Une des questions qui reviennent souvent est celle de l'utilisation du contenu qu'on trouve sur le Web pour entrainer les grands modèles, sa légitimité, la charge qu'elle induit pour les serveurs, les moyens de la contrôler, etc. Un colloque avait été organisé (https://www.ietf.org/blog/impressions-ai-control-workshop/) par l'IAB en septembre 2024 sur ces questions et ce RFC en est le compte-rendu. Ce colloque avait lancé le projet IETF aipref (https://datatracker.ietf.org/wg/aipref/).