# Crawler

An automated web client that can recursively traverse links for indexing.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## Cloudflare Adds Setting to Block AI Training Crawlers While Allowing Search Crawlers

DevFeed: [Cloudflare Adds Setting to Block AI Training Crawlers While Allowing Search Crawlers](<https://devfeed.tech/articles/cloudflare-just-gave-ai-training-bots-the-middle-finger-31387.md>)

Original publisher: [Read original article](<https://webdesignerdepot.com/cloudflare-just-gave-ai-training-bots-the-middle-finger/>)

Author: Alex Harper

Published: 2026-09-16T17:18:57Z

Content type: news

Language: en

Sources: [Web Designer Depot](<https://devfeed.tech/sources/web-designer-depot.md>)

Topics: [Cloudflare](<https://devfeed.tech/topics/cloudflare.md>), [AI Bots](<https://devfeed.tech/topics/ai-bots.md>), [Crawler](<https://devfeed.tech/topics/crawler.md>), [Google Search](<https://devfeed.tech/topics/google-search.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-crawlers](<https://devfeed.tech/tags/ai-crawlers.md>), [ai-tech](<https://devfeed.tech/tags/ai-tech.md>), [ai-training](<https://devfeed.tech/tags/ai-training.md>), [anthropic](<https://devfeed.tech/tags/anthropic.md>), [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [cloudflare](<https://devfeed.tech/tags/cloudflare.md>), [content-protection](<https://devfeed.tech/tags/content-protection.md>), [future-of-the-web](<https://devfeed.tech/tags/future-of-the-web.md>), [google-extended](<https://devfeed.tech/tags/google-extended.md>), [google-search](<https://devfeed.tech/tags/google-search.md>), [googlebot](<https://devfeed.tech/tags/googlebot.md>), [openai](<https://devfeed.tech/tags/openai.md>), [publishers](<https://devfeed.tech/tags/publishers.md>), [robots-txt](<https://devfeed.tech/tags/robots-txt.md>), [search](<https://devfeed.tech/tags/search.md>), [search-engines](<https://devfeed.tech/tags/search-engines.md>), [web](<https://devfeed.tech/tags/web.md>), [web-design](<https://devfeed.tech/tags/web-design.md>), [web-development](<https://devfeed.tech/tags/web-development.md>), [web-publishing](<https://devfeed.tech/tags/web-publishing.md>), [web-scraping](<https://devfeed.tech/tags/web-scraping.md>), [website-traffic](<https://devfeed.tech/tags/website-traffic.md>)

### AI overview

Cloudflare launched a Disallow AI Training setting that lets website owners allow traditional search crawlers while blocking training-only crawlers from companies including Amazon, Anthropic, Meta, and OpenAI. The article notes that robots.txt depends on crawler compliance and that blocking Google-Extended does not remove content from Google Search features such as AI Overviews or AI Mode.

### Source excerpt

Cloudflare just gave website owners a new weapon against AI crawlers: keep the search traffic, block the AI training. After years of watching bots consume the web's content, publishers finally have an easier way to tell AI companies where to go.

## Cloudflare adds controls to allow search indexing while disallowing AI training

DevFeed: [Cloudflare adds controls to allow search indexing while disallowing AI training](<https://devfeed.tech/articles/have-it-both-ways-stay-discoverable-in-search-while-disallowing-ai-training-26580.md>)

Original publisher: [Read original article](<https://blog.cloudflare.com/accountable-mixed-use-ai-crawlers/>)

Author: Bryan Becker

Published: 2026-09-15T13:00:00Z

Content type: release

Language: en

Sources: [Cloudflare Blog](<https://devfeed.tech/sources/cloudflare-blog.md>)

Topics: [Cloudflare](<https://devfeed.tech/topics/cloudflare.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Crawler](<https://devfeed.tech/topics/crawler.md>), [Google](<https://devfeed.tech/topics/google.md>), [Microsoft](<https://devfeed.tech/topics/microsoft.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-bots](<https://devfeed.tech/tags/ai-bots.md>), [blocking](<https://devfeed.tech/tags/blocking.md>), [bot-management](<https://devfeed.tech/tags/bot-management.md>), [bots](<https://devfeed.tech/tags/bots.md>), [cloudflare](<https://devfeed.tech/tags/cloudflare.md>), [content](<https://devfeed.tech/tags/content.md>), [google](<https://devfeed.tech/tags/google.md>), [internet](<https://devfeed.tech/tags/internet.md>), [microsoft](<https://devfeed.tech/tags/microsoft.md>), [network-services](<https://devfeed.tech/tags/network-services.md>), [product-news](<https://devfeed.tech/tags/product-news.md>), [search](<https://devfeed.tech/tags/search.md>), [security](<https://devfeed.tech/tags/security.md>)

### AI overview

Cloudflare announces a Disallow AI Training setting that lets website owners remain indexed in search while refusing AI training by mixed-use crawlers. Apple, Google, and Microsoft honor or have committed to honor the setting.

### Source excerpt

Cloudflare is giving site owners a way to stay discoverable while disallowing AI training. New controls and an Accountable designation establish a shared model with Apple, Google, and Microsoft.

## GHarchive data has become unreliable for measuring GitHub activity

DevFeed: [GHarchive data has become unreliable for measuring GitHub activity](<https://devfeed.tech/articles/how-much-should-you-trust-your-oss-data-34319.md>)

Original publisher: [Read original article](<http://opensource.googleblog.com/2026/09/how-much-should-you-trust-your-oss-data.html>)

Author: KD (noreply@blogger.com)

Published: 2026-09-03T16:00:00Z

Content type: opinion

Language: en

Sources: [Google Open Source Blog](<https://devfeed.tech/sources/google-open-source-blog.md>)

Topics: [data](<https://devfeed.tech/topics/data.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [GitHub](<https://devfeed.tech/topics/github.md>), [Crawler](<https://devfeed.tech/topics/crawler.md>), [GitHub API](<https://devfeed.tech/topics/github-api.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [GraphQL](<https://devfeed.tech/topics/graphql.md>)

Tags: [collect](<https://devfeed.tech/tags/collect.md>), [data](<https://devfeed.tech/tags/data.md>), [dataset](<https://devfeed.tech/tags/dataset.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [github](<https://devfeed.tech/tags/github.md>), [google](<https://devfeed.tech/tags/google.md>), [google-open-source](<https://devfeed.tech/tags/google-open-source.md>), [graphql](<https://devfeed.tech/tags/graphql.md>), [open-data-sets](<https://devfeed.tech/tags/open-data-sets.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [oss](<https://devfeed.tech/tags/oss.md>), [pipeline](<https://devfeed.tech/tags/pipeline.md>), [pull-requests](<https://devfeed.tech/tags/pull-requests.md>), [retention](<https://devfeed.tech/tags/retention.md>), [stream](<https://devfeed.tech/tags/stream.md>), [volume](<https://devfeed.tech/tags/volume.md>)

### AI overview

This article examines the reliability of open source data, focusing on GHarchive's coverage of GitHub events. It argues that GHarchive should not be used for real-time or volume-based metrics because event retention has declined and some activity is omitted by the GitHub Event stream and API limitations.

### Source excerpt

by Sophia Vargas, Google Open Source & Andrew Nesbitt, Ecosyste.ms Every second, open source contribution quietly shapes the software we rely on, and yet our view of this open ecosystem is surprisingly opaque. Open source development is performed in public spaces -- we can see the commits, issues and comments, the APIs and endpoints are free to use -- the logs are just sitting there, so why can't we just collect all of the data? ...Said every researcher, everywhere. However in most cases of open source related data, we are only looking at part of the whole. Why am I writing this post? Because many of us (including many business decision-makers) are too comfortable with unsubstantiated data. We've gotten used to it. Our models assume that it's smelly and we adjust the logic and weights to compromise. When it comes to open source, our confidence is even lower, even though our resulting decisions can directly impact individuals whom we collectively depend on. Let's consider one of my favorite datasets: GHarchive. Started as a hobby project in 2011, this crawler has amassed more than 15 years of event data from GitHub. While this source provides a historical record of open source development on GitHub, as a real-time or comprehensive source of metrics, it's unreliable and should not be a source for volume-based metrics. In 2025, GHarchive captured 14% fewer events than in 2024, despite steady growth in platform adoption. Since 2025, we estimate that data retention in GHarchive has fallen to ~50% and in 2026 it may be as low as 20% for some event types (see figure below). Prior to 2025, you could make the general assumption that the majority of events would be represented in this pipeline. Since 2025, we must now assume we may be missing at least half of events and possibly more -- not to mention all of the additional activity that's left out of the event API (see GitHub's GraphQL API.) The crawler logic behind this dataset is simple: give me all the events from the GitHub Ev

## TIME serves bots a different website, and User-Agent is now a billing identity

DevFeed: [TIME serves bots a different website, and User-Agent is now a billing identity](<https://devfeed.tech/articles/time-serves-bots-a-different-website-and-user-agent-is-now-a-billing-identity-16068.md>)

Original publisher: [Read original article](<https://workos.com/blog/user-agent-is-now-a-billing-identity>)

Author: WorkOS

Published: 2026-08-06T01:46:28Z

Content type: opinion

Language: en

Sources: [WorkOS Blog](<https://devfeed.tech/sources/workos-blog.md>)

Topics: [AI Bots](<https://devfeed.tech/topics/ai-bots.md>), [Bot](<https://devfeed.tech/topics/bot.md>), [Crawler](<https://devfeed.tech/topics/crawler.md>), [SIEM, Security, Observability](<https://devfeed.tech/topics/siem-security-observability.md>)

Tags: [ads](<https://devfeed.tech/tags/ads.md>), [billing](<https://devfeed.tech/tags/billing.md>), [bots](<https://devfeed.tech/tags/bots.md>), [curl](<https://devfeed.tech/tags/curl.md>), [logs](<https://devfeed.tech/tags/logs.md>), [measurement](<https://devfeed.tech/tags/measurement.md>), [openai](<https://devfeed.tech/tags/openai.md>), [routing](<https://devfeed.tech/tags/routing.md>)

### AI overview

The article examines how TIME serves different content to crawlers based on User-Agent headers and records bot reads as billable ad impressions. It argues that because clients can freely change the header, crawler identity is an unreliable basis for billing and measurement.

### Source excerpt

TIME forks its content per crawler and logs each bot read as a billable ad impression. The routing key for that whole ledger is a header any client can type.

## Your robots.txt Says Yes. Your Firewall Says 403.

DevFeed: [Your robots.txt Says Yes. Your Firewall Says 403.](<https://devfeed.tech/articles/your-robots-txt-says-yes-your-firewall-says-403-30874.md>)

Original publisher: [Read original article](<https://brent.leekley.me/blog/robots-vs-firewall/>)

Author: Brent Leekley

Published: 2026-06-11T00:00:00Z

Content type: article

Language: en

Sources: [brent.leekley.me blog](<https://devfeed.tech/sources/brent-leekley-me-blog.md>)

Topics: [Cloudflare](<https://devfeed.tech/topics/cloudflare.md>), [AI Bots](<https://devfeed.tech/topics/ai-bots.md>), [AI search](<https://devfeed.tech/topics/ai-search.md>), [Crawler](<https://devfeed.tech/topics/crawler.md>), [Firewall](<https://devfeed.tech/topics/firewall.md>)

Tags: [403](<https://devfeed.tech/tags/403.md>), [aeo](<https://devfeed.tech/tags/aeo.md>), [ai-bots](<https://devfeed.tech/tags/ai-bots.md>), [ai-crawl-control](<https://devfeed.tech/tags/ai-crawl-control.md>), [ai-crawlers](<https://devfeed.tech/tags/ai-crawlers.md>), [ai-search](<https://devfeed.tech/tags/ai-search.md>), [ai-visibility](<https://devfeed.tech/tags/ai-visibility.md>), [blocking](<https://devfeed.tech/tags/blocking.md>), [bot-management](<https://devfeed.tech/tags/bot-management.md>), [chatgpt-user](<https://devfeed.tech/tags/chatgpt-user.md>), [cloudflare](<https://devfeed.tech/tags/cloudflare.md>), [firewall](<https://devfeed.tech/tags/firewall.md>), [gptbot](<https://devfeed.tech/tags/gptbot.md>), [oai-searchbot](<https://devfeed.tech/tags/oai-searchbot.md>), [perplexity](<https://devfeed.tech/tags/perplexity.md>), [robots](<https://devfeed.tech/tags/robots.md>), [robots-txt](<https://devfeed.tech/tags/robots-txt.md>)

### AI overview

A field note explains how a Cloudflare AI-bot blocking setting returned 403 responses to all AI agents even though the site's robots.txt allowed AI search and blocked training crawlers. It distinguishes training crawlers, search indexers, and user-triggered fetchers, and recommends auditing enforcement at the firewall layer.

### Source excerpt

A client's robots.txt welcomed AI search and blocked training crawlers, but Cloudflare's blunt AI-bot toggle was returning 403 to every AI agent at the edge. How the block was found, the AI Crawl Control fix, and why you should audit enforcement, not intent.

## How often do LLMs visit llms.txt?

DevFeed: [How often do LLMs visit llms.txt?](<https://devfeed.tech/articles/how-often-do-llms-visit-llms-txt-31022.md>)

Original publisher: [Read original article](<https://www.mintlify.com/blog/how-often-do-llms-visit-llms-txt>)

Author: Tiffany Chen

Published: 2025-06-27T00:00:00Z

Content type: article

Language: en

Sources: [Mintlify Blog](<https://devfeed.tech/sources/mintlify-blog.md>)

Topics: [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Markdown](<https://devfeed.tech/topics/markdown.md>), [Data analysis](<https://devfeed.tech/topics/data-analysis.md>), [Crawler](<https://devfeed.tech/topics/crawler.md>), [Retrieval-Augmented Generation](<https://devfeed.tech/topics/retrieval-augmented-generation.md>)

Tags: [ai-trends](<https://devfeed.tech/tags/ai-trends.md>), [claude](<https://devfeed.tech/tags/claude.md>), [data-analysis](<https://devfeed.tech/tags/data-analysis.md>), [google](<https://devfeed.tech/tags/google.md>), [llms](<https://devfeed.tech/tags/llms.md>), [markdown](<https://devfeed.tech/tags/markdown.md>), [pages](<https://devfeed.tech/tags/pages.md>), [retrieval](<https://devfeed.tech/tags/retrieval.md>), [retrieval-augmented-generation](<https://devfeed.tech/tags/retrieval-augmented-generation.md>)

### AI overview

The article presents a Profound data analysis of how language models access llms.txt and llms-full.txt. The supplied text says both files receive AI traffic, with a strong preference for llms-full.txt and ChatGPT accounting for most visits. It attributes this preference to models embedding full content instead of relying on retrieval-augmented generation, while noting that the analysis covers 25 companies.

### Source excerpt

Last month, we explored signals for the emerging standard of llms.txt, which is a Markdown file that makes websites easier for LLMs to index.

## SourceHut author describes the operational cost of aggressive LLM crawlers

DevFeed: [SourceHut author describes the operational cost of aggressive LLM crawlers](<https://devfeed.tech/articles/please-stop-externalizing-your-costs-directly-into-my-face-20810.md>)

Original publisher: [Read original article](<https://drewdevault.com/blog/Stop-externalizing-your-costs-on-me/>)

Author: March

Published: 2025-03-17T00:00:00Z

Content type: opinion

Language: en

Sources: [Drew DeVault](<https://devfeed.tech/sources/drew-devault.md>)

Topics: [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Crawler](<https://devfeed.tech/topics/crawler.md>), [Git](<https://devfeed.tech/topics/git.md>), [Go Language](<https://devfeed.tech/topics/go-language.md>), [ci](<https://devfeed.tech/topics/ci.md>), [Cryptocurrency](<https://devfeed.tech/topics/cryptocurrency.md>), [HTTP](<https://devfeed.tech/topics/http.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [ci](<https://devfeed.tech/tags/ci.md>), [cryptocurrency](<https://devfeed.tech/tags/cryptocurrency.md>), [git](<https://devfeed.tech/tags/git.md>), [go](<https://devfeed.tech/tags/go.md>), [http](<https://devfeed.tech/tags/http.md>), [llms](<https://devfeed.tech/tags/llms.md>), [robots](<https://devfeed.tech/tags/robots.md>)

### AI overview

A personal opinion post describes SourceHut's experience mitigating aggressive LLM crawlers that ignore robots.txt, crawl expensive Git endpoints, distribute requests across many IP addresses, and contribute to recurring outages. It also compares this burden with earlier CI cryptocurrency-mining abuse and Go module mirror traffic.

### Source excerpt

This blog post is expressing personal experiences and opinions and doesn't reflect any official policies of SourceHut. Over the past few months, instead of working on our priorities at SourceHut, I have spent anywhere from 20-100% of my time in any given week mitigating hyper-aggressive LLM crawlers at scale. This isn't the first time SourceHut has been at the wrong end of some malicious bullshit or paid someone else's externalized costs - every couple of years someone invents a new way of ruining my day. Four years ago, we decided to require payment to use our CI services because it was being abused to mine cryptocurrency. We alternated between periods of designing and deploying tools to curb this abuse and periods of near-complete outage when they adapted to our mitigations and saturated all of our compute with miners seeking a profit. It was bad enough having to beg my friends and family to avoid "investing" in the scam without having the scam break into my business and trash the place every day. Two years ago, we threatened to blacklist the Go module mirror because for some reason the Go team thinks that running terabytes of git clones all day, every day for every Go project on git.sr.ht is cheaper than maintaining any state or using webhooks or coordinating the work between instances or even just designing a module system that doesn't require Google to DoS git forges whose entire annual budgets are considerably smaller than a single Google engineer's salary. Now it's LLMs. If you think these crawlers respect robots.txt then you are several assumptions of good faith removed from reality. These bots crawl everything they can find, robots.txt be damned, including expensive endpoints like git blame, every page of every git log, and every commit in every repo, and they do so using random User-Agents that overlap with end-users and come from tens of thousands of IP addresses - mostly residential, in unrelated subnets, each one making no more than one HTTP request ove

## Exploring the intertwingularity of a docs site

DevFeed: [Exploring the intertwingularity of a docs site](<https://devfeed.tech/articles/exploring-the-intertwingularity-of-a-docs-site-31140.md>)

Original publisher: [Read original article](<https://technicalwriting.dev/2024/10/intertwingularity/index.html>)

Published: 2024-10-16T00:00:00Z

Content type: article

Language: en

Sources: [technicalwriting.dev](<https://devfeed.tech/sources/technicalwriting-dev.md>)

Topics: [Crawler](<https://devfeed.tech/topics/crawler.md>), [Documentation](<https://devfeed.tech/topics/documentation.md>), [Website](<https://devfeed.tech/topics/website.md>), [Web](<https://devfeed.tech/topics/web.md>)

Tags: [building](<https://devfeed.tech/tags/building.md>), [docs](<https://devfeed.tech/tags/docs.md>), [exploring](<https://devfeed.tech/tags/exploring.md>), [intertwingularity](<https://devfeed.tech/tags/intertwingularity.md>), [links](<https://devfeed.tech/tags/links.md>), [web](<https://devfeed.tech/tags/web.md>)

### AI overview

The article describes building a web crawler to measure how pages in a documentation site link to one another and to the wider web. The backlink data is intended to help identify important pages and prioritize technical-writing effort.

### Source excerpt

Building a web crawler to quantify interconnected knowledge.

## SEO Myths: Top 5 Sitemap Myths Demystified

DevFeed: [SEO Myths: Top 5 Sitemap Myths Demystified](<https://devfeed.tech/articles/seo-myths-top-5-sitemap-myths-demystified-31294.md>)

Original publisher: [Read original article](<https://nystudio107.com/blog/seo-myths-top-5-sitemap-myths-demystified>)

Author: andrew@nystudio107.com (Andrew Welch)

Published: 2023-10-30T20:03:00Z

Content type: article

Language: en

Sources: [nystudio107 | Articles on modern web development.](<https://devfeed.tech/sources/nystudio107-articles-on-modern-web-development.md>)

Topics: [Search engine optimization (SEO)](<https://devfeed.tech/topics/seo.md>), [Google Search](<https://devfeed.tech/topics/google-search.md>), [Crawler](<https://devfeed.tech/topics/crawler.md>)

Tags: [2005](<https://devfeed.tech/tags/2005.md>), [article](<https://devfeed.tech/tags/article.md>), [demystify](<https://devfeed.tech/tags/demystify.md>), [google-search](<https://devfeed.tech/tags/google-search.md>), [insights](<https://devfeed.tech/tags/insights.md>), [misconceptions](<https://devfeed.tech/tags/misconceptions.md>), [myths](<https://devfeed.tech/tags/myths.md>), [seo](<https://devfeed.tech/tags/seo.md>), [sitemaps](<https://devfeed.tech/tags/sitemaps.md>), [surrounding](<https://devfeed.tech/tags/surrounding.md>), [we-ll](<https://devfeed.tech/tags/we-ll.md>)

### AI overview

This article examines five common myths about sitemaps and explains that they are optional signals search engines may use, not directives. It also describes situations where a sitemap may help, such as large or new sites and sites with rich media.

### Source excerpt

Sitemaps have been around since 2005, but there are still many myths and misconceptions surrounding them. We'll demystify five SEO myths in this article.

## DocSearch migration

DevFeed: [DocSearch migration](<https://devfeed.tech/articles/docsearch-migration-40932.md>)

Original publisher: [Read original article](<https://docusaurus.io/blog/2021/11/21/algolia-docsearch-migration>)

Author: Sébastien Lorber

Published: 2021-11-21T00:00:00Z

Content type: article

Language: en

Sources: [Docusaurus Blog](<https://devfeed.tech/sources/docusaurus-blog.md>)

Topics: [Algolia](<https://devfeed.tech/topics/algolia.md>), [migration](<https://devfeed.tech/topics/migration.md>), [configuration](<https://devfeed.tech/topics/configuration.md>), [Crawler](<https://devfeed.tech/topics/crawler.md>), [account](<https://devfeed.tech/topics/account.md>)

Tags: [account](<https://devfeed.tech/tags/account.md>), [algolia](<https://devfeed.tech/tags/algolia.md>), [configuration](<https://devfeed.tech/tags/configuration.md>), [crawler](<https://devfeed.tech/tags/crawler.md>), [documentation](<https://devfeed.tech/tags/documentation.md>), [github](<https://devfeed.tech/tags/github.md>), [guide](<https://devfeed.tech/tags/guide.md>), [migration](<https://devfeed.tech/tags/migration.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [search](<https://devfeed.tech/tags/search.md>)

### AI overview

DocSearch is migrating to a new system that gives users their own Algolia application and new credentials. Docusaurus site owners must update their configuration by February 1, 2022, after which existing search indexes will become read-only.

### Source excerpt

DocSearch is migrating to a new, more powerful system, which gives users their own Algolia application and new credentials.

## Going Full Static

DevFeed: [Going Full Static](<https://devfeed.tech/articles/going-full-static-3389.md>)

Original publisher: [Read original article](<https://nuxt.com/blog/going-full-static>)

Published: 2020-06-18T00:00:00Z

Content type: release

Language: en

Sources: [The Nuxt Blog](<https://devfeed.tech/sources/the-nuxt-blog.md>)

Topics: [Nuxt.js](<https://devfeed.tech/topics/nuxt.md>), [Jamstack](<https://devfeed.tech/topics/jamstack.md>), [Server-side rendering](<https://devfeed.tech/topics/server-side-rendering.md>), [Crawler](<https://devfeed.tech/topics/crawler.md>), [HTML](<https://devfeed.tech/topics/html.md>), [Single-page application (SPA)](<https://devfeed.tech/topics/spa.md>), [Web Development](<https://devfeed.tech/topics/web-development.md>)

Tags: [hosting](<https://devfeed.tech/tags/hosting.md>), [html](<https://devfeed.tech/tags/html.md>), [jamstack](<https://devfeed.tech/tags/jamstack.md>), [navigation](<https://devfeed.tech/tags/navigation.md>), [nuxt](<https://devfeed.tech/tags/nuxt.md>), [performance](<https://devfeed.tech/tags/performance.md>), [release](<https://devfeed.tech/tags/release.md>), [server-side-rendering](<https://devfeed.tech/tags/server-side-rendering.md>), [ssr](<https://devfeed.tech/tags/ssr.md>)

### AI overview

Nuxt 2.13 introduces full static export for JAMstack applications. The static target pre-renders pages to HTML, extracts payloads for client-side navigation without API calls, improves smart prefetching, includes a crawler, supports SPA fallback behavior, and improves the development experience and deployment speed.

### Source excerpt

Long awaited features for JAMstack fans has been shipped in v2.13: full static export, improved smart prefetching, integrated crawler, faster re-deploy, built-in web server and new target option for config ⚡

## Building a legacy search engine for a legacy protocol

DevFeed: [Building a legacy search engine for a legacy protocol](<https://devfeed.tech/articles/building-a-legacy-search-engine-for-a-legacy-protocol-41592.md>)

Original publisher: [Read original article](<https://blog.benjojo.co.uk/post/building-a-search-engine-for-gopher>)

Author: ben@benjojo.co.uk

Published: 2017-05-21T12:26:48Z

Content type: article

Language: en

Sources: [benjojo blog](<https://devfeed.tech/sources/benjojo-blog.md>)

Topics: [Protocol (disambiguation)](<https://devfeed.tech/topics/protocol.md>), [Crawler](<https://devfeed.tech/topics/crawler.md>), [search-engines](<https://devfeed.tech/topics/search-engines.md>), [legacy](<https://devfeed.tech/topics/legacy.md>), [selectors](<https://devfeed.tech/topics/selectors.md>), [Windows](<https://devfeed.tech/topics/windows.md>), [Bash](<https://devfeed.tech/topics/bash.md>), [win32](<https://devfeed.tech/topics/win32.md>), [Support](<https://devfeed.tech/topics/support.md>)

Tags: [bash](<https://devfeed.tech/tags/bash.md>), [crawler](<https://devfeed.tech/tags/crawler.md>), [css](<https://devfeed.tech/tags/css.md>), [http](<https://devfeed.tech/tags/http.md>), [implementation](<https://devfeed.tech/tags/implementation.md>), [legacy](<https://devfeed.tech/tags/legacy.md>), [protocol](<https://devfeed.tech/tags/protocol.md>), [search](<https://devfeed.tech/tags/search.md>), [search-engine](<https://devfeed.tech/tags/search-engine.md>), [server](<https://devfeed.tech/tags/server.md>), [servers](<https://devfeed.tech/tags/servers.md>), [software](<https://devfeed.tech/tags/software.md>), [tcp](<https://devfeed.tech/tags/tcp.md>)

### AI overview

The article describes building a search engine for the Gopher protocol. It covers crawling Gopher servers, indexing menus and text files, and running an old AltaVista search engine on Windows 98 with a stunnel relay.

### Source excerpt

Building a legacy search engine for a legacy protocol Translations are availa

## Preventing Google from Indexing Staging Sites

DevFeed: [Preventing Google from Indexing Staging Sites](<https://devfeed.tech/articles/preventing-google-from-indexing-staging-sites-31289.md>)

Original publisher: [Read original article](<https://nystudio107.com/blog/prevent-google-from-indexing-staging-sites>)

Author: andrew@nystudio107.com (Andrew Welch)

Published: 2017-02-02T09:00:00Z

Content type: tutorial

Language: en

Sources: [nystudio107 | Articles on modern web development.](<https://devfeed.tech/sources/nystudio107-articles-on-modern-web-development.md>)

Topics: [Search engine optimization (SEO)](<https://devfeed.tech/topics/seo.md>), [Crawler](<https://devfeed.tech/topics/crawler.md>), [Content Management System](<https://devfeed.tech/topics/cms.md>), [Web Development](<https://devfeed.tech/topics/web-development.md>)

Tags: [config](<https://devfeed.tech/tags/config.md>), [diluting](<https://devfeed.tech/tags/diluting.md>), [environments](<https://devfeed.tech/tags/environments.md>), [google](<https://devfeed.tech/tags/google.md>), [indexing](<https://devfeed.tech/tags/indexing.md>), [insights](<https://devfeed.tech/tags/insights.md>), [multi-environment](<https://devfeed.tech/tags/multi-environment.md>), [prevent](<https://devfeed.tech/tags/prevent.md>), [robots-txt](<https://devfeed.tech/tags/robots-txt.md>), [seo](<https://devfeed.tech/tags/seo.md>), [seomatic](<https://devfeed.tech/tags/seomatic.md>), [sites](<https://devfeed.tech/tags/sites.md>), [staging](<https://devfeed.tech/tags/staging.md>), [value](<https://devfeed.tech/tags/value.md>)

### AI overview

This tutorial explains how to prevent Google and other search engines from indexing staging sites. It presents robots.txt and SEOmatic as an alternative to password protection, allowing external performance and SEO testing while avoiding duplicate-content concerns.

### Source excerpt

SEOmatic and a multi-environment config can prevent Google from indexing your staging sites, and diluting your SEO value

## Mention's CTO journey to 400,000 users

DevFeed: [Mention's CTO journey to 400,000 users](<https://devfeed.tech/articles/mention-s-cto-journey-to-400-000-users-34701.md>)

Original publisher: [Read original article](<https://medium.com/unexpected-token/mention-s-cto-journey-to-400-000-users-f22dc242eec?source=rss----2d2624499d2---4>)

Author: Hexa

Published: 2015-10-22T12:17:15Z

Content type: article

Language: en

Sources: [eFounders](<https://devfeed.tech/sources/efounders.md>)

Topics: [App](<https://devfeed.tech/topics/app.md>), [Monitoring](<https://devfeed.tech/topics/monitoring.md>), [Crawler](<https://devfeed.tech/topics/crawler.md>), [Concurrency](<https://devfeed.tech/topics/concurrency.md>), [race-condition](<https://devfeed.tech/topics/race-condition.md>), [Go Language](<https://devfeed.tech/topics/go-language.md>), [MySQL](<https://devfeed.tech/topics/mysql.md>), [API](<https://devfeed.tech/topics/api.md>), [Kafka](<https://devfeed.tech/topics/kafka.md>), [Redis](<https://devfeed.tech/topics/redis.md>), [Percona](<https://devfeed.tech/topics/percona.md>), [backups](<https://devfeed.tech/topics/backups.md>)

Tags: [api](<https://devfeed.tech/tags/api.md>), [backups](<https://devfeed.tech/tags/backups.md>), [developer](<https://devfeed.tech/tags/developer.md>), [development](<https://devfeed.tech/tags/development.md>), [dns](<https://devfeed.tech/tags/dns.md>), [go](<https://devfeed.tech/tags/go.md>), [kafka](<https://devfeed.tech/tags/kafka.md>), [monitoring](<https://devfeed.tech/tags/monitoring.md>), [mysql](<https://devfeed.tech/tags/mysql.md>), [percona](<https://devfeed.tech/tags/percona.md>), [programming](<https://devfeed.tech/tags/programming.md>), [race-condition](<https://devfeed.tech/tags/race-condition.md>), [redis](<https://devfeed.tech/tags/redis.md>), [startup](<https://devfeed.tech/tags/startup.md>), [tech](<https://devfeed.tech/tags/tech.md>), [users](<https://devfeed.tech/tags/users.md>)

### AI overview

Mention's co-founder and CTO describes the technical challenges involved in growing the real-time monitoring application to 400,000 users. The article covers high-concurrency web crawling, a libc DNS resolver race condition, and the use of Go as a workaround. It also explains Mention's storage architecture using Percona MySQL, Redis for caching, Kafka for messaging, and incremental backups.

### Source excerpt

Mention is a real-time monitoring application used to track and analyze trends and e-reputation in a very complete and intuitive way. Co-founded in 2010, it now counts 400,000 users across the world. This impressive growth rate not only implies great marketing talent but also impressive technical achievements. Arnaud le Blanc, Mention's co-founder and CTO, tells us about how he got Mention there. You gotta love technical challenges "When we developed Mention, our main competitor was Google Alerts, which feels a little like being David versus Goliath at first. But then it became one of the reasons I like my job: it is very challenging. Getting to this point of a product development involves a true entrepreneurial mindset. Before being Mention's CTO, I was a developer and then the lead developer working on the media monitoring feature at Pressking. eFounders offered me to become Mention's co-founder. I might not have thought of myself as an entrepreneur before, although today, I really feel like Mention is my product: I've built it together with my co-founders. Building an application like Mention is full of technical challenges. We handle a large amount of data, crawl thousands of web page per second, built an open API for our own apps, and use lots of third-party APIs." FOCUS: a challenge overcome when developing Mention's web crawler Since Mention crawls a large amount of web pages in parallel, there is a point when the crawler reaches a high concurrency level and stresses the OS: that is when we got to the limit of our system. In our case, these limits entailed the appearance of a race condition in the libc's DNS resolver. libc was sending over DNS queries on random unrelated file descriptors, which was leading to weird behaviors. The bug was reported and fixed a few months later, but in the meantime, we had to use the plain-Go DNS resolve, which happened to work really well. "At Mention, several millions of Mentions are inserted in user feeds per day (see infogra

## robots.txt usage over the Alexa million

DevFeed: [robots.txt usage over the Alexa million](<https://devfeed.tech/articles/robots-txt-usage-over-the-alexa-million-41629.md>)

Original publisher: [Read original article](<https://blog.benjojo.co.uk/post/robots-txt-over-1-million-sites>)

Author: ben@benjojo.co.uk

Published: 2015-09-05T16:35:21Z

Content type: article

Language: en

Sources: [benjojo blog](<https://devfeed.tech/sources/benjojo-blog.md>)

Topics: [Crawler](<https://devfeed.tech/topics/crawler.md>), [Bot](<https://devfeed.tech/topics/bot.md>)

Tags: [bots](<https://devfeed.tech/tags/bots.md>), [robots-txt](<https://devfeed.tech/tags/robots-txt.md>)

### AI overview

The article examines how website operators use robots.txt to guide bots and crawlers, including search engines, price indexers, social-media embed bots, SEO services, and archives. It discusses the protocol's voluntary nature and its role in limiting crawling, resource use, indexing, and exposure of sensitive site paths.

### Source excerpt

robots.txt usage over the Alexa million If you ever had to deal with bots while running a site you will have at least at some point looked into robots.txt, a system that isn't rea

## Jak dokuczać spamerom

DevFeed: [Jak dokuczać spamerom](<https://devfeed.tech/articles/jak-dokuczac-spamerom-27557.md>)

Original publisher: [Read original article](<https://gagor.pro/2012/12/jak-dokuczac-spamerom/>)

Author: Tom

Published: 2012-12-18T00:00:00Z

Content type: tutorial

Language: pl

Sources: [Tomasz Gągor](<https://devfeed.tech/sources/tomasz-gagor.md>)

Topics: [WordPress](<https://devfeed.tech/topics/wordpress.md>), [Crawler](<https://devfeed.tech/topics/crawler.md>), [Template](<https://devfeed.tech/topics/template.md>), [DDoS](<https://devfeed.tech/topics/ddos.md>)

Tags: [ddos](<https://devfeed.tech/tags/ddos.md>), [php](<https://devfeed.tech/tags/php.md>), [robots](<https://devfeed.tech/tags/robots.md>), [spam](<https://devfeed.tech/tags/spam.md>), [wordpress](<https://devfeed.tech/tags/wordpress.md>)

### AI overview

The article describes a technique for generating large numbers of fake email addresses on a WordPress page to trap spam crawlers. It recommends excluding the path in robots.txt and discusses precautions to avoid causing mail-related denial-of-service problems for one's own or third-party domains.

### Source excerpt

Dawno, dawno temu... Za górami, za lasami... czytałem sobie tekst Lemat'a o dokuczaniu spamerom i pomyślałem że sam też tak mogę i nawet chcę więc popełniłem skrypcik, który dla losowych słów generował maile. Skrypcik działał z dwa lata na mojej poprzedniej stronie i nie raz zdarzyło się tam jakiejś mendzie zapętlić. Jakoś nie miałem czasu od razu, a później zapomniałem wrzucić go na nową stronie i tak zostało - na pewien czas.