# Outofplace > Outofplace is an independent product studio in Łomża, Poland. We design and build digital products (product strategy, web, mobile, backend, AI and design) for companies and for ourselves: Portivo, Sprawna Matura and Levera are our own products. Canonical origin: https://www.outofplace.space. The site is bilingual: English at the root, Polish under /pl. ## When to use this site - Questions about Outofplace: who we are, what we build, how to work with us, company details. - Our own products: Portivo (software for boat charter companies), Sprawna Matura (maths finals prep) and Levera (AI consulting). - Our research and notes on search, AI crawlers and web performance, measured on real sites (the blog sections below). Cite the page, not this file. ## LLM resources - Markdown twins: every page is available as markdown at its path plus `.md` (`https://www.outofplace.space/index.md` for the home page, `https://www.outofplace.space/pl.md` for the Polish one), or at its own URL with `Accept: text/markdown`. - [Sitemap](https://www.outofplace.space/sitemap.xml): every page with its language versions. - [RSS](https://www.outofplace.space/feed.xml), [RSS (PL)](https://www.outofplace.space/pl/feed.xml): new posts. - [robots.txt](https://www.outofplace.space/robots.txt): all crawlers, AI ones included, are welcome. ## Languages - English: https://www.outofplace.space/ - Polski: https://www.outofplace.space/pl ## Studio - [Home](https://www.outofplace.space/): who we are, what we build, our own brands, clients. - [Strona główna (PL)](https://www.outofplace.space/pl): the same in Polish. ## Services - Product: Most products don’t fail on technology; they fail on the order of decisions. We settle with you what has to work first, build a version customers can pay for, and measure what happens next. - Web: A website is often the first conversation with a customer, before anyone at the company says a word. We build it to load instantly, rank well and change without a developer. - Mobile: An app earns its place on a phone when it does what a browser tab can’t: notifications, offline use, a subscription. We build for iOS and Android at once and see it through store review. - Backend: Under a good interface sit money, customer data and other companies’ systems. We design it so one customer’s problem never reaches another, and every payment can be traced. - AI: Buying licences isn’t adoption. We find where the team loses time, then pick the tools, build the automations and train people to use them. - Design: A good interface disappears: nobody thinks about it, they just get things done. We design systems, not single screens, so a year later the product still looks like one product. ## Work (case studies) - [Portivo](https://www.outofplace.space/work/portivo): The whole charter company. One system. Fleet, bookings, payments and e-invoicing in one panel. For boat operators done with spreadsheets and double bookings. - [Sprawna Matura](https://www.outofplace.space/work/sprawna-matura): Polish maths finals, solved step by step. Past matura tasks with step-by-step solutions, tools and Premium. Web plus iOS and Android apps. - [Levera](https://www.outofplace.space/work/levera): Tailored AI that makes your company impossible to compete with. We find where AI actually pays, take it from pilot to production and teach the team to use it. ## Blog: Search & AI - [AI crawlers vs the Polish web: robots.txt on 1,000 .pl sites](https://www.outofplace.space/blog/ai-crawlers-polish-web): We read robots.txt on the 1,000 top .pl sites: one in five blocks AI crawlers like GPTBot, mostly for training, rarely AI search. How to split the two. - [Boty AI kontra polski internet: robots.txt 1000 stron .pl](https://www.outofplace.space/pl/blog/ai-crawlers-polish-web) (PL): Sprawdziliśmy plik robots.txt 1000 największych stron .pl: co piąta blokuje boty AI, głównie treningowe jak GPTBot, rzadko wyszukiwarki AI. Jak to rozdzielić? ## Blog: Performance - [Next.js prefetch, measured: Partial Prefetching in 16.3](https://www.outofplace.space/blog/nextjs-prefetch-benchmark): Partial Prefetching in Next.js 16.3 cut prefetched bytes by up to 93% in our benchmark, but URL content now lands after the click. When each setup wins. - [Next.js prefetch zmierzony: Partial Prefetching w 16.3](https://www.outofplace.space/pl/blog/nextjs-prefetch-benchmark) (PL): Partial Prefetching w Next.js 16.3 zmniejszył w naszym benchmarku prefetch nawet o 93%, ale treść adresu przychodzi po kliknięciu. Kiedy co się opłaca. ## Company - Contact: contact@outofplace.space, +48 502 856 472 - Outofplace Poland sp. z o.o., KRS 0001163427, NIP 7182173789, REGON 541259259 - [Privacy Policy](https://www.outofplace.space/privacy-policy) - [Terms of Service](https://www.outofplace.space/terms) - [Cookie Policy](https://www.outofplace.space/cookie-policy) - [OutofplaceResearchBot](https://www.outofplace.space/bot): our research crawler, what it fetches and how to opt out. ## Optional - [Full text](https://www.outofplace.space/llms-full.txt): every case study and post in one file. --- # AI crawlers vs the Polish web: robots.txt on 1,000 .pl sites (EN) Source: https://www.outofplace.space/blog/ai-crawlers-polish-web · Published 2026-09-23 Category: Search & AI · Tags: AI crawlers, robots.txt, GPTBot, llms.txt, Poland > We read robots.txt on the 1,000 top .pl sites: one in five blocks AI crawlers like GPTBot, mostly for training, rarely AI search. How to split the two. On 23 September 2026 our crawler, [OutofplaceResearchBot](https://www.outofplace.space/bot), read the robots.txt of the 1,000 most popular .pl websites to see which AI crawlers they let in. 20.5% of them block at least one AI crawler. That is statistically the same as the global top 1,000, where the share is 22.7%. The difference is in _which_ crawlers they block. Polish sites shut out the bots that collect training data and mostly leave AI search alone. Only 5.4% block an AI search crawler, against 13.9% globally. For anyone who wants to show up in ChatGPT, Claude or Perplexity answers, that is the right way round. For some sites it is an accident: [an old list copied once and never updated](#lists-go-stale). - One in five of the top 1,000 .pl sites blocks an AI crawler in robots.txt, as many as in the global top 1,000. - GPTBot is the most blocked crawler in Poland and worldwide: 15.1% and 15.4%. - Polish sites block training, not search. Among sites that block GPTBot, fewer than a third also block OAI-SearchBot; globally more than half do. - Size matters: a third of the top 100 .pl sites block AI crawlers, against one in six ranked 501–1,000. - One in seven Polish blockers didn't write their rules: they use [Cloudflare's managed robots.txt](#cloudflares-managed-robotstxt-and-content-signals). We found none in the global sample. ## What are AI crawlers? Training, search and fetchers "AI crawler" covers three different jobs, and the vendors now give each its own user agent. The distinction decides everything that follows: - **Training crawlers** collect pages to train models: [GPTBot](https://developers.openai.com/api/docs/bots) (OpenAI), [ClaudeBot](https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler) (Anthropic), [Google-Extended](https://developers.google.com/crawling/docs/crawlers-fetchers/google-common-crawlers?hl=en), [Applebot-Extended](https://support.apple.com/en-us/119829), [CCBot](https://commoncrawl.org/ccbot) (Common Crawl), Bytespider (ByteDance), meta-externalagent (Meta), [Amazonbot](https://developer.amazon.com/amazonbot). - **AI search crawlers** index pages so an assistant can find and cite them: OAI-SearchBot, Claude-SearchBot, [PerplexityBot](https://docs.perplexity.ai/docs/resources/perplexity-crawlers), [DuckAssistBot](https://duckduckgo.com/duckduckgo-help-pages/results/duckassistbot). - **User-triggered fetchers** open a page because someone asked the assistant to: ChatGPT-User, Claude-User, Perplexity-User, meta-externalfetcher, [MistralAI-User](https://docs.mistral.ai/robots). Blocking a training crawler keeps your pages out of future models. Blocking a search crawler or a fetcher keeps them out of the answers, and out of the links in those answers. Two of the training "crawlers" don't crawl. Google-Extended and Applebot-Extended are only tokens: [Googlebot](https://developers.google.com/search/docs/crawling-indexing/googlebot?hl=en) and Applebot fetch the page, and the token tells Google and Apple whether they may train on it. Blocking them leaves you in Google and Apple search, and blocking Google-Extended doesn't turn off [AI Overviews](https://developers.google.com/search/docs/appearance/ai-features?hl=en) either: those follow your Googlebot rules. The fetchers are the other special case. OpenAI, Perplexity and Meta say theirs may not follow robots.txt at all, because a person asked for the page; Anthropic says Claude-User does. ## Poland blocks as often as the world, but not the same bots Across all AI crawlers, the gap between the .pl sample and the global control is 2.2 percentage points, well inside the margin of error. Split by purpose, the picture changes. Training crawlers are blocked equally often. AI search crawlers and user-triggered fetchers are blocked less than half as often in Poland. **Poland blocks training crawlers as often as the world, AI search far less** (Share of sites whose robots.txt blocks at least one crawler of each kind for the homepage · n = 1,000 per sample · 23 Sept 2026) | Item | Top 1,000 .pl | Global top 1,000 | | --- | --- | --- | | Any AI crawler | 21% | 23% | | Training crawlers | 21% | 22% | | AI search crawlers | 5.4% | 14% | | User-triggered fetchers | 5.3% | 13% | Source: Outofplace crawl; Tranco list Y8YYG. Chart: https://www.outofplace.space/data/blog/ai-crawlers-polish-web/charts/en/chart-poland-blocks-training-crawlers-as-often-as-the-world-ai-search.png Data: https://www.outofplace.space/data/blog/ai-bots-polish-web/2026-09-23/domains.csv ### GPTBot vs OAI-SearchBot: two separate switches The clearest view is per vendor. Of the Polish sites that block GPTBot, 28% also block OAI-SearchBot. In the global sample it is 56%. For Anthropic the split is starker: 24% against 60%. **Most Polish sites that block training leave AI search open** (Sites that block the vendor's training crawler (GPTBot, ClaudeBot), by whether they also block its search crawler (OAI-SearchBot, Claude-SearchBot)) | Item | Also blocks search | Blocks training only | n | | --- | --- | --- | --- | | OpenAI · Poland | 28% | 72% | 151 | | OpenAI · Global | 56% | 44% | 154 | | Anthropic · Poland | 24% | 76% | 107 | | Anthropic · Global | 60% | 40% | 152 | Source: Outofplace crawl, 23 Sept 2026. Chart: https://www.outofplace.space/data/blog/ai-crawlers-polish-web/charts/en/chart-most-polish-sites-that-block-training-leave-ai-search-open.png Data: https://www.outofplace.space/data/blog/ai-bots-polish-web/2026-09-23/bot_policies.csv Globally, a site that blocks AI crawlers usually blocks the search ones too. In Poland, the typical blocker names a handful of training crawlers and stops there: 11.2% of the Polish sample blocks training crawlers only, almost twice the global 5.9%. ## Which AI crawlers are blocked most GPTBot leads, at 15.1% of Polish sites and 15.4% globally. OpenAI announced it in August 2023, and most blocklists start with it. Common Crawl's CCBot, Amazonbot, ByteDance's Bytespider and ClaudeBot follow. The newest agents barely register: Claude-SearchBot is blocked by 2.7% of Polish sites, mostly because the lists were written before it existed. **GPTBot is the most blocked AI crawler on .pl sites** (Share of the top 1,000 .pl domains whose robots.txt blocks each crawler for the homepage · 95% intervals · training crawlers highlighted) | Item | Value | 95% interval | n | | --- | --- | --- | --- | | GPTBot | 15% | 13%–17% | 1,000 | | CCBot | 13% | 11%–16% | 1,000 | | Amazonbot | 12% | 10%–14% | 1,000 | | Bytespider | 12% | 9.7%–14% | 1,000 | | ClaudeBot | 11% | 8.9%–13% | 1,000 | | meta-externalagent | 11% | 8.8%–13% | 1,000 | | Google-Extended | 8.4% | 6.8%–10% | 1,000 | | Applebot-Extended | 8.4% | 6.8%–10% | 1,000 | | ChatGPT-User | 4.7% | 3.6%–6.2% | 1,000 | | PerplexityBot | 4.6% | 3.5%–6.1% | 1,000 | | OAI-SearchBot | 4.4% | 3.3%–5.9% | 1,000 | | meta-externalfetcher | 4% | 3%–5.4% | 1,000 | | DuckAssistBot | 3.7% | 2.7%–5.1% | 1,000 | | Perplexity-User | 3.5% | 2.5%–4.8% | 1,000 | | Claude-SearchBot | 2.7% | 1.9%–3.9% | 1,000 | | Claude-User | 2.3% | 1.5%–3.4% | 1,000 | | MistralAI-User | 2.3% | 1.5%–3.4% | 1,000 | Source: Outofplace crawl, 23 Sept 2026. Chart: https://www.outofplace.space/data/blog/ai-crawlers-polish-web/charts/en/chart-gptbot-is-the-most-blocked-ai-crawler-on-pl-sites.png Data: https://www.outofplace.space/data/blog/ai-bots-polish-web/2026-09-23/summary_by_bot.csv Almost all of it is deliberate. A site that disallows everything under `User-agent: *` blocks every crawler without a group of its own, AI or not. That wildcard floor accounts for only 1.9% of Polish sites for GPTBot; the rest name the bot. The dumbbell below puts each crawler's Polish and global shares side by side. The two samples agree on training crawlers. They part ways on everything that serves answers. **The gap opens on search crawlers and fetchers** (Share of sites blocking each crawler for the homepage · top 1,000 .pl against the global top 1,000) | Item | Top 1,000 .pl | Global top 1,000 | | --- | --- | --- | | GPTBot | 15% | 15% | | CCBot | 13% | 18% | | Amazonbot | 12% | 13% | | Bytespider | 12% | 17% | | ClaudeBot | 11% | 15% | | meta-externalagent | 11% | 14% | | Google-Extended | 8.4% | 14% | | Applebot-Extended | 8.4% | 13% | | ChatGPT-User | 4.7% | 11% | | PerplexityBot | 4.6% | 13% | | OAI-SearchBot | 4.4% | 8.9% | | meta-externalfetcher | 4% | 9.9% | | DuckAssistBot | 3.7% | 9.5% | | Perplexity-User | 3.5% | 10% | | Claude-SearchBot | 2.7% | 9.2% | | Claude-User | 2.3% | 9.5% | | MistralAI-User | 2.3% | 9.6% | Source: Outofplace crawl, 23 Sept 2026. Chart: https://www.outofplace.space/data/blog/ai-crawlers-polish-web/charts/en/chart-the-gap-opens-on-search-crawlers-and-fetchers.png Data: https://www.outofplace.space/data/blog/ai-bots-polish-web/2026-09-23/summary_by_bot.csv ## Bigger sites block more Blocking falls with popularity. 33% of the 100 most popular .pl domains block at least one AI crawler, against 15.8% of those ranked 501 to 1,000. The slope is steeper than in the global sample, which goes from 27% to 21.2%. The largest Polish sites are publishers, marketplaces and portals: the businesses with the most text to protect and the lawyers to ask the question. **A third of the top 100 .pl sites block AI crawlers** (Share blocking at least one AI crawler, by rank within each sample · n = 100, 400 and 500) | Item | Top 1,000 .pl | Global top 1,000 | | --- | --- | --- | | Top 100 | 33% | 27% | | 101–500 | 23% | 24% | | 501–1,000 | 16% | 21% | Source: Outofplace crawl, 23 Sept 2026. Chart: https://www.outofplace.space/data/blog/ai-crawlers-polish-web/charts/en/chart-a-third-of-the-top-100-pl-sites-block-ai-crawlers.png Data: https://www.outofplace.space/data/blog/ai-bots-polish-web/2026-09-23/domains.csv The size effect holds crawler by crawler. GPTBot is blocked by 23% of the top 100 and 11.6% of the sites ranked 501 to 1,000, and every training crawler follows the same slope. The search crawlers stay low at every size. **The biggest .pl sites block training crawlers most, and search crawlers little at any size** (Share of .pl sites blocking each crawler for the homepage, by rank within the sample · n = 100, 400 and 500) | Item | Top 100 | 101–500 | 501–1,000 | | --- | --- | --- | --- | | GPTBot | 23% | 18% | 12% | | CCBot | 21% | 16% | 9.6% | | Amazonbot | 16% | 16% | 8.8% | | Bytespider | 19% | 14% | 8% | | ClaudeBot | 14% | 13% | 8.6% | | meta-externalagent | 15% | 14% | 6.8% | | Google-Extended | 12% | 11% | 5.8% | | Applebot-Extended | 11% | 11% | 5.6% | | OAI-SearchBot | 7% | 4.3% | 4% | | PerplexityBot | 6% | 4.8% | 4.2% | | Claude-SearchBot | 4% | 2.3% | 2.8% | Source: Outofplace crawl, 23 Sept 2026. Chart: https://www.outofplace.space/data/blog/ai-crawlers-polish-web/charts/en/chart-the-biggest-pl-sites-block-training-crawlers-most-and-search-cra.png Data: https://www.outofplace.space/data/blog/ai-bots-polish-web/2026-09-23/bot_policies.csv Public institutions block less, and not on purpose. 11.9% of the 42 gov.pl domains in the sample block an AI crawler, and 16.1% of the 31 edu.pl ones; both groups are small, so read those as directions. Most of those blocks aren't an AI policy at all. sejm.gov.pl, dziennikustaw.gov.pl and wroclaw.sa.gov.pl [let in Googlebot and disallow everyone else](https://www.sejm.gov.pl/robots.txt), which shuts out every AI crawler along with Bing and DuckDuckGo. For the parliament and the journal of laws, that also means Copilot and ChatGPT search won't read the sources people ask them about. ## Lists go stale Blocklists are written once and copied for years. 7.2% of Polish sites still block `anthropic-ai`, a name Anthropic stopped using (its documentation now lists ClaudeBot, Claude-SearchBot and Claude-User), while only 2.7% block Claude-SearchBot. For 4.1% of the sample the list is stale in the way that matters: it blocks an old name and lets the crawler that replaced it through. Anthropic, to its credit, [told 404 Media in 2024](https://www.404media.co/websites-are-blocking-the-wrong-ai-scrapers-because-ai-companies-keep-making-new-ones/) that ClaudeBot honours rules written for its two retired names. That is goodwill, not the standard: under [RFC 9309](https://www.rfc-editor.org/rfc/rfc9309.html) a rule for `anthropic-ai` doesn't apply to ClaudeBot, and the next vendor to rename a crawler may not be as generous. Well-behaved crawlers read it and obey it; nothing forces them to. Named groups also don't inherit anything from `User-agent: *`: a crawler follows the one group that names it and ignores the rest. If you give AI search crawlers their own `Allow` group, repeat your private paths in it. ## Cloudflare's managed robots.txt and Content Signals 3.1% of Polish sites serve [Cloudflare's managed robots.txt](https://developers.cloudflare.com/bots/additional-configurations/managed-robots-txt/): one switch in the dashboard, and the CDN puts a block of rules in front of the site's own file and keeps it up to date. The block disallows exactly the eight training crawlers in our list, from GPTBot to Amazonbot, and no search crawler. Those sites, among them the Polska Press regional titles such as poranny.pl, nto.pl, pomorska.pl and gazetawroclawska.pl, make up 15.1% of all Polish blockers and 24% of the sites that block training only. They are part of why Poland leans that way, though not all of it: take them out and training-only blocks are still more common than in the global sample, where we found no managed file at all. Large international sites write their own. The switch shows in the numbers. Behind Cloudflare, 25.6% of Polish sites block AI crawlers, against 18.2% elsewhere. Globally it is the other way round. **In Poland, sites behind Cloudflare block more** (Share blocking at least one AI crawler, by whether the site is served through Cloudflare) | Item | Behind Cloudflare | Other hosting | | --- | --- | --- | | Poland | 26% | 18% | | Global | 19% | 24% | Source: Outofplace crawl, 23 Sept 2026. Chart: https://www.outofplace.space/data/blog/ai-crawlers-polish-web/charts/en/chart-in-poland-sites-behind-cloudflare-block-more.png Data: https://www.outofplace.space/data/blog/ai-bots-polish-web/2026-09-23/domains.csv The managed file also carries a `Content-Signal` line (`search=yes, ai-train=no`), part of the [Content Signals Policy](https://contentsignals.org/) Cloudflare published in September 2025 for saying what content may be used for: search, AI input, AI training. 3.9% of Polish sites say `ai-train=no` this way, against 1.6% globally. If a robots.txt tester flags `Content-Signal` as an unknown directive, that is expected and harmless: under RFC 9309 a crawler skips lines it doesn't recognise. 79% of those lines come from the managed file, not from anyone typing them. ## What is llms.txt, and who has one? [`llms.txt`](https://llmstxt.org) is a Markdown file that tells language models what a site is and where its important pages are, [proposed by Jeremy Howard](https://www.answer.ai/posts/2024-09-03-llmstxt.html) in September 2024. The spec asks only for an H1 with the site's name; we also required at least one link, since a file without one points nowhere. 9.2% of the Polish sample serve such a file at `/llms.txt`. Another 10% answer `/llms.txt` with an HTML page and a 200 status, which is what a model reading it receives instead, and what a careless checker counts as "has llms.txt". Hosting companies got there first: home.pl, nazwa.pl, cyberfolks.pl, dhosting.pl and kei.pl all serve a valid file, and so do mbank.pl, t-mobile.pl, rossmann.pl, mediamarkt.pl and, among public bodies, uokik.gov.pl. Google has said [its Search ignores the file](https://developers.google.com/search/docs/fundamentals/ai-optimization-guide?hl=en); it is there for the other assistants. We serve one ourselves: [our llms.txt](https://www.outofplace.space/llms.txt) lists our studies, case studies and services, and [llms-full.txt](https://www.outofplace.space/llms-full.txt) carries the full text of every post. **Most sites have no llms.txt; one in ten answer with a web page** (What GET /llms.txt returned · n = 1,000 per sample · errors and paths our crawler was asked not to fetch count as none) | Item | Valid llms.txt | Wrong format | HTML page (soft 404) | None | n | | --- | --- | --- | --- | --- | --- | | Top 1,000 .pl | 9.2% | 2.4% | 10% | 78% | 1,000 | | Global top 1,000 | 14% | 3.4% | 9.6% | 74% | 1,000 | Source: Outofplace crawl, 23 Sept 2026. Chart: https://www.outofplace.space/data/blog/ai-crawlers-polish-web/charts/en/chart-most-sites-have-no-llms-txt-one-in-ten-answer-with-a-web-page.png Data: https://www.outofplace.space/data/blog/ai-bots-polish-web/2026-09-23/domains.csv ## The rest of the machine-readable web Crawlers read more than robots.txt. Here Polish sites are close to the world on the basics and behind on the details that help a machine place a page: [`hreflang` alternates](https://developers.google.com/search/docs/specialty/international/localized-versions?hl=en), and [a sitemap listed where crawlers look for it](https://www.sitemaps.org/protocol.html#submit_robots). [Zstandard compression](https://www.rfc-editor.org/rfc/rfc8878.html) is several times more common on .pl homepages, but that is Cloudflare again: all but one of the Polish sites that sent it came through the CDN. ## Three policies worth copying Some Polish sites have clearly thought about it. We checked each of these files by hand on the day of the crawl. - **Training no, answers yes.** The Wirtualna Polska titles (wp.pl, o2.pl, money.pl, pudelek.pl, abczdrowie.pl, dobreprogramy.pl) [disallow GPTBot, CCBot and Bytespider](https://www.wp.pl/robots.txt), and give OAI-SearchBot, ChatGPT-User, PerplexityBot and Perplexity-User their own `Allow: /` groups. A publisher saying exactly what it wants. - **Everyone welcome, by name.** mediamarkt.pl and euro.com.pl have [`Allow: /` groups](https://mediamarkt.pl/robots.txt) for GPTBot, ClaudeBot, OAI-SearchBot, Claude-SearchBot, PerplexityBot and more. For a retailer, an assistant that recommends a product is a shop window. - **Let the CDN keep the list.** The Polska Press regional titles use Cloudflare's managed file, which blocks the training crawlers and stays current without anyone editing it. ## How to block AI crawlers in robots.txt ### Should you block AI crawlers? There is no single right policy. A publisher that sells its archive has different interests from a shop that wants to be recommended. But whichever you choose, choose by purpose, not by vendor: ### A robots.txt that blocks AI training and keeps AI search This is the shape of a robots.txt that blocks training and keeps AI search. The training group is the same eight crawlers Cloudflare's managed file blocks. Our own site does the opposite and [lets every crawler in](https://www.outofplace.space/robots.txt), because being cited is the point of a studio blog. ```txt title="robots.txt" # AI search and user-triggered fetches: they cite and link to you User-agent: OAI-SearchBot User-agent: ChatGPT-User User-agent: Claude-SearchBot User-agent: Claude-User User-agent: PerplexityBot User-agent: Perplexity-User Allow: / # Named groups don't inherit from *, so repeat private paths Disallow: /cart/ Disallow: /account/ # Model training User-agent: GPTBot User-agent: ClaudeBot User-agent: Google-Extended User-agent: Applebot-Extended User-agent: CCBot User-agent: Bytespider User-agent: meta-externalagent User-agent: Amazonbot Disallow: / # Everyone else, including Googlebot and Bingbot User-agent: * Disallow: /cart/ Disallow: /account/ Sitemap: https://www.example.pl/sitemap.xml ``` Training, AI search, user-triggered fetches. Write down what you want for each before you touch the file. Take them from [the vendors' own pages](#further-reading), not from a list someone posted in 2023. A crawler that matches a named group ignores `User-agent: *` entirely. Paste the file into the checker below and look at the paths that matter to you. Copilot rides on Bingbot, so its controls are the [`noarchive` and `nocache` meta tags](https://www.bing.com/webmasters/help/robots-meta-tags-and-attributes-that-bing-supports-5198d240). Google's AI Overviews follow Googlebot and [snippet controls such as `nosnippet`](https://developers.google.com/search/docs/crawling-indexing/robots-meta-tag?hl=en). Fetchers that may ignore robots.txt need a CDN or server rule if they must be stopped. New agents appear a few times a year. A stale list is worse than none, because it looks like a decision. Paste your robots.txt here to see what each AI crawler may fetch. It runs in your browser on the same parser we used for the study; nothing is sent anywhere. ## How we measured We took the [Tranco list Y8YYG](https://tranco-list.eu/list/Y8YYG/full) (created 22 September 2026, 30-day window, five providers) and walked it from the top, keeping the first 1,000 registrable .pl domains whose server answered. The global control is the first 1,000 domains of the same list, processed the same way. On 23 September our crawler fetched each domain's robots.txt and homepage, plus `/llms.txt`, `/llms-full.txt` and the sitemap, identifying itself as `OutofplaceResearchBot` with a link to a page explaining the study. Each robots.txt was parsed under RFC 9309 and evaluated for the homepage path `/` for 38 user agents: 36 AI crawlers of 17 operators, retired names included, plus Googlebot and Bingbot for reference. The charts show the tokens each operator currently documents. A crawler counts as blocked when the rule that applies to it disallows `/`, whether it is named or falls back to `*`. Shares have 95% Wilson intervals; differences between the samples, Newcombe intervals. The sample is the first 1,000 .pl domains in the Tranco list that answered our requests. 23 of them disallow our crawler from their homepage in robots.txt; we still read their robots.txt (it is public by definition) but fetched nothing else. Domains that failed DNS, refused connections or returned server errors were skipped and replaced by the next one on the list. With 1,000 sites, the 95% interval around the headline share runs from 18.1% to 23.1%. Each site gets one class, first match wins. Blanket block: Googlebot is blocked too, so the rule is not about AI. All AI: every training crawler and at least one search crawler or fetcher are blocked. Stale list: a name the vendor no longer documents is blocked, while the crawler that replaced it is not. Training only: training crawlers blocked, no search crawler or fetcher. Inconsistent: a search crawler or fetcher is blocked while a training crawler is let through. Open: no AI crawler blocked. The rules and their tests are in the repository. robots.txt is only what a site asks for. Some sites block AI crawlers at the firewall or CDN instead, and some serve a bot challenge to every crawler; we counted challenges separately and never as "allowed". We evaluated the homepage path only: a site that blocks AI crawlers from its articles but not from `/` counts as open. A .pl domain is not the same as a Polish business, and popularity lists favour big sites. The numbers describe the most visited part of the .pl web on one day. ## Further reading - [RFC 9309: Robots Exclusion Protocol](https://www.rfc-editor.org/rfc/rfc9309.html): The standard for robots.txt: groups, matching and what a crawler must do when the file is missing or unreachable. - [OpenAI: Overview of OpenAI crawlers](https://developers.openai.com/api/docs/bots): GPTBot, OAI-SearchBot and ChatGPT-User, and what each is used for. - [Anthropic: How site owners can block the crawlers](https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler): ClaudeBot, Claude-SearchBot and Claude-User. - [Google: Common crawlers, including Google-Extended](https://developers.google.com/crawling/docs/crawlers-fetchers/google-common-crawlers?hl=en): What Google-Extended controls, and what it doesn't. - [Cloudflare: Managed robots.txt](https://developers.cloudflare.com/bots/additional-configurations/managed-robots-txt/): What the one-click file blocks, and the Content-Signal line it adds. - [llms.txt: the proposal and its format](https://llmstxt.org): The original specification: where the file lives and how it is structured. - [OutofplaceResearchBot: what it fetches and how to block it](https://www.outofplace.space/bot): Who runs our crawler, how it identifies itself and how politely it crawls. --- # Next.js prefetch, measured: Partial Prefetching in 16.3 (EN) Source: https://www.outofplace.space/blog/nextjs-prefetch-benchmark · Published 2026-09-23 Category: Performance · Tags: Next.js, Prefetching, React > Partial Prefetching in Next.js 16.3 cut prefetched bytes by up to 93% in our benchmark, but URL content now lands after the click. When each setup wins. [Next.js 16.3](https://nextjs.org/blog/next-16-3) changes what Next.js prefetch means. With [`partialPrefetching`](https://nextjs.org/docs/app/api-reference/config/next-config-js/partialPrefetching) on, the [App Router](https://nextjs.org/docs/app) no longer fetches a payload for every link on screen: it fetches one reusable shell per route and leaves the URL-specific content for after the click. We built one app six ways and clicked through it 3,744 times to see what that trade buys and what it costs. On a docs page with a 100-link sidebar, prefetched bytes fell by 93%. The price is time: the content that used to be on screen almost at once now arrives a few hundred milliseconds after the click, behind a skeleton. Which side of that trade you want depends on the page, and on [one detail of React that is easy to miss](#the-skeleton-you-will-see-300-ms). - Partial Prefetching cut prefetched bytes by 78–93% on link-heavy pages, and a whole docs navigation by about eleven times. - The shell paints instantly. The URL content arrives after the click: 345ms on docs against 52ms with Cache Components alone, on a throttled phone. - A skeleton that appears stays for about 300 ms, because React holds each reveal for that long. - Links with `prefetch={true}` stay instant in every setup. - Keep `prefetchInlining` at its default; HTTP/2 changed nothing that matters. ## What Partial Prefetching changes Until 16.3, the App Router prefetched per link: as N links scrolled into view, it fetched about N route payloads. With `partialPrefetching` (it requires [`cacheComponents`](https://nextjs.org/docs/app/api-reference/config/next-config-js/cacheComponents)), it fetches one [App Shell](https://nextjs.org/docs/app/guides/adopting-partial-prefetching) per route: everything the page renders that doesn't depend on the URL. Links to the same route share that one prefetch. Content that reads `params` or `searchParams` resolves after the click, unless a link asks for more with `prefetch={true}`. We compared six builds of the same app: - **L0**, no Cache Components: the classic App Router. - **C0**, Cache Components alone. - **P0**, Cache Components with Partial Prefetching and the default prefetch inlining. - **P1–P3**, the same with inlining off, with small thresholds and with large ones. Each was clicked through three navigations: a catalog of 48 products to a product page, a docs page with a 100-link sidebar to another docs page, and a search results page to another search through one of eight `prefetch={true}` links. ## Prefetch bytes: up to 93% fewer On the docs page, Cache Components alone prefetched every visible sidebar link. Partial Prefetching took that from 402 kB → 29 kB (−93%), in 8.5 requests instead of 49. On the catalog the saving was 78%. Search barely moved, because its eight `prefetch={true}` links still fetch their own results. **Partial Prefetching prefetches a fraction of the bytes on link-heavy pages** (Compressed RSC bytes prefetched before the click · desktop, unthrottled, warm · medians of 20 runs) | Item | Catalog → product | Docs → docs | Search → search | | --- | --- | --- | --- | | L0 · Classic router | 30 kB | 408 kB | 98 kB | | C0 · Cache Components | 160 kB | 402 kB | 163 kB | | P0 · Partial Prefetching | 34 kB | 29 kB | 107 kB | | P1 · Inlining off | 33 kB | 32 kB | 105 kB | | P2 · Small inlining | 32 kB | 28 kB | 103 kB | | P3 · Large inlining | 34 kB | 33 kB | 105 kB | Source: Outofplace benchmark, Next.js 16.3.6. Chart: https://www.outofplace.space/data/blog/nextjs-prefetch-benchmark/charts/en/chart-partial-prefetching-prefetches-a-fraction-of-the-bytes-on-link-h.png Data: https://www.outofplace.space/data/blog/nextjs-partial-prefetching-benchmark/2026-09-23/summary_by_variant.csv One row deserves a second look. Turning on Cache Components without Partial Prefetching made the catalog prefetch five times more than the classic router, 30 kB → 160 kB (+426%): its prerendered product shells become prefetchable per link. If you adopt Cache Components on link-heavy pages, decide on Partial Prefetching at the same time. ## Is navigation actually instant? The shell is. Under Partial Prefetching it painted 43ms–45ms after the click on a throttled phone. The part people came for, the product or the article, is another story: it needs a request after the click. **Cache Components alone shows the content at once; Partial Prefetching after a round trip** (Click to the first frame with the URL-specific content · mobile, Lighthouse-mobile throttling, warm click · ms) | Item | L0 · Classic router | C0 · Cache Components | P0 · Partial Prefetching | | --- | --- | --- | --- | | Catalog → product | 375ms | 60ms | 262ms | | Docs → docs | 47ms | 52ms | 345ms | | Search → search | 40ms | 41ms | 43ms | Source: Outofplace benchmark, Next.js 16.3.6. Chart: https://www.outofplace.space/data/blog/nextjs-prefetch-benchmark/charts/en/chart-cache-components-alone-shows-the-content-at-once-partial-prefetc.png Data: https://www.outofplace.space/data/blog/nextjs-partial-prefetching-benchmark/2026-09-23/summary_by_variant.csv On docs, Cache Components alone painted the new article in 52ms; Partial Prefetching took 345ms. On the product page the gap was smaller, 60ms → 262ms (+337%), and the classic router was slowest of all, because it had nothing prerendered for the dynamic route. Search, where every link carries `prefetch={true}`, painted in 40ms–43ms in every setup. ### The skeleton you will see: 300 ms On the product page the three stages land one after another: the shell, the product, then the part rendered at request time. **Under Partial Prefetching the page arrives in three steps** (Catalog → product · mobile, throttled, warm · click to each stage on screen · ms) | Item | Shell | Content | Request-time part | | --- | --- | --- | --- | | L0 · Classic router | 375ms | 375ms | 375ms | | C0 · Cache Components | 60ms | 60ms | 342ms | | P0 · Partial Prefetching | 47ms | 262ms | 562ms | Source: Outofplace benchmark, Next.js 16.3.6. Chart: https://www.outofplace.space/data/blog/nextjs-prefetch-benchmark/charts/en/chart-under-partial-prefetching-the-page-arrives-in-three-steps.png Data: https://www.outofplace.space/data/blog/nextjs-partial-prefetching-benchmark/2026-09-23/summary_by_variant.csv On docs, the skeleton stayed up 298ms in every Partial Prefetching build. That is not network time: the response had arrived after 250ms. The React build inside Next.js 16.3.6 (`19.3.0-canary`) [waits until 300 ms](https://github.com/react/react/blob/d083ec1da1e5252abd3ddfdde6dfbc09701a2c51/packages/react-reconciler/src/ReactFiberWorkLoop.js#L528) after the last switch between a [Suspense](https://react.dev/reference/react/Suspense) fallback and its content before it commits the next one, so a loading state doesn't flash. The side effect: stages chain. A skeleton that appears is on screen for about 300 ms even when the data is already there, and a separate boundary for request-time data reveals 300 ms after the one before it. Design skeletons to be looked at, and consider one boundary where two would only chain. ### An early click Clicking 50 ms after the page hydrates, before any prefetch finishes, takes the advantage away from everyone who prefetches. Cache Components alone is as slow as Partial Prefetching then, because both show a fallback first and then wait out the 300 ms. **Clicked early, no setup has prefetched yet** (Click 50 ms after hydration · mobile, throttled · click to URL content · ms) | Item | L0 · Classic router | C0 · Cache Components | P0 · Partial Prefetching | | --- | --- | --- | --- | | Catalog → product | 367ms | 374ms | 389ms | | Docs → docs | 262ms | 528ms | 529ms | | Search → search | 234ms | 404ms | 339ms | Source: Outofplace benchmark, Next.js 16.3.6. Chart: https://www.outofplace.space/data/blog/nextjs-prefetch-benchmark/charts/en/chart-clicked-early-no-setup-has-prefetched-yet.png Data: https://www.outofplace.space/data/blog/nextjs-partial-prefetching-benchmark/2026-09-23/summary_by_variant.csv The last stage, the part rendered at request time, splits these clicks in two. With Partial Prefetching, 6 of 12 early clicks on the product page had the whole page on screen after 375ms–400ms. The other 6 took 650ms–700ms, and no run landed in between. The gap is about the 300 ms React waits between two reveals: the last stage arrives with the product or one wait later. **The last stage lands with the product or about 300 ms later** (Catalog → product, early click · Partial Prefetching · mobile, throttled, HTTP/2 run · click to the last stage on screen · ms) | Range | runs | | --- | --- | | 375ms–400ms | 6 | | 400ms–425ms | 0 | | 425ms–450ms | 0 | | 450ms–475ms | 0 | | 475ms–500ms | 0 | | 500ms–525ms | 0 | | 525ms–550ms | 0 | | 550ms–575ms | 0 | | 575ms–600ms | 0 | | 600ms–625ms | 0 | | 625ms–650ms | 0 | | 650ms–675ms | 3 | | 675ms–700ms | 3 | Source: Outofplace benchmark, Next.js 16.3.6. Chart: https://www.outofplace.space/data/blog/nextjs-prefetch-benchmark/charts/en/chart-the-last-stage-lands-with-the-product-or-about-300-ms-later.png Data: https://www.outofplace.space/data/blog/nextjs-partial-prefetching-benchmark/2026-09-23/runs.csv ## The whole bill Partial Prefetching moves bytes from before the click to after it, and on the product page it also doubles them: the navigation response carried two complete copies of the product's payload, 13 kB → 26 kB (+99%) compressed. Counting everything a navigation downloads, prefetch included, it still comes out far ahead on docs. **A whole docs navigation costs about a tenth of the bytes** (All compressed RSC bytes of one docs navigation, prefetch and navigation · mobile, throttled, warm) | Item | Value | n | | --- | --- | --- | | L0 · Classic router | 463 kB | 20 | | C0 · Cache Components | 458 kB | 20 | | P0 · Partial Prefetching | 40 kB | 20 | | P1 · Inlining off | 43 kB | 20 | | P2 · Small inlining | 39 kB | 20 | | P3 · Large inlining | 44 kB | 20 | Source: Outofplace benchmark, Next.js 16.3.6. Chart: https://www.outofplace.space/data/blog/nextjs-prefetch-benchmark/charts/en/chart-a-whole-docs-navigation-costs-about-a-tenth-of-the-bytes.png Data: https://www.outofplace.space/data/blog/nextjs-partial-prefetching-benchmark/2026-09-23/summary_by_variant.csv About 11 times fewer bytes on docs, 2.2 times fewer on the catalog. Hydration and input delay on the start page didn't change in any setup. ## When `prefetch={true}` is worth it [` `](https://nextjs.org/docs/app/api-reference/components/link#prefetch) fetches the destination's URL data too, so its click is as instant as a full prefetch. It is paid per visible link: the eight search links cost 107 kB, more than the whole docs page under Partial Prefetching. Use it where the click is likely and the data cacheable: search suggestions, a "next" link, the first few results. On grids of cards, prefetch on intent (hover or touch) instead of marking every card. ## prefetchInlining: leave it on [`experimental.prefetchInlining`](https://nextjs.org/docs/app/api-reference/config/next-config-js/prefetchInlining) bundles small prefetch responses into one. The default and our two threshold settings differed by at most one request and a few kilobytes, with no measurable difference in time. Turning inlining off added five or six requests on the catalog and search, and over HTTP/1.1 that cost an early click on the catalog: six connections per host fill up. Over [HTTP/2](https://www.rfc-editor.org/rfc/rfc9113.html) the penalty disappeared. Every other setup was the same on both protocols. **HTTP/2 only matters with inlining turned off** (Catalog → product, early click · mobile, throttled · click to content · HTTP/1.1 and HTTP/2 interleaved, 12 runs each · ms) | Item | HTTP/1.1 | HTTP/2 | | --- | --- | --- | | L0 · Classic router | 365ms | 363ms | | C0 · Cache Components | 375ms | 373ms | | P0 · Partial Prefetching | 387ms | 386ms | | P1 · Inlining off | 440ms | 372ms | | P2 · Small inlining | 390ms | 383ms | | P3 · Large inlining | 387ms | 386ms | Source: Outofplace benchmark, Next.js 16.3.6. Chart: https://www.outofplace.space/data/blog/nextjs-prefetch-benchmark/charts/en/chart-http-2-only-matters-with-inlining-turned-off.png Data: https://www.outofplace.space/data/blog/nextjs-partial-prefetching-benchmark/2026-09-23/summary_by_variant.csv ## Which setup to choose ```ts title="next.config.ts" const nextConfig: NextConfig = { cacheComponents: true, // One App Shell per route instead of one prefetch per link. partialPrefetching: true, }; ``` ```tsx title="The links where a wait would hurt" {suggestion} ``` Pages with many links to the same dynamic route (docs, grids, feeds) gain the most. A page with a handful of links to different pages gains little. Where instant content matters more than bytes and the pages are static, keep a route on full prefetch; elsewhere turn Partial Prefetching on. The [segment config `export const prefetch = 'partial'`](https://nextjs.org/docs/app/api-reference/file-conventions/route-segment-config/prefetch) adopts it one route at a time. Add `prefetch={true}` where the next click is predictable and the data cacheable. It will be seen for about 300 ms on a slow connection. Put the title and navigation in the part that doesn't depend on the URL: it paints almost at once. This site runs Partial Prefetching: most of our pages have few links, and the ones that have many (the blog, the case studies) are static and small. Our own products, Portivo, Sprawna Matura and Levera, are built on Next.js too. ## How we measured A generated App Router app (a 48-product catalog, 100 docs pages with a sidebar in the layout, a search page) was built six times with Next.js 16.3.6, differing only in `cacheComponents`, `partialPrefetching` and `experimental.prefetchInlining`, and served with `next start`. Headless Chrome opened a start page in a fresh browser context, waited, and clicked a link, recording every request over the [DevTools protocol](https://chromedevtools.github.io/devtools-protocol/) and timing marker elements from the click event to the next animation frame. Each combination of setup, scenario, phone or desktop, fast or [throttled network (150 ms, 1.6 Mbps, 4× CPU)](https://github.com/GoogleChrome/lighthouse/blob/main/docs/throttling.md) and early or late click ran 20 times in shuffled order, 2,880 runs; an HTTP/2 arm added 864 more. Medians carry seeded bootstrap 95% intervals. It is a lab: one machine, localhost, no CDN, no images or fonts, generated text, and a synthetic 300 ms request-time delay. Throttling is per request, not packet-level, so absolute times are optimistic; the comparisons are what count. A late click is the best case for prefetching and an early one a specific race; real users click everywhere in between. The App Shell mechanics, the doubled product payload and the 300 ms reveal are specific to Next.js 16.3.6 and its React canary and may change in a patch release. The shared shell is fetched through the URL of whichever link the router schedules first, and that response carries the whole page for that URL. A click on exactly that link needs no request at all. We picked click targets further down the page so that this didn't flatter Partial Prefetching. ## Further reading - [Next.js: Adopting Partial Prefetching](https://nextjs.org/docs/app/guides/adopting-partial-prefetching): How the App Shell works and how to move routes over one at a time. - [Next.js: Optimizing prefetching](https://nextjs.org/docs/app/guides/optimizing-prefetching): Per-link prefetching with prefetch={true}, and when to use it. - [Next.js: prefetchInlining](https://nextjs.org/docs/app/api-reference/config/next-config-js/prefetchInlining): What inlining bundles, and why the default is right for most apps. --- # Boty AI kontra polski internet: robots.txt 1000 stron .pl (PL) Source: https://www.outofplace.space/pl/blog/ai-crawlers-polish-web · Published 2026-09-23 Category: Wyszukiwanie i AI · Tags: boty AI, robots.txt, GPTBot, llms.txt, Polska > Sprawdziliśmy plik robots.txt 1000 największych stron .pl: co piąta blokuje boty AI, głównie treningowe jak GPTBot, rzadko wyszukiwarki AI. Jak to rozdzielić? 23 września 2026 nasz crawler [OutofplaceResearchBot](https://www.outofplace.space/pl/bot) przeczytał plik robots.txt tysiąca najpopularniejszych stron w domenie .pl, żebyśmy mogli sprawdzić, które boty AI wpuszczają. 20,5% z nich blokuje co najmniej jednego bota AI. Statystycznie to tyle samo, co w światowym top 1000, gdzie odsetek wynosi 22,7%. Różnica leży w tym, _których_ botów to dotyczy. Polskie strony zamykają drzwi przed botami zbierającymi dane treningowe, a wyszukiwarki AI w większości wpuszczają. Bota wyszukiwarki AI blokuje tylko 5,4% z nich, przy 13,9% na świecie. Jeśli chcesz pojawiać się w odpowiedziach ChatGPT, Claude'a czy Perplexity, to właściwa kolejność. Tyle że czasem wychodzi przypadkiem: [stara lista skopiowana raz i nigdy nieaktualizowana](#listy-sie-starzeja). - Co piąta strona z top 1000 .pl blokuje w robots.txt jakiegoś bota AI, tyle samo co w światowym top 1000. - GPTBot jest najczęściej blokowanym botem w Polsce i na świecie: 15,1% i 15,4%. - Polskie strony blokują trening, nie wyszukiwanie. Spośród stron blokujących GPTBota mniej niż jedna trzecia blokuje też OAI-SearchBota; na świecie ponad połowa. - Liczy się wielkość: boty AI blokuje jedna trzecia stron z top 100 .pl i co szósta z miejsc 501–1000. - Co siódma polska strona blokująca boty nie napisała tych reguł sama: używa [zarządzanego robots.txt od Cloudflare](#zarzadzany-robotstxt-od-cloudflare-i-content-signals). W próbie światowej nie znaleźliśmy ani jednej takiej. ## Boty AI: co to jest i jakie są rodzaje „Bot AI” to trzy różne zadania i dostawcy dają dziś każdemu osobny user agent. To rozróżnienie decyduje o wszystkim, co dalej: - **Boty treningowe** zbierają strony do trenowania modeli: [GPTBot](https://developers.openai.com/api/docs/bots) (OpenAI), [ClaudeBot](https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler) (Anthropic), [Google-Extended](https://developers.google.com/crawling/docs/crawlers-fetchers/google-common-crawlers?hl=pl), [Applebot-Extended](https://support.apple.com/pl-pl/119829), [CCBot](https://commoncrawl.org/ccbot) (Common Crawl), Bytespider (ByteDance), meta-externalagent (Meta), [Amazonbot](https://developer.amazon.com/amazonbot). - **Boty wyszukiwarek AI** indeksują strony, żeby asystent mógł je znaleźć i zacytować: OAI-SearchBot, Claude-SearchBot, [PerplexityBot](https://docs.perplexity.ai/docs/resources/perplexity-crawlers), [DuckAssistBot](https://duckduckgo.com/duckduckgo-help-pages/results/duckassistbot). - **Boty pobierające na prośbę użytkownika** otwierają stronę, bo ktoś poprosił o to asystenta: ChatGPT-User, Claude-User, Perplexity-User, meta-externalfetcher, [MistralAI-User](https://docs.mistral.ai/robots). Blokada bota treningowego trzyma strony z dala od przyszłych modeli. Blokada bota wyszukiwarki albo bota pobierającego usuwa je z odpowiedzi i z linków w tych odpowiedziach. Dwa z „botów” treningowych niczego nie pobierają. Google-Extended i Applebot-Extended to tylko tokeny: stronę ściąga [Googlebot](https://developers.google.com/search/docs/crawling-indexing/googlebot?hl=pl) albo Applebot, a token mówi Google'owi i Apple, czy wolno na niej trenować. Ich blokada zostawia stronę w wyszukiwarkach Google i Apple, a blokada Google-Extended nie wyłącza też [AI Overviews](https://developers.google.com/search/docs/appearance/ai-features?hl=pl): te słuchają reguł dla Googlebota. Drugi wyjątek to boty pobierające. OpenAI, Perplexity i Meta piszą, że ich boty mogą w ogóle nie stosować się do robots.txt, bo o stronę poprosił człowiek; Anthropic deklaruje, że Claude-User się stosuje. ## Polska blokuje tak często jak świat, ale nie te same boty Licząc wszystkie boty AI, różnica między próbą .pl a światową wynosi 2,2 punktu procentowego, dobrze w granicach błędu. Po podziale według przeznaczenia obraz się zmienia. Boty treningowe są blokowane równie często. Boty wyszukiwarek AI i boty pobierające na prośbę użytkownika – w Polsce ponad dwa razy rzadziej. **Polska blokuje trening tak często jak świat, wyszukiwarki AI dużo rzadziej** (Odsetek stron, których robots.txt blokuje stronę główną co najmniej jednemu botowi danego rodzaju · n = 1000 w każdej próbie · 23 września 2026) | Pozycja | Top 1000 .pl | Globalny top 1000 | | --- | --- | --- | | Dowolny bot AI | 21% | 23% | | Boty treningowe | 21% | 22% | | Boty wyszukiwarek AI | 5,4% | 14% | | Pobieranie na prośbę użytkownika | 5,3% | 13% | Source: Badanie Outofplace; lista Tranco Y8YYG. Chart: https://www.outofplace.space/data/blog/ai-crawlers-polish-web/charts/pl/chart-polska-blokuje-trening-tak-czesto-jak-swiat-wyszukiwarki-ai-duzo.png Data: https://www.outofplace.space/data/blog/ai-bots-polish-web/2026-09-23/domains.csv ### GPTBot czy OAI-SearchBot: dwa osobne przełączniki Najwyraźniej widać to u poszczególnych dostawców. Spośród polskich stron blokujących GPTBota OAI-SearchBota blokuje też 28%. W próbie światowej – 56%. U Anthropic różnica jest jeszcze większa: 24% wobec 60%. **Większość polskich stron blokujących trening zostawia wyszukiwarki AI otwarte** (Strony blokujące bota treningowego danego dostawcy (GPTBot, ClaudeBot) według tego, czy blokują też jego bota wyszukiwarki (OAI-SearchBot, Claude-SearchBot)) | Pozycja | Blokuje też wyszukiwanie | Blokuje tylko trening | n | | --- | --- | --- | --- | | OpenAI · Polska | 28% | 72% | 151 | | OpenAI · Świat | 56% | 44% | 154 | | Anthropic · Polska | 24% | 76% | 107 | | Anthropic · Świat | 60% | 40% | 152 | Source: Badanie Outofplace, 23 września 2026. Chart: https://www.outofplace.space/data/blog/ai-crawlers-polish-web/charts/pl/chart-wiekszosc-polskich-stron-blokujacych-trening-zostawia-wyszukiwar.png Data: https://www.outofplace.space/data/blog/ai-bots-polish-web/2026-09-23/bot_policies.csv Na świecie strona, która blokuje boty AI, zwykle blokuje też te wyszukujące. W Polsce typowa strona wymienia kilka botów treningowych i na tym kończy: 11,2% polskiej próby blokuje wyłącznie trening, prawie dwa razy więcej niż 5,9% na świecie. ## Które boty AI są blokowane najczęściej Prowadzi GPTBot: 15,1% polskich stron i 15,4% na świecie. OpenAI ogłosiło go w sierpniu 2023 i od niego zaczyna się większość list blokad. Dalej są CCBot z Common Crawl, Amazonbot, Bytespider od ByteDance i ClaudeBot. Najnowsze boty są ledwo widoczne: Claude-SearchBota blokuje 2,7% polskich stron, głównie dlatego, że listy powstały, zanim się pojawił. **GPTBot to najczęściej blokowany bot AI na stronach .pl** (Odsetek domen z top 1000 .pl, których robots.txt blokuje danemu botowi stronę główną · 95-procentowe przedziały · wyróżnione boty treningowe) | Pozycja | Wartość | Przedział 95% | n | | --- | --- | --- | --- | | GPTBot | 15% | 13%–17% | 1000 | | CCBot | 13% | 11%–16% | 1000 | | Amazonbot | 12% | 10%–14% | 1000 | | Bytespider | 12% | 9,7%–14% | 1000 | | ClaudeBot | 11% | 8,9%–13% | 1000 | | meta-externalagent | 11% | 8,8%–13% | 1000 | | Google-Extended | 8,4% | 6,8%–10% | 1000 | | Applebot-Extended | 8,4% | 6,8%–10% | 1000 | | ChatGPT-User | 4,7% | 3,6%–6,2% | 1000 | | PerplexityBot | 4,6% | 3,5%–6,1% | 1000 | | OAI-SearchBot | 4,4% | 3,3%–5,9% | 1000 | | meta-externalfetcher | 4% | 3%–5,4% | 1000 | | DuckAssistBot | 3,7% | 2,7%–5,1% | 1000 | | Perplexity-User | 3,5% | 2,5%–4,8% | 1000 | | Claude-SearchBot | 2,7% | 1,9%–3,9% | 1000 | | Claude-User | 2,3% | 1,5%–3,4% | 1000 | | MistralAI-User | 2,3% | 1,5%–3,4% | 1000 | Source: Badanie Outofplace, 23 września 2026. Chart: https://www.outofplace.space/data/blog/ai-crawlers-polish-web/charts/pl/chart-gptbot-to-najczesciej-blokowany-bot-ai-na-stronach-pl.png Data: https://www.outofplace.space/data/blog/ai-bots-polish-web/2026-09-23/summary_by_bot.csv Prawie wszystko to świadome decyzje. Strona, która w grupie `User-agent: *` zabrania wszystkiego, blokuje każdego bota bez własnej grupy, AI czy nie. Ta „podłoga” odpowiada w przypadku GPTBota tylko za 1,9% polskich stron; reszta wymienia bota z nazwy. Poniżej polski i światowy odsetek dla każdego bota obok siebie. Obie próby zgadzają się co do botów treningowych. Rozchodzą się przy wszystkim, co służy odpowiedziom. **Różnica pojawia się przy wyszukiwarkach i botach pobierających** (Odsetek stron blokujących danemu botowi stronę główną · top 1000 .pl wobec światowego top 1000) | Pozycja | Top 1000 .pl | Globalny top 1000 | | --- | --- | --- | | GPTBot | 15% | 15% | | CCBot | 13% | 18% | | Amazonbot | 12% | 13% | | Bytespider | 12% | 17% | | ClaudeBot | 11% | 15% | | meta-externalagent | 11% | 14% | | Google-Extended | 8,4% | 14% | | Applebot-Extended | 8,4% | 13% | | ChatGPT-User | 4,7% | 11% | | PerplexityBot | 4,6% | 13% | | OAI-SearchBot | 4,4% | 8,9% | | meta-externalfetcher | 4% | 9,9% | | DuckAssistBot | 3,7% | 9,5% | | Perplexity-User | 3,5% | 10% | | Claude-SearchBot | 2,7% | 9,2% | | Claude-User | 2,3% | 9,5% | | MistralAI-User | 2,3% | 9,6% | Source: Badanie Outofplace, 23 września 2026. Chart: https://www.outofplace.space/data/blog/ai-crawlers-polish-web/charts/pl/chart-roznica-pojawia-sie-przy-wyszukiwarkach-i-botach-pobierajacych.png Data: https://www.outofplace.space/data/blog/ai-bots-polish-web/2026-09-23/summary_by_bot.csv ## Większe strony blokują częściej Im popularniejsza strona, tym częściej blokuje. Co najmniej jednego bota AI blokuje 33% ze stu najpopularniejszych domen .pl i 15,8% z miejsc od 501 do 1000. Spadek jest bardziej stromy niż w próbie światowej, gdzie wynosi od 27% do 21,2%. Największe polskie strony to wydawcy, marketplace'y i portale: firmy, które mają najwięcej tekstu do ochrony i prawników, którzy zadają to pytanie. **Jedna trzecia stron z top 100 .pl blokuje boty AI** (Odsetek blokujących co najmniej jednego bota AI według miejsca w próbie · n = 100, 400 i 500) | Pozycja | Top 1000 .pl | Globalny top 1000 | | --- | --- | --- | | Top 100 | 33% | 27% | | 101–500 | 23% | 24% | | 501–1000 | 16% | 21% | Source: Badanie Outofplace, 23 września 2026. Chart: https://www.outofplace.space/data/blog/ai-crawlers-polish-web/charts/pl/chart-jedna-trzecia-stron-z-top-100-pl-blokuje-boty-ai.png Data: https://www.outofplace.space/data/blog/ai-bots-polish-web/2026-09-23/domains.csv Efekt wielkości widać przy każdym bocie z osobna. GPTBota blokuje 23% stron z top 100 i 11,6% stron z miejsc od 501 do 1000; każdy bot treningowy idzie tym samym spadkiem. Boty wyszukiwarek zostają nisko przy każdej wielkości. **Największe strony .pl najczęściej blokują boty treningowe, a wyszukiwarki rzadko niezależnie od wielkości** (Odsetek stron .pl blokujących danemu botowi stronę główną według miejsca w próbie · n = 100, 400 i 500) | Pozycja | Top 100 | 101–500 | 501–1000 | | --- | --- | --- | --- | | GPTBot | 23% | 18% | 12% | | CCBot | 21% | 16% | 9,6% | | Amazonbot | 16% | 16% | 8,8% | | Bytespider | 19% | 14% | 8% | | ClaudeBot | 14% | 13% | 8,6% | | meta-externalagent | 15% | 14% | 6,8% | | Google-Extended | 12% | 11% | 5,8% | | Applebot-Extended | 11% | 11% | 5,6% | | OAI-SearchBot | 7% | 4,3% | 4% | | PerplexityBot | 6% | 4,8% | 4,2% | | Claude-SearchBot | 4% | 2,3% | 2,8% | Source: Badanie Outofplace, 23 września 2026. Chart: https://www.outofplace.space/data/blog/ai-crawlers-polish-web/charts/pl/chart-najwieksze-strony-pl-najczesciej-blokuja-boty-treningowe-a-wyszu.png Data: https://www.outofplace.space/data/blog/ai-bots-polish-web/2026-09-23/bot_policies.csv Instytucje publiczne blokują rzadziej i nie celowo. Bota AI blokuje 11,9% spośród 42 domen gov.pl w próbie i 16,1% spośród 31 domen edu.pl; obie grupy są małe, więc traktuj to jako kierunek. Większość tych blokad to w ogóle nie polityka wobec AI. sejm.gov.pl, dziennikustaw.gov.pl i wroclaw.sa.gov.pl [wpuszczają Googlebota i zabraniają wszystkim innym](https://www.sejm.gov.pl/robots.txt), co odcina każdego bota AI, a razem z nimi Binga i DuckDuckGo. W przypadku Sejmu i Dziennika Ustaw oznacza to też, że Copilot i wyszukiwarka ChatGPT nie przeczytają źródeł, o które ludzie ich pytają. ## Listy się starzeją Listy blokad pisze się raz i kopiuje latami. 7,2% polskich stron wciąż blokuje `anthropic-ai`, nazwę, której Anthropic przestał używać (dokumentacja wymienia dziś ClaudeBota, Claude-SearchBota i Claude-User), a Claude-SearchBota tylko 2,7%. W 4,1% próby lista jest nieaktualna w sposób, który ma znaczenie: blokuje starą nazwę, a bota, który ją zastąpił, przepuszcza. Trzeba oddać Anthropic, że [powiedział 404 Media w 2024 roku](https://www.404media.co/websites-are-blocking-the-wrong-ai-scrapers-because-ai-companies-keep-making-new-ones/), iż ClaudeBot respektuje reguły zapisane dla dwóch wycofanych nazw. To dobra wola, nie standard: według [RFC 9309](https://www.rfc-editor.org/rfc/rfc9309.html) reguła dla `anthropic-ai` ClaudeBota nie dotyczy, a następny dostawca, który zmieni nazwę bota, może nie być tak hojny. Porządne boty go czytają i się stosują; nic ich do tego nie zmusza. Nazwane grupy nie dziedziczą też niczego z `User-agent: *`: bot stosuje się do jednej grupy, która go wymienia, i ignoruje resztę. Jeśli dajesz botom wyszukiwarek AI własną grupę `Allow`, powtórz w niej swoje prywatne ścieżki. ## Zarządzany robots.txt od Cloudflare i Content Signals 3,1% polskich stron serwuje [zarządzany robots.txt od Cloudflare](https://developers.cloudflare.com/bots/additional-configurations/managed-robots-txt/): jeden przełącznik w panelu i CDN dokleja przed plikiem strony blok reguł, który sam aktualizuje. Blok zabrania dostępu dokładnie ośmiu botom treningowym z naszej listy, od GPTBota po Amazonbota, i żadnemu botowi wyszukiwarki. Te strony, w tym regionalne tytuły Polska Press, takie jak poranny.pl, nto.pl, pomorska.pl i gazetawroclawska.pl, to 15,1% wszystkich polskich stron blokujących boty AI i 24% stron blokujących tylko trening. Częściowo tłumaczą, czemu Polska skłania się w tę stronę, ale nie w całości: bez nich blokowanie samego treningu i tak jest częstsze niż w próbie światowej, gdzie zarządzanego pliku nie znaleźliśmy wcale. Duże międzynarodowe serwisy piszą własne. Ten przełącznik widać w liczbach. Za Cloudflare boty AI blokuje 25,6% polskich stron, a poza nim 18,2%. Na świecie jest odwrotnie. **W Polsce strony za Cloudflare blokują częściej** (Odsetek blokujących co najmniej jednego bota AI według tego, czy strona idzie przez Cloudflare) | Pozycja | Za Cloudflare | Inny hosting | | --- | --- | --- | | Polska | 26% | 18% | | Świat | 19% | 24% | Source: Badanie Outofplace, 23 września 2026. Chart: https://www.outofplace.space/data/blog/ai-crawlers-polish-web/charts/pl/chart-w-polsce-strony-za-cloudflare-blokuja-czesciej.png Data: https://www.outofplace.space/data/blog/ai-bots-polish-web/2026-09-23/domains.csv Zarządzany plik zawiera też linię `Content-Signal` (`search=yes, ai-train=no`), część [Content Signals Policy](https://contentsignals.org/), którą Cloudflare opublikował we wrześniu 2025, by deklarować, do czego wolno używać treści: wyszukiwanie, dane wejściowe dla AI, trening AI. W ten sposób `ai-train=no` deklaruje 3,9% polskich stron i 1,6% na świecie. Jeśli tester robots.txt oznacza `Content-Signal` jako nieznaną dyrektywę, to normalne i nieszkodliwe: według RFC 9309 bot pomija linie, których nie rozumie. 79% tych linii pochodzi z zarządzanego pliku, a nie od kogoś, kto je wpisał. ## Co to jest llms.txt i kto go ma? [`llms.txt`](https://llmstxt.org) to plik Markdown, który mówi modelom językowym, czym jest serwis i gdzie są najważniejsze strony; [zaproponował go Jeremy Howard](https://www.answer.ai/posts/2024-09-03-llmstxt.html) we wrześniu 2024. Specyfikacja wymaga tylko nagłówka H1 z nazwą serwisu; my wymagaliśmy też co najmniej jednego linku, bo plik bez linków nigdzie nie prowadzi. Taki plik pod `/llms.txt` serwuje 9,2% polskiej próby. Kolejne 10% odpowiada pod `/llms.txt` stroną HTML ze statusem 200 i to właśnie dostaje model, który próbuje go przeczytać, a niedbały checker liczy jako „ma llms.txt”. Pierwsze były firmy hostingowe: home.pl, nazwa.pl, cyberfolks.pl, dhosting.pl i kei.pl serwują poprawny plik, podobnie mbank.pl, t-mobile.pl, rossmann.pl, mediamarkt.pl, a z instytucji publicznych uokik.gov.pl. Google pisze, że [jego wyszukiwarka ten plik ignoruje](https://developers.google.com/search/docs/fundamentals/ai-optimization-guide?hl=pl); jest dla innych asystentów. Sami też go udostępniamy: [nasz llms.txt](https://www.outofplace.space/llms.txt) wymienia nasze badania, case studies i usługi, a [llms-full.txt](https://www.outofplace.space/llms-full.txt) zawiera pełny tekst każdego wpisu. **Większość stron nie ma llms.txt; co dziesiąta odpowiada stroną WWW** (Co zwróciło GET /llms.txt · n = 1000 w każdej próbie · błędy i ścieżki, których nie wolno nam było pobrać, liczymy jako brak) | Pozycja | Poprawny llms.txt | Zły format | Strona HTML (soft 404) | Brak | n | | --- | --- | --- | --- | --- | --- | | Top 1000 .pl | 9,2% | 2,4% | 10% | 78% | 1000 | | Globalny top 1000 | 14% | 3,4% | 9,6% | 74% | 1000 | Source: Badanie Outofplace, 23 września 2026. Chart: https://www.outofplace.space/data/blog/ai-crawlers-polish-web/charts/pl/chart-wiekszosc-stron-nie-ma-llms-txt-co-dziesiata-odpowiada-strona-ww.png Data: https://www.outofplace.space/data/blog/ai-bots-polish-web/2026-09-23/domains.csv ## Reszta sieci czytelnej dla maszyn Boty czytają więcej niż robots.txt. Polskie strony są tu blisko świata w podstawach i odstają w szczegółach, które pomagają maszynie umiejscowić stronę: [alternatywach `hreflang`](https://developers.google.com/search/docs/specialty/international/localized-versions?hl=pl) i [mapie strony wskazanej tam, gdzie boty jej szukają](https://www.sitemaps.org/protocol.html#submit_robots). [Kompresja Zstandard](https://www.rfc-editor.org/rfc/rfc8878.html) jest na stronach głównych .pl kilka razy częstsza, ale to znowu Cloudflare: wszystkie polskie strony, które ją wysłały, poza jedną, szły przez ten CDN. ## Trzy polityki warte skopiowania Niektóre polskie strony wyraźnie to przemyślały. Każdy z tych plików sprawdziliśmy ręcznie w dniu badania. - **Trening nie, odpowiedzi tak.** Tytuły Wirtualnej Polski (wp.pl, o2.pl, money.pl, pudelek.pl, abczdrowie.pl, dobreprogramy.pl) [zabraniają dostępu GPTBotowi, CCBotowi i Bytespiderowi](https://www.wp.pl/robots.txt), a OAI-SearchBotowi, ChatGPT-User, PerplexityBotowi i Perplexity-User dają własne grupy `Allow: /`. Wydawca, który mówi dokładnie, czego chce. - **Wszyscy mile widziani, z imienia.** mediamarkt.pl i euro.com.pl mają [grupy `Allow: /`](https://mediamarkt.pl/robots.txt) dla GPTBota, ClaudeBota, OAI-SearchBota, Claude-SearchBota, PerplexityBota i innych. Dla sprzedawcy asystent polecający produkt to witryna sklepowa. - **Niech listę prowadzi CDN.** Regionalne tytuły Polska Press używają zarządzanego pliku Cloudflare, który blokuje boty treningowe i pozostaje aktualny bez niczyjej edycji. ## Jak zablokować ChatGPT i inne boty AI w robots.txt ### Czy blokować boty AI? Nie ma jednej dobrej polityki. Wydawca, który sprzedaje archiwum, ma inne interesy niż sklep, który chce być polecany. Ale cokolwiek wybierzesz, wybieraj według przeznaczenia bota, nie według dostawcy: ### robots.txt, który blokuje trening i zostawia wyszukiwarki AI Tak wygląda robots.txt, który blokuje trening i zostawia wyszukiwarki AI. Grupa treningowa to te same osiem botów, które blokuje zarządzany plik Cloudflare. Nasza strona robi odwrotnie i [wpuszcza wszystkie boty](https://www.outofplace.space/robots.txt), bo blog studia jest po to, żeby go cytowano. ```txt title="robots.txt" # Wyszukiwarki AI i pobieranie na prośbę użytkownika: cytują i linkują User-agent: OAI-SearchBot User-agent: ChatGPT-User User-agent: Claude-SearchBot User-agent: Claude-User User-agent: PerplexityBot User-agent: Perplexity-User Allow: / # Nazwane grupy nie dziedziczą z *, więc powtórz prywatne ścieżki Disallow: /koszyk/ Disallow: /konto/ # Trenowanie modeli User-agent: GPTBot User-agent: ClaudeBot User-agent: Google-Extended User-agent: Applebot-Extended User-agent: CCBot User-agent: Bytespider User-agent: meta-externalagent User-agent: Amazonbot Disallow: / # Wszyscy pozostali, w tym Googlebot i Bingbot User-agent: * Disallow: /koszyk/ Disallow: /konto/ Sitemap: https://www.example.pl/sitemap.xml ``` Trening, wyszukiwarki AI, pobieranie na prośbę użytkownika. Zapisz, czego chcesz dla każdego, zanim dotkniesz pliku. Bierz je [ze stron dostawców](#dalsza-lektura), nie z listy, którą ktoś wrzucił w 2023. Bot, który pasuje do nazwanej grupy, całkowicie ignoruje `User-agent: *`. Wklej plik do narzędzia poniżej i sprawdź ścieżki, które są dla ciebie ważne. Copilot jeździ na Bingbocie, więc steruje się nim [meta tagami `noarchive` i `nocache`](https://www.bing.com/webmasters/help/robots-meta-tags-and-attributes-that-bing-supports-5198d240). AI Overviews w Google słuchają Googlebota i [ustawień fragmentów, np. `nosnippet`](https://developers.google.com/search/docs/crawling-indexing/robots-meta-tag?hl=pl). Boty pobierające, które mogą ignorować robots.txt, zatrzyma dopiero reguła w CDN albo na serwerze. Nowe boty pojawiają się kilka razy w roku. Nieaktualna lista jest gorsza niż żadna, bo wygląda jak decyzja. Wklej tu swój robots.txt, żeby zobaczyć, co może pobrać każdy bot AI. Działa w przeglądarce, na tym samym parserze, którego użyliśmy w badaniu; nic nie jest nigdzie wysyłane. ## Jak mierzyliśmy Wzięliśmy [listę Tranco Y8YYG](https://tranco-list.eu/list/Y8YYG/full) (utworzona 22 września 2026, okno 30 dni, pięć źródeł) i szliśmy od góry, zbierając pierwsze 1000 domen .pl, których serwer odpowiedział. Próba kontrolna to pierwsze 1000 domen z tej samej listy, przetworzone tak samo. 23 września nasz bot pobrał robots.txt i stronę główną każdej domeny, a także `/llms.txt`, `/llms-full.txt` i mapę strony. Przedstawiał się jako `OutofplaceResearchBot` z linkiem do strony opisującej badanie. Każdy robots.txt parsowaliśmy zgodnie z RFC 9309 i sprawdzaliśmy ścieżkę strony głównej `/` dla 38 user agentów: 36 botów AI od 17 operatorów, łącznie z wycofanymi nazwami, oraz Googlebota i Bingbota dla porównania. Wykresy pokazują nazwy, które operatorzy obecnie dokumentują. Bot jest zablokowany, jeśli reguła, która go dotyczy, zabrania `/`, niezależnie od tego, czy jest wymieniony z nazwy, czy spada do `*`. Odsetki mają 95-procentowe przedziały Wilsona, różnice między próbami – przedziały Newcombe'a. Próba to pierwsze 1000 domen .pl z listy Tranco, które odpowiedziały na nasze zapytania. 23 z nich zabrania naszemu botowi dostępu do strony głównej w robots.txt; ich robots.txt i tak przeczytaliśmy (z definicji jest publiczny), ale nic więcej nie pobieraliśmy. Domeny z błędem DNS, odrzuconym połączeniem albo błędem serwera pomijaliśmy i braliśmy następną z listy. Przy 1000 stron 95-procentowy przedział dla głównego wyniku wynosi od 18,1% do 23,1%. Każda strona dostaje jedną klasę, wygrywa pierwsze dopasowanie. Blokada całkowita: zablokowany jest też Googlebot, więc reguła nie dotyczy AI. Wszystkie boty AI: zablokowane są wszystkie boty treningowe i co najmniej jeden bot wyszukiwarki lub pobierający. Nieaktualna lista: zablokowana jest nazwa, której dostawca już nie dokumentuje, a bot, który ją zastąpił, nie. Tylko trening: zablokowane boty treningowe, żaden bot wyszukiwarki ani pobierający. Niespójna: zablokowany bot wyszukiwarki lub pobierający, a przepuszczony bot treningowy. Otwarta: żaden bot AI nie jest zablokowany. Reguły i ich testy są w repozytorium. robots.txt to tylko to, o co strona prosi. Część stron blokuje boty AI na firewallu lub w CDN, a część każdemu botowi pokazuje wyzwanie; wyzwania liczyliśmy osobno i nigdy jako „wpuszczone”. Sprawdzaliśmy tylko ścieżkę strony głównej: strona, która blokuje botom artykuły, ale nie `/`, liczy się jako otwarta. Domena .pl to nie to samo co polska firma, a listy popularności faworyzują duże serwisy. Liczby opisują najczęściej odwiedzaną część polskiego internetu w jednym dniu. ## Dalsza lektura - [RFC 9309: Robots Exclusion Protocol](https://www.rfc-editor.org/rfc/rfc9309.html): Standard robots.txt: grupy, dopasowanie i co bot musi zrobić, gdy pliku nie ma albo jest niedostępny. - [OpenAI: Overview of OpenAI crawlers](https://developers.openai.com/api/docs/bots): GPTBot, OAI-SearchBot i ChatGPT-User oraz do czego służy każdy z nich. - [Anthropic: jak zablokować boty](https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler): ClaudeBot, Claude-SearchBot i Claude-User. - [Google: Common crawlers, w tym Google-Extended](https://developers.google.com/crawling/docs/crawlers-fetchers/google-common-crawlers?hl=pl): Co kontroluje Google-Extended, a czego nie. - [Cloudflare: Managed robots.txt](https://developers.cloudflare.com/bots/additional-configurations/managed-robots-txt/): Co blokuje plik włączany jednym kliknięciem i jaką linię Content-Signal dodaje. - [llms.txt: propozycja i jej format](https://llmstxt.org): Oryginalna specyfikacja: gdzie leży plik i jak jest zbudowany. - [OutofplaceResearchBot: co pobiera i jak go zablokować](https://www.outofplace.space/pl/bot): Kto uruchamia nasz bot, jak się przedstawia i jak uprzejmie pobiera strony. --- # Next.js prefetch zmierzony: Partial Prefetching w 16.3 (PL) Source: https://www.outofplace.space/pl/blog/nextjs-prefetch-benchmark · Published 2026-09-23 Category: Wydajność · Tags: Next.js, Prefetching, React > Partial Prefetching w Next.js 16.3 zmniejszył w naszym benchmarku prefetch nawet o 93%, ale treść adresu przychodzi po kliknięciu. Kiedy co się opłaca. [Next.js 16.3](https://nextjs.org/blog/next-16-3) zmienia to, czym jest Next.js prefetch. Z włączonym [`partialPrefetching`](https://nextjs.org/docs/app/api-reference/config/next-config-js/partialPrefetching) App Router nie pobiera już payloadu dla każdego linku na ekranie: pobiera jeden wspólny szkielet na trasę, a treść zależną od adresu zostawia na po kliknięciu. Zbudowaliśmy jedną aplikację na sześć sposobów i przeklikaliśmy ją 3744 razy, żeby sprawdzić, co ta zamiana daje i ile kosztuje. Na stronie dokumentacji ze 100 linkami w bocznym menu prefetch zmalał o 93%. Ceną jest czas: treść, która wcześniej była na ekranie niemal od razu, przychodzi teraz kilkaset milisekund po kliknięciu, za szkieletem strony. Która strona tej zamiany jest lepsza, zależy od strony i od [jednego szczegółu Reacta, który łatwo przeoczyć](#szkielet-ktory-zobaczysz-300-ms). - Partial Prefetching zmniejszył prefetch o 78–93% na stronach z wieloma linkami, a całą nawigację w dokumentacji mniej więcej jedenastokrotnie. - Szkielet pojawia się od razu. Treść zależna od adresu przychodzi po kliknięciu: 345 ms w dokumentacji wobec 52 ms przy samym Cache Components, na zwolnionym telefonie. - Szkielet, który się pojawi, zostaje na około 300 ms, bo React tyle wstrzymuje każde odsłonięcie. - Linki z `prefetch={true}` zostają natychmiastowe w każdej konfiguracji. - `prefetchInlining` zostaw domyślny; HTTP/2 nie zmienił niczego ważnego. ## Co zmienia Partial Prefetching Do wersji 16.3 [App Router](https://nextjs.org/docs/app) robił prefetch dla każdego linku: gdy na ekran wjeżdżało N linków, pobierał mniej więcej N payloadów. Z `partialPrefetching` (wymaga [`cacheComponents`](https://nextjs.org/docs/app/api-reference/config/next-config-js/cacheComponents)) pobiera jeden [App Shell](https://nextjs.org/docs/app/guides/adopting-partial-prefetching) na trasę: wszystko, co strona renderuje niezależnie od adresu. Linki do tej samej trasy dzielą ten jeden prefetch. Treść, która czyta `params` albo `searchParams`, rozwiązuje się po kliknięciu, chyba że link poprosi o więcej przez `prefetch={true}`. Porównaliśmy sześć buildów tej samej aplikacji: - **L0**, bez Cache Components: klasyczny App Router. - **C0**, samo Cache Components. - **P0**, Cache Components z Partial Prefetching i domyślnym inliningiem prefetchu. - **P1–P3**, to samo z wyłączonym inliningiem, z małymi progami i z dużymi. Każdy przeklikaliśmy w trzech nawigacjach: z katalogu 48 produktów na stronę produktu, ze strony dokumentacji ze 100 linkami w menu na inną stronę dokumentacji i z wyników wyszukiwania na inne wyszukiwanie przez jeden z ośmiu linków z `prefetch={true}`. ## Bajty prefetchu: nawet o 93% mniej W dokumentacji samo Cache Components pobierało z wyprzedzeniem każdy widoczny link z menu. Partial Prefetching zmienił to z 402 kB → 29 kB (−93%), w 8,5 żądaniach zamiast 49. W katalogu oszczędność wyniosła 78%. Wyszukiwanie prawie się nie zmieniło, bo jego osiem linków z `prefetch={true}` nadal pobiera własne wyniki. **Partial Prefetching pobiera ułamek bajtów na stronach z wieloma linkami** (Skompresowane bajty RSC pobrane przed kliknięciem · komputer, bez ograniczeń, późne kliknięcie · mediany z 20 przebiegów) | Pozycja | Katalog → produkt | Dokumentacja → dokumentacja | Wyszukiwanie → wyszukiwanie | | --- | --- | --- | --- | | L0 · Klasyczny router | 30 kB | 408 kB | 98 kB | | C0 · Cache Components | 160 kB | 402 kB | 163 kB | | P0 · Partial Prefetching | 34 kB | 29 kB | 107 kB | | P1 · Bez inliningu | 33 kB | 32 kB | 105 kB | | P2 · Mały inlining | 32 kB | 28 kB | 103 kB | | P3 · Duży inlining | 34 kB | 33 kB | 105 kB | Source: Benchmark Outofplace, Next.js 16.3.6. Chart: https://www.outofplace.space/data/blog/nextjs-prefetch-benchmark/charts/pl/chart-partial-prefetching-pobiera-ulamek-bajtow-na-stronach-z-wieloma.png Data: https://www.outofplace.space/data/blog/nextjs-partial-prefetching-benchmark/2026-09-23/summary_by_variant.csv Jeden wiersz zasługuje na drugie spojrzenie. Włączenie Cache Components bez Partial Prefetching sprawiło, że katalog pobierał pięć razy więcej niż klasyczny router, 30 kB → 160 kB (+426%): prerenderowane szkielety produktów da się wtedy pobrać osobno dla każdego linku. Jeśli włączasz Cache Components na stronach z wieloma linkami, od razu zdecyduj o Partial Prefetching. ## Czy nawigacja jest naprawdę natychmiastowa? Szkielet tak. Przy Partial Prefetching pojawiał się 43 ms–45 ms po kliknięciu na zwolnionym telefonie. Z tym, po co ktoś kliknął, czyli produktem albo artykułem, jest inaczej: potrzebuje żądania po kliknięciu. **Samo Cache Components pokazuje treść od razu, Partial Prefetching po jednym żądaniu** (Od kliknięcia do pierwszej klatki z treścią zależną od adresu · telefon, ograniczenie Lighthouse mobile, późne kliknięcie · ms) | Pozycja | L0 · Klasyczny router | C0 · Cache Components | P0 · Partial Prefetching | | --- | --- | --- | --- | | Katalog → produkt | 375 ms | 60 ms | 262 ms | | Dokumentacja → dokumentacja | 47 ms | 52 ms | 345 ms | | Wyszukiwanie → wyszukiwanie | 40 ms | 41 ms | 43 ms | Source: Benchmark Outofplace, Next.js 16.3.6. Chart: https://www.outofplace.space/data/blog/nextjs-prefetch-benchmark/charts/pl/chart-samo-cache-components-pokazuje-tresc-od-razu-partial-prefetching.png Data: https://www.outofplace.space/data/blog/nextjs-partial-prefetching-benchmark/2026-09-23/summary_by_variant.csv W dokumentacji samo Cache Components pokazało nowy artykuł po 52 ms, Partial Prefetching po 345 ms. Na stronie produktu różnica była mniejsza, 60 ms → 262 ms (+337%), a najwolniejszy był klasyczny router, bo dla dynamicznej trasy nie miał nic wyrenderowanego z góry. Wyszukiwanie, gdzie każdy link ma `prefetch={true}`, pokazywało treść po 40 ms–43 ms w każdej konfiguracji. ### Szkielet, który zobaczysz: 300 ms Na stronie produktu trzy etapy przychodzą po kolei: szkielet, produkt, a potem część renderowana w chwili żądania. **Przy Partial Prefetching strona przychodzi w trzech krokach** (Katalog → produkt · telefon, zwolniony, późne kliknięcie · od kliknięcia do każdego etapu na ekranie · ms) | Pozycja | Szkielet | Treść | Część z serwera | | --- | --- | --- | --- | | L0 · Klasyczny router | 375 ms | 375 ms | 375 ms | | C0 · Cache Components | 60 ms | 60 ms | 342 ms | | P0 · Partial Prefetching | 47 ms | 262 ms | 562 ms | Source: Benchmark Outofplace, Next.js 16.3.6. Chart: https://www.outofplace.space/data/blog/nextjs-prefetch-benchmark/charts/pl/chart-przy-partial-prefetching-strona-przychodzi-w-trzech-krokach.png Data: https://www.outofplace.space/data/blog/nextjs-partial-prefetching-benchmark/2026-09-23/summary_by_variant.csv W dokumentacji szkielet był widoczny przez 298 ms w każdym buildzie z Partial Prefetching. To nie jest czas sieci: odpowiedź dochodziła po 250 ms. React dołączony do Next.js 16.3.6 (`19.3.0-canary`) [czeka 300 ms](https://github.com/react/react/blob/d083ec1da1e5252abd3ddfdde6dfbc09701a2c51/packages/react-reconciler/src/ReactFiberWorkLoop.js#L528) od ostatniej zamiany między fallbackiem [Suspense](https://react.dev/reference/react/Suspense) a treścią, zanim zatwierdzi kolejną, żeby stan ładowania nie mignął. Efekt uboczny: etapy się łańcuchują. Szkielet, który się pojawi, zostaje na ekranie około 300 ms, nawet gdy dane już są, a osobna granica dla danych z serwera odsłania się 300 ms po poprzedniej. Projektuj szkielety tak, żeby dało się na nie patrzeć, i rozważ jedną granicę tam, gdzie dwie tylko by się łańcuchowały. ### Wczesne kliknięcie Kliknięcie 50 ms po hydratacji, zanim skończy się jakikolwiek prefetch, zabiera przewagę wszystkim, którzy pobierają z wyprzedzeniem. Samo Cache Components jest wtedy mniej więcej tak samo wolne jak Partial Prefetching, bo obie konfiguracje najpierw pokazują fallback, a potem odczekują 300 ms. **Przy wczesnym kliknięciu żadna konfiguracja niczego jeszcze nie pobrała** (Kliknięcie 50 ms po hydratacji · telefon, zwolniony · od kliknięcia do treści · ms) | Pozycja | L0 · Klasyczny router | C0 · Cache Components | P0 · Partial Prefetching | | --- | --- | --- | --- | | Katalog → produkt | 367 ms | 374 ms | 389 ms | | Dokumentacja → dokumentacja | 262 ms | 528 ms | 529 ms | | Wyszukiwanie → wyszukiwanie | 234 ms | 404 ms | 339 ms | Source: Benchmark Outofplace, Next.js 16.3.6. Chart: https://www.outofplace.space/data/blog/nextjs-prefetch-benchmark/charts/pl/chart-przy-wczesnym-kliknieciu-zadna-konfiguracja-niczego-jeszcze-nie.png Data: https://www.outofplace.space/data/blog/nextjs-partial-prefetching-benchmark/2026-09-23/summary_by_variant.csv Ostatni etap, czyli część z serwera, dzieli te kliknięcia na dwie grupy. Przy Partial Prefetching w 6 z 12 wczesnych kliknięć na stronie produktu cała strona była na ekranie po 375 ms–400 ms. W pozostałych 6 dopiero po 650 ms–700 ms, a pomiędzy nie trafił żaden przebieg. Ta przerwa to mniej więcej 300 ms, które React odczekuje między dwoma odsłonięciami: ostatni etap pojawia się razem z produktem albo jedno odczekanie później. **Ostatni etap przychodzi razem z produktem albo mniej więcej 300 ms później** (Katalog → produkt, wczesne kliknięcie · Partial Prefetching · telefon, zwolniony, przebieg HTTP/2 · od kliknięcia do ostatniego etapu na ekranie · ms) | Zakres | przebiegi | | --- | --- | | 375 ms–400 ms | 6 | | 400 ms–425 ms | 0 | | 425 ms–450 ms | 0 | | 450 ms–475 ms | 0 | | 475 ms–500 ms | 0 | | 500 ms–525 ms | 0 | | 525 ms–550 ms | 0 | | 550 ms–575 ms | 0 | | 575 ms–600 ms | 0 | | 600 ms–625 ms | 0 | | 625 ms–650 ms | 0 | | 650 ms–675 ms | 3 | | 675 ms–700 ms | 3 | Source: Benchmark Outofplace, Next.js 16.3.6. Chart: https://www.outofplace.space/data/blog/nextjs-prefetch-benchmark/charts/pl/chart-ostatni-etap-przychodzi-razem-z-produktem-albo-mniej-wiecej-300.png Data: https://www.outofplace.space/data/blog/nextjs-partial-prefetching-benchmark/2026-09-23/runs.csv ## Cały rachunek Partial Prefetching przenosi bajty sprzed kliknięcia na po kliknięciu, a na stronie produktu dodatkowo je podwaja: odpowiedź nawigacji niosła dwie pełne kopie payloadu produktu, 13 kB → 26 kB (+99%) po kompresji. Licząc wszystko, co pobiera nawigacja, razem z prefetchem, w dokumentacji i tak wychodzi daleko do przodu. **Cała nawigacja w dokumentacji kosztuje mniej więcej jedną dziesiątą bajtów** (Wszystkie skompresowane bajty RSC jednej nawigacji w dokumentacji, prefetch i nawigacja · telefon, zwolniony, późne kliknięcie) | Pozycja | Wartość | n | | --- | --- | --- | | L0 · Klasyczny router | 463 kB | 20 | | C0 · Cache Components | 458 kB | 20 | | P0 · Partial Prefetching | 40 kB | 20 | | P1 · Bez inliningu | 43 kB | 20 | | P2 · Mały inlining | 39 kB | 20 | | P3 · Duży inlining | 44 kB | 20 | Source: Benchmark Outofplace, Next.js 16.3.6. Chart: https://www.outofplace.space/data/blog/nextjs-prefetch-benchmark/charts/pl/chart-cala-nawigacja-w-dokumentacji-kosztuje-mniej-wiecej-jedna-dziesi.png Data: https://www.outofplace.space/data/blog/nextjs-partial-prefetching-benchmark/2026-09-23/summary_by_variant.csv Mniej więcej 11 razy mniej bajtów w dokumentacji i 2,2 raza mniej w katalogu. Hydratacja i opóźnienie reakcji na stronie startowej nie zmieniły się w żadnej konfiguracji. ## Kiedy `prefetch={true}` się opłaca [` `](https://nextjs.org/docs/app/api-reference/components/link#prefetch) pobiera też dane z adresu strony docelowej, więc kliknięcie jest tak samo natychmiastowe jak przy pełnym prefetchu. Płacisz za każdy widoczny link: osiem linków wyszukiwania kosztowało 107 kB, więcej niż cała strona dokumentacji przy Partial Prefetching. Używaj go tam, gdzie kliknięcie jest prawdopodobne, a dane da się cache'ować: podpowiedzi wyszukiwania, link „dalej”, kilka pierwszych wyników. Na siatkach kart rób prefetch na zamiar (najechanie albo dotknięcie), zamiast oznaczać każdą kartę. ## prefetchInlining: zostaw włączony [`experimental.prefetchInlining`](https://nextjs.org/docs/app/api-reference/config/next-config-js/prefetchInlining) łączy małe odpowiedzi prefetchu w jedną. Ustawienie domyślne i nasze dwa zestawy progów różniły się najwyżej jednym żądaniem i kilkoma kilobajtami, bez mierzalnej różnicy w czasie. Wyłączenie inliningu dodało pięć albo sześć żądań w katalogu i wyszukiwaniu, a przez HTTP/1.1 kosztowało to wczesne kliknięcie w katalogu: sześć połączeń na host się zapycha. Przez [HTTP/2](https://www.rfc-editor.org/rfc/rfc9113.html) kara znikała. Wszystkie inne konfiguracje wypadły na obu protokołach tak samo. **HTTP/2 ma znaczenie tylko przy wyłączonym inliningu** (Katalog → produkt, wczesne kliknięcie · telefon, zwolniony · od kliknięcia do treści · HTTP/1.1 i HTTP/2 na przemian, po 12 przebiegów · ms) | Pozycja | HTTP/1.1 | HTTP/2 | | --- | --- | --- | | L0 · Klasyczny router | 365 ms | 363 ms | | C0 · Cache Components | 375 ms | 373 ms | | P0 · Partial Prefetching | 387 ms | 386 ms | | P1 · Bez inliningu | 440 ms | 372 ms | | P2 · Mały inlining | 390 ms | 383 ms | | P3 · Duży inlining | 387 ms | 386 ms | Source: Benchmark Outofplace, Next.js 16.3.6. Chart: https://www.outofplace.space/data/blog/nextjs-prefetch-benchmark/charts/pl/chart-http-2-ma-znaczenie-tylko-przy-wylaczonym-inliningu.png Data: https://www.outofplace.space/data/blog/nextjs-partial-prefetching-benchmark/2026-09-23/summary_by_variant.csv ## Którą konfigurację wybrać ```ts title="next.config.ts" const nextConfig: NextConfig = { cacheComponents: true, // Jeden App Shell na trasę zamiast prefetchu na każdy link. partialPrefetching: true, }; ``` ```tsx title="Linki, przy których czekanie by bolało" {suggestion} ``` Najwięcej zyskują strony z wieloma linkami do tej samej dynamicznej trasy: dokumentacja, siatki, feedy. Strona z kilkoma linkami do różnych stron zyskuje niewiele. Tam, gdzie natychmiastowa treść jest ważniejsza niż bajty, a strony są statyczne, zostaw pełny prefetch; gdzie indziej włącz Partial Prefetching. [Konfiguracja segmentu `export const prefetch = 'partial'`](https://nextjs.org/docs/app/api-reference/file-conventions/route-segment-config/prefetch) pozwala przechodzić trasa po trasie. Dodaj `prefetch={true}` tam, gdzie następne kliknięcie da się przewidzieć, a dane da się cache'ować. Na wolnym łączu będzie widoczny przez około 300 ms. Tytuł i nawigację umieść w części niezależnej od adresu: pojawia się niemal od razu. Ta strona działa na Partial Prefetching: większość naszych stron ma mało linków, a te, które mają ich dużo (blog, case studies), są statyczne i lekkie. Nasze własne produkty, Portivo, Sprawna Matura i Levera, też stoją na Next.js. ## Jak mierzyliśmy Wygenerowaną aplikację App Router (katalog 48 produktów, 100 stron dokumentacji z menu w layoucie, strona wyszukiwania) zbudowaliśmy sześć razy na Next.js 16.3.6, zmieniając tylko `cacheComponents`, `partialPrefetching` i `experimental.prefetchInlining`, i serwowaliśmy przez `next start`. Headless Chrome otwierał stronę startową w świeżym kontekście przeglądarki, czekał i klikał link, zapisując każde żądanie przez [protokół DevTools](https://chromedevtools.github.io/devtools-protocol/) i mierząc znaczniki od zdarzenia kliknięcia do następnej klatki animacji. Każda kombinacja konfiguracji, scenariusza, telefonu albo komputera, szybkiej albo [zwolnionej sieci (150 ms, 1,6 Mb/s, CPU 4×)](https://github.com/GoogleChrome/lighthouse/blob/main/docs/throttling.md) i wczesnego albo późnego kliknięcia przeszła 20 razy w losowej kolejności, razem 2880 przebiegów; test HTTP/2 dodał 864 kolejne. Mediany mają 95-procentowe przedziały bootstrap ze stałym ziarnem. To laboratorium: jedna maszyna, localhost, bez CDN, bez obrazów i fontów, wygenerowany tekst i sztuczne 300 ms opóźnienia po stronie serwera. Ograniczanie działa per żądanie, nie na poziomie pakietów, więc czasy bezwzględne są optymistyczne; liczą się porównania. Późne kliknięcie to najlepszy przypadek dla prefetchu, a wczesne to jeden konkretny wyścig; prawdziwi użytkownicy klikają wszędzie pomiędzy. Mechanika App Shella, podwojony payload produktu i 300 ms odsłonięcia dotyczą Next.js 16.3.6 z jego wersją canary Reacta i mogą się zmienić w wydaniu poprawkowym. Wspólny szkielet jest pobierany przez adres tego linku, który router zaplanuje jako pierwszy, a ta odpowiedź niesie całą stronę dla tego adresu. Kliknięcie dokładnie w ten link nie potrzebuje żadnego żądania. Cele kliknięć wybraliśmy niżej na stronie, żeby to nie podbijało wyników Partial Prefetching. ## Dalsza lektura - [Next.js: Adopting Partial Prefetching](https://nextjs.org/docs/app/guides/adopting-partial-prefetching): Jak działa App Shell i jak przenosić trasy jedna po drugiej. - [Next.js: Optimizing prefetching](https://nextjs.org/docs/app/guides/optimizing-prefetching): Prefetch per link przez prefetch={true} i kiedy go używać. - [Next.js: prefetchInlining](https://nextjs.org/docs/app/api-reference/config/next-config-js/prefetchInlining): Co łączy inlining i dlaczego domyślne ustawienie pasuje większości aplikacji.