---
title: "nosnippet on the Polish web: almost nobody limits Google"
description: "7 of 7,785 .pl homepages limit Google's snippet, while a quarter ask for bigger previews. None of the 1,015 that block AI crawlers in robots.txt use nosnippet."
canonical_url: "https://www.outofplace.space/blog/snippet-controls-polish-web"
language: "en"
last_updated: "2026-10-10"
alternates:
  pl: "https://www.outofplace.space/pl/blog/nosnippet-na-polskich-stronach.md"
---

# nosnippet on the Polish web: almost nobody limits Google

> 7 of 7,785 .pl homepages limit Google's snippet, while a quarter ask for bigger previews. None of the 1,015 that block AI crawlers in robots.txt use nosnippet.

Published 2026-10-10 · Category: Search & AI · Tags: Meta robots, AI Overviews, AI crawlers, Poland

Almost no Polish site uses `nosnippet` or any other rule that limits what Google and AI answers may quote. We read the homepages of the 7,785 .pl sites in [the Tranco list of popular sites](https://tranco-list.eu/list/Y8YYG/full) that answered with HTML, and only 7 limit their text snippet. Far more ask for the opposite: 25% tell Google it may show large image previews, snippets of any length or both, mostly because WordPress or an SEO plugin writes that by default.

This is the on-page side of the question we asked in [our study of AI crawlers on the Polish web](https://www.outofplace.space/blog/ai-crawlers-polish-web). robots.txt decides which crawlers may fetch a page; snippet rules decide what a search engine may show of a page it already has, and Google says `nosnippet` also keeps a page's text out of AI Overviews and AI Mode. The sites that block AI crawlers in robots.txt don't use it: of the 1,015 homepages whose robots.txt blocks an AI training crawler, 0 limit their snippet.

- Only 7 of 7,785 .pl homepages limit Google's text snippet with `nosnippet` or `max-snippet`; 19 limit any preview.
- 25% ask Google for more: large image previews, unlimited snippets or both. On Yoast SEO sites the share is 96%.
- `noai` appears in no meta tag at all. All 22 `noai` headers come from shops on one platform, AtomStore.
- None of the 1,015 sites that block an AI training crawler in robots.txt limit their snippet; the AI controls they do use on the page are Bing's `noarchive` and `nocache`, `data-nosnippet` and `noai`.

## What nosnippet, max-snippet and max-image-preview do

Search engines read a page's robots rules from a ` ` tag in the HTML or from an `X-Robots-Tag` header in the server's response. A tag named `googlebot` or `bingbot` addresses one crawler only. Most rules people know decide whether a page is indexed (`noindex`) or its links followed (`nofollow`). A second group decides how much of an indexed page a search engine may show. [Google's robots meta tag documentation](https://developers.google.com/search/docs/crawling-indexing/robots-meta-tag) (last updated 24 March 2026) lists them:

- `nosnippet`: no text snippet or video preview. It applies "to all forms of search results", AI Overviews and AI Mode included, and "will also prevent the content from being used as a direct input for AI Overviews and AI Mode".
- `max-snippet:[number]`: at most that many characters of text; `0` is the same as `nosnippet`, `-1` lets Google choose.
- `max-image-preview:none|standard|large`: no image preview, a default-sized one, or one as wide as the screen.
- `max-video-preview:[number]`: at most that many seconds of video preview.
- `data-nosnippet`: an HTML attribute that keeps one part of a page out of snippets.

When rules conflict, Google applies the more restrictive one. Bing gives two older rules a meaning for its AI answers, and a third family, `noai` and `noimageai`, has no standard behind it. We read all of them on every homepage.

## A quarter of Polish homepages ask Google for more

Just over half the homepages carry a robots meta tag at all (54%), and many of those say only `index, follow`, which is what search engines assume anyway. Of the meta tag and header rules that change what Google shows, the common ones all widen it. Each of those rules that would narrow it is on fewer than one homepage in fifty; the `data-nosnippet` attribute, which works inside the page, comes later.

**Polish homepages ask Google for bigger previews, almost never smaller ones** (Share of .pl homepages stating each robots rule in a meta tag or the X-Robots-Tag header · rules that widen what Google shows in cool tones, rules that limit it in warm ones · n = 7,785 · 95% intervals)

| Item | Value | 95% interval | n |
| --- | --- | --- | --- |
| max-image-preview:large | 25% | 24%–26% | 7,785 |
| max-snippet:-1 | 17% | 17%–18% | 7,785 |
| max-video-preview:-1 | 17% | 16%–18% | 7,785 |
| noarchive | 1.6% | 1.3%–1.9% | 7,785 |
| noindex | 1.1% | 0.9%–1.4% | 7,785 |
| noai | 0.3% | 0.2%–0.4% | 7,785 |
| nocache | 0.2% | 0.1%–0.4% | 7,785 |
| max-image-preview:standard | 0.1% | 0.1%–0.3% | 7,785 |
| max-snippet:n | 0.1% | 0%–0.1% | 7,785 |
| nosnippet | 0% | 0%–0.1% | 7,785 |
| max-snippet:0 | 0% | 0%–0.1% | 7,785 |
| max-image-preview:none | 0% | 0%–0.1% | 7,785 |
| noimageai | 0% | 0%–0.1% | 7,785 |

Source: Outofplace census, 24 Sept 2026; Tranco list Y8YYG. Chart: https://www.outofplace.space/data/blog/snippet-controls-polish-web/charts/en/chart-polish-homepages-ask-google-for-bigger-previews-almost-never-sma.png Data: https://www.outofplace.space/data/blog/snippet-controls-polish-web/2026-09-24/rules.csv

Put together, 19 homepages limit any preview, text, image or video, and 7 limit the text, with `nosnippet` or a `max-snippet` length. Another 84 homepages carry `noindex` that Google would apply and keep out of search results altogether. These counts are the rules Google would follow; the chart and the table at the end count every mention, including tags addressed to other crawlers.

## Where max-image-preview:large comes from

`max-image-preview:large` is a default of the most common software. [WordPress 5.7](https://make.wordpress.org/core/2021/02/19/robots-api-and-max-image-preview-directive-in-wordpress-5-7/) (March 2021) adds it to every site that allows search engines, and [Yoast SEO's specification](https://developer.yoast.com/features/seo-tags/meta-robots/functional-specification/) says the plugin writes `max-snippet:-1, max-image-preview:large, max-video-preview:-1` on each public page by default. The census shows the defaults at work: the share asking for large previews rises and falls with the software behind the page.

**Large image previews follow the software behind the site** (What .pl homepages tell Google about image previews, by the SEO plugin or CMS behind the page · Yoast n = 907, Rank Math n = 239, All in One SEO n = 100, other WordPress n = 591, not WordPress n = 5,948)

| Item | Limited (standard or none) | No rule (default size) | Large previews asked for | n |
| --- | --- | --- | --- | --- |
| Yoast SEO | 0.1% | 4% | 96% | 907 |
| Rank Math | 0% | 5.4% | 95% | 239 |
| All in One SEO | 0% | 10% | 90% | 100 |
| Other WordPress | 0.5% | 33% | 66% | 591 |
| Not WordPress | 0.1% | 94% | 6% | 5,948 |

Source: Outofplace census, 24 Sept 2026. Chart: https://www.outofplace.space/data/blog/snippet-controls-polish-web/charts/en/chart-large-image-previews-follow-the-software-behind-the-site.png Data: https://www.outofplace.space/data/blog/snippet-controls-polish-web/2026-09-24/software.csv

Nine in ten or more homepages with Yoast SEO, Rank Math or All in One SEO ask for large previews. Other WordPress sites ask less often, 66%, perhaps because some switch the core default off or run a theme that prints its own tag; outside WordPress, 6% do. Of all the homepages asking for large previews, 61% run an SEO plugin we recognise.

That is a reasonable default for a site that wants its pictures seen. It does mean that most of the requests for bigger previews match what the software writes by default, and that the opposite, limiting what is shown, is something almost nobody asks for.

## noai: one shop platform's header

`noai` and `noimageai` are unofficial rules that ask AI services not to use a page's text or images. They are not among the rules in Google's documentation. In the census they appear in no meta tag at all. All of them are `X-Robots-Tag: noai` headers, and all come from shops on AtomStore: 22 of the 36 AtomStore shops in our sample send it, which points to a platform setting rather than separate decisions.

## noarchive and nocache: Bing's switches for AI answers

In September 2023 Bing announced [new options to control the use of content in Bing Chat](https://blogs.bing.com/webmaster/september-2023/Announcing-new-options-for-webmasters-to-control-usage-of-their-content-in-Bing-Chat), under which two old rules now govern its AI: content tagged `noarchive` "will not be included in Bing Chat answers" nor used to train Microsoft's models, and with `nocache` Bing shows only the URL, title and snippet. Two months later Microsoft [renamed Bing Chat to Copilot at Ignite 2023](https://blogs.microsoft.com/blog/2023/11/15/microsoft-ignite-2023-ai-transformation-and-the-technology-driving-change/). Google Search, for its part, [no longer uses `noarchive`](https://developers.google.com/search/docs/crawling-indexing/robots-meta-tag#noarchive), since it dropped cached links.

101 Polish homepages give Bing `noarchive` and 17 `nocache`, 1.5% together. These are the rules Bing would apply, with a page carrying both counted as `nocache`, so they differ from the table's count of every mention. Whether they meant to keep out of Copilot or kept a rule from the days of Google's cached pages, the tag can't say.

## What does data-nosnippet do?

`data-nosnippet` hides one part of a page from snippets while the rest stays quotable: a disclaimer, a price that changes, a block of boilerplate. Google reads it only on `span`, `div` and `section` elements. 4.6% of Polish homepages use it, more than all the other limiting rules together, but mostly not on content.

Of the homepages with `data-nosnippet`, 64% put it on a cookie banner. Most of those banners come from two WordPress plugins in [our study of cookie consent tools](https://www.outofplace.space/blog/cookie-consent-polish-web): Complianz, on 64% of them, and Cookie Law Info, on 28%. They mark their banners so that "We use cookies" never turns up as the page's description in Google.

On content the attribute appears on 1.7% of homepages, and 9% of the homepages using it put it on an element other than `span`, `div` or `section`, where Google doesn't read it.

## robots.txt says no to AI crawlers, the page says nothing

robots.txt and snippet rules do different jobs. Blocking GPTBot or ClaudeBot in robots.txt stops those crawlers fetching the site. It does nothing about Google's AI Overviews, which are built from Googlebot's index, and Google's [list of its crawlers](https://developers.google.com/crawling/docs/crawlers-fetchers/google-common-crawlers) says Google-Extended, the token for Gemini, "does not impact a site's inclusion in Google Search" (read on 9 October 2026). For AI Overviews and AI Mode, a site's levers are Googlebot itself and the snippet rules.

We ran the same rules as our AI crawler study over the census: 13% of the homepages' robots.txt files block at least one AI training crawler, and 228 block Google-Extended. None of them limits its snippet. Snippet limits are so rare that this zero sets the blockers apart from no one: among the other homepages only 7 limit theirs. What it does show is that blocking AI crawlers in robots.txt doesn't come with a snippet rule on the page.

**Sites that block AI crawlers in robots.txt don't limit their snippets either** (Share of .pl homepages with each on-page control, by whether their robots.txt blocks an AI training crawler · blocks one n = 1,015, blocks none n = 6,770)

| Item | robots.txt blocks an AI training crawler | robots.txt blocks none |
| --- | --- | --- |
| Snippet limit (nosnippet, max-snippet) | 0% | 0.1% |
| noarchive or nocache (Bing) | 3.5% | 1.2% |
| data-nosnippet on content | 4.6% | 1.3% |
| noai or noimageai | 1.8% | 0.1% |
| noindex | 1.6% | 1% |
| Any of the AI controls | 7.5% | 2.7% |

Source: Outofplace census, 24 Sept 2026. Chart: https://www.outofplace.space/data/blog/snippet-controls-polish-web/charts/en/chart-sites-that-block-ai-crawlers-in-robots-txt-don-t-limit-their-sni.png Data: https://www.outofplace.space/data/blog/snippet-controls-polish-web/2026-09-24/robots_txt.csv

The sites that block AI crawlers do use AI controls on the page more often than the rest, 7.5% against 2.7%, but the extra comes from `data-nosnippet`, Bing's `noarchive` and `nocache`, and AtomStore's `noai`, which 18 of the shops sending it pair with an AI block in robots.txt. The one rule Google names for AI Overviews is missing from all of them.

## How to limit what Google and AI answers quote

A snippet rule is a trade. `nosnippet` keeps a page's text out of AI Overviews, and out of the snippet under its blue link in ordinary results too, which usually costs clicks. Decide per page, not per site.

Open the page's source and search for `name="robots"`, then look at the response headers in the browser's network panel for `X-Robots-Tag`. If you run WordPress with Yoast or Rank Math, you will most likely find `max-image-preview:large, max-snippet:-1`: you have asked for more, not less.

For a page whose text you don't want quoted, set a rule in its head. `max-snippet` keeps a short snippet; `nosnippet` removes it:

```html title="head"

<!-- or, to keep the text out of snippets and AI Overviews entirely -->

```

Wrap a paragraph or block in a `div`, `span` or `section` with the attribute; other elements are ignored:

```html
 Prices change daily; check the product page.
```

PDFs and images have no HTML head, so their rules go in a header. In nginx:

```nginx title="nginx.conf"
location ~* \.pdf$ {
  add_header X-Robots-Tag "nosnippet" always;
}
```

In Apache, `Header set X-Robots-Tag "nosnippet"` inside a ` ` block does the same.

`noarchive` keeps a page out of Bing's AI answers; `nocache` lets Bing show only its URL, title and snippet there. Neither affects Google.

A crawler blocked in robots.txt never reads the page's own rules. To keep a page's text out of AI Overviews, let Googlebot fetch it and tell it `nosnippet`.

## How we measured

The homepages come from our census of 24 September 2026 by `OutofplaceResearchBot` (see [what our crawler fetches and how to block it](https://www.outofplace.space/bot)), which reads robots.txt first and fetches only the homepage: 8,023 of the 9,208 .pl domains in the Tranco list answered with HTML, and 7,785 distinct sites remained after merging domains that redirect to the same address.

From each homepage we read every robots meta tag (`robots` and tags naming a crawler, such as `googlebot` or `bingbot`) and the `X-Robots-Tag` header of the final response, without running any script. We combined them the way Google's documentation describes: the tags and header rules addressed to all crawlers or to Googlebot, the more restrictive rule winning. Bing's view takes the tags and header rules for all crawlers, Bingbot or msnbot. `noai` and `noimageai` count wherever they appear. For `data-nosnippet` we found every element carrying it and filed it as a cookie banner when the element, or one it sits in, is named for cookies or consent.

robots.txt comes from the same census and is read with the parser and crawler list of our AI crawler study: a homepage's robots.txt blocks a crawler when it disallows `/` for that crawler, by name or through `*`. "An AI training crawler" is any of the current training crawlers on that list.

We saw the homepage only, as the server sent it. Article and product pages may carry other rules (publishers often set them per article), and a rule added by JavaScript or sent only for other paths is invisible. A homepage longer than 1 MiB was cut there (331 of them). We read what the sites ask; whether Google, Bing and AI services do as asked is their documentation's claim, which we didn't test. A cookie banner whose markup doesn't name cookies or consent is counted as content.

## Further reading

- [Robots meta tag, data-nosnippet and X-Robots-Tag specifications](https://developers.google.com/search/docs/crawling-indexing/robots-meta-tag): Google's documentation of every snippet and indexing rule, including how nosnippet applies to AI Overviews and AI Mode.

- [Bing: new options to control usage of content in Bing Chat](https://blogs.bing.com/webmaster/september-2023/Announcing-new-options-for-webmasters-to-control-usage-of-their-content-in-Bing-Chat): Bing's announcement of what noarchive and nocache mean for its AI answers and model training, September 2023.

- [Google's common crawlers](https://developers.google.com/crawling/docs/crawlers-fetchers/google-common-crawlers): Google's list of its crawlers and user agent tokens, including what Google-Extended does and doesn't control.

- [AI crawlers vs the Polish web: robots.txt on 1,000 .pl sites](https://www.outofplace.space/blog/ai-crawlers-polish-web): Our study of robots.txt on the largest .pl sites: which AI crawlers they block, and how.
