Features

Power your results with AI, data, and teamwork. All you need for web operations in one platform

User Support

Your go-to support page for troubleshooting and getting the most out of MONJI+

Blog
Jul 27, 2026
WebOps

“Should we just block the AI bots?”: why we stopped answering yes-or-no

“Should we just block the AI bots?”

We’ve been hearing this question more and more from the site operators we work with.

For a long time, we could only answer it as a binary. Block everything and AI can’t find you; allow everything and you’re fed to training and automated agents alike. That’s how it looked.

This piece lays out how to set an AI-bot policy without deciding “all or nothing” — drawing the line by purpose, and recording it. It’s for the operators and agencies who set the AI-bot policy for their own sites and their clients’.

Why teams drift to “all or nothing” (the structure of the problem)

Here’s the answer up front.
The honest wish on the ground is “we want search to find us, but not to permit training or automated action” — yet the tool only offered a binary.

There’s a long-standing way to control bot traffic to a site: robots.txt. But what you can write there is, roughly, “come in” or “stay out.” AI bots have all kinds of goals — reading for search, acting for the task at hand, collecting for training — and splitting those with a single line was always a stretch.

So when the call gets hard, you drift toward blocking all or allowing all. That was the ceiling of the old approach.

Cloudflare split AI into three purposes

On July 1, 2026, Cloudflare brought in a way to split by purpose. It’s available to everyone, including the free plan.

You can now handle AI access in three buckets:

  • Search: collecting and organizing content to answer questions later.
  • Agent: moving to finish a task on the spot, on someone’s behalf.
  • Training: taking content into a model’s training. Once it’s in, it stays.

“Let search through, because we want to be found. But stop training and automated action.” That honest wish now drops straight into a setting. Not all-or-nothing, but a line drawn by purpose. That’s the big shift.

From September 15, the default itself changes

There’s a timing note, too.

From September 15, 2026, domains newly registered with Cloudflare get a different default. On pages that show ads, training and agents are blocked from the start, and search is let through. That becomes the do-nothing starting state.

Leave it alone and one policy still gets chosen for you. If you’re launching a new site, keep this date in mind.

You can attach a three-level “how you’d like to be used” in robots.txt

Drawing the line isn’t only allow-or-deny. Once you allow, you can also state how you’d like to be used. Cloudflare prepared a new signal to attach to robots.txt, expressing your preference in three levels:

  1. Respond in the moment is fine, but don’t store or reuse it.
  2. Indexing, quoting a part, and linking back is fine (this is the default).
  3. Summarizing, even reproducing the whole thing, is allowed.

Set automatically, the middle one — “index and quote, with a link back” — applies. Not allowing everything, not refusing everything, but attaching a preference for how it’s used.

Four steps to draw the line by purpose (what we do on the ground)

Nothing flashy. Here are the steps we actually follow.

Step 1: Decide what to allow per page type

Don’t treat every page the same. Here’s the kind of lines we draw:

  • Company and service pages, article indexes: we want these found, so we let search through.
  • Pricing and contact — the pages we most want people to reach: definitely keep these in search.
  • Purchase and account flows, where automated action on a person’s behalf would be a problem: here we’re wary of agents.
  • Original materials and photos we spent real time on: anything we don’t want swallowed whole into training, we lean toward blocking training.

Step 2: Record the policy and the reason where the team can see it

So it carries over when the owner changes, we keep not just the policy but why we chose it. If the reason isn’t kept, the next person just resets to block-all or allow-all again. To keep these operational decisions and the history of checks in one place, we built MONJI+.

Step 3: Add “AI-bot policy” to the pre-launch checklist for new domains

Since the default shifts for domains registered from September 15, 2026, make it a standing pre-launch check item.

Step 4: After a change, watch how behavior shifts

When you change a policy, watch whether site traffic or crawl behavior moved. Looking at the change and the numbers in the same flow makes the next call easier.

What actually pays off

None of this is flashy. But once we put this order in place, a few things changed.

  • Because the reason is kept, the policy carries over even when the owner changes.
  • The drift back to “just block/allow everything” dropped.
  • Even for new domains, the policy now gets checked as a line item before launch.

One setting alone doesn’t prove an outcome — search and AI ranking depend on many factors. Still, keeping the policy, the reason, and the numbers in one place made the next decision faster.

A note on the limits: it’s a request, not a rule

So you don’t over-expect, here are the limits, honestly.

The preference you write in robots.txt is a statement of intent. Whether it’s honored is left to whoever comes to read, and bots that ignore it do exist. If you seriously want to stop something, stating a preference isn’t enough — you need a separate measure that watches the access and enforces allow-or-deny technically.

Why care this much? In Cloudflare’s data, for every time AI comes to read your content, the number of people who actually visit from it ranged from 118-to-1 to, at the high end, 50,000-to-1, depending on the site. Read a lot, and only a trickle comes back.

FAQ

Q. Should I just block all the AI bots?
A. Block everything and search can’t find you either. Splitting by purpose is more practical: let search through, and stop training or automated action (agents) as needed. On Cloudflare, this purpose-by-purpose control has been available even on the free plan since July 1, 2026.

Q. If I write a preference in robots.txt, can I stop training?
A. It’s a preference, not a rule. Some bots ignore it, so seriously stopping traffic needs a separate measure that controls access technically. Keep “writing a preference” and “actually stopping it” as two different things.

Q. What changes from September 15, 2026?
A. For domains newly registered with Cloudflare, ad-showing pages default to blocking training and agents while letting search through. Leave it alone and one policy still gets chosen, so keep this date in mind for new sites.

Wrap-up

  • The honest wish on the ground is “we want search to find us, but not to permit training or automated action,” and robots.txt’s block-all-or-allow-all binary couldn’t express it.
  • On July 1, 2026, Cloudflare made AI access something you can handle across three purposes — search, agent, and training — available to everyone, including the free plan.
  • From September 15, 2026, newly registered domains default to blocking training and agents on ad-showing pages while letting search through. Leave it alone and one policy still gets chosen.
  • In robots.txt, on top of allow-or-deny, you can now attach a three-level preference for how you’re used. The default is “index and quote, with a link back.”
  • But it’s a preference, not a rule. Since some bots ignore it, seriously stopping traffic needs a separate technical measure.
  • What you can do on the ground: decide a policy per page type, record the reasons with the team, add it to the pre-launch check for new domains, and watch behavior after a change.

MONJI+, a Collaborative AI WebOps Platform

Around AI, the range you have to watch in running a website keeps widening.

What pays off in that is
keeping the policy you decided, and the reason for it,
somewhere everyone on the team can see in the same place.

MONJI+ is a WebOps platform that brings these operational decisions and the history of checks into one place.
Share a policy with the team, line it up as a pre-launch checklist item, and watch — in the same place — whether the site’s state shifted after a change.

▼ About MONJI+
https://monji.tech/plus/

▼ Start free
Try every feature with a 30-day free trial. There’s also a Free plan that stays free forever.
https://monji.tech/plus/trial/

▼ We’d love to hear your voice
https://monji.tech/plus/co-creation/

check the list.