How to Edit robots.txt in HubSpot (and Check It Works)

You add a line to hubspot robots.txt, hit Save, and assume Google will stop showing that page. Then it keeps appearing in search. HubSpot’s own guidance explains why: you cannot stop content from being indexed and shown in search results with a robots.txt file. Robots.txt is a crawling control, not a “remove this URL from Google” switch.

Before you touch anything, make sure you’re working on the right domain. HubSpot is explicit that only content hosted on a domain connected to HubSpot can be blocked in your robots.txt file. If the page you care about lives elsewhere, editing robots.txt inside HubSpot won’t change what bots can access.

This guide sticks to what HubSpot documents: where to edit robots.txt in your portal, how default settings differ from a domain override, what User-agent and Disallow do, and the one HubSpot-specific bot (HubSpotContentSearchBot) you may need to account for if you use HubSpot site search. You’ll also get a practical way to sanity-check your rules before you wait for crawlers to revisit.

Which HubSpot Domains Can Be Blocked in Robots.txt?

Before you edit hubspot robots.txt, make sure you are editing a file that can actually affect the content you care about. HubSpot is explicit about scope: only content hosted on a domain connected to HubSpot can be blocked in your robots.txt file. If a page lives on a domain that is not connected to HubSpot, changing robots.txt inside HubSpot will not control crawler access to that page.

That scope matters in real accounts because many portals have more than one domain connected (for example, a primary website domain plus a blog subdomain). When you edit robots.txt in HubSpot, you pick the specific domain you want to affect, or you apply defaults across connected domains. The next section shows the exact click-path for how to edit robots.txt HubSpot provides.

HubSpot System Domains (hs-sites) Are Always No-Index

HubSpot also calls out an important exception: content on HubSpot system domains containing hs-sites is always set as no-index in a robots.txt file. In practice, this means you should not expect to manage indexing for those system domains through your own custom rules. HubSpot already sets them to no-index.

If you are troubleshooting why something shows up in search, start by confirming which host the URL uses. If the URL uses a connected domain, you can manage crawler access through HubSpot’s robots.txt settings for that domain. If the URL uses an hs-sites system domain, HubSpot’s default no-index behavior applies. If the URL uses an entirely different domain that is not connected to HubSpot, you need to manage robots.txt on that other platform or hosting provider.

How to Edit Robots.txt in HubSpot (Step-by-Step)

Once you have confirmed the URL uses a domain connected to HubSpot, you can edit hubspot robots.txt directly in your portal. HubSpot documents robots.txt management inside your website settings, per domain, under the SEO tools.

  1. Open HubSpot Settings.
  2. Go to Content > Pages. (This is the area HubSpot uses for website-level settings.)
  3. Click the SEO & Crawlers tab. HubSpot places the robots.txt editor here.
  4. Open the “Choose a domain to edit its settings” dropdown. Select the connected domain whose robots.txt you want to change.
  5. Pick your editing mode:
    • Default settings for all domains to apply your robots.txt rules across connected domains.
    • Select a specific domain, then click Override default settings if you need domain-specific rules. HubSpot notes that this option overrides any default robots.txt settings for that domain.
  6. Make your robots.txt edits. (HubSpot’s robots.txt format uses User-agent and Disallow, which we break down later.)
  7. Click Save in the bottom left.

If you manage multiple brands or environments (for example, a main domain and a subdomain), slow down at the domain dropdown. Most “edit robots.txt HubSpot” mistakes come from changing the wrong domain’s settings and then checking a different hostname in the browser.

HubSpot’s official step list for editing robots.txt appears in its knowledge base article Prevent content from appearing in search results. If you cannot find the SEO & Crawlers tab or the dropdown, use that article to confirm you are in the documented location.

Default Settings vs Override: Which Should You Use?

The choice between hubspot robots.txt defaults and a domain-specific override comes down to one question: do you want the same crawl rules applied everywhere, or do you need different rules per connected domain?

HubSpot’s knowledge base explains that you can edit robots.txt for all connected domains by selecting Default settings for all domains. Use this when your portal has multiple connected domains (for example, a main site domain and a blog subdomain) and you want consistent crawl restrictions across all of them.

Pick a specific domain in the Choose a domain to edit its settings dropdown when that domain needs its own rules. This is the common case when different parts of your site live on different connected domains, and you want to block crawling in one area without affecting the others.

What “Override Default Settings” Changes in HubSpot

When you select a specific domain, HubSpot may show an Override default settings option. HubSpot’s documentation states that Override default settings will override any robots.txt default settings for that domain.

That “override” detail is the part people miss. An override does not layer extra rules on top of the default robots.txt. It replaces the default robots.txt settings for that one domain. If you rely on a default rule (for example, a Disallow line you expect to apply everywhere), you need to make sure the override version for that domain includes whatever you still want enforced.

Use this quick decision guide:

  • Choose “Default settings for all domains” when one robots.txt policy fits every connected domain.
  • Choose a specific domain + “Override default settings” when one connected domain needs different crawl rules than the rest.

If you are unsure, start with the default settings first, then move to a domain override only when you can explain why that single domain needs different crawler access.

What Do User-Agent and Disallow Mean in HubSpot’s Robots.txt?

When you edit hubspot robots.txt, you are mainly editing two directives HubSpot calls out in its documentation: User-agent and Disallow. Understanding what each line does makes it easier to use the default settings safely, and harder to accidentally block the wrong part of your site.

HubSpot defines User-agent as the line that specifies which search engine or web bot a rule applies to. HubSpot also notes that the default User-agent is set to include all search engines and shows as an asterisk (*). In plain terms, User-agent: * means “apply the rules below to every crawler that reads robots.txt.”

HubSpot defines Disallow as telling a search engine not to crawl and index any files or pages using a specific URL slug. In HubSpot’s own example, you add a page to robots.txt by entering Disallow: /url-slug. That format matters: you are blocking by path (the part after your domain), not by a full URL.

User-Agent and Disallow in HubSpot Robots.txt: How the Lines Work Together

A robots.txt file groups rules under a User-agent. The simplest pattern you will use when you edit robots.txt in HubSpot looks like this:

  • User-agent: * (the rule targets all bots)
  • Disallow: /url-slug (the rule blocks crawling for that slug pattern)

Think in “if-then” terms: if the crawler matches the User-agent, then it follows the Disallow lines underneath. If you pick the wrong User-agent, you can end up blocking a tool you did not mean to block, or failing to block the bots you care about.

Also keep your scope straight. A Disallow line like Disallow: /url-slug applies to the domain you selected in HubSpot’s “Choose a domain to edit its settings” dropdown, and it matches paths on that domain. If you need different rules for a blog subdomain versus a main website domain, you handle that by selecting the right domain in HubSpot, then editing the User-agent and Disallow lines for that domain’s robots.txt.

Do You Use HubSpot Site Search? Add HubSpotContentSearchBot

Robots.txt rules in HubSpot apply per domain and per bot, so the hubspot robots.txt file can affect more than Googlebot. If you use HubSpot’s site search module, HubSpot says you need to add HubSpotContentSearchBot as a separate User-agent so the search feature can crawl your pages. HubSpot documents this requirement in its article Prevent content from appearing in search results.

This point matters when you edit robots.txt HubSpot settings to block areas of your site. A broad rule under User-agent: * can block the crawler HubSpot uses to populate on-site search results. Adding HubSpotContentSearchBot as its own user-agent gives you a place to allow crawling for site search, while keeping stricter rules for other bots.

How To Add HubSpotContentSearchBot In HubSpot Robots.txt

Use the same editor you already use for domain-level robots.txt changes in HubSpot:

  1. Go to Settings in HubSpot.
  2. Navigate to Content > Pages.
  3. Open the SEO & Crawlers tab.
  4. In Choose a domain to edit its settings, select the connected domain where you use site search.
  5. Edit the robots.txt content to include HubSpotContentSearchBot as its own User-agent. HubSpot states that including it as a separate user-agent allows the search feature to crawl your pages.
  6. Click Save (bottom left).

Keep the domain selection tight. If your blog lives on a subdomain and your main site lives on the root domain, add HubSpotContentSearchBot on the domain where the searchable pages actually live. The robots.txt file only applies to the domain you picked in HubSpot’s dropdown.

Robots.txt vs Noindex: When Should You Use Each?

When you edit hubspot robots.txt, treat it as a crawling control, not a guaranteed way to remove a URL from Google. HubSpot’s own guidance says you cannot stop content from being indexed and shown in search results with a robots.txt file. That matters any time you are trying to hide a page that already appears in search.

Use robots.txt when your goal is to tell crawlers where they should not go on a HubSpot-connected domain. Use noindex when your goal is to keep a page out of search results, especially if search engines already discovered it.

When To Use Robots.txt vs Noindex In HubSpot

A practical way to decide: think in “indexed vs not indexed,” not “crawlable vs not crawlable.” Robots.txt controls crawling. Noindex targets indexing.

Keep one constraint in mind when you plan your fix. HubSpot says you should not combine the noindex meta tag method with the robots.txt method because blocking a page in robots.txt prevents search engines from seeing the noindex tag. If you block crawling first, crawlers may never fetch the page to read the noindex directive in the head.

If you are troubleshooting an “edit robots.txt HubSpot” change that did not remove a URL from search, this is usually why. You gave crawl instructions, but you did not give an indexing directive that search engines could actually read.

Sanity-Check Your Rules With These Free Robots.txt Tools

Screenshot of workspace Balzac

Robots.txt changes often “fail” for a simple reason: you edited hubspot robots.txt, saved it, then checked the wrong URL, the wrong domain, or a bot that follows a different rule block. A quick sanity check catches those mistakes before you wait on crawlers to revisit your site.

These three free Balzac tools help you validate syntax and spot access issues after you edit robots.txt in HubSpot:

How to Verify a HubSpot Robots.txt Change Before and After You Save

Use this workflow whenever you edit robots.txt HubSpot settings, especially when you use domain overrides:

  1. Confirm the exact robots.txt URL you plan to check. Robots.txt lives at the root of the domain (for example, https://www.hubspot.com/robots.txt). Make sure you check the same hostname you selected in HubSpot’s domain dropdown.
  2. Copy your current robots.txt text into the Robots.txt Checker. Focus on obvious mistakes: missing User-agent lines, extra characters, or a Disallow path that is broader than you intended.
  3. Test the specific paths you care about. Use real examples like /pricing, /blog/, or the exact slug you added (HubSpot’s documented format is Disallow: /url-slug).
  4. If you use HubSpot site search, check your bot blocks. HubSpot recommends adding HubSpotContentSearchBot as a separate user-agent so site search can crawl pages. Verify you did not block it with a catch-all rule.
  5. After you click Save in HubSpot, re-check the live file. Load /robots.txt in a browser and run the updated content through the checker again. This confirms the published file matches what you meant to ship.

If you are drafting rules from scratch, use the Robots.txt Generator first, then paste the output into HubSpot’s robots.txt editor. It is faster than hand-typing directives and helps you avoid formatting errors.

If your concern includes AI bots, run the AI Crawler Checker before and after the change. It gives you a practical view of crawler access without guessing which user-agent string a tool uses.

FAQ: HubSpot Robots.txt Questions People Ask Most

When you check whether a crawler can access a URL, the next questions usually come fast: where is the file, do you even need it, and what exactly did your rule block? This FAQ answers the common “edit robots.txt HubSpot” questions using HubSpot’s own documentation and blog guidance.

Where does hubspot robots.txt live? Robots.txt lives at the root of your domain. HubSpot’s blog says the robots.txt file always sits at the root domain, for example https://www.hubspot.com/robots.txt.

Is a robots.txt file required? No. HubSpot states a robots.txt file is not required for a website. You still may choose to use one to guide crawlers away from certain paths.

Which domain should I edit in HubSpot? Edit the domain that actually hosts the content you want to control. HubSpot states that only content hosted on a domain connected to HubSpot can be blocked in your robots.txt file. In HubSpot, you pick that domain from the Choose a domain to edit its settings dropdown under Settings, Content > Pages, SEO & Crawlers.

Can I change robots.txt for hs-sites system domains? You generally do not manage those the same way. HubSpot says content on HubSpot system domains containing hs-sites is always set as no-index in a robots.txt file.

What Does Disallow Block in HubSpot Robots.txt?

What does Disallow actually do? HubSpot defines Disallow as telling a search engine not to crawl and index any files or pages using a specific URL slug. HubSpot’s example format is Disallow: /url-slug, which blocks by path (slug), on the domain you selected in HubSpot.

Will blocking in robots.txt remove a page from Google results? Not reliably. HubSpot’s blog says you cannot stop content from being indexed and shown in search results with a robots.txt file. If search engines already indexed the page, HubSpot says you can add a noindex meta tag to the content’s head HTML, and HubSpot also says you should not combine noindex with robots.txt blocking because crawlers cannot see the noindex tag if you block the page.

Quick Wrap-Up and Next Step

If you edit hubspot robots.txt, keep the goal straight: you are controlling crawling on a connected domain, not forcing a URL to disappear from search results. HubSpot’s own guidance also matters operationally: if you need a page out of Google and it is already indexed, use the noindex approach in the page head HTML, and do not pair it with robots.txt blocking because crawlers cannot see the noindex tag if you block the page.

When you want to edit robots.txt HubSpot settings safely, use this flow:

  1. Confirm the hostname. Make sure the content lives on a domain connected to HubSpot, since HubSpot says only connected domains can be blocked in robots.txt.
  2. Open the documented editor. Go to HubSpot’s “Prevent content from appearing in search results” guidance if you need to confirm the location: Settings, then Content > Pages, then the SEO & Crawlers tab.
  3. Select the right domain. Use the “Choose a domain to edit its settings” dropdown before you change anything.
  4. Choose default or override on purpose. Use “Default settings for all domains” when one policy fits every connected domain. Use a domain selection plus “Override default settings” when that domain needs different rules (and remember it replaces defaults for that domain).
  5. Edit only what you mean to block. HubSpot’s robots.txt format centers on User-agent and Disallow, with path-based entries like Disallow: /url-slug.
  6. If you use HubSpot site search, protect it. HubSpot recommends adding HubSpotContentSearchBot as a separate user-agent so the search feature can crawl pages.
  7. Save, then verify the live file. HubSpot publishes changes when you click Save (bottom left). Check the root /robots.txt on the same domain you edited.

If you want a fast final check before you walk away, run your rules through the free Robots.txt Checker, draft clean blocks with the Robots.txt Generator, and confirm AI bot access with the AI Crawler Checker.

New Balzac signups get 3 free articles with no card.

Sources

← All articles

Free AI visibility check

Is ChatGPT citing your site?

See if ChatGPT search and Google AI Overviews cite you for your top searches, and who they cite instead. About a minute, no account.

Want content like this on autopilot?

Balzac researches, writes and publishes articles like this one to your site. Every week.

3 free articles · No card needed