robots.txt and sitemap.xml Explained: A Beginner's Guide
Two small files that tell search engines (and AI crawlers) what to crawl and what exists on your site. How robots.txt and sitemap.xml work, examples, the robots.txt mistake that hides your whole site, why Disallow doesn't remove pages from Google, and how to generate both in Next.js.
Two small text files sit at the root of most websites and quietly shape how search engines see them: robots.txt and sitemap.xml. One says where crawlers may go; the other lists the pages you want found. Both take minutes to set up, and getting robots.txt wrong can hide your entire site from Google.
What a crawler is
Search engines discover pages with crawlers (also called bots or spiders) — programs that fetch pages and follow links. Google's is Googlebot. AI companies run their own crawlers too. Before crawling, well-behaved bots read your robots.txt.
robots.txt: the rules for crawlers
It lives at exactly https://yourdomain.com/robots.txt. A typical one:
User-agent: *
Disallow: /admin/
Disallow: /api/
Sitemap: https://yourdomain.com/sitemap.xml
User-agent: *— these rules apply to all crawlers.Disallow: /admin/— don't crawl anything under/admin/.Sitemap:— where your sitemap is.
You can write rules for specific bots by name, for example to allow or block particular AI crawlers:
User-agent: GPTBot
Disallow: /
User-agent: *
Allow: /
The mistake that hides your whole site
User-agent: *
Disallow: /
This blocks everything. It's often added on a staging site to keep it out of Google — then accidentally copied to production. If your site has vanished from search, check this first.
robots.txt is not security or a "remove from Google" button
Two important misunderstandings:
- It's a polite request. Well-behaved crawlers obey it; malicious ones ignore it. Never rely on it to hide private pages — protect them with a login. (And listing
/secret-admin-panel/inrobots.txtadvertises it.) - Disallow doesn't de-index. Blocking a page stops Google crawling it, but if other sites link to it, the URL can still appear in results. To keep a page out of search results, let it be crawled and add a
noindextag:
<meta name="robots" content="noindex">
sitemap.xml: the list of your pages
A sitemap is a list of the URLs you want search engines to know about:
<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<url>
<loc>https://yourdomain.com/</loc>
<lastmod>2026-10-01</lastmod>
</url>
<url>
<loc>https://yourdomain.com/pricing</loc>
<lastmod>2026-09-15</lastmod>
</url>
</urlset>
Things worth knowing:
- Use full, canonical URLs — the exact version you want indexed (https, with or without www consistently; see www vs non-www).
lastmodhelps if it's accurate. Google uses it when it reflects real changes. Don't set every page to today's date.priorityandchangefreqare ignored by Google. You'll see them in old examples; you can leave them out.- Limits: 50,000 URLs or 50 MB per sitemap file. Larger sites use a sitemap index pointing at several sitemaps.
- Only include pages you want indexed — no redirects, 404s, or
noindexpages.
A sitemap doesn't guarantee indexing; it helps search engines find pages, especially on new sites with few links.
Generating them in Next.js
Next.js (App Router) can generate both from code — handy because the sitemap updates itself when you add pages:
// app/robots.ts
import type { MetadataRoute } from "next";
export default function robots(): MetadataRoute.Robots {
return {
rules: { userAgent: "*", allow: "/", disallow: ["/admin/", "/api/"] },
sitemap: "https://yourdomain.com/sitemap.xml",
};
}
// app/sitemap.ts
import type { MetadataRoute } from "next";
export default function sitemap(): MetadataRoute.Sitemap {
return [
{ url: "https://yourdomain.com/", lastModified: new Date("2026-10-01") },
{ url: "https://yourdomain.com/pricing", lastModified: new Date("2026-09-15") },
];
}
For other frameworks, plugins or a static file in public/ work fine.
Submit your sitemap
Add your site to Google Search Console, then submit the sitemap URL under Sitemaps. Search Console will tell you how many URLs it found and any problems. See how to get your website on Google.
The summary
robots.txttells crawlers where they may go;sitemap.xmllists the pages you want found.Disallow: /blocks your entire site — check for it if you've vanished from Google.- robots.txt isn't security, and Disallow isn't removal; use
noindexto keep pages out of results. - Keep sitemaps to canonical, indexable URLs with honest
lastmoddates, and submit them in Search Console.
EasySpawn serves your app on your own domain with HTTPS from day one, so your robots.txt and sitemap have a real, public, canonical address from the start — and Claude Code can generate both and check them on the live site. See how it works or join the waitlist.
Related: SEO Basics for Your App · Anatomy of a URL · Subdomain vs Subdirectory · Dev, Staging, and Production Explained
Keep reading
How to Add a Favicon to Your Website (HTML, Next.js, and Vite)
The small set of favicon files a modern site actually needs, the HTML tags to add, the Next.js and Vite shortcuts, the sizes Google and iPhones use, and why your new favicon still isn't showing.
What Is Vercel? What It Does, What It Costs You, and When to Use Something Else
Vercel is a hosting platform built around frontend frameworks, especially Next.js: push to GitHub and your site is live. What it actually does, what serverless functions mean for your app, where the limits are, and when a regular server is a better fit.