Database Seeding Explained: Test Data That Makes Development Easier
Seeding fills your database with starter data so the app is usable the moment you run it. What seed data is, the difference between reference data and sample data, how to write an idempotent seed script, using Faker, and how to keep seeds away from production.
You clone a project, start the app, and every page is empty: no products, no users, nothing to click. To see anything, you'd have to create it all by hand first. Seeding solves that — a script that fills the database with useful starter data in one command.
What seeding is
A seed script inserts a known set of data into a database. Run it after creating the tables (with migrations), and the app immediately has something to show.
npm run db:seed
Two kinds of seed data
It helps to separate them, because they're treated very differently:
1. Reference data — data the app needs to work, in every environment, including production:
- the list of countries or currencies,
- subscription plan names,
- default roles like
adminandmember, - product categories.
2. Sample (development) data — fake data that makes development and testing pleasant, and must never reach production:
- a demo user you can log in as,
- 50 fake products with images,
- orders in different states (paid, refunded, cancelled),
- edge cases: a very long name, an empty cart, a user with no orders.
Many projects keep these in separate scripts, or make the sample part run only in development.
A simple seed script
Here's a Node.js example using Postgres. The important idea is that it's idempotent — safe to run more than once without creating duplicates:
// scripts/seed.ts
import { db } from "../src/db";
async function main() {
// Reference data: insert if missing
for (const name of ["admin", "member"]) {
await db.query(
"INSERT INTO roles (name) VALUES ($1) ON CONFLICT (name) DO NOTHING",
[name]
);
}
// Sample data: development only
if (process.env.NODE_ENV !== "production") {
await db.query(
`INSERT INTO users (email, name, role)
VALUES ('demo@example.com', 'Demo User', 'admin')
ON CONFLICT (email) DO NOTHING`
);
}
}
main().then(() => process.exit(0));
ON CONFLICT ... DO NOTHING means "if it's already there, skip it". That relies on a unique constraint on name and email — see primary key vs foreign key. The $1 placeholder keeps values safe from SQL injection.
ORMs have seeding built in or documented: Prisma has a seed command configured in its settings, Django uses fixtures, Rails has db/seeds.rb, Laravel has seeders.
Realistic fake data with Faker
Typing out 50 fake users is tedious. Libraries like Faker (@faker-js/faker in JavaScript, Faker in Python) generate plausible names, emails, addresses and text:
import { faker } from "@faker-js/faker";
faker.seed(42); // same "random" data every run
for (let i = 0; i < 50; i++) {
await db.query("INSERT INTO products (name, price) VALUES ($1, $2)", [
faker.commerce.productName(),
faker.commerce.price({ min: 5, max: 200 }),
]);
}
faker.seed(42) makes the output repeatable, so everyone on the team gets the same data and bugs are reproducible.
Good seed data includes awkward cases
Seed data is a chance to see problems early. Include:
- names with accents, apostrophes and emoji (
O'Brien,Zoë,🚀 Rocket Co), - very long text that might break layouts,
- empty states — a user with no orders,
- every status your app has,
- dates in different time zones (dates and time zones),
- enough rows to make pagination appear (API pagination).
Keep it away from production
Sample data in production is embarrassing at best ("Test Product 1" on your live store) and a security hole at worst (a demo admin with a known password).
- Guard the sample section with an environment check, as above.
- Never seed a demo admin account in production.
- Don't copy real production data into development as "seed data" — it contains real people's personal information. If you need realistic volumes, generate it or anonymise it. (GDPR basics for app builders.)
A reset command
During development it's handy to wipe and rebuild everything:
npm run db:reset # drop, migrate, seed
Make absolutely sure this can't run against production — check the database URL or environment before dropping anything. AI agents in particular should never have a command like this pointed at a real database. (How to stop an AI agent from deleting your production database.)
The summary
- A seed script fills the database with starter data in one command.
- Reference data belongs everywhere; sample data only in development.
- Make seeds idempotent with
ON CONFLICT DO NOTHING(or your ORM's equivalent). - Use Faker with a fixed seed for realistic, repeatable fake data, including awkward cases.
- Guard against seeding — or resetting — production.
EasySpawn servers come with PostgreSQL ready, so Claude Code can write your seed script, run it, and check the app with realistic data in it — with daily backups behind your real data. See how it works or join the waitlist.
Related: How to Design Your First Database · Dev, Staging, and Production Explained · How to Test Your App Before Launch · How to View Your Postgres Database
Keep reading
What Is MongoDB? A Beginner's Guide to Document Databases
MongoDB stores data as flexible JSON-like documents instead of tables. How it works, what collections and documents are, where it shines, where Postgres is the better pick, and why AI tools sometimes reach for it.
What Is Supabase? A Beginner's Guide to the Backend Behind Many AI-Built Apps
Supabase gives your app a Postgres database, logins, file storage and serverless functions from one dashboard. What each part does, how the publishable and secret keys work, why row-level security matters, free-plan limits, and when to use something else.