Blog
4 min read

Database Seeding Explained: Test Data That Makes Development Easier

Seeding fills your database with starter data so the app is usable the moment you run it. What seed data is, the difference between reference data and sample data, how to write an idempotent seed script, using Faker, and how to keep seeds away from production.

You clone a project, start the app, and every page is empty: no products, no users, nothing to click. To see anything, you'd have to create it all by hand first. Seeding solves that — a script that fills the database with useful starter data in one command.

What seeding is

A seed script inserts a known set of data into a database. Run it after creating the tables (with migrations), and the app immediately has something to show.

npm run db:seed

Two kinds of seed data

It helps to separate them, because they're treated very differently:

1. Reference data — data the app needs to work, in every environment, including production:

  • the list of countries or currencies,
  • subscription plan names,
  • default roles like admin and member,
  • product categories.

2. Sample (development) data — fake data that makes development and testing pleasant, and must never reach production:

  • a demo user you can log in as,
  • 50 fake products with images,
  • orders in different states (paid, refunded, cancelled),
  • edge cases: a very long name, an empty cart, a user with no orders.

Many projects keep these in separate scripts, or make the sample part run only in development.

A simple seed script

Here's a Node.js example using Postgres. The important idea is that it's idempotent — safe to run more than once without creating duplicates:

// scripts/seed.ts
import { db } from "../src/db";

async function main() {
  // Reference data: insert if missing
  for (const name of ["admin", "member"]) {
    await db.query(
      "INSERT INTO roles (name) VALUES ($1) ON CONFLICT (name) DO NOTHING",
      [name]
    );
  }

  // Sample data: development only
  if (process.env.NODE_ENV !== "production") {
    await db.query(
      `INSERT INTO users (email, name, role)
       VALUES ('demo@example.com', 'Demo User', 'admin')
       ON CONFLICT (email) DO NOTHING`
    );
  }
}

main().then(() => process.exit(0));

ON CONFLICT ... DO NOTHING means "if it's already there, skip it". That relies on a unique constraint on name and email — see primary key vs foreign key. The $1 placeholder keeps values safe from SQL injection.

ORMs have seeding built in or documented: Prisma has a seed command configured in its settings, Django uses fixtures, Rails has db/seeds.rb, Laravel has seeders.

Realistic fake data with Faker

Typing out 50 fake users is tedious. Libraries like Faker (@faker-js/faker in JavaScript, Faker in Python) generate plausible names, emails, addresses and text:

import { faker } from "@faker-js/faker";

faker.seed(42); // same "random" data every run

for (let i = 0; i < 50; i++) {
  await db.query("INSERT INTO products (name, price) VALUES ($1, $2)", [
    faker.commerce.productName(),
    faker.commerce.price({ min: 5, max: 200 }),
  ]);
}

faker.seed(42) makes the output repeatable, so everyone on the team gets the same data and bugs are reproducible.

Good seed data includes awkward cases

Seed data is a chance to see problems early. Include:

  • names with accents, apostrophes and emoji (O'Brien, Zoë, 🚀 Rocket Co),
  • very long text that might break layouts,
  • empty states — a user with no orders,
  • every status your app has,
  • dates in different time zones (dates and time zones),
  • enough rows to make pagination appear (API pagination).

Keep it away from production

Sample data in production is embarrassing at best ("Test Product 1" on your live store) and a security hole at worst (a demo admin with a known password).

  • Guard the sample section with an environment check, as above.
  • Never seed a demo admin account in production.
  • Don't copy real production data into development as "seed data" — it contains real people's personal information. If you need realistic volumes, generate it or anonymise it. (GDPR basics for app builders.)

A reset command

During development it's handy to wipe and rebuild everything:

npm run db:reset   # drop, migrate, seed

Make absolutely sure this can't run against production — check the database URL or environment before dropping anything. AI agents in particular should never have a command like this pointed at a real database. (How to stop an AI agent from deleting your production database.)

The summary

  • A seed script fills the database with starter data in one command.
  • Reference data belongs everywhere; sample data only in development.
  • Make seeds idempotent with ON CONFLICT DO NOTHING (or your ORM's equivalent).
  • Use Faker with a fixed seed for realistic, repeatable fake data, including awkward cases.
  • Guard against seeding — or resetting — production.

EasySpawn servers come with PostgreSQL ready, so Claude Code can write your seed script, run it, and check the app with realistic data in it — with daily backups behind your real data. See how it works or join the waitlist.

Related: How to Design Your First Database · Dev, Staging, and Production Explained · How to Test Your App Before Launch · How to View Your Postgres Database

Keep reading