Blog
5 min read

Linux Namespaces Explained: The Kernel Feature Containers Are Made Of

A container is a process with its own namespaces. What each of the eight Linux namespaces isolates — mount, PID, network, UTS, IPC, user, cgroup, time — how to build a container-like process by hand with unshare, inspect one with nsenter and /proc, and what namespaces do not protect against.

There's no "container" object in the Linux kernel. A container is an ordinary process started with a particular combination of kernel features: namespaces to change what it can see, cgroups to limit what it can use, and capabilities, seccomp and LSMs to restrict what it can do. (Container CPU and memory limits, hardening with seccomp and AppArmor)

This post is about the first part. Once you've built a namespace by hand, Docker stops being magic.

What a namespace is

A namespace wraps a global system resource so that processes inside it get their own isolated instance. Processes in different namespaces of the same type see different things: different process lists, different network interfaces, different mount trees.

Every process belongs to exactly one namespace of each type. You can see them:

ls -l /proc/self/ns
# cgroup -> 'cgroup:[4026531835]'
# ipc    -> 'ipc:[4026531839]'
# mnt    -> 'mnt:[4026531841]'
# net    -> 'net:[4026531840]'
# pid    -> 'pid:[4026531836]'
# time   -> 'time:[4026531834]'
# user   -> 'user:[4026531837]'
# uts    -> 'uts:[4026531838]'

Two processes with the same inode number share that namespace.

The eight namespaces

Namespace Isolates Container effect
mnt Mount points Own root filesystem and mounts
pid Process IDs Own PID 1; can't see host processes
net Network stack Own interfaces, IPs, routes, iptables, ports
uts Hostname, domain name Own hostname
ipc System V IPC, POSIX message queues Own shared memory segments
user UIDs/GIDs, capabilities Root inside maps to unprivileged outside
cgroup View of the cgroup hierarchy Sees its own cgroup as root
time CLOCK_MONOTONIC/BOOTTIME offsets Own uptime clock (used for checkpoint/restore)

The three syscalls behind it all: clone() (create a process in new namespaces), unshare() (move the caller into new namespaces), setns() (join an existing namespace).

Build a container by hand

unshare exposes this from the shell. As root (or with a user namespace, see below):

sudo unshare --pid --fork --mount-proc --uts --ipc --net --mount bash

Inside:

hostname sandbox && hostname     # sandbox — host unaffected (UTS)
ps aux                           # only bash and ps; bash is PID 1 (PID)
ip link                          # just a down loopback interface (net)

--mount-proc remounts /proc inside the new mount namespace so ps reads the new PID namespace's view. --fork is needed because the caller of unshare can't change its own PID namespace; its child becomes PID 1.

Add a root filesystem and you're most of the way to a container:

mkdir -p /tmp/rootfs
docker export $(docker create alpine) | tar -C /tmp/rootfs -xf -
sudo unshare --pid --fork --mount --uts --ipc --net --root=/tmp/rootfs /bin/sh -c 'mount -t proc proc /proc; exec /bin/sh'

Real runtimes use pivot_root rather than chroot-style root changes (it's harder to escape), set up cgroups, drop capabilities and apply seccomp — but the namespace part is exactly this. (Docker image layers and OverlayFS covers where that root filesystem comes from.)

Network namespaces

A fresh network namespace has only a loopback device, down. Container networking is wiring between namespaces:

sudo ip netns add demo
sudo ip link add veth-host type veth peer name veth-demo
sudo ip link set veth-demo netns demo
sudo ip addr add 10.10.0.1/24 dev veth-host && sudo ip link set veth-host up
sudo ip netns exec demo ip addr add 10.10.0.2/24 dev veth-demo
sudo ip netns exec demo ip link set veth-demo up
sudo ip netns exec demo ping -c1 10.10.0.1

Docker does this with a bridge, NAT rules and an embedded DNS server. (Container networking internals)

User namespaces: root that isn't root

A user namespace maps IDs inside to different IDs outside. UID 0 in the container can map to UID 100000 on the host: it has full capabilities within its namespaces (it can mount things there, configure its network), but to the host kernel it's an unprivileged user.

unshare --user --map-root-user --pid --fork --mount-proc bash
id    # uid=0(root) — but only inside

No sudo needed: unprivileged user namespaces let normal users create the others. That's what rootless containers build on — and also why some distributions restrict unprivileged user namespaces, since they expose more kernel surface to unprivileged users. (Rootless containers and user namespaces)

Inspecting and entering namespaces

lsns                                 # list namespaces and their processes
PID=$(docker inspect -f '{{.State.Pid}}' mycontainer)
sudo nsenter -t "$PID" -n ip addr     # run host's `ip` inside the container's net namespace
sudo nsenter -t "$PID" -a sh          # enter all its namespaces

nsenter -n is a superb debugging trick: use host tools (tcpdump, ss, curl) inside a minimal container's network namespace without installing anything in the image.

What namespaces don't do

Namespaces change visibility, not privilege or resource use:

  • One kernel. Every container shares the host kernel. A kernel vulnerability reachable via a syscall can break out regardless of namespaces. That's why seccomp filters syscalls, and why gVisor and microVMs exist. (Firecracker vs gVisor vs containers)
  • No resource limits. A process in its own namespaces can still eat all CPU and memory — cgroups handle that.
  • Not everything is namespaced. The system clock (CLOCK_REALTIME), kernel modules, many /proc and /sys entries, the kernel keyring and devices need other controls; runtimes mask sensitive paths.
  • Root in a non-user-namespaced container is real root for anything that crosses a boundary (a mounted host path, a device, an extra capability).

For untrusted, multi-tenant workloads, namespaces are a necessary layer, not a sufficient one. (Docker vs Linux users for multi-tenant isolation, run AI-generated code safely)

The summary

  • A container is a process with its own namespaces, plus cgroups and security filters.
  • Eight namespaces: mount, PID, network, UTS, IPC, user, cgroup, time.
  • unshare creates them, nsenter enters them, /proc/<pid>/ns and lsns show them.
  • User namespaces make container root unprivileged on the host.
  • Namespaces isolate views, not the kernel — combine with cgroups, seccomp and stronger boundaries for untrusted code.

EasySpawn gives every server its own virtual machine with a dedicated kernel — a stronger boundary than namespaces alone — so AI agents and apps on one account can't see another's. See how it works or join the waitlist.

Related: Rootless Containers and User Namespaces · Container Networking Internals · Containers vs Virtual Machines · How KVM Virtualization Works

Keep reading