Asterrr's Handbook

Containers and runtimes

What a container really is, images and layers, tags and digests, the three OCI specs, registries, the CRI, containerd and CRI-O, low-level runtimes, and sandboxed runtimes with RuntimeClass.

Exam tasks: 1.4 (Containerization: containers, images, registries, standards and runtimes)

The decision: at which layer does a tool or standard sit (build, image, registry, high-level runtime, low-level runtime), and which one does the question describe?

The stack from kubelet to kernel

What a container is

  • A container is an ordinary process that the kernel isolates. There is no "container" object in Linux.
  • Namespaces limit what the process can see: pid (its own process tree), net (its own interfaces and IP), mnt (its own filesystem view), uts (hostname), ipc, user (UID mapping) and cgroup.
  • cgroups limit what it can use: CPU, memory, I/O, number of processes. Kubernetes requests and limits become cgroup settings.
  • Containers share the host kernel. That's why they start in milliseconds and are small, and also why a kernel exploit can escape them. Virtual machines each boot their own kernel.
ContainerVirtual machine
KernelShared with the hostIts own guest kernel
Start timeMilliseconds to secondsTens of seconds or more
SizeMegabytesGigabytes
IsolationProcess-level (namespaces, cgroups)Hardware-level (hypervisor)

Exam signal

"Isolates what a process can see" is namespaces. "Limits how much CPU and memory a process can use" is cgroups. Don't confuse Linux namespaces with Kubernetes namespaces, which only partition API objects.

Images, layers, tags and digests

  • An image is a stack of read-only layers plus a config (entrypoint, env, user). Each build step that changes files adds a layer; identical layers are shared between images and cached on the node.
  • At run time the runtime adds a thin writable layer on top. Anything written there disappears with the container, so persistent data needs a volume.
  • A tag (:3.2.0, :latest) is a movable pointer. A digest (@sha256:9f1c...) identifies exact content and never changes. Pin digests when you need to know exactly what runs.
  • Smaller images (multi-stage builds, distroless or minimal bases) pull faster and expose less attack surface.
# Build stage
FROM golang:1.25 AS build
WORKDIR /src
COPY . .
RUN CGO_ENABLED=0 go build -o /out/tidewatch ./cmd/tidewatch

# Run stage: only the binary ships
FROM gcr.io/distroless/static-debian12
COPY --from=build /out/tidewatch /tidewatch
USER 65532
ENTRYPOINT ["/tidewatch"]

imagePullPolicy defaults

If the image tag is :latest or missing, the default imagePullPolicy is Always. With any other tag (or a digest) it's IfNotPresent. Pushing a new image under the same fixed tag therefore may not reach nodes that already have the old one cached.

The OCI specs

SpecDefinesLets you
Image specImage format: manifest, config, layersBuild with one tool, run with another
Runtime specHow to run an unpacked bundle: config.json plus a root filesystemSwap runc for crun or a sandbox
Distribution specThe HTTP API for pushing and pullingUse any compliant registry
  • The Open Container Initiative (under the Linux Foundation) maintains these standards. runc is its reference runtime implementation.
  • Because of the image spec, images built with Docker, Podman, Buildah or BuildKit run on any Kubernetes runtime.

Registries

  • A registry stores and serves images: Docker Hub, GitHub Container Registry, cloud registries, or self-hosted ones such as Harbor (a CNCF graduated project that adds scanning, signing and replication).
  • Image names read registry/namespace/repository:tag. With no registry, Docker Hub (docker.io) is assumed.
  • Private registries need credentials, which Pods reference through imagePullSecrets.

The CRI and high-level runtimes

  • The Container Runtime Interface is a gRPC API between the kubelet and the runtime. Any runtime that implements it can be plugged in without changing Kubernetes.
  • containerd: CNCF graduated, widely used default on managed and kubeadm clusters. Also the runtime under Docker Engine.
  • CRI-O: a CNCF graduated runtime built specifically for Kubernetes and nothing else. Default on OpenShift.
  • Both use an OCI low-level runtime (runc or crun) to actually create the container.
  • crictl is the CLI for talking to any CRI runtime on a node, for debugging.

Legacy: use containerd or CRI-O through the CRI instead

Docker Engine never implemented the CRI, so Kubernetes shipped an adapter called dockershim. It was removed in v1.24. Clusters now use containerd or CRI-O directly. Images built with Docker are unaffected, because they are OCI images.

Kubernetes stopped supporting Docker images

The dockershim removal changed only the runtime the kubelet talks to. Docker as a build tool and the images it produces still work everywhere. An option claiming you must rebuild all images is wrong.

Sandboxed runtimes and RuntimeClass

  • gVisor (runtime runsc) intercepts system calls in a user-space kernel, so the container never talks to the host kernel directly.
  • Kata Containers runs each Pod inside a lightweight virtual machine with its own kernel.
  • Both trade some performance for stronger isolation, for untrusted or multi-tenant code.
  • A RuntimeClass object names a runtime handler configured on the nodes. Pods opt in with spec.runtimeClassName:
apiVersion: node.k8s.io/v1
kind: RuntimeClass
metadata:
  name: sandboxed
handler: runsc          # must match the handler configured in containerd or CRI-O
---
apiVersion: v1
kind: Pod
metadata:
  name: plugin-runner
spec:
  runtimeClassName: sandboxed
  containers:
    - name: runner
      image: registry.example.com/plugin-runner:0.9.3

Exam signal

"Run untrusted customer code with VM-like isolation while still using Kubernetes" points to a sandboxed runtime (gVisor or Kata Containers) selected with a RuntimeClass, not to a separate namespace or a privileged container.

How the CRI, CNI and CSI compare is covered in CKA extension interfaces.

Scenarios

Scenario
Which interface does the kubelet use to ask containerd or CRI-O to start a Pod's containers?
Scenario
A team at Harborlight is told their cluster is being upgraded to a Kubernetes version without dockershim. Their images are built with docker build and pushed to a private registry. What do they need to change about their images?
Scenario
Which Linux kernel feature limits how much memory a container can consume?

Further reading

On this page