Containers and runtimes
What a container really is, images and layers, tags and digests, the three OCI specs, registries, the CRI, containerd and CRI-O, low-level runtimes, and sandboxed runtimes with RuntimeClass.
Exam tasks: 1.4 (Containerization: containers, images, registries, standards and runtimes)
The decision: at which layer does a tool or standard sit (build, image, registry, high-level runtime, low-level runtime), and which one does the question describe?
The stack from kubelet to kernel
What a container is
- A container is an ordinary process that the kernel isolates. There is no "container" object in Linux.
- Namespaces limit what the process can see:
pid(its own process tree),net(its own interfaces and IP),mnt(its own filesystem view),uts(hostname),ipc,user(UID mapping) andcgroup. - cgroups limit what it can use: CPU, memory, I/O, number of processes. Kubernetes requests and limits become cgroup settings.
- Containers share the host kernel. That's why they start in milliseconds and are small, and also why a kernel exploit can escape them. Virtual machines each boot their own kernel.
| Container | Virtual machine | |
|---|---|---|
| Kernel | Shared with the host | Its own guest kernel |
| Start time | Milliseconds to seconds | Tens of seconds or more |
| Size | Megabytes | Gigabytes |
| Isolation | Process-level (namespaces, cgroups) | Hardware-level (hypervisor) |
Exam signal
"Isolates what a process can see" is namespaces. "Limits how much CPU and memory a process can use" is cgroups. Don't confuse Linux namespaces with Kubernetes namespaces, which only partition API objects.
Images, layers, tags and digests
- An image is a stack of read-only layers plus a config (entrypoint, env, user). Each build step that changes files adds a layer; identical layers are shared between images and cached on the node.
- At run time the runtime adds a thin writable layer on top. Anything written there disappears with the container, so persistent data needs a volume.
- A tag (
:3.2.0,:latest) is a movable pointer. A digest (@sha256:9f1c...) identifies exact content and never changes. Pin digests when you need to know exactly what runs. - Smaller images (multi-stage builds, distroless or minimal bases) pull faster and expose less attack surface.
# Build stage
FROM golang:1.25 AS build
WORKDIR /src
COPY . .
RUN CGO_ENABLED=0 go build -o /out/tidewatch ./cmd/tidewatch
# Run stage: only the binary ships
FROM gcr.io/distroless/static-debian12
COPY --from=build /out/tidewatch /tidewatch
USER 65532
ENTRYPOINT ["/tidewatch"]imagePullPolicy defaults
If the image tag is :latest or missing, the default imagePullPolicy is Always. With any other tag (or a
digest) it's IfNotPresent. Pushing a new image under the same fixed tag therefore may not reach nodes that
already have the old one cached.
The OCI specs
| Spec | Defines | Lets you |
|---|---|---|
| Image spec | Image format: manifest, config, layers | Build with one tool, run with another |
| Runtime spec | How to run an unpacked bundle: config.json plus a root filesystem | Swap runc for crun or a sandbox |
| Distribution spec | The HTTP API for pushing and pulling | Use any compliant registry |
- The Open Container Initiative (under the Linux Foundation) maintains these standards. runc is its reference runtime implementation.
- Because of the image spec, images built with Docker, Podman, Buildah or BuildKit run on any Kubernetes runtime.
Registries
- A registry stores and serves images: Docker Hub, GitHub Container Registry, cloud registries, or self-hosted ones such as Harbor (a CNCF graduated project that adds scanning, signing and replication).
- Image names read
registry/namespace/repository:tag. With no registry, Docker Hub (docker.io) is assumed. - Private registries need credentials, which Pods reference through
imagePullSecrets.
The CRI and high-level runtimes
- The Container Runtime Interface is a gRPC API between the kubelet and the runtime. Any runtime that implements it can be plugged in without changing Kubernetes.
- containerd: CNCF graduated, widely used default on managed and kubeadm clusters. Also the runtime under Docker Engine.
- CRI-O: a CNCF graduated runtime built specifically for Kubernetes and nothing else. Default on OpenShift.
- Both use an OCI low-level runtime (runc or crun) to actually create the container.
crictlis the CLI for talking to any CRI runtime on a node, for debugging.
Legacy: use containerd or CRI-O through the CRI instead
Docker Engine never implemented the CRI, so Kubernetes shipped an adapter called dockershim. It was removed in v1.24. Clusters now use containerd or CRI-O directly. Images built with Docker are unaffected, because they are OCI images.
Kubernetes stopped supporting Docker images
The dockershim removal changed only the runtime the kubelet talks to. Docker as a build tool and the images it produces still work everywhere. An option claiming you must rebuild all images is wrong.
Sandboxed runtimes and RuntimeClass
- gVisor (runtime
runsc) intercepts system calls in a user-space kernel, so the container never talks to the host kernel directly. - Kata Containers runs each Pod inside a lightweight virtual machine with its own kernel.
- Both trade some performance for stronger isolation, for untrusted or multi-tenant code.
- A RuntimeClass object names a runtime handler configured on the nodes. Pods opt in with
spec.runtimeClassName:
apiVersion: node.k8s.io/v1
kind: RuntimeClass
metadata:
name: sandboxed
handler: runsc # must match the handler configured in containerd or CRI-O
---
apiVersion: v1
kind: Pod
metadata:
name: plugin-runner
spec:
runtimeClassName: sandboxed
containers:
- name: runner
image: registry.example.com/plugin-runner:0.9.3Exam signal
"Run untrusted customer code with VM-like isolation while still using Kubernetes" points to a sandboxed runtime (gVisor or Kata Containers) selected with a RuntimeClass, not to a separate namespace or a privileged container.
How the CRI, CNI and CSI compare is covered in CKA extension interfaces.
Scenarios
The Container Runtime Interface is the kubelet-to-runtime API. CNI configures Pod networking, CSI attaches storage, and the distribution spec describes how registries serve images.
Removing dockershim changed the runtime on the nodes, not the image format. Docker produces OCI-compliant images that any CRI runtime can pull and run. There's no CRI image format, rebuilding with another tool changes nothing, and runc is already the default low-level runtime.
cgroups account for and cap CPU, memory and other resources, and Kubernetes limits are enforced through them. Namespaces (mount, network and the rest) control visibility, not consumption. chroot only changes the apparent root directory.
Further reading
Scheduling
How kube-scheduler filters and scores nodes, and how resource requests, nodeSelector, node and Pod affinity, taints and tolerations, and priority steer where Pods land.
Domain 2 · Container orchestration
28% of the exam. How Pods reach each other and the outside world, how a cluster is secured, how storage attaches to Pods, and how you tell what broke.