Domain 4 · Cloud native architecture
12% of the exam. Observability signals and tools, the principles behind cloud native design and autoscaling, the CNCF projects and their maturity, and how the community governs Kubernetes.
Domain 4 steps back from individual Kubernetes objects. Questions ask why cloud native systems are built the way they are, which project does a job, and who decides how projects and Kubernetes evolve. Expect short recall questions where every option is a real project or body, and only one does the job named.
| Competency | What it's really asking | Pages |
|---|---|---|
| 4.1 Observability | Logs vs metrics vs traces, what Prometheus, OpenTelemetry, Jaeger and Fluent Bit each do, the metrics APIs behind kubectl top and the HPA, SLOs and cost | Observability |
| 4.2 Cloud native ecosystem and principles | Loose coupling, immutability and declarative APIs, which autoscaler fits, serverless, and which CNCF project at which maturity level does the job | Cloud native principles, CNCF ecosystem |
| 4.3 Cloud native community and collaboration | TOC vs TAGs vs SIGs, how a KEP becomes a feature, the release cadence, and the open standards (OCI, CRI, CNI, CSI, OpenTelemetry) | Community and governance |
What connects Domain 4 to the rest of the exam:
- Runtimes and the OCI and CRI specs are covered in depth in Containers and runtimes.
- CNI, CSI and service meshes show up as mechanisms in Networking, Storage and Service mesh.
- Argo CD, Flux and Helm are the delivery tools in GitOps and CI/CD and Packaging and rollouts.
Right category, wrong project
Distractors in this domain are usually real CNCF projects from a neighbouring category: Jaeger offered for a metrics question, Helm for a GitOps question, Envoy for a DNS question. Name the job first (store metrics, sync from Git, resolve names), then pick the project built for exactly that job.
Debugging applications
Debugging a running application on Kubernetes with logs, events, kubectl exec, port-forward, kubectl debug and ephemeral containers, and reading why a Deployment rollout is stuck.
Observability
Logs, metrics and traces in Kubernetes, how Prometheus, OpenTelemetry, Jaeger and Fluent Bit fit together, the metrics APIs behind autoscaling, and SLOs and cost signals.