Cloud native principles
What "cloud native" means in the CNCF definition, microservices vs monoliths, immutability and declarative APIs, autoscaling with HPA, VPA, Cluster Autoscaler and KEDA, and serverless with Knative.
Exam tasks: 4.2 (cloud native ecosystem and principles: architecture characteristics, autoscaling, serverless)
The decision: which design property (loose coupling, immutability, declarative state, elastic scaling) does the scenario need, and which Kubernetes mechanism or CNCF project delivers it?
The CNCF definition in one line
The CNCF's official definition says cloud native technology helps you build and run scalable applications in dynamic environments: public, private and hybrid clouds. Its usual building blocks are containers, microservices, service meshes, declarative APIs and immutable infrastructure. The aim is loosely coupled systems you can operate, observe and recover easily, paired with automation so engineers can ship frequent, predictable changes with little toil.
Exam signal
If an option says cloud native "means running in a public cloud", it's wrong. Cloud native is about how applications are built and operated. It applies equally on premises.
Characteristics worth recognising
| Principle | What it means in practice | Kubernetes expression |
|---|---|---|
| Declarative over imperative | You state the end result; a controller works out the steps | YAML manifests applied to the API, reconciled by controllers |
| Immutable infrastructure | Never patch a running instance; build a new image and replace it | New image tag, rolling update replaces Pods |
| Disposable, stateless processes | Any instance can die and be replaced without losing data | Deployments with state kept in databases or volumes |
| Self-healing | The system notices drift and corrects it | ReplicaSets recreate Pods, probes restart containers |
| Elastic | Capacity follows demand, both ways | HPA, VPA, Cluster Autoscaler, KEDA |
| Resilient by design | Expect failure: retries, timeouts, circuit breakers, multiple replicas across zones | Topology spread, PodDisruptionBudgets, service mesh policies |
| Observable | You can understand the system from its telemetry | Prometheus, OpenTelemetry |
- The phrase cattle, not pets sums up disposability: servers and Pods are numbered and replaceable, not named and nursed back to health.
- The Twelve-Factor App guidelines predate Kubernetes but map well onto it: config in the environment (ConfigMaps, Secrets), logs as event streams (stdout), stateless processes, and dev/prod parity.
Monoliths and microservices
| Monolith | Microservices | |
|---|---|---|
| Deploy unit | One artifact for the whole app | One per service, released independently |
| Scaling | Scale everything together | Scale only the hot service |
| Failure blast radius | One bug can take down the whole app | Contained to a service if designed well |
| Data | One shared database | Each service owns its data |
| Complexity lives in | The codebase | The network: discovery, latency, partial failure, tracing |
- Microservices trade code complexity for operational complexity. That's why they come with service discovery, service meshes and distributed tracing.
- A small team with a simple app may be better off with a well-structured monolith. The exam rewards knowing the trade-off, not "microservices always".
Microservices sharing one database
Splitting an app into services that all read and write the same tables keeps the coupling and adds network hops. Loose coupling means each service owns its data and talks to others through APIs or events.
Autoscaling
| Autoscaler | Scales | Driven by | Where it comes from |
|---|---|---|---|
| HPA | Replicas of a Deployment or StatefulSet | CPU and memory utilization, or custom and external metrics | Built into Kubernetes (autoscaling/v2) |
| VPA | CPU and memory requests of Pods | Observed usage history | Add-on from the Kubernetes autoscaler project |
| Cluster Autoscaler | Nodes in node groups | Pods stuck Pending for lack of capacity, and underused nodes | Add-on, works with cloud provider node groups |
| Karpenter | Nodes, picked to fit pending Pods | Pending Pods, consolidation | Kubernetes SIG Autoscaling subproject |
| KEDA | Replicas, including to and from zero | Event sources: queue depth, Kafka lag, cron, Prometheus queries | CNCF graduated project, drives an HPA under the hood |
- HPA utilization targets are a percentage of the Pod's requests. With no CPU request set, a CPU target has nothing to compare against.
- HPA and VPA should not both act on the same CPU or memory metric for one workload; they fight each other.
- Node autoscalers react to Pending Pods, not to high CPU on existing nodes. Pod autoscaling creates the Pods; node autoscaling makes room for them.
- In-place Pod resize (changing a running container's CPU and memory without recreating the Pod) became stable in Kubernetes v1.35, which lets vertical scaling avoid restarts in many cases.
Exam signal
"Scale workers to zero when the queue is empty and back up when messages arrive" is KEDA. Plain HPA keeps at least one replica by default.
Serverless and Knative
- Serverless means you supply code or a container and the platform handles servers, scaling (including to zero) and often per-request billing. Functions as a Service (FaaS) is one form.
- Knative (graduated) brings serverless to Kubernetes. Knative Serving runs request-driven containers with revisions, traffic splitting and scale to zero. Knative Eventing routes events between producers and consumers.
- CloudEvents (graduated) is the CNCF specification for describing event data in a common way, so events can move between platforms. Knative Eventing uses it.
- Trade-offs: cold starts after scaling to zero, limits on run time, and less control over the environment.
Scenarios
The CNCF definition focuses on loosely coupled, resilient, manageable and observable systems in public, private or hybrid clouds. It doesn't require a public cloud or serverless. Patching in place goes against immutable infrastructure.
Pending Pods for lack of resources are exactly what Cluster Autoscaler (or Karpenter) reacts to by adding nodes. VPA changes Pod size, which would make the shortage worse if it raised requests. A higher maxReplicas creates more Pending Pods. The scheduler can only place Pods on nodes that exist.
Immutable infrastructure means you build a new image and replace the running instances, so every copy is identical and reproducible. Editing live containers creates drift that disappears on the next restart. The other principles are real but aren't what the change violates.
Further reading
Observability
Logs, metrics and traces in Kubernetes, how Prometheus, OpenTelemetry, Jaeger and Fluent Bit fit together, the metrics APIs behind autoscaling, and SLOs and cost signals.
CNCF ecosystem and projects
The CNCF landscape, the sandbox, incubating and graduated maturity levels and what each requires, and which graduated project does which job.