Observability
Logs, metrics and traces in Kubernetes, how Prometheus, OpenTelemetry, Jaeger and Fluent Bit fit together, the metrics APIs behind autoscaling, and SLOs and cost signals.
Exam tasks: 4.1 (observability: telemetry signals, Prometheus, OpenTelemetry, tracing, logging, cost management)
The decision: which signal answers the question (a log line, a metric or a trace), and which CNCF project collects, stores or shows it?
Three signals, three questions
| Signal | Answers | Shape | Typical CNCF tool |
|---|---|---|---|
| Logs | What exactly happened in this process at this moment? | Timestamped text or structured events | Fluentd, Fluent Bit |
| Metrics | How much, how often, how fast, over time? | Numeric samples with labels, cheap to store | Prometheus |
| Traces | Where did this one request spend its time across services? | A tree of spans sharing a trace ID | Jaeger, OpenTelemetry |
- Monitoring watches for failure modes you predicted (an alert on error rate). Observability lets you ask new questions about failures you didn't predict, by correlating all three signals.
- Some newer material adds profiles (continuous CPU and memory profiling) as a fourth signal. OpenTelemetry is adding profiling support, but the exam focuses on the three above.
Exam signal
"Which microservice made the checkout request slow?" is a trace question. "Is the error rate above 2% for the last 10 minutes?" is a metric question. "What stack trace did the payment Pod print before it crashed?" is a log question.
Logs in Kubernetes
- Containers should write to stdout and stderr. The container runtime writes those streams to files on the
node, and the kubelet serves them to
kubectl logsand rotates them. kubectl logs <pod> -c <container>picks a container,--previousshows the last crashed instance, and-ffollows the stream.- Kubernetes has no built-in log storage. When a Pod is deleted or its node dies, its logs go with it. That's why clusters ship logs elsewhere.
| Pattern | How it works | Trade-off |
|---|---|---|
| Node-level agent (most common) | A DaemonSet such as Fluent Bit tails every container log file on each node | One agent per node, no app changes |
| Sidecar | A helper container in the Pod streams or ships the app's log file | Works for apps that only log to files, costs a container per Pod |
| App pushes directly | The app sends logs to a backend over the network | Couples the app to the backend |
- Fluentd (graduated) is the original unified logging layer. Fluent Bit is its lightweight sibling under the same CNCF project, preferred as the per-node agent.
- Structured logs (JSON with fields like
levelandtrace_id) can be filtered and linked to traces.
kubectl logs as the logging strategy
kubectl logs reads what's still on the node. It can't search across Pods, and it shows nothing after the Pod is
gone. An answer that relies on it for audits or post-incident review is wrong; centralized shipping is the fix.
Metrics and Prometheus
- Prometheus (graduated, the second project to join CNCF after Kubernetes) pulls (scrapes) metrics over
HTTP from
/metricsendpoints at an interval and stores them as time series: a metric name plus labels. - Kubernetes components (API server, kubelet, scheduler, controller manager) expose metrics in Prometheus format already. Software that can't, gets an exporter that translates (for example node_exporter for host metrics).
- Short-lived batch jobs can't be scraped reliably, so they push to the Pushgateway, which Prometheus scrapes.
- You query with PromQL, for example
rate(http_requests_total[5m]). - Alertmanager receives alerts from Prometheus rules and handles routing, grouping, deduplication and silencing. Prometheus itself evaluates the rules; it doesn't send emails or pages.
- Prometheus is built for a single server. Thanos and Cortex (both incubating) add long-term storage and a global view across many Prometheus servers.
| Metric type | Behaviour | Example |
|---|---|---|
| Counter | Only goes up (resets on restart) | Total requests served |
| Gauge | Goes up and down | Current memory use, queue length |
| Histogram | Counts observations in buckets, so you can compute percentiles on the server | Request latency |
| Summary | Computes quantiles in the client | Request latency, pre-aggregated |
Grafana is the CNCF dashboard project
Grafana is the usual dashboard on top of Prometheus, but it's a Grafana Labs project, not a CNCF project. If the question asks which CNCF graduated project collects and stores metrics, the answer is Prometheus.
The metrics APIs behind autoscaling
- metrics-server collects CPU and memory usage from each kubelet and serves the Resource Metrics API
(
metrics.k8s.io). It keeps only current values in memory: it isn't a monitoring system. kubectl top podsandkubectl top nodesfail until metrics-server (or another provider) is installed.- The HPA reads resource metrics, or custom and external metrics served by an adapter.
- kube-state-metrics is different: it turns the state of objects (desired vs available replicas, Pod phase) into Prometheus metrics. Usage comes from metrics-server; object state comes from kube-state-metrics.
Traces and OpenTelemetry
- A trace is one request's journey. Each step is a span with a start time, duration and parent. The trace ID travels between services in request headers, usually in the W3C Trace Context format.
- OpenTelemetry (OTel, graduated) is the vendor-neutral standard for producing and moving telemetry: APIs, SDKs, auto-instrumentation, the OTLP protocol and the Collector. It covers traces, metrics and logs.
- OTel is not a backend. It sends data to Jaeger, Prometheus or a commercial vendor. Switching vendors means changing the Collector's exporter, not re-instrumenting code.
- The Collector pipeline is receivers, then processors (batching, sampling, dropping sensitive attributes), then exporters. It runs as an agent per node, a gateway Deployment, or both.
- Jaeger (graduated) stores and visualizes traces. Current Jaeger is built on the OTel Collector and accepts OTLP natively.
Legacy: use OpenTelemetry instead
OpenTracing (a CNCF project) and OpenCensus (from Google) merged into OpenTelemetry in 2019. Both are archived. Jaeger's own client libraries are also retired in favour of the OpenTelemetry SDKs.
SLOs, alerting and cost
| Term | Meaning | Example |
|---|---|---|
| SLI | A measured indicator | Share of checkout requests answered under 300 ms |
| SLO | Your internal target for the SLI | 99.5% over 30 days |
| SLA | A contract with a penalty, looser than the SLO | 99% or the customer gets credits |
| Error budget | What the SLO lets you miss | 0.5% of requests per 30 days |
- Alert on symptoms users feel (SLO burn rate, error rate, latency) rather than on every cause (CPU at 80%).
- Cost is an observability signal too. Requests drive scheduling and cost, so oversized requests waste nodes even when usage is low. Compare requested vs used resources and right-size.
- OpenCost (incubating) allocates cluster cost to namespaces, workloads and labels from usage and
cloud prices. Labels such as
teamandcost-centermake the breakdown useful.
Scenarios
A trace links every hop of one request and shows each span's duration, so the slow service stands out. Logs show what each service printed but need manual correlation. Node CPU metrics describe hosts, not a request path. Events describe object lifecycle changes such as scheduling and image pulls.
kubectl top reads the Resource Metrics API (metrics.k8s.io), which metrics-server provides. kube-state-metrics
exposes object state for Prometheus, not usage through that API. Alertmanager only routes alerts, and the OTel
Collector moves telemetry to backends.
OpenTelemetry is the vendor-neutral instrumentation standard, and the Collector can export to any backend, so a switch is a configuration change. Jaeger's own clients are retired and tie you to one backend. Prometheus exporters produce metrics, not traces. Fluentd is a log pipeline.
Further reading
Domain 4 · Cloud native architecture
12% of the exam. Observability signals and tools, the principles behind cloud native design and autoscaling, the CNCF projects and their maturity, and how the community governs Kubernetes.
Cloud native principles
What "cloud native" means in the CNCF definition, microservices vs monoliths, immutability and declarative APIs, autoscaling with HPA, VPA, Cluster Autoscaler and KEDA, and serverless with Knative.