Service mesh
What a service mesh adds to Kubernetes networking, data plane vs control plane, sidecar vs Istio ambient mode, mTLS and workload identity, traffic management, Istio and Linkerd, and the Gateway API GAMMA initiative.
Exam tasks: 2.1 (networking: service mesh)
The decision: does this requirement (encrypt all service-to-service traffic, retry failed calls, shift 5% of traffic to a new version, see per-call latency) need a service mesh, and if so, which data plane model?
What a mesh adds
Kubernetes Services give you discovery and simple load balancing. A service mesh moves the rest of service-to-service concerns out of application code and into infrastructure:
| Area | What the mesh does | Without a mesh |
|---|---|---|
| Security | Mutual TLS between workloads, identity-based authorization policies | Each app implements TLS and auth itself |
| Traffic management | Retries, timeouts, circuit breaking, weighted routing, mirroring, fault injection | Client libraries in every language |
| Observability | Golden metrics (requests, errors, latency) per call, distributed tracing headers, access logs | Instrument each service by hand |
- The mesh handles east-west traffic (service to service inside the cluster). North-south traffic still enters through a Gateway or Ingress, which the mesh can also secure.
- Apps don't change. The mesh intercepts traffic transparently.
Data plane and control plane
- The data plane is the set of proxies that actually carry traffic.
- The control plane programs those proxies, issues workload certificates and watches the Kubernetes API for Services and endpoints. It is not in the request path.
Sidecar vs ambient
| Sidecar | Istio ambient | |
|---|---|---|
| Where the proxy runs | One proxy container in every Pod | ztunnel per node (DaemonSet) for L4, optional waypoint proxies for L7 |
| Joining the mesh | Inject the sidecar, then restart Pods | Label the namespace; no Pod restart |
| Resource cost | Proxy CPU and memory per Pod | Shared per node; L7 cost only where waypoints are deployed |
| L7 features | Always available | Only for workloads behind a waypoint |
| Status | The classic model for Istio and Linkerd | GA since Istio 1.24 |
- ztunnel handles mTLS, L4 authorization and TCP telemetry. A waypoint (an Envoy proxy, usually per namespace) adds HTTP routing, retries and L7 policy.
- Kubernetes also has native sidecar containers (init containers with
restartPolicy: Always), which start before and stop after the app container. Meshes use them to fix startup and Job-completion ordering problems.
Exam signal
"Add a mesh without injecting a proxy into every Pod" or "reduce per-Pod proxy overhead" points to ambient mode (ztunnel plus waypoints). "Each Pod has an extra container that intercepts its traffic" describes the sidecar model.
mTLS and workload identity
- In plain TLS only the server proves who it is. In mutual TLS, both sides present certificates, so the server also knows which workload is calling.
- The mesh control plane acts as a certificate authority. It issues short-lived certificates to each workload, usually tied to the Pod's ServiceAccount, and rotates them automatically.
- Istio identities follow SPIFFE (
spiffe://<trust-domain>/ns/<namespace>/sa/<serviceaccount>). SPIFFE and its reference implementation SPIRE are CNCF graduated. - Authorization policies can then say "only the
checkoutServiceAccount may callbilling", which is stronger than IP-based rules because Pod IPs change. - Linkerd turns on mTLS automatically for all meshed TCP traffic. Istio supports permissive mode (accept plain text and mTLS during migration) and strict mode (mTLS only).
NetworkPolicy as encryption
A NetworkPolicy decides whether two Pods may connect. It doesn't encrypt anything and doesn't know workload identity. "Encrypt all Pod-to-Pod traffic and verify both sides" needs mTLS from a service mesh (or a CNI with transparent encryption), not a NetworkPolicy.
Traffic management
- Retries and timeouts: the proxy retries failed calls and gives up after a deadline, so apps don't need their own retry logic.
- Circuit breaking: stop sending traffic to an instance that keeps failing (outlier detection), so one bad replica doesn't drag down callers.
- Weighted routing: send 90% to
v1and 10% tov2for a canary, independent of replica counts. - Traffic mirroring: copy live requests to a new version and discard its responses.
- Fault injection: add delays or errors on purpose to test resilience.
The projects
| Istio | Linkerd | |
|---|---|---|
| CNCF status | Graduated (2023) | Graduated (2021), the first service mesh to graduate |
| Data plane | Envoy (sidecar), or ztunnel plus Envoy waypoints (ambient) | linkerd2-proxy, a small purpose-built Rust proxy (sidecar) |
| Strengths | Very broad feature set, ambient mode, multi-cluster | Simplicity, low overhead, mTLS on by default |
- Envoy (CNCF graduated) is a general-purpose L7 proxy. It's the data plane of Istio and of many Gateway API controllers. It's a proxy, not a mesh by itself.
- Cilium offers mesh features (mTLS, L7 policy) using eBPF and per-node Envoy, without sidecars.
Standards: from SMI to GAMMA
Legacy: use Gateway API (GAMMA) instead
The Service Mesh Interface (SMI) was a CNCF sandbox spec for a common mesh API (traffic split, access control, metrics). It was archived in 2023. Its role passed to the Gateway API GAMMA initiative.
- GAMMA (Gateway API for Mesh Management and Administration) uses the same
HTTPRouteyou use for ingress, but attaches it to a Service instead of a Gateway. The route then applies to traffic going to that Service inside the mesh. - Mesh support is part of the Gateway API standard channel since v1.1. Istio, Linkerd and Cilium implement it.
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
name: recommender-canary
namespace: shop
spec:
parentRefs:
- group: ""
kind: Service
name: recommender
port: 8080
rules:
- backendRefs:
- name: recommender-v1
port: 8080
weight: 90
- name: recommender-v2
port: 8080
weight: 10Scenarios
A service mesh gives every workload a certificate and enforces mutual TLS in the proxies, so traffic is encrypted and both sides are authenticated with no code changes. NetworkPolicies restrict connections but don't encrypt or identify workloads. An Ingress controller terminates TLS for incoming north-south traffic only. A shared API key needs code changes and doesn't identify individual callers.
Ambient mode replaces per-Pod sidecars with a per-node ztunnel for L4 and mTLS, and enrolling a namespace doesn't require Pod restarts. Sidecar injection still adds a container per Pod and needs restarts. Linkerd does use a proxy, in a sidecar. A generic NGINX DaemonSet doesn't provide mesh identity or mTLS.
GAMMA extends Gateway API to east-west traffic: an HTTPRoute whose parent is a Service configures mesh routing in a portable way, and Istio, Linkerd and Cilium implement it. SMI is archived. Ingress annotations are controller-specific and only cover incoming traffic. Raw Envoy config is tied to one data plane.
Further reading
Kubernetes networking
The Kubernetes networking model, CNI plugins, Service types and kube-proxy, cluster DNS names, Ingress vs Gateway API, and NetworkPolicies for KCNA.
Kubernetes security
The 4Cs of cloud native security, the API request path (authentication, authorization, admission), RBAC and ServiceAccounts, Pod Security Standards, Secrets, and image and supply chain security for KCNA.