CRDs and operators
Writing a CustomResourceDefinition with a schema, versions and printer columns, working with custom resources through kubectl, and installing and checking an operator.
Exam tasks: 1.8 (understand CRDs, install and configure operators)
The decision: what does a CRD add to the API, what does it take for a custom resource to actually do something, and how do you install, inspect and configure the operator that makes it happen?
CRD, custom resource, controller
| Piece | What it is | Without it |
|---|---|---|
| CRD | A cluster-scoped object that registers a new resource type, its schema and versions | kubectl apply of the custom kind fails with "no matches for kind" |
| Custom resource (CR) | An instance of that type, stored in etcd like any object | Nothing to act on |
| Controller | Code that watches the CRs and reconciles the real world toward .spec, reporting in .status | CRs are stored and listed, but nothing happens |
| Operator | A controller (plus its CRDs) that encodes how to run one piece of software: install, upgrade, back up, fail over |
Writing a CRD
apiVersion: apiextensions.k8s.io/v1
kind: CustomResourceDefinition
metadata:
name: backups.vault.lab.io # must be <plural>.<group>
spec:
group: vault.lab.io
scope: Namespaced # or Cluster
names:
plural: backups
singular: backup
kind: Backup
shortNames: ["bk"]
versions:
- name: v1
served: true
storage: true # exactly one version is the storage version
schema:
openAPIV3Schema:
type: object
properties:
spec:
type: object
required: ["schedule", "target"]
properties:
schedule:
type: string
target:
type: string
keepLast:
type: integer
minimum: 1
default: 7
x-kubernetes-validations:
- rule: "self.target.startsWith('pvc/')"
message: "target must look like pvc/NAME"
status:
type: object
properties:
lastRun:
type: string
subresources:
status: {} # .status is written through /status only
additionalPrinterColumns:
- name: Schedule
type: string
jsonPath: .spec.schedule
- name: Last-Run
type: string
jsonPath: .status.lastRunapiextensions.k8s.io/v1requires a structural schema for every version. Unknown fields are pruned unless you mark a subtreex-kubernetes-preserve-unknown-fields: true.x-kubernetes-validationsholds CEL rules the API server checks on create and update. Defaults (default:) are applied on read and write.- Several versions can be
served, but only one isstorage. Conversion between versions usesNone(fields must be compatible) or a conversion webhook.
Deleting a CRD
Deleting the CustomResourceDefinition deletes every custom resource of that type in every namespace. If a task says to remove an operator but keep its data, delete the operator Deployment, not the CRD.
Working with custom resources
apiVersion: vault.lab.io/v1
kind: Backup
metadata:
name: ledger-nightly
namespace: ledger
spec:
schedule: "0 2 * * *"
target: pvc/ledger-datak get crd | grep vault.lab.io
k api-resources --api-group=vault.lab.io # kind, short names, namespaced or not
k explain backup.spec # schema-driven docs, same as built-in kinds
k apply -f ledger-nightly.yaml
k get bk -n ledger # printer columns show up here
k get backups.vault.lab.io -A -o yaml # fully qualified name avoids clashes
k describe backup ledger-nightly -n ledger # status, events from the controllerExam signal
Tasks often ask you to find something about a CRD rather than write one: its group, its versions, whether
it's namespaced, or which fields its spec has. k get crd <name> -o yaml, k api-resources --api-group=... and
k explain <kind>.spec --recursive answer those without reading the operator's docs.
Installing an operator
| Method | Looks like | Check with |
|---|---|---|
| Plain manifests | k apply -f https://.../release/v1.4.0/install.yaml (CRDs, namespace, RBAC, Deployment) | k get crd, k get deploy -n <operator-ns> |
| Helm chart | helm install ... --set crds.enabled=true or a chart that ships crds/ | helm list -A, k get crd |
| Operator Lifecycle Manager | A Subscription to a catalog; OLM installs and upgrades the operator | k get csv -A, k get subscription -A |
An operator install always gives you three things to verify:
- CRDs registered:
k get crd | grep <group>. - RBAC: a ServiceAccount with a ClusterRole that lets it watch its CRs and manage what it creates.
- The controller Deployment running:
k get pods -n <operator-ns>, andk logs deploy/<name> -n <operator-ns>when CRs don't reconcile.
- Configure the operator itself through its Deployment arguments, a ConfigMap or Helm values (for example which namespaces it watches). Configure what it manages through the custom resources.
- Install CRDs before CRs. Applying both in one file can fail on the first run because the CRD isn't established
yet; rerunning the apply, or
k wait --for=condition=Established crd/<name>, fixes it. - Helm installs CRDs from a chart's
crds/folder only on first install and never upgrades or deletes them. Upgrade those CRDs withkubectl applyas the operator's docs describe.
Custom resources that never change status
If CRs are accepted but .status stays empty and nothing gets created, the CRD is fine and the controller isn't:
it isn't running, is watching other namespaces, or lacks RBAC. Read the operator's logs for forbidden errors
before touching the CR.
Scenarios
"No matches for kind" comes from API discovery: the group and version combination isn't served. With the CRD present, the version name (for example v1beta1 vs v1) is the usual mismatch. A stopped operator still lets you create CRs, a missing namespace gives a NotFound error, and RBAC failures say forbidden.
Stopping the controller stops reconciliation while the CRs stay stored in etcd. Deleting the CRD would cascade to every Database object. CRDs are cluster-scoped, so they aren't inside any namespace, but deleting the namespace removes the operator anyway and any CRs in that namespace. Deleting and restoring CRs risks the operator tearing down the actual databases.
Drill
kubectl config use-context cka-crd
- An operator's install manifest is at
/opt/op/beacon-operator.yaml. Install it and wait until its CRDbeacons.signal.lab.iois established and the operator Pod is running in namespacebeacon-system. - Create a
Beaconnamednorthin namespacerelaywithspec.intervalSecondsset to 30 andspec.endpointset tohttp://relay-api.relay:8080. - Write the CRD's scope (Namespaced or Cluster) and its storage version to
/opt/answers/beacon.txt.
k apply -f /opt/op/beacon-operator.yaml
k wait --for=condition=Established crd/beacons.signal.lab.io --timeout=60s
k get pods -n beacon-system # Running
k explain beacon.spec # confirm field names
k api-resources --api-group=signal.lab.io # shortnames, namespacedk create ns relay --dry-run=client -o yaml | k apply -f -
cat <<'EOF' | k apply -f -
apiVersion: signal.lab.io/v1
kind: Beacon
metadata:
name: north
namespace: relay
spec:
intervalSeconds: 30
endpoint: http://relay-api.relay:8080
EOF(Use the version from k api-resources if it isn't v1.)
k get crd beacons.signal.lab.io \
-o jsonpath='{.spec.scope}{"\n"}{range .spec.versions[?(@.storage==true)]}{.name}{"\n"}{end}' \
> /opt/answers/beacon.txt
cat /opt/answers/beacon.txtVerify:
k get beacon north -n relay -o yaml # status filled in by the operator after a few seconds
k logs -n beacon-system deploy/beacon-operator --tail=20Further reading
Extension interfaces (CRI, CNI, CSI)
What the kubelet delegates to container runtimes, network plugins and storage drivers, where each one is configured on a node, and how to tell which plugin is installed.
Domain 2 · Workloads and scheduling
15% of the exam. Rolling out and rolling back Deployments, injecting configuration, autoscaling, self-healing controllers, and controlling where Pods land and how much they may use.