← Certified Kubernetes Administrator (CKA)
Test yourself →

Cluster Architecture, Installation & Configuration

## Kubernetes Cluster Architecture

A Kubernetes cluster consists of a set of control plane components and worker nodes. The control plane manages the cluster, while worker nodes run the actual applications (pods).

## Control Plane Components

These components make global decisions about the cluster (e.g., scheduling) and detect and respond to cluster events.

  • kube-apiserver: The front-end of the Kubernetes control plane. It exposes the Kubernetes API, which all other components and users interact with. It's the only control plane component directly exposed.
  • etcd: A consistent and highly available key-value store that stores all cluster data, configuration, and state. It's the single source of truth for the cluster.
  • kube-scheduler: Watches for newly created Pods with no assigned node and selects a node for them to run on, based on resource requirements, policies, and affinity/anti-affinity specifications.
  • kube-controller-manager: Runs various controller processes. Each controller is a separate process, but they are compiled into a single binary to reduce complexity. Examples include the Node Controller, Replication Controller, Endpoints Controller, and Service Account & Token Controllers.
  • cloud-controller-manager: (Optional) Interacts with the underlying cloud provider to manage cloud-specific resources like load balancers, persistent volumes, and node lifecycle.

## Worker Node Components

These components run on each worker node and are responsible for maintaining running pods and providing the Kubernetes runtime environment.

  • kubelet: An agent that runs on each node in the cluster. It ensures that containers are running in a Pod as specified by the control plane. It registers the node with the API server.
  • kube-proxy: A network proxy that runs on each node. It maintains network rules on nodes, allowing network communication to your Pods from network sessions inside or outside the cluster. It handles Service abstraction.
  • Container Runtime: The software responsible for running containers. Kubernetes supports various runtimes through the Container Runtime Interface (CRI), such as containerd, CRI-O, or Docker (via containerd shim).

## Cluster Installation & Configuration with kubeadm

kubeadm is a tool designed to easily bootstrap a minimum viable Kubernetes cluster. It handles the setup of control plane components and allows worker nodes to join.

  • `kubeadm init`: Initializes a control plane node. It performs pre-flight checks, generates certificates, sets up the kube-apiserver, etcd, kube-scheduler, and kube-controller-manager as static pods or systemd services.
  • `kubeadm join`: Allows worker nodes to join an existing cluster. It takes a token and the control plane's IP address to securely connect.
  • Pod Network Add-on (CNI): After `kubeadm init`, a Container Network Interface (CNI) plugin (e.g., Calico, Flannel) must be installed to enable pod-to-pod communication. Without it, pods will remain in a `Pending` state or fail to communicate.
  • kubeconfig: A file that stores cluster access information (cluster details, user credentials, context) for `kubectl` and other tools.
  • The **kube-apiserver** is the central management entity and entry point for the Kubernetes control plane.
  • **etcd** is the distributed key-value store that holds the entire state and configuration of the Kubernetes cluster.
  • The **kubelet** is the primary agent on each worker node, ensuring containers run in pods.
  • **kube-proxy** manages network rules to enable communication to and from Kubernetes Services.
  • **kubeadm** is the recommended tool for bootstrapping a Kubernetes cluster.
  • A **CNI plugin** is essential for inter-pod networking after cluster initialization.
  • Control plane components often run as **static pods** or systemd services.
  • **kubeconfig** files contain the necessary credentials and configuration to connect to a cluster.
What is the primary function of the **kube-apiserver**?
It exposes the Kubernetes API and is the front-end for the control plane, handling all communication.
tap to reveal
Which component stores the entire state of the Kubernetes cluster?
**etcd**
tap to reveal
What is the role of the **kubelet** on a worker node?
It's an agent that ensures containers are running in a Pod as specified by the control plane.
tap to reveal
What does `kubeadm init` accomplish?
It initializes a Kubernetes control-plane node, setting up all necessary control plane components and certificates.
tap to reveal
Why is a **CNI plugin** necessary for a Kubernetes cluster?
It provides pod networking, allowing pods to communicate with each other across nodes.
tap to reveal
What is the purpose of **kube-proxy**?
It maintains network rules (e.g., iptables or IPVS) on nodes to enable network communication to Pods via Services.
tap to reveal
Where are the credentials and configuration for accessing a Kubernetes cluster stored?
In **kubeconfig** files.
tap to reveal
Name two components of the Kubernetes control plane.
kube-apiserver, etcd, kube-scheduler, kube-controller-manager (any two).
tap to reveal

Workloads & Scheduling

## Workloads & Scheduling

Kubernetes Pods are the smallest deployable units, encapsulating one or more containers, storage resources, a unique network IP, and options that govern how the containers run. Multi-container Pods often use patterns like sidecar containers for logging or monitoring, and init containers for setup tasks that must complete before the main application containers start.

Pod Management Controllers

Deployments are the primary way to manage stateless applications. They define the desired state for a set of Pods and their underlying ReplicaSets. Deployments handle rolling updates and rollbacks, allowing zero-downtime updates and easy reversion to previous versions. ReplicaSets ensure a specified number of identical Pods are running at all times. DaemonSets ensure that all (or some) nodes run a copy of a specific Pod, useful for cluster-level services like logging agents or monitoring daemons.

Batch Workloads

For tasks that run to completion, Kubernetes provides Jobs. A Job creates one or more Pods and ensures that a specified number of them successfully terminate. CronJobs automate Jobs on a recurring schedule, similar to cron in Linux, enabling scheduled tasks like backups or report generation.

Scheduling and Resource Management

Kubernetes' scheduler places Pods on nodes based on various constraints. Resource Requests define the minimum CPU and memory a Pod needs, ensuring it gets scheduled on a node with sufficient capacity. Resource Limits define the maximum CPU and memory a Pod can consume, preventing resource exhaustion on a node.

Advanced scheduling includes Node Selectors for simple node assignment based on labels. Node Affinity provides more flexible rules (e.g., "prefer this node," "must run on this node"). Taints on nodes repel Pods unless those Pods have a matching Toleration. This is crucial for dedicating nodes or handling special hardware. Pod Disruption Budgets (PDBs) ensure that a minimum number of Pods for a given application remain running during voluntary disruptions (e.g., node maintenance), enhancing application availability. Liveness Probes detect if a container is running correctly and restart it if it fails. Readiness Probes determine if a container is ready to serve traffic, removing it from service endpoints if not. Startup Probes delay liveness and readiness checks until the application has started, preventing premature restarts for slow-starting apps.

  • Pods are the smallest deployable unit, encapsulating containers, storage, and network for an application.
  • Deployments manage stateless applications, enabling rolling updates and rollbacks via ReplicaSets.
  • DaemonSets ensure a single Pod runs on every (or selected) node, useful for cluster-level services.
  • Jobs run tasks to completion, while CronJobs schedule them periodically.
  • Resource Requests guarantee minimum resources for scheduling; Resource Limits cap maximum consumption.
  • Taints repel Pods from nodes unless the Pod has a matching Toleration.
  • Liveness, Readiness, and Startup Probes ensure container health and proper traffic routing.
  • Node Affinity and Node Selectors control where Pods are scheduled based on node labels.
  • Pod Disruption Budgets (PDBs) maintain application availability during voluntary node disruptions.
What is the primary purpose of a Kubernetes Pod?
To encapsulate one or more containers, storage, network IP, and options, forming the smallest deployable unit.
tap to reveal
How do you manage stateless applications with rolling updates and rollbacks in Kubernetes?
Using a **Deployment**.
tap to reveal
What Kubernetes object ensures that exactly one copy of a Pod runs on every (or selected) node?
A **DaemonSet**.
tap to reveal
Explain the difference between `requests` and `limits` for CPU/memory in a Pod.
**Requests** guarantee minimum resources for scheduling; **Limits** cap maximum consumption to prevent resource exhaustion.
tap to reveal
What is the function of a **Taint** on a node?
A Taint repels Pods from being scheduled on that node unless the Pod has a matching **Toleration**.
tap to reveal
What is the purpose of a **Liveness Probe**?
To detect if a container is running correctly and restart it if it fails, ensuring application health.
tap to reveal
How do you run a one-off task to completion in Kubernetes?
Using a **Job**.
tap to reveal
What does a **Pod Disruption Budget (PDB)** achieve?
It ensures that a minimum number of Pods for an application remain available during voluntary disruptions (e.g., node maintenance).
tap to reveal

Services & Networking

## Services

A Service is a Kubernetes object that defines a logical set of Pods and a policy by which to access them. Services provide a stable network endpoint (IP address and DNS name) for a dynamic set of Pods, which are often ephemeral. Services use label selectors to identify the Pods they target.

There are several Service Types:

  • ClusterIP: The default type. Exposes the Service on an internal IP address within the cluster. It's only reachable from inside the cluster.
  • NodePort: Exposes the Service on each Node's IP at a static port (the `NodePort`). This makes the Service accessible from outside the cluster via `<NodeIP>:<NodePort>`.
  • LoadBalancer: Exposes the Service externally using a cloud provider's load balancer. This type is only effective when running on a cloud provider that supports it.
  • ExternalName: Maps the Service to a DNS name by returning a `CNAME` record. No proxying is involved.

kube-proxy is a network proxy that runs on each node in the cluster, implementing the Service concept. It maintains network rules (e.g., iptables or IPVS) to forward traffic to the correct backend Pods. EndpointSlices are a more scalable and efficient way to track network endpoints for Services, especially in large clusters, replacing the older Endpoints resource.

## Ingress

Ingress is an API object that manages external access to services in a cluster, typically HTTP/HTTPS. It provides routing rules for traffic based on hostnames and paths to backend Services. An Ingress Controller is essential for Ingress resources to function; it's a specialized pod that watches the Ingress API and configures a load balancer or proxy accordingly (e.g., Nginx, HAProxy). Ingress can also be configured for TLS termination using Kubernetes Secrets.

## Network Policies

NetworkPolicy is an API object that allows you to specify how groups of Pods are allowed to communicate with each other and other network endpoints. By default, Pods are non-isolated, meaning they can communicate freely. Network Policies provide pod-level isolation. They use podSelectors to select the target Pods and define rules for Ingress (incoming traffic) and Egress (outgoing traffic). Rules can specify communication based on Pods, Namespaces, or IP blocks (`ipBlock`). A Container Network Interface (CNI) plugin that supports NetworkPolicy (e.g., Calico) is required for them to be enforced.

## DNS

CoreDNS is the default DNS server in Kubernetes, responsible for Service Discovery. Pods can discover Services using DNS names like `my-service.my-namespace.svc.cluster.local` (fully qualified), `my-service.my-namespace` (within the same cluster), or `my-service` (within the same namespace). Each Pod gets its own DNS configuration, inheriting from the cluster's CoreDNS.

  • Services provide stable network endpoints for dynamic Pods using label selectors.
  • ClusterIP is for internal cluster access, NodePort exposes services via node IPs, LoadBalancer uses cloud-provider LBs, and ExternalName maps to an external DNS.
  • kube-proxy implements the Service concept by managing network rules (iptables/IPVS) on each node.
  • Ingress manages external HTTP/HTTPS access to services and requires an Ingress Controller to function.
  • NetworkPolicy objects control pod-to-pod communication based on ingress/egress rules and require a CNI plugin that supports them.
  • CoreDNS is the cluster's DNS server, enabling service discovery via DNS names.
  • Pods are non-isolated by default; NetworkPolicy enables isolation.
  • Ingress can handle TLS termination for external traffic using Kubernetes Secrets.
What are the four main Service types in Kubernetes?
ClusterIP, NodePort, LoadBalancer, ExternalName.
tap to reveal
Which Kubernetes component implements the Service concept by managing network rules on each node?
kube-proxy.
tap to reveal
What is required for an Ingress resource to function correctly and route external traffic to Services?
An Ingress Controller (e.g., Nginx Ingress Controller).
tap to reveal
How do you restrict network access between pods in Kubernetes?
By creating NetworkPolicy resources.
tap to reveal
What is the fully qualified domain name (FQDN) format for a Service named `my-service` in the `default` namespace?
`my-service.default.svc.cluster.local`.
tap to reveal
What is the default isolation state for pods in Kubernetes regarding network communication?
Pods are non-isolated by default, meaning they can communicate with any other pod.
tap to reveal
Which Service type is used to expose a Service on a static port on each Node's IP address?
NodePort.
tap to reveal
What is the purpose of EndpointSlices?
To provide a more scalable and efficient way to track network endpoints for Services, especially in large clusters.
tap to reveal

Storage

## Kubernetes Storage Overview

Kubernetes abstracts underlying storage with PersistentVolumes (PVs), PersistentVolumeClaims (PVCs), and StorageClasses, crucial for stateful applications.

## PersistentVolumes (PVs)

A PersistentVolume (PV) is a cluster-wide storage resource, provisioned by an administrator or dynamically. It represents actual storage (e.g., NFS, cloud disks) and is independent of Pods. PVs define access modes and a reclaim policy.

## PersistentVolumeClaims (PVCs)

A PersistentVolumeClaim (PVC) is a user's request for storage. Kubernetes binds a PVC to an available PV that matches its requirements (size, access mode). If no suitable PV exists, the PVC remains pending or triggers dynamic provisioning.

## StorageClasses & Dynamic Provisioning

A StorageClass defines a "class" of storage, specifying a provisioner (e.g., `kubernetes.io/aws-ebs`) and parameters. It enables dynamic provisioning: when a PVC requests a StorageClass, the provisioner automatically creates a new PV on the backend storage to fulfill the claim, automating storage setup.

## Access Modes & Reclaim Policy

Access Modes dictate how storage can be mounted:

  • ReadWriteOnce (RWO): Mounted read-write by a single node.
  • ReadOnlyMany (ROX): Mounted read-only by many nodes.
  • ReadWriteMany (RWX): Mounted read-write by many nodes (e.g., NFS).

The Reclaim Policy defines PV behavior after its PVC is deleted:

  • Retain: PV and data preserved; manual cleanup required. Safest for data.
  • Delete: Underlying storage asset and PV are deleted. Common with dynamic provisioning.

## Using Storage in Pods

Pods use persistent storage by referencing a PVC in their `volumes` section and then mounting that volume into containers via `volumeMounts`. The Pod interacts with the PVC, which is bound to a PV.

```yaml

apiVersion: v1

kind: Pod

metadata:

name: my-app

spec:

containers:

  • name: my-container

image: nginx

volumeMounts:

  • name: my-storage

mountPath: /data

volumes:

  • name: my-storage

persistentVolumeClaim:

claimName: my-pvc

```

  • **PersistentVolumes (PVs)** are cluster-wide storage resources, independent of Pods.
  • **PersistentVolumeClaims (PVCs)** are user requests for storage that consume PVs.
  • **StorageClasses** enable **dynamic provisioning** of PVs based on PVC requests.
  • Common **Access Modes** include `ReadWriteOnce`, `ReadOnlyMany`, and `ReadWriteMany`.
  • The **Retain** reclaim policy prevents data loss by preserving the PV after PVC deletion.
  • Pods reference a **PVC** in their `volumes` section to use persistent storage.
  • Dynamic provisioning automates the creation of underlying storage resources.
  • A PVC must bind to a suitable PV (or have one dynamically created) before it can be used by a Pod.
What is a PersistentVolume (PV)?
A cluster-wide resource representing a piece of storage provisioned for use by Kubernetes.
tap to reveal
What is a PersistentVolumeClaim (PVC)?
A request for storage by a user, which binds to an available PersistentVolume (PV).
tap to reveal
What is the primary purpose of a StorageClass?
To define classes of storage and enable dynamic provisioning of PersistentVolumes (PVs).
tap to reveal
Name the three main Access Modes for PersistentVolumes.
ReadWriteOnce (RWO), ReadOnlyMany (ROX), ReadWriteMany (RWX).
tap to reveal
What is the "Retain" reclaim policy for a PV?
The PV and its data are preserved after the PVC is deleted, requiring manual cleanup.
tap to reveal
How does a Pod specify that it wants to use persistent storage?
By referencing a PersistentVolumeClaim (PVC) in its `volumes` section.
tap to reveal
What is dynamic provisioning in Kubernetes storage?
The automatic creation of a PersistentVolume (PV) by a StorageClass provisioner when a PVC requests storage.
tap to reveal
What happens if a PVC cannot find a suitable PV to bind to?
The PVC will remain in a "Pending" state until a matching PV becomes available or is dynamically provisioned.
tap to reveal

Application Troubleshooting

## Application Troubleshooting in Kubernetes

Troubleshooting applications in Kubernetes is a critical skill for a CKA. It involves systematically diagnosing issues from the pod level up to service exposure.

## Common Pod Troubleshooting Scenarios

  • Pending: A pod stuck in `Pending` state often indicates a scheduling issue. Check resource requests (CPU/memory) against available node capacity using `kubectl describe pod <pod-name>` for events like `FailedScheduling`. Also, verify node selectors, taints and tolerations, or unbound PersistentVolumeClaims (PVCs).
  • CrashLoopBackOff: This state means your container is repeatedly starting and crashing. The most common causes are application errors (e.g., incorrect configuration, missing environment variables, database connection issues), an incorrect command or args in the pod definition, or missing dependencies. Use `kubectl logs <pod-name>` to inspect application output, and `kubectl describe pod <pod-name>` for restart counts and events.
  • ImagePullBackOff: Occurs when Kubernetes cannot pull the container image. Verify the image name and tag are correct. If using a private registry, ensure imagePullSecrets are correctly configured and have valid credentials. Network connectivity issues to the registry can also cause this.
  • Running (but unhealthy): If a pod is `Running` but your application isn't working, check Liveness and Readiness probes. A failing liveness probe will cause the container to restart; a failing readiness probe will remove the pod from service endpoints. Use `kubectl describe pod` to see probe status and events.

## Essential Troubleshooting Commands

  • `kubectl describe <resource-type> <resource-name>`: Provides a wealth of information, including status, events, conditions, resource requests/limits, and associated objects. Indispensable for initial diagnosis.
  • `kubectl logs <pod-name> [-c <container-name>]`: Retrieves logs from a container, crucial for understanding application behavior and errors. Use `-f` for follow, `--previous` for logs from a previous instance.
  • `kubectl exec -it <pod-name> -- bash`: Allows you to execute commands inside a running container, enabling interactive debugging, checking file systems, or running diagnostic tools.
  • `kubectl get events`: Shows cluster-level events, which can reveal issues affecting multiple pods or nodes.
  • `kubectl top pod <pod-name>` / `kubectl top node <node-name>`: Displays resource usage (CPU/memory) for pods and nodes, helping identify resource bottlenecks.
  • `kubectl get pods -o wide`: Shows which node a pod is running on, useful for node-specific issues.

## Service and Network Troubleshooting

If pods are healthy but inaccessible, investigate Services. Use `kubectl describe service <service-name>` to check its selector and ensure it matches the labels of your target pods. Verify that the Service's targetPort matches the container port. Check `kubectl get endpoints <service-name>` to confirm that healthy pods are correctly registered as endpoints. Finally, review any NetworkPolicies that might be inadvertently blocking traffic.

  • `kubectl describe` is the primary command for detailed pod status, events, and conditions.
  • `CrashLoopBackOff` typically indicates an application error, incorrect command, or missing dependencies.
  • `ImagePullBackOff` points to issues with the image name, registry access, or `imagePullSecrets`.
  • A `Pending` pod often means insufficient resources, node selectors/taints, or unbound PVCs.
  • `kubectl logs` is crucial for viewing application output and identifying runtime errors.
  • Failing Liveness probes trigger container restarts, while failing Readiness probes remove pods from service endpoints.
  • `kubectl exec -it <pod-name> -- bash` allows interactive debugging inside a container.
  • For Service issues, verify the selector matches pod labels and `targetPort` aligns with the container port.
What is the primary command to inspect detailed pod status, events, and conditions?
`kubectl describe pod <pod-name>`
tap to reveal
A pod is in `CrashLoopBackOff` state. What are common causes?
Application errors, incorrect `command`/`args` in the pod definition, or missing configuration/dependencies.
tap to reveal
How do you access logs from a specific container within a pod?
`kubectl logs <pod-name> -c <container-name>`
tap to reveal
What does `ImagePullBackOff` indicate?
Kubernetes failed to pull the container image, often due to a wrong image name/tag, private registry authentication issues, or network problems.
tap to reveal
What is the purpose of a Liveness probe?
To determine if an application within a container is running and healthy; if it fails, the container is restarted.
tap to reveal
How can you interactively debug inside a running container?
`kubectl exec -it <pod-name> -- bash` (or `sh`)
tap to reveal
A pod is stuck in `Pending`. What are potential causes?
Insufficient resources (CPU/memory), node selectors/taints, or an unbound PersistentVolumeClaim (PVC).
tap to reveal
What command helps verify if a Service's selector is correctly matching pods and if endpoints are registered?
`kubectl describe service <service-name>` and `kubectl get endpoints <service-name>`
tap to reveal

Cluster Component Troubleshooting

## Cluster Component Troubleshooting

Troubleshooting Kubernetes cluster components requires a systematic approach, often starting with identifying the affected component and then examining its logs, status, and configuration. The primary components include the Control Plane (kube-apiserver, etcd, kube-scheduler, kube-controller-manager) and Worker Node components (kubelet, kube-proxy, container runtime).

Control Plane Components

  • `kube-apiserver`: This is the central hub. If it's down, `kubectl` commands will fail.
  • Diagnosis: Check `systemctl status kube-apiserver`, `journalctl -u kube-apiserver`, and its static pod logs (e.g., `/etc/kubernetes/manifests/kube-apiserver.yaml` if deployed as a static pod). Look for network issues, certificate problems, or etcd connectivity errors.
  • `etcd`: The cluster's backend store. If `etcd` is unhealthy or unreachable, the API server cannot function.
  • Diagnosis: Use `etcdctl --endpoints=<etcd-ip>:2379 member list` and `etcdctl --endpoints=<etcd-ip>:2379 endpoint health`. Check `systemctl status etcd` and `journalctl -u etcd`. Ensure the data directory (`--data-dir`) is accessible and not full.
  • `kube-scheduler`: Responsible for assigning pods to nodes. If down, new pods will remain in `Pending` state.
  • Diagnosis: Check `systemctl status kube-scheduler` and `journalctl -u kube-scheduler`. Look for errors related to API server connectivity or resource constraints.
  • `kube-controller-manager`: Runs various controllers (e.g., Node, Replication, Endpoint). If down, many cluster operations (like creating services, managing endpoints) will fail or be delayed.
  • Diagnosis: Check `systemctl status kube-controller-manager` and `journalctl -u kube-controller-manager`. Similar to the scheduler, look for API server connectivity issues.

Worker Node Components

  • `kubelet`: The agent that runs on each node, responsible for registering the node and managing pods.
  • Diagnosis: Check `systemctl status kubelet` and `journalctl -u kubelet`. Look for errors related to container runtime connectivity, CNI issues, or incorrect `--register-node` configuration. A node might appear `NotReady` if `kubelet` is unhealthy.
  • `kube-proxy`: Maintains network rules on nodes, enabling communication to pods and services.
  • Diagnosis: Check `systemctl status kube-proxy` and `journalctl -u kube-proxy`. Verify `iptables -L -t nat` or `ipvsadm -L` rules. Issues often manifest as service connectivity problems.
  • Container Runtime (e.g., `containerd`, `Docker`): Responsible for running containers.
  • Diagnosis: Check `systemctl status containerd` (or `docker`) and `journalctl -u containerd`. Pods will fail to start or remain in `ContainerCreating` if the runtime is unhealthy. Use `crictl ps` (for `containerd`) or `docker ps` to verify running containers.

General Troubleshooting Steps

1. Check Logs: Always start with `journalctl -u <component>` or `kubectl logs -n kube-system <pod-name>`.

2. Check Events: `kubectl get events -A` can provide high-level insights into what's failing.

3. Resource Usage: Ensure components have sufficient CPU/memory.

4. Networking: Verify connectivity between components (e.g., API server to etcd, kubelet to API server). Use `netstat`, `ss`, `ping`, `curl`.

5. Certificates: Ensure all components have valid and correctly configured TLS certificates.

  • `kube-apiserver` is the single point of entry; if it's down, `kubectl` fails.
  • `etcd` is the cluster's database; its health is critical for the entire control plane.
  • `kubelet` is the agent on each node; a `NotReady` node often indicates `kubelet` issues.
  • `kube-scheduler` failure results in new pods stuck in `Pending` state.
  • `kube-proxy` manages network rules for service connectivity via `iptables` or `ipvs`.
  • Always check `systemctl status <component>` and `journalctl -u <component>` first for component health.
  • `kubectl get events -A` provides a cluster-wide overview of recent issues.
  • Container runtime issues (e.g., `containerd`) prevent pods from entering `Running` state.
What is the primary indicator that the `kube-apiserver` might be down?
`kubectl` commands fail to connect to the cluster.
tap to reveal
How do you check the health of `etcd` members from the command line?
`etcdctl --endpoints=<etcd-ip>:2379 member list` and `etcdctl --endpoints=<etcd-ip>:2379 endpoint health`.
tap to reveal
What is the most common symptom if the `kube-scheduler` is not functioning correctly?
New pods remain in the `Pending` state indefinitely.
tap to reveal
Which command is typically used to check the status and logs of a systemd-managed Kubernetes component on a node?
`systemctl status <component>` and `journalctl -u <component>`.
tap to reveal
A node is showing `NotReady`. Which component is most likely having issues on that node?
The `kubelet` agent.
tap to reveal
What is the role of `kube-proxy` and what common tool does it use to implement its function?
`kube-proxy` maintains network rules for service communication, primarily using `iptables` (or `ipvs`).
tap to reveal
If pods are stuck in `ContainerCreating` state, which component should you investigate on the node?
The container runtime (e.g., `containerd` or `Docker`).
tap to reveal
How can you view cluster-wide events that might indicate component failures or issues?
`kubectl get events -A`.
tap to reveal

Cluster Maintenance & Upgrades

## Cluster Maintenance & Upgrades

Maintaining a healthy Kubernetes cluster involves regular upgrades and robust backup/restore strategies. The CKA exam focuses on manual upgrades using `kubeadm` and `etcd` snapshot management.

## Upgrading Kubernetes Clusters

Manual upgrades typically involve updating the `kubeadm`, `kubelet`, and `kubectl` binaries. It's crucial to upgrade these components to the same minor version across the cluster to ensure compatibility.

Upgrade Process for Control Plane Nodes:

1. Backup etcd: This is the most critical step before any upgrade. See the backup section below.

2. Drain the node: Prepare the control plane node for maintenance by evicting all pods: `kubectl drain <node-name> --ignore-daemonsets`.

3. Upgrade kubeadm: Update the `kubeadm` package to the target version (e.g., `apt-get install -y kubeadm=1.28.x-00` or `yum install -y kubeadm-1.28.x`).

4. Plan the upgrade: Check for potential issues: `kubeadm upgrade plan`.

5. Apply the upgrade: Execute the upgrade on the control plane: `kubeadm upgrade apply v1.28.x`.

6. Upgrade kubelet and kubectl: Update the `kubelet` and `kubectl` packages to the same target version (e.g., `apt-get install -y kubelet=1.28.x-00 kubectl=1.28.x-00`).

7. Restart kubelet: `systemctl restart kubelet`.

8. Uncordon the node: Allow pods to be scheduled again: `kubectl uncordon <node-name>`.

Upgrade Process for Worker Nodes:

1. Drain the node: `kubectl drain <node-name> --ignore-daemonsets`.

2. Upgrade kubeadm: Update the `kubeadm` package.

3. Run kubeadm upgrade node: This command updates the worker node's configuration: `kubeadm upgrade node`.

4. Upgrade kubelet and kubectl: Update the `kubelet` and `kubectl` packages.

5. Restart kubelet: `systemctl restart kubelet`.

6. Uncordon the node: `kubectl uncordon <node-name>`.

Repeat for all worker nodes.

## Backup and Restore Operations

The etcd data store holds the entire state of your Kubernetes cluster. Backing it up is paramount for disaster recovery.

Backup etcd:

Use the `etcdctl snapshot save` command. You'll need to specify the API version, etcd endpoints, and TLS certificates if etcd is secured (which it typically is in a production cluster).

```bash

# Ensure ETCDCTL_API is set to 3

export ETCDCTL_API=3

# Example command (adjust paths for your setup)

etcdctl snapshot save /tmp/etcd-snapshot.db \

--endpoints=https://127.0.0.1:2379 \

--cacert=/etc/kubernetes/pki/etcd/ca.crt \

--cert=/etc/kubernetes/pki/etcd/peer.crt \

--key=/etc/kubernetes/pki/etcd/peer.key

```

Restore etcd:

Restoring etcd involves stopping the Kubernetes control plane components, restoring the snapshot, and then restarting the components.

1. Stop Kubernetes control plane components: On the control plane node, stop the `kube-apiserver`, `kube-scheduler`, and `kube-controller-manager` static pods. This often involves moving their manifest files out of `/etc/kubernetes/manifests/` or stopping the `kubelet` service.

2. Restore the snapshot: Create a *new* data directory for etcd and restore the snapshot into it.

```bash

export ETCDCTL_API=3

etcdctl snapshot restore /tmp/etcd-snapshot.db \

--data-dir=/var/lib/etcd-restored-data

```

3. Update etcd static pod manifest: Modify the `etcd.yaml` manifest file (usually `/etc/kubernetes/manifests/etcd.yaml`) to point to the new `--data-dir` (`/var/lib/etcd-restored-data` in the example).

4. Restart Kubernetes control plane components: Move the manifest files back or restart the `kubelet` service. Verify cluster health.

Always ensure you have backups of your Kubernetes configuration files (`/etc/kubernetes/`) and certificates (`/etc/kubernetes/pki/`) as well.

    Security & RBAC

    ## Security & RBAC in Kubernetes

    Security in Kubernetes is multifaceted, covering authentication, authorization, network controls, and workload isolation. Role-Based Access Control (RBAC) is the primary authorization mechanism, defining what authenticated users and processes can do.

    ## Authentication and Authorization

    Authentication verifies the identity of a user or process (e.g., client certificates, service account tokens, OIDC). Once authenticated, Authorization determines if the identity is permitted to perform a requested action on a resource. The Kubernetes API server handles both.

    ## Role-Based Access Control (RBAC)

    RBAC uses four key objects:

    • Role: Defines a set of permissions *within a specific namespace*. Permissions are defined by verbs (e.g., `get`, `list`, `create`, `delete`) on resources (e.g., `pods`, `deployments`) within API groups.
    • ClusterRole: Similar to a Role, but defines permissions that are *cluster-scoped* (e.g., `nodes`) or apply to *non-namespaced resources*, or grant access across *all namespaces*.
    • RoleBinding: Grants the permissions defined in a Role to a user, group, or Service Account *within a specific namespace*.
    • ClusterRoleBinding: Grants the permissions defined in a ClusterRole to a user, group, or Service Account *cluster-wide*.

    Use `kubectl auth can-i <verb> <resource> --namespace <namespace>` to check permissions.

    ## Service Accounts

    Service Accounts provide an identity for processes running in pods. Each namespace has a default service account. Pods are automatically assigned the `default` service account of their namespace unless specified otherwise. A token for the service account is mounted into the pod at `/var/run/secrets/kubernetes.io/serviceaccount/token`. You can disable this with `automountServiceAccountToken: false` in the pod or service account specification.

    ## Secrets

    Secrets are used to store sensitive information like passwords, OAuth tokens, and SSH keys. They are Base64 encoded, not encrypted by default at rest (unless configured with an encryption provider). Secrets can be mounted as data volumes or exposed as environment variables within pods. Common types include `Opaque`, `kubernetes.io/dockerconfigjson`, and `kubernetes.io/tls`.

    ## Network Policies

    Network Policies specify how groups of pods are allowed to communicate with each other and with external network endpoints. They define ingress (incoming) and egress (outgoing) rules based on pod selectors and namespace selectors. Network Policies are enforced by a Network Policy Controller, which is typically part of the Container Network Interface (CNI) plugin (e.g., Calico, Cilium).

    ## Pod Security Standards (PSS)

    Pod Security Standards (PSS) replaced Pod Security Policies (PSPs) to enforce security best practices for pods. PSS defines three levels:

    • Privileged: Unrestricted, allowing known escalations.
    • Baseline: Minimally restrictive, preventing known escalations.
    • Restricted: Heavily restricted, following current hardening best practices.

    PSS are enforced by the PodSecurity admission controller, which can be configured to `enforce`, `audit`, or `warn` based on security levels for specific namespaces using labels like `pod-security.kubernetes.io/enforce: restricted`.

    • RBAC uses Roles/ClusterRoles and RoleBindings/ClusterRoleBindings for authorization.
    • Service Accounts provide an identity for processes running inside pods.
    • Secrets are Base64 encoded, not encrypted at rest by default, and store sensitive data.
    • Network Policies control pod communication and require a CNI plugin with a Network Policy Controller.
    • Pod Security Standards (PSS) replaced PSPs and are enforced by the PodSecurity admission controller.
    • The `kubectl auth can-i` command is used to check authorization for specific actions.
    • Setting `automountServiceAccountToken: false` prevents a service account token from being mounted into a pod.
    • A Role defines permissions within a namespace, while a ClusterRole defines cluster-wide permissions.
    What is the primary authorization mechanism in Kubernetes?
    RBAC (Role-Based Access Control).
    tap to reveal
    Which Kubernetes object grants permissions defined in a Role to a user within a specific namespace?
    RoleBinding.
    tap to reveal
    What is the purpose of a Service Account?
    To provide an identity for processes running in a Pod.
    tap to reveal
    How are Kubernetes Secrets stored by default?
    Base64 encoded (not encrypted at rest unless explicitly configured).
    tap to reveal
    Which Kubernetes object controls network traffic between pods?
    NetworkPolicy.
    tap to reveal
    What replaced Pod Security Policies (PSPs) for enforcing pod security standards?
    Pod Security Standards (PSS), enforced by the PodSecurity admission controller.
    tap to reveal
    Which `kubectl` command helps check if a user or service account can perform a specific action?
    `kubectl auth can-i`.
    tap to reveal
    What is the difference between a Role and a ClusterRole?
    A Role defines permissions within a specific namespace, while a ClusterRole defines cluster-wide permissions or for non-namespaced resources.
    tap to reveal