## Kubernetes Cluster Architecture
A Kubernetes cluster consists of a set of control plane components and worker nodes. The control plane manages the cluster, while worker nodes run the actual applications (pods).
## Control Plane Components
These components make global decisions about the cluster (e.g., scheduling) and detect and respond to cluster events.
## Worker Node Components
These components run on each worker node and are responsible for maintaining running pods and providing the Kubernetes runtime environment.
## Cluster Installation & Configuration with kubeadm
kubeadm is a tool designed to easily bootstrap a minimum viable Kubernetes cluster. It handles the setup of control plane components and allows worker nodes to join.
## Workloads & Scheduling
Kubernetes Pods are the smallest deployable units, encapsulating one or more containers, storage resources, a unique network IP, and options that govern how the containers run. Multi-container Pods often use patterns like sidecar containers for logging or monitoring, and init containers for setup tasks that must complete before the main application containers start.
Deployments are the primary way to manage stateless applications. They define the desired state for a set of Pods and their underlying ReplicaSets. Deployments handle rolling updates and rollbacks, allowing zero-downtime updates and easy reversion to previous versions. ReplicaSets ensure a specified number of identical Pods are running at all times. DaemonSets ensure that all (or some) nodes run a copy of a specific Pod, useful for cluster-level services like logging agents or monitoring daemons.
For tasks that run to completion, Kubernetes provides Jobs. A Job creates one or more Pods and ensures that a specified number of them successfully terminate. CronJobs automate Jobs on a recurring schedule, similar to cron in Linux, enabling scheduled tasks like backups or report generation.
Kubernetes' scheduler places Pods on nodes based on various constraints. Resource Requests define the minimum CPU and memory a Pod needs, ensuring it gets scheduled on a node with sufficient capacity. Resource Limits define the maximum CPU and memory a Pod can consume, preventing resource exhaustion on a node.
Advanced scheduling includes Node Selectors for simple node assignment based on labels. Node Affinity provides more flexible rules (e.g., "prefer this node," "must run on this node"). Taints on nodes repel Pods unless those Pods have a matching Toleration. This is crucial for dedicating nodes or handling special hardware. Pod Disruption Budgets (PDBs) ensure that a minimum number of Pods for a given application remain running during voluntary disruptions (e.g., node maintenance), enhancing application availability. Liveness Probes detect if a container is running correctly and restart it if it fails. Readiness Probes determine if a container is ready to serve traffic, removing it from service endpoints if not. Startup Probes delay liveness and readiness checks until the application has started, preventing premature restarts for slow-starting apps.
## Services
A Service is a Kubernetes object that defines a logical set of Pods and a policy by which to access them. Services provide a stable network endpoint (IP address and DNS name) for a dynamic set of Pods, which are often ephemeral. Services use label selectors to identify the Pods they target.
There are several Service Types:
kube-proxy is a network proxy that runs on each node in the cluster, implementing the Service concept. It maintains network rules (e.g., iptables or IPVS) to forward traffic to the correct backend Pods. EndpointSlices are a more scalable and efficient way to track network endpoints for Services, especially in large clusters, replacing the older Endpoints resource.
## Ingress
Ingress is an API object that manages external access to services in a cluster, typically HTTP/HTTPS. It provides routing rules for traffic based on hostnames and paths to backend Services. An Ingress Controller is essential for Ingress resources to function; it's a specialized pod that watches the Ingress API and configures a load balancer or proxy accordingly (e.g., Nginx, HAProxy). Ingress can also be configured for TLS termination using Kubernetes Secrets.
## Network Policies
NetworkPolicy is an API object that allows you to specify how groups of Pods are allowed to communicate with each other and other network endpoints. By default, Pods are non-isolated, meaning they can communicate freely. Network Policies provide pod-level isolation. They use podSelectors to select the target Pods and define rules for Ingress (incoming traffic) and Egress (outgoing traffic). Rules can specify communication based on Pods, Namespaces, or IP blocks (`ipBlock`). A Container Network Interface (CNI) plugin that supports NetworkPolicy (e.g., Calico) is required for them to be enforced.
## DNS
CoreDNS is the default DNS server in Kubernetes, responsible for Service Discovery. Pods can discover Services using DNS names like `my-service.my-namespace.svc.cluster.local` (fully qualified), `my-service.my-namespace` (within the same cluster), or `my-service` (within the same namespace). Each Pod gets its own DNS configuration, inheriting from the cluster's CoreDNS.
## Kubernetes Storage Overview
Kubernetes abstracts underlying storage with PersistentVolumes (PVs), PersistentVolumeClaims (PVCs), and StorageClasses, crucial for stateful applications.
## PersistentVolumes (PVs)
A PersistentVolume (PV) is a cluster-wide storage resource, provisioned by an administrator or dynamically. It represents actual storage (e.g., NFS, cloud disks) and is independent of Pods. PVs define access modes and a reclaim policy.
## PersistentVolumeClaims (PVCs)
A PersistentVolumeClaim (PVC) is a user's request for storage. Kubernetes binds a PVC to an available PV that matches its requirements (size, access mode). If no suitable PV exists, the PVC remains pending or triggers dynamic provisioning.
## StorageClasses & Dynamic Provisioning
A StorageClass defines a "class" of storage, specifying a provisioner (e.g., `kubernetes.io/aws-ebs`) and parameters. It enables dynamic provisioning: when a PVC requests a StorageClass, the provisioner automatically creates a new PV on the backend storage to fulfill the claim, automating storage setup.
## Access Modes & Reclaim Policy
Access Modes dictate how storage can be mounted:
The Reclaim Policy defines PV behavior after its PVC is deleted:
## Using Storage in Pods
Pods use persistent storage by referencing a PVC in their `volumes` section and then mounting that volume into containers via `volumeMounts`. The Pod interacts with the PVC, which is bound to a PV.
```yaml
apiVersion: v1
kind: Pod
metadata:
name: my-app
spec:
containers:
image: nginx
volumeMounts:
mountPath: /data
volumes:
persistentVolumeClaim:
claimName: my-pvc
```
## Application Troubleshooting in Kubernetes
Troubleshooting applications in Kubernetes is a critical skill for a CKA. It involves systematically diagnosing issues from the pod level up to service exposure.
## Common Pod Troubleshooting Scenarios
## Essential Troubleshooting Commands
## Service and Network Troubleshooting
If pods are healthy but inaccessible, investigate Services. Use `kubectl describe service <service-name>` to check its selector and ensure it matches the labels of your target pods. Verify that the Service's targetPort matches the container port. Check `kubectl get endpoints <service-name>` to confirm that healthy pods are correctly registered as endpoints. Finally, review any NetworkPolicies that might be inadvertently blocking traffic.
## Cluster Component Troubleshooting
Troubleshooting Kubernetes cluster components requires a systematic approach, often starting with identifying the affected component and then examining its logs, status, and configuration. The primary components include the Control Plane (kube-apiserver, etcd, kube-scheduler, kube-controller-manager) and Worker Node components (kubelet, kube-proxy, container runtime).
1. Check Logs: Always start with `journalctl -u <component>` or `kubectl logs -n kube-system <pod-name>`.
2. Check Events: `kubectl get events -A` can provide high-level insights into what's failing.
3. Resource Usage: Ensure components have sufficient CPU/memory.
4. Networking: Verify connectivity between components (e.g., API server to etcd, kubelet to API server). Use `netstat`, `ss`, `ping`, `curl`.
5. Certificates: Ensure all components have valid and correctly configured TLS certificates.
## Cluster Maintenance & Upgrades
Maintaining a healthy Kubernetes cluster involves regular upgrades and robust backup/restore strategies. The CKA exam focuses on manual upgrades using `kubeadm` and `etcd` snapshot management.
## Upgrading Kubernetes Clusters
Manual upgrades typically involve updating the `kubeadm`, `kubelet`, and `kubectl` binaries. It's crucial to upgrade these components to the same minor version across the cluster to ensure compatibility.
Upgrade Process for Control Plane Nodes:
1. Backup etcd: This is the most critical step before any upgrade. See the backup section below.
2. Drain the node: Prepare the control plane node for maintenance by evicting all pods: `kubectl drain <node-name> --ignore-daemonsets`.
3. Upgrade kubeadm: Update the `kubeadm` package to the target version (e.g., `apt-get install -y kubeadm=1.28.x-00` or `yum install -y kubeadm-1.28.x`).
4. Plan the upgrade: Check for potential issues: `kubeadm upgrade plan`.
5. Apply the upgrade: Execute the upgrade on the control plane: `kubeadm upgrade apply v1.28.x`.
6. Upgrade kubelet and kubectl: Update the `kubelet` and `kubectl` packages to the same target version (e.g., `apt-get install -y kubelet=1.28.x-00 kubectl=1.28.x-00`).
7. Restart kubelet: `systemctl restart kubelet`.
8. Uncordon the node: Allow pods to be scheduled again: `kubectl uncordon <node-name>`.
Upgrade Process for Worker Nodes:
1. Drain the node: `kubectl drain <node-name> --ignore-daemonsets`.
2. Upgrade kubeadm: Update the `kubeadm` package.
3. Run kubeadm upgrade node: This command updates the worker node's configuration: `kubeadm upgrade node`.
4. Upgrade kubelet and kubectl: Update the `kubelet` and `kubectl` packages.
5. Restart kubelet: `systemctl restart kubelet`.
6. Uncordon the node: `kubectl uncordon <node-name>`.
Repeat for all worker nodes.
## Backup and Restore Operations
The etcd data store holds the entire state of your Kubernetes cluster. Backing it up is paramount for disaster recovery.
Backup etcd:
Use the `etcdctl snapshot save` command. You'll need to specify the API version, etcd endpoints, and TLS certificates if etcd is secured (which it typically is in a production cluster).
```bash
# Ensure ETCDCTL_API is set to 3
export ETCDCTL_API=3
# Example command (adjust paths for your setup)
etcdctl snapshot save /tmp/etcd-snapshot.db \
--endpoints=https://127.0.0.1:2379 \
--cacert=/etc/kubernetes/pki/etcd/ca.crt \
--cert=/etc/kubernetes/pki/etcd/peer.crt \
--key=/etc/kubernetes/pki/etcd/peer.key
```
Restore etcd:
Restoring etcd involves stopping the Kubernetes control plane components, restoring the snapshot, and then restarting the components.
1. Stop Kubernetes control plane components: On the control plane node, stop the `kube-apiserver`, `kube-scheduler`, and `kube-controller-manager` static pods. This often involves moving their manifest files out of `/etc/kubernetes/manifests/` or stopping the `kubelet` service.
2. Restore the snapshot: Create a *new* data directory for etcd and restore the snapshot into it.
```bash
export ETCDCTL_API=3
etcdctl snapshot restore /tmp/etcd-snapshot.db \
--data-dir=/var/lib/etcd-restored-data
```
3. Update etcd static pod manifest: Modify the `etcd.yaml` manifest file (usually `/etc/kubernetes/manifests/etcd.yaml`) to point to the new `--data-dir` (`/var/lib/etcd-restored-data` in the example).
4. Restart Kubernetes control plane components: Move the manifest files back or restart the `kubelet` service. Verify cluster health.
Always ensure you have backups of your Kubernetes configuration files (`/etc/kubernetes/`) and certificates (`/etc/kubernetes/pki/`) as well.
## Security & RBAC in Kubernetes
Security in Kubernetes is multifaceted, covering authentication, authorization, network controls, and workload isolation. Role-Based Access Control (RBAC) is the primary authorization mechanism, defining what authenticated users and processes can do.
## Authentication and Authorization
Authentication verifies the identity of a user or process (e.g., client certificates, service account tokens, OIDC). Once authenticated, Authorization determines if the identity is permitted to perform a requested action on a resource. The Kubernetes API server handles both.
## Role-Based Access Control (RBAC)
RBAC uses four key objects:
Use `kubectl auth can-i <verb> <resource> --namespace <namespace>` to check permissions.
## Service Accounts
Service Accounts provide an identity for processes running in pods. Each namespace has a default service account. Pods are automatically assigned the `default` service account of their namespace unless specified otherwise. A token for the service account is mounted into the pod at `/var/run/secrets/kubernetes.io/serviceaccount/token`. You can disable this with `automountServiceAccountToken: false` in the pod or service account specification.
## Secrets
Secrets are used to store sensitive information like passwords, OAuth tokens, and SSH keys. They are Base64 encoded, not encrypted by default at rest (unless configured with an encryption provider). Secrets can be mounted as data volumes or exposed as environment variables within pods. Common types include `Opaque`, `kubernetes.io/dockerconfigjson`, and `kubernetes.io/tls`.
## Network Policies
Network Policies specify how groups of pods are allowed to communicate with each other and with external network endpoints. They define ingress (incoming) and egress (outgoing) rules based on pod selectors and namespace selectors. Network Policies are enforced by a Network Policy Controller, which is typically part of the Container Network Interface (CNI) plugin (e.g., Calico, Cilium).
## Pod Security Standards (PSS)
Pod Security Standards (PSS) replaced Pod Security Policies (PSPs) to enforce security best practices for pods. PSS defines three levels:
PSS are enforced by the PodSecurity admission controller, which can be configured to `enforce`, `audit`, or `warn` based on security levels for specific namespaces using labels like `pod-security.kubernetes.io/enforce: restricted`.