
Table of Contents
By Khimananda Oli | Last reviewed: September 2026
Most Kubernetes teams assume their clusters will self-heal during failures, but assumptions cause outages. Chaos Engineering with LitmusChaos: Running Your First GameDay transforms those assumptions into verified resilience data through controlled fault injection. Instead of waiting for a 3 AM incident to expose weak points, you systematically break components in a safe environment while measuring recovery against defined SLOs. This guide walks you through installing LitmusChaos, designing safe experiments, and executing a GameDay that actually improves production reliability.
What Is Chaos Engineering with LitmusChaos and Why Run a GameDay?
Chaos Engineering with LitmusChaos is the practice of injecting managed faults into Kubernetes workloads using a CNCF-graduated framework designed specifically for cloud-native environments. Unlike generic stress testing tools, LitmusChaos integrates directly with the Kubernetes API server and uses Custom Resource Definitions (CRDs) to declare experiments as code. A GameDay is not just running random tests; it is a scheduled, collaborative event where engineering, operations, and sometimes product teams gather to execute predefined scenarios, observe system behavior in real-time, and document gaps between expected and actual resilience.
In my experience helping teams across Nepal and globally prepare for SOC 2 audits, I have found that documented GameDays serve as powerful evidence of operational maturity. Auditors look for proactive risk management, and a well-executed chaos engineering test plan demonstrates that you verify recovery procedures rather than just hoping they work. Before you inject any faults, however, you must understand the architecture of the tool itself to avoid creating accidental production incidents.
The diagram above illustrates how these components interact during a GameDay. The Control Plane orchestrates workflows, pulling experiment definitions from ChaosHub and applying them to the target cluster via the Chaos Operator. Crucially, the observability stack sits alongside this flow, capturing metrics that determine whether the experiment passed or failed based on your defined SLIs and SLOs. Without this feedback loop, you are just breaking things without learning.
How Do You Install and Configure LitmusChaos for Safe Experimentation?
Before running Chaos Engineering with LitmusChaos: Running Your First GameDay, you need a properly configured environment. Never start experimenting directly in production. Set up a staging cluster that mirrors production topology, including resource limits, network policies, and storage classes. If you are evaluating local options, check our comparison of Minikube vs Kind for local Kubernetes to choose the right sandbox.
Install the Litmus Operator via Helm
Helm is the recommended installation method for 2026 releases because it handles CRD lifecycle management and dependency resolution automatically. Add the official chart repository and install the control plane along with the chaos operator in a dedicated namespace:
helm repo add litmuschaos https://litmuschaos.github.io/litmus-helm/
helm repo update
kubectl create namespace litmus
helm install chaos litmuschaos/litmus \
--namespace litmus \
--set portal.server.endpoint=http://localhost:9091 \
--wait Verify that all pods reach the Ready state. The chaos-operator manages the lifecycle of experiments, while the chaos-exporter exposes Prometheus metrics. If pods remain pending, check resource limits and requests to ensure your cluster has sufficient capacity for both the chaos infrastructure and the overhead of fault injection.
Configure RBAC and Namespace Isolation
A common mistake is granting the chaos service account cluster-admin privileges. This violates least-privilege principles and creates audit findings. Instead, scope permissions to specific namespaces where experiments will run:
kubectl apply -f - <<EOF
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
name: chaos-role
namespace: app-staging
rules:
- apiGroups: [""]
resources: ["pods", "services", "configmaps", "secrets"]
verbs: ["get", "list", "watch", "delete", "patch"]
- apiGroups: ["apps"]
resources: ["deployments", "replicasets", "statefulsets"]
verbs: ["get", "list", "watch", "patch"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: chaos-rolebinding
namespace: app-staging
subjects:
- kind: ServiceAccount
name: chaos-service-account
namespace: litmus
roleRef:
kind: Role
name: chaos-role
apiGroup: rbac.authorization.k8s.io
EOF This configuration allows LitmusChaos to inject faults only within app-staging. For multi-tenant environments or compliance-scoped clusters, replicate this pattern per namespace. Document these bindings as part of your security review evidence.
Which Chaos Experiments Should You Run During Your First GameDay?
Your first GameDay should focus on foundational resilience patterns, not exotic failure modes. Select experiments that validate core assumptions about your application's ability to handle common Kubernetes disruptions. ChaosHub provides curated, versioned experiment specifications that reduce configuration errors.
- Pod Delete: Validates that your application recovers when individual replicas terminate unexpectedly. This tests liveness probes, readiness gates, and graceful shutdown handlers.
- Container Kill: Simulates container runtime crashes without node-level impact. Useful for verifying restart policies and state persistence in databases like PostgreSQL or MongoDB.
- Network Latency: Injects configurable delay between services. Essential for microservices architectures to validate timeout configurations and circuit breaker thresholds.
- DNS Chaos: Disrupts DNS resolution for targeted domains. Tests fallback mechanisms and caching behavior in service discovery layers.
- Node Drain: Evicts all pods from a specific node. Validates PodDisruptionBudgets, anti-affinity rules, and autoscaler responsiveness.
Start with Pod Delete and Container Kill. These have predictable blast radii and clear success criteria. Avoid Node-level experiments until you have validated application-layer resilience and confirmed that your monitoring stack captures disruption events accurately. Refer to the four golden signals of monitoring to ensure your dashboards reflect user-facing health during chaos.
This sequential workflow prevents the most common GameDay failure: executing experiments without clear success criteria or abort conditions. Each stage gates the next. If your dry run reveals missing metrics or broken probes, fix them before scheduling the live event. Skipping validation turns a learning exercise into an unplanned outage.
How Do You Execute and Monitor a LitmusChaos GameDay Safely?
Execution requires discipline. Create a ChaosEngine manifest that explicitly defines the target application, experiment parameters, and probe configurations. Below is a minimal Pod Delete experiment targeting a deployment named payment-api in the staging namespace:
apiVersion: litmuschaos.io/v1alpha1
kind: ChaosEngine
metadata:
name: payment-pod-delete
namespace: app-staging
spec:
appinfo:
appns: app-staging
applabel: "app=payment-api"
appkind: deployment
engineState: active
chaosServiceAccount: chaos-service-account
experiments:
- name: pod-delete
spec:
components:
env:
- name: TOTAL_CHAOS_DURATION
value: "60"
- name: CHAOS_INTERVAL
value: "10"
- name: FORCE
value: "false" Note that FORCE is set to false. This ensures pods receive SIGTERM and can execute graceful shutdown hooks. Setting this to true bypasses termination grace periods and may corrupt in-flight transactions. Always validate graceful shutdown behavior first; forceful kills are advanced tests for later GameDays.
Integrate Resilience Probes
Experiments without probes are just vandalism. Add HTTP or CMD probes to the ChaosEngine to automatically validate application health during and after fault injection:
probes:
- name: check-payment-health
type: http
httpProbe/inputs:
url: "http://payment-api.app-staging.svc.cluster.local/healthz"
responseTimeout: 5s
method:
get:
criteria: "=="
responseCode: "200"
mode: Continuous
runProperties:
probeTimeout: 10
retry: 3
interval: 5 The Continuous mode runs this probe throughout the chaos duration. If the health endpoint returns non-200 responses beyond the retry threshold, the experiment fails immediately. This automated verdict removes subjective interpretation and ties results directly to your SLO definitions.
Monitor with Purpose
During the GameDay, keep three dashboards open: application error rate, latency percentiles, and chaos experiment status. Use Prometheus and Grafana to correlate chaos events with metric anomalies. The Litmus Chaos Exporter exposes litmuschaos_experiment_verdict and litmuschaos_probe_success_percentage metrics that should be overlaid on application dashboards. If you cannot see the exact moment chaos started and ended on your graphs, your observability instrumentation is insufficient for reliable GameDays.
| Criterion | First GameDay (Safe) | Mature GameDay (Advanced) |
|---|---|---|
| Environment | Staging / Non-prod mirror | Production (with guardrails) |
| Blast Radius | Single replica, single namespace | Multi-node, cross-service cascades |
| Experiment Type | Pod Delete, Container Kill | Network Partition, Node Taint, DNS |
| Abort Mechanism | Manual kill switch + auto-probe | Automated SLO-based rollback |
| Team Presence | On-call engineer + observer | Cross-functional war room |
| Success Metric | Recovery within RTO | No SLO budget burn during chaos |
This progression table prevents teams from attempting advanced scenarios prematurely. Mature GameDays earn their complexity through repeated success at foundational levels. Jumping to network partitions before validating pod recovery creates noise, not signal.
The contrast is stark. Teams on the left treat chaos as a destructive event; teams on the right treat it as a validation mechanism. The difference is not tooling—it is process discipline. Every element on the right side maps directly to steps in the workflow diagram above. Skip none.
Start Your Chaos Engineering Journey with Intention
Chaos Engineering with LitmusChaos: Running Your First GameDay succeeds when you prioritize learning over spectacle. Install the operator with scoped RBAC, select foundational experiments, define measurable success criteria via probes, and monitor with purpose. Document every GameDay as both operational improvement and compliance evidence. When you are ready to expand beyond pod-level faults, revisit your SLOs and ensure your observability stack can capture the increased complexity. If your team needs guidance designing a GameDay program that aligns with audit requirements or production safety standards, reach out to discuss your resilience testing strategy.