
Table of Contents
By Khimananda Oli | Last reviewed: October 2026
AI Kubernetes troubleshooting with k8sgpt is two separate things: a set of rule-based analyzers that read object state and need no AI at all, and an optional --explain step that sends their findings to an LLM. The analyzers reliably catch broken state — bad images, crash loops, unschedulable pods, empty Services — and are blind to anything that leaves state healthy, like an app returning HTTP 500 behind green probes.
What does k8sgpt actually do in AI Kubernetes troubleshooting?
The name suggests a model reading your cluster. That is not what happens, and the difference decides how you should use it. When you run k8sgpt analyze, a fixed set of analyzers — ordinary Go code, one per resource kind — list objects through the Kubernetes API and apply rules: is this container waiting with an error reason, does this Service have endpoints, is this PersistentVolumeClaim stuck in Pending. Each rule that fires produces a short text finding. No model is involved.
Only when you add --explain do those findings go to an AI backend, wrapped in a prompt that asks for an explanation and a fix. So k8sgpt is a deterministic detector with an optional AI summariser bolted on, not an AI that investigates. That has three practical consequences this guide is built around: its coverage is exactly the analyzer rules, nothing more; without --explain, nothing leaves your machine; and the AI half only ever sees what the rules extracted.
Installing the in-cluster operator is covered in best AI tools for DevOps engineers in 2026, and the manual techniques this tool sits on top of are in the Kubernetes troubleshooting field guide. Neither is repeated here. This piece is about judging the tool itself.
What does k8sgpt catch, and what does it miss?
Because detection is code, coverage can be read rather than guessed. The table below comes from the analyzer source at the v0.4.39 release tag (September 2026), not from the README — which matters, because the README's descriptions had already drifted from the code on at least one point covered in the next section. Fourteen analyzers run by default; seventeen more run only when you add them with k8sgpt filters add.
| Failure | Analyzer | Verdict | What you actually get |
|---|---|---|---|
| Bad image tag or registry | Pod · default | Caught | The ErrImagePull / ImagePullBackOff message |
| CrashLoopBackOff | Pod · default | Caught, shallowly | The last termination reason and exit code — not the application's own error, which is in the logs |
| OOMKilled | Pod · default | Caught | Reported as a crash loop whose last termination reason is OOMKilled |
| Unschedulable (CPU, memory, taints) | Pod · default | Caught | The scheduler's "0/N nodes are available" message |
| Missing ConfigMap key or Secret | Pod · default | Caught | The CreateContainerConfigError message naming what is missing |
| Readiness probe failing | Pod · default | Usually | Only when the pod's latest event is Unhealthy; a newer, unrelated event hides it |
| Service with no endpoints | Service · default | Caught | "Service has no endpoints, expected label app=…" — selector typos stand out immediately |
| PVC stuck Pending | PVC · default | Usually | Only when an event says ProvisioningFailed; a claim waiting silently is not reported |
| Ingress to a missing Service, class or TLS secret | Ingress · default | Caught | One finding per broken reference |
| Node NotReady or under pressure | Node · default | Caught | The condition type, reason and message |
| App errors in logs, pod healthy | Log · optional | Only if enabled | Lines matching error|exception|fail in the last 100 log lines |
| NetworkPolicy matching no pods | NetworkPolicy · optional | Only if enabled | Flags the policy; says nothing about whether traffic is actually blocked |
| HPA targeting a missing workload | HPA · optional | Only if enabled | The missing scaleTargetRef |
| HTTP 500s behind green probes | — | Invisible | Nothing: every object is in a healthy state |
| Latency or saturation without restarts | — | Invisible | Nothing: analyzers read state, not metrics |
| Valid but wrong config (wrong DB host) | — | Invisible | Nothing, unless the app logs errors and the Log analyzer is on |
The detection core is small enough to quote. This is the list of container waiting reasons the Pod analyzer treats as errors:
// pkg/analyzer/pod.go, k8sgpt v0.4.39
failureReasons := []string{
"CrashLoopBackOff", "ImagePullBackOff", "CreateContainerConfigError",
"PreCreateHookError", "CreateContainerError", "PreStartHookError",
"RunContainerError", "ImageInspectError", "ErrImagePull",
"ErrImageNeverPull", "InvalidImageName",
} If a pod is not in one of those states, not Pending, not Evicted, not terminated with a non-zero exit, and not Running-but-unready, the Pod analyzer has nothing to say about it.
Why the misses are structural, not a model problem
Swapping in a better LLM changes none of the right-hand column. A pod that is Running and Ready while returning HTTP 500s has no faulty field anywhere in the API, so no analyzer fires, so there is nothing to explain. The same holds for a deploy that is slow but up, a connection pool that is exhausted, or a configuration value that is valid YAML pointing at the wrong database. Those are found with metrics, traces and the techniques in the guides for debugging a CrashLoopBackOff, OOMKilled debugging and ephemeral containers on distroless pods. k8sgpt's honest job is to clear the left-hand column in seconds so you can spend your time on the right.
Is --anonymize enough to send findings to a cloud LLM?
This is where reading the code paid off. The README has long said that anonymization does not apply to Pod, Events, Log, ReplicaSet and PersistentVolumeClaim results. That was true until recently, and it was a real leak: pod and event failures copy messages straight from Kubernetes, and those messages contain the very names --anonymize is meant to hide (tracked as issue #560).
Release v0.4.39 fixed the general case. For every result, k8sgpt now derives masking pairs from the result's own identity — the namespace, pod and container segments of its path — and replaces them in the text before it is sent, then restores them in the answer. So on a current release, the names of the broken resource itself are masked for every analyzer. That is a meaningful improvement, and if you are on an older release you should upgrade before relying on the flag at all.
What it still does not mask is everything else in the text. The masking is identity-based, not content-based:
- Other objects named in a message. A
CreateContainerConfigErrorreads couldn't find key UPSTREAM_URL in ConfigMap payments/gateway-config. The namespace is masked because it is part of the pod's identity;gateway-configand the key name are not. - Log content. The Log analyzer masks the pod name and sends the matching lines verbatim — internal hostnames, ports, database users, and anything a developer ever printed in an error path.
- Hostnames, image references and URLs that appear in scheduler, kubelet or Ingress messages.
The practical rule follows directly. If you only run default analyzers on a current release, --anonymize with a cloud backend is a defensible choice for many teams. The moment you enable the Log analyzer, treat everything as unmasked and point --explain at a local model instead. The broader reasoning on what may leave your network is in protecting PII and secrets in LLM apps; and remember that log lines are attacker-influenced input, which matters if a model's answer is ever acted on automatically — see prompt injection attacks and defenses.
What does a good AI Kubernetes troubleshooting workflow with k8sgpt look like?
1. Run it without --explain first
Plain analysis is deterministic, fast, and sends nothing anywhere — the safest first step in any AI Kubernetes troubleshooting session. On an incident it is the quickest way to answer "is anything in this namespace visibly broken?":
# Rule-based findings only — no AI backend, no egress
k8sgpt analyze --namespace payments
# Narrow to the analyzers you care about
k8sgpt analyze -n payments --filter Pod,Service,Ingress
# One resource, with the analyzer timings
k8sgpt analyze -n payments --resource Deployment/gateway --with-stat An empty result is information too: it moves the incident firmly into the "state is healthy, behaviour isn't" column, and you stop looking at pods and start looking at metrics.
2. Explain with a local model
The explanation step is a small task — the default prompt asks for an error summary and a fix "in no more than 280 characters" — so a local 7–8B model is enough, and nothing leaves the machine. Configure an Ollama backend once:
# one-time: register a local backend (Ollama on its default port)
k8sgpt auth add --backend ollama --model qwen2.5:7b-instruct \
--baseurl http://localhost:11434
k8sgpt analyze -n payments --explain --backend ollama --anonymize Keep --anonymize on even for a local model; it costs nothing and protects you the day someone switches the default backend. Setup for the model side is in running local LLMs with Ollama for DevOps workflows.
3. Enable optional analyzers deliberately
k8sgpt filters list # what is active vs available
k8sgpt filters add NetworkPolicy,HorizontalPodAutoscaler
k8sgpt filters add Log # only with a local backend
k8sgpt filters remove Log NetworkPolicy and HPA findings are cheap and catch configuration typos that cause hours of confusion. The Log analyzer is the one that changes your data exposure, so turn it on per investigation rather than leaving it enabled.
4. Put the output where people already look
# JSON is stable enough to script against
k8sgpt analyze -n payments -o json \
| jq -r '.results[]? | "\(.kind) \(.name): \(.error[0].Text)"' Piping that into the incident channel as the first message of a page — "here is what is visibly broken" — saves the first five minutes of every Kubernetes incident. It fits naturally into the triage flow in AI incident response: triaging alerts and runbooks with LLMs.
When should AI Kubernetes troubleshooting go beyond k8sgpt?
Two signals mean you have reached the edge of the tool. The first is an empty or irrelevant result during a real incident: the failure lives in behaviour, not state, and you need metrics, traces and someone who knows the service. The second is a finding whose explanation is too shallow to act on — usually a crash loop, where k8sgpt can tell you the container exited with code 1 but not why. There the next step is the application's own logs and a deliberate investigation, either manually or with a bounded evidence bundle handed to a model as in AI Linux troubleshooting with an LLM.
Recent releases also expose k8sgpt as an MCP server (k8sgpt serve --mcp), so an AI agent can call its analyzers as tools. That is a genuinely good fit — the analyzers are read-only and deterministic, exactly what you want an agent to lean on — provided the service account behind it has read-only RBAC and the agent is not also handed write access to the cluster. If your NetworkPolicy findings start turning into real incidents, Kubernetes network policies explained covers what the analyzer cannot: whether traffic is actually being blocked.
Used for what it is — a fast, deterministic sweep for broken state with optional plain-language summaries — k8sgpt is one of the few AI Kubernetes troubleshooting tools that earns a permanent place in an on-call toolkit. Used as an AI that "finds the problem", it will quietly miss the incidents that cost the most. If you want AI-assisted Kubernetes operations set up properly, with local models, read-only RBAC and the alerting wired in, my DevOps and cloud consulting services cover exactly that.