AI Kubernetes Troubleshooting with k8sgpt: What It Catches and Misses (2026)

Khimananda Oli 12 min read Database, Virtualization
AI Kubernetes Troubleshooting with k8sgpt: What It Catches and Misses (2026)

By Khimananda Oli | Last reviewed: October 2026

AI Kubernetes troubleshooting with k8sgpt is two separate things: a set of rule-based analyzers that read object state and need no AI at all, and an optional --explain step that sends their findings to an LLM. The analyzers reliably catch broken state — bad images, crash loops, unschedulable pods, empty Services — and are blind to anything that leaves state healthy, like an app returning HTTP 500 behind green probes.

What does k8sgpt actually do in AI Kubernetes troubleshooting?

The name suggests a model reading your cluster. That is not what happens, and the difference decides how you should use it. When you run k8sgpt analyze, a fixed set of analyzers — ordinary Go code, one per resource kind — list objects through the Kubernetes API and apply rules: is this container waiting with an error reason, does this Service have endpoints, is this PersistentVolumeClaim stuck in Pending. Each rule that fires produces a short text finding. No model is involved.

Only when you add --explain do those findings go to an AI backend, wrapped in a prompt that asks for an explanation and a fix. So k8sgpt is a deterministic detector with an optional AI summariser bolted on, not an AI that investigates. That has three practical consequences this guide is built around: its coverage is exactly the analyzer rules, nothing more; without --explain, nothing leaves your machine; and the AI half only ever sees what the rules extracted.

Installing the in-cluster operator is covered in best AI tools for DevOps engineers in 2026, and the manual techniques this tool sits on top of are in the Kubernetes troubleshooting field guide. Neither is repeated here. This piece is about judging the tool itself.

Where the AI actually is in k8sgptDetection is rules. The model only rewrites what the rules found.Kubernetes APIpods, events,services, PVCs…AnalyzersGo rules per kind14 default · 17 optionalno AI involvedFindingsshort text linesprinted locally--anonymizemasks resourcenames onlyLLM backendexplain in≤ 280 charsonly with --explaink8sgpt analyzeRules run against the API and print findings.Same input gives the same output, every time.Nothing is sent anywhere.Most of the value lives here.k8sgpt analyze --explainSame findings, then one prompt per resultsent to your configured backend.Output varies by model and run.Useful wording, no new detection.
In AI Kubernetes troubleshooting with k8sgpt, detection is done by rule-based analyzers; the LLM only sees the findings, and only when you pass --explain.

What does k8sgpt catch, and what does it miss?

Because detection is code, coverage can be read rather than guessed. The table below comes from the analyzer source at the v0.4.39 release tag (September 2026), not from the README — which matters, because the README's descriptions had already drifted from the code on at least one point covered in the next section. Fourteen analyzers run by default; seventeen more run only when you add them with k8sgpt filters add.

FailureAnalyzerVerdictWhat you actually get
Bad image tag or registryPod · defaultCaughtThe ErrImagePull / ImagePullBackOff message
CrashLoopBackOffPod · defaultCaught, shallowlyThe last termination reason and exit code — not the application's own error, which is in the logs
OOMKilledPod · defaultCaughtReported as a crash loop whose last termination reason is OOMKilled
Unschedulable (CPU, memory, taints)Pod · defaultCaughtThe scheduler's "0/N nodes are available" message
Missing ConfigMap key or SecretPod · defaultCaughtThe CreateContainerConfigError message naming what is missing
Readiness probe failingPod · defaultUsuallyOnly when the pod's latest event is Unhealthy; a newer, unrelated event hides it
Service with no endpointsService · defaultCaught"Service has no endpoints, expected label app=…" — selector typos stand out immediately
PVC stuck PendingPVC · defaultUsuallyOnly when an event says ProvisioningFailed; a claim waiting silently is not reported
Ingress to a missing Service, class or TLS secretIngress · defaultCaughtOne finding per broken reference
Node NotReady or under pressureNode · defaultCaughtThe condition type, reason and message
App errors in logs, pod healthyLog · optionalOnly if enabledLines matching error|exception|fail in the last 100 log lines
NetworkPolicy matching no podsNetworkPolicy · optionalOnly if enabledFlags the policy; says nothing about whether traffic is actually blocked
HPA targeting a missing workloadHPA · optionalOnly if enabledThe missing scaleTargetRef
HTTP 500s behind green probes—InvisibleNothing: every object is in a healthy state
Latency or saturation without restarts—InvisibleNothing: analyzers read state, not metrics
Valid but wrong config (wrong DB host)—InvisibleNothing, unless the app logs errors and the Log analyzer is on

The detection core is small enough to quote. This is the list of container waiting reasons the Pod analyzer treats as errors:

// pkg/analyzer/pod.go, k8sgpt v0.4.39
failureReasons := []string{
    "CrashLoopBackOff", "ImagePullBackOff", "CreateContainerConfigError",
    "PreCreateHookError", "CreateContainerError", "PreStartHookError",
    "RunContainerError", "ImageInspectError", "ErrImagePull",
    "ErrImageNeverPull", "InvalidImageName",
}

If a pod is not in one of those states, not Pending, not Evicted, not terminated with a non-zero exit, and not Running-but-unready, the Pod analyzer has nothing to say about it.

The line k8sgpt cannot crossIt sees broken objects. It cannot see a healthy object doing the wrong thing.Default analyzersruns with no setup· bad image / pull errors· CrashLoopBackOff, OOMKilled· unschedulable pods· missing ConfigMap / Secret· Service without endpoints· PVC provisioning failed· broken Ingress references· node conditionsState is visibly wrongOptional analyzersk8sgpt filters add …· Log: error lines in last 100· NetworkPolicy selectors· HPA scale targets· PodDisruptionBudgets· Gateway API routes· Security, StorageOff by default for a reason:Log sends raw log linesConfiguration looks wrongInvisibleno analyzer can see it· HTTP 500 behind green probes· latency and saturation· valid but wrong config· slow memory leak, no restart yet· DNS or CNI packet loss· a bad deploy that still servesNeeds metrics, traces and ahuman — or a wider AI workflowState is healthy, behaviour isn'tMost real incidents start in the right-hand column. Use k8sgpt to rule out the left one fast.
What k8sgpt catches by default, what needs optional analyzers, and the incidents that stay invisible to AI Kubernetes troubleshooting built on object state.

Why the misses are structural, not a model problem

Swapping in a better LLM changes none of the right-hand column. A pod that is Running and Ready while returning HTTP 500s has no faulty field anywhere in the API, so no analyzer fires, so there is nothing to explain. The same holds for a deploy that is slow but up, a connection pool that is exhausted, or a configuration value that is valid YAML pointing at the wrong database. Those are found with metrics, traces and the techniques in the guides for debugging a CrashLoopBackOff, OOMKilled debugging and ephemeral containers on distroless pods. k8sgpt's honest job is to clear the left-hand column in seconds so you can spend your time on the right.

Is --anonymize enough to send findings to a cloud LLM?

This is where reading the code paid off. The README has long said that anonymization does not apply to Pod, Events, Log, ReplicaSet and PersistentVolumeClaim results. That was true until recently, and it was a real leak: pod and event failures copy messages straight from Kubernetes, and those messages contain the very names --anonymize is meant to hide (tracked as issue #560).

Release v0.4.39 fixed the general case. For every result, k8sgpt now derives masking pairs from the result's own identity — the namespace, pod and container segments of its path — and replaces them in the text before it is sent, then restores them in the answer. So on a current release, the names of the broken resource itself are masked for every analyzer. That is a meaningful improvement, and if you are on an older release you should upgrade before relying on the flag at all.

What it still does not mask is everything else in the text. The masking is identity-based, not content-based:

  • Other objects named in a message. A CreateContainerConfigError reads couldn't find key UPSTREAM_URL in ConfigMap payments/gateway-config. The namespace is masked because it is part of the pod's identity; gateway-config and the key name are not.
  • Log content. The Log analyzer masks the pod name and sends the matching lines verbatim — internal hostnames, ports, database users, and anything a developer ever printed in an error path.
  • Hostnames, image references and URLs that appear in scheduler, kubelet or Ingress messages.
What --anonymize actually removes (v0.4.39)Finding as k8sgpt builds itcouldn't find key UPSTREAM_URL in ConfigMap payments/gateway-configERROR dial tcp db-primary.corp.lan:5432: connection refused user=settle_svcresult: payments/gateway-7d9f/gw · payments/settlement-5c2a/appWhat is sent with --anonymizecouldn't find key UPSTREAM_URL in ConfigMap ▇▇▇▇/gateway-configERROR dial tcp db-primary.corp.lan:5432: connection refused user=settle_svcmasked: namespace, pod and container segments of each result's own nameMasked· the broken resource's namespace· its pod and container names· names an analyzer declares itselfStill sent in clear· other objects named in messages· log lines: hosts, ports, users· hostnames, images, URLs, keys
k8sgpt --anonymize in v0.4.39 masks each result's own identity; other identifiers inside messages and log lines still reach the AI backend unmasked.

The practical rule follows directly. If you only run default analyzers on a current release, --anonymize with a cloud backend is a defensible choice for many teams. The moment you enable the Log analyzer, treat everything as unmasked and point --explain at a local model instead. The broader reasoning on what may leave your network is in protecting PII and secrets in LLM apps; and remember that log lines are attacker-influenced input, which matters if a model's answer is ever acted on automatically — see prompt injection attacks and defenses.

What does a good AI Kubernetes troubleshooting workflow with k8sgpt look like?

1. Run it without --explain first

Plain analysis is deterministic, fast, and sends nothing anywhere — the safest first step in any AI Kubernetes troubleshooting session. On an incident it is the quickest way to answer "is anything in this namespace visibly broken?":

# Rule-based findings only — no AI backend, no egress
k8sgpt analyze --namespace payments

# Narrow to the analyzers you care about
k8sgpt analyze -n payments --filter Pod,Service,Ingress

# One resource, with the analyzer timings
k8sgpt analyze -n payments --resource Deployment/gateway --with-stat

An empty result is information too: it moves the incident firmly into the "state is healthy, behaviour isn't" column, and you stop looking at pods and start looking at metrics.

2. Explain with a local model

The explanation step is a small task — the default prompt asks for an error summary and a fix "in no more than 280 characters" — so a local 7–8B model is enough, and nothing leaves the machine. Configure an Ollama backend once:

# one-time: register a local backend (Ollama on its default port)
k8sgpt auth add --backend ollama --model qwen2.5:7b-instruct \
  --baseurl http://localhost:11434

k8sgpt analyze -n payments --explain --backend ollama --anonymize

Keep --anonymize on even for a local model; it costs nothing and protects you the day someone switches the default backend. Setup for the model side is in running local LLMs with Ollama for DevOps workflows.

3. Enable optional analyzers deliberately

k8sgpt filters list                 # what is active vs available
k8sgpt filters add NetworkPolicy,HorizontalPodAutoscaler
k8sgpt filters add Log              # only with a local backend
k8sgpt filters remove Log

NetworkPolicy and HPA findings are cheap and catch configuration typos that cause hours of confusion. The Log analyzer is the one that changes your data exposure, so turn it on per investigation rather than leaving it enabled.

4. Put the output where people already look

# JSON is stable enough to script against
k8sgpt analyze -n payments -o json \
  | jq -r '.results[]? | "\(.kind) \(.name): \(.error[0].Text)"'

Piping that into the incident channel as the first message of a page — "here is what is visibly broken" — saves the first five minutes of every Kubernetes incident. It fits naturally into the triage flow in AI incident response: triaging alerts and runbooks with LLMs.

When should AI Kubernetes troubleshooting go beyond k8sgpt?

Two signals mean you have reached the edge of the tool. The first is an empty or irrelevant result during a real incident: the failure lives in behaviour, not state, and you need metrics, traces and someone who knows the service. The second is a finding whose explanation is too shallow to act on — usually a crash loop, where k8sgpt can tell you the container exited with code 1 but not why. There the next step is the application's own logs and a deliberate investigation, either manually or with a bounded evidence bundle handed to a model as in AI Linux troubleshooting with an LLM.

Recent releases also expose k8sgpt as an MCP server (k8sgpt serve --mcp), so an AI agent can call its analyzers as tools. That is a genuinely good fit — the analyzers are read-only and deterministic, exactly what you want an agent to lean on — provided the service account behind it has read-only RBAC and the agent is not also handed write access to the cluster. If your NetworkPolicy findings start turning into real incidents, Kubernetes network policies explained covers what the analyzer cannot: whether traffic is actually being blocked.

Used for what it is — a fast, deterministic sweep for broken state with optional plain-language summaries — k8sgpt is one of the few AI Kubernetes troubleshooting tools that earns a permanent place in an on-call toolkit. Used as an AI that "finds the problem", it will quietly miss the incidents that cost the most. If you want AI-assisted Kubernetes operations set up properly, with local models, read-only RBAC and the alerting wired in, my DevOps and cloud consulting services cover exactly that.

Frequently Asked Questions

An open-source CLI and operator that scans a Kubernetes cluster with rule-based analyzers, one per resource kind, and reports what is visibly broken. It can optionally send those findings to an LLM backend with --explain for a plain-language explanation and suggested fix.

No. Detection is done by deterministic Go analyzers that read object state through the Kubernetes API. The AI is only used afterwards, when you pass --explain, to rewrite the findings the analyzers already produced.

Only when you run analyze with --explain, and only to whichever backend you configured. Without --explain nothing leaves your machine. With a local backend such as Ollama, nothing leaves your network either.

Image pull errors, CrashLoopBackOff and OOMKilled containers, unschedulable pods, missing ConfigMap keys and Secrets, failing readiness probes, Services without endpoints, PVCs that failed provisioning, broken Ingress references, unavailable Deployment replicas and unhealthy node conditions.

Anything that leaves object state healthy: HTTP 500s behind green health probes, latency and saturation, slow memory leaks before a restart, valid configuration pointing at the wrong dependency, and most DNS or CNI packet loss. Those need metrics, traces and human investigation.

Partly. Since v0.4.39 it masks each broken resource's own namespace, pod and container names for every analyzer. It does not mask other identifiers inside messages, such as referenced ConfigMap names, hostnames or image URLs, and it sends Log analyzer lines verbatim apart from the pod name.

On v0.4.39 and later, yes for the failing resource's own identity. Older releases left Pod and event-derived findings unmasked, which the project tracked as issue 560. The README's older wording on this is out of date, so upgrade before relying on the flag.

Yes. Register a backend with k8sgpt auth add --backend ollama, a model name and the Ollama base URL, then run analyze --explain --backend ollama. Because the default prompt asks for a short explanation, a 7–8B local model is usually sufficient.

Run k8sgpt filters add Log. It reads the last 100 lines of each container and reports lines matching error, exception or fail. Those lines are sent to the backend nearly verbatim, so only enable it with a local model, and remove it after the investigation.

The default prompt template asks the model to explain the error and give a step-by-step solution in no more than 280 characters. That keeps answers scannable but means it rarely goes deeper than restating the finding with an obvious next step.

The CLI is the right start: it uses your kubeconfig, runs on demand and needs no installation in the cluster. The operator suits continuous scanning that feeds dashboards or alerts. Either way, give it a read-only service account.

Yes. k8sgpt serve --mcp exposes its analyzers over the Model Context Protocol so an agent can call them as tools. It is a good fit because the analyzers are read-only and deterministic, provided the agent's credentials are also read-only.

Run analyze with -o json and extract kind, name and the first error text with jq, then post that as the first message of an incident. It answers "what is visibly broken" in seconds and lets responders move on to metrics if the list is empty.

No. Detection is entirely the analyzer rules, so a larger or newer model only changes how findings are worded. Coverage grows only by enabling optional analyzers or writing custom ones.

Yes, as a fast deterministic sweep that rules out visibly broken state at the start of an incident, with optional local-model summaries. It is not a replacement for observability or investigation, and treating it as an AI that finds the root cause will cause it to miss the costliest incidents.