A Go agent that runs as a pod inside a Kubernetes cluster, receives alerts from Alertmanager, decides using declarative runbooks, restrains itself with guardrails, and reports what it repaired and what it could not.
It works with or without AI. Everything below runs deterministically with no API key; adding one lets the agent interpret runbooks written in prose, explain root causes, and write its own reports.
This is a demo. The timings and limits are tuned to fit inside a presentation, not production.
┌──────────┐ scrape ┌────────────┐ alert ┌──────────────┐
│ mock apps│◄─────────│ Prometheus │─────────►│ Alertmanager │
│ (demo) │ └────────────┘ └──────┬───────┘
└────▲─────┘ │ webhook
│ ▼
│ restart / scale ┌────────────────────────────────────┐
│ (demo namespace only) │ selfheal-agent (Go, in a pod) │
└─────────────────────────┤ │
│ 1. observe the alert │
│ 2. decide runbook (Markdown) │
│ 3. authorize guardrails (Go) │
│ 4. act client-go │
│ 5. VERIFY did it work? │
│ 6. report │
└───────────────┬────────────────────┘
▼
demo console · Telegram · logs
✅ repaired / 🚨 escalated
Step 5 is what makes this an agent rather than a cron job running
kubectl delete. Without verification there is no way to know whether the
action helped, and reporting "repaired" without checking is worse than saying
nothing.
Requirements: docker, minikube, kubectl, helm. (go 1.25+ only for the
tests — the three binaries are compiled inside Docker.)
Step-by-step installation guide: INSTALLATION.md · 🇪🇸 en español: INSTALACION.md
Use recent
minikubeandkubectl. The kube-prometheus-stack chart installs CRDs that older clients do not handle well, and a minikube profile created more than a year ago starts with expired cluster certificates (minikube deleteand recreating it fixes that).
# Windows
.\scripts\setup.ps1# Linux / macOS / Git Bash
./scripts/setup.shThen open http://localhost:8088 — the demo console. Everything is driven from there.
The first run takes 5–10 minutes (it pulls the Prometheus chart and builds three images). The script leaves port-forwards running for the console, Prometheus and Grafana.
Alertmanager runs but is not forwarded. It is essential — it groups the alerts and routes them to the agent's webhook — but its UI only shows anything during the ~90 seconds an alert is firing. To look at it anyway:
kubectl port-forward -n monitoring svc/kube-prom-kube-prometheus-alertmanager 9093:9093Two providers work: Anthropic directly, or OpenRouter, which has free models that support tool calling — so the whole AI half of this demo, chat included, can be shown without a paid account.
.\scripts\setup.ps1 -AnthropicApiKey "sk-ant-..."
.\scripts\setup.ps1 -OpenRouterApiKey "sk-or-v1-..."ANTHROPIC_API_KEY=sk-ant-... ./scripts/setup.sh
OPENROUTER_API_KEY=sk-or-v1-... ./scripts/setup.shThe provider is inferred from the key prefix, so nothing else changes. A key pasted into the console works the same way.
The console shows the same reports, so this is only for demonstrating the real channel. Create a bot with @BotFather, message it, then:
.\scripts\telegram-chatid.ps1 -Token "123456:ABC-DEF..."
.\scripts\setup.ps1 -TelegramToken "123456:ABC-DEF..." -TelegramChatId "-1001234567890"./scripts/telegram-chatid.sh "123456:ABC-DEF..."
TELEGRAM_BOT_TOKEN="123456:ABC-DEF..." TELEGRAM_CHAT_ID="-1001234567890" ./scripts/setup.shhttp://localhost:8088
Left: buttons that break things for real. Right: the agent's reports, rendered as the chat channel they would land in. Each button says up front what should happen, so the demo explains itself.
The header shows the three facts that change how a scenario plays out: which cluster, whether the agent is in dry-run, and whether AI is on.
Every scenario is also available from the terminal — .\scripts\demo.ps1 oom,
./scripts/demo.sh queue, and so on.
| Scenario | What breaks | What the agent does |
|---|---|---|
| Memory leak | The API leaks until the kernel OOMKills it | ✅ Captures logs, deletes the pod, verifies |
| Stuck queue | The transcoder stops draining its backlog | ✅ Restarts the Deployment, verifies |
| Ingest surge | More jobs arrive than the workers can absorb | ✅ Scales to 3 workers, verifies |
| Full disk | The PVC fills up | 🚨 Diagnoses and escalates |
| Corrupt manifests | Malformed HLS in the catalog | 🚨 Escalates — needs AI to read its runbook |
| Origin latency | P95 degrades | 🚨 Escalates — no runbook exists |
The difference is not arbitrary, and it is the single most important idea in the demo: where the broken state lives.
- The OOM marker lives on an
emptyDir. It survives the kernel killing the container — same pod, so the app fails again and a real CrashLoop forms — but not the pod being recreated. Restarting genuinely fixes it. - The disk and manifest markers live on a PVC. They survive everything. Restarting frees nothing and rewrites nothing; it would only hide the problem for a few minutes and burn remediation budget.
The stuck queue is the interesting middle case: nothing is down. The pod stays Ready, probes pass, there are no restarts. Only a business metric moves. If the agent only watched infrastructure health, the cluster would look perfect while the catalog fell behind.
One file carries both halves. The frontmatter is what runs; the prose is what a human reads — and what the AI reads when it is enabled.
---
name: restart-crashlooping-pod
match:
alertname: MockAppCrashLooping
guardrails:
allowedNamespaces: [demo]
maxActionsPerHour: 3
cooldown: 3m
minHealthyReplicas: 1
maxAlertAge: 15m
steps: # optional — see below
- action: capture_logs
continueOnError: true
- action: restart_pod
- action: verify_healthy
timeout: 3m
---
# Pod in CrashLoopBackOff from memory exhaustion
The process accumulates memory until the kernel kills it…If steps is present, the runbook is deterministic and always works.
If steps is absent, the procedure exists only in prose, and interpreting it
requires AI. Without a key the agent does not guess — it escalates saying it has
a written procedure for this case but needs AI to read it. That is the whole
degradation contract, and runbooks/vod-corrupt-manifests.md demonstrates it.
Actions are a closed set (capture_logs, describe_pod, restart_pod,
rollout_restart, scale_deployment, verify_healthy, escalate). Neither a
runbook nor a language model can invent verbs or run arbitrary commands.
The loader rejects dangerous runbooks at startup rather than mid-incident:
- mutating the cluster without a
verify_healthystep, - mutating the cluster without declaring
allowedNamespaces, - no
match(it would apply to every alert), - an unknown action,
- prose-only with no
allowedNamespacesdeclared up front.
Runbooks are editable data; guardrails live in Go. Editing a ConfigMap cannot loosen a safety limit. Where a runbook and the global config disagree, the stricter one wins.
| Guardrail | What it prevents |
|---|---|
DRY_RUN (defaults to true) |
A fresh install touching the cluster before anyone opts in |
ALLOWED_NAMESPACES (global) |
A bad runbook acting outside its blast radius |
allowedNamespaces (per runbook) |
The same, declared per procedure |
maxActionsPerHour |
Repairing a bug in a loop that restarting will never fix |
cooldown |
Stomping on a repair that is still settling |
minHealthyReplicas |
Killing the last healthy replica and turning degradation into an outage |
maxReplicas |
An agent with a scaling verb turning a traffic spike into an invoice |
maxAlertAge |
Acting on a stale alert that no longer describes the present |
MAX_ACTIONS_PER_HOUR_GLOBAL |
An alert storm becoming an action storm |
DISABLED_ACTIONS |
Leaving a verb off during a change freeze |
Three decisions worth pointing at:
The rate limit follows the Deployment, not the pod. Counting per pod would mean each restart produced a new name and the counter never reached its ceiling — exactly the infinite loop the guardrail exists to prevent.
Read-only actions cost nothing and are never skipped. Reading breaks nothing, and denying diagnostics to whoever receives the escalation is the worst thing the agent could do. When a guardrail blocks a repair, the agent still gathers logs and events and ships them with the report.
minHealthyReplicas looks at the state after acting, not before. This is
subtle, and it is where the guardrail breaks if implemented the obvious way: a pod
in CrashLoop is not Ready, so asking "are there at least N healthy replicas right
now?" would block precisely the incident the agent exists to fix. What the rule
must prevent is killing the only pod still answering, not restarting one that
is already down.
With a key set — Anthropic's or OpenRouter's — the agent uses a model for four things:
| Use | What changes | If it fails |
|---|---|---|
| Interpret prose runbooks | Markdown prose becomes executable steps | Escalates; never guesses |
| Root-cause analysis | 40 lines of stack trace become one useful sentence | Report ships without it |
| Write the report | Natural prose instead of the template | Falls back to the template |
| Suggest for unknown alerts | A hint attached to the escalation | Escalates without it |
The security boundary does not move. The model proposes; the guardrails
dispose. Interpreted steps go through runbook.ValidateSteps — the same
function that validates hand-written ones — so the model cannot invent an action,
skip verify_healthy, widen a namespace allowlist, or raise a rate limit. And
underneath all of it the RBAC Role still says what the agent is able to do at
all.
Every AI call is best-effort and time-boxed. A slow or unavailable model can make a report less useful; it cannot change what happened to the cluster.
Which vendor answers is decided in one place. internal/ai splits into a
thin provider interface with two implementations — the Anthropic SDK, and any
OpenAI-compatible endpoint, which is what OpenRouter serves. Everything above
that line is shared: the prompts, the parsing, and above all ValidateSteps.
Swapping vendors cannot open a path around the guardrails, because the
guardrails never see a provider at all.
The one deliberate difference is the effort hint, which the OpenAI-compatible provider drops rather than translates. OpenRouter forwards parameters to hundreds of models and rejects requests the target does not accept, so guessing at a mapping would trade a slightly better answer for an intermittent failure that depends on which model somebody picked.
Paste an Anthropic key into the console and the #incidents panel becomes a
two-way channel. You can ask the agent what is happening, and you can ask it to
do things.
The key is held in the agent's memory only: never written to disk, never stored in a Secret, gone when the pod restarts. The agent validates it against the real API before accepting it, so a typo fails immediately rather than mid-incident.
What makes this worth demonstrating is what happens when you ask for something you should not get:
> restart the etcd pod in kube-system
🚫 I can't. Namespace "kube-system" is not in my allowlist — the only
namespace I'm permitted to act in is "demo". That limit is in my code,
not in a config I can talk myself out of.
> ok, restart mockapp then
✅ Deleted pod mockapp-7d9cb-4jbkd. 1 of 3 hourly remediations used.
A chat request carries no runbook, so it is authorized by
guardrail.EvaluateAction — the same function that authorizes a runbook's steps.
The consequences are concrete and testable:
- The namespace allowlist applies identically.
- A human's restarts spend the same hourly budget as the agent's own. Ask three times and the fourth is refused, even if no alert ever fired.
minHealthyReplicasstill refuses to kill the last serving pod.- In dry-run, the agent refuses everything and says why.
- Read-only tools are never rate-limited, so you can always ask what is going on.
The model reaches the cluster only through the same closed catalog of actions the
runbooks use. There is no shell, no kubectl passthrough, and no way to name a
verb outside that list. When a guardrail refuses, the refusal is returned to the
model as a tool error, which is why it can explain it in plain language instead
of failing opaquely.
This endpoint has no authentication, and neither does the key one. That is fine behind a
127.0.0.1port-forward on a local demo cluster and nowhere else. Putting this agent on a shared cluster would mean authenticating both first.
The guardrails say what the agent is willing to do. The Kubernetes Role says what it is able to do even if the code is wrong or a model proposes something it should not.
It is a Role scoped to the demo namespace, not a ClusterRole. It has
delete on pods and patch on deployments, but never create or delete of
workloads. The images are distroless with no shell, read-only root filesystem,
and run as nonroot: a container holding permissions over the cluster is exactly
where you do not want a shell.
The CrashLoop rule looks like it should be built on the restart counter. The obvious version is wrong, and it fails in the worst possible direction:
increase(kube_pod_container_status_restarts_total[3m]) >= 2 # DON'T
CrashLoopBackOff restarts back off exponentially — 10s, 20s, 40s, 80s, 160s, then a five-minute ceiling. After a handful of crashes Kubernetes only retries the pod every five minutes, so a three-minute window sees zero restarts and the alert resolves itself while the pod is still down and serving nothing.
The longer the outage lasts, the less likely the alert is to notice it. Measured on a pod that had been down for half an hour:
pod: CrashLoopBackOff, 9 restarts, 0/1 Ready
increase(restarts[3m]) 0 ← the rule saw nothing
kube_pod_container_status_waiting_reason 1 ← the pod was still broken
The fix is to describe the state rather than the rate:
max by (namespace, pod, container) (
kube_pod_container_status_waiting_reason{reason="CrashLoopBackOff"}
) > 0
That series is 1 for as long as the container is genuinely stuck, whatever the backoff has stretched to, and clears the moment a healthy pod replaces it.
verify_healthy does not settle for the Deployment reporting Ready:
- Waits for the controller to have observed the change
(
observedGeneration >= generation). - Waits for every replica to be Ready and updated.
- Checks no container is
Waiting(a pod can be Ready between two crashes). - Compares the restart count before and after. If it went up, the repair did not work, and the incident escalates even though the pod looks fine.
Step 1 is easy to leave out and quietly ruins the rest. After a
rollout_restart the Deployment still reports the old ReplicaSet as fully
Ready and updated for a moment, so a verification that starts immediately passes
in milliseconds — having verified the state we were trying to replace. It is the
exact failure mode this whole design exists to prevent, hiding inside the step
meant to prevent it.
# Observation mode: decides and reports, changes nothing
kubectl set env -n monitoring deploy/selfheal-agent DRY_RUN=true
# Back to repairing
kubectl set env -n monitoring deploy/selfheal-agent DRY_RUN=false
# Turn AI on or off without redeploying
kubectl create secret generic selfheal-ai -n monitoring \
--from-literal=api-key="sk-ant-..." --dry-run=client -o yaml | kubectl apply -f -
kubectl rollout restart deploy/selfheal-agent -n monitoring
# Change runbooks without recompiling
kubectl create configmap selfheal-runbooks -n monitoring \
--from-file=runbooks/ --dry-run=client -o yaml | kubectl apply -f -
kubectl rollout restart deploy/selfheal-agent -n monitoring
# What is loaded right now
curl http://localhost:8088/api/statusWhoever watches also needs watching. On /metrics:
| Metric | Answers |
|---|---|
selfheal_incidents_started_total |
How many incidents arrived |
selfheal_incidents_finished_total{outcome} |
How many it repaired vs. escalated |
selfheal_incident_duration_seconds |
How long it takes to close one |
selfheal_actions_total{action} |
Which actions it runs, how often |
selfheal_ai_calls_total{purpose} |
What the AI is being used for (zero when off) |
The SelfhealAgentDown rule alerts when the agent stops responding — and routes
to a separate receiver, since sending the agent an alert about the agent being
down would not help.
cmd/agent/ the agent
cmd/mockapp/ the victim apps (APP_ROLE=api | transcoder)
cmd/console/ the demo console (Go server + embedded page)
internal/
ai/ the optional Claude layer
alertmanager/ the webhook payload
runbook/ loads, validates and matches Markdown runbooks
guardrail/ the limits, in code
k8s/ client-go, narrowed to what the agent needs
engine/ observe → decide → authorize → act → verify → report
notify/ Telegram, console, and the demo-console webhook
runbooks/ the runbooks that ship
deploy/ Kubernetes manifests and Helm values
scripts/ setup, demo, port-forward, teardown
go test ./... # includes a test that validates the shipped runbooksThe tests cover the full flow against a fake cluster (client-go/fake):
verified repair, escalation without touching anything, the prose-runbook
degradation when AI is off, dry-run, a guardrail blocking, and re-send
deduplication. ValidateSteps is tested directly against unsafe plans, because
that is the function standing between a model's output and the cluster.
Worth saying plainly, because these are the questions that will come up:
- Guardrail history is in memory. If the agent restarts, the counters reset — and a CrashLoop of the agent itself would reset its own limits. This belongs in a shared store (Redis, or a ConfigMap).
- A single replica. Two replicas would remediate the same incident twice; this needs leader election.
- No persistent audit trail. Incidents live in logs and metrics. A queryable record of "what did the agent do, and why" is the first thing a platform team asks for.
- Demo timings.
for: 30son the alerts andgroup_wait: 10sin Alertmanager exist so a presentation flows. In production those are far more conservative, precisely so the agent does not remediate on noise. minHealthyReplicasis blunt for Deployment-wide actions. It refuses them outright rather than reading the Deployment's rollout strategy, so arollout_restartthatmaxSurgewould have made perfectly safe is still blocked. Refusing is the right default while the agent cannot see the strategy; reading it is the fix.- An AI-relayed refusal is not the refusal. The guardrail's decision is
enforced in code and cannot be talked around — but the sentence the model
writes about it can be wrong. Pushed on a refusal it once invented a
mechanism, complete with the
maxSurgesemantics it thought applied, and none of it was in the agent. The prompt now forbids explaining internals it cannot see, and the refusals say what the agent does not look at, so there is less to fill in. Neither is a guarantee. The action was still blocked either way, which is the part that matters, but anyone reading an explanation from the chat should treat it as commentary rather than as the rule. - No optional human approval. For anything riskier than restarting a pod, the natural next step is a button in the chat message that authorizes the remediation instead of executing it directly.
- AI cost and latency are unbounded per incident. There is a timeout, but no budget ceiling across incidents, and no caching of interpretations for a runbook whose prose has not changed.