Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

Self-healing monitoring agent

A Go agent that runs as a pod inside a Kubernetes cluster, receives alerts from Alertmanager, decides using declarative runbooks, restrains itself with guardrails, and reports what it repaired and what it could not.

It works with or without AI. Everything below runs deterministically with no API key; adding one lets the agent interpret runbooks written in prose, explain root causes, and write its own reports.

This is a demo. The timings and limits are tuned to fit inside a presentation, not production.


The idea in one picture

 ┌──────────┐  scrape  ┌────────────┐  alert   ┌──────────────┐
 │ mock apps│◄─────────│ Prometheus │─────────►│ Alertmanager │
 │  (demo)  │          └────────────┘          └──────┬───────┘
 └────▲─────┘                                         │ webhook
      │                                               ▼
      │  restart / scale        ┌────────────────────────────────────┐
      │  (demo namespace only)  │   selfheal-agent   (Go, in a pod)  │
      └─────────────────────────┤                                    │
                                │  1. observe    the alert           │
                                │  2. decide     runbook (Markdown)  │
                                │  3. authorize  guardrails (Go)     │
                                │  4. act        client-go           │
                                │  5. VERIFY     did it work?        │
                                │  6. report                         │
                                └───────────────┬────────────────────┘
                                                ▼
                                   demo console · Telegram · logs
                                   ✅ repaired  /  🚨 escalated

Step 5 is what makes this an agent rather than a cron job running kubectl delete. Without verification there is no way to know whether the action helped, and reporting "repaired" without checking is worse than saying nothing.


Quick start

Requirements: docker, minikube, kubectl, helm. (go 1.25+ only for the tests — the three binaries are compiled inside Docker.)

Step-by-step installation guide: INSTALLATION.md · 🇪🇸 en español: INSTALACION.md

Use recent minikube and kubectl. The kube-prometheus-stack chart installs CRDs that older clients do not handle well, and a minikube profile created more than a year ago starts with expired cluster certificates (minikube delete and recreating it fixes that).

# Windows
.\scripts\setup.ps1
# Linux / macOS / Git Bash
./scripts/setup.sh

Then open http://localhost:8088 — the demo console. Everything is driven from there.

The first run takes 5–10 minutes (it pulls the Prometheus chart and builds three images). The script leaves port-forwards running for the console, Prometheus and Grafana.

Alertmanager runs but is not forwarded. It is essential — it groups the alerts and routes them to the agent's webhook — but its UI only shows anything during the ~90 seconds an alert is firing. To look at it anyway:

kubectl port-forward -n monitoring svc/kube-prom-kube-prometheus-alertmanager 9093:9093

With AI enabled

Two providers work: Anthropic directly, or OpenRouter, which has free models that support tool calling — so the whole AI half of this demo, chat included, can be shown without a paid account.

.\scripts\setup.ps1 -AnthropicApiKey "sk-ant-..."
.\scripts\setup.ps1 -OpenRouterApiKey "sk-or-v1-..."
ANTHROPIC_API_KEY=sk-ant-... ./scripts/setup.sh
OPENROUTER_API_KEY=sk-or-v1-... ./scripts/setup.sh

The provider is inferred from the key prefix, so nothing else changes. A key pasted into the console works the same way.

With a real Telegram bot (optional)

The console shows the same reports, so this is only for demonstrating the real channel. Create a bot with @BotFather, message it, then:

.\scripts\telegram-chatid.ps1 -Token "123456:ABC-DEF..."
.\scripts\setup.ps1 -TelegramToken "123456:ABC-DEF..." -TelegramChatId "-1001234567890"
./scripts/telegram-chatid.sh "123456:ABC-DEF..."
TELEGRAM_BOT_TOKEN="123456:ABC-DEF..." TELEGRAM_CHAT_ID="-1001234567890" ./scripts/setup.sh

The demo console

http://localhost:8088

Left: buttons that break things for real. Right: the agent's reports, rendered as the chat channel they would land in. Each button says up front what should happen, so the demo explains itself.

The header shows the three facts that change how a scenario plays out: which cluster, whether the agent is in dry-run, and whether AI is on.

Every scenario is also available from the terminal — .\scripts\demo.ps1 oom, ./scripts/demo.sh queue, and so on.


The six scenarios

Scenario What breaks What the agent does
Memory leak The API leaks until the kernel OOMKills it ✅ Captures logs, deletes the pod, verifies
Stuck queue The transcoder stops draining its backlog ✅ Restarts the Deployment, verifies
Ingest surge More jobs arrive than the workers can absorb ✅ Scales to 3 workers, verifies
Full disk The PVC fills up 🚨 Diagnoses and escalates
Corrupt manifests Malformed HLS in the catalog 🚨 Escalates — needs AI to read its runbook
Origin latency P95 degrades 🚨 Escalates — no runbook exists

Why some repair and others do not

The difference is not arbitrary, and it is the single most important idea in the demo: where the broken state lives.

  • The OOM marker lives on an emptyDir. It survives the kernel killing the container — same pod, so the app fails again and a real CrashLoop forms — but not the pod being recreated. Restarting genuinely fixes it.
  • The disk and manifest markers live on a PVC. They survive everything. Restarting frees nothing and rewrites nothing; it would only hide the problem for a few minutes and burn remediation budget.

The stuck queue is the interesting middle case: nothing is down. The pod stays Ready, probes pass, there are no restarts. Only a business metric moves. If the agent only watched infrastructure health, the cluster would look perfect while the catalog fell behind.


How it is built

Runbooks: Markdown, with the executable part in frontmatter

One file carries both halves. The frontmatter is what runs; the prose is what a human reads — and what the AI reads when it is enabled.

---
name: restart-crashlooping-pod
match:
  alertname: MockAppCrashLooping
guardrails:
  allowedNamespaces: [demo]
  maxActionsPerHour: 3
  cooldown: 3m
  minHealthyReplicas: 1
  maxAlertAge: 15m
steps:                        # optional — see below
  - action: capture_logs
    continueOnError: true
  - action: restart_pod
  - action: verify_healthy
    timeout: 3m
---

# Pod in CrashLoopBackOff from memory exhaustion

The process accumulates memory until the kernel kills it…

If steps is present, the runbook is deterministic and always works.

If steps is absent, the procedure exists only in prose, and interpreting it requires AI. Without a key the agent does not guess — it escalates saying it has a written procedure for this case but needs AI to read it. That is the whole degradation contract, and runbooks/vod-corrupt-manifests.md demonstrates it.

Actions are a closed set (capture_logs, describe_pod, restart_pod, rollout_restart, scale_deployment, verify_healthy, escalate). Neither a runbook nor a language model can invent verbs or run arbitrary commands.

The loader rejects dangerous runbooks at startup rather than mid-incident:

  • mutating the cluster without a verify_healthy step,
  • mutating the cluster without declaring allowedNamespaces,
  • no match (it would apply to every alert),
  • an unknown action,
  • prose-only with no allowedNamespaces declared up front.

Guardrails: code, not data

Runbooks are editable data; guardrails live in Go. Editing a ConfigMap cannot loosen a safety limit. Where a runbook and the global config disagree, the stricter one wins.

Guardrail What it prevents
DRY_RUN (defaults to true) A fresh install touching the cluster before anyone opts in
ALLOWED_NAMESPACES (global) A bad runbook acting outside its blast radius
allowedNamespaces (per runbook) The same, declared per procedure
maxActionsPerHour Repairing a bug in a loop that restarting will never fix
cooldown Stomping on a repair that is still settling
minHealthyReplicas Killing the last healthy replica and turning degradation into an outage
maxReplicas An agent with a scaling verb turning a traffic spike into an invoice
maxAlertAge Acting on a stale alert that no longer describes the present
MAX_ACTIONS_PER_HOUR_GLOBAL An alert storm becoming an action storm
DISABLED_ACTIONS Leaving a verb off during a change freeze

Three decisions worth pointing at:

The rate limit follows the Deployment, not the pod. Counting per pod would mean each restart produced a new name and the counter never reached its ceiling — exactly the infinite loop the guardrail exists to prevent.

Read-only actions cost nothing and are never skipped. Reading breaks nothing, and denying diagnostics to whoever receives the escalation is the worst thing the agent could do. When a guardrail blocks a repair, the agent still gathers logs and events and ships them with the report.

minHealthyReplicas looks at the state after acting, not before. This is subtle, and it is where the guardrail breaks if implemented the obvious way: a pod in CrashLoop is not Ready, so asking "are there at least N healthy replicas right now?" would block precisely the incident the agent exists to fix. What the rule must prevent is killing the only pod still answering, not restarting one that is already down.

The AI layer is optional, and it never holds the keys

With a key set — Anthropic's or OpenRouter's — the agent uses a model for four things:

Use What changes If it fails
Interpret prose runbooks Markdown prose becomes executable steps Escalates; never guesses
Root-cause analysis 40 lines of stack trace become one useful sentence Report ships without it
Write the report Natural prose instead of the template Falls back to the template
Suggest for unknown alerts A hint attached to the escalation Escalates without it

The security boundary does not move. The model proposes; the guardrails dispose. Interpreted steps go through runbook.ValidateSteps — the same function that validates hand-written ones — so the model cannot invent an action, skip verify_healthy, widen a namespace allowlist, or raise a rate limit. And underneath all of it the RBAC Role still says what the agent is able to do at all.

Every AI call is best-effort and time-boxed. A slow or unavailable model can make a report less useful; it cannot change what happened to the cluster.

Which vendor answers is decided in one place. internal/ai splits into a thin provider interface with two implementations — the Anthropic SDK, and any OpenAI-compatible endpoint, which is what OpenRouter serves. Everything above that line is shared: the prompts, the parsing, and above all ValidateSteps. Swapping vendors cannot open a path around the guardrails, because the guardrails never see a provider at all.

The one deliberate difference is the effort hint, which the OpenAI-compatible provider drops rather than translates. OpenRouter forwards parameters to hundreds of models and rejects requests the target does not accept, so guessing at a mapping would trade a slightly better answer for an intermittent failure that depends on which model somebody picked.

Talking to the agent — and being told no

Paste an Anthropic key into the console and the #incidents panel becomes a two-way channel. You can ask the agent what is happening, and you can ask it to do things.

The key is held in the agent's memory only: never written to disk, never stored in a Secret, gone when the pod restarts. The agent validates it against the real API before accepting it, so a typo fails immediately rather than mid-incident.

What makes this worth demonstrating is what happens when you ask for something you should not get:

> restart the etcd pod in kube-system

🚫 I can't. Namespace "kube-system" is not in my allowlist — the only
   namespace I'm permitted to act in is "demo". That limit is in my code,
   not in a config I can talk myself out of.

> ok, restart mockapp then

✅ Deleted pod mockapp-7d9cb-4jbkd. 1 of 3 hourly remediations used.

A chat request carries no runbook, so it is authorized by guardrail.EvaluateAction — the same function that authorizes a runbook's steps. The consequences are concrete and testable:

  • The namespace allowlist applies identically.
  • A human's restarts spend the same hourly budget as the agent's own. Ask three times and the fourth is refused, even if no alert ever fired.
  • minHealthyReplicas still refuses to kill the last serving pod.
  • In dry-run, the agent refuses everything and says why.
  • Read-only tools are never rate-limited, so you can always ask what is going on.

The model reaches the cluster only through the same closed catalog of actions the runbooks use. There is no shell, no kubectl passthrough, and no way to name a verb outside that list. When a guardrail refuses, the refusal is returned to the model as a tool error, which is why it can explain it in plain language instead of failing opaquely.

This endpoint has no authentication, and neither does the key one. That is fine behind a 127.0.0.1 port-forward on a local demo cluster and nowhere else. Putting this agent on a shared cluster would mean authenticating both first.

RBAC: defence in depth

The guardrails say what the agent is willing to do. The Kubernetes Role says what it is able to do even if the code is wrong or a model proposes something it should not.

It is a Role scoped to the demo namespace, not a ClusterRole. It has delete on pods and patch on deployments, but never create or delete of workloads. The images are distroless with no shell, read-only root filesystem, and run as nonroot: a container holding permissions over the cluster is exactly where you do not want a shell.

Alert on state, not on rate

The CrashLoop rule looks like it should be built on the restart counter. The obvious version is wrong, and it fails in the worst possible direction:

increase(kube_pod_container_status_restarts_total[3m]) >= 2   # DON'T

CrashLoopBackOff restarts back off exponentially — 10s, 20s, 40s, 80s, 160s, then a five-minute ceiling. After a handful of crashes Kubernetes only retries the pod every five minutes, so a three-minute window sees zero restarts and the alert resolves itself while the pod is still down and serving nothing.

The longer the outage lasts, the less likely the alert is to notice it. Measured on a pod that had been down for half an hour:

pod:                                     CrashLoopBackOff, 9 restarts, 0/1 Ready
increase(restarts[3m])                   0        ← the rule saw nothing
kube_pod_container_status_waiting_reason 1        ← the pod was still broken

The fix is to describe the state rather than the rate:

max by (namespace, pod, container) (
  kube_pod_container_status_waiting_reason{reason="CrashLoopBackOff"}
) > 0

That series is 1 for as long as the container is genuinely stuck, whatever the backoff has stretched to, and clears the moment a healthy pod replaces it.

Verification, not optimism

verify_healthy does not settle for the Deployment reporting Ready:

  1. Waits for the controller to have observed the change (observedGeneration >= generation).
  2. Waits for every replica to be Ready and updated.
  3. Checks no container is Waiting (a pod can be Ready between two crashes).
  4. Compares the restart count before and after. If it went up, the repair did not work, and the incident escalates even though the pod looks fine.

Step 1 is easy to leave out and quietly ruins the rest. After a rollout_restart the Deployment still reports the old ReplicaSet as fully Ready and updated for a moment, so a verification that starts immediately passes in milliseconds — having verified the state we were trying to replace. It is the exact failure mode this whole design exists to prevent, hiding inside the step meant to prevent it.


Operating it

# Observation mode: decides and reports, changes nothing
kubectl set env -n monitoring deploy/selfheal-agent DRY_RUN=true

# Back to repairing
kubectl set env -n monitoring deploy/selfheal-agent DRY_RUN=false

# Turn AI on or off without redeploying
kubectl create secret generic selfheal-ai -n monitoring \
    --from-literal=api-key="sk-ant-..." --dry-run=client -o yaml | kubectl apply -f -
kubectl rollout restart deploy/selfheal-agent -n monitoring

# Change runbooks without recompiling
kubectl create configmap selfheal-runbooks -n monitoring \
    --from-file=runbooks/ --dry-run=client -o yaml | kubectl apply -f -
kubectl rollout restart deploy/selfheal-agent -n monitoring

# What is loaded right now
curl http://localhost:8088/api/status

The agent's own metrics

Whoever watches also needs watching. On /metrics:

Metric Answers
selfheal_incidents_started_total How many incidents arrived
selfheal_incidents_finished_total{outcome} How many it repaired vs. escalated
selfheal_incident_duration_seconds How long it takes to close one
selfheal_actions_total{action} Which actions it runs, how often
selfheal_ai_calls_total{purpose} What the AI is being used for (zero when off)

The SelfhealAgentDown rule alerts when the agent stops responding — and routes to a separate receiver, since sending the agent an alert about the agent being down would not help.


Layout

cmd/agent/          the agent
cmd/mockapp/        the victim apps (APP_ROLE=api | transcoder)
cmd/console/        the demo console (Go server + embedded page)
internal/
  ai/               the optional Claude layer
  alertmanager/     the webhook payload
  runbook/          loads, validates and matches Markdown runbooks
  guardrail/        the limits, in code
  k8s/              client-go, narrowed to what the agent needs
  engine/           observe → decide → authorize → act → verify → report
  notify/           Telegram, console, and the demo-console webhook
runbooks/           the runbooks that ship
deploy/             Kubernetes manifests and Helm values
scripts/            setup, demo, port-forward, teardown
go test ./...   # includes a test that validates the shipped runbooks

The tests cover the full flow against a fake cluster (client-go/fake): verified repair, escalation without touching anything, the prose-runbook degradation when AI is off, dry-run, a guardrail blocking, and re-send deduplication. ValidateSteps is tested directly against unsafe plans, because that is the function standing between a model's output and the cluster.


What it would need for production

Worth saying plainly, because these are the questions that will come up:

  • Guardrail history is in memory. If the agent restarts, the counters reset — and a CrashLoop of the agent itself would reset its own limits. This belongs in a shared store (Redis, or a ConfigMap).
  • A single replica. Two replicas would remediate the same incident twice; this needs leader election.
  • No persistent audit trail. Incidents live in logs and metrics. A queryable record of "what did the agent do, and why" is the first thing a platform team asks for.
  • Demo timings. for: 30s on the alerts and group_wait: 10s in Alertmanager exist so a presentation flows. In production those are far more conservative, precisely so the agent does not remediate on noise.
  • minHealthyReplicas is blunt for Deployment-wide actions. It refuses them outright rather than reading the Deployment's rollout strategy, so a rollout_restart that maxSurge would have made perfectly safe is still blocked. Refusing is the right default while the agent cannot see the strategy; reading it is the fix.
  • An AI-relayed refusal is not the refusal. The guardrail's decision is enforced in code and cannot be talked around — but the sentence the model writes about it can be wrong. Pushed on a refusal it once invented a mechanism, complete with the maxSurge semantics it thought applied, and none of it was in the agent. The prompt now forbids explaining internals it cannot see, and the refusals say what the agent does not look at, so there is less to fill in. Neither is a guarantee. The action was still blocked either way, which is the part that matters, but anyone reading an explanation from the chat should treat it as commentary rather than as the rule.
  • No optional human approval. For anything riskier than restarting a pod, the natural next step is a button in the chat message that authorizes the remediation instead of executing it directly.
  • AI cost and latency are unbounded per incident. There is a timeout, but no budget ceiling across incidents, and no caching of interpretations for a runbook whose prose has not changed.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages