Prometheus Alert Health Checks
Prometheus alert health checks evaluate PromQL expressions from an alert profile during a Krkn run. They let you verify cluster SLOs and automatically identify health regressions caused by a chaos scenario.
Prerequisites
- Prometheus must be reachable from the machine running Krkn.
- On OpenShift, Krkn can discover the Prometheus route and bearer token. On Kubernetes, set
prometheus_urlandprometheus_bearer_token. - An alert profile must contain one or more PromQL expressions.
Configuration
Configure alert health checks under performance_monitoring:
performance_monitoring:
prometheus_url: "http://prometheus.example.com"
prometheus_bearer_token: ""
enable_alerts: True
alert_profile: config/alerts.yaml
run_during: ["pre", "during", "post"]
exit_on_failure: True
only_failures: False
| Field | Description | Default |
|---|---|---|
prometheus_url |
Prometheus API URL. | Auto-detected on OpenShift |
prometheus_bearer_token |
Bearer token used to authenticate with Prometheus. | Auto-detected on OpenShift |
enable_alerts |
Enable alert-profile evaluation. | False |
alert_profile |
Path or URL to the alert profile. | config/alerts.yaml |
run_during |
Phase or phases when expressions are evaluated. | during |
exit_on_failure |
Fail the run when a blocking alert fails. | False |
only_failures |
Include only failed evaluations in telemetry and reports. | False |
Alert Profile
Each entry contains a PromQL expression, a description, and an optional severity:
- expr: sum(rate(apiserver_request_total{code=~"5.."}[5m])) > 0
description: Kubernetes API server returned errors
severity: critical
- expr: increase(etcd_server_leader_changes_seen_total[2m]) > 0
description: etcd leader changes observed
severity: warning
Krkn ships example profiles in the Krkn config directory. Copy and adapt alerts.yaml for the cluster and SLOs being tested.
Evaluation Phases
Set run_during to one phase or a list of phases:
| Value | Behavior |
|---|---|
pre |
Evaluate each expression once before chaos starts. |
during |
Evaluate once after chaos and wait_duration, using a range covering the chaos and wait window. It is not polled on the regular health-check interval. |
post |
Evaluate each expression once after chaos completes. |
["pre", "during", "post"] |
Evaluate at all three phases. |
Use a pre-check to establish a baseline, a post-check to verify recovery, or both:
performance_monitoring:
enable_alerts: True
alert_profile: config/alerts.yaml
run_during: ["pre", "post"]
exit_on_failure: True
Severity and Failure Behavior
Alert severity controls how a failed expression is reported:
infoandwarningevaluations are reported but are non-blocking.errorandcriticalevaluations are blocking whenexit_on_failure: True.- With
only_failures: True, passing evaluations are omitted from telemetry and reports.
The phase and evaluation result are included in alert telemetry. Alert results are also shown in the generated HTML and PDF reports.
Configuration by Runner
The same alert health-check settings are available through Krkn, Krkn-Hub, and Krknctl:
Configure the performance_monitoring section in config.yaml:
performance_monitoring:
prometheus_url: "http://prometheus.example.com"
prometheus_bearer_token: ""
enable_alerts: True
alert_profile: config/alerts.yaml
run_during: ["pre", "during", "post"]
exit_on_failure: True
only_failures: False
Set these environment variables before starting the scenario container:
export PROMETHEUS_URL="http://prometheus.example.com"
export PROMETHEUS_TOKEN=""
export ENABLE_ALERTS=True
export ALERTS_PATH=config/alerts.yaml
export ALERTS_RUN_DURING='[pre, during, post]'
export ALERTS_EXIT_ON_FAILURE=True
export ALERTS_ONLY_FAILURES=False
Pass the global alert options to the scenario command:
krknctl run pod-scenarios \
--prometheus-url http://prometheus.example.com \
--prometheus-token "" \
--enable-alerts True \
--alerts-path config/alerts.yaml \
--alerts-run-during '["pre", "during", "post"]' \
--alerts-exit-on-failure True \
--alerts-only-failures False
The alert profile must be available at the path supplied to the container, or be provided as a URL supported by Krkn.