Krkn-Hub All Scenarios Variables
These variables are to be used for the top level configuration template that are shared by all the scenarios in Krkn-hub.
Each section below corresponds to a section in the Krkn config reference. Set variables on the host running the container:
export <parameter_name>=<value>
Kraken
Signal and status publishing settings. See Kraken config for full details.
| Parameter | Description | Default |
|---|---|---|
| KRKN_KUBE_CONFIG | Path to the kubeconfig file for cluster access | /home/krkn/.kube/config |
| DISABLE_IMAGE_SIGNATURE | Disables workload image signature verification before deployment | False |
| PUBLISH_KRAKEN_STATUS | Publish kraken status to the signal address | False |
| SIGNAL_ADDRESS | Address to publish kraken status to | 0.0.0.0 |
| PORT | Port to publish kraken status to | 8081 |
| SIGNAL_STATE | Waits for the RUN signal when set to PAUSE before running the scenarios, refer docs for more details | RUN |
| REPORT_FORMATS | List of report formats to generate after each scenario (e.g., pdf, html) |
[pdf,html] |
| KRKN_DEBUG | Enables debug mode for Krkn | False |
Cerberus
Cluster health monitoring integration. See Cerberus config for full details.
| Parameter | Description | Default |
|---|---|---|
| CERBERUS_ENABLED | Set this to true if cerberus is running and monitoring the cluster | False |
| CERBERUS_URL | URL to poll for the go/no-go signal | http://0.0.0.0:8080 |
Performance Monitoring
Prometheus metrics collection and alert evaluation. See Performance Monitoring config for full details.
| Parameter | Description | Default |
|---|---|---|
| UUID | UUID for the run; auto-generated if not set | "" |
| PROMETHEUS_URL | URL to Prometheus instance; auto-detected on OpenShift, required for Kubernetes | "" |
| PROMETHEUS_TOKEN | Bearer token for Prometheus authentication; auto-detected on OpenShift, required for Kubernetes | "" |
| CAPTURE_METRICS | Captures metrics as specified in the profile from in-cluster prometheus. Default metrics captures are listed here | False |
| ENABLE_ALERTS | Evaluates expressions from in-cluster prometheus and exits 0 or 1 based on the severity set. Default profile. | False |
| ALERTS_PATH | Path to the alerts file to use when ENABLE_ALERTS is set | config/alerts.yaml |
| ALERTS_RUN_DURING | Prometheus alert health-check phases: pre, during, post, or a list of phases | during |
| ALERTS_EXIT_ON_FAILURE | Fail the run for critical or error alert evaluations | False |
| ALERTS_ONLY_FAILURES | Include only failed alert evaluations in telemetry and reports | False |
| METRICS_PATH | Path to the metrics profile to use when CAPTURE_METRICS is set | config/metrics-aggregated.yaml |
| CHECK_CRITICAL_ALERTS | When enabled will check prometheus for critical alerts firing post chaos | False |
Resiliency Score
Resiliency scoring configuration. See Resiliency Score config for full details.
| Parameter | Description | Default |
|---|---|---|
| RESILIENCY_RUN_MODE | Resiliency scoring mode: standalone embeds score in telemetry, detailed prints JSON report to stdout, disabled turns off scoring |
standalone |
| RESILIENCY_FILE | Path to a YAML file containing SLO definitions; defaults to the alerts profile or config/alerts.yaml |
config/alerts.yaml |
Elastic
Elasticsearch storage for telemetry and metrics. See Elastic config for full details.
| Parameter | Description | Default |
|---|---|---|
| ENABLE_ES | Enable Elasticsearch integration | False |
| ES_SERVER | URL of the Elasticsearch instance | http://0.0.0.0 |
| ES_PORT | Port of the Elasticsearch instance | 443 |
| ES_USERNAME | Username for Elasticsearch authentication | elastic |
| ES_PASSWORD | Password for Elasticsearch authentication | - |
| ES_VERIFY_CERTS | Verify SSL certificates when connecting to Elasticsearch | False |
| ES_RUN_TAG | Tag to identify the run in Elasticsearch | "" |
| ES_METRICS_INDEX | Elasticsearch index for metrics data | krkn-metrics |
| ES_ALERTS_INDEX | Elasticsearch index for alerts data | krkn-alerts |
| ES_TELEMETRY_INDEX | Elasticsearch index for telemetry data | krkn-telemetry |
Tunings
Execution timing and iteration controls. See Tunings config for full details.
| Parameter | Description | Default |
|---|---|---|
| WAIT_DURATION | Duration in seconds to wait between each chaos scenario | 60 |
| ITERATIONS | Number of times to execute the scenarios | 1 |
| DAEMON_MODE | Iterations are set to infinity which means that the kraken will cause chaos forever | False |
Telemetry
Run data collection and upload settings. See Telemetry config for full details.
| Parameter | Description | Default |
|---|---|---|
| TELEMETRY_ENABLED | Enable/disables the telemetry collection feature | False |
| TELEMETRY_API_URL | Telemetry service endpoint | https://ulnmf9xv7j.execute-api.us-west-2.amazonaws.com/produ... |
| TELEMETRY_USERNAME | Telemetry service username | redhat-chaos |
| TELEMETRY_PASSWORD | Telemetry service password | - |
| TELEMETRY_PROMETHEUS_BACKUP | Enables/disables prometheus data collection | True |
| TELEMETRY_FULL_PROMETHEUS_BACKUP | If set to False only the /prometheus/wal folder will be downloaded | False |
| TELEMETRY_BACKUP_THREADS | Number of telemetry download/upload threads | 5 |
| TELEMETRY_ARCHIVE_PATH | Local path where the archive files will be temporarily stored | /tmp |
| TELEMETRY_MAX_RETRIES | Maximum number of upload retries (if 0 will retry forever) | 0 |
| TELEMETRY_RUN_TAG | If set, this will be appended to the run folder in the bucket (useful to group the runs) | chaos |
| TELEMETRY_GROUP | If set will archive the telemetry in the S3 bucket on a folder named after the value | default |
| TELEMETRY_ARCHIVE_SIZE | The size of the prometheus data archive in KB | 1000 |
| TELEMETRY_LOGS_BACKUP | Logs backup to S3 | False |
| TELEMETRY_FILTER_PATTERN | Filter logs based on certain timestamp patterns | ["(\\w{3}\\s\\d{1,2}\\s\\d{2}:\\d{2}:\\d{2}\\.\\d+).+","kini... |
| TELEMETRY_CLI_PATH | OC CLI path, if not specified will be searched in $PATH | "" |
| TELEMETRY_EVENTS_BACKUP | Enables/disables events backup to S3 | True |
Note
For setting theTELEMETRY_ARCHIVE_SIZE, the lower the value the higher the number of archive files produced and uploaded (processed by TELEMETRY_BACKUP_THREADS simultaneously). For unstable or slow connections, keep this value low and increase TELEMETRY_BACKUP_THREADS so that on upload failure only the failed chunk is retried.
Health Checks
Application endpoint monitoring during chaos. See Health Checks config for full details.
| Parameter | Description | Default |
|---|---|---|
| HEALTH_CHECK_INTERVAL | Interval in seconds at which to run health checks | 2 |
| HEALTH_CHECK_URL | URL to continually check and detect downtimes | "" |
| HEALTH_CHECK_BEARER_TOKEN | Bearer token used for authenticating into health check URL | "" |
| HEALTH_CHECK_AUTH | Tuple of (username, password) used for authenticating into health check URL | "" |
| HEALTH_CHECK_EXIT_ON_FAILURE | If True, exits when health check fails for application | "" |
| HEALTH_CHECK_VERIFY | Health check URL SSL validation | False |
| HEALTH_CHECK_RUN_DURING | When to run health checks: pre (before chaos), during (continuous), post (after chaos), or combination like [pre,post] | during |
| HEALTH_CHECK_ONLY_FAILURES | Only create telemetry records for failed health checks | False |
Virt Checks
KubeVirt VMI SSH connection monitoring during chaos. See Virt Checks config for full details.
| Parameter | Description | Default |
|---|---|---|
| KUBE_VIRT_CHECK_INTERVAL | Interval in seconds at which to test kubevirt connections | 2 |
| KUBE_VIRT_NAMESPACE | Namespace to find VMIs in and watch | "" |
| KUBE_VIRT_NAME | Regex style name to match VMIs to watch | "" |
| KUBE_VIRT_LABEL_SELECTOR | Label selector to filter VMs in KubeVirt | "" |
| KUBE_VIRT_FAILURES | If True, will only report when ssh connections fail to VMI | False |
| KUBE_VIRT_DISCONNECTED | Use disconnected check by passing cluster API | False |
| KUBE_VIRT_SSH_NODE | If set, will be a backup way to SSH to a node. Should be a node not targeted in chaos | "" |
| KUBE_VIRT_NODE_NAME | Filter only VMI’s running a specific node name | "" |
| KUBE_VIRT_EXIT_ON_FAIL | Fails run if VMs still have false status at end of run | False |
Other
Parameters found in the source that no section above covers.
| Parameter | Description | Default |
|---|---|---|
| RETRY_WAIT | 120 |
Triggers
Parameters found in the source that no section above covers.
| Parameter | Description | Default |
|---|---|---|
| TRIGGER_COMMAND | Shell command to evaluate before chaos starts. Chaos begins only after this command exits with the expected code. Leave empty to disable triggers. | "" |
| TRIGGER_EXPECTED_RC | Expected shell command exit code for the trigger to be considered satisfied | 0 |
| TRIGGER_HTTP_URL | URL to poll for the HTTP trigger | "" |
| TRIGGER_HTTP_METHOD | HTTP method to use (GET, POST, etc.) | GET |
| TRIGGER_HTTP_EXPECTED_STATUS | Expected HTTP status code | 200 |
| TRIGGER_HTTP_BEARER_TOKEN | Bearer token for HTTP authentication | "" |
| TRIGGER_HTTP_BODY_CONTAINS | Substring expected in the HTTP response body | "" |
| TRIGGER_K8S_API_VERSION | Kubernetes API version for the resource to watch (e.g. apps/v1, kubevirt.io/v1). Required when using a k8s trigger condition. | "" |
| TRIGGER_K8S_KIND | Kubernetes resource kind to watch (e.g. Deployment, Pod, VirtualMachineInstanceMigration) | "" |
| TRIGGER_K8S_NAME | Name of the specific Kubernetes resource to watch | "" |
| TRIGGER_K8S_NAMESPACE | Namespace of the resource to watch. Leave empty for cluster-scoped resources like Nodes. | "" |
| TRIGGER_K8S_CONDITION | Condition expression to evaluate against the resource (e.g. status.phase == Running, status.readyReplicas >= 1) | "" |
| TRIGGER_PROM_QUERY | PromQL expression for the prometheus trigger. Chaos begins only after the query returns a non-empty result. Leave empty to disable. | "" |
| TRIGGER_PROM_URL | Prometheus API URL used by the prometheus trigger condition. Required when trigger-prom-query is set. | "" |
| TRIGGER_PROM_TOKEN | Optional bearer token for authenticating the prometheus trigger against prometheus-url | "" |
| TRIGGERS_MODE | How multiple trigger conditions are combined | all_of |
| TRIGGERS_TIMEOUT | Maximum time in seconds to wait for trigger conditions before applying on-timeout behavior | 300 |
| TRIGGERS_INTERVAL | Polling interval in seconds while waiting for trigger conditions | 5 |
| TRIGGERS_ON_TIMEOUT | Behavior when trigger conditions are not met before timeout | skip |
Object State Checks
Parameters found in the source that no section above covers.
| Parameter | Description | Default |
|---|---|---|
| OBJECT_STATE_CHECK_INTERVAL | How often to check object states (seconds) | 2 |
| OBJECT_STATE_CHECK_RUN_DURING | When to run checks: pre (before chaos), during (continuous), post (after chaos), or combination like [pre,during] | during |
| OBJECT_STATE_CHECK_EXIT_ON_FAILURE | Exit on failure when object state check fails | False |
| OBJECT_STATE_CHECK_ONLY_FAILURES | Only create telemetry records for failed object state checks | False |
| OBJECT_STATE_CHECK_NAME | Name of the object state check (e.g., etcd-ready-pods) | "" |
| OBJECT_STATE_CHECK_KIND | Kubernetes resource kind to check (Pod, Deployment, StatefulSet, DaemonSet, ReplicaSet, Service, Job, etc.) | Pod |
| OBJECT_STATE_CHECK_OBJECT_NAME | Name or regex pattern of objects to check (leave empty for all) | "" |
| OBJECT_STATE_CHECK_NAMESPACE | Kubernetes namespace to check objects in | "" |
| OBJECT_STATE_CHECK_LABEL_SELECTOR | Kubernetes label selector to filter objects (e.g., app=nginx) | "" |
| OBJECT_STATE_CHECK_CONDITION_TYPE | Kubernetes condition type to check (Ready, Available, Progressing, etc.) | Ready |
| OBJECT_STATE_CHECK_CONDITION_STATUS | Expected status for the condition (True, False, or Unknown) | True |