Krkn-Hub All Scenarios Variables

These variables are to be used for the top level configuration template that are shared by all the scenarios in Krkn-hub.

Each section below corresponds to a section in the Krkn config reference. Set variables on the host running the container:

export <parameter_name>=<value>

Kraken

Signal and status publishing settings. See Kraken config for full details.

Parameter Description Default
KRKN_KUBE_CONFIG Path to the kubeconfig file for cluster access /home/krkn/.kube/config
DISABLE_IMAGE_SIGNATURE Disables workload image signature verification before deployment False
PUBLISH_KRAKEN_STATUS Publish kraken status to the signal address False
SIGNAL_ADDRESS Address to publish kraken status to 0.0.0.0
PORT Port to publish kraken status to 8081
SIGNAL_STATE Waits for the RUN signal when set to PAUSE before running the scenarios, refer docs for more details RUN
REPORT_FORMATS List of report formats to generate after each scenario (e.g., pdf, html) [pdf,html]
KRKN_DEBUG Enables debug mode for Krkn False

Cerberus

Cluster health monitoring integration. See Cerberus config for full details.

Parameter Description Default
CERBERUS_ENABLED Set this to true if cerberus is running and monitoring the cluster False
CERBERUS_URL URL to poll for the go/no-go signal http://0.0.0.0:8080

Performance Monitoring

Prometheus metrics collection and alert evaluation. See Performance Monitoring config for full details.

Parameter Description Default
UUID UUID for the run; auto-generated if not set ""
PROMETHEUS_URL URL to Prometheus instance; auto-detected on OpenShift, required for Kubernetes ""
PROMETHEUS_TOKEN Bearer token for Prometheus authentication; auto-detected on OpenShift, required for Kubernetes ""
CAPTURE_METRICS Captures metrics as specified in the profile from in-cluster prometheus. Default metrics captures are listed here False
ENABLE_ALERTS Evaluates expressions from in-cluster prometheus and exits 0 or 1 based on the severity set. Default profile. False
ALERTS_PATH Path to the alerts file to use when ENABLE_ALERTS is set config/alerts.yaml
ALERTS_RUN_DURING Prometheus alert health-check phases: pre, during, post, or a list of phases during
ALERTS_EXIT_ON_FAILURE Fail the run for critical or error alert evaluations False
ALERTS_ONLY_FAILURES Include only failed alert evaluations in telemetry and reports False
METRICS_PATH Path to the metrics profile to use when CAPTURE_METRICS is set config/metrics-aggregated.yaml
CHECK_CRITICAL_ALERTS When enabled will check prometheus for critical alerts firing post chaos False

Resiliency Score

Resiliency scoring configuration. See Resiliency Score config for full details.

Parameter Description Default
RESILIENCY_RUN_MODE Resiliency scoring mode: standalone embeds score in telemetry, detailed prints JSON report to stdout, disabled turns off scoring standalone
RESILIENCY_FILE Path to a YAML file containing SLO definitions; defaults to the alerts profile or config/alerts.yaml config/alerts.yaml

Elastic

Elasticsearch storage for telemetry and metrics. See Elastic config for full details.

Parameter Description Default
ENABLE_ES Enable Elasticsearch integration False
ES_SERVER URL of the Elasticsearch instance http://0.0.0.0
ES_PORT Port of the Elasticsearch instance 443
ES_USERNAME Username for Elasticsearch authentication elastic
ES_PASSWORD Password for Elasticsearch authentication -
ES_VERIFY_CERTS Verify SSL certificates when connecting to Elasticsearch False
ES_RUN_TAG Tag to identify the run in Elasticsearch ""
ES_METRICS_INDEX Elasticsearch index for metrics data krkn-metrics
ES_ALERTS_INDEX Elasticsearch index for alerts data krkn-alerts
ES_TELEMETRY_INDEX Elasticsearch index for telemetry data krkn-telemetry

Tunings

Execution timing and iteration controls. See Tunings config for full details.

Parameter Description Default
WAIT_DURATION Duration in seconds to wait between each chaos scenario 60
ITERATIONS Number of times to execute the scenarios 1
DAEMON_MODE Iterations are set to infinity which means that the kraken will cause chaos forever False

Telemetry

Run data collection and upload settings. See Telemetry config for full details.

Parameter Description Default
TELEMETRY_ENABLED Enable/disables the telemetry collection feature False
TELEMETRY_API_URL Telemetry service endpoint https://ulnmf9xv7j.execute-api.us-west-2.amazonaws.com/produ...
TELEMETRY_USERNAME Telemetry service username redhat-chaos
TELEMETRY_PASSWORD Telemetry service password -
TELEMETRY_PROMETHEUS_BACKUP Enables/disables prometheus data collection True
TELEMETRY_FULL_PROMETHEUS_BACKUP If set to False only the /prometheus/wal folder will be downloaded False
TELEMETRY_BACKUP_THREADS Number of telemetry download/upload threads 5
TELEMETRY_ARCHIVE_PATH Local path where the archive files will be temporarily stored /tmp
TELEMETRY_MAX_RETRIES Maximum number of upload retries (if 0 will retry forever) 0
TELEMETRY_RUN_TAG If set, this will be appended to the run folder in the bucket (useful to group the runs) chaos
TELEMETRY_GROUP If set will archive the telemetry in the S3 bucket on a folder named after the value default
TELEMETRY_ARCHIVE_SIZE The size of the prometheus data archive in KB 1000
TELEMETRY_LOGS_BACKUP Logs backup to S3 False
TELEMETRY_FILTER_PATTERN Filter logs based on certain timestamp patterns ["(\\w{3}\\s\\d{1,2}\\s\\d{2}:\\d{2}:\\d{2}\\.\\d+).+","kini...
TELEMETRY_CLI_PATH OC CLI path, if not specified will be searched in $PATH ""
TELEMETRY_EVENTS_BACKUP Enables/disables events backup to S3 True

Health Checks

Application endpoint monitoring during chaos. See Health Checks config for full details.

Parameter Description Default
HEALTH_CHECK_INTERVAL Interval in seconds at which to run health checks 2
HEALTH_CHECK_URL URL to continually check and detect downtimes ""
HEALTH_CHECK_BEARER_TOKEN Bearer token used for authenticating into health check URL ""
HEALTH_CHECK_AUTH Tuple of (username, password) used for authenticating into health check URL ""
HEALTH_CHECK_EXIT_ON_FAILURE If True, exits when health check fails for application ""
HEALTH_CHECK_VERIFY Health check URL SSL validation False
HEALTH_CHECK_RUN_DURING When to run health checks: pre (before chaos), during (continuous), post (after chaos), or combination like [pre,post] during
HEALTH_CHECK_ONLY_FAILURES Only create telemetry records for failed health checks False

Virt Checks

KubeVirt VMI SSH connection monitoring during chaos. See Virt Checks config for full details.

Parameter Description Default
KUBE_VIRT_CHECK_INTERVAL Interval in seconds at which to test kubevirt connections 2
KUBE_VIRT_NAMESPACE Namespace to find VMIs in and watch ""
KUBE_VIRT_NAME Regex style name to match VMIs to watch ""
KUBE_VIRT_LABEL_SELECTOR Label selector to filter VMs in KubeVirt ""
KUBE_VIRT_FAILURES If True, will only report when ssh connections fail to VMI False
KUBE_VIRT_DISCONNECTED Use disconnected check by passing cluster API False
KUBE_VIRT_SSH_NODE If set, will be a backup way to SSH to a node. Should be a node not targeted in chaos ""
KUBE_VIRT_NODE_NAME Filter only VMI’s running a specific node name ""
KUBE_VIRT_EXIT_ON_FAIL Fails run if VMs still have false status at end of run False

Other

Parameters found in the source that no section above covers.

Parameter Description Default
RETRY_WAIT 120

Triggers

Parameters found in the source that no section above covers.

Parameter Description Default
TRIGGER_COMMAND Shell command to evaluate before chaos starts. Chaos begins only after this command exits with the expected code. Leave empty to disable triggers. ""
TRIGGER_EXPECTED_RC Expected shell command exit code for the trigger to be considered satisfied 0
TRIGGER_HTTP_URL URL to poll for the HTTP trigger ""
TRIGGER_HTTP_METHOD HTTP method to use (GET, POST, etc.) GET
TRIGGER_HTTP_EXPECTED_STATUS Expected HTTP status code 200
TRIGGER_HTTP_BEARER_TOKEN Bearer token for HTTP authentication ""
TRIGGER_HTTP_BODY_CONTAINS Substring expected in the HTTP response body ""
TRIGGER_K8S_API_VERSION Kubernetes API version for the resource to watch (e.g. apps/v1, kubevirt.io/v1). Required when using a k8s trigger condition. ""
TRIGGER_K8S_KIND Kubernetes resource kind to watch (e.g. Deployment, Pod, VirtualMachineInstanceMigration) ""
TRIGGER_K8S_NAME Name of the specific Kubernetes resource to watch ""
TRIGGER_K8S_NAMESPACE Namespace of the resource to watch. Leave empty for cluster-scoped resources like Nodes. ""
TRIGGER_K8S_CONDITION Condition expression to evaluate against the resource (e.g. status.phase == Running, status.readyReplicas >= 1) ""
TRIGGER_PROM_QUERY PromQL expression for the prometheus trigger. Chaos begins only after the query returns a non-empty result. Leave empty to disable. ""
TRIGGER_PROM_URL Prometheus API URL used by the prometheus trigger condition. Required when trigger-prom-query is set. ""
TRIGGER_PROM_TOKEN Optional bearer token for authenticating the prometheus trigger against prometheus-url ""
TRIGGERS_MODE How multiple trigger conditions are combined all_of
TRIGGERS_TIMEOUT Maximum time in seconds to wait for trigger conditions before applying on-timeout behavior 300
TRIGGERS_INTERVAL Polling interval in seconds while waiting for trigger conditions 5
TRIGGERS_ON_TIMEOUT Behavior when trigger conditions are not met before timeout skip

Object State Checks

Parameters found in the source that no section above covers.

Parameter Description Default
OBJECT_STATE_CHECK_INTERVAL How often to check object states (seconds) 2
OBJECT_STATE_CHECK_RUN_DURING When to run checks: pre (before chaos), during (continuous), post (after chaos), or combination like [pre,during] during
OBJECT_STATE_CHECK_EXIT_ON_FAILURE Exit on failure when object state check fails False
OBJECT_STATE_CHECK_ONLY_FAILURES Only create telemetry records for failed object state checks False
OBJECT_STATE_CHECK_NAME Name of the object state check (e.g., etcd-ready-pods) ""
OBJECT_STATE_CHECK_KIND Kubernetes resource kind to check (Pod, Deployment, StatefulSet, DaemonSet, ReplicaSet, Service, Job, etc.) Pod
OBJECT_STATE_CHECK_OBJECT_NAME Name or regex pattern of objects to check (leave empty for all) ""
OBJECT_STATE_CHECK_NAMESPACE Kubernetes namespace to check objects in ""
OBJECT_STATE_CHECK_LABEL_SELECTOR Kubernetes label selector to filter objects (e.g., app=nginx) ""
OBJECT_STATE_CHECK_CONDITION_TYPE Kubernetes condition type to check (Ready, Available, Progressing, etc.) Ready
OBJECT_STATE_CHECK_CONDITION_STATUS Expected status for the condition (True, False, or Unknown) True