Pod Network Scenarios

Pod outage

Scenario to block the traffic (Ingress/Egress) of a pod matching the labels for the specified duration of time to understand the behavior of the service/other services which depend on it during downtime. This helps with planning the requirements accordingly, be it improving the timeouts or tweaking the alerts etc. With the current network policies, it is not possible to explicitly block ports which are enabled by allowed network policy rule. This chaos scenario addresses this issue by using OVS flow rules to block ports related to the pod. It supports OpenShiftSDN and OVNKubernetes based networks.

Excluding Pods from Network Outage

The pod outage scenario now supports excluding specific pods from chaos testing using the exclude_label parameter. This allows you to target a namespace or group of pods with your chaos testing while deliberately preserving certain critical workloads.

Why Use Pod Exclusion?

This feature addresses several common use cases:

  • Testing resiliency of an application while keeping critical monitoring pods operational
  • Preserving designated “control plane” pods within a microservice architecture
  • Allowing targeted chaos without affecting auxiliary services in the same namespace
  • Enabling more precise pod selection when network policies require all related services to be in the same namespace
How to Use the exclude_label Parameter

The exclude_label parameter works alongside existing pod selection parameters (label_selector and pod_name). The system will:

  1. Identify all pods in the target namespace
  2. Exclude pods matching the exclude_label criteria (in format “key=value”)
  3. Apply the existing filters (label_selector or pod_name)
  4. Apply the chaos scenario to the resulting pod list
Example Configurations

Basic exclude configuration:

- id: pod_network_outage
  config:
    namespace: my-application
    label_selector: "app=my-service"
    exclude_label: "critical=true"
    direction:
      - egress
    test_duration: 600

In this example, network disruption is applied to all pods with the label app=my-service in the my-application namespace, except for those that also have the label critical=true.

Complete scenario example:

- id: pod_network_outage
  config:
    namespace: openshift-console
    direction:
      - ingress
    ingress_ports:
      - 8443
    label_selector: 'component=ui'
    exclude_label: 'excluded=true'
    test_duration: 600

This scenario blocks ingress traffic on port 8443 for pods matching component=ui label in the openshift-console namespace, but will skip any pods labeled with excluded=true.

The exclude_label parameter is also supported in the pod network shaping scenarios (pod_egress_shaping and pod_ingress_shaping), allowing for the same selective application of network latency, packet loss, and bandwidth restriction.

How to Run Pod Network Scenarios

Choose your preferred method to run pod network scenarios:

Example scenario file: pod_network_outage.yml

Sample scenario config (using a plugin)
- id: pod_network_outage
  config:
    namespace: openshift-console   # Required - Namespace of the pod to which filter need to be applied
    direction:                     # Optional - List of directions to apply filters
        - ingress                  # Blocks ingress traffic, Default both egress and ingress
    ingress_ports:                 # Optional - List of ports to block traffic on
        - 8443                     # Blocks 8443, Default [], i.e. all ports.
    label_selector: 'component=ui' # Blocks access to openshift console
    exclude_label: 'critical=true' # Optional - Pods matching this label will be excluded from the chaos
    image: quay.io/krkn-chaos/krkn:tools

Pod Network shaping

Scenario to introduce network latency, packet loss, and bandwidth restriction in the Pod’s network interface. The purpose of this scenario is to observe faults caused by random variations in the network.

Sample scenario config for egress traffic shaping (using plugin)
- id: pod_egress_shaping
  config:
    namespace: openshift-console   # Required - Namespace of the pod to which filter need to be applied.
    label_selector: 'component=ui' # Applies traffic shaping to access openshift console.
    exclude_label: 'critical=true' # Optional - Pods matching this label will be excluded from the chaos
    network_params:
        latency: 500ms             # Add 500ms latency to egress traffic from the pod.
    image: quay.io/krkn-chaos/krkn:tools
Sample scenario config for ingress traffic shaping (using plugin)
- id: pod_ingress_shaping
  config:
    namespace: openshift-console   # Required - Namespace of the pod to which filter need to be applied.
    label_selector: 'component=ui' # Applies traffic shaping to access openshift console.
    exclude_label: 'critical=true' # Optional - Pods matching this label will be excluded from the chaos
    network_params:
        latency: 500ms             # Add 500ms latency to egress traffic from the pod.
    image: quay.io/krkn-chaos/krkn:tools
Steps
  • Pick the pods to introduce the network anomaly either from label_selector or pod_name.
  • Identify the pod interface name on the node.
  • Set traffic shaping config on pod’s interface using tc and netem.
  • Wait for the duration time.
  • Remove traffic shaping config on pod’s interface.
  • Remove the job that spawned the pod.

How to Use Plugin Name

Add the plugin name to the list of chaos_scenarios section in the config/config.yaml file

kraken:
    kubeconfig_path: ~/.kube/config                     # Path to kubeconfig
    ..
    chaos_scenarios:
        - pod_network_scenarios:
            - scenarios/<scenario_name>.yaml

Run

python run_kraken.py --config config/config.yaml

This scenario runs network chaos at the pod level on a Kubernetes/OpenShift cluster.

Run

If enabling Cerberus to monitor the cluster and pass/fail the scenario post chaos, refer docs. Make sure to start it before injecting the chaos and set CERBERUS_ENABLED environment variable for the chaos injection container to autoconnect.

$ podman run \
  --name=<container_name> \
  --net=host \
  --pull=always \
  --env-host=true \
  -v <path-to-kube-config>:/home/krkn/.kube/config:Z \
  -d containers.krkn-chaos.dev/krkn-chaos/krkn-hub:pod-network-chaos
$ podman logs -f <container_name or container_id> # Streams Kraken logs
$ podman inspect <container-name or container-id> \
  --format "{{.State.ExitCode}}" # Outputs exit code which can considered as pass/fail for the scenario
$ docker run $(./get_docker_params.sh) \
  --name=<container_name> \
  --net=host \
  --pull=always \
  -v <path-to-kube-config>:/home/krkn/.kube/config:Z \
  -d containers.krkn-chaos.dev/krkn-chaos/krkn-hub:pod-network-chaos
$ docker run \
  -e <VARIABLE>=<value> \
  --name=<container_name> \
  --net=host \
  --pull=always \
  -v <path-to-kube-config>:/home/krkn/.kube/config:Z \
  -d containers.krkn-chaos.dev/krkn-chaos/krkn-hub:pod-network-chaos

$ docker logs -f <container_name or container_id> # Streams Kraken logs
$ docker inspect <container-name or container-id> \
  --format "{{.State.ExitCode}}" # Outputs exit code which can considered as pass/fail for the scenario

Supported parameters

The following environment variables can be set on the host running the container to tweak the scenario/faults being injected:

Example if –env-host is used:

export <parameter_name>=<value>

OR on the command line like example:

-e <VARIABLE>=<value>

See list of variables that apply to all scenarios here that can be used/set in addition to these scenario specific variables

Parameter DescriptionType Default
NAMESPACE Required - Namespace of the pod to which filter need to be appliedstring ""
TRAFFIC_TYPE List of directions to apply filters - egress/ingress ( needs to be a list )string [ingress, egress]
INGRESS_PORTS Ingress ports to block. Empty list ([]) blocks all ingress traffic; otherwise provide port numbers as a bracket list (e.g. [8443] or [22,80]).string []
EGRESS_PORTS Egress ports to block. Empty list ([]) blocks all egress traffic; otherwise provide port numbers as a bracket list (e.g. [8443] or [22,80]).string []
WAIT_DURATION The duration (in seconds) that the network chaos (traffic shaping, packet loss, etc.) persists on the target pods. This is the actual time window where the network disruption is active. It must be longer than TEST_DURATION to ensure the fault is active for the entire test.number 360
LABEL_SELECTOR Label of the pod(s) to targetstring ""
EXCLUDE_LABEL Pods matching this label will be excluded from the chaos even if they match other criteriastring ""
POD_NAME When label_selector is not specified, pod matching the name will be selected for the chaos scenariostring ""
INSTANCE_COUNT Number of pods to perform action/select that match the label selectornumber 1
TEST_DURATION Duration of the test run (e.g. workload or verification)number 120
SCENARIO_TYPE Plugin key krkn dispatches the scenario on. Fixed by the container image, not a knob to change.- pod_network_scenarios
SCENARIO_FILE Path to the scenario file baked into the container image. Fixed by the image, not a knob to change.- scenarios/pod_network_scenario.yaml
IMAGE Image used to disrupt network on a podstring quay.io/krkn-chaos/krkn:tools

Parameter Dependencies

  • INGRESS_PORTS / EGRESS_PORTS: When left empty ([]), all ports are blocked for that traffic direction. Specify port numbers as a bracket list (e.g. [8443] or [22,80]) to restrict the filter to only those ports.
  • WAIT_DURATION: Must be longer than TEST_DURATION to ensure the network disruption is active for the entire test run.

For example:

$ podman run \
  --name=<container_name> \
  --net=host \
  --pull=always \
  --env-host=true \
  -v <path-to-custom-metrics-profile>:/home/krkn/kraken/config/metrics-aggregated.yaml \
  -v <path-to-custom-alerts-profile>:/home/krkn/kraken/config/alerts \
  -v <path-to-kube-config>:/home/krkn/.kube/config:Z \
  -d containers.krkn-chaos.dev/krkn-chaos/krkn-hub:pod-network-chaos
krknctl run pod-network-chaos [--<parameter> <value>]

Can also set any global variable listed here

Scenario specific parameters:

Parameter DescriptionType RequiredDefault
--namespace Namespace of the pod to which filter need to be appliedstring Yes -
--image Image used to disrupt network on a podstring No quay.io/krkn-chaos/krkn:tools
--label-selector When pod_name is not specified, pod matching the label will be selected for the chaos scenariostring No ""
--exclude-label Pods matching this label will be excluded from the chaos even if they match other criteriastring No ""
--pod-name When label_selector is not specified, pod matching the name will be selected for the chaos scenariostring No ""
--instance-count Targeted instance count matching the label selectornumber No 1
--traffic-type List of directions to apply filters - egress/ingress ( needs to be a list )string No [ingress,egress]
--ingress-ports Ingress ports to block. Use [] to block all ingress ports, or a bracket list of port numbers (e.g. [22,80]).string No []
--egress-ports Egress ports to block. Use [] to block all egress ports, or a bracket list of port numbers (e.g. [22,80]).string No []
--wait-duration Ensure that it is at least about twice of test_durationnumber No 300
--test-duration Duration of the test runnumber No 120

Parameter Dependencies

  • --ingress-ports / --egress-ports: When left empty, all ports are blocked for that traffic direction. Specify port numbers to restrict the filter to only those ports.
  • --wait-duration: Must be at least 2× --test-duration to allow the network to stabilize before verification.

To see all available scenario options

krknctl run pod-network-chaos --help