java-topology/docs/tickets/kubernetes-0001-job-exit-code-policy-slices-contains.md
russell@unturf.com 9934133dcf whitepaper: 312 sites / 151 ecosystems — wave2+3 defect tables and PDF rebuild
Add 88 new defect entries to HIGH and MEDIUM tables:
  HIGH: mysql-0001/0002, mariadb-0001, redis-0001/0002, valkey-0001/0002, openvpn-0001,
        vlc-0001, prometheus-0001, otel-collector-0001, cockroachdb-0001..0004,
        tidb-0001..0008, kubernetes-0001/0002, go-0001, kotlin-0002, scala-0001,
        allegro5-0001, sdl2-0001, grafana-0001, clickhouse-0001, duckdb-0001,
        mongodb-0001, envoy-0001, istio-0001, cilium-0001, linkerd2-0001,
        linux-0001/0002/0003, tor-0002/0003, curl-0001, julia-0001, lua-0001,
        perl5-0001, nats-0001, spring-0003/0004, tomcat-0001, onos-0002, odl-0002

  MEDIUM: helm-0001, mariadb-0002, openssl-0001/0002, memcached-0001,
          cassandra-0001..0004, flink-0001, storm-0001/0002, zookeeper-0001..0003,
          pip-0001, gradle-0001, nginx-0001, haproxy-0001, caddy-0001, varnish-0001,
          ffmpeg-0001, gstreamer-0001, raylib-0001, love2d-0001, php-0001/0002,
          r-source-0001, cpython-0002, ruby-0001, rabbitmq-0003/0004, activemq-0001,
          ovs-0001, onos-0003, odl-0002, jetty-0001

PDF: 976K
2026-03-27 15:23:43 -04:00

2.2 KiB
Raw Blame History

kubernetes-0001: Job pod failure policy — O(n) exit-code set membership per container per rule

Target: kubernetes/kubernetes File: pkg/controller/job/pod_failure_policy.go Severity: MEDIUM Pattern: CWE-407 — linear membership test inside a loop

Description

isOnExitCodesOperatorMatching calls slices.Contains(requirement.Values, exitCode) to check whether a container exit code matches an In or NotIn policy rule. requirement.Values is an unsorted []int32 slice; every call performs an O(V) linear scan.

This function is called from getMatchingContainerFromList, which iterates over every container status in a pod — meaning every failing pod triggers O(C × R × V) work, where:

  • C = number of containers in the pod
  • R = number of OnExitCodes rules in the policy
  • V = number of values in a rule's Values list

In large batch jobs with many containers and complex failure policies this degrades to O(n²) per pod failure event processed by the job controller.

Call chain

managedJob (job_controller.go)
  → matchPodFailurePolicy (pod_failure_policy.go:44)
    → matchOnExitCodes (pod_failure_policy.go:86)
      → getMatchingContainerFromList (pod_failure_policy.go:107)
        → isOnExitCodesOperatorMatching (pod_failure_policy.go:123)
          → slices.Contains(requirement.Values, exitCode)   ← O(V) each call

Slow path

// O(V) per container per rule — called inside for _, containerStatus := range containerStatuses
func isOnExitCodesOperatorMatching(exitCode int32, requirement *...) bool {
    switch requirement.Operator {
    case ..OpIn:
        return slices.Contains(requirement.Values, exitCode)   // LINEAR
    case ..OpNotIn:
        return !slices.Contains(requirement.Values, exitCode)  // LINEAR
    }
}

Fix

Sort requirement.Values once at admission time (already validated as non-empty) and use sort.SearchInt32s / binary search at match time: O(log V). Alternatively, build a map[int32]struct{} from Values once and reuse it throughout the rule evaluation — O(1) per lookup, O(V) build amortised across all containers.

Patch

defects/kubernetes/patch/kubernetes-0001.patch

Unit test

defects/kubernetes/unit/KubernetesTest.java