chaos: bump bloom-storm bloom_send_rate ceiling 30 → 40

The bloom-storm scenario's bloom_send_rate ceiling has been bumped
from 30 to 40 sends per node over the trailing 30 s window. A 59-rep
characterization run under `github-runner-equivalent` pressure with
per-container CPU pinning to `cpuset=0,1` (mimicking a 2-core
`ubuntu-latest` runner) measured n04 (the structural max-spike node)
at mean 24.4, P99 29, max 30. The original ceiling of 30 sat at the
lab's structural max, leaving no headroom for the asymmetric transient
spikes observed on GitHub Actions (n04=34 on master CI run 25933972365,
re-fired on run 26008950865). GHA fires do not reproduce on this lab
host even with the cpuset sidecar applied.

Rationale: lab max + ~2σ ≈ 39.4 → round to 40, giving 33 % margin over
the lab maximum while staying well below the deployment-scale storm
rate (~480× steady state). The companion `min_parent_switches` guard
is unchanged.

See README.md alongside this file for the updated threshold derivation.
This commit is contained in:
Johnathan Corgan
2026-05-24 01:39:15 +00:00
parent ce0eb71722
commit dae33d4fd1
2 changed files with 31 additions and 1 deletions

View File

@@ -106,6 +106,36 @@ If `link_swap.interval_secs` or the netem delta is changed,
recalibrate. The threshold is calibrated against the values in
`scenarios/bloom-storm.yaml` as committed and the `seed: 31` pin.
### Lab-data ceiling re-tune (2026-05-24)
The ceiling was bumped from 30 to 40 after a 59-rep characterization run
under the `github-runner-equivalent` pressure profile with per-container
CPU pinning to `cpuset=0,1` (mimicking a 2-core `ubuntu-latest` GitHub
runner). Combined pinned distribution on n04 (the structural max-spike
node — flap target at depth 2):
| metric | value |
| ------ | ----: |
| mean | 24.4 |
| sd | 4.7 |
| P90 | 28 |
| P95 | 29 |
| P99 | 29 |
| max | 30 |
The lab's structural ceiling at 30 corresponds to the bloom-advertise
rate-limit token bucket's steady-state cap of ~1 send per second over the
trailing 30 s assertion window. GHA fires at n04=34 represent transient
release of queued bloom-sends during flap-recovery windows and do not
reproduce on this lab host even with CPU-pinning sidecar (`cpuset=0,1`)
applied to every chaos-spawned container.
Rationale for ceiling = 40: lab max 30 + ~2σ headroom (≈ 39.4) rounds to
40, giving 33 % margin over the observed lab maximum while still firing
loud on a regression-class storm (the original `0caef2a` regression
scaled mesh-wide bloom traffic ~480× above steady state, far above any
plausible jitter band).
## Limitations
- The bloom-storm regression has not been confirmed-failing here

View File

@@ -115,7 +115,7 @@ assertions:
# observed steady-state ceiling on fixed master and well below
# the storm-state value on the regressed binary.
window_secs: 30
max_per_node: 30
max_per_node: 40
min_parent_switches:
# Sanity guard: detect a misconfigured harness where the flap
# inducer fires but the topology never produces a real parent