Files
fips/testing/acl-allowlist
Johnathan Corgan e4a854f6b0 Stop concurrent local CI runs from clobbering each other's containers
Two runs on one host destroyed each other's containers, producing
mid-test "No such container" failures that look like real defects. The
automated builder runs a full local CI on the same box every few minutes,
so the machine is contended almost always and this has red-ed both a hand
run and an automated gate.

Two independent causes are fixed here. Container names were hardcoded, and
docker names are global rather than scoped by compose project, so two runs
collided on the same name; every name now takes an optional suffix that
the harness sets from the run id. And the cleanup sweep matched a label
shared by every run, so one run's teardown force-removed another's
containers; resources now also carry a per-run label and the sweep can be
narrowed to it.

A third hazard turned up that was not in the original report: the sidecar
suite passes explicit compose project names, which override the run-scoped
project and put it outside the shared prefix entirely. Its project names,
network, and derived container references are now scoped too.

Both are default-off. With the suffix unset, names render exactly as they
do today and a bare compose invocation is unchanged, which is what keeps
the hosted CI and the documentation correct. A cleanup run with no run id
still reaps everything, which is what a manual "clear the box" wants.

The literal-name sweep was not sufficient: six scripts build container
names dynamically from node labels, and two suites create their own
containers outside compose. Those are handled at their construction sites.

Verified: syntax check on all modified scripts; compose validation on
every modified file with the suffix both set and unset; and a synthetic
two-run reproduction that shows the old cleanup destroying a bystander run
and the new one leaving it alone.

Known gap: subnets are still hardcoded, so two concurrent full runs will
still collide on address-pool overlap. That fix reaches into topology
configs, chaos scenarios, diagrams and production source, so it is left
for its own change rather than half-done here.
2026-07-19 06:41:25 +00:00
..

ACL Allowlist Test

Six Docker nodes use per-node ACL files mounted at the hardcoded runtime paths:

  • node-a and node-b carry the insider allowlist (node-a, node-b, node-e, node-f)
  • node-c and node-d each carry a broad allowlist containing every node alias
  • node-e and node-f do not mount any ACL files locally
  • every node gets a generated /etc/fips/hosts with aliases for node-a through node-f

This lets us test three different node behaviors at once:

  • insiders (a, b) explicitly allow a, b, e, and f
  • outsiders (c, d) allow everyone locally, but still cannot join because insiders reject them
  • allowed remotes (e, f) rely on the insider ACLs and do not need local ACL files

Test Identities

Allowed:

  • node-a
    • npub1sjlh2c3x9w7kjsqg2ay080n2lff2uvt325vpan33ke34rn8l5jcqawh57m
    • 0102030405060708090a0b0c0d0e0f101112131415161718191a1b1c1d1e1f20
  • node-b
    • npub1tdwa4vjrjl33pcjdpf2t4p027nl86xrx24g4d3avg4vwvayr3g8qhd84le
    • b102030405060708090a0b0c0d0e0f101112131415161718191a1b1c1d1e1fb0

Denied:

  • node-c
    • npub1cld9yay0u24davpu6c35l4vldrhzvaq66pcqtg9a0j2cnjrn9rtsxx2pe6
    • c102030405060708090a0b0c0d0e0f101112131415161718191a1b1c1d1e1fc0
  • node-d
    • npub1n9lpnv0592cc2ps6nm0ca3qls642vx7yjsv35rkxqzj2vgds52sqgpverl
    • d102030405060708090a0b0c0d0e0f101112131415161718191a1b1c1d1e1fd0

Additional allowed:

  • node-e
    • npub1x5z9rwzzm26q9verutx4aajhf2zw2pyp34c6whhde2zduxqav40qgq36l6
    • nsec1egyrmekfw3u4l88v8zhrak9uht503s2kvn9v49tqgp6c5l2yuxgsv386l0
  • node-f
    • npub1ytrut7gjncn2zfnhn56c0zgftf0w6p99gf6fu8j73hzw5603zglqc9av6c
    • nsec1afh3nysthqh47awpdewcw59wvvp499f8dvlyclmnv4gvpxdk56dsa6eqsn

The generated fips.key fixtures use a mix of bare hex and nsec1... values. FIPS accepts either format in key files.

Run

Build the Linux binaries and test image:

./testing/scripts/build.sh --no-docker

Start the ACL test mesh:

./testing/acl-allowlist/generate-configs.sh
docker compose -f testing/acl-allowlist/docker-compose.yml up -d --build

Or run the full integration check:

./testing/acl-allowlist/test.sh

test.sh regenerates the ACL fixtures automatically before starting Docker. The generated ACL files use alias names, and the generated hosts file makes those aliases resolvable at runtime.

The ACL harness pins the expected test entrypoint explicitly so it does not accidentally reuse an older fips-test:latest image with a different startup script.

Docker service/container/hostname identifiers in this harness intentionally use service-*, fips-acl-container-*, and host-* names so they do not collide with the logical FIPS aliases node-a through node-f. For data-plane checks and operator examples, use the explicit FIPS names such as node-a.fips and node-d.fips.

ACL paths are fixed in this branch:

  • /etc/fips/peers.allow
  • /etc/fips/peers.deny

Mounted ACL files in this harness:

  • node-a and node-b: insider allowlist plus ALL deny fallback
  • node-c and node-d: broad local allowlist used by outsider nodes trying to blend in
  • node-e and node-f: no ACL files mounted
  • all nodes: /etc/fips/hosts aliases for node-a through node-f

Generated fixture location:

  • testing/acl-allowlist/generated-configs/

Inspect peer state:

docker exec fips-acl-container-a fipsctl show peers
docker exec fips-acl-container-b fipsctl show peers
docker exec fips-acl-container-c fipsctl show peers
docker exec fips-acl-container-d fipsctl show peers
docker exec fips-acl-container-e fipsctl show peers
docker exec fips-acl-container-f fipsctl show peers

Inspect the loaded ACL state directly:

docker exec fips-acl-container-a fipsctl acl show

Use explicit .fips names when checking reachability through the FIPS overlay:

docker exec fips-acl-container-a ping node-d.fips
docker exec fips-acl-container-a ping npub1n9lpnv0592cc2ps6nm0ca3qls642vx7yjsv35rkxqzj2vgds52sqgpverl.fips

The output shows both the original ACL file entries and the resolved effective npub entries.

Expected:

  • node-a sees node-b, node-e, and node-f
  • node-b sees node-a
  • node-c sees no peers
  • node-d sees no peers
  • node-e sees node-a
  • node-f sees node-a

Visible rejection logs:

docker compose -f testing/acl-allowlist/docker-compose.yml logs -f service-a service-b service-c service-d service-e service-f

On startup, node-c and node-d immediately try their configured outbound static connection to node-a. Their own ACLs permit that attempt, but the insider ACL on node-a still rejects both peers. Because node-a also has static peer stanzas for node-c and node-d, you may see both outbound_connect and inbound_handshake rejection messages during startup. The outsider-initiated path emits messages like:

Rejected peer by ACL ... context=inbound_handshake decision=denylist match

Those messages are now emitted at debug level. This harness enables RUST_LOG=info,fips::node=debug so the ACL rejection details stay visible in test logs, and operators can temporarily raise log level the same way when diagnosing ACL issues locally.

A later ping6 from node-c.fips does not emit a new inbound_handshake message. The ping uses the data-plane session path, and since no peer session to node-a.fips was established, it just times out.

Stop and clean up:

docker compose -f testing/acl-allowlist/docker-compose.yml down