Compare commits
429 Commits
v0.2.1-rel
...
master
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
80c956a6fd | ||
|
|
75d7077880 | ||
|
|
a47ddbd5a5 | ||
|
|
15db6471db | ||
|
|
146d19a8d8 | ||
|
|
bda327b5f5 | ||
|
|
ea74cd7e58 | ||
|
|
78377208af | ||
|
|
37adb13d5b | ||
|
|
26a579b1c9 | ||
|
|
93a5b71728 | ||
|
|
3ebb14eda4 | ||
|
|
4ff7de4d81 | ||
|
|
e4a854f6b0 | ||
|
|
0c3d9a0b73 | ||
|
|
281ed132f1 | ||
|
|
c5492f4572 | ||
|
|
7a97599921 | ||
|
|
7fe1d75637 | ||
|
|
e21e09d7e6 | ||
|
|
e7537929ba | ||
|
|
347cbe60bd | ||
|
|
11ec16777c | ||
|
|
5b09e22956 | ||
|
|
a382b17931 | ||
|
|
a90049d3a1 | ||
|
|
791b35c221 | ||
|
|
6a80790742 | ||
|
|
60b8acf716 | ||
|
|
4fc295d90a | ||
|
|
b3f2018fce | ||
|
|
56bbc81a40 | ||
|
|
b38f8c6ffb | ||
|
|
cf62cff5f4 | ||
|
|
7fe3388f2f | ||
|
|
252d16fab9 | ||
|
|
94d7b91244 | ||
|
|
7790eb86bd | ||
|
|
cf1c957336 | ||
|
|
5ccd95cf3f | ||
|
|
119b85d28e | ||
|
|
74245e80ac | ||
|
|
e42598a86e | ||
|
|
054d17aac5 | ||
|
|
fbb4fb8879 | ||
|
|
87399795f8 | ||
|
|
765819f52b | ||
|
|
5021197f5c | ||
|
|
31f5a8c1b7 | ||
|
|
3e7ca90212 | ||
|
|
9588c50063 | ||
|
|
f698da50b6 | ||
|
|
1f765cfd8f | ||
|
|
3b99a416ad | ||
|
|
e064c96df3 | ||
|
|
a70c725e48 | ||
|
|
c8077967cd | ||
|
|
5dfa571908 | ||
|
|
6bebca88ac | ||
|
|
5d5da69a5b | ||
|
|
e05b868cf8 | ||
|
|
0ebd1b44c0 | ||
|
|
800cfb23e3 | ||
|
|
e9112cc1bb | ||
|
|
c80a7fdea5 | ||
|
|
0bf031dd32 | ||
|
|
4a0584a5e9 | ||
|
|
59155df4e3 | ||
|
|
fcaee74ec0 | ||
|
|
56e3d56c25 | ||
|
|
7b0590f70e | ||
|
|
b93a127623 | ||
|
|
85a4983dbe | ||
|
|
5090ab7851 | ||
|
|
03ced618ce | ||
|
|
bf81f422ea | ||
|
|
a0cf593580 | ||
|
|
5d08d27d3c | ||
|
|
b676c9d83a | ||
|
|
a45eefb58a | ||
|
|
d61d189572 | ||
|
|
d6ca632251 | ||
|
|
6c5fd3f4b0 | ||
|
|
434b9726aa | ||
|
|
26d70ebb59 | ||
|
|
567e6a535e | ||
|
|
6011d233c1 | ||
|
|
cb5a32693e | ||
|
|
9b46b6fa85 | ||
|
|
cbc089b820 | ||
|
|
4c95be0000 | ||
|
|
e362ab67a6 | ||
|
|
6c9f55ea80 | ||
|
|
89a31fd555 | ||
|
|
ab0a46f2c0 | ||
|
|
e839aead7a | ||
|
|
6d6889d0f6 | ||
|
|
196d9492da | ||
|
|
0f2e91b479 | ||
|
|
32475d859e | ||
|
|
f2e6b8befb | ||
|
|
1aacdfa086 | ||
|
|
a7dfe47663 | ||
|
|
8aab71af86 | ||
|
|
2cffc10520 | ||
|
|
81e4207631 | ||
|
|
1c1ed0d939 | ||
|
|
1c41f73931 | ||
|
|
39ad4d2e67 | ||
|
|
4d2504f59d | ||
|
|
1208f6a5c2 | ||
|
|
3f80530cc5 | ||
|
|
3b401a0cbd | ||
|
|
0b2212e1e8 | ||
|
|
b2ce7cd3c8 | ||
|
|
e3e03f6a5d | ||
|
|
2b009196b5 | ||
|
|
5d13090d8f | ||
|
|
6538731176 | ||
|
|
9697026c81 | ||
|
|
a2400d823f | ||
|
|
dc9334e725 | ||
|
|
309a91d293 | ||
|
|
c1ddbf053c | ||
|
|
b53db662c3 | ||
|
|
4ad5940114 | ||
|
|
4ed674ea8b | ||
|
|
a67801099d | ||
|
|
50a595a0ed | ||
|
|
4802792e38 | ||
|
|
9ea57b483a | ||
|
|
e03b206f62 | ||
|
|
1dbfefc9d0 | ||
|
|
1d277e67c7 | ||
|
|
793f844448 | ||
|
|
2491091868 | ||
|
|
243bd7985a | ||
|
|
965de26239 | ||
|
|
ab915d0479 | ||
|
|
30c5808e09 | ||
|
|
8f30924fc7 | ||
|
|
3c9a629ad4 | ||
|
|
d5ee526f0e | ||
|
|
22a5b3e5c6 | ||
|
|
3ea7ca1fd1 | ||
|
|
262d98a8eb | ||
|
|
274b09d4ff | ||
|
|
0f1fd18c25 | ||
|
|
225fab29ab | ||
|
|
effd69bd53 | ||
|
|
3749853716 | ||
|
|
289e5f8571 | ||
|
|
3733349d33 | ||
|
|
9a9e90a32c | ||
|
|
759f199518 | ||
|
|
d3cf1d6f25 | ||
|
|
e03a1ac50b | ||
|
|
507086e39d | ||
|
|
3e0d9f5726 | ||
|
|
4e3890a780 | ||
|
|
a308e71ca1 | ||
|
|
3d771c6688 | ||
|
|
4e43cb81e9 | ||
|
|
fb8bb4fb97 | ||
|
|
5eac3a98f3 | ||
|
|
a4802ccf9e | ||
|
|
7a74fa8ca2 | ||
|
|
f5f4ebe76f | ||
|
|
fd30ab0994 | ||
|
|
5fc2359432 | ||
|
|
81cd10d5db | ||
|
|
063c3a194a | ||
|
|
c77e564462 | ||
|
|
1f457d84f9 | ||
|
|
bdf571a2b2 | ||
|
|
2eea20a216 | ||
|
|
ea9c7f2d8d | ||
|
|
e09d9f8412 | ||
|
|
f3eb5bf4c2 | ||
|
|
44f7451828 | ||
|
|
c2fb12d997 | ||
|
|
bca981b79f | ||
|
|
03f7511a0e | ||
|
|
d364933ca5 | ||
|
|
42011a9a2f | ||
|
|
1b7528ce89 | ||
|
|
d548add18d | ||
|
|
87bf17dd4d | ||
|
|
79b945b93d | ||
|
|
974e146bb9 | ||
|
|
180950badf | ||
|
|
e5372cbe0f | ||
|
|
9dcc421f6f | ||
|
|
86c043cc94 | ||
|
|
555d00cfa6 | ||
|
|
dd4074249c | ||
|
|
43ad2ae946 | ||
|
|
8fd515e81f | ||
|
|
bf4e0df8c5 | ||
|
|
c7218d8486 | ||
|
|
3bc8e5611c | ||
|
|
0ce9bb5b99 | ||
|
|
e7349202b5 | ||
|
|
f29c2e65fa | ||
|
|
de327e4527 | ||
|
|
0b7daeb380 | ||
|
|
4af3730be6 | ||
|
|
36c830edfd | ||
|
|
22a41cb1a0 | ||
|
|
25fe87ff60 | ||
|
|
d9a4a7807c | ||
|
|
08b8b3908e | ||
|
|
2d0e8de8c8 | ||
|
|
5987b54730 | ||
|
|
53c6c78721 | ||
|
|
3c5d9fd4f2 | ||
|
|
7d7b551ca1 | ||
|
|
da0d9d39a0 | ||
|
|
d672ed865f | ||
|
|
0bb9ce09c6 | ||
|
|
66732e89c1 | ||
|
|
8d94c0f29c | ||
|
|
e6e2a06879 | ||
|
|
6dee6dfe27 | ||
|
|
2809f0351e | ||
|
|
f6429c19d2 | ||
|
|
d575c1f986 | ||
|
|
5b229c03bf | ||
|
|
6991a152e6 | ||
|
|
d4687e5d30 | ||
|
|
df43ac79b9 | ||
|
|
0cfc85c154 | ||
|
|
18f5c12ab9 | ||
|
|
c4c3fdd94b | ||
|
|
ffd78440a8 | ||
|
|
00bd849ee1 | ||
|
|
18297283ad | ||
|
|
4d5380604a | ||
|
|
5dfbd05fe8 | ||
|
|
f396d71826 | ||
|
|
cc7f967128 | ||
|
|
dae33d4fd1 | ||
|
|
ce0eb71722 | ||
|
|
de78c94d58 | ||
|
|
050483f3bf | ||
|
|
9c0dcd0f59 | ||
|
|
7e424f34bc | ||
|
|
3fc0178192 | ||
|
|
6e5cb8965f | ||
|
|
13c9bdacac | ||
|
|
66020bc318 | ||
|
|
7a1365fb9e | ||
|
|
57a089f6c3 | ||
|
|
0a5c367edc | ||
|
|
6e7e44c8ff | ||
|
|
d418106034 | ||
|
|
79ae430725 | ||
|
|
c0ccedb491 | ||
|
|
a83342cce8 | ||
|
|
647b8155af | ||
|
|
2bc9dd557a | ||
|
|
6bd40640bf | ||
|
|
306e455513 | ||
|
|
f51dde647f | ||
|
|
49bd210480 | ||
|
|
b1af151aef | ||
|
|
59225ccfe1 | ||
|
|
b05c80e5f5 | ||
|
|
09eb5ad6bf | ||
|
|
87d1af0269 | ||
|
|
d9ab58a285 | ||
|
|
ab1e248ff4 | ||
|
|
7f518731c8 | ||
|
|
80fb086071 | ||
|
|
e9dd3167f2 | ||
|
|
4f3d2f8471 | ||
|
|
2e54edb920 | ||
|
|
7bd8d3b7a0 | ||
|
|
538ce077df | ||
|
|
6533276eda | ||
|
|
32a3b58d1f | ||
|
|
aa8f276069 | ||
|
|
9bf9701d92 | ||
|
|
212432a9c6 | ||
|
|
32697a16f0 | ||
|
|
627fd3627b | ||
|
|
1617f6ec1c | ||
|
|
025ab49d26 | ||
|
|
733ee512d3 | ||
|
|
b6bd28f77c | ||
|
|
0e57216d98 | ||
|
|
77fdd52fe0 | ||
|
|
42b88c9bb8 | ||
|
|
eaba693b18 | ||
|
|
d52d7debb7 | ||
|
|
77ecfda1a1 | ||
|
|
9b1016ffaf | ||
|
|
5cda4a9a55 | ||
|
|
8094a51a82 | ||
|
|
0cc3de3daa | ||
|
|
2d18d019d6 | ||
|
|
64cc30df12 | ||
|
|
6ce1406664 | ||
|
|
e81fd4b477 | ||
|
|
fac4450694 | ||
|
|
253dddabe3 | ||
|
|
cd56fee7cf | ||
|
|
f0bb29ff6e | ||
|
|
e471807239 | ||
|
|
c412646498 | ||
|
|
53ad528f7d | ||
|
|
b3a1fb464f | ||
|
|
6807a3213b | ||
|
|
b547dd70f5 | ||
|
|
1d7d0d2522 | ||
|
|
67e660d813 | ||
|
|
c255e3f4a2 | ||
|
|
f32bc83034 | ||
|
|
e4f37082c2 | ||
|
|
7daca6bcf1 | ||
|
|
0fcf0f6f8f | ||
|
|
9112c8f7f0 | ||
|
|
db5b6b10bd | ||
|
|
18019bb1b5 | ||
|
|
5abf9a9325 | ||
|
|
4cdf382038 | ||
|
|
a62a0a6cf4 | ||
|
|
7fc890b7a2 | ||
|
|
a78f670a6a | ||
|
|
2f95929862 | ||
|
|
bcc9c525d3 | ||
|
|
f66be793b8 | ||
|
|
5c92fffa2b | ||
|
|
43639fecb9 | ||
|
|
33a2063672 | ||
|
|
5611e976ad | ||
|
|
d822ee8b3c | ||
|
|
ff40966832 | ||
|
|
b8b1bb03a0 | ||
|
|
616010f8c8 | ||
|
|
1d8e698b57 | ||
|
|
e08f42e3cc | ||
|
|
81c0547bdf | ||
|
|
5ed2d36464 | ||
|
|
9204888a54 | ||
|
|
00f4a4c7af | ||
|
|
9c96c9193d | ||
|
|
c86dc32197 | ||
|
|
037a965a93 | ||
|
|
953137ede7 | ||
|
|
ae607431eb | ||
|
|
996a591001 | ||
|
|
da5d23ccb7 | ||
|
|
8448e38510 | ||
|
|
a41f80a776 | ||
|
|
ab2edec2c6 | ||
|
|
239cbdc4ba | ||
|
|
96c6b7dea8 | ||
|
|
e641eb5b5f | ||
|
|
37c2973e2f | ||
|
|
674c7fe1ff | ||
|
|
23c6609a6e | ||
|
|
3092c95d54 | ||
|
|
c8502cdb97 | ||
|
|
6def31bcf6 | ||
|
|
bf77ececad | ||
|
|
34e00b9f6e | ||
|
|
1e3b2c319e | ||
|
|
f16b837a12 | ||
|
|
cbc78091ab | ||
|
|
be0708ac9b | ||
|
|
ad5ad53848 | ||
|
|
ed312ac6f2 | ||
|
|
03f6db58e8 | ||
|
|
c83e14ac97 | ||
|
|
745b523ac6 | ||
|
|
5cdcff7386 | ||
|
|
83b20b3078 | ||
|
|
5087ef9a95 | ||
|
|
213c0e87c3 | ||
|
|
c009eb7514 | ||
|
|
7780dffa93 | ||
|
|
7b3c2daa12 | ||
|
|
5abae0859e | ||
|
|
2d342a4e47 | ||
|
|
5029b40d49 | ||
|
|
6698c4d669 | ||
|
|
2a943e6695 | ||
|
|
5208d3222a | ||
|
|
fe205e74de | ||
|
|
774e33fd27 | ||
|
|
7494ed058d | ||
|
|
68dafbc72a | ||
|
|
7258469b18 | ||
|
|
0d4ffc61f0 | ||
|
|
5645284893 | ||
|
|
e9da598f8a | ||
|
|
6196307f0e | ||
|
|
e693f4fb7e | ||
|
|
1e4f375dcc | ||
|
|
60e5fefb1f | ||
|
|
51119347c3 | ||
|
|
aac96510d0 | ||
|
|
4370441e48 | ||
|
|
864a8bcc9e | ||
|
|
adfbeb2348 | ||
|
|
6633d22132 | ||
|
|
0382642d1e | ||
|
|
59f21ca185 | ||
|
|
5d27efb179 | ||
|
|
7224ce34f6 | ||
|
|
8e38d889fa | ||
|
|
bd08505002 | ||
|
|
4bc30d2b8a | ||
|
|
75466ae4e8 | ||
|
|
db9549885a | ||
|
|
8f1494853a | ||
|
|
cb6f263a1d | ||
|
|
d801fd0052 | ||
|
|
88fcf57067 | ||
|
|
89352d3218 | ||
|
|
d3385b902a | ||
|
|
79a10a2700 | ||
|
|
fc8c0dce15 | ||
|
|
71e2955da0 | ||
|
|
9519dc1cf4 | ||
|
|
b8fbecc575 | ||
|
|
0ff3f029ed | ||
|
|
9d9e2b05a1 |
@@ -1,2 +1,24 @@
|
||||
[profile.ci]
|
||||
junit = { path = "junit.xml" }
|
||||
junit = { path = "junit.xml" }
|
||||
# Synthetic node tests build 250-edge meshes with one-shot UDP
|
||||
# handshakes; on shared CI runners the localhost stack still drops the
|
||||
# occasional msg1 under burst load even with the per-edge repair loop.
|
||||
# Allow a retry rather than failing the whole CI run on a single
|
||||
# dropped packet.
|
||||
retries = 2
|
||||
|
||||
[test-groups]
|
||||
node-synthetic = { max-threads = 1 }
|
||||
|
||||
# nextest runs each test in a separate process, so in-process Tokio mutexes
|
||||
# can't serialize the synthetic localhost UDP node tests on CI. Those tests
|
||||
# send one-shot handshakes without production reconnect timers; under runner
|
||||
# load even small topologies drop the lone msg1. Group all node tests so
|
||||
# they run mutually exclusive — slower CI, reliable assertions.
|
||||
[[profile.default.overrides]]
|
||||
filter = 'test(node::tests::)'
|
||||
test-group = 'node-synthetic'
|
||||
|
||||
[[profile.ci.overrides]]
|
||||
filter = 'test(node::tests::)'
|
||||
test-group = 'node-synthetic'
|
||||
|
||||
@@ -1,2 +1,5 @@
|
||||
# rustfmt bulk reformat (maint)
|
||||
13c0b70dc3111cf94fef217b0f8b5fdbe469d3eb
|
||||
|
||||
# rustfmt master-only code
|
||||
e9da598f8ab13de5dea3a1496531d675af6a0b94
|
||||
|
||||
5
.gitattributes
vendored
Normal file
@@ -0,0 +1,5 @@
|
||||
# Keep snapshot test fixtures LF-only across all platforms so that
|
||||
# Windows checkouts with core.autocrlf=true don't convert them to CRLF
|
||||
# (which would mismatch the actual JSON serialization output during
|
||||
# snapshot comparison).
|
||||
src/control/snapshots/*.json text eol=lf
|
||||
57
.github/workflows/aur-publish-git.yml
vendored
Normal file
@@ -0,0 +1,57 @@
|
||||
name: AUR Publish (fips-git)
|
||||
|
||||
on:
|
||||
workflow_dispatch:
|
||||
push:
|
||||
branches:
|
||||
- master
|
||||
paths:
|
||||
- 'packaging/aur/PKGBUILD-git'
|
||||
- 'packaging/aur/fips.sysusers'
|
||||
- 'packaging/aur/fips.tmpfiles'
|
||||
- 'packaging/aur/fips.install'
|
||||
|
||||
jobs:
|
||||
aur-publish-fips-git:
|
||||
name: Publish fips-git to AUR
|
||||
runs-on: ubuntu-latest
|
||||
|
||||
steps:
|
||||
- uses: actions/checkout@v6
|
||||
|
||||
- name: Patch PKGBUILD-git b2sums for local assets
|
||||
run: |
|
||||
set -euo pipefail
|
||||
SYSUSERS_SUM=$(b2sum packaging/aur/fips.sysusers | awk '{print $1}')
|
||||
TMPFILES_SUM=$(b2sum packaging/aur/fips.tmpfiles | awk '{print $1}')
|
||||
if [ -z "$SYSUSERS_SUM" ] || [ -z "$TMPFILES_SUM" ]; then
|
||||
echo "Failed to compute asset b2sums"; exit 1
|
||||
fi
|
||||
awk -v s1="$SYSUSERS_SUM" -v s2="$TMPFILES_SUM" '
|
||||
/^b2sums=\(/ { in_block=1; count=0 }
|
||||
in_block {
|
||||
count++
|
||||
if (count == 2) sub(/[a-f0-9]{128}/, s1)
|
||||
if (count == 3) sub(/[a-f0-9]{128}/, s2)
|
||||
if ($0 ~ /\)/) in_block=0
|
||||
}
|
||||
{ print }
|
||||
' packaging/aur/PKGBUILD-git > packaging/aur/PKGBUILD-git.new
|
||||
mv packaging/aur/PKGBUILD-git.new packaging/aur/PKGBUILD-git
|
||||
echo "Patched PKGBUILD-git b2sums:"
|
||||
awk '/^b2sums=\(/,/\)$/' packaging/aur/PKGBUILD-git
|
||||
|
||||
- name: Publish to AUR
|
||||
uses: KSXGitHub/github-actions-deploy-aur@v4.1.2
|
||||
with:
|
||||
pkgname: fips-git
|
||||
pkgbuild: packaging/aur/PKGBUILD-git
|
||||
updpkgsums: false
|
||||
assets: |
|
||||
packaging/aur/fips.sysusers
|
||||
packaging/aur/fips.tmpfiles
|
||||
packaging/aur/fips.install
|
||||
commit_username: ${{ github.repository_owner }}
|
||||
commit_email: ${{ secrets.AUR_EMAIL }}
|
||||
ssh_private_key: ${{ secrets.AUR_SSH_PRIVATE_KEY }}
|
||||
commit_message: "Update PKGBUILD-git (${{ github.sha }})"
|
||||
191
.github/workflows/aur-publish.yml
vendored
@@ -2,30 +2,197 @@ name: AUR Publish
|
||||
|
||||
on:
|
||||
push:
|
||||
branches:
|
||||
- master
|
||||
- maint
|
||||
- next
|
||||
tags:
|
||||
- 'v*'
|
||||
pull_request:
|
||||
workflow_dispatch:
|
||||
inputs:
|
||||
tag:
|
||||
description: 'Release tag to publish (e.g. v0.4.0). Defaults to the tag the workflow was dispatched from.'
|
||||
required: false
|
||||
default: ''
|
||||
pkgrel:
|
||||
description: 'AUR pkgrel to publish. Use 2+ for packaging-only republishes of an existing tag.'
|
||||
required: false
|
||||
default: '1'
|
||||
|
||||
jobs:
|
||||
aur-publish-fips:
|
||||
name: Publish fips to AUR
|
||||
# ───────────────────────────────────────────────────────────────────────────
|
||||
# Build + lint the AUR package on every trigger, matching the coverage the
|
||||
# other package workflows (linux/macos/windows/openwrt) give their artifacts:
|
||||
# branch pushes, pull requests, tags, and manual dispatch. Uses makepkg +
|
||||
# namcap in an Arch container (neither tool exists on ubuntu-latest) and builds
|
||||
# the *checked-out tree* from a local git-archive tarball, so it works for
|
||||
# branch/PR builds and unreleased rc tags whose GitHub source archive does not
|
||||
# exist yet. This job never publishes.
|
||||
# ───────────────────────────────────────────────────────────────────────────
|
||||
aur-build:
|
||||
name: Build and lint fips AUR package
|
||||
runs-on: ubuntu-latest
|
||||
continue-on-error: true
|
||||
if: "!contains(github.ref_name, '-')"
|
||||
container: archlinux:base-devel
|
||||
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
|
||||
- name: Update pkgver in PKGBUILD
|
||||
- name: Install build and lint tooling
|
||||
run: |
|
||||
VERSION="${GITHUB_REF_NAME#v}"
|
||||
sed -i "s/^pkgver=.*/pkgver=${VERSION}/" packaging/aur/PKGBUILD
|
||||
set -euo pipefail
|
||||
pacman -Sy --noconfirm --needed base-devel namcap git curl
|
||||
|
||||
- uses: actions/checkout@v6
|
||||
|
||||
- name: Resolve package version
|
||||
id: ver
|
||||
env:
|
||||
INPUT_TAG: ${{ inputs.tag }}
|
||||
INPUT_PKGREL: ${{ inputs.pkgrel }}
|
||||
run: |
|
||||
set -euo pipefail
|
||||
if [ -n "${INPUT_TAG:-}" ]; then
|
||||
RAW="${INPUT_TAG#v}"
|
||||
elif [ "${GITHUB_REF_TYPE:-}" = "tag" ]; then
|
||||
RAW="${GITHUB_REF_NAME#v}"
|
||||
else
|
||||
# Branch push / PR: derive the version from the crate manifest.
|
||||
RAW=$(grep -m1 '^version' Cargo.toml | sed -E 's/.*"([^"]+)".*/\1/')
|
||||
fi
|
||||
# makepkg forbids '-' in pkgver; map e.g. 0.4.0-rc1 -> 0.4.0rc1,
|
||||
# 0.4.0-dev -> 0.4.0dev. The build only needs an internally consistent
|
||||
# pkgver (it matches the git-archive prefix below); this is not the
|
||||
# value the real publish uses.
|
||||
VERSION="${RAW//-/}"
|
||||
PKGREL="${INPUT_PKGREL:-1}"
|
||||
case "$PKGREL" in
|
||||
''|*[!0-9]*|0) echo "pkgrel '$PKGREL' must be a positive integer"; exit 1 ;;
|
||||
esac
|
||||
echo "version=${VERSION}" >> "$GITHUB_OUTPUT"
|
||||
echo "pkgrel=${PKGREL}" >> "$GITHUB_OUTPUT"
|
||||
echo "Resolved AUR pkgver=${VERSION} pkgrel=${PKGREL}"
|
||||
|
||||
- name: Create non-root build user and fix ownership
|
||||
run: |
|
||||
set -euo pipefail
|
||||
# makepkg refuses to run as root; create an unprivileged build user
|
||||
# with passwordless sudo (needed for pacman dep installs during -s).
|
||||
useradd -m -s /bin/bash builder
|
||||
echo 'builder ALL=(ALL) NOPASSWD: ALL' > /etc/sudoers.d/builder
|
||||
chmod 0440 /etc/sudoers.d/builder
|
||||
# The checkout is owned by root; hand it to the build user.
|
||||
chown -R builder:builder "$GITHUB_WORKSPACE"
|
||||
|
||||
- name: Build a local source tarball of the checkout
|
||||
env:
|
||||
VERSION: ${{ steps.ver.outputs.version }}
|
||||
run: |
|
||||
set -euo pipefail
|
||||
# The PKGBUILD source= points at GitHub archive/<tag>.tar.gz, which does
|
||||
# not exist for a branch push, a PR, or an unreleased rc tag and would
|
||||
# 404. Instead build the checked-out tree by packing it into a local
|
||||
# tarball whose top-level directory matches what the PKGBUILD expects
|
||||
# ("fips-<pkgver>/"); patch-pkgbuild.sh repoints source= at it.
|
||||
TARBALL="packaging/aur/fips-${VERSION}.tar.gz"
|
||||
git config --global --add safe.directory "$GITHUB_WORKSPACE"
|
||||
git -C "$GITHUB_WORKSPACE" archive --format=tar.gz \
|
||||
--prefix="fips-${VERSION}/" -o "$TARBALL" HEAD
|
||||
chown builder:builder "$TARBALL"
|
||||
ls -l "$TARBALL"
|
||||
|
||||
- name: Patch PKGBUILD
|
||||
env:
|
||||
TAG: v${{ steps.ver.outputs.version }}
|
||||
VERSION: ${{ steps.ver.outputs.version }}
|
||||
PKGREL: ${{ steps.ver.outputs.pkgrel }}
|
||||
run: |
|
||||
set -euo pipefail
|
||||
LOCAL_TARBALL="packaging/aur/fips-${VERSION}.tar.gz" \
|
||||
bash packaging/aur/patch-pkgbuild.sh
|
||||
chown builder:builder packaging/aur/PKGBUILD
|
||||
|
||||
- name: makepkg build and namcap lint (as build user)
|
||||
run: |
|
||||
set -euo pipefail
|
||||
sudo -u builder bash -euo pipefail -c '
|
||||
cd packaging/aur
|
||||
echo "::group::namcap PKGBUILD"
|
||||
namcap PKGBUILD
|
||||
echo "::endgroup::"
|
||||
echo "::group::makepkg build"
|
||||
# --nocheck: skip the PKGBUILD check() (cargo test --lib); the test
|
||||
# suite is already covered by ci.yml. This job validates packaging.
|
||||
makepkg -s --noconfirm --nocheck
|
||||
echo "::endgroup::"
|
||||
echo "::group::namcap built package"
|
||||
for pkg in *.pkg.tar.*; do
|
||||
echo "namcap $pkg"
|
||||
namcap "$pkg"
|
||||
done
|
||||
echo "::endgroup::"
|
||||
'
|
||||
|
||||
# ───────────────────────────────────────────────────────────────────────────
|
||||
# Publish to the AUR. Runs only on a real (non-prerelease) release tag push,
|
||||
# or a manual dispatch (packaging-only republish with explicit tag + pkgrel).
|
||||
# Branch pushes and pull requests build+lint above but never reach this job.
|
||||
# Gated on aur-build so a package that fails to build/lint is never published.
|
||||
# ───────────────────────────────────────────────────────────────────────────
|
||||
aur-publish-fips:
|
||||
name: Publish fips to AUR
|
||||
needs: aur-build
|
||||
runs-on: ubuntu-latest
|
||||
if: >-
|
||||
github.event_name == 'workflow_dispatch'
|
||||
|| (github.event_name == 'push'
|
||||
&& startsWith(github.ref, 'refs/tags/v')
|
||||
&& !contains(github.ref_name, '-'))
|
||||
|
||||
steps:
|
||||
- name: Resolve release tag
|
||||
id: tag
|
||||
env:
|
||||
INPUT_TAG: ${{ inputs.tag }}
|
||||
INPUT_PKGREL: ${{ inputs.pkgrel }}
|
||||
run: |
|
||||
set -euo pipefail
|
||||
TAG="${INPUT_TAG:-$GITHUB_REF_NAME}"
|
||||
PKGREL="${INPUT_PKGREL:-1}"
|
||||
case "$TAG" in
|
||||
v*) ;;
|
||||
*) echo "Tag '$TAG' does not look like a release tag (vX.Y.Z)"; exit 1 ;;
|
||||
esac
|
||||
case "$PKGREL" in
|
||||
''|*[!0-9]*|0) echo "pkgrel '$PKGREL' must be a positive integer"; exit 1 ;;
|
||||
esac
|
||||
case "$TAG" in
|
||||
*-*)
|
||||
if [ "$GITHUB_EVENT_NAME" != "workflow_dispatch" ]; then
|
||||
echo "Pre-release tag '$TAG' — skipping AUR publish"
|
||||
exit 1
|
||||
fi
|
||||
;;
|
||||
esac
|
||||
echo "tag=${TAG}" >> "$GITHUB_OUTPUT"
|
||||
echo "version=${TAG#v}" >> "$GITHUB_OUTPUT"
|
||||
echo "pkgrel=${PKGREL}" >> "$GITHUB_OUTPUT"
|
||||
|
||||
- uses: actions/checkout@v6
|
||||
with:
|
||||
ref: ${{ steps.tag.outputs.tag }}
|
||||
|
||||
- name: Patch PKGBUILD with pkgver, pkgrel, conflicts, and b2sums
|
||||
env:
|
||||
TAG: ${{ steps.tag.outputs.tag }}
|
||||
VERSION: ${{ steps.tag.outputs.version }}
|
||||
PKGREL: ${{ steps.tag.outputs.pkgrel }}
|
||||
run: bash packaging/aur/patch-pkgbuild.sh
|
||||
|
||||
- name: Publish to AUR
|
||||
uses: KSXGitHub/github-actions-deploy-aur@v4.1.1
|
||||
uses: KSXGitHub/github-actions-deploy-aur@v4.1.2
|
||||
with:
|
||||
pkgname: fips
|
||||
pkgbuild: packaging/aur/PKGBUILD
|
||||
updpkgsums: true
|
||||
updpkgsums: false
|
||||
assets: |
|
||||
packaging/aur/fips.sysusers
|
||||
packaging/aur/fips.tmpfiles
|
||||
@@ -33,4 +200,4 @@ jobs:
|
||||
commit_username: ${{ github.repository_owner }}
|
||||
commit_email: ${{ secrets.AUR_EMAIL }}
|
||||
ssh_private_key: ${{ secrets.AUR_SSH_PRIVATE_KEY }}
|
||||
commit_message: "Update to ${{ github.ref_name }}"
|
||||
commit_message: "Update to ${{ steps.tag.outputs.tag }}"
|
||||
|
||||
622
.github/workflows/ci.yml
vendored
@@ -11,6 +11,10 @@ on:
|
||||
type: boolean
|
||||
default: false
|
||||
|
||||
concurrency:
|
||||
group: ${{ github.workflow }}-${{ github.ref }}
|
||||
cancel-in-progress: true
|
||||
|
||||
permissions:
|
||||
checks: write
|
||||
contents: read
|
||||
@@ -20,23 +24,112 @@ env:
|
||||
RUST_BACKTRACE: 1
|
||||
SOURCE_DATE_EPOCH: 0 # overridden per-step after checkout
|
||||
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# CI parity invariant
|
||||
#
|
||||
# This GitHub integration matrix and the local default suite set
|
||||
# (testing/ci-local.sh) MUST run the same integration suites, EXCEPT for the
|
||||
# deliberate local-only entries below. Adding a suite to one runner without
|
||||
# the other means "local green" and "GitHub green" stop being equivalent.
|
||||
# testing/check-ci-parity.sh enforces this and fails on unexpected drift.
|
||||
#
|
||||
# Deliberate local-only (NOT on the GitHub gate), with reason:
|
||||
# tor-socks5 — requires live Tor network; opt-in via --with-tor,
|
||||
# unreliable on GitHub-hosted runners.
|
||||
# tor-directory — same; live Tor dependency.
|
||||
#
|
||||
# Granularity-only differences (same coverage, different matrix shape —
|
||||
# NOT a divergence):
|
||||
# deb-install — split here into per-distro legs (debian12/debian13/
|
||||
# ubuntu22/ubuntu24/ubuntu26) for parallelism; local runs the
|
||||
# same distro set in one suite.
|
||||
# dns-resolver — single leg here; runs all scenarios (same as local).
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Job 1 – Build matrix
|
||||
#
|
||||
# Builds on Linux x86_64 and Linux aarch64. macOS and Windows are in a
|
||||
# separate ci-compat.yml workflow so their failures don't mark this run red.
|
||||
# Builds on Linux x86_64, Linux aarch64, and macOS.
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
jobs:
|
||||
fmt:
|
||||
name: Format check
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
- uses: dtolnay/rust-toolchain@stable
|
||||
- uses: actions/checkout@v6
|
||||
- uses: actions-rust-lang/setup-rust-toolchain@v1
|
||||
with:
|
||||
components: rustfmt
|
||||
cache: false
|
||||
rustflags: ''
|
||||
- run: cargo fmt --check
|
||||
|
||||
clippy:
|
||||
name: Clippy
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
- uses: actions/checkout@v6
|
||||
- name: Install system dependencies
|
||||
run: sudo apt-get update && sudo apt-get install -y libdbus-1-dev
|
||||
- uses: actions-rust-lang/setup-rust-toolchain@v1
|
||||
with:
|
||||
components: clippy
|
||||
cache: false
|
||||
rustflags: ''
|
||||
- name: Cache Cargo registry + build
|
||||
uses: actions/cache@v5
|
||||
with:
|
||||
path: |
|
||||
~/.cargo/registry
|
||||
~/.cargo/git
|
||||
target
|
||||
key: ${{ runner.os }}-cargo-clippy-${{ hashFiles('**/Cargo.lock') }}
|
||||
restore-keys: |
|
||||
${{ runner.os }}-cargo-
|
||||
- run: cargo clippy --all-targets --all-features -- -D warnings
|
||||
|
||||
# ───────────────────────────────────────────────────────────────────────────
|
||||
# Android cross-check
|
||||
#
|
||||
# FIPS runs on Android as an embedded library — the host app owns the TUN
|
||||
# (an Android VpnService), so there are no daemon binaries to package, unlike
|
||||
# the desktop targets. This job only cross-compiles the library for the
|
||||
# android target to guard the android-only cfg paths (and the `not(android)`
|
||||
# exclusions) from silently bit-rotting; nothing else in CI compiles them.
|
||||
# cargo-ndk wires the NDK toolchain, which is required even for a check
|
||||
# because `ring` compiles C at build time.
|
||||
# ───────────────────────────────────────────────────────────────────────────
|
||||
android-check:
|
||||
name: Android cross-check (aarch64)
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
- uses: actions/checkout@v6
|
||||
- name: Install Rust toolchain (+ Android target)
|
||||
uses: actions-rust-lang/setup-rust-toolchain@v1
|
||||
with:
|
||||
target: aarch64-linux-android
|
||||
components: clippy
|
||||
cache: false
|
||||
rustflags: ''
|
||||
- name: Cache Cargo registry + build
|
||||
uses: actions/cache@v5
|
||||
with:
|
||||
path: |
|
||||
~/.cargo/registry
|
||||
~/.cargo/git
|
||||
target
|
||||
key: ${{ runner.os }}-cargo-android-${{ hashFiles('**/Cargo.lock') }}
|
||||
restore-keys: |
|
||||
${{ runner.os }}-cargo-
|
||||
- name: Install cargo-ndk
|
||||
uses: taiki-e/install-action@v2
|
||||
with:
|
||||
tool: cargo-ndk
|
||||
- name: Clippy the library for Android
|
||||
run: |
|
||||
export ANDROID_NDK_HOME="${ANDROID_NDK_HOME:-$ANDROID_NDK_LATEST_HOME}"
|
||||
cargo ndk -t arm64-v8a clippy --lib -- -D warnings
|
||||
|
||||
build:
|
||||
name: Build (${{ matrix.os }})
|
||||
runs-on: ${{ matrix.os }}
|
||||
@@ -47,18 +140,39 @@ jobs:
|
||||
include:
|
||||
- os: ubuntu-latest
|
||||
- os: ubuntu-24.04-arm
|
||||
- os: macos-latest
|
||||
- os: windows-latest
|
||||
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
- uses: actions/checkout@v6
|
||||
|
||||
- name: Set SOURCE_DATE_EPOCH from git
|
||||
- name: Set SOURCE_DATE_EPOCH from git (Unix)
|
||||
if: runner.os != 'Windows'
|
||||
run: echo "SOURCE_DATE_EPOCH=$(git log -1 --format=%ct)" >> "$GITHUB_ENV"
|
||||
|
||||
- name: Set SOURCE_DATE_EPOCH from git (Windows)
|
||||
if: runner.os == 'Windows'
|
||||
shell: pwsh
|
||||
run: |
|
||||
$epoch = git log -1 --format=%ct
|
||||
echo "SOURCE_DATE_EPOCH=$epoch" >> $env:GITHUB_ENV
|
||||
|
||||
- name: Install system dependencies (Linux only)
|
||||
if: runner.os == 'Linux'
|
||||
run: sudo apt-get update && sudo apt-get install -y libdbus-1-dev nftables
|
||||
|
||||
- name: Validate fips.nft syntax (Linux only)
|
||||
if: runner.os == 'Linux'
|
||||
run: sudo nft -c -f packaging/common/fips.nft
|
||||
|
||||
- name: Install Rust toolchain
|
||||
uses: dtolnay/rust-toolchain@stable
|
||||
uses: actions-rust-lang/setup-rust-toolchain@v1
|
||||
with:
|
||||
cache: false
|
||||
rustflags: ''
|
||||
|
||||
- name: Cache Cargo registry + build
|
||||
uses: actions/cache@v4
|
||||
uses: actions/cache@v5
|
||||
with:
|
||||
path: |
|
||||
~/.cargo/registry
|
||||
@@ -71,19 +185,30 @@ jobs:
|
||||
- name: Build
|
||||
run: cargo build --release
|
||||
|
||||
- name: SHA-256 hashes
|
||||
run: sha256sum target/release/fips target/release/fipsctl target/release/fipstop
|
||||
- name: SHA-256 hashes (Linux)
|
||||
if: runner.os == 'Linux'
|
||||
run: sha256sum target/release/fips target/release/fipsctl target/release/fipstop target/release/fips-gateway
|
||||
|
||||
- name: SHA-256 hashes (macOS)
|
||||
if: runner.os == 'macOS'
|
||||
run: shasum -a 256 target/release/fips target/release/fipsctl target/release/fipstop
|
||||
|
||||
- name: SHA-256 hashes (Windows)
|
||||
if: runner.os == 'Windows'
|
||||
shell: pwsh
|
||||
run: Get-FileHash target\release\fips.exe, target\release\fipsctl.exe, target\release\fipstop.exe -Algorithm SHA256
|
||||
|
||||
# Upload the Linux binary so integration jobs can use it without rebuilding
|
||||
- name: Upload Linux binary
|
||||
if: matrix.os == 'ubuntu-latest'
|
||||
uses: actions/upload-artifact@v4
|
||||
uses: actions/upload-artifact@v7
|
||||
with:
|
||||
name: fips-linux
|
||||
path: |
|
||||
target/release/fips
|
||||
target/release/fipsctl
|
||||
target/release/fipstop
|
||||
target/release/fips-gateway
|
||||
retention-days: 1
|
||||
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
@@ -97,16 +222,22 @@ jobs:
|
||||
runs-on: ubuntu-latest
|
||||
needs: [build]
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
- uses: actions/checkout@v6
|
||||
|
||||
- name: Set SOURCE_DATE_EPOCH from git
|
||||
run: echo "SOURCE_DATE_EPOCH=$(git log -1 --format=%ct)" >> "$GITHUB_ENV"
|
||||
|
||||
- name: Install system dependencies
|
||||
run: sudo apt-get update && sudo apt-get install -y libdbus-1-dev
|
||||
|
||||
- name: Install Rust toolchain
|
||||
uses: dtolnay/rust-toolchain@stable
|
||||
uses: actions-rust-lang/setup-rust-toolchain@v1
|
||||
with:
|
||||
cache: false
|
||||
rustflags: ''
|
||||
|
||||
- name: Cache Cargo registry + build
|
||||
uses: actions/cache@v4
|
||||
uses: actions/cache@v5
|
||||
with:
|
||||
path: |
|
||||
~/.cargo/registry
|
||||
@@ -139,6 +270,103 @@ jobs:
|
||||
check_name: Unit Tests Summary
|
||||
fail_on_failure: false
|
||||
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Job 2b – Unit tests (macOS)
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
test-macos:
|
||||
name: Unit tests (macOS)
|
||||
runs-on: macos-latest
|
||||
needs: [build]
|
||||
steps:
|
||||
- uses: actions/checkout@v6
|
||||
|
||||
- name: Set SOURCE_DATE_EPOCH from git
|
||||
run: echo "SOURCE_DATE_EPOCH=$(git log -1 --format=%ct)" >> "$GITHUB_ENV"
|
||||
|
||||
- name: Install Rust toolchain
|
||||
uses: actions-rust-lang/setup-rust-toolchain@v1
|
||||
with:
|
||||
cache: false
|
||||
rustflags: ''
|
||||
|
||||
- name: Cache Cargo registry + build
|
||||
uses: actions/cache@v5
|
||||
with:
|
||||
path: |
|
||||
~/.cargo/registry
|
||||
~/.cargo/git
|
||||
target
|
||||
key: ${{ runner.os }}-cargo-${{ hashFiles('**/Cargo.lock') }}
|
||||
restore-keys: |
|
||||
${{ runner.os }}-cargo-
|
||||
|
||||
- name: Install cargo-nextest
|
||||
uses: taiki-e/install-action@nextest
|
||||
|
||||
- name: Run unit tests
|
||||
run: cargo nextest run --all --profile ci
|
||||
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Job 2c – Unit tests (Windows)
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
test-windows:
|
||||
name: Unit tests (Windows)
|
||||
runs-on: windows-latest
|
||||
steps:
|
||||
- uses: actions/checkout@v6
|
||||
|
||||
- name: Install Rust toolchain
|
||||
uses: actions-rust-lang/setup-rust-toolchain@v1
|
||||
with:
|
||||
cache: false
|
||||
rustflags: ''
|
||||
|
||||
- name: Cache Cargo registry + build
|
||||
uses: actions/cache@v5
|
||||
with:
|
||||
path: |
|
||||
~/.cargo/registry
|
||||
~/.cargo/git
|
||||
target
|
||||
key: ${{ runner.os }}-cargo-${{ hashFiles('**/Cargo.lock') }}
|
||||
restore-keys: |
|
||||
${{ runner.os }}-cargo-
|
||||
|
||||
- name: Install cargo-nextest
|
||||
uses: taiki-e/install-action@nextest
|
||||
|
||||
- name: Run unit tests
|
||||
run: cargo nextest run --all --profile ci
|
||||
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Job 2d – PowerShell lint (Windows packaging scripts)
|
||||
#
|
||||
# Runs PSScriptAnalyzer against the operator-facing installer/build
|
||||
# scripts shipped in the Windows ZIP package. Settings live in
|
||||
# packaging/windows/PSScriptAnalyzerSettings.psd1 (each suppressed rule
|
||||
# is documented there). Pre-installed on windows-latest runners; no
|
||||
# Install-Module step needed.
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
windows-lint:
|
||||
name: PowerShell lint (Windows packaging)
|
||||
runs-on: windows-latest
|
||||
steps:
|
||||
- uses: actions/checkout@v6
|
||||
|
||||
- name: Run PSScriptAnalyzer
|
||||
shell: pwsh
|
||||
run: |
|
||||
$results = Invoke-ScriptAnalyzer `
|
||||
-Path packaging/windows/*.ps1 `
|
||||
-Settings packaging/windows/PSScriptAnalyzerSettings.psd1
|
||||
if ($results) {
|
||||
$results | Format-Table -AutoSize
|
||||
Write-Error "PSScriptAnalyzer found $($results.Count) issue(s)"
|
||||
exit 1
|
||||
} else {
|
||||
Write-Host "PSScriptAnalyzer: no issues"
|
||||
}
|
||||
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Job 3 – Integration tests (static mesh + chaos simulation)
|
||||
#
|
||||
@@ -169,6 +397,25 @@ jobs:
|
||||
- suite: rekey
|
||||
type: rekey
|
||||
topology: rekey
|
||||
- suite: rekey-accept-off
|
||||
type: rekey-accept-off
|
||||
topology: rekey-accept-off
|
||||
- suite: rekey-outbound-only
|
||||
type: rekey-outbound-only
|
||||
topology: rekey-outbound-only
|
||||
# ── Inbound max_peers admission-cap test ───────────────────────
|
||||
- suite: admission-cap
|
||||
type: admission-cap
|
||||
topology: mesh
|
||||
- suite: acl-allowlist
|
||||
type: acl-allowlist
|
||||
# ── Firewall baseline (fips0 nftables default-deny) ────────────
|
||||
- suite: firewall
|
||||
type: firewall
|
||||
# ── Outbound LAN gateway integration test ──────────────────────
|
||||
- suite: gateway
|
||||
type: gateway
|
||||
topology: gateway
|
||||
# ── Chaos / stochastic scenarios ───────────────────────────────────
|
||||
- suite: chaos-smoke-10
|
||||
type: chaos
|
||||
@@ -207,16 +454,77 @@ jobs:
|
||||
- suite: congestion-stress
|
||||
type: chaos
|
||||
scenario: congestion-stress
|
||||
- suite: bloom-storm
|
||||
type: chaos
|
||||
scenario: bloom-storm
|
||||
# ── Sidecar deployment ──────────────────────────────────────────
|
||||
- suite: sidecar
|
||||
type: sidecar
|
||||
# ── NAT traversal lab (Nostr/STUN UDP hole punch) ───────────────
|
||||
- suite: nat-cone
|
||||
type: nat
|
||||
scenario: cone
|
||||
- suite: nat-symmetric
|
||||
type: nat
|
||||
scenario: symmetric
|
||||
- suite: nat-lan
|
||||
type: nat
|
||||
scenario: lan
|
||||
# ── Nostr overlay advert publish/consume round-trip ─────────────
|
||||
# Two FIPS daemons + the existing strfry relay; covers Phase 1
|
||||
# (A→B publish/consume), Phase 2 (B→A reverse), and Phase 3
|
||||
# (malformed advert injected to relay; consumers must reject
|
||||
# without crashing). UDP transport baseline for v0.3.0.
|
||||
- suite: nostr-publish-consume
|
||||
type: nostr-publish-consume
|
||||
# ── STUN fault-injection ───────────────────────────────────────
|
||||
# One FIPS daemon + a netns-sharing shim that injects tc/iptables
|
||||
# faults against UDP egress to the in-lab STUN server. Three
|
||||
# phases: 100% drop, ~5s delay then clear, then full STUN
|
||||
# container kill. Asserts the daemon notices each fault,
|
||||
# recovers from delay, and never panics.
|
||||
- suite: stun-faults
|
||||
type: stun-faults
|
||||
# ── Real-deb install across target distros ─────────────────────
|
||||
# Boots a privileged systemd container per distro, runs
|
||||
# `apt install ./fips_*.deb` with the locally-built package,
|
||||
# then asserts end-to-end `.fips` resolution + the
|
||||
# gateway/daemon default-pairing. The most thorough single
|
||||
# test surface — exercises packaging, maintainer scripts,
|
||||
# systemd unit ordering, real TUN, and the DNS responder
|
||||
# filter on a per-distro resolver backend.
|
||||
- suite: deb-install-debian12
|
||||
type: deb-install
|
||||
scenario: debian12
|
||||
- suite: deb-install-debian13
|
||||
type: deb-install
|
||||
scenario: debian13
|
||||
- suite: deb-install-ubuntu22
|
||||
type: deb-install
|
||||
scenario: ubuntu22
|
||||
- suite: deb-install-ubuntu24
|
||||
type: deb-install
|
||||
scenario: ubuntu24
|
||||
- suite: deb-install-ubuntu26
|
||||
type: deb-install
|
||||
scenario: ubuntu26
|
||||
# ── DNS resolver multi-backend coverage ────────────────────────
|
||||
# Exercises every fips-dns-setup backend (resolved, dnsmasq,
|
||||
# NM+dnsmasq, dns-delegate, no-resolver) across five distros,
|
||||
# plus end-to-end scenarios that boot a real fips daemon with a
|
||||
# real TUN and assert `dig @127.0.0.53 AAAA <npub>.fips`
|
||||
# returns AAAA. Pins the production DNS bind path that
|
||||
# ISSUE-2026-0002 lived in. Single matrix entry runs all 13
|
||||
# scenarios sequentially; ~7-12 min warm, ~12-15 min cold.
|
||||
- suite: dns-resolver
|
||||
type: dns-resolver
|
||||
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
- uses: actions/checkout@v6
|
||||
|
||||
# Fetch the pre-built Linux binary from job 1
|
||||
- name: Download Linux binary
|
||||
uses: actions/download-artifact@v4
|
||||
uses: actions/download-artifact@v8
|
||||
with:
|
||||
name: fips-linux
|
||||
path: _bin
|
||||
@@ -226,9 +534,11 @@ jobs:
|
||||
run: |
|
||||
chmod +x _bin/fips _bin/fipsctl
|
||||
[ -f _bin/fipstop ] && chmod +x _bin/fipstop || true
|
||||
[ -f _bin/fips-gateway ] && chmod +x _bin/fips-gateway || true
|
||||
cp _bin/fips testing/docker/fips
|
||||
cp _bin/fipsctl testing/docker/fipsctl
|
||||
[ -f _bin/fipstop ] && cp _bin/fipstop testing/docker/fipstop || true
|
||||
[ -f _bin/fips-gateway ] && cp _bin/fips-gateway testing/docker/fips-gateway || true
|
||||
docker build -t fips-test:latest testing/docker
|
||||
docker build -t fips-test-app:latest -f testing/docker/Dockerfile.app testing/docker
|
||||
|
||||
@@ -288,6 +598,107 @@ jobs:
|
||||
docker compose -f testing/static/docker-compose.yml \
|
||||
--profile rekey down --volumes --remove-orphans
|
||||
|
||||
# ── Rekey + accept_connections=false variant ──────────────────────────
|
||||
- name: Generate and inject configs (rekey-accept-off)
|
||||
if: matrix.type == 'rekey-accept-off'
|
||||
env:
|
||||
REKEY_TOPOLOGY: rekey-accept-off
|
||||
REKEY_ACCEPT_OFF_NODES: b
|
||||
run: |
|
||||
bash testing/static/scripts/generate-configs.sh rekey-accept-off
|
||||
bash testing/static/scripts/rekey-test.sh inject-config
|
||||
|
||||
- name: Start containers (rekey-accept-off)
|
||||
if: matrix.type == 'rekey-accept-off'
|
||||
run: |
|
||||
docker compose -f testing/static/docker-compose.yml \
|
||||
--profile rekey-accept-off up -d
|
||||
|
||||
- name: Run rekey test (accept-off variant)
|
||||
if: matrix.type == 'rekey-accept-off'
|
||||
env:
|
||||
REKEY_TOPOLOGY: rekey-accept-off
|
||||
REKEY_ACCEPT_OFF_NODES: b
|
||||
run: bash testing/static/scripts/rekey-test.sh
|
||||
|
||||
- name: Collect logs on failure (rekey-accept-off)
|
||||
if: matrix.type == 'rekey-accept-off' && failure()
|
||||
run: |
|
||||
docker compose -f testing/static/docker-compose.yml \
|
||||
--profile rekey-accept-off logs --no-color | tail -300
|
||||
|
||||
- name: Stop containers (rekey-accept-off)
|
||||
if: matrix.type == 'rekey-accept-off' && always()
|
||||
run: |
|
||||
docker compose -f testing/static/docker-compose.yml \
|
||||
--profile rekey-accept-off down --volumes --remove-orphans
|
||||
|
||||
# ── Rekey + udp.outbound_only=true variant ─────────────────────────────
|
||||
- name: Generate and inject configs (rekey-outbound-only)
|
||||
if: matrix.type == 'rekey-outbound-only'
|
||||
env:
|
||||
REKEY_TOPOLOGY: rekey-outbound-only
|
||||
REKEY_OUTBOUND_ONLY_NODES: b
|
||||
run: |
|
||||
bash testing/static/scripts/generate-configs.sh rekey-outbound-only
|
||||
bash testing/static/scripts/rekey-test.sh inject-config
|
||||
|
||||
- name: Start containers (rekey-outbound-only)
|
||||
if: matrix.type == 'rekey-outbound-only'
|
||||
run: |
|
||||
docker compose -f testing/static/docker-compose.yml \
|
||||
--profile rekey-outbound-only up -d
|
||||
|
||||
- name: Run rekey test (outbound-only variant)
|
||||
if: matrix.type == 'rekey-outbound-only'
|
||||
env:
|
||||
REKEY_TOPOLOGY: rekey-outbound-only
|
||||
REKEY_OUTBOUND_ONLY_NODES: b
|
||||
run: bash testing/static/scripts/rekey-test.sh
|
||||
|
||||
- name: Collect logs on failure (rekey-outbound-only)
|
||||
if: matrix.type == 'rekey-outbound-only' && failure()
|
||||
run: |
|
||||
docker compose -f testing/static/docker-compose.yml \
|
||||
--profile rekey-outbound-only logs --no-color | tail -300
|
||||
|
||||
- name: Stop containers (rekey-outbound-only)
|
||||
if: matrix.type == 'rekey-outbound-only' && always()
|
||||
run: |
|
||||
docker compose -f testing/static/docker-compose.yml \
|
||||
--profile rekey-outbound-only down --volumes --remove-orphans
|
||||
|
||||
# ── ACL allowlist integration test ─────────────────────────────────────
|
||||
- name: Run ACL allowlist integration test
|
||||
if: matrix.type == 'acl-allowlist'
|
||||
run: bash testing/acl-allowlist/test.sh --skip-build --keep-up
|
||||
|
||||
- name: Collect logs on failure (acl-allowlist)
|
||||
if: matrix.type == 'acl-allowlist' && failure()
|
||||
run: |
|
||||
docker compose -f testing/acl-allowlist/docker-compose.yml logs --no-color
|
||||
|
||||
- name: Stop containers (acl-allowlist)
|
||||
if: matrix.type == 'acl-allowlist' && always()
|
||||
run: |
|
||||
docker compose -f testing/acl-allowlist/docker-compose.yml down --volumes --remove-orphans
|
||||
|
||||
# ── Firewall baseline integration test ─────────────────────────────────
|
||||
- name: Run firewall baseline integration test
|
||||
if: matrix.type == 'firewall'
|
||||
run: bash testing/firewall/test.sh --skip-build --keep-up
|
||||
|
||||
- name: Collect logs on failure (firewall)
|
||||
if: matrix.type == 'firewall' && failure()
|
||||
run: |
|
||||
docker compose -f testing/firewall/docker-compose.yml logs --no-color
|
||||
docker exec fips-fw-container-b nft list table inet fips || true
|
||||
|
||||
- name: Stop containers (firewall)
|
||||
if: matrix.type == 'firewall' && always()
|
||||
run: |
|
||||
docker compose -f testing/firewall/docker-compose.yml down --volumes --remove-orphans
|
||||
|
||||
# ── Chaos simulation ───────────────────────────────────────────────────
|
||||
- name: Install Python deps (chaos)
|
||||
if: matrix.type == 'chaos'
|
||||
@@ -299,7 +710,7 @@ jobs:
|
||||
|
||||
- name: Upload sim results on failure (chaos)
|
||||
if: matrix.type == 'chaos' && failure()
|
||||
uses: actions/upload-artifact@v4
|
||||
uses: actions/upload-artifact@v7
|
||||
with:
|
||||
name: sim-results-${{ matrix.scenario }}
|
||||
path: testing/chaos/sim-results/
|
||||
@@ -318,3 +729,180 @@ jobs:
|
||||
docker logs "sidecar-${node}-fips-1" 2>&1 || true
|
||||
echo ""
|
||||
done
|
||||
|
||||
# ── NAT traversal lab ───────────────────────────────────────────────
|
||||
- name: Run NAT lab scenario
|
||||
if: matrix.type == 'nat'
|
||||
run: bash testing/nat/scripts/nat-test.sh ${{ matrix.scenario }}
|
||||
|
||||
- name: Collect logs on failure (nat)
|
||||
if: matrix.type == 'nat' && failure()
|
||||
run: |
|
||||
docker compose -f testing/nat/docker-compose.yml \
|
||||
--profile ${{ matrix.scenario }} logs --no-color
|
||||
|
||||
- name: Stop containers (nat)
|
||||
if: matrix.type == 'nat' && always()
|
||||
run: |
|
||||
docker compose -f testing/nat/docker-compose.yml \
|
||||
--profile cone --profile symmetric --profile lan \
|
||||
down --volumes --remove-orphans
|
||||
|
||||
# ── Nostr overlay advert publish/consume ───────────────────────────
|
||||
- name: Run Nostr publish/consume test
|
||||
if: matrix.type == 'nostr-publish-consume'
|
||||
run: bash testing/nat/scripts/nostr-relay-test.sh
|
||||
|
||||
- name: Collect logs on failure (nostr-publish-consume)
|
||||
if: matrix.type == 'nostr-publish-consume' && failure()
|
||||
run: |
|
||||
docker compose -f testing/nat/docker-compose.yml \
|
||||
--profile nostr-publish-consume logs --no-color | tail -300
|
||||
|
||||
- name: Stop containers (nostr-publish-consume)
|
||||
if: matrix.type == 'nostr-publish-consume' && always()
|
||||
run: |
|
||||
docker compose -f testing/nat/docker-compose.yml \
|
||||
--profile nostr-publish-consume down --volumes --remove-orphans
|
||||
|
||||
# ── STUN fault-injection ───────────────────────────────────────────
|
||||
- name: Run STUN fault-injection test
|
||||
if: matrix.type == 'stun-faults'
|
||||
run: bash testing/nat/scripts/stun-faults-test.sh
|
||||
|
||||
- name: Collect logs on failure (stun-faults)
|
||||
if: matrix.type == 'stun-faults' && failure()
|
||||
run: |
|
||||
docker compose -f testing/nat/docker-compose.yml \
|
||||
--profile stun-faults logs --no-color | tail -300
|
||||
|
||||
- name: Stop containers (stun-faults)
|
||||
if: matrix.type == 'stun-faults' && always()
|
||||
run: |
|
||||
docker compose -f testing/nat/docker-compose.yml \
|
||||
--profile stun-faults down --volumes --remove-orphans
|
||||
|
||||
# ── Outbound LAN gateway integration test ──────────────────────────
|
||||
- name: Generate configs (gateway)
|
||||
if: matrix.type == 'gateway'
|
||||
run: bash testing/static/scripts/generate-configs.sh gateway gateway-test
|
||||
|
||||
- name: Inject gateway config (gateway)
|
||||
if: matrix.type == 'gateway'
|
||||
run: bash testing/static/scripts/gateway-test.sh inject-config
|
||||
|
||||
- name: Start containers (gateway)
|
||||
if: matrix.type == 'gateway'
|
||||
run: |
|
||||
docker compose -f testing/static/docker-compose.yml \
|
||||
--profile gateway up -d
|
||||
|
||||
- name: Run gateway test
|
||||
if: matrix.type == 'gateway'
|
||||
run: bash testing/static/scripts/gateway-test.sh
|
||||
|
||||
- name: Collect logs on failure (gateway)
|
||||
if: matrix.type == 'gateway' && failure()
|
||||
run: |
|
||||
docker compose -f testing/static/docker-compose.yml \
|
||||
--profile gateway logs --no-color | tail -300
|
||||
|
||||
- name: Stop containers (gateway)
|
||||
if: matrix.type == 'gateway' && always()
|
||||
run: |
|
||||
docker compose -f testing/static/docker-compose.yml \
|
||||
--profile gateway down --volumes --remove-orphans
|
||||
|
||||
# ── Inbound max_peers admission-cap integration test ────────────────
|
||||
# Lowers node.max_peers on one mesh node and asserts the inbound cap
|
||||
# holds under sustained retry pressure: denied peers keep retrying but
|
||||
# are never promoted to an active session. The admission-cap-test.sh
|
||||
# assertions are tailored per link-layer handshake variant; the leg
|
||||
# itself is uniform. Static-style harness on the shared mesh profile.
|
||||
- name: Generate configs (admission-cap)
|
||||
if: matrix.type == 'admission-cap'
|
||||
run: bash testing/static/scripts/generate-configs.sh mesh
|
||||
|
||||
- name: Inject admission-cap config (admission-cap)
|
||||
if: matrix.type == 'admission-cap'
|
||||
run: bash testing/static/scripts/admission-cap-test.sh inject-config
|
||||
|
||||
- name: Start containers (admission-cap)
|
||||
if: matrix.type == 'admission-cap'
|
||||
run: |
|
||||
docker compose -f testing/static/docker-compose.yml \
|
||||
--profile mesh up -d
|
||||
|
||||
- name: Run admission-cap test
|
||||
if: matrix.type == 'admission-cap'
|
||||
run: bash testing/static/scripts/admission-cap-test.sh
|
||||
|
||||
- name: Collect logs on failure (admission-cap)
|
||||
if: matrix.type == 'admission-cap' && failure()
|
||||
run: |
|
||||
docker compose -f testing/static/docker-compose.yml \
|
||||
--profile mesh logs --no-color | tail -300
|
||||
|
||||
- name: Stop containers (admission-cap)
|
||||
if: matrix.type == 'admission-cap' && always()
|
||||
run: |
|
||||
docker compose -f testing/static/docker-compose.yml \
|
||||
--profile mesh down --volumes --remove-orphans
|
||||
|
||||
# ── Real-deb install integration ────────────────────────────────────
|
||||
# The deb-install harness builds its own .deb from source in a
|
||||
# cargo-deb builder image; the pre-built Linux binary from the
|
||||
# build job is intentionally not used here so the test exercises
|
||||
# the full packaging pipeline. ~5-7 min cold-cache on a fresh
|
||||
# runner (.deb build dominates), ~1-2 min warm-cache.
|
||||
- name: Run deb-install scenario
|
||||
if: matrix.type == 'deb-install'
|
||||
timeout-minutes: 25
|
||||
run: bash testing/deb-install/test.sh ${{ matrix.scenario }}
|
||||
|
||||
- name: Collect logs on failure (deb-install)
|
||||
if: matrix.type == 'deb-install' && failure()
|
||||
run: |
|
||||
docker ps -a --filter "name=fips-deb-test-${{ matrix.scenario }}" --format '{{.Names}}' | while read -r c; do
|
||||
echo "--- ${c} fips.service ---"
|
||||
docker exec "$c" journalctl -u fips.service --no-pager 2>&1 | tail -100 || true
|
||||
echo "--- ${c} fips-dns.service ---"
|
||||
docker exec "$c" journalctl -u fips-dns.service --no-pager 2>&1 | tail -100 || true
|
||||
echo "--- ${c} fips-gateway.service ---"
|
||||
docker exec "$c" journalctl -u fips-gateway.service --no-pager 2>&1 | tail -100 || true
|
||||
done
|
||||
|
||||
- name: Stop containers (deb-install)
|
||||
if: matrix.type == 'deb-install' && always()
|
||||
run: |
|
||||
docker ps -a --filter "name=fips-deb-test-${{ matrix.scenario }}" --format '{{.Names}}' | while read -r c; do
|
||||
docker rm -f "$c" >/dev/null 2>&1 || true
|
||||
done
|
||||
|
||||
# ── DNS resolver multi-backend integration ──────────────────────────
|
||||
# The dns-resolver harness builds its own fips binary from source in a
|
||||
# Debian 12 builder image (shared cache layout with deb-install). Runs
|
||||
# all 13 scenarios in a single job: dummy-TUN backend-detection tests
|
||||
# plus real-fips end-to-end queries through systemd-resolved across
|
||||
# five distros. ~7-12 min warm, ~12-15 min cold.
|
||||
- name: Run dns-resolver test
|
||||
if: matrix.type == 'dns-resolver'
|
||||
timeout-minutes: 30
|
||||
run: bash testing/dns-resolver/test.sh
|
||||
|
||||
- name: Collect logs on failure (dns-resolver)
|
||||
if: matrix.type == 'dns-resolver' && failure()
|
||||
run: |
|
||||
docker ps -a --filter "name=fips-dns-test-" --format '{{.Names}}' | while read -r c; do
|
||||
echo "--- ${c} fips.service ---"
|
||||
docker exec "$c" journalctl -u fips.service --no-pager 2>&1 | tail -100 || true
|
||||
echo "--- ${c} fips-dns.service ---"
|
||||
docker exec "$c" journalctl -u fips-dns.service --no-pager 2>&1 | tail -100 || true
|
||||
done
|
||||
|
||||
- name: Stop containers (dns-resolver)
|
||||
if: matrix.type == 'dns-resolver' && always()
|
||||
run: |
|
||||
docker ps -a --filter "name=fips-dns-test-" --format '{{.Names}}' | while read -r c; do
|
||||
docker rm -f "$c" >/dev/null 2>&1 || true
|
||||
done
|
||||
|
||||
20
.github/workflows/package-linux.yml
vendored
@@ -1,8 +1,13 @@
|
||||
name: Linux Package
|
||||
on:
|
||||
push:
|
||||
branches:
|
||||
- master
|
||||
- maint
|
||||
- next
|
||||
tags:
|
||||
- "v*"
|
||||
pull_request:
|
||||
workflow_dispatch:
|
||||
|
||||
env:
|
||||
@@ -14,7 +19,7 @@ jobs:
|
||||
outputs:
|
||||
linux_package_version: ${{ steps.linux_version.outputs.linux_package_version }}
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
- uses: actions/checkout@v6
|
||||
with:
|
||||
fetch-depth: 0
|
||||
|
||||
@@ -56,7 +61,7 @@ jobs:
|
||||
deb_arch: arm64
|
||||
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
- uses: actions/checkout@v6
|
||||
with:
|
||||
fetch-depth: 0
|
||||
|
||||
@@ -67,11 +72,14 @@ jobs:
|
||||
run: sudo apt-get update && sudo apt-get install -y --no-install-recommends libdbus-1-dev llvm
|
||||
|
||||
- name: Install Rust toolchain
|
||||
uses: dtolnay/rust-toolchain@stable
|
||||
uses: actions-rust-lang/setup-rust-toolchain@v1
|
||||
with:
|
||||
cache: false
|
||||
rustflags: ''
|
||||
|
||||
- name: Cache Cargo registry + build
|
||||
if: ${{ env.ACT != 'true' }}
|
||||
uses: actions/cache@v4
|
||||
uses: actions/cache@v5
|
||||
with:
|
||||
path: |
|
||||
~/.cargo/registry
|
||||
@@ -132,7 +140,7 @@ jobs:
|
||||
|
||||
- name: Upload artifact (GitHub only)
|
||||
if: ${{ env.ACT != 'true' }}
|
||||
uses: actions/upload-artifact@v4
|
||||
uses: actions/upload-artifact@v7
|
||||
with:
|
||||
name: fips_${{ needs.determine-versioning.outputs.linux_package_version }}_${{ matrix.artifact_arch }}_linux
|
||||
path: |
|
||||
@@ -156,7 +164,7 @@ jobs:
|
||||
|
||||
steps:
|
||||
- name: Download Linux artifacts
|
||||
uses: actions/download-artifact@v4
|
||||
uses: actions/download-artifact@v8
|
||||
with:
|
||||
path: dist
|
||||
merge-multiple: true
|
||||
|
||||
361
.github/workflows/package-macos.yml
vendored
Normal file
@@ -0,0 +1,361 @@
|
||||
name: macOS Package
|
||||
on:
|
||||
push:
|
||||
branches:
|
||||
- master
|
||||
- maint
|
||||
- next
|
||||
tags:
|
||||
- "v*"
|
||||
pull_request:
|
||||
workflow_dispatch:
|
||||
|
||||
env:
|
||||
CARGO_TERM_COLOR: always
|
||||
|
||||
jobs:
|
||||
determine-versioning:
|
||||
runs-on: macos-latest
|
||||
outputs:
|
||||
macos_package_version: ${{ steps.macos_version.outputs.macos_package_version }}
|
||||
steps:
|
||||
- uses: actions/checkout@v6
|
||||
with:
|
||||
fetch-depth: 0
|
||||
|
||||
- name: Derive macOS package version
|
||||
id: macos_version
|
||||
shell: bash
|
||||
run: |
|
||||
: ${GITHUB_OUTPUT:=/tmp/github_output}
|
||||
|
||||
BASE_VERSION=$(grep '^version' Cargo.toml | head -1 | sed 's/.*"\(.*\)"/\1/')
|
||||
if [[ "$GITHUB_REF" == refs/tags/* ]]; then
|
||||
VERSION="${GITHUB_REF_NAME#v}"
|
||||
else
|
||||
BRANCH=$(echo "$GITHUB_REF_NAME" | sed 's|[^A-Za-z0-9]|.|g; s/\.\.+/./g; s/^\.//; s/\.$//')
|
||||
HEIGHT=$(git rev-list --count HEAD)
|
||||
HASH=$(git rev-parse --short HEAD)
|
||||
if [[ -z "$BRANCH" ]]; then
|
||||
BRANCH="ref"
|
||||
fi
|
||||
VERSION="${BASE_VERSION}+${BRANCH}.${HEIGHT}.${HASH}"
|
||||
fi
|
||||
|
||||
echo "macos_package_version=${VERSION}" >> "$GITHUB_OUTPUT"
|
||||
|
||||
build:
|
||||
name: Build macOS package (${{ matrix.arch }})
|
||||
runs-on: ${{ matrix.os }}
|
||||
needs: determine-versioning
|
||||
|
||||
strategy:
|
||||
fail-fast: false
|
||||
matrix:
|
||||
include:
|
||||
- os: macos-latest
|
||||
arch: arm64
|
||||
target: aarch64-apple-darwin
|
||||
- os: macos-latest
|
||||
arch: x86_64
|
||||
target: x86_64-apple-darwin
|
||||
|
||||
steps:
|
||||
- uses: actions/checkout@v6
|
||||
with:
|
||||
fetch-depth: 0
|
||||
|
||||
- name: Set SOURCE_DATE_EPOCH from git
|
||||
run: echo "SOURCE_DATE_EPOCH=$(git log -1 --format=%ct)" >> "$GITHUB_ENV"
|
||||
|
||||
- name: Install Rust toolchain
|
||||
uses: actions-rust-lang/setup-rust-toolchain@v1
|
||||
with:
|
||||
target: ${{ matrix.target }}
|
||||
cache: false
|
||||
rustflags: ''
|
||||
|
||||
- name: Cache Cargo registry + build
|
||||
uses: actions/cache@v5
|
||||
with:
|
||||
path: |
|
||||
~/.cargo/registry
|
||||
~/.cargo/git
|
||||
target
|
||||
key: macos-release-${{ matrix.arch }}-${{ hashFiles('**/Cargo.lock') }}
|
||||
restore-keys: |
|
||||
macos-release-${{ matrix.arch }}-
|
||||
|
||||
- name: Build release binaries
|
||||
run: cargo build --release --target ${{ matrix.target }}
|
||||
|
||||
- name: Build macOS package
|
||||
run: |
|
||||
packaging/macos/build-pkg.sh \
|
||||
--version "${{ needs.determine-versioning.outputs.macos_package_version }}" \
|
||||
--target ${{ matrix.target }} \
|
||||
--no-build
|
||||
|
||||
- name: Resolve macOS asset path
|
||||
id: macos-assets
|
||||
shell: bash
|
||||
run: |
|
||||
: ${GITHUB_OUTPUT:=/tmp/github_output}
|
||||
set -euo pipefail
|
||||
|
||||
# build-pkg.sh names the package from the build target, so each
|
||||
# matrix leg produces a distinctly named, arch-correct asset.
|
||||
# Assert that here: a regression in that naming then fails loudly
|
||||
# at the build stage instead of as a silent collision when the
|
||||
# release job merges both artifacts into one directory.
|
||||
EXPECTED="deploy/fips-${{ needs.determine-versioning.outputs.macos_package_version }}-macos-${{ matrix.arch }}.pkg"
|
||||
if [[ ! -f "$EXPECTED" ]]; then
|
||||
echo "Expected package $EXPECTED was not produced" >&2
|
||||
echo "deploy/ contains:" >&2
|
||||
ls -la deploy >&2 || true
|
||||
exit 1
|
||||
fi
|
||||
|
||||
echo "pkg=$EXPECTED" >> "$GITHUB_OUTPUT"
|
||||
|
||||
- name: Verify .pkg structural correctness
|
||||
shell: bash
|
||||
run: |
|
||||
set -euo pipefail
|
||||
PKG="${{ steps.macos-assets.outputs.pkg }}"
|
||||
EXPAND_DIR="$(mktemp -d)/expanded"
|
||||
PAYLOAD_DIR="$(mktemp -d)/payload"
|
||||
fail=0
|
||||
|
||||
echo "==> Verifying $PKG"
|
||||
|
||||
# 1) Flat-package expansion
|
||||
if pkgutil --expand "$PKG" "$EXPAND_DIR"; then
|
||||
echo "PASS: pkgutil --expand"
|
||||
else
|
||||
echo "FAIL: pkgutil --expand"
|
||||
fail=1
|
||||
fi
|
||||
|
||||
# Extract the cpio.gz Payload so we can inspect installed file layout
|
||||
PAYLOAD_FILE="$(find "$EXPAND_DIR" -name Payload -type f | head -n 1)"
|
||||
if [[ -z "$PAYLOAD_FILE" ]]; then
|
||||
echo "FAIL: no Payload file inside expanded pkg"
|
||||
fail=1
|
||||
else
|
||||
mkdir -p "$PAYLOAD_DIR"
|
||||
(cd "$PAYLOAD_DIR" && gzip -dc "$PAYLOAD_FILE" | cpio -i --quiet)
|
||||
echo "PASS: extracted Payload to $PAYLOAD_DIR"
|
||||
fi
|
||||
|
||||
# 2) Binary at canonical install path (./usr/local/bin/fips inside payload)
|
||||
BIN_PATH="$PAYLOAD_DIR/usr/local/bin/fips"
|
||||
if [[ -f "$BIN_PATH" ]]; then
|
||||
echo "PASS: binary present at usr/local/bin/fips"
|
||||
else
|
||||
echo "FAIL: binary missing at usr/local/bin/fips"
|
||||
echo " fallback search:"
|
||||
find "$PAYLOAD_DIR" -name fips -type f -print || true
|
||||
fail=1
|
||||
fi
|
||||
for extra in fipsctl fipstop; do
|
||||
if [[ -f "$PAYLOAD_DIR/usr/local/bin/$extra" ]]; then
|
||||
echo "PASS: binary present at usr/local/bin/$extra"
|
||||
else
|
||||
echo "FAIL: binary missing at usr/local/bin/$extra"
|
||||
fail=1
|
||||
fi
|
||||
done
|
||||
|
||||
# 3) LaunchDaemon plist at canonical location
|
||||
PLIST_PATH="$PAYLOAD_DIR/Library/LaunchDaemons/com.fips.daemon.plist"
|
||||
if [[ -f "$PLIST_PATH" ]]; then
|
||||
echo "PASS: plist present at Library/LaunchDaemons/com.fips.daemon.plist"
|
||||
else
|
||||
echo "FAIL: plist missing at Library/LaunchDaemons/com.fips.daemon.plist"
|
||||
echo " fallback search:"
|
||||
find "$PAYLOAD_DIR" -name '*.plist' -print || true
|
||||
fail=1
|
||||
fi
|
||||
|
||||
# 4) plutil -lint on the plist
|
||||
if [[ -f "$PLIST_PATH" ]]; then
|
||||
if plutil -lint "$PLIST_PATH"; then
|
||||
echo "PASS: plutil -lint"
|
||||
else
|
||||
echo "FAIL: plutil -lint"
|
||||
fail=1
|
||||
fi
|
||||
fi
|
||||
|
||||
if [[ "$fail" -ne 0 ]]; then
|
||||
echo "==> .pkg verification FAILED"
|
||||
exit 1
|
||||
fi
|
||||
echo "==> .pkg verification PASSED"
|
||||
|
||||
- name: SHA-256 hash and sidecar
|
||||
shell: bash
|
||||
run: |
|
||||
set -euo pipefail
|
||||
PKG="${{ steps.macos-assets.outputs.pkg }}"
|
||||
echo "==> macOS release asset:"
|
||||
# Capture the SHA-256 of the verified .pkg on the macOS runner and
|
||||
# write it to a sidecar file next to the .pkg, in the standard
|
||||
# `<hash> <basename>` shasum format. The verify-handoff and release
|
||||
# jobs re-check the downloaded bytes against this value, so any
|
||||
# corruption introduced after this point is detected before
|
||||
# publication.
|
||||
( cd "$(dirname "$PKG")" && shasum -a 256 "$(basename "$PKG")" | tee "$(basename "$PKG").sha256" )
|
||||
|
||||
- name: Upload artifact
|
||||
uses: actions/upload-artifact@v7
|
||||
with:
|
||||
name: fips_${{ needs.determine-versioning.outputs.macos_package_version }}_${{ matrix.arch }}_macos
|
||||
path: |
|
||||
${{ steps.macos-assets.outputs.pkg }}
|
||||
${{ steps.macos-assets.outputs.pkg }}.sha256
|
||||
retention-days: 30
|
||||
|
||||
- name: Build summary
|
||||
run: |
|
||||
echo "Build Summary for macOS/${{ matrix.arch }}:"
|
||||
echo " Package: ${{ steps.macos-assets.outputs.pkg }}"
|
||||
|
||||
verify-handoff:
|
||||
name: Verify macOS package handoff integrity
|
||||
runs-on: ubuntu-latest
|
||||
needs: build
|
||||
|
||||
steps:
|
||||
- name: Download macOS artifacts
|
||||
uses: actions/download-artifact@v8
|
||||
with:
|
||||
path: dist
|
||||
merge-multiple: true
|
||||
|
||||
- name: Verify .pkg integrity across the handoff
|
||||
shell: bash
|
||||
run: |
|
||||
set -euo pipefail
|
||||
cd dist
|
||||
|
||||
pkgs=$(find . -maxdepth 1 -type f -name '*.pkg' | LC_ALL=C sort)
|
||||
if [[ -z "$pkgs" ]]; then
|
||||
echo "FAIL: no .pkg artifacts were downloaded" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
fail=0
|
||||
while IFS= read -r pkg; do
|
||||
base=$(basename "$pkg")
|
||||
sidecar="${pkg}.sha256"
|
||||
if [[ ! -f "$sidecar" ]]; then
|
||||
echo "FAIL: missing SHA-256 sidecar for $base" >&2
|
||||
fail=1
|
||||
continue
|
||||
fi
|
||||
expected=$(awk '{print $1}' "$sidecar")
|
||||
actual=$(sha256sum "$pkg" | awk '{print $1}')
|
||||
if [[ "$expected" != "$actual" ]]; then
|
||||
echo "FAIL: $base SHA-256 mismatch across the artifact handoff" >&2
|
||||
echo " expected (macOS runner): $expected" >&2
|
||||
echo " actual (downloaded): $actual" >&2
|
||||
fail=1
|
||||
continue
|
||||
fi
|
||||
echo "PASS: $base matches the macOS-runner SHA-256 ($actual)"
|
||||
done <<<"$pkgs"
|
||||
|
||||
if [[ "$fail" -ne 0 ]]; then
|
||||
echo "==> macOS package handoff verification FAILED" >&2
|
||||
exit 1
|
||||
fi
|
||||
echo "==> macOS package handoff verification PASSED"
|
||||
|
||||
release:
|
||||
name: Publish macOS assets to GitHub Release
|
||||
runs-on: ubuntu-latest
|
||||
needs: [build, verify-handoff]
|
||||
if: startsWith(github.ref, 'refs/tags/')
|
||||
permissions:
|
||||
contents: write
|
||||
|
||||
steps:
|
||||
- name: Download macOS artifacts
|
||||
uses: actions/download-artifact@v8
|
||||
with:
|
||||
path: dist
|
||||
merge-multiple: true
|
||||
|
||||
- name: Validate .pkg bytes before publishing
|
||||
shell: bash
|
||||
run: |
|
||||
set -euo pipefail
|
||||
cd dist
|
||||
|
||||
pkgs=$(find . -maxdepth 1 -type f -name '*.pkg' | LC_ALL=C sort)
|
||||
if [[ -z "$pkgs" ]]; then
|
||||
echo "FAIL: no .pkg artifacts were downloaded" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
fail=0
|
||||
while IFS= read -r pkg; do
|
||||
base=$(basename "$pkg")
|
||||
sidecar="${pkg}.sha256"
|
||||
if [[ ! -f "$sidecar" ]]; then
|
||||
echo "FAIL: missing SHA-256 sidecar for $base" >&2
|
||||
fail=1
|
||||
continue
|
||||
fi
|
||||
expected=$(awk '{print $1}' "$sidecar")
|
||||
actual=$(sha256sum "$pkg" | awk '{print $1}')
|
||||
if [[ "$expected" != "$actual" ]]; then
|
||||
echo "FAIL: $base SHA-256 mismatch on the bytes about to be published" >&2
|
||||
echo " expected (macOS runner): $expected" >&2
|
||||
echo " actual (downloaded): $actual" >&2
|
||||
fail=1
|
||||
continue
|
||||
fi
|
||||
echo "PASS: $base matches the macOS-runner SHA-256 ($actual)"
|
||||
done <<<"$pkgs"
|
||||
|
||||
if [[ "$fail" -ne 0 ]]; then
|
||||
echo "==> pre-publish .pkg verification FAILED; not publishing" >&2
|
||||
exit 1
|
||||
fi
|
||||
echo "==> pre-publish .pkg verification PASSED"
|
||||
|
||||
- name: Generate macOS release checksums
|
||||
run: |
|
||||
cd dist
|
||||
find . -maxdepth 1 -type f -name '*.pkg' -printf '%P\n' \
|
||||
| LC_ALL=C sort \
|
||||
| xargs sha256sum \
|
||||
> checksums-macos.txt
|
||||
|
||||
- name: Wait for tag release
|
||||
env:
|
||||
GH_TOKEN: ${{ github.token }}
|
||||
run: |
|
||||
for attempt in $(seq 1 20); do
|
||||
if gh release view "${GITHUB_REF_NAME}" --repo "${GITHUB_REPOSITORY}" >/dev/null 2>&1; then
|
||||
exit 0
|
||||
fi
|
||||
echo "Release ${GITHUB_REF_NAME} not available yet; waiting..."
|
||||
sleep 15
|
||||
done
|
||||
|
||||
echo "Timed out waiting for release ${GITHUB_REF_NAME}" >&2
|
||||
exit 1
|
||||
|
||||
- name: Upload macOS assets
|
||||
env:
|
||||
GH_TOKEN: ${{ github.token }}
|
||||
run: |
|
||||
gh release upload "${GITHUB_REF_NAME}" \
|
||||
dist/*.pkg \
|
||||
dist/checksums-macos.txt \
|
||||
--clobber \
|
||||
--repo "${GITHUB_REPOSITORY}"
|
||||
667
.github/workflows/package-openwrt.yml
vendored
@@ -1,6 +1,13 @@
|
||||
name: OpenWrt Package
|
||||
on:
|
||||
push:
|
||||
branches:
|
||||
- master
|
||||
- maint
|
||||
- next
|
||||
tags:
|
||||
- "v*"
|
||||
pull_request:
|
||||
workflow_dispatch:
|
||||
inputs:
|
||||
arch:
|
||||
@@ -17,9 +24,10 @@ jobs:
|
||||
runs-on: ubuntu-latest
|
||||
outputs:
|
||||
package_version: ${{ steps.version.outputs.package_version }}
|
||||
apk_version: ${{ steps.version.outputs.apk_version }}
|
||||
release_channel: ${{ steps.channel.outputs.release_channel }}
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
- uses: actions/checkout@v6
|
||||
with:
|
||||
fetch-depth: 0
|
||||
|
||||
@@ -28,13 +36,20 @@ jobs:
|
||||
shell: bash
|
||||
run: |
|
||||
: ${GITHUB_OUTPUT:=/tmp/github_output}
|
||||
# package_version is the human-readable label used in artifact
|
||||
# filenames; apk_version is the apk-tools-compatible string embedded
|
||||
# in the .apk metadata. apk_version is built directly from the same
|
||||
# structured inputs (tag, or commit height) — no reparse of the
|
||||
# flattened package_version. See packaging/openwrt-apk/apk-version.sh.
|
||||
if [[ "$GITHUB_REF" == refs/tags/* ]]; then
|
||||
echo "package_version=${GITHUB_REF_NAME}" >> "$GITHUB_OUTPUT"
|
||||
echo "apk_version=$(sh packaging/openwrt-apk/apk-version.sh tag "${GITHUB_REF_NAME}")" >> "$GITHUB_OUTPUT"
|
||||
else
|
||||
BRANCH=$(echo "$GITHUB_REF_NAME" | sed 's|/|-|g')
|
||||
HEIGHT=$(git rev-list --count HEAD)
|
||||
HASH=$(git rev-parse --short HEAD)
|
||||
echo "package_version=${BRANCH}.${HEIGHT}.${HASH}" >> "$GITHUB_OUTPUT"
|
||||
echo "apk_version=$(sh packaging/openwrt-apk/apk-version.sh dev "${HEIGHT}")" >> "$GITHUB_OUTPUT"
|
||||
fi
|
||||
|
||||
- name: Determine release channel
|
||||
@@ -57,8 +72,8 @@ jobs:
|
||||
echo "release_channel=dev" >> "$GITHUB_OUTPUT"
|
||||
fi
|
||||
|
||||
build:
|
||||
name: Build .ipk (${{ matrix.openwrt_arch }})
|
||||
compile-binaries:
|
||||
name: Cross-compile (${{ matrix.openwrt_arch }})
|
||||
runs-on: ubuntu-latest
|
||||
needs: determine-versioning
|
||||
|
||||
@@ -71,7 +86,9 @@ jobs:
|
||||
rust_target: aarch64-unknown-linux-musl
|
||||
rust_channel: stable
|
||||
# MT3000, MT6000, Flint 2, RPi 3/4/5
|
||||
# MIPS disabled: 32-bit MIPS lacks AtomicU64; needs portable-atomic crate
|
||||
# MIPS disabled: nostr-relay-pool 0.44 uses std::sync::atomic::AtomicU64
|
||||
# directly (fips's own atomics already use portable_atomic). Re-enable
|
||||
# once an upstream portable-atomic patch lands (or via [patch.crates-io]).
|
||||
# - build_arch: mipsel
|
||||
# openwrt_arch: mipsel_24kc
|
||||
# rust_target: mipsel-unknown-linux-musl
|
||||
@@ -87,23 +104,17 @@ jobs:
|
||||
# x86 routers / VMs
|
||||
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
- uses: actions/checkout@v6
|
||||
with:
|
||||
fetch-depth: 0
|
||||
|
||||
- name: Set SOURCE_DATE_EPOCH from git
|
||||
run: echo "SOURCE_DATE_EPOCH=$(git log -1 --format=%ct)" >> "$GITHUB_ENV"
|
||||
|
||||
- name: Initialize
|
||||
run: |
|
||||
PACKAGE_FILENAME=${{ env.PACKAGE_NAME }}_${{ needs.determine-versioning.outputs.package_version }}_${{ matrix.openwrt_arch }}.ipk
|
||||
echo "PACKAGE_FILENAME=$PACKAGE_FILENAME" >> $GITHUB_ENV
|
||||
|
||||
- name: Install Rust toolchain (stable)
|
||||
if: matrix.rust_channel == 'stable'
|
||||
uses: dtolnay/rust-toolchain@stable
|
||||
uses: actions-rust-lang/setup-rust-toolchain@v1
|
||||
with:
|
||||
targets: ${{ matrix.rust_target }}
|
||||
target: ${{ matrix.rust_target }}
|
||||
cache: false
|
||||
rustflags: ''
|
||||
|
||||
- name: Install Rust toolchain (nightly, Tier 3)
|
||||
if: matrix.rust_channel == 'nightly'
|
||||
@@ -113,7 +124,7 @@ jobs:
|
||||
|
||||
- name: Cache Cargo registry + build
|
||||
if: ${{ env.ACT != 'true' }}
|
||||
uses: actions/cache@v4
|
||||
uses: actions/cache@v5
|
||||
with:
|
||||
path: |
|
||||
~/.cargo/registry
|
||||
@@ -142,6 +153,64 @@ jobs:
|
||||
- name: Install llvm-strip
|
||||
run: sudo apt-get update && sudo apt-get install -y --no-install-recommends llvm
|
||||
|
||||
# Cross-compile + strip once; both the .ipk and .apk packagers consume
|
||||
# these artifacts via --bin-dir, so the Rust build runs a single time
|
||||
# per architecture instead of once per package format.
|
||||
- name: Cross-compile and strip binaries
|
||||
run: |
|
||||
set -euo pipefail
|
||||
cargo zigbuild --release --target ${{ matrix.rust_target }} \
|
||||
--bin fips --bin fipsctl --bin fipstop --bin fips-gateway
|
||||
RELEASE_DIR="target/${{ matrix.rust_target }}/release"
|
||||
mkdir -p out
|
||||
for b in fips fipsctl fipstop fips-gateway; do
|
||||
llvm-strip "$RELEASE_DIR/$b" 2>/dev/null || true
|
||||
cp "$RELEASE_DIR/$b" "out/$b"
|
||||
done
|
||||
ls -lh out/
|
||||
|
||||
- name: Upload binaries artifact
|
||||
uses: actions/upload-artifact@v7
|
||||
with:
|
||||
name: fips-bins-${{ matrix.openwrt_arch }}
|
||||
path: out/
|
||||
retention-days: 1
|
||||
|
||||
build:
|
||||
name: Build .ipk (${{ matrix.openwrt_arch }})
|
||||
runs-on: ubuntu-latest
|
||||
needs: [determine-versioning, compile-binaries]
|
||||
|
||||
strategy:
|
||||
fail-fast: false
|
||||
matrix:
|
||||
# Must be a subset of compile-binaries' arches (this job consumes those
|
||||
# binary artifacts). Currently both ship aarch64 + x86_64.
|
||||
include:
|
||||
- build_arch: aarch64
|
||||
openwrt_arch: aarch64_cortex-a53
|
||||
- build_arch: x86_64
|
||||
openwrt_arch: x86_64
|
||||
|
||||
steps:
|
||||
- uses: actions/checkout@v6
|
||||
with:
|
||||
fetch-depth: 0
|
||||
|
||||
- name: Set SOURCE_DATE_EPOCH from git
|
||||
run: echo "SOURCE_DATE_EPOCH=$(git log -1 --format=%ct)" >> "$GITHUB_ENV"
|
||||
|
||||
- name: Initialize
|
||||
run: |
|
||||
PACKAGE_FILENAME=${{ env.PACKAGE_NAME }}_${{ needs.determine-versioning.outputs.package_version }}_${{ matrix.openwrt_arch }}.ipk
|
||||
echo "PACKAGE_FILENAME=$PACKAGE_FILENAME" >> $GITHUB_ENV
|
||||
|
||||
- name: Download prebuilt binaries
|
||||
uses: actions/download-artifact@v8
|
||||
with:
|
||||
name: fips-bins-${{ matrix.openwrt_arch }}
|
||||
path: bins
|
||||
|
||||
- name: Install nak
|
||||
shell: bash
|
||||
run: |
|
||||
@@ -190,19 +259,202 @@ jobs:
|
||||
- name: Build .ipk
|
||||
env:
|
||||
PKG_VERSION: ${{ needs.determine-versioning.outputs.package_version }}
|
||||
LLVM_STRIP: llvm-strip
|
||||
run: ./packaging/openwrt-ipk/build-ipk.sh --arch ${{ matrix.build_arch }}
|
||||
run: ./packaging/openwrt-ipk/build-ipk.sh --arch ${{ matrix.build_arch }} --bin-dir "$GITHUB_WORKSPACE/bins"
|
||||
|
||||
- name: Install shellcheck (if missing)
|
||||
shell: bash
|
||||
run: |
|
||||
if ! command -v shellcheck >/dev/null 2>&1; then
|
||||
sudo apt-get update && sudo apt-get install -y --no-install-recommends shellcheck
|
||||
fi
|
||||
shellcheck --version
|
||||
|
||||
- name: Lint shipped shell scripts
|
||||
shell: bash
|
||||
run: |
|
||||
set -euo pipefail
|
||||
FILES_DIR=packaging/openwrt-ipk/files
|
||||
# Scripts shipped inside the .ipk. The init scripts use the OpenWrt
|
||||
# `#!/bin/sh /etc/rc.common` shebang; tell shellcheck to treat them
|
||||
# as POSIX sh and silence the unrecognized-shebang warning (SC1008).
|
||||
# SC2317 (unreachable command) fires on rc.common's externally-invoked
|
||||
# start_service/stop_service/reload_service hooks.
|
||||
TARGETS=(
|
||||
"$FILES_DIR/etc/init.d/fips"
|
||||
"$FILES_DIR/etc/init.d/fips-gateway"
|
||||
"$FILES_DIR/etc/fips/firewall.sh"
|
||||
"$FILES_DIR/etc/hotplug.d/net/99-fips"
|
||||
"$FILES_DIR/etc/uci-defaults/90-fips-setup"
|
||||
"$FILES_DIR/usr/bin/fips-mesh-setup"
|
||||
)
|
||||
fail=0
|
||||
for f in "${TARGETS[@]}"; do
|
||||
if [ ! -f "$f" ]; then
|
||||
echo "FAIL: missing $f"
|
||||
fail=1
|
||||
continue
|
||||
fi
|
||||
echo "==> shellcheck $f"
|
||||
if shellcheck --shell=sh --exclude=SC1008,SC2317,SC2034,SC3043,SC2086,SC2089,SC2090 "$f"; then
|
||||
echo " PASS"
|
||||
else
|
||||
echo " FAIL"
|
||||
fail=1
|
||||
fi
|
||||
done
|
||||
if [ "$fail" -ne 0 ]; then
|
||||
echo "shellcheck FAILED"
|
||||
exit 1
|
||||
fi
|
||||
echo "shellcheck PASS (${#TARGETS[@]} scripts)"
|
||||
|
||||
- name: Sysctl drop-in syntax check
|
||||
shell: bash
|
||||
run: |
|
||||
set -euo pipefail
|
||||
FILES_DIR=packaging/openwrt-ipk/files
|
||||
TARGETS=(
|
||||
"$FILES_DIR/etc/sysctl.d/fips-gateway.conf"
|
||||
"$FILES_DIR/etc/sysctl.d/fips-bridge.conf"
|
||||
)
|
||||
fail=0
|
||||
for conf in "${TARGETS[@]}"; do
|
||||
if [ ! -f "$conf" ]; then
|
||||
echo "FAIL: missing $conf"
|
||||
fail=1
|
||||
continue
|
||||
fi
|
||||
echo "==> validating $conf"
|
||||
lineno=0
|
||||
file_ok=1
|
||||
while IFS= read -r line || [ -n "$line" ]; do
|
||||
lineno=$((lineno + 1))
|
||||
# Skip comments and blank lines.
|
||||
case "$line" in
|
||||
''|\#*) continue ;;
|
||||
esac
|
||||
# Match: <key> = <value> where key is dotted lower-id and value is
|
||||
# an integer (sysctl drop-ins shipped here are all numeric toggles).
|
||||
if ! [[ "$line" =~ ^[a-z0-9_.-]+[[:space:]]*=[[:space:]]*-?[0-9]+[[:space:]]*$ ]]; then
|
||||
echo " FAIL line $lineno: $line"
|
||||
file_ok=0
|
||||
fi
|
||||
done < "$conf"
|
||||
if [ "$file_ok" -eq 1 ]; then
|
||||
echo " PASS"
|
||||
else
|
||||
fail=1
|
||||
fi
|
||||
done
|
||||
if [ "$fail" -ne 0 ]; then
|
||||
echo "sysctl drop-in syntax check FAILED"
|
||||
exit 1
|
||||
fi
|
||||
echo "sysctl drop-in syntax check PASS"
|
||||
|
||||
- name: Verify ipk structural integrity
|
||||
shell: bash
|
||||
run: |
|
||||
set -euo pipefail
|
||||
IPK="dist/${{ env.PACKAGE_FILENAME }}"
|
||||
if [ ! -f "$IPK" ]; then
|
||||
echo "FAIL: produced ipk not found at $IPK"
|
||||
exit 1
|
||||
fi
|
||||
echo "==> file type:"
|
||||
file "$IPK"
|
||||
|
||||
# OpenWrt .ipk = tar.gz containing debian-binary + control.tar.gz +
|
||||
# data.tar.gz (NOT an ar archive like Debian's .deb).
|
||||
WORK=$(mktemp -d)
|
||||
trap 'rm -rf "$WORK"' EXIT
|
||||
|
||||
tar -xzf "$IPK" -C "$WORK"
|
||||
echo "==> top-level entries:"
|
||||
ls -la "$WORK"
|
||||
|
||||
# Top-level structural assertions.
|
||||
fail=0
|
||||
for entry in debian-binary control.tar.gz data.tar.gz; do
|
||||
if [ ! -f "$WORK/$entry" ]; then
|
||||
echo "FAIL: missing top-level $entry"
|
||||
fail=1
|
||||
else
|
||||
echo " PASS top-level: $entry"
|
||||
fi
|
||||
done
|
||||
if [ "$fail" -ne 0 ]; then exit 1; fi
|
||||
|
||||
# debian-binary content sanity.
|
||||
dbin_content=$(cat "$WORK/debian-binary" | tr -d '[:space:]')
|
||||
if [ "$dbin_content" != "2.0" ]; then
|
||||
echo "FAIL: debian-binary content is '$dbin_content' (expected 2.0)"
|
||||
exit 1
|
||||
fi
|
||||
echo " PASS debian-binary content: 2.0"
|
||||
|
||||
# Inspect data.tar.gz contents.
|
||||
DATA_LIST="$WORK/data.list"
|
||||
tar -tzf "$WORK/data.tar.gz" > "$DATA_LIST"
|
||||
echo "==> data.tar.gz entry count: $(wc -l < "$DATA_LIST")"
|
||||
|
||||
# Required filesystem entries inside data.tar.gz. Entries are
|
||||
# produced with a leading "./" by build-ipk.sh.
|
||||
REQUIRED=(
|
||||
./usr/bin/fips
|
||||
./usr/bin/fipsctl
|
||||
./usr/bin/fipstop
|
||||
./usr/bin/fips-gateway
|
||||
./usr/bin/fips-mesh-setup
|
||||
./etc/init.d/fips
|
||||
./etc/init.d/fips-gateway
|
||||
./etc/fips/fips.yaml
|
||||
./etc/fips/firewall.sh
|
||||
./etc/dnsmasq.d/fips.conf
|
||||
./etc/sysctl.d/fips-gateway.conf
|
||||
./etc/sysctl.d/fips-bridge.conf
|
||||
./etc/hotplug.d/net/99-fips
|
||||
./etc/uci-defaults/90-fips-setup
|
||||
./lib/upgrade/keep.d/fips
|
||||
)
|
||||
for path in "${REQUIRED[@]}"; do
|
||||
if grep -Fxq "$path" "$DATA_LIST"; then
|
||||
echo " PASS data: $path"
|
||||
else
|
||||
echo " FAIL data: missing $path"
|
||||
fail=1
|
||||
fi
|
||||
done
|
||||
|
||||
# Inspect control.tar.gz: must contain control file + maintainer scripts.
|
||||
CTRL_LIST="$WORK/control.list"
|
||||
tar -tzf "$WORK/control.tar.gz" > "$CTRL_LIST"
|
||||
echo "==> control.tar.gz entry count: $(wc -l < "$CTRL_LIST")"
|
||||
for path in ./control ./conffiles ./postinst ./prerm; do
|
||||
if grep -Fxq "$path" "$CTRL_LIST"; then
|
||||
echo " PASS control: $path"
|
||||
else
|
||||
echo " FAIL control: missing $path"
|
||||
fail=1
|
||||
fi
|
||||
done
|
||||
|
||||
if [ "$fail" -ne 0 ]; then
|
||||
echo "ipk structural verification FAILED"
|
||||
exit 1
|
||||
fi
|
||||
echo "ipk structural verification PASS"
|
||||
|
||||
- name: SHA-256 hashes
|
||||
run: |
|
||||
echo "==> Binaries:"
|
||||
sha256sum target/${{ matrix.rust_target }}/release/fips target/${{ matrix.rust_target }}/release/fipsctl target/${{ matrix.rust_target }}/release/fipstop
|
||||
sha256sum bins/fips bins/fipsctl bins/fipstop
|
||||
echo "==> Package:"
|
||||
sha256sum dist/${{ env.PACKAGE_FILENAME }}
|
||||
|
||||
- name: Upload artifact (GitHub only)
|
||||
if: ${{ env.ACT != 'true' }}
|
||||
uses: actions/upload-artifact@v4
|
||||
uses: actions/upload-artifact@v7
|
||||
with:
|
||||
name: ${{ env.PACKAGE_FILENAME }}
|
||||
path: dist/${{ env.PACKAGE_FILENAME }}
|
||||
@@ -210,6 +462,7 @@ jobs:
|
||||
|
||||
- name: Upload to Blossom
|
||||
id: blossom_upload
|
||||
continue-on-error: true
|
||||
shell: bash
|
||||
env:
|
||||
BLOSSOM_SERVER: "https://blossom.primal.net"
|
||||
@@ -217,17 +470,30 @@ jobs:
|
||||
run: |
|
||||
: ${GITHUB_OUTPUT:=/tmp/github_output}
|
||||
|
||||
UPLOAD_RESPONSE=$(nak blossom upload \
|
||||
--server "$BLOSSOM_SERVER" \
|
||||
--sec "$NSEC" \
|
||||
"dist/${{ env.PACKAGE_FILENAME }}" < /dev/null)
|
||||
FILE_HASH=""
|
||||
for attempt in 1 2 3; do
|
||||
if UPLOAD_RESPONSE=$(nak blossom upload \
|
||||
--server "$BLOSSOM_SERVER" \
|
||||
--sec "$NSEC" \
|
||||
"dist/${{ env.PACKAGE_FILENAME }}" < /dev/null); then
|
||||
echo "Upload response (attempt $attempt):"
|
||||
echo "$UPLOAD_RESPONSE"
|
||||
FILE_HASH=$(echo "$UPLOAD_RESPONSE" | jq -r '.sha256')
|
||||
if [ -n "$FILE_HASH" ] && [ "$FILE_HASH" != "null" ]; then
|
||||
break
|
||||
fi
|
||||
echo "Upload response had no sha256 (attempt $attempt)"
|
||||
else
|
||||
echo "Blossom upload timed out or failed (attempt $attempt)"
|
||||
fi
|
||||
FILE_HASH=""
|
||||
[ "$attempt" -lt 3 ] && sleep $((attempt * 10))
|
||||
done
|
||||
|
||||
echo "Upload response:"
|
||||
echo "$UPLOAD_RESPONSE"
|
||||
|
||||
FILE_HASH=$(echo "$UPLOAD_RESPONSE" | jq -r '.sha256')
|
||||
if [ -z "$FILE_HASH" ] || [ "$FILE_HASH" = "null" ]; then
|
||||
echo "Failed to extract hash from upload response"
|
||||
echo "Blossom upload did not succeed after 3 attempts; non-fatal."
|
||||
echo "The package still ships as a GitHub release artifact; only the"
|
||||
echo "supplementary Blossom/nostr distribution is skipped this run."
|
||||
exit 1
|
||||
fi
|
||||
|
||||
@@ -238,9 +504,10 @@ jobs:
|
||||
|
||||
- name: Publish NIP-94 release event
|
||||
id: publish
|
||||
if: steps.blossom_upload.outcome == 'success'
|
||||
shell: bash
|
||||
env:
|
||||
RELAYS: "wss://relay.damus.io wss://nos.lol wss://nostr.mom wss://relay.primal.net"
|
||||
RELAYS: "wss://relay.damus.io wss://nos.lol wss://nostr.mom wss://offchain.pub"
|
||||
NSEC: ${{ steps.keys.outputs.nsec }}
|
||||
run: |
|
||||
: ${GITHUB_OUTPUT:=/tmp/github_output}
|
||||
@@ -259,6 +526,7 @@ jobs:
|
||||
--tag A="${{ matrix.openwrt_arch }}" \
|
||||
--tag v="$VERSION" \
|
||||
--tag n="${{ env.PACKAGE_NAME }}" \
|
||||
--tag format="ipk" \
|
||||
--tag compression="none" \
|
||||
> event.json 2> event.err
|
||||
|
||||
@@ -288,7 +556,7 @@ jobs:
|
||||
if: ${{ steps.publish.outputs.eventId != '' }}
|
||||
env:
|
||||
EVENT_ID: ${{ steps.publish.outputs.eventId }}
|
||||
RELAYS: "wss://relay.damus.io wss://nos.lol wss://nostr.mom wss://relay.primal.net"
|
||||
RELAYS: "wss://relay.damus.io wss://nos.lol wss://nostr.mom wss://offchain.pub"
|
||||
run: |
|
||||
echo "Verifying event $EVENT_ID on relays..."
|
||||
FOUND=0
|
||||
@@ -316,23 +584,352 @@ jobs:
|
||||
echo " Release EventId: ${{ steps.publish.outputs.eventId }}"
|
||||
echo " Blossom URL: ${{ steps.blossom_upload.outputs.url }}"
|
||||
|
||||
build-apk:
|
||||
name: Build .apk (${{ matrix.openwrt_arch }})
|
||||
runs-on: ubuntu-latest
|
||||
needs: [determine-versioning, compile-binaries]
|
||||
|
||||
strategy:
|
||||
fail-fast: false
|
||||
matrix:
|
||||
# Must be a subset of compile-binaries' arches.
|
||||
include:
|
||||
- build_arch: aarch64
|
||||
openwrt_arch: aarch64_cortex-a53
|
||||
# MT3000, MT6000, Flint 2, RPi 3/4/5 on OpenWrt 25+
|
||||
- build_arch: x86_64
|
||||
openwrt_arch: x86_64
|
||||
# x86 routers / VMs on OpenWrt 25+
|
||||
|
||||
env:
|
||||
# apk-tools commit OpenWrt pins for the .apk (ADB) format. Keep in sync
|
||||
# with package/system/apk/Makefile in the targeted OpenWrt release so the
|
||||
# packages we produce are readable by the apk on the device.
|
||||
APK_TOOLS_VERSION: "3.0.5"
|
||||
APK_TOOLS_COMMIT: "b5a31c0d865342ad80be10d68f1bb3d3ad9b0866"
|
||||
|
||||
steps:
|
||||
- uses: actions/checkout@v6
|
||||
with:
|
||||
fetch-depth: 0
|
||||
|
||||
- name: Set SOURCE_DATE_EPOCH from git
|
||||
run: echo "SOURCE_DATE_EPOCH=$(git log -1 --format=%ct)" >> "$GITHUB_ENV"
|
||||
|
||||
- name: Initialize
|
||||
run: |
|
||||
PACKAGE_FILENAME=${{ env.PACKAGE_NAME }}_${{ needs.determine-versioning.outputs.package_version }}_${{ matrix.openwrt_arch }}.apk
|
||||
echo "PACKAGE_FILENAME=$PACKAGE_FILENAME" >> $GITHUB_ENV
|
||||
|
||||
- name: Download prebuilt binaries
|
||||
uses: actions/download-artifact@v8
|
||||
with:
|
||||
name: fips-bins-${{ matrix.openwrt_arch }}
|
||||
path: bins
|
||||
|
||||
- name: Install fakeroot
|
||||
run: sudo apt-get update && sudo apt-get install -y --no-install-recommends fakeroot
|
||||
|
||||
# apk mkpkg lives in apk-tools v3, which is not packaged for Ubuntu, so we
|
||||
# build the pinned release from source. This is the SDK-free equivalent of
|
||||
# how the .ipk path uses plain tar — one small C tool, no OpenWrt SDK.
|
||||
- name: Build apk-tools (${{ env.APK_TOOLS_VERSION }}) from source
|
||||
run: |
|
||||
set -euo pipefail
|
||||
sudo apt-get install -y --no-install-recommends \
|
||||
git ca-certificates build-essential meson ninja-build pkg-config \
|
||||
zlib1g-dev libssl-dev libzstd-dev liblzma-dev lua5.4-dev scdoc
|
||||
git clone --quiet https://gitlab.alpinelinux.org/alpine/apk-tools.git /tmp/apk-tools
|
||||
cd /tmp/apk-tools
|
||||
git checkout --quiet "${APK_TOOLS_COMMIT}"
|
||||
meson setup build
|
||||
ninja -C build src/apk
|
||||
APK_BIN=/tmp/apk-tools/build/src/apk
|
||||
"$APK_BIN" --version 2>/dev/null || "$APK_BIN" version 2>/dev/null || true
|
||||
echo "APK_BIN=$APK_BIN" >> "$GITHUB_ENV"
|
||||
|
||||
- name: Build .apk
|
||||
env:
|
||||
PKG_VERSION: ${{ needs.determine-versioning.outputs.package_version }}
|
||||
APK_VERSION: ${{ needs.determine-versioning.outputs.apk_version }}
|
||||
run: ./packaging/openwrt-apk/build-apk.sh --arch ${{ matrix.build_arch }} --bin-dir "$GITHUB_WORKSPACE/bins"
|
||||
|
||||
- name: Verify apk structural integrity
|
||||
shell: bash
|
||||
run: |
|
||||
set -euo pipefail
|
||||
APK="dist/${{ env.PACKAGE_FILENAME }}"
|
||||
if [ ! -s "$APK" ]; then
|
||||
echo "FAIL: produced apk not found or empty at $APK"
|
||||
exit 1
|
||||
fi
|
||||
echo "==> file type:"; file "$APK"
|
||||
|
||||
# apk v3 packages are ADB containers; dump the whole manifest with the
|
||||
# apk-tools we just built. Print it in full so the exact schema is
|
||||
# always visible in the log if an assertion needs adjusting.
|
||||
DUMP=$(mktemp)
|
||||
"$APK_BIN" adbdump "$APK" > "$DUMP" 2>/dev/null || {
|
||||
echo "FAIL: 'apk adbdump' could not read $APK"; exit 1; }
|
||||
echo "==> full adbdump:"; cat "$DUMP"
|
||||
|
||||
fail=0
|
||||
# Package metadata (flat keys under info:).
|
||||
for needle in "name: fips" "version: ${{ needs.determine-versioning.outputs.apk_version }}" "arch: ${{ matrix.openwrt_arch }}"; do
|
||||
if grep -qF "$needle" "$DUMP"; then
|
||||
echo " PASS meta: $needle"
|
||||
else
|
||||
echo " FAIL meta: missing '$needle'"; fail=1
|
||||
fi
|
||||
done
|
||||
|
||||
# installed-size reflects the bundled binaries (4 stripped Rust
|
||||
# binaries, several MB). A payload regression that drops them shows up
|
||||
# here regardless of how the path tree is formatted.
|
||||
SIZE=$(awk '/^[[:space:]]*installed-size:/ {print $2; exit}' "$DUMP")
|
||||
echo " installed-size: ${SIZE:-unknown}"
|
||||
if [ -z "${SIZE:-}" ] || [ "$SIZE" -lt 1000000 ]; then
|
||||
echo " FAIL: installed-size implausibly small (binaries missing?)"; fail=1
|
||||
else
|
||||
echo " PASS: installed-size >= 1MB"
|
||||
fi
|
||||
|
||||
# The adbdump paths: block is hierarchical. Each directory is a
|
||||
# top-level list item "- name: <full relative dir>"; its files are
|
||||
# "- name: <basename>" nested one indent level deeper under "files:".
|
||||
# (There are no "path:" keys.) Reconstruct full file paths by keying
|
||||
# off the indentation of the directory-level list items.
|
||||
RECON=$(awk '
|
||||
/^paths:/ {p=1; diri=-1; next}
|
||||
p && /^[^ #-]/ {p=0} # a new top-level key ends paths:
|
||||
!p {next}
|
||||
match($0, /^ *- /) {
|
||||
ind=RLENGTH; rest=substr($0, RLENGTH+1)
|
||||
if (diri==-1) diri=ind # first list item = directory indent
|
||||
if (ind==diri) { # directory entry (or the root acl: entry)
|
||||
if (rest ~ /^name: /) { dir=rest; sub(/^name: /,"",dir) } else dir=""
|
||||
next
|
||||
}
|
||||
if (rest ~ /^name: /) { # deeper item = a file under files:
|
||||
f=rest; sub(/^name: /,"",f); print (dir==""?f:dir"/"f)
|
||||
}
|
||||
}
|
||||
' "$DUMP")
|
||||
echo "==> reconstructed paths:"; printf '%s\n' "$RECON"
|
||||
|
||||
for path in \
|
||||
usr/bin/fips usr/bin/fipsctl usr/bin/fipstop usr/bin/fips-gateway \
|
||||
usr/bin/fips-mesh-setup \
|
||||
etc/init.d/fips etc/init.d/fips-gateway \
|
||||
etc/fips/fips.yaml etc/fips/firewall.sh etc/dnsmasq.d/fips.conf \
|
||||
etc/sysctl.d/fips-gateway.conf etc/sysctl.d/fips-bridge.conf \
|
||||
etc/hotplug.d/net/99-fips etc/uci-defaults/90-fips-setup \
|
||||
lib/upgrade/keep.d/fips; do
|
||||
if printf '%s\n' "$RECON" | grep -qxF "$path"; then
|
||||
echo " PASS path: $path"
|
||||
else
|
||||
echo " FAIL path: missing $path"; fail=1
|
||||
fi
|
||||
done
|
||||
|
||||
if [ "$fail" -ne 0 ]; then
|
||||
echo "apk structural verification FAILED"
|
||||
exit 1
|
||||
fi
|
||||
echo "apk structural verification PASS"
|
||||
|
||||
- name: SHA-256 hashes
|
||||
run: |
|
||||
echo "==> Binaries:"
|
||||
sha256sum bins/fips bins/fipsctl bins/fipstop
|
||||
echo "==> Package:"
|
||||
sha256sum dist/${{ env.PACKAGE_FILENAME }}
|
||||
|
||||
- name: Upload artifact (GitHub only)
|
||||
if: ${{ env.ACT != 'true' }}
|
||||
uses: actions/upload-artifact@v7
|
||||
with:
|
||||
name: ${{ env.PACKAGE_FILENAME }}
|
||||
path: dist/${{ env.PACKAGE_FILENAME }}
|
||||
retention-days: 30
|
||||
|
||||
- name: Install nak
|
||||
shell: bash
|
||||
run: |
|
||||
NAK_VERSION="0.16.2"
|
||||
ARCH=$(uname -m)
|
||||
case "$ARCH" in
|
||||
x86_64|amd64) NAK_ARCH="amd64" ;;
|
||||
aarch64|arm64) NAK_ARCH="arm64" ;;
|
||||
*) echo "Unsupported architecture: $ARCH"; exit 1 ;;
|
||||
esac
|
||||
curl -fsSL "https://github.com/fiatjaf/nak/releases/download/v${NAK_VERSION}/nak-v${NAK_VERSION}-linux-${NAK_ARCH}" \
|
||||
-o /usr/local/bin/nak
|
||||
chmod +x /usr/local/bin/nak
|
||||
nak --version
|
||||
|
||||
- name: Install jq
|
||||
run: |
|
||||
if ! command -v jq &>/dev/null; then
|
||||
sudo apt-get update && sudo apt-get install -y jq
|
||||
fi
|
||||
|
||||
# Priority: HIVE_CI_NSEC from env (loom job) > repo secret > generate ephemeral
|
||||
- name: Resolve signing key
|
||||
id: keys
|
||||
shell: bash
|
||||
env:
|
||||
SECRET_NSEC: ${{ secrets.HIVE_CI_NSEC }}
|
||||
run: |
|
||||
: ${GITHUB_OUTPUT:=/tmp/github_output}
|
||||
if [ -n "${HIVE_CI_NSEC:-}" ]; then
|
||||
echo "Using HIVE_CI_NSEC from loom job environment"
|
||||
NSEC="$HIVE_CI_NSEC"
|
||||
elif [ -n "$SECRET_NSEC" ]; then
|
||||
echo "Using HIVE_CI_NSEC from repository secrets"
|
||||
NSEC="$SECRET_NSEC"
|
||||
else
|
||||
echo "No nsec provided -- generating ephemeral keypair"
|
||||
NSEC=$(nak key generate)
|
||||
fi
|
||||
|
||||
PUBKEY=$(echo "$NSEC" | nak key public)
|
||||
echo "::add-mask::$NSEC"
|
||||
echo "nsec=$NSEC" >> "$GITHUB_OUTPUT"
|
||||
echo "pubkey=$PUBKEY" >> "$GITHUB_OUTPUT"
|
||||
echo "Publisher pubkey (hex): $PUBKEY"
|
||||
|
||||
- name: Upload to Blossom
|
||||
id: blossom_upload
|
||||
continue-on-error: true
|
||||
shell: bash
|
||||
env:
|
||||
BLOSSOM_SERVER: "https://blossom.primal.net"
|
||||
NSEC: ${{ steps.keys.outputs.nsec }}
|
||||
run: |
|
||||
: ${GITHUB_OUTPUT:=/tmp/github_output}
|
||||
|
||||
FILE_HASH=""
|
||||
for attempt in 1 2 3; do
|
||||
if UPLOAD_RESPONSE=$(nak blossom upload \
|
||||
--server "$BLOSSOM_SERVER" \
|
||||
--sec "$NSEC" \
|
||||
"dist/${{ env.PACKAGE_FILENAME }}" < /dev/null); then
|
||||
echo "Upload response (attempt $attempt):"
|
||||
echo "$UPLOAD_RESPONSE"
|
||||
FILE_HASH=$(echo "$UPLOAD_RESPONSE" | jq -r '.sha256')
|
||||
if [ -n "$FILE_HASH" ] && [ "$FILE_HASH" != "null" ]; then
|
||||
break
|
||||
fi
|
||||
echo "Upload response had no sha256 (attempt $attempt)"
|
||||
else
|
||||
echo "Blossom upload timed out or failed (attempt $attempt)"
|
||||
fi
|
||||
FILE_HASH=""
|
||||
[ "$attempt" -lt 3 ] && sleep $((attempt * 10))
|
||||
done
|
||||
|
||||
if [ -z "$FILE_HASH" ] || [ "$FILE_HASH" = "null" ]; then
|
||||
echo "Blossom upload did not succeed after 3 attempts; non-fatal."
|
||||
echo "The package still ships as a GitHub release artifact; only the"
|
||||
echo "supplementary Blossom/nostr distribution is skipped this run."
|
||||
exit 1
|
||||
fi
|
||||
|
||||
BLOSSOM_URL="${BLOSSOM_SERVER}/${FILE_HASH}"
|
||||
echo "url=$BLOSSOM_URL" >> "$GITHUB_OUTPUT"
|
||||
echo "hash=$FILE_HASH" >> "$GITHUB_OUTPUT"
|
||||
echo "Uploaded to Blossom: $BLOSSOM_URL"
|
||||
|
||||
- name: Publish NIP-94 release event
|
||||
id: publish
|
||||
if: steps.blossom_upload.outcome == 'success'
|
||||
shell: bash
|
||||
env:
|
||||
RELAYS: "wss://relay.damus.io wss://nos.lol wss://nostr.mom wss://offchain.pub"
|
||||
NSEC: ${{ steps.keys.outputs.nsec }}
|
||||
run: |
|
||||
: ${GITHUB_OUTPUT:=/tmp/github_output}
|
||||
set -e
|
||||
|
||||
VERSION="${{ needs.determine-versioning.outputs.package_version }}"
|
||||
CHANNEL="${{ needs.determine-versioning.outputs.release_channel }}"
|
||||
|
||||
nak event --sec "$NSEC" -k 1063 \
|
||||
-c "FIPS Package: ${{ env.PACKAGE_NAME }} for ${{ matrix.openwrt_arch }} (apk)" \
|
||||
--tag url="${{ steps.blossom_upload.outputs.url }}" \
|
||||
--tag m="application/octet-stream" \
|
||||
--tag x="${{ steps.blossom_upload.outputs.hash }}" \
|
||||
--tag ox="${{ steps.blossom_upload.outputs.hash }}" \
|
||||
--tag filename="${{ env.PACKAGE_FILENAME }}" \
|
||||
--tag A="${{ matrix.openwrt_arch }}" \
|
||||
--tag v="$VERSION" \
|
||||
--tag n="${{ env.PACKAGE_NAME }}" \
|
||||
--tag format="apk" \
|
||||
--tag compression="none" \
|
||||
> event.json 2> event.err
|
||||
|
||||
if [ ! -s event.json ]; then
|
||||
echo "Failed to create event"
|
||||
cat event.err 2>/dev/null || true
|
||||
exit 1
|
||||
fi
|
||||
|
||||
echo "=== Event JSON ==="
|
||||
cat event.json
|
||||
echo "=================="
|
||||
|
||||
EVENT_ID=$(jq -r '.id' event.json)
|
||||
if [ -z "$EVENT_ID" ] || [ "$EVENT_ID" = "null" ]; then
|
||||
echo "Failed to extract event ID"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# Publish to relays
|
||||
cat event.json | nak event $RELAYS 2>&1
|
||||
|
||||
echo "eventId=$EVENT_ID" >> "$GITHUB_OUTPUT"
|
||||
echo "Published NIP-94 event: $EVENT_ID"
|
||||
|
||||
- name: Build Summary
|
||||
run: |
|
||||
echo "Build Summary for ${{ matrix.openwrt_arch }} (apk):"
|
||||
echo " Package: ${{ env.PACKAGE_FILENAME }}"
|
||||
echo " apk version: ${{ needs.determine-versioning.outputs.apk_version }}"
|
||||
echo " Release EventId: ${{ steps.publish.outputs.eventId }}"
|
||||
echo " Blossom URL: ${{ steps.blossom_upload.outputs.url }}"
|
||||
|
||||
release:
|
||||
name: Publish GitHub Release (github only)
|
||||
runs-on: ubuntu-latest
|
||||
needs: build
|
||||
needs: [build, build-apk]
|
||||
if: startsWith(github.ref, 'refs/tags/')
|
||||
permissions:
|
||||
contents: write
|
||||
|
||||
steps:
|
||||
- name: Download all .ipk artifacts
|
||||
uses: actions/download-artifact@v4
|
||||
- name: Download package artifacts
|
||||
uses: actions/download-artifact@v8
|
||||
with:
|
||||
# Only the .ipk/.apk packages (named fips_<ver>_<arch>.*), not the
|
||||
# fips-bins-* raw-binary artifacts shared between the build jobs.
|
||||
pattern: fips_*
|
||||
path: dist
|
||||
merge-multiple: true
|
||||
|
||||
- name: Generate OpenWrt release checksums
|
||||
run: |
|
||||
cd dist
|
||||
find . -maxdepth 1 -type f \( -name '*.ipk' -o -name '*.apk' \) -printf '%P\n' \
|
||||
| LC_ALL=C sort \
|
||||
| xargs sha256sum \
|
||||
> checksums-openwrt.txt
|
||||
|
||||
- name: Create release
|
||||
uses: softprops/action-gh-release@v2
|
||||
with:
|
||||
files: dist/*.ipk
|
||||
files: |
|
||||
dist/*.ipk
|
||||
dist/*.apk
|
||||
dist/checksums-openwrt.txt
|
||||
generate_release_notes: true
|
||||
|
||||
209
.github/workflows/package-windows.yml
vendored
Normal file
@@ -0,0 +1,209 @@
|
||||
name: Windows Package
|
||||
on:
|
||||
push:
|
||||
branches:
|
||||
- master
|
||||
- maint
|
||||
- next
|
||||
tags:
|
||||
- "v*"
|
||||
pull_request:
|
||||
workflow_dispatch:
|
||||
|
||||
env:
|
||||
CARGO_TERM_COLOR: always
|
||||
|
||||
jobs:
|
||||
determine-versioning:
|
||||
runs-on: windows-latest
|
||||
outputs:
|
||||
package_version: ${{ steps.version.outputs.package_version }}
|
||||
steps:
|
||||
- uses: actions/checkout@v6
|
||||
with:
|
||||
fetch-depth: 0
|
||||
|
||||
- name: Derive package version
|
||||
id: version
|
||||
shell: pwsh
|
||||
run: |
|
||||
$cargoToml = Get-Content Cargo.toml -Raw
|
||||
if ($cargoToml -match 'version\s*=\s*"([^"]+)"') {
|
||||
$baseVersion = $Matches[1]
|
||||
} else {
|
||||
throw "Could not determine version from Cargo.toml"
|
||||
}
|
||||
|
||||
if ($env:GITHUB_REF -like "refs/tags/*") {
|
||||
$version = $env:GITHUB_REF_NAME -replace '^v', ''
|
||||
} else {
|
||||
$branch = $env:GITHUB_REF_NAME -replace '[^A-Za-z0-9]', '.' -replace '\.\.+', '.' -replace '^\.|\.$$', ''
|
||||
$height = git rev-list --count HEAD
|
||||
$hash = git rev-parse --short HEAD
|
||||
if (-not $branch) { $branch = "ref" }
|
||||
$version = "$baseVersion+$branch.$height.$hash"
|
||||
}
|
||||
|
||||
echo "package_version=$version" >> $env:GITHUB_OUTPUT
|
||||
|
||||
build:
|
||||
name: Build Windows package
|
||||
runs-on: windows-latest
|
||||
needs: determine-versioning
|
||||
|
||||
steps:
|
||||
- uses: actions/checkout@v6
|
||||
with:
|
||||
fetch-depth: 0
|
||||
|
||||
- name: Set SOURCE_DATE_EPOCH from git
|
||||
shell: pwsh
|
||||
run: |
|
||||
$epoch = git log -1 --format=%ct
|
||||
echo "SOURCE_DATE_EPOCH=$epoch" >> $env:GITHUB_ENV
|
||||
|
||||
- name: Install Rust toolchain
|
||||
uses: actions-rust-lang/setup-rust-toolchain@v1
|
||||
with:
|
||||
cache: false
|
||||
rustflags: ''
|
||||
|
||||
- name: Cache Cargo registry + build
|
||||
uses: actions/cache@v5
|
||||
with:
|
||||
path: |
|
||||
~/.cargo/registry
|
||||
~/.cargo/git
|
||||
target
|
||||
key: windows-release-x86_64-${{ hashFiles('**/Cargo.lock') }}
|
||||
restore-keys: |
|
||||
windows-release-x86_64-
|
||||
|
||||
- name: Build release binaries
|
||||
run: cargo build --release
|
||||
|
||||
- name: Build Windows package
|
||||
shell: pwsh
|
||||
run: |
|
||||
powershell -File packaging/windows/build-zip.ps1 `
|
||||
-Version "${{ needs.determine-versioning.outputs.package_version }}" `
|
||||
-NoBuild
|
||||
|
||||
- name: Verify ZIP structural correctness
|
||||
shell: pwsh
|
||||
run: |
|
||||
$zip = Get-ChildItem deploy\fips-*-windows-*.zip | Select-Object -First 1
|
||||
if (-not $zip) { Write-Error "No ZIP artifact found in deploy\"; exit 1 }
|
||||
Write-Host "Verifying: $($zip.Name)"
|
||||
|
||||
$extractDir = "verify-extract"
|
||||
if (Test-Path $extractDir) { Remove-Item -Recurse -Force $extractDir }
|
||||
Expand-Archive -Path $zip.FullName -DestinationPath $extractDir
|
||||
|
||||
# Expected top-level files (flat ZIP, no wrapper directory).
|
||||
# Source of truth: packaging/windows/build-zip.ps1
|
||||
$expected = @(
|
||||
"fips.exe",
|
||||
"fipsctl.exe",
|
||||
"fipstop.exe",
|
||||
"fips.yaml",
|
||||
"hosts",
|
||||
"install-service.ps1",
|
||||
"uninstall-service.ps1",
|
||||
"README.txt"
|
||||
)
|
||||
|
||||
$missing = @()
|
||||
foreach ($f in $expected) {
|
||||
$path = Join-Path $extractDir $f
|
||||
if (Test-Path -LiteralPath $path -PathType Leaf) {
|
||||
Write-Host "PASS: $f"
|
||||
} else {
|
||||
Write-Host "FAIL: $f"
|
||||
$missing += $f
|
||||
}
|
||||
}
|
||||
|
||||
Write-Host ""
|
||||
Write-Host "ZIP contents:"
|
||||
Get-ChildItem -Path $extractDir -Recurse | ForEach-Object {
|
||||
Write-Host " $($_.FullName.Substring((Resolve-Path $extractDir).Path.Length + 1))"
|
||||
}
|
||||
|
||||
if ($missing.Count -gt 0) {
|
||||
Write-Error "Missing expected files: $($missing -join ', ')"
|
||||
exit 1
|
||||
}
|
||||
Write-Host ""
|
||||
Write-Host "All expected files present."
|
||||
|
||||
- name: SHA-256 hash
|
||||
shell: pwsh
|
||||
run: |
|
||||
Write-Host "==> Windows release asset:"
|
||||
Get-ChildItem deploy\fips-*-windows-*.zip | ForEach-Object {
|
||||
Get-FileHash $_.FullName -Algorithm SHA256 | Format-Table -AutoSize
|
||||
}
|
||||
|
||||
- name: Upload artifact
|
||||
uses: actions/upload-artifact@v7
|
||||
with:
|
||||
name: fips_${{ needs.determine-versioning.outputs.package_version }}_x86_64_windows
|
||||
path: deploy/fips-*-windows-*.zip
|
||||
retention-days: 30
|
||||
|
||||
- name: Build summary
|
||||
shell: pwsh
|
||||
run: |
|
||||
$pkg = Get-ChildItem deploy\fips-*-windows-*.zip | Select-Object -First 1
|
||||
Write-Host "Build Summary for Windows/x86_64:"
|
||||
Write-Host " Package: $($pkg.Name)"
|
||||
Write-Host " Size: $([math]::Round($pkg.Length / 1MB, 2)) MB"
|
||||
|
||||
release:
|
||||
name: Publish Windows assets to GitHub Release
|
||||
runs-on: ubuntu-latest
|
||||
needs: build
|
||||
if: startsWith(github.ref, 'refs/tags/')
|
||||
permissions:
|
||||
contents: write
|
||||
|
||||
steps:
|
||||
- name: Download Windows artifacts
|
||||
uses: actions/download-artifact@v8
|
||||
with:
|
||||
path: dist
|
||||
merge-multiple: true
|
||||
|
||||
- name: Generate Windows release checksums
|
||||
run: |
|
||||
cd dist
|
||||
find . -maxdepth 1 -type f -name '*.zip' -printf '%P\n' \
|
||||
| LC_ALL=C sort \
|
||||
| xargs sha256sum \
|
||||
> checksums-windows.txt
|
||||
|
||||
- name: Wait for tag release
|
||||
env:
|
||||
GH_TOKEN: ${{ github.token }}
|
||||
run: |
|
||||
for attempt in $(seq 1 20); do
|
||||
if gh release view "${GITHUB_REF_NAME}" --repo "${GITHUB_REPOSITORY}" >/dev/null 2>&1; then
|
||||
exit 0
|
||||
fi
|
||||
echo "Release ${GITHUB_REF_NAME} not available yet; waiting..."
|
||||
sleep 15
|
||||
done
|
||||
|
||||
echo "Timed out waiting for release ${GITHUB_REF_NAME}" >&2
|
||||
exit 1
|
||||
|
||||
- name: Upload Windows assets
|
||||
env:
|
||||
GH_TOKEN: ${{ github.token }}
|
||||
run: |
|
||||
gh release upload "${GITHUB_REF_NAME}" \
|
||||
dist/*.zip \
|
||||
dist/checksums-windows.txt \
|
||||
--clobber \
|
||||
--repo "${GITHUB_REPOSITORY}"
|
||||
16
.gitignore
vendored
@@ -11,13 +11,16 @@
|
||||
.vscode/
|
||||
.idea/
|
||||
|
||||
# Claude Code
|
||||
# AI Agents
|
||||
.claude/
|
||||
AGENTS.md
|
||||
CLAUDE.md
|
||||
agents/
|
||||
|
||||
deploy/
|
||||
vps.env
|
||||
|
||||
reference/
|
||||
/reference/
|
||||
|
||||
dist/
|
||||
*.ipk
|
||||
@@ -28,4 +31,11 @@ sim-results/
|
||||
__pycache__/
|
||||
*.py[cod]
|
||||
*.egg-info/
|
||||
*.egg
|
||||
*.egg
|
||||
|
||||
# Runtime artifacts from running fips in-tree during local testing.
|
||||
# Root-anchored so legitimately-tracked fips.yaml under packaging/ and
|
||||
# examples/ stays included.
|
||||
/fips.key
|
||||
/fips.pub
|
||||
/fips.yaml
|
||||
|
||||
1488
CHANGELOG.md
268
CONTRIBUTING.md
@@ -1,33 +1,267 @@
|
||||
# Contributing to FIPS
|
||||
|
||||
## Getting Started
|
||||
<!-- markdownlint-disable MD013 -->
|
||||
|
||||
Clone the repo and verify your setup:
|
||||
FIPS is a mesh routing protocol for Nostr identities over arbitrary
|
||||
transports. The architecture is layered, top to bottom:
|
||||
|
||||
```
|
||||
- **IPv6 TUN compatibility layer** — presents the mesh as a local
|
||||
network interface (`fips0`) so unmodified applications can use it.
|
||||
Applications send IPv6 packets to `fd::/8` addresses derived from
|
||||
Nostr pubkeys; the daemon converts between IPv6 packets and FSP
|
||||
datagrams.
|
||||
- **FSP** (FIPS Session Protocol) — end-to-end encrypted sessions
|
||||
between identities, with periodic rekey.
|
||||
- **FMP** (FIPS Mesh Protocol) — peer management, spanning tree,
|
||||
bloom filters, routing and forwarding, and link encryption.
|
||||
- **Transport** — the actual wire: UDP, TCP, Tor, Bluetooth LE,
|
||||
Ethernet, and so on. Each transport plugs into FMP via a trait.
|
||||
|
||||
Most non-trivial changes affect behavior visible across the mesh —
|
||||
how nodes find each other, how packets route, how sessions rekey, how
|
||||
peers recover from failure. A single-node `cargo test` run is
|
||||
necessary but not sufficient for that class of change; the integration
|
||||
harness in [testing/](testing/) is where regressions actually surface.
|
||||
This document covers the workflow assuming that context. Protocol
|
||||
depth lives in [docs/design/](docs/design/).
|
||||
|
||||
## Quick start
|
||||
|
||||
```bash
|
||||
git clone https://github.com/jmcorgan/fips.git
|
||||
cd fips
|
||||
cargo build
|
||||
cargo test
|
||||
```
|
||||
|
||||
Read [docs/design/](docs/design/) for protocol understanding, starting with
|
||||
[fips-intro.md](docs/design/fips-intro.md).
|
||||
The pinned toolchain in [rust-toolchain.toml](rust-toolchain.toml) is
|
||||
used for deterministic builds. On Linux, a source build requires
|
||||
`libclang` (`sudo apt install libclang-dev` on Debian/Ubuntu): the LAN
|
||||
gateway's nftables bindings are generated by `bindgen` at build time
|
||||
and fail without it. BLE-capable builds additionally need `bluez`,
|
||||
`libdbus-1-dev`, and `pkg-config` installed; the default build picks
|
||||
up BLE if those are present and skips it cleanly if not.
|
||||
|
||||
## Filing Issues
|
||||
On Nix, `nix develop` provides the pinned toolchain and all of these
|
||||
build prerequisites without any manual install; see the Nix / NixOS
|
||||
section of [packaging/README.md](packaging/README.md).
|
||||
|
||||
- Search existing issues before opening a new one.
|
||||
- Include FIPS version, Rust version, and OS.
|
||||
- For bugs: steps to reproduce, expected vs actual behavior.
|
||||
For multi-node integration runs, Docker is required. The harness
|
||||
under [testing/](testing/) starts containerized topologies and
|
||||
exercises real mesh behavior; see [testing/README.md](testing/README.md)
|
||||
for the suite catalog.
|
||||
|
||||
## Pull Requests
|
||||
For a guided first-run that joins the public test mesh, see
|
||||
[docs/tutorials/join-the-test-mesh.md](docs/tutorials/join-the-test-mesh.md).
|
||||
Pointing your local daemon at a `test-*` node is the cheapest way to
|
||||
dogfood a change end-to-end before opening a PR.
|
||||
|
||||
- All PRs must pass `cargo build`, `cargo test`, and `cargo clippy` with no
|
||||
warnings.
|
||||
- Keep commits focused — one logical change per commit.
|
||||
- Add tests for new functionality.
|
||||
- Reference relevant design docs if the change touches protocol behavior.
|
||||
## Choosing a branch to target
|
||||
|
||||
## Questions
|
||||
FIPS uses three long-lived branches, each a superset of the previous:
|
||||
|
||||
Open a GitHub issue for design or implementation questions.
|
||||
- **`maint`** — bug fixes for the latest released version.
|
||||
- **`master`** — compatible work for the next feature release.
|
||||
- **`next`** — wire-format-breaking and API-breaking work, staged for
|
||||
the next forklift release.
|
||||
|
||||
Pick the branch that matches the scope of your change:
|
||||
|
||||
| Your change | Target |
|
||||
| --- | --- |
|
||||
| Bug fix in a feature that shipped in the latest release | `maint` |
|
||||
| Bug fix in code added on `master` since the last release | `master` |
|
||||
| Bug fix in `next`-only code (wire-format-breaking work) | `next` |
|
||||
| New feature, no wire-format or API break | `master` |
|
||||
| Wire-format-breaking or API-breaking change | `next` |
|
||||
| Documentation, CI, contributor-facing changes | `maint` if they apply to released material, else `master` |
|
||||
|
||||
When in doubt, ask in the issue. The maintainer can retarget if
|
||||
needed. The full release workflow, version conventions, and
|
||||
merge-direction rationale are in [docs/branching.md](docs/branching.md).
|
||||
|
||||
## Reporting bugs
|
||||
|
||||
Search [open issues](https://github.com/jmcorgan/fips/issues) before
|
||||
filing a new one — duplicates are common in a young project.
|
||||
|
||||
When you open a bug report, please include:
|
||||
|
||||
- **FIPS version** (`fipsctl --version`)
|
||||
- **Rust toolchain version** (`rustc --version`)
|
||||
- **OS / distro** (Linux distro + kernel, or macOS / Windows version)
|
||||
- **What you expected to happen** — your mental model of the
|
||||
behavior, ideally referencing the relevant docs or config field.
|
||||
- **What actually happened** — the observed behavior, including the
|
||||
surprise.
|
||||
- **Reproduction steps** — minimal and deterministic if you can.
|
||||
Multi-node bugs should include the topology and per-node config
|
||||
excerpts.
|
||||
- **Evidence** — relevant log excerpts (`journalctl -u fips` or stdout
|
||||
with `RUST_LOG=info` or `debug`), `fipsctl show` output if relevant
|
||||
(`peers`, `links`, `status`), and any visible mesh state.
|
||||
|
||||
One issue per bug. Don't bundle unrelated symptoms even if you
|
||||
suspect they share a root cause — the maintainer will link them if
|
||||
they turn out to be related.
|
||||
|
||||
## Submitting pull requests
|
||||
|
||||
### Scope discipline
|
||||
|
||||
Every PR should make one logical change. The reviewer should be able
|
||||
to read the whole diff and trace every line back to the PR's stated
|
||||
purpose.
|
||||
|
||||
- No drive-by reformatting of unrelated files.
|
||||
- No unrelated refactors folded into a bug fix or a feature PR.
|
||||
- No "while I was in there" cleanups in files outside the change's
|
||||
natural footprint. Send them as separate PRs; they'll usually land
|
||||
faster on their own.
|
||||
- Pre-existing lint warnings in files you didn't touch are not yours
|
||||
to fix in this PR.
|
||||
|
||||
### Required before opening any PR
|
||||
|
||||
Run these locally and confirm they all pass:
|
||||
|
||||
```bash
|
||||
cargo fmt --check
|
||||
cargo build
|
||||
cargo clippy --all-targets -- -D warnings
|
||||
cargo test
|
||||
```
|
||||
|
||||
`fmt` and `clippy -D warnings` are CI gates — PRs with formatting
|
||||
drift or new clippy warnings will fail CI and be sent back.
|
||||
|
||||
Then run the integration suite that exercises your change:
|
||||
|
||||
```bash
|
||||
./testing/ci-local.sh --only <suite>
|
||||
```
|
||||
|
||||
See [testing/README.md](testing/README.md) for the available suites
|
||||
and what each covers. Routing, discovery, rekey, NAT, gateway, and
|
||||
transport changes all have specific suites; pick the narrowest one
|
||||
that touches your code path.
|
||||
|
||||
**Recommended before opening**: the full local CI run.
|
||||
|
||||
```bash
|
||||
./testing/ci-local.sh
|
||||
```
|
||||
|
||||
This is the same matrix that runs on GitHub Actions. Catching a
|
||||
regression locally is much cheaper than catching it in CI.
|
||||
|
||||
### Self-review against the project review checklist
|
||||
|
||||
The 13-criteria checklist the maintainer runs on every incoming PR is
|
||||
published at [PR-REVIEW.md](PR-REVIEW.md). Run your own change through
|
||||
it before opening — or hand the document to your coding agent with
|
||||
"review my branch against this checklist" and let it do the pass. The
|
||||
checklist covers PR hygiene (body, commit shape, base freshness), diff
|
||||
content (does the change do what the description says, does it fit the
|
||||
codebase as a natural extension), and cross-cutting concerns (tests,
|
||||
docs, dependencies, security, contributor-conventional Rust patterns).
|
||||
|
||||
This is the first thing the maintainer does on any submission, so
|
||||
running it yourself saves a review round trip.
|
||||
|
||||
### Additional requirements for feature PRs
|
||||
|
||||
- **New CI coverage.** Features added without a test that exercises
|
||||
them won't be reviewed. Either extend an existing integration
|
||||
suite or add a new one under `testing/`. Coverage of just the
|
||||
happy path is fine for an initial PR; edge cases can land as
|
||||
follow-ups.
|
||||
- **Documentation updated alongside the code.** Protocol changes
|
||||
update the relevant [docs/design/](docs/design/) page. Config
|
||||
changes update the operator-facing docs in [docs/](docs/) and the
|
||||
reference config. Behavior visible to operators updates
|
||||
[README.md](README.md) and any tutorial it touches.
|
||||
|
||||
### Additional requirements for bug-fix PRs
|
||||
|
||||
- **A regression test** where practical. If a regression test isn't
|
||||
tractable (some bugs only surface under timing or scale that's hard
|
||||
to encode), say so in the PR description with a one-paragraph
|
||||
explanation.
|
||||
- **Commit message references the bug**: the symptom, the root cause
|
||||
in one sentence, and the fix shape.
|
||||
|
||||
### Merge mechanics
|
||||
|
||||
PRs are merged via **squash-merge**. One logical change per PR
|
||||
becomes one commit on the destination branch, which keeps `git
|
||||
bisect` useful across the integration suite. Your in-PR commit
|
||||
history doesn't matter for the final landed history — the maintainer
|
||||
rewrites the commit message at merge time.
|
||||
|
||||
## AI coding assistant policy
|
||||
|
||||
Use of AI coding assistants (Claude Code, Copilot, Cursor, Aider, and
|
||||
similar) in preparing a contribution is welcome. These tools are
|
||||
force multipliers and we have no objection in principle to their use
|
||||
in writing code, tests, documentation, or PR descriptions.
|
||||
|
||||
What we require is that the contributor does a thorough manual review
|
||||
and editorial pass over the output before submission. Concretely:
|
||||
|
||||
- Verify that the code does what it claims, not just that it
|
||||
compiles.
|
||||
- Verify that any tests the agent wrote actually test something
|
||||
useful, not just that they pass.
|
||||
- Verify that any documentation matches the behavior.
|
||||
- Spot-check the diff for nothing-surprising: no unrelated files
|
||||
modified, no fabricated APIs, no references to symbols that don't
|
||||
exist, no version bumps you didn't intend, no churn outside the
|
||||
change's natural footprint.
|
||||
- Be ready to discuss the design choices in the PR as if you wrote
|
||||
every line, because for the purposes of accountability you did.
|
||||
|
||||
The coding agent is a tool. The contributor is the author of record
|
||||
and is accountable for whatever they submit. PRs are reviewed on
|
||||
what they contain, not on who or what wrote them.
|
||||
|
||||
**Review effort scales with submission effort.** A submission that
|
||||
shows signs of being unreviewed agent output — irrelevant edits
|
||||
scattered across the tree, hallucinated function names, mismatched
|
||||
test/behavior pairs, fabricated API references, ChatGPT-style summary
|
||||
prose in comments — will receive an AI-coding-agent reply in turn,
|
||||
without human review. If you want a human reviewer's attention, do
|
||||
the editorial pass yourself first.
|
||||
|
||||
Repeated submissions of unreviewed AI output will result in the
|
||||
contributor being asked to step back and may result in account
|
||||
restrictions.
|
||||
|
||||
## Where the conversation happens
|
||||
|
||||
- **GitHub issues** — bugs, feature requests, design discussions
|
||||
that don't fit on a specific PR.
|
||||
- **GitHub PRs** — design discussion specific to a change in
|
||||
flight. Comment threads on the diff are the right place to push
|
||||
back on a decision.
|
||||
- **[fips.network](https://fips.network)** — community page, podcast,
|
||||
and the project's Nostr account. Broader project conversation and
|
||||
announcements happen here.
|
||||
|
||||
For implementation questions specific to your PR, ask in the PR
|
||||
itself. For design or roadmap questions that don't have a clear PR
|
||||
home yet, file a GitHub issue with the `design` label.
|
||||
|
||||
## Further reading
|
||||
|
||||
- [PR-REVIEW.md](PR-REVIEW.md) — the 13-criteria PR review checklist
|
||||
the maintainer runs on every incoming PR; run it yourself before
|
||||
opening to save a round trip.
|
||||
- [docs/design/README.md](docs/design/README.md) — protocol design tree.
|
||||
- [docs/branching.md](docs/branching.md) — full release workflow and
|
||||
merge-direction rationale.
|
||||
- [docs/getting-started.md](docs/getting-started.md) — operator
|
||||
walkthrough for a new node.
|
||||
- [docs/tutorials/join-the-test-mesh.md](docs/tutorials/join-the-test-mesh.md)
|
||||
— how to dogfood your change against the public test mesh.
|
||||
- [testing/README.md](testing/README.md) — integration suite catalog.
|
||||
|
||||
1871
Cargo.lock
generated
79
Cargo.toml
@@ -1,24 +1,25 @@
|
||||
[package]
|
||||
name = "fips"
|
||||
version = "0.2.1"
|
||||
version = "0.5.0-dev"
|
||||
edition = "2024"
|
||||
description = "A distributed, decentralized network routing protocol for mesh nodes connecting over arbitrary transports"
|
||||
license = "MIT"
|
||||
authors = ["Johnathan Corgan <jcorgan@corganlabs.com>"]
|
||||
repository = "https://github.com/jmcorgan/fips"
|
||||
homepage = "https://fips.network"
|
||||
readme = "README.md"
|
||||
|
||||
[features]
|
||||
default = ["tui"]
|
||||
tui = ["dep:ratatui"]
|
||||
keywords = ["mesh", "p2p", "decentralized", "overlay-network", "nostr"]
|
||||
categories = ["network-programming", "command-line-utilities", "cryptography"]
|
||||
|
||||
[dependencies]
|
||||
ratatui = { version = "0.30", optional = true }
|
||||
ratatui = "0.30"
|
||||
secp256k1 = { version = "0.30", features = ["rand", "global-context"] }
|
||||
sha2 = "0.10"
|
||||
hkdf = "0.12"
|
||||
chacha20poly1305 = "0.10"
|
||||
rand = "0.10.0"
|
||||
ring = "0.17"
|
||||
libm = "0.2"
|
||||
rand = "0.10.1"
|
||||
crossbeam-channel = "0.5"
|
||||
thiserror = "2.0"
|
||||
bech32 = "0.11"
|
||||
serde = { version = "1.0", features = ["derive"] }
|
||||
@@ -26,17 +27,37 @@ serde_json = "1.0"
|
||||
serde_yaml = "0.9"
|
||||
dirs = "6.0"
|
||||
hex = "0.4"
|
||||
clap = { version = "4.5", features = ["derive"] }
|
||||
clap = { version = "4.6", features = ["derive"] }
|
||||
tracing = "0.1"
|
||||
tracing-subscriber = { version = "0.3", features = ["env-filter"] }
|
||||
tun = { version = "0.8.5", features = ["async"] }
|
||||
libc = "0.2"
|
||||
rtnetlink = "0.20.0"
|
||||
tokio = { version = "1", features = ["rt", "macros", "signal", "sync", "net", "time"] }
|
||||
tokio = { version = "1", features = ["rt", "macros", "signal", "sync", "net", "time", "process", "io-util"] }
|
||||
futures = "0.3"
|
||||
simple-dns = "0.11.2"
|
||||
mdns-sd = "0.19"
|
||||
socket2 = { version = "0.6.2", features = ["all"] }
|
||||
tokio-socks = "0.5"
|
||||
portable-atomic = { version = "1", features = ["std"] }
|
||||
|
||||
nostr = { version = "0.44", features = ["std", "nip59"] }
|
||||
nostr-sdk = "0.44"
|
||||
arc-swap = "1"
|
||||
|
||||
[target.'cfg(unix)'.dependencies]
|
||||
tun = { version = "0.8.7", features = ["async"] }
|
||||
libc = "0.2"
|
||||
|
||||
[target.'cfg(target_os = "linux")'.dependencies]
|
||||
rtnetlink = "0.21.0"
|
||||
rustables = "0.8.7"
|
||||
procfs = { version = "0.18", default-features = false }
|
||||
|
||||
# bluer/BlueZ needs glibc — see build.rs `bluer_available` cfg gate.
|
||||
[target.'cfg(all(target_os = "linux", not(target_env = "musl")))'.dependencies]
|
||||
bluer = { version = "0.17", features = ["bluetoothd", "l2cap"] }
|
||||
|
||||
[target.'cfg(windows)'.dependencies]
|
||||
wintun = "0.5"
|
||||
windows-service = "0.8.1"
|
||||
|
||||
[package.metadata.deb]
|
||||
maintainer = "Johnathan Corgan <jcorgan@corganlabs.com>"
|
||||
@@ -44,34 +65,52 @@ copyright = "2026 Johnathan Corgan"
|
||||
license-file = ["LICENSE", "0"]
|
||||
section = "net"
|
||||
priority = "optional"
|
||||
depends = "libc6, systemd"
|
||||
depends = "libc6, systemd, libdbus-1-3"
|
||||
recommends = "bluez"
|
||||
extended-description = """\
|
||||
FIPS is a distributed, decentralized network routing protocol for mesh \
|
||||
nodes connecting over arbitrary transports including UDP, TCP, and Ethernet. \
|
||||
It provides encrypted peer-to-peer connectivity with automatic key management, \
|
||||
TUN-based virtual networking, and .fips DNS resolution."""
|
||||
nodes connecting over arbitrary transports including UDP, TCP, Ethernet, \
|
||||
Tor, and Bluetooth (BLE). It provides encrypted peer-to-peer connectivity \
|
||||
with automatic key management, TUN-based virtual networking, and .fips DNS \
|
||||
resolution."""
|
||||
maintainer-scripts = "packaging/debian/"
|
||||
assets = [
|
||||
["target/release/fips", "/usr/bin/", "755"],
|
||||
["target/release/fipsctl", "/usr/bin/", "755"],
|
||||
["target/release/fipstop", "/usr/bin/", "755"],
|
||||
["packaging/common/fips.yaml", "/etc/fips/fips.yaml", "600"],
|
||||
["packaging/common/fips.yaml", "/usr/share/fips/fips.yaml.example", "644"],
|
||||
["packaging/common/hosts", "/etc/fips/hosts", "644"],
|
||||
["packaging/common/fips.nft", "/etc/fips/fips.nft", "644"],
|
||||
["packaging/debian/fips.service", "/lib/systemd/system/fips.service", "644"],
|
||||
["packaging/debian/fips-dns.service", "/lib/systemd/system/fips-dns.service", "644"],
|
||||
["packaging/debian/fips-firewall.service", "/lib/systemd/system/fips-firewall.service", "644"],
|
||||
["packaging/common/fips-dns-setup", "/usr/lib/fips/fips-dns-setup", "755"],
|
||||
["packaging/common/fips-dns-teardown", "/usr/lib/fips/fips-dns-teardown", "755"],
|
||||
["packaging/debian/fips.tmpfiles", "/usr/lib/tmpfiles.d/fips.conf", "644"],
|
||||
["target/release/fips-gateway", "/usr/bin/", "755"],
|
||||
["packaging/debian/fips-gateway.service", "/lib/systemd/system/fips-gateway.service", "644"],
|
||||
["docs/design/fips-security.md", "/usr/share/doc/fips/fips-security.md", "644"],
|
||||
]
|
||||
conf-files = ["/etc/fips/fips.yaml", "/etc/fips/hosts"]
|
||||
conf-files = ["/etc/fips/hosts", "/etc/fips/fips.nft"]
|
||||
|
||||
[dev-dependencies]
|
||||
tempfile = "3.15"
|
||||
criterion = { version = "0.8.2", features = ["html_reports"] }
|
||||
tokio = { version = "1", features = ["test-util"] }
|
||||
|
||||
[[bin]]
|
||||
name = "fipsctl"
|
||||
path = "src/bin/fipsctl.rs"
|
||||
|
||||
[[bin]]
|
||||
name = "fips-gateway"
|
||||
path = "src/bin/fips-gateway.rs"
|
||||
|
||||
[[bin]]
|
||||
name = "fipstop"
|
||||
path = "src/bin/fipstop/main.rs"
|
||||
required-features = ["tui"]
|
||||
|
||||
[[bench]]
|
||||
name = "routing_next_hop"
|
||||
path = "benches/routing_next_hop.rs"
|
||||
harness = false
|
||||
|
||||
199
PR-REVIEW.md
Normal file
@@ -0,0 +1,199 @@
|
||||
# PR Review Checklist
|
||||
|
||||
<!-- markdownlint-disable MD013 -->
|
||||
|
||||
This is the 13-criteria checklist the maintainer runs against every
|
||||
incoming PR. The first pass on any submission is exactly this list,
|
||||
so executing it yourself before opening — or after pushing a fresh
|
||||
revision — saves a review round trip and surfaces problems faster.
|
||||
|
||||
The document is also written so you can hand it to a coding agent
|
||||
(Claude Code, Copilot, Cursor, Aider, etc.) with "review my branch
|
||||
against this checklist" and get a structured pass. The agent gets
|
||||
better results than a free-form "review my PR" because every concern
|
||||
the maintainer cares about is enumerated below.
|
||||
|
||||
## Step 1 — Should this even be reviewed?
|
||||
|
||||
Skip the review (and say so) if the PR is:
|
||||
|
||||
- closed, merged, or marked draft
|
||||
- automated (bot author, dependabot, etc.) and trivially OK
|
||||
- so small and obviously correct (typo fix, single-line doc tweak)
|
||||
that a thirteen-point pass is overkill — a one-paragraph informal
|
||||
review is better in that case
|
||||
|
||||
## Step 2 — Gather context
|
||||
|
||||
Read these *before* analyzing the diff so the review is grounded:
|
||||
|
||||
1. PR metadata. Title, body, author, head ref, base ref, head SHA,
|
||||
base SHA, mergeable status, CI rollup, commit list.
|
||||
|
||||
```bash
|
||||
gh pr view <num> --json title,body,author,headRefName,baseRefName,headRefOid,baseRefOid,mergeable,statusCheckRollup,commits
|
||||
```
|
||||
|
||||
2. The diff.
|
||||
|
||||
```bash
|
||||
gh pr diff <num>
|
||||
```
|
||||
|
||||
3. Base-branch freshness. How many commits have landed on the PR's
|
||||
base since the PR forked from it.
|
||||
4. Project guidance. Read [CLAUDE.md](CLAUDE.md) at the repo root and
|
||||
any nested `CLAUDE.md` in directories the diff touches. These
|
||||
describe project-specific conventions and constraints not visible
|
||||
from the diff alone.
|
||||
5. Related work on GitHub. Skim the [open issues](https://github.com/jmcorgan/fips/issues)
|
||||
and other [open PRs](https://github.com/jmcorgan/fips/pulls) for
|
||||
work that overlaps, duplicates, partially addresses, or is unblocked
|
||||
by this PR.
|
||||
6. For "this looks wrong" observations later: `git blame` the modified
|
||||
lines and read recent commit history on the same files for context
|
||||
before flagging something as a problem. What looks like a bug at
|
||||
first glance is often a deliberate workaround documented in a prior
|
||||
commit message.
|
||||
|
||||
## Step 3 — The 13 criteria
|
||||
|
||||
The review must address all 13 criteria below at some point. They
|
||||
group naturally into PR hygiene, diff content, and cross-cutting
|
||||
concerns — but the report itself is *not* organized this way; see
|
||||
Step 4.
|
||||
|
||||
### Group A — PR hygiene (structural review)
|
||||
|
||||
1. **PR body and issue cross-reference**. Does the body accurately
|
||||
describe the change (feature added or bug fixed) and match what
|
||||
the diff actually does? Is there an associated issue that
|
||||
should be referenced via `Closes #N` / `Fixes #N`?
|
||||
2. **Commit hygiene and base freshness**. Is the PR a clean set of
|
||||
commits (or a single commit) representing appropriately chunked
|
||||
development items, or are there intermediate "WIP" / "fix typo" /
|
||||
"address review" commits that should have been squashed? Is the
|
||||
branch based off a recent `maint` / `master` / `next`, or has the
|
||||
base diverged far enough that rebase work is needed?
|
||||
3. **Commit message quality**. Are the commit messages well-structured
|
||||
(subject + body where the change warrants), accurately referencing
|
||||
everything actually in each commit, and free of extraneous footers
|
||||
— particularly coding-assistant attribution (`Generated with
|
||||
Claude Code`, `Co-Authored-By: Claude`, similar from other AI
|
||||
tools)?
|
||||
|
||||
### Group B — Diff content
|
||||
|
||||
4. **Does it do what it says it does**. Walk each claimed behavior
|
||||
from the PR body against the actual diff lines.
|
||||
5. **Coherent whole**. Are all parts of the diff in service of the
|
||||
stated goal, or are there drive-by formatting changes, unrelated
|
||||
touch-ups, or scope creep?
|
||||
6. **Fits the codebase as a natural extension**. Does the new code
|
||||
use existing idioms, helpers, error types, and patterns, or does
|
||||
it introduce new ones where existing ones would have served?
|
||||
|
||||
### Group C — Cross-cutting concerns
|
||||
|
||||
7. **New dependency surface**. Any new crates, system deps,
|
||||
build-time requirements, or external-service dependencies?
|
||||
8. **New test coverage**. Are the new code paths covered, are the
|
||||
tests scoped correctly (unit / integration / end-to-end), and
|
||||
are there obvious test gaps? Don't reflag anything CI already
|
||||
enforces (formatting, lint, type errors, unit-test pass/fail).
|
||||
9. **Documentation impact**. Does this need a CHANGELOG entry,
|
||||
rustdoc updates, design-doc changes
|
||||
([docs/design/](docs/design/)), README adjustments, or operator
|
||||
doc updates in [docs/](docs/)?
|
||||
10. **Security vulnerabilities**. Any new attack surface,
|
||||
untrusted-input parsing, `unsafe` blocks, panic-on-untrusted
|
||||
paths, secret-handling concerns, or side-channel exposure?
|
||||
11. **Rust and OSS best practices**. Idiomatic error handling, no
|
||||
silently-swallowed errors, no `unwrap` / `expect` on untrusted
|
||||
input, no `#[allow]` without justification, appropriate
|
||||
visibility (`pub` vs `pub(crate)` vs private), naming, and
|
||||
module shape.
|
||||
12. **Overlap with existing work**. Cross-check open issues and
|
||||
other open PRs (and recently closed/merged ones) for related
|
||||
work that overlaps, duplicates, partially addresses, or is
|
||||
unblocked by this PR.
|
||||
13. **Other concerns**. Anything not captured above — wire-format
|
||||
implications, branch-flow questions (`maint` vs `master` vs
|
||||
`next`; see [docs/branching.md](docs/branching.md)),
|
||||
deployment / packaging impact, contributor coordination needs,
|
||||
fragility notes for future maintainers.
|
||||
|
||||
## Step 4 — Compose the review
|
||||
|
||||
The review report is **not** a Q&A walk through the 13 criteria.
|
||||
Write it as natural prose in a coherent, integrated narrative that
|
||||
reads start-to-finish. All 13 criteria must be addressed at some
|
||||
point in the body, but ordering, grouping, and emphasis follow the
|
||||
actual shape of THIS PR — lead with what matters most for this PR,
|
||||
not a fixed template.
|
||||
|
||||
A typical shape that often falls out naturally:
|
||||
|
||||
- **Opening paragraph**: what the PR does and the headline
|
||||
observations (subsumes criteria 1 and 4).
|
||||
- **Substantive body**: diff analysis, design fit, cross-cutting
|
||||
concerns, surprises, fragilities, missing coverage,
|
||||
cross-PR/issue overlap, anything unusual. Don't reference
|
||||
criterion numbers in the prose.
|
||||
- **Closing**: short summary and a proposed disposition — *land*,
|
||||
*land-with-followups* (list them), *request-changes* (with the
|
||||
blocking items called out), or *hold-for-thematic-batch*.
|
||||
|
||||
Short subheadings are fine where they aid scanning. Bullets are fine
|
||||
for enumerable items (test names, file paths, follow-up actions).
|
||||
Avoid bullets that just enumerate criterion responses.
|
||||
|
||||
## Step 5 — Filter aggressively
|
||||
|
||||
Quality over quantity. Do not flag:
|
||||
|
||||
- Pre-existing issues on lines the PR did not modify
|
||||
- Issues that linter, type-checker, formatter, or CI would catch
|
||||
- Pedantic style nitpicks a senior engineer would not call out
|
||||
- Likely intentional changes related to the broader goal
|
||||
- Things explicitly silenced by an `#[allow]` with justification
|
||||
- Stylistic preferences not anchored in `CLAUDE.md` or the
|
||||
surrounding codebase's idioms
|
||||
|
||||
When in doubt about whether something is worth surfacing: would a
|
||||
senior maintainer skim past it, or would they want it raised?
|
||||
Skim-past items don't belong in the report.
|
||||
|
||||
For every issue you *do* surface, include a concrete fix suggestion
|
||||
inline ("rename X to Y", "extract this into the existing helper at
|
||||
`foo.rs:42`", "add a test exercising the `Err` branch") so the
|
||||
author can act without a round-trip.
|
||||
|
||||
## Step 6 — Citation discipline
|
||||
|
||||
When the review references a specific code location, use full-SHA
|
||||
GitHub permalinks so the link survives future history rewrites:
|
||||
|
||||
```text
|
||||
https://github.com/jmcorgan/fips/blob/<full-40-char-sha>/<path>#L<start>-L<end>
|
||||
```
|
||||
|
||||
For multi-line ranges include at least one line of context before
|
||||
and after the line(s) being discussed. After `gh pr checkout <num>`,
|
||||
use `git rev-parse HEAD` to grab the full SHA — never partial SHAs
|
||||
in permalinks.
|
||||
|
||||
## Notes
|
||||
|
||||
- The review is one human's read of the PR. Confidence calibration
|
||||
matters: distinguish "this is a blocker" from "this is worth asking
|
||||
about" from "this is a fragility note for future maintainers." The
|
||||
closing disposition makes the action explicit.
|
||||
- If a re-review is triggered after the author pushes new commits,
|
||||
lead with the delta from the prior review rather than re-walking
|
||||
the whole PR.
|
||||
- This checklist exists to surface problems, not to assign blame.
|
||||
If you're running it as the author or via an agent, treat each
|
||||
finding as "would the maintainer ask about this?" — and either fix
|
||||
it before opening, or pre-empt it in the PR body so the maintainer
|
||||
doesn't have to ask.
|
||||
437
README.md
@@ -3,275 +3,270 @@
|
||||

|
||||
[](LICENSE)
|
||||
[](https://www.rust-lang.org/)
|
||||
[](#status--roadmap)
|
||||
[](#status--roadmap)
|
||||
|
||||
A distributed, decentralized network routing protocol for mesh nodes
|
||||
connecting over arbitrary transports.
|
||||
A self-organizing encrypted mesh network built on Nostr identities,
|
||||
capable of operating over arbitrary transports without central
|
||||
infrastructure.
|
||||
|
||||
> FIPS is under active development. The protocol and APIs are not yet stable.
|
||||
> See [Status & Roadmap](#status--roadmap) below.
|
||||
> FIPS is under active development. The protocol and APIs are not
|
||||
> yet stable. See [Status & roadmap](#status--roadmap) below.
|
||||
|
||||
## Overview
|
||||
## What FIPS does
|
||||
|
||||
FIPS is a self-organizing mesh network that operates natively over a variety
|
||||
of physical and logical media — local area networks, Bluetooth, serial links,
|
||||
radio, or the existing internet as an overlay. Nodes generate their own
|
||||
identities, discover each other, and route traffic without any central
|
||||
authority or global topology knowledge.
|
||||
A machine running FIPS becomes a node in the mesh with a
|
||||
self-generated cryptographic identity (a Nostr keypair). There are
|
||||
two equally-supported deployment modes.
|
||||
|
||||
FIPS uses Nostr keypairs (secp256k1/schnorr) as native node identities,
|
||||
allowing users to generate their own persistent or ephemeral node addresses.
|
||||
Nodes address each other by npub, and the same cryptographic identity serves
|
||||
as both the routing address and the basis for end-to-end encrypted sessions
|
||||
across the mesh.
|
||||
**As an overlay** on top of existing IP networks, FIPS lets your
|
||||
node reach any other FIPS node wherever it sits — behind a NAT, on
|
||||
a different ISP, on a phone over cellular, on a laptop with only
|
||||
Bluetooth in range, or behind a Tor onion. The mesh forwards IPv6
|
||||
traffic transparently and end-to-end encrypted, with no central VPN
|
||||
concentrator or coordinating server.
|
||||
|
||||
FIPS allows existing TCP/IP based network software to use the FIPS mesh
|
||||
network by generating a local IP address from the node npub and tunnelling
|
||||
IP packets to other endpoints transparently knowing only their npub. Native
|
||||
FIPS-aware applications do not need this IP tunneling or emulation capability.
|
||||
**Ground up** over raw Ethernet, WiFi, or Bluetooth, FIPS provides
|
||||
a complete permissionless network without any pre-existing IP
|
||||
infrastructure, ISP, or DNS. Any node that joins the link gets
|
||||
routable IPv6 addresses, peer discovery, and a path to every other
|
||||
node automatically.
|
||||
|
||||
All traffic over the FIPS mesh is encrypted and authenticated both
|
||||
hop-to-hop between peers and independently end-to-end between FIPS
|
||||
endpoints.
|
||||
Either way, existing networking software runs over it unchanged —
|
||||
SSH, HTTP servers, file transfer, anything IPv6-native works the
|
||||
same way it would on a local network.
|
||||
|
||||
## Features
|
||||
|
||||
- **Self-organizing mesh routing** — spanning tree coordinates with bloom
|
||||
filter guided discovery, no global routing tables
|
||||
- **Multi-transport** — UDP, TCP, Ethernet, and Tor today; designed for
|
||||
Bluetooth, serial, and radio
|
||||
- **Noise encryption** — hop-by-hop link encryption (IK) plus independent
|
||||
end-to-end session encryption (XK), with periodic rekey for forward secrecy
|
||||
- **Nostr-native identity** — secp256k1 keypairs as node addresses, no
|
||||
registration or central authority
|
||||
- **IPv6 adaptation** — TUN interface maps npubs to fd00::/8 addresses for
|
||||
unmodified IP applications; static hostname mapping (`/etc/fips/hosts`)
|
||||
- **Metrics Measurement Protocol** — per-link RTT, loss, jitter, and goodput
|
||||
measurement with mesh size estimation
|
||||
- **ECN congestion signaling** — hop-by-hop CE flag relay with RFC 3168 IPv6
|
||||
marking, transport kernel drop detection
|
||||
- **Operator visibility** — `fipsctl` CLI and `fipstop` TUI dashboard for
|
||||
runtime inspection and runtime peer management
|
||||
- **Zero configuration** — sensible defaults; a node can start with no config
|
||||
file, though peer addresses are needed to join a network
|
||||
- **Self-organizing mesh routing.** Spanning-tree coordinates with
|
||||
bloom-filter-guided discovery; no global routing tables, no
|
||||
flooding.
|
||||
- **Multi-transport.** UDP, TCP, Ethernet, Tor, Nym, and Bluetooth
|
||||
(BLE L2CAP) ship today; transports compose on a single mesh and a
|
||||
node may run several at once.
|
||||
- **Two-layer encryption.** Noise IK between peers (hop-by-hop) and
|
||||
Noise XK between mesh endpoints (independent end-to-end), with
|
||||
periodic rekey for forward secrecy.
|
||||
- **Nostr-native identity.** secp256k1 / schnorr keypairs as node
|
||||
addresses; self-generated, no registration, no central authority.
|
||||
- **IPv6 adapter.** A TUN interface maps each remote npub to an
|
||||
`fd00::/8` address, so unmodified IPv6 software reaches mesh
|
||||
peers as `<npub>.fips`. Built-in `.fips` DNS resolver, with
|
||||
optional static name mapping via `/etc/fips/hosts`.
|
||||
- **Nostr-mediated discovery and NAT traversal.** Peers publish
|
||||
endpoint adverts on public Nostr relays, exchange candidates via
|
||||
NIP-59 gift-wrapped offers and answers, and establish direct
|
||||
paths through NATs using STUN-assisted hole punching. On the local
|
||||
network, mDNS LAN discovery finds peers directly without relays.
|
||||
- **LAN gateway.** Optional `fips-gateway` service folds an entire
|
||||
unmodified LAN into the mesh: outbound (LAN clients reach mesh
|
||||
destinations through a DNS-allocated virtual IPv6 pool and
|
||||
nftables NAT) and inbound (LAN-side services exposed to the mesh
|
||||
through 1:1 port forwards).
|
||||
- **Per-link metrics.** RTT, loss, jitter, and goodput on every
|
||||
hop, plus mesh-size estimation, via the Metrics Measurement
|
||||
Protocol.
|
||||
- **ECN congestion signaling.** Hop-by-hop CE-flag relay with RFC
|
||||
3168 IPv6 marking and transport kernel-drop detection.
|
||||
- **Mesh-interface security baseline.** Optional default-deny
|
||||
nftables policy for `fips0` shipped as a packaged conffile
|
||||
(`/etc/fips/fips.nft`) with an operator drop-in directory
|
||||
(`/etc/fips/fips.d/`) and a disabled-by-default
|
||||
`fips-firewall.service`. The baseline polices only the mesh
|
||||
interface, leaving Docker, Tor, and the host firewall untouched.
|
||||
- **Operator visibility.** `fipsctl` CLI for control and inspection
|
||||
with time-series stats history queryable for any metric,
|
||||
`fipstop` TUI for live status with inline sparkline dashboards,
|
||||
and a JSON-line control socket on each binary for direct
|
||||
programmatic access.
|
||||
- **Reproducible builds** with toolchain pinning and
|
||||
`SOURCE_DATE_EPOCH`.
|
||||
|
||||
## Building
|
||||
## Quick start
|
||||
|
||||
The shortest path on Debian / Ubuntu:
|
||||
|
||||
```bash
|
||||
git clone https://github.com/jmcorgan/fips.git
|
||||
cd fips
|
||||
cargo build --release
|
||||
```
|
||||
|
||||
Requires Rust 1.85+ (edition 2024) and Linux with TUN support.
|
||||
|
||||
## Installation
|
||||
|
||||
After building, choose one of the following methods to install.
|
||||
|
||||
### Debian / Ubuntu (.deb)
|
||||
|
||||
Requires [cargo-deb](https://crates.io/crates/cargo-deb):
|
||||
|
||||
```bash
|
||||
cargo install cargo-deb
|
||||
cargo deb
|
||||
sudo dpkg -i target/debian/fips_*.deb
|
||||
```
|
||||
|
||||
This installs the daemon, CLI tools, systemd units, and a default
|
||||
configuration. Edit `/etc/fips/fips.yaml` before starting:
|
||||
|
||||
```bash
|
||||
sudo nano /etc/fips/fips.yaml
|
||||
sudo systemctl start fips
|
||||
```
|
||||
|
||||
The service is enabled at boot automatically. To use `fipsctl` and
|
||||
`fipstop` without sudo, add your user to the `fips` group:
|
||||
This installs the daemon, CLI tools (`fipsctl`, `fipstop`), the
|
||||
optional `fips-gateway` service, systemd units, and a default
|
||||
`/etc/fips/fips.yaml` you can edit before starting.
|
||||
|
||||
For macOS, Windows, OpenWrt, the systemd tarball, a Nix flake, or a
|
||||
from-source build, see [docs/getting-started.md](docs/getting-started.md)
|
||||
for the full multi-platform installation guide.
|
||||
|
||||
To join a live mesh and reach your first peer, follow the new-user
|
||||
tutorial progression starting at
|
||||
[docs/tutorials/join-the-test-mesh.md](docs/tutorials/join-the-test-mesh.md).
|
||||
|
||||
### Building from source
|
||||
|
||||
```bash
|
||||
sudo usermod -aG fips $USER # log out and back in to take effect
|
||||
cargo build --release
|
||||
```
|
||||
|
||||
Remove with `sudo dpkg -r fips` (preserves config) or
|
||||
`sudo dpkg -P fips` (removes everything including identity keys).
|
||||
Requires Rust 1.94.1+ (edition 2024). Linux, macOS, and Windows run as
|
||||
standalone daemons; Android is supported as an embedded library (the host
|
||||
app owns the TUN, e.g. a `VpnService`). Transport availability varies by
|
||||
platform.
|
||||
|
||||
### Generic Linux (systemd tarball)
|
||||
| Transport | Linux | macOS | Windows | Android | OpenWrt |
|
||||
|-----------|:-----:|:-----:|:-------:|:-------:|:-------:|
|
||||
| UDP | ✅ | ✅ | ✅ | ✅ | ✅ |
|
||||
| TCP | ✅ | ✅ | ✅ | ✅ | ✅ |
|
||||
| Ethernet | ✅ | ✅ | ❌ | ❌ | ✅ |
|
||||
| Tor | ✅ | ✅ | ✅ | ❌ | ✅ |
|
||||
| Nym | ✅ | ✅ | ✅ | ❌ | ❌ |
|
||||
| BLE | ✅ | ❌ | ❌ | ❌ | ❌ |
|
||||
|
||||
```bash
|
||||
./packaging/systemd/build-tarball.sh
|
||||
tar xzf deploy/fips-*-linux-*.tar.gz
|
||||
cd fips-*-linux-*/
|
||||
sudo ./install.sh
|
||||
```
|
||||
On Linux, a source build requires `libclang` — the LAN gateway's
|
||||
nftables bindings are generated by `bindgen` at build time, which
|
||||
needs `libclang.so` on the build host. Install it before building
|
||||
(`sudo apt install libclang-dev` on Debian / Ubuntu); without it the
|
||||
build fails inside the `rustables` crate with an "Unable to find
|
||||
libclang" error. This is a build-time prerequisite only — it is not a
|
||||
runtime dependency, and the pre-built `.deb` artifacts do not need it.
|
||||
|
||||
See [packaging/systemd/README.install.md](packaging/systemd/README.install.md)
|
||||
for the full installation and configuration guide.
|
||||
BLE is optional and, on Linux, requires BlueZ and libdbus
|
||||
(`sudo apt install bluez libdbus-1-dev` on Debian / Ubuntu). It is
|
||||
gated on a build-script probe — install the dependencies first and
|
||||
the `cargo build` line above picks it up. The OpenWrt ipk omits
|
||||
BLE because libdbus is not available on the target.
|
||||
|
||||
## Configuration
|
||||
Nym (mixnet) transport builds on all desktop platforms. The OpenWrt
|
||||
❌ is provisional, pending verification of `nym-socks5-client`
|
||||
availability on the target; it will flip to ✅ only if confirmed
|
||||
buildable there.
|
||||
|
||||
The default configuration file is installed at `/etc/fips/fips.yaml`:
|
||||
|
||||
```yaml
|
||||
# FIPS Node Configuration
|
||||
|
||||
node:
|
||||
identity:
|
||||
# By default, a new ephemeral keypair is generated on each start.
|
||||
# Uncomment persistent to keep the same identity across restarts;
|
||||
# on first start a keypair is saved to fips.key/fips.pub next to
|
||||
# this config file (mode 0600/0644).
|
||||
# persistent: true
|
||||
#
|
||||
# Or set an explicit key (overrides persistent):
|
||||
# nsec: "nsec1..."
|
||||
|
||||
tun:
|
||||
enabled: true
|
||||
name: fips0
|
||||
mtu: 1280
|
||||
|
||||
dns:
|
||||
enabled: true
|
||||
bind_addr: "127.0.0.1"
|
||||
port: 5354
|
||||
|
||||
transports:
|
||||
udp:
|
||||
bind_addr: "0.0.0.0:2121"
|
||||
|
||||
tcp:
|
||||
# Accepts inbound connections. No static outbound peers.
|
||||
bind_addr: "0.0.0.0:8443"
|
||||
|
||||
# Ethernet transport — uncomment and set your interface name.
|
||||
# ethernet:
|
||||
# interface: "eth0"
|
||||
# discovery: true
|
||||
# announce: true
|
||||
# auto_connect: true
|
||||
# accept_connections: true
|
||||
|
||||
peers:
|
||||
# Static peers for bootstrapping (UDP or TCP):
|
||||
- npub: "npub1qmc3cvfz0yu2hx96nq3gp55zdan2qclealn7xshgr448d3nh6lks7zel98"
|
||||
alias: "fips-test-node"
|
||||
addresses:
|
||||
- transport: udp
|
||||
addr: "217.77.8.91:2121"
|
||||
connect_policy: auto_connect
|
||||
```
|
||||
|
||||
See [docs/design/fips-configuration.md](docs/design/fips-configuration.md)
|
||||
for the full reference.
|
||||
|
||||
## Usage
|
||||
|
||||
### DNS Resolution
|
||||
|
||||
FIPS includes a DNS resolver (enabled by default, port 5354) that maps
|
||||
`.fips` names to fd00::/8 IPv6 addresses. With systemd-resolved:
|
||||
|
||||
```bash
|
||||
sudo resolvectl dns fips0 127.0.0.1:5354
|
||||
sudo resolvectl domain fips0 ~fips
|
||||
```
|
||||
|
||||
Then reach any FIPS node by npub with standard IPv6 tools:
|
||||
|
||||
```bash
|
||||
ping6 npub1bbb....fips
|
||||
ssh npub1bbb....fips
|
||||
```
|
||||
|
||||
### Monitoring
|
||||
|
||||
Use `fipsctl` to query a running node:
|
||||
|
||||
```bash
|
||||
fipsctl show status # Node status overview
|
||||
fipsctl show peers # Authenticated peers
|
||||
fipsctl show links # Active links
|
||||
fipsctl show tree # Spanning tree state
|
||||
fipsctl show sessions # End-to-end sessions
|
||||
fipsctl show transports # Transport instances
|
||||
fipsctl show routing # Routing table summary
|
||||
```
|
||||
|
||||
`fipstop` provides an interactive TUI dashboard with live-updating
|
||||
views of node status, peers, links, sessions, tree state, transports,
|
||||
and routing:
|
||||
|
||||
```bash
|
||||
fipstop # connect to local daemon
|
||||
fipstop -r 1 # 1-second refresh interval
|
||||
```
|
||||
|
||||
### Service Management
|
||||
|
||||
```bash
|
||||
sudo systemctl start fips
|
||||
sudo systemctl stop fips
|
||||
sudo systemctl restart fips
|
||||
sudo journalctl -u fips -f
|
||||
```
|
||||
|
||||
### Testing
|
||||
|
||||
See [testing/](testing/) for Docker-based integration test harnesses
|
||||
including static topology tests and stochastic chaos simulation.
|
||||
Alternatively, the repo ships a [Nix flake](flake.nix): `nix develop`
|
||||
drops you into a shell with the pinned toolchain and every build
|
||||
prerequisite (libclang, dbus, pkg-config) already provided, and
|
||||
`nix build .#fips` builds all four binaries with no host setup. See the
|
||||
Nix / NixOS section of [packaging/README.md](packaging/README.md).
|
||||
|
||||
## Documentation
|
||||
|
||||
Protocol design documentation is in [docs/design/](docs/design/), organized as
|
||||
a layered protocol specification. Start with
|
||||
[fips-intro.md](docs/design/fips-intro.md) for the full protocol overview.
|
||||
`docs/` is organised by reader purpose:
|
||||
|
||||
## Project Structure
|
||||
- **[Tutorials](docs/tutorials/)** — hand-held walk-throughs from
|
||||
a fresh install through to a participating mesh node, plus
|
||||
advanced deployments (gateway on OpenWrt, hosting services,
|
||||
ground-up two-device mesh).
|
||||
- **[How-to guides](docs/how-to/)** — operator recipes for
|
||||
specific tasks: firewall activation, Nostr discovery, Tor onion
|
||||
service, Bluetooth peering, LAN gateway deployment and
|
||||
troubleshooting, MTU diagnostics, host aliases, persistent
|
||||
identity, unprivileged-user setup, UDP buffer tuning.
|
||||
- **[Reference](docs/reference/)** — `fips.yaml` configuration,
|
||||
wire formats, control-socket protocol, CLI references for each
|
||||
binary, security posture matrix, Nostr events catalog, transport
|
||||
statistics inventory.
|
||||
- **[Design](docs/design/)** — protocol-level architecture and
|
||||
layer specifications. Start with
|
||||
[fips-concepts.md](docs/design/fips-concepts.md) for the framing,
|
||||
then [fips-architecture.md](docs/design/fips-architecture.md) for
|
||||
the protocol stack.
|
||||
|
||||
If you want to contribute, see [CONTRIBUTING.md](CONTRIBUTING.md)
|
||||
and [testing/README.md](testing/README.md).
|
||||
|
||||
## Examples
|
||||
|
||||
- **[examples/sidecar-nostr-relay/](examples/sidecar-nostr-relay/)** —
|
||||
Run a [strfry](https://github.com/hoytech/strfry) Nostr relay
|
||||
reachable exclusively over the FIPS mesh. The relay container
|
||||
shares the FIPS sidecar's network namespace and is isolated from
|
||||
the host network.
|
||||
- **[examples/sidecar-nostr-mixnet-relay/](examples/sidecar-nostr-mixnet-relay/)** —
|
||||
Single-container demo of FIPS peering through a **mixnet**
|
||||
(implemented with [Nym](https://nym.com/)): the FIPS daemon, the mixnet
|
||||
proxy, and a strfry Nostr relay all in one isolated container, with
|
||||
the direct route to the peer firewalled off so traffic provably
|
||||
crosses the mixnet.
|
||||
- **[examples/k8s-sidecar/](examples/k8s-sidecar/)** — Run FIPS as
|
||||
a Kubernetes Pod sidecar. The sidecar creates `fips0` in the
|
||||
Pod's shared network namespace so every other container in the
|
||||
Pod gets mesh access without modification.
|
||||
- **[examples/wireguard-sidecar-macos/](examples/wireguard-sidecar-macos/)** —
|
||||
Reach the FIPS mesh from a macOS host through a local Docker
|
||||
container over a WireGuard tunnel. Only traffic destined for
|
||||
`fd00::/8` transits the sidecar; regular internet traffic
|
||||
continues to use the host network.
|
||||
|
||||
## Project structure
|
||||
|
||||
```text
|
||||
src/ Rust source (library + fips/fipsctl/fipstop binaries)
|
||||
packaging/ Debian, systemd tarball, and shared packaging files
|
||||
docs/design/ Protocol design specifications
|
||||
testing/ Docker-based integration test harnesses
|
||||
src/ Rust source: library + fips, fipsctl, fipstop, fips-gateway binaries
|
||||
docs/ Documentation: tutorials, how-to, reference, design
|
||||
packaging/ Debian, macOS .pkg, Windows ZIP, OpenWrt ipk, AUR, systemd tarball
|
||||
examples/ Deployment examples (Nostr relay, K8s sidecar, macOS WireGuard)
|
||||
testing/ Docker-based integration test harnesses + chaos simulation
|
||||
```
|
||||
|
||||
## Status & Roadmap
|
||||
## Status & roadmap
|
||||
|
||||
FIPS is at **v0.2.1**. The core protocol works end-to-end over UDP, TCP,
|
||||
Ethernet, and Tor with a small live mesh of deployed nodes.
|
||||
FIPS is at **v0.5.0-dev** on the `master` branch.
|
||||
[v0.4.1](https://github.com/jmcorgan/fips/releases/tag/v0.4.1) has
|
||||
shipped; this development line continues the testing-and-polishing
|
||||
track toward v0.5.0. The core protocol works end-to-end over
|
||||
UDP, TCP, Ethernet, Tor, Nym, and Bluetooth on a global, public test
|
||||
mesh of thousands of nodes. v0.4.0 added the Nym mixnet transport and
|
||||
mDNS LAN discovery alongside the existing Nostr-mediated peer discovery,
|
||||
UDP NAT traversal, peer ACL, and packaging hardening. New wire-format work
|
||||
continues to be staged on the `next` branch for the subsequent
|
||||
release line.
|
||||
|
||||
### What works today
|
||||
|
||||
- Spanning tree construction with greedy coordinate routing
|
||||
- Bloom filter guided discovery (no flooding, single-path with retry)
|
||||
- Noise IK (link layer) and Noise XK (session layer) encryption
|
||||
- Periodic Noise rekey with hitless cutover for forward secrecy (FMP + FSP)
|
||||
- Persistent node identity with key file management
|
||||
- IPv6 TUN adapter with DNS resolution of `.fips` names
|
||||
- Static hostname mapping (`/etc/fips/hosts`) with auto-reload
|
||||
- Per-link metrics (RTT, loss, jitter, goodput) and mesh size estimation
|
||||
- ECN congestion signaling (hop-by-hop CE relay, IPv6 CE marking, kernel drop detection)
|
||||
- UDP, TCP, Ethernet, and Tor transports (SOCKS5 outbound + directory-mode onion service inbound)
|
||||
- Runtime inspection and peer management via `fipsctl` and `fipstop`
|
||||
- Reproducible builds with toolchain pinning and SOURCE_DATE_EPOCH
|
||||
- Debian and systemd tarball packaging
|
||||
- Docker-based integration and chaos testing
|
||||
- Spanning-tree construction with greedy coordinate routing.
|
||||
- Bloom-filter-guided destination discovery (no flooding,
|
||||
single-path with retry).
|
||||
- Two-layer Noise encryption (IK at the link, XK at the session)
|
||||
with periodic hitless rekey for forward secrecy at both layers.
|
||||
- Persistent or ephemeral node identity with key-file management.
|
||||
- IPv6 TUN adapter with built-in `.fips` DNS resolver and
|
||||
multi-backend auto-configuration (systemd dns-delegate,
|
||||
systemd-resolved, dnsmasq, NetworkManager).
|
||||
- Static hostname mapping (`/etc/fips/hosts`) with auto-reload.
|
||||
- Per-link metrics (RTT, loss, jitter, goodput) and mesh size
|
||||
estimation.
|
||||
- ECN congestion signaling (hop-by-hop CE relay, IPv6 CE marking,
|
||||
kernel-drop detection).
|
||||
- UDP, TCP, Ethernet, Tor, Nym (mixnet), and BLE transports (BLE
|
||||
via L2CAP CoC with per-link MTU negotiation).
|
||||
- Nostr-mediated overlay endpoint discovery and UDP hole punching
|
||||
for NAT traversal, plus mDNS LAN discovery for local peers.
|
||||
- LAN gateway (`fips-gateway`) with both outbound (LAN-to-mesh)
|
||||
and inbound (mesh-to-LAN port-forwarding) modes.
|
||||
- Peer ACL: per-npub allow / deny admission control at the link
|
||||
layer; opt-in mesh-firewall baseline at `fips0` ingress.
|
||||
- Runtime inspection and peer management via `fipsctl` and
|
||||
`fipstop`.
|
||||
- Reproducible builds with toolchain pinning and
|
||||
`SOURCE_DATE_EPOCH`.
|
||||
- Linux (Debian, systemd tarball, OpenWrt, AUR), macOS (`.pkg`),
|
||||
and Windows (ZIP, service) packaging.
|
||||
- Docker-based integration and chaos testing.
|
||||
|
||||
### Near-term priorities
|
||||
|
||||
- Peer discovery via Nostr relays (bootstrap without static peer lists)
|
||||
- Native API for FIPS-aware applications (npub:port addressing)
|
||||
- Additional transports (Bluetooth)
|
||||
- Security audit of cryptographic protocols
|
||||
- Native API for FIPS-aware applications (npub:port addressing
|
||||
without the IPv6-shim path).
|
||||
- Security audit of the cryptographic protocols.
|
||||
|
||||
### Longer-term
|
||||
|
||||
- Mobile platform support
|
||||
- Bandwidth-aware routing and QoS
|
||||
- Protocol stability and versioned wire format
|
||||
- Published crate
|
||||
- Mobile platform support.
|
||||
- Bandwidth-aware routing and QoS.
|
||||
- Protocol stability and a versioned wire format.
|
||||
- Published crate.
|
||||
|
||||
## License
|
||||
|
||||
|
||||
229
RELEASE-NOTES.md
@@ -1,109 +1,136 @@
|
||||
# FIPS v0.2.1
|
||||
# FIPS v0.4.1
|
||||
|
||||
**Released**: 2026-05-11
|
||||
**Released**: 2026-07-19
|
||||
|
||||
v0.2.1 is a maintenance release on the v0.2.x line. No new features
|
||||
and no wire-format changes; operators running v0.2.0 can upgrade in
|
||||
place. The release rolls up bug fixes and operational hardening for
|
||||
issues surfaced in v0.2.0 deployments, plus a bloom-filter fill-ratio
|
||||
validation that protects mesh-size estimates from saturated-filter
|
||||
inputs.
|
||||
v0.4.1 is a maintenance release on the v0.4.x line. It raises the default
|
||||
antipoison cap on inbound bloom filter announcements, removes a redundant
|
||||
spanning-tree metric counter, fixes two convergence and path-MTU bugs, and
|
||||
cuts per-packet CPU in the bloom and identity paths. There is no wire
|
||||
format change and no new feature surface.
|
||||
|
||||
v0.4.1 is wire-compatible with v0.4.0. Nodes can be upgraded one at a time
|
||||
with no coordinated restart, though one behavior change below is worth
|
||||
reading before you start a rolling upgrade.
|
||||
|
||||
## At a glance
|
||||
|
||||
- 22 commits since v0.2.0, 5 committers plus 2 issue reporters.
|
||||
- All changes are backwards-compatible with v0.2.0 on the wire.
|
||||
- Bloom filter fill-ratio validation hardens the FilterAnnounce
|
||||
ingress path.
|
||||
- TreeAnnounce ancestry validation tightened to match the
|
||||
spanning-tree specification.
|
||||
- Signed-tarball + `.deb` artifact workflow added for tagged
|
||||
releases; AUR auto-publish on stable tags.
|
||||
- `node.bloom.max_inbound_fpr` default moves from `0.10` to `0.20`.
|
||||
- The `parent_switched` metric counter is gone. Use `parent_switches`.
|
||||
- Spanning tree no longer serves stale coordinates after a parent link is
|
||||
lost through peer removal.
|
||||
- Discovery no longer loosens a path MTU clamp it had correctly tightened.
|
||||
- Bloom probing and identity operations do measurably less work per call,
|
||||
with identical results.
|
||||
|
||||
## Behavior changes worth flagging
|
||||
|
||||
- **Bloom filter fill-ratio validation** runs on every inbound
|
||||
`FilterAnnounce`. Filters whose derived false-positive rate exceeds
|
||||
`node.bloom.max_inbound_fpr` (new config field, default `0.05`) are
|
||||
rejected silently on the wire, logged at WARN, and counted in a
|
||||
new `bloom.fill_exceeded` counter. A rate-limited WARN also fires
|
||||
when the local outgoing filter exceeds the cap.
|
||||
`BloomFilter::estimated_count` now takes `max_fpr` and returns
|
||||
`Option<f64>`, returning `None` for saturated filters; this
|
||||
propagates through `compute_mesh_size` into `estimated_mesh_size`.
|
||||
- **TreeAnnounce ancestry validation** is now run before tree-state
|
||||
mutation, enforcing ancestry-self-match, root-single-entry,
|
||||
parent-second-entry, and root-is-minimum-NodeAddr. Non-conforming
|
||||
announces are rejected with a WARN. Mixed v0.2.0 / v0.2.1 meshes
|
||||
may produce WARN log lines on the v0.2.1 side until all peers
|
||||
upgrade; behavior is correct, log noise only.
|
||||
### The inbound filter FPR cap default doubles again
|
||||
|
||||
`node.bloom.max_inbound_fpr` goes from `0.10` to `0.20`. The cap rejects
|
||||
inbound `FilterAnnounce` frames whose advertised false positive rate
|
||||
exceeds it. On the fixed 1 KB, k=5 filter, `0.10` corresponds to a fill of
|
||||
0.631 and roughly 1,630 reachable entries, and the busiest nodes'
|
||||
aggregates had started reaching that ceiling as the mesh grew. `0.20`
|
||||
corresponds to a fill of 0.7248 and roughly 2,114 entries.
|
||||
|
||||
Be aware that this is the second time in two releases that this default
|
||||
has doubled, for the same reason both times. That is worth stating plainly
|
||||
rather than repeating the previous release's framing: raising the cap buys
|
||||
headroom, it does not fix anything. The real constraint is the fixed 1 KB
|
||||
filter size, which is a protocol constant. The structural remedy is the v2
|
||||
filter work, where filter capacity scales with the mesh instead of being
|
||||
pinned. This release is an interim step to keep legitimate aggregates from
|
||||
being rejected until that lands. It is not the start of a pattern of
|
||||
raising the cap once per release, and if you are sizing capacity planning
|
||||
around this number, plan against the v2 work rather than against a third
|
||||
raise.
|
||||
|
||||
The antipoison property the cap exists for is preserved. A saturated or
|
||||
deliberately poisoned filter still presents an FPR near 100% and is still
|
||||
rejected.
|
||||
|
||||
**This matters during a rolling upgrade.** A v0.4.1 node accepts a
|
||||
`FilterAnnounce` with a derived FPR between 0.10 and 0.20; a v0.4.0 node
|
||||
drops the same frame, and the drop is silent on the wire with no NACK. The
|
||||
cap also gates the mesh size estimator, which declines to produce a value
|
||||
when any contributing filter is over the cap. So while a mesh is partly
|
||||
upgraded, upgraded and not-yet-upgraded nodes can legitimately report
|
||||
different mesh sizes, or one can report a size while the other reports
|
||||
unknown. This resolves once every node is on v0.4.1. If you want to avoid
|
||||
the window entirely, set `node.bloom.max_inbound_fpr: 0.10` explicitly in
|
||||
your config before upgrading and remove it after the last node is done.
|
||||
|
||||
### The `parent_switched` counter is removed
|
||||
|
||||
`parent_switched` was incremented on the line immediately before
|
||||
`parent_switches` at every site and never independently, so the two
|
||||
counters always held the same value. `parent_switched` is now gone from
|
||||
the tree metrics, the control socket snapshot, and the `fipstop` tree
|
||||
view. `parent_switches` remains and is unchanged.
|
||||
|
||||
If you scrape the control socket, or have dashboards or alerts referencing
|
||||
`parent_switched`, point them at `parent_switches`. Anything still asking
|
||||
for `parent_switched` will find nothing rather than a zero.
|
||||
|
||||
## Notable bug fixes
|
||||
|
||||
- **Control socket path detection** in `fipsctl` and `fipstop` now
|
||||
checks for the `/run/fips/` directory instead of the socket file
|
||||
inside it. Users not yet in the `fips` group get a clear
|
||||
"Permission denied" error instead of a misleading "No such file"
|
||||
fallback to `$XDG_RUNTIME_DIR`
|
||||
([#30](https://github.com/jmcorgan/fips/issues/30), reported by
|
||||
[@Sebastix](https://github.com/Sebastix)).
|
||||
- **`fd00::/8` routing protected from Tailscale interception.** The
|
||||
daemon installs an IPv6 routing-policy rule
|
||||
(`ip -6 rule to fd00::/8 lookup main priority 5265`) at TUN setup,
|
||||
so Tailscale's table 52 default route can no longer divert mesh
|
||||
traffic.
|
||||
- **Bloom filter routing greedy-tree fallback.** `find_next_hop` no
|
||||
longer returns `NoRoute` when the bloom candidate set is non-empty
|
||||
but no candidate is strictly closer than the current node; it
|
||||
falls through to greedy tree routing instead. Previously, this
|
||||
caused dropped packets in topologies where the tree parent was
|
||||
closer but not a bloom candidate.
|
||||
- **Auto-connect peers reconnect after a graceful Disconnect.**
|
||||
Previously, a clean upstream shutdown left the auto-connect peer
|
||||
orphaned; only the link-dead, decrypt-fail, and peer-restart paths
|
||||
scheduled a reconnect
|
||||
([#60](https://github.com/jmcorgan/fips/issues/60), reported by
|
||||
[@SwapMarket](https://github.com/SwapMarket)).
|
||||
- **`fipsctl connect` rejects FIPS mesh addresses** (`fd00::/8`) for
|
||||
`udp`, `tcp`, and `ethernet` transports with a clear error message
|
||||
instead of echoing success while the daemon silently failed the
|
||||
bind with `EAFNOSUPPORT`
|
||||
([#61](https://github.com/jmcorgan/fips/issues/61), reported by
|
||||
[@SwapMarket](https://github.com/SwapMarket)).
|
||||
- **OpenWrt ipk** cross-compiles cleanly again after excluding the
|
||||
BLE feature that requires D-Bus, which is unavailable on OpenWrt
|
||||
targets.
|
||||
### Stale coordinates after losing a parent through peer removal
|
||||
|
||||
## Packaging
|
||||
When a node's parent link dropped via peer removal, the node correctly
|
||||
reparented or self-rooted, but skipped the coordinate cache invalidation
|
||||
that every other position-change path performs. Cached entries for
|
||||
downstream destinations kept the node's old coordinate prefix. This did
|
||||
not self-correct the way a stale cache entry normally would: routing
|
||||
access refreshes an entry's TTL, so an entry that was actively being
|
||||
routed through never expired, and was only fixed by an unrelated fresh
|
||||
insert. Both invalidation classes now run on this path, matching the
|
||||
loop-detection branch.
|
||||
|
||||
- **Linux release artifact workflow** builds x86_64 and aarch64
|
||||
tarballs and `.deb` packages on `v*` tag push, with SHA-256
|
||||
checksums, and publishes them to the GitHub release page.
|
||||
- **AUR publish workflow** auto-publishes the `fips` PKGBUILD on
|
||||
stable `v*` tags.
|
||||
### Discovery could loosen a tightened path MTU clamp
|
||||
|
||||
An originator handling a `LookupResponse` overwrote its cached path MTU
|
||||
unconditionally. If a reactive `MtuExceeded` or `PathMtuNotification` had
|
||||
already taught it a tighter value, a later, looser discovery estimate
|
||||
would clobber that and re-loosen the clamp, risking a return to dropped
|
||||
oversized packets. The cached and received values are now compared and the
|
||||
tighter one is kept.
|
||||
|
||||
## Upgrade notes
|
||||
|
||||
Operator-actionable items when moving from v0.2.0 to v0.2.1:
|
||||
This is a drop-in upgrade from v0.4.0 with no wire format change, no
|
||||
config migration, and no coordinated restart. Upgrade nodes in whatever
|
||||
order you like.
|
||||
|
||||
- **Bloom filter fill-ratio cap (default 0.05).** Inbound
|
||||
`FilterAnnounce` messages whose derived FPR exceeds the cap are
|
||||
rejected silently on the wire. Operators with unusually saturated
|
||||
filters in the field may want to confirm that the default applies
|
||||
cleanly to their deployment; check the new `bloom.fill_exceeded`
|
||||
counter if mesh-size estimates drift after upgrade.
|
||||
- **TreeAnnounce ancestry tightening.** Mixed v0.2.0 / v0.2.1 meshes
|
||||
may produce WARN log lines on the v0.2.1 side until all peers
|
||||
upgrade. Behavior is correct, log noise only.
|
||||
Two things to do rather than assume:
|
||||
|
||||
## Getting v0.2.1
|
||||
1. If you monitor `parent_switched`, move to `parent_switches` before
|
||||
upgrading, or your dashboards will go blank rather than error.
|
||||
2. During the rolling window, expect upgraded and not-yet-upgraded nodes
|
||||
to potentially disagree about mesh size, per the FPR cap section above.
|
||||
This is expected and self-resolves. Do not chase it as a bug unless it
|
||||
persists after every node reports `0.4.1`.
|
||||
|
||||
If you have pinned `node.bloom.max_inbound_fpr` explicitly in your config,
|
||||
your setting is honored and nothing changes for you. The change only
|
||||
affects nodes taking the default.
|
||||
|
||||
Downgrading to v0.4.0 is supported and needs no special handling.
|
||||
|
||||
## Getting v0.4.1
|
||||
|
||||
- **Linux x86_64 / aarch64**: `.deb` and tarball at the
|
||||
[v0.2.1 release page](https://github.com/jmcorgan/fips/releases/tag/v0.2.1).
|
||||
[v0.4.1 release page](https://github.com/jmcorgan/fips/releases/tag/v0.4.1).
|
||||
- **Arch Linux**: `fips` from the AUR.
|
||||
- **OpenWrt**: `.ipk` at the v0.2.1 release page.
|
||||
- **From source**: `cargo build --release` from a checkout of the
|
||||
v0.2.1 tag.
|
||||
- **macOS**: `.pkg` at the v0.4.1 release page.
|
||||
- **Windows**: ZIP at the v0.4.1 release page.
|
||||
- **OpenWrt**: `.ipk` (OpenWrt 24.x and earlier) or `.apk` (OpenWrt 25+)
|
||||
at the v0.4.1 release page.
|
||||
- **From source**: `cargo build --release` from a checkout of the v0.4.1
|
||||
tag (Rust 1.94.1 per `rust-toolchain.toml`; `libclang-dev` is a
|
||||
required Linux build prerequisite).
|
||||
- **Nix / NixOS**: `nix build .#fips` from a checkout of the v0.4.1 tag
|
||||
builds the binaries from source with the pinned toolchain and no manual
|
||||
prerequisites (see the Nix section of `packaging/README.md`).
|
||||
|
||||
The full per-commit changelog lives in
|
||||
[`CHANGELOG.md`](../../CHANGELOG.md). Issues and discussion at
|
||||
@@ -111,31 +138,9 @@ The full per-commit changelog lives in
|
||||
|
||||
## Contributors
|
||||
|
||||
Thanks to everyone who contributed code or bug reports to this
|
||||
release.
|
||||
Thanks to everyone who contributed code, packaging work, bug reports, or
|
||||
reviews to this release.
|
||||
|
||||
**Code and packaging**:
|
||||
|
||||
- [@jcorgan](https://github.com/jmcorgan): release shepherd, bloom
|
||||
fill-ratio validation, auto-connect reconnect fix, `fipsctl`
|
||||
mesh-address rejection, control-socket path detection,
|
||||
Tailscale-vs-`fd00::/8` routing policy, bloom routing greedy
|
||||
fallback, rustfmt baseline.
|
||||
- [@Origami74](https://github.com/Origami74): OpenWrt ipk
|
||||
BLE-feature build fix.
|
||||
- [@jodobear](https://github.com/jodobear): Linux release-artifact
|
||||
workflow and target-aware build scripts.
|
||||
- [@dskvr](https://github.com/dskvr): AUR publish workflow.
|
||||
- [@SatsAndSports](https://github.com/SatsAndSports): TreeAnnounce
|
||||
semantic validation.
|
||||
|
||||
**Issue reports that drove fixes in this release**:
|
||||
|
||||
- [@Sebastix](https://github.com/Sebastix): `fipsctl` / `fipstop`
|
||||
control-socket path detection
|
||||
([#30](https://github.com/jmcorgan/fips/issues/30)).
|
||||
- [@SwapMarket](https://github.com/SwapMarket): auto-connect
|
||||
reconnect after graceful disconnect
|
||||
([#60](https://github.com/jmcorgan/fips/issues/60)) and
|
||||
`fipsctl` mesh-address rejection
|
||||
([#61](https://github.com/jmcorgan/fips/issues/61)).
|
||||
- [@jcorgan](https://github.com/jmcorgan): release shepherd, spanning-tree
|
||||
and discovery fixes, bloom and identity performance work, antipoison cap
|
||||
change, and testing.
|
||||
|
||||
365
benches/routing_next_hop.rs
Normal file
@@ -0,0 +1,365 @@
|
||||
//! Micro-benchmark quantifying the per-forwarded-packet heap-allocation cost
|
||||
//! of the routing next-hop candidate-assembly path.
|
||||
//!
|
||||
//! `find_next_hop` runs once per forwarded data packet. Its sans-IO core
|
||||
//! assembles a `Vec<Candidate>` by enumerating every peer through the
|
||||
//! `RoutingView` seam: `peer_addrs()` materializes a `Vec<NodeAddr>` of all
|
||||
//! peers, the survivors are snapshotted (each cloning its `TreeCoordinate`),
|
||||
//! and the result is collected into a second `Vec`. This bench measures that
|
||||
//! per-call allocation against a fused zero-alloc reference that iterates the
|
||||
//! peer map directly and borrows coordinates instead of cloning.
|
||||
//!
|
||||
//! Visibility caveat: the production `routing_candidates` / `select_best_candidate`
|
||||
//! / `RoutingView` / `Candidate` are `pub(crate)` (src/proto/routing/core.rs)
|
||||
//! and are not re-exported at the crate root, so an external bench crate cannot
|
||||
//! name them. Rather than change production visibility, this file reproduces
|
||||
//! that path verbatim over the real public `NodeAddr` / `TreeCoordinate` /
|
||||
//! `CoordEntry` / `BloomFilter` types with the same iterator chain and the same
|
||||
//! `HashMap`-backed view the shell uses (src/node/mod.rs NodeRoutingView). The
|
||||
//! allocation behavior is therefore identical to production by construction;
|
||||
//! only the symbol identity differs.
|
||||
|
||||
use std::alloc::{GlobalAlloc, Layout, System};
|
||||
use std::collections::HashMap;
|
||||
use std::hint::black_box;
|
||||
use std::sync::atomic::{AtomicUsize, Ordering};
|
||||
|
||||
use criterion::{BenchmarkId, Criterion, criterion_group, criterion_main};
|
||||
use fips::{BloomFilter, NodeAddr, TreeCoordinate};
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Counting global allocator: bumps a process-global counter on every heap
|
||||
// allocation operation (alloc / alloc_zeroed / realloc). Sampled tightly and
|
||||
// single-threaded in `report_allocs` so no unrelated allocations are captured.
|
||||
// ---------------------------------------------------------------------------
|
||||
struct CountingAlloc;
|
||||
|
||||
static ALLOCS: AtomicUsize = AtomicUsize::new(0);
|
||||
|
||||
unsafe impl GlobalAlloc for CountingAlloc {
|
||||
unsafe fn alloc(&self, layout: Layout) -> *mut u8 {
|
||||
ALLOCS.fetch_add(1, Ordering::Relaxed);
|
||||
unsafe { System.alloc(layout) }
|
||||
}
|
||||
unsafe fn dealloc(&self, ptr: *mut u8, layout: Layout) {
|
||||
unsafe { System.dealloc(ptr, layout) }
|
||||
}
|
||||
unsafe fn alloc_zeroed(&self, layout: Layout) -> *mut u8 {
|
||||
ALLOCS.fetch_add(1, Ordering::Relaxed);
|
||||
unsafe { System.alloc_zeroed(layout) }
|
||||
}
|
||||
unsafe fn realloc(&self, ptr: *mut u8, layout: Layout, new_size: usize) -> *mut u8 {
|
||||
ALLOCS.fetch_add(1, Ordering::Relaxed);
|
||||
unsafe { System.realloc(ptr, layout, new_size) }
|
||||
}
|
||||
}
|
||||
|
||||
#[global_allocator]
|
||||
static GLOBAL: CountingAlloc = CountingAlloc;
|
||||
|
||||
const PEER_COUNTS: [usize; 4] = [8, 32, 128, 256];
|
||||
/// Fraction of peers whose bloom filter reports the destination reachable.
|
||||
const REACH_NUMERATOR: usize = 1;
|
||||
const REACH_DENOMINATOR: usize = 2;
|
||||
/// Tree depth for synthetic coordinates (self..root), a realistic mesh depth.
|
||||
const COORD_DEPTH: usize = 8;
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Reproduction of the pub(crate) routing seam (src/proto/routing/core.rs).
|
||||
// ---------------------------------------------------------------------------
|
||||
trait RoutingView {
|
||||
fn peer_addrs(&self) -> Vec<NodeAddr>;
|
||||
fn peer_may_reach(&self, peer: &NodeAddr, dest: &NodeAddr) -> bool;
|
||||
fn peer_can_send(&self, peer: &NodeAddr) -> bool;
|
||||
fn peer_link_cost(&self, peer: &NodeAddr) -> f64;
|
||||
fn peer_coords(&self, peer: &NodeAddr) -> Option<TreeCoordinate>;
|
||||
}
|
||||
|
||||
struct Candidate {
|
||||
addr: NodeAddr,
|
||||
can_send: bool,
|
||||
link_cost: f64,
|
||||
coords: Option<TreeCoordinate>,
|
||||
}
|
||||
|
||||
/// Verbatim from `routing::routing_candidates` (core.rs). Allocates the
|
||||
/// `peer_addrs` Vec, clones each survivor's coords, and collects into a Vec.
|
||||
fn routing_candidates(rv: &impl RoutingView, dest: &NodeAddr) -> Vec<Candidate> {
|
||||
rv.peer_addrs()
|
||||
.into_iter()
|
||||
.filter(|peer| rv.peer_may_reach(peer, dest))
|
||||
.map(|peer| Candidate {
|
||||
can_send: rv.peer_can_send(&peer),
|
||||
link_cost: rv.peer_link_cost(&peer),
|
||||
coords: rv.peer_coords(&peer),
|
||||
addr: peer,
|
||||
})
|
||||
.collect()
|
||||
}
|
||||
|
||||
/// Verbatim from `routing::select_best_candidate` (core.rs). Pure, no alloc.
|
||||
fn select_best_candidate(
|
||||
candidates: &[Candidate],
|
||||
dest_coords: &TreeCoordinate,
|
||||
my_coords: &TreeCoordinate,
|
||||
) -> Option<NodeAddr> {
|
||||
let my_distance = my_coords.distance_to(dest_coords);
|
||||
let mut best: Option<(&Candidate, f64, usize)> = None;
|
||||
for candidate in candidates {
|
||||
if !candidate.can_send {
|
||||
continue;
|
||||
}
|
||||
let cost = candidate.link_cost;
|
||||
let dist = candidate
|
||||
.coords
|
||||
.as_ref()
|
||||
.map(|pc| pc.distance_to(dest_coords))
|
||||
.unwrap_or(usize::MAX);
|
||||
if dist >= my_distance {
|
||||
continue;
|
||||
}
|
||||
let dominated = match &best {
|
||||
None => true,
|
||||
Some((_, best_cost, best_dist)) => {
|
||||
cost < *best_cost
|
||||
|| (cost == *best_cost && dist < *best_dist)
|
||||
|| (cost == *best_cost
|
||||
&& dist == *best_dist
|
||||
&& candidate.addr < best.as_ref().unwrap().0.addr)
|
||||
}
|
||||
};
|
||||
if dominated {
|
||||
best = Some((candidate, cost, dist));
|
||||
}
|
||||
}
|
||||
best.map(|(candidate, _, _)| candidate.addr)
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Bench-local view, HashMap-backed exactly like src/node/mod.rs NodeRoutingView.
|
||||
// ---------------------------------------------------------------------------
|
||||
struct BenchPeer {
|
||||
bloom: BloomFilter,
|
||||
can_send: bool,
|
||||
link_cost: f64,
|
||||
}
|
||||
|
||||
struct BenchView {
|
||||
peers: HashMap<NodeAddr, BenchPeer>,
|
||||
coords: HashMap<NodeAddr, TreeCoordinate>,
|
||||
}
|
||||
|
||||
impl RoutingView for BenchView {
|
||||
fn peer_addrs(&self) -> Vec<NodeAddr> {
|
||||
self.peers.keys().copied().collect()
|
||||
}
|
||||
fn peer_may_reach(&self, peer: &NodeAddr, dest: &NodeAddr) -> bool {
|
||||
self.peers.get(peer).is_some_and(|p| p.bloom.contains(dest))
|
||||
}
|
||||
fn peer_can_send(&self, peer: &NodeAddr) -> bool {
|
||||
self.peers.get(peer).is_some_and(|p| p.can_send)
|
||||
}
|
||||
fn peer_link_cost(&self, peer: &NodeAddr) -> f64 {
|
||||
self.peers.get(peer).map_or(f64::INFINITY, |p| p.link_cost)
|
||||
}
|
||||
fn peer_coords(&self, peer: &NodeAddr) -> Option<TreeCoordinate> {
|
||||
self.coords.get(peer).cloned()
|
||||
}
|
||||
}
|
||||
|
||||
/// Zero-alloc reference: what an iterator/visitor seam would do. Iterates the
|
||||
/// peer map directly, fuses the may_reach + can_send filters, borrows coords
|
||||
/// instead of cloning, and tracks the best hop inline. No Vec, no coord clone.
|
||||
fn resolve_next_hop_zeroalloc(
|
||||
view: &BenchView,
|
||||
dest: &NodeAddr,
|
||||
dest_coords: &TreeCoordinate,
|
||||
my_coords: &TreeCoordinate,
|
||||
) -> Option<NodeAddr> {
|
||||
let my_distance = my_coords.distance_to(dest_coords);
|
||||
let mut best: Option<(NodeAddr, f64, usize)> = None;
|
||||
for (addr, peer) in &view.peers {
|
||||
if !peer.bloom.contains(dest) {
|
||||
continue;
|
||||
}
|
||||
if !peer.can_send {
|
||||
continue;
|
||||
}
|
||||
let cost = peer.link_cost;
|
||||
let dist = view
|
||||
.coords
|
||||
.get(addr)
|
||||
.map(|pc| pc.distance_to(dest_coords))
|
||||
.unwrap_or(usize::MAX);
|
||||
if dist >= my_distance {
|
||||
continue;
|
||||
}
|
||||
let dominated = match &best {
|
||||
None => true,
|
||||
Some((best_addr, best_cost, best_dist)) => {
|
||||
cost < *best_cost
|
||||
|| (cost == *best_cost && dist < *best_dist)
|
||||
|| (cost == *best_cost && dist == *best_dist && *addr < *best_addr)
|
||||
}
|
||||
};
|
||||
if dominated {
|
||||
best = Some((*addr, cost, dist));
|
||||
}
|
||||
}
|
||||
best.map(|(addr, _, _)| addr)
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Scenario construction.
|
||||
// ---------------------------------------------------------------------------
|
||||
fn addr(tag: u8, i: u16) -> NodeAddr {
|
||||
let mut b = [0u8; 16];
|
||||
b[0] = tag;
|
||||
b[1..3].copy_from_slice(&i.to_le_bytes());
|
||||
NodeAddr::from_bytes(b)
|
||||
}
|
||||
|
||||
/// A depth-`COORD_DEPTH` coordinate whose leaf is `leaf`, sharing a fixed
|
||||
/// interior path and root with `shared_tag`. Peers built with the dest's
|
||||
/// shared_tag sit close to the destination (distance 2); a distinct shared_tag
|
||||
/// sits far (near the root), modeling our own position.
|
||||
fn coord(leaf: NodeAddr, shared_tag: u8) -> TreeCoordinate {
|
||||
let mut path = Vec::with_capacity(COORD_DEPTH);
|
||||
path.push(leaf);
|
||||
for level in 1..(COORD_DEPTH - 1) {
|
||||
path.push(addr(shared_tag, level as u16));
|
||||
}
|
||||
path.push(addr(9, 0)); // common root
|
||||
TreeCoordinate::from_addrs(path).expect("valid coord path")
|
||||
}
|
||||
|
||||
struct Scenario {
|
||||
view: BenchView,
|
||||
dest: NodeAddr,
|
||||
dest_coords: TreeCoordinate,
|
||||
my_coords: TreeCoordinate,
|
||||
}
|
||||
|
||||
impl Scenario {
|
||||
fn new(n: usize) -> Self {
|
||||
let dest = addr(2, 0);
|
||||
// Destination path uses interior tag 4; peers reuse tag 4 so survivors
|
||||
// are close to the destination. Our own coords use tag 5 (far).
|
||||
let dest_coords = coord(dest, 4);
|
||||
let my_coords = coord(addr(6, 0), 5);
|
||||
|
||||
let mut peers = HashMap::new();
|
||||
let mut coords = HashMap::new();
|
||||
for i in 0..n {
|
||||
let paddr = addr(1, i as u16);
|
||||
let mut bloom = BloomFilter::new();
|
||||
// Realistic fill: a handful of unrelated reachable addrs.
|
||||
for f in 0..4u16 {
|
||||
bloom.insert(&addr(7, i as u16 * 4 + f));
|
||||
}
|
||||
// A controlled fraction advertise the destination as reachable.
|
||||
if (i % REACH_DENOMINATOR) < REACH_NUMERATOR {
|
||||
bloom.insert(&dest);
|
||||
}
|
||||
peers.insert(
|
||||
paddr,
|
||||
BenchPeer {
|
||||
bloom,
|
||||
can_send: true,
|
||||
link_cost: 1.0 + (i as f64) * 0.01,
|
||||
},
|
||||
);
|
||||
// Peers share the destination's interior path (tag 4) → close.
|
||||
coords.insert(paddr, coord(paddr, 4));
|
||||
}
|
||||
|
||||
Self {
|
||||
view: BenchView { peers, coords },
|
||||
dest,
|
||||
dest_coords,
|
||||
my_coords,
|
||||
}
|
||||
}
|
||||
|
||||
fn survivors(&self) -> usize {
|
||||
self.view
|
||||
.peers
|
||||
.values()
|
||||
.filter(|p| p.bloom.contains(&self.dest))
|
||||
.count()
|
||||
}
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Allocation-per-call report (printed once, before criterion timing).
|
||||
// ---------------------------------------------------------------------------
|
||||
fn count_allocs<T>(iters: usize, mut f: impl FnMut() -> T) -> f64 {
|
||||
for _ in 0..8 {
|
||||
black_box(f());
|
||||
}
|
||||
let start = ALLOCS.load(Ordering::Relaxed);
|
||||
for _ in 0..iters {
|
||||
black_box(f());
|
||||
}
|
||||
let end = ALLOCS.load(Ordering::Relaxed);
|
||||
(end - start) as f64 / iters as f64
|
||||
}
|
||||
|
||||
fn report_allocs() {
|
||||
const ITERS: usize = 2000;
|
||||
println!("\n=== allocations per call (heap alloc ops: alloc+alloc_zeroed+realloc) ===");
|
||||
println!(
|
||||
"{:>6} {:>10} {:>16} {:>16}",
|
||||
"peers", "survivors", "current/call", "zero-alloc/call"
|
||||
);
|
||||
for &n in &PEER_COUNTS {
|
||||
let s = Scenario::new(n);
|
||||
let survivors = s.survivors();
|
||||
let current = count_allocs(ITERS, || {
|
||||
let cands = routing_candidates(&s.view, &s.dest);
|
||||
select_best_candidate(&cands, &s.dest_coords, &s.my_coords)
|
||||
});
|
||||
let zero = count_allocs(ITERS, || {
|
||||
resolve_next_hop_zeroalloc(&s.view, &s.dest, &s.dest_coords, &s.my_coords)
|
||||
});
|
||||
println!("{n:>6} {survivors:>10} {current:>16.2} {zero:>16.2}");
|
||||
}
|
||||
println!();
|
||||
}
|
||||
|
||||
fn bench_next_hop(c: &mut Criterion) {
|
||||
report_allocs();
|
||||
|
||||
let mut group = c.benchmark_group("find_next_hop");
|
||||
for &n in &PEER_COUNTS {
|
||||
let scenario = Scenario::new(n);
|
||||
group.bench_with_input(BenchmarkId::new("current_alloc", n), &n, |b, _| {
|
||||
b.iter(|| {
|
||||
let cands = routing_candidates(&scenario.view, &scenario.dest);
|
||||
black_box(select_best_candidate(
|
||||
&cands,
|
||||
&scenario.dest_coords,
|
||||
&scenario.my_coords,
|
||||
))
|
||||
});
|
||||
});
|
||||
group.bench_with_input(BenchmarkId::new("zero_alloc_ref", n), &n, |b, _| {
|
||||
b.iter(|| {
|
||||
black_box(resolve_next_hop_zeroalloc(
|
||||
&scenario.view,
|
||||
&scenario.dest,
|
||||
&scenario.dest_coords,
|
||||
&scenario.my_coords,
|
||||
))
|
||||
});
|
||||
});
|
||||
}
|
||||
group.finish();
|
||||
}
|
||||
|
||||
criterion_group! {
|
||||
name = benches;
|
||||
config = Criterion::default().sample_size(50);
|
||||
targets = bench_next_hop
|
||||
}
|
||||
criterion_main!(benches);
|
||||
10
build.rs
@@ -36,4 +36,14 @@ fn main() {
|
||||
|
||||
// Support reproducible builds (Debian packaging)
|
||||
println!("cargo:rerun-if-env-changed=SOURCE_DATE_EPOCH");
|
||||
|
||||
// bluer/BlueZ is glibc-linux only: musl cross-compiles (OpenWrt) can't
|
||||
// satisfy libdbus-sys's pkg-config cross-compile requirement, and musl
|
||||
// router targets don't run BlueZ by default anyway.
|
||||
println!("cargo:rustc-check-cfg=cfg(bluer_available)");
|
||||
let target_os = std::env::var("CARGO_CFG_TARGET_OS").unwrap_or_default();
|
||||
let target_env = std::env::var("CARGO_CFG_TARGET_ENV").unwrap_or_default();
|
||||
if target_os == "linux" && target_env != "musl" {
|
||||
println!("cargo:rustc-cfg=bluer_available");
|
||||
}
|
||||
}
|
||||
|
||||
@@ -1,5 +1,54 @@
|
||||
# FIPS Documentation
|
||||
|
||||
| Directory | Description |
|
||||
|-----------|-------------|
|
||||
| [design/](design/) | Protocol design specifications and analysis |
|
||||
FIPS (Free Internetworking Peering System) is a self-organizing
|
||||
encrypted mesh network built on Nostr identities, capable of
|
||||
operating over arbitrary transports — local networks, the public
|
||||
internet, Tor, Bluetooth, or point-to-point links — without central
|
||||
infrastructure.
|
||||
|
||||
With FIPS, your machine becomes a node in the mesh with a
|
||||
self-generated cryptographic identity. There are two ways to
|
||||
deploy it.
|
||||
|
||||
**As an overlay** on top of existing IP networks, FIPS lets
|
||||
your node reach any other FIPS node wherever it sits — behind a NAT, on a
|
||||
different ISP, on a phone over cellular, on a laptop with only
|
||||
Bluetooth in range, or behind a Tor onion. The mesh forwards
|
||||
IPv6 traffic transparently and end-to-end encrypted, with no
|
||||
central VPN concentrator or coordinating server.
|
||||
|
||||
**From the ground up** over raw Ethernet, WiFi, or Bluetooth,
|
||||
FIPS provides a complete permissionless network
|
||||
without any pre-existing IP infrastructure, ISP, or DNS. Any
|
||||
node that joins the link gets routable IPv6 addresses, peer
|
||||
discovery, and a path to every other node automatically.
|
||||
|
||||
Either way, existing networking software runs over it unchanged:
|
||||
SSH, HTTP servers, file transfer, anything IPv6-native works the
|
||||
same way it would on a local network.
|
||||
|
||||
New to FIPS? Start with the [Getting Started](getting-started.md)
|
||||
guide.
|
||||
|
||||
## Documentation Sections
|
||||
|
||||
### [Tutorials](tutorials/)
|
||||
|
||||
If you are starting from scratch and want a guided path to a
|
||||
working mesh, go here.
|
||||
|
||||
### [How-To Guides](how-to/)
|
||||
|
||||
If you have a specific task in mind — enabling a feature,
|
||||
deploying a component, diagnosing a problem — go here.
|
||||
|
||||
### [Reference](reference/)
|
||||
|
||||
If you need to look up wire formats, configuration keys, command
|
||||
flags, or counter inventories, go here.
|
||||
|
||||
### [Design](design/)
|
||||
|
||||
If you want to understand how the mesh self-organizes, why FIPS
|
||||
makes the choices it does, or how the pieces fit together, go
|
||||
here.
|
||||
|
||||
143
docs/branching.md
Normal file
@@ -0,0 +1,143 @@
|
||||
# FIPS Branching and Merging Strategy
|
||||
|
||||
<!-- markdownlint-disable MD013 -->
|
||||
|
||||
This document explains how the three long-lived branches relate, when
|
||||
to target each one, and how merges propagate fixes and features. For
|
||||
the day-to-day "how do I send a PR" workflow, see
|
||||
[CONTRIBUTING.md](../CONTRIBUTING.md).
|
||||
|
||||
## Branch Structure
|
||||
|
||||
Three long-lived branches track parallel development streams:
|
||||
|
||||
```text
|
||||
next ──●──●──●──●──●──────────────●──●── (wire-format-breaking work)
|
||||
\ /
|
||||
master ────●──●──●──●──●──●──────●──●──●── (compatible features, latest release line)
|
||||
\ /
|
||||
maint ────────●──●──●──●──●────────────── (bug fixes for the latest release)
|
||||
```
|
||||
|
||||
### maint
|
||||
|
||||
- Reset to each minor release tag at release time
|
||||
- Accepts only bug fixes for functionality that shipped in the
|
||||
latest release
|
||||
- No new features, no API changes, no wire-format changes
|
||||
- Patch releases tag from here (e.g., `v0.3.1`, `v0.3.2`)
|
||||
- Periodically merged forward into `master` so fixes propagate
|
||||
|
||||
### master
|
||||
|
||||
- Compatible development for the next feature release
|
||||
- Multiple feature releases may ship from master (`v0.4.0`, `v0.5.0`)
|
||||
before `next` promotes
|
||||
- No wire-format breaking changes; no API breaks
|
||||
- Receives merges from `maint` so released-line fixes flow forward
|
||||
- Periodically merged forward into `next`
|
||||
|
||||
### next
|
||||
|
||||
- Accumulates work that breaks wire format, API, or compatibility
|
||||
- Receives merges from `master` so it stays current with bug fixes
|
||||
and compatible feature work
|
||||
- Cargo version on `next` is the expected release version with a
|
||||
`-dev` suffix, updated if `master` ships additional minor
|
||||
releases first
|
||||
- Becomes the new `master` at the next breaking release; at the same
|
||||
point the old `master` becomes the new `maint`
|
||||
|
||||
## Versioning
|
||||
|
||||
While the project is in the `0.x` era, semver treats minor bumps as
|
||||
potentially breaking. Both `master` and `next` bump the minor version;
|
||||
the distinction between compatible and breaking is captured in the
|
||||
changelog and in which branch the work landed on.
|
||||
|
||||
The `-dev` suffix in `Cargo.toml` indicates an unreleased development
|
||||
state on the branch.
|
||||
|
||||
## Merge Direction
|
||||
|
||||
Fixes and features flow in **one direction only**: `maint → master → next`.
|
||||
Never merge backward (`next` into `master`, or `master` into `maint`).
|
||||
|
||||
```text
|
||||
maint ──→ master ──→ next
|
||||
```
|
||||
|
||||
This guarantees:
|
||||
|
||||
- Bug fixes shipped in a release reach all subsequent branches
|
||||
- Compatible features reach `next`
|
||||
- Wire-format-breaking work stays isolated on `next` until release
|
||||
|
||||
If you submit a PR on `next` that should also be on master or maint
|
||||
(rare, since the criteria for needing it on multiple branches are
|
||||
usually mutually exclusive), the PR stays on its target; the
|
||||
maintainer either backports as a separate commit on the upstream
|
||||
branch or asks you to.
|
||||
|
||||
## Choosing a Branch for Your PR
|
||||
|
||||
Pick the branch that matches the scope of your change:
|
||||
|
||||
| Your change | Target branch | Why |
|
||||
| --- | --- | --- |
|
||||
| Bug fix in a feature that shipped in the latest release | `maint` | Fix forward-merges to `master` and `next` |
|
||||
| Bug fix in code added on `master` since the last release (not in any released version) | `master` | The released v0.x.y line is unaffected, so `maint` does not need the change |
|
||||
| Bug fix in code added on `next` (wire-format-breaking work) | `next` | The bug only exists where the breaking work exists |
|
||||
| New feature that does not break wire format or API | `master` | Becomes part of the next compatible release |
|
||||
| Wire-format breaking change, API break, or fundamental protocol shape change | `next` | Stays isolated until the next forklift release |
|
||||
| Documentation, CI, or contributor-facing changes | `maint` if they apply to released material, else `master` | Forward-merges propagate naturally |
|
||||
|
||||
If you are not sure, ask in the related issue. The safest defaults
|
||||
are `master` for new features and `maint` for bug fixes; the
|
||||
maintainer will retarget the PR if needed.
|
||||
|
||||
## Release Workflow
|
||||
|
||||
### Bug fix release (from `maint`)
|
||||
|
||||
1. Fix on `maint`
|
||||
2. Bump patch version, tag (e.g., `v0.3.1`)
|
||||
3. Merge `maint` into `master`
|
||||
4. Merge `master` into `next`
|
||||
|
||||
### Compatible feature release (from `master`)
|
||||
|
||||
1. Finalize features on `master`
|
||||
2. Merge `maint` into `master` to pick up any pending fixes
|
||||
3. Set version, tag (e.g., `v0.4.0`)
|
||||
4. Reset `maint` to the new tag
|
||||
5. Bump `master` to the next `-dev` version
|
||||
6. Merge `master` into `next`
|
||||
|
||||
### Breaking release (from `next`)
|
||||
|
||||
1. Finalize features on `next`
|
||||
2. Merge `master` into `next` to pick up pending fixes and features
|
||||
3. Assign version as the next minor after `master`'s last release, tag
|
||||
4. `master` becomes the new `maint`
|
||||
5. `next` becomes the new `master`
|
||||
6. Create a new `next` branch from `master`
|
||||
|
||||
## Practical Guidelines
|
||||
|
||||
- **Commit to the appropriate branch for the scope of the change.**
|
||||
Do not commit bug fixes to `master` when they apply to the latest
|
||||
release — put them on `maint` and let the forward-merge propagate.
|
||||
- **Feature branches base off the long-lived branch they target.**
|
||||
Create with `git checkout -b my-feature maint` (or `master` or
|
||||
`next`), not `git checkout -b my-feature origin/maint`. The
|
||||
`origin/`-prefixed form auto-sets the new branch's upstream to
|
||||
the source ref, which can cause `git push` to land on the wrong
|
||||
ref under some configurations.
|
||||
- **When in doubt about whether a change is compatible**, target
|
||||
`next`. The maintainer can advise on retargeting.
|
||||
- **Resolve merge conflicts on the receiving branch**, preserving
|
||||
both the inherited fix and the new development.
|
||||
- **PRs are merged via squash-merge.** One logical change per PR
|
||||
becomes one commit on the destination branch, making bisect
|
||||
clean across the integration suite.
|
||||
@@ -1,51 +1,65 @@
|
||||
# FIPS Design Documents
|
||||
# FIPS Design
|
||||
|
||||
Protocol design specifications for the Federated Interoperable Peering
|
||||
System — a self-organizing encrypted mesh network built on Nostr identities.
|
||||
Architectural and protocol-level explanations for FIPS — the *why*
|
||||
and the *how* behind the wire and the system. For wire formats and
|
||||
configuration keys, see [reference/](../reference/). For task
|
||||
recipes, see [how-to/](../how-to/). For end-to-end lessons, see
|
||||
[tutorials/](../tutorials/).
|
||||
|
||||
## Reading Order
|
||||
|
||||
Start with the introduction, then follow the protocol stack from bottom to
|
||||
top. After the stack, the mesh operation document explains how all the
|
||||
pieces work together. Supporting references provide deeper dives into
|
||||
specific topics.
|
||||
Start with [fips-concepts.md](fips-concepts.md) for the
|
||||
novice-friendly framing of what FIPS is and why, then move to
|
||||
[fips-architecture.md](fips-architecture.md) for the protocol stack,
|
||||
identity model, and two-layer encryption walkthrough. From there,
|
||||
follow the protocol stack from bottom to top. After the stack,
|
||||
[fips-mesh-operation.md](fips-mesh-operation.md) explains how the
|
||||
pieces work together at runtime. Cross-cutting and supporting
|
||||
documents cover specific subsystems in detail.
|
||||
|
||||
### Foundations
|
||||
|
||||
| Document | Description |
|
||||
| -------- | ----------- |
|
||||
| [fips-concepts.md](fips-concepts.md) | What FIPS is, why it exists, mental model |
|
||||
| [fips-architecture.md](fips-architecture.md) | Protocol stack, identity, two-layer encryption |
|
||||
| [fips-prior-work.md](fips-prior-work.md) | Designs and protocols FIPS builds on |
|
||||
|
||||
### Protocol Stack
|
||||
|
||||
| Document | Description |
|
||||
| -------- | ----------- |
|
||||
| [fips-intro.md](fips-intro.md) | Protocol introduction: goals, architecture, layer model |
|
||||
| [fips-transport-layer.md](fips-transport-layer.md) | Transport layer: datagram delivery over arbitrary media |
|
||||
| [fips-mesh-layer.md](fips-mesh-layer.md) | FIPS Mesh Protocol (FMP): peer authentication, link encryption, forwarding |
|
||||
| [fips-session-layer.md](fips-session-layer.md) | FIPS Session Protocol (FSP): end-to-end encryption, sessions |
|
||||
| [fips-ipv6-adapter.md](fips-ipv6-adapter.md) | IPv6 adaptation: TUN interface, DNS, MTU enforcement |
|
||||
|
||||
### Cross-Cutting
|
||||
|
||||
| Document | Description |
|
||||
| -------- | ----------- |
|
||||
| [fips-mmp.md](fips-mmp.md) | Metrics Measurement Protocol (link + session) |
|
||||
| [fips-mtu.md](fips-mtu.md) | Path MTU model, encapsulation overhead, PMTUD |
|
||||
| [fips-security.md](fips-security.md) | `fips0` interface threat model and default-deny baseline |
|
||||
|
||||
### Mesh Behavior
|
||||
|
||||
| Document | Description |
|
||||
| -------- | ----------- |
|
||||
| [fips-mesh-operation.md](fips-mesh-operation.md) | How the mesh operates: routing, discovery, error recovery |
|
||||
| [fips-wire-formats.md](fips-wire-formats.md) | Wire format reference for all message types |
|
||||
| [fips-nostr-discovery.md](fips-nostr-discovery.md) | Optional Nostr-mediated peer discovery and UDP NAT hole-punch |
|
||||
| [port-advertisement-and-nat-traversal.md](port-advertisement-and-nat-traversal.md) | Nostr-signaled port advertisement and UDP NAT-traversal protocol; generic, with FIPS as an example implementation |
|
||||
|
||||
### Supporting References
|
||||
### Deeper Dives
|
||||
|
||||
| Document | Description |
|
||||
| -------- | ----------- |
|
||||
| [fips-spanning-tree.md](fips-spanning-tree.md) | Spanning tree algorithms: root discovery, parent selection, coordinates |
|
||||
| [fips-bloom-filters.md](fips-bloom-filters.md) | Bloom filter math: FPR analysis, size classes, split-horizon |
|
||||
|
||||
### Implementation
|
||||
|
||||
| Document | Description |
|
||||
| -------- | ----------- |
|
||||
| [fips-configuration.md](fips-configuration.md) | YAML configuration reference |
|
||||
|
||||
### Supplemental
|
||||
|
||||
| Document | Description |
|
||||
| -------- | ----------- |
|
||||
| [fips-bloom-filters.md](fips-bloom-filters.md) | Bloom filter properties: FPR analysis, size classes, split-horizon |
|
||||
| [spanning-tree-dynamics.md](spanning-tree-dynamics.md) | Spanning tree walkthroughs: convergence scenarios, worked examples |
|
||||
|
||||
## Document Relationships
|
||||
### Adjacent Components
|
||||
|
||||

|
||||
| Document | Description |
|
||||
| -------- | ----------- |
|
||||
| [fips-gateway.md](fips-gateway.md) | `fips-gateway` service: outbound (LAN-to-mesh) DNS-proxy + virtual-IP NAT and inbound (mesh-to-LAN) port-forwarding, sharing one nftables table |
|
||||
|
||||
@@ -82,7 +82,7 @@
|
||||
<text x="610" y="223" text-anchor="middle" class="alabel" fill="#5080c0">LookupRequest</text>
|
||||
|
||||
<!-- Bloom filter annotation -->
|
||||
<text x="430" y="260" text-anchor="middle" class="annot">guided by bloom filters at each hop</text>
|
||||
<text x="430" y="260" text-anchor="middle" class="annot">guided by bloom filters at each hop; transits do not cache</text>
|
||||
|
||||
<!-- ═══ Phase separator ═══ -->
|
||||
<line x1="70" y1="285" x2="830" y2="285" class="sep"/>
|
||||
@@ -99,23 +99,19 @@
|
||||
<polygon points="534,335 524,340 534,345" class="resp"/>
|
||||
<text x="610" y="333" text-anchor="middle" class="alabel" fill="#d0a040">LookupResponse + coords</text>
|
||||
|
||||
<!-- Cache box at C -->
|
||||
<rect x="490" y="358" width="60" height="22" class="cache"/>
|
||||
<text x="520" y="373" text-anchor="middle" class="clabel">cache D</text>
|
||||
<!-- C → B LookupResponse (transit forwards without caching) -->
|
||||
<line x1="510" y1="380" x2="352" y2="380" class="resp"/>
|
||||
<polygon points="354,375 344,380 354,385" class="resp"/>
|
||||
<text x="430" y="373" text-anchor="middle" class="alabel" fill="#d0a040">LookupResponse + coords</text>
|
||||
|
||||
<!-- C → B LookupResponse -->
|
||||
<line x1="510" y1="395" x2="352" y2="395" class="resp"/>
|
||||
<polygon points="354,390 344,395 354,400" class="resp"/>
|
||||
<text x="430" y="388" text-anchor="middle" class="alabel" fill="#d0a040">LookupResponse + coords</text>
|
||||
<!-- B → A LookupResponse (transit forwards without caching) -->
|
||||
<line x1="330" y1="420" x2="172" y2="420" class="resp"/>
|
||||
<polygon points="174,415 164,420 174,425" class="resp"/>
|
||||
<text x="250" y="413" text-anchor="middle" class="alabel" fill="#d0a040">LookupResponse + coords</text>
|
||||
|
||||
<!-- Cache box at B -->
|
||||
<rect x="310" y="413" width="60" height="22" class="cache"/>
|
||||
<text x="340" y="428" text-anchor="middle" class="clabel">cache D</text>
|
||||
|
||||
<!-- B → A LookupResponse -->
|
||||
<line x1="330" y1="448" x2="172" y2="448" class="resp"/>
|
||||
<polygon points="174,443 164,448 174,453" class="resp"/>
|
||||
<text x="250" y="441" text-anchor="middle" class="alabel" fill="#d0a040">LookupResponse + coords</text>
|
||||
<!-- Cache box at A — only the originator caches on LookupResponse -->
|
||||
<rect x="130" y="438" width="60" height="22" class="cache"/>
|
||||
<text x="160" y="453" text-anchor="middle" class="clabel">cache D</text>
|
||||
|
||||
<!-- ═══ Phase separator ═══ -->
|
||||
<line x1="70" y1="475" x2="830" y2="475" class="sep"/>
|
||||
@@ -125,32 +121,34 @@
|
||||
<!-- ═══════════════════════════════════════════════ -->
|
||||
|
||||
<text x="40" y="508" text-anchor="middle" class="plabel">Phase 3</text>
|
||||
<text x="40" y="524" text-anchor="middle" class="annot">Routing</text>
|
||||
<text x="40" y="524" text-anchor="middle" class="annot">Data flow</text>
|
||||
|
||||
<!-- A → B Data -->
|
||||
<!-- A → B Data with coords (CP flag / SessionSetup) -->
|
||||
<line x1="170" y1="530" x2="328" y2="530" class="data"/>
|
||||
<polygon points="326,525 336,530 326,535" class="data"/>
|
||||
<text x="250" y="523" text-anchor="middle" class="alabel" fill="#40a060">Data</text>
|
||||
<text x="250" y="523" text-anchor="middle" class="alabel" fill="#40a060">Data + coords</text>
|
||||
|
||||
<!-- Cached coords note at B -->
|
||||
<text x="340" y="553" text-anchor="middle" class="annot">cached coords</text>
|
||||
<!-- Cache box at B — warmed in-band from CP-flagged data -->
|
||||
<rect x="310" y="540" width="60" height="22" class="cache"/>
|
||||
<text x="340" y="555" text-anchor="middle" class="clabel">cache D</text>
|
||||
|
||||
<!-- B → C Data -->
|
||||
<line x1="350" y1="565" x2="508" y2="565" class="data"/>
|
||||
<polygon points="506,560 516,565 506,570" class="data"/>
|
||||
<text x="430" y="558" text-anchor="middle" class="alabel" fill="#40a060">Data</text>
|
||||
<!-- B → C Data with coords -->
|
||||
<line x1="350" y1="572" x2="508" y2="572" class="data"/>
|
||||
<polygon points="506,567 516,572 506,577" class="data"/>
|
||||
<text x="430" y="565" text-anchor="middle" class="alabel" fill="#40a060">Data + coords</text>
|
||||
|
||||
<!-- Cached coords note at C -->
|
||||
<text x="520" y="588" text-anchor="middle" class="annot">cached coords</text>
|
||||
<!-- Cache box at C — warmed in-band -->
|
||||
<rect x="490" y="582" width="60" height="22" class="cache"/>
|
||||
<text x="520" y="597" text-anchor="middle" class="clabel">cache D</text>
|
||||
|
||||
<!-- C → D Data -->
|
||||
<line x1="530" y1="600" x2="688" y2="600" class="data"/>
|
||||
<polygon points="686,595 696,600 686,605" class="data"/>
|
||||
<text x="610" y="593" text-anchor="middle" class="alabel" fill="#40a060">Data</text>
|
||||
<line x1="530" y1="614" x2="688" y2="614" class="data"/>
|
||||
<polygon points="686,609 696,614 686,619" class="data"/>
|
||||
<text x="610" y="607" text-anchor="middle" class="alabel" fill="#40a060">Data + coords</text>
|
||||
|
||||
<!-- Efficient forwarding annotation -->
|
||||
<text x="430" y="622" text-anchor="middle" class="annot">cached coords enable efficient forwarding — no re-discovery needed</text>
|
||||
<text x="430" y="636" text-anchor="middle" class="annot">transits cache coords from in-flight data; subsequent traffic forwards without re-discovery</text>
|
||||
|
||||
<!-- ═══ Caption ═══ -->
|
||||
<text x="430" y="660" text-anchor="middle" class="caption">Each transit node caches coordinates from the LookupResponse return path</text>
|
||||
<text x="430" y="668" text-anchor="middle" class="caption">LookupResponse caches coords at the originator only; transit caches warm during the subsequent data flow</text>
|
||||
</svg>
|
||||
|
||||
|
Before Width: | Height: | Size: 7.8 KiB After Width: | Height: | Size: 8.1 KiB |
@@ -71,7 +71,7 @@
|
||||
<!-- Arrow down from pubkey -->
|
||||
<line x1="400" y1="120" x2="400" y2="170" class="derive"/>
|
||||
<polygon points="396,168 400,176 404,168" class="dhead"/>
|
||||
<text x="416" y="148" class="op">one-way hash</text>
|
||||
<text x="416" y="148" class="op">SHA-256, truncate to 16 bytes</text>
|
||||
|
||||
<!-- node_addr box -->
|
||||
<rect x="280" y="176" width="240" height="50" class="box derived"/>
|
||||
@@ -95,7 +95,7 @@
|
||||
<!-- Arrow down from node_addr -->
|
||||
<line x1="400" y1="226" x2="400" y2="276" class="derive"/>
|
||||
<polygon points="396,274 400,282 404,274" class="dhead"/>
|
||||
<text x="416" y="254" class="op">add fd00::/8 prefix</text>
|
||||
<text x="416" y="254" class="op">0xfd + node_addr[0..15]</text>
|
||||
|
||||
<!-- IPv6 address box -->
|
||||
<rect x="280" y="282" width="240" height="50" class="box compat"/>
|
||||
|
||||
|
Before Width: | Height: | Size: 6.2 KiB After Width: | Height: | Size: 6.3 KiB |
@@ -82,48 +82,44 @@
|
||||
<!-- === Transport layer === -->
|
||||
|
||||
<!-- Overlay transports -->
|
||||
<rect x="80" y="336" width="150" height="60" class="cat"/>
|
||||
<text x="155" y="349" text-anchor="middle" class="cat">Overlay</text>
|
||||
<rect x="80" y="336" width="216" height="60" class="cat"/>
|
||||
<text x="188" y="349" text-anchor="middle" class="cat">Overlay</text>
|
||||
|
||||
<rect x="92" y="356" width="60" height="30" class="layer xport"/>
|
||||
<text x="122" y="371" text-anchor="middle" font-size="10" font-weight="bold" fill="#e0e0e0">UDP</text>
|
||||
<text x="122" y="381" text-anchor="middle" class="sub">IP</text>
|
||||
|
||||
<rect x="158" y="356" width="60" height="30" class="layer xport"/>
|
||||
<text x="188" y="371" text-anchor="middle" font-size="10" font-weight="bold" fill="#e0e0e0">Tor</text>
|
||||
<text x="188" y="381" text-anchor="middle" class="sub">.onion</text>
|
||||
<text x="188" y="371" text-anchor="middle" font-size="10" font-weight="bold" fill="#e0e0e0">TCP</text>
|
||||
<text x="188" y="381" text-anchor="middle" class="sub">IP</text>
|
||||
|
||||
<rect x="224" y="356" width="60" height="30" class="layer xport"/>
|
||||
<text x="254" y="371" text-anchor="middle" font-size="10" font-weight="bold" fill="#e0e0e0">Tor</text>
|
||||
<text x="254" y="381" text-anchor="middle" class="sub">.onion</text>
|
||||
|
||||
<!-- Shared medium transports -->
|
||||
<rect x="240" y="336" width="300" height="60" class="cat"/>
|
||||
<text x="390" y="349" text-anchor="middle" class="cat">Shared Medium</text>
|
||||
<rect x="306" y="336" width="234" height="60" class="cat"/>
|
||||
<text x="423" y="349" text-anchor="middle" class="cat">Shared Medium</text>
|
||||
|
||||
<rect x="254" y="356" width="60" height="30" class="layer xport"/>
|
||||
<text x="284" y="371" text-anchor="middle" font-size="10" font-weight="bold" fill="#e0e0e0">Ether</text>
|
||||
<text x="284" y="381" text-anchor="middle" class="sub">802.3</text>
|
||||
<rect x="318" y="356" width="60" height="30" class="layer xport"/>
|
||||
<text x="348" y="371" text-anchor="middle" font-size="10" font-weight="bold" fill="#e0e0e0">Ether</text>
|
||||
<text x="348" y="381" text-anchor="middle" class="sub">802.3</text>
|
||||
|
||||
<rect x="320" y="356" width="60" height="30" class="layer xport"/>
|
||||
<text x="350" y="371" text-anchor="middle" font-size="10" font-weight="bold" fill="#e0e0e0">WiFi</text>
|
||||
<text x="350" y="381" text-anchor="middle" class="sub">802.11</text>
|
||||
<rect x="384" y="356" width="60" height="30" class="layer xport"/>
|
||||
<text x="414" y="371" text-anchor="middle" font-size="10" font-weight="bold" fill="#e0e0e0">BLE</text>
|
||||
<text x="414" y="381" text-anchor="middle" class="sub">L2CAP</text>
|
||||
|
||||
<rect x="386" y="356" width="60" height="30" class="layer xport"/>
|
||||
<text x="416" y="371" text-anchor="middle" font-size="10" font-weight="bold" fill="#e0e0e0">BT</text>
|
||||
<text x="416" y="381" text-anchor="middle" class="sub">RFCOMM</text>
|
||||
|
||||
<rect x="452" y="356" width="60" height="30" class="layer xport"/>
|
||||
<text x="482" y="371" text-anchor="middle" font-size="10" font-weight="bold" fill="#e0e0e0">Radio</text>
|
||||
<text x="482" y="381" text-anchor="middle" class="sub">Sat, ...</text>
|
||||
<rect x="450" y="356" width="80" height="30" class="layer xport"/>
|
||||
<text x="490" y="371" text-anchor="middle" font-size="10" font-weight="bold" fill="#e0e0e0">Radio ...</text>
|
||||
<text x="490" y="381" text-anchor="middle" class="sub">future</text>
|
||||
|
||||
<!-- Point-to-point transports -->
|
||||
<rect x="550" y="336" width="150" height="60" class="cat"/>
|
||||
<text x="625" y="349" text-anchor="middle" class="cat">Point-to-Point</text>
|
||||
|
||||
<rect x="562" y="356" width="60" height="30" class="layer xport"/>
|
||||
<text x="592" y="371" text-anchor="middle" font-size="10" font-weight="bold" fill="#e0e0e0">Serial</text>
|
||||
<text x="592" y="381" text-anchor="middle" class="sub">UART</text>
|
||||
|
||||
<rect x="628" y="356" width="60" height="30" class="layer xport"/>
|
||||
<text x="658" y="371" text-anchor="middle" font-size="10" font-weight="bold" fill="#e0e0e0">...</text>
|
||||
<text x="658" y="381" text-anchor="middle" class="sub"></text>
|
||||
<rect x="562" y="356" width="126" height="30" class="layer xport"/>
|
||||
<text x="625" y="371" text-anchor="middle" font-size="10" font-weight="bold" fill="#e0e0e0">Serial ...</text>
|
||||
<text x="625" y="381" text-anchor="middle" class="sub">future</text>
|
||||
|
||||
<!-- === Peer networks below node box === -->
|
||||
<text x="155" y="436" text-anchor="middle" class="sub">Internet / Overlay Peers</text>
|
||||
|
||||
|
Before Width: | Height: | Size: 8.1 KiB After Width: | Height: | Size: 7.8 KiB |
@@ -19,7 +19,7 @@
|
||||
<rect x="50" y="25" width="775" height="100" class="layer app"/>
|
||||
<text x="75" y="58" class="name">Application Layer Interface</text>
|
||||
<text x="75" y="78" class="desc">Native FIPS API — for FIPS-aware applications</text>
|
||||
<text x="75" y="98" class="desc">IPv6 Shim — for traditional IP application backward compatibility</text>
|
||||
<text x="75" y="98" class="desc">IPv6 adapter — for traditional IP application backward compatibility</text>
|
||||
|
||||
<!-- FSP layer -->
|
||||
<rect x="50" y="150" width="775" height="120" class="layer fsp"/>
|
||||
|
||||
|
Before Width: | Height: | Size: 2.8 KiB After Width: | Height: | Size: 2.8 KiB |
@@ -76,54 +76,54 @@
|
||||
<text x="372" y="364" class="branch">No</text>
|
||||
|
||||
<!-- ============================================================ -->
|
||||
<!-- STEP 3: Bloom filter hit? -->
|
||||
<!-- STEP 3: Coords known? -->
|
||||
<!-- ============================================================ -->
|
||||
<text x="240" y="395" text-anchor="end" class="step">3</text>
|
||||
<polygon points="360,380 440,420 360,460 280,420" class="diamond"/>
|
||||
<text x="360" y="416" text-anchor="middle" class="decision">Bloom filter</text>
|
||||
<text x="360" y="430" text-anchor="middle" class="decision">hit?</text>
|
||||
<text x="360" y="416" text-anchor="middle" class="decision">Coords</text>
|
||||
<text x="360" y="430" text-anchor="middle" class="decision">known?</text>
|
||||
|
||||
<!-- Yes → right to interim step -->
|
||||
<!-- No → right to error outcome -->
|
||||
<line x1="440" y1="420" x2="520" y2="420" class="arrow" marker-end="url(#arrowhead)"/>
|
||||
<text x="475" y="412" text-anchor="middle" class="branch">Yes</text>
|
||||
<rect x="520" y="396" width="176" height="48" class="interim"/>
|
||||
<text x="608" y="412" text-anchor="middle" class="action">Rank candidates by</text>
|
||||
<text x="608" y="426" text-anchor="middle" class="action">tree distance and</text>
|
||||
<text x="608" y="440" text-anchor="middle" class="action">link performance</text>
|
||||
<text x="475" y="412" text-anchor="middle" class="branch">No</text>
|
||||
<rect x="528" y="402" width="160" height="36" class="error"/>
|
||||
<text x="608" y="425" text-anchor="middle" class="action">No route → error signal</text>
|
||||
|
||||
<!-- Arrow from interim → final outcome -->
|
||||
<line x1="608" y1="444" x2="608" y2="470" class="arrow" marker-end="url(#arrowhead)"/>
|
||||
<rect x="528" y="478" width="160" height="36" class="outcome"/>
|
||||
<text x="608" y="501" text-anchor="middle" class="action">Forward to 'best'</text>
|
||||
|
||||
<!-- No → down -->
|
||||
<!-- Yes → down -->
|
||||
<line x1="360" y1="460" x2="360" y2="530" class="arrow" marker-end="url(#arrowhead)"/>
|
||||
<text x="372" y="500" class="branch">No</text>
|
||||
<text x="372" y="500" class="branch">Yes</text>
|
||||
|
||||
<!-- ============================================================ -->
|
||||
<!-- STEP 4: Coords known? -->
|
||||
<!-- STEP 4: Bloom filter hit? -->
|
||||
<!-- ============================================================ -->
|
||||
<text x="240" y="545" text-anchor="end" class="step">4</text>
|
||||
<polygon points="360,530 440,570 360,610 280,570" class="diamond"/>
|
||||
<text x="360" y="566" text-anchor="middle" class="decision">Coords</text>
|
||||
<text x="360" y="580" text-anchor="middle" class="decision">known?</text>
|
||||
<text x="360" y="566" text-anchor="middle" class="decision">Bloom filter</text>
|
||||
<text x="360" y="580" text-anchor="middle" class="decision">hit?</text>
|
||||
|
||||
<!-- Yes → right to outcome -->
|
||||
<!-- Yes → right to interim step -->
|
||||
<line x1="440" y1="570" x2="520" y2="570" class="arrow" marker-end="url(#arrowhead)"/>
|
||||
<text x="475" y="562" text-anchor="middle" class="branch">Yes</text>
|
||||
<rect x="528" y="552" width="160" height="36" class="outcome"/>
|
||||
<text x="608" y="575" text-anchor="middle" class="action">Greedy tree forward</text>
|
||||
<rect x="520" y="546" width="176" height="48" class="interim"/>
|
||||
<text x="608" y="562" text-anchor="middle" class="action">Rank candidates by</text>
|
||||
<text x="608" y="576" text-anchor="middle" class="action">tree distance and</text>
|
||||
<text x="608" y="590" text-anchor="middle" class="action">link performance</text>
|
||||
|
||||
<!-- Arrow from interim → final outcome -->
|
||||
<line x1="608" y1="594" x2="608" y2="620" class="arrow" marker-end="url(#arrowhead)"/>
|
||||
<rect x="528" y="628" width="160" height="36" class="outcome"/>
|
||||
<text x="608" y="651" text-anchor="middle" class="action">Forward to 'best'</text>
|
||||
|
||||
<!-- No → down -->
|
||||
<line x1="360" y1="610" x2="360" y2="660" class="arrow" marker-end="url(#arrowhead)"/>
|
||||
<text x="372" y="640" class="branch">No</text>
|
||||
|
||||
<!-- ============================================================ -->
|
||||
<!-- STEP 5: No route → error signal -->
|
||||
<!-- STEP 5: Greedy tree forward -->
|
||||
<!-- ============================================================ -->
|
||||
<text x="240" y="683" text-anchor="end" class="step">5</text>
|
||||
<rect x="260" y="668" width="200" height="36" class="error"/>
|
||||
<text x="360" y="691" text-anchor="middle" class="action">No route → error signal</text>
|
||||
<rect x="260" y="668" width="200" height="36" class="outcome"/>
|
||||
<text x="360" y="691" text-anchor="middle" class="action">Greedy tree forward</text>
|
||||
|
||||
<!-- ============================================================ -->
|
||||
<!-- Legend -->
|
||||
@@ -149,5 +149,5 @@
|
||||
<text x="442" y="835" font-size="11" fill="#a0a0b0">Control flow direction</text>
|
||||
|
||||
<!-- Caption -->
|
||||
<text x="360" y="896" text-anchor="middle" class="caption">Each hop evaluates destinations in priority order 1–4, falling through on miss</text>
|
||||
<text x="360" y="896" text-anchor="middle" class="caption">Each hop checks 1–4 in priority order, falling through to greedy tree (5) when bloom yields no candidate; missing coords is the only error path</text>
|
||||
</svg>
|
||||
|
||||
|
Before Width: | Height: | Size: 8.5 KiB After Width: | Height: | Size: 8.6 KiB |
@@ -1,100 +0,0 @@
|
||||
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 720 520" font-family="system-ui, -apple-system, sans-serif" font-size="13">
|
||||
<defs>
|
||||
<marker id="arrow" viewBox="0 0 10 10" refX="10" refY="5" markerWidth="8" markerHeight="8" orient="auto-start-reverse">
|
||||
<path d="M 0 0 L 10 5 L 0 10 z" fill="#555"/>
|
||||
</marker>
|
||||
</defs>
|
||||
|
||||
<!-- Background -->
|
||||
<rect width="720" height="520" rx="8" fill="#fafafa" stroke="#ddd" stroke-width="1"/>
|
||||
|
||||
<!-- Title -->
|
||||
<text x="360" y="32" text-anchor="middle" font-size="16" font-weight="600" fill="#333">Document Relationships</text>
|
||||
|
||||
<!-- fips-intro -->
|
||||
<rect x="260" y="50" width="200" height="32" rx="6" fill="#e3f2fd" stroke="#90caf9"/>
|
||||
<text x="360" y="71" text-anchor="middle" font-weight="500" fill="#1565c0">fips-intro.md</text>
|
||||
|
||||
<!-- Arrows from intro -->
|
||||
<line x1="310" y1="82" x2="130" y2="120" stroke="#555" marker-end="url(#arrow)"/>
|
||||
<line x1="360" y1="82" x2="360" y2="120" stroke="#555" marker-end="url(#arrow)"/>
|
||||
|
||||
<!-- fips-transport-layer -->
|
||||
<rect x="30" y="120" width="200" height="32" rx="6" fill="#e8f5e9" stroke="#a5d6a7"/>
|
||||
<text x="130" y="141" text-anchor="middle" font-weight="500" fill="#2e7d32">fips-transport-layer.md</text>
|
||||
|
||||
<!-- fips-mesh-operation -->
|
||||
<rect x="270" y="120" width="200" height="32" rx="6" fill="#fff3e0" stroke="#ffcc80"/>
|
||||
<text x="370" y="141" text-anchor="middle" font-weight="500" fill="#e65100">fips-mesh-operation.md</text>
|
||||
|
||||
<!-- Arrow: transport -> mesh-layer -->
|
||||
<line x1="130" y1="152" x2="130" y2="190" stroke="#555" marker-end="url(#arrow)"/>
|
||||
|
||||
<!-- fips-mesh-layer -->
|
||||
<rect x="30" y="190" width="200" height="32" rx="6" fill="#e8f5e9" stroke="#a5d6a7"/>
|
||||
<text x="130" y="211" text-anchor="middle" font-weight="500" fill="#2e7d32">fips-mesh-layer.md</text>
|
||||
|
||||
<!-- Arrow: mesh-operation -> mesh-layer -->
|
||||
<line x1="270" y1="145" x2="230" y2="200" stroke="#555" marker-end="url(#arrow)"/>
|
||||
|
||||
<!-- Arrow: mesh-layer -> session-layer -->
|
||||
<line x1="130" y1="222" x2="130" y2="260" stroke="#555" marker-end="url(#arrow)"/>
|
||||
|
||||
<!-- fips-session-layer -->
|
||||
<rect x="30" y="260" width="200" height="32" rx="6" fill="#e8f5e9" stroke="#a5d6a7"/>
|
||||
<text x="130" y="281" text-anchor="middle" font-weight="500" fill="#2e7d32">fips-session-layer.md</text>
|
||||
|
||||
<!-- Arrow: session-layer -> ipv6-adapter -->
|
||||
<line x1="130" y1="292" x2="130" y2="330" stroke="#555" marker-end="url(#arrow)"/>
|
||||
|
||||
<!-- fips-ipv6-adapter -->
|
||||
<rect x="30" y="330" width="200" height="32" rx="6" fill="#e8f5e9" stroke="#a5d6a7"/>
|
||||
<text x="130" y="351" text-anchor="middle" font-weight="500" fill="#2e7d32">fips-ipv6-adapter.md</text>
|
||||
|
||||
<!-- Arrow: mesh-operation -> spanning-tree -->
|
||||
<line x1="470" y1="145" x2="550" y2="190" stroke="#555" marker-end="url(#arrow)"/>
|
||||
|
||||
<!-- Arrow: mesh-operation -> bloom-filters -->
|
||||
<line x1="470" y1="148" x2="550" y2="260" stroke="#555" marker-end="url(#arrow)"/>
|
||||
|
||||
<!-- fips-spanning-tree -->
|
||||
<rect x="500" y="190" width="200" height="32" rx="6" fill="#f3e5f5" stroke="#ce93d8"/>
|
||||
<text x="600" y="211" text-anchor="middle" font-weight="500" fill="#7b1fa2">fips-spanning-tree.md</text>
|
||||
|
||||
<!-- Arrow: intro -> spanning-tree -->
|
||||
<line x1="460" y1="72" x2="560" y2="190" stroke="#555" marker-end="url(#arrow)"/>
|
||||
|
||||
<!-- fips-bloom-filters -->
|
||||
<rect x="500" y="260" width="200" height="32" rx="6" fill="#f3e5f5" stroke="#ce93d8"/>
|
||||
<text x="600" y="281" text-anchor="middle" font-weight="500" fill="#7b1fa2">fips-bloom-filters.md</text>
|
||||
|
||||
<!-- Arrow: spanning-tree -> bloom (dependency) -->
|
||||
<line x1="600" y1="222" x2="600" y2="260" stroke="#999" stroke-dasharray="4,3" marker-end="url(#arrow)"/>
|
||||
|
||||
<!-- fips-wire-formats -->
|
||||
<rect x="260" y="400" width="200" height="32" rx="6" fill="#fce4ec" stroke="#ef9a9a"/>
|
||||
<text x="360" y="421" text-anchor="middle" font-weight="500" fill="#c62828">fips-wire-formats.md</text>
|
||||
<text x="360" y="448" text-anchor="middle" font-size="11" fill="#888">(referenced by all layer docs)</text>
|
||||
|
||||
<!-- fips-configuration -->
|
||||
<rect x="30" y="400" width="200" height="32" rx="6" fill="#f5f5f5" stroke="#bdbdbd"/>
|
||||
<text x="130" y="421" text-anchor="middle" font-weight="500" fill="#424242">fips-configuration.md</text>
|
||||
<text x="130" y="448" text-anchor="middle" font-size="11" fill="#888">(standalone reference)</text>
|
||||
|
||||
<!-- spanning-tree-dynamics -->
|
||||
<rect x="500" y="330" width="200" height="32" rx="6" fill="#f3e5f5" stroke="#ce93d8"/>
|
||||
<text x="600" y="351" text-anchor="middle" font-weight="500" fill="#7b1fa2">spanning-tree-dynamics.md</text>
|
||||
<text x="600" y="378" text-anchor="middle" font-size="11" fill="#888">(companion to fips-spanning-tree)</text>
|
||||
|
||||
<!-- Legend -->
|
||||
<rect x="30" y="472" width="14" height="14" rx="3" fill="#e8f5e9" stroke="#a5d6a7"/>
|
||||
<text x="50" y="484" font-size="11" fill="#666">Protocol stack</text>
|
||||
<rect x="155" y="472" width="14" height="14" rx="3" fill="#fff3e0" stroke="#ffcc80"/>
|
||||
<text x="175" y="484" font-size="11" fill="#666">Mesh behavior</text>
|
||||
<rect x="295" y="472" width="14" height="14" rx="3" fill="#f3e5f5" stroke="#ce93d8"/>
|
||||
<text x="315" y="484" font-size="11" fill="#666">Supporting references</text>
|
||||
<rect x="460" y="472" width="14" height="14" rx="3" fill="#fce4ec" stroke="#ef9a9a"/>
|
||||
<text x="480" y="484" font-size="11" fill="#666">Wire formats</text>
|
||||
<rect x="580" y="472" width="14" height="14" rx="3" fill="#f5f5f5" stroke="#bdbdbd"/>
|
||||
<text x="600" y="484" font-size="11" fill="#666">Implementation</text>
|
||||
</svg>
|
||||
|
Before Width: | Height: | Size: 5.4 KiB |
266
docs/design/fips-architecture.md
Normal file
@@ -0,0 +1,266 @@
|
||||
# FIPS Architecture
|
||||
|
||||
The protocol architecture, identity system, and two-layer encryption
|
||||
model. For the higher-level "what is FIPS and why" framing, see
|
||||
[fips-concepts.md](fips-concepts.md). For prior art and academic
|
||||
citations, see [fips-prior-work.md](fips-prior-work.md).
|
||||
|
||||
## Protocol Architecture
|
||||
|
||||
FIPS is organized in three protocol layers, each with distinct
|
||||
responsibilities and clean service boundaries. No layer depends on
|
||||
the specifics of the layers above or below it — transport plugins
|
||||
know nothing about sessions, the routing layer knows nothing about
|
||||
application addressing, and applications know nothing about which
|
||||
physical media carry their traffic. This separation means new
|
||||
transports, protocol features, and application interfaces can be
|
||||
added independently.
|
||||
|
||||

|
||||
|
||||
### Mapping to Traditional Networking
|
||||
|
||||
Readers familiar with the OSI model or TCP/IP networking may find it
|
||||
helpful to see how FIPS concepts relate to traditional layers:
|
||||
|
||||

|
||||
|
||||
Note that FMP spans what would traditionally be separate link and
|
||||
network layers. This is intentional — in a self-organizing mesh, the
|
||||
same layer that authenticates peers also makes routing decisions,
|
||||
because routing depends on authenticated peer state (spanning tree
|
||||
positions, bloom filters).
|
||||
|
||||
### Layer Responsibilities
|
||||
|
||||
**Transport layer**: Delivers datagrams between endpoints over a
|
||||
specific medium. Each transport type (UDP socket, Ethernet interface,
|
||||
radio modem) implements the same abstract interface: send and receive
|
||||
datagrams, report MTU. The transport layer knows nothing about FIPS
|
||||
identities, routing, or encryption. It provides raw datagram delivery
|
||||
to FMP above.
|
||||
|
||||
See [fips-transport-layer.md](fips-transport-layer.md) for the
|
||||
transport layer specification.
|
||||
|
||||
**FIPS Mesh Protocol (FMP)**: Manages peer connections, authenticates
|
||||
peers via Noise IK handshakes, and encrypts all traffic on each link.
|
||||
FMP is where the mesh organizes itself — nodes exchange spanning tree
|
||||
announcements and bloom filters with their direct peers, and FMP
|
||||
makes forwarding decisions for transit traffic. FMP provides
|
||||
authenticated, encrypted forwarding to FSP above.
|
||||
|
||||
See [fips-mesh-layer.md](fips-mesh-layer.md) for the FMP specification
|
||||
and [fips-mesh-operation.md](fips-mesh-operation.md) for how FMP's
|
||||
routing and self-organization work in practice.
|
||||
|
||||
**FIPS Session Protocol (FSP)**: Provides end-to-end authenticated
|
||||
encryption between any two nodes, regardless of how many intermediate
|
||||
hops separate them. FSP manages session lifecycle (setup, data
|
||||
transfer, teardown), caches destination coordinates for efficient
|
||||
routing, and handles the warmup strategy that keeps transit node
|
||||
caches populated. Session dispatch uses index-based routing inspired
|
||||
by [WireGuard](https://www.wireguard.com/), enabling O(1) packet
|
||||
demultiplexing. FSP provides a datagram service to applications above.
|
||||
|
||||
See [fips-session-layer.md](fips-session-layer.md) for the FSP
|
||||
specification.
|
||||
|
||||
**IPv6 adaptation layer**: Sits above FSP as a service on port 256,
|
||||
adapting the FIPS datagram service for unmodified IPv6 applications.
|
||||
Provides DNS resolution (npub → fd00::/8 address), identity cache
|
||||
management, IPv6 header compression, MTU enforcement, and a TUN
|
||||
interface. This is the primary way existing applications use the FIPS
|
||||
mesh.
|
||||
|
||||
See [fips-ipv6-adapter.md](fips-ipv6-adapter.md) for the IPv6 adapter.
|
||||
|
||||
### Node Architecture
|
||||
|
||||
Application services sit at the top of the stack, dispatched by FSP
|
||||
port number: the IPv6 TUN adapter (port 256) maps npubs to `fd00::/8`
|
||||
addresses with header compression so unmodified IP applications can
|
||||
use the network transparently, while the native datagram API
|
||||
addresses destinations directly by npub.
|
||||
|
||||

|
||||
|
||||
The mesh routes application traffic across heterogeneous transports
|
||||
transparently. A packet may traverse WiFi, Ethernet, UDP/IP, and Tor
|
||||
links on its way from source to destination — the application never
|
||||
needs to know which transports are involved. Each hop is independently
|
||||
encrypted at the link layer, while a single end-to-end session
|
||||
protects the payload across the entire path.
|
||||
|
||||

|
||||
|
||||

|
||||
|
||||
## Identity System
|
||||
|
||||
FIPS uses [Nostr](https://github.com/nostr-protocol/nips) keypairs
|
||||
(secp256k1) as node identities. The public key identifies the node;
|
||||
the private key signs protocol messages and establishes encrypted
|
||||
sessions.
|
||||
|
||||
The public key (or its bech32-encoded npub form) is the primary means
|
||||
for application-layer software to identify communication endpoints.
|
||||
Internally, the protocol derives a `node_addr` (a 16-byte SHA-256 hash
|
||||
of the pubkey) used as the routing identifier in packet headers, and
|
||||
an IPv6 address derived from the node_addr for the TUN adapter.
|
||||
Applications use the pubkey or npub; the routing layer uses node_addr;
|
||||
unmodified IPv6 applications use the derived `fd00::/8` address. All
|
||||
three are deterministically derived from the same keypair.
|
||||
|
||||
### FIPS Identity Handling
|
||||
|
||||

|
||||
|
||||
The pubkey is the node's cryptographic identity, used in Noise
|
||||
handshakes for both link encryption (IK) and session encryption (XK).
|
||||
It is never exposed beyond the endpoints of an encrypted channel. The node_addr, a one-way
|
||||
SHA-256 hash truncated to 16 bytes, serves as the routing identifier
|
||||
in packet headers and bloom filters. Intermediate routers see only
|
||||
node_addrs — they can forward traffic without learning the Nostr
|
||||
identities of the endpoints. An observer can verify "does this
|
||||
node_addr belong to pubkey X?" if they already know the pubkey, but
|
||||
cannot enumerate communicating identities by inspecting traffic. The
|
||||
IPv6 address prepends `fd` to the first 15 bytes of the node_addr,
|
||||
providing a ULA overlay address for unmodified IP applications via the
|
||||
TUN interface.
|
||||
|
||||
Below the FIPS identity layer, each transport uses its own native
|
||||
addressing — IP:port or hostname:port addresses, MAC addresses,
|
||||
.onion identifiers. These **link addresses** are opaque to everything
|
||||
above FMP and discarded once link authentication completes.
|
||||
|
||||
### Identity Verification
|
||||
|
||||
The Noise Protocol Framework mutually authenticates both peer-to-peer
|
||||
link connections (at FMP) and end-to-end session traffic (at FSP),
|
||||
proving each party controls the private key for their claimed
|
||||
identity.
|
||||
|
||||
See [fips-mesh-layer.md](fips-mesh-layer.md) for peer authentication
|
||||
and [fips-session-layer.md](fips-session-layer.md) for end-to-end
|
||||
session establishment.
|
||||
|
||||
Key rotation changes the node's identity — a new keypair produces a
|
||||
new node_addr and IPv6 address, requiring all sessions to be
|
||||
re-established. Migration mechanisms that allow a node to announce a
|
||||
successor key are a future consideration.
|
||||
|
||||
## Two-Layer Encryption
|
||||
|
||||
FIPS uses independent encryption at two protocol layers:
|
||||
|
||||
| Layer | Scope | Pattern | Purpose |
|
||||
| ----- | ----- | ------- | ------- |
|
||||
| **FMP (Mesh)** | Hop-by-hop | Noise IK | Encrypt all traffic on each peer link |
|
||||
| **FSP (Session)** | End-to-end | Noise XK | Encrypt application payload between endpoints |
|
||||
|
||||
### Link Layer (Hop-by-Hop)
|
||||
|
||||
When two nodes establish a direct connection, they perform a [Noise
|
||||
IK](https://noiseprotocol.org/) handshake. This authenticates both
|
||||
parties and establishes symmetric keys for encrypting all traffic on
|
||||
that link. Every packet between direct peers is encrypted — gossip
|
||||
messages, routing queries, and forwarded session datagrams alike.
|
||||
|
||||
The IK pattern is used because outbound connections know the peer's
|
||||
npub from configuration, while inbound connections learn the
|
||||
initiator's identity from the first handshake message.
|
||||
|
||||
### Session Layer (End-to-End)
|
||||
|
||||
FIPS establishes end-to-end encrypted sessions between any two
|
||||
communicating nodes using Noise XK, regardless of how many hops
|
||||
separate them. The initiator knows the destination's npub (required
|
||||
for XK's pre-message); the responder learns the initiator's identity
|
||||
from the third handshake message. Unlike the link-layer IK pattern
|
||||
where the initiator's identity is revealed in msg1, XK delays
|
||||
identity disclosure until msg3, providing stronger initiator identity
|
||||
protection for traffic traversing untrusted intermediate nodes.
|
||||
|
||||
A packet from A to D through intermediate nodes B and C:
|
||||
|
||||
1. A encrypts payload with A↔D session key (FSP)
|
||||
2. A wraps in SessionDatagram, encrypts with A↔B link key (FMP),
|
||||
sends to B
|
||||
3. B decrypts link layer, reads destination node_addr, re-encrypts
|
||||
with B↔C link key, forwards to C
|
||||
4. C decrypts link layer, re-encrypts with C↔D link key, forwards
|
||||
to D
|
||||
5. D decrypts link layer, then decrypts session layer to get payload
|
||||
|
||||
Intermediate nodes route based on destination node_addr but cannot
|
||||
read session-layer payloads. Each hop strips one link encryption and
|
||||
applies the next — the session-layer ciphertext passes through
|
||||
untouched.
|
||||
|
||||
Both layers always apply, even between adjacent peers — a packet to a
|
||||
direct neighbor is still encrypted twice. This uniform model means no
|
||||
special cases for local vs remote destinations, and topology changes
|
||||
(a direct peer becomes reachable only through intermediaries) don't
|
||||
affect existing sessions.
|
||||
|
||||
See [fips-mesh-layer.md](fips-mesh-layer.md) for link encryption and
|
||||
[fips-session-layer.md](fips-session-layer.md) for session encryption.
|
||||
|
||||
## Routing and Mesh Operation
|
||||
|
||||
Forwarding decisions are local. Each node combines spanning-tree
|
||||
coordinates with peer bloom filters to choose a next hop, falling back
|
||||
to greedy tree routing when bloom filters have not converged. Discovery
|
||||
warms transit node caches with destination coordinates, and three
|
||||
explicit error signals (CoordsRequired, PathBroken, MtuExceeded) drive
|
||||
recovery when forwarding fails. The full routing decision process,
|
||||
discovery protocol, and error-recovery integration view live in
|
||||
[fips-mesh-operation.md](fips-mesh-operation.md).
|
||||
|
||||
## Transport Abstraction
|
||||
|
||||
FIPS treats the communication medium as a pluggable component. UDP,
|
||||
TCP, raw Ethernet, Tor, BLE, and Nym all implement the same small
|
||||
datagram interface (send, receive, report MTU) and feed peers into a
|
||||
single FMP routing layer; radio and serial transports are in the
|
||||
planned set. Nym (an outbound-only mixnet transport) and Tor are
|
||||
privacy-oriented deployment modes rather than failover paths.
|
||||
Multi-transport nodes bridge between networks transparently. The
|
||||
transport-layer specification — including per-transport categories,
|
||||
the trait surface, the connection model, and implementation status —
|
||||
is in [fips-transport-layer.md](fips-transport-layer.md).
|
||||
|
||||
## Security
|
||||
|
||||
FIPS defends against four adversary classes (transport observers,
|
||||
active transport attackers, intermediate routers, and adversarial
|
||||
mesh nodes) through layered controls: hop-by-hop FMP link encryption,
|
||||
end-to-end FSP session encryption with stronger initiator identity
|
||||
protection, signed and replay-protected gossip, and rate-limited
|
||||
handshake processing. The threat-model details and per-layer
|
||||
mitigations are in [fips-mesh-layer.md](fips-mesh-layer.md), and the
|
||||
operator-facing controls (default-deny baseline, peer ACLs,
|
||||
filesystem permissions, cryptographic primitives) are consolidated in
|
||||
[fips-security.md](fips-security.md) and
|
||||
[../reference/security.md](../reference/security.md).
|
||||
|
||||
## MTU as a Cross-Cutting Concern
|
||||
|
||||
MTU is not owned by any single layer. The transport layer reports
|
||||
per-link MTU, FMP carries `path_mtu` in SessionDatagram and
|
||||
LookupResponse to track the minimum along a path, FSP echoes the
|
||||
observed forward-path MTU back to the source, and the IPv6 adapter
|
||||
enforces the resulting effective MTU at the TUN with ICMP Packet Too
|
||||
Big and TCP MSS clamping. The unified design — encapsulation overhead
|
||||
budget, proactive PMTUD, reactive MtuExceeded, and per-destination
|
||||
storage — is in [fips-mtu.md](fips-mtu.md).
|
||||
|
||||
## Approaches Considered but Rejected
|
||||
|
||||
One design alternative evaluated and ruled out during the architecture
|
||||
pass was onion routing, rejected because it requires the sender to
|
||||
know the full path upfront (incompatible with self-organizing
|
||||
routing) and prevents per-hop error feedback (incompatible with
|
||||
CoordsRequired/PathBroken recovery). The canonical mention lives in
|
||||
[fips-mesh-operation.md](fips-mesh-operation.md#privacy-considerations).
|
||||
@@ -153,6 +153,15 @@ this node, and this node thinks it can reach the same destination through Q.
|
||||
Split-horizon is computed per-peer: the outbound filter for peer Q merges
|
||||
all tree peer inbound filters except Q's.
|
||||
|
||||
### Filter Propagation Diagram
|
||||
|
||||

|
||||
|
||||
The outbound filter for peer Q merges this node's identity with tree
|
||||
peer inbound filters except Q's (split-horizon exclusion). Upward
|
||||
filters (child → parent) contain the child's subtree, while downward
|
||||
filters (parent → child) contain the complement.
|
||||
|
||||
### Directional Asymmetry
|
||||
|
||||
Because merge is restricted to tree peers, outgoing filters exhibit
|
||||
@@ -171,7 +180,7 @@ network with no overlap (excluding the node itself at the split point).
|
||||
|
||||
All peers — including non-tree mesh shortcuts — still **receive**
|
||||
FilterAnnounce messages and **store** received filters locally. These
|
||||
stored filters are consulted during routing (step 3 of `find_next_hop()`)
|
||||
stored filters are consulted during routing (step 4 of `find_next_hop()`)
|
||||
for single-hop shortcut discovery. However, mesh peer filters contain
|
||||
only the mesh peer's own tree-propagated information, not transitive
|
||||
entries from the broader network.
|
||||
@@ -202,9 +211,10 @@ Filter updates are event-driven, not periodic:
|
||||
|
||||
### Rate Limiting
|
||||
|
||||
Updates are rate-limited at 500ms minimum interval per peer to prevent
|
||||
storms during topology changes. Multiple pending changes within the
|
||||
cooldown period are coalesced into a single announcement.
|
||||
Updates are rate-limited at a 500ms minimum interval per peer
|
||||
(`node.bloom.update_debounce_ms`) to prevent storms during topology
|
||||
changes. Multiple pending changes within the cooldown period are
|
||||
coalesced into a single announcement.
|
||||
|
||||
### Propagation Scope
|
||||
|
||||
@@ -249,21 +259,14 @@ Where `filter_bits = 8 × (512 << size_class)` — 8,192 for v1.
|
||||
|
||||
## Wire Format
|
||||
|
||||
FilterAnnounce messages are carried inside encrypted link-layer frames:
|
||||
|
||||
| Offset | Field | Size | Description |
|
||||
| ------ | ----- | ---- | ----------- |
|
||||
| 0 | msg_type | 1 byte | 0x20 |
|
||||
| 1 | sequence | 8 bytes LE | Monotonic counter for freshness |
|
||||
| 9 | hash_count | 1 byte | Number of hash functions (5 in v1) |
|
||||
| 10 | size_class | 1 byte | Filter size: `512 << size_class` bytes |
|
||||
| 11 | filter_bits | 1,024 bytes | Bloom filter bit array (v1) |
|
||||
|
||||
**v1 total**: 1,035 bytes payload, 1,064 bytes with link encryption
|
||||
overhead.
|
||||
|
||||
See [fips-wire-formats.md](fips-wire-formats.md) for the complete wire
|
||||
format reference.
|
||||
The FilterAnnounce byte layout (`msg_type 0x20`, sequence, hash_count,
|
||||
size_class, filter_bits) lives in
|
||||
[../reference/wire-formats.md](../reference/wire-formats.md). The
|
||||
v1 plaintext payload is 1,035 bytes (11-byte header + 1,024-byte
|
||||
filter); link encryption adds 36 bytes of FMP framing (16-byte outer
|
||||
header + 4-byte inner timestamp + 16-byte AEAD tag), bringing the
|
||||
on-the-wire size to roughly 1,071 bytes before the underlying
|
||||
transport's per-packet overhead.
|
||||
|
||||
## Scale and Size Classes
|
||||
|
||||
@@ -320,9 +323,8 @@ The envisioned approach is that hub nodes near the root — which carry the
|
||||
largest downward filters — would use larger size classes, while leaf nodes
|
||||
and resource-constrained nodes continue with smaller filters. A node
|
||||
receiving a filter larger than its own size class folds it down locally.
|
||||
The mechanism by which heterogeneous filter sizes propagate through the
|
||||
tree is a future design direction not specified in v1. See
|
||||
[IDEA-0043](../../ideas/IDEA-0043-heterogeneous-filter-propagation.md).
|
||||
The mechanism by which heterogeneous filter sizes propagate through
|
||||
the tree is a future design direction not specified in v1.
|
||||
|
||||
### Folding
|
||||
|
||||
@@ -335,6 +337,43 @@ The hash function design supports folding: membership tests at a smaller
|
||||
size use `hash(item, i) % smaller_bit_count`, which maps to the same bit
|
||||
positions that folding produces.
|
||||
|
||||
## Mesh Size Estimation
|
||||
|
||||
A filter's saturation can be inverted into an estimated entry count
|
||||
via the standard formula `n ≈ -(m/k) · ln(1 − X/m)`, where `m` is the
|
||||
filter size in bits, `k` is the hash count, and `X` is the population
|
||||
count. Rather than estimate per-filter and sum, the node first builds
|
||||
an **OR-union of every connected peer's inbound filter** — all routing
|
||||
peers, including cross-links, not just the tree parent and children —
|
||||
inserts its own address into the union, and inverts the cardinality
|
||||
**once on the resulting union**. Because filter propagation is
|
||||
split-horizon (each outgoing filter excludes the peer it routes back
|
||||
to), every routing peer advertises a near-complete "whole mesh minus
|
||||
my subtree" view, so the union covers the network. OR-ing is
|
||||
idempotent, so overlapping bits deduplicate instead of over-counting,
|
||||
and folding in all peers rather than only the tree neighborhood damps
|
||||
the count flap on a parent switch (the cross-links still carry the
|
||||
upward coverage) and removes any dependence on tree-declaration cache
|
||||
freshness. The result is cached on the node and exposed through the
|
||||
control socket and `fipstop` dashboard. (See `compute_mesh_size()` in
|
||||
`src/node/mod.rs`.)
|
||||
|
||||
The estimator refuses to produce a value when any contributing filter
|
||||
is above the antipoison FPR cap (`node.bloom.max_inbound_fpr`,
|
||||
default `0.20`); a partial aggregate would silently underestimate.
|
||||
Consumers handle the resulting `None` by displaying an "unknown"
|
||||
state rather than a misleading number.
|
||||
|
||||
## Antipoison: Inbound FPR Cap
|
||||
|
||||
Inbound `FilterAnnounce` payloads are checked against
|
||||
`node.bloom.max_inbound_fpr` (default `0.20`). Filters whose
|
||||
estimated false positive rate exceeds the cap are dropped silently
|
||||
(no NACK on the wire) — they would otherwise inflate downstream
|
||||
candidate evaluation cost without contributing useful discrimination.
|
||||
The cap also gates filters from feeding into mesh size estimation,
|
||||
as described above.
|
||||
|
||||
## Implementation Status
|
||||
|
||||
| Feature | Status |
|
||||
@@ -349,6 +388,8 @@ positions that folding produces.
|
||||
| 500ms rate limiting | **Implemented** |
|
||||
| FilterAnnounce gossip (all peers) | **Implemented** |
|
||||
| Filter cardinality logging | **Implemented** |
|
||||
| Mesh size estimation (OR-union of peer filters) | **Implemented** |
|
||||
| Inbound FPR cap (antipoison) | **Implemented** |
|
||||
| Size class negotiation | Future direction |
|
||||
| Folding support | Future direction |
|
||||
| Adaptive filter sizing | Future direction |
|
||||
@@ -357,6 +398,7 @@ positions that folding produces.
|
||||
|
||||
- [fips-mesh-operation.md](fips-mesh-operation.md) — How bloom filters fit
|
||||
into routing
|
||||
- [fips-wire-formats.md](fips-wire-formats.md) — FilterAnnounce wire format
|
||||
- [../reference/wire-formats.md](../reference/wire-formats.md) —
|
||||
FilterAnnounce wire format
|
||||
- [fips-spanning-tree.md](fips-spanning-tree.md) — The coordinate system
|
||||
that bloom filter candidates are ranked by
|
||||
|
||||
123
docs/design/fips-concepts.md
Normal file
@@ -0,0 +1,123 @@
|
||||
# FIPS Concepts
|
||||
|
||||
A novice-friendly introduction to what FIPS is, why it exists, and the
|
||||
mental model behind a self-organizing mesh. For the protocol stack,
|
||||
identity system, and encryption walkthrough, see
|
||||
[fips-architecture.md](fips-architecture.md). For prior art and
|
||||
academic citations, see [fips-prior-work.md](fips-prior-work.md).
|
||||
|
||||
## What is FIPS?
|
||||
|
||||
FIPS is a self-organizing mesh network that can operate natively over a
|
||||
variety of physical and logical media, such as local area networks,
|
||||
Bluetooth, serial links, or the existing internet as an overlay. The
|
||||
long-term goal is infrastructure that can function alongside or
|
||||
ultimately replace dependence on the Internet itself. Systems running
|
||||
FIPS establish peer connections, authenticate each other, and route
|
||||
traffic for each other without any central authority or global topology
|
||||
knowledge, and allow end-to-end encrypted sessions between any two
|
||||
nodes regardless of how many hops separate them.
|
||||
|
||||
Nodes in the mesh route traffic for each other using Nostr identities
|
||||
(npubs) as network addresses. Applications can access the mesh through
|
||||
a native FIPS datagram service, or through an IPv6 adaptation layer
|
||||
that presents each node as an IPv6 endpoint for compatibility with
|
||||
existing IP-based applications.
|
||||
|
||||
## Why FIPS?
|
||||
|
||||
**Self-sovereign identity**: FIPS nodes generate their own addresses,
|
||||
node IDs, and security credentials without coordination with any
|
||||
central authority. These identities can be long-term fixed or may be
|
||||
ephemeral, changed at any time. These identities are not visible to
|
||||
the FIPS network itself — they are used only at the application layer
|
||||
and for end-to-end session encryption.
|
||||
|
||||
**Infrastructure independence**: The internet depends on centralized
|
||||
infrastructure — ISPs, backbone providers, DNS, certificate
|
||||
authorities. FIPS works over any transport that can carry packets: a
|
||||
serial connection, onion-routed connections through Tor, local area
|
||||
networking, radio links between remote sites, or the existing internet
|
||||
as an overlay. When the internet is unavailable, unreliable, or
|
||||
untrusted, the mesh still works.
|
||||
|
||||
**Privacy by design**: FIPS provides secure, authenticated, and
|
||||
encrypted communication between any two nodes in the mesh, independent
|
||||
of the mix of transports used along the routed path between them.
|
||||
Furthermore, the mesh itself is designed to minimize metadata exposure
|
||||
— intermediate nodes route packets without learning the identities of
|
||||
the endpoints.
|
||||
|
||||
**Zero configuration**: Nodes discover each other and build routing
|
||||
automatically. Connect to one peer and you can reach the entire mesh.
|
||||
The network self-heals around failures and adapts to changing topology.
|
||||
|
||||
## A Self-Organizing Mesh
|
||||
|
||||
Traditional networks are built top-down. A central authority assigns
|
||||
addresses, configures routing tables, provisions hardware, and manages
|
||||
the topology. If the authority disappears or the infrastructure fails,
|
||||
the network fails with it. Nodes cannot reach each other without
|
||||
infrastructure mediating the connection.
|
||||
|
||||
FIPS inverts this model. There is no central authority, no address
|
||||
assignment service, no routing table pushed from above. Each node
|
||||
generates its own identity from a cryptographic keypair. Each node
|
||||
independently decides which peers to connect to and which transports
|
||||
to use. From these local decisions alone, the network self-organizes:
|
||||
|
||||
- A **spanning tree** forms through distributed parent selection,
|
||||
giving every node a coordinate in the network without any node
|
||||
knowing the full topology
|
||||
- **Bloom filters** propagate through gossip, so each node learns
|
||||
which peers can reach which destinations — again without global
|
||||
knowledge
|
||||
- **Routing decisions** are made locally at each hop, using only the
|
||||
node's immediate peers and cached coordinate information
|
||||
|
||||
Each peer link and end-to-end session actively measures RTT, loss,
|
||||
jitter, and goodput through a lightweight in-band Metrics Measurement
|
||||
Protocol (MMP), providing operator visibility and a foundation for
|
||||
quality-aware routing.
|
||||
|
||||
The result is a network that builds itself from the bottom up, heals
|
||||
around failures automatically, and scales without central coordination.
|
||||
Adding a node is as simple as connecting to one existing peer — the
|
||||
network integrates the new node through its normal mesh protocols.
|
||||
|
||||
## Specific Design Goals
|
||||
|
||||
- **Nostr-native identity and cryptography** — Use Nostr keypairs as
|
||||
node identities and leverage secp256k1, Schnorr signatures, and
|
||||
SHA-256
|
||||
- **Transport agnostic** — Support overlay, shared medium, and
|
||||
point-to-point transports transparently
|
||||
- **Self-organizing** — Automatic topology discovery and route
|
||||
optimization
|
||||
- **Privacy preserving** — Minimize metadata leakage across untrusted
|
||||
links
|
||||
- **Resilient** — Self-healing with graceful degradation
|
||||
|
||||
Non-goals include:
|
||||
|
||||
- **Reliable delivery** — FIPS provides a best-effort datagram
|
||||
service; retransmission and ordering are left to applications or
|
||||
higher-layer protocols
|
||||
- **Anonymity** — Direct peers learn each other's identity; FIPS
|
||||
minimizes metadata exposure but is not an anonymity network like Tor
|
||||
- **Congestion control** — FIPS measures link quality but does not
|
||||
implement flow control or congestion avoidance at the mesh layer
|
||||
|
||||
## Where to Read Next
|
||||
|
||||
- [fips-architecture.md](fips-architecture.md) — protocol stack,
|
||||
identity system, two-layer encryption, MTU as a cross-cutting
|
||||
concern
|
||||
- [fips-spanning-tree.md](fips-spanning-tree.md) — how the tree forms
|
||||
and reconverges
|
||||
- [fips-bloom-filters.md](fips-bloom-filters.md) — how reachability
|
||||
information propagates
|
||||
- [fips-mesh-operation.md](fips-mesh-operation.md) — how the pieces
|
||||
work together at runtime
|
||||
- [fips-prior-work.md](fips-prior-work.md) — designs and protocols
|
||||
FIPS builds on
|
||||
541
docs/design/fips-gateway.md
Normal file
@@ -0,0 +1,541 @@
|
||||
# FIPS Gateway
|
||||
|
||||
The FIPS gateway lets unmodified IPv6 hosts on a LAN exchange traffic
|
||||
with the mesh without running any FIPS software themselves. It is a
|
||||
niche feature — most operators will never enable it. The gateway
|
||||
runs most conveniently on a system that is already providing network
|
||||
services (DHCP, DNS, RA) to a LAN segment, since hosts on that
|
||||
segment already get IP assignment and a default route from that box.
|
||||
The canonical example is an OpenWrt-based WiFi access point: every
|
||||
client that associates with the AP already has the AP as default
|
||||
router and DNS server, which is exactly the placement the gateway
|
||||
needs. The OpenWrt ipk ships with the `gateway:` block of
|
||||
`/etc/fips/fips.yaml` pre-populated and the integration glue
|
||||
(dnsmasq forwarding, RA route for the virtual pool, global-scope
|
||||
IPv6 prefix on `br-lan`) automated by the init script —
|
||||
[`packaging/openwrt-ipk/files/etc/init.d/fips-gateway`](https://github.com/jmcorgan/fips/blob/master/packaging/openwrt-ipk/files/etc/init.d/fips-gateway).
|
||||
The operator only needs to enable and start the service. Running the
|
||||
gateway on a non-OpenWrt LAN-edge host (a Linux router/server, for
|
||||
example) is technically possible but requires manual integration:
|
||||
distributing a route to the virtual-IP pool, wiring DNS forwarding so
|
||||
LAN clients send `.fips` queries to the gateway, configuring sysctls
|
||||
and capabilities. That path is supported but tedious; it is the
|
||||
secondary path.
|
||||
|
||||
The feature has two halves that share common machinery and have
|
||||
their own unique parts.
|
||||
|
||||
The **outbound half** carries traffic from LAN to mesh. A non-FIPS
|
||||
LAN workstation resolves `<npub>.fips` (or a `.fips` host alias) via
|
||||
the gateway's DNS proxy, which returns a virtual IPv6 address from a
|
||||
managed pool. The kernel routes the LAN packet to that virtual IP via
|
||||
a route to the pool CIDR (RA-advertised, statically distributed, or
|
||||
on-link via the default route). The gateway runs nftables NAT so the
|
||||
packet appears on the mesh as if it had originated from the gateway's
|
||||
own FIPS identity: prerouting DNAT rewrites the destination from the
|
||||
virtual IP to the real `fd00::/8` mesh address, and postrouting
|
||||
masquerade rewrites the source from the LAN host's address to the
|
||||
gateway's `fips0` address. Return traffic follows the conntrack
|
||||
reverse path back to the originating LAN host, with postrouting SNAT
|
||||
restoring the virtual IP as source so the client sees a response from
|
||||
the address it connected to.
|
||||
|
||||
The **inbound half** carries traffic from mesh to LAN. A
|
||||
configuration entry in `gateway.port_forwards[]` exposes a LAN
|
||||
service (`host:port`) on a port of the gateway's mesh-side `fips0`
|
||||
address. Mesh peers reach it as `<gateway-npub>.fips:<listen_port>`.
|
||||
A prerouting DNAT rule keyed on `(iif=fips0, l4proto, dport)`
|
||||
rewrites the destination to the LAN target; a LAN-side masquerade in
|
||||
postrouting rewrites the mesh peer's source so the LAN target sees a
|
||||
reachable LAN address and conntrack steers replies back through the
|
||||
gateway. This is the inverse of port-forwarding on a conventional NAT
|
||||
router.
|
||||
|
||||
The two halves are independent and can be configured separately.
|
||||
Inbound port-forwards work without any outbound configuration (just
|
||||
a port-forward list and the table); outbound works without any
|
||||
inbound forwards. They share the same nftables table, the same
|
||||
binary, the same control socket, and the same atomic-rebuild
|
||||
strategy. That shared machinery is what makes them halves of one
|
||||
feature rather than two separate features.
|
||||
|
||||
## Architecture
|
||||
|
||||
### The `fips-gateway` Service
|
||||
|
||||
The gateway is a separate binary, [`fips-gateway`](https://github.com/jmcorgan/fips/blob/master/src/bin/fips-gateway.rs),
|
||||
not part of the FIPS daemon. It reads the same `/etc/fips/fips.yaml`
|
||||
the daemon reads (via `--config`, or the standard search path), but
|
||||
acts on the `gateway.*` block. It needs `CAP_NET_ADMIN` to install
|
||||
nftables rules, manage proxy NDP entries, and add the pool route.
|
||||
The CLI is documented in
|
||||
[../reference/cli-fips-gateway.md](../reference/cli-fips-gateway.md).
|
||||
|
||||
The gateway connects to the daemon indirectly. The outbound half
|
||||
forwards `.fips` DNS queries to the daemon's built-in resolver
|
||||
(default `[::1]:5354`); the daemon resolves the name to a mesh
|
||||
address and primes its identity cache as a side effect. The inbound
|
||||
half does not require any daemon plumbing at all — packets that
|
||||
arrive on `fips0` after the daemon's TUN injection path are matched
|
||||
by the nftables rules on `fips0` ingress. There is no shared memory,
|
||||
no IPC channel, and no startup ordering coupling beyond "the daemon's
|
||||
DNS responder must be reachable before the gateway starts serving
|
||||
LAN queries", which the gateway enforces with a bounded reachability
|
||||
probe at startup.
|
||||
|
||||
### nftables Table Layout
|
||||
|
||||
All gateway rules live in a single nftables table, `inet
|
||||
fips_gateway`, with two chains:
|
||||
|
||||
- `prerouting` — `type nat hook prerouting priority dstnat (-100)`,
|
||||
for both LAN→mesh DNAT (per virtual-IP mapping) and mesh→LAN DNAT
|
||||
(per port-forward).
|
||||
- `postrouting` — `type nat hook postrouting priority srcnat (100)`,
|
||||
for both the always-on `oifname fips0` masquerade, the per-mapping
|
||||
return-path SNAT, and (when any port-forward is configured) the
|
||||
LAN-side masquerade for inbound traffic.
|
||||
|
||||
The table is rebuilt atomically on every change. The rebuild
|
||||
sequence — delete the existing table (ignore `ENOENT` on first
|
||||
call), then create a new table with chains and the full rule set in
|
||||
a single netlink batch — avoids reliance on kernel rule-handle
|
||||
tracking, which the rustables crate does not expose. The table stays
|
||||
small (one always-on masquerade plus two rules per active outbound
|
||||
mapping plus one rule per inbound forward, with one extra masquerade
|
||||
when any forward is present), so rebuilds are cheap.
|
||||
|
||||
### Control Socket
|
||||
|
||||
`fips-gateway` exposes a Unix-domain control socket at
|
||||
`/run/fips/gateway.sock` (`root:fips`, mode `0770`) with two
|
||||
commands: `show_gateway` and `show_mappings`. The protocol is the
|
||||
same line-delimited JSON used by the daemon's control socket. The
|
||||
shapes are documented in the
|
||||
[Gateway command catalog](../reference/control-socket.md#gateway-command-catalog).
|
||||
There is no `fipsctl gateway` subcommand; clients (including
|
||||
`fipstop`'s gateway view) talk to the socket directly.
|
||||
|
||||
### Diagram
|
||||
|
||||
```text
|
||||
LAN clients
|
||||
│
|
||||
DNS query (.fips) │ IPv6 packet
|
||||
for outbound │ to virtual IP
|
||||
│ or mesh peer
|
||||
▼
|
||||
┌───────────────────────────────────┐
|
||||
│ fips-gateway │
|
||||
│ │
|
||||
│ ┌──────────────┐ ┌───────────┐ │
|
||||
│ │ DNS proxy │ │ Virtual │ │
|
||||
│ │ ([::1]:5353) │─▶│ IP pool │ │
|
||||
│ │ .fips only │ │ (state │ │
|
||||
│ └──────┬───────┘ │ machine) │ │
|
||||
│ │ └─────┬─────┘ │
|
||||
│ │ │ │
|
||||
│ forward to │ pool │
|
||||
│ daemon resolver │ events │
|
||||
│ ([::1]:5354) ▼ │
|
||||
│ │ ┌───────────┐ │
|
||||
│ │ │ NAT │ │
|
||||
│ │ │ manager │ │
|
||||
│ │ │ (rebuild │ │
|
||||
│ │ │ inet │ │
|
||||
│ │ │ fips_ │ │
|
||||
│ │ │ gateway) │ │
|
||||
│ │ └─────┬─────┘ │
|
||||
│ │ │ │
|
||||
│ │ ┌─────▼─────┐ │
|
||||
│ │ │ net │ │
|
||||
│ │ │ setup │ │
|
||||
│ │ │ (proxy │ │
|
||||
│ │ │ NDP, lo │ │
|
||||
│ │ │ route) │ │
|
||||
│ │ └───────────┘ │
|
||||
│ │ │
|
||||
│ │ control socket │
|
||||
│ │ /run/fips/ │
|
||||
│ │ gateway.sock │
|
||||
└─────────┼─────────────────────────┘
|
||||
│
|
||||
▼
|
||||
FIPS daemon resolver
|
||||
([::1]:5354)
|
||||
│
|
||||
▼
|
||||
fips0 TUN interface
|
||||
│
|
||||
▼
|
||||
the mesh
|
||||
```
|
||||
|
||||
The DNS proxy and the virtual IP pool are exclusive to the outbound
|
||||
half. The NAT manager and the kernel-side machinery (nftables table,
|
||||
`fips0` and LAN interfaces, conntrack) are shared. The inbound half
|
||||
contributes per-port-forward rules to the same table without
|
||||
involving the DNS proxy or the pool.
|
||||
|
||||
## The Outbound Half (LAN → Mesh)
|
||||
|
||||
### DNS Resolution Flow
|
||||
|
||||
1. A LAN client sends a DNS query to the gateway's listener (default
|
||||
`[::1]:5353`, configurable via `gateway.dns.listen`). The default
|
||||
is loopback-only on an unprivileged port: the canonical deployment
|
||||
has another resolver on the host (dnsmasq, systemd-resolved, BIND)
|
||||
holding port 53 and forwarding `.fips` queries to the gateway over
|
||||
loopback. Operators on a host without a pre-existing resolver on
|
||||
53 can override the listen value to `"[::]:53"` to let LAN clients
|
||||
query the gateway directly.
|
||||
2. If the question is not for a `.fips` domain, the gateway replies
|
||||
`REFUSED`. The proxy is intentionally narrow — it does not resolve
|
||||
public DNS, and the LAN's primary resolver should hold port 53 on
|
||||
the gateway host (the OpenWrt init script wires dnsmasq to forward
|
||||
`.fips` queries to the loopback listener automatically).
|
||||
3. The gateway forwards the query to the daemon resolver
|
||||
(`gateway.dns.upstream`, default `[::1]:5354`). The daemon must
|
||||
match: an IPv6 socket bound to `[::1]` does not accept v4-mapped
|
||||
traffic, so a `127.0.0.1:5354` upstream cannot reach a daemon
|
||||
bound on `[::1]:5354`.
|
||||
4. If the daemon is unreachable or times out (5 s), the gateway
|
||||
replies `SERVFAIL`. If the daemon returns `NXDOMAIN` or a
|
||||
non-`AAAA` answer, the gateway forwards the response unchanged.
|
||||
5. The gateway extracts the AAAA (`fd00::/8`) record from the
|
||||
daemon's response. This resolution primes the daemon's identity
|
||||
cache as a side effect — a prerequisite for `fips0` routing,
|
||||
because the daemon needs the cache entry to map the mesh address
|
||||
back to a `NodeAddr` for forwarding.
|
||||
6. The gateway allocates a virtual IP from the pool for that mesh
|
||||
address (idempotent: an existing mapping is reused and its TTL
|
||||
refreshed).
|
||||
7. If a new mapping was created, the pool emits `MappingCreated`,
|
||||
which the main loop turns into `add_mapping` calls on the NAT
|
||||
manager and `add_proxy_ndp` on the network setup.
|
||||
8. The gateway returns an `AAAA` response containing the virtual IP,
|
||||
with the configured TTL (default 60 s).
|
||||
|
||||
### Virtual IP Pool
|
||||
|
||||
The pool allocates IPv6 addresses from a required CIDR (commonly
|
||||
`fd01::/112`). Each address maps to one mesh destination, keyed by
|
||||
`NodeAddr` rather than by hostname — different `.fips` aliases for
|
||||
the same node share a virtual IP. Address 0 (the network-equivalent)
|
||||
is reserved; the rest are allocatable. The pool is capped at 2^16
|
||||
addresses regardless of prefix length, to bound memory.
|
||||
|
||||
The pool tracks state per address:
|
||||
|
||||
```text
|
||||
Allocated ──→ Active ──→ Draining ──→ Free
|
||||
│ ▲
|
||||
└──────────────────────────────────┘
|
||||
(TTL expired, no sessions)
|
||||
```
|
||||
|
||||
| State | Meaning |
|
||||
| ----- | ------- |
|
||||
| Allocated | DNS query created the mapping; no NAT sessions yet. |
|
||||
| Active | Conntrack reports at least one session for this virtual IP. |
|
||||
| Draining | TTL has expired; sessions may still be in progress, or grace period is running after sessions ended. |
|
||||
| Free | Reclaimed and available for new allocations. |
|
||||
|
||||
Transitions:
|
||||
|
||||
- **Allocated → Active**: conntrack sessions count goes above zero.
|
||||
- **Allocated → Free**: TTL expires before any session is ever
|
||||
observed.
|
||||
- **Active → Draining**: TTL expires (sessions may or may not still
|
||||
be present).
|
||||
- **Draining → Free**: session count is zero and the grace period
|
||||
has elapsed since draining began.
|
||||
|
||||
Timing:
|
||||
|
||||
- **TTL** (`gateway.dns.ttl`, default 60 s) is both the DNS TTL
|
||||
returned to the client and the mapping's idle lifetime. Repeated
|
||||
DNS queries for the same destination refresh the
|
||||
`last_referenced` timestamp.
|
||||
- **Grace period** (`gateway.pool_grace_period`, default 60 s) is
|
||||
the dwell time after the last session ends before the address is
|
||||
recycled. It prevents immediate reuse from confusing hosts with
|
||||
cached DNS responses.
|
||||
- **Tick interval**: the pool re-evaluates state every 10 s.
|
||||
|
||||
Active session counts come from `/proc/net/nf_conntrack`: an entry
|
||||
counts as a session if its original destination is the virtual IP.
|
||||
|
||||
If the pool is exhausted, new DNS queries return `SERVFAIL`.
|
||||
Existing mappings are never evicted prematurely — the correctness of
|
||||
in-flight sessions takes precedence over fresh allocations.
|
||||
|
||||
### NAT Pipeline (Outbound)
|
||||
|
||||
Three rule classes in `inet fips_gateway` together implement the
|
||||
LAN→mesh path:
|
||||
|
||||
**Prerouting DNAT (per mapping)** rewrites the destination from the
|
||||
virtual IP to the corresponding mesh address:
|
||||
|
||||
```text
|
||||
match: nfproto ipv6 && ip6 daddr == <virtual_ip>
|
||||
action: dnat to <mesh_addr>
|
||||
```
|
||||
|
||||
After DNAT, the kernel routes the packet through `fips0` via the
|
||||
standard routing table.
|
||||
|
||||
**Postrouting masquerade (`oifname fips0`)** rewrites the source of
|
||||
all traffic exiting via `fips0` to the gateway's own `fips0` address:
|
||||
|
||||
```text
|
||||
match: oifname == "fips0"
|
||||
action: masquerade
|
||||
```
|
||||
|
||||
This rule is critical. Without it, LAN client source addresses (for
|
||||
example `fd02::20` from the LAN's RA-advertised prefix, or virtual
|
||||
addresses from another forwarding domain) would appear as the source
|
||||
on the mesh. Those addresses are meaningless to mesh nodes, so
|
||||
return traffic would be black-holed. Masquerade ensures all mesh
|
||||
traffic appears to originate from the gateway's own FIPS identity.
|
||||
|
||||
**Postrouting SNAT (per mapping)** rewrites the source of return
|
||||
traffic from the mesh address back to the virtual IP:
|
||||
|
||||
```text
|
||||
match: nfproto ipv6 && ip6 saddr == <mesh_addr>
|
||||
action: snat to <virtual_ip>
|
||||
```
|
||||
|
||||
Without it, the LAN client would see replies from the raw
|
||||
`fd00::/8` mesh address rather than from the virtual IP it had
|
||||
originally connected to, breaking application-layer assumptions about
|
||||
the destination address.
|
||||
|
||||
### Network Requirements (Outbound)
|
||||
|
||||
The gateway host needs IPv6 forwarding enabled
|
||||
(`net.ipv6.conf.all.forwarding=1`), proxy NDP enabled on the LAN
|
||||
interface, `CAP_NET_ADMIN` for `fips-gateway`, and a `local
|
||||
<pool-cidr> dev lo` route so the kernel accepts packets to the pool
|
||||
as locally owned and runs them through the NAT chains. LAN clients
|
||||
need a route to the pool via the gateway and DNS resolution that
|
||||
forwards `.fips` queries there. On OpenWrt the init script handles
|
||||
all of this; on other Linux hosts the operator handles it manually.
|
||||
Full setup is documented in
|
||||
[../how-to/deploy-gateway.md](../how-to/deploy-gateway.md).
|
||||
|
||||
## The Inbound Half (Mesh → LAN)
|
||||
|
||||
### Configuration Shape
|
||||
|
||||
Inbound port-forwards live in `gateway.port_forwards[]`. Each entry
|
||||
is a triple:
|
||||
|
||||
| Field | Type | Notes |
|
||||
| ----- | ---- | ----- |
|
||||
| `listen_port` | `u16` | Port on the gateway's `fips0` address. Must be non-zero. |
|
||||
| `proto` | `tcp` \| `udp` | Match protocol. |
|
||||
| `target` | `[ipv6]:port` | LAN destination. IPv4 targets are rejected at parse time by `SocketAddrV6`. |
|
||||
|
||||
Validation runs at startup and on every config reload:
|
||||
`(listen_port, proto)` must be unique across the list, and zero
|
||||
listen ports are rejected. Forwards are independent of outbound
|
||||
configuration: a gateway with no `pool` consumers can still expose
|
||||
inbound services (the pool route and DNS proxy still run, since they
|
||||
are part of the same binary, but they sit idle).
|
||||
|
||||
### NAT Pipeline (Inbound)
|
||||
|
||||
For each port-forward, a single prerouting DNAT rule matches
|
||||
mesh-originated traffic landing on the gateway's `fips0` address
|
||||
and rewrites it to the LAN target:
|
||||
|
||||
```text
|
||||
match: iifname == "fips0" && nfproto ipv6
|
||||
&& l4proto == <tcp|udp> && th dport == <listen_port>
|
||||
action: dnat to <target_ip>:<target_port>
|
||||
```
|
||||
|
||||
The match clause is deliberately narrow:
|
||||
|
||||
- **`iifname == "fips0"`** restricts the rule to traffic that
|
||||
arrived from the mesh. LAN-side ingress is never subject to
|
||||
inbound forwarding.
|
||||
- **`nfproto ipv6`** is enforced both here and at config-load time
|
||||
(`SocketAddrV6` rejects IPv4 targets); FIPS is IPv6-only end to
|
||||
end.
|
||||
- **`l4proto + dport`** narrows the match to one
|
||||
`(listen_port, proto)` pair per rule. Unique-tuple validation
|
||||
ensures no two rules contend for the same packet.
|
||||
|
||||
When *any* port-forward is configured, a single LAN-side masquerade
|
||||
is added to postrouting:
|
||||
|
||||
```text
|
||||
match: iifname == "fips0" && oifname == <lan_interface>
|
||||
&& nfproto ipv6
|
||||
action: masquerade
|
||||
```
|
||||
|
||||
Without this rule, the LAN target would attempt to reply directly to
|
||||
the mesh peer's `fd00::/8` source address, which is not reachable on
|
||||
the LAN. Masquerade rewrites the source to the gateway's LAN-side
|
||||
address so the target sees a reachable peer and conntrack routes
|
||||
the reply back through the gateway.
|
||||
|
||||
This LAN-side masquerade is independent of the `oifname fips0`
|
||||
masquerade in the outbound pipeline; the two have disjoint match
|
||||
clauses (different `iifname`/`oifname` combinations) and coexist
|
||||
without interaction when both directions are active.
|
||||
|
||||
### Independence From Outbound
|
||||
|
||||
The inbound half does not require:
|
||||
|
||||
- A virtual-IP pool. Mesh peers connect directly to the gateway's
|
||||
own `fips0` address, which the FIPS daemon already owns.
|
||||
- DNS resolution. Mesh peers reach the gateway as
|
||||
`<gateway-npub>.fips:<port>` using their own resolver (or a
|
||||
numeric mesh address); the gateway's DNS proxy is not in the path.
|
||||
- A daemon-side identity cache for the LAN target. The target is a
|
||||
LAN-side IPv6 address, not a mesh address; no `fd00::/8` lookup
|
||||
happens for it.
|
||||
|
||||
A gateway configured with port-forwards but with no LAN clients ever
|
||||
issuing `.fips` DNS queries will have an empty pool and zero
|
||||
outbound mappings, but its inbound forwards work normally. The
|
||||
inverse is also true: a gateway that serves only outbound LAN→mesh
|
||||
traffic has zero entries in the port-forwards list and no LAN-side
|
||||
masquerade.
|
||||
|
||||
## Atomic Table Rebuild (Common)
|
||||
|
||||
Both halves contribute rules to the same `inet fips_gateway` table,
|
||||
and that table is rebuilt as one unit on every state change —
|
||||
mapping added, mapping removed, port-forwards updated. The rebuild
|
||||
sequence is:
|
||||
|
||||
1. Delete the existing table in its own batch (ignore `ENOENT`).
|
||||
2. In a fresh batch: add the table; add the `prerouting` and
|
||||
`postrouting` chains; add the always-on `oifname fips0`
|
||||
masquerade; add per-mapping DNAT/SNAT rules for every active
|
||||
pool entry; add per-port-forward DNAT rules; add the LAN-side
|
||||
masquerade if any port-forwards exist.
|
||||
3. Send the batch as a single netlink transaction.
|
||||
|
||||
The rustables crate does not expose rule-handle tracking, so
|
||||
incremental update of individual rules is not available. Atomic
|
||||
rebuild was chosen for simplicity and correctness: it eliminates an
|
||||
entire class of partial-update inconsistency bugs at the cost of
|
||||
repeating the (cheap) rule construction on every change. The total
|
||||
rule count is bounded by the pool capacity (2 per mapping, capped
|
||||
at 2^16) and the port-forward count, both of which are small in
|
||||
practice.
|
||||
|
||||
## Configuration Reference
|
||||
|
||||
The full `gateway.*` block — pool CIDR, LAN interface, DNS
|
||||
listen/upstream/TTL, pool grace period, conntrack timeouts, and
|
||||
inbound port-forwards — is documented in the
|
||||
[Gateway section](../reference/configuration.md#gateway-gateway)
|
||||
of the configuration reference. The same block governs both halves;
|
||||
fields specific to one half (`pool`, `dns.*` for outbound;
|
||||
`port_forwards[]` for inbound) are simply unused when the other
|
||||
half is not in play.
|
||||
|
||||
## Operations and Troubleshooting
|
||||
|
||||
- [../tutorials/deploy-fips-gateway.md](../tutorials/deploy-fips-gateway.md)
|
||||
— end-to-end walkthrough on OpenWrt.
|
||||
- [../how-to/deploy-gateway.md](../how-to/deploy-gateway.md) —
|
||||
recipe for non-OpenWrt Linux hosts and inbound-port-forwarding
|
||||
configuration.
|
||||
- [../how-to/troubleshoot-gateway.md](../how-to/troubleshoot-gateway.md)
|
||||
— diagnostic recipes (DNS failures, ping working but TCP not,
|
||||
conntrack inspection, pool exhaustion, port-53 conflicts,
|
||||
port-forward verification).
|
||||
- [../reference/cli-fips-gateway.md](../reference/cli-fips-gateway.md)
|
||||
— command-line interface.
|
||||
- [../reference/control-socket.md](../reference/control-socket.md#gateway-command-catalog)
|
||||
— `show_gateway` and `show_mappings` commands.
|
||||
|
||||
## Security Considerations
|
||||
|
||||
### Outbound
|
||||
|
||||
- **LAN trust boundary.** The DNS listener and the virtual-IP pool
|
||||
are reachable by every host on the LAN. Any LAN host that can
|
||||
resolve `.fips` and route to the pool CIDR can reach mesh
|
||||
destinations. There is no per-client authentication; access
|
||||
restriction is a network-level concern, enforced with firewall
|
||||
rules on the LAN interface or on the gateway host itself.
|
||||
- **Identity masking.** All outbound LAN traffic appears on the
|
||||
mesh under the gateway's own FIPS identity. Mesh nodes cannot
|
||||
determine which LAN host originated a connection. This provides
|
||||
privacy for LAN hosts but means the gateway's reputation covers
|
||||
all of its clients — and that abusive behavior from one LAN host
|
||||
is attributed to the gateway, not to the host.
|
||||
- **Plaintext between client and gateway.** Traffic between the LAN
|
||||
client and the gateway is unencrypted at the IP layer. FIPS
|
||||
encryption (FSP) protects the segment between the gateway and the
|
||||
destination mesh node; application-layer encryption (TLS, SSH,
|
||||
Noise) is the only thing that provides true end-to-end protection
|
||||
through the gateway.
|
||||
- **Pool addresses are ephemeral.** Virtual IPs are allocated
|
||||
dynamically and recycled. They are not authenticated and not
|
||||
bound to client identity — a LAN host connecting to a virtual IP
|
||||
is trusting the gateway's recent DNS response.
|
||||
- **DNS upstream trust.** The outbound half's correctness depends
|
||||
on the FIPS daemon's resolver returning honest `fd00::/8`
|
||||
answers; a compromised daemon could redirect LAN clients to
|
||||
arbitrary mesh nodes.
|
||||
|
||||
### Inbound
|
||||
|
||||
- **Port exposure.** Each entry in `port_forwards[]` exposes the
|
||||
matched `(listen_port, proto)` on the gateway's mesh-side
|
||||
address to every reachable mesh peer. Inbound port-forwards are
|
||||
not gated by any peer ACL beyond what FMP normally enforces;
|
||||
treat them with the same care as a public-internet port forward.
|
||||
- **Mesh peer trust.** The LAN target sees connections that have
|
||||
been masqueraded to the gateway's LAN address. The target cannot
|
||||
distinguish one mesh peer from another, and there is no
|
||||
authenticated peer identity available to the LAN target — any
|
||||
application-layer authentication or rate-limiting must run on
|
||||
the target itself.
|
||||
- **Return-path masquerade exposes the gateway's LAN address.**
|
||||
The LAN-side masquerade rewrites the mesh peer's source to the
|
||||
gateway's LAN address. A malicious or buggy LAN target can use
|
||||
this to send unsolicited traffic back at the gateway, or to
|
||||
probe other LAN hosts via the gateway's network position; LAN
|
||||
segmentation (VLANs, host firewalls) is the right control.
|
||||
|
||||
### Common
|
||||
|
||||
- **No client identity verification.** The gateway authenticates
|
||||
neither LAN clients nor mesh peers beyond what the underlying
|
||||
layers already do — `fips0` ingress carries an FSP-authenticated
|
||||
payload, the LAN side is whoever the LAN admits.
|
||||
|
||||
## References
|
||||
|
||||
- [fips-ipv6-adapter.md](fips-ipv6-adapter.md) — IPv6 adapter and
|
||||
TUN interface design.
|
||||
- [fips-architecture.md](fips-architecture.md) — protocol layer
|
||||
architecture.
|
||||
- [fips-concepts.md](fips-concepts.md) — protocol overview.
|
||||
- [../reference/configuration.md](../reference/configuration.md) —
|
||||
configuration reference.
|
||||
- [../reference/cli-fips-gateway.md](../reference/cli-fips-gateway.md)
|
||||
— `fips-gateway` CLI.
|
||||
- [../reference/control-socket.md](../reference/control-socket.md) —
|
||||
control-socket protocol and command catalog.
|
||||
- [../how-to/deploy-gateway.md](../how-to/deploy-gateway.md) —
|
||||
gateway host and LAN client setup.
|
||||
- [../how-to/troubleshoot-gateway.md](../how-to/troubleshoot-gateway.md)
|
||||
— diagnostic recipes.
|
||||
- [../tutorials/deploy-fips-gateway.md](../tutorials/deploy-fips-gateway.md)
|
||||
— OpenWrt walkthrough.
|
||||
@@ -1,817 +0,0 @@
|
||||
# FIPS: Free Internetworking Peering System
|
||||
|
||||
## What is FIPS?
|
||||
|
||||
FIPS is a self-organizing mesh network that can operate natively over a
|
||||
variety of physical and logical media, such as local area networks,
|
||||
Bluetooth, serial links, or the existing internet as an overlay. The
|
||||
long-term goal is infrastructure that can function alongside or ultimately
|
||||
replace dependence on the Internet itself. Systems running FIPS establish
|
||||
peer connections, authenticate each other, and route traffic for each other
|
||||
without any central authority or global topology knowledge, and allow
|
||||
end-to-end encrypted sessions between any two nodes regardless of how many
|
||||
hops separate them.
|
||||
|
||||
Nodes in the mesh route traffic for each other using Nostr identities
|
||||
(npubs) as network addresses. Applications can access the mesh through a
|
||||
native FIPS datagram service, or through an IPv6 adaptation layer that
|
||||
presents each node as an IPv6 endpoint for compatibility with existing
|
||||
IP-based applications.
|
||||
|
||||
## Why FIPS?
|
||||
|
||||
**Self-sovereign identity**: FIPS nodes generate their own addresses, node
|
||||
IDs, and security credentials without coordination with any central
|
||||
authority. These identities can be long-term fixed or may be ephemeral,
|
||||
changed at any time. These identities are not visible to the FIPS network
|
||||
itself — they are used only at the application layer and for end-to-end
|
||||
session encryption.
|
||||
|
||||
**Infrastructure independence**: The internet depends on centralized
|
||||
infrastructure — ISPs, backbone providers, DNS, certificate authorities.
|
||||
FIPS works over any transport that can carry packets: a serial connection,
|
||||
onion-routed connections through Tor, local area networking, radio links
|
||||
between remote sites, or the existing internet as an overlay. When the
|
||||
internet is unavailable, unreliable, or untrusted, the mesh still works.
|
||||
|
||||
**Privacy by design**: FIPS provides secure, authenticated, and encrypted
|
||||
communication between any two nodes in the mesh, independent of the mix of
|
||||
transports used along the routed path between them. Furthermore, the mesh
|
||||
itself is designed to minimize metadata exposure — intermediate nodes route
|
||||
packets without learning the identities of the endpoints.
|
||||
|
||||
**Zero configuration**: Nodes discover each other and build routing
|
||||
automatically. Connect to one peer and you can reach the entire mesh. The
|
||||
network self-heals around failures and adapts to changing topology.
|
||||
|
||||
## A Self-Organizing Mesh
|
||||
|
||||
Traditional networks are built top-down. A central authority assigns
|
||||
addresses, configures routing tables, provisions hardware, and manages the
|
||||
topology. If the authority disappears or the infrastructure fails, the
|
||||
network fails with it. Nodes cannot reach each other without infrastructure
|
||||
mediating the connection.
|
||||
|
||||
FIPS inverts this model. There is no central authority, no address
|
||||
assignment service, no routing table pushed from above. Each node generates
|
||||
its own identity from a cryptographic keypair. Each node independently
|
||||
decides which peers to connect to and which transports to use. From these
|
||||
local decisions alone, the network self-organizes:
|
||||
|
||||
- A **spanning tree** forms through distributed parent selection, giving
|
||||
every node a coordinate in the network without any node knowing the full
|
||||
topology
|
||||
- **Bloom filters** propagate through gossip, so each node learns which
|
||||
peers can reach which destinations — again without global knowledge
|
||||
- **Routing decisions** are made locally at each hop, using only the node's
|
||||
immediate peers and cached coordinate information
|
||||
|
||||
Each peer link and end-to-end session actively measures RTT, loss, jitter,
|
||||
and goodput through a lightweight in-band Metrics Measurement Protocol
|
||||
(MMP), providing operator visibility and a foundation for quality-aware
|
||||
routing.
|
||||
|
||||
The result is a network that builds itself from the bottom up, heals around
|
||||
failures automatically, and scales without central coordination. Adding a
|
||||
node is as simple as connecting to one existing peer — the network
|
||||
integrates the new node through its normal mesh protocols.
|
||||
|
||||
## Specific Design Goals
|
||||
|
||||
- **Nostr-native identity and cryptography** — Use Nostr keypairs as node
|
||||
identities and leverage secp256k1, Schnorr signatures, and SHA-256
|
||||
- **Transport agnostic** — Support overlay, shared medium, and
|
||||
point-to-point transports transparently
|
||||
- **Self-organizing** — Automatic topology discovery and route optimization
|
||||
- **Privacy preserving** — Minimize metadata leakage across untrusted links
|
||||
- **Resilient** — Self-healing with graceful degradation
|
||||
|
||||
Non-goals include:
|
||||
|
||||
- **Reliable delivery** — FIPS provides a best-effort datagram service;
|
||||
retransmission and ordering are left to applications or higher-layer
|
||||
protocols
|
||||
- **Anonymity** — Direct peers learn each other's identity; FIPS minimizes
|
||||
metadata exposure but is not an anonymity network like Tor
|
||||
- **Congestion control** — FIPS measures link quality but does not implement
|
||||
flow control or congestion avoidance at the mesh layer
|
||||
|
||||
---
|
||||
|
||||
## Protocol Architecture
|
||||
|
||||
FIPS is organized in three protocol layers, each with distinct
|
||||
responsibilities and clean service boundaries. No layer depends on the
|
||||
specifics of the layers above or below it — transport plugins know nothing
|
||||
about sessions, the routing layer knows nothing about application addressing,
|
||||
and applications know nothing about which physical media carry their traffic.
|
||||
This separation means new transports, protocol features, and application
|
||||
interfaces can be added independently.
|
||||
|
||||

|
||||
|
||||
### Mapping to Traditional Networking
|
||||
|
||||
Readers familiar with the OSI model or TCP/IP networking may find it helpful
|
||||
to see how FIPS concepts relate to traditional layers:
|
||||
|
||||

|
||||
|
||||
Note that FMP spans what would traditionally be separate link and network
|
||||
layers. This is intentional — in a self-organizing mesh, the same layer that
|
||||
authenticates peers also makes routing decisions, because routing depends on
|
||||
authenticated peer state (spanning tree positions, bloom filters).
|
||||
|
||||
### Layer Responsibilities
|
||||
|
||||
**Transport layer**: Delivers datagrams between endpoints over a specific
|
||||
medium. Each transport type (UDP socket, Ethernet interface, radio modem)
|
||||
implements the same abstract interface: send and receive datagrams, report
|
||||
MTU. The transport layer knows nothing about FIPS identities, routing, or
|
||||
encryption. It provides raw datagram delivery to FMP above.
|
||||
|
||||
See [fips-transport-layer.md](fips-transport-layer.md) for the transport layer
|
||||
specification.
|
||||
|
||||
**FIPS Mesh Protocol (FMP)**: Manages peer connections, authenticates peers
|
||||
via Noise IK handshakes, and encrypts all traffic on each link. FMP is where
|
||||
the mesh organizes itself — nodes exchange spanning tree announcements and
|
||||
bloom filters with their direct peers, and FMP makes forwarding decisions
|
||||
for transit traffic. FMP provides authenticated, encrypted forwarding to FSP
|
||||
above.
|
||||
|
||||
See [fips-mesh-layer.md](fips-mesh-layer.md) for the FMP specification and
|
||||
[fips-mesh-operation.md](fips-mesh-operation.md) for how FMP's routing and
|
||||
self-organization work in practice.
|
||||
|
||||
**FIPS Session Protocol (FSP)**: Provides end-to-end authenticated
|
||||
encryption between any two nodes, regardless of how many intermediate hops
|
||||
separate them. FSP manages session lifecycle (setup, data transfer,
|
||||
teardown), caches destination coordinates for efficient routing, and handles
|
||||
the warmup strategy that keeps transit node caches populated. Session
|
||||
dispatch uses index-based routing inspired by
|
||||
[WireGuard](https://www.wireguard.com/), enabling O(1) packet
|
||||
demultiplexing. FSP provides a datagram service to applications above.
|
||||
|
||||
See [fips-session-layer.md](fips-session-layer.md) for the FSP specification.
|
||||
|
||||
**IPv6 adaptation layer**: Sits above FSP as a service on port 256, adapting
|
||||
the FIPS datagram service for unmodified IPv6 applications. Provides DNS
|
||||
resolution (npub → fd00::/8 address), identity cache management, IPv6 header
|
||||
compression, MTU enforcement, and a TUN interface. This is the primary way
|
||||
existing applications use the FIPS mesh.
|
||||
|
||||
See [fips-ipv6-adapter.md](fips-ipv6-adapter.md) for the IPv6 adapter.
|
||||
|
||||
### Node Architecture
|
||||
|
||||
Application services sit at the top of the stack, dispatched by FSP port
|
||||
number: the IPv6 TUN adapter (port 256) maps npubs to `fd00::/8` addresses
|
||||
with header compression so unmodified IP applications can use the network
|
||||
transparently, while the native datagram API addresses destinations directly
|
||||
by npub.
|
||||
|
||||

|
||||
|
||||
The mesh routes application traffic across heterogeneous transports
|
||||
transparently. A packet may traverse WiFi, Ethernet, UDP/IP, and Tor links
|
||||
on its way from source to destination — the application never needs to know
|
||||
which transports are involved. Each hop is independently encrypted at the
|
||||
link layer, while a single end-to-end session protects the payload across
|
||||
the entire path.
|
||||
|
||||

|
||||
|
||||
---
|
||||
|
||||
## Identity System
|
||||
|
||||
FIPS uses [Nostr](https://github.com/nostr-protocol/nips) keypairs
|
||||
(secp256k1) as node identities. The public key identifies the node; the
|
||||
private key signs protocol messages and establishes encrypted sessions.
|
||||
|
||||
The public key (or its bech32-encoded npub form) is the primary means for
|
||||
application-layer software to identify communication endpoints. Internally,
|
||||
the protocol derives a `node_addr` (a 16-byte SHA-256 hash of the pubkey)
|
||||
used as the routing identifier in packet headers, and an IPv6 address derived
|
||||
from the node_addr for the TUN adapter. Applications use the pubkey or npub;
|
||||
the routing layer uses node_addr; unmodified IPv6 applications use the
|
||||
derived `fd00::/8` address. All three are deterministically derived from the
|
||||
same keypair.
|
||||
|
||||
### FIPS Identity Handling
|
||||
|
||||

|
||||
|
||||
The pubkey is the node's cryptographic identity, used in Noise IK handshakes
|
||||
for both link and session encryption. It is never exposed beyond the
|
||||
endpoints of an encrypted channel. The node_addr, a one-way SHA-256 hash
|
||||
truncated to 16 bytes, serves as the routing identifier in packet headers
|
||||
and bloom filters. Intermediate routers see only node_addrs — they can
|
||||
forward traffic without learning the Nostr identities of the endpoints. An
|
||||
observer can verify "does this node_addr belong to pubkey X?" if they already
|
||||
know the pubkey, but cannot enumerate communicating identities by inspecting
|
||||
traffic. The IPv6
|
||||
address prepends `fd` to the first 15 bytes of the node_addr, providing a
|
||||
ULA overlay address for unmodified IP applications via the TUN interface.
|
||||
|
||||
Below the FIPS identity layer, each transport uses its own native addressing
|
||||
— IP:port or hostname:port addresses, MAC addresses, .onion identifiers. These **link
|
||||
addresses** are opaque to everything above FMP and discarded once link
|
||||
authentication completes.
|
||||
|
||||
### Identity Verification
|
||||
|
||||
The Noise Protocol Framework mutually authenticates both peer-to-peer link
|
||||
connections (at FMP) and end-to-end session traffic (at FSP), proving each
|
||||
party controls the private key for their claimed identity.
|
||||
|
||||
See [fips-mesh-layer.md](fips-mesh-layer.md) for peer authentication and
|
||||
[fips-session-layer.md](fips-session-layer.md) for end-to-end session
|
||||
establishment.
|
||||
|
||||
Key rotation changes the node's identity — a new keypair produces a new
|
||||
node_addr and IPv6 address, requiring all sessions to be re-established.
|
||||
Migration mechanisms that allow a node to announce a successor key are a
|
||||
future consideration.
|
||||
|
||||
---
|
||||
|
||||
## Two-Layer Encryption
|
||||
|
||||
FIPS uses independent encryption at two protocol layers:
|
||||
|
||||
| Layer | Scope | Pattern | Purpose |
|
||||
| ----- | ----- | ------- | ------- |
|
||||
| **FMP (Mesh)** | Hop-by-hop | Noise IK | Encrypt all traffic on each peer link |
|
||||
| **FSP (Session)** | End-to-end | Noise XK | Encrypt application payload between endpoints |
|
||||
|
||||
### Link Layer (Hop-by-Hop)
|
||||
|
||||
When two nodes establish a direct connection, they perform a [Noise
|
||||
IK](https://noiseprotocol.org/) handshake. This authenticates both parties
|
||||
and establishes symmetric keys for encrypting all traffic on that link.
|
||||
Every packet between direct peers is encrypted — gossip messages, routing
|
||||
queries, and forwarded session datagrams alike.
|
||||
|
||||
The IK pattern is used because outbound connections know the peer's npub
|
||||
from configuration, while inbound connections learn the initiator's identity
|
||||
from the first handshake message.
|
||||
|
||||
### Session Layer (End-to-End)
|
||||
|
||||
FIPS establishes end-to-end encrypted sessions between any two communicating
|
||||
nodes using Noise XK, regardless of how many hops separate them. The
|
||||
initiator knows the destination's npub (required for XK's pre-message);
|
||||
the responder learns the initiator's identity from the third handshake
|
||||
message. Unlike the link-layer IK pattern where the initiator's identity
|
||||
is revealed in msg1, XK delays identity disclosure until msg3, providing
|
||||
stronger initiator identity protection for traffic traversing untrusted
|
||||
intermediate nodes.
|
||||
|
||||
A packet from A to D through intermediate nodes B and C:
|
||||
|
||||
1. A encrypts payload with A↔D session key (FSP)
|
||||
2. A wraps in SessionDatagram, encrypts with A↔B link key (FMP), sends to B
|
||||
3. B decrypts link layer, reads destination node_addr, re-encrypts with B↔C
|
||||
link key, forwards to C
|
||||
4. C decrypts link layer, re-encrypts with C↔D link key, forwards to D
|
||||
5. D decrypts link layer, then decrypts session layer to get payload
|
||||
|
||||
Intermediate nodes route based on destination node_addr but cannot read
|
||||
session-layer payloads. Each hop strips one link encryption and applies the
|
||||
next — the session-layer ciphertext passes through untouched.
|
||||
|
||||
Both layers always apply, even between adjacent peers — a packet to a direct
|
||||
neighbor is still encrypted twice. This uniform model means no special cases
|
||||
for local vs remote destinations, and topology changes (a direct peer
|
||||
becomes reachable only through intermediaries) don't affect existing
|
||||
sessions.
|
||||
|
||||
See [fips-mesh-layer.md](fips-mesh-layer.md) for link encryption and
|
||||
[fips-session-layer.md](fips-session-layer.md) for session encryption.
|
||||
|
||||
---
|
||||
|
||||
## Routing and Mesh Operation
|
||||
|
||||
Each node makes forwarding decisions using only local information — its
|
||||
immediate peers, their bloom filters, and cached coordinates — rather than
|
||||
centrally distributed routing tables or global topology knowledge. Two
|
||||
complementary mechanisms provide the information each node needs.
|
||||
|
||||
### Spanning Tree: The Coordinate System
|
||||
|
||||

|
||||
|
||||
Nodes self-organize into a spanning tree through gossip — each node
|
||||
exchanges announcements with its direct peers and independently selects a
|
||||
parent. Because every node applies the same rule (prefer the root with the
|
||||
smallest node_addr), the network converges on a single agreed-upon root
|
||||
without any voting or coordination. This is the same principle behind the
|
||||
[Spanning Tree
|
||||
Protocol](https://en.wikipedia.org/wiki/Spanning_Tree_Protocol) used in
|
||||
Ethernet bridging since the 1980s: purely local decisions that converge to
|
||||
consistent global state. The resulting tree gives every node a
|
||||
**coordinate** — its path from itself to the root. Using tree coordinates
|
||||
for routing is adapted from
|
||||
[Yggdrasil](https://yggdrasil-network.github.io/)'s
|
||||
[Ironwood](https://github.com/Arceliar/ironwood) routing library.
|
||||
|
||||
These coordinates enable distance calculations between any two nodes: the
|
||||
distance is the number of hops from each node to their lowest common
|
||||
ancestor in the tree. This provides a metric for routing decisions without
|
||||
any node needing to know the full network topology.
|
||||
|
||||
The tree maintains itself through gossip — nodes exchange TreeAnnounce
|
||||
messages with their peers, propagating parent selections and ancestry
|
||||
chains. Changes cascade through the tree proportional to depth, not network
|
||||
size. If the network partitions, each segment converges to its own new root
|
||||
through the same process and reconverges automatically when segments rejoin.
|
||||
|
||||
See [fips-spanning-tree.md](fips-spanning-tree.md) for the tree algorithms
|
||||
and [spanning-tree-dynamics.md](spanning-tree-dynamics.md) for detailed
|
||||
convergence walkthroughs.
|
||||
|
||||
### Bloom Filters: Candidate Selection
|
||||
|
||||
The spanning tree provides a coordinate system for distance-based routing,
|
||||
but on its own each node would only know about its immediate neighbors.
|
||||
Bloom filters complement the tree by distributing reachability knowledge
|
||||
across the entire mesh — each node learns which destinations are reachable
|
||||
through which peers, without any node needing a complete view of the
|
||||
network.
|
||||
|
||||
Each node's peer-advertised [bloom
|
||||
filter](https://en.wikipedia.org/wiki/Bloom_filter) is a compact, fixed-size
|
||||
data structure that answers one question: "can this peer possibly reach
|
||||
destination D?" The answer is either "no" (definitive) or "maybe"
|
||||
(probabilistic — false positives are possible). Because the filter size is
|
||||
constant regardless of how many destinations it represents, bloom filters
|
||||
scale efficiently as the network grows. This is candidate selection for
|
||||
routing — bloom filters narrow the set of peers worth considering, and the
|
||||
actual forwarding decision ranks those candidates by tree distance and link
|
||||
quality.
|
||||
|
||||
Filters propagate transitively through tree edges, with each node computing
|
||||
outbound filters by merging the filters received from its tree peers (parent
|
||||
and children) using a
|
||||
[split-horizon](https://en.wikipedia.org/wiki/Split_horizon_route_advertisement)
|
||||
technique borrowed from distance-vector routing. All peers — including
|
||||
non-tree mesh shortcuts — receive FilterAnnounce messages, but only tree
|
||||
peers' filters are merged into outgoing computation. This prevents filter
|
||||
saturation where mesh shortcuts would cause every filter to converge toward
|
||||
the full network.
|
||||
|
||||
See [fips-bloom-filters.md](fips-bloom-filters.md) for filter parameters and
|
||||
mathematical properties.
|
||||
|
||||

|
||||
|
||||
The outbound filter for peer Q merges this node's identity with tree peer
|
||||
inbound filters except Q's (split-horizon exclusion). This creates
|
||||
directional asymmetry: upward filters (child → parent) contain the child's
|
||||
subtree, while downward filters (parent → child) contain the complement.
|
||||
Mesh peers receive filters but their inbound filters are not merged
|
||||
transitively — they provide single-hop shortcut visibility only.
|
||||
|
||||
A node with multiple peers receives genuinely different filters from each.
|
||||
In the diagram, R receives {B, D, E} from B and {C, F} from C — two disjoint
|
||||
subtrees. When R needs to reach F, only C's filter matches. This is where
|
||||
bloom filters provide real candidate selection: a node with several peers
|
||||
can narrow the forwarding choice before consulting tree coordinates. Leaf
|
||||
nodes like D have only one peer, so their single inbound filter is
|
||||
necessarily near-complete (everything except themselves) and offers no
|
||||
selection — but leaf nodes have no choice to make anyway.
|
||||
|
||||
Bloom filter sizing (bit count and hash functions) requires further analysis
|
||||
based on actual deployment scenarios. The FMP wire format is versioned to
|
||||
accommodate future parameter changes as operational experience accumulates.
|
||||
|
||||
### Routing Decisions
|
||||
|
||||
At each hop, FMP makes a local forwarding decision using the following
|
||||
priority chain:
|
||||
|
||||
1. **Local delivery** — the destination is this node
|
||||
2. **Direct peer** — the destination is an authenticated neighbor
|
||||
3. **Bloom-guided candidate selection** — bloom filters identify peers that
|
||||
can reach the destination; tree coordinates rank them by distance and
|
||||
link quality
|
||||
4. **[Greedy routing](https://en.wikipedia.org/wiki/Greedy_embedding)** —
|
||||
fallback when bloom filters haven't converged; forward to the peer that
|
||||
minimizes tree distance to the destination
|
||||
5. **No route** — destination unreachable; send error signal to source
|
||||
|
||||
All multi-hop routing depends on knowing the destination's tree coordinates.
|
||||
These are cached at each node after being learned through discovery
|
||||
(LookupRequest/LookupResponse) or session establishment (SessionSetup). The
|
||||
coordinate cache is the critical piece that enables efficient forwarding.
|
||||
|
||||

|
||||
|
||||
### Coordinate Caching and Discovery
|
||||
|
||||
When a node first needs to reach an unknown destination, it sends a
|
||||
LookupRequest that propagates through the network guided by bloom filters
|
||||
and loop prevention. The destination responds with its coordinates, which
|
||||
the source and intermediate nodes along the return path cache. Subsequent
|
||||
traffic routes efficiently using the cached coordinates.
|
||||
|
||||
Session establishment (SessionSetup) also carries coordinates, warming
|
||||
transit node caches along the path so that data packets can be forwarded
|
||||
without individual discovery at each hop.
|
||||
|
||||

|
||||
|
||||
### Error Recovery
|
||||
|
||||
When routing fails — because cached coordinates are stale, a path has
|
||||
broken, or a packet exceeds a link's MTU — transit nodes signal the source:
|
||||
|
||||
- **CoordsRequired**: A transit node lacks the destination's coordinates.
|
||||
The source re-initiates discovery and resets its coordinate warmup
|
||||
strategy.
|
||||
- **PathBroken**: Greedy routing reached a dead end. The source re-discovers
|
||||
the destination's current coordinates.
|
||||
- **MtuExceeded**: A transit node cannot forward a packet because it exceeds
|
||||
the next-hop link MTU. The source adjusts its path MTU estimate.
|
||||
|
||||
All three signals trigger active recovery, and are rate-limited to prevent
|
||||
storms during topology changes.
|
||||
|
||||
See [fips-mesh-operation.md](fips-mesh-operation.md) for the complete
|
||||
routing and mesh behavior description.
|
||||
|
||||
### Metrics Measurement Protocol (MMP)
|
||||
|
||||
Each peer link runs an instance of the Metrics Measurement Protocol, which
|
||||
measures link quality through in-band report exchange. MMP computes smoothed
|
||||
round-trip time (SRTT), packet loss rate, interarrival jitter, goodput, and
|
||||
one-way delay trend — all derived from counter and timestamp fields already
|
||||
present in the FMP wire format, with no additional probing traffic required.
|
||||
|
||||
MMP operates in three modes. **Full** mode exchanges both SenderReports and
|
||||
ReceiverReports to compute all metrics including RTT. **Lightweight** mode
|
||||
exchanges only ReceiverReports, providing loss and jitter but not RTT — useful
|
||||
for constrained links. **Minimal** mode disables reports entirely, relying
|
||||
only on spin bit and congestion echo flags in the frame header.
|
||||
|
||||
Reports are sent at RTT-adaptive intervals (clamped to 100 ms–2 s), so
|
||||
high-latency links don't generate excessive measurement traffic while
|
||||
low-latency links converge quickly. Each metric carries both short-term and
|
||||
long-term exponentially weighted moving averages, enabling detection of
|
||||
quality changes against a stable baseline.
|
||||
|
||||
MMP serves dual roles: operator visibility and cost-based parent selection.
|
||||
Periodic log lines report per-link RTT, loss, jitter, and goodput. MMP
|
||||
computes an Expected Transmission Count (ETX) from bidirectional delivery
|
||||
ratios, which feeds into cost-based parent selection where each node
|
||||
evaluates `effective_depth = depth + link_cost` using
|
||||
`link_cost = etx * (1.0 + srtt_ms / 100.0)`. ETX is not yet used in
|
||||
`find_next_hop()` candidate ranking for data forwarding.
|
||||
|
||||
See [fips-mesh-layer.md](fips-mesh-layer.md) for MMP operating modes, report
|
||||
scheduling, and the spin bit design.
|
||||
|
||||
---
|
||||
|
||||
## Transport Abstraction
|
||||
|
||||
FIPS treats the communication medium as a pluggable component. Every transport
|
||||
— whether a UDP socket, an Ethernet interface, a Tor circuit, or a radio modem
|
||||
— implements the same simple interface: send a datagram to an address, receive
|
||||
datagrams, and report the link MTU. The rest of the protocol stack sees no
|
||||
difference between them.
|
||||
|
||||
A **transport** is a driver for a particular medium. A **link** is a peer
|
||||
connection established over a transport. Transport addresses (IP:port, MAC
|
||||
address, .onion) are opaque to all layers above FMP — they exist only to
|
||||
deliver datagrams and are discarded once FMP has authenticated the peer via
|
||||
the Noise IK handshake. From that point on, the peer is identified solely by
|
||||
its cryptographic identity.
|
||||
|
||||
Transports fall into three categories based on their connectivity model:
|
||||
|
||||
| Category | Examples | Characteristics |
|
||||
| -------- | -------- | --------------- |
|
||||
| Overlay | UDP/IP, Tor | Tunnels FIPS over existing networks |
|
||||
| Shared medium | Ethernet, WiFi, Bluetooth, Radio | Local broadcast, peer discovery |
|
||||
| Point-to-point | Serial, dialup | Fixed connections, no discovery |
|
||||
|
||||
These categories differ in addressing, MTU, reliability, and whether they
|
||||
support local discovery, but FMP handles all of them uniformly. A node
|
||||
running multiple transports simultaneously bridges between those networks
|
||||
automatically — peers from all transports feed into a single spanning tree,
|
||||
and the router selects the best path regardless of which medium carries it.
|
||||
If one transport fails, traffic reroutes through alternatives without
|
||||
application involvement.
|
||||
|
||||
Some transports support an optional discovery capability — the ability to
|
||||
broadcast and listen for announcements indicating the availability of FIPS
|
||||
endpoints on the local medium. Shared media like Ethernet, WiFi, Bluetooth,
|
||||
and radio are natural fits for this, as they can reach nearby devices without
|
||||
prior configuration. When discovery is available, nodes can automatically
|
||||
find and peer with other FIPS nodes on the same medium. Transports that
|
||||
lack discovery (such as configured UDP endpoints) simply skip this step and
|
||||
connect directly to configured addresses. Additionally, endpoint discovery
|
||||
using Nostr relays and signed events is planned, allowing internet-reachable
|
||||
nodes to publish their transport addresses for other FIPS nodes to find.
|
||||
|
||||
NAT traversal is not currently addressed by the protocol.
|
||||
Internet-connected nodes behind NAT must be reachable through port
|
||||
forwarding, a publicly addressed peer, or relay through other mesh nodes.
|
||||
UDP hole punching and relay-assisted NAT traversal are potential future
|
||||
mechanisms but are not part of the current design.
|
||||
|
||||
> **Implementation status**: UDP/IP, TCP/IP, Ethernet, and Tor
|
||||
> (SOCKS5 outbound + directory-mode inbound via onion service)
|
||||
> transports are implemented. All others are future directions.
|
||||
|
||||
See [fips-transport-layer.md](fips-transport-layer.md) for the full transport
|
||||
layer specification.
|
||||
|
||||
---
|
||||
|
||||
## Security
|
||||
|
||||
FIPS is designed around four classes of adversary, each addressed by a
|
||||
different layer of the protocol.
|
||||
|
||||
### Transport Observers
|
||||
|
||||
A passive observer on the underlying transport — someone monitoring a WiFi
|
||||
network, tapping an Ethernet segment, or inspecting UDP traffic — sees only
|
||||
encrypted packets. The FMP link-layer Noise IK session encrypts all traffic
|
||||
between direct peers, including routing gossip and forwarded session
|
||||
datagrams. The observer can infer timing, packet sizes, and which transport
|
||||
endpoints are exchanging traffic, but cannot read content or determine
|
||||
FIPS-level node identities from the encrypted packets. Traffic analysis —
|
||||
correlating timing and volume patterns across multiple vantage points to
|
||||
infer communication relationships — is not defended against (see
|
||||
[Specific Design Goals](#specific-design-goals)).
|
||||
|
||||
### Active Attackers on the Transport
|
||||
|
||||
An adversary who can inject, modify, drop, or replay packets on the
|
||||
transport is also defeated by the FMP link-layer Noise IK session. Mutual
|
||||
authentication prevents impersonation, AEAD encryption detects tampering,
|
||||
and counter-based nonces with a sliding replay window reject replayed
|
||||
packets.
|
||||
|
||||
### Other FIPS Nodes (Intermediate Routers)
|
||||
|
||||
The most important adversary class is the operators of other nodes in the
|
||||
mesh — the peers that forward your traffic. FIPS treats every intermediate
|
||||
router as potentially adversarial. The FSP session layer establishes a
|
||||
completely independent Noise XK session between the communicating endpoints,
|
||||
so intermediate nodes cannot read application payloads even though they
|
||||
decrypt and re-encrypt the link-layer envelope at each hop.
|
||||
|
||||
Routing headers expose only the destination's node_addr — an opaque
|
||||
SHA-256 hash of the actual public key. Intermediate routers can forward
|
||||
traffic without learning which Nostr identities are communicating. An
|
||||
observer can verify "does this node_addr belong to pubkey X?" if they
|
||||
already know the pubkey, but cannot enumerate communicating identities by
|
||||
inspecting routed traffic.
|
||||
|
||||
| Entity | Can See |
|
||||
| ------ | ------- |
|
||||
| Transport observer | Encrypted packets, timing, packet sizes |
|
||||
| Direct peer | Your npub, traffic volume, timing |
|
||||
| Intermediate router | Source and destination node_addrs, packet size |
|
||||
| Destination | Your npub, payload content |
|
||||
|
||||
### Adversarial Nodes Disrupting the Mesh
|
||||
|
||||
Beyond passive observation, a malicious node could attempt to disrupt
|
||||
routing by injecting false spanning tree announcements, advertising bogus
|
||||
bloom filters, or claiming invalid tree positions. FMP mitigates these
|
||||
through signed TreeAnnounce messages verified by direct peers, transitive
|
||||
ancestry chain validation, replay protection via sequence numbers, and
|
||||
discretionary peering — node operators choose who to peer with, so an
|
||||
attacker with many identities still needs real nodes to accept their
|
||||
connections. Handshake rate limiting further constrains how fast an attacker
|
||||
can establish new links. In fully open networks with automatic peer
|
||||
discovery, Sybil resistance relies primarily on rate limiting; discretionary
|
||||
peering provides stronger resistance in curated deployments where operators
|
||||
vet their peers. An attacker who controls all of a target node's direct
|
||||
peers can completely control its view of the network (an eclipse attack);
|
||||
diverse peering across independent operators and transports is the primary
|
||||
mitigation.
|
||||
|
||||
---
|
||||
|
||||
## Prior Work
|
||||
|
||||
FIPS builds on proven designs rather than inventing new cryptography or routing
|
||||
algorithms. Nearly every major design decision has deployed precedent.
|
||||
|
||||
### Spanning Tree Self-Organization
|
||||
|
||||
The idea that distributed nodes can build a spanning tree through purely local
|
||||
decisions — each node selecting a parent based on announcements from its
|
||||
neighbors — dates to the
|
||||
[IEEE 802.1D Spanning Tree Protocol](https://en.wikipedia.org/wiki/Spanning_Tree_Protocol)
|
||||
(STP, 1985). STP demonstrated that a network-wide tree emerges from a simple
|
||||
deterministic rule (lowest bridge ID wins root election) applied independently
|
||||
at each node. FIPS uses the same principle — lowest node address determines the
|
||||
root — adapted from an Ethernet bridging context to a general-purpose overlay
|
||||
mesh.
|
||||
|
||||
### Tree Coordinate Routing
|
||||
|
||||
The spanning tree coordinates, bloom filter candidate selection, and greedy
|
||||
routing algorithms are adapted from
|
||||
[Yggdrasil v0.5](https://yggdrasil-network.github.io/2023/10/22/upcoming-v05-release.html)
|
||||
and its [Ironwood](https://github.com/Arceliar/ironwood) routing library.
|
||||
Yggdrasil's key insight was using the tree path from root to node as a
|
||||
routable coordinate, enabling greedy forwarding without global routing tables.
|
||||
FIPS adapts these algorithms for multi-transport operation, Nostr identity
|
||||
integration, and constrained MTU environments.
|
||||
|
||||
The theoretical foundation for greedy routing on tree embeddings draws on
|
||||
[Kleinberg's work](https://www.cs.cornell.edu/home/kleinber/swn.pdf) on
|
||||
navigable small-world networks, which showed that greedy forwarding succeeds
|
||||
in O(log² n) steps when the network has hierarchical structure. Thorup-Zwick
|
||||
compact routing schemes separately demonstrated that sublinear routing state
|
||||
is achievable with bounded stretch, motivating the use of tree coordinates
|
||||
rather than full routing tables.
|
||||
|
||||
### Split-Horizon Bloom Filter Propagation
|
||||
|
||||
FIPS distributes reachability information using bloom filters computed with a
|
||||
split-horizon rule: when advertising to a peer, exclude that peer's own
|
||||
contributions. This technique is borrowed from distance-vector routing
|
||||
protocols — [RIP](https://en.wikipedia.org/wiki/Routing_Information_Protocol)
|
||||
(1988) and [Babel](https://www.irif.fr/~jch/software/babel/) use split-horizon
|
||||
to prevent routing loops by not advertising a route back to the neighbor it was
|
||||
learned from. FIPS applies the same principle to probabilistic set
|
||||
advertisements rather than distance-vector tables.
|
||||
|
||||
### Cryptographic Identity as Network Address
|
||||
|
||||
FIPS nodes are identified by their Nostr public keys (secp256k1). The network
|
||||
address *is* the cryptographic identity — there is no separate address
|
||||
assignment or registration step.
|
||||
[CJDNS](https://github.com/cjdelisle/cjdns) pioneered this approach in
|
||||
overlay meshes, deriving IPv6 addresses from the double-SHA-512 of each node's
|
||||
public key. Tor [.onion addresses](https://spec.torproject.org/rend-spec-v3)
|
||||
and the IETF
|
||||
[Host Identity Protocol](https://en.wikipedia.org/wiki/Host_Identity_Protocol)
|
||||
(HIP) follow the same principle. FIPS uses Nostr's existing key infrastructure
|
||||
rather than introducing a new identity scheme.
|
||||
|
||||
### Dual-Layer Encryption
|
||||
|
||||
FIPS encrypts traffic twice: FMP provides hop-by-hop link encryption
|
||||
(protecting against transport-layer observers), while FSP provides independent
|
||||
end-to-end session encryption (protecting against intermediate FIPS nodes).
|
||||
This layered approach mirrors [Tor](https://www.torproject.org/), where each
|
||||
relay peels one layer of encryption (hop-by-hop) while the innermost layer
|
||||
protects end-to-end payload. [I2P](https://geti2p.net/) uses a similar
|
||||
garlic routing scheme with tunnel-layer and end-to-end encryption. Unlike Tor
|
||||
and I2P, FIPS does not provide anonymity — its dual encryption protects
|
||||
confidentiality and integrity rather than hiding traffic patterns.
|
||||
|
||||
### Noise Protocol Framework
|
||||
|
||||
FIPS uses the [Noise Protocol Framework](https://noiseprotocol.org/) at both
|
||||
protocol layers, with different handshake patterns chosen for each layer's
|
||||
threat model. FMP link encryption uses **Noise IK**, providing mutual
|
||||
authentication with a single round trip where the initiator knows the
|
||||
responder's static key in advance.
|
||||
[WireGuard](https://www.wireguard.com/) uses the same IK base pattern
|
||||
(extended with a pre-shared key as IKpsk2) for VPN tunnels. FSP session
|
||||
encryption uses **Noise XK**, the same pattern used by the
|
||||
[Lightning Network](https://github.com/lightning/bolts/blob/master/08-transport.md),
|
||||
where the initiator's static key is transmitted in a third message rather
|
||||
than the first. XK provides stronger initiator identity hiding at the cost
|
||||
of an additional round trip — a worthwhile tradeoff for session-layer traffic
|
||||
that traverses untrusted intermediate nodes. At the link layer, where both
|
||||
peers are configured and directly connected, IK's single round trip is
|
||||
preferred.
|
||||
|
||||
### Index-Based Session Dispatch
|
||||
|
||||
FIPS uses locally-assigned 32-bit session indices to demultiplex incoming
|
||||
packets to the correct cryptographic session in O(1) time, without parsing
|
||||
source addresses or performing expensive lookups. This directly follows
|
||||
[WireGuard's](https://www.wireguard.com/papers/wireguard.pdf) receiver index
|
||||
approach, where each peer assigns a random index during handshake and the
|
||||
remote side includes it in every packet header.
|
||||
|
||||
### Transport-Agnostic Overlay Mesh
|
||||
|
||||
FIPS is designed to operate over any datagram-capable transport — UDP, raw
|
||||
Ethernet, Bluetooth, radio, serial — through a uniform transport abstraction.
|
||||
Several mesh overlays have demonstrated transport-agnostic design:
|
||||
[CJDNS](https://github.com/cjdelisle/cjdns) runs over UDP and Ethernet,
|
||||
[Yggdrasil](https://yggdrasil-network.github.io/) supports TCP and TLS
|
||||
transports, and [Tor](https://www.torproject.org/) can use pluggable
|
||||
transports to tunnel through various media. FIPS extends this pattern to
|
||||
shared-medium transports (radio, BLE) with per-transport MTU and discovery
|
||||
capabilities.
|
||||
|
||||
### Metrics Measurement Protocol
|
||||
|
||||
MMP's design assembles well-established measurement techniques into a unified
|
||||
per-link protocol. The SenderReport/ReceiverReport exchange structure follows
|
||||
[RTCP](https://www.rfc-editor.org/rfc/rfc3550) (RFC 3550), which uses the
|
||||
same report pairing for media stream quality monitoring in RTP sessions. MMP's
|
||||
jitter computation uses the RTCP interarrival jitter algorithm directly.
|
||||
|
||||
The smoothed RTT estimator uses the Jacobson/Karels algorithm
|
||||
([RFC 6298](https://www.rfc-editor.org/rfc/rfc6298)), the same SRTT
|
||||
computation used in TCP for retransmission timeout calculation since 1988.
|
||||
MMP derives RTT from timestamp-echo in ReceiverReports with dwell-time
|
||||
compensation, rather than from packet round-trips.
|
||||
|
||||
The spin bit in the FMP frame header follows the
|
||||
[QUIC](https://www.rfc-editor.org/rfc/rfc9000) spin bit
|
||||
([RFC 9312](https://www.rfc-editor.org/rfc/rfc9312)) — a single bit that
|
||||
alternates each round trip, enabling passive latency measurement. FIPS
|
||||
implements the spin bit state machine but relies on timestamp-echo for SRTT,
|
||||
as irregular mesh traffic makes spin bit RTT unreliable.
|
||||
|
||||
The Expected Transmission Count (ETX) metric, computed from bidirectional
|
||||
delivery ratios, was introduced by
|
||||
[De Couto et al. (2003)](https://pdos.csail.mit.edu/papers/grid:mobicom03/paper.pdf)
|
||||
for wireless mesh routing and is used in protocols including
|
||||
[OLSR](https://en.wikipedia.org/wiki/Optimized_Link_State_Routing_Protocol)
|
||||
and [Babel](https://www.irif.fr/~jch/software/babel/). FIPS computes ETX
|
||||
per-link from MMP loss measurements for future use in candidate ranking.
|
||||
|
||||
The CE (Congestion Experienced) echo flag provides hop-by-hop
|
||||
[ECN](https://en.wikipedia.org/wiki/Explicit_Congestion_Notification)
|
||||
signaling, following the TCP/IP ECN echo pattern (RFC 3168). Transit nodes
|
||||
detect congestion via MMP loss/ETX metrics or kernel buffer drops and set
|
||||
the CE flag on forwarded frames; destination nodes mark ECN-capable IPv6
|
||||
packets accordingly.
|
||||
|
||||
### Cryptographic Primitives
|
||||
|
||||
FIPS reuses [Nostr's](https://github.com/nostr-protocol/nips) cryptographic
|
||||
stack — secp256k1 for identity keys, Schnorr signatures for authentication,
|
||||
SHA-256 for hashing, and ChaCha20-Poly1305 for authenticated encryption. This
|
||||
is the same primitive set used across Bitcoin, Nostr, and a growing ecosystem
|
||||
of self-sovereign identity systems. No novel cryptography is introduced.
|
||||
|
||||
---
|
||||
|
||||
## Further Reading
|
||||
|
||||
### Protocol Layers
|
||||
|
||||
| Document | Description |
|
||||
| -------- | ----------- |
|
||||
| [fips-transport-layer.md](fips-transport-layer.md) | Transport layer: abstraction, types, services provided to FMP |
|
||||
| [fips-mesh-layer.md](fips-mesh-layer.md) | FMP: peer authentication, link encryption, forwarding |
|
||||
| [fips-session-layer.md](fips-session-layer.md) | FSP: end-to-end encryption, session lifecycle |
|
||||
| [fips-ipv6-adapter.md](fips-ipv6-adapter.md) | IPv6 adaptation: DNS, TUN interface, MTU enforcement |
|
||||
|
||||
### Mesh Behavior and Wire Formats
|
||||
|
||||
| Document | Description |
|
||||
| -------- | ----------- |
|
||||
| [fips-mesh-operation.md](fips-mesh-operation.md) | How the mesh operates: routing, discovery, error recovery |
|
||||
| [fips-wire-formats.md](fips-wire-formats.md) | Complete wire format reference for all protocol layers |
|
||||
|
||||
### Supporting References
|
||||
|
||||
| Document | Description |
|
||||
| -------- | ----------- |
|
||||
| [fips-spanning-tree.md](fips-spanning-tree.md) | Spanning tree algorithms and data structures |
|
||||
| [fips-bloom-filters.md](fips-bloom-filters.md) | Bloom filter parameters, math, and computation |
|
||||
| [spanning-tree-dynamics.md](spanning-tree-dynamics.md) | Scenario walkthroughs: convergence, partitions, recovery |
|
||||
|
||||
### Implementation
|
||||
|
||||
| Document | Description |
|
||||
| -------- | ----------- |
|
||||
| [fips-configuration.md](fips-configuration.md) | YAML configuration reference |
|
||||
|
||||
### External References
|
||||
|
||||
- [IEEE 802.1D Spanning Tree Protocol](https://en.wikipedia.org/wiki/Spanning_Tree_Protocol)
|
||||
- [Yggdrasil Network](https://yggdrasil-network.github.io/)
|
||||
- [Yggdrasil v0.5 Release Notes](https://yggdrasil-network.github.io/2023/10/22/upcoming-v05-release.html)
|
||||
- [Ironwood Routing Library](https://github.com/Arceliar/ironwood)
|
||||
- [Kleinberg — The Small-World Phenomenon](https://www.cs.cornell.edu/home/kleinber/swn.pdf)
|
||||
- [CJDNS](https://github.com/cjdelisle/cjdns)
|
||||
- [Tor Project](https://www.torproject.org/)
|
||||
- [I2P](https://geti2p.net/)
|
||||
- [Host Identity Protocol (HIP)](https://en.wikipedia.org/wiki/Host_Identity_Protocol)
|
||||
- [Babel Routing Protocol](https://www.irif.fr/~jch/software/babel/)
|
||||
- [Noise Protocol Framework](https://noiseprotocol.org/)
|
||||
- [WireGuard](https://www.wireguard.com/)
|
||||
- [WireGuard Whitepaper](https://www.wireguard.com/papers/wireguard.pdf)
|
||||
- [Lightning Network BOLT #8 — Transport](https://github.com/lightning/bolts/blob/master/08-transport.md)
|
||||
- [QUIC (RFC 9000)](https://www.rfc-editor.org/rfc/rfc9000)
|
||||
- [QUIC Spin Bit (RFC 9312)](https://www.rfc-editor.org/rfc/rfc9312)
|
||||
- [RTCP (RFC 3550)](https://www.rfc-editor.org/rfc/rfc3550)
|
||||
- [TCP SRTT / RTO (RFC 6298)](https://www.rfc-editor.org/rfc/rfc6298)
|
||||
- [ECN (RFC 3168)](https://www.rfc-editor.org/rfc/rfc3168)
|
||||
- [ETX — De Couto et al. 2003](https://pdos.csail.mit.edu/papers/grid:mobicom03/paper.pdf)
|
||||
- [OLSR](https://en.wikipedia.org/wiki/Optimized_Link_State_Routing_Protocol)
|
||||
- [Nostr Protocol](https://github.com/nostr-protocol/nips)
|
||||
@@ -69,6 +69,24 @@ Known cache population mechanisms:
|
||||
- **Inbound traffic**: Authenticated sessions from other nodes populate the
|
||||
cache with their identity information
|
||||
|
||||
### Mesh-Interface Query Filter
|
||||
|
||||
The DNS responder is intended for local applications resolving `.fips`
|
||||
names; queries arriving over the mesh interface itself are dropped. The
|
||||
daemon records the index of the TUN interface at startup and compares
|
||||
it against the arrival interface of each incoming UDP DNS query. When
|
||||
they match — meaning the query came from another mesh node, not from a
|
||||
local socket — the responder discards the query without replying.
|
||||
|
||||
The check is implemented in
|
||||
[`is_mesh_interface_query`](../../src/upper/dns.rs) and prevents two
|
||||
classes of misbehaviour: a peer asking the daemon to resolve `.fips`
|
||||
names on its behalf (which would let one node use another as an
|
||||
identity-cache priming proxy), and accidental query loops where a
|
||||
misconfigured resolver forwards `.fips` queries back into the mesh.
|
||||
Local applications binding to the host's loopback or non-mesh
|
||||
interfaces are unaffected.
|
||||
|
||||
## IPv6 Address Derivation
|
||||
|
||||
FIPS addresses use the IPv6 Unique Local Address (ULA) prefix `fd00::/8`:
|
||||
@@ -125,41 +143,24 @@ entry hasn't been evicted by memory pressure.
|
||||
|
||||
## MTU Enforcement
|
||||
|
||||
FIPS does not provide fragmentation or reassembly at the session or mesh
|
||||
protocol layers — every datagram must fit in a single transport-layer packet.
|
||||
Some transports may perform fragmentation and reassembly internally (e.g., BLE
|
||||
L2CAP) and can advertise a larger virtual MTU than the physical medium
|
||||
supports, but this is transparent to FIPS. The mesh layer provides two
|
||||
facilities to manage MTU across heterogeneous paths: route discovery can
|
||||
constrain results to paths that support a required minimum MTU, and transit
|
||||
nodes that cannot forward an oversized datagram send an MtuExceeded error
|
||||
signal back to the source. The adapter must ensure that IPv6 packets from
|
||||
applications fit within the FIPS encapsulation budget after all layers of
|
||||
wrapping.
|
||||
The adapter sits at the boundary between the host's IPv6 stack and the
|
||||
FIPS encapsulation budget. Its job is to keep IPv6 packets small
|
||||
enough that they fit through the FIPS protocol envelope on every link
|
||||
along the path. The cross-cutting MTU model — proactive
|
||||
SessionDatagram `path_mtu` annotation, reactive MtuExceeded signals,
|
||||
end-to-end PathMtuNotification echo, and per-destination MTU storage
|
||||
— is documented in [fips-mtu.md](fips-mtu.md). What the adapter
|
||||
contributes is the IPv6-specific overhead accounting and the TUN-side
|
||||
enforcement integration.
|
||||
|
||||
### Encapsulation Overhead
|
||||
### IPv6-Specific Overhead
|
||||
|
||||
| Layer | Overhead | Purpose |
|
||||
| ----- | -------- | ------- |
|
||||
| Link encryption | 37 bytes | 16-byte outer header + 5-byte inner header (timestamp + msg_type) + 16-byte AEAD tag |
|
||||
| SessionDatagram body | 35 bytes | ttl + path_mtu + src_addr + dest_addr (msg_type counted in inner header) |
|
||||
| FSP header | 12 bytes | 4-byte prefix + 8-byte counter (used as AEAD AAD) |
|
||||
| FSP inner header | 6 bytes | 4-byte timestamp + 1-byte msg_type + 1-byte inner_flags (inside AEAD) |
|
||||
| Session AEAD tag | 16 bytes | ChaCha20-Poly1305 tag on session-encrypted payload |
|
||||
| **Protocol envelope** | **106 bytes** | `FIPS_OVERHEAD` constant |
|
||||
| Port header | 4 bytes | src_port + dst_port (DataPacket service dispatch) |
|
||||
| IPv6 compression | −33 bytes | 40-byte IPv6 header → 7-byte format + residual |
|
||||
| **IPv6 data path total** | **77 bytes** | `FIPS_IPV6_OVERHEAD` constant |
|
||||
|
||||
Coordinate piggybacking (CP flag) adds variable overhead: `2 + entries × 16`
|
||||
per coordinate, with both src and dst coords sent. The send path skips the
|
||||
CP flag if adding coords would exceed the transport MTU.
|
||||
|
||||
The `FIPS_OVERHEAD` constant (106 bytes) represents the base protocol
|
||||
envelope overhead (link encryption + routing + session encryption). For IPv6
|
||||
traffic, FSP port multiplexing adds 4 bytes (port header) while IPv6 header
|
||||
compression saves 33 bytes (40-byte header → 7-byte format + residual),
|
||||
yielding a net `FIPS_IPV6_OVERHEAD` of 77 bytes.
|
||||
For IPv6 traffic, FSP port multiplexing adds 4 bytes (port header)
|
||||
while IPv6 header compression saves 33 bytes (40-byte header →
|
||||
7-byte format + residual), yielding a net `FIPS_IPV6_OVERHEAD` of
|
||||
77 bytes on top of the base `FIPS_OVERHEAD` (106 bytes) protocol
|
||||
envelope. The full encapsulation breakdown lives in
|
||||
[fips-mtu.md](fips-mtu.md#encapsulation-overhead).
|
||||
|
||||
### Effective IPv6 MTU
|
||||
|
||||
@@ -183,48 +184,51 @@ transport path MTU for the IPv6 adapter is therefore:
|
||||
1280 + 77 = 1357 bytes
|
||||
```
|
||||
|
||||
Transports with smaller MTUs (radio at ~250 bytes, serial at 256 bytes) cannot
|
||||
support the IPv6 adapter without some form of internal fragmentation and
|
||||
reassembly. Otherwise, applications on those transports must use the native
|
||||
FIPS datagram API.
|
||||
Transports with smaller MTUs (radio at ~250 bytes, serial at 256
|
||||
bytes) cannot support the IPv6 adapter without some form of internal
|
||||
fragmentation and reassembly. Otherwise, applications on those
|
||||
transports must use the native FIPS datagram API.
|
||||
|
||||
### ICMP Packet Too Big
|
||||
### TUN-Side ICMP Packet Too Big
|
||||
|
||||
When an outbound packet at the TUN exceeds the effective IPv6 MTU, the adapter
|
||||
generates an ICMPv6 Packet Too Big message and delivers it back to the
|
||||
application via the TUN. This triggers the kernel's Path MTU Discovery (PMTUD)
|
||||
mechanism, which adjusts TCP segment sizes for subsequent transmissions.
|
||||
When an outbound packet at the TUN exceeds the effective IPv6 MTU,
|
||||
the adapter generates an ICMPv6 Packet Too Big message and delivers
|
||||
it back to the application via the TUN. This triggers the kernel's
|
||||
Path MTU Discovery mechanism, which adjusts TCP segment sizes for
|
||||
subsequent transmissions.
|
||||
|
||||
ICMP Packet Too Big generation is rate-limited per source address (100ms
|
||||
interval) to prevent storms from applications sending many oversized packets.
|
||||
ICMP Packet Too Big generation is rate-limited per source address
|
||||
(100ms interval) to prevent storms from applications sending many
|
||||
oversized packets. The ICMP response is delivered locally back through
|
||||
the TUN; no network traversal is needed, so delivery is reliable.
|
||||
|
||||
The ICMP response is delivered locally (back through the TUN to the kernel) —
|
||||
no network traversal is needed, so delivery is reliable.
|
||||
### TUN-Side TCP MSS Clamping
|
||||
|
||||
### TCP MSS Clamping
|
||||
|
||||
The adapter intercepts TCP SYN and SYN-ACK packets at the TUN interface and
|
||||
clamps the Maximum Segment Size (MSS) option:
|
||||
The adapter intercepts TCP SYN and SYN-ACK packets at the TUN
|
||||
interface and clamps the Maximum Segment Size (MSS) option:
|
||||
|
||||
```text
|
||||
clamped_mss = effective_ipv6_mtu - 40 (IPv6 header) - 20 (TCP header)
|
||||
```
|
||||
|
||||
This prevents TCP connections from negotiating segment sizes that would exceed
|
||||
the FIPS path MTU. Clamping is applied in two places:
|
||||
Clamping is applied in two places:
|
||||
|
||||
- **TUN reader** (outbound): Clamps MSS on outbound SYN packets
|
||||
- **TUN writer** (inbound): Clamps MSS on inbound SYN-ACK packets
|
||||
|
||||
Together, these ensure both directions of a TCP connection use appropriately
|
||||
sized segments from the start, avoiding the initial oversized packet loss
|
||||
that would occur with ICMP Packet Too Big alone.
|
||||
Together, these ensure both directions of a TCP connection use
|
||||
appropriately sized segments from the start, avoiding the initial
|
||||
oversized packet loss that would occur with ICMP Packet Too Big
|
||||
alone. The conditional clamp (per-flow lookup with cold-flow
|
||||
fallback) and the rationale for `max_mss` semantics are in
|
||||
[fips-mtu.md](fips-mtu.md#tcp-mss-clamping).
|
||||
|
||||
### ICMP Rate Limiting
|
||||
|
||||
ICMPv6 error generation is rate-limited per source address using a token bucket
|
||||
(100ms interval). This matches the standard ICMP rate limiting approach and
|
||||
prevents amplification when an application sends a burst of oversized packets.
|
||||
ICMPv6 error generation is rate-limited per source address using a
|
||||
token bucket (100ms interval). This matches the standard ICMP rate
|
||||
limiting approach and prevents amplification when an application sends
|
||||
a burst of oversized packets.
|
||||
|
||||
## TUN Interface
|
||||
|
||||
@@ -287,20 +291,44 @@ path.
|
||||
|
||||
### Configuration
|
||||
|
||||
```yaml
|
||||
tun:
|
||||
enabled: true
|
||||
name: fips0
|
||||
mtu: 1280
|
||||
```
|
||||
The TUN block (`tun.*`) is documented in
|
||||
[../reference/configuration.md](../reference/configuration.md).
|
||||
|
||||
### Privileges
|
||||
|
||||
TUN device creation requires `CAP_NET_ADMIN`. Options:
|
||||
TUN device creation requires `CAP_NET_ADMIN`. The shipped Debian
|
||||
systemd unit runs the daemon as `root` by default; for the
|
||||
alternative — running under a dedicated unprivileged service
|
||||
account with the capability granted on the binary — see
|
||||
[../how-to/run-as-unprivileged-user.md](../how-to/run-as-unprivileged-user.md).
|
||||
|
||||
- Run as root
|
||||
- Set capability: `sudo setcap cap_net_admin+ep ./target/debug/fips`
|
||||
- Pre-created persistent TUN device
|
||||
### App-Owned TUN (embedded hosts)
|
||||
|
||||
On platforms where FIPS is embedded rather than run as a daemon — notably
|
||||
Android, where the `VpnService` owns the TUN fd and the app has no
|
||||
`CAP_NET_ADMIN` — FIPS does not create `fips0` itself. Instead the embedder owns
|
||||
the fd and exchanges IPv6 packet bytes with FIPS over channels.
|
||||
|
||||
`Node::enable_app_owned_tun()` sets this up. It is called after `Node::new` and
|
||||
before `start()` (and before the node is moved into a background task), mirroring
|
||||
`control_read_handle()`, and returns two app-side channel ends:
|
||||
|
||||
- **app → mesh** — the embedder pushes IPv6 packets read from its fd into
|
||||
`app_outbound_tx`. These are drained by `run_rx_loop` into `handle_tun_outbound`
|
||||
and routed exactly as the Reader Thread's output would be.
|
||||
- **mesh → app** — inbound mesh traffic on port 256 is reconstructed and written
|
||||
to the node's `tun_tx` (the same sink the Writer Thread reads); the embedder
|
||||
pulls from `app_inbound_rx` and writes to its fd.
|
||||
|
||||
With the channels installed, `start()` skips system-TUN creation (it gates on
|
||||
`tun_tx` being unset), so FIPS does no `CAP_NET_ADMIN` operations.
|
||||
|
||||
Because packets enter via `app_outbound_tx` rather than the Reader Thread, they
|
||||
**bypass `handle_tun_packet`** — the `fd00::/8` destination filter, the ICMPv6
|
||||
Destination Unreachable for off-mesh dests (see [Reader Thread](#reader-thread)),
|
||||
and the [TUN-Side TCP MSS Clamping](#tun-side-tcp-mss-clamping). The embedder is
|
||||
therefore responsible for routing only `fd00::/8` to its TUN (so only mesh-bound
|
||||
packets arrive) and for clamping TCP MSS on outbound SYNs.
|
||||
|
||||
## Implementation Status
|
||||
|
||||
@@ -314,6 +342,7 @@ TUN device creation requires `CAP_NET_ADMIN`. Options:
|
||||
| ICMP rate limiting (per-source) | **Implemented** |
|
||||
| TCP MSS clamping (SYN + SYN-ACK) | **Implemented** |
|
||||
| DNS service (.fips domain) | **Implemented** |
|
||||
| DNS responder mesh-interface filter | **Implemented** |
|
||||
| Port-based service multiplexing (port 256) | **Implemented** |
|
||||
| IPv6 header compression (format 0x00) | **Implemented** |
|
||||
| Per-destination route MTU (netlink) | Planned |
|
||||
@@ -324,38 +353,27 @@ TUN device creation requires `CAP_NET_ADMIN`. Options:
|
||||
|
||||
## Design Considerations
|
||||
|
||||
### Path MTU Discovery
|
||||
### Path MTU Discovery and No-Fragmentation Policy
|
||||
|
||||
Two complementary mechanisms support full PMTUD:
|
||||
|
||||
1. **Proactive**: The `path_mtu` field (2 bytes) in the SessionDatagram envelope
|
||||
is implemented at the FMP level. The source sets it to its outbound link MTU
|
||||
minus overhead; each transit node applies
|
||||
`min(current, own_outbound_mtu - overhead)`. The destination receives the
|
||||
forward-path minimum. PathMtuNotification is handled at the session layer;
|
||||
the destination sends the observed forward-path MTU back to the source,
|
||||
which applies it with decrease-immediate / increase-requires-3-consecutive
|
||||
hysteresis.
|
||||
|
||||
2. **Reactive**: When a transit node cannot forward a packet (MTU exceeded), it
|
||||
sends an error signal back to the source. This handles the in-flight gap
|
||||
between a path MTU decrease and the source learning via the echo.
|
||||
|
||||
Both are needed: proactive handles steady state; reactive handles the transient
|
||||
window when oversized packets hit a new bottleneck before the source adapts.
|
||||
|
||||
### No Fragmentation
|
||||
|
||||
FIPS remains a pure datagram service with no fragmentation at transit nodes.
|
||||
Session-layer encryption is end-to-end — the AEAD tag authenticates the entire
|
||||
plaintext. Fragmenting encrypted datagrams would require either exposing
|
||||
plaintext structure to transit nodes (unacceptable) or reassembly before
|
||||
decryption (opens attack surface).
|
||||
Path MTU Discovery (proactive `path_mtu` annotation, reactive
|
||||
MtuExceeded, end-to-end PathMtuNotification) and the no-fragmentation
|
||||
policy that drives the design both live in the unified MTU treatment
|
||||
at [fips-mtu.md](fips-mtu.md). The adapter is a consumer of that
|
||||
model — its job is to enforce the resulting effective IPv6 MTU at the
|
||||
TUN with ICMP Packet Too Big and TCP MSS clamping.
|
||||
|
||||
## References
|
||||
|
||||
- [fips-intro.md](fips-intro.md) — Protocol overview and architecture
|
||||
- [fips-concepts.md](fips-concepts.md) — Protocol overview
|
||||
- [fips-architecture.md](fips-architecture.md) — Layer architecture and
|
||||
identity model
|
||||
- [fips-session-layer.md](fips-session-layer.md) — FSP (below the adapter)
|
||||
- [fips-wire-formats.md](fips-wire-formats.md) — FSP and SessionDatagram wire
|
||||
formats
|
||||
- [fips-configuration.md](fips-configuration.md) — TUN configuration parameters
|
||||
- [fips-mtu.md](fips-mtu.md) — Unified path MTU model (proactive,
|
||||
reactive, hysteresis, no-fragmentation)
|
||||
- [../reference/wire-formats.md](../reference/wire-formats.md) — FSP and
|
||||
SessionDatagram wire formats
|
||||
- [../reference/configuration.md](../reference/configuration.md) — TUN
|
||||
configuration parameters
|
||||
- [../how-to/run-as-unprivileged-user.md](../how-to/run-as-unprivileged-user.md)
|
||||
— privilege options for the daemon, including the unprivileged
|
||||
service-account path
|
||||
|
||||
@@ -215,8 +215,8 @@ The plaintext inside the encrypted frame begins with a 5-byte inner header
|
||||
(4-byte session-relative timestamp followed by a message type byte), then the
|
||||
message-specific payload.
|
||||
|
||||
See [fips-wire-formats.md](fips-wire-formats.md) for the complete wire format
|
||||
specification.
|
||||
See [../reference/wire-formats.md](../reference/wire-formats.md) for
|
||||
the complete wire format specification.
|
||||
|
||||
### What Encryption Provides
|
||||
|
||||
@@ -292,6 +292,12 @@ Roaming is most useful for UDP, where source addresses can change due to NAT
|
||||
rebinding or network changes. For connection-oriented transports, "roaming"
|
||||
manifests as reconnection rather than mid-session address change.
|
||||
|
||||
Roaming addresses *mid-session* NAT rebinding. Establishing the initial UDP
|
||||
path through NAT is a separate concern, addressed by the optional
|
||||
Nostr-mediated overlay discovery and STUN-assisted hole punching feature
|
||||
(see [fips-transport-layer.md](fips-transport-layer.md) and
|
||||
[../reference/configuration.md](../reference/configuration.md)).
|
||||
|
||||
## Replay Protection
|
||||
|
||||
Each link session maintains per-direction counters:
|
||||
@@ -335,7 +341,9 @@ Additional protections:
|
||||
memory usage
|
||||
- **Handshake timeout**: Stale pending handshakes are cleaned up after a
|
||||
configurable timeout
|
||||
- **Allowlist/blocklist**: Optional peer filtering before handshake processing
|
||||
- **Peer ACL**: Optional allowlist / denylist filtering of peer npubs
|
||||
before handshake processing (loaded from `/etc/fips/peers.allow` and
|
||||
`/etc/fips/peers.deny`, mtime-watched and reloaded automatically)
|
||||
|
||||
## Disconnect
|
||||
|
||||
@@ -355,6 +363,72 @@ links.
|
||||
On node shutdown, Disconnect is sent to all active peers before transports are
|
||||
stopped.
|
||||
|
||||
## Rekey
|
||||
|
||||
FMP periodically negotiates a fresh Noise session over each established link
|
||||
to bound forward-secrecy exposure: limiting AEAD nonce reuse risk, bounding
|
||||
the volume of ciphertext recoverable from a stolen long-term static key, and
|
||||
rotating the session indices that addressed packets carry on the wire.
|
||||
|
||||
A rekey is initiated when either threshold is reached on the link's current
|
||||
session: `node.rekey.after_secs` (default 120) elapsed since the link came
|
||||
up or last rekeyed, or `node.rekey.after_messages` (default 65536) frames
|
||||
sent. Either side can be the initiator independently. Rekey is on by
|
||||
default and can be disabled via `node.rekey.enabled: false` (the
|
||||
configuration tree is documented in
|
||||
[../reference/configuration.md](../reference/configuration.md)).
|
||||
|
||||
### Mechanism
|
||||
|
||||
A rekey reuses the Noise IK pattern of the initial handshake, but the two
|
||||
messages travel over the existing link as ordinary encrypted FMP frames
|
||||
rather than as plaintext bootstrap packets. The initiator builds a fresh
|
||||
`HandshakeState`, generates msg1, and sends it through the current session;
|
||||
the responder consumes msg1, builds msg2, and replies. After both sides
|
||||
have exchanged messages and finalised the new keys, traffic transitions
|
||||
from the old session to the new one.
|
||||
|
||||
Cutover is signalled in-band by the **K-bit** in the FMP flags byte. Each
|
||||
side starts emitting frames under the new session with K set; on receipt
|
||||
of the first K-marked frame the peer accepts the cutover and follows
|
||||
suit. A new pair of session indices is allocated as part of the new
|
||||
session, replacing the old indices on subsequent frames (see
|
||||
[Index Properties](#index-properties)).
|
||||
|
||||
### Drain Window
|
||||
|
||||
To absorb in-flight reordering across the cutover, the old session is not
|
||||
discarded immediately. Each peer retains it in a `previous_session` slot
|
||||
on the active-peer state for `DRAIN_WINDOW_SECS = 10` seconds (a
|
||||
compile-time constant in `src/node/handlers/rekey.rs`). During the
|
||||
window, decrypt attempts fall back to `previous_session` when the new
|
||||
session rejects a frame, so a packet sent under the old keys that
|
||||
arrives a few hundred milliseconds late still decrypts. After the
|
||||
window expires, the old session is dropped.
|
||||
|
||||
### Dual-Initiation Race
|
||||
|
||||
On high-latency links, both sides' rekey timers can fire close enough
|
||||
together that each peer's msg1 crosses the other in flight. Without
|
||||
arbitration, each side would act as both initiator and responder, end
|
||||
up with two different Noise sessions, and lose connectivity at cutover.
|
||||
FMP arbitrates with a deterministic tie-breaker: the peer with the
|
||||
**numerically smaller `NodeAddr`** wins the role of initiator and
|
||||
discards any inbound msg1 it sees during the race; the larger-`NodeAddr`
|
||||
peer abandons its own initiation and processes the inbound msg1 as
|
||||
responder. The same tie-breaker is applied to cross-connection races
|
||||
during initial handshake.
|
||||
|
||||
### Operator Visibility
|
||||
|
||||
Successful cutover is reported at INFO level on the K-bit observation;
|
||||
intermediate steps (handshake start, msg1/msg2 exchange, drain-window
|
||||
fallback decrypts) log at DEBUG/TRACE. Failures (handshake error,
|
||||
drain-window expiry without cutover) log at WARN.
|
||||
|
||||
The end-to-end rekey at the session layer follows a parallel design;
|
||||
see [fips-session-layer.md](fips-session-layer.md).
|
||||
|
||||
## Liveness Detection
|
||||
|
||||
FMP detects link liveness through a combination of explicit heartbeats and
|
||||
@@ -362,8 +436,8 @@ traffic observation.
|
||||
|
||||
### Heartbeat
|
||||
|
||||
A Heartbeat message (0x51) is sent to each active peer every
|
||||
`node.heartbeat_interval_secs` (default 10s). The heartbeat is a minimal
|
||||
A Heartbeat message (0x51) is sent to each active peer at a configurable
|
||||
interval (`node.heartbeat_interval_secs`). The heartbeat is a minimal
|
||||
encrypted frame with no payload beyond the standard inner header (timestamp +
|
||||
message type). Any successfully decrypted frame — data, gossip, MMP report,
|
||||
or heartbeat — resets the peer's last-receive timestamp tracked by the MMP
|
||||
@@ -371,8 +445,8 @@ receiver.
|
||||
|
||||
### Dead Timeout
|
||||
|
||||
When no traffic (of any kind) is received from a peer for
|
||||
`node.link_dead_timeout_secs` (default 30s), the peer is declared dead and
|
||||
When no traffic (of any kind) is received from a peer for the
|
||||
configured `node.link_dead_timeout_secs` window, the peer is declared dead and
|
||||
removed via `remove_active_peer()`. This triggers the full teardown cascade:
|
||||
spanning tree parent reselection (if the dead peer was the parent),
|
||||
TreeAnnounce propagation, coordinate cache flush, and bloom filter recompute.
|
||||
@@ -380,156 +454,74 @@ TreeAnnounce propagation, coordinate cache flush, and bloom filter recompute.
|
||||
If the dead peer is eligible for auto-reconnect (see [Auto-Reconnect]
|
||||
(#auto-reconnect)), reconnection is scheduled immediately after removal.
|
||||
|
||||
The heartbeat is independent of MMP — it is needed because idle links in
|
||||
Lightweight MMP mode have no guaranteed periodic traffic (gossip is
|
||||
event-driven, and MMP reports require at least one side running Full mode).
|
||||
The heartbeat is independent of MMP. Gossip is event-driven, Lightweight
|
||||
produces receiver reports only when traffic arrives, and Minimal emits no
|
||||
reports at all — so on a fully idle link no MMP-mode combination
|
||||
guarantees periodic activity. The heartbeat is the always-on liveness
|
||||
signal.
|
||||
|
||||
## Link Message Types
|
||||
|
||||
FMP defines eight message types carried inside encrypted frames:
|
||||
FMP defines several encrypted message types carried inside the
|
||||
established-frame envelope. They group naturally by purpose:
|
||||
|
||||
| Type | Name | Purpose |
|
||||
| ---- | ---- | ------- |
|
||||
| 0x10 | TreeAnnounce | Spanning tree state announcements between peers |
|
||||
| 0x20 | FilterAnnounce | Bloom filter reachability updates |
|
||||
| 0x30 | LookupRequest | Coordinate discovery — flood toward destination |
|
||||
| 0x31 | LookupResponse | Coordinate discovery — response with coordinates |
|
||||
| 0x00 | SessionDatagram | Encapsulated session-layer payload for forwarding |
|
||||
| 0x01 | SenderReport | MMP sender-side metrics report |
|
||||
| 0x02 | ReceiverReport | MMP receiver-side metrics report |
|
||||
| 0x50 | Disconnect | Orderly link teardown with reason code |
|
||||
| 0x51 | Heartbeat | Link liveness probe |
|
||||
- **Routing gossip**: TreeAnnounce carries spanning-tree announcements
|
||||
between direct peers; FilterAnnounce carries bloom-filter
|
||||
reachability updates between direct peers. Both are peer-to-peer
|
||||
(not forwarded).
|
||||
- **Discovery**: LookupRequest is forwarded through tree peers under
|
||||
bloom-filter guidance to find a destination's coordinates;
|
||||
LookupResponse routes back to the requester via reverse-path lookup
|
||||
in `recent_requests`.
|
||||
- **Forwarded payload**: SessionDatagram carries a session-layer
|
||||
payload hop-by-hop toward the destination.
|
||||
- **Metrics**: SenderReport and ReceiverReport carry the link-layer
|
||||
MMP report stream peer-to-peer.
|
||||
- **Liveness and lifecycle**: Heartbeat is a minimal frame sent
|
||||
peer-to-peer to keep the link alive; Disconnect carries an orderly
|
||||
teardown reason code peer-to-peer.
|
||||
|
||||
Additionally, handshake messages (phase 0x1 msg1, phase 0x2 msg2) are sent
|
||||
unencrypted before the link session is established.
|
||||
Handshake messages (phase 0x1 msg1, phase 0x2 msg2) travel before
|
||||
encryption is established and are identified by the FMP common-prefix
|
||||
`phase` field rather than a `msg_type` byte.
|
||||
|
||||
TreeAnnounce and FilterAnnounce are exchanged between direct peers only — they
|
||||
are not forwarded. LookupRequest and LookupResponse are forwarded through the
|
||||
mesh (flooded with deduplication). SessionDatagram is forwarded hop-by-hop
|
||||
toward the destination. Disconnect is peer-to-peer.
|
||||
|
||||
See [fips-mesh-operation.md](fips-mesh-operation.md) for how these messages
|
||||
work together to build and maintain the mesh, and
|
||||
[fips-wire-formats.md](fips-wire-formats.md) for byte-level message layouts.
|
||||
See [../reference/wire-formats.md](../reference/wire-formats.md) for
|
||||
byte-level message layouts and the canonical FMP message type
|
||||
catalog, and [fips-mesh-operation.md](fips-mesh-operation.md) for how
|
||||
these messages work together to build and maintain the mesh.
|
||||
|
||||
## Metrics Measurement Protocol (MMP)
|
||||
|
||||
Each active peer link runs an instance of the Metrics Measurement Protocol,
|
||||
providing per-link quality metrics to the operator and to the spanning tree
|
||||
layer for cost-based parent selection.
|
||||
MMP runs on every active link to provide per-link quality metrics
|
||||
(SRTT, loss, jitter, goodput, OWD trend, ETX) to the operator and to
|
||||
the spanning tree layer for cost-based parent selection. Reports are
|
||||
exchanged peer-to-peer between direct neighbors at RTT-adaptive
|
||||
intervals clamped to `[1s, 5s]`, with a 200 ms cold-start floor for
|
||||
the first five SRTT samples.
|
||||
|
||||
### Metrics Tracked
|
||||
The CE (Congestion Experienced) bit in the FMP flags byte carries
|
||||
hop-by-hop ECN signaling: transit nodes detect congestion on outgoing
|
||||
links (via MMP loss/ETX or `SO_RXQ_OVFL` kernel drops) and set CE on
|
||||
forwarded packets, which the destination then mirrors to the IPv6
|
||||
Traffic Class for ECN-capable flows.
|
||||
|
||||
MMP computes the following metrics from the per-frame counter and timestamp
|
||||
fields in the FMP wire format:
|
||||
|
||||
- **SRTT** — Smoothed round-trip time (Jacobson/RFC 6298, α=1/8). Derived
|
||||
from timestamp-echo in ReceiverReports with dwell-time compensation.
|
||||
- **Loss rate** — Bidirectional loss inferred from counter gaps. Tracked as
|
||||
both instantaneous (per-interval) and long-term EWMA.
|
||||
- **Jitter** — Interarrival jitter (RFC 3550 algorithm) in microseconds.
|
||||
- **Goodput** — Bytes per second of payload data (excludes MMP reports).
|
||||
- **OWD trend** — One-way delay trend (µs/s, signed). Indicates congestion
|
||||
buildup before loss occurs.
|
||||
- **ETX** — Expected Transmission Count, computed from bidirectional delivery
|
||||
ratios. Used in cost-based parent selection via
|
||||
`link_cost = etx * (1.0 + srtt_ms / 100.0)`; not yet used in
|
||||
`find_next_hop()` candidate ranking.
|
||||
- **Dual EWMA trends** — Short-term (α=1/4) and long-term (α=1/32) trend
|
||||
indicators for both RTT and loss, enabling change detection.
|
||||
|
||||
### Operating Modes
|
||||
|
||||
MMP supports three modes, configured via `node.mmp.mode`:
|
||||
|
||||
| Mode | Reports Exchanged | Metrics Available |
|
||||
| ---- | ----------------- | ----------------- |
|
||||
| **Full** (default) | SenderReport + ReceiverReport | All metrics including RTT, loss, jitter, goodput, OWD trend |
|
||||
| **Lightweight** | ReceiverReport only | Loss (from counter gaps), jitter, OWD trend. No RTT. |
|
||||
| **Minimal** | None | Spin bit and CE echo flags only. No computed metrics. |
|
||||
|
||||
### Report Scheduling
|
||||
|
||||
Reports are sent at RTT-adaptive intervals, clamped to [100ms, 2s]. A
|
||||
cold-start interval of 500ms is used before SRTT converges. The interval
|
||||
formula is `clamp(2 × SRTT, 100ms, 2000ms)`.
|
||||
|
||||
### Spin Bit and RTT
|
||||
|
||||
The SP (spin bit) flag in the FMP inner header follows the QUIC spin bit
|
||||
pattern: reflected on receive, toggled on send when the reflected value
|
||||
matches the last sent value. The spin bit state machine runs for TX
|
||||
reflection, but **RTT samples from the spin bit are discarded**. In a mesh
|
||||
protocol where frames are sent irregularly (tree announces, bloom filters,
|
||||
MMP reports on different timers), inter-frame processing delays inflate spin
|
||||
bit RTT measurements unpredictably. Timestamp-echo from ReceiverReports
|
||||
(with dwell-time compensation) is the sole SRTT source.
|
||||
|
||||
### ECN Congestion Signaling
|
||||
|
||||
The CE (Congestion Experienced) flag (bit 1 in the FMP flags byte) provides
|
||||
hop-by-hop congestion signaling through the mesh. Transit nodes detect
|
||||
congestion on outgoing links and set CE on forwarded packets; once set, the
|
||||
flag stays set for all subsequent hops to the destination.
|
||||
|
||||
**Congestion detection** (`detect_congestion()`) triggers on any of:
|
||||
|
||||
- Outgoing link MMP loss rate ≥ `node.ecn.loss_threshold` (default 5%)
|
||||
- Outgoing link MMP ETX ≥ `node.ecn.etx_threshold` (default 3.0)
|
||||
- Kernel receive buffer drops detected on any local transport (via
|
||||
`SO_RXQ_OVFL` on UDP)
|
||||
|
||||
**CE relay**: The forwarding path computes `outgoing_ce = incoming_ce ||
|
||||
local_congestion`. The `send_encrypted_link_message_with_ce()` method ORs
|
||||
`FLAG_CE` into the FMP header flags when ce is true. The original
|
||||
`send_encrypted_link_message()` delegates with `ce_flag=false`, leaving the
|
||||
20+ existing call sites unchanged.
|
||||
|
||||
**IPv6 ECN-CE marking**: When a CE-flagged DataPacket arrives at its final
|
||||
destination, the IPv6 Traffic Class ECN bits are marked CE (0b11) before
|
||||
TUN delivery — but only for ECN-capable packets (ECT(0) or ECT(1)). Not-ECT
|
||||
packets are never marked per RFC 3168. The host TCP stack then echoes ECE in
|
||||
ACKs, triggering sender cwnd reduction through standard congestion control.
|
||||
|
||||
**Session-layer tracking**: The `ecn_ce_count` field in MMP ReceiverReports
|
||||
tracks CE-flagged packets received per link, providing end-to-end visibility
|
||||
into congestion propagation.
|
||||
|
||||
**Monitoring**: `CongestionStats` tracks four counters — `ce_forwarded`,
|
||||
`ce_received`, `congestion_detected`, and `kernel_drop_events` — exposed via
|
||||
`fipsctl show routing` (congestion block) and `fipstop` (routing tab).
|
||||
Rate-limited warn logging (5s interval) alerts on congestion detection events.
|
||||
|
||||
See `node.ecn.*` in
|
||||
[fips-configuration.md](fips-configuration.md#ecn-signaling-nodeecn) for
|
||||
tuning parameters.
|
||||
|
||||
### Operator Logging
|
||||
|
||||
MMP emits periodic link metrics at info level (configurable via
|
||||
`node.mmp.log_interval_secs`, default 30s):
|
||||
|
||||
```text
|
||||
MMP link metrics peer=node-b rtt=2.3ms loss=0.2% jitter=0.1ms goodput=76.0MB/s tx_pkts=1234 rx_pkts=5678
|
||||
```
|
||||
|
||||
Teardown logs include final SRTT, loss rate, jitter, ETX, goodput, and
|
||||
cumulative tx/rx packet and byte counts.
|
||||
For the full MMP design — operating modes, report scheduling, spin
|
||||
bit interaction, ECN, and the algorithmic details shared with
|
||||
session-layer MMP — see [fips-mmp.md](fips-mmp.md). For the
|
||||
SenderReport and ReceiverReport byte layouts, see
|
||||
[../reference/wire-formats.md](../reference/wire-formats.md).
|
||||
Configuration knobs live under `node.mmp.*` and `node.ecn.*` in
|
||||
[../reference/configuration.md](../reference/configuration.md).
|
||||
|
||||
## Security Properties
|
||||
|
||||
### Threat Resistance
|
||||
|
||||
| Threat | Mitigation |
|
||||
| ------ | ---------- |
|
||||
| Connection exhaustion | Token bucket rate limit + connection count limit |
|
||||
| CPU exhaustion (msg1 flood) | Rate limit before crypto operations |
|
||||
| Replay attacks | Counter-based nonces with sliding window |
|
||||
| State confusion | Strict handshake state machine validation |
|
||||
| Spoofed encrypted packets | Index lookup + AEAD verification |
|
||||
| Spoofed msg2 | Index lookup + Noise ephemeral key binding |
|
||||
| Address spoofing | Cryptographic authority, not address-based |
|
||||
| Session correlation | Index rotation on rekey |
|
||||
The link-layer threat-resistance matrix (connection exhaustion, CPU
|
||||
exhaustion, replay, state confusion, spoofing variants, address
|
||||
spoofing, session correlation) is consolidated in
|
||||
[../reference/security.md](../reference/security.md) along with the
|
||||
session-layer matrix and operator-facing controls.
|
||||
|
||||
### Unauthenticated Attack Surface
|
||||
|
||||
@@ -574,13 +566,17 @@ an attacker sends invalid packets to elicit responses.
|
||||
| Metrics Measurement Protocol (MMP) | **Implemented** |
|
||||
| ECN congestion signaling (CE relay, IPv6 marking) | **Implemented** |
|
||||
| Rekey with index rotation | **Implemented** |
|
||||
| Allowlist/blocklist | Planned |
|
||||
| Peer ACL (allowlist / denylist) | **Implemented** |
|
||||
|
||||
## References
|
||||
|
||||
- [fips-intro.md](fips-intro.md) — Protocol overview and architecture
|
||||
- [fips-concepts.md](fips-concepts.md) — Protocol overview
|
||||
- [fips-architecture.md](fips-architecture.md) — Layer architecture and
|
||||
identity model
|
||||
- [fips-transport-layer.md](fips-transport-layer.md) — Transport layer (below FMP)
|
||||
- [fips-session-layer.md](fips-session-layer.md) — FSP (above FMP)
|
||||
- [fips-mmp.md](fips-mmp.md) — Metrics Measurement Protocol (link + session)
|
||||
- [fips-mesh-operation.md](fips-mesh-operation.md) — How FMP's routing and
|
||||
self-organization work in practice
|
||||
- [fips-wire-formats.md](fips-wire-formats.md) — Byte-level wire format reference
|
||||
- [../reference/wire-formats.md](../reference/wire-formats.md) — Byte-level
|
||||
wire format reference
|
||||
|
||||
@@ -31,183 +31,51 @@ and self-healing.
|
||||
|
||||
## Spanning Tree Formation and Maintenance
|
||||
|
||||
### What the Spanning Tree Provides
|
||||
For routing purposes, the spanning tree provides each node with a
|
||||
coordinate (its ancestry path from itself to the root) plus a way to
|
||||
compute distance between any two nodes (hops to their lowest common
|
||||
ancestor). The strictly-decreasing distance invariant gives greedy
|
||||
forwarding its loop-freedom.
|
||||
|
||||
The spanning tree gives each node a **coordinate**: its ancestry path from
|
||||
itself to the root, expressed as a sequence of node_addrs. These coordinates
|
||||
enable:
|
||||
The tree forms through distributed parent selection — root is the
|
||||
smallest node_addr (no election), and each node picks the peer with
|
||||
the lowest `effective_depth = depth + link_cost`. Cost-aware parent
|
||||
selection lets the tree trade hop count for link quality once MMP has
|
||||
accumulated SRTT and ETX metrics. Hysteresis (20% improvement
|
||||
required to switch) and hold-down (suppress non-mandatory
|
||||
re-evaluation after a switch) keep the tree stable under metric
|
||||
noise. Partitions self-resolve — each segment converges to its own
|
||||
root and reconverges to the smallest reachable root when segments
|
||||
rejoin.
|
||||
|
||||
- **Distance calculation**: The tree distance between two nodes is the number
|
||||
of hops from each to their lowest common ancestor (LCA). This provides a
|
||||
routing metric without any node knowing the full topology.
|
||||
- **Greedy routing**: At each hop, forward to the peer that minimizes tree
|
||||
distance to the destination. The strictly-decreasing distance invariant
|
||||
guarantees loop-free forwarding.
|
||||
Liveness is detected via FMP heartbeats; dead-peer removal triggers
|
||||
tree reconvergence and bloom filter recomputation for the affected
|
||||
subtree. The heartbeat and dead-timeout mechanism lives at the link
|
||||
layer; see [fips-mesh-layer.md](fips-mesh-layer.md#liveness-detection).
|
||||
|
||||
### How the Tree Forms
|
||||
|
||||
Nodes self-organize into a spanning tree through distributed parent selection:
|
||||
|
||||
1. **Root discovery**: The node with the smallest node_addr becomes the root.
|
||||
No election protocol — this is a consequence of each node independently
|
||||
preferring lower-addressed roots.
|
||||
2. **Parent selection**: Each node selects a single parent from among its
|
||||
direct peers based on which offers the lowest effective depth (tree depth
|
||||
weighted by local link cost).
|
||||
3. **Coordinate computation**: Once a node has a parent, its coordinate is
|
||||
computed from its ancestry path.
|
||||
|
||||
### How the Tree Maintains Itself
|
||||
|
||||
Nodes exchange **TreeAnnounce** messages with their direct peers (not
|
||||
forwarded — peer-to-peer only). Each TreeAnnounce carries the sender's
|
||||
current ancestry chain and a sequence number.
|
||||
|
||||
Changes cascade through the tree:
|
||||
|
||||
- A node that changes its parent recomputes its coordinates and announces to
|
||||
all peers
|
||||
- Each receiving peer evaluates whether the change affects its own parent
|
||||
selection
|
||||
- Only nodes that actually change their coordinates (root or depth changed)
|
||||
propagate further
|
||||
|
||||
TreeAnnounce propagation is rate-limited at 500ms minimum interval per peer.
|
||||
A tree of depth D reconverges in roughly D×0.5s to D×1.0s.
|
||||
|
||||
### How the Tree Adapts to Link Quality
|
||||
|
||||
The initial tree forms based on hop count alone — all links default to a
|
||||
cost of 1.0 before measurements are available. As the Metrics Measurement
|
||||
Protocol (MMP) accumulates bidirectional delivery ratios and round-trip
|
||||
time estimates, each node computes a per-link cost:
|
||||
|
||||
```text
|
||||
link_cost = ETX × (1.0 + SRTT_ms / 100.0)
|
||||
```
|
||||
|
||||
ETX (Expected Transmission Count) captures loss — a perfect link has
|
||||
ETX = 1.0, while 10% loss in each direction yields ETX ≈ 1.23. The SRTT
|
||||
term weights latency so that a low-loss but high-latency link (e.g., a
|
||||
satellite hop) costs more than a low-loss, low-latency link.
|
||||
|
||||
Parent selection uses **effective depth** rather than raw hop count:
|
||||
|
||||
```text
|
||||
effective_depth = peer.depth + link_cost_to_peer
|
||||
```
|
||||
|
||||
This allows a node to trade a shorter but lossy path for a longer but
|
||||
higher-quality one. A node two hops from the root over clean links
|
||||
(effective depth ≈ 3.0) is preferred over a node one hop away over a
|
||||
degraded link (effective depth ≈ 4.5).
|
||||
|
||||
Parent reselection is triggered by three paths:
|
||||
|
||||
1. **TreeAnnounce**: When a peer announces a new tree position, the node
|
||||
re-evaluates using current link costs
|
||||
2. **Periodic re-evaluation**: Every 60s (configurable), the node
|
||||
re-evaluates its parent choice using the latest MMP metrics, catching
|
||||
gradual link degradation that doesn't trigger TreeAnnounce
|
||||
3. **Parent loss**: When the current parent is removed, the node
|
||||
immediately selects the best alternative
|
||||
|
||||
To prevent oscillation from metric noise, parent switches are subject to
|
||||
**hysteresis**: a candidate must offer an effective depth at least 20%
|
||||
better than the current parent to trigger a switch. A **hold-down period**
|
||||
(default 30s) suppresses non-mandatory re-evaluation after a switch,
|
||||
allowing MMP metrics to stabilize on the new link before reconsidering.
|
||||
|
||||
### Flap Dampening
|
||||
|
||||
Unstable links that repeatedly connect and disconnect can cause cascading
|
||||
tree reconvergence. The spanning tree uses flap dampening with hysteresis
|
||||
and hold-down periods to suppress rapid parent oscillation. Links that flap
|
||||
above a configurable threshold are temporarily penalized, preventing them
|
||||
from being selected as parent until the link stabilizes.
|
||||
|
||||
### Link Liveness
|
||||
|
||||
Each node sends a dedicated **Heartbeat** message (0x51, 1 byte, no
|
||||
payload) to every peer at a fixed interval (default 10s). Any
|
||||
authenticated encrypted frame — heartbeat, MMP report, TreeAnnounce,
|
||||
data packet — resets the peer's liveness timer. On an idle link with no
|
||||
application data or topology changes, the heartbeat is the only traffic
|
||||
that keeps the link alive.
|
||||
|
||||
Peers that are silent for a configurable dead timeout (default 30s) are
|
||||
considered dead and removed from the peer table. With the default 10s
|
||||
heartbeat interval, a peer must miss three consecutive heartbeats before
|
||||
removal. This triggers tree reconvergence and bloom filter recomputation
|
||||
for the affected subtree.
|
||||
|
||||
### Partition Handling
|
||||
|
||||
If the network partitions, each segment independently rediscovers its own
|
||||
root (the smallest node_addr in the segment) and reconverges. When segments
|
||||
rejoin, nodes discover the globally-smallest root through TreeAnnounce
|
||||
exchange and reconverge to a single tree.
|
||||
|
||||
See [fips-spanning-tree.md](fips-spanning-tree.md) for algorithm details
|
||||
and [spanning-tree-dynamics.md](spanning-tree-dynamics.md) for convergence
|
||||
walkthroughs.
|
||||
For the parent-selection algorithm, hold-down/hysteresis details, and
|
||||
the convergence walkthroughs, see
|
||||
[fips-spanning-tree.md](fips-spanning-tree.md) and
|
||||
[spanning-tree-dynamics.md](spanning-tree-dynamics.md).
|
||||
|
||||
## Bloom Filter Gossip and Propagation
|
||||
|
||||
### What Bloom Filters Provide
|
||||
For routing purposes, each node maintains a bloom filter per peer
|
||||
that answers "can peer P possibly reach destination D?" — either "no"
|
||||
(definitive) or "maybe" (probabilistic). Because filters propagate
|
||||
along tree edges with split-horizon exclusion, a bloom hit on a tree
|
||||
peer reliably indicates which subtree contains the destination, and
|
||||
tree-coordinate distance ranks competing matches.
|
||||
|
||||
Each node maintains a bloom filter per peer, answering: "can peer P possibly
|
||||
reach destination D?" The answer is either "no" (definitive) or "maybe"
|
||||
(probabilistic — false positives are possible).
|
||||
FilterAnnounce updates are event-driven (peer changes, tree
|
||||
restructuring, local identity changes) and rate-limited to prevent
|
||||
storms. False positives at large scale never cause loops — the
|
||||
self-distance check at each hop guarantees forward progress, and
|
||||
mismatched bloom matches fall through to greedy tree routing.
|
||||
|
||||
Because filters propagate along tree edges with split-horizon exclusion,
|
||||
they encode directional reachability: a bloom hit on a tree peer reliably
|
||||
indicates which subtree contains the destination. When multiple peers match,
|
||||
tree coordinate distance ranks them.
|
||||
|
||||
### How Filters Propagate
|
||||
|
||||
Nodes exchange **FilterAnnounce** messages with all direct peers. Each
|
||||
FilterAnnounce replaces the previous filter for that peer — there is no
|
||||
incremental update.
|
||||
|
||||
Filter computation uses **tree-only merge with split-horizon exclusion**:
|
||||
the outbound filter for peer Q is computed by merging the local node's own
|
||||
identity, its leaf-only dependents (if any), and the inbound filters from
|
||||
tree peers (parent and children) *except* Q. Filters from non-tree mesh
|
||||
peers are stored locally for routing queries but are not merged into
|
||||
outgoing filters. This prevents saturation where mesh shortcuts cause
|
||||
filters to converge toward the full network.
|
||||
|
||||
The restriction creates **directional asymmetry**: upward filters
|
||||
(child → parent) contain the child's subtree, while downward filters
|
||||
(parent → child) contain the complement. Together they cover the entire
|
||||
network.
|
||||
|
||||
Filters propagate transitively through tree edges. At steady state, every
|
||||
reachable destination appears in at least one tree peer's filter.
|
||||
|
||||
### Update Triggers
|
||||
|
||||
Filter updates are event-driven, not periodic:
|
||||
|
||||
- Peer connects or disconnects
|
||||
- A peer's incoming filter changes (triggers recomputation for other peers)
|
||||
- Tree relationship changes (new parent, new child, parent switch)
|
||||
- Local state changes (new identity, leaf-only dependent changes)
|
||||
|
||||
Updates are rate-limited at 500ms to prevent storms during topology changes.
|
||||
|
||||
### Scale Properties
|
||||
|
||||
At moderate network sizes, bloom filters are highly accurate. At larger
|
||||
scales (~1M nodes), hub nodes with many peers may see elevated false positive
|
||||
rates (7–15% for nodes with 20+ peers). False positives may cause a packet
|
||||
to be forwarded toward the wrong subtree, but the self-distance check at
|
||||
each hop prevents loops and the packet falls through to greedy tree routing.
|
||||
|
||||
See [fips-bloom-filters.md](fips-bloom-filters.md) for filter parameters,
|
||||
FPR calculations, and size class folding.
|
||||
For the filter computation, split-horizon merge rules, FPR analysis,
|
||||
size classes, and folding, see
|
||||
[fips-bloom-filters.md](fips-bloom-filters.md).
|
||||
|
||||
## Routing Decision Process
|
||||
|
||||
@@ -222,34 +90,33 @@ priority chain. This is the core routing algorithm.
|
||||
2. **Direct peer** — The destination is an authenticated neighbor. Forward
|
||||
directly. No coordinates or bloom filters needed.
|
||||
|
||||
3. **Bloom-guided routing** — One or more peers' bloom filters contain the
|
||||
3. **Coordinate cache check** — Multi-hop forwarding requires the
|
||||
destination's tree coordinates to be in the local cache. On miss,
|
||||
`find_next_hop()` returns None immediately — bloom filters are never
|
||||
consulted — and the source receives a CoordsRequired error signal.
|
||||
|
||||
4. **Bloom-guided routing** — One or more peers' bloom filters contain the
|
||||
destination. Select the best peer by composite key:
|
||||
`(link_cost, tree_distance, node_addr)`. This requires the destination's
|
||||
tree coordinates to be in the local coordinate cache.
|
||||
`(link_cost, tree_distance, node_addr)`.
|
||||
|
||||
4. **Greedy tree routing** — Fallback when bloom filters haven't converged
|
||||
for this destination. Forward to the peer that minimizes tree distance.
|
||||
Also requires destination coordinates.
|
||||
5. **Greedy tree routing** — Fall-through when bloom yields no candidate.
|
||||
Forward to the peer that minimizes tree distance. If the tree has no
|
||||
next hop closer to the destination, the source receives a PathBroken
|
||||
error signal.
|
||||
|
||||
5. **No route** — Destination unreachable. Generate an error signal
|
||||
(CoordsRequired or PathBroken) back to the source.
|
||||
### Convergence Requirements
|
||||
|
||||
### The Coordinate Requirement
|
||||
|
||||
All multi-hop routing (steps 3–4) requires the destination's tree coordinates
|
||||
to be in the local coordinate cache. Without coordinates, `find_next_hop()`
|
||||
returns None immediately — bloom filters are never even consulted.
|
||||
|
||||
This creates two simultaneous convergence requirements for multi-hop routing:
|
||||
Multi-hop routing depends on two propagation processes that must run
|
||||
to convergence simultaneously:
|
||||
|
||||
1. **Bloom convergence**: Filters must propagate so peers advertise
|
||||
reachability
|
||||
2. **Coordinate availability**: Destination coordinates must be cached at
|
||||
every transit node on the path
|
||||
|
||||
Both must be satisfied simultaneously. Bloom convergence without coordinates
|
||||
causes a coordinate cache miss. Coordinates without bloom convergence falls
|
||||
through to greedy tree routing (functional but suboptimal).
|
||||
Bloom convergence without coordinates trips step 3 (coord-cache miss →
|
||||
CoordsRequired). Coordinates without bloom convergence falls through to
|
||||
greedy tree routing — functional but suboptimal.
|
||||
|
||||
### Candidate Ranking
|
||||
|
||||
@@ -270,6 +137,10 @@ A peer with a bloom filter hit but no entry in the peer ancestry table
|
||||
(missing TreeAnnounce) defaults to maximum distance and is effectively
|
||||
invisible to routing.
|
||||
|
||||
### Routing Decision Flowchart
|
||||
|
||||

|
||||
|
||||
### Loop Prevention
|
||||
|
||||
The routing decision enforces strict progress: a packet is only forwarded
|
||||
@@ -284,48 +155,12 @@ PathBroken error.
|
||||
|
||||
## Coordinate Caching
|
||||
|
||||
The coordinate cache maps `NodeAddr → TreeCoordinate` and is the critical
|
||||
data structure for multi-hop routing. Without it, forwarding decisions cannot
|
||||
be made.
|
||||
|
||||
### Unified Cache
|
||||
|
||||
The coordinate cache is a single unified cache. All sources — SessionSetup
|
||||
transit, CP-flagged data packets, LookupResponse — write to the same cache.
|
||||
|
||||
### Population Sources
|
||||
|
||||
| Source | When | What |
|
||||
| ------ | ---- | ---- |
|
||||
| SessionSetup transit | Session establishment | Both src and dest coordinates |
|
||||
| SessionAck transit | Session establishment | Both src and dest coordinates |
|
||||
| CP-flagged data packet | Warmup or recovery | Both src and dest coordinates (cleartext) |
|
||||
| LookupResponse | Discovery | Target's coordinates |
|
||||
|
||||
### Eviction
|
||||
|
||||
- **TTL-based**: Entries expire after 300s (configurable)
|
||||
- **Refresh on use**: Active routing refreshes the TTL, keeping hot entries
|
||||
alive
|
||||
- **LRU**: When full, least recently used entries are evicted first
|
||||
- **Flush on parent change**: When the local node's tree parent changes, the
|
||||
entire cache is flushed. Parent changes mean the node's own coordinates
|
||||
have changed, making relative distance calculations with cached coordinates
|
||||
potentially invalid. Flushing is preferred over stale routing: the cost of
|
||||
re-discovery is lower than routing packets to dead ends.
|
||||
|
||||
### Cache and Session Timer Ordering
|
||||
|
||||
Timer values are ordered so that idle sessions tear down before transit
|
||||
caches expire:
|
||||
|
||||
| Timer | Default | Purpose |
|
||||
| ----- | ------- | ------- |
|
||||
| Session idle | 90s | Session teardown |
|
||||
| Coordinate cache TTL | 300s | Coordinate expiration |
|
||||
|
||||
When traffic stops, the session tears down at 90s. When traffic resumes, a
|
||||
fresh SessionSetup re-warms transit caches (still within their 300s TTL).
|
||||
The coordinate cache maps `NodeAddr → TreeCoordinate` and is the
|
||||
critical data structure for multi-hop routing. The session layer owns
|
||||
this cache (its eviction policy, TTL/refresh semantics, parent-change
|
||||
flush, and timer ordering with session idle timeout); see
|
||||
[fips-session-layer.md](fips-session-layer.md#coordinate-cache) for
|
||||
the canonical treatment.
|
||||
|
||||
## Discovery Protocol
|
||||
|
||||
@@ -375,29 +210,30 @@ where a request might arrive via both tree and fallback paths.
|
||||
|
||||
Single-path forwarding is more fragile than flooding — if any transit node
|
||||
on the path has a stale bloom filter or loses a link, the request fails.
|
||||
To compensate, the originator retries:
|
||||
To compensate, each discovery is a sequence of attempts with growing
|
||||
per-attempt timeouts. The default sequence is `[1s, 2s, 4s, 8s]`
|
||||
(configurable via `node.discovery.attempt_timeouts_secs`); the destination
|
||||
is declared unreachable only after the full sequence is exhausted (15s
|
||||
total at default).
|
||||
|
||||
- **T=0**: Initial lookup sent
|
||||
- **T=5s**: Retry if no response (configurable via `retry_interval_secs`)
|
||||
- **T=10s**: Timeout, fail (configurable via `timeout_secs`)
|
||||
When the current attempt's deadline elapses without a `LookupResponse`,
|
||||
the originator sends another `LookupRequest` with a **fresh `request_id`**
|
||||
and the next entry in the sequence as its deadline. Fresh `request_id`s
|
||||
let each attempt take a different forwarding path as the bloom and tree
|
||||
state evolve, which is particularly useful during cold-start convergence.
|
||||
|
||||
The default `max_attempts` is 2 (initial + one retry). Each retry generates
|
||||
a fresh `request_id` and re-evaluates bloom filter matches, so it can take
|
||||
a different path if the tree has restructured.
|
||||
### Originator Backoff (optional, off by default)
|
||||
|
||||
### Originator Backoff
|
||||
|
||||
After a lookup times out or no peer's bloom filter contains the target, the
|
||||
originator enters **exponential backoff** before re-attempting discovery for
|
||||
the same target:
|
||||
|
||||
- **Base delay**: 30s (configurable via `backoff_base_secs`)
|
||||
- **Multiplier**: 2x per consecutive failure
|
||||
- **Cap**: 300s (configurable via `backoff_max_secs`)
|
||||
|
||||
Backoff is **reset on topology changes** that might make previously
|
||||
unreachable targets reachable: parent switch, new peer connection, first
|
||||
RTT measurement from MMP, or peer reconnection.
|
||||
After the per-attempt sequence is exhausted, the originator can additionally
|
||||
suppress further fresh lookups for the same target with exponential
|
||||
post-failure backoff. This is **disabled by default** (`backoff_base_secs:
|
||||
0`); the per-attempt sequence is the only retry pacing in the standard
|
||||
configuration. Operators may opt in via `node.discovery.backoff_base_secs`
|
||||
and `node.discovery.backoff_max_secs` if their deployment has chatty apps
|
||||
generating repeated lookups for genuinely unreachable destinations. When
|
||||
enabled, backoff is **reset on topology changes** that might make
|
||||
previously unreachable targets reachable: parent switch, new peer
|
||||
connection, first RTT measurement from MMP, or peer reconnection.
|
||||
|
||||
### Bloom Filter Pre-Check
|
||||
|
||||
@@ -450,6 +286,10 @@ verification at the source confirms the target holds the claimed position.
|
||||
The `path_mtu` field is excluded from the proof because it is a transit
|
||||
annotation modified at each hop.
|
||||
|
||||
### Coordinate Discovery Sequence
|
||||
|
||||

|
||||
|
||||
### Discovery Outcome
|
||||
|
||||
On receiving a verified LookupResponse, the source caches the target's
|
||||
@@ -461,50 +301,19 @@ If discovery times out (no response after all retry attempts), queued
|
||||
packets receive ICMPv6 Destination Unreachable and the target enters
|
||||
backoff.
|
||||
|
||||
## SessionSetup Self-Bootstrapping
|
||||
## Coordinate Cache Warming
|
||||
|
||||
SessionSetup is the mechanism that warms transit node coordinate caches
|
||||
along a path, enabling subsequent data packets to route efficiently.
|
||||
|
||||
### How It Works
|
||||
|
||||
SessionSetup carries plaintext coordinates (outside the Noise handshake
|
||||
payload, visible to transit nodes):
|
||||
|
||||
- **src_coords**: Source's current tree coordinates
|
||||
- **dest_coords**: Destination's tree coordinates (learned from discovery)
|
||||
|
||||
As the SessionSetup transits each intermediate node:
|
||||
|
||||
1. The transit node extracts both coordinate sets
|
||||
2. Caches `src_addr → src_coords` and `dest_addr → dest_coords` in its
|
||||
coordinate cache
|
||||
3. Forwards the message using the cached destination coordinates
|
||||
|
||||
SessionAck returns along the reverse path, carrying both the responder's
|
||||
and initiator's coordinates and warming caches in the other direction. This
|
||||
ensures return-path transit nodes can route even when the reverse path
|
||||
diverges from the forward path (e.g., after tree reconvergence).
|
||||
|
||||
### Result
|
||||
|
||||
After the handshake completes, the entire forward and reverse paths have
|
||||
cached coordinates for both endpoints. Subsequent data packets use minimal
|
||||
headers (no coordinates) and route efficiently through the warmed caches.
|
||||
|
||||
## Hybrid Coordinate Warmup (CP + CoordsWarmup)
|
||||
|
||||
The CP flag in the FSP common prefix and the standalone CoordsWarmup message
|
||||
(0x14) together provide a hybrid cache-warming mechanism that complements
|
||||
SessionSetup. See [fips-session-layer.md](fips-session-layer.md) for the
|
||||
full warmup strategy.
|
||||
|
||||
Transit nodes parse the CP flag from the FSP header and extract source and
|
||||
destination coordinates from the cleartext section between the header and
|
||||
ciphertext — no decryption needed. This is the same caching operation
|
||||
performed for SessionSetup coordinates. CoordsWarmup messages use the same
|
||||
CP-flag format and are handled identically by transit nodes via the existing
|
||||
`try_warm_coord_cache()` path.
|
||||
SessionSetup carries plaintext source and destination coordinates,
|
||||
which transit nodes cache as the message travels — warming the
|
||||
forward path. SessionAck carries them back along the reverse path,
|
||||
warming return-path caches. Steady-state data packets piggyback
|
||||
coordinates via the FSP CP flag during the warmup window, falling
|
||||
back to standalone CoordsWarmup messages when piggybacking would
|
||||
exceed the transport MTU. See
|
||||
[fips-session-layer.md](fips-session-layer.md#hybrid-coordinate-warmup-strategy)
|
||||
for the canonical hybrid-warmup design (SessionSetup
|
||||
self-bootstrapping plus CP-flag piggyback plus standalone
|
||||
CoordsWarmup).
|
||||
|
||||
## Error Recovery
|
||||
|
||||
@@ -704,21 +513,10 @@ routing decisions but retains its own end-to-end encryption and identity.
|
||||
|
||||
## Packet Type Summary
|
||||
|
||||
| Message | Typical Size | When | Forwarded? |
|
||||
| ------- | ------------ | ---- | ---------- |
|
||||
| TreeAnnounce | Variable (depth-dependent) | Topology changes | No (peer-to-peer) |
|
||||
| FilterAnnounce | ~1 KB | Topology changes | No (peer-to-peer) |
|
||||
| LookupRequest | ~300 bytes | First contact, recovery | Yes (bloom-guided tree) |
|
||||
| LookupResponse | ~400 bytes | Response to discovery | Yes (greedy routed) |
|
||||
| SessionDatagram + SessionSetup | ~232–402 bytes | Session establishment | Yes (routed) |
|
||||
| SessionDatagram + SessionAck | ~170 bytes | Session confirmation | Yes (routed) |
|
||||
| SessionDatagram + Data (minimal) | 77 bytes + IPv6 payload | Bulk IPv6 traffic (compressed) | Yes (routed) |
|
||||
| SessionDatagram + Data (with CP) | 77 + coords + IPv6 payload | Warmup/recovery (compressed) | Yes (routed) |
|
||||
| SessionDatagram + CoordsRequired | 70 bytes | Cache miss error | Yes (routed) |
|
||||
| SessionDatagram + PathBroken | 70+ bytes | Dead-end error | Yes (routed) |
|
||||
| Disconnect | 2 bytes | Link teardown | No (peer-to-peer) |
|
||||
|
||||
See [fips-wire-formats.md](fips-wire-formats.md) for byte-level layouts.
|
||||
For typical sizes, forwarding category, and the byte-level layouts
|
||||
of each FMP and FSP message type, see
|
||||
[../reference/wire-formats.md](../reference/wire-formats.md). The
|
||||
canonical Packet Type Summary table lives there.
|
||||
|
||||
## Privacy Considerations
|
||||
|
||||
@@ -772,11 +570,14 @@ recovery).
|
||||
|
||||
## References
|
||||
|
||||
- [fips-intro.md](fips-intro.md) — Protocol overview
|
||||
- [fips-concepts.md](fips-concepts.md) — Protocol overview
|
||||
- [fips-architecture.md](fips-architecture.md) — Layer architecture and
|
||||
identity model
|
||||
- [fips-mesh-layer.md](fips-mesh-layer.md) — FMP specification
|
||||
- [fips-spanning-tree.md](fips-spanning-tree.md) — Tree algorithms and data
|
||||
structures
|
||||
- [fips-bloom-filters.md](fips-bloom-filters.md) — Filter parameters and math
|
||||
- [fips-wire-formats.md](fips-wire-formats.md) — Wire format reference
|
||||
- [../reference/wire-formats.md](../reference/wire-formats.md) — Wire
|
||||
format reference
|
||||
- [spanning-tree-dynamics.md](spanning-tree-dynamics.md) — Convergence
|
||||
walkthroughs
|
||||
|
||||
223
docs/design/fips-mmp.md
Normal file
@@ -0,0 +1,223 @@
|
||||
# Metrics Measurement Protocol (MMP)
|
||||
|
||||
The Metrics Measurement Protocol provides per-link and per-session
|
||||
quality metrics — SRTT, loss, jitter, goodput, ETX, and one-way delay
|
||||
trend — using only counter and timestamp fields already present in
|
||||
the FMP and FSP wire formats. No additional probing traffic is
|
||||
required. The same algorithms and report message format are used at
|
||||
both layers; only the routing scope and configuration namespace
|
||||
differ.
|
||||
|
||||
This document is the canonical home for the MMP design. For the
|
||||
link-layer instance's role inside FMP, see
|
||||
[fips-mesh-layer.md](fips-mesh-layer.md). For the session-layer
|
||||
instance's role inside FSP, see
|
||||
[fips-session-layer.md](fips-session-layer.md). For the byte-level
|
||||
SenderReport and ReceiverReport layouts, see
|
||||
[../reference/wire-formats.md](../reference/wire-formats.md).
|
||||
|
||||
## Two Layers, One Protocol
|
||||
|
||||
MMP runs at two layers:
|
||||
|
||||
- **Link-layer MMP**: One instance per active FMP peer link. Reports
|
||||
are exchanged peer-to-peer between direct neighbors and measure the
|
||||
quality of that single hop.
|
||||
- **Session-layer MMP**: One instance per established FSP session.
|
||||
Reports are encrypted end-to-end and forwarded through every transit
|
||||
link, measuring end-to-end quality independent of hop count.
|
||||
|
||||
The algorithms (SRTT estimation, jitter computation, loss inference,
|
||||
ETX) are identical at both layers. The differences are configuration
|
||||
namespace, report intervals, and routing scope. See
|
||||
[Layer Differences](#layer-differences) below.
|
||||
|
||||
## Metrics Tracked
|
||||
|
||||
MMP computes the following metrics from the per-frame counter and
|
||||
timestamp fields:
|
||||
|
||||
- **SRTT** — Smoothed round-trip time (Jacobson/RFC 6298, α=1/8).
|
||||
Derived from timestamp-echo in ReceiverReports with dwell-time
|
||||
compensation.
|
||||
- **Loss rate** — Bidirectional loss inferred from counter gaps.
|
||||
Tracked as both instantaneous (per-interval) and long-term EWMA.
|
||||
- **Jitter** — Interarrival jitter (RFC 3550 algorithm) in
|
||||
microseconds.
|
||||
- **Goodput** — Bytes per second of payload data (excludes MMP
|
||||
reports).
|
||||
- **OWD trend** — One-way delay trend (µs/s, signed). Indicates
|
||||
congestion buildup before loss occurs.
|
||||
- **ETX** — Expected Transmission Count, computed from bidirectional
|
||||
delivery ratios. Used in cost-based parent selection via
|
||||
`link_cost = etx * (1.0 + srtt_ms / 100.0)`, and in bloom-filter
|
||||
candidate ranking inside `find_next_hop()` (the same `link_cost`
|
||||
is the primary key when choosing among bloom-filter peers, with
|
||||
tree distance as the tie-breaker).
|
||||
- **Dual EWMA trends** — Short-term (α=1/4) and long-term (α=1/32)
|
||||
trend indicators for both RTT and loss, enabling change detection.
|
||||
|
||||
Session-layer MMP additionally tracks the observed forward-path MTU;
|
||||
see [fips-mtu.md](fips-mtu.md) for the end-to-end path-MTU mechanism.
|
||||
|
||||
## Operating Modes
|
||||
|
||||
MMP supports three modes:
|
||||
|
||||
| Mode | Reports Exchanged | Metrics Available |
|
||||
| ---- | ----------------- | ----------------- |
|
||||
| **Full** (default) | SenderReport + ReceiverReport | All metrics including RTT, loss, jitter, goodput, OWD trend |
|
||||
| **Lightweight** | ReceiverReport only | Loss (from counter gaps), jitter, OWD trend. No RTT. |
|
||||
| **Minimal** | None | Spin bit and CE echo flags only. No computed metrics. |
|
||||
|
||||
The mode is configured per layer (`node.mmp.mode` and
|
||||
`node.session_mmp.mode`).
|
||||
|
||||
## Report Scheduling
|
||||
|
||||
Reports are sent at RTT-adaptive intervals computed as
|
||||
`clamp(2 × SRTT, low, high)`. A cold-start interval is used until SRTT
|
||||
has converged.
|
||||
|
||||
| Layer | Adaptive bounds | Cold-start |
|
||||
| ----- | --------------- | ---------- |
|
||||
| Link | `[1s, 5s]` | 200 ms (first 5 samples) |
|
||||
| Session | `[500ms, 10s]` | 1 s |
|
||||
|
||||
The session-layer bounds are higher because session reports are
|
||||
encrypted and forwarded through every transit link, so bandwidth cost
|
||||
is proportional to path length.
|
||||
|
||||
## Spin Bit and RTT
|
||||
|
||||
The SP (spin bit) flag in the FMP inner header follows the QUIC spin
|
||||
bit pattern: reflected on receive, toggled on send when the reflected
|
||||
value matches the last sent value. The spin bit state machine runs
|
||||
for TX reflection, but **RTT samples from the spin bit are
|
||||
discarded**. In a mesh protocol where frames are sent irregularly
|
||||
(tree announces, bloom filters, MMP reports on different timers),
|
||||
inter-frame processing delays inflate spin bit RTT measurements
|
||||
unpredictably. Timestamp-echo from ReceiverReports (with dwell-time
|
||||
compensation) is the sole SRTT source.
|
||||
|
||||
Duplicate or regressed ReceiverReports are ignored before any RTT, loss,
|
||||
goodput, or ETX update. If receiver-side dwell time exceeds the wire
|
||||
field, the report keeps its counters but sends a zero timestamp echo so
|
||||
the sender cannot form an invalid RTT sample.
|
||||
|
||||
The spin bit lives in the link-layer FMP inner header, so this
|
||||
mechanism applies to link-layer MMP only. Session-layer MMP carries
|
||||
its spin bit in the FSP encrypted inner header but uses it the same
|
||||
way: reflected for diagnostic visibility, not used for SRTT.
|
||||
|
||||
## ECN Congestion Signaling
|
||||
|
||||
The CE (Congestion Experienced) flag (bit 1 in the FMP flags byte)
|
||||
provides hop-by-hop congestion signaling through the mesh. Transit
|
||||
nodes detect congestion on outgoing links and set CE on forwarded
|
||||
packets; once set, the flag stays set for all subsequent hops to the
|
||||
destination.
|
||||
|
||||
**Congestion detection** triggers on any of:
|
||||
|
||||
- Outgoing link MMP loss rate ≥ `node.ecn.loss_threshold` (default 5%)
|
||||
- Outgoing link MMP ETX ≥ `node.ecn.etx_threshold` (default 3.0)
|
||||
- Kernel receive buffer drops detected on any local transport (via
|
||||
`SO_RXQ_OVFL` on UDP)
|
||||
|
||||
**CE relay**: The forwarding path computes
|
||||
`outgoing_ce = incoming_ce || local_congestion`. Once CE is set on a
|
||||
packet, it remains set for the rest of the forward path.
|
||||
|
||||
**IPv6 ECN-CE marking**: When a CE-flagged DataPacket arrives at its
|
||||
final destination, the IPv6 Traffic Class ECN bits are marked CE
|
||||
(0b11) before TUN delivery — but only for ECN-capable packets (ECT(0)
|
||||
or ECT(1)). Not-ECT packets are never marked per RFC 3168. The host
|
||||
TCP stack then echoes ECE in ACKs, triggering sender cwnd reduction
|
||||
through standard congestion control.
|
||||
|
||||
**Session-layer tracking**: The `ecn_ce_count` field in MMP
|
||||
ReceiverReports tracks CE-flagged packets received per link, providing
|
||||
end-to-end visibility into congestion propagation.
|
||||
|
||||
ECN signaling is a link-layer mechanism. Session-layer MMP only
|
||||
observes the CE counter as part of the report stream; CE marking is
|
||||
not generated end-to-end. Tuning parameters live under `node.ecn.*`
|
||||
in [../reference/configuration.md](../reference/configuration.md).
|
||||
|
||||
## Send Failure Backoff (Session Layer Only)
|
||||
|
||||
When a session MMP report cannot be delivered (destination unreachable,
|
||||
no route), the sender applies exponential backoff to the probe
|
||||
interval — a standard distributed-systems pattern for transient
|
||||
failure handling:
|
||||
|
||||
- Each consecutive failure doubles the interval: 2x, 4x, 8x, 16x, 32x
|
||||
- Backoff caps at 32x the base interval (5 consecutive failures)
|
||||
- A successful send resets to the normal SRTT-based interval
|
||||
- Debug logging is suppressed after 3 consecutive failures; a summary
|
||||
is logged when the destination becomes reachable again
|
||||
|
||||
This prevents wasted CPU and log noise when a session's remote
|
||||
endpoint has departed the network but the local session has not yet
|
||||
timed out. Link-layer MMP has no equivalent — link-layer reports are
|
||||
peer-to-peer over an authenticated link, so delivery failure is
|
||||
indistinguishable from link death and the link-liveness mechanism
|
||||
takes over.
|
||||
|
||||
## Layer Differences
|
||||
|
||||
| Aspect | Link layer | Session layer |
|
||||
| ------ | ---------- | ------------- |
|
||||
| Routing scope | Peer-to-peer (one hop) | End-to-end (forwarded through every hop) |
|
||||
| Configuration namespace | `node.mmp.*` | `node.session_mmp.*` |
|
||||
| Report bounds | `[1s, 5s]` | `[500ms, 10s]` |
|
||||
| Cold-start interval | 200 ms (first 5 samples) | 1 s |
|
||||
| Bandwidth cost | One link | Proportional to path length |
|
||||
| Send-failure backoff | Not applicable | Yes |
|
||||
| Path-MTU echo | Not applicable | PathMtuNotification (see [fips-mtu.md](fips-mtu.md)) |
|
||||
| Idle-timeout interaction | None | Reports do **not** reset session idle timer |
|
||||
|
||||
## Idle Timeout Interaction (Session Layer Only)
|
||||
|
||||
MMP reports (SenderReport, ReceiverReport) and PathMtuNotification do
|
||||
**not** reset the session idle timer. Only application data
|
||||
(DataPacket, type 0x10) resets `last_activity`. This ensures sessions
|
||||
with no application traffic tear down after
|
||||
`node.session.idle_timeout_secs` (default 90s), while MMP continues
|
||||
providing measurement data up to the teardown moment.
|
||||
|
||||
## Operator Logging
|
||||
|
||||
Both layers emit periodic metrics at info level. The interval is
|
||||
`node.mmp.log_interval_secs` for link-layer (default 30s) and
|
||||
`node.session_mmp.log_interval_secs` for session-layer (default 30s).
|
||||
|
||||
Link-layer:
|
||||
|
||||
```text
|
||||
MMP link metrics peer=node-b rtt=2.3ms loss=0.2% jitter=0.1ms goodput=76.0MB/s tx_pkts=1234 rx_pkts=5678
|
||||
```
|
||||
|
||||
Session-layer:
|
||||
|
||||
```text
|
||||
MMP session metrics session=npub1tdwa...84le rtt=4.3ms loss=0.6% jitter=0.2ms goodput=71.3MB/s mtu=1472 tx_pkts=1234 rx_pkts=5678
|
||||
```
|
||||
|
||||
Teardown logs include final SRTT, loss rate, jitter, ETX, goodput,
|
||||
and cumulative tx/rx packet and byte counts.
|
||||
|
||||
## See also
|
||||
|
||||
- [fips-mesh-layer.md](fips-mesh-layer.md) — link-layer MMP integration
|
||||
inside FMP
|
||||
- [fips-session-layer.md](fips-session-layer.md) — session-layer MMP
|
||||
integration inside FSP
|
||||
- [fips-mtu.md](fips-mtu.md) — PathMtuNotification, the session-only
|
||||
end-to-end path-MTU echo
|
||||
- [../reference/wire-formats.md](../reference/wire-formats.md) —
|
||||
SenderReport (0x01 / 0x11) and ReceiverReport (0x02 / 0x12) byte
|
||||
layouts
|
||||
- [../reference/configuration.md](../reference/configuration.md) —
|
||||
full `node.mmp.*`, `node.session_mmp.*`, and `node.ecn.*` knob tables
|
||||
316
docs/design/fips-mtu.md
Normal file
@@ -0,0 +1,316 @@
|
||||
# FIPS Path MTU and Encapsulation Overhead
|
||||
|
||||
MTU is a cross-cutting concern in FIPS. No single layer owns it: the
|
||||
transport reports per-link MTU, FMP propagates `path_mtu` along
|
||||
forward and reverse paths, FSP echoes the observed path MTU end-to-end
|
||||
back to the source, and the IPv6 adapter enforces the resulting
|
||||
effective MTU at the TUN interface. This document is the canonical
|
||||
home for the unified MTU model.
|
||||
|
||||
For operator-facing diagnostic recipes (interpreting `MtuExceeded`
|
||||
counters, tuning IPv6 application MSS, troubleshooting cold-flow
|
||||
oversize), see the relevant how-to under `docs/how-to/`.
|
||||
|
||||
## The MTU Problem in FIPS
|
||||
|
||||
A FIPS path can traverse heterogeneous link types — UDP/IP (1280
|
||||
default, IPv6 minimum), Ethernet (interface MTU − 3, typically 1497),
|
||||
BLE (negotiated ATT_MTU per link), Tor stream (1400 default), radio
|
||||
(51–222) — within a single end-to-end session.
|
||||
The minimum MTU along the path determines the largest datagram a
|
||||
session can deliver. Several properties make this harder than in
|
||||
classic IP networks:
|
||||
|
||||
- **No fragmentation.** FIPS does not fragment at transit nodes (see
|
||||
[No fragmentation policy](#no-fragmentation-policy)). A datagram
|
||||
that exceeds the next-hop link MTU is dropped, and the source is
|
||||
signaled.
|
||||
- **Forward/reverse path asymmetry.** After tree reconvergence the
|
||||
return path may diverge from the forward path, so the bottleneck
|
||||
on each direction can differ.
|
||||
- **First-flow race.** The very first SessionDatagram races
|
||||
destination discovery — the source has not yet learned the path MTU
|
||||
but must pick a payload size for the queued packet.
|
||||
- **Variable per-link MTU.** Some transports (BLE, TCP via
|
||||
`TCP_MAXSEG`) report different MTUs for different links rather than
|
||||
a single transport-wide value.
|
||||
|
||||
The unified MTU model below combines proactive and reactive
|
||||
mechanisms to converge on a working effective MTU within the first
|
||||
few packets of a session, then maintain it across topology changes.
|
||||
|
||||
## Encapsulation Overhead
|
||||
|
||||
The byte budget for a FIPS-encapsulated packet:
|
||||
|
||||
| Layer | Overhead | Purpose |
|
||||
| ----- | -------- | ------- |
|
||||
| Link encryption | 37 bytes | 16-byte outer header + 5-byte inner header (timestamp + msg_type) + 16-byte AEAD tag |
|
||||
| SessionDatagram body | 35 bytes | ttl + path_mtu + src_addr + dest_addr (msg_type counted in inner header) |
|
||||
| FSP header | 12 bytes | 4-byte prefix + 8-byte counter (used as AEAD AAD) |
|
||||
| FSP inner header | 6 bytes | 4-byte timestamp + 1-byte msg_type + 1-byte inner_flags (inside AEAD) |
|
||||
| Session AEAD tag | 16 bytes | ChaCha20-Poly1305 tag on session-encrypted payload |
|
||||
| **Protocol envelope** | **106 bytes** | `FIPS_OVERHEAD` constant — the base payload budget for any service |
|
||||
|
||||
`FIPS_OVERHEAD = 106` is the constant the rest of the system reasons
|
||||
about. Coordinate piggybacking via the CP flag adds variable extra
|
||||
overhead — `2 + entries × 16` bytes per coordinate, with both source
|
||||
and destination coordinates carried — and the send path skips the CP
|
||||
flag if adding coords would exceed the transport MTU.
|
||||
|
||||
Service-specific overheads layer on top of `FIPS_OVERHEAD`:
|
||||
|
||||
| Service | Overhead | Note |
|
||||
| ------- | -------- | ---- |
|
||||
| DataPacket port header | +4 bytes | Always present for port-multiplexed services |
|
||||
| IPv6 compression | −33 bytes | 40-byte IPv6 header → 7-byte format + residual |
|
||||
| **IPv6 effective overhead** | **77 bytes** | `FIPS_IPV6_OVERHEAD` constant |
|
||||
|
||||
See [fips-ipv6-adapter.md](fips-ipv6-adapter.md) for the IPv6
|
||||
compression scheme that lets the adapter reach `FIPS_IPV6_OVERHEAD`.
|
||||
|
||||
## Per-Link MTU Reporting
|
||||
|
||||
Each transport implements two MTU methods on its trait:
|
||||
|
||||
- `mtu() -> u16` — Transport-wide default MTU.
|
||||
- `link_mtu(addr: &TransportAddr) -> u16` — Per-link MTU for a
|
||||
specific remote address. The default implementation falls back to
|
||||
`mtu()`, so transports with uniform MTU (UDP, raw Ethernet) need
|
||||
not override it.
|
||||
|
||||
FMP uses `link_mtu()` when it needs to reason about a specific
|
||||
outbound link — typically for `path_mtu` annotation in
|
||||
SessionDatagram and LookupResponse. Per-transport defaults:
|
||||
|
||||
| Transport | Default MTU | Per-link MTU source |
|
||||
| --------- | ----------- | ------------------- |
|
||||
| UDP | 1280 (IPv6 minimum) | uniform (`mtu()` fallback) |
|
||||
| Ethernet | interface MTU − 3 (typically 1497) | uniform |
|
||||
| TCP | 1400 | derived from `TCP_MAXSEG` per connection |
|
||||
| Tor | 1400 | uniform |
|
||||
| BLE | 2048 default; negotiated ATT_MTU per link | per-link (overrides `mtu()`) |
|
||||
|
||||
For TCP, the per-connection `TCP_MAXSEG` query lets FMP discover the
|
||||
actual MSS the kernel negotiated for each connection, rather than
|
||||
assuming a single value across all TCP peers.
|
||||
|
||||
## Proactive PMTUD: SessionDatagram path_mtu
|
||||
|
||||
Every SessionDatagram and LookupResponse carries a 2-byte `path_mtu`
|
||||
field. The source initializes it to its outbound link MTU; each
|
||||
transit node applies `min(current, link_mtu(next_hop))` before
|
||||
forwarding. The destination receives the forward-path minimum.
|
||||
|
||||
For SessionDatagram, the receiver of the forward-path minimum is the
|
||||
session-layer destination, which then echoes the value back to the
|
||||
source via PathMtuNotification (see
|
||||
[End-to-end echo](#end-to-end-echo-pathmtunotification)).
|
||||
|
||||
For LookupResponse, the receiver is the original requester, and the
|
||||
annotation is reverse-path-only: the LookupResponse path is the
|
||||
return path of the lookup, so the annotated `path_mtu` reflects what
|
||||
the requester can use to reach the discovered destination over the
|
||||
discovered path.
|
||||
|
||||
Because the field is initialized by the source and mins as it travels,
|
||||
it converges to the bottleneck without any additional probing. The
|
||||
first SessionDatagram on a fresh session may carry an over-estimate
|
||||
(the source has not yet been told a smaller min), which is what makes
|
||||
the reactive MtuExceeded path necessary.
|
||||
|
||||
## Reactive PMTUD: MtuExceeded
|
||||
|
||||
When a transit node receives a SessionDatagram whose total wire size
|
||||
exceeds the next-hop `link_mtu`, it cannot forward without
|
||||
fragmentation. Instead:
|
||||
|
||||
1. The transit node generates a SessionDatagram addressed back to the
|
||||
source carrying an `MtuExceeded` payload (msg_type 0x22). The
|
||||
payload identifies the destination, the reporting router, and the
|
||||
bottleneck MTU.
|
||||
2. The error is routed via `find_next_hop(src_addr)`. If the source
|
||||
is also unreachable, the error is dropped silently (no cascading
|
||||
errors).
|
||||
3. The original oversized packet is dropped.
|
||||
|
||||
The source's FSP layer applies the reported bottleneck immediately —
|
||||
unlike the increase case (see hysteresis below), decrease is always
|
||||
take-the-lower-value because the original packet has already been
|
||||
dropped. The source can then reduce payload sizes on subsequent
|
||||
SessionDatagrams.
|
||||
|
||||
MtuExceeded is the reactive complement to the proactive `path_mtu`
|
||||
field. The proactive field tracks the minimum along the forward path
|
||||
under steady-state convergence; MtuExceeded handles the in-flight gap
|
||||
when an oversized packet hits a new bottleneck (forward path shifted,
|
||||
peer's outbound MTU dropped, BLE renegotiated) before the source has
|
||||
adapted.
|
||||
|
||||
Error generation is rate-limited at 100ms per destination at the
|
||||
transit node to prevent storms during topology changes.
|
||||
|
||||
## End-to-End Echo: PathMtuNotification
|
||||
|
||||
PathMtuNotification (msg_type 0x13, session-layer) provides
|
||||
end-to-end path MTU feedback, adapting RFC 1191 Path MTU Discovery
|
||||
for overlay networks — the transit-node `min()` propagation replaces
|
||||
ICMP Packet Too Big.
|
||||
|
||||
Mechanism:
|
||||
|
||||
1. The source sets `path_mtu` in each SessionDatagram envelope to its
|
||||
outbound link MTU.
|
||||
2. Each transit node applies `min(current, transport.link_mtu(addr))`
|
||||
before forwarding.
|
||||
3. The destination receives the forward-path minimum and sends a
|
||||
PathMtuNotification (2-byte body: `u16 LE path_mtu`) back to the
|
||||
source.
|
||||
4. The source applies the notification with hysteresis:
|
||||
- **Decrease**: immediate (take lower value).
|
||||
- **Increase**: requires 3 consecutive higher-value notifications
|
||||
spanning at least 2 × notification interval.
|
||||
5. Notifications are sent on first measurement, on any decrease, and
|
||||
periodically at `max(10s, 5 × SRTT)`.
|
||||
|
||||
The hysteresis on increase prevents oscillation when the path MTU
|
||||
fluctuates around a boundary; the immediate decrease prevents
|
||||
delivering oversized packets after a path has narrowed.
|
||||
|
||||
PathMtuNotification is wrapped in a session-layer encrypted message
|
||||
and travels back to the source via the session's normal forwarding
|
||||
path. It is part of the session-layer MMP report stream's traffic
|
||||
budget and (along with SenderReport and ReceiverReport) does not
|
||||
reset the session idle timer.
|
||||
|
||||
## Per-Destination MTU Storage
|
||||
|
||||
Two storage locations track per-destination MTU, serving different
|
||||
consumers:
|
||||
|
||||
- **Session-canonical** (`MmpSessionState.path_mtu`, type
|
||||
`PathMtuState`). Holds the running end-to-end path MTU for an
|
||||
established FSP session. Updated by both `PathMtuNotification`
|
||||
(proactive, end-to-end echo) and reactive `MtuExceeded` from
|
||||
transit routers. Read by the session layer when constructing
|
||||
outbound `SessionDatagram` envelopes.
|
||||
|
||||
- **TCP-clamp mirror** (`path_mtu_lookup`, a
|
||||
`HashMap<FipsAddress, u16>` on the Node). Read by the
|
||||
TUN-side TCP MSS clamp (`per_flow_max_mss` in
|
||||
`src/upper/tun.rs`) at first-SYN time so outbound TCP flows
|
||||
are clamped to the per-destination MTU rather than a generic
|
||||
ceiling. Written from four sites, all using tighter-only
|
||||
semantics — the clamp is never loosened:
|
||||
- Discovery's `LookupResponse` handler — reverse-path
|
||||
annotated value carried back by the discovery target.
|
||||
- `seed_path_mtu_for_link_peer` when a peer is promoted to
|
||||
an active link, seeding with the new link's `link_mtu`
|
||||
so traffic to that peer immediately uses the per-link
|
||||
value rather than a generic default.
|
||||
- The reactive `MtuExceeded` handler, mirroring the
|
||||
bottleneck reported by a transit router.
|
||||
- The proactive `PathMtuNotification` handler, mirroring
|
||||
the new effective end-to-end value so a fresh TCP flow
|
||||
benefits immediately from PMTU knowledge the session has
|
||||
already acquired.
|
||||
|
||||
All four writers apply the same tighter-only rule, so the mirror
|
||||
converges to the smallest MTU any signal has reported for that
|
||||
destination and a subsequent looser observation cannot widen it.
|
||||
|
||||
## TCP MSS Clamping
|
||||
|
||||
The IPv6 adapter intercepts TCP SYN and SYN-ACK packets at the TUN
|
||||
interface and clamps the Maximum Segment Size (MSS) option to:
|
||||
|
||||
```text
|
||||
clamped_mss = effective_ipv6_mtu - 40 (IPv6 header) - 20 (TCP header)
|
||||
```
|
||||
|
||||
Clamping is applied in two places:
|
||||
|
||||
- **TUN reader** (outbound): clamps MSS on outbound SYN packets
|
||||
- **TUN writer** (inbound): clamps MSS on inbound SYN-ACK packets
|
||||
|
||||
Together these ensure both directions of a TCP connection use
|
||||
appropriately-sized segments from the start, avoiding the initial
|
||||
oversized-packet loss that would occur if the adapter relied on ICMP
|
||||
Packet Too Big alone.
|
||||
|
||||
Clamping is **conditional**: when `per_flow_max_mss` already has an
|
||||
entry for the flow, that entry is used; otherwise the clamp falls
|
||||
back to a ceiling derived from the most pessimistic effective IPv6
|
||||
MTU the adapter knows about (1143 with the typical 1280 transport
|
||||
floor). The fallback handles cold-flow first-SYN traffic — the very
|
||||
first SYN of a flow may arrive before the MMP path-MTU echo and any
|
||||
per-flow lookup has been populated, so the conservative ceiling
|
||||
prevents the SYN-ACK chain from negotiating a too-large MSS that
|
||||
would later drop.
|
||||
|
||||
The adapter integrates with the MTU subsystem rather than owning it.
|
||||
The "why we clamp and what `max_mss` means" lives here in the MTU
|
||||
design; the "how the clamp is implemented at the TUN" lives in the
|
||||
[IPv6 adapter](fips-ipv6-adapter.md#tun-side-tcp-mss-clamping) doc.
|
||||
|
||||
## ICMP Packet Too Big
|
||||
|
||||
When an outbound packet at the TUN exceeds the effective IPv6 MTU,
|
||||
the adapter generates an ICMPv6 Packet Too Big message and delivers
|
||||
it back to the application via the TUN. This triggers the kernel's
|
||||
Path MTU Discovery mechanism for non-TCP traffic and for any TCP flow
|
||||
where MSS clamping was insufficient.
|
||||
|
||||
ICMPv6 Packet Too Big generation is rate-limited per source address
|
||||
(100ms interval) to prevent storms from applications sending many
|
||||
oversized packets. The ICMP response is delivered locally back
|
||||
through the TUN; no network traversal is needed, so delivery is
|
||||
reliable.
|
||||
|
||||
## No Fragmentation Policy
|
||||
|
||||
FIPS does not perform fragmentation at transit nodes:
|
||||
|
||||
- **Why no transit fragmentation.** Session-layer encryption is
|
||||
end-to-end — the AEAD tag authenticates the entire plaintext.
|
||||
Fragmenting an encrypted SessionDatagram would require either
|
||||
exposing plaintext structure to transit nodes (unacceptable) or
|
||||
reassembling before decryption (opens an attack surface — a transit
|
||||
node could replay or withhold fragments to influence reassembly).
|
||||
- **Why no source-side fragmentation.** The source doesn't need
|
||||
fragmentation because the proactive `path_mtu` field plus the
|
||||
reactive MtuExceeded signal converge on a working size within the
|
||||
first few packets. Applications that need oversized payloads run
|
||||
TCP over the IPv6 adapter, which has its own segmentation under
|
||||
MSS clamping.
|
||||
|
||||
Some transports may perform fragmentation and reassembly internally
|
||||
(e.g., BLE L2CAP) and can advertise a larger virtual MTU than the
|
||||
physical medium supports — this is transparent to FIPS.
|
||||
|
||||
## Operational Considerations
|
||||
|
||||
Diagnosing MTU-related symptoms (handshakes succeed but bulk
|
||||
transfers stall, ssh hangs after `Welcome` banner, sporadic
|
||||
`MtuExceeded` spikes during topology changes) requires inspecting
|
||||
per-link MTU, per-session MTU, and the per-destination
|
||||
`path_mtu_lookup` table. See
|
||||
[../how-to/diagnose-mtu-issues.md](../how-to/diagnose-mtu-issues.md)
|
||||
for the operator recipes. The relevant control-socket queries are
|
||||
`fipsctl show sessions` (per-session MTU), `fipsctl show transports`
|
||||
(per-link MTU), and `fipsctl show identity-cache` (with adapter MTU
|
||||
context).
|
||||
|
||||
## See also
|
||||
|
||||
- [fips-transport-layer.md](fips-transport-layer.md) — the `mtu()` /
|
||||
`link_mtu()` trait surface and per-transport defaults
|
||||
- [fips-mesh-layer.md](fips-mesh-layer.md) — SessionDatagram and the
|
||||
MtuExceeded error signal
|
||||
- [fips-session-layer.md](fips-session-layer.md) — session-layer
|
||||
PathMtuNotification echo, applied with hysteresis
|
||||
- [fips-ipv6-adapter.md](fips-ipv6-adapter.md) — TUN-side ICMPv6 PTB
|
||||
generation, MSS clamping integration, IPv6-specific overhead table
|
||||
- [../reference/wire-formats.md](../reference/wire-formats.md) —
|
||||
SessionDatagram, LookupResponse, MtuExceeded, PathMtuNotification
|
||||
byte layouts
|
||||
592
docs/design/fips-nostr-discovery.md
Normal file
@@ -0,0 +1,592 @@
|
||||
# FIPS Discovery: Nostr-Mediated and LAN/mDNS
|
||||
|
||||
FIPS nodes have two discovery mechanisms beyond the static `peers[]`
|
||||
list. The bulk of this document describes **Nostr-mediated discovery**,
|
||||
which works across the internet using public Nostr relays as a
|
||||
signaling channel and can punch through UDP NAT. A second, much
|
||||
simpler mechanism — **LAN/mDNS discovery** — finds peers on the same
|
||||
local link with no relay, STUN, or NAT traversal at all; it is
|
||||
described in its own section near the end. The two are independent: a
|
||||
node can enable either, both, or neither.
|
||||
|
||||
Nostr-mediated discovery lets FIPS nodes find each other, and if
|
||||
necessary, punch through UDP NAT, using public Nostr relays as the
|
||||
signaling channel. A node publishes its reachable transport endpoints to
|
||||
a small set of relays under its own Nostr identity (which is also its
|
||||
FIPS identity), and peers resolve those endpoints at dial time by npub.
|
||||
For peers behind UDP NAT, the same relay channel carries an encrypted
|
||||
offer/answer exchange, and STUN supplies the reflexive address used for
|
||||
a coordinated hole-punch.
|
||||
|
||||
Nostr discovery is unconditionally compiled into the `fips` binary on
|
||||
every supported platform and ships in every stock packaging artifact
|
||||
(`.deb`, AUR, systemd tarball, OpenWrt `.ipk`, macOS `.pkg`, Windows
|
||||
`.zip`). It is runtime-opt-in: the YAML configuration defaults to
|
||||
disabled (`node.discovery.nostr.enabled: false`), so the discovery
|
||||
runtime stays dormant — and opens no relay connections — until an
|
||||
operator flips the flag. Default relay and STUN-server lists ship in
|
||||
the config; both are optional overrides. When disabled, nodes behave
|
||||
exactly as before: only the static `peers[]` addresses are used.
|
||||
|
||||
## Role
|
||||
|
||||
The feature adds three capabilities on top of FIPS's static peer model:
|
||||
|
||||
- **Advertising.** A node publishes the transport endpoints it wants
|
||||
peers to use (direct UDP, direct TCP, a Tor onion, or the special
|
||||
`udp:nat` rendezvous token) as a signed Nostr event. The advert is
|
||||
anchored to the node's FIPS identity key — a peer that knows the npub
|
||||
knows the advert is authentic.
|
||||
- **Lookup.** When dialing a configured peer marked `via_nostr`, or any
|
||||
peer in `policy: open` mode, the node fetches that peer's advert from
|
||||
the configured relays and appends the advertised endpoints to its
|
||||
dial list. Static addresses are always tried first.
|
||||
- **UDP NAT hole-punch.** When both sides of a connection have UDP NAT
|
||||
endpoints, the advert carries enough information to run a STUN-based
|
||||
offer/answer exchange over encrypted ([NIP-59](https://github.com/nostr-protocol/nips/blob/master/59.md))
|
||||
Nostr events. Each side observes its reflexive address via STUN,
|
||||
exchanges candidate pairs through the relay, and both sides send UDP
|
||||
probes at a shared punch time. On the first successful probe, the
|
||||
punch socket is handed to FMP and becomes a normal UDP transport.
|
||||
|
||||
## When to use it
|
||||
|
||||
- **You run a public node** and want peers who know your npub to reach
|
||||
you without you distributing an address list out-of-band.
|
||||
- **You want to reach a peer behind UDP NAT** without deploying a relay
|
||||
or running Tor on both sides. The peer advertises `udp:nat` and you
|
||||
dial by npub.
|
||||
- **You want zero-touch peer discovery** within a known application
|
||||
namespace (`policy: open`), subject to an admission budget.
|
||||
- **You want to advertise a Tor onion** so peers don't need to know the
|
||||
`.onion` address out-of-band.
|
||||
|
||||
Skip the feature when every peer is already reachable through a stable
|
||||
static address (a LAN mesh, a pre-configured test bed, or a deployment
|
||||
where operators distribute `peers[]` blocks directly). The feature adds
|
||||
relay dependencies, STUN round-trips for NAT cases, and a small ambient
|
||||
background of relay traffic; none of that is useful when you already
|
||||
know where peers are.
|
||||
|
||||
## Scenarios and configuration
|
||||
|
||||
For end-to-end operator recipes — each of the five activation scenarios
|
||||
(advertise a directly-reachable UDP node, advertise a Tor onion node,
|
||||
look up a configured peer by npub without advertising, NAT hole-punch
|
||||
between two configured peers, and open discovery within an `app`
|
||||
namespace) — see
|
||||
[../how-to/enable-nostr-discovery.md](../how-to/enable-nostr-discovery.md).
|
||||
The full configuration knob tables, per-transport keys, and startup
|
||||
validation rules live in
|
||||
[../reference/configuration.md](../reference/configuration.md) under
|
||||
`node.discovery.nostr.*`. The Kind 37195 advert event format is in
|
||||
[../reference/nostr-events.md](../reference/nostr-events.md). The rest
|
||||
of this document covers the design of the discovery runtime itself.
|
||||
|
||||
## Under the covers
|
||||
|
||||
The rest of this document describes how the feature works inside the
|
||||
node. For the generic protocol shape (event tags, NIP usage, on-the-
|
||||
wire offer/answer schema, failure-suppression machinery), see
|
||||
[port-advertisement-and-nat-traversal.md](port-advertisement-and-nat-traversal.md).
|
||||
|
||||
### Overview
|
||||
|
||||
The discovery runtime is a background task group started during node
|
||||
initialization when `nostr.enabled` is true. It maintains a single
|
||||
`nostr-sdk` client connected to the union of `advert_relays` and
|
||||
`dm_relays`, and runs four loops: advert publication, advert
|
||||
subscription (for open discovery and cache warming), DM subscription
|
||||
(for incoming offers and answers), and a periodic advert-cache prune.
|
||||
Discovery has no CLI surface; all operations are driven by the
|
||||
configuration and by connection attempts made by the rest of the node.
|
||||
|
||||
```text
|
||||
+-----------------------+
|
||||
| Discovery runtime |
|
||||
+-----------------------+
|
||||
| | |
|
||||
advert publish | | DM sub (offers, answers)
|
||||
| |
|
||||
v v
|
||||
+-------------------------+
|
||||
| Nostr relay pool | (advert_relays ∪ dm_relays)
|
||||
+-------------------------+
|
||||
^ ^
|
||||
advert fetch/cache | | encrypted signaling
|
||||
| |
|
||||
+----------------+ | | +--------------------+
|
||||
| connect_peer |--+ +->| offer / answer |
|
||||
| (node side) | | handler |
|
||||
+----------------+ +--------------------+
|
||||
| |
|
||||
v v
|
||||
+---------+ +--------------+
|
||||
| STUN |<-- same socket --->| UDP punch |
|
||||
+---------+ +--------------+
|
||||
|
|
||||
v
|
||||
adopt_established_traversal()
|
||||
|
|
||||
v
|
||||
FMP IK handshake
|
||||
on adopted socket
|
||||
```
|
||||
|
||||
### Phase 1 — Advertisement
|
||||
|
||||
Adverts are published as Nostr kind `37195` parameterized replaceable
|
||||
events (FIPS-specific, in the application-defined replaceable range
|
||||
`30000–39999`; the digits visually spell `FIPS` — 7=F, 1=I, 9=P, 5=S).
|
||||
The `d` tag is hardcoded to the wire-format identifier
|
||||
`fips-overlay-v1` (or `fips-overlay-v1-next` on the `next` branch),
|
||||
so each node has a single, in-place-updatable advert under its
|
||||
identity. The configurable `app` value populates a separate
|
||||
`protocol` tag, which scopes adverts within a relay set without
|
||||
splitting them across multiple `d`-tag streams. The event is signed
|
||||
with the node's FIPS identity key; there is no separate Nostr key. A
|
||||
NIP-40 `expiration` tag is set to now + `advert_ttl_secs`, and a
|
||||
`version` tag carries the protocol version. The advert content is a
|
||||
JSON document shaped as `OverlayAdvert` (see
|
||||
[../reference/nostr-events.md](../reference/nostr-events.md) for the
|
||||
schema).
|
||||
|
||||
Publication happens on startup, again whenever the set of advertised
|
||||
endpoints changes (for example, when a Tor onion hostname first
|
||||
becomes available), and on a refresh timer every `advert_refresh_secs`.
|
||||
If the `advertise` flag is turned off, the previous advert event is
|
||||
deleted using a NIP-9 kind 5 delete event. Advert publication is
|
||||
fan-out: the same event is sent to every relay in `advert_relays` with
|
||||
no explicit failover — relay redundancy is implicit.
|
||||
|
||||
For a UDP or TCP transport with `public: true`, the address advertised
|
||||
follows a fixed precedence: an operator-supplied `external_addr` wins;
|
||||
otherwise a non-wildcard bound `local_addr` is used directly;
|
||||
otherwise — only for UDP — the runtime asks `stun_servers` for the
|
||||
reflexive address of the bound socket and advertises that. TCP has no
|
||||
STUN equivalent, so wildcard-bound TCP without `external_addr`
|
||||
produces a loud WARN and the endpoint is omitted from the advert.
|
||||
|
||||
### Phase 2 — Lookup
|
||||
|
||||
When the node decides to dial a peer that is eligible for Nostr
|
||||
resolution (a `via_nostr` peer, or any peer under `policy: open`), it
|
||||
issues a Nostr REQ filtered by `author = peer_pubkey`, `kind = 37195`,
|
||||
`#d = fips-overlay-v1`. The fetch is time-bounded (~2 s) and runs
|
||||
against all configured `advert_relays` in parallel. The first valid
|
||||
advert wins; adverts whose `protocol` tag does not match the local
|
||||
`app` value are rejected at validation.
|
||||
|
||||
Results are kept in an in-memory cache keyed by author npub. Cache
|
||||
entries carry the advert's expiration time; a periodic prune drops
|
||||
expired entries, and an LRU-by-expiry eviction enforces
|
||||
`advert_cache_max_entries`. A parallel long-lived subscription on the
|
||||
advert relays populates the cache passively, so open-discovery
|
||||
candidates do not require per-dial fetches.
|
||||
|
||||
On cache hit, advert endpoints are appended to the peer's static
|
||||
address list with lower priority; the static list is tried first.
|
||||
|
||||
### Phase 3 — Offer/Answer signaling
|
||||
|
||||
For any endpoint shaped as `udp:nat`, dialing triggers an
|
||||
offer/answer exchange before the first packet is sent. Signaling events
|
||||
are Nostr kind `21059` (ephemeral, not stored by conforming relays),
|
||||
gift-wrapped per [NIP-59](https://github.com/nostr-protocol/nips/blob/master/59.md)
|
||||
and encrypted with [NIP-44](https://github.com/nostr-protocol/nips/blob/master/44.md),
|
||||
so only the intended recipient can decrypt the payload.
|
||||
|
||||
The initiator performs STUN first (see Phase 4), then builds a
|
||||
`TraversalOffer` containing:
|
||||
|
||||
- A unique `sessionId` and a random `nonce` (used to correlate the
|
||||
answer).
|
||||
- Its reflexive address (if STUN succeeded).
|
||||
- Its list of local (private) addresses for same-LAN paths.
|
||||
- The STUN server it used, for informational reporting only.
|
||||
- An `expiresAt` equal to now + `signal_ttl_secs`.
|
||||
|
||||
The offer is sealed to the recipient's npub and published to the peer's
|
||||
preferred signaling relays — the node first tries to resolve the peer's
|
||||
NIP-17 DM relay list (kind 10050), and falls back to `dm_relays` if
|
||||
the inbox-relays fetch fails. Each side also publishes its own inbox
|
||||
relay list on startup so dialers can discover it.
|
||||
|
||||
On the receiving side, an inbound semaphore bounds concurrent offer
|
||||
processing at `max_concurrent_incoming_offers`. When the semaphore is
|
||||
full, the offer is dropped with a warn log; this is the primary guard
|
||||
against offer-spam from a misbehaving or compromised relay. A
|
||||
`sessionId` replay cache (bounded by `seen_sessions_max_entries`, with
|
||||
entries valid for `replay_window_secs`) rejects duplicates.
|
||||
|
||||
The responder runs its own STUN query and replies with a
|
||||
`TraversalAnswer` carrying its reflexive and local addresses plus a
|
||||
`PunchHint { startAtMs, intervalMs, durationMs }` that tells both sides
|
||||
when to begin probing and how aggressively. If the responder has no
|
||||
usable addresses at all, it replies with `accepted: false` and a
|
||||
`reason` string.
|
||||
|
||||
### Phase 4 — UDP hole-punch
|
||||
|
||||
Each side runs STUN (parsing XOR-MAPPED-ADDRESS from the response, all
|
||||
other attributes ignored) on the *same* UDP socket it will later use
|
||||
for punching and for the adopted FMP transport. This is critical: NAT
|
||||
state is per-socket, so the punch has to reuse the socket that taught
|
||||
the NAT about this binding.
|
||||
|
||||
Given its own reflexive + local addresses and the peer's, each side
|
||||
builds a candidate-pair plan that tries, in priority order:
|
||||
|
||||
1. **Reflexive ↔ reflexive.** The classic STUN path. Tried first because
|
||||
it is the only candidate that's reliable across arbitrary network
|
||||
topologies — host candidates from one peer that happen to be
|
||||
reachable from the other (via a corporate VPN, a Tailscale subnet
|
||||
route, or overlapping private address space) will succeed at the
|
||||
socket layer in the punch but fail in the FMP handshake when the
|
||||
return path doesn't match.
|
||||
2. **LAN ↔ LAN.** If both sides share a /24 prefix, same-subnet private
|
||||
addresses are likely reachable directly. Only fires when both peers
|
||||
shared local host candidates (which requires `share_local_candidates`
|
||||
to be enabled — off by default).
|
||||
3. **Mixed.** Reflexive on one side, local on the other — catches
|
||||
hairpin and one-side-public scenarios.
|
||||
|
||||
At `startAtMs` both sides begin sending 24-byte probe packets on the
|
||||
candidate pair(s) at `intervalMs` cadence for up to `durationMs`. A
|
||||
probe carries a 4-byte magic (`NPTC`), a 4-byte sequence, and the
|
||||
first 16 bytes of `SHA256(sessionId)`; both sides can compute the same
|
||||
session hash independently from the public `sessionId`, so no shared
|
||||
secret is needed on the punch path itself. On receiving a valid probe,
|
||||
a side replies with an `NPTA` ack. The first valid probe or ack seen
|
||||
from the far side records the working remote address and completes the
|
||||
attempt.
|
||||
|
||||
On timeout (`attempt_timeout_secs` as overall bound,
|
||||
`punch_duration_ms` as probe window), both sides issue NIP-9 deletes
|
||||
for their offer and answer events and report failure up to the
|
||||
discovery runtime's `BootstrapEvent::Failed` channel.
|
||||
|
||||
### Phase 5 — Adoption
|
||||
|
||||
On success, the discovery runtime emits `BootstrapEvent::Established`
|
||||
carrying the session id, the punch socket, and the learned remote
|
||||
address. `adopt_established_traversal()` in the node lifecycle takes
|
||||
the socket, registers it with the UDP transport layer as a new
|
||||
transport instance, and calls `initiate_connection()` with the peer's
|
||||
FIPS identity as the expected remote. FMP's Noise IK handshake runs on
|
||||
the same socket — there is no "promote link" step between punch and
|
||||
handshake; the punch socket *is* the FMP socket.
|
||||
|
||||
From that moment on, the connection is a normal FMP link and is
|
||||
subject to the usual liveness (MMP heartbeats), rekey, and removal
|
||||
behavior. A link-dead event does not re-enter the discovery runtime
|
||||
automatically; reconnection relies on `auto_reconnect` and the same
|
||||
dial path that triggered the original punch.
|
||||
|
||||
### Auto-connect semantics
|
||||
|
||||
Discovery does not itself initiate connections. It only supplies
|
||||
addresses. Dial attempts originate from the existing peer-connection
|
||||
machinery:
|
||||
|
||||
- **Configured peers** (`peers[]` with `connect_policy: auto_connect`)
|
||||
are dialed on startup and on retry. When `via_nostr` is set, advert
|
||||
endpoints are appended to the dial list with lower priority than
|
||||
static entries.
|
||||
- **Open discovery peers** are assembled from the advert cache, fenced
|
||||
by the peer ACL, and enqueued into a bounded retry queue sized by
|
||||
`open_discovery_max_pending`. There is no event-driven
|
||||
"connect on every advert" — a peer re-enters the queue only when its
|
||||
prior attempt has drained.
|
||||
- **Manual dials** (`fipsctl connect`) can target any configured peer
|
||||
and use the same dial path, including Nostr resolution if configured.
|
||||
|
||||
### Rate limits and safeguards
|
||||
|
||||
| Mechanism | Default | What it prevents | Behavior at limit |
|
||||
| --- | --- | --- | --- |
|
||||
| Offer semaphore (`max_concurrent_incoming_offers`) | 16 | CPU and memory exhaustion from offer spam on DM relays. | Warn log, offer dropped. |
|
||||
| Advert cache (`advert_cache_max_entries`) | 2048 | Memory growth from ambient advert traffic under `policy: open`. | LRU-by-expiry eviction. |
|
||||
| Seen-sessions (`seen_sessions_max_entries`) | 2048 | Replay of stale `sessionId` values. | Oldest entry evicted. |
|
||||
| Signal TTL (`signal_ttl_secs`) | 120 s | Indefinite in-flight offers on relays. | Expired offers rejected at validation. |
|
||||
| Open discovery queue (`open_discovery_max_pending`) | 64 | Unbounded retry queue under ambient advert load. | New candidates skipped until the queue drains. |
|
||||
| Punch window (`punch_duration_ms`) | 10 s | Endless probe traffic after one side has given up. | Attempt declared failed; sockets discarded. |
|
||||
| Failure-streak threshold (`failure_streak_threshold`) | 5 | Repeated traversal attempts against a peer that keeps failing. | Peer enters extended cooldown. |
|
||||
| Extended cooldown (`extended_cooldown_secs`) | 1800 s | Tight retry loops after a failure streak. | Per-peer suppression for the cooldown window. |
|
||||
| WARN log throttle (`warn_log_interval_secs`) | 300 s | Log floods from a peer that fails on every attempt. | One WARN per peer per interval; the rest demote to debug. |
|
||||
| Failure-state cap (`failure_state_max_entries`) | 4096 | Memory growth from per-peer failure tracking. | LRU eviction. |
|
||||
|
||||
The load-shedding mechanisms (`max_concurrent_incoming_offers` and the
|
||||
failure-streak / extended-cooldown pair) are deliberately conservative
|
||||
so that a misbehaving relay cannot flood the node with offers and a
|
||||
chronically unreachable peer cannot keep the traversal pipeline
|
||||
saturated. The remaining rows are capacity bounds.
|
||||
|
||||
Adverts also undergo a stale-advert sweep: cached entries whose
|
||||
`expiresAt` has passed are evicted on the periodic prune tick. Inbound
|
||||
signaling tolerates ±60 s of clock skew between sender and receiver,
|
||||
and the runtime maintains an NTP-style skew estimate per remote so
|
||||
that consistently-skewed relays don't trip the freshness check.
|
||||
|
||||
### Relay model
|
||||
|
||||
All configured relays (advert + DM) are opened on a single
|
||||
`nostr-sdk::Client` at startup. Publication is fan-out: the same event
|
||||
is sent to every relay in the target list, with no explicit retry or
|
||||
relay selection. Redundancy is implicit — a downed relay simply means
|
||||
its copy of the advert or signal is unavailable, while other relays
|
||||
still serve the same data.
|
||||
|
||||
For signaling specifically, the node prefers the recipient's NIP-17
|
||||
DM relays when available (the recipient publishes its DM relay list as
|
||||
a kind 10050 event to its own DM relays on startup) and falls back to
|
||||
the local `dm_relays` list otherwise. This keeps the common case
|
||||
off the sender's DM relays when those are different from the
|
||||
recipient's, at the cost of one extra NIP-17 fetch per offer.
|
||||
|
||||
There is no per-relay rate limiting or health check. The relay model
|
||||
assumes that an operator chooses relays they trust to be best-effort
|
||||
available and that outright misbehavior is handled at the offer
|
||||
semaphore and replay-cache layers downstream.
|
||||
|
||||
## Security and threat model
|
||||
|
||||
- **Relay operators can observe metadata.** They see which npubs
|
||||
publish adverts, to whom offers are sent, and the timing of that
|
||||
traffic. The *contents* of offer and answer events are
|
||||
NIP-59/NIP-44 sealed — only the intended recipient decrypts them.
|
||||
Adverts are public by design.
|
||||
- **STUN servers see the node's public IP and port.** Only the STUN
|
||||
servers listed in the node's own `stun_servers` are ever contacted
|
||||
for reflexive discovery. Peer-advertised STUN values are
|
||||
informational; a malicious peer cannot steer this node to a
|
||||
chosen STUN target. See the doc comment on
|
||||
`node.discovery.nostr.stun_servers`.
|
||||
- **The FIPS identity key signs adverts.** Compromise of
|
||||
`fips.key` is compromise of the node's Nostr identity — an attacker
|
||||
can publish adverts on behalf of the node. The recovery path is
|
||||
the same as for any identity compromise: rotate the key and
|
||||
re-advertise. There is no separate Nostr keypair to rotate
|
||||
independently.
|
||||
- **Tor advertising leaks timing via clearnet relays.** When a
|
||||
Tor-only node advertises its onion address, the advert itself is
|
||||
published on clearnet WebSocket relays. Operators who want full
|
||||
unlinkability between the advertising identity and the node's
|
||||
IP must route relay traffic through Tor as well — for example by
|
||||
running `fips` inside a network namespace with a Tor SOCKS
|
||||
proxy as its only egress, or by pointing `advert_relays` and
|
||||
`dm_relays` at onion relay endpoints.
|
||||
- **Open discovery accepts anyone publishing on the same `app`.**
|
||||
Admission control is the peer ACL, not the discovery layer. Verify
|
||||
the ACL before enabling `policy: open`, and consider using a
|
||||
non-default `app` value to scope visibility.
|
||||
- **Nothing about discovery bypasses FMP.** A successful punch yields
|
||||
a UDP socket with a claimed remote identity. That identity is not
|
||||
trusted until FMP's Noise IK handshake completes. A peer whose
|
||||
advert says "I am npub X at 1.2.3.4:5678" but whose FMP handshake
|
||||
presents a different static key is rejected at the mesh layer.
|
||||
|
||||
## LAN/mDNS discovery
|
||||
|
||||
LAN discovery is a separate, link-local discovery mechanism that finds
|
||||
peers on the same broadcast domain using mDNS / DNS-SD
|
||||
([RFC 6762](https://www.rfc-editor.org/rfc/rfc6762) /
|
||||
[RFC 6763](https://www.rfc-editor.org/rfc/rfc6763)). Unlike
|
||||
Nostr-mediated discovery, it contacts no relay, runs no STUN
|
||||
observation, and performs no NAT traversal: an endpoint learned from a
|
||||
LAN advert is by construction routable from the consumer's own link.
|
||||
The result is sub-second peer pairing on the same LAN.
|
||||
|
||||
It is unrelated to the "LAN candidate" terminology used in the
|
||||
NAT-traversal sections above (which refers to a host's own
|
||||
locally-bound address offered as a hole-punch candidate). LAN/mDNS
|
||||
discovery is a distinct subsystem under `src/discovery/lan/`.
|
||||
|
||||
### Role
|
||||
|
||||
LAN discovery adds two capabilities, both confined to the local link:
|
||||
|
||||
- **Advertising.** The node publishes a `_fips._udp.local.` DNS-SD
|
||||
service advert carrying its `npub`, its protocol version, and (if
|
||||
configured) a discovery scope. The advert is multicast on the local
|
||||
link only; it does not leave the broadcast domain unless the
|
||||
operator's network bridges mDNS.
|
||||
- **Browsing.** The node concurrently browses for the same service
|
||||
type, learns the endpoints of other FIPS nodes on the link, and
|
||||
initiates a normal FMP link to each newly-seen peer.
|
||||
|
||||
The mDNS service type is `_fips._udp.local.`
|
||||
(`src/discovery/lan/mod.rs:45`). Per RFC 6763 the `_udp` label denotes
|
||||
the IP transport used for the advert, not the FIPS upper protocol —
|
||||
both UDP and TCP FIPS endpoints announce under the same service type
|
||||
because the link-layer handshake travels over UDP either way. (In
|
||||
practice LAN discovery dials only over a UDP transport; see the
|
||||
handshake subsection.)
|
||||
|
||||
### When to use it
|
||||
|
||||
- **You run several FIPS nodes on one LAN** (a lab bench, an office
|
||||
segment, a home network) and want them to find each other without
|
||||
hand-maintaining `peers[]` blocks or standing up Nostr discovery.
|
||||
- **You want the lowest-latency pairing path.** Same-link pairing
|
||||
completes in well under a second with no relay round-trip.
|
||||
|
||||
Skip it when nodes are not on a shared broadcast domain (mDNS does not
|
||||
cross routed boundaries), or when you do not want the node to multicast
|
||||
its identity on the local link. LAN discovery is **opt-in and disabled
|
||||
by default**, so doing nothing leaves it off.
|
||||
|
||||
### How it works
|
||||
|
||||
The LAN discovery runtime (`src/discovery/lan/mod.rs`) is started
|
||||
during node initialization when `node.discovery.lan.enabled` is true.
|
||||
It is independent of Nostr discovery and runs even when Nostr is
|
||||
disabled (`src/node/lifecycle.rs:1159-1162`). Startup requires an
|
||||
operational UDP transport: the node advertises the port of its
|
||||
lowest-`TransportId` operational, non-bootstrap UDP transport, chosen
|
||||
deterministically so the advertised port is stable across restarts
|
||||
(`src/node/lifecycle.rs:1169-1180`). If no such port exists, the
|
||||
runtime returns `NoAdvertisedPort` and LAN discovery does not start
|
||||
(`src/discovery/lan/mod.rs:156-158`).
|
||||
|
||||
The runtime does two things concurrently:
|
||||
|
||||
1. **Responder.** It registers a DNS-SD service with instance name
|
||||
`fips-<first-16-chars-of-npub>` and a TXT record carrying the keys
|
||||
below. `mdns-sd`'s address auto-detection appends every non-loopback
|
||||
interface address, with `127.0.0.1` seeded so same-host peers and
|
||||
integration tests can still resolve the advert
|
||||
(`src/discovery/lan/mod.rs:182-203`).
|
||||
2. **Browser.** A background pump receives `ServiceResolved` events for
|
||||
the same service type. For each resolved advert it extracts the
|
||||
`npub` and `scope` TXT values, drops adverts that echo the node's own
|
||||
npub, drops cross-scope adverts (see scope filtering), drops records
|
||||
without an `npub`, and surfaces one `LanDiscoveredPeer` per routable
|
||||
interface address (`src/discovery/lan/mod.rs:212-299`). IPv6
|
||||
unicast link-local addresses without an interface scope id are
|
||||
skipped, since they cannot be dialed unambiguously
|
||||
(`src/discovery/lan/mod.rs:348-365`).
|
||||
|
||||
The TXT record carries three keys (`src/discovery/lan/mod.rs:47-55`):
|
||||
|
||||
| TXT key | Contents |
|
||||
| --- | --- |
|
||||
| `npub` | bech32-encoded npub of the advertising node |
|
||||
| `scope` | the node's discovery scope, if one is configured (omitted otherwise) |
|
||||
| `v` | FIPS protocol version (the same `PROTOCOL_VERSION` used by the Nostr advert) |
|
||||
|
||||
Once per node tick, the node drains browser events and acts on them in
|
||||
`poll_lan_discovery()` (`src/node/lifecycle.rs:907`, called from
|
||||
`src/node/dataplane/rx_loop.rs:266`). For each discovered peer it finds
|
||||
a UDP transport whose family matches the peer address, parses the
|
||||
`npub` into a `PeerIdentity`, skips peers it is already connected to or
|
||||
currently connecting to, and otherwise initiates a connection.
|
||||
|
||||
### Handshake: Noise IK
|
||||
|
||||
LAN-discovered peers are dialed through the standard FMP outbound link
|
||||
path. `poll_lan_discovery()` calls `initiate_connection()`
|
||||
(`src/node/lifecycle.rs:380`), which, for connectionless transports
|
||||
such as UDP, allocates a link and **starts the Noise IK handshake**
|
||||
(documented at `src/node/lifecycle.rs:373-374`). This is the same
|
||||
link-layer handshake used by every other FMP connection — IK at the
|
||||
link layer per the FIPS architecture — not a different pattern for LAN
|
||||
peers.
|
||||
|
||||
The mDNS advert is **unauthenticated**: anyone on the link can
|
||||
multicast a TXT claiming any `npub`. Identity is proven end-to-end by
|
||||
the Noise IK handshake against the observed endpoint. A spoofed advert
|
||||
carrying another node's npub fails the handshake — the impostor does
|
||||
not hold the matching static key — and the half-open link is dropped.
|
||||
The mDNS advert is therefore a routing hint, never an identity
|
||||
assertion, exactly as a Nostr advert is treated (a successful contact
|
||||
is not trusted until FMP's Noise IK handshake completes).
|
||||
|
||||
> Note: a stale source doc-comment at `src/node/lifecycle.rs:904-906`
|
||||
> describes this path as a "Noise XX" handshake. That comment is
|
||||
> inaccurate — the path uses Noise IK as described above. The comment
|
||||
> is flagged for a separate source fix and does not reflect actual
|
||||
> behavior.
|
||||
|
||||
### Scope filtering
|
||||
|
||||
When a discovery scope is configured, the advert carries it in the
|
||||
`scope` TXT entry and the browser surfaces only peers whose advert
|
||||
carries a matching scope. Nodes on the same physical LAN but configured
|
||||
for different mesh networks therefore do not cross-feed each other.
|
||||
|
||||
The scope is resolved by `lan_discovery_scope()`
|
||||
(`src/node/lifecycle.rs:880-902`): the explicit
|
||||
`node.discovery.lan.scope`, if non-empty, is used directly. Otherwise
|
||||
the node falls back to deriving a scope from the Nostr discovery `app`
|
||||
tag (stripping the `fips-overlay-v1:` prefix when present). This lets
|
||||
an application keep its public, relay-visible Nostr `app` tag generic
|
||||
while still isolating LAN discovery per private network, or share one
|
||||
value across both. A node with no scope on either side surfaces all
|
||||
adverts it sees on the link.
|
||||
|
||||
### Configuration
|
||||
|
||||
LAN discovery is configured under `node.discovery.lan.*`
|
||||
(`src/config/node.rs:222-227`, `src/discovery/lan/mod.rs:88-129`):
|
||||
|
||||
| Key | Type | Default | Meaning |
|
||||
| --- | --- | --- | --- |
|
||||
| `node.discovery.lan.enabled` | bool | `false` | Master switch. LAN discovery is opt-in; default-off avoids an unexpected per-link identity multicast on upgrade. |
|
||||
| `node.discovery.lan.service_type` | string | `_fips._udp.local.` | DNS-SD service type. Overridable mainly so integration tests can isolate multiple services on one loopback interface. |
|
||||
| `node.discovery.lan.scope` | string (optional) | unset | Application/network scope carried in the LAN-only `scope` TXT record. Kept deliberately separate from the public Nostr `app` tag. When unset, the scope falls back to the derived Nostr `app` value. |
|
||||
|
||||
The identity surface published over mDNS (`npub`, version, optional
|
||||
scope) is a strict subset of what `nostr.advertise` already publishes
|
||||
publicly, so enabling LAN discovery adds no marginal privacy cost
|
||||
beyond making the node's presence observable on its own local link.
|
||||
|
||||
### Relationship to Nostr discovery
|
||||
|
||||
The two mechanisms are complementary and independent:
|
||||
|
||||
| | Nostr-mediated | LAN/mDNS |
|
||||
| --- | --- | --- |
|
||||
| Reach | Internet-wide, via relays | Same broadcast domain only |
|
||||
| Signaling channel | Public Nostr relays | mDNS multicast on the local link |
|
||||
| NAT traversal | STUN + UDP hole-punch for `udp:nat` peers | None — endpoint is link-routable by construction |
|
||||
| Identity carrier | signed kind 37195 advert (authenticated at publish) | unauthenticated mDNS TXT (routing hint only) |
|
||||
| Identity proof | FMP Noise IK on the connection | FMP Noise IK on the connection |
|
||||
| Default | disabled (`nostr.enabled: false`) | disabled (`lan.enabled: false`) |
|
||||
| Scope key | `app` tag (public) | `scope` TXT (link-local), falls back to `app` |
|
||||
|
||||
Both ultimately converge on the same trust boundary: discovery only
|
||||
supplies candidate endpoints, and no peer is trusted until FMP's Noise
|
||||
IK handshake confirms the claimed identity. A node may run both at
|
||||
once — for example, advertising globally over Nostr while also pairing
|
||||
instantly with same-LAN peers — with no interaction between the two
|
||||
beyond the shared scope fallback.
|
||||
|
||||
## See also
|
||||
|
||||
- [../how-to/enable-nostr-discovery.md](../how-to/enable-nostr-discovery.md)
|
||||
— operator activation recipes grouped under three capabilities
|
||||
(resolve, advertise, open) across five scenarios.
|
||||
- [../tutorials/resolve-peers-via-nostr.md](../tutorials/resolve-peers-via-nostr.md),
|
||||
[../tutorials/advertise-your-node.md](../tutorials/advertise-your-node.md),
|
||||
and [../tutorials/open-discovery.md](../tutorials/open-discovery.md)
|
||||
— hand-held walkthroughs of the three capabilities, in
|
||||
pedagogical order.
|
||||
- [../reference/configuration.md](../reference/configuration.md) — full
|
||||
configuration reference, including all surrounding keys elided from
|
||||
the scenarios above.
|
||||
- [../reference/nostr-events.md](../reference/nostr-events.md) — Kind
|
||||
37195 (overlay advert), Kind 21059 (gift-wrapped traversal
|
||||
signaling), Kind 10050 (NIP-17 inbox relay list).
|
||||
- [../reference/security.md](../reference/security.md) — consolidated
|
||||
security reference, including how the FIPS identity key signs both
|
||||
adverts and Noise handshakes.
|
||||
- [fips-transport-layer.md](fips-transport-layer.md) — UDP, TCP, and
|
||||
Tor transport mechanics; the punch socket is adopted as a normal
|
||||
UDP transport after handoff.
|
||||
- [fips-mesh-layer.md](fips-mesh-layer.md) — FMP Noise IK handshake
|
||||
that runs on the adopted socket.
|
||||
- [port-advertisement-and-nat-traversal.md](port-advertisement-and-nat-traversal.md)
|
||||
— generic protocol reference (event tags, NIP usage, on-the-wire
|
||||
offer/answer schema, failure-suppression machinery), with the
|
||||
FIPS-specific values called out as worked examples.
|
||||
341
docs/design/fips-prior-work.md
Normal file
@@ -0,0 +1,341 @@
|
||||
# FIPS Prior Work and References
|
||||
|
||||
FIPS builds on proven designs rather than inventing new cryptography or
|
||||
routing algorithms. Nearly every major design decision has deployed
|
||||
precedent. This document collects the relevant prior art, organized by
|
||||
the FIPS subsystem that draws on it, and gathers the academic and
|
||||
standards references cited from the per-subsystem design docs.
|
||||
|
||||
## Spanning Tree Self-Organization
|
||||
|
||||
The idea that distributed nodes can build a spanning tree through
|
||||
purely local decisions — each node selecting a parent based on
|
||||
announcements from its neighbors — dates to the
|
||||
[IEEE 802.1D Spanning Tree Protocol](https://en.wikipedia.org/wiki/Spanning_Tree_Protocol)
|
||||
(STP, 1985). STP demonstrated that a network-wide tree emerges from a
|
||||
simple deterministic rule (lowest bridge ID wins root election)
|
||||
applied independently at each node. FIPS uses the same principle —
|
||||
lowest node address determines the root — adapted from an Ethernet
|
||||
bridging context to a general-purpose overlay mesh.
|
||||
|
||||
## Tree Coordinate Routing
|
||||
|
||||
The spanning tree coordinates, bloom filter candidate selection, and
|
||||
greedy routing algorithms are adapted from
|
||||
[Yggdrasil v0.5](https://yggdrasil-network.github.io/2023/10/22/upcoming-v05-release.html)
|
||||
and its [Ironwood](https://github.com/Arceliar/ironwood) routing
|
||||
library. Yggdrasil's key insight was using the tree path from root to
|
||||
node as a routable coordinate, enabling greedy forwarding without
|
||||
global routing tables. FIPS adapts these algorithms for
|
||||
multi-transport operation, Nostr identity integration, and constrained
|
||||
MTU environments.
|
||||
|
||||
The theoretical foundation for greedy routing on tree embeddings draws
|
||||
on [Kleinberg's work](https://www.cs.cornell.edu/home/kleinber/swn.pdf)
|
||||
on navigable small-world networks, which showed that greedy forwarding
|
||||
succeeds in O(log² n) steps when the network has hierarchical
|
||||
structure. Thorup-Zwick compact routing schemes separately demonstrated
|
||||
that sublinear routing state is achievable with bounded stretch,
|
||||
motivating the use of tree coordinates rather than full routing tables.
|
||||
|
||||
## Split-Horizon Bloom Filter Propagation
|
||||
|
||||
FIPS distributes reachability information using bloom filters computed
|
||||
with a split-horizon rule: when advertising to a peer, exclude that
|
||||
peer's own contributions. This technique is borrowed from
|
||||
distance-vector routing protocols —
|
||||
[RIP](https://en.wikipedia.org/wiki/Routing_Information_Protocol)
|
||||
(1988) and [Babel](https://www.irif.fr/~jch/software/babel/) use
|
||||
split-horizon to prevent routing loops by not advertising a route back
|
||||
to the neighbor it was learned from. FIPS applies the same principle
|
||||
to probabilistic set advertisements rather than distance-vector tables.
|
||||
|
||||
## Cryptographic Identity as Network Address
|
||||
|
||||
FIPS nodes are identified by their Nostr public keys (secp256k1). The
|
||||
network address *is* the cryptographic identity — there is no separate
|
||||
address assignment or registration step.
|
||||
[CJDNS](https://github.com/cjdelisle/cjdns) pioneered this approach in
|
||||
overlay meshes, deriving IPv6 addresses from the double-SHA-512 of
|
||||
each node's public key. Tor [.onion
|
||||
addresses](https://spec.torproject.org/rend-spec-v3) and the IETF
|
||||
[Host Identity Protocol](https://en.wikipedia.org/wiki/Host_Identity_Protocol)
|
||||
(HIP) follow the same principle. FIPS uses Nostr's existing key
|
||||
infrastructure rather than introducing a new identity scheme.
|
||||
|
||||
## Dual-Layer Encryption
|
||||
|
||||
FIPS encrypts traffic twice: FMP provides hop-by-hop link encryption
|
||||
(protecting against transport-layer observers), while FSP provides
|
||||
independent end-to-end session encryption (protecting against
|
||||
intermediate FIPS nodes). This layered approach mirrors
|
||||
[Tor](https://www.torproject.org/), where each relay peels one layer
|
||||
of encryption (hop-by-hop) while the innermost layer protects
|
||||
end-to-end payload. [I2P](https://geti2p.net/) uses a similar garlic
|
||||
routing scheme with tunnel-layer and end-to-end encryption. Unlike Tor
|
||||
and I2P, FIPS does not provide anonymity — its dual encryption
|
||||
protects confidentiality and integrity rather than hiding traffic
|
||||
patterns.
|
||||
|
||||
## Noise Protocol Framework
|
||||
|
||||
FIPS uses the [Noise Protocol Framework](https://noiseprotocol.org/)
|
||||
at both protocol layers, with different handshake patterns chosen for
|
||||
each layer's threat model. FMP link encryption uses **Noise IK**,
|
||||
providing mutual authentication with a single round trip where the
|
||||
initiator knows the responder's static key in advance.
|
||||
[WireGuard](https://www.wireguard.com/) uses the same IK base pattern
|
||||
(extended with a pre-shared key as IKpsk2) for VPN tunnels. FSP
|
||||
session encryption uses **Noise XK**, the same pattern used by the
|
||||
[Lightning Network](https://github.com/lightning/bolts/blob/master/08-transport.md),
|
||||
where the initiator's static key is transmitted in a third message
|
||||
rather than the first. XK provides stronger initiator identity hiding
|
||||
at the cost of an additional round trip — a worthwhile tradeoff for
|
||||
session-layer traffic that traverses untrusted intermediate nodes. At
|
||||
the link layer, where both peers are configured and directly
|
||||
connected, IK's single round trip is preferred.
|
||||
|
||||
Specific Noise references and adapted constructions:
|
||||
|
||||
- Perrin, T. ["The Noise Protocol Framework"](https://noiseprotocol.org/noise.html).
|
||||
Revision 34, 2018. *Framework for building crypto protocols using
|
||||
Diffie-Hellman key agreement and AEAD ciphers. FSP uses the XK
|
||||
handshake pattern.*
|
||||
|
||||
- Donenfeld, J.A. ["WireGuard: Next Generation Kernel Network Tunnel"](https://www.wireguard.com/papers/wireguard.pdf).
|
||||
NDSS 2017. *Transport-independent cryptographic sessions bound to
|
||||
identity keys rather than network addresses; AEAD-only authentication
|
||||
model.*
|
||||
|
||||
## Index-Based Session Dispatch
|
||||
|
||||
FIPS uses locally-assigned 32-bit session indices to demultiplex
|
||||
incoming packets to the correct cryptographic session in O(1) time,
|
||||
without parsing source addresses or performing expensive lookups.
|
||||
This directly follows
|
||||
[WireGuard's](https://www.wireguard.com/papers/wireguard.pdf) receiver
|
||||
index approach, where each peer assigns a random index during
|
||||
handshake and the remote side includes it in every packet header.
|
||||
|
||||
## Replay Protection Over Unreliable Transports
|
||||
|
||||
FSP and FMP both use explicit per-packet counters with a sliding
|
||||
bitmap window for replay protection — the standard DTLS approach,
|
||||
chosen because implicit nonce counters desynchronize permanently under
|
||||
UDP packet loss or reordering.
|
||||
|
||||
- Rescorla, E., Modadugu, N. [RFC 6347](https://datatracker.ietf.org/doc/html/rfc6347):
|
||||
"Datagram Transport Layer Security Version 1.2". 2012. *Explicit
|
||||
sequence numbers with sliding bitmap window for replay protection
|
||||
over unreliable transports.*
|
||||
|
||||
## Transport-Agnostic Overlay Mesh
|
||||
|
||||
FIPS is designed to operate over any datagram-capable transport — UDP,
|
||||
raw Ethernet, Bluetooth, radio, serial — through a uniform transport
|
||||
abstraction. Several mesh overlays have demonstrated transport-agnostic
|
||||
design: [CJDNS](https://github.com/cjdelisle/cjdns) runs over UDP and
|
||||
Ethernet, [Yggdrasil](https://yggdrasil-network.github.io/) supports
|
||||
TCP and TLS transports, and [Tor](https://www.torproject.org/) can use
|
||||
pluggable transports to tunnel through various media. FIPS extends
|
||||
this pattern to shared-medium transports (radio, BLE) with
|
||||
per-transport MTU and discovery capabilities.
|
||||
|
||||
## Metrics Measurement Protocol
|
||||
|
||||
MMP's design assembles well-established measurement techniques into a
|
||||
unified per-link protocol. The SenderReport/ReceiverReport exchange
|
||||
structure follows [RTCP](https://www.rfc-editor.org/rfc/rfc3550)
|
||||
(RFC 3550), which uses the same report pairing for media stream
|
||||
quality monitoring in RTP sessions. MMP's jitter computation uses the
|
||||
RTCP interarrival jitter algorithm directly.
|
||||
|
||||
The smoothed RTT estimator uses the Jacobson/Karels algorithm
|
||||
([RFC 6298](https://www.rfc-editor.org/rfc/rfc6298)), the same SRTT
|
||||
computation used in TCP for retransmission timeout calculation since
|
||||
1988. MMP derives RTT from timestamp-echo in ReceiverReports with
|
||||
dwell-time compensation, rather than from packet round-trips.
|
||||
|
||||
The spin bit in the FMP frame header follows the
|
||||
[QUIC](https://www.rfc-editor.org/rfc/rfc9000) spin bit
|
||||
([RFC 9312](https://www.rfc-editor.org/rfc/rfc9312)) — a single bit
|
||||
that alternates each round trip, enabling passive latency measurement.
|
||||
FIPS implements the spin bit state machine but relies on
|
||||
timestamp-echo for SRTT, as irregular mesh traffic makes spin bit RTT
|
||||
unreliable.
|
||||
|
||||
The Expected Transmission Count (ETX) metric, computed from
|
||||
bidirectional delivery ratios, was introduced by
|
||||
[De Couto et al. (2003)](https://pdos.csail.mit.edu/papers/grid:mobicom03/paper.pdf)
|
||||
for wireless mesh routing and is used in protocols including
|
||||
[OLSR](https://en.wikipedia.org/wiki/Optimized_Link_State_Routing_Protocol)
|
||||
and [Babel](https://www.irif.fr/~jch/software/babel/). FIPS computes
|
||||
ETX per-link from MMP loss measurements for future use in candidate
|
||||
ranking.
|
||||
|
||||
The CE (Congestion Experienced) echo flag provides hop-by-hop
|
||||
[ECN](https://en.wikipedia.org/wiki/Explicit_Congestion_Notification)
|
||||
signaling, following the TCP/IP ECN echo pattern (RFC 3168). Transit
|
||||
nodes detect congestion via MMP loss/ETX metrics or kernel buffer
|
||||
drops and set the CE flag on forwarded frames; destination nodes mark
|
||||
ECN-capable IPv6 packets accordingly.
|
||||
|
||||
## Path MTU Discovery
|
||||
|
||||
FSP adapts RFC 1191 Path MTU Discovery for overlay networks. The
|
||||
classic ICMP Packet Too Big mechanism is replaced by a transit-node
|
||||
`min()` propagation in SessionDatagram and LookupResponse plus an
|
||||
end-to-end PathMtuNotification echo back to the source.
|
||||
|
||||
- Mogul, J., Deering, S. [RFC 1191](https://datatracker.ietf.org/doc/html/rfc1191):
|
||||
"Path MTU Discovery". 1990. *End-to-end path MTU discovery; FSP
|
||||
adapts this for overlay networks using transit-node min()
|
||||
propagation.*
|
||||
|
||||
## Session Restart and Simultaneous Initiation
|
||||
|
||||
FSP's epoch-based peer restart detection mirrors IKEv2's
|
||||
INITIAL_CONTACT notification, and its lowest-address-wins
|
||||
simultaneous-initiation tie-breaker mirrors IKEv2's resolution rule.
|
||||
|
||||
- Kaufman, C., Hoffman, P., Nir, Y., Eronen, P., Kivinen, T.
|
||||
[RFC 7296](https://datatracker.ietf.org/doc/html/rfc7296):
|
||||
"Internet Key Exchange Protocol Version 2 (IKEv2)". 2014.
|
||||
*Simultaneous initiation resolution (§2.8) and INITIAL_CONTACT peer
|
||||
restart detection (§2.4).*
|
||||
|
||||
## Hybrid Coordinate Warmup
|
||||
|
||||
FSP's hybrid coordinate warmup (CP flag piggybacking + standalone
|
||||
CoordsWarmup) draws on Yggdrasil's approach of embedding coordinates
|
||||
in session traffic to keep transit caches populated.
|
||||
|
||||
- [Yggdrasil Network](https://yggdrasil-network.github.io/).
|
||||
*Coordinate-based overlay routing with session traffic used to warm
|
||||
transit node coordinate caches.*
|
||||
|
||||
## Cryptographic Primitives
|
||||
|
||||
FIPS reuses [Nostr's](https://github.com/nostr-protocol/nips)
|
||||
cryptographic stack — secp256k1 for identity keys, Schnorr signatures
|
||||
for authentication, SHA-256 for hashing, and ChaCha20-Poly1305 for
|
||||
authenticated encryption. This is the same primitive set used across
|
||||
Bitcoin, Nostr, and a growing ecosystem of self-sovereign identity
|
||||
systems. No novel cryptography is introduced.
|
||||
|
||||
## Spanning-Tree Dynamics: Foundations
|
||||
|
||||
The CRDT framing, gossip dissemination, failure detection, link
|
||||
metrics, and route stability mechanisms in
|
||||
[spanning-tree-dynamics.md](spanning-tree-dynamics.md) draw on a body
|
||||
of academic and standards work, summarized below.
|
||||
|
||||
### Virtual Coordinate Routing
|
||||
|
||||
- Rao, A., Ratnasamy, S., Papadimitriou, C., Shenker, S., Stoica, I.
|
||||
["Geographic Routing without Location Information"](https://people.eecs.berkeley.edu/~sylvia/papers/p327-rao.pdf).
|
||||
MobiCom 2003. *Established virtual coordinate routing using network
|
||||
topology.*
|
||||
|
||||
### Greedy Embedding Theory
|
||||
|
||||
- Kleinberg, R.
|
||||
["Geographic Routing Using Hyperbolic Space"](https://www.semanticscholar.org/paper/Geographic-Routing-Using-Hyperbolic-Space-Kleinberg/f506b2ddb142d2ec539400297ba53383d958abef).
|
||||
IEEE INFOCOM 2007. *Proved every connected graph has a greedy
|
||||
embedding in hyperbolic space; showed spanning trees enable
|
||||
coordinate assignment.*
|
||||
|
||||
- Cvetkovski, A., Crovella, M.
|
||||
["Hyperbolic Embedding and Routing for Dynamic Graphs"](https://www.cs.bu.edu/faculty/crovella/paper-archive/infocom09-hyperbolic.pdf).
|
||||
IEEE INFOCOM 2009. *Dynamic embedding for nodes joining/leaving;
|
||||
introduced Gravity-Pressure routing for failure recovery.*
|
||||
|
||||
- Crovella, M. et al.
|
||||
["On the Choice of a Spanning Tree for Greedy Embedding"](https://www.cs.bu.edu/faculty/crovella/paper-archive/networking-science13.pdf).
|
||||
Networking Science 2013. *Analysis of how tree structure affects
|
||||
routing stretch.*
|
||||
|
||||
- Bläsius, T. et al.
|
||||
["Hyperbolic Embeddings for Near-Optimal Greedy Routing"](https://dl.acm.org/doi/10.1145/3381751).
|
||||
ACM Journal of Experimental Algorithmics 2020. *Achieved 100%
|
||||
success ratio with 6% stretch on Internet graph.*
|
||||
|
||||
### Link Metrics
|
||||
|
||||
- De Couto, D., Aguayo, D., Bicket, J., Morris, R.
|
||||
"A High-Throughput Path Metric for Multi-Hop Wireless Routing".
|
||||
MobiCom 2003. *Introduced ETX (Expected Transmission Count) as a
|
||||
link quality metric for wireless mesh networks.*
|
||||
|
||||
### Routing Protocol Stability
|
||||
|
||||
- IEEE 802.1D. "IEEE Standard for Local and Metropolitan Area
|
||||
Networks: Media Access Control (MAC) Bridges". *Spanning Tree
|
||||
Protocol (STP) — root election via bridge ID, BPDU exchange.*
|
||||
|
||||
- Moy, J. [RFC 2328](https://datatracker.ietf.org/doc/html/rfc2328):
|
||||
"OSPF Version 2". 1998. *Link-state routing with cumulative path
|
||||
costs and SPF computation. FIPS's local-only cost approach is
|
||||
contrasted with OSPF's cumulative model in
|
||||
[spanning-tree-dynamics.md §8](spanning-tree-dynamics.md#8-parent-selection).*
|
||||
|
||||
### Distributed Systems Primitives
|
||||
|
||||
- Shapiro, M., Preguiça, N., Baquero, C., Zawirski, M.
|
||||
"Conflict-free Replicated Data Types". SSS 2011. *Formal definition
|
||||
of CRDTs enabling coordination-free consistency.*
|
||||
|
||||
- Das, A., Gupta, I., Motivala, A.
|
||||
["SWIM: Scalable Weakly-consistent Infection-style Process Group Membership"](https://www.cs.cornell.edu/projects/Quicksilver/public_pdfs/SWIM.pdf).
|
||||
IPDPS 2002. *O(1) failure detection, O(log N) dissemination via
|
||||
gossip.*
|
||||
|
||||
- Kermarrec, A-M.
|
||||
["Gossiping in Distributed Systems"](https://www.distributed-systems.net/my-data/papers/2007.osr.pdf).
|
||||
ACM SIGOPS Operating Systems Review 2007. *Framework for
|
||||
gossip-based protocols achieving O(log N) propagation.*
|
||||
|
||||
## FIPS Contributions
|
||||
|
||||
The protocol builds on these foundations and adds several new elements:
|
||||
|
||||
- Cost-aware parent selection using local-only link metrics
|
||||
(`effective_depth = depth + link_cost`), replacing Yggdrasil's
|
||||
depth-only selection
|
||||
- Combined ETX + SRTT link cost formula with MMP-measured components
|
||||
- Flap dampening with mandatory switch bypass
|
||||
- Announcement suppression for transient state changes
|
||||
- Tree-only bloom filter merge with split-horizon exclusion
|
||||
- Hybrid coordinate warmup (CP flag piggybacking plus standalone
|
||||
CoordsWarmup) layered on top of SessionSetup self-bootstrapping
|
||||
- Bloom-guided tree routing for discovery (vs. flooding)
|
||||
- Reverse-path routing for LookupResponse via `recent_requests`
|
||||
|
||||
## External Reference Index
|
||||
|
||||
| Reference | Used by |
|
||||
| --------- | ------- |
|
||||
| [IEEE 802.1D STP](https://en.wikipedia.org/wiki/Spanning_Tree_Protocol) | spanning tree, root election |
|
||||
| [Yggdrasil v0.5](https://yggdrasil-network.github.io/2023/10/22/upcoming-v05-release.html) | tree coordinates, greedy routing |
|
||||
| [Ironwood](https://github.com/Arceliar/ironwood) | tree coordinates, candidate ranking |
|
||||
| [Kleinberg, Small-world](https://www.cs.cornell.edu/home/kleinber/swn.pdf) | greedy routing on tree embeddings |
|
||||
| [CJDNS](https://github.com/cjdelisle/cjdns) | cryptographic-identity-as-address |
|
||||
| [Tor](https://www.torproject.org/) | onion address scheme, dual-layer encryption |
|
||||
| [I2P](https://geti2p.net/) | dual-layer encryption (garlic routing) |
|
||||
| [HIP](https://en.wikipedia.org/wiki/Host_Identity_Protocol) | identity-as-address |
|
||||
| [Babel](https://www.irif.fr/~jch/software/babel/) | split-horizon, ETX |
|
||||
| [RIP](https://en.wikipedia.org/wiki/Routing_Information_Protocol) | split-horizon |
|
||||
| [Noise Framework](https://noiseprotocol.org/) | FMP IK, FSP XK |
|
||||
| [WireGuard](https://www.wireguard.com/) | IK pattern, receiver-index dispatch, identity-bound sessions |
|
||||
| [Lightning BOLT #8](https://github.com/lightning/bolts/blob/master/08-transport.md) | XK pattern |
|
||||
| [QUIC (RFC 9000)](https://www.rfc-editor.org/rfc/rfc9000) | spin bit, transport design |
|
||||
| [QUIC Spin Bit (RFC 9312)](https://www.rfc-editor.org/rfc/rfc9312) | passive RTT measurement |
|
||||
| [RTCP (RFC 3550)](https://www.rfc-editor.org/rfc/rfc3550) | sender/receiver report structure, jitter algorithm |
|
||||
| [TCP SRTT/RTO (RFC 6298)](https://www.rfc-editor.org/rfc/rfc6298) | Jacobson/Karels SRTT |
|
||||
| [ECN (RFC 3168)](https://www.rfc-editor.org/rfc/rfc3168) | CE echo |
|
||||
| [DTLS 1.2 (RFC 6347)](https://datatracker.ietf.org/doc/html/rfc6347) | replay window |
|
||||
| [IKEv2 (RFC 7296)](https://datatracker.ietf.org/doc/html/rfc7296) | INITIAL_CONTACT, simultaneous-initiation tie-breaker |
|
||||
| [PMTUD (RFC 1191)](https://datatracker.ietf.org/doc/html/rfc1191) | adapted PMTUD |
|
||||
| [ETX paper, De Couto et al.](https://pdos.csail.mit.edu/papers/grid:mobicom03/paper.pdf) | ETX metric |
|
||||
| [OLSR](https://en.wikipedia.org/wiki/Optimized_Link_State_Routing_Protocol) | ETX in mesh routing |
|
||||
| [Nostr](https://github.com/nostr-protocol/nips) | identity stack |
|
||||
215
docs/design/fips-security.md
Normal file
@@ -0,0 +1,215 @@
|
||||
# FIPS Mesh-Interface Security
|
||||
|
||||
This document describes the threat model and design rationale for the
|
||||
operator-facing security posture of the `fips0` mesh interface on Linux.
|
||||
The default-deny nftables baseline shipped as `/etc/fips/fips.nft` is the
|
||||
artifact discussed below; for the operator activation steps and drop-in
|
||||
extension recipes, see [enable-mesh-firewall.md](../how-to/enable-mesh-firewall.md).
|
||||
|
||||
The baseline is a documented operator conffile, not an auto-loaded
|
||||
package side-effect. Activation is an explicit one-liner. The
|
||||
rationale for that design follows.
|
||||
|
||||
## Threat Model for `fips0`
|
||||
|
||||
The mesh is a flat layer-3 segment. Every mesh node that can route to
|
||||
you can deliver packets to your `fips0` address — your direct peers
|
||||
forward traffic from non-peer mesh nodes onto your `fips0` the same
|
||||
way any router forwards transit traffic. Identity on the mesh is the
|
||||
originating node's npub — the FMP link layer authenticates direct
|
||||
peers with Noise IK and the FSP session layer authenticates session
|
||||
endpoints with Noise XK — but identity is **not** authorization.
|
||||
Knowing who sent a packet does not, by itself, decide whether the
|
||||
local host should accept it.
|
||||
|
||||
That means: any service on a mesh host that binds to a wildcard
|
||||
address (`0.0.0.0`, `[::]`, or any IPv6 address that includes the
|
||||
`fips0` interface in its scope) is reachable from every mesh node
|
||||
that can route to you by default, not only from your direct peers.
|
||||
There is no NAT, no perimeter firewall, no "local-only" address
|
||||
space between you and an arbitrary mesh node. The mesh is closer
|
||||
to a shared LAN than to the public internet.
|
||||
|
||||
Compare to the corresponding internet trust assumptions:
|
||||
|
||||
| Surface | Public internet | FIPS mesh (no baseline) |
|
||||
|---|---|---|
|
||||
| Reachability from arbitrary mesh node | Mediated by NAT, firewalls, ISPs | Direct |
|
||||
| Default identity | None | Originating node's npub (authenticated) |
|
||||
| Default authorization | None | None |
|
||||
| Accidental exposure cost | Low (NAT hides you) | High (every mesh node sees you) |
|
||||
|
||||
The third row is the gap this document closes. The default-deny
|
||||
baseline removes "accidental exposure" from the failure modes an
|
||||
operator has to think about.
|
||||
|
||||
## The Default-Deny Baseline
|
||||
|
||||
The shipped baseline is `/etc/fips/fips.nft`. It defines a single
|
||||
nftables table, `inet fips`, with one chain hooked at `input`. The
|
||||
chain:
|
||||
|
||||
1. Returns immediately for any packet not arriving on `fips0`. This
|
||||
makes the table a no-op for every other interface — Docker, Tor,
|
||||
the host's main filter table, OPNsense, anything.
|
||||
2. Accepts packets that conntrack identifies as `established` or
|
||||
`related`. Replies to outbound flows initiated from the mesh host
|
||||
come back; ICMPv6 errors related to existing flows (Packet Too
|
||||
Big, Destination Unreachable) come back.
|
||||
3. Accepts ICMPv6 echo-request, so `ping6` reachability tests work.
|
||||
4. Includes operator drop-ins from `/etc/fips/fips.d/*.nft`. An empty
|
||||
directory is fine — the include glob simply matches nothing.
|
||||
5. Falls through to `counter drop`. Every dropped packet increments
|
||||
the counter, visible via `nft list table inet fips`.
|
||||
|
||||
Outbound from `fips0` is unrestricted. The baseline is concerned only
|
||||
with what the mesh host accepts, not what it sends.
|
||||
|
||||
The file is a documented dpkg conffile. Operator edits to
|
||||
`/etc/fips/fips.nft` are preserved across upgrades, the same way
|
||||
edits to `/etc/fips/fips.yaml` and `/etc/fips/hosts` are preserved.
|
||||
If the packaged baseline is ever updated upstream, dpkg prompts the
|
||||
operator on upgrade rather than silently overwriting local changes.
|
||||
|
||||
The canonical artifact is the file itself; read it for the inline
|
||||
documentation that the rest of this document references.
|
||||
|
||||
## Why no auto-load on package install
|
||||
|
||||
The `postinst` script does **not** enable `fips-firewall.service`.
|
||||
This is deliberate. Quietly mutating host firewall state on package
|
||||
install is hostile on every axis that matters: it surprises operators
|
||||
who already have their own nftables ruleset, it can collide with
|
||||
podman/Docker/OPNsense integrations even though the early-return
|
||||
makes it technically safe, and it converts an explicit security
|
||||
decision into an invisible one. The mesh-interface filter belongs to
|
||||
the operator, not to the package's `postinst`.
|
||||
|
||||
The activation gesture is one short, well-formed command. The
|
||||
rationale is documented in the file's inline header and in this
|
||||
document. That is enough; auto-loading would trade discoverability
|
||||
for no real gain.
|
||||
|
||||
## Coexistence with other firewalls
|
||||
|
||||
The `inet fips` table only matches packets arriving on `fips0`.
|
||||
Anything else returns from the chain on the first rule. Specifically:
|
||||
|
||||
- **Docker / containerd** install nftables rules in the `ip` and `ip6`
|
||||
families and operate on `docker0`, `br-*`, and `veth*`
|
||||
interfaces. They do not touch `fips0`. The two tables coexist
|
||||
without interference.
|
||||
- **Tor** runs in user space and does not install firewall rules. The
|
||||
baseline is independent of Tor's onion-service and SOCKS listeners.
|
||||
- **OPNsense** is an upstream perimeter device. The baseline runs on
|
||||
the local host and applies only to traffic that has already reached
|
||||
the host's `fips0` interface. They do not interact.
|
||||
- **The host's main `/etc/nftables.conf`** typically defines a
|
||||
separate `inet filter` table. nftables allows multiple tables in
|
||||
the same family to coexist; both run in parallel at hook
|
||||
`input`/priority 0 and the `iifname != "fips0" return` rule keeps
|
||||
the `inet fips` table from interfering with anything outside the
|
||||
mesh interface.
|
||||
- **`inet fips_gateway`**, when `fips-gateway` is running, manages
|
||||
DNAT/SNAT on the LAN-facing interface to translate virtual IPs to
|
||||
mesh addresses. It is a separate concern owned by the gateway
|
||||
binary and is unrelated to this baseline. See the section below.
|
||||
|
||||
## Coexistence with `inet fips_gateway`
|
||||
|
||||
When `fips-gateway` is running, it manages a separate nftables
|
||||
table, `inet fips_gateway`, containing the DNAT and masquerade rules
|
||||
that translate between the gateway's virtual-IP pool and mesh
|
||||
addresses on the LAN-facing interface. That table is created and
|
||||
torn down by the gateway binary at runtime and is not an operator
|
||||
artifact in the same sense as `inet fips`.
|
||||
|
||||
The two tables do not interfere:
|
||||
|
||||
- `inet fips` filters inbound on `fips0`.
|
||||
- `inet fips_gateway` performs NAT on the LAN interface.
|
||||
|
||||
They operate on different interfaces and at different hook points
|
||||
(`input` filter vs. `prerouting`/`postrouting` NAT). Both can be
|
||||
loaded simultaneously on a gateway host, and that is the intended
|
||||
deployment shape. See [fips-gateway.md](fips-gateway.md) for the
|
||||
gateway table's structure.
|
||||
|
||||
## What the Baseline Does Not Cover
|
||||
|
||||
The baseline is one half of a defense-in-depth posture. It is
|
||||
explicitly not:
|
||||
|
||||
- **Outbound filtering.** Anything the mesh host originates on
|
||||
`fips0` is unrestricted. If you need to constrain what the host
|
||||
can send to the mesh, add rules to a separate chain hooked at
|
||||
`output` — out of scope for the baseline.
|
||||
- **Application-layer authorization.** The baseline decides whether
|
||||
a packet reaches a service. It does not decide whether the
|
||||
originating mesh node's npub is allowed to use that service. That
|
||||
is the application's responsibility (e.g., an `authorized_keys`
|
||||
file for SSH, an ACL in the application's configuration).
|
||||
- **ACL on the mesh handshake.** The FMP Noise IK handshake
|
||||
authenticates the peer's npub and, on both inbound and outbound
|
||||
paths, consults the peer ACL (`peers.allow` / `peers.deny`) before
|
||||
promoting the connection. The ACL evaluates in TCP-Wrappers order:
|
||||
an `allow` match permits, otherwise a `deny` match rejects,
|
||||
otherwise the connection is permitted. A strict allowlist posture
|
||||
therefore requires an explicit `ALL` entry in `peers.deny`; a
|
||||
populated `peers.allow` alone does not turn the ACL into a strict
|
||||
allowlist. Mesh-level ACLs are a separate concern from the inbound
|
||||
packet filter described here; see the peer ACL section in
|
||||
[../reference/security.md](../reference/security.md).
|
||||
- **Compromised peers.** A peer whose key has been stolen or whose
|
||||
host has been taken over is, by mesh-level identity, still that
|
||||
peer. Source-address filtering in drop-ins operates on the source
|
||||
mesh address of inbound traffic regardless of whether that source
|
||||
is a direct peer or a multi-hop mesh node, and so can limit damage
|
||||
from a known-compromised mesh address; but the baseline cannot
|
||||
revoke trust on its own.
|
||||
|
||||
Treat the baseline as removing the "wide-open by default" failure
|
||||
mode. Higher-layer authorization decisions are the operator's and
|
||||
the application's, the same as on any other shared network.
|
||||
|
||||
## Future Work
|
||||
|
||||
The current baseline is Linux-only. Parallel work for other targets:
|
||||
|
||||
- **macOS PF baseline.** macOS uses Packet Filter (PF), inherited
|
||||
from OpenBSD. PF maps cleanly onto the same conceptual model as
|
||||
nftables: stateful inspection (`keep state` ≈ `ct state
|
||||
established,related`), default policy, anchor-based modular rule
|
||||
loading. A `packaging/macos/fips.pf` will land alongside the
|
||||
Linux baseline with the same posture: documented asset, no
|
||||
auto-load, operator opts in via launchd. The macOS interface name
|
||||
is `utunN` rather than `fips0`, so the rule template needs runtime
|
||||
substitution or a PF interface group assigned at TUN bring-up;
|
||||
this is being worked through with the macOS port.
|
||||
- **OpenWrt fw4 path.** OpenWrt's fw4 already drives nftables under
|
||||
the hood, but rules go into `/etc/nftables.d/` includes or UCI
|
||||
entries in `/etc/config/firewall`, not a free-standing
|
||||
`fips.nft`. The ipk will ship a layout-compatible variant or
|
||||
document the operator setup separately, decided when the OpenWrt
|
||||
packaging is updated.
|
||||
- **Cross-OS gateway abstraction.** `fips-gateway` is currently
|
||||
Linux-only because `src/gateway/nat.rs` uses the `rustables`
|
||||
netlink API directly. macOS gateway support requires a PF-backed
|
||||
equivalent behind a shared backend trait. This is a larger lift
|
||||
than the static baseline and is tracked separately under the same
|
||||
cross-OS thread.
|
||||
|
||||
When those land, this document will grow per-OS sections describing
|
||||
each baseline's load mechanism and extension points. The threat
|
||||
model and the operator-extension principle are the same on every OS;
|
||||
only the filter syntax and the activation gesture differ.
|
||||
|
||||
## See also
|
||||
|
||||
- [enable-mesh-firewall.md](../how-to/enable-mesh-firewall.md) — operator
|
||||
activation steps, drop-in recipes, drop visibility and debugging
|
||||
- [../reference/security.md](../reference/security.md) — consolidated
|
||||
security reference (cryptographic primitives, peer ACL format,
|
||||
filesystem permissions, default network exposures)
|
||||
- [fips-gateway.md](fips-gateway.md) — `fips-gateway` service and the
|
||||
separate `inet fips_gateway` table
|
||||
@@ -1,10 +1,10 @@
|
||||
# FIPS Session Protocol (FSP)
|
||||
|
||||
The FIPS Session Protocol is the top protocol layer in the FIPS stack. It sits
|
||||
above the FIPS Mesh Protocol (FMP) and below applications (native FIPS API or
|
||||
IPv6 adapter). FSP provides end-to-end authenticated, encrypted datagram
|
||||
delivery between any two FIPS nodes, regardless of how many intermediate hops
|
||||
separate them.
|
||||
The FIPS Session Protocol is the topmost layer of the FIPS protocol stack.
|
||||
It sits above the FIPS Mesh Protocol (FMP) and below applications (native
|
||||
FIPS API or IPv6 adapter). FSP provides end-to-end authenticated, encrypted
|
||||
datagram delivery between any two FIPS nodes, regardless of how many
|
||||
intermediate hops separate them.
|
||||
|
||||
## Role
|
||||
|
||||
@@ -99,9 +99,9 @@ FMP signals routing failures asynchronously:
|
||||
a standalone CoordsWarmup (rate-limited), re-discovering the destination's
|
||||
current coordinates, and resetting the warmup counter.
|
||||
- **MtuExceeded**: A transit node cannot forward a SessionDatagram because
|
||||
the packet exceeds the next-hop link MTU. FSP uses the reported bottleneck
|
||||
MTU to adjust its session-layer path MTU estimate. MtuExceeded is the
|
||||
reactive complement to the proactive `path_mtu` field in SessionDatagram.
|
||||
the packet exceeds the next-hop link MTU; FSP adjusts its
|
||||
session-layer path MTU estimate from the reported bottleneck. See
|
||||
[fips-mtu.md](fips-mtu.md) for the full forward/reverse MTU model.
|
||||
|
||||
All three signals are generated by transit nodes (not the destination) and
|
||||
travel back to the source inside a new SessionDatagram. They are plaintext
|
||||
@@ -123,9 +123,10 @@ a destination with no existing session.
|
||||
FSP uses Noise XK for session key agreement (Noise Protocol Framework;
|
||||
Perrin 2018). The initiator knows the destination's npub (required for
|
||||
XK's pre-message `s` token); the responder learns the initiator's
|
||||
identity from msg3 (not msg1, unlike IK at the link layer). This provides stronger initiator identity hiding
|
||||
— the initiator's static key is encrypted under the established shared
|
||||
secret rather than under only the responder's static key.
|
||||
identity from msg3 (not msg1, unlike IK at the link layer). This
|
||||
provides stronger initiator identity hiding — the initiator's static
|
||||
key is encrypted under the established shared secret rather than under
|
||||
only the responder's static key.
|
||||
|
||||
The handshake is a three-message flow carried in SessionSetup, SessionAck,
|
||||
and SessionMsg3:
|
||||
@@ -244,17 +245,9 @@ operations) rather than under only the responder's static key.
|
||||
|
||||
### Cryptographic Primitives
|
||||
|
||||
| Component | Choice | Notes |
|
||||
| --------- | ------ | ----- |
|
||||
| Curve | secp256k1 | Nostr-native |
|
||||
| DH | ECDH on secp256k1 | Standard EC Diffie-Hellman |
|
||||
| Cipher | ChaCha20-Poly1305 | AEAD, same as NIP-44 |
|
||||
| Hash | SHA-256 | Nostr-native |
|
||||
| Key derivation | HKDF-SHA256 | Standard Noise KDF |
|
||||
|
||||
These choices prioritize compatibility with the Nostr cryptographic stack —
|
||||
secp256k1 + ChaCha20-Poly1305 + SHA-256 aligns with the NIP-44 encrypted
|
||||
messaging standard.
|
||||
FSP uses ChaCha20-Poly1305 with secp256k1 ECDH; see
|
||||
[../reference/security.md](../reference/security.md) for the full
|
||||
primitive table shared with the link layer.
|
||||
|
||||
### secp256k1 Parity Normalization
|
||||
|
||||
@@ -391,22 +384,10 @@ signal generation (100ms per destination).
|
||||
|
||||
## Identity Cache
|
||||
|
||||
The identity cache maps FIPS address prefix (15 bytes, the `fd00::/8` IPv6
|
||||
address minus the `fd` prefix) to `(NodeAddr, PublicKey)`. This cache is
|
||||
needed only when using the IPv6 adapter — the native FIPS API provides the
|
||||
public key directly.
|
||||
|
||||
The mapping is deterministic (derived from the public key via SHA-256) and
|
||||
never becomes stale. The cache uses LRU-only eviction bounded by a
|
||||
configurable size (default 10K entries). There is no TTL — entries are evicted
|
||||
only when the cache is full and space is needed for a new entry.
|
||||
|
||||
Cache population mechanisms:
|
||||
|
||||
- **DNS lookup**: The primary path. Resolving `npub1xxx...xxx.fips` derives
|
||||
the IPv6 address and populates the identity cache.
|
||||
- **Inbound traffic**: Authenticated sessions from other nodes populate the
|
||||
cache with their identity information.
|
||||
The IPv6 adapter requires an identity cache to map `fd00::/8` addresses
|
||||
back to `(NodeAddr, PublicKey)` for routing; see
|
||||
[fips-ipv6-adapter.md](fips-ipv6-adapter.md#identity-cache) for the
|
||||
cache rationale, eviction policy, and population mechanics.
|
||||
|
||||
## Coordinate Cache
|
||||
|
||||
@@ -451,84 +432,30 @@ node caches (still within their 300s TTL) are re-warmed.
|
||||
|
||||
## Session-Layer MMP
|
||||
|
||||
Each established session runs its own Metrics Measurement Protocol instance,
|
||||
providing end-to-end quality metrics independent of the number of hops.
|
||||
FSP runs an MMP instance per established session for end-to-end metrics
|
||||
independent of hop count. Reports are encrypted and forwarded through
|
||||
every transit link, so bandwidth cost is proportional to path length;
|
||||
the session-layer report intervals are correspondingly higher than the
|
||||
link-layer intervals (clamped to `[500ms, 10s]` vs. `[1s, 5s]`).
|
||||
|
||||
### Relationship to Link-Layer MMP
|
||||
The session-layer instance shares its algorithms (SRTT, jitter, loss,
|
||||
ETX) and report wire format with link-layer MMP. The differences —
|
||||
configuration namespace (`node.session_mmp.*`), routing scope,
|
||||
send-failure backoff, idle-timeout interaction, and the
|
||||
PathMtuNotification mechanism — are documented in the unified MMP
|
||||
treatment at [fips-mmp.md](fips-mmp.md). For the end-to-end path-MTU
|
||||
echo specifically, see [fips-mtu.md](fips-mtu.md). Reports and
|
||||
PathMtuNotification do **not** reset the session idle timer, so a
|
||||
session carrying only MMP traffic still tears down at the configured
|
||||
idle threshold.
|
||||
|
||||
Session-layer MMP uses the same report wire format (SenderReport 0x11,
|
||||
ReceiverReport 0x12) and identical algorithms as link-layer MMP, but with
|
||||
two key differences:
|
||||
### MtuExceeded Handling
|
||||
|
||||
1. **End-to-end routing**: Session reports are encrypted and forwarded through
|
||||
every transit link. A 3-hop session generates report traffic on all 3 links,
|
||||
making bandwidth cost proportional to path length.
|
||||
2. **Independent configuration**: The `node.session_mmp.*` parameters are
|
||||
separate from `node.mmp.*`, allowing operators to run a lighter mode for
|
||||
sessions (e.g., Lightweight) while keeping Full mode on links.
|
||||
|
||||
### Metrics
|
||||
|
||||
The same metrics as link-layer MMP: SRTT, loss rate, jitter, goodput, OWD
|
||||
trend, ETX, and dual EWMA trends. Session MMP additionally tracks observed
|
||||
path MTU.
|
||||
|
||||
### Report Intervals
|
||||
|
||||
Session-layer report intervals are higher than link-layer to account for
|
||||
bandwidth cost: clamped to [500ms, 10s] with a cold-start interval of 1s
|
||||
(vs. link-layer [100ms, 2s] with 500ms cold-start).
|
||||
|
||||
### Path MTU Tracking
|
||||
|
||||
PathMtuNotification (message type 0x13) provides end-to-end path MTU
|
||||
feedback, adapting RFC 1191 Path MTU Discovery for overlay networks — the
|
||||
transit-node `min()` propagation replaces ICMP Packet Too Big:
|
||||
|
||||
1. The source sets `path_mtu` in each SessionDatagram envelope to its
|
||||
outbound link MTU.
|
||||
2. Each transit node applies `min(current, transport.link_mtu(addr))` before
|
||||
forwarding.
|
||||
3. The destination receives the forward-path minimum and sends a
|
||||
PathMtuNotification (2-byte body: u16 LE path_mtu) back to the source.
|
||||
4. The source applies the notification with hysteresis:
|
||||
- **Decrease**: immediate (take lower value).
|
||||
- **Increase**: requires 3 consecutive higher-value notifications spanning
|
||||
at least 2 × notification interval.
|
||||
5. Notifications are sent on first measurement, on any decrease, and
|
||||
periodically at `max(10s, 5 × SRTT)`.
|
||||
|
||||
### Send Failure Backoff
|
||||
|
||||
When a session MMP report cannot be delivered (destination unreachable, no
|
||||
route), the sender applies exponential backoff to the probe interval (a
|
||||
standard distributed systems pattern for transient failure handling):
|
||||
|
||||
- Each consecutive failure doubles the interval: 2x, 4x, 8x, 16x, 32x
|
||||
- Backoff caps at 32x the base interval (5 consecutive failures)
|
||||
- A successful send resets to the normal SRTT-based interval
|
||||
- Debug logging is suppressed after 3 consecutive failures; a summary is
|
||||
logged when the destination becomes reachable again
|
||||
|
||||
This prevents wasted CPU and log noise when a session's remote endpoint has
|
||||
departed the network but the local session has not yet timed out.
|
||||
|
||||
### Idle Timeout Interaction
|
||||
|
||||
MMP reports (SenderReport, ReceiverReport) and PathMtuNotification do **not**
|
||||
reset the session idle timer. Only application data (DataPacket, type 0x10)
|
||||
resets `last_activity`. This ensures sessions with no application traffic
|
||||
tear down after `node.session.idle_timeout_secs` (default 90s), while MMP
|
||||
continues providing measurement data up to the teardown moment.
|
||||
|
||||
### Operator Logging
|
||||
|
||||
Session metrics are logged at info level (configurable via
|
||||
`node.session_mmp.log_interval_secs`, default 30s):
|
||||
|
||||
```text
|
||||
MMP session metrics session=npub1tdwa...84le rtt=4.3ms loss=0.6% jitter=0.2ms goodput=71.3MB/s mtu=1472 tx_pkts=1234 rx_pkts=5678
|
||||
```
|
||||
When FMP signals MtuExceeded (a transit node could not forward a
|
||||
SessionDatagram because it exceeded the next-hop link MTU), FSP uses
|
||||
the reported bottleneck MTU to adjust its session-layer path MTU
|
||||
estimate immediately. See [fips-mtu.md](fips-mtu.md) for the full
|
||||
reactive PMTUD mechanism.
|
||||
|
||||
## Implementation Status
|
||||
|
||||
@@ -559,36 +486,27 @@ MMP session metrics session=npub1tdwa...84le rtt=4.3ms loss=0.6% jitter=0.2ms go
|
||||
|
||||
### FIPS Internal Documentation
|
||||
|
||||
- [fips-intro.md](fips-intro.md) — Protocol overview and architecture
|
||||
- [fips-concepts.md](fips-concepts.md) — Protocol overview
|
||||
- [fips-architecture.md](fips-architecture.md) — Layer architecture and
|
||||
identity model
|
||||
- [fips-mesh-layer.md](fips-mesh-layer.md) — FMP specification (below FSP)
|
||||
- [fips-ipv6-adapter.md](fips-ipv6-adapter.md) — IPv6 adaptation layer (above FSP)
|
||||
- [fips-mesh-operation.md](fips-mesh-operation.md) — Routing, discovery, and
|
||||
error recovery
|
||||
- [fips-wire-formats.md](fips-wire-formats.md) — Wire format reference for all
|
||||
session message types
|
||||
- [fips-ipv6-adapter.md](fips-ipv6-adapter.md) — IPv6 adaptation layer
|
||||
(above FSP)
|
||||
- [fips-mesh-operation.md](fips-mesh-operation.md) — Routing, discovery,
|
||||
and error recovery
|
||||
- [fips-mmp.md](fips-mmp.md) — Metrics Measurement Protocol (link + session)
|
||||
- [fips-mtu.md](fips-mtu.md) — Path MTU model (PathMtuNotification,
|
||||
MtuExceeded, hysteresis)
|
||||
- [fips-prior-work.md](fips-prior-work.md) — Noise XK, WireGuard,
|
||||
DTLS replay window, IKEv2 simultaneous initiation, hybrid coordinate
|
||||
warmup citations
|
||||
- [../reference/wire-formats.md](../reference/wire-formats.md) — Wire
|
||||
format reference for all session message types
|
||||
- [../reference/security.md](../reference/security.md) — Cryptographic
|
||||
primitives and rekey defaults
|
||||
|
||||
### External References
|
||||
|
||||
- Perrin, T. ["The Noise Protocol Framework"](https://noiseprotocol.org/noise.html).
|
||||
Revision 34, 2018. *Framework for building crypto protocols using Diffie-Hellman
|
||||
key agreement and AEAD ciphers. FSP uses the XK handshake pattern.*
|
||||
|
||||
- Donenfeld, J.A. ["WireGuard: Next Generation Kernel Network Tunnel"](https://www.wireguard.com/papers/wireguard.pdf).
|
||||
NDSS 2017. *Transport-independent cryptographic sessions bound to identity keys
|
||||
rather than network addresses; AEAD-only authentication model.*
|
||||
|
||||
- Rescorla, E., Modadugu, N. [RFC 6347](https://datatracker.ietf.org/doc/html/rfc6347):
|
||||
"Datagram Transport Layer Security Version 1.2". 2012. *Explicit sequence numbers
|
||||
with sliding bitmap window for replay protection over unreliable transports.*
|
||||
|
||||
- Kaufman, C., Hoffman, P., Nir, Y., Eronen, P., Kivinen, T.
|
||||
[RFC 7296](https://datatracker.ietf.org/doc/html/rfc7296):
|
||||
"Internet Key Exchange Protocol Version 2 (IKEv2)". 2014. *Simultaneous
|
||||
initiation resolution (§2.8) and INITIAL_CONTACT peer restart detection (§2.4).*
|
||||
|
||||
- Mogul, J., Deering, S. [RFC 1191](https://datatracker.ietf.org/doc/html/rfc1191):
|
||||
"Path MTU Discovery". 1990. *End-to-end path MTU discovery; FSP adapts this for
|
||||
overlay networks using transit-node min() propagation.*
|
||||
|
||||
- [Yggdrasil Network](https://yggdrasil-network.github.io/). *Coordinate-based
|
||||
overlay routing with session traffic used to warm transit node coordinate caches.*
|
||||
|
||||
@@ -71,8 +71,11 @@ quality.
|
||||
2. Compute **effective depth** for each candidate peer:
|
||||
`effective_depth = peer.depth + link_cost`, where
|
||||
`link_cost = etx * (1.0 + srtt_ms / 100.0)` using locally measured MMP
|
||||
metrics. When MMP metrics have not yet converged, `link_cost` defaults to
|
||||
1.0, preserving pure depth-based behavior as a graceful fallback.
|
||||
metrics. During cold start (no peer has MMP data yet), candidates without
|
||||
measurements default to `link_cost = 1.0`, preserving pure depth-based
|
||||
behavior. Once any peer has MMP data, unmeasured candidates are excluded
|
||||
so that a freshly connected peer cannot win parent selection on the
|
||||
default cost alone.
|
||||
3. Apply **hysteresis**: switch parents only when the best candidate's
|
||||
effective depth is significantly better than the current parent's:
|
||||
`best_eff_depth < current_eff_depth * (1.0 - parent_hysteresis)`
|
||||
@@ -109,6 +112,9 @@ immediate parent reselection:
|
||||
(ETX and SRTT). No cumulative path costs are propagated, avoiding
|
||||
the trust problems inherent in self-reported cost metrics in a
|
||||
permissionless network.
|
||||
- **Loop rejection**: Candidates whose advertised ancestry already contains
|
||||
the local node are skipped, preventing two nodes from selecting each
|
||||
other as parent and entering an alternating coordinate loop.
|
||||
|
||||
### After Parent Change
|
||||
|
||||
@@ -180,9 +186,14 @@ A node re-announces (propagates) only when its own state changes:
|
||||
|
||||
- **Root changed**: Always propagate — this is a significant topology event
|
||||
- **Depth changed**: Always propagate — affects routing distance calculations
|
||||
- **Mid-chain ancestor swap**: A reroute that replaces an interior ancestor
|
||||
without changing the root or the path length still alters the node's
|
||||
coordinate path, so it propagates. Without this, downstream peers would
|
||||
route into a phantom intermediate that no longer appears on the parent's
|
||||
tree.
|
||||
- **Sequence-only refresh**: Does NOT propagate beyond depth 1 — peers that
|
||||
receive a sequence-only update do not re-announce, because their own root
|
||||
and depth have not changed
|
||||
receive a sequence-only update do not re-announce, because their own root,
|
||||
depth, and address path have not changed
|
||||
|
||||
This means TreeAnnounce cascades through the tree proportional to depth,
|
||||
not network size. A change at depth D affects at most D nodes along the
|
||||
@@ -303,12 +314,15 @@ Example: In a 1000-node network with depth 10 and 5 peers, a node stores
|
||||
| Rate limiting (500ms per peer) | **Implemented** |
|
||||
| Coord cache flush on parent change | **Implemented** |
|
||||
| Flap dampening (extended hold-down on rapid switches) | **Implemented** |
|
||||
| Loop rejection (ancestry self-check in `evaluate_parent`) | **Implemented** |
|
||||
| Mid-chain ancestor swap propagation | **Implemented** |
|
||||
| Per-ancestry-entry signatures | Future direction |
|
||||
|
||||
## References
|
||||
|
||||
- [fips-mesh-operation.md](fips-mesh-operation.md) — How the spanning tree
|
||||
fits into mesh routing
|
||||
- [fips-wire-formats.md](fips-wire-formats.md) — TreeAnnounce wire format
|
||||
- [../reference/wire-formats.md](../reference/wire-formats.md) —
|
||||
TreeAnnounce wire format
|
||||
- [spanning-tree-dynamics.md](spanning-tree-dynamics.md) — Convergence
|
||||
scenario walkthroughs
|
||||
|
||||
@@ -1,5 +1,7 @@
|
||||
# FIPS Transport Layer
|
||||
|
||||
<!-- markdownlint-disable MD024 -->
|
||||
|
||||
The transport layer is the bottom of the FIPS protocol stack. It delivers
|
||||
datagrams between transport-specific endpoints over arbitrary physical or
|
||||
logical media. Everything above — peer authentication, routing, encryption,
|
||||
@@ -50,7 +52,7 @@ determine how much payload can fit in a single packet after link-layer
|
||||
encryption overhead.
|
||||
|
||||
MTU is fundamentally a per-link property. A transport with a fixed MTU
|
||||
(Ethernet: 1500, UDP configured at 1472) returns the same value for every
|
||||
(Ethernet effective 1499, UDP default 1280) returns the same value for every
|
||||
link — this is the degenerate case. Transports that negotiate MTU
|
||||
per-connection (e.g., BLE ATT_MTU) report the negotiated value for each
|
||||
link individually.
|
||||
@@ -68,7 +70,7 @@ forwarding and LookupResponse transit annotation.
|
||||
### Connection Lifecycle
|
||||
|
||||
For connection-oriented transports, manage the underlying connection: TCP
|
||||
handshake, Tor circuit establishment, Bluetooth pairing. FMP cannot begin
|
||||
handshake, Tor circuit establishment, BLE pairing. FMP cannot begin
|
||||
the Noise IK link handshake until the transport-layer connection is
|
||||
established.
|
||||
|
||||
@@ -117,8 +119,8 @@ for internet connectivity:
|
||||
| --------- | ---------- | --- | ----------- | ----- |
|
||||
| UDP/IP | host:port | 1280–1472 | Unreliable | Primary internet transport |
|
||||
| TCP/IP | host:port | Stream | Reliable | Requires length-prefix framing |
|
||||
| WebSocket | URL | Stream | Reliable | Browser-compatible |
|
||||
| Tor | .onion | Stream | Reliable | High latency, strong anonymity |
|
||||
| Nym | host:port | Stream | Reliable | Mixnet, outbound-only, strong anonymity |
|
||||
|
||||
**Shared medium transports** operate over broadcast- or multicast-capable
|
||||
media:
|
||||
@@ -127,7 +129,6 @@ media:
|
||||
| --------- | ---------- | --- | ----------- | ----- |
|
||||
| Ethernet | MAC | 1500 | Unreliable | Raw AF_PACKET frames |
|
||||
| WiFi | MAC | 1500 | Unreliable | Infrastructure mode = Ethernet |
|
||||
| Bluetooth | BD_ADDR | 672–64K | Reliable | L2CAP |
|
||||
| BLE | BD_ADDR | 23–517 | Reliable | Negotiated ATT_MTU |
|
||||
| Radio | Device addr | 51–222 | Unreliable | Low bandwidth, long range |
|
||||
|
||||
@@ -157,9 +158,9 @@ require connection setup before FMP can begin the Noise IK link handshake,
|
||||
adding startup latency.
|
||||
|
||||
**Stream vs. datagram**: Datagram transports have natural packet boundaries.
|
||||
Stream transports (TCP, WebSocket, Tor) require framing to delineate FIPS
|
||||
packets within the byte stream. The FMP common prefix includes a payload
|
||||
length field that provides this framing directly, replacing the need for a
|
||||
Stream transports (TCP, Tor) require framing to delineate FIPS packets
|
||||
within the byte stream. The FMP common prefix includes a payload length
|
||||
field that provides this framing directly, replacing the need for a
|
||||
separate length-prefix layer.
|
||||
|
||||
**Addressing opacity**: Transport addresses are opaque byte vectors. FMP
|
||||
@@ -189,9 +190,8 @@ proceed.
|
||||
| Transport | Connection Setup |
|
||||
| --------- | ---------------- |
|
||||
| TCP/IP | TCP three-way handshake |
|
||||
| WebSocket | HTTP upgrade + TCP |
|
||||
| Tor | Circuit establishment (500ms–5s) |
|
||||
| Bluetooth | L2CAP connection |
|
||||
| Tor | Circuit establishment (typically 10–60s, default timeout 120s) |
|
||||
| Nym | SOCKS5 connect through mixnet (minutes possible, default timeout 300s) |
|
||||
| BLE | L2CAP CoC or GATT connection |
|
||||
| Serial | Physical connection (static) |
|
||||
|
||||
@@ -203,9 +203,9 @@ Connected → Disconnected. Failure can occur during connection setup, adding
|
||||
error handling paths that connectionless transports don't have.
|
||||
|
||||
**Startup latency**: Connection-oriented transports add delay before a peer
|
||||
becomes usable. This ranges from milliseconds (TCP) to seconds (Tor
|
||||
circuit). Peer timeout configuration must account for transport-specific
|
||||
setup times.
|
||||
becomes usable. This ranges from milliseconds (TCP) to tens of seconds
|
||||
(Tor circuit). Peer timeout configuration must account for
|
||||
transport-specific setup times.
|
||||
|
||||
**Framing**: Stream transports must delimit FIPS packets within the byte
|
||||
stream. The FMP common prefix includes a payload length field that provides
|
||||
@@ -229,45 +229,29 @@ NAT devices and firewalls, limiting deployment to networks without NAT.
|
||||
|
||||
### Socket Buffer Sizing
|
||||
|
||||
The default Linux UDP receive buffer (`net.core.rmem_default`, typically
|
||||
212 KB) is insufficient for high-throughput forwarding. At ~85 MB/s, a 212 KB
|
||||
buffer fills in ~2.5 ms; any stall in the async receive loop (decryption,
|
||||
routing, forwarding overhead) causes the kernel to silently drop incoming
|
||||
datagrams.
|
||||
The default Linux UDP receive buffer (`net.core.rmem_default`,
|
||||
typically 212 KB) is insufficient for high-throughput forwarding. At
|
||||
~85 MB/s, a 212 KB buffer fills in ~2.5 ms; any stall in the async
|
||||
receive loop (decryption, routing, forwarding overhead) causes the
|
||||
kernel to silently drop incoming datagrams.
|
||||
|
||||
FIPS uses `socket2::Socket` wrapped in `tokio::io::unix::AsyncFd` for the
|
||||
UDP receive path. This replaces `tokio::UdpSocket` and enables direct
|
||||
`libc::recvmsg()` calls with ancillary data parsing — specifically the
|
||||
`SO_RXQ_OVFL` socket option, which delivers a cumulative kernel receive
|
||||
buffer drop counter on every received packet. The drop counter feeds into
|
||||
the ECN congestion detection system (see
|
||||
[fips-mesh-layer.md](fips-mesh-layer.md#ecn-congestion-signaling)).
|
||||
FIPS uses `socket2::Socket` wrapped in `tokio::io::unix::AsyncFd` for
|
||||
the UDP receive path. This replaces `tokio::UdpSocket` and enables
|
||||
direct `libc::recvmsg()` calls with ancillary data parsing —
|
||||
specifically the `SO_RXQ_OVFL` socket option, which delivers a
|
||||
cumulative kernel receive buffer drop counter on every received
|
||||
packet. The drop counter feeds into the ECN congestion detection
|
||||
system (see [fips-mmp.md](fips-mmp.md#ecn-congestion-signaling)).
|
||||
|
||||
Socket buffers are configured at bind time via `socket2`:
|
||||
|
||||
| Parameter | Default | Description |
|
||||
| ---------------- | ------- | ------------------------------------ |
|
||||
| `recv_buf_size` | 2 MB | `SO_RCVBUF` — kernel receive buffer |
|
||||
| `send_buf_size` | 2 MB | `SO_SNDBUF` — kernel send buffer |
|
||||
|
||||
Linux internally doubles the requested value (to account for kernel
|
||||
bookkeeping overhead), so requesting 2 MB yields 4 MB actual buffer space.
|
||||
The kernel silently clamps to `net.core.rmem_max` if the request exceeds it.
|
||||
|
||||
**Host requirement**: `net.core.rmem_max` and `net.core.wmem_max` must be
|
||||
set to at least the requested buffer size on the host. For Docker containers,
|
||||
this must be configured on the Docker host (containers share the host kernel).
|
||||
Verify with:
|
||||
|
||||
```text
|
||||
sysctl net.core.rmem_max net.core.wmem_max
|
||||
```
|
||||
|
||||
Actual buffer sizes are logged at startup:
|
||||
|
||||
```text
|
||||
UDP transport started local_addr=0.0.0.0:2121 recv_buf=4194304 send_buf=4194304
|
||||
```
|
||||
Socket buffers (`recv_buf_size`, `send_buf_size`) are configured at
|
||||
bind time via `socket2`. Linux internally doubles the requested value
|
||||
(to account for kernel bookkeeping overhead) and silently clamps to
|
||||
`net.core.rmem_max` / `net.core.wmem_max` if the request exceeds the
|
||||
host kernel limits. The full UDP transport configuration is in
|
||||
[../reference/configuration.md](../reference/configuration.md). The
|
||||
host-side sysctl requirements and how to set them persistently live
|
||||
in
|
||||
[../how-to/tune-udp-buffers.md](../how-to/tune-udp-buffers.md).
|
||||
|
||||
## Ethernet: The Local Network Transport
|
||||
|
||||
@@ -278,7 +262,7 @@ UDP (1500 vs 1472 MTU).
|
||||
- **No IP dependency**: Operates below the IP layer. Nodes on the same
|
||||
Ethernet segment can communicate without IP addresses or routing
|
||||
infrastructure
|
||||
- **Broadcast discovery**: Nodes discover each other via periodic beacon
|
||||
- **Broadcast neighbor detection**: Nodes discover each other via periodic beacon
|
||||
broadcasts on the shared medium, with no static peer configuration required
|
||||
- **Higher MTU**: Standard Ethernet frames carry 1500 bytes of payload,
|
||||
yielding an effective FIPS MTU of 1499 after the frame type prefix
|
||||
@@ -309,7 +293,7 @@ socket.
|
||||
| Addressing | 6-byte MAC address |
|
||||
| Platform | Linux only (`CAP_NET_RAW` required) |
|
||||
|
||||
### Beacon Discovery
|
||||
### Neighbor Beacons
|
||||
|
||||
Ethernet nodes discover peers via broadcast beacons sent to
|
||||
ff:ff:ff:ff:ff:ff. Each beacon is a 34-byte frame containing the sender's
|
||||
@@ -317,24 +301,22 @@ x-only public key. Receiving nodes extract the MAC source address from the
|
||||
frame and the public key from the payload, then report the discovered peer
|
||||
to FMP.
|
||||
|
||||
Four configuration flags control discovery behavior:
|
||||
Four configuration flags control neighbor behavior — `listen`
|
||||
(listen for beacons), `announce` (broadcast beacons), `auto_connect`
|
||||
(initiate handshakes to discovered peers), and `accept_connections`
|
||||
(accept inbound handshakes). The flag table and per-flag defaults
|
||||
live in [../reference/configuration.md](../reference/configuration.md)
|
||||
under `transports.ethernet.*`.
|
||||
|
||||
| Flag | Default | Description |
|
||||
| ---- | ------- | ----------- |
|
||||
| `discovery` | true | Listen for beacons from other nodes |
|
||||
| `announce` | false | Broadcast beacons periodically |
|
||||
| `auto_connect` | false | Initiate handshakes to discovered peers |
|
||||
| `accept_connections` | false | Accept inbound handshake attempts |
|
||||
|
||||
A typical discoverable node sets `announce: true`, `auto_connect: true`, and
|
||||
`accept_connections: true`. A passive listener uses just `discovery: true` to
|
||||
observe the network without announcing itself.
|
||||
A typical discoverable node sets `announce`, `auto_connect`, and
|
||||
`accept_connections` all true. A passive listener uses just
|
||||
`listen: true` to observe the network without announcing itself.
|
||||
|
||||
### WiFi Compatibility
|
||||
|
||||
WiFi interfaces in infrastructure (managed) mode work transparently for
|
||||
unicast — the mac80211 subsystem handles frame translation between 802.11
|
||||
and 802.3. Broadcast beacon discovery is unreliable in managed mode because
|
||||
and 802.3. Broadcast neighbor detection is unreliable in managed mode because
|
||||
access points commonly isolate clients from each other's broadcast traffic.
|
||||
|
||||
Startup logging:
|
||||
@@ -343,10 +325,12 @@ Startup logging:
|
||||
Ethernet transport started name=eth0 interface=eth0 mac=aa:bb:cc:dd:ee:ff mtu=1499 if_mtu=1500
|
||||
```
|
||||
|
||||
## TCP/IP: Firewall Traversal Transport
|
||||
## TCP/IP: Transport for UDP-Filtered Networks
|
||||
|
||||
For networks where UDP is blocked but TCP port 443 is open, the TCP
|
||||
transport provides an alternative path.
|
||||
For peers whose networks filter outbound UDP, the TCP transport
|
||||
provides an alternative datagram path between public endpoints. TCP
|
||||
is not a NAT-traversal mechanism — there is no `tcp:nat` analogue to
|
||||
the UDP hole-punch flow.
|
||||
|
||||
FIPS protocols (FMP, FSP, MMP) are all unreliable datagrams. Running them
|
||||
over TCP introduces head-of-line blocking, which adds latency jitter. MMP
|
||||
@@ -425,21 +409,12 @@ removes it from the pool and aborts its receive task.
|
||||
|
||||
### Configuration
|
||||
|
||||
```yaml
|
||||
transports:
|
||||
tcp:
|
||||
bind_addr: "0.0.0.0:8443" # Listen address (omit for outbound-only)
|
||||
mtu: 1400 # Default MTU
|
||||
connect_timeout_ms: 5000 # Outbound connect timeout
|
||||
nodelay: true # TCP_NODELAY (disable Nagle)
|
||||
keepalive_secs: 30 # TCP keepalive interval (0 = disabled)
|
||||
recv_buf_size: 2097152 # SO_RCVBUF (2 MB)
|
||||
send_buf_size: 2097152 # SO_SNDBUF (2 MB)
|
||||
max_inbound_connections: 256 # Resource protection limit
|
||||
```
|
||||
|
||||
If `bind_addr` is configured, the transport accepts inbound connections.
|
||||
Without it, the transport operates in outbound-only mode (no listener
|
||||
The TCP transport configuration block (`transports.tcp.*` — bind
|
||||
address, MTU, connect timeout, TCP_NODELAY, keepalive, socket buffer
|
||||
sizes, max inbound connections) is documented in
|
||||
[../reference/configuration.md](../reference/configuration.md). If
|
||||
`bind_addr` is configured, the transport accepts inbound connections;
|
||||
without it, the transport operates in outbound-only mode (no listener
|
||||
socket is created).
|
||||
|
||||
## Tor: The Anonymity Transport
|
||||
@@ -531,22 +506,13 @@ connections arrive from `127.0.0.1` (Tor daemon's local forwarding); peer
|
||||
identity is resolved during the Noise IK handshake, not from the transport
|
||||
address.
|
||||
|
||||
Configuration requires coordinating `torrc` and `fips.yaml`:
|
||||
|
||||
```text
|
||||
# torrc
|
||||
HiddenServiceDir /var/lib/tor/fips
|
||||
HiddenServicePort 8443 127.0.0.1:8444
|
||||
|
||||
# fips.yaml tor section
|
||||
mode: "directory"
|
||||
directory_service:
|
||||
hostname_file: "/var/lib/tor/fips/hostname"
|
||||
bind_addr: "127.0.0.1:8444"
|
||||
```
|
||||
|
||||
The `HiddenServicePort` external port (8443) is what peers connect to.
|
||||
The bind_addr must match the `HiddenServicePort` target address.
|
||||
Configuration requires coordinating `torrc` and `fips.yaml`. The
|
||||
operator setup — torrc directives, `fips.yaml` `tor` section,
|
||||
HiddenServiceDir permissions, and `Sandbox 1` notes — is in
|
||||
[../how-to/deploy-tor-onion.md](../how-to/deploy-tor-onion.md). In
|
||||
brief: the `HiddenServicePort` external port is what peers connect
|
||||
to, and `tor.directory_service.bind_addr` must match the
|
||||
`HiddenServicePort` target address.
|
||||
|
||||
### Session Independence
|
||||
|
||||
@@ -569,10 +535,9 @@ anonymous node's IP.
|
||||
|
||||
### Latency Characteristics
|
||||
|
||||
Tor adds 200ms–2s RTT per circuit. First-packet latency after connection
|
||||
is higher (~2.8s) due to circuit warm-up. MMP measures this elevated
|
||||
latency, and cost-based parent selection penalizes Tor links (high SRTT
|
||||
→ high link cost). ETX is 1.0 since TCP handles retransmission.
|
||||
Tor adds 200ms–2s RTT per circuit. MMP measures this elevated latency,
|
||||
and cost-based parent selection penalizes Tor links (high SRTT → high
|
||||
link cost). ETX is 1.0 since TCP handles retransmission.
|
||||
|
||||
Tor throughput is typically 1–5 Mbps — adequate for control plane and
|
||||
moderate data transfer, not for bulk transfer.
|
||||
@@ -600,37 +565,22 @@ connections (`/run/tor/control`) are preferred over TCP for security.
|
||||
|
||||
### Configuration
|
||||
|
||||
```yaml
|
||||
transports:
|
||||
tor:
|
||||
mode: "socks5" # "socks5", "control_port", or "directory"
|
||||
socks5_addr: "127.0.0.1:9050" # SOCKS5 proxy address
|
||||
connect_timeout_ms: 120000 # Connect timeout (120s for Tor circuits)
|
||||
mtu: 1400 # Default MTU
|
||||
# control_port mode: monitoring via Tor control port (no inbound)
|
||||
# control_addr: "/run/tor/control" # Unix socket (preferred) or host:port
|
||||
# control_auth: "cookie" # "cookie" or "password:<secret>"
|
||||
# cookie_path: "/var/run/tor/control.authcookie"
|
||||
# directory mode: inbound via Tor-managed HiddenServiceDir
|
||||
# directory_service:
|
||||
# hostname_file: "/var/lib/tor/fips/hostname"
|
||||
# bind_addr: "127.0.0.1:8444"
|
||||
# max_inbound_connections: 64
|
||||
```
|
||||
|
||||
Three modes are available:
|
||||
The Tor transport block (`transports.tor.*`) is documented in
|
||||
[../reference/configuration.md](../reference/configuration.md). Three
|
||||
modes are available:
|
||||
|
||||
- **`socks5`** (default): Outbound-only through a SOCKS5 proxy. No
|
||||
control port, no inbound connections.
|
||||
- **`control_port`**: Outbound via SOCKS5 plus control port connection
|
||||
for Tor daemon monitoring. No inbound connections.
|
||||
- **`directory`** (recommended for inbound): Outbound via SOCKS5 plus
|
||||
inbound via Tor-managed `HiddenServiceDir` onion service. Optionally
|
||||
connects to the control port for monitoring when `control_addr` is set.
|
||||
Enables Tor's `Sandbox 1` for maximum security.
|
||||
inbound via Tor-managed `HiddenServiceDir` onion service.
|
||||
Optionally connects to the control port for monitoring when
|
||||
`control_addr` is set. Enables Tor's `Sandbox 1` for maximum
|
||||
security.
|
||||
|
||||
The Tor transport requires an external Tor daemon. Named instances are
|
||||
supported for multiple proxy endpoints.
|
||||
The Tor transport requires an external Tor daemon. Named instances
|
||||
are supported for multiple proxy endpoints.
|
||||
|
||||
### Implementation Roadmap
|
||||
|
||||
@@ -645,21 +595,125 @@ supported for multiple proxy endpoints.
|
||||
|
||||
### Statistics
|
||||
|
||||
The transport tracks per-instance statistics:
|
||||
The Tor transport exposes per-instance counters covering successful
|
||||
send/receive, send/receive errors, connection establishment,
|
||||
SOCKS5-level errors, MTU rejections, accepted/rejected inbound
|
||||
connections, and Tor control-port errors. The full counter table
|
||||
lives in [../reference/transports.md](../reference/transports.md).
|
||||
|
||||
| Counter | Description |
|
||||
| ------- | ----------- |
|
||||
| `packets_sent` / `bytes_sent` | Successful sends |
|
||||
| `packets_recv` / `bytes_recv` | Successful receives |
|
||||
| `send_errors` / `recv_errors` | Send/receive failures |
|
||||
| `connections_established` | Successful SOCKS5 connections |
|
||||
| `connect_timeouts` | Connection timeout count |
|
||||
| `connect_refused` | Connection refused count |
|
||||
| `socks5_errors` | SOCKS5 protocol errors |
|
||||
| `mtu_exceeded` | Packets rejected for MTU violation |
|
||||
| `connections_accepted` | Accepted inbound connections via onion service |
|
||||
| `connections_rejected` | Rejected inbound connections (limit exceeded) |
|
||||
| `control_errors` | Tor control port errors |
|
||||
## Nym: The Mixnet Transport
|
||||
|
||||
The Nym transport routes FIPS traffic through the Nym mixnet, providing
|
||||
network-level anonymity via Sphinx packet routing and timing
|
||||
obfuscation. It uses the "mixnet-as-proxy" pattern: a node connects
|
||||
outbound through a local `nym-socks5-client` SOCKS5 proxy, which carries
|
||||
the traffic into the mixnet. The `nym-socks5-client` runs as a separate
|
||||
process alongside the fips daemon and must be started independently.
|
||||
|
||||
Like Tor, Nym is a privacy-oriented deployment mode chosen for the
|
||||
anonymity properties of the mixnet, not a failover for other transports.
|
||||
Like TCP and Tor, it is connection-oriented and reliable; the same
|
||||
TCP-over-TCP considerations apply, and cost-based parent selection
|
||||
naturally deprioritizes the high-latency Nym links.
|
||||
|
||||
### Architecture
|
||||
|
||||
The Nym transport is a separate `NymTransport` implementation. It reuses
|
||||
the FMP header-based stream reader (`tcp/stream.rs`) for packet framing
|
||||
on the underlying byte stream, and follows the same connection-pool
|
||||
pattern as the TCP and Tor transports.
|
||||
|
||||
It maintains two pools: a `ConnectingPool` for background SOCKS5
|
||||
connection attempts, and an established pool of `NymConnection` entries.
|
||||
Each `NymConnection` holds a write half, a per-connection receive task,
|
||||
the configured MTU, and a connection timestamp.
|
||||
|
||||
| Property | Value |
|
||||
| -------- | ----- |
|
||||
| Addressing | IP:port or hostname:port |
|
||||
| Default MTU | 1400 bytes |
|
||||
| Framing | FMP header-based (shared with TCP) |
|
||||
| Connection model | Outbound-only, non-blocking connect through SOCKS5 |
|
||||
| Platform | Cross-platform (requires external nym-socks5-client) |
|
||||
|
||||
### Outbound-Only
|
||||
|
||||
The Nym transport is strictly outbound. It supports no inbound service:
|
||||
`accept_connections()` returns `false` and `discover()` returns no
|
||||
peers. A node using the Nym transport can initiate links to remote peers
|
||||
through the mixnet, but cannot accept inbound connections over Nym. (A
|
||||
node can still accept inbound links over other transports it runs.)
|
||||
|
||||
### Address Types
|
||||
|
||||
The Nym transport accepts two address formats, parsed into an internal
|
||||
target address:
|
||||
|
||||
- **IP:port** — a numeric IP and port, sent to the SOCKS5 proxy as a
|
||||
numeric target.
|
||||
- **Hostname:port** — the hostname is passed through SOCKS5 so it is
|
||||
resolved on the exit side rather than locally.
|
||||
|
||||
Both forms are routed through the same SOCKS5 proxy.
|
||||
|
||||
### Connection Establishment
|
||||
|
||||
Connection setup follows the same non-blocking pattern as the TCP and
|
||||
Tor transports. When FMP needs to reach a peer, the node initiates a
|
||||
background connect (`connect_async`). The transport spawns a background
|
||||
tokio task that opens a SOCKS5 connection through the local
|
||||
`nym-socks5-client`, configures the socket (including TCP keepalive),
|
||||
splits the stream, and spawns a per-connection receive loop using the
|
||||
shared FMP stream reader. The call returns immediately while the connect
|
||||
proceeds in the background.
|
||||
|
||||
SOCKS5 connection setup through the mixnet can take much longer than a
|
||||
direct TCP connection because each connection traverses multiple mix
|
||||
nodes with timing obfuscation. Accordingly the connect timeout defaults
|
||||
to 300 seconds (`connect_timeout_ms`). Non-blocking connect is essential
|
||||
here — a blocking connect would stall the FMP event loop for the
|
||||
duration of mixnet setup. As a fallback, `send_async(addr, data)`
|
||||
performs a connect-on-send if no connection to the address yet exists.
|
||||
|
||||
Each outbound packet is checked against the configured MTU before being
|
||||
written; an oversized packet is rejected with an MTU-exceeded error
|
||||
rather than being sent.
|
||||
|
||||
### Startup Readiness
|
||||
|
||||
At startup the transport validates the configured `socks5_addr` and then
|
||||
probes the SOCKS5 port to wait for `nym-socks5-client` to become ready,
|
||||
using exponential backoff (starting at 1 second, capped at 10 seconds
|
||||
between attempts) up to `startup_timeout_secs` (default 120 seconds). If
|
||||
the proxy does not become reachable within that window, the transport
|
||||
logs a warning and starts anyway; outbound connections then fail until
|
||||
the `nym-socks5-client` becomes available.
|
||||
|
||||
### Session Independence
|
||||
|
||||
Same as TCP and Tor: loss of a Nym connection does **not** tear down the
|
||||
FIPS peer. Noise keys, MMP state, and FSP sessions survive reconnection.
|
||||
|
||||
### Configuration
|
||||
|
||||
The Nym transport block (`transports.nym.*`) has the following fields:
|
||||
|
||||
| Field | Default | Description |
|
||||
| ----- | ------- | ----------- |
|
||||
| `socks5_addr` | `127.0.0.1:1080` | Address (host:port) of the local nym-socks5-client SOCKS5 proxy |
|
||||
| `connect_timeout_ms` | `300000` | Outbound SOCKS5 connect timeout in milliseconds (300s) |
|
||||
| `mtu` | `1400` | Maximum FIPS packet size for Nym connections, in bytes |
|
||||
| `startup_timeout_secs` | `120` | Seconds to wait for nym-socks5-client to become ready at startup |
|
||||
|
||||
The Nym transport requires an external `nym-socks5-client`. Named
|
||||
instances are supported for multiple proxy endpoints. Unknown
|
||||
configuration keys are rejected.
|
||||
|
||||
### Statistics
|
||||
|
||||
The Nym transport exposes per-instance counters covering successful
|
||||
send/receive, send/receive errors, connection establishment, SOCKS5-level
|
||||
errors, connect timeouts, and MTU rejections.
|
||||
|
||||
## Discovery
|
||||
|
||||
@@ -695,7 +749,7 @@ X." FMP does not need to distinguish beacons from query responses.
|
||||
| Radio | Beacon | Shared RF channel, natural fit |
|
||||
| BLE | Advertising | GATT service UUID |
|
||||
|
||||
### Nostr Relay Discovery *(future direction)*
|
||||
### Nostr Relay Discovery
|
||||
|
||||
For internet-reachable transports, a node publishes a signed Nostr event
|
||||
containing its FIPS discovery information — public key and reachable
|
||||
@@ -707,6 +761,12 @@ feeds addresses to other transports. A node discovers via Nostr that a peer
|
||||
is reachable at UDP 1.2.3.4:9735, then establishes the link over the UDP
|
||||
transport.
|
||||
|
||||
For NAT'd UDP endpoints, a node may advertise `addr: "nat"` instead of a
|
||||
concrete address, signaling that peers should initiate STUN-assisted UDP
|
||||
hole punching. Offer/answer exchange uses Nostr gift-wrap (NIP-59) events
|
||||
on the configured DM relays; the resulting punched socket is adopted into
|
||||
the standard UDP transport via the bootstrap handoff path.
|
||||
|
||||
Key properties:
|
||||
|
||||
- Identity is built in — Nostr events are signed, so discovery information
|
||||
@@ -723,8 +783,11 @@ Key properties:
|
||||
> broadcast — the `discover()` trait method returns newly seen endpoints,
|
||||
> and per-transport `auto_connect()` / `accept_connections()` policies
|
||||
> control whether discovered peers are connected automatically or require
|
||||
> explicit configuration. TCP and Tor have no discovery mechanism.
|
||||
> Nostr relay discovery is not yet implemented.
|
||||
> explicit configuration. TCP and Tor have no built-in discovery mechanism.
|
||||
> Nostr relay discovery and STUN-assisted UDP hole punching are
|
||||
> implemented and toggled via configuration; see
|
||||
> [../reference/configuration.md](../reference/configuration.md) for the
|
||||
> `node.discovery.nostr.*` configuration tree.
|
||||
|
||||
## Transport Interface
|
||||
|
||||
@@ -778,7 +841,8 @@ TransportType {
|
||||
}
|
||||
```
|
||||
|
||||
Predefined types exist for UDP, TCP, Ethernet, WiFi, Tor, and Serial.
|
||||
Predefined types exist for UDP, TCP, Ethernet, WiFi, Tor, Nym, BLE, and
|
||||
Serial.
|
||||
|
||||
### Congestion Reporting
|
||||
|
||||
@@ -803,13 +867,15 @@ on all forwarded datagrams.
|
||||
| UDP | `SO_RXQ_OVFL` kernel drop counter | `recvmsg()` ancillary data on every packet |
|
||||
| TCP | Not implemented | Returns `None` (TCP handles congestion internally) |
|
||||
| Tor | Not implemented | Returns `None` (TCP handles congestion internally) |
|
||||
| Nym | Not implemented | Returns `None` (TCP handles congestion internally) |
|
||||
| Ethernet | Not implemented | Returns `None` |
|
||||
|
||||
### Transport Addresses
|
||||
|
||||
Transport addresses (`TransportAddr`) are opaque byte vectors. The transport
|
||||
layer interprets them (e.g., UDP/TCP resolve "host:port" strings (IP fast path, DNS fallback with 60s cache for UDP)); all layers above
|
||||
treat them as opaque handles passed back to the transport for sending.
|
||||
layer interprets them — e.g. UDP and TCP resolve `host:port` strings (IP
|
||||
fast path, DNS fallback with a 60s cache on UDP). All layers above treat
|
||||
them as opaque handles passed back to the transport for sending.
|
||||
|
||||
### Transport State Machine
|
||||
|
||||
@@ -829,10 +895,11 @@ transitions through `Starting` to `Up` (operational). `stop()` moves to
|
||||
| --------- | ------ | ----- |
|
||||
| UDP/IP | **Implemented** | Primary transport, AsyncFd/recvmsg, SO_RXQ_OVFL kernel drop detection |
|
||||
| TCP/IP | **Implemented** | FMP header-based framing, non-blocking connect, per-connection MSS MTU |
|
||||
| Ethernet | **Implemented** | AF_PACKET SOCK_DGRAM, EtherType 0x2121, beacon discovery, Linux only |
|
||||
| WiFi | Future direction | Infrastructure mode = Ethernet driver |
|
||||
| Ethernet | **Implemented** | AF_PACKET SOCK_DGRAM, EtherType 0x2121, neighbor beacons, Linux only |
|
||||
| WiFi | **Implemented** (via Ethernet transport, infrastructure mode) | mac80211 translates 802.11↔802.3; broadcast beacons unreliable through APs |
|
||||
| Tor | **Implemented** | Outbound SOCKS5, inbound via onion service, .onion and clearnet addressing |
|
||||
| BLE | Future direction | ATT_MTU negotiation, per-link MTU |
|
||||
| Nym | **Implemented** | Outbound-only SOCKS5 through nym-socks5-client, mixnet anonymity, IP/hostname addressing |
|
||||
| BLE | **Implemented** (Linux/glibc only; experimental) | L2CAP CoC, ATT_MTU negotiation, per-link MTU; musl/macOS/Windows skip |
|
||||
| Radio | Future direction | Constrained MTU (51–222 bytes) |
|
||||
| Serial | Future direction | SLIP/COBS framing, point-to-point |
|
||||
|
||||
@@ -840,7 +907,7 @@ transitions through `Starting` to `Up` (operational). `stop()` moves to
|
||||
|
||||
### TCP-over-TCP Avoidance
|
||||
|
||||
Running TCP application traffic over a reliable transport (TCP, WebSocket)
|
||||
Running TCP application traffic over a reliable transport (TCP, Tor)
|
||||
creates a layering violation where retransmission and congestion control
|
||||
operate at both levels. When the inner TCP detects loss (which may just be
|
||||
transport-layer retransmission delay), it retransmits, creating more traffic
|
||||
@@ -875,6 +942,15 @@ quality difference is significant. Link cost is not yet used in
|
||||
|
||||
## References
|
||||
|
||||
- [fips-intro.md](fips-intro.md) — Protocol overview and layer architecture
|
||||
- [fips-mesh-layer.md](fips-mesh-layer.md) — FMP specification (the layer above)
|
||||
- [fips-wire-formats.md](fips-wire-formats.md) — Transport framing details
|
||||
- [fips-concepts.md](fips-concepts.md) — Protocol overview
|
||||
- [fips-architecture.md](fips-architecture.md) — Layer architecture
|
||||
- [fips-mesh-layer.md](fips-mesh-layer.md) — FMP specification (the
|
||||
layer above)
|
||||
- [fips-mtu.md](fips-mtu.md) — How transport-reported `link_mtu`
|
||||
feeds the unified path-MTU model
|
||||
- [../reference/wire-formats.md](../reference/wire-formats.md) —
|
||||
Transport framing details
|
||||
- [../reference/configuration.md](../reference/configuration.md) —
|
||||
Per-transport configuration blocks
|
||||
- [../reference/transports.md](../reference/transports.md) —
|
||||
Per-transport statistics counter inventory
|
||||
|
||||
594
docs/design/port-advertisement-and-nat-traversal.md
Normal file
@@ -0,0 +1,594 @@
|
||||
# Port Advertisement and NAT Traversal via Nostr
|
||||
|
||||
## Abstract
|
||||
|
||||
This document describes two related-but-independent mechanisms that an
|
||||
application protocol can build on top of Nostr relays:
|
||||
|
||||
1. **Port advertisement.** A node publishes a parameterized replaceable
|
||||
event describing the application protocol it speaks, the version, and
|
||||
the endpoint(s) at which it can be reached. Other nodes discover the
|
||||
advert by querying relays.
|
||||
2. **NAT traversal.** When the advertised endpoint indicates that the
|
||||
responder is behind NAT, the two peers exchange ephemeral
|
||||
gift-wrapped offer/answer events through Nostr relays, run STUN
|
||||
against a public server to learn their reflexive addresses, and
|
||||
coordinate UDP hole punching so they can exchange application traffic
|
||||
over a direct UDP path.
|
||||
|
||||
The two mechanisms compose naturally — an advert that includes a
|
||||
`<protocol>:nat` endpoint signals "reach me by running the traversal
|
||||
protocol" — but they are independently useful. An advert with only
|
||||
public-IP endpoints needs no traversal. A pair of peers that already
|
||||
know each other's pubkeys but want to coordinate a traversal can do so
|
||||
without ever publishing a public advert.
|
||||
|
||||
The protocol described here is generic. Any application protocol can
|
||||
adopt it by picking its own kind number, `d`-tag scope, and endpoint
|
||||
schema. [FIPS](https://github.com/jmcorgan/fips) (the Free
|
||||
Internetworking Peering System) is used throughout the document as an
|
||||
example implementation; FIPS-specific values appear in clearly marked
|
||||
example blocks and do not affect the generic protocol shape.
|
||||
|
||||
No WebRTC, DTLS, or ICE stack is required. The protocol operates at
|
||||
the raw UDP level, using Nostr solely for ephemeral signaling and STUN
|
||||
solely for reflexive address discovery.
|
||||
|
||||
---
|
||||
|
||||
## Terminology
|
||||
|
||||
- **Application protocol.** The protocol that runs on top of the
|
||||
punched UDP channel after this document's procedures complete.
|
||||
- **Initiator.** The peer that discovers the responder's advert and
|
||||
begins the traversal exchange.
|
||||
- **Responder.** The peer that publishes a service advertisement and
|
||||
is willing to be dialled.
|
||||
- **Reflexive address.** The public `IP:port` tuple that a STUN server
|
||||
observes for a UDP socket — i.e., the NAT's external mapping for
|
||||
that socket.
|
||||
- **Punch socket.** The single UDP socket a peer uses for STUN, for
|
||||
the offer/answer exchange's address fields, for the punch packets
|
||||
themselves, and for the application traffic that follows. The same
|
||||
socket must be used across all phases of one traversal attempt.
|
||||
|
||||
### Socket lifecycle
|
||||
|
||||
The protocol assumes **per-peer, per-attempt punch sockets**:
|
||||
|
||||
- Each outbound traversal attempt allocates a fresh UDP socket bound
|
||||
to `0.0.0.0:0` (OS-assigned port).
|
||||
- That socket is owned by exactly one remote peer and exactly one
|
||||
traversal session.
|
||||
- STUN, the offer/answer reflexive-address fields, the punch packets,
|
||||
and the eventual adopted application transport all share that
|
||||
socket for the lifetime of the attempt.
|
||||
- If the attempt fails, the socket is discarded. A retry allocates a
|
||||
new socket and obtains a fresh reflexive address.
|
||||
- A long-lived application listener (for example, a fixed UDP port
|
||||
shared across peers) must **not** be reused as the punch socket —
|
||||
doing so couples NAT mappings and retry state across peers.
|
||||
|
||||
This rule is not optional: closing or rebinding the socket between
|
||||
phases invalidates the NAT mapping that the rest of the protocol
|
||||
depends on.
|
||||
|
||||
---
|
||||
|
||||
## Part 1: Service Advertisement
|
||||
|
||||
### Event shape
|
||||
|
||||
The advert is a NIP-01 parameterized replaceable event whose kind
|
||||
falls in the application-defined replaceable range
|
||||
`30000–39999`. The event carries:
|
||||
|
||||
- A `d` tag scoping the advert (so the same pubkey can publish
|
||||
multiple distinct adverts under different scopes).
|
||||
- A `protocol` tag carrying the application protocol's name, used as
|
||||
a discovery filter for peers that don't already know the
|
||||
responder's pubkey.
|
||||
- A `version` tag carrying the application protocol version.
|
||||
- An optional `expiration` tag (NIP-40) so a relay garbage-collects
|
||||
the advert when the responder goes offline without explicitly
|
||||
deleting it.
|
||||
- An optional `relays` tag listing relays where the responder
|
||||
subscribes for incoming signaling messages (used by Part 2).
|
||||
- An optional `stun` tag listing STUN servers the responder
|
||||
recommends.
|
||||
- A `content` field carrying the application-specific payload —
|
||||
typically the endpoint set, capability flags, and any encryption
|
||||
keys the application layer needs. The content may be plaintext or
|
||||
NIP-44-encrypted; encryption requires the consumer to already know
|
||||
the responder's pubkey.
|
||||
|
||||
The replaceable semantics let the responder update the advert in
|
||||
place under the same `d` tag. A NIP-09 deletion event removes the
|
||||
advert when the responder permanently retires.
|
||||
|
||||
```json
|
||||
{
|
||||
"kind": <application-specific>,
|
||||
"pubkey": "<responder_pubkey>",
|
||||
"created_at": <unix_seconds>,
|
||||
"tags": [
|
||||
["d", "<application-defined-scope>"],
|
||||
["protocol", "<application_protocol_name>"],
|
||||
["version", "<protocol_version>"],
|
||||
["relays", "wss://relay1.example.com", "wss://relay2.example.com"],
|
||||
["stun", "stun.l.google.com:19302"],
|
||||
["expiration", "<unix_seconds + ttl>"]
|
||||
],
|
||||
"content": "<application payload, optionally NIP-44 encrypted>",
|
||||
"sig": "<signature>"
|
||||
}
|
||||
```
|
||||
|
||||
### Endpoint schema
|
||||
|
||||
The `content` field is application-defined. Its structure typically
|
||||
includes a list of endpoints describing how the responder can be
|
||||
reached. Endpoint entries should distinguish:
|
||||
|
||||
- **Direct public endpoints** (transport + address + port) where any
|
||||
initiator can connect without traversal.
|
||||
- **NAT-mapped endpoints** that signal "I can be reached by running
|
||||
the traversal protocol against this transport on my pubkey."
|
||||
- **Anonymity-network endpoints** (e.g. Tor onion services) where
|
||||
the addressing scheme implies its own connection semantics.
|
||||
|
||||
#### FIPS example: kind 37195 advertisement
|
||||
|
||||
FIPS uses **kind `37195`** (the digits visually spell `FIPS` —
|
||||
7=F, 1=I, 9=P, 5=S). The `d` tag is hardcoded to
|
||||
`fips-overlay-v1`; the configurable `app` value populates the
|
||||
separate `protocol` tag, scoping adverts within a relay set
|
||||
without splitting them across multiple `d`-tag streams.
|
||||
|
||||
The advert content is a JSON document carrying a list of endpoint
|
||||
entries, each shaped as `{transport, addr}`. The transport string
|
||||
takes one of:
|
||||
|
||||
- `udp:host:port` — direct public UDP endpoint.
|
||||
- `udp:nat` — NAT-mapped UDP endpoint; reach via Part 2 traversal.
|
||||
- `tcp:host:port` — direct public TCP endpoint, for peers whose
|
||||
networks filter outbound UDP. Public-only; there is no
|
||||
`tcp:nat` analogue.
|
||||
- `tor:<onion>:<port>` — Tor onion-service endpoint.
|
||||
|
||||
FIPS publishes the advert with `expiration` set to `now +
|
||||
advert_ttl_secs` (default 1 hour) and refreshes it every
|
||||
`advert_refresh_secs` (default 30 minutes).
|
||||
|
||||
### Public-IP discovery on advertisement
|
||||
|
||||
A responder behind a NAT or wildcard-bound to a non-routable address
|
||||
needs to determine what external address to put in its advert. The
|
||||
responder uses a fixed precedence:
|
||||
|
||||
1. An operator-supplied external address override (FIPS:
|
||||
`transports.{udp,tcp}.external_addr`) wins.
|
||||
2. A non-wildcard `local_addr` is used directly.
|
||||
3. For a wildcard-bound UDP listener with an explicit "publish this"
|
||||
flag (FIPS: `public: true`), the runtime queries STUN against
|
||||
the configured servers and publishes the reflexive address.
|
||||
4. For a wildcard-bound TCP listener, no STUN equivalent exists.
|
||||
Implementations should refuse to silently advertise an unreachable
|
||||
endpoint; FIPS emits a loud WARN and omits the endpoint.
|
||||
|
||||
This precedence keeps adverts honest: an endpoint that appears in
|
||||
the published content is one the responder believes is reachable.
|
||||
|
||||
### Discovery (consumer side)
|
||||
|
||||
A consumer queries one or more relays for an advert it can act on.
|
||||
Two filter shapes are typical:
|
||||
|
||||
By author, when the responder's pubkey is already known:
|
||||
|
||||
```json
|
||||
["REQ", "<sub_id>", {
|
||||
"kinds": [<advert_kind>],
|
||||
"authors": ["<responder_pubkey>"],
|
||||
"#d": ["<application-defined-scope>"]
|
||||
}]
|
||||
```
|
||||
|
||||
By application protocol, for "open discovery" of any peer running
|
||||
the same application:
|
||||
|
||||
```json
|
||||
["REQ", "<sub_id>", {
|
||||
"kinds": [<advert_kind>],
|
||||
"#protocol": ["<application_protocol_name>"]
|
||||
}]
|
||||
```
|
||||
|
||||
Adverts whose `protocol` tag does not match the consumer's expected
|
||||
value, or whose `expiration` tag has elapsed, are rejected at
|
||||
validation. Consumers cache adverts in memory keyed by author npub
|
||||
and respect the embedded expiration.
|
||||
|
||||
#### FIPS example: discovery filters
|
||||
|
||||
The FIPS daemon issues both filter shapes: by-author for peers it
|
||||
intends to dial directly, and by-`#protocol` when an operator has
|
||||
opted into open discovery against the same application namespace.
|
||||
Cached adverts persist until their `expiration` lapses; a periodic
|
||||
prune drops expired entries.
|
||||
|
||||
---
|
||||
|
||||
## Part 2: NAT Traversal
|
||||
|
||||
The traversal protocol coordinates UDP hole punching between two
|
||||
peers via gift-wrapped Nostr signaling. It is invoked when the
|
||||
initiator decides to dial a NAT-mapped endpoint advertised by the
|
||||
responder.
|
||||
|
||||
### Signaling event shape
|
||||
|
||||
Signaling messages are ephemeral kinds in the range `20000–29999`,
|
||||
NIP-44-encrypted to the recipient, and NIP-59 gift-wrapped so the
|
||||
outer event is signed by an ephemeral keypair rather than the
|
||||
sender's long-term identity. The wrap carries a `p` tag pointing at
|
||||
the recipient's pubkey and an NIP-40 `expiration` tag bounding how
|
||||
long the relay should retain it.
|
||||
|
||||
#### FIPS example: signaling kind 21059
|
||||
|
||||
FIPS signaling uses **kind `21059`**. Wraps are addressed by `p`
|
||||
tag and published to the responder's NIP-17 inbox relay list (kind
|
||||
`10050`) when one is available, falling back to the local
|
||||
`dm_relays` configuration otherwise. Each side publishes its own
|
||||
inbox relay list on startup so dialers can discover it.
|
||||
|
||||
### Phase 1: Initiator STUN binding
|
||||
|
||||
Before constructing any signaling message, the initiator:
|
||||
|
||||
1. Allocates a fresh UDP punch socket bound to `0.0.0.0:0`.
|
||||
2. Sends a STUN Binding Request (RFC 8489) to one of its locally
|
||||
configured STUN servers.
|
||||
3. Parses the Binding Response, extracts the
|
||||
`XOR-MAPPED-ADDRESS` attribute, and records that as its
|
||||
reflexive address. Other STUN attributes are ignored.
|
||||
4. Records local-candidate addresses for the same socket port:
|
||||
active private non-loopback interface addresses (RFC1918 IPv4,
|
||||
IPv6 ULA) and probed local egress addresses.
|
||||
|
||||
The punch socket must remain open across all subsequent phases.
|
||||
Closing or rebinding it discards the NAT mapping.
|
||||
|
||||
### Phase 2: Initiator sends offer
|
||||
|
||||
The initiator constructs an offer payload containing its reflexive
|
||||
address, its local-candidate addresses, an opaque session
|
||||
identifier, freshness timestamps, and any application-specific
|
||||
parameters. The payload is NIP-44-encrypted to the responder's
|
||||
pubkey, wrapped with NIP-59, and published to the responder's
|
||||
signaling relays. The initiator also subscribes by `p` tag on
|
||||
those relays to receive the answer.
|
||||
|
||||
```json
|
||||
{
|
||||
"type": "offer",
|
||||
"sessionId": "<random_hex_32>",
|
||||
"issuedAt": <unix_millis>,
|
||||
"expiresAt": <unix_millis>,
|
||||
"nonce": "<random_nonce>",
|
||||
"senderNpub": "<initiator_npub>",
|
||||
"recipientNpub": "<responder_npub>",
|
||||
"reflexiveAddress": {"protocol":"udp","ip":"<ip>","port":<port>},
|
||||
"localAddresses": [{"protocol":"udp","ip":"<ip>","port":<port>}],
|
||||
"stunServer": "<host>:<port>",
|
||||
"app_params": { ... }
|
||||
}
|
||||
```
|
||||
|
||||
- `sessionId` is a random identifier correlating offer and answer.
|
||||
- `reflexiveAddress` is the address STUN observed in Phase 1.
|
||||
- `localAddresses` enables a same-LAN fast path when both peers
|
||||
happen to share a private subnet.
|
||||
- `stunServer` is informational, recording which server the
|
||||
initiator used.
|
||||
- `issuedAt` / `expiresAt` bound the freshness window — the
|
||||
responder rejects stale offers, since a NAT mapping that has not
|
||||
been refreshed in tens of seconds may already be gone.
|
||||
|
||||
### Phase 3: Responder validates and answers
|
||||
|
||||
The responder maintains a standing `p`-tagged subscription on its
|
||||
advertised signaling relays. On receiving an offer:
|
||||
|
||||
1. Decrypts the wrap and recovers the offer payload.
|
||||
2. Validates freshness (rejects if outside the configured window;
|
||||
see *Skew tolerance* below).
|
||||
3. Rejects replays — if the `sessionId` is in a recently-seen
|
||||
cache, drop the offer.
|
||||
4. Allocates its own punch socket (`0.0.0.0:0`) and runs its own
|
||||
STUN query.
|
||||
5. Constructs an answer payload that echoes `sessionId`, carries
|
||||
the responder's reflexive and local addresses, includes a
|
||||
`PunchHint { startAtMs, intervalMs, durationMs }` telling both
|
||||
sides when to begin probing and how aggressively, and is
|
||||
wrapped, encrypted, and published the same way as the offer.
|
||||
|
||||
```json
|
||||
{
|
||||
"type": "answer",
|
||||
"sessionId": "<same as offer>",
|
||||
"issuedAt": <unix_millis>,
|
||||
"expiresAt": <unix_millis>,
|
||||
"nonce": "<random_nonce>",
|
||||
"senderNpub": "<responder_npub>",
|
||||
"recipientNpub": "<initiator_npub>",
|
||||
"inReplyTo": "<offer_event_id>",
|
||||
"accepted": true,
|
||||
"reflexiveAddress": {"protocol":"udp","ip":"<ip>","port":<port>},
|
||||
"localAddresses": [{"protocol":"udp","ip":"<ip>","port":<port>}],
|
||||
"stunServer": "<host>:<port>",
|
||||
"punch": {"startAtMs": <ms>, "intervalMs": <ms>, "durationMs": <ms>},
|
||||
"offerReceivedAt": <unix_millis>,
|
||||
"app_params": { ... }
|
||||
}
|
||||
```
|
||||
|
||||
If the responder has no usable addresses, it returns
|
||||
`accepted: false` with an explanatory `reason` and no `punch`.
|
||||
|
||||
The optional `offerReceivedAt` field carries the responder's
|
||||
wall-clock at the moment the offer arrived. The initiator can
|
||||
combine its own `T1` (offer-publish time), `T2 = offerReceivedAt`,
|
||||
`T3` (answer's `issuedAt`), and `T4` (answer-receive time) into the
|
||||
NTP-style estimate `((T2 − T1) + (T3 − T4)) / 2`, giving a per-peer
|
||||
clock-skew measurement that's useful for tuning freshness windows
|
||||
and for telemetry.
|
||||
|
||||
**Immediately after publishing the answer**, the responder begins
|
||||
Phase 4 punching without waiting for any acknowledgement that the
|
||||
initiator received the answer. NAT mappings are decaying and time
|
||||
is the binding constraint.
|
||||
|
||||
The responder must bind the inner JSON `senderNpub` /
|
||||
`recipientNpub` fields to the actual Nostr pubkeys that delivered
|
||||
the gift wrap, rather than treating those JSON fields as
|
||||
independently trustworthy. The wrap pubkey is the authentication
|
||||
ground-truth.
|
||||
|
||||
### Phase 4: Hole punching
|
||||
|
||||
Both peers now know each other's reflexive and local addresses.
|
||||
Both begin sending UDP packets from their respective punch sockets:
|
||||
|
||||
1. Send punch packets every **`intervalMs`** (typically 200 ms)
|
||||
across each planned target path:
|
||||
- reflexive-to-reflexive
|
||||
- private-subnet local-address paths (when subnet-compatible)
|
||||
- mixed local/reflexive fallbacks
|
||||
2. Each punch packet carries a fixed magic header so transit and
|
||||
peer code can distinguish it from stray UDP traffic:
|
||||
|
||||
```text
|
||||
Bytes 0–3: <PROBE_MAGIC> (application-defined u32)
|
||||
Bytes 4–7: sequence number (u32, big-endian, starting at 0)
|
||||
Bytes 8–23: first 16 bytes of SHA-256(sessionId)
|
||||
```
|
||||
|
||||
3. On receiving a valid punch packet (magic matches, session-id
|
||||
hash matches), the peer records the source address as the
|
||||
confirmed peer address and replies with an acknowledgement
|
||||
packet under a different magic value:
|
||||
|
||||
```text
|
||||
Bytes 0–3: <ACK_MAGIC> (application-defined u32)
|
||||
Bytes 4–7: echoed sequence number
|
||||
Bytes 8–23: first 16 bytes of SHA-256(sessionId)
|
||||
```
|
||||
|
||||
4. On receiving an acknowledgement, the peer considers the path
|
||||
punched and transitions to Phase 5.
|
||||
|
||||
If both peers advertised compatible local-subnet candidates, the
|
||||
local-address path will typically punch through faster than the
|
||||
reflexive path. The first path to acknowledge wins.
|
||||
|
||||
### Phase 5: Application protocol takeover
|
||||
|
||||
Once the path has acknowledged in both directions:
|
||||
|
||||
- The application protocol takes over the punch socket.
|
||||
- The signaling subscription can be closed.
|
||||
- The application is responsible for sending keepalive traffic at
|
||||
least every 15 seconds to refresh the NAT mapping. A flow that
|
||||
goes idle longer risks losing its mapping and having to retraverse.
|
||||
|
||||
### Phase 6: Cleanup
|
||||
|
||||
After the attempt completes (success or failure):
|
||||
|
||||
1. Close the relay subscription used for signaling.
|
||||
2. Optionally publish a NIP-09 deletion event referencing any
|
||||
signaling events the peer published. Because the wraps were
|
||||
ephemeral kinds with NIP-40 expiration tags, well-behaved relays
|
||||
will discard them automatically without explicit deletion.
|
||||
3. Discard the per-attempt punch socket if the attempt failed; a
|
||||
retry must allocate a new socket and a fresh reflexive address.
|
||||
|
||||
If the responder is going offline permanently it should also
|
||||
delete its kind-37195 (or equivalent) advert.
|
||||
|
||||
### Timeouts and retries
|
||||
|
||||
- If the initiator publishes an offer and receives no answer
|
||||
within a configured window (e.g. 10 s from offer publish), the
|
||||
attempt has failed. Causes: responder offline, advert stale,
|
||||
responder relay unreachable.
|
||||
- If the answer arrives but no valid punch acknowledgement is
|
||||
observed within `durationMs` (typically 10 s), the attempt has
|
||||
failed. Causes: symmetric NAT on either side, firewall
|
||||
interference, stale reflexive addresses.
|
||||
|
||||
The initiator may retry with a fresh STUN query, a fresh punch
|
||||
socket, and a new offer. Repeated failures against the same
|
||||
responder should be suppressed by the application layer; see
|
||||
*Application-specific failure handling* below.
|
||||
|
||||
---
|
||||
|
||||
## Security
|
||||
|
||||
### Authentication
|
||||
|
||||
Offer and answer payloads are NIP-44-encrypted to the recipient and
|
||||
NIP-59 gift-wrapped, so only the intended recipient can decrypt.
|
||||
Authentication of the sender comes from the inner-wrap signature
|
||||
(the rumour signed by the sender's long-term identity inside the
|
||||
NIP-59 seal), **not** from the outer wrap signature (which is the
|
||||
ephemeral pubkey).
|
||||
|
||||
The inner JSON `senderNpub` / `recipientNpub` fields must be bound
|
||||
to the actual signing pubkey of the inner rumour. Treating those
|
||||
JSON fields as independently trustworthy is a vulnerability —
|
||||
implementations must compare them against the unwrapped signature.
|
||||
|
||||
Once the UDP path is punched, the raw UDP channel has **no inherent
|
||||
authentication or encryption**. The application layer is responsible
|
||||
for establishing its own security on the punched channel — for
|
||||
example, a Noise Protocol handshake keyed from the Nostr identity,
|
||||
or an application-specific authenticated-encryption layer. FIPS
|
||||
runs its FMP Noise IK handshake immediately after adoption; the
|
||||
identity proven by the Noise handshake is the same Nostr pubkey
|
||||
that signed the inner offer/answer rumour, so a man-in-the-middle on
|
||||
the relay cannot impersonate the responder.
|
||||
|
||||
### Replay protection
|
||||
|
||||
The `sessionId` and `issuedAt` / `expiresAt` fields together
|
||||
defeat replays at the signaling layer. The responder must keep a
|
||||
bounded cache of recently-seen `sessionId` values and reject
|
||||
duplicates within the freshness window.
|
||||
|
||||
### Skew tolerance
|
||||
|
||||
Strict freshness checks fail under modest clock skew between
|
||||
peers. Implementations should accept offers and answers whose
|
||||
timestamps are off by a small absolute amount (FIPS uses ±60 s),
|
||||
and feed observed skew into a per-peer estimate for telemetry and
|
||||
tuning. Outright rejection should be reserved for grossly stale or
|
||||
future-dated messages.
|
||||
|
||||
### Metadata exposure
|
||||
|
||||
Even though signaling content is encrypted, the gift-wrap metadata
|
||||
reveals that the initiator's ephemeral pubkey contacted the
|
||||
responder's pubkey at a particular time, through a particular
|
||||
relay. The advert itself is public and reveals the responder's
|
||||
pubkey and the application protocol it speaks.
|
||||
|
||||
If metadata privacy is required, the advert content can be
|
||||
encrypted (consumers must already know the responder's pubkey),
|
||||
both peers can use ephemeral Nostr identities rather than their
|
||||
long-term keys, and the operator can run a private relay.
|
||||
|
||||
### NAT mapping integrity
|
||||
|
||||
If too much wall-clock time elapses between STUN discovery and the
|
||||
hole-punch attempt, the reflexive address goes stale. Both peers
|
||||
should complete the entire signaling exchange within tens of
|
||||
seconds of their respective STUN queries. Relay latency is the
|
||||
primary risk factor. Implementations targeting flaky relays should
|
||||
prefer relays known to deliver ephemeral events sub-second.
|
||||
|
||||
---
|
||||
|
||||
## Relay requirements
|
||||
|
||||
The protocol works best with relays that:
|
||||
|
||||
- Support ephemeral event kinds (`20000–29999`) and do not persist
|
||||
them.
|
||||
- Honor NIP-40 `expiration` tags and garbage-collect expired
|
||||
events.
|
||||
- Deliver events with low latency (sub-second WebSocket push).
|
||||
- Support NIP-09 deletion requests.
|
||||
|
||||
Relays that do not support ephemeral kinds will store the
|
||||
signaling events as regular events. The encrypted content remains
|
||||
opaque, but persisted wraps are wasteful and expose metadata
|
||||
unnecessarily. Operators deploying this protocol at scale should
|
||||
prefer relays that handle ephemeral kinds correctly, or run their
|
||||
own.
|
||||
|
||||
---
|
||||
|
||||
## Failure modes
|
||||
|
||||
| Failure | Symptom | Mitigation |
|
||||
| --- | --- | --- |
|
||||
| Symmetric NAT (one side) | Punch timeout | Retry with port-prediction heuristics; otherwise fall back to an application-level relay |
|
||||
| Symmetric NAT (both sides) | Punch timeout | Application-level relay required |
|
||||
| Relay latency > 60 s | Stale reflexive address | Use low-latency relays; consider self-hosted relay |
|
||||
| Relay does not support ephemeral kinds | Signaling events persist | Use NIP-40 expiration + NIP-09 deletion as fallback |
|
||||
| Responder offline | No answer received | Initiator times out after configurable period |
|
||||
| Stale advert (responder no longer up) | Offer reaches no listener | Application-level failure suppression (see below) |
|
||||
| STUN server unreachable | No reflexive address | Fall back to alternate STUN server; fail if none reachable |
|
||||
| Firewall blocks outbound UDP | STUN fails entirely | NAT-traversal does not apply; reachable peers are limited to those that publish a non-UDP transport (e.g. TCP) and accept inbound |
|
||||
|
||||
### Application-specific failure handling
|
||||
|
||||
Repeated traversal failures against the same responder are common
|
||||
in practice — the responder may be offline, the advert may be
|
||||
stale, or the responder may be on a network that doesn't admit
|
||||
incoming UDP. A naive implementation that retries on every dial
|
||||
attempt floods the relay layer and the operator's logs.
|
||||
|
||||
Implementations should layer per-peer suppression on top of the
|
||||
basic retry. The shape of that suppression is application-specific.
|
||||
|
||||
#### FIPS example: failure suppression
|
||||
|
||||
FIPS layers the following suppression machinery on the basic retry
|
||||
loop:
|
||||
|
||||
- **Per-npub WARN log rate-limit** (`warn_log_interval_secs`,
|
||||
default 5 minutes). Subsequent failures inside the window log
|
||||
at debug level instead.
|
||||
- **Per-npub consecutive-failure counter and extended cooldown.**
|
||||
After `failure_streak_threshold` (default 5) consecutive
|
||||
failures, the per-peer retry deadline is pushed past
|
||||
`extended_cooldown_secs` (default 30 minutes). Open-discovery
|
||||
sweeps consult the cooldown so they don't immediately re-enqueue
|
||||
the same peer.
|
||||
- **Stale-advert eviction on streak transition.** When a peer
|
||||
hits the failure-streak threshold, the daemon actively
|
||||
re-fetches its advert from the configured advert relays. If the
|
||||
advert has been removed or replaced, the cache entry is evicted
|
||||
and the streak resets; if the advert is unchanged, the cooldown
|
||||
applies.
|
||||
- **Per-peer skew estimate.** The NTP-style skew computed from
|
||||
`offerReceivedAt` is recorded so consistently-skewed peers don't
|
||||
trip the freshness check on every attempt.
|
||||
- **Bounded failure-state cache** (`failure_state_max_entries`,
|
||||
default 4096) with LRU eviction so the suppression machinery
|
||||
itself does not grow unbounded.
|
||||
|
||||
These knobs are documented in
|
||||
[FIPS configuration reference](https://github.com/jmcorgan/fips/blob/master/docs/reference/configuration.md)
|
||||
under `node.discovery.nostr`.
|
||||
|
||||
---
|
||||
|
||||
## References
|
||||
|
||||
- **RFC 8489** — Session Traversal Utilities for NAT (STUN)
|
||||
- **RFC 8445** — Interactive Connectivity Establishment (ICE)
|
||||
- **RFC 4787** — NAT Behavioral Requirements for Unicast UDP
|
||||
- **NIP-01** — Basic Nostr protocol flow
|
||||
- **NIP-09** — Event deletion request
|
||||
- **NIP-17** — Inbox relay list (kind `10050`) for direct-message
|
||||
routing
|
||||
- **NIP-40** — Expiration timestamp
|
||||
- **NIP-44** — Versioned encryption
|
||||
- **NIP-59** — Gift wrap
|
||||
- **NIP-78** — Application-specific data
|
||||
@@ -1,14 +1,20 @@
|
||||
# FIPS Spanning Tree Protocol Dynamics
|
||||
|
||||
A detailed study of the gossip-based spanning tree protocol, focusing on
|
||||
operational behavior under various mesh conditions. This document complements
|
||||
[fips-intro.md](fips-intro.md) with step-by-step walkthroughs of protocol
|
||||
dynamics rather than message formats and data structures.
|
||||
A detailed study of the gossip-based spanning tree protocol, focusing
|
||||
on operational behavior under various mesh conditions. This document
|
||||
complements [fips-concepts.md](fips-concepts.md) and
|
||||
[fips-architecture.md](fips-architecture.md) with step-by-step
|
||||
walkthroughs of protocol dynamics rather than message formats and
|
||||
data structures.
|
||||
|
||||
For wire formats, see [fips-wire-formats.md](fips-wire-formats.md) (TreeAnnounce section).
|
||||
For spanning tree algorithms and data structures, see
|
||||
[fips-spanning-tree.md](fips-spanning-tree.md). For how the spanning tree fits
|
||||
into mesh routing, see [fips-mesh-operation.md](fips-mesh-operation.md).
|
||||
For wire formats, see
|
||||
[../reference/wire-formats.md](../reference/wire-formats.md)
|
||||
(TreeAnnounce section). For spanning tree algorithms and data
|
||||
structures, see [fips-spanning-tree.md](fips-spanning-tree.md). For
|
||||
how the spanning tree fits into mesh routing, see
|
||||
[fips-mesh-operation.md](fips-mesh-operation.md). For the academic
|
||||
foundations and references that underpin this document, see
|
||||
[fips-prior-work.md](fips-prior-work.md).
|
||||
|
||||
## Contents
|
||||
|
||||
@@ -93,7 +99,7 @@ When a node starts with no peers, it bootstraps as a single-node network.
|
||||
**T0: Node A starts.**
|
||||
|
||||
- Generates or loads keypair `(npub_A, nsec_A)`
|
||||
- Computes `node_addr_A = SHA-256(npub_A)`
|
||||
- Computes `node_addr_A = SHA-256(pubkey_A)[..16]` (128 bits)
|
||||
- Initializes empty TreeState
|
||||
- Sets `parent = self` (A is its own root), `sequence = 1`
|
||||
- Records current timestamp
|
||||
@@ -518,6 +524,11 @@ converge to the same "link failed" state, though B detects it up to
|
||||
## 8. Parent Selection
|
||||
|
||||
Parent selection determines tree structure and routing efficiency.
|
||||
The algorithm itself (effective-depth ranking, hold-down, hysteresis,
|
||||
mandatory-switch bypass) is canonically documented in
|
||||
[fips-spanning-tree.md](fips-spanning-tree.md); this section walks
|
||||
through what re-selection looks like under specific dynamic
|
||||
conditions and the rationale for the local-only cost metric.
|
||||
|
||||
### Cost-Based Selection with Effective Depth
|
||||
|
||||
@@ -537,10 +548,15 @@ purely by tree depth without link quality consideration.
|
||||
2. **Compute effective depth for each candidate.** For every peer whose
|
||||
announced root matches the smallest root, the algorithm calculates
|
||||
`effective_depth = peer.depth + link_cost`, where `link_cost` comes from
|
||||
`peer_costs` (MMP-derived) or defaults to 1.0 when metrics have not yet
|
||||
converged. The best candidate is the peer with the lowest effective depth,
|
||||
with ties broken by numerically smallest `NodeAddr`. If the best candidate
|
||||
is already the current parent, no switch is needed.
|
||||
`peer_costs` (MMP-derived). During cold start, when no peer has MMP data
|
||||
yet (`peer_costs` is empty), unmeasured candidates default to 1.0; once
|
||||
any peer has MMP data, unmeasured candidates are skipped so a freshly
|
||||
connected peer cannot win on its default cost. Candidates whose ancestry
|
||||
already contains the local node are also rejected, preventing an
|
||||
alternating two-node loop. The best candidate is the peer with the
|
||||
lowest effective depth, with ties broken by numerically smallest
|
||||
`NodeAddr`. If the best candidate is already the current parent, no
|
||||
switch is needed.
|
||||
|
||||
3. **Check for mandatory switches.** Two conditions bypass all stability
|
||||
mechanisms and trigger an immediate parent change: the current parent is no
|
||||
@@ -571,10 +587,10 @@ re-evaluation independent of TreeAnnounce traffic).
|
||||
|
||||
Where ETX (Expected Transmission Count, from De Couto et al., "A
|
||||
High-Throughput Path Metric for Multi-Hop Wireless Routing", 2003) comes from
|
||||
bidirectional MMP delivery ratios and SRTT (Smoothed Round-Trip Time) from MMP
|
||||
timestamp-echo. When MMP
|
||||
metrics have not yet converged, `link_cost` defaults to 1.0, preserving
|
||||
depth-only behavior as a graceful fallback.
|
||||
bidirectional MMP delivery ratios and SRTT (Smoothed Round-Trip Time) from
|
||||
MMP timestamp-echo. During cold start, before any peer has MMP data, the
|
||||
default cost of 1.0 is used and the algorithm reduces to depth-only
|
||||
selection.
|
||||
|
||||
**What this means for tree structure**: The algorithm can prefer a deeper parent
|
||||
with a better link over a shallower parent with a poor link, when the effective
|
||||
@@ -942,28 +958,16 @@ costs to form efficient tree structures.
|
||||
|
||||
### Prior Art and FIPS Contributions
|
||||
|
||||
The protocol builds on established foundations and adds several new elements:
|
||||
|
||||
**Derived from prior work**:
|
||||
|
||||
- Spanning tree coordinate routing (Yggdrasil/Ironwood, building on Kleinberg
|
||||
2007 and Cvetkovski/Crovella 2009)
|
||||
- Deterministic root discovery via smallest identifier (Yggdrasil; echoes
|
||||
IEEE 802.1D STP bridge ID selection)
|
||||
- CRDT-based distributed state (Shapiro et al. 2011)
|
||||
- Gossip dissemination (epidemic model; Kermarrec 2007)
|
||||
- Heartbeat-based failure detection (SWIM; Das et al. 2002)
|
||||
- ETX link metric (De Couto et al. 2003)
|
||||
- Hysteresis and hold-down for route stability (OSPF, BGP, IS-IS)
|
||||
|
||||
**FIPS additions**:
|
||||
|
||||
- Cost-aware parent selection using local-only link metrics (effective depth =
|
||||
tree depth + link cost), replacing Yggdrasil's depth-only selection
|
||||
- Combined ETX + SRTT link cost formula with MMP-measured components
|
||||
- Flap dampening with mandatory switch bypass
|
||||
- Announcement suppression for transient state changes
|
||||
- Tree-only bloom filter merge with split-horizon exclusion
|
||||
The protocol builds on established foundations (Yggdrasil/Ironwood
|
||||
tree-coordinate routing, IEEE 802.1D STP root election, CRDT-based
|
||||
distributed state, SWIM-style failure detection, ETX, OSPF-style
|
||||
hysteresis and hold-down) and adds several new elements (cost-aware
|
||||
parent selection on local-only metrics, the combined ETX + SRTT cost
|
||||
formula, flap dampening with mandatory-switch bypass, announcement
|
||||
suppression, and tree-only bloom filter merge with split-horizon).
|
||||
Both the prior-art map and the FIPS contributions list are
|
||||
consolidated in
|
||||
[fips-prior-work.md](fips-prior-work.md#fips-contributions).
|
||||
|
||||
---
|
||||
|
||||
@@ -971,75 +975,17 @@ The protocol builds on established foundations and adds several new elements:
|
||||
|
||||
### FIPS Internal Documentation
|
||||
|
||||
- [fips-spanning-tree.md](fips-spanning-tree.md) — Spanning tree algorithms and data structures
|
||||
- [fips-mesh-operation.md](fips-mesh-operation.md) — How the spanning tree fits into mesh routing
|
||||
- [fips-wire-formats.md](fips-wire-formats.md) — TreeAnnounce wire format
|
||||
- [fips-spanning-tree.md](fips-spanning-tree.md) — Spanning tree
|
||||
algorithms and data structures
|
||||
- [fips-mesh-operation.md](fips-mesh-operation.md) — How the spanning
|
||||
tree fits into mesh routing
|
||||
- [../reference/wire-formats.md](../reference/wire-formats.md) —
|
||||
TreeAnnounce wire format
|
||||
|
||||
### Yggdrasil Documentation
|
||||
### Prior Art and Academic Foundations
|
||||
|
||||
- [Yggdrasil v0.5 Release Notes](https://yggdrasil-network.github.io/2023/10/22/upcoming-v05-release.html)
|
||||
- [Ironwood Routing Library](https://github.com/Arceliar/ironwood)
|
||||
- [The World Tree (Yggdrasil Blog)](https://yggdrasil-network.github.io/2018/07/17/world-tree.html)
|
||||
- [Yggdrasil Implementation Overview](https://yggdrasil-network.github.io/implementation.html)
|
||||
|
||||
### Academic Foundations
|
||||
|
||||
#### Virtual Coordinate Routing
|
||||
|
||||
- Rao, A., Ratnasamy, S., Papadimitriou, C., Shenker, S., Stoica, I.
|
||||
["Geographic Routing without Location Information"](https://people.eecs.berkeley.edu/~sylvia/papers/p327-rao.pdf).
|
||||
MobiCom 2003. *Established virtual coordinate routing using network topology.*
|
||||
|
||||
#### Greedy Embedding Theory
|
||||
|
||||
- Kleinberg, R.
|
||||
["Geographic Routing Using Hyperbolic Space"](https://www.semanticscholar.org/paper/Geographic-Routing-Using-Hyperbolic-Space-Kleinberg/f506b2ddb142d2ec539400297ba53383d958abef).
|
||||
IEEE INFOCOM 2007. *Proved every connected graph has a greedy embedding in
|
||||
hyperbolic space; showed spanning trees enable coordinate assignment.*
|
||||
|
||||
- Cvetkovski, A., Crovella, M.
|
||||
["Hyperbolic Embedding and Routing for Dynamic Graphs"](https://www.cs.bu.edu/faculty/crovella/paper-archive/infocom09-hyperbolic.pdf).
|
||||
IEEE INFOCOM 2009. *Dynamic embedding for nodes joining/leaving; introduced
|
||||
Gravity-Pressure routing for failure recovery.*
|
||||
|
||||
- Crovella, M. et al.
|
||||
["On the Choice of a Spanning Tree for Greedy Embedding"](https://www.cs.bu.edu/faculty/crovella/paper-archive/networking-science13.pdf).
|
||||
Networking Science 2013. *Analysis of how tree structure affects routing stretch.*
|
||||
|
||||
- Bläsius, T. et al.
|
||||
["Hyperbolic Embeddings for Near-Optimal Greedy Routing"](https://dl.acm.org/doi/10.1145/3381751).
|
||||
ACM Journal of Experimental Algorithmics 2020. *Achieved 100% success ratio
|
||||
with 6% stretch on Internet graph.*
|
||||
|
||||
#### Link Metrics
|
||||
|
||||
- De Couto, D., Aguayo, D., Bicket, J., Morris, R.
|
||||
"A High-Throughput Path Metric for Multi-Hop Wireless Routing".
|
||||
MobiCom 2003. *Introduced ETX (Expected Transmission Count) as a link
|
||||
quality metric for wireless mesh networks.*
|
||||
|
||||
#### Routing Protocol Stability
|
||||
|
||||
- IEEE 802.1D. "IEEE Standard for Local and Metropolitan Area
|
||||
Networks: Media Access Control (MAC) Bridges". *Spanning Tree
|
||||
Protocol (STP) — root election via bridge ID, BPDU exchange.*
|
||||
|
||||
- Moy, J. [RFC 2328](https://datatracker.ietf.org/doc/html/rfc2328):
|
||||
"OSPF Version 2". 1998. *Link-state routing with cumulative path
|
||||
costs and SPF computation. FIPS's local-only cost approach is
|
||||
contrasted with OSPF's cumulative model in §8.*
|
||||
|
||||
#### Distributed Systems Primitives
|
||||
|
||||
- Shapiro, M., Preguiça, N., Baquero, C., Zawirski, M.
|
||||
"Conflict-free Replicated Data Types". SSS 2011.
|
||||
*Formal definition of CRDTs enabling coordination-free consistency.*
|
||||
|
||||
- Das, A., Gupta, I., Motivala, A.
|
||||
["SWIM: Scalable Weakly-consistent Infection-style Process Group Membership"](https://www.cs.cornell.edu/projects/Quicksilver/public_pdfs/SWIM.pdf).
|
||||
IPDPS 2002. *O(1) failure detection, O(log N) dissemination via gossip.*
|
||||
|
||||
- Kermarrec, A-M.
|
||||
["Gossiping in Distributed Systems"](https://www.distributed-systems.net/my-data/papers/2007.osr.pdf).
|
||||
ACM SIGOPS Operating Systems Review 2007. *Framework for gossip-based
|
||||
protocols achieving O(log N) propagation.*
|
||||
The Yggdrasil documentation and the academic-foundations bibliography
|
||||
(virtual coordinate routing, greedy embedding theory, link metrics,
|
||||
routing-protocol stability, and distributed systems primitives) are
|
||||
collected in
|
||||
[fips-prior-work.md](fips-prior-work.md#spanning-tree-dynamics-foundations).
|
||||
|
||||
259
docs/getting-started.md
Normal file
@@ -0,0 +1,259 @@
|
||||
# Getting Started with FIPS
|
||||
|
||||
FIPS (Free Internetworking Peering System) is a self-organizing
|
||||
encrypted mesh network built on Nostr identities. Your machine
|
||||
becomes a node in the mesh with a self-generated cryptographic
|
||||
identity, and existing networking software — SSH, web servers,
|
||||
file transfer, anything IPv6-native — runs over the mesh
|
||||
unchanged.
|
||||
|
||||
There are two common ways to deploy FIPS, and the rest of this
|
||||
guide and the linked docs branch accordingly:
|
||||
|
||||
- **As an overlay** on top of existing IP networks (Ethernet,
|
||||
WiFi, the public internet, Tor), FIPS lets your node reach
|
||||
any other peer regardless of NAT, ISP, or physical location.
|
||||
- **From the ground up** over non-IP transports — raw Ethernet,
|
||||
WiFi, Bluetooth — FIPS provides a complete permissionless
|
||||
network without any pre-existing IP infrastructure, ISP, or
|
||||
DNS.
|
||||
|
||||
The two paths share a lot of common ground — install, identity,
|
||||
configuration. They diverge mainly in transport setup and the
|
||||
deployment topology you choose.
|
||||
|
||||
There is no central server. Any node can run; any pair of
|
||||
running nodes can mesh.
|
||||
|
||||
## What you'll need
|
||||
|
||||
- A Linux, macOS, or Windows host. Linux is the most exercised
|
||||
platform; macOS and Windows installers are available.
|
||||
- The pre-built installer for your platform (see the project
|
||||
README's [Quick start](../README.md#quick-start) section for
|
||||
download links), **or** a source checkout if you want to build
|
||||
the installer yourself.
|
||||
- For the source-build path only: a working Rust toolchain (the
|
||||
version pinned in `rust-toolchain.toml` is auto-installed by
|
||||
rustup), and the platform-specific build dependencies listed in
|
||||
[packaging/README.md](../packaging/README.md).
|
||||
|
||||
## Install
|
||||
|
||||
FIPS is installed by running a binary installer for your
|
||||
platform. The installer drops the daemon and CLI tools into
|
||||
system locations, installs systemd / launchd / Windows-service
|
||||
unit files, places a default `fips.yaml`, and creates the `fips`
|
||||
system group. There is no `cargo install` path: the daemon needs
|
||||
more than just binaries copied into place.
|
||||
|
||||
You can either build the installer yourself from source, or
|
||||
download a pre-built one from the release distribution. Both
|
||||
paths produce the same installer artifacts and the same
|
||||
post-install state.
|
||||
|
||||
### From the release distribution
|
||||
|
||||
The most direct path. The release distribution carries a
|
||||
per-platform installer:
|
||||
|
||||
- Debian/Ubuntu — `.deb` package
|
||||
- Arch Linux — `fips` AUR package
|
||||
- OpenWrt — `.ipk` package
|
||||
- macOS — `.pkg` installer
|
||||
- Windows — `.zip` with service-install scripts
|
||||
- Generic systemd Linux — `.tar.gz` with an `install.sh` script
|
||||
|
||||
See the [project README's Quick start section](../README.md#quick-start)
|
||||
for download links and per-platform invocations.
|
||||
|
||||
### From source
|
||||
|
||||
For development, custom builds, or unsupported architectures.
|
||||
The `packaging/` tree builds the same installer formats locally;
|
||||
you then apply the resulting installer the same way you would a
|
||||
downloaded one.
|
||||
|
||||
```sh
|
||||
git clone https://github.com/jmcorgan/fips.git
|
||||
cd fips/packaging
|
||||
make deb # or: tarball, ipk, aur, pkg, zip, all
|
||||
```
|
||||
|
||||
The resulting installer lands in `deploy/` at the project root.
|
||||
Apply it the same way you would a downloaded one (for example
|
||||
`sudo dpkg -i deploy/fips_*.deb` on Debian/Ubuntu).
|
||||
|
||||
See [packaging/README.md](../packaging/README.md) for per-format
|
||||
build details, cross-target options, and the full `make` target
|
||||
list.
|
||||
|
||||
### With Nix (flake)
|
||||
|
||||
On Nix/NixOS, a [flake](../flake.nix) at the project root builds the
|
||||
binaries from source with the pinned toolchain and no manual
|
||||
prerequisite install:
|
||||
|
||||
```sh
|
||||
nix build .#fips # all four binaries, into ./result/bin
|
||||
nix develop # dev shell with the toolchain + build deps
|
||||
```
|
||||
|
||||
This path produces binaries only — it does not run the installer, so
|
||||
there are no systemd units, no `fips` group, and no default `fips.yaml`.
|
||||
On NixOS, wire the daemon in through your system configuration using the
|
||||
flake's `packages.<system>.fips` output instead. See the Nix / NixOS
|
||||
section of [packaging/README.md](../packaging/README.md).
|
||||
|
||||
## What's installed and running
|
||||
|
||||
Here's what the installer leaves on your machine, what's
|
||||
running, and what you'll need to set up yourself.
|
||||
|
||||
**Binaries installed system-wide:**
|
||||
|
||||
- `fips` (daemon)
|
||||
- `fipsctl` (control-socket client)
|
||||
- `fipstop` (live-status TUI)
|
||||
- `fips-gateway`
|
||||
|
||||
**Files placed on disk:**
|
||||
|
||||
- `/etc/fips/fips.yaml` — default daemon config (preserved on
|
||||
upgrade).
|
||||
- `/etc/fips/fips.nft` — mesh-interface nftables baseline (used
|
||||
only when the firewall service is enabled).
|
||||
- `/etc/fips/fips.d/` — empty drop-in directory for operator
|
||||
nftables additions.
|
||||
- Systemd, launchd, or Windows-service unit files for the four
|
||||
fips services.
|
||||
|
||||
**System changes:**
|
||||
|
||||
- A `fips` system group is created. Add your user to it
|
||||
(`sudo usermod -aG fips $USER`, then re-login) to run
|
||||
`fipsctl` and `fipstop` without `sudo`.
|
||||
- The runtime directory `/run/fips/` exists with mode
|
||||
`0750 root:fips`.
|
||||
|
||||
**Services enabled and started on boot:**
|
||||
|
||||
- `fips.service` — the daemon. Brings up the `fips0` TUN
|
||||
adapter, listens on the configured transports, and exposes
|
||||
the control socket at `/run/fips/control.sock`.
|
||||
- `fips-dns.service` — wires `.fips` hostname resolution into
|
||||
the host resolver (a `/etc/systemd/resolved.conf.d/` drop-in
|
||||
pointing at `[::1]:5354` on systemd hosts).
|
||||
|
||||
**Services installed but not enabled** (operator opt-in):
|
||||
|
||||
- `fips-firewall.service` — applies `/etc/fips/fips.nft` to
|
||||
the mesh interface. See
|
||||
[how-to/enable-mesh-firewall.md](how-to/enable-mesh-firewall.md).
|
||||
|
||||
**What's working out of the box:**
|
||||
|
||||
- The daemon is running with a fresh **ephemeral** identity —
|
||||
a new Nostr keypair is generated on every start.
|
||||
- The `fips0` TUN adapter exists with the daemon's mesh address.
|
||||
- The daemon's transport listeners are up: UDP `0.0.0.0:2121`
|
||||
and TCP `0.0.0.0:8443`. They are inert at this point because
|
||||
no other node knows your daemon's npub yet — see "What's not
|
||||
yet configured" below.
|
||||
- `.fips` hostname resolution is plumbed into the host
|
||||
resolver.
|
||||
|
||||
**What's not yet configured** — these are what guide your next
|
||||
steps:
|
||||
|
||||
- **No peers.** The daemon has nobody to talk to until you add
|
||||
a static peer entry, enable Nostr-mediated discovery, or
|
||||
bring up a transport (Ethernet, Bluetooth) where peers find
|
||||
each other automatically on the same physical link.
|
||||
- **Ephemeral identity.** Your node's npub changes every
|
||||
restart. The
|
||||
[persistent-identity tutorial](tutorials/persistent-identity.md)
|
||||
walks through pinning the daemon to a stable Nostr keypair
|
||||
for any node others will reference by name.
|
||||
- **Mesh firewall not active.** Inbound exposure on `fips0`
|
||||
follows the host's existing firewall rules until you enable
|
||||
the baseline service.
|
||||
|
||||
## Reaching mesh nodes by name
|
||||
|
||||
A FIPS node is identified by its Nostr public key (`npub1...`).
|
||||
For ordinary IP software running over the mesh — SSH, web
|
||||
browsers, `ping`, file transfer — use the form `<npub>.fips`
|
||||
as the destination; the local `.fips` resolver translates that
|
||||
to the corresponding mesh IPv6 address so the FIPS node can be
|
||||
found. The resolver runs entirely on your machine and does not
|
||||
generate any external DNS traffic.
|
||||
|
||||
For shorter forms, the resolver also consults two host maps
|
||||
before falling back to direct npub lookup: `/etc/fips/hosts`
|
||||
(shipped pre-populated with the public test mesh roster, and
|
||||
freely editable for your own entries) and the `alias:` field
|
||||
on configured peers in `fips.yaml`. So `test-us01.fips`,
|
||||
`my-laptop.fips`, or any other shortname you map resolves the
|
||||
same way `<npub>.fips` does. See
|
||||
[how-to/host-aliases.md](how-to/host-aliases.md) for the full
|
||||
mechanics.
|
||||
|
||||
## Join the test mesh
|
||||
|
||||
The fastest way to see FIPS in action is to connect your daemon
|
||||
to the public FIPS test mesh. The
|
||||
[Join the Test Mesh](tutorials/join-the-test-mesh.md) tutorial
|
||||
walks through adding a single static peer entry, watching the
|
||||
link come up, and reaching both that peer and a second mesh node
|
||||
forwarded through it — a ten-minute exercise that demonstrates
|
||||
the central FIPS guarantee that one good peer connects you to
|
||||
the rest of the mesh.
|
||||
|
||||
## Where to go next
|
||||
|
||||
Documentation is organised into four sections, each with a different
|
||||
job. Pick the one that matches what you want to do.
|
||||
|
||||
### [Tutorials](tutorials/)
|
||||
|
||||
Step-by-step lessons that take you from zero to a working setup.
|
||||
Read these end-to-end. Start with
|
||||
[Join the Test Mesh](tutorials/join-the-test-mesh.md) and follow
|
||||
with
|
||||
[ipv6-adapter-walkthrough](tutorials/ipv6-adapter-walkthrough.md)
|
||||
to understand what each piece does, then move on to
|
||||
[persistent-identity](tutorials/persistent-identity.md) and
|
||||
the three Nostr-discovery tutorials —
|
||||
[resolve-peers-via-nostr](tutorials/resolve-peers-via-nostr.md),
|
||||
[advertise-your-node](tutorials/advertise-your-node.md), and
|
||||
[open-discovery](tutorials/open-discovery.md) — to give your
|
||||
node a stable npub, look up peer endpoints, publish your
|
||||
own, and join the ambient discovery namespace. Then [host-a-service](tutorials/host-a-service.md) for hosting
|
||||
a service on your node, and [ground-up-mesh](tutorials/ground-up-mesh.md)
|
||||
for the second deployment mode where two devices peer over
|
||||
Ethernet, WiFi, or Bluetooth with no IP between them.
|
||||
|
||||
### [How-To Guides](how-to/)
|
||||
|
||||
Task-oriented recipes for operators with a specific goal: enable a
|
||||
firewall, deploy the LAN gateway, set up Bluetooth peering,
|
||||
diagnose an MTU problem, configure persistent identity. Each guide
|
||||
takes the shortest correct path from "I want to do X" to "X is done".
|
||||
|
||||
### [Reference](reference/)
|
||||
|
||||
Lookup material consulted on demand: wire formats, configuration
|
||||
keys, command-line flags, control-socket commands. Austere by
|
||||
design; no guidance on when to use a feature.
|
||||
|
||||
### [Design](design/)
|
||||
|
||||
Architectural and protocol-level explanations: the mesh layer, the
|
||||
session layer, the spanning tree, Bloom-filter discovery, the
|
||||
unified MTU model, the IPv6 adapter. Read these to understand *why*
|
||||
FIPS makes the choices it does.
|
||||
|
||||
The design section's
|
||||
[fips-concepts.md](design/fips-concepts.md) is a good entry point if
|
||||
you want the mental model before touching any commands.
|
||||
29
docs/how-to/README.md
Normal file
@@ -0,0 +1,29 @@
|
||||
# How-To Guides
|
||||
|
||||
Task-oriented, step-by-step recipes for operators with a specific
|
||||
goal in mind. Each guide assumes the reader already knows what FIPS
|
||||
is and wants to get a particular thing done — enable a feature,
|
||||
deploy a component, troubleshoot a class of problem.
|
||||
|
||||
How-to guides do not teach concepts (that is the role of design/)
|
||||
and do not enumerate options (that is the role of reference/). They
|
||||
take the reader along the shortest correct path from "I want to do
|
||||
X" to "X is done".
|
||||
|
||||
## Available Guides
|
||||
|
||||
| Guide | Goal |
|
||||
| ----- | ---- |
|
||||
| [enable-mesh-firewall.md](enable-mesh-firewall.md) | Activate the default-deny nftables baseline on `fips0` |
|
||||
| [enable-nostr-discovery.md](enable-nostr-discovery.md) | Turn on Nostr-mediated discovery (3 capabilities — resolve, advertise, open — across 5 scenarios) |
|
||||
| [deploy-tor-onion.md](deploy-tor-onion.md) | Run a Tor onion service for inbound FIPS connections |
|
||||
| [tune-udp-buffers.md](tune-udp-buffers.md) | Set host sysctls so FIPS UDP sockets don't get clamped |
|
||||
| [tune-file-descriptors.md](tune-file-descriptors.md) | Raise `RLIMIT_NOFILE` so a busy node doesn't exhaust file descriptors (`EMFILE`) as peer count grows |
|
||||
| [run-as-unprivileged-user.md](run-as-unprivileged-user.md) | Run the daemon under a dedicated unprivileged service account (drops the default-root posture) |
|
||||
| [deploy-gateway.md](deploy-gateway.md) | Manually deploy `fips-gateway` on a non-OpenWrt Linux host (LAN-to-mesh outbound + mesh-to-LAN inbound port-forwards). For the OpenWrt path, see the gateway tutorial. |
|
||||
| [troubleshoot-gateway.md](troubleshoot-gateway.md) | Diagnostic recipes for the gateway, organised by half (outbound, inbound, common) |
|
||||
| [persistent-identity.md](persistent-identity.md) | Provision a stable Nostr keypair so the node keeps the same npub across restarts |
|
||||
| [host-aliases.md](host-aliases.md) | Use shortnames (`test-us01.fips`, `my-laptop.fips`) instead of full npubs by editing `/etc/fips/hosts` or setting peer aliases |
|
||||
| [set-up-bluetooth-peer.md](set-up-bluetooth-peer.md) | Configure a Bluetooth Low Energy peer link |
|
||||
| [set-up-80211s-mesh-backhaul.md](set-up-80211s-mesh-backhaul.md) | Link OpenWrt FIPS routers over an open 802.11s radio backhaul (FIPS provides encryption, authentication, and routing) |
|
||||
| [diagnose-mtu-issues.md](diagnose-mtu-issues.md) | Triage MTU-shaped failures and rule out their imposters (bufferbloat, transport saturation) |
|
||||
458
docs/how-to/deploy-gateway.md
Normal file
@@ -0,0 +1,458 @@
|
||||
# Deploy `fips-gateway` (Manual Linux-Host Setup)
|
||||
|
||||
`fips-gateway` is a separate service that runs alongside the FIPS
|
||||
daemon and bridges a non-FIPS LAN to the FIPS mesh in two
|
||||
independent directions: **outbound** (LAN clients reach mesh
|
||||
services through DNS proxy + virtual-IP NAT) and **inbound** (mesh
|
||||
peers reach LAN services through 1:1 port forwards on `fips0`).
|
||||
This guide covers the **manual Linux-host** deployment path —
|
||||
wiring DNS forwarding, route distribution, and firewall integration
|
||||
on a server or non-OpenWrt router by hand.
|
||||
|
||||
> **Running OpenWrt?** Use the
|
||||
> [tutorial](../tutorials/deploy-fips-gateway.md) instead. The OpenWrt
|
||||
> ipk ships with the `gateway:` block pre-populated and the init
|
||||
> script automates dnsmasq forwarding, RA route distribution, and the
|
||||
> global IPv6 prefix on `br-lan`. The OpenWrt path is the canonical
|
||||
> deployment of this feature; this how-to is the secondary path for
|
||||
> operators with a different LAN-edge box (a Linux server already
|
||||
> serving DHCP/DNS, a custom router distribution, etc.).
|
||||
|
||||
For the gateway design (NAT pipeline, virtual IP pool lifecycle, DNS
|
||||
resolution flow), see [../design/fips-gateway.md](../design/fips-gateway.md).
|
||||
For the full `gateway.*` configuration block, see the
|
||||
[Gateway section](../reference/configuration.md#gateway-gateway) of
|
||||
the configuration reference. For the `fips-gateway` binary's CLI
|
||||
flags, see [../reference/cli-fips-gateway.md](../reference/cli-fips-gateway.md).
|
||||
|
||||
## The two halves
|
||||
|
||||
The gateway exposes two independent features that share a common
|
||||
control plane (the same binary, the same nftables table `inet
|
||||
fips_gateway`, the same control socket `/run/fips/gateway.sock`, the
|
||||
same `gateway.*` config block). You can configure either half on its
|
||||
own or both together.
|
||||
|
||||
- **Outbound gateway** (LAN → mesh). Non-FIPS LAN workstations resolve
|
||||
`<npub>.fips` names against the gateway's DNS listener and receive
|
||||
AAAA answers from the gateway's virtual-IP pool. Outbound traffic
|
||||
to those addresses is DNAT'd to the real mesh address and SNAT'd
|
||||
(masqueraded) onto `fips0` under the gateway's mesh identity. The
|
||||
audience is unmodified LAN clients.
|
||||
|
||||
- **Inbound gateway** (mesh → LAN). A static `(listen_port, proto)
|
||||
→ [target_addr]:target_port` table — configured in
|
||||
`gateway.port_forwards[]` — exposes selected LAN services to the
|
||||
mesh as `<gateway-npub>.fips:<listen_port>`. Mesh peers connect to
|
||||
the gateway's mesh address; the gateway DNATs to the LAN target
|
||||
and masquerades on the LAN side so return traffic flows through
|
||||
conntrack. The audience is mesh peers reaching a service that
|
||||
happens to live on this LAN.
|
||||
|
||||
The two halves are independent. Configure the outbound half if you
|
||||
want LAN clients to *reach* the mesh; configure the inbound half if
|
||||
you want mesh peers to *reach into* the LAN; configure both if you
|
||||
want both.
|
||||
|
||||
## Common gateway-host setup
|
||||
|
||||
Both halves require the same host preparation. Work through this
|
||||
section first, then jump to whichever half (or both) you need.
|
||||
|
||||
### FIPS daemon prerequisites
|
||||
|
||||
The gateway runs alongside a `fips` daemon on the same host:
|
||||
|
||||
- The daemon must be running with the TUN adapter enabled (the
|
||||
`fips0` interface must exist).
|
||||
- The daemon's DNS resolver must be enabled (`dns.enabled: true`,
|
||||
default) and reachable from `fips-gateway`. By default that means
|
||||
`[::1]:5354` (IPv6 loopback). The gateway's default
|
||||
`dns.upstream` matches this; a v4 upstream like `127.0.0.1:5354`
|
||||
cannot reach a daemon bound on `[::1]:5354` because Linux IPv6
|
||||
sockets bound to explicit `::1` do not accept v4-mapped traffic.
|
||||
|
||||
If the daemon is not yet running with these features, set up the
|
||||
daemon first — see [persistent-identity.md](persistent-identity.md)
|
||||
and [../reference/configuration.md](../reference/configuration.md).
|
||||
|
||||
### Kernel sysctls
|
||||
|
||||
```sh
|
||||
sudo sysctl -w net.ipv6.conf.all.forwarding=1
|
||||
sudo sysctl -w net.ipv6.conf.all.proxy_ndp=1
|
||||
```
|
||||
|
||||
`forwarding` lets the host route IPv6 packets between the LAN
|
||||
interface and `fips0`. `proxy_ndp` lets the gateway answer Neighbor
|
||||
Solicitation requests for virtual-pool addresses so LAN clients can
|
||||
resolve their link-layer addresses (only relevant for the outbound
|
||||
half, but harmless if you only run the inbound half).
|
||||
|
||||
Persist via a drop-in:
|
||||
|
||||
```sh
|
||||
sudo tee /etc/sysctl.d/60-fips-gateway.conf <<'EOF'
|
||||
net.ipv6.conf.all.forwarding = 1
|
||||
net.ipv6.conf.all.proxy_ndp = 1
|
||||
EOF
|
||||
sudo sysctl --system
|
||||
```
|
||||
|
||||
### Capability
|
||||
|
||||
`fips-gateway` requires `CAP_NET_ADMIN` to manage its nftables table
|
||||
(`inet fips_gateway`) and proxy-NDP entries. The packaged systemd
|
||||
unit (`fips-gateway.service`) runs as root, which satisfies this. For
|
||||
non-package installs, set the file capability:
|
||||
|
||||
```sh
|
||||
sudo setcap cap_net_admin+ep /usr/bin/fips-gateway
|
||||
```
|
||||
|
||||
### Pool route
|
||||
|
||||
At startup `fips-gateway` adds `local <pool-cidr> dev lo` to the
|
||||
local routing table. This tells the kernel to accept packets
|
||||
destined for pool addresses as locally-owned, enabling the NAT
|
||||
processing path. The route is cleaned up on shutdown. You do not
|
||||
need to install it manually; if you see "destination unreachable"
|
||||
errors for pool addresses on the gateway host, verify the route is
|
||||
present:
|
||||
|
||||
```sh
|
||||
ip -6 route show table local | grep <pool-cidr>
|
||||
```
|
||||
|
||||
### Minimum configuration
|
||||
|
||||
In `/etc/fips/fips.yaml`, populate the `gateway` block with at minimum
|
||||
`enabled: true`, `pool`, and `lan_interface`:
|
||||
|
||||
```yaml
|
||||
gateway:
|
||||
enabled: true
|
||||
pool: "fd01::/112"
|
||||
lan_interface: "enp3s0"
|
||||
```
|
||||
|
||||
Pick a pool CIDR that does **not** overlap with any address space in
|
||||
use on the LAN or in the mesh (the FIPS mesh occupies `fd00::/8`;
|
||||
pick a different `fdXX::/N`). The `/112` size yields 65 536 virtual
|
||||
IPs, which is the gateway's hard cap regardless of CIDR width.
|
||||
|
||||
This minimum config is enough to start the gateway. The `dns.*` block
|
||||
is optional and defaults to `listen: "[::1]:5353"` and
|
||||
`upstream: "[::1]:5354"`. The full block — including `dns.*`,
|
||||
`pool_grace_period`, `conntrack.*`, and `port_forwards[]` — is
|
||||
documented in
|
||||
[../reference/configuration.md#gateway-gateway](../reference/configuration.md#gateway-gateway).
|
||||
|
||||
### Start the service
|
||||
|
||||
```sh
|
||||
sudo systemctl enable --now fips-gateway
|
||||
```
|
||||
|
||||
Verify the unit came up:
|
||||
|
||||
```sh
|
||||
sudo systemctl status fips-gateway
|
||||
sudo journalctl -u fips-gateway -e
|
||||
```
|
||||
|
||||
The startup log will report `Gateway config loaded`,
|
||||
`DNS upstream is reachable`, `Created nftables table 'fips_gateway'`,
|
||||
and finally `fips-gateway running`. The unit's `ExecStartPre` waits up
|
||||
to 30 s for `fips0` to appear, which covers the cold-boot race where
|
||||
the daemon is still bringing up its TUN.
|
||||
|
||||
## Configure the outbound half
|
||||
|
||||
The outbound half lets LAN clients resolve `.fips` names and reach
|
||||
mesh destinations. Three operator decisions are involved: pool CIDR,
|
||||
DNS listen address, and how LAN clients learn the route to the pool
|
||||
and the resolver address.
|
||||
|
||||
### Choose the pool CIDR
|
||||
|
||||
```yaml
|
||||
gateway:
|
||||
pool: "fd01::/112"
|
||||
```
|
||||
|
||||
Constraints:
|
||||
|
||||
- Must not overlap with `fd00::/8` (the FIPS mesh address space).
|
||||
- Must not overlap with any LAN-side IPv6 prefix already in use.
|
||||
- `/112` is the practical width — wider just wastes address space
|
||||
because the pool is hard-capped at 65 536 entries. Narrower is
|
||||
fine if you want a smaller pool, but you'll reject DNS lookups
|
||||
faster under churn.
|
||||
|
||||
### Choose the DNS listen address
|
||||
|
||||
```yaml
|
||||
gateway:
|
||||
dns:
|
||||
listen: "[::1]:5353"
|
||||
upstream: "[::1]:5354"
|
||||
ttl: 60
|
||||
```
|
||||
|
||||
Common cases:
|
||||
|
||||
- **Another resolver on the host (the canonical case):** the default
|
||||
`listen: "[::1]:5353"` is loopback-only on an unprivileged port,
|
||||
so it never conflicts with dnsmasq, systemd-resolved, or BIND
|
||||
holding 53. Configure the existing resolver to forward `.fips`
|
||||
queries to `[::1]:5353` and you are done — this is what the
|
||||
OpenWrt ipk does automatically.
|
||||
- **No other resolver on the host:** set `listen: "[::]:53"`
|
||||
explicitly and LAN clients can query the gateway directly.
|
||||
- **systemd-resolved is on port 53:** the default already side-steps
|
||||
this — leave the listen address at `[::1]:5353` and configure the
|
||||
stub or a small forwarder to delegate `.fips` to the gateway. If
|
||||
you would rather have the gateway on 53 directly, disable the
|
||||
systemd stub listener (`DNSStubListener=no` in
|
||||
`/etc/systemd/resolved.conf`) and switch `listen` to `"[::]:53"`.
|
||||
See
|
||||
[troubleshoot-gateway.md](troubleshoot-gateway.md#port-conflict-on-the-dns-listen-port).
|
||||
- **Bind on the LAN address only:** `listen: "192.168.1.1:53"`
|
||||
exposes the resolver only to LAN clients, not loopback.
|
||||
|
||||
The gateway returns `REFUSED` for any non-`.fips` query — clients
|
||||
that point at it directly need a fallback resolver, or you should
|
||||
front it with a stub forwarder.
|
||||
|
||||
### Distribute the route to LAN clients
|
||||
|
||||
Each LAN client must route the gateway's pool CIDR to the gateway's
|
||||
LAN-side IPv6 address. Three options, in order of preference for
|
||||
production:
|
||||
|
||||
- **RA Route Information Option** (RFC 4191). If the LAN's RA daemon
|
||||
(`radvd`, `dnsmasq --enable-ra`, OpenWrt's `odhcpd`) supports
|
||||
publishing route options, configure it to advertise the pool CIDR
|
||||
with the gateway as next-hop. Clients pick this up automatically.
|
||||
|
||||
- **Static route on the LAN router**. If clients route through a
|
||||
central LAN router, add a static route entry there — the router
|
||||
then handles forwarding to the gateway. The exact syntax depends
|
||||
on the router OS.
|
||||
|
||||
- **Per-host static route** (testing or single-client deployments):
|
||||
|
||||
```sh
|
||||
sudo ip -6 route add fd01::/112 via fe80::<gateway-link-local>%<iface>
|
||||
# or, if the gateway has a stable global LAN address:
|
||||
sudo ip -6 route add fd01::/112 via <gateway-lan-addr>
|
||||
```
|
||||
|
||||
### Distribute the resolver to LAN clients
|
||||
|
||||
LAN clients also need to send `.fips` queries to the gateway. Two
|
||||
patterns:
|
||||
|
||||
- **Forward `.fips` from the LAN's main resolver.** If the LAN runs
|
||||
Pi-hole, Unbound, dnsmasq, or systemd-resolved as the central
|
||||
resolver, configure a conditional forward for `fips.`. Unbound
|
||||
example:
|
||||
|
||||
```text
|
||||
forward-zone:
|
||||
name: "fips."
|
||||
forward-addr: <gateway-lan-addr>@53
|
||||
```
|
||||
|
||||
dnsmasq example:
|
||||
|
||||
```text
|
||||
server=/fips/<gateway-lan-addr>
|
||||
```
|
||||
|
||||
Clients keep their existing DNS settings; only `.fips` queries are
|
||||
diverted.
|
||||
|
||||
- **Point clients directly at the gateway.** Simpler for testing,
|
||||
but the gateway returns `REFUSED` for non-`.fips` queries, so each
|
||||
client must also have a fallback resolver configured.
|
||||
|
||||
### Verify the outbound path
|
||||
|
||||
From a LAN client:
|
||||
|
||||
```sh
|
||||
dig @<gateway-lan-addr> hostname.fips AAAA
|
||||
# Expect an AAAA from the pool CIDR
|
||||
|
||||
ping6 hostname.fips
|
||||
# Should succeed via the gateway
|
||||
```
|
||||
|
||||
If either step fails, see
|
||||
[troubleshoot-gateway.md](troubleshoot-gateway.md#outbound-half-diagnostics).
|
||||
|
||||
## Configure the inbound half
|
||||
|
||||
The inbound half exposes a LAN-side service to mesh peers. Configured
|
||||
under `gateway.port_forwards[]`:
|
||||
|
||||
```yaml
|
||||
gateway:
|
||||
port_forwards:
|
||||
- listen_port: 8080
|
||||
proto: tcp
|
||||
target: "[fd12:3456::10]:80"
|
||||
- listen_port: 2222
|
||||
proto: tcp
|
||||
target: "[fd12:3456::20]:22"
|
||||
- listen_port: 5353
|
||||
proto: udp
|
||||
target: "[fd12:3456::10]:53"
|
||||
```
|
||||
|
||||
Field reference:
|
||||
|
||||
- `listen_port` — port on the gateway's `fips0` mesh-side address
|
||||
that mesh peers connect to. Must be non-zero. Each
|
||||
`(listen_port, proto)` pair must be unique across the list (the
|
||||
same port on TCP and UDP is allowed; the same port twice on the
|
||||
same proto is rejected at config-load time).
|
||||
- `proto` — `tcp` or `udp`.
|
||||
- `target` — IPv6 LAN destination as `[addr]:port`. IPv4 targets are
|
||||
rejected at parse time by the YAML deserializer (the field is
|
||||
typed `SocketAddrV6`). If the LAN host is reachable only by IPv4,
|
||||
put a small IPv6-aware reverse proxy in front of it on the gateway
|
||||
itself.
|
||||
|
||||
### Worked example: HTTP and DNS
|
||||
|
||||
Suppose the gateway runs on a LAN with an HTTP server at
|
||||
`[fd12:3456::10]:80` and a recursive resolver at
|
||||
`[fd12:3456::10]:53`, and you want mesh peers to reach them as
|
||||
`<gateway-npub>.fips:8080` (HTTP) and `<gateway-npub>.fips:5353`
|
||||
(DNS). Add to the gateway's `fips.yaml`:
|
||||
|
||||
```yaml
|
||||
gateway:
|
||||
port_forwards:
|
||||
- listen_port: 8080
|
||||
proto: tcp
|
||||
target: "[fd12:3456::10]:80"
|
||||
- listen_port: 5353
|
||||
proto: udp
|
||||
target: "[fd12:3456::10]:53"
|
||||
```
|
||||
|
||||
Reload:
|
||||
|
||||
```sh
|
||||
sudo systemctl restart fips-gateway
|
||||
```
|
||||
|
||||
From any mesh peer (the host name `gateway` is whatever the gateway's
|
||||
npub maps to in the local `hosts` file or via Nostr advert):
|
||||
|
||||
```sh
|
||||
curl http://gateway.fips:8080/
|
||||
dig @gateway.fips -p 5353 example.com A
|
||||
```
|
||||
|
||||
Each mesh-side request enters `fips0` on the listen port, gets DNAT'd
|
||||
to the LAN target, and the LAN-side masquerade rule rewrites the
|
||||
source to the gateway's LAN address so return traffic flows back
|
||||
through conntrack.
|
||||
|
||||
### Compose with the mesh firewall
|
||||
|
||||
`gateway.port_forwards[]` opens *mesh-side* listeners on `fips0`. If
|
||||
the host's mesh firewall is enabled (see
|
||||
[enable-mesh-firewall.md](enable-mesh-firewall.md)), inbound TCP/UDP
|
||||
on `fips0` for these ports must be permitted in the baseline or via
|
||||
a drop-in. The default baseline allows established/related and
|
||||
ICMPv6 only, so without an explicit allow rule, mesh peers will see
|
||||
TCP RSTs or silent drops on the listen port.
|
||||
|
||||
A typical drop-in for the worked example:
|
||||
|
||||
```nft
|
||||
# /etc/fips/fips.d/gateway-inbound.nft
|
||||
tcp dport 8080 accept
|
||||
udp dport 5353 accept
|
||||
```
|
||||
|
||||
Reload the firewall:
|
||||
|
||||
```sh
|
||||
sudo systemctl reload-or-restart fips-firewall.service
|
||||
```
|
||||
|
||||
If the inbound half doesn't need access control beyond the listen
|
||||
port itself, no source filter is needed. To restrict to specific
|
||||
mesh peers, follow the `ip6 saddr <addr> tcp dport <port> accept`
|
||||
pattern from the firewall guide.
|
||||
|
||||
### Verify the inbound path
|
||||
|
||||
From a mesh peer (any FIPS node):
|
||||
|
||||
```sh
|
||||
curl -v http://<gateway-npub>.fips:8080/
|
||||
```
|
||||
|
||||
A successful response confirms the full path: mesh ingress on
|
||||
`fips0`, DNAT to the LAN target, LAN-side masquerade, and conntrack-
|
||||
tracked return. If it fails, see
|
||||
[troubleshoot-gateway.md](troubleshoot-gateway.md#inbound-half-diagnostics).
|
||||
|
||||
## Operate and verify
|
||||
|
||||
`fips-gateway` exposes its own control socket at
|
||||
`/run/fips/gateway.sock`, separate from the daemon's
|
||||
`/run/fips/control.sock`. There is no `fipsctl gateway` subcommand —
|
||||
talk to it directly:
|
||||
|
||||
```sh
|
||||
echo '{"command":"show_gateway"}' | sudo nc -U /run/fips/gateway.sock
|
||||
echo '{"command":"show_mappings"}' | sudo nc -U /run/fips/gateway.sock
|
||||
```
|
||||
|
||||
`show_gateway` returns pool counters (`pool_total`, `pool_allocated`,
|
||||
`pool_active`, `pool_draining`, `pool_free`), `nat_mappings`,
|
||||
`dns_listen`, `uptime_secs`, and the active config snapshot.
|
||||
`show_mappings` returns the per-allocation list with virtual IP, mesh
|
||||
address, npub-derived `node_addr`, dns name, state (`Allocated`,
|
||||
`Active`, `Draining`), session count, and ages. For the full schema
|
||||
see [../reference/control-socket.md#gateway-command-catalog](../reference/control-socket.md#gateway-command-catalog).
|
||||
|
||||
The journal is the other primary signal:
|
||||
|
||||
```sh
|
||||
sudo systemctl status fips-gateway
|
||||
sudo journalctl -u fips-gateway -e
|
||||
```
|
||||
|
||||
Expect `MappingCreated`/`MappingRemoved` debug lines as DNS-driven
|
||||
allocations come and go (run with `--log-level debug` to see them),
|
||||
and `Final pool status` on shutdown. Errors in adding NAT rules or
|
||||
proxy-NDP entries surface here.
|
||||
|
||||
## See also
|
||||
|
||||
- [../tutorials/deploy-fips-gateway.md](../tutorials/deploy-fips-gateway.md) —
|
||||
the canonical, package-driven OpenWrt deployment path.
|
||||
- [../design/fips-gateway.md](../design/fips-gateway.md) — gateway
|
||||
design, NAT pipeline, virtual IP pool lifecycle, security
|
||||
considerations.
|
||||
- [Gateway section](../reference/configuration.md#gateway-gateway) of
|
||||
the configuration reference — full `gateway.*` block.
|
||||
- [../reference/cli-fips-gateway.md](../reference/cli-fips-gateway.md) —
|
||||
`fips-gateway` binary CLI flags.
|
||||
- [Gateway command catalog](../reference/control-socket.md#gateway-command-catalog)
|
||||
in the control-socket reference — JSON schema for `show_gateway`
|
||||
and `show_mappings`.
|
||||
- [troubleshoot-gateway.md](troubleshoot-gateway.md) — diagnostic
|
||||
recipes grouped by half.
|
||||
- [enable-mesh-firewall.md](enable-mesh-firewall.md) — mesh-firewall
|
||||
baseline and drop-ins (needed when exposing inbound ports).
|
||||
217
docs/how-to/deploy-tor-onion.md
Normal file
@@ -0,0 +1,217 @@
|
||||
# Deploy a Tor Onion Service for FIPS
|
||||
|
||||
This guide covers running a Tor onion service that accepts inbound
|
||||
FIPS peer connections.
|
||||
|
||||
For the Tor transport's design and the bridge-node pattern (running
|
||||
Tor and UDP simultaneously), see
|
||||
[../design/fips-transport-layer.md](../design/fips-transport-layer.md).
|
||||
For the full `transports.tor.*` config knob inventory, see
|
||||
[../reference/configuration.md](../reference/configuration.md).
|
||||
|
||||
## Inbound modes
|
||||
|
||||
FIPS supports two inbound Tor modes. (A third mode, `socks5`, is
|
||||
outbound-only and not covered here.)
|
||||
|
||||
- **`directory` mode** *(recommended)*. Tor manages the onion
|
||||
service via `HiddenServiceDir` and `HiddenServicePort` directives
|
||||
in `torrc`. FIPS reads the resulting `.onion` hostname from a
|
||||
file and binds a local TCP listener for Tor to forward inbound
|
||||
connections to. No control-port interaction is required, which
|
||||
makes this mode compatible with Tor's `Sandbox 1` seccomp-bpf
|
||||
hardening.
|
||||
- **`torrc` requires:** `HiddenServiceDir` + `HiddenServicePort`.
|
||||
- **`control_port` mode**. FIPS speaks to Tor's control port to
|
||||
create an ephemeral onion service at startup (`ADD_ONION`). The
|
||||
onion key lives only for the lifetime of the FIPS daemon's
|
||||
control-port session. This mode is **incompatible** with
|
||||
`Sandbox 1` — the sandbox forbids control-port-driven onion
|
||||
service management.
|
||||
- **`torrc` requires:** `ControlPort` (typically the Unix socket
|
||||
`/run/tor/control`) and a usable auth method
|
||||
(`CookieAuthentication 1` is the common choice).
|
||||
|
||||
Pick `directory` unless you have a specific reason to prefer
|
||||
`control_port`. The rest of this guide covers `directory` mode
|
||||
end-to-end.
|
||||
|
||||
## Prerequisites
|
||||
|
||||
- Tor daemon installed and running (Debian/Ubuntu: `apt install tor`)
|
||||
- FIPS daemon configured and able to start
|
||||
- Operator access to `/etc/tor/torrc` (or a drop-in under
|
||||
`/etc/tor/torrc.d/`)
|
||||
|
||||
## Step 1: Configure Tor's HiddenServiceDir
|
||||
|
||||
Add the following to `/etc/tor/torrc`:
|
||||
|
||||
```text
|
||||
HiddenServiceDir /var/lib/tor/fips
|
||||
HiddenServicePort 8443 127.0.0.1:8444
|
||||
```
|
||||
|
||||
`HiddenServiceDir` tells Tor where to store the onion service's
|
||||
private key and `hostname` file. `HiddenServicePort` declares that
|
||||
inbound TCP traffic to port 8443 of the onion address should be
|
||||
forwarded to `127.0.0.1:8444` on the local host — that is where FIPS
|
||||
will bind its listener.
|
||||
|
||||
The external port (`8443` here) is what peers will connect to over
|
||||
Tor; the internal target (`127.0.0.1:8444`) is purely local and is
|
||||
not directly reachable from the network.
|
||||
|
||||
## Step 2: Reload Tor and read the onion hostname
|
||||
|
||||
```sh
|
||||
sudo systemctl reload tor@default # or `tor` on systems without instance support
|
||||
```
|
||||
|
||||
After Tor processes the new config, the hostname file appears:
|
||||
|
||||
```sh
|
||||
sudo cat /var/lib/tor/fips/hostname
|
||||
# xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx.onion
|
||||
```
|
||||
|
||||
Tor regenerates the onion key only on first run (or if you remove
|
||||
`HiddenServiceDir`). The `hostname` value is stable across daemon
|
||||
restarts as long as `HiddenServiceDir` is preserved.
|
||||
|
||||
## Step 3: Verify HiddenServiceDir permissions
|
||||
|
||||
The directory must be readable only by the Tor user (Tor refuses to
|
||||
start otherwise):
|
||||
|
||||
```sh
|
||||
ls -la /var/lib/tor/fips
|
||||
# drwx------ debian-tor debian-tor ...
|
||||
```
|
||||
|
||||
With the shipped Debian systemd unit, FIPS runs as root and reads
|
||||
the `hostname` file directly — no permission adjustment is needed.
|
||||
|
||||
### Non-default deployments
|
||||
|
||||
If you run FIPS as an unprivileged user (custom packaging,
|
||||
hardened deployment, etc.), the FIPS daemon user needs read access
|
||||
to `hostname`. Options:
|
||||
|
||||
- Add the FIPS user to the `debian-tor` group and loosen group
|
||||
read on `HiddenServiceDir` (Tor still requires the directory
|
||||
itself to be `0700`, so this typically means making `hostname`
|
||||
itself group-readable rather than the directory).
|
||||
- Read `hostname` once at startup as root, then drop privileges.
|
||||
- Copy the hostname into a path the FIPS user can read, refreshed
|
||||
whenever the onion key changes.
|
||||
|
||||
## Step 4: Configure the FIPS Tor transport
|
||||
|
||||
In `/etc/fips/fips.yaml`, configure `transports.tor` with `mode:
|
||||
directory`:
|
||||
|
||||
```yaml
|
||||
transports:
|
||||
tor:
|
||||
mode: directory
|
||||
socks5_addr: "127.0.0.1:9050"
|
||||
connect_timeout_ms: 120000
|
||||
mtu: 1400
|
||||
advertised_port: 8443
|
||||
directory_service:
|
||||
hostname_file: "/var/lib/tor/fips/hostname"
|
||||
bind_addr: "127.0.0.1:8444"
|
||||
```
|
||||
|
||||
The `bind_addr` must match the *target* of the `HiddenServicePort`
|
||||
directive in `torrc`. The `hostname_file` path must match
|
||||
`HiddenServiceDir` plus `/hostname`.
|
||||
|
||||
`advertised_port` is the *virtual* onion port peers dial — i.e. the
|
||||
first number on the `HiddenServicePort` line, **not** the local
|
||||
target. The default is `443`; this guide uses `8443` on both sides
|
||||
to match the `HiddenServicePort 8443 127.0.0.1:8444` example
|
||||
above. Setting this explicitly is important if you ever flip
|
||||
`advertise_on_nostr: true`: the published advert otherwise
|
||||
defaults to `tor:<hash>.onion:443`, which won't match the actual
|
||||
onion port.
|
||||
|
||||
The `socks5_addr` is the Tor SOCKS5 proxy used for *outbound*
|
||||
connections to other onion services or clearnet endpoints (separate
|
||||
from inbound onion service handling).
|
||||
|
||||
Optional monitoring knobs: `control_addr` and `control_auth` (e.g.
|
||||
`/run/tor/control` and `cookie`) let the daemon read Tor's status
|
||||
through the control port even in `directory` mode. They are
|
||||
non-fatal on failure — the onion service still works without them.
|
||||
See [../reference/configuration.md](../reference/configuration.md)
|
||||
for the full key list and examples.
|
||||
|
||||
## Step 5: Reload the FIPS daemon
|
||||
|
||||
```sh
|
||||
sudo systemctl reload-or-restart fips
|
||||
```
|
||||
|
||||
At startup the daemon reads the `.onion` hostname from
|
||||
`hostname_file`, binds `127.0.0.1:8444`, and announces the onion
|
||||
endpoint internally. From this point inbound connections to
|
||||
`<your-onion>.onion:8443` arrive at FIPS over Tor.
|
||||
|
||||
## Step 6: Verify
|
||||
|
||||
Check that the FIPS daemon log shows the onion endpoint at startup:
|
||||
|
||||
```sh
|
||||
sudo journalctl -u fips -e | grep -i 'onion\|directory'
|
||||
```
|
||||
|
||||
You should see a line indicating the onion address FIPS will accept
|
||||
inbound connections on, and that the local bind on `127.0.0.1:8444`
|
||||
succeeded.
|
||||
|
||||
From another node configured with the Tor transport in `socks5` or
|
||||
`directory` mode, attempt to dial:
|
||||
|
||||
```sh
|
||||
fipsctl connect <peer-npub-or-hostname> <your-onion>.onion:8443 tor
|
||||
```
|
||||
|
||||
A successful `fipsctl show peers` afterwards on the inbound side
|
||||
shows the new peer with `transport=tor`.
|
||||
|
||||
## Optional: advertise the onion endpoint via Nostr discovery
|
||||
|
||||
If `node.discovery.nostr.enabled: true`, set
|
||||
`transports.tor.advertise_on_nostr: true` so the onion endpoint
|
||||
appears in this node's published advert. See
|
||||
[enable-nostr-discovery.md](enable-nostr-discovery.md) Scenario 2.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
- **Tor refuses to start with `Sandbox 1` and onion-service errors.**
|
||||
`Sandbox 1` requires `directory` mode and forbids creating onion
|
||||
services through the control port. Verify your `torrc` uses
|
||||
`HiddenServiceDir` (this guide), not `ADD_ONION` via control port.
|
||||
- **FIPS daemon fails to bind `127.0.0.1:8444`.** Another process is
|
||||
already bound to that port. Either stop the conflicting process or
|
||||
pick a different port and update both `torrc`'s
|
||||
`HiddenServicePort` target and `fips.yaml`'s `bind_addr` to match.
|
||||
- **Onion hostname is empty or missing.** Check `journalctl -u tor`
|
||||
for permission errors on `HiddenServiceDir`. The directory must be
|
||||
owned by the Tor user with mode `0700`.
|
||||
- **FIPS daemon cannot read `hostname_file`.** File is owned by the
|
||||
Tor user and not readable by the FIPS daemon user. Adjust
|
||||
permissions, or copy the hostname into a path the FIPS user can
|
||||
read.
|
||||
|
||||
## See also
|
||||
|
||||
- [../design/fips-transport-layer.md](../design/fips-transport-layer.md)
|
||||
— Tor transport design, three modes (`socks5`, `control_port`,
|
||||
`directory`), bridge-node pattern
|
||||
- [../reference/configuration.md](../reference/configuration.md) —
|
||||
full `transports.tor.*` configuration knob table
|
||||
- [enable-nostr-discovery.md](enable-nostr-discovery.md) — Scenario 2
|
||||
for advertising the onion endpoint to peers via Nostr
|
||||
211
docs/how-to/diagnose-mtu-issues.md
Normal file
@@ -0,0 +1,211 @@
|
||||
# Diagnose MTU Issues
|
||||
|
||||
MTU symptoms in FIPS look like ordinary network failures: handshakes
|
||||
succeed but bulk transfers hang, ssh connects but stalls after the
|
||||
banner, an HTTP request times out on the first response. This guide
|
||||
walks through the diagnostic surfaces that FIPS exposes so you can
|
||||
distinguish a real MTU problem from its frequent imposters
|
||||
(bufferbloat, transport saturation, transient packet loss).
|
||||
|
||||
For the underlying model — encapsulation overhead, proactive vs
|
||||
reactive PMTUD, the per-destination MTU storage layout — read
|
||||
[../design/fips-mtu.md](../design/fips-mtu.md) first.
|
||||
|
||||
## Symptom map
|
||||
|
||||
| Application symptom | Likely cause |
|
||||
| ------------------- | ------------ |
|
||||
| `iperf3 -c <host.fips>` control socket closes immediately after `Connecting to host`. | Forward-path MTU smaller than the negotiated MSS on the control connection. |
|
||||
| `ssh user@<host.fips>` shows the SSH banner then hangs forever. | First post-banner exchange exceeds the path MTU; SYN MSS clamp did not engage in time, or the path narrowed mid-session. |
|
||||
| `curl http://<host.fips>/` connects, then times out before the first response byte. | Same shape as the SSH-banner case, applied to the first server-to-client large packet. |
|
||||
| Throughput bursts then drops to zero, recovers, drops again, in seconds-long cycles. | Bufferbloat masquerading as MTU failure — usually the upload of the underlay link is saturated. See [Distinguishing bufferbloat](#distinguishing-bufferbloat-from-mtu-drops). |
|
||||
| `MtuExceeded` counters tick up under topology change but settle in seconds. | Normal: the reactive MTU mechanism doing its job. No action needed. |
|
||||
| `MtuExceeded` counters tick continuously under steady state. | Forward-path MTU smaller than what the source learned via `path_mtu` echo. After `mmp.path_mtu` has settled, this is a bug — see [File a bug](#file-a-bug). |
|
||||
|
||||
The first three are MTU candidates; the fourth is usually not. The
|
||||
fifth is benign. The sixth is the bug shape worth filing.
|
||||
|
||||
## Diagnostic toolkit
|
||||
|
||||
### `fipsctl show sessions`
|
||||
|
||||
The authoritative end-to-end MTU for an established session:
|
||||
|
||||
```sh
|
||||
fipsctl show sessions | jq '.sessions[] | {display_name, state, mmp: .mmp.path_mtu}'
|
||||
```
|
||||
|
||||
`mmp.path_mtu` is the value the session-layer MMP currently believes
|
||||
is in force end-to-end. It updates on each PathMtuNotification echo
|
||||
from the destination — immediately on decrease, with hysteresis on
|
||||
increase. A field that starts at `1280` (the IPv6 floor) and then
|
||||
climbs to a higher value as echoes arrive is healthy; one that
|
||||
oscillates between two values may indicate a flapping path.
|
||||
|
||||
### `fipsctl show transports`
|
||||
|
||||
Per-transport MTU. The `mtu` field reports the transport-wide
|
||||
default; for BLE, individual links may have a smaller negotiated
|
||||
ATT_MTU.
|
||||
|
||||
```sh
|
||||
fipsctl show transports | jq '.transports[] | {type, mtu}'
|
||||
```
|
||||
|
||||
### `fipsctl show cache`
|
||||
|
||||
The coordinate cache carries reverse-path-annotated MTU per
|
||||
destination — the freshest "what fit on the way back from the
|
||||
discovery target" estimate, consulted before the session has any
|
||||
PathMtuNotification feedback.
|
||||
|
||||
```sh
|
||||
fipsctl show cache | jq '.entries[] | {display_name, depth, path_mtu}'
|
||||
```
|
||||
|
||||
Entries without a `path_mtu` field are pre-discovery or were
|
||||
populated through a path that did not annotate the MTU.
|
||||
|
||||
### `fipsctl show peers`
|
||||
|
||||
Per-peer link state, including the link-layer MMP metrics. Useful
|
||||
mostly for ruling out underlying loss (loss rate near zero, SRTT
|
||||
sane) before chasing an MTU explanation.
|
||||
|
||||
```sh
|
||||
fipsctl show peers | jq '.peers[] | {display_name, mmp: .mmp}'
|
||||
```
|
||||
|
||||
### Trace logging
|
||||
|
||||
Module-scoped trace logging on the TUN reader and the MMP handler
|
||||
shows the per-packet decisions. The `tracing` macros default the
|
||||
target to the emitting module path, so the filter targets are the
|
||||
fully-qualified module paths under the `fips` crate.
|
||||
|
||||
```sh
|
||||
sudo systemctl edit fips
|
||||
# Add:
|
||||
# [Service]
|
||||
# Environment=RUST_LOG=info,fips::upper::tun=trace,fips::node::handlers::mmp=debug
|
||||
sudo systemctl restart fips
|
||||
sudo journalctl -u fips -f
|
||||
```
|
||||
|
||||
### tcpdump on `fips0`
|
||||
|
||||
Capturing on the TUN reveals the IPv6 packets the daemon hands the
|
||||
kernel and vice-versa. Two important caveats live in the design doc
|
||||
and are worth restating here:
|
||||
|
||||
- TX direction (outbound from a local app): tcpdump sees the packet
|
||||
**before** the daemon's TCP MSS clamp at the TUN boundary. The
|
||||
packet may be larger than the daemon will let leave the node.
|
||||
- RX direction (inbound to a local app): tcpdump sees the packet
|
||||
**after** the daemon's MSS clamp on inbound SYN-ACKs. The clamp
|
||||
fires only when `max_mss < kernel-natural-MSS`; otherwise it is a
|
||||
silent no-op.
|
||||
|
||||
```sh
|
||||
sudo tcpdump -ni fips0 -w /tmp/fips0.pcap port 22 or port 80
|
||||
# in another terminal, reproduce the symptom, then Ctrl-C
|
||||
```
|
||||
|
||||
Open the pcap in Wireshark and check segment sizes against what the
|
||||
session's `path_mtu` reports.
|
||||
|
||||
## Distinguishing bufferbloat from MTU drops
|
||||
|
||||
WAN bufferbloat (sustained upload saturation on a cable or DSL link)
|
||||
produces a retransmit signature that looks remarkably like
|
||||
oversized-packet drops. Both manifest as long stalls in TCP flows,
|
||||
both clear when you stop pushing data, both can ramp the loss-rate
|
||||
counter without obvious cause.
|
||||
|
||||
Two ways to disambiguate:
|
||||
|
||||
1. **Saturate the underlay first.** Run a reference upload outside
|
||||
FIPS (`iperf3 -c <internet-target>`) until it stabilises, then
|
||||
measure latency to the underlay's first hop with a separate `ping`.
|
||||
If RTT shoots up by hundreds of ms during the upload, the
|
||||
underlay buffer is the culprit, not FIPS MTU. Apply CAKE / fq_codel
|
||||
on the underlay router before continuing.
|
||||
|
||||
2. **Watch the FIPS counters during the symptom.** A real MTU
|
||||
problem ticks `MtuExceeded` (visible in `fipsctl show routing`'s
|
||||
`error_signals` block) and shifts the session's `mmp.path_mtu`
|
||||
downward. Bufferbloat ticks loss rate and RTT but leaves
|
||||
`path_mtu` and `MtuExceeded` alone.
|
||||
|
||||
If both signatures fire together, you have both problems.
|
||||
|
||||
## Cold-flow first-SYN
|
||||
|
||||
The MMP echo populates path-MTU state only after the first
|
||||
end-to-end exchange, but the TUN reader has to size the very first
|
||||
SYN before any echo has arrived. The cold-flow ceiling is the
|
||||
1143-byte conservative fallback derived from the 1280-byte IPv6
|
||||
floor. The first SYN may therefore be smaller than what the path
|
||||
ultimately supports; once MMP echoes arrive, subsequent flows use
|
||||
the larger learned value.
|
||||
|
||||
If the first SYN of a flow is still oversized relative to the path,
|
||||
the receiving transit node generates an `MtuExceeded`, the source
|
||||
shrinks immediately, and the next packet of the flow fits. This is
|
||||
expected for one round trip; it becomes a problem only if it
|
||||
persists.
|
||||
|
||||
## Fixes
|
||||
|
||||
The operator's choices, in rough order of preference:
|
||||
|
||||
### Pin a per-transport MTU floor in config
|
||||
|
||||
If a known link in the path has a small MTU that discovery does not
|
||||
pick up promptly (e.g., a Tor hop with an unusually tight cap), set
|
||||
a transport-level MTU floor on the relevant `transports.*` block.
|
||||
See [../reference/configuration.md](../reference/configuration.md)
|
||||
for the per-transport MTU keys.
|
||||
|
||||
### Tune host UDP buffers
|
||||
|
||||
For UDP transports specifically, undersized kernel buffers can drop
|
||||
oversized datagrams in a way that looks identical to MTU failure.
|
||||
See [tune-udp-buffers.md](tune-udp-buffers.md).
|
||||
|
||||
### Accept the floor on intrinsically small links
|
||||
|
||||
Tor and BLE link MTUs are properties of the medium, not tunables.
|
||||
For sessions that cross those links, the path MTU will be small; the
|
||||
fix is to design applications around it (smaller TCP windows, fewer
|
||||
large RTTs) rather than fight the transport.
|
||||
|
||||
### File a bug
|
||||
|
||||
The bug shape worth filing is session `mmp.path_mtu` itself
|
||||
oscillating, or `MtuExceeded` ticking *within* an established
|
||||
session after `mmp.path_mtu` has settled. The TCP-clamp mirror
|
||||
(`path_mtu_lookup`) is now updated on every successful proactive
|
||||
`PathMtuNotification` apply (tighter-only) as well as by the
|
||||
reactive `MtuExceeded` handler, so a steady-state divergence
|
||||
between the per-session `mmp.path_mtu` and the mirror used for
|
||||
new TCP flows is itself a defect, not an expected behavior.
|
||||
|
||||
Capture `fipsctl show sessions`, `fipsctl show cache`, `fipsctl
|
||||
show routing` (for the `error_signals` block), and a tcpdump from
|
||||
`fips0` covering the symptom window. See
|
||||
[../design/fips-mtu.md](../design/fips-mtu.md#per-destination-mtu-storage)
|
||||
for the per-destination MTU storage layout.
|
||||
|
||||
## See also
|
||||
|
||||
- [../design/fips-mtu.md](../design/fips-mtu.md) — encapsulation
|
||||
overhead, the proactive `path_mtu` field, the reactive
|
||||
`MtuExceeded` mechanism, MSS clamping, the no-fragmentation
|
||||
policy.
|
||||
- [../design/fips-mmp.md](../design/fips-mmp.md) — what the MMP
|
||||
metrics mean and how they are computed.
|
||||
- [../design/fips-ipv6-adapter.md](../design/fips-ipv6-adapter.md) —
|
||||
TUN-side ICMPv6 PTB generation and the MSS clamp.
|
||||
- [tune-udp-buffers.md](tune-udp-buffers.md) — host sysctl recipes
|
||||
that rule out kernel-buffer drops as a confounder.
|
||||
192
docs/how-to/enable-mesh-firewall.md
Normal file
@@ -0,0 +1,192 @@
|
||||
# Enable the Mesh-Interface Firewall
|
||||
|
||||
FIPS ships a default-deny nftables baseline at `/etc/fips/fips.nft` that
|
||||
restricts inbound traffic on the `fips0` mesh interface to conntrack
|
||||
replies and ICMPv6 echo. The baseline is **not** enabled by default — see
|
||||
[../design/fips-security.md](../design/fips-security.md) for the threat
|
||||
model and the rationale behind keeping activation explicit. This guide
|
||||
covers the operator steps to load the baseline, extend it with per-host
|
||||
allowances, and inspect drops.
|
||||
|
||||
## Activate the baseline
|
||||
|
||||
The package ships `fips-firewall.service`, a systemd oneshot that runs
|
||||
`nft -f /etc/fips/fips.nft` on start and removes the `inet fips` table
|
||||
on stop. To activate:
|
||||
|
||||
```sh
|
||||
sudo systemctl enable --now fips-firewall.service
|
||||
```
|
||||
|
||||
This loads the table now and arranges for it to load on every subsequent
|
||||
boot. To disable and tear it down:
|
||||
|
||||
```sh
|
||||
sudo systemctl disable --now fips-firewall.service
|
||||
```
|
||||
|
||||
To reload after editing `/etc/fips/fips.nft` or adding a drop-in under
|
||||
`/etc/fips/fips.d/`:
|
||||
|
||||
```sh
|
||||
sudo systemctl reload-or-restart fips-firewall.service
|
||||
```
|
||||
|
||||
The file is idempotent — it begins with `add table inet fips; flush
|
||||
table inet fips;` so re-running it replaces the live ruleset atomically.
|
||||
Equivalently:
|
||||
|
||||
```sh
|
||||
sudo nft -f /etc/fips/fips.nft
|
||||
```
|
||||
|
||||
## Folding the baseline into the host's main nftables
|
||||
|
||||
If you prefer to load the baseline from your existing
|
||||
`/etc/nftables.conf` rather than via the systemd unit, include it
|
||||
directly:
|
||||
|
||||
```nft
|
||||
# in /etc/nftables.conf
|
||||
include "/etc/fips/fips.nft"
|
||||
```
|
||||
|
||||
In that case do **not** enable `fips-firewall.service` — the host's main
|
||||
nftables setup owns the loading. The two paths are mutually exclusive.
|
||||
|
||||
## Extend with per-host allowances via drop-ins
|
||||
|
||||
The baseline drops everything inbound on `fips0` except conntrack
|
||||
replies and ICMPv6 echo. To open specific services to specific mesh
|
||||
nodes, drop a file into `/etc/fips/fips.d/` ending in `.nft`. Each
|
||||
file is included inline into the `inbound` chain at the marked point
|
||||
and may contain any nftables rule lines valid in that context.
|
||||
|
||||
Reload after editing:
|
||||
|
||||
```sh
|
||||
sudo systemctl reload-or-restart fips-firewall.service
|
||||
# or: sudo nft -f /etc/fips/fips.nft
|
||||
```
|
||||
|
||||
### Allow inbound SSH from a specific mesh node
|
||||
|
||||
```nft
|
||||
# /etc/fips/fips.d/ssh-from-bastion.nft
|
||||
ip6 saddr fd97:1234:5678:9abc:def0:1234:5678:9abc tcp dport 22 accept
|
||||
```
|
||||
|
||||
The source filter is the node's mesh address. To find a node's mesh
|
||||
address, look in their `fips.pub` (which contains the npub) and derive
|
||||
the `fd97:...` address from it, or query the running daemon:
|
||||
|
||||
```sh
|
||||
fipsctl show identity-cache
|
||||
fipsctl show peers
|
||||
```
|
||||
|
||||
### Allow inbound DNS broadly
|
||||
|
||||
Some services need to be reachable from any mesh node (a public DNS
|
||||
resolver, a public bootstrap node):
|
||||
|
||||
```nft
|
||||
# /etc/fips/fips.d/dns-public.nft
|
||||
udp dport 53 accept
|
||||
tcp dport 53 accept
|
||||
```
|
||||
|
||||
Omit the source filter only when the service is intended to be
|
||||
universally reachable on the mesh. The baseline's purpose is to make
|
||||
"universally reachable" an explicit decision rather than the default.
|
||||
|
||||
### Multiple nodes, one service
|
||||
|
||||
```nft
|
||||
# /etc/fips/fips.d/git-from-trusted.nft
|
||||
ip6 saddr {
|
||||
fd97:1111:2222:3333:4444:5555:6666:7777,
|
||||
fd97:8888:9999:aaaa:bbbb:cccc:dddd:eeee
|
||||
} tcp dport 9418 accept
|
||||
```
|
||||
|
||||
Set syntax keeps multi-node rules readable and is more efficient than a
|
||||
chain of individual rules.
|
||||
|
||||
## Verify with fipstop
|
||||
|
||||
`fipstop`'s Node tab carries a **Listening on fips0** panel
|
||||
(right-half of the Traffic block) that pairs each local IPv6
|
||||
listener with its current baseline-filter classification. After
|
||||
adding or editing a drop-in and reloading, this is the fastest
|
||||
way to confirm the rule landed correctly without manually
|
||||
parsing `nft list table inet fips`.
|
||||
|
||||
| Panel state | Reading |
|
||||
| ----------- | ------- |
|
||||
| Service row in **default White** with `OPEN` in the State column | The chain has a canonical, unrestricted accept rule for this (proto, port). The service is reachable from any mesh node. |
|
||||
| Service row in **DarkGray** with `filt` | No matching accept rule; the chain falls through to `counter drop`. The service is not reachable from the mesh. |
|
||||
| Service row in **DarkGray** with `filt?` | A rule references the port but uses matchers the panel cannot fully decompose (saddr filter, jump, daddr filter). The intent is operator-defined; inspect with `sudo nft list table inet fips` to see the actual rule. |
|
||||
| **Yellow banner** above the panel: "fips-firewall.service inactive — all listeners exposed" | The `inet fips` table is not loaded. Every listener is mesh-reachable (subject only to whatever ACL you have at the peer layer). |
|
||||
|
||||
A common workflow when extending the baseline is to keep `fipstop`
|
||||
open on the Node tab in one terminal while editing
|
||||
`/etc/fips/fips.d/` in another. After each
|
||||
`sudo systemctl reload-or-restart fips-firewall.service`, the panel
|
||||
re-classifies on the next poll tick and the affected row's State
|
||||
column flips. A row staying `filt` after you expected `OPEN`
|
||||
usually means the drop-in failed to load (syntax error in any file
|
||||
under `/etc/fips/fips.d/` aborts the whole reload) or carries a
|
||||
saddr filter that triggers `filt?` rather than `OPEN`.
|
||||
|
||||
The classifier is conservative: it recognizes only the canonical
|
||||
unrestricted shapes (`tcp dport N accept`, `udp dport N accept`,
|
||||
`dport { ... } accept`, `dport A-B accept`). Source-restricted
|
||||
accepts intentionally render as `filt?` rather than `OPEN` —
|
||||
the panel is a security screen, and any rule that varies by
|
||||
source is an operator decision the panel will not silently bless
|
||||
as fully open.
|
||||
|
||||
## Inspect drops
|
||||
|
||||
The baseline counter increments on every dropped packet. Inspect it:
|
||||
|
||||
```sh
|
||||
sudo nft list table inet fips
|
||||
```
|
||||
|
||||
Look for the `counter packets N bytes M drop` line at the bottom of the
|
||||
`inbound` chain. A non-zero counter means mesh nodes are sending
|
||||
traffic that hits the default-deny — usually benign (probes, neighbor
|
||||
discovery) but occasionally a misconfigured drop-in.
|
||||
|
||||
To see which packets are being dropped, uncomment the `log` line near
|
||||
the bottom of `/etc/fips/fips.nft`:
|
||||
|
||||
```nft
|
||||
log prefix "fips drop: " level info limit rate 10/minute
|
||||
```
|
||||
|
||||
Reload:
|
||||
|
||||
```sh
|
||||
sudo nft -f /etc/fips/fips.nft
|
||||
```
|
||||
|
||||
Then tail the kernel log:
|
||||
|
||||
```sh
|
||||
sudo journalctl -k -f -g "fips drop:"
|
||||
```
|
||||
|
||||
The rate-limit prevents flooding the journal under sustained probing.
|
||||
Adjust the rate, log level, or prefix as needed for the situation.
|
||||
Re-comment the rule when you are done; production hosts do not need
|
||||
the log line on by default.
|
||||
|
||||
## See also
|
||||
|
||||
- [../design/fips-security.md](../design/fips-security.md) — threat
|
||||
model, baseline design, and coexistence with other firewalls
|
||||
- [../reference/security.md](../reference/security.md) — consolidated
|
||||
security reference
|
||||
362
docs/how-to/enable-nostr-discovery.md
Normal file
@@ -0,0 +1,362 @@
|
||||
# Enable Nostr-Mediated Discovery and NAT Traversal
|
||||
|
||||
Nostr-mediated discovery lets FIPS nodes find each other (and punch
|
||||
through UDP NAT) using public Nostr relays as the signaling channel.
|
||||
The feature ships in every stock packaging artifact but is **off by
|
||||
default** — it activates when an operator sets
|
||||
`node.discovery.nostr.enabled: true`. Default relay and STUN-server
|
||||
lists ship in the config; both are optional overrides. See
|
||||
[../design/fips-nostr-discovery.md](../design/fips-nostr-discovery.md)
|
||||
for the design and rationale; see
|
||||
[../reference/configuration.md](../reference/configuration.md) for the
|
||||
full knob inventory.
|
||||
|
||||
Nostr discovery provides three independent capabilities. They can be
|
||||
enabled separately; most deployments end up using two or three of
|
||||
them together.
|
||||
|
||||
1. **Resolve a known peer's address by npub.** Your daemon consumes
|
||||
adverts from the relays to look up the current network endpoint
|
||||
for a peer you have configured by npub. You don't have to know
|
||||
their IP / port / transport in advance.
|
||||
2. **Publish your own endpoint so others can resolve you.** Your
|
||||
daemon publishes a signed advert listing the transports it will
|
||||
accept connections on. Has two sub-shapes depending on your
|
||||
network topology: UDP (using NAT traversal if needed) or TCP.
|
||||
Running a Tor onion service is a separate deployment mode,
|
||||
covered in its own section below.
|
||||
3. **Discover peers without prior configuration.** Your daemon
|
||||
subscribes to all adverts on a chosen application namespace and
|
||||
treats any publisher as a connection candidate. The most
|
||||
permissive posture; useful for ambient mesh participation.
|
||||
|
||||
Each capability is covered below as one or more scenarios with the
|
||||
minimal YAML fragment that enables it. Only keys relevant to Nostr
|
||||
discovery are shown; surrounding node, transport, TUN, DNS, and peer
|
||||
configuration follows the usual shape.
|
||||
|
||||
All scenarios assume `node.identity` is set to a persistent key — an
|
||||
ephemeral identity would invalidate any advert the moment the node
|
||||
restarts. See [persistent-identity.md](persistent-identity.md) for
|
||||
the persistent-key setup.
|
||||
|
||||
For hand-held walkthroughs of each capability, see the
|
||||
[resolve-peers-via-nostr](../tutorials/resolve-peers-via-nostr.md),
|
||||
[advertise-your-node](../tutorials/advertise-your-node.md),
|
||||
and [open-discovery](../tutorials/open-discovery.md)
|
||||
tutorials.
|
||||
|
||||
## Capability 1: Resolve a known peer's address by npub
|
||||
|
||||
The node does not publish any advert of its own. It only consumes
|
||||
adverts for peers it has explicitly listed with `via_nostr: true`.
|
||||
This is the right shape for a client that wants Nostr-mediated
|
||||
resolution without becoming a rendezvous target itself.
|
||||
|
||||
```yaml
|
||||
node:
|
||||
identity:
|
||||
persistent: true
|
||||
discovery:
|
||||
nostr:
|
||||
enabled: true
|
||||
advertise: false
|
||||
policy: configured_only
|
||||
|
||||
transports:
|
||||
udp:
|
||||
bind_addr: "0.0.0.0:2121"
|
||||
|
||||
peers:
|
||||
- npub: "npub1peer..."
|
||||
alias: "remote-node"
|
||||
via_nostr: true
|
||||
connect_policy: auto_connect
|
||||
```
|
||||
|
||||
What this achieves: dial endpoints for this peer are taken from
|
||||
the peer's published Nostr advert. `configured_only` is the
|
||||
default — it is shown here for clarity.
|
||||
|
||||
> **Note:** You can also supply a static address alongside
|
||||
> `via_nostr: true` (for example, while testing, or as a
|
||||
> known-good fallback if the advert is stale). Add an `addresses`
|
||||
> block to the peer entry; static addresses are tried first on
|
||||
> dial and Nostr-resolved endpoints are appended as additional
|
||||
> candidates.
|
||||
|
||||
## Capability 2: Publish your own endpoint so others can resolve you
|
||||
|
||||
This capability has three sub-scenarios depending on the network
|
||||
shape your node sits behind.
|
||||
|
||||
### Sub-scenario 2a: UDP (using NAT traversal if needed)
|
||||
|
||||
The node has a public IP (or a stable port-forward) and binds UDP on
|
||||
a known port. It publishes `udp:host:port` to the advert relays. Any
|
||||
peer that knows this node's npub and has Nostr discovery enabled can
|
||||
dial it without knowing the address out-of-band.
|
||||
|
||||
When UDP is wildcard-bound (`0.0.0.0:2121`, the default), the daemon
|
||||
needs help knowing what IP to put in the advert. There are two ways:
|
||||
STUN auto-discovery (`public: true`) or an explicit override
|
||||
(`external_addr`). Both are first-class options; pick the one that
|
||||
fits the deployment.
|
||||
|
||||
```yaml
|
||||
node:
|
||||
identity:
|
||||
persistent: true
|
||||
discovery:
|
||||
nostr:
|
||||
enabled: true
|
||||
advertise: true
|
||||
|
||||
transports:
|
||||
udp:
|
||||
bind_addr: "0.0.0.0:2121"
|
||||
advertise_on_nostr: true
|
||||
public: true # ← STUN auto-discovery
|
||||
```
|
||||
|
||||
Or, when the public IP is known up front (static residential IP,
|
||||
cloud Elastic IP behind 1:1 NAT, etc.):
|
||||
|
||||
```yaml
|
||||
transports:
|
||||
udp:
|
||||
bind_addr: "0.0.0.0:2121"
|
||||
advertise_on_nostr: true
|
||||
public: true # ← required, master switch
|
||||
external_addr: "203.0.113.45:2121" # ← explicit address
|
||||
```
|
||||
|
||||
`external_addr` accepts a bare IP (combined with the bind port) or a
|
||||
full `host:port`. `public: true` is the master switch that gates UDP
|
||||
advertisement; inside that branch, the daemon picks the advertised
|
||||
address in precedence order: explicit `external_addr` (no STUN
|
||||
observation), a non-wildcard `bind_addr`, or STUN auto-discovery.
|
||||
Setting `external_addr` alongside `public: true` skips STUN entirely
|
||||
— there is no logging cross-check. If UDP is bound directly to a
|
||||
public IP rather than to a wildcard, neither `external_addr` nor STUN
|
||||
is needed — but `advertise_on_nostr: true` and `public: true` are
|
||||
still both required for the daemon to publish the endpoint.
|
||||
|
||||
What this achieves: the node publishes a single
|
||||
`udp:<public-ip>:2121` endpoint to the three default advert relays
|
||||
(`wss://relay.damus.io`, `wss://nos.lol`, `wss://offchain.pub`).
|
||||
|
||||
What the other side needs: either a static `addresses` entry for this
|
||||
peer, or a peer entry with `via_nostr: true` and an empty (or
|
||||
omitted) `addresses` list — the advert-resolved endpoint will be used
|
||||
at dial time. Static and Nostr-resolved addresses can also be
|
||||
combined: when both are present, static addresses are tried first and
|
||||
Nostr-resolved endpoints are appended as fallback.
|
||||
|
||||
#### When the node is behind NAT
|
||||
|
||||
If this node doesn't have a stable public UDP endpoint, advertise
|
||||
`udp:nat`. The daemon runs the STUN + offer/answer exchange with
|
||||
the peer and punches through the NAT to establish a direct UDP
|
||||
link. The peer can either have a public endpoint of its own or
|
||||
also be behind NAT — both shapes work, as long as at least one
|
||||
side has a NAT type compatible with hole-punching.
|
||||
|
||||
```yaml
|
||||
node:
|
||||
identity:
|
||||
persistent: true
|
||||
discovery:
|
||||
nostr:
|
||||
enabled: true
|
||||
advertise: true
|
||||
dm_relays: # overrides the default three-relay
|
||||
- "wss://relay.damus.io" # set with two for demonstration;
|
||||
- "wss://nos.lol" # omit this block to keep the defaults
|
||||
stun_servers:
|
||||
- "stun:stun.l.google.com:19302"
|
||||
- "stun:stun.cloudflare.com:3478"
|
||||
|
||||
transports:
|
||||
udp:
|
||||
bind_addr: "0.0.0.0:2121"
|
||||
advertise_on_nostr: true
|
||||
public: false
|
||||
|
||||
peers:
|
||||
- npub: "npub1peer..."
|
||||
alias: "nat-peer"
|
||||
addresses:
|
||||
- transport: udp
|
||||
addr: "nat"
|
||||
via_nostr: true
|
||||
connect_policy: auto_connect
|
||||
```
|
||||
|
||||
What this achieves: the node publishes a `udp:nat` endpoint plus its
|
||||
signaling relays in the advert. When either side initiates, an
|
||||
encrypted offer is sealed to the peer's npub, a matching answer
|
||||
comes back, and both sides punch at the negotiated time. On success,
|
||||
the punch socket is adopted as an FMP UDP transport and Noise IK
|
||||
proceeds normally.
|
||||
|
||||
> **Validation:** `advertise_on_nostr: true` with `public: false` on
|
||||
> UDP requires `dm_relays` and `stun_servers` to be non-empty. Both
|
||||
> ship with non-empty defaults (three relays and three STUN servers
|
||||
> respectively), so the default config passes. The node fails
|
||||
> startup only if the operator has explicitly emptied either list —
|
||||
> a `udp:nat` advert without signaling relays or STUN servers is
|
||||
> unreachable by construction.
|
||||
|
||||
Hole-punching is best-effort. It works reliably when both sides are
|
||||
full-cone or port-restricted NATs. Symmetric NAT on either side
|
||||
typically defeats the punch — the public port a peer sees varies per
|
||||
remote endpoint, so the address learned via STUN does not match the
|
||||
mapping the peer actually needs. The punch attempt times out after
|
||||
`punch_duration_ms`. `udp:nat` is the only NAT-traversal mechanism
|
||||
in FIPS; when it can't succeed, there's no in-protocol substitute.
|
||||
Being reachable then becomes a deployment-prerequisite question
|
||||
rather than a transport question — a publicly reachable port (UDP
|
||||
or TCP — both require the same kind of network resource) published
|
||||
as a direct advert per Sub-scenario 2a or 2b.
|
||||
|
||||
### Sub-scenario 2b: TCP
|
||||
|
||||
The node has a public IP (or a stable port-forward) and accepts
|
||||
inbound TCP. It publishes `tcp:host:port` to the advert relays.
|
||||
|
||||
TCP endpoints exist to serve peers whose networks filter outbound
|
||||
UDP (corporate LANs, restrictive guest WiFi). NAT traversal does
|
||||
not apply: the publishing node is publicly reachable on TCP, and
|
||||
the dialing peer's network only needs to permit outbound TCP to
|
||||
the advertised port.
|
||||
|
||||
```yaml
|
||||
node:
|
||||
identity:
|
||||
persistent: true
|
||||
discovery:
|
||||
nostr:
|
||||
enabled: true
|
||||
advertise: true
|
||||
|
||||
transports:
|
||||
tcp:
|
||||
bind_addr: "0.0.0.0:8443"
|
||||
advertise_on_nostr: true
|
||||
external_addr: "203.0.113.45:8443"
|
||||
```
|
||||
|
||||
`external_addr` is typically required on cloud setups (AWS Elastic
|
||||
IP, etc.) where binding directly to the public IP returns
|
||||
`EADDRNOTAVAIL`. When TCP is bound directly to a public IP, the
|
||||
override is unnecessary.
|
||||
|
||||
What this achieves: the node publishes a `tcp:<public-ip>:8443`
|
||||
endpoint to the advert relays. Peers with Nostr discovery enabled
|
||||
dial by npub without out-of-band address exchange.
|
||||
|
||||
### Tor onion node
|
||||
|
||||
A separate deployment mode for nodes that want anonymity and
|
||||
censorship-resistance properties on the data plane. Functionally
|
||||
this still uses Capability 2 (publishing an endpoint to advert
|
||||
relays) — the difference is that the published endpoint is a Tor
|
||||
hidden service rather than a public IP.
|
||||
|
||||
The node runs a Tor onion service in directory mode (Tor-managed
|
||||
`HiddenServiceDir`) and advertises the `.onion` address. Peers dial
|
||||
via their local Tor SOCKS5 proxy without ever knowing the onion
|
||||
string out-of-band. For the Tor daemon side of this setup, including the inbound-mode
|
||||
trade-offs and the `torrc` directives each requires, see
|
||||
[deploy-tor-onion.md](deploy-tor-onion.md).
|
||||
|
||||
```yaml
|
||||
node:
|
||||
identity:
|
||||
persistent: true
|
||||
discovery:
|
||||
nostr:
|
||||
enabled: true
|
||||
advertise: true
|
||||
|
||||
transports:
|
||||
tor:
|
||||
mode: directory
|
||||
socks5_addr: "127.0.0.1:9050"
|
||||
advertised_port: 8443
|
||||
directory_service:
|
||||
hostname_file: "/var/lib/tor/fips/hostname"
|
||||
bind_addr: "127.0.0.1:8444"
|
||||
advertise_on_nostr: true
|
||||
```
|
||||
|
||||
What this achieves: the node publishes a `tor:<hash>.onion:8443`
|
||||
endpoint alongside any other advertised transports. The advert itself
|
||||
is still published over clearnet WebSocket relays — Tor protects the
|
||||
data plane, not the discovery plane. See the security and threat
|
||||
model section in
|
||||
[../design/fips-nostr-discovery.md](../design/fips-nostr-discovery.md#security-and-threat-model)
|
||||
for the trade-off and how to route relay traffic through Tor as well.
|
||||
|
||||
## Capability 3: Discover peers without prior configuration
|
||||
|
||||
Under `policy: open`, any node that publishes an advert under the
|
||||
same `app` namespace becomes a candidate. Discovered peers are queued
|
||||
for connection attempts subject to `open_discovery_max_pending`.
|
||||
|
||||
```yaml
|
||||
node:
|
||||
identity:
|
||||
persistent: true
|
||||
discovery:
|
||||
nostr:
|
||||
enabled: true
|
||||
advertise: true
|
||||
policy: open
|
||||
open_discovery_max_pending: 64
|
||||
app: "my-experiment.v1"
|
||||
|
||||
transports:
|
||||
udp:
|
||||
bind_addr: "0.0.0.0:2121"
|
||||
advertise_on_nostr: true
|
||||
public: true
|
||||
|
||||
peers: []
|
||||
```
|
||||
|
||||
What this achieves: peers are discovered entirely through ambient
|
||||
advert traffic on the configured relays. Setting a non-default `app`
|
||||
value (replacing `fips-overlay-v1`) scopes the discovery set to
|
||||
participants who opt into the same experiment and avoids being joined
|
||||
to unrelated overlays that happen to share the default namespace.
|
||||
|
||||
> **Scope warning:** Open discovery is an admission-free mode. Any
|
||||
> node that publishes on the same `app` name and passes the peer-ACL
|
||||
> check becomes a connection candidate. If you rely on peer ACLs for
|
||||
> admission control, verify that list is set correctly before
|
||||
> enabling this mode. See
|
||||
> [../reference/security.md](../reference/security.md) for the peer
|
||||
> ACL format.
|
||||
|
||||
## See also
|
||||
|
||||
- [../tutorials/resolve-peers-via-nostr.md](../tutorials/resolve-peers-via-nostr.md)
|
||||
— hand-held walkthrough of capability 1
|
||||
- [../tutorials/advertise-your-node.md](../tutorials/advertise-your-node.md)
|
||||
— hand-held walkthrough of capability 2 (publish, plus a
|
||||
short section on `udp:nat` NAT traversal)
|
||||
- [../tutorials/open-discovery.md](../tutorials/open-discovery.md)
|
||||
— hand-held walkthrough of capability 3 (open ambient
|
||||
discovery, the additive policy: open mode)
|
||||
- [../design/fips-nostr-discovery.md](../design/fips-nostr-discovery.md)
|
||||
— discovery runtime design, security model
|
||||
- [../reference/configuration.md](../reference/configuration.md) —
|
||||
full `node.discovery.nostr.*` and per-transport
|
||||
`advertise_on_nostr`/`public` table
|
||||
- [../reference/nostr-events.md](../reference/nostr-events.md) — Kind
|
||||
37195 advert format, Kind 21059 traversal signaling, Kind 10050
|
||||
inbox relay list
|
||||
- [deploy-tor-onion.md](deploy-tor-onion.md) — Tor daemon-side setup
|
||||
for advertising onion endpoints
|
||||
173
docs/how-to/host-aliases.md
Normal file
@@ -0,0 +1,173 @@
|
||||
# Use Shortnames Instead of Long Npubs
|
||||
|
||||
A FIPS node's canonical address is `<npub>.fips`. The npub is
|
||||
63 characters of bech32 — fine for the daemon, awkward to type
|
||||
or fit in a docs example. The local DNS resolver consults a
|
||||
host map before falling back to direct-npub resolution, so
|
||||
short names like `test-us01.fips` work as substitutes wherever
|
||||
`<npub>.fips` would.
|
||||
|
||||
This guide covers the two ways to populate that map and when
|
||||
to use which.
|
||||
|
||||
## When to use which
|
||||
|
||||
Two independent mechanisms feed the same DNS responder:
|
||||
|
||||
| Mechanism | Source | Scope | Reload |
|
||||
|-----------|--------|-------|--------|
|
||||
| Hosts file | `/etc/fips/hosts` | Node-local, intended for shared rosters | Auto on mtime change |
|
||||
| Peer alias | `alias:` field on a `peers:` entry | Node-local, scoped to configured peers | Daemon restart |
|
||||
|
||||
Pick the hosts file when:
|
||||
|
||||
- The shortname refers to a peer your operator-team agrees to
|
||||
call by that name across machines (the public test mesh
|
||||
ships this way).
|
||||
- You want the destination's name to resolve in DNS or appear
|
||||
in `fipsctl show peers` display even though it isn't in your
|
||||
`peers:` block — e.g., a mesh node you reach transitively
|
||||
through your direct peers. The hosts-file entry is for name
|
||||
resolution and display only; it does not stand in for the
|
||||
npub a peer-config entry requires.
|
||||
|
||||
Pick the peer alias when:
|
||||
|
||||
- The shortname is just a label *you* use locally for a peer
|
||||
that's already in your `peers:` block.
|
||||
- You want the alias to live with the rest of the peer config
|
||||
(one place to look) rather than in a separate file.
|
||||
|
||||
The two coexist. If both reference the same shortname,
|
||||
`/etc/fips/hosts` wins — the file is treated as the
|
||||
authoritative shared roster.
|
||||
|
||||
## What ships in the default `/etc/fips/hosts`
|
||||
|
||||
The installer drops `/etc/fips/hosts` populated with the
|
||||
public test mesh roster:
|
||||
|
||||
```text
|
||||
test-us01 npub1qmc3cvfz0yu2hx96nq3gp55zdan2qclealn7xshgr448d3nh6lks7zel98
|
||||
test-us02 npub10yffd020a4ag8zcy75f9pruq3rnghvvhd5hphl9s62zgp35s560qrksp9u
|
||||
test-us03 npub136yqae6na688fs75g95ppps3lxe07fvxefj77938zf47uhm6074sxw8ctm
|
||||
test-us03-next npub15m6c4ghuegx4pcde6tra8f7smn8vfv2wundyxwhkjynuerkrzmgsy09sh3
|
||||
test-us04 npub1gd7ye2qp2lphhzx75fynnjzaxx4dqanddecet0wtt5ss5ek8h9ps62wdkf
|
||||
test-de01 npub1260n42s06vzc7796w0fh3ny7zcpw6tlk4gq3940gmfrzl5c9pv2s3657q8
|
||||
test-es01 npub17lpmzulpc98d8ff727k6e98atxn3phzupzsqqwe54ytduym747ws4tw5zm
|
||||
test-uk01 npub1u0z26dc4qeneu5rvwvmpfhtwh3522ed6rlgxr9jarrfnjrc6ew4qxjysrs
|
||||
```
|
||||
|
||||
These resolve out of the box — `ping6 test-us01.fips` works
|
||||
even before you've added any peer to your config, as long as
|
||||
the destination is reachable through your mesh links.
|
||||
|
||||
If you don't intend to interact with the public test mesh,
|
||||
the entries are safe to comment out or delete. They are
|
||||
plain hosts-file lines, not protocol participants — removing
|
||||
them only changes name resolution on your machine.
|
||||
|
||||
## Add an entry to `/etc/fips/hosts`
|
||||
|
||||
Append a line to `/etc/fips/hosts`:
|
||||
|
||||
```text
|
||||
my-laptop npub1abc...xyz
|
||||
```
|
||||
|
||||
Format rules:
|
||||
|
||||
- One hostname and one npub per line, separated by
|
||||
whitespace.
|
||||
- Hostnames are lowercase letters, digits, and hyphens; max
|
||||
63 characters.
|
||||
- Comments start with `#` and continue to end of line; blank
|
||||
lines are ignored.
|
||||
- On duplicate hostnames, the last entry wins.
|
||||
|
||||
The daemon picks up the change on the next DNS query — no
|
||||
restart required (the file's mtime is checked on each query).
|
||||
Verify:
|
||||
|
||||
```sh
|
||||
dig my-laptop.fips AAAA +short
|
||||
```
|
||||
|
||||
Expect one `fd97:...` AAAA record.
|
||||
|
||||
`/etc/fips/hosts` is shipped as a dpkg conffile (and the AUR
|
||||
equivalent), so package upgrades preserve your edits. The
|
||||
file is `0644 root:root` — readable by anyone, writable by
|
||||
root.
|
||||
|
||||
## Add a peer alias
|
||||
|
||||
In `/etc/fips/fips.yaml`, set the `alias:` field on the peer
|
||||
entry:
|
||||
|
||||
```yaml
|
||||
peers:
|
||||
- npub: "npub1abc...xyz"
|
||||
alias: "my-laptop"
|
||||
addresses:
|
||||
- transport: udp
|
||||
addr: "192.0.2.10:2121"
|
||||
connect_policy: auto_connect
|
||||
```
|
||||
|
||||
Restart the daemon for the alias to take effect:
|
||||
|
||||
```sh
|
||||
sudo systemctl restart fips
|
||||
dig my-laptop.fips AAAA +short
|
||||
```
|
||||
|
||||
The alias also shows up in `fipsctl show peers` `display_name`
|
||||
column, so log entries and CLI output reference the peer by
|
||||
shortname instead of truncated npub.
|
||||
|
||||
## Resolution order
|
||||
|
||||
When the DNS responder receives a query for `<name>.fips`:
|
||||
|
||||
1. **Hosts file lookup.** If `<name>` matches an entry in
|
||||
`/etc/fips/hosts`, the daemon returns the AAAA record
|
||||
derived from that entry's npub.
|
||||
2. **Peer alias lookup.** If `<name>` matches the `alias`
|
||||
field on a configured peer, return that peer's AAAA.
|
||||
3. **Direct npub resolution.** If `<name>` is itself a valid
|
||||
bech32 npub (the canonical 63-char `npub1...` form), the
|
||||
daemon returns the AAAA derived from that npub directly.
|
||||
4. **NXDOMAIN.** If none of the above match, the query
|
||||
returns no answer.
|
||||
|
||||
The order means the hosts file overrides peer aliases on
|
||||
conflict. That's deliberate: the file represents
|
||||
operator-shared naming, the peer alias is a node-local label.
|
||||
|
||||
## Cross-references and ACLs
|
||||
|
||||
Aliases interact with the peer ACL — if you maintain
|
||||
`peers.allow` or `peers.deny` lists keyed on hostnames rather
|
||||
than npubs, those names go through the same hosts-file
|
||||
resolution. See
|
||||
[../reference/security.md](../reference/security.md) for the
|
||||
ACL format and the alias-resolution semantics.
|
||||
|
||||
`fipsctl connect` and `fipsctl disconnect` accept a shortname
|
||||
where they expect an npub. Resolution for these commands goes
|
||||
through `/etc/fips/hosts` only — peer-config `alias:` entries
|
||||
are not loaded by `fipsctl`, so a shortname that exists only as
|
||||
a peer alias must still be referenced by full npub on the CLI.
|
||||
See [../reference/cli-fipsctl.md](../reference/cli-fipsctl.md).
|
||||
|
||||
## See also
|
||||
|
||||
- [../reference/configuration.md § Host Mapping](../reference/configuration.md#host-mapping)
|
||||
— minimal reference entry for the host-map mechanism.
|
||||
- [../reference/cli-fipsctl.md](../reference/cli-fipsctl.md)
|
||||
— `fipsctl` arguments that accept shortnames.
|
||||
- [../reference/security.md](../reference/security.md)
|
||||
— peer ACL semantics with aliased entries.
|
||||
- [../design/fips-ipv6-adapter.md](../design/fips-ipv6-adapter.md)
|
||||
— the DNS resolver design and the npub-to-IPv6 derivation.
|
||||
232
docs/how-to/persistent-identity.md
Normal file
@@ -0,0 +1,232 @@
|
||||
# Provision a Persistent Identity
|
||||
|
||||
A FIPS node's identity is a Nostr keypair. Its public key (npub)
|
||||
determines the node's `fd00::/8` mesh address; peers and configs
|
||||
reference the node by that npub. Out of the box the daemon generates
|
||||
a fresh identity on every start (`node.identity.persistent: false`),
|
||||
which is fine for one-off testing but useless when other nodes need
|
||||
to refer to this one across restarts.
|
||||
|
||||
This guide covers the three ways to give a node a stable identity.
|
||||
For the configuration keys involved, see
|
||||
[../reference/configuration.md](../reference/configuration.md).
|
||||
|
||||
> **First time?** If you have just installed FIPS and want a
|
||||
> hand-held walkthrough of the package-default path (set
|
||||
> `persistent: true`, restart, observe the keys land), the
|
||||
> [persistent-identity tutorial](../tutorials/persistent-identity.md)
|
||||
> is the gentler entry point. This guide assumes an operator
|
||||
> picking among Options A/B/C for a deployment.
|
||||
|
||||
## When to use
|
||||
|
||||
Use a persistent identity for any node that:
|
||||
|
||||
- Other operators reference by npub (in their `peers` lists, `hosts`
|
||||
files, or ACL allow-lists).
|
||||
- Acts as a discoverable bootstrap or rendezvous (Nostr advert,
|
||||
static peer entry, gateway).
|
||||
- Is expected to keep its `fd00::/8` mesh address across restarts.
|
||||
|
||||
Stay with the ephemeral default for throw-away clients, sandbox
|
||||
nodes, and tests where you actively want a fresh identity per run.
|
||||
|
||||
## Option A: Let the package do it
|
||||
|
||||
The Debian/Ubuntu `.deb` and the Arch `fips` AUR package both ship a
|
||||
default `/etc/fips/fips.yaml` with `node.identity.persistent` left as
|
||||
the upstream default (false), so the daemon writes a fresh keypair to
|
||||
`/etc/fips/fips.{key,pub}` on every start until you set
|
||||
`persistent: true`. To pin the current keypair:
|
||||
|
||||
1. Install the package and start the daemon once so it generates
|
||||
`fips.key` / `fips.pub`:
|
||||
|
||||
```sh
|
||||
sudo systemctl start fips
|
||||
sudo systemctl status fips # confirm it came up
|
||||
```
|
||||
|
||||
2. Edit `/etc/fips/fips.yaml` and set:
|
||||
|
||||
```yaml
|
||||
node:
|
||||
identity:
|
||||
persistent: true
|
||||
```
|
||||
|
||||
3. Restart the daemon and verify the identity is reused:
|
||||
|
||||
```sh
|
||||
sudo systemctl restart fips
|
||||
fipsctl show status | grep -E '"npub"|"node_addr"'
|
||||
cat /etc/fips/fips.pub
|
||||
```
|
||||
|
||||
The npub printed by `fipsctl show status` should match
|
||||
`/etc/fips/fips.pub` and remain stable across subsequent restarts.
|
||||
|
||||
The package's `postinst` script does **not** generate the keypair —
|
||||
the daemon does, on first start. This means the keypair is only
|
||||
present after the first successful daemon start. If the daemon never
|
||||
came up cleanly (config error, permission problem), the key files
|
||||
will be missing.
|
||||
|
||||
### File layout and permissions
|
||||
|
||||
| Path | Mode | Owner | Contents |
|
||||
| ---- | ---- | ----- | -------- |
|
||||
| `/etc/fips/fips.key` | `0600` | `root:root` | Bech32 `nsec` (one line). |
|
||||
| `/etc/fips/fips.pub` | `0644` | `root:root` | Bech32 `npub` (one line). |
|
||||
|
||||
Both files live next to the highest-priority `fips.yaml` the daemon
|
||||
loaded. For non-systemd installs that use a different config path,
|
||||
the key files are placed in that config's directory.
|
||||
|
||||
## Option B: Generate manually
|
||||
|
||||
For from-source installs, custom config paths, or any deployment
|
||||
where you want to mint the keypair before the daemon ever runs.
|
||||
|
||||
### With `fipsctl keygen`
|
||||
|
||||
```sh
|
||||
sudo fipsctl keygen --dir /etc/fips
|
||||
```
|
||||
|
||||
This writes `/etc/fips/fips.key` (mode `0600`) and
|
||||
`/etc/fips/fips.pub` (mode `0644`), prints the new npub on stderr,
|
||||
and reminds you to set `persistent: true`. Add `--force` to overwrite
|
||||
an existing `fips.key`. Add `--stdout` to print `nsec` then `npub`
|
||||
to stdout instead of writing files.
|
||||
|
||||
To put the keypair in a non-default directory (e.g., a per-deployment
|
||||
config tree), pass `--dir` and point your `fips.yaml` search at the
|
||||
matching directory.
|
||||
|
||||
### Without the daemon installed
|
||||
|
||||
If you cannot run `fipsctl` (e.g., scripting on a build host), any
|
||||
nostr-tools-equivalent that emits a bech32 `nsec` works. Write the
|
||||
nsec to `fips.key` (mode `0600`) and the corresponding `npub` to
|
||||
`fips.pub` (mode `0644`).
|
||||
|
||||
### Hooking the keypair into the config
|
||||
|
||||
```yaml
|
||||
node:
|
||||
identity:
|
||||
persistent: true
|
||||
```
|
||||
|
||||
`persistent: true` plus a `fips.key` next to the loaded config is the
|
||||
intended steady-state setup.
|
||||
|
||||
## Option C: Provision from an existing nsec
|
||||
|
||||
To migrate an existing Nostr identity into a FIPS node — for example,
|
||||
re-using a personal npub for a node you operate.
|
||||
|
||||
1. Obtain the bech32 `nsec` for the identity.
|
||||
2. Write it to the config-adjacent key file:
|
||||
|
||||
```sh
|
||||
sudo install -m 0600 -o root -g root /dev/null /etc/fips/fips.key
|
||||
sudo bash -c 'printf "%s\n" nsec1... > /etc/fips/fips.key'
|
||||
```
|
||||
|
||||
3. Derive the matching `npub` and write `fips.pub`:
|
||||
|
||||
```sh
|
||||
# compute the npub with any nostr tool, then:
|
||||
sudo bash -c 'printf "%s\n" npub1... > /etc/fips/fips.pub'
|
||||
sudo chmod 0644 /etc/fips/fips.pub
|
||||
```
|
||||
|
||||
4. Set `persistent: true` and restart:
|
||||
|
||||
```yaml
|
||||
node:
|
||||
identity:
|
||||
persistent: true
|
||||
```
|
||||
|
||||
```sh
|
||||
sudo systemctl restart fips
|
||||
fipsctl show status | grep '"npub"'
|
||||
```
|
||||
|
||||
The reported npub should match the one you wrote to `fips.pub`.
|
||||
|
||||
## Verifying
|
||||
|
||||
The daemon prints the resolved identity at startup; the same value is
|
||||
queryable via the control socket:
|
||||
|
||||
```sh
|
||||
fipsctl show status | jq '{npub, node_addr, ipv6_addr}'
|
||||
cat /etc/fips/fips.pub
|
||||
```
|
||||
|
||||
The `npub` field of `show status` and the contents of `fips.pub`
|
||||
should match. The `node_addr` is the SHA-256 prefix used internally
|
||||
by FMP/FSP; the `ipv6_addr` is the routable `fd00::/8` mesh address
|
||||
derived from the node addr. Together they are stable for the lifetime
|
||||
of the keypair.
|
||||
|
||||
The journal also records the source on every start:
|
||||
|
||||
```text
|
||||
INFO Loaded persistent identity from key file path=/etc/fips/fips.key
|
||||
```
|
||||
|
||||
(`Generated persistent identity, saved to key file` on the first
|
||||
start; `Using ephemeral identity (new keypair each start)` when
|
||||
persistence is off.)
|
||||
|
||||
## Rotating
|
||||
|
||||
Key rotation is a destructive operation: every cached
|
||||
`(node_addr → npub)` mapping on every other node points at the old
|
||||
key, every Nostr advert and every static peer entry references the
|
||||
old npub, and every existing FSP session was authenticated under the
|
||||
old keypair. There is no in-protocol "key change" message.
|
||||
|
||||
To rotate:
|
||||
|
||||
1. Stop the daemon.
|
||||
|
||||
```sh
|
||||
sudo systemctl stop fips
|
||||
```
|
||||
|
||||
2. Remove the existing key files.
|
||||
|
||||
```sh
|
||||
sudo rm /etc/fips/fips.key /etc/fips/fips.pub
|
||||
```
|
||||
|
||||
3. Start the daemon. With `persistent: true`, the daemon generates a
|
||||
new keypair and writes new `fips.key` / `fips.pub`.
|
||||
|
||||
```sh
|
||||
sudo systemctl start fips
|
||||
cat /etc/fips/fips.pub # the new npub
|
||||
```
|
||||
|
||||
4. Update every downstream reference: peer configs that name this
|
||||
node by npub, `hosts` files, ACL allow-lists, Nostr adverts
|
||||
pinned by other operators.
|
||||
|
||||
There is no recovery from a lost `fips.key` — the npub is gone with
|
||||
the secret. Treat key rotation as a coordinated event; do not rotate
|
||||
production identities ad hoc.
|
||||
|
||||
## See also
|
||||
|
||||
- [../reference/configuration.md](../reference/configuration.md) —
|
||||
`node.identity.*` keys.
|
||||
- [../reference/cli-fipsctl.md](../reference/cli-fipsctl.md) —
|
||||
`fipsctl keygen`.
|
||||
- [../design/fips-architecture.md](../design/fips-architecture.md) —
|
||||
identity model, npub-to-NodeAddr derivation.
|
||||
195
docs/how-to/run-as-unprivileged-user.md
Normal file
@@ -0,0 +1,195 @@
|
||||
# Run the FIPS Daemon as an Unprivileged User
|
||||
|
||||
By default, the FIPS daemon runs as `root` — the shipped Debian
|
||||
systemd unit configures this, and no further setup is required.
|
||||
The trade-off is that the daemon has full root authority,
|
||||
including outside its actual network needs. Acceptable for
|
||||
single-purpose hosts; less desirable for shared hosts.
|
||||
|
||||
This guide covers the alternative: drop privileges and run the
|
||||
daemon under a dedicated unprivileged user account. The TUN
|
||||
device that the FIPS IPv6 adapter creates requires
|
||||
`CAP_NET_ADMIN` on Linux; the recipe below grants that privilege
|
||||
via a file capability on the binary, plus everything else the
|
||||
daemon needs to keep working without root: a service user
|
||||
account, file permissions on the config directory, and a systemd
|
||||
unit override to drop privileges.
|
||||
|
||||
For the design context (why the adapter needs a TUN, how the
|
||||
adapter integrates with the kernel routing table), see
|
||||
[../design/fips-ipv6-adapter.md](../design/fips-ipv6-adapter.md).
|
||||
|
||||
## Prerequisites
|
||||
|
||||
- FIPS package installed (the postinst already creates the `fips`
|
||||
system group used for control-socket access).
|
||||
- `setcap` available (`apt install libcap2-bin` on Debian/Ubuntu;
|
||||
it is a standard utility on most distributions).
|
||||
- Operator access to systemd unit overrides (`systemctl edit`).
|
||||
|
||||
## Step 1: Create a `fips` system user
|
||||
|
||||
The package creates a `fips` system *group* but no matching user.
|
||||
Add a system user that belongs to the `fips` group:
|
||||
|
||||
```sh
|
||||
sudo useradd --system --gid fips --no-create-home --shell /usr/sbin/nologin fips
|
||||
```
|
||||
|
||||
The user has no home directory and no login shell — this account
|
||||
exists only to run the daemon.
|
||||
|
||||
## Step 2: Grant `CAP_NET_ADMIN` to the binary
|
||||
|
||||
Apply the file capability so the daemon can create the TUN device
|
||||
without root authority:
|
||||
|
||||
```sh
|
||||
sudo setcap cap_net_admin+ep /usr/bin/fips
|
||||
```
|
||||
|
||||
Verify:
|
||||
|
||||
```sh
|
||||
getcap /usr/bin/fips
|
||||
# /usr/bin/fips cap_net_admin=ep
|
||||
```
|
||||
|
||||
The binary can now create TUN devices when run by any user.
|
||||
|
||||
**File-capability caveats:**
|
||||
|
||||
- The capability is attached to the binary file. **Re-applying
|
||||
the capability after every package upgrade is required**,
|
||||
because package upgrades replace the binary file and lose the
|
||||
cap. The systemd override in Step 4 includes an `ExecStartPre`
|
||||
line that automates this.
|
||||
- File capabilities are stripped when the binary is copied across
|
||||
most filesystems and when it is downloaded via web tooling. If
|
||||
you build from source and install manually, remember to
|
||||
re-`setcap` after each rebuild.
|
||||
- `LD_LIBRARY_PATH` and similar environment-driven loader
|
||||
controls are stripped at exec time when file capabilities are
|
||||
present; this is normally what you want, but development
|
||||
workflows that rely on custom library paths may be surprised.
|
||||
|
||||
## Step 3: Adjust config-file permissions
|
||||
|
||||
The shipped `/etc/fips/fips.yaml` is mode `0600` and owned by
|
||||
`root:root`. The daemon needs to read it and, if persistent
|
||||
identity is enabled, write `/etc/fips/fips.key` into the same
|
||||
directory.
|
||||
|
||||
```sh
|
||||
sudo chown -R fips:fips /etc/fips
|
||||
sudo chmod 0640 /etc/fips/fips.yaml
|
||||
```
|
||||
|
||||
If `node.identity.persistent: true` is set and `fips.key` does
|
||||
not exist yet, leave `/etc/fips` itself writable by the `fips`
|
||||
user so the daemon can create it on first start. After the key
|
||||
file exists, you can tighten further:
|
||||
|
||||
```sh
|
||||
sudo chmod 0600 /etc/fips/fips.key
|
||||
```
|
||||
|
||||
## Step 4: Drop privileges in the systemd unit
|
||||
|
||||
Create an override:
|
||||
|
||||
```sh
|
||||
sudo systemctl edit fips.service
|
||||
```
|
||||
|
||||
Add:
|
||||
|
||||
```ini
|
||||
[Service]
|
||||
User=fips
|
||||
Group=fips
|
||||
AmbientCapabilities=CAP_NET_ADMIN
|
||||
NoNewPrivileges=no
|
||||
ExecStartPre=/sbin/setcap cap_net_admin+ep /usr/bin/fips
|
||||
```
|
||||
|
||||
`User=` / `Group=` set the service identity.
|
||||
`AmbientCapabilities=` ensures the file capability granted in
|
||||
Step 2 actually carries into the daemon's process tree.
|
||||
`NoNewPrivileges=no` is required for file-capability execution
|
||||
to work — systemd defaults this to `yes` for hardened units,
|
||||
which would block the `setcap` from taking effect.
|
||||
`ExecStartPre=` re-applies the capability before each start,
|
||||
which makes the package-upgrade path self-heal.
|
||||
|
||||
The unit's `RuntimeDirectory=fips` directive already arranges
|
||||
for `/run/fips/` to be created with the right ownership at
|
||||
service start, now as `fips:fips 0750` instead of
|
||||
`root:fips 0750`.
|
||||
|
||||
Reload and restart:
|
||||
|
||||
```sh
|
||||
sudo systemctl daemon-reload
|
||||
sudo systemctl restart fips
|
||||
```
|
||||
|
||||
## Step 5: Verify
|
||||
|
||||
Confirm the daemon is running as `fips`:
|
||||
|
||||
```sh
|
||||
ps -eo user,cmd | grep '[/]usr/bin/fips'
|
||||
# fips /usr/bin/fips --config /etc/fips/fips.yaml
|
||||
```
|
||||
|
||||
Confirm the TUN device came up (the `setcap` worked):
|
||||
|
||||
```sh
|
||||
ip link show fips0
|
||||
# fips0: <POINTOPOINT,UP,...> mtu 1280 ...
|
||||
```
|
||||
|
||||
Confirm the control socket is bound and accessible to the `fips`
|
||||
group:
|
||||
|
||||
```sh
|
||||
ls -la /run/fips/control.sock
|
||||
# srwxrwx--- 1 fips fips ... /run/fips/control.sock
|
||||
```
|
||||
|
||||
Add yourself to the `fips` group so you can use `fipsctl` /
|
||||
`fipstop` without `sudo`:
|
||||
|
||||
```sh
|
||||
sudo usermod -aG fips $USER
|
||||
# log out and back in for the group change to take effect
|
||||
```
|
||||
|
||||
Then:
|
||||
|
||||
```sh
|
||||
fipsctl show status
|
||||
```
|
||||
|
||||
## Caveats
|
||||
|
||||
- **`fips-firewall.service` still runs as root.** Loading nftables
|
||||
rules into the kernel requires root regardless. The firewall
|
||||
unit is intentionally separate from the daemon unit.
|
||||
- **Bluetooth peers (`transports.ble.*`)** require additional
|
||||
privileges the `CAP_NET_ADMIN` setcap doesn't cover. If you use
|
||||
the BLE transport, you'll likely need to keep running as root
|
||||
or layer additional capability/D-Bus configuration; that path
|
||||
is not covered here.
|
||||
|
||||
## See also
|
||||
|
||||
- [persistent-identity.md](persistent-identity.md) — how the
|
||||
daemon manages `/etc/fips/fips.key`
|
||||
- [../design/fips-ipv6-adapter.md](../design/fips-ipv6-adapter.md)
|
||||
— IPv6 adapter design, TUN interface architecture
|
||||
- [../reference/security.md](../reference/security.md) —
|
||||
consolidated security surface
|
||||
- [../reference/configuration.md](../reference/configuration.md)
|
||||
— `tun.*` configuration block
|
||||
247
docs/how-to/set-up-80211s-mesh-backhaul.md
Normal file
@@ -0,0 +1,247 @@
|
||||
# Set Up an 802.11s Mesh Backhaul (OpenWrt)
|
||||
|
||||
Link FIPS routers over radio — no cables, no APs, no shared
|
||||
infrastructure — by running the Ethernet transport on an open 802.11s
|
||||
mesh interface. The radio layer provides nothing but L2 frames to
|
||||
direct neighbors; FIPS provides everything else: encryption and
|
||||
authentication (Noise IK), peer discovery (Ethernet beacons), and
|
||||
routing (the spanning tree).
|
||||
|
||||
For the transport design, see
|
||||
[../design/fips-transport-layer.md](../design/fips-transport-layer.md).
|
||||
For all `transports.ethernet.*` configuration keys, see
|
||||
[../reference/configuration.md](../reference/configuration.md).
|
||||
|
||||
## Why open, why forwarding off
|
||||
|
||||
Two deliberate choices distinguish this from a stock 802.11s setup:
|
||||
|
||||
- **`encryption none`** — the mesh is open on purpose. Every FIPS peer
|
||||
link is already authenticated and encrypted by the Noise IK
|
||||
handshake, so SAE at L2 would duplicate that work, add a shared
|
||||
credential to provision across routers, and (on ath10k) force the
|
||||
firmware into its slower raw Tx/Rx mode. A stranger can form an
|
||||
802.11s peering with your router, but their frames die at the FIPS
|
||||
handshake — the same security model as mDNS and BLE discovery, where
|
||||
the advert is only a hint and the handshake is the authentication.
|
||||
What you concede: L2 metadata (MAC addresses, frame sizes) is
|
||||
visible in the air, and a hostile radio can burn airtime — both true
|
||||
of any radio link regardless of L2 encryption.
|
||||
- **`mesh_fwding 0`** — disables 802.11s's own HWMP routing so each
|
||||
mesh link is a plain neighbor link. FIPS is the routing layer; two
|
||||
routing layers would fight, and broadcast discovery beacons would
|
||||
flood the whole mesh instead of reaching direct neighbors only.
|
||||
|
||||
The interface is **not** bridged into `br-lan` — the FIPS Ethernet
|
||||
transport binds it directly.
|
||||
|
||||
## When to use
|
||||
|
||||
- Two or more OpenWrt FIPS routers within radio range of each other,
|
||||
where running cable is impractical.
|
||||
- You want the mesh segment to keep working with zero shared
|
||||
credentials or per-site configuration ("flash and drop in").
|
||||
|
||||
It is **not** for connecting phones or laptops — client devices
|
||||
cannot join an 802.11s mesh. They enter the mesh through a normal AP
|
||||
on the same router (see constraints below), or over BLE.
|
||||
|
||||
## Requirements
|
||||
|
||||
- OpenWrt 22.03+ with the FIPS package installed.
|
||||
- A radio whose driver supports mesh point interfaces. Check with:
|
||||
|
||||
```sh
|
||||
iw list | grep -A 10 "Supported interface modes" | grep "mesh point"
|
||||
```
|
||||
|
||||
The mainstream OpenWrt chips (ath9k, ath10k, mt76) all qualify.
|
||||
- Ideally a dual- or tri-band router, so one band can be dedicated to
|
||||
the backhaul (see constraints).
|
||||
|
||||
## Step 1 — create the mesh interface(s)
|
||||
|
||||
On **each** router, run the helper once per radio you want in the
|
||||
backhaul:
|
||||
|
||||
```sh
|
||||
fips-mesh-setup radio1
|
||||
```
|
||||
|
||||
This creates an open 802.11s interface with mesh ID `fips-mesh` and
|
||||
HWMP forwarding off, attaches it to an unmanaged netifd interface (no
|
||||
IP configuration — none is needed), uncomments the matching `meshN`
|
||||
transport entry in `/etc/fips/fips.yaml` (see Step 2), and reloads the
|
||||
radio. Interfaces are named by radio index: `radio0` → `fips-mesh0`,
|
||||
`radio1` → `fips-mesh1`. Pass a second argument to use a different
|
||||
mesh ID.
|
||||
|
||||
Note: the helper runs `wifi reload`, which re-applies the whole
|
||||
wireless config and so briefly drops every client AP on all radios for
|
||||
a few seconds. `fips-mesh-setup remove` reloads the same way. Expect
|
||||
the blip if clients are connected.
|
||||
|
||||
On dual-band routers, meshing **both** bands is worth it: 2.4 GHz
|
||||
reaches further at lower rates, 5 GHz carries more over shorter
|
||||
links. Note this is **failover, not multipath**: FIPS keeps one
|
||||
active link per peer, so traffic uses one band at a time — the other
|
||||
is a standby that re-establishes the peer if the active link dies
|
||||
(detection via keepalive timeout, so a cutover takes seconds, not
|
||||
milliseconds):
|
||||
|
||||
```sh
|
||||
fips-mesh-setup radio0
|
||||
fips-mesh-setup radio1
|
||||
```
|
||||
|
||||
**Pin the same channel on every backhaul router, per band.** Mesh
|
||||
points only peer on the same channel, and the mesh inherits whatever
|
||||
the radio is set to — with `channel 'auto'` (the default on many
|
||||
devices) each router picks its own and the mesh silently never forms.
|
||||
The script prints the radio's current band and channel and warns on
|
||||
`auto`:
|
||||
|
||||
```sh
|
||||
uci set wireless.radio1.channel='36'
|
||||
uci commit wireless && wifi reload
|
||||
```
|
||||
|
||||
Prefer a non-DFS channel (36–48 on 5 GHz): on DFS channels the radio
|
||||
must wait ~60 s in CAC before transmitting after every reload.
|
||||
|
||||
Equivalent manual UCI (per radio), if you prefer to see what it does:
|
||||
|
||||
```sh
|
||||
uci batch <<'EOF'
|
||||
set wireless.fips_mesh_radio1=wifi-iface
|
||||
set wireless.fips_mesh_radio1.device='radio1'
|
||||
set wireless.fips_mesh_radio1.mode='mesh'
|
||||
set wireless.fips_mesh_radio1.mesh_id='fips-mesh'
|
||||
set wireless.fips_mesh_radio1.encryption='none'
|
||||
set wireless.fips_mesh_radio1.mesh_fwding='0'
|
||||
set wireless.fips_mesh_radio1.ifname='fips-mesh1'
|
||||
set wireless.fips_mesh_radio1.network='fips_mesh_radio1'
|
||||
set network.fips_mesh_radio1=interface
|
||||
set network.fips_mesh_radio1.proto='none'
|
||||
EOF
|
||||
uci commit
|
||||
wifi reload
|
||||
```
|
||||
|
||||
## Step 2 — check the FIPS transport binding
|
||||
|
||||
The `fips.yaml` shipped in the OpenWrt package carries one transport
|
||||
entry per radio, but **commented out** — so a stock install that never
|
||||
runs this helper logs no per-boot "interface missing" warning.
|
||||
`fips-mesh-setup` uncommented the matching `meshN` entry in Step 1, so
|
||||
there is normally nothing to do here. If you maintain your own config
|
||||
(or ran the manual UCI above instead of the helper), make sure the
|
||||
entries are present and uncommented:
|
||||
|
||||
```yaml
|
||||
transports:
|
||||
ethernet:
|
||||
mesh0:
|
||||
interface: "fips-mesh0"
|
||||
discovery: true
|
||||
announce: true
|
||||
auto_connect: true
|
||||
accept_connections: true
|
||||
mesh1:
|
||||
interface: "fips-mesh1"
|
||||
discovery: true
|
||||
announce: true
|
||||
auto_connect: true
|
||||
accept_connections: true
|
||||
```
|
||||
|
||||
## Step 3 — restart the daemon (order matters)
|
||||
|
||||
```sh
|
||||
/etc/init.d/fips restart
|
||||
```
|
||||
|
||||
Restart fips **after** the mesh interface is up. A transport whose
|
||||
interface is missing at startup is logged and skipped, not retried —
|
||||
so if the daemon comes up before the radio, the mesh transport stays
|
||||
dead until the next restart. (An interface that *vanishes and
|
||||
returns* after startup is recovered automatically; only the missing-
|
||||
at-startup case needs this ordering.)
|
||||
|
||||
## Verify
|
||||
|
||||
L2 first — the 802.11s peering, with a second configured router in
|
||||
range:
|
||||
|
||||
```sh
|
||||
iw dev fips-mesh0 station dump
|
||||
```
|
||||
|
||||
You should see one station entry per neighbor router, with signal
|
||||
levels. No entries means a radio problem, not a FIPS problem — triage
|
||||
in this order:
|
||||
|
||||
1. **Channel mismatch** (the most common cause): compare
|
||||
`iw dev fips-mesh0 info` on both routers — mesh ID *and* channel
|
||||
must match exactly.
|
||||
2. **The mesh interface never joined** — `iw dev fips-meshX info`
|
||||
shows `type mesh point` but **no channel line**, and `station dump`
|
||||
is empty. Usual cause: a client (`sta`) interface on the same
|
||||
radio. A STA must follow its upstream AP's channel, the whole
|
||||
radio follows the STA, and a mesh pinned to a different channel
|
||||
silently stays down. Check for a STA sharing the radio
|
||||
(`iw dev`, look for `type managed` on the same phy), compare
|
||||
`iw dev <sta-iface> info | grep channel`, and re-pin the mesh
|
||||
channel to match — on every backhaul router.
|
||||
3. **Is the other router transmitting at all?**
|
||||
|
||||
```sh
|
||||
iw dev fips-mesh0 scan | grep -i -B4 "MESH ID"
|
||||
```
|
||||
|
||||
Its mesh ID visible → transmission works, peering is failing
|
||||
(mesh ID typo, or one side has encryption set). Nothing visible →
|
||||
check `wifi status` on the other router, remember the ~60 s DFS
|
||||
CAC wait, and confirm the country code is set
|
||||
(`uci get wireless.radio1.country`) — an unset regdomain can
|
||||
block channels entirely.
|
||||
4. `logread | grep -iE "mesh|fips-mesh0"` on both sides.
|
||||
|
||||
Then the FIPS layer on top:
|
||||
|
||||
```sh
|
||||
logread | grep -i beacon # beacons flowing on the new transport
|
||||
fipsctl show peers # neighbor authenticated and connected
|
||||
fipsctl show links # link on the 'ethernet' transport
|
||||
```
|
||||
|
||||
Discovery is automatic: each node beacons its pubkey every few
|
||||
seconds, and `auto_connect` initiates the Noise handshake on first
|
||||
sight.
|
||||
|
||||
## Constraints
|
||||
|
||||
- **Airtime is shared per radio.** All virtual interfaces on one
|
||||
radio (AP + mesh) share one channel, and multi-hop forwarding on a
|
||||
single radio roughly halves throughput per hop. On dual/tri-band
|
||||
hardware, dedicate one band to `fips-mesh0` and serve clients on
|
||||
the others.
|
||||
- **AP + mesh coexistence is driver-dependent.** It works on the
|
||||
mainstream chips (this is the standard Freifunk/Gluon setup), but
|
||||
check `iw list` under "valid interface combinations" for your
|
||||
hardware.
|
||||
- **Clients can't join.** Phones and laptops reach the mesh through
|
||||
the router's normal AP or via BLE — never through the 802.11s
|
||||
interface.
|
||||
- **Radio links are lossy.** A neighbor at the edge of range will
|
||||
form an 802.11s peering yet deliver a fraction of its frames.
|
||||
Expect link-quality effects that don't exist on wired Ethernet.
|
||||
- **A client (STA) uplink on the same radio owns the channel.** The
|
||||
STA must follow whatever channel its upstream AP uses; every other
|
||||
interface on that radio follows the STA. A mesh pinned to a
|
||||
different channel silently never joins, and it does **not** recover
|
||||
when the STA disconnects — a `wifi reload` (plus a fips restart) is
|
||||
needed. A *roaming* uplink (travel-router / hotspot-chasing setups)
|
||||
is fundamentally incompatible with a fixed-channel mesh on the same
|
||||
radio: dedicate the mesh to the radio the STA never uses, and treat
|
||||
any mesh sharing a STA radio as best-effort.
|
||||
299
docs/how-to/set-up-bluetooth-peer.md
Normal file
@@ -0,0 +1,299 @@
|
||||
# Set Up a Bluetooth (BLE) Peer Link
|
||||
|
||||
FIPS supports Bluetooth Low Energy as a transport for short-range
|
||||
mesh extension — same room, same building, no IP infrastructure
|
||||
between the two endpoints. The BLE transport runs as L2CAP
|
||||
Connection-Oriented Channels on a configurable PSM and reports
|
||||
per-link MTU back to the mesh layer for path-MTU computation.
|
||||
|
||||
For the design rationale and per-link MTU model, see
|
||||
[../design/fips-transport-layer.md](../design/fips-transport-layer.md).
|
||||
For all `transports.ble.*` configuration keys, see
|
||||
[../reference/configuration.md](../reference/configuration.md).
|
||||
|
||||
> **Experimental.** The BLE transport works but is still maturing.
|
||||
> Expect rougher edges than UDP or TCP — particularly around link
|
||||
> stability under interference and MTU negotiation on older
|
||||
> controllers. Treat it as you would any experimental transport in a
|
||||
> production deployment.
|
||||
|
||||
## When to use
|
||||
|
||||
BLE is the right transport when:
|
||||
|
||||
- Two nodes are within roughly 10 metres line-of-sight (more with
|
||||
external antennas, less through walls).
|
||||
- You want a self-contained mesh segment with no shared WiFi or
|
||||
Ethernet between the participants.
|
||||
- You can work within practical L2CAP CoC throughput (1-2 Mbps in
|
||||
good conditions, often substantially less under interference or
|
||||
at range) and the higher latency variance compared to WiFi.
|
||||
|
||||
It is **not** the right transport for backbone links between rooms
|
||||
where WiFi or Ethernet exists, for high-throughput data, or for any
|
||||
deployment where range matters more than infrastructure-freedom.
|
||||
|
||||
## Platform support
|
||||
|
||||
The BLE transport is **Linux-only** in the current implementation.
|
||||
The runtime depends on BlueZ via the `bluer` crate, which in turn
|
||||
needs `glibc` (musl builds skip BLE; the build script gates the
|
||||
crate accordingly).
|
||||
|
||||
| Platform | BLE transport |
|
||||
| -------- | -------------- |
|
||||
| Linux (glibc) | Supported. |
|
||||
| Linux (musl, OpenWrt) | Disabled at build time. |
|
||||
| macOS | Not supported. |
|
||||
| Windows | Not supported. |
|
||||
|
||||
The Debian package `Recommends: bluez`; install it explicitly if you
|
||||
opted out:
|
||||
|
||||
```sh
|
||||
sudo apt install bluez
|
||||
```
|
||||
|
||||
## Prerequisites
|
||||
|
||||
Both endpoints need:
|
||||
|
||||
1. A BLE-capable HCI adapter visible to BlueZ. Confirm with:
|
||||
|
||||
```sh
|
||||
sudo bluetoothctl show
|
||||
```
|
||||
|
||||
Note the controller name (typically `hci0`).
|
||||
|
||||
2. The `bluetoothd` service running and the adapter powered on:
|
||||
|
||||
```sh
|
||||
sudo systemctl enable --now bluetooth
|
||||
sudo bluetoothctl power on
|
||||
```
|
||||
|
||||
3. Sufficient privileges for the FIPS daemon. There are two
|
||||
independent privilege concerns; the BLE-only deployment case
|
||||
(mesh router with `tun.enabled: false`) needs only the second.
|
||||
|
||||
- **TUN adapter (always required when `tun.enabled: true`).**
|
||||
The daemon needs `CAP_NET_ADMIN` to create and configure the
|
||||
TUN device. The shipped systemd unit handles this by running
|
||||
as root; if you prefer to drop privileges, see
|
||||
[run-as-unprivileged-user.md](run-as-unprivileged-user.md).
|
||||
|
||||
- **BLE access (required for this how-to).** BlueZ exposes
|
||||
L2CAP and D-Bus paths under either group membership or
|
||||
`CAP_NET_RAW`. Pick one:
|
||||
|
||||
- Run the daemon as root. The shipped systemd unit takes
|
||||
this route.
|
||||
- Run as an unprivileged user that is a member of the
|
||||
`bluetooth` group. No additional capability is needed for
|
||||
the BLE side.
|
||||
- Run as an unprivileged user with no group membership, and
|
||||
grant the binary `CAP_NET_RAW`:
|
||||
|
||||
```sh
|
||||
sudo setcap cap_net_raw+ep $(which fips)
|
||||
```
|
||||
|
||||
This bypasses BlueZ's polkit/group check by holding
|
||||
`CAP_NET_RAW` directly. If you also need `CAP_NET_ADMIN`
|
||||
for TUN, combine them:
|
||||
|
||||
```sh
|
||||
sudo setcap cap_net_admin,cap_net_raw+ep $(which fips)
|
||||
```
|
||||
|
||||
4. The same L2CAP PSM on both endpoints. The default is `0x0085`
|
||||
(133); override only if you need to coexist with another L2CAP
|
||||
service on that PSM.
|
||||
|
||||
## Configuration
|
||||
|
||||
Add a `ble` block under `transports` in `fips.yaml`. A minimum BLE-
|
||||
active node looks like this:
|
||||
|
||||
```yaml
|
||||
transports:
|
||||
ble:
|
||||
adapter: "hci0"
|
||||
advertise: true
|
||||
scan: true
|
||||
auto_connect: true
|
||||
accept_connections: true
|
||||
```
|
||||
|
||||
Note: `auto_connect: true` is intentionally non-default (the default
|
||||
is `false`). For a symmetric ground-up discovery flow where either
|
||||
side may dial, both ends must opt in explicitly.
|
||||
|
||||
| Key | Purpose |
|
||||
| --- | ------- |
|
||||
| `adapter` | HCI controller name. Default: `hci0`. |
|
||||
| `psm` | L2CAP PSM. Default: `0x0085` (must match on both ends). |
|
||||
| `mtu` | Default L2CAP CoC MTU. Default: `2048`. The kernel may negotiate lower per link. |
|
||||
| `max_connections` | Concurrent BLE connections. Default: `7` (Bluetooth controllers typically support up to ~7 simultaneous L2CAP CoCs). |
|
||||
| `advertise` | Broadcast our BLE adverts so other FIPS nodes discover us. Default: `true`. |
|
||||
| `scan` | Listen for other FIPS nodes' BLE adverts. Default: `true`. |
|
||||
| `auto_connect` | Initiate a BLE connection to discovered FIPS adverts. Default: `false`. |
|
||||
| `accept_connections` | Accept inbound L2CAP connections. Default: `true`. |
|
||||
| `connect_timeout_ms` | Outbound L2CAP connect timeout. Default: `10000`. |
|
||||
| `probe_cooldown_secs` | After probing a BD_ADDR (success or failure), wait this long before probing it again. Default: `30`. |
|
||||
|
||||
Two pairing patterns are common:
|
||||
|
||||
**Symmetric auto-discovery.** Both nodes advertise, scan, and
|
||||
auto-connect. Whichever side completes the L2CAP connection first
|
||||
wins; the other side aborts its in-flight attempt. This is the
|
||||
"toss two devices in the same room" setup.
|
||||
|
||||
```yaml
|
||||
# Both nodes
|
||||
transports:
|
||||
ble:
|
||||
adapter: "hci0"
|
||||
advertise: true
|
||||
scan: true
|
||||
auto_connect: true
|
||||
accept_connections: true
|
||||
```
|
||||
|
||||
**Asymmetric peripheral / central.** One node only listens
|
||||
(peripheral), the other actively dials (central). Useful when one
|
||||
endpoint is a dedicated bootstrap and the other is mobile.
|
||||
|
||||
```yaml
|
||||
# Listener
|
||||
transports:
|
||||
ble:
|
||||
adapter: "hci0"
|
||||
advertise: true
|
||||
scan: false
|
||||
auto_connect: false
|
||||
accept_connections: true
|
||||
```
|
||||
|
||||
```yaml
|
||||
# Dialer
|
||||
transports:
|
||||
ble:
|
||||
adapter: "hci0"
|
||||
advertise: false
|
||||
scan: true
|
||||
auto_connect: true
|
||||
accept_connections: false
|
||||
```
|
||||
|
||||
After editing, restart the daemon on each side:
|
||||
|
||||
```sh
|
||||
sudo systemctl restart fips
|
||||
```
|
||||
|
||||
## Verify
|
||||
|
||||
On each endpoint, confirm the transport came up:
|
||||
|
||||
```sh
|
||||
fipsctl show transports
|
||||
```
|
||||
|
||||
Look for an entry of type `ble` in the `state: Running` (or
|
||||
equivalent) state. The `mtu` field reports the configured default;
|
||||
per-link MTU is reported separately.
|
||||
|
||||
Confirm the link is established:
|
||||
|
||||
```sh
|
||||
fipsctl show peers
|
||||
```
|
||||
|
||||
The peer entry for the BLE-attached neighbour should report
|
||||
`transport_type: "ble"` and a non-zero `last_seen_ms`.
|
||||
|
||||
BLE peering is auto-discovery only: there is no `fipsctl connect`
|
||||
path for BLE (the command accepts `udp`, `tcp`, `tor`, and
|
||||
`ethernet` only). Links come up via advert/scan; if you don't see
|
||||
the peer here, the configuration above is the only knob.
|
||||
|
||||
To watch the link in real time, use `fipstop`'s **Peers** and
|
||||
**Transports** tabs:
|
||||
|
||||
```sh
|
||||
fipstop
|
||||
```
|
||||
|
||||
The Performance tab reports the per-link MMP metrics — SRTT, loss
|
||||
rate, ETX — which on BLE typically run an order of magnitude worse
|
||||
than over UDP, with much higher jitter.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
### Transport never comes up
|
||||
|
||||
Check the BlueZ side first:
|
||||
|
||||
```sh
|
||||
systemctl status bluetooth
|
||||
sudo bluetoothctl show
|
||||
```
|
||||
|
||||
If `bluetoothctl show` reports `Powered: no`, fix that before
|
||||
debugging FIPS. The FIPS daemon will log a warning if it cannot
|
||||
acquire the adapter.
|
||||
|
||||
If the FIPS log contains `bluer` D-Bus errors, the daemon usually
|
||||
lacks permission. Run as root or grant `CAP_NET_ADMIN` and add the
|
||||
fips user to the `bluetooth` group.
|
||||
|
||||
### Peers see each other but never connect
|
||||
|
||||
Verify `accept_connections` is true on at least one side and
|
||||
`auto_connect` is true on at least one side. Two listen-only nodes
|
||||
will discover each other but never establish an L2CAP connection.
|
||||
|
||||
Check `psm` matches on both ends. A mismatch presents as adverts
|
||||
visible (in `fipstop` discovery counters) but every connect attempt
|
||||
fails.
|
||||
|
||||
### Link comes up but throughput is poor
|
||||
|
||||
Practical L2CAP CoC throughput in good conditions reaches
|
||||
1-2 Mbps, but interference, range, and controller capability all
|
||||
push it lower. If throughput is well below that range, check the
|
||||
negotiated ATT_MTU — a small ATT_MTU (default 23 bytes when
|
||||
extended ATT MTU is not negotiated) caps per-PDU payload
|
||||
regardless of radio conditions. The per-link MTU reported in
|
||||
`fipsctl show transports` reveals what was negotiated.
|
||||
|
||||
If MTU is unexpectedly low, both endpoints must support and have
|
||||
negotiated the BlueZ L2CAP `cocmode=2` extension. Older Bluetooth
|
||||
controllers cap MTU regardless.
|
||||
|
||||
### Unstable links / repeated reconnects
|
||||
|
||||
Bluetooth in busy 2.4 GHz environments suffers from WiFi
|
||||
interference. Switch the adapter to a less crowded channel (kernel
|
||||
side, not configurable from FIPS) or add an external antenna. The
|
||||
`probe_cooldown_secs` tunable backs off retry attempts; raise it if
|
||||
the daemon log shows many short-lived probes.
|
||||
|
||||
### Permission errors on socket open
|
||||
|
||||
Most modern systemd installs do not allow non-root processes to
|
||||
open raw L2CAP sockets without an explicit policy. Run the daemon
|
||||
as root (the shipped systemd unit does this) or add a `polkit`
|
||||
rule for the `bluetooth` group.
|
||||
|
||||
## See also
|
||||
|
||||
- [../design/fips-transport-layer.md](../design/fips-transport-layer.md)
|
||||
— per-transport MTU reporting and the BLE row of the supported-
|
||||
transports table.
|
||||
- [../reference/configuration.md](../reference/configuration.md) —
|
||||
full `transports.ble.*` reference.
|
||||
- [run-as-unprivileged-user.md](run-as-unprivileged-user.md) —
|
||||
adjacent privilege handling for the daemon process.
|
||||
453
docs/how-to/troubleshoot-gateway.md
Normal file
@@ -0,0 +1,453 @@
|
||||
# Troubleshoot `fips-gateway`
|
||||
|
||||
Diagnostic recipes for `fips-gateway`, grouped by which half of the
|
||||
gateway is failing. For gateway design and deployment, see
|
||||
[../design/fips-gateway.md](../design/fips-gateway.md) and
|
||||
[deploy-gateway.md](deploy-gateway.md). For OpenWrt-specific
|
||||
deployment problems, see the
|
||||
[OpenWrt deployment tutorial](../tutorials/deploy-fips-gateway.md);
|
||||
most of the recipes below apply on OpenWrt as well, but paths and
|
||||
service names differ.
|
||||
|
||||
## Inspect gateway state via the control socket
|
||||
|
||||
Before digging into nftables or conntrack, ask the gateway directly
|
||||
whether it has the mapping or session you expect. `fips-gateway`
|
||||
exposes a separate control socket (`/run/fips/gateway.sock`) with its
|
||||
own command set; there is no `fipsctl gateway` subcommand — talk to
|
||||
the socket directly with `nc -U`. Each request is a single line of
|
||||
JSON terminated with a newline; the connection is closed after one
|
||||
response.
|
||||
|
||||
Pool summary, listen address, NAT counters, uptime, and the loaded
|
||||
config snapshot:
|
||||
|
||||
```sh
|
||||
echo '{"command":"show_gateway"}' | sudo nc -U /run/fips/gateway.sock
|
||||
```
|
||||
|
||||
Per-mapping virtual-IP state (allocated, active, draining):
|
||||
|
||||
```sh
|
||||
echo '{"command":"show_mappings"}' | sudo nc -U /run/fips/gateway.sock
|
||||
```
|
||||
|
||||
If either command returns `gateway not yet initialized`, the gateway
|
||||
is still in early startup; wait a moment and retry. If a mapping you
|
||||
expect is not in the list, the DNS path didn't allocate it — fall
|
||||
through to the outbound DNS recipes below. If the mapping exists in
|
||||
`state: Active` but mesh traffic still fails, the problem is
|
||||
downstream of the allocation (firewall, route, masquerade); see the
|
||||
recipes that follow.
|
||||
|
||||
For the full command catalog and JSON shapes, see
|
||||
[../reference/control-socket.md#gateway-command-catalog](../reference/control-socket.md#gateway-command-catalog).
|
||||
|
||||
## Common (either-half) issues
|
||||
|
||||
These break both halves at once because they affect the gateway
|
||||
process itself or the shared NAT machinery.
|
||||
|
||||
### "No gateway section in configuration"
|
||||
|
||||
`fips-gateway` is normally launched by the systemd unit shipped with
|
||||
the package:
|
||||
|
||||
```sh
|
||||
sudo systemctl restart fips-gateway
|
||||
sudo journalctl -u fips-gateway -e
|
||||
```
|
||||
|
||||
The unit reads the standard FIPS config search paths (typically
|
||||
`/etc/fips/fips.yaml`). If the unit logs "no gateway section in
|
||||
configuration" or "Gateway section exists but is not enabled", confirm
|
||||
the section is present and `enabled: true`:
|
||||
|
||||
```sh
|
||||
grep -A1 '^gateway:' /etc/fips/fips.yaml
|
||||
```
|
||||
|
||||
For one-off debugging outside systemd, run the binary directly and
|
||||
point it at a specific config file:
|
||||
|
||||
```sh
|
||||
sudo fips-gateway --config /etc/fips/fips.yaml --log-level debug
|
||||
```
|
||||
|
||||
This is useful to capture stderr in a terminal, but the systemd unit
|
||||
is the supported entry point in production. See
|
||||
[../reference/cli-fips-gateway.md](../reference/cli-fips-gateway.md)
|
||||
for the full flag list.
|
||||
|
||||
### Port conflict on the DNS listen port
|
||||
|
||||
Symptom: gateway fails to start with "address already in use" on
|
||||
the configured `gateway.dns.listen` address.
|
||||
|
||||
The default `[::1]:5353` is loopback-only on an unprivileged port and
|
||||
should not collide with any standard resolver. If you have overridden
|
||||
`dns.listen` to bind port 53 (or a LAN-side address) and another DNS
|
||||
server (systemd-resolved, dnsmasq, BIND) is already bound there,
|
||||
identify it:
|
||||
|
||||
```sh
|
||||
sudo ss -tulnp | grep ':53'
|
||||
```
|
||||
|
||||
Two options:
|
||||
|
||||
- **Stay on the loopback default.** Drop the override and let the
|
||||
gateway use `[::1]:5353`. Configure the existing resolver to
|
||||
forward `.fips` queries to it (the canonical OpenWrt deployment
|
||||
works this way out of the box).
|
||||
|
||||
- **Relocate the conflicting resolver.** Move it to a different port
|
||||
(or disable it if not needed) and let the gateway bind 53.
|
||||
Practical for systemd-resolved (set `DNSStubListener=no` in
|
||||
`/etc/systemd/resolved.conf`); rarely worth it for production
|
||||
resolvers.
|
||||
|
||||
### IPv6 forwarding disabled
|
||||
|
||||
Symptom: gateway exits at startup with
|
||||
"IPv6 forwarding is disabled. Enable with: sysctl -w
|
||||
net.ipv6.conf.all.forwarding=1".
|
||||
|
||||
The gateway is completely non-functional without forwarding — packets
|
||||
cannot traverse the NAT pipeline. Enable it:
|
||||
|
||||
```sh
|
||||
sudo sysctl -w net.ipv6.conf.all.forwarding=1
|
||||
```
|
||||
|
||||
Persist via the drop-in shown in
|
||||
[deploy-gateway.md](deploy-gateway.md#kernel-sysctls). The same
|
||||
section lists `proxy_ndp`, which is also required for the outbound
|
||||
half.
|
||||
|
||||
### nftables table missing or not loaded
|
||||
|
||||
Symptom: `show_gateway` reports an active gateway but
|
||||
`nft list table inet fips_gateway` errors with "No such file or
|
||||
directory".
|
||||
|
||||
The table is created by the gateway at startup and rebuilt atomically
|
||||
on every mapping change and on every `set_port_forwards` call. If the
|
||||
table is missing while the gateway claims to be running, something
|
||||
else (a host firewall script, a `nft flush ruleset` from another
|
||||
service) deleted it after creation. Restart the gateway to recreate
|
||||
it:
|
||||
|
||||
```sh
|
||||
sudo systemctl restart fips-gateway
|
||||
```
|
||||
|
||||
If a peer service is repeatedly clobbering the table, switch that
|
||||
service to use `add table` / `flush table <name>` for its own table
|
||||
rather than `flush ruleset`, which destroys every table on the host.
|
||||
|
||||
### Control socket permission errors
|
||||
|
||||
Symptom: `nc -U /run/fips/gateway.sock` fails with "Permission
|
||||
denied" or "No such file or directory".
|
||||
|
||||
The socket is owned by root with mode `0770` (group `fips`). Either
|
||||
run `nc` as root (`sudo nc -U ...`) or add your user to the `fips`
|
||||
group and re-login. If the file does not exist at all, the gateway
|
||||
either failed to start (check `journalctl -u fips-gateway`) or
|
||||
failed to bind the socket and continued without it (the warning
|
||||
`Failed to bind gateway control socket — continuing without it` is
|
||||
in the journal in that case).
|
||||
|
||||
## Outbound-half diagnostics
|
||||
|
||||
Symptoms in this section all involve a LAN client trying to reach a
|
||||
mesh destination through the gateway and failing.
|
||||
|
||||
### DNS queries fail
|
||||
|
||||
Symptom: LAN clients get `SERVFAIL` or no response when querying
|
||||
`.fips` names; or the gateway log shows DNS upstream timeouts.
|
||||
|
||||
**Step 1.** Verify the daemon resolver is running and reachable from
|
||||
the gateway host:
|
||||
|
||||
```sh
|
||||
dig @::1 -p 5354 hostname.fips AAAA
|
||||
```
|
||||
|
||||
If this returns no answer or fails, the FIPS daemon's DNS resolver is
|
||||
not running or not enabled. Check that the daemon config has
|
||||
`dns.enabled: true` (the default) and the daemon is healthy:
|
||||
`fipsctl show status`.
|
||||
|
||||
**Step 2.** Verify the gateway is listening on its DNS port:
|
||||
|
||||
```sh
|
||||
sudo ss -tulnp | grep -E ':(53|5353)\b'
|
||||
```
|
||||
|
||||
If nothing is listening on the configured `dns.listen` address, the
|
||||
gateway either failed to start or is bound to a different address.
|
||||
Check the gateway log: `sudo journalctl -u fips-gateway -e`.
|
||||
|
||||
**Step 3.** Verify the LAN client can reach the gateway's DNS port:
|
||||
|
||||
```sh
|
||||
# from the LAN client
|
||||
dig @<gateway-lan-addr> hostname.fips AAAA
|
||||
```
|
||||
|
||||
If this hangs, the LAN-side firewall is blocking DNS, or the LAN
|
||||
route to the gateway is missing.
|
||||
|
||||
### Ping works but TCP does not
|
||||
|
||||
Symptom: `ping6 <virtual-ip>` succeeds from a LAN client, but TCP
|
||||
connections (SSH, HTTP) hang or time out.
|
||||
|
||||
This usually means the `fips0`-side masquerade rule is missing or
|
||||
misconfigured. Inspect the gateway's nftables table:
|
||||
|
||||
```sh
|
||||
sudo nft list table inet fips_gateway
|
||||
```
|
||||
|
||||
In the `postrouting` chain, look for a rule matching
|
||||
`oifname "fips0"` with a `masquerade` verdict. Without masquerade,
|
||||
the destination mesh node sees a source address (from the virtual
|
||||
pool) it cannot route replies to, and return packets are
|
||||
black-holed.
|
||||
|
||||
If the rule is missing, restart the gateway — the table is rebuilt
|
||||
atomically on every mapping change and on startup.
|
||||
|
||||
### Connection timeout to a virtual IP
|
||||
|
||||
Symptom: any traffic to a virtual pool address times out, including
|
||||
ping.
|
||||
|
||||
**Step 1.** Verify IPv6 forwarding is still enabled:
|
||||
|
||||
```sh
|
||||
sysctl net.ipv6.conf.all.forwarding
|
||||
# Expect: net.ipv6.conf.all.forwarding = 1
|
||||
```
|
||||
|
||||
**Step 2.** Verify the pool route exists:
|
||||
|
||||
```sh
|
||||
ip -6 route show table local | grep <pool-cidr>
|
||||
```
|
||||
|
||||
If the route is missing, the kernel does not recognize pool
|
||||
addresses as locally-owned and drops the packets before NAT can
|
||||
process them. The gateway adds this route at startup; if it's
|
||||
missing, check the gateway log for startup errors.
|
||||
|
||||
**Step 3.** Verify the destination mesh address actually exists in
|
||||
the FIPS daemon's identity cache:
|
||||
|
||||
```sh
|
||||
fipsctl show identity-cache | grep <fd00-mesh-addr>
|
||||
```
|
||||
|
||||
If the entry is missing, the DNS-side mapping never primed the
|
||||
identity cache, which means the daemon resolver did not actually
|
||||
resolve the `.fips` name. Re-test the DNS path:
|
||||
|
||||
```sh
|
||||
dig @::1 -p 5354 hostname.fips AAAA
|
||||
```
|
||||
|
||||
### Virtual IP unreachable from a LAN client
|
||||
|
||||
Symptom: client cannot reach the virtual IP at all (no ping, no
|
||||
ARP/ND response).
|
||||
|
||||
**Step 1.** Verify the client has a route to the pool via the
|
||||
gateway:
|
||||
|
||||
```sh
|
||||
# from the LAN client
|
||||
ip -6 route get <virtual-ip>
|
||||
```
|
||||
|
||||
The output should show the gateway as the next-hop. If it shows
|
||||
something else (or "unreachable"), fix the LAN-side route — see
|
||||
[deploy-gateway.md](deploy-gateway.md#distribute-the-route-to-lan-clients).
|
||||
|
||||
**Step 2.** On the gateway, verify proxy NDP entries exist for
|
||||
allocated virtual IPs:
|
||||
|
||||
```sh
|
||||
ip -6 neigh show proxy
|
||||
```
|
||||
|
||||
If proxy NDP entries are missing, the gateway cannot answer Neighbor
|
||||
Solicitation requests for virtual IPs on the LAN, so clients cannot
|
||||
resolve the link-layer address and packets never leave the client's
|
||||
NIC.
|
||||
|
||||
The gateway adds these entries when a mapping is created (i.e., when
|
||||
a `.fips` DNS query allocates a virtual IP). If they're absent,
|
||||
trigger a DNS query first:
|
||||
|
||||
```sh
|
||||
dig @<gateway-lan-addr> hostname.fips AAAA
|
||||
```
|
||||
|
||||
Then re-check `ip -6 neigh show proxy`.
|
||||
|
||||
**Step 3.** Verify `proxy_ndp` is enabled in the kernel:
|
||||
|
||||
```sh
|
||||
sysctl net.ipv6.conf.all.proxy_ndp
|
||||
# Expect: net.ipv6.conf.all.proxy_ndp = 1
|
||||
```
|
||||
|
||||
If 0, enable it (see
|
||||
[deploy-gateway.md](deploy-gateway.md#kernel-sysctls)).
|
||||
|
||||
## Inbound-half diagnostics
|
||||
|
||||
Symptoms in this section all involve a mesh peer trying to reach a
|
||||
LAN-side service through the gateway and failing.
|
||||
|
||||
### Mesh peer can't reach `<gateway-npub>.fips:<listen_port>`
|
||||
|
||||
Walk the path from the mesh-side ingress to the LAN target:
|
||||
|
||||
**Step 1.** Verify the port-forward rule is loaded. On the gateway:
|
||||
|
||||
```sh
|
||||
sudo nft list table inet fips_gateway
|
||||
```
|
||||
|
||||
Look in the `prerouting` chain for a rule of the form
|
||||
|
||||
```text
|
||||
iif "fips0" meta nfproto ipv6 meta l4proto <tcp|udp> \
|
||||
<th> dport <listen_port> dnat ip6 to [<target_addr>]:<target_port>
|
||||
```
|
||||
|
||||
and, in the `postrouting` chain, a rule of the form
|
||||
|
||||
```text
|
||||
iif "fips0" oif "<lan_interface>" meta nfproto ipv6 masquerade
|
||||
```
|
||||
|
||||
The port-forward DNAT and the LAN-side masquerade come from
|
||||
`gateway.port_forwards[]` and the active `lan_interface` setting.
|
||||
The masquerade is emitted only when at least one port-forward exists.
|
||||
If either rule is missing, restart the gateway — the table is rebuilt
|
||||
atomically on config load.
|
||||
|
||||
**Step 2.** Verify the mesh firewall is not blocking the listen
|
||||
port. If `fips-firewall.service` is enabled, the default baseline
|
||||
drops everything inbound on `fips0` except established/related and
|
||||
ICMPv6. Add an explicit allow rule under `/etc/fips/fips.d/`:
|
||||
|
||||
```nft
|
||||
# /etc/fips/fips.d/gateway-inbound.nft
|
||||
tcp dport <listen_port> accept
|
||||
```
|
||||
|
||||
(See [enable-mesh-firewall.md](enable-mesh-firewall.md) for the
|
||||
full drop-in pattern, including source-address restrictions.)
|
||||
Without an allow rule, mesh peers see TCP RSTs (the firewall drops
|
||||
on the way in) or silent UDP loss.
|
||||
|
||||
**Step 3.** Verify the LAN target is reachable from the gateway
|
||||
itself:
|
||||
|
||||
```sh
|
||||
ping6 <target_addr>
|
||||
curl -v http://[<target_addr>]:<target_port>/ # for TCP HTTP
|
||||
```
|
||||
|
||||
If the target is unreachable from the gateway, the DNAT rule will
|
||||
fire but the inner connection attempt will fail. Fix LAN-side
|
||||
routing or the target service before going further.
|
||||
|
||||
**Step 4.** Verify conntrack is tracking the inbound flow. Try the
|
||||
connection from a mesh peer once:
|
||||
|
||||
```sh
|
||||
curl -v http://<gateway-npub>.fips:<listen_port>/
|
||||
```
|
||||
|
||||
Then on the gateway:
|
||||
|
||||
```sh
|
||||
sudo conntrack -L | grep -E '<listen_port>|<target_port>'
|
||||
```
|
||||
|
||||
You should see a flow tuple in both directions (orig and reply) with
|
||||
the mesh peer's source on `fips0` and the gateway's LAN address as
|
||||
the masqueraded source on the LAN side. No conntrack entry suggests
|
||||
the prerouting DNAT didn't match — recheck step 1.
|
||||
|
||||
**Step 5.** Check the gateway log for nftables or rule install
|
||||
errors:
|
||||
|
||||
```sh
|
||||
sudo journalctl -u fips-gateway -e | grep -E 'port_forward|nftables'
|
||||
```
|
||||
|
||||
A "Failed to install port-forward rules" log line at startup means
|
||||
the rule batch was rejected by netlink — usually a transient
|
||||
condition during a config edit, but persistent failures warrant
|
||||
inspecting the rule with `nft -d`.
|
||||
|
||||
### Config rejected: IPv4 target
|
||||
|
||||
Symptom: `fips-gateway` exits at startup with a deserialization
|
||||
error referencing `port_forwards[N].target` and an invalid IPv6
|
||||
literal.
|
||||
|
||||
The `target` field is typed as `SocketAddrV6` and rejects IPv4
|
||||
literals at parse time:
|
||||
|
||||
```yaml
|
||||
# fails at config load
|
||||
- listen_port: 8080
|
||||
proto: tcp
|
||||
target: "192.168.1.10:80"
|
||||
```
|
||||
|
||||
Either re-address the LAN service to be reachable on IPv6, or front
|
||||
it with a small IPv6-aware reverse proxy on the gateway and point
|
||||
the `target` at that proxy.
|
||||
|
||||
### Config rejected: zero or duplicate listen_port
|
||||
|
||||
Symptom: `fips-gateway` exits at startup with
|
||||
"Invalid gateway.port_forwards: …".
|
||||
|
||||
`validate_port_forwards()` enforces:
|
||||
|
||||
- `listen_port` must be non-zero.
|
||||
- The pair `(listen_port, proto)` must be unique across the list
|
||||
(the same port on TCP and UDP simultaneously is allowed; the same
|
||||
port twice on the same proto is not).
|
||||
|
||||
Fix the offending entry and reload.
|
||||
|
||||
## See also
|
||||
|
||||
- [../tutorials/deploy-fips-gateway.md](../tutorials/deploy-fips-gateway.md) —
|
||||
canonical OpenWrt deployment.
|
||||
- [../design/fips-gateway.md](../design/fips-gateway.md) — gateway
|
||||
design, NAT pipeline, virtual IP pool lifecycle, security
|
||||
considerations.
|
||||
- [deploy-gateway.md](deploy-gateway.md) — manual Linux-host setup.
|
||||
- [Gateway section](../reference/configuration.md#gateway-gateway) of
|
||||
the configuration reference — full `gateway.*` block.
|
||||
- [../reference/cli-fips-gateway.md](../reference/cli-fips-gateway.md) —
|
||||
`fips-gateway` binary CLI flags.
|
||||
- [Gateway command catalog](../reference/control-socket.md#gateway-command-catalog)
|
||||
in the control-socket reference — JSON schema for `show_gateway`
|
||||
and `show_mappings`.
|
||||
- [enable-mesh-firewall.md](enable-mesh-firewall.md) — mesh-firewall
|
||||
baseline and drop-ins.
|
||||
144
docs/how-to/tune-file-descriptors.md
Normal file
@@ -0,0 +1,144 @@
|
||||
# Tune the File-Descriptor Limit for FIPS
|
||||
|
||||
A busy FIPS node opens many file descriptors, and the count grows with
|
||||
the number of peers it serves. On most systemd distributions the daemon
|
||||
inherits a soft `RLIMIT_NOFILE` of 1024, which a well-connected node can
|
||||
exhaust — at which point peer admission, handshakes, and discovery start
|
||||
failing with `EMFILE` ("Too many open files").
|
||||
|
||||
This guide explains the FD budget, shows how to raise the limit on
|
||||
systemd and on OpenWrt, and how to verify the result.
|
||||
|
||||
## Why FIPS is FD-hungry
|
||||
|
||||
Unlike a service that multiplexes all traffic over one socket, the FIPS
|
||||
data plane allocates descriptors **per peer**. The dominant term is:
|
||||
|
||||
```text
|
||||
fds ≈ 3·P + fixed overhead (~30)
|
||||
```
|
||||
|
||||
where `P` is the number of established UDP peers. Each such peer consumes
|
||||
**3 file descriptors**:
|
||||
|
||||
- one `connect()`-ed UDP socket dedicated to that peer, plus
|
||||
- a 2-FD self-pipe owned by that peer's receive-drain worker (used to
|
||||
wake and stop the worker cleanly).
|
||||
|
||||
The remaining consumers are bounded and do not scale with peer count:
|
||||
|
||||
- the TUN device (one descriptor, process-lifetime),
|
||||
- the wildcard UDP listen socket(s) (one per bound UDP transport),
|
||||
- TCP and Tor transport listeners and the Tor control socket,
|
||||
- Nostr relay websockets (one per configured discovery relay),
|
||||
- the control socket (one `UnixListener`, plus short-lived per-request
|
||||
client connections for `fipsctl` / `fipstop`),
|
||||
- and base runtime descriptors (epoll, eventfd, logs).
|
||||
|
||||
Together these add a roughly flat overhead of about 30 descriptors. The
|
||||
per-peer term is what drives the daemon toward the FD ceiling.
|
||||
|
||||
## The symptom
|
||||
|
||||
The systemd and distro default **soft** `RLIMIT_NOFILE` is **1024**.
|
||||
With the `3·P` budget above, that ceiling is reached near **~320 peers**
|
||||
(3 × 320 ≈ 960, plus the fixed overhead). Once the process is out of
|
||||
descriptors, every syscall that allocates one — `socket()`, `accept()`,
|
||||
`open()`, `pipe()` — fails with `EMFILE`, which surfaces as:
|
||||
|
||||
- failed peer admission (new peers cannot be accepted),
|
||||
- failed handshakes (the daemon cannot open the per-peer socket), and
|
||||
- dropped discovery (relay or probe sockets cannot be created).
|
||||
|
||||
These symptoms appear only under load, once the node has accumulated
|
||||
enough peers to cross the ceiling, so they can be easy to misattribute.
|
||||
|
||||
## Raise the limit on systemd
|
||||
|
||||
Create a drop-in override for the service:
|
||||
|
||||
```sh
|
||||
sudo systemctl edit fips.service
|
||||
```
|
||||
|
||||
Add:
|
||||
|
||||
```ini
|
||||
[Service]
|
||||
LimitNOFILE=65535
|
||||
```
|
||||
|
||||
A single `LimitNOFILE=` value sets **both** the soft and the hard limit,
|
||||
so no separate soft/hard syntax is needed here.
|
||||
|
||||
Reload systemd and restart the daemon so the new limit takes effect:
|
||||
|
||||
```sh
|
||||
sudo systemctl daemon-reload
|
||||
sudo systemctl restart fips
|
||||
```
|
||||
|
||||
`65535` (2¹⁶ − 1) is the conventional headroom value for network
|
||||
daemons. With the `3·P` budget, it supports roughly **~21,800 peers**
|
||||
before the FD ceiling binds — well beyond any plausible single-node FIPS
|
||||
mesh degree. Past that point other limits (threads, memory, CPU) bind
|
||||
first, so raising `LimitNOFILE` higher buys nothing.
|
||||
|
||||
## Raise the limit on OpenWrt
|
||||
|
||||
OpenWrt uses procd, not systemd, so `LimitNOFILE` does not apply.
|
||||
Set the equivalent limit in the init script at `/etc/init.d/fips`,
|
||||
inside the block that starts the service:
|
||||
|
||||
```sh
|
||||
procd_set_param limits nofile="65535 65535"
|
||||
```
|
||||
|
||||
The two values are the soft and hard limits respectively; setting them
|
||||
equal mirrors the single-value systemd behaviour above.
|
||||
|
||||
Restart the service to apply:
|
||||
|
||||
```sh
|
||||
/etc/init.d/fips restart
|
||||
```
|
||||
|
||||
## Verify
|
||||
|
||||
Compare the live descriptor count against the established peer count:
|
||||
|
||||
```sh
|
||||
ls /proc/$(pidof fips)/fd | wc -l
|
||||
fipsctl show peers | wc -l
|
||||
```
|
||||
|
||||
At steady state, expect a stable ratio of about **3 descriptors per
|
||||
peer** plus the flat ~30-descriptor overhead. A ratio that holds steady
|
||||
as peers come and go confirms healthy, bounded scaling.
|
||||
|
||||
If the descriptor count climbs steadily while the peer count stays flat,
|
||||
that would indicate a descriptor leak rather than legitimate scaling —
|
||||
the limit bump would only delay the wall. The current data plane has
|
||||
been audited as leak-free (every per-peer descriptor has a guaranteed
|
||||
close on every teardown path), so a climbing ratio at fixed peer count
|
||||
would be a regression worth investigating.
|
||||
|
||||
## A note on deployment lines
|
||||
|
||||
The per-peer connected-UDP socket — the amplifier behind the
|
||||
`3·P` term — is present on the master and next data planes. It is **not
|
||||
yet present on the maintenance line**. On maintenance-only deployments
|
||||
the 3-descriptor-per-peer term does not apply, and FD pressure comes
|
||||
only from the fixed consumers listed above. Raising `LimitNOFILE` there
|
||||
is still worthwhile as forward-looking headroom, and harmless where the
|
||||
amplifier is absent.
|
||||
|
||||
## See also
|
||||
|
||||
- [tune-udp-buffers.md](tune-udp-buffers.md) — host sysctls so FIPS UDP
|
||||
sockets don't get clamped
|
||||
- [run-as-unprivileged-user.md](run-as-unprivileged-user.md) — run the
|
||||
daemon under a dedicated service account
|
||||
- [../reference/configuration.md](../reference/configuration.md) —
|
||||
transport and discovery configuration that influences the fixed FD
|
||||
overhead
|
||||
125
docs/how-to/tune-udp-buffers.md
Normal file
@@ -0,0 +1,125 @@
|
||||
# Tune Host UDP Socket Buffers for FIPS
|
||||
|
||||
The FIPS UDP transport requests larger send and receive socket
|
||||
buffers (default 2 MB each, doubled by the kernel to 4 MB actual)
|
||||
than the Linux defaults provide. The kernel silently clamps the
|
||||
request to `net.core.rmem_max` and `net.core.wmem_max` if those
|
||||
sysctls are smaller than the requested size — which causes silent
|
||||
packet drops under high throughput. For the design context (why FIPS
|
||||
requests larger buffers and how `SO_RXQ_OVFL` feeds ECN congestion
|
||||
detection), see
|
||||
[../design/fips-transport-layer.md](../design/fips-transport-layer.md#socket-buffer-sizing).
|
||||
|
||||
This guide covers the host-side sysctl setup needed before deploying
|
||||
a high-throughput FIPS node.
|
||||
|
||||
## Why this matters
|
||||
|
||||
The default Linux UDP receive buffer (`net.core.rmem_default`,
|
||||
typically 212 KB) fills in roughly 2.5 ms at ~85 MB/s. Any stall in
|
||||
the FIPS receive loop (decryption, routing, forwarding) causes the
|
||||
kernel to drop incoming datagrams without notification — they don't
|
||||
appear in `recv` errors, they don't trigger any application-visible
|
||||
event. The drops show up only in `SO_RXQ_OVFL` on subsequent
|
||||
packets, where FIPS surfaces them as congestion-detection events.
|
||||
|
||||
Setting `rmem_max` and `wmem_max` to at least the requested buffer
|
||||
size prevents the kernel clamp and the silent drop loss it causes.
|
||||
|
||||
## Step 1: Check current limits
|
||||
|
||||
```sh
|
||||
sysctl net.core.rmem_max net.core.wmem_max
|
||||
```
|
||||
|
||||
Typical defaults on stock Linux distributions are 212992 bytes
|
||||
(212 KB). FIPS requests 2 MB by default, which the kernel doubles
|
||||
internally to 4 MB; for the request to succeed without clamping, both
|
||||
sysctls must be at least 4194304 (4 MB).
|
||||
|
||||
## Step 2: Set the limits temporarily
|
||||
|
||||
```sh
|
||||
sudo sysctl -w net.core.rmem_max=4194304
|
||||
sudo sysctl -w net.core.wmem_max=4194304
|
||||
```
|
||||
|
||||
Verify:
|
||||
|
||||
```sh
|
||||
sysctl net.core.rmem_max net.core.wmem_max
|
||||
```
|
||||
|
||||
These changes take effect immediately for new socket binds but do
|
||||
not survive a reboot.
|
||||
|
||||
## Step 3: Make the limits persistent
|
||||
|
||||
Drop a file under `/etc/sysctl.d/`:
|
||||
|
||||
```sh
|
||||
sudo tee /etc/sysctl.d/60-fips.conf <<'EOF'
|
||||
# FIPS UDP transport requests 2 MB socket buffers, kernel doubles to 4 MB.
|
||||
# Avoid silent receive-buffer drops under load.
|
||||
net.core.rmem_max = 4194304
|
||||
net.core.wmem_max = 4194304
|
||||
EOF
|
||||
```
|
||||
|
||||
Apply:
|
||||
|
||||
```sh
|
||||
sudo sysctl --system
|
||||
```
|
||||
|
||||
The drop-in is loaded automatically on every boot.
|
||||
|
||||
## Step 4: Restart FIPS and verify the actual buffer size
|
||||
|
||||
After raising the host limits, restart the FIPS daemon so the next
|
||||
socket bind picks up the new ceiling:
|
||||
|
||||
```sh
|
||||
sudo systemctl restart fips
|
||||
```
|
||||
|
||||
The daemon logs the actual buffer sizes at startup:
|
||||
|
||||
```text
|
||||
UDP transport started local_addr=0.0.0.0:2121 recv_buf=4194304 send_buf=4194304
|
||||
```
|
||||
|
||||
If `recv_buf` or `send_buf` shows a smaller number than expected, the
|
||||
host sysctl is still clamping. Recheck `sysctl net.core.rmem_max
|
||||
net.core.wmem_max` and confirm the drop-in file is being loaded
|
||||
(`sudo sysctl --system` prints the loaded files).
|
||||
|
||||
## Docker and other container hosts
|
||||
|
||||
Containers share the host kernel, so sysctls apply to the host, not
|
||||
the container. If you run FIPS inside Docker, set
|
||||
`net.core.rmem_max` / `net.core.wmem_max` on the **Docker host**, not
|
||||
inside the container. Container privileges (cap_sys_admin) and
|
||||
`--sysctl` flags do not let you raise these particular limits from
|
||||
inside a container — they are global to the host network namespace.
|
||||
|
||||
For Kubernetes deployments, the host-level sysctl tuning is the same;
|
||||
node-level configuration (DaemonSet with `privileged: true`, or a
|
||||
node-init script) is the typical mechanism.
|
||||
|
||||
## Tuning higher
|
||||
|
||||
The 4 MB ceiling is a conservative starting point. For very high
|
||||
throughput (multi-gigabit per second), raise both sysctls and the
|
||||
corresponding `transports.udp.recv_buf_size` /
|
||||
`transports.udp.send_buf_size` config values together. Setting a config value
|
||||
larger than the host ceiling silently clamps to the ceiling, so both
|
||||
must move in lockstep.
|
||||
|
||||
## See also
|
||||
|
||||
- [../design/fips-transport-layer.md](../design/fips-transport-layer.md)
|
||||
— UDP transport design, why FIPS requests larger buffers,
|
||||
`SO_RXQ_OVFL` and ECN integration
|
||||
- [../reference/configuration.md](../reference/configuration.md) —
|
||||
`transports.udp.recv_buf_size` and `send_buf_size` defaults
|
||||
25
docs/reference/README.md
Normal file
@@ -0,0 +1,25 @@
|
||||
# Reference
|
||||
|
||||
Information-oriented technical descriptions for lookup on demand.
|
||||
Reference content describes *what is*: wire formats, configuration
|
||||
keys, command-line flags, control-socket commands, default values,
|
||||
file paths, exit codes. It is consulted, not read end-to-end.
|
||||
|
||||
Reference is austere by design: minimal narrative, no opinions, no
|
||||
guidance on when to use a feature. The "why" lives in design/; the
|
||||
"how do I accomplish X" lives in how-to/.
|
||||
|
||||
## Available Reference
|
||||
|
||||
| Document | Scope |
|
||||
| -------- | ----- |
|
||||
| [wire-formats.md](wire-formats.md) | All FMP and FSP message byte layouts, encapsulation walkthrough |
|
||||
| [configuration.md](configuration.md) | Full YAML configuration reference for the daemon and gateway |
|
||||
| [security.md](security.md) | nftables baseline, peer ACL, cryptographic primitives, rekey defaults, threat-resistance matrix |
|
||||
| [nostr-events.md](nostr-events.md) | Kind 37195 advert, Kind 21059 traversal signaling, Kind 10050 inbox relays |
|
||||
| [transports.md](transports.md) | Per-transport statistics counter inventory |
|
||||
| [control-socket.md](control-socket.md) | Line-delimited JSON control protocol for the daemon and gateway |
|
||||
| [cli-fips.md](cli-fips.md) | `fips` daemon CLI: options, exit codes, environment, files |
|
||||
| [cli-fipsctl.md](cli-fipsctl.md) | `fipsctl` control-client: subcommands, options, exit codes |
|
||||
| [cli-fipstop.md](cli-fipstop.md) | `fipstop` live-status TUI: tabs, keybindings |
|
||||
| [cli-fips-gateway.md](cli-fips-gateway.md) | `fips-gateway` service CLI: options, exit codes, files |
|
||||
125
docs/reference/cli-fips-gateway.md
Normal file
@@ -0,0 +1,125 @@
|
||||
# `fips-gateway`
|
||||
|
||||
Long-running service that bridges a LAN segment into the FIPS mesh.
|
||||
|
||||
## Synopsis
|
||||
|
||||
```text
|
||||
fips-gateway [-c FILE] [-l LEVEL]
|
||||
```
|
||||
|
||||
## Description
|
||||
|
||||
`fips-gateway` runs alongside `fips` on the same host, reads the same
|
||||
`fips.yaml`, and exposes two complementary functions to the LAN it
|
||||
fronts:
|
||||
|
||||
- **Outbound (LAN -> mesh).** Allocates a virtual IPv6 from a managed
|
||||
pool when a LAN client resolves `<npub>.fips`, installs nftables
|
||||
DNAT/SNAT/masquerade rules so the client's traffic is rewritten and
|
||||
carried into the mesh through the daemon's `fips0` adapter.
|
||||
- **Inbound (mesh -> LAN).** Installs nftables DNAT and LAN-side
|
||||
masquerade rules so mesh-side traffic arriving on `fips0` for the
|
||||
configured listen ports is rewritten to a LAN `host:port`, per the
|
||||
`gateway.port_forwards[]` block.
|
||||
|
||||
The service runs alongside `fips`, not as a replacement for it:
|
||||
the daemon must be running on the same host with the TUN adapter
|
||||
and DNS resolver enabled. The gateway is read-only with respect to the
|
||||
daemon's state, and connects to the daemon's resolver only — it is
|
||||
not a peer. For the architecture, see
|
||||
[../design/fips-gateway.md](../design/fips-gateway.md).
|
||||
|
||||
`fips-gateway` is **Linux-only**. The binary errors out and exits with
|
||||
status `1` on any other platform, since the NAT pipeline is built on
|
||||
nftables and proxy NDP. See
|
||||
[Configuration](#configuration) for the platform notes that follow
|
||||
from this.
|
||||
|
||||
## Options
|
||||
|
||||
| Flag | Argument | Default | Description |
|
||||
| ---- | -------- | ------- | ----------- |
|
||||
| `-c`, `--config` | `FILE` | *(default search paths)* | Use `FILE` as the configuration. Skips the default search paths. |
|
||||
| `-l`, `--log-level` | `LEVEL` | `info` | Tracing level: `trace`, `debug`, `info`, `warn`, `error`. Overridden by `RUST_LOG` if set (see [Environment](#environment)). |
|
||||
| `-V` | — | — | Print the short version. |
|
||||
| `--version` | — | — | Print the long version (short version plus build target triple). |
|
||||
| `-h`, `--help` | — | — | Print usage and exit. |
|
||||
|
||||
## Configuration
|
||||
|
||||
`fips-gateway` reads the same `fips.yaml` as `fips`; the gateway is
|
||||
configured under the top-level `gateway:` block. The block must
|
||||
include at minimum `enabled: true`, `pool`, and `lan_interface`. For
|
||||
each field — pool, LAN interface, DNS listener, conntrack overrides,
|
||||
and inbound `port_forwards[]` — see the
|
||||
[Gateway section](configuration.md#gateway-gateway) of the
|
||||
configuration reference.
|
||||
|
||||
The same default search paths apply as for `fips`
|
||||
(see [`fips`](cli-fips.md#files)); `-c FILE` overrides the search.
|
||||
The gateway must be able to read the same configuration file the
|
||||
daemon is reading, or the two will disagree about pool, DNS port,
|
||||
and LAN interface.
|
||||
|
||||
For deployment recipes, see
|
||||
[../how-to/deploy-gateway.md](../how-to/deploy-gateway.md) (manual
|
||||
Linux host) and
|
||||
[../tutorials/deploy-fips-gateway.md](../tutorials/deploy-fips-gateway.md)
|
||||
(OpenWrt walk-through).
|
||||
|
||||
## Exit Codes
|
||||
|
||||
| Code | Meaning |
|
||||
| ---- | ------- |
|
||||
| `0` | Clean shutdown after `SIGINT` / `SIGTERM`. |
|
||||
| `1` | Non-Linux platform, configuration load failure, missing or invalid `gateway:` block, NAT/network setup failure, or control-socket bind failure. The reason is printed to stderr or the log before exit. |
|
||||
|
||||
## Environment
|
||||
|
||||
| Variable | Description |
|
||||
| -------- | ----------- |
|
||||
| `RUST_LOG` | Tracing filter directive. Takes precedence over `--log-level`. Examples: `info`, `debug`, `fips=trace,fips::gateway=debug`. |
|
||||
|
||||
## Files
|
||||
|
||||
| Path | Purpose |
|
||||
| ---- | ------- |
|
||||
| `/etc/fips/fips.yaml` | Gateway configuration (top-level `gateway:` block). Same file the daemon reads. |
|
||||
| `/run/fips/gateway.sock` | Gateway control socket. Hardcoded path; chowned to group `fips` (mode `0770`) at startup so members of that group can query without sudo. |
|
||||
| `inet fips_gateway` (nftables) | NAT table the gateway installs and tears down. View with `nft list table inet fips_gateway`. |
|
||||
|
||||
The gateway also adds and removes a `local <pool-cidr> dev lo` route
|
||||
in the local routing table so the kernel accepts pool addresses as
|
||||
locally-owned.
|
||||
|
||||
## Control Socket
|
||||
|
||||
`fips-gateway` exposes a JSON line-protocol control socket separate
|
||||
from the daemon's. The command set (`show_gateway`, `show_mappings`)
|
||||
and JSON shapes are documented in the
|
||||
[Gateway Command Catalog](control-socket.md#gateway-command-catalog).
|
||||
|
||||
There is no `fipsctl` subcommand for the gateway — query the socket
|
||||
directly with `nc -U`, or watch the **Gateway** tab in
|
||||
[`fipstop`](cli-fipstop.md), which polls the gateway socket
|
||||
automatically.
|
||||
|
||||
## See also
|
||||
|
||||
- [`fips`](cli-fips.md) — the daemon. Required to be running on the
|
||||
same host.
|
||||
- [`fipstop`](cli-fipstop.md) — the live-status TUI; its Gateway tab
|
||||
polls the gateway control socket.
|
||||
- [configuration.md § Gateway](configuration.md#gateway-gateway) —
|
||||
full `gateway.*` block reference.
|
||||
- [control-socket.md § Gateway Command Catalog](control-socket.md#gateway-command-catalog)
|
||||
— wire protocol for the gateway socket.
|
||||
- [../design/fips-gateway.md](../design/fips-gateway.md) — design,
|
||||
NAT pipeline, virtual IP pool lifecycle.
|
||||
- [../how-to/deploy-gateway.md](../how-to/deploy-gateway.md) — manual
|
||||
Linux deployment.
|
||||
- [../how-to/troubleshoot-gateway.md](../how-to/troubleshoot-gateway.md)
|
||||
— diagnostic recipes.
|
||||
- [../tutorials/deploy-fips-gateway.md](../tutorials/deploy-fips-gateway.md)
|
||||
— OpenWrt walk-through.
|
||||
93
docs/reference/cli-fips.md
Normal file
@@ -0,0 +1,93 @@
|
||||
# `fips`
|
||||
|
||||
The FIPS mesh network daemon.
|
||||
|
||||
## Synopsis
|
||||
|
||||
```text
|
||||
fips [-c FILE]
|
||||
```
|
||||
|
||||
On Windows the same binary additionally accepts `--install-service`,
|
||||
`--uninstall-service`, and (used internally by the service control
|
||||
manager) `--service`.
|
||||
|
||||
## Description
|
||||
|
||||
`fips` is the FIPS daemon. It loads a YAML configuration, resolves an
|
||||
identity, brings up the TUN adapter, listens on configured transports,
|
||||
authenticates peers, maintains the spanning tree, and forwards mesh
|
||||
traffic. There is one daemon per node.
|
||||
|
||||
The daemon stays in the foreground, logging to stderr, until it
|
||||
receives `SIGINT` or `SIGTERM`. On Windows, the service variant is
|
||||
controlled through the standard service control manager.
|
||||
|
||||
## Options
|
||||
|
||||
| Flag | Argument | Description |
|
||||
| ---- | -------- | ----------- |
|
||||
| `-c`, `--config` | `FILE` | Use `FILE` as the configuration. Skips the default search paths. |
|
||||
| `-V` | — | Print the short version (e.g. `0.4.0 (rev abcdef1)`). |
|
||||
| `--version` | — | Print the long version: short version plus build target triple. |
|
||||
| `-h`, `--help` | — | Print usage and exit. |
|
||||
| `--install-service` | — | (Windows only) Install `fips` as a Windows service. Requires Administrator. |
|
||||
| `--uninstall-service` | — | (Windows only) Uninstall the Windows service. Requires Administrator. |
|
||||
| `--service` | — | (Windows only, internal) Run as a Windows service. Invoked by the service control manager — not for direct use. |
|
||||
|
||||
There are no other CLI flags; all daemon behaviour is governed by the
|
||||
YAML configuration. See [configuration.md](configuration.md).
|
||||
|
||||
## Exit Codes
|
||||
|
||||
| Code | Meaning |
|
||||
| ---- | ------- |
|
||||
| `0` | Clean shutdown after `SIGINT` / `SIGTERM`. |
|
||||
| `1` | Failed to load configuration, resolve identity, construct the node, or start the node. The reason is printed to stderr before exit. |
|
||||
|
||||
## Environment
|
||||
|
||||
| Variable | Description |
|
||||
| -------- | ----------- |
|
||||
| `RUST_LOG` | Tracing filter directive. Overrides `node.log_level` from the config. Examples: `info`, `debug`, `fips=trace,fips::node::handlers::mmp=debug`. |
|
||||
| `XDG_RUNTIME_DIR` | Used to derive the default control-socket path when `/run/fips` does not exist. See [control-socket.md](control-socket.md). |
|
||||
| `FIPS_CONFIG` | (Windows service mode only) Path to the configuration file when the daemon runs under the service control manager. |
|
||||
|
||||
The daemon also clamps the `nostr_relay_pool`, `nostr_sdk`, and `nostr`
|
||||
log targets to `info` whenever the effective log level is below
|
||||
`trace`, so that `RUST_LOG=debug` does not flood the journal with raw
|
||||
relay frames. To see those frames, set the level to `trace`.
|
||||
|
||||
## Files
|
||||
|
||||
`fips` looks for `fips.yaml` in the following locations, lowest to
|
||||
highest priority. All present files are merged in priority order; the
|
||||
highest-priority value wins.
|
||||
|
||||
| Priority | Path | Purpose |
|
||||
| -------- | ---- | ------- |
|
||||
| 1 | `/etc/fips/fips.yaml` | System-wide defaults |
|
||||
| 2 | `~/.config/fips/fips.yaml` | User preferences |
|
||||
| 3 | `~/.fips.yaml` | Legacy user config |
|
||||
| 4 | `./fips.yaml` | Deployment-specific overrides |
|
||||
|
||||
Adjacent to the highest-priority config file the daemon reads (or
|
||||
writes, on first start) the identity files:
|
||||
|
||||
| File | Mode | Purpose |
|
||||
| ---- | ---- | ------- |
|
||||
| `fips.key` | `0600` | Bech32 nsec for the persistent identity (Unix only; Windows inherits parent ACLs). |
|
||||
| `fips.pub` | `0644` | Bech32 npub corresponding to `fips.key`. |
|
||||
|
||||
When `node.identity.persistent` is `false` (the default), a fresh
|
||||
keypair is written to these files on every start.
|
||||
|
||||
The control socket path is derived per
|
||||
[control-socket.md](control-socket.md).
|
||||
|
||||
## See also
|
||||
|
||||
- [`fipsctl`](cli-fipsctl.md) — control-socket client.
|
||||
- [`fipstop`](cli-fipstop.md) — live-status TUI.
|
||||
- [configuration.md](configuration.md) — YAML reference.
|
||||
- [control-socket.md](control-socket.md) — control-socket protocol.
|
||||
149
docs/reference/cli-fipsctl.md
Normal file
@@ -0,0 +1,149 @@
|
||||
# `fipsctl`
|
||||
|
||||
Command-line client for the FIPS daemon's control socket.
|
||||
|
||||
## Synopsis
|
||||
|
||||
```text
|
||||
fipsctl [-s SOCKET] <subcommand> [args...]
|
||||
```
|
||||
|
||||
## Description
|
||||
|
||||
`fipsctl` connects to a running daemon over its control socket
|
||||
(Unix domain socket on Linux/macOS, TCP loopback on Windows), sends
|
||||
one JSON request, and pretty-prints the response. Exits with a
|
||||
non-zero status if the socket cannot be reached, the daemon returns an
|
||||
error, or the request times out.
|
||||
|
||||
`fipsctl keygen` is a special case: it does not contact the daemon and
|
||||
operates purely on local files.
|
||||
|
||||
For the line-delimited JSON wire protocol, see
|
||||
[control-socket.md](control-socket.md). For the YAML configuration
|
||||
that defines the socket location, see
|
||||
[configuration.md](configuration.md).
|
||||
|
||||
## Global Options
|
||||
|
||||
| Flag | Argument | Description |
|
||||
| ---- | -------- | ----------- |
|
||||
| `-s`, `--socket` | `PATH` | Override the control-socket path (Linux/macOS) or TCP port (Windows). |
|
||||
| `-V`, `--version` | — | Print the short version. |
|
||||
| `--version` | — | Print the long version. |
|
||||
| `-h`, `--help` | — | Print usage and exit. Per-subcommand help via `fipsctl <subcommand> --help`. |
|
||||
|
||||
## Subcommands
|
||||
|
||||
### `show <what>`
|
||||
|
||||
Read-only queries against the daemon. Each subcommand maps 1:1 to a
|
||||
control-socket query (see [control-socket.md](control-socket.md)) and
|
||||
prints the response's `data` object as pretty JSON.
|
||||
|
||||
| Subcommand | Control-socket command | Returns |
|
||||
| ---------- | ---------------------- | ------- |
|
||||
| `show status` | `show_status` | Node-level status: identity, version, peer/link/session counts, TUN state, recent sparklines. |
|
||||
| `show peers` | `show_peers` | Authenticated peer list with link IDs, transport addresses, MMP metrics, Noise/rekey state. |
|
||||
| `show links` | `show_links` | Active links (one per FMP-authenticated peer): direction, state, byte counters. |
|
||||
| `show tree` | `show_tree` | Spanning-tree state: root, my coordinates, parent, peer declarations. |
|
||||
| `show sessions` | `show_sessions` | End-to-end FSP sessions: state, traffic counters, session-MMP metrics, path MTU. |
|
||||
| `show bloom` | `show_bloom` | Bloom-filter state: own filter sequence, leaf dependents, per-peer filter summaries. |
|
||||
| `show mmp` | `show_mmp` | MMP metrics summary: per-peer link-layer metrics and per-session session-layer metrics. |
|
||||
| `show cache` | `show_cache` | Coordinate cache: TTL, fill ratio, per-destination coords and path MTU. |
|
||||
| `show connections` | `show_connections` | Pending handshake connections: state, idle time, resend count. |
|
||||
| `show transports` | `show_transports` | Transport instances: type, state, MTU, local address, per-transport stats. |
|
||||
| `show routing` | `show_routing` | Routing summary: pending lookups, retry state, forwarding/discovery/error/congestion counters. |
|
||||
| `show identity-cache` | `show_identity_cache` | Cached `(node_addr → npub)` entries with last-seen timestamps. |
|
||||
|
||||
### `acl <what>`
|
||||
|
||||
| Subcommand | Control-socket command | Returns |
|
||||
| ---------- | ---------------------- | ------- |
|
||||
| `acl show` | `show_acl` | Loaded peer-ACL state: allow/deny files, effective mode, default decision, entry counts. |
|
||||
|
||||
### `stats <what>`
|
||||
|
||||
Time-series metrics from the in-process history rings.
|
||||
|
||||
| Subcommand | Control-socket command | Description |
|
||||
| ---------- | ---------------------- | ----------- |
|
||||
| `stats list` | `show_stats_list` | Enumerate available metrics, their units, and the per-ring retention windows. |
|
||||
| `stats metrics` | `show_metrics` | Dump current counter values for every protocol metric family (`forwarding`, `discovery`, `tree`, `bloom`, `congestion`, `errors`). |
|
||||
| `stats peers` | `show_stats_peers` | List peers tracked in stats history (active or recently active). |
|
||||
| `stats history <metric> [options]` | `show_stats_history` | Fetch a time-series window for one metric. |
|
||||
|
||||
`stats history` options:
|
||||
|
||||
| Flag | Argument | Default | Description |
|
||||
| ---- | -------- | ------- | ----------- |
|
||||
| `--peer` | `npub` or hostname | *(none)* | Required for per-peer metrics; resolves through `/etc/fips/hosts` if not an npub. |
|
||||
| `--window` | `<N>s` / `<N>m` / `<N>h` | `10m` | Window duration. |
|
||||
| `--granularity` | `1s` or `1m` | `1s` | Ring resolution. `1s` uses the fast ring; `1m` uses the slow ring. |
|
||||
| `--plot` | — | off | Render a Unicode-block sparkline to stdout instead of JSON. |
|
||||
|
||||
### `keygen [options]`
|
||||
|
||||
Generate a new FIPS identity keypair locally. Does not contact the
|
||||
daemon.
|
||||
|
||||
| Flag | Argument | Default | Description |
|
||||
| ---- | -------- | ------- | ----------- |
|
||||
| `-d`, `--dir` | `DIR` | `/etc/fips` (Unix), `%APPDATA%\fips` (Windows) | Output directory for `fips.key` and `fips.pub`. |
|
||||
| `-f`, `--force` | — | off | Overwrite an existing `fips.key`. |
|
||||
| `-s`, `--stdout` | — | off | Print `nsec` then `npub` to stdout instead of writing files. |
|
||||
|
||||
`fips.key` is written with mode `0600` and `fips.pub` with mode `0644`
|
||||
on Unix. After running `keygen`, set `node.identity.persistent: true`
|
||||
in `fips.yaml` or the daemon will overwrite the keys on next start.
|
||||
|
||||
### `connect <peer> <address> <transport>`
|
||||
|
||||
Tell the daemon to dial a peer over a specific transport.
|
||||
|
||||
| Argument | Description |
|
||||
| -------- | ----------- |
|
||||
| `peer` | npub (bech32) or hostname from `/etc/fips/hosts`. |
|
||||
| `address` | Transport endpoint, e.g. `192.168.1.10:2121`, `[2001:db8::1]:2121`, or a Tor onion. FIPS-mesh ULAs (`fd00::/8`) are rejected for the IP-based transports (udp, tcp, ethernet). |
|
||||
| `transport` | One of `udp`, `tcp`, `tor`, `nym`, `ethernet`. The named transport must be configured and running. |
|
||||
|
||||
### `disconnect <peer>`
|
||||
|
||||
Tell the daemon to drop a peer link.
|
||||
|
||||
| Argument | Description |
|
||||
| -------- | ----------- |
|
||||
| `peer` | npub (bech32) or hostname from `/etc/fips/hosts`. |
|
||||
|
||||
## Exit Codes
|
||||
|
||||
| Code | Meaning |
|
||||
| ---- | ------- |
|
||||
| `0` | Daemon returned `{"status":"ok",...}`. |
|
||||
| `1` | Argument parse failure, control-socket connection failure, daemon returned `{"status":"error",...}`, or local I/O failure (keygen). The error message is printed to stderr. |
|
||||
|
||||
## Environment
|
||||
|
||||
| Variable | Description |
|
||||
| -------- | ----------- |
|
||||
| `XDG_RUNTIME_DIR` | Used to derive the default control-socket path when `/run/fips` is absent. |
|
||||
|
||||
`fipsctl` does not consume `RUST_LOG`; logging is for the daemon.
|
||||
|
||||
## Files
|
||||
|
||||
| Path | Purpose |
|
||||
| ---- | ------- |
|
||||
| `/etc/fips/hosts` | Maps hostnames to npubs for the `connect`, `disconnect`, and `--peer` arguments. See [configuration.md](configuration.md). |
|
||||
| Control socket (default) | Same resolution as the daemon: `/run/fips/control.sock` if present, else `$XDG_RUNTIME_DIR/fips/control.sock`, else `/tmp/fips-control.sock` (Unix); TCP `localhost:21210` (Windows). |
|
||||
|
||||
If you get `Permission denied` connecting to the socket on Linux,
|
||||
add your user to the `fips` group (`sudo usermod -aG fips $USER`)
|
||||
and log out and back in.
|
||||
|
||||
## See also
|
||||
|
||||
- [`fips`](cli-fips.md) — the daemon.
|
||||
- [`fipstop`](cli-fipstop.md) — live-status TUI.
|
||||
- [control-socket.md](control-socket.md) — wire protocol.
|
||||
- [configuration.md](configuration.md) — YAML reference.
|
||||
198
docs/reference/cli-fipstop.md
Normal file
@@ -0,0 +1,198 @@
|
||||
# `fipstop`
|
||||
|
||||
Live-status terminal UI for a running FIPS daemon.
|
||||
|
||||
## Synopsis
|
||||
|
||||
```text
|
||||
fipstop [-s SOCKET] [--gateway-socket PATH] [-r SECONDS]
|
||||
```
|
||||
|
||||
## Description
|
||||
|
||||
`fipstop` is a `ratatui`-based dashboard. It opens the daemon control
|
||||
socket, polls a small set of `show_*` queries on a timer, and renders
|
||||
the state in a tabbed full-screen UI. A separate poll runs against the
|
||||
gateway control socket when the Gateway tab is active.
|
||||
|
||||
`fipstop` is almost entirely read-only: the only state-mutating action
|
||||
it offers is disconnecting a peer (`Del` on a selected Peers row, with
|
||||
a confirmation prompt — see [Keybindings](#keybindings)). For
|
||||
`connect` and other mutating commands, use
|
||||
[`fipsctl`](cli-fipsctl.md).
|
||||
|
||||
## Options
|
||||
|
||||
| Flag | Argument | Default | Description |
|
||||
| ---- | -------- | ------- | ----------- |
|
||||
| `-s`, `--socket` | `PATH` | (auto) | Daemon control-socket path / port. Same default as `fipsctl`. |
|
||||
| `--gateway-socket` | `PATH` | (auto) | `fips-gateway` control-socket path / port. Default: `/run/fips/gateway.sock` (Unix), TCP port `21211` (Windows). |
|
||||
| `-r`, `--refresh` | `SECONDS` | `2` | Poll interval. |
|
||||
| `-V`, `--version` | — | — | Print short version. |
|
||||
| `--version` | — | — | Print long version. |
|
||||
| `-h`, `--help` | — | — | Print usage and exit. |
|
||||
|
||||
## Tabs
|
||||
|
||||
Tabs cycle in this order. Each tab issues the listed control-socket
|
||||
query on its first activation and on every refresh tick while active.
|
||||
|
||||
| Tab | Query | Shows |
|
||||
| --- | ----- | ----- |
|
||||
| **Node** | `show_status` (+ `show_listening_sockets`) | Identity, version, uptime, peer/link/session counts, sparklines for mesh size, tree depth, peer count, bytes, loss. The Traffic block on this tab is split: TUN counters on the left, the **Listening on fips0** panel on the right (see below). |
|
||||
| **Peers** | `show_peers` (+ `show_links`, `show_transports` cross-refs) | Authenticated peers in a table. Selecting a row and pressing Enter opens a detail view. |
|
||||
| **Transports** | `show_transports` (+ `show_links`, `show_peers` cross-refs) | Tree of transport instances with per-link children when expanded. |
|
||||
| **Sessions** | `show_sessions` | End-to-end FSP sessions. |
|
||||
| **Tree** | `show_tree` | Spanning-tree state and per-peer coordinates. |
|
||||
| **Filters** | `show_bloom` | Per-peer Bloom-filter state. |
|
||||
| **Performance** | `show_mmp` | Link-layer and session-layer MMP metrics. |
|
||||
| **Routing** | `show_routing` (+ `show_cache` cross-ref) | Forwarding/discovery counters, pending lookups, retry state. |
|
||||
| **Graphs** | `show_stats_history` family + `show_stats_peers` | Stacked time-series plots. Three modes: node-level metrics, one metric across peers, all metrics for one peer. |
|
||||
| **Gateway** | `show_gateway` and `show_mappings` against the gateway socket | Pool utilisation and per-mapping state when `fips-gateway` is running. Empty when the gateway socket is unreachable. |
|
||||
|
||||
The cycle order in the UI is: Node → Peers → Transports → Sessions →
|
||||
Tree → Filters → Performance → Routing → Graphs → Gateway. The Links
|
||||
and Cache tabs are not in the cycle but are fetched as cross-references
|
||||
to populate Peers, Transports, and Routing detail views.
|
||||
|
||||
## Listening on fips0 panel (Node tab)
|
||||
|
||||
The right half of the Node tab's Traffic block lists local IPv6
|
||||
listening sockets reachable from `fips0`, paired with the current
|
||||
`inet fips` baseline filter classification for each (proto, port).
|
||||
The panel exists to remind the operator which local services are
|
||||
exposed to the mesh and which of those are admitted by the
|
||||
default-deny firewall.
|
||||
|
||||
| Column | Meaning |
|
||||
| ------ | ------- |
|
||||
| **Proto** | `tcp` or `udp`. IPv4 listeners are not enumerated; `fips0` is IPv6-only. |
|
||||
| **Port** | Listening port number. |
|
||||
| **Process** | `comm(pid)` resolved by walking `/proc/<pid>/fd/`. A trailing `*` marks wildcard binds (`local_addr == ::`) — the bind is not fips0-specific, so the operator sees that the service is exposed across every interface, not just the mesh. |
|
||||
| **State** | `OPEN` (default White) — the baseline filter has a canonical accept rule for this (proto, port). `filt` (DarkGray) — chain falls through to `counter drop`. `filt?` (DarkGray) — a rule references the port but uses matchers (saddr filter, jump, daddr) the panel cannot fully decompose; operator should `nft list table inet fips` to confirm. |
|
||||
|
||||
When `fips-firewall.service` is **not** active, the `inet fips`
|
||||
table is absent. The panel renders every row in default White and
|
||||
replaces the title with a yellow banner reading
|
||||
"`Listening on fips0 fips-firewall.service inactive — all listeners exposed`".
|
||||
|
||||
The panel is read-only and unselectable. It refreshes on the same
|
||||
poll tick as the rest of the Node tab. Sockets owned by other users
|
||||
that the daemon could not resolve to a PID render as `?` in the
|
||||
Process column; this only happens if the daemon itself is running
|
||||
without root privileges (an unusual dev setup), since walking
|
||||
`/proc/<pid>/fd/` for processes the daemon does not own requires
|
||||
elevated capabilities.
|
||||
|
||||
The panel is Linux-only; on non-Linux daemons the query returns an
|
||||
empty list and the panel hides.
|
||||
|
||||
## Keybindings
|
||||
|
||||
Press `?` at any time for an in-app help overlay. The overlay and the
|
||||
status-bar hint footer both read from a single keybinding registry
|
||||
keyed by `(tab, mode)`, so the always-visible hints describe exactly
|
||||
the keys the current context accepts.
|
||||
|
||||
### Global
|
||||
|
||||
| Key | Action |
|
||||
| --- | ------ |
|
||||
| `q`, `Ctrl-C` | Quit. |
|
||||
| `Tab` | Next tab. |
|
||||
| `Shift-Tab` | Previous tab. |
|
||||
| `g` | Jump to the Graphs tab. |
|
||||
| `?` | Toggle the help overlay. |
|
||||
| `Esc` | Close an open detail view; otherwise deselect the active table row. |
|
||||
|
||||
### Table tabs (Peers, Sessions, Transports, Gateway)
|
||||
|
||||
| Key | Action |
|
||||
| --- | ------ |
|
||||
| `Up`, `Down` | Move row selection. |
|
||||
| `Enter` | Open detail view for the selected row. |
|
||||
| `Esc` | Deselect the row (return to the tab's overview state). |
|
||||
|
||||
### Peers tab (extra)
|
||||
|
||||
| Key | Action |
|
||||
| --- | ------ |
|
||||
| `Del` | Disconnect the selected peer. Opens a `Y`/`N` confirmation modal first; this is the only state-mutating action in `fipstop`. |
|
||||
|
||||
### Transports tab (extra)
|
||||
|
||||
| Key | Action |
|
||||
| --- | ------ |
|
||||
| `Right`, `Space` | Expand the selected transport row to show its links. |
|
||||
| `Left` | Collapse the selected transport row. |
|
||||
| `e` | Expand all transports. |
|
||||
| `c` | Collapse all transports. |
|
||||
|
||||
### Multi-pane scrolling tabs (Tree, Filters, Routing)
|
||||
|
||||
Each lays out stacked panes that scroll independently.
|
||||
|
||||
| Key | Action |
|
||||
| --- | ------ |
|
||||
| `f` | Move focus to the next pane. |
|
||||
| `Up`, `Down` | Scroll the focused pane by one row. |
|
||||
| `PageUp`, `PageDown` | Scroll the focused pane by ten rows. |
|
||||
| `Home`, `End` | Jump to the top / bottom of the focused pane. |
|
||||
|
||||
### Performance tab (extra)
|
||||
|
||||
The Performance tab lays out two panes (Link MMP, Session MMP).
|
||||
|
||||
| Key | Action |
|
||||
| --- | ------ |
|
||||
| `f` | Move focus between the Link and Session MMP panes. |
|
||||
| `Up`, `Down` | Scroll the focused pane. |
|
||||
| `PageUp`, `PageDown` | Scroll the focused pane by ten rows. |
|
||||
| `Home`, `End` | Jump to the top / bottom of the focused pane. |
|
||||
| `s` | Cycle the sort column of the focused pane. |
|
||||
| `Shift-S` | Toggle the sort direction of the focused pane. |
|
||||
|
||||
### Graphs tab (extra)
|
||||
|
||||
| Key | Action |
|
||||
| --- | ------ |
|
||||
| `Up`, `Down` | Scroll the stacked plots; in `MetricByPeer` mode, move the by-peer selection (and follow it when the by-peer detail is open). |
|
||||
| `Right`, `Space` | Next time window. Cycles `1m / 1s` → `10m / 1s` → `1h / 1s` → `24h / 1m`. |
|
||||
| `Left` | Previous time window. |
|
||||
| `Enter` | In `MetricByPeer` mode, expand the selected peer summary into a full-pane plot. |
|
||||
| `m` | Cycle view mode: `Node` (stacked node metrics) → `MetricByPeer` (one per-peer metric across all peers) → `PeerByMetric` (all per-peer metrics for one peer). |
|
||||
| `n` | Next selector (next per-peer metric in MetricByPeer; next peer in PeerByMetric). |
|
||||
| `Shift-N` | Previous selector. |
|
||||
| `s` | Cycle the sort column of the by-peer summary list. |
|
||||
| `Shift-S` | Toggle the sort direction of the by-peer summary list. |
|
||||
|
||||
## Exit Codes
|
||||
|
||||
| Code | Meaning |
|
||||
| ---- | ------- |
|
||||
| `0` | Normal quit. |
|
||||
| `1` | Failed to initialise the terminal. The reason is printed to stderr. |
|
||||
|
||||
A failure to reach the daemon socket is **not** fatal: the dashboard
|
||||
displays "Disconnected" in the status bar and retries on every refresh
|
||||
tick.
|
||||
|
||||
## Environment
|
||||
|
||||
| Variable | Description |
|
||||
| -------- | ----------- |
|
||||
| `XDG_RUNTIME_DIR` | Used to derive the default control-socket and gateway-socket paths when `/run/fips` is absent. |
|
||||
|
||||
## Files
|
||||
|
||||
Same control-socket resolution rules as
|
||||
[`fipsctl`](cli-fipsctl.md#files). The gateway socket follows the same
|
||||
pattern with `gateway.sock` in place of `control.sock`, falling back
|
||||
to `/tmp/fips-gateway.sock` if neither system path nor
|
||||
`XDG_RUNTIME_DIR` is available.
|
||||
|
||||
## See also
|
||||
|
||||
- [`fipsctl`](cli-fipsctl.md) — issue mutating commands.
|
||||
- [`fips`](cli-fips.md) — the daemon.
|
||||
- [control-socket.md](control-socket.md) — wire protocol fipstop polls.
|
||||
@@ -53,15 +53,29 @@ peers: # Static peer list
|
||||
|
||||
| Parameter | Type | Default | Description |
|
||||
|-----------|------|---------|-------------|
|
||||
| `node.control.enabled` | bool | `true` | Enable the Unix domain control socket |
|
||||
| `node.control.socket_path` | string | *(auto)* | Socket file path. Default: `$XDG_RUNTIME_DIR/fips/control.sock`, then `/run/fips/control.sock` (if root), then `/tmp/fips-control.sock` |
|
||||
| `node.control.enabled` | bool | `true` | Enable the control socket |
|
||||
| `node.control.socket_path` | string | *(auto)* | **Linux:** Socket file path. Resolved at daemon startup: `$XDG_RUNTIME_DIR/fips/control.sock` if `XDG_RUNTIME_DIR` is set, else `/run/fips/control.sock` if `/run/fips` can be created (typical when running under the shipped systemd unit), else `/tmp/fips-control.sock`. (Note: the `fipsctl` / `fipstop` clients use a different fallback order — `/run/fips` first if it already exists, then `XDG_RUNTIME_DIR`, then `/tmp` — so when both schemes apply, set this field explicitly to avoid mismatch.) **Windows:** TCP port number (default: `21210`); the control socket listens on `127.0.0.1` at this port. |
|
||||
|
||||
The control socket provides access to node state and runtime management
|
||||
via the `fipsctl` command-line tool. In addition to read-only status
|
||||
queries, `fipsctl connect` and `fipsctl disconnect` enable runtime peer
|
||||
management. See the project [README](../../README.md#inspect) for the
|
||||
management. See the [`fipsctl` reference](cli-fipsctl.md) for the
|
||||
command list.
|
||||
|
||||
On Linux, the control socket is a Unix domain socket with filesystem
|
||||
permissions (mode 0770, group `fips`). On Windows, it is a TCP listener
|
||||
on localhost. TCP does not provide filesystem-level ACLs, so any local
|
||||
user can connect to the control port.
|
||||
|
||||
> **Security note (Windows):** The TCP control socket on Windows is a
|
||||
> known limitation. Any process running on the local machine can connect
|
||||
> to the control port and issue commands, including `disconnect`,
|
||||
> `connect`, and `inject-config`. This is acceptable for single-user
|
||||
> workstations but may be inappropriate for shared machines. Future
|
||||
> improvements may include named pipe support (with Windows ACLs) or an
|
||||
> authentication token mechanism. On shared Windows systems, consider
|
||||
> using firewall rules to restrict access to the control port.
|
||||
|
||||
All tunable protocol parameters live under `node.*`, organized as sysctl-style
|
||||
dotted paths. The top-level sections (`tun`, `dns`, `transports`, `peers`)
|
||||
handle infrastructure concerns only.
|
||||
@@ -95,6 +109,7 @@ to the highest-priority config file for operator visibility, even in ephemeral m
|
||||
| `node.base_rtt_ms` | u64 | `100` | Initial RTT estimate for new links before measurements converge |
|
||||
| `node.heartbeat_interval_secs` | u64 | `10` | Heartbeat send interval per peer for liveness detection |
|
||||
| `node.link_dead_timeout_secs` | u64 | `30` | No-traffic timeout before a peer is declared dead and removed |
|
||||
| `node.log_level` | string | `"info"` | Tracing filter default. Case-insensitive; one of `trace`, `debug`, `info`, `warn`, `error`. Overridden by the `RUST_LOG` environment variable when set |
|
||||
|
||||
### Resource Limits (`node.limits.*`)
|
||||
|
||||
@@ -151,14 +166,98 @@ Controls bloom-guided node discovery (LookupRequest/LookupResponse).
|
||||
| Parameter | Type | Default | Description |
|
||||
|-----------|------|---------|-------------|
|
||||
| `node.discovery.ttl` | u8 | `64` | Hop limit for LookupRequest forwarding |
|
||||
| `node.discovery.timeout_secs` | u64 | `10` | Lookup completion timeout |
|
||||
| `node.discovery.attempt_timeouts_secs` | array<u64> | `[1, 2, 4, 8]` | Per-attempt timeouts. Each entry is the deadline for one `LookupRequest` before sending the next attempt with a fresh `request_id`. Length determines total attempt count; default gives 4 attempts and a 15s total budget |
|
||||
| `node.discovery.recent_expiry_secs` | u64 | `10` | Dedup cache expiry for recent request IDs |
|
||||
| `node.discovery.retry_interval_secs` | u64 | `5` | Retry interval within the timeout window; after this interval without a response, resend the lookup |
|
||||
| `node.discovery.max_attempts` | u8 | `2` | Max attempts per lookup (1 = no retry, 2 = one retry) |
|
||||
| `node.discovery.backoff_base_secs` | u64 | `30` | Base for exponential backoff after lookup failure; doubles per consecutive failure |
|
||||
| `node.discovery.backoff_max_secs` | u64 | `300` | Cap on exponential backoff (5 minutes) |
|
||||
| `node.discovery.backoff_base_secs` | u64 | `0` | Optional post-failure suppression base in seconds; doubles per consecutive failure. `0` disables (default) — the per-attempt sequence is the only retry pacing |
|
||||
| `node.discovery.backoff_max_secs` | u64 | `0` | Cap on optional post-failure backoff |
|
||||
| `node.discovery.forward_min_interval_secs` | u64 | `2` | Transit-side rate limiting: minimum interval between forwarded lookups for the same target |
|
||||
|
||||
#### Nostr Overlay Discovery (`node.discovery.nostr.*`)
|
||||
|
||||
Optional Nostr-mediated overlay discovery. This layer publishes replaceable
|
||||
endpoint adverts (`fips-overlay-v1`), consumes advert-derived endpoint
|
||||
fallbacks for configured peers, and can optionally discover non-configured
|
||||
peers (`policy: open`). `udp:nat` remains the trigger for NAT traversal
|
||||
offer/answer + punch-through, after which the established UDP socket is handed
|
||||
into the normal FIPS transport/session stack.
|
||||
Inbox-relay discovery falls back to the local DM relay list if remote relay
|
||||
metadata cannot be fetched.
|
||||
The Nostr discovery runtime is compiled into every build of the crate; it
|
||||
is enabled at runtime via `node.discovery.nostr.enabled: true` and stays
|
||||
inert otherwise.
|
||||
|
||||
| Parameter | Type | Default | Description |
|
||||
|-----------|------|---------|-------------|
|
||||
| `node.discovery.nostr.enabled` | bool | `false` | Enable Nostr-mediated overlay discovery |
|
||||
| `node.discovery.nostr.policy` | string | `"configured_only"` | Advert discovery policy: `disabled`, `configured_only`, `open` |
|
||||
| `node.discovery.nostr.open_discovery_max_pending` | usize | `64` | Max open-discovery peers queued in outbound retry/connection state at once |
|
||||
| `node.discovery.nostr.max_concurrent_incoming_offers` | usize | `16` | Max concurrent inbound traversal offers processed at once (rate limit against offer spam) |
|
||||
| `node.discovery.nostr.advert_cache_max_entries` | usize | `2048` | Max cached overlay adverts retained from relay traffic |
|
||||
| `node.discovery.nostr.seen_sessions_max_entries` | usize | `2048` | Max seen-session IDs retained for replay detection |
|
||||
| `node.discovery.nostr.advertise` | bool | `true` | Publish local endpoint adverts |
|
||||
| `node.discovery.nostr.advert_relays` | list[string] | `["wss://relay.damus.io", "wss://nos.lol", "wss://offchain.pub"]` | Relays used for service adverts |
|
||||
| `node.discovery.nostr.dm_relays` | list[string] | `["wss://relay.damus.io", "wss://nos.lol", "wss://offchain.pub"]` | Relays used for encrypted signaling events |
|
||||
| `node.discovery.nostr.stun_servers` | list[string] | `["stun:stun.l.google.com:19302", "stun:stun.cloudflare.com:3478", "stun:global.stun.twilio.com:3478"]` | STUN servers used for local reflexive address discovery |
|
||||
| `node.discovery.nostr.share_local_candidates` | bool | `false` | Whether to advertise local (RFC 1918 / ULA) interface addresses as host candidates in the traversal offer. Off by default: in most deployments peers aren't on the same broadcast domain, and sharing private host candidates causes misleading punch successes when an asymmetric L3 path (VPN, Tailscale subnet route, overlapping address space) makes a peer's private IP one-way reachable. Enable only when peers are on the same physical LAN |
|
||||
| `node.discovery.nostr.app` | string | `"fips-overlay-v1"` | Traversal application namespace, published in the advert's `protocol` tag (the `d` tag itself is hardcoded to `fips-overlay-v1`) |
|
||||
| `node.discovery.nostr.signal_ttl_secs` | u64 | `120` | Signaling TTL in seconds |
|
||||
| `node.discovery.nostr.attempt_timeout_secs` | u64 | `10` | Overall traversal attempt timeout in seconds |
|
||||
| `node.discovery.nostr.replay_window_secs` | u64 | `300` | Replay tracking retention window in seconds |
|
||||
| `node.discovery.nostr.punch_start_delay_ms` | u64 | `2000` | Delay before punch traffic starts |
|
||||
| `node.discovery.nostr.punch_interval_ms` | u64 | `200` | Interval between punch packets |
|
||||
| `node.discovery.nostr.punch_duration_ms` | u64 | `10000` | How long to keep punching before failure |
|
||||
| `node.discovery.nostr.advert_ttl_secs` | u64 | `3600` | Advert TTL in seconds |
|
||||
| `node.discovery.nostr.advert_refresh_secs` | u64 | `1800` | How often adverts are refreshed in seconds |
|
||||
| `node.discovery.nostr.startup_sweep_delay_secs` | u64 | `5` | Settle delay after Nostr discovery starts before the one-shot startup advert sweep runs (only used under `policy: open`). Allows the relay subscription backlog to populate the in-memory advert cache before the sweep fires |
|
||||
| `node.discovery.nostr.startup_sweep_max_age_secs` | u64 | `3600` | Maximum advert age (`now - created_at`) considered by the one-shot startup sweep (only used under `policy: open`). Adverts older than this are skipped on startup; the per-tick sweep still considers them up to `valid_until_ms` |
|
||||
| `node.discovery.nostr.failure_streak_threshold` | u32 | `5` | Consecutive NAT-traversal failures against a peer before an extended cooldown is applied. At this threshold the daemon also actively re-fetches the peer's advert from `advert_relays` to evict cache entries for peers that have gone away |
|
||||
| `node.discovery.nostr.extended_cooldown_secs` | u64 | `1800` | Cooldown applied to a peer once `failure_streak_threshold` is hit. Suppresses both open-discovery sweep enqueues and per-attempt retry firings until elapsed (30 minutes default) |
|
||||
| `node.discovery.nostr.warn_log_interval_secs` | u64 | `300` | Minimum interval between `NAT traversal failed` WARN log lines for the same peer. Subsequent failures inside the window log at DEBUG to reduce log spam on public-test nodes with many cache-learned peers |
|
||||
| `node.discovery.nostr.failure_state_max_entries` | usize | `4096` | Maximum entries retained in the per-npub failure-state map. Bounds memory under high cache turnover; oldest entries (by last failure time) are evicted when the cap is exceeded |
|
||||
| `node.discovery.nostr.protocol_mismatch_cooldown_secs` | u64 | `86400` | Cooldown applied after observing a fatal protocol mismatch on a Nostr-adopted bootstrap transport (e.g. `Unknown FMP version` from a peer running a different FMP-protocol version). Independent of `extended_cooldown_secs` and much longer (24 hours default) because the mismatch is structural — re-traversing is wasted effort until one side upgrades |
|
||||
|
||||
If `stun_servers` is omitted, the built-in default list above is used. If it is
|
||||
specified in YAML, the configured list fully overrides the defaults.
|
||||
Initiators use only this local list for outbound STUN queries; peer-advertised
|
||||
STUN values are published for diagnostics/interoperability but are not used as
|
||||
arbitrary egress targets.
|
||||
The built-in advert and DM relay defaults point at widely-operated public
|
||||
relays (Damus, nos.lol, Primal) as best-effort endpoints; operators are
|
||||
encouraged to override them with their own relay preferences for production
|
||||
deployments.
|
||||
Advert freshness is enforced semantically: events with expired NIP-40
|
||||
`expiration` tags are dropped, and adverts are also bounded by a created-at
|
||||
staleness window derived from `advert_ttl_secs` (with a grace multiplier).
|
||||
The current in-tree STUN parser handles IPv4 and IPv6 mapped-address
|
||||
attributes. Local traversal candidates include active non-loopback private
|
||||
interface addresses (RFC1918 IPv4 and IPv6 ULA) plus probed local egress
|
||||
addresses for the punch socket port.
|
||||
During punching, compatible private-subnet candidates and reflexive candidates
|
||||
are attempted in parallel; the first successful path wins.
|
||||
|
||||
#### LAN Discovery (`node.discovery.lan.*`)
|
||||
|
||||
Peer discovery on the local link via mDNS / DNS-SD (RFC 6762 / RFC
|
||||
6763). When enabled, the node publishes a `_fips._udp.local.` service
|
||||
advert carrying its `npub` (and optional scope) and concurrently
|
||||
browses for the same service type to learn same-broadcast-domain peers.
|
||||
The result is sub-second peer pairing with no Nostr-relay roundtrip,
|
||||
STUN observation, or NAT traversal: the observed endpoint is by
|
||||
construction routable from the consumer's LAN.
|
||||
|
||||
mDNS adverts are unauthenticated, so a LAN advert is treated only as a
|
||||
routing hint. Identity is still proven end-to-end by the Noise XX
|
||||
handshake the node initiates against the observed endpoint; a spoofed
|
||||
advert carrying another peer's npub fails the handshake and is dropped.
|
||||
LAN discovery requires an active UDP transport (peers dial the
|
||||
advertised UDP port to begin the handshake).
|
||||
|
||||
| Parameter | Type | Default | Description |
|
||||
|-----------|------|---------|-------------|
|
||||
| `node.discovery.lan.enabled` | bool | `false` | Master switch. Opt-in: enable for sub-second same-LAN pairing. Default-off avoids reintroducing a per-LAN identity broadcast on nodes that have deliberately disabled other discovery channels |
|
||||
| `node.discovery.lan.service_type` | string | `"_fips._udp.local."` | DNS-SD service type. Primarily an override for integration tests running multiple isolated services on one loopback interface; leave at the default in production |
|
||||
| `node.discovery.lan.scope` | string | *(none)* | Optional application/network scope carried in a `scope=<name>` TXT entry. Browsers with a scope set only surface peers advertising the same scope, so nodes on the same physical LAN configured for different mesh networks do not cross-feed. Intentionally separate from `node.discovery.nostr.app` so relay-visible adverts can stay generic while LAN discovery is isolated per private network |
|
||||
|
||||
### Spanning Tree (`node.tree.*`)
|
||||
|
||||
Controls tree construction and parent selection.
|
||||
@@ -178,6 +277,7 @@ Controls tree construction and parent selection.
|
||||
| Parameter | Type | Default | Description |
|
||||
|-----------|------|---------|-------------|
|
||||
| `node.bloom.update_debounce_ms` | u64 | `500` | Debounce interval for filter update propagation |
|
||||
| `node.bloom.max_inbound_fpr` | f64 | `0.20` | Antipoison cap: reject inbound `FilterAnnounce` frames whose advertised false-positive rate exceeds this value. Valid range `(0.0, 1.0)`. The default `0.20` corresponds to fill 0.7248 at k=5 (≈2,114 entries on the 1 KB filter); a saturated/poisoned filter is still ~100% FPR and rejected |
|
||||
|
||||
Bloom filter size (1 KB), hash count (5), and size classes are protocol
|
||||
constants and not configurable.
|
||||
@@ -234,7 +334,7 @@ configurable.
|
||||
### Link-Layer MMP (`node.mmp.*`)
|
||||
|
||||
Metrics Measurement Protocol for per-peer link measurement. See
|
||||
[fips-mesh-layer.md](fips-mesh-layer.md) for behavioral details.
|
||||
[../design/fips-mesh-layer.md](../design/fips-mesh-layer.md) for behavioral details.
|
||||
|
||||
| Parameter | Type | Default | Description |
|
||||
|-----------|------|---------|-------------|
|
||||
@@ -282,7 +382,7 @@ with the node for routing.
|
||||
| Parameter | Type | Default | Description |
|
||||
|-----------|------|---------|-------------|
|
||||
| `dns.enabled` | bool | `true` | Enable DNS responder |
|
||||
| `dns.bind_addr` | string | `"127.0.0.1"` | Bind address |
|
||||
| `dns.bind_addr` | string | `"::1"` | Bind address. Default is IPv6 loopback only; the shipped `fips-dns-setup` configures systemd-resolved to forward `.fips` queries to `[::1]:5354`. To expose the responder to mesh peers (or to the gateway over IPv4), override (e.g., `"::"` for all interfaces). |
|
||||
| `dns.port` | u16 | `5354` | Listen port |
|
||||
| `dns.ttl` | u32 | `300` | AAAA record TTL in seconds |
|
||||
|
||||
@@ -299,19 +399,30 @@ The host map is populated from two sources:
|
||||
2. **Hosts file** — `/etc/fips/hosts`, one `hostname npub1...` per line.
|
||||
Blank lines and `#` comments are allowed.
|
||||
|
||||
The hosts file is auto-reloaded on modification (mtime change) without
|
||||
On conflict, hosts-file entries take precedence over peer aliases. The
|
||||
hosts file is auto-reloaded on modification (mtime change) without
|
||||
restarting the daemon. Hostnames are case-insensitive.
|
||||
|
||||
The installer ships `/etc/fips/hosts` pre-populated with the public test
|
||||
mesh roster (`test-us01` … `test-uk01`). Operator-style guide for
|
||||
adding entries and the precedence rules:
|
||||
[../how-to/host-aliases.md](../how-to/host-aliases.md).
|
||||
|
||||
## Transports (`transports.*`)
|
||||
|
||||
### UDP (`transports.udp.*`)
|
||||
|
||||
| Parameter | Type | Default | Description |
|
||||
|-----------|------|---------|-------------|
|
||||
| `transports.udp.bind_addr` | string | `"0.0.0.0:2121"` | UDP bind address and port |
|
||||
| `transports.udp.bind_addr` | string | `"0.0.0.0:2121"` | UDP bind address and port. Ignored when `outbound_only: true` (kernel-assigned ephemeral port is used regardless). |
|
||||
| `transports.udp.mtu` | u16 | `1280` | Transport MTU |
|
||||
| `transports.udp.recv_buf_size` | usize | `2097152` | UDP socket receive buffer size in bytes (2 MB). Linux kernel doubles the requested value internally. Host `net.core.rmem_max` must be >= this value. |
|
||||
| `transports.udp.send_buf_size` | usize | `2097152` | UDP socket send buffer size in bytes (2 MB). Host `net.core.wmem_max` must be >= this value. |
|
||||
| `transports.udp.advertise_on_nostr` | bool | `false` | Include this UDP transport in Nostr endpoint adverts. Implicitly forced false when `outbound_only: true`. |
|
||||
| `transports.udp.public` | bool | `false` | If advertised: `true` publishes direct `host:port`; `false` publishes `udp:nat` rendezvous |
|
||||
| `transports.udp.external_addr` | string | *(none)* | Explicit advertise-as override. Bare IP (`"203.0.113.45"` — bind port is appended) or full `host:port`. Takes precedence over the bound address and STUN autodiscovery. Useful when the public IP isn't on a local interface (cloud 1:1 NAT, EIP) or to skip STUN for a deterministic value. |
|
||||
| `transports.udp.outbound_only` | bool | `false` | Pure-client posture. When `true`, the transport binds to `0.0.0.0:0` (kernel-assigned ephemeral port) regardless of `bind_addr`, refuses inbound handshake msg1, and is never advertised on Nostr regardless of `advertise_on_nostr`. |
|
||||
| `transports.udp.accept_connections` | bool | `true` | Accept inbound handshake msg1 from new peers. Combine with `outbound_only: false` and `accept_connections: false` (plus `auto_connect` on peer entries) for a node that initiates outbound links but rejects fresh inbound handshakes. The handshake handler carves out msg1 from peers already established on this transport so rekey continues to work. |
|
||||
|
||||
### Ethernet (`transports.ethernet.*`)
|
||||
|
||||
@@ -325,7 +436,7 @@ Requires `CAP_NET_RAW` or running as root. Linux only.
|
||||
| `mtu` | u16 | *(auto)* | Override MTU. Default: interface MTU minus 3 (for frame type + length prefix) |
|
||||
| `recv_buf_size` | usize | `2097152` | Socket receive buffer size in bytes (2 MB) |
|
||||
| `send_buf_size` | usize | `2097152` | Socket send buffer size in bytes (2 MB) |
|
||||
| `discovery` | bool | `true` | Listen for discovery beacons from other nodes |
|
||||
| `listen` | bool | `true` | Listen for neighbor beacons from other nodes |
|
||||
| `announce` | bool | `false` | Broadcast announcement beacons on the LAN |
|
||||
| `auto_connect` | bool | `false` | Auto-connect to discovered peers |
|
||||
| `accept_connections` | bool | `false` | Accept incoming connection attempts from discovered peers |
|
||||
@@ -339,7 +450,7 @@ transports:
|
||||
ethernet:
|
||||
lan:
|
||||
interface: "eth0"
|
||||
discovery: true
|
||||
listen: true
|
||||
announce: true
|
||||
backbone:
|
||||
interface: "eth1"
|
||||
@@ -347,7 +458,7 @@ transports:
|
||||
```
|
||||
|
||||
Each named instance operates independently with its own socket and
|
||||
discovery state. The instance name is used in log messages and the
|
||||
neighbor state. The instance name is used in log messages and the
|
||||
`name()` method on the Transport trait.
|
||||
|
||||
### TCP (`transports.tcp.*`)
|
||||
@@ -366,6 +477,8 @@ overhead.
|
||||
| `transports.tcp.recv_buf_size` | usize | `2097152` | Socket receive buffer size in bytes (2 MB) |
|
||||
| `transports.tcp.send_buf_size` | usize | `2097152` | Socket send buffer size in bytes (2 MB) |
|
||||
| `transports.tcp.max_inbound_connections` | usize | `256` | Maximum simultaneous inbound connections |
|
||||
| `transports.tcp.advertise_on_nostr` | bool | `false` | Include this TCP transport in Nostr endpoint adverts |
|
||||
| `transports.tcp.external_addr` | string | *(none)* | Explicit advertise-as override. Bare IP or full `host:port`. **Required** when `bind_addr` is wildcard (e.g. `"0.0.0.0:443"`) and `advertise_on_nostr: true`, since TCP has no STUN equivalent for autodiscovery. Common on cloud 1:1 NAT / EIP setups where the public IP isn't bindable on the host. |
|
||||
|
||||
**Named instances.** Like other transports, multiple TCP instances can
|
||||
be configured with named sub-keys:
|
||||
@@ -399,6 +512,7 @@ Requires an external Tor daemon providing a SOCKS5 proxy. Three modes:
|
||||
| `transports.tor.max_inbound_connections` | usize | `64` | Maximum inbound connections via onion service. |
|
||||
| `transports.tor.directory_service.hostname_file` | string | `"/var/lib/tor/fips_onion_service/hostname"` | Path to Tor-managed hostname file containing the `.onion` address. |
|
||||
| `transports.tor.directory_service.bind_addr` | string | `"127.0.0.1:8443"` | Local bind address for the listener that Tor forwards inbound connections to. Must match `HiddenServicePort` target in `torrc`. |
|
||||
| `transports.tor.advertised_port` | u16 | `443` | Public-facing onion port published in Nostr overlay adverts. Must match the virtual port in torrc's `HiddenServicePort <port> 127.0.0.1:<bind_port>` directive — that is the port other peers will use to reach this onion. |
|
||||
|
||||
**Named instances.** Like other transports, multiple Tor instances can
|
||||
be configured with named sub-keys for different SOCKS5 proxy endpoints.
|
||||
@@ -485,6 +599,101 @@ HiddenServiceDir /var/lib/tor/fips
|
||||
HiddenServicePort 8443 127.0.0.1:8444
|
||||
```
|
||||
|
||||
### Nym (`transports.nym.*`)
|
||||
|
||||
Nym transport routes FIPS traffic through the Nym mixnet for
|
||||
metadata-resistant anonymity. Outbound-only: connections are made
|
||||
through a `nym-socks5-client` SOCKS5 proxy that must be running
|
||||
separately (e.g. as a service running alongside the fips daemon or as a
|
||||
container). There is no inbound listener — a Nym-only node initiates
|
||||
outbound links but is not reachable for unsolicited inbound handshakes.
|
||||
|
||||
| Parameter | Type | Default | Description |
|
||||
|-----------|------|---------|-------------|
|
||||
| `transports.nym.socks5_addr` | string | `"127.0.0.1:1080"` | `nym-socks5-client` SOCKS5 proxy address (host:port) |
|
||||
| `transports.nym.connect_timeout_ms` | u64 | `300000` | Outbound connect timeout in milliseconds. Mixnet SOCKS5 connections traverse 3 mix nodes with timing obfuscation and can take several minutes, so this is generous (300s). |
|
||||
| `transports.nym.mtu` | u16 | `1400` | Default MTU |
|
||||
| `transports.nym.startup_timeout_secs` | u64 | `120` | Seconds to wait for `nym-socks5-client` to become ready at startup before giving up |
|
||||
|
||||
**Named instances.** Like other transports, multiple Nym instances can
|
||||
be configured with named sub-keys for different SOCKS5 proxy endpoints.
|
||||
|
||||
### BLE (`transports.ble.*`)
|
||||
|
||||
Bluetooth Low Energy transport using L2CAP Connection-Oriented Channels.
|
||||
Linux + glibc only — at build time, `build.rs` probes for the BlueZ /
|
||||
`bluer` crate dependencies and sets the `bluer_available` `cfg`; the BLE
|
||||
runtime is gated behind `#[cfg(bluer_available)]`. There is no Cargo
|
||||
feature flag to toggle. On non-glibc Linux (musl) or non-Linux platforms,
|
||||
BLE config still parses but the transport runtime is absent and config
|
||||
entries become no-ops. Communicates with BlueZ via D-Bus through the
|
||||
`bluer` crate.
|
||||
|
||||
| Parameter | Type | Default | Description |
|
||||
|-----------|------|---------|-------------|
|
||||
| `transports.ble.adapter` | string | `"hci0"` | HCI adapter name |
|
||||
| `transports.ble.psm` | u16 | `0x0085` (133) | L2CAP Protocol/Service Multiplexer |
|
||||
| `transports.ble.mtu` | u16 | `2048` | Default MTU. Actual MTU is negotiated per-link during L2CAP connection setup. |
|
||||
| `transports.ble.max_connections` | usize | `7` | Maximum concurrent BLE connections |
|
||||
| `transports.ble.connect_timeout_ms` | u64 | `10000` | Outbound connect timeout in milliseconds |
|
||||
| `transports.ble.advertise` | bool | `true` | Broadcast BLE beacon advertisements for peer discovery |
|
||||
| `transports.ble.scan` | bool | `true` | Listen for BLE beacon advertisements from other nodes |
|
||||
| `transports.ble.auto_connect` | bool | `false` | Automatically connect to discovered peers |
|
||||
| `transports.ble.accept_connections` | bool | `true` | Accept incoming L2CAP connections |
|
||||
| `transports.ble.probe_cooldown_secs` | u64 | `30` | Cooldown before re-probing the same BLE address |
|
||||
|
||||
**Address format.** BLE peer addresses use the form
|
||||
`"adapter/device_address"` — for example, `"hci0/AA:BB:CC:DD:EE:FF"`.
|
||||
|
||||
**Advertising and scanning.** When `advertise` is enabled, the transport
|
||||
advertises the FIPS service UUID continuously so that nearby nodes can
|
||||
discover and connect via L2CAP. When `scan` is enabled, the transport
|
||||
continuously scans for other FIPS nodes' advertisements. Discovered
|
||||
peers are probed immediately (L2CAP connect + pubkey exchange) with a
|
||||
cooldown (`probe_cooldown_secs`) to prevent rapid re-probing of the same
|
||||
address. If two nodes probe each other at the same time (cross-probe),
|
||||
a deterministic tie-breaker based on NodeAddr comparison ensures only
|
||||
one connection is established.
|
||||
|
||||
**Connection pool.** The `max_connections` parameter limits the number of
|
||||
concurrent BLE connections. When the pool is full, the least-recently-used
|
||||
connection is evicted to make room for new connections.
|
||||
|
||||
### BLE Example
|
||||
|
||||
A node using BLE for local mesh discovery alongside UDP for internet peers:
|
||||
|
||||
```yaml
|
||||
node:
|
||||
identity:
|
||||
persistent: true
|
||||
|
||||
tun:
|
||||
enabled: true
|
||||
|
||||
transports:
|
||||
udp:
|
||||
bind_addr: "0.0.0.0:2121"
|
||||
ble:
|
||||
adapter: "hci0"
|
||||
advertise: true
|
||||
scan: true
|
||||
auto_connect: true
|
||||
accept_connections: true
|
||||
|
||||
peers:
|
||||
- npub: "npub1abc..."
|
||||
alias: "internet-peer"
|
||||
addresses:
|
||||
- transport: udp
|
||||
addr: "203.0.113.5:2121"
|
||||
connect_policy: auto_connect
|
||||
```
|
||||
|
||||
BLE peers on the local radio range are discovered automatically via
|
||||
beacons — no static peer entries needed. Internet peers still require
|
||||
explicit configuration.
|
||||
|
||||
## Peers (`peers[]`)
|
||||
|
||||
Static peer list. Each entry defines a peer to connect to.
|
||||
@@ -493,11 +702,107 @@ Static peer list. Each entry defines a peer to connect to.
|
||||
|-----------|------|---------|-------------|
|
||||
| `peers[].npub` | string | *(required)* | Peer's Nostr public key (npub-encoded) |
|
||||
| `peers[].alias` | string | *(none)* | Human-readable name for logging |
|
||||
| `peers[].addresses[].transport` | string | *(required)* | Transport type: `udp`, `tcp`, `ethernet`, or `tor` |
|
||||
| `peers[].addresses[].addr` | string | *(required)* | Transport address. UDP/TCP: `"host:port"` (IP or DNS hostname). Ethernet: `"interface/mac"` (e.g., `"eth0/aa:bb:cc:dd:ee:ff"`). Tor: `".onion:port"` or `"host:port"` |
|
||||
| `peers[].addresses` | list | `[]` | Transport addresses for the peer. May be left empty (or omitted) when `via_nostr: true`, in which case the daemon resolves endpoints from the peer's Nostr advert at dial time. |
|
||||
| `peers[].addresses[].transport` | string | *(required)* | Transport type: `udp`, `tcp`, `ethernet`, `tor`, or `ble` |
|
||||
| `peers[].addresses[].addr` | string | *(required)* | Transport address. UDP/TCP: `"host:port"` (IP or DNS hostname). Ethernet: `"interface/mac"` (e.g., `"eth0/aa:bb:cc:dd:ee:ff"`). BLE: `"adapter/device_address"` (e.g., `"hci0/AA:BB:CC:DD:EE:FF"`). Tor: `".onion:port"` or `"host:port"` |
|
||||
| `peers[].addresses[].priority` | u8 | `100` | Address priority (lower = preferred) |
|
||||
| `peers[].connect_policy` | string | `"auto_connect"` | Connection policy: `auto_connect`, `on_demand`, or `manual` |
|
||||
| `peers[].connect_policy` | string | `"auto_connect"` | Connection policy: `auto_connect`, `on_demand`, or `manual`. Note: `on_demand` and `manual` are reserved for future use; the only policy currently honored at runtime is `auto_connect`. |
|
||||
| `peers[].auto_reconnect` | bool | `true` | Automatically reconnect after MMP link-dead removal (exponential backoff, unlimited retries) |
|
||||
| `peers[].via_nostr` | bool | `false` | Append Nostr advert-derived endpoints after static addresses for this peer |
|
||||
|
||||
## Gateway (`gateway.*`)
|
||||
|
||||
The `gateway.*` block configures the optional `fips-gateway`
|
||||
service, which lets unmodified LAN hosts reach mesh destinations
|
||||
through DNS proxy + virtual-IP NAT (and, optionally, exposes
|
||||
LAN-side services back into the mesh through inbound port forwards).
|
||||
The gateway is a separate service from the FIPS daemon but reads the
|
||||
same `fips.yaml` file. The block is read only when `fips-gateway` is
|
||||
running; the `fips` daemon ignores it. Linux only — the field is
|
||||
gated behind `#[cfg(target_os = "linux")]`. For setup, see
|
||||
[../how-to/deploy-gateway.md](../how-to/deploy-gateway.md); for the
|
||||
end-to-end design, see
|
||||
[../design/fips-gateway.md](../design/fips-gateway.md).
|
||||
|
||||
| Parameter | Type | Default | Description |
|
||||
|-----------|------|---------|-------------|
|
||||
| `gateway.enabled` | bool | `false` | Enable the gateway. Must be `true` for `fips-gateway` to start. |
|
||||
| `gateway.pool` | string | *(required)* | Virtual IPv6 pool CIDR (e.g., `"fd01::/112"`). Must not overlap with the FIPS mesh address space (`fd00::/8`) or any address space already in use on the LAN. The `/112` size yields 65 536 virtual IPs, which is the gateway's hard cap regardless of CIDR width. |
|
||||
| `gateway.lan_interface` | string | *(required)* | LAN-facing network interface name (e.g., `"enp3s0"`). Used for proxy-NDP entry installation so LAN clients can resolve the link-layer address of allocated virtual IPs. |
|
||||
| `gateway.pool_grace_period` | u64 | `60` | Seconds a virtual-IP allocation is retained after its last referencing session ends, before the address is returned to the free pool. Larger values reduce churn for short-lived flows; smaller values reclaim addresses faster. |
|
||||
|
||||
### Gateway DNS (`gateway.dns.*`)
|
||||
|
||||
Settings for the gateway's DNS listener and its upstream link to the
|
||||
FIPS daemon's `.fips` resolver. The gateway proxies `.fips` queries to
|
||||
the daemon's resolver, which returns mesh addresses; the gateway then
|
||||
allocates a virtual IP from the pool and rewrites the response.
|
||||
Non-`.fips` queries are answered with `REFUSED`.
|
||||
|
||||
| Parameter | Type | Default | Description |
|
||||
|-----------|------|---------|-------------|
|
||||
| `gateway.dns.listen` | string | `"[::1]:5353"` | DNS listen address. The default binds IPv6 loopback on an unprivileged port, matching the canonical deployment where another resolver on the host (dnsmasq, systemd-resolved, BIND) holds port 53 and forwards `.fips` queries to the gateway over loopback. Bind on the LAN-side IP (e.g., `"192.168.1.1:53"`) or wildcard (`"[::]:53"`) only on hosts with no other resolver on 53 and where LAN clients query the gateway directly. See [../how-to/troubleshoot-gateway.md](../how-to/troubleshoot-gateway.md). |
|
||||
| `gateway.dns.upstream` | string | `"[::1]:5354"` | Upstream FIPS daemon resolver. **Must match the daemon's `dns.bind_addr` and `dns.port`.** Defaults match the daemon defaults (`::1:5354`). A v4 upstream (`"127.0.0.1:5354"`) cannot reach a daemon bound on `[::1]:5354` — Linux IPv6 sockets bound to explicit `::1` do not accept v4-mapped traffic. If you change the daemon's `dns.bind_addr`, update this field accordingly. |
|
||||
| `gateway.dns.ttl` | u32 | `60` | TTL in seconds on AAAA responses returned to LAN clients. Smaller values let the gateway recycle pool addresses faster; larger values reduce LAN-side query traffic. |
|
||||
|
||||
### Conntrack (`gateway.conntrack.*`)
|
||||
|
||||
Linux conntrack timeout overrides for the gateway's NAT table. These
|
||||
adjust the kernel-default timeouts for NAT sessions installed by the
|
||||
gateway. All values are in seconds; omit any field to inherit the
|
||||
gateway's built-in default (which itself usually matches the kernel
|
||||
default for that protocol).
|
||||
|
||||
| Parameter | Type | Default | Description |
|
||||
|-----------|------|---------|-------------|
|
||||
| `gateway.conntrack.tcp_established` | u64 | `432000` | TCP established-state timeout (5 days). Long-lived TCP flows (SSH, persistent HTTP) keep their NAT mapping alive for at least this long without traffic. |
|
||||
| `gateway.conntrack.udp_timeout` | u64 | `30` | UDP unreplied timeout. Applied until reply traffic is observed in the reverse direction. |
|
||||
| `gateway.conntrack.udp_assured` | u64 | `180` | UDP assured (bidirectional) timeout. Applied once reply traffic has been observed. |
|
||||
| `gateway.conntrack.icmp_timeout` | u64 | `30` | ICMP echo / error timeout. |
|
||||
|
||||
### Inbound Port Forwards (`gateway.port_forwards[]`)
|
||||
|
||||
Optional list of inbound port-forward rules. Each rule maps a TCP or
|
||||
UDP port on the gateway's `fips0` mesh-side address to a `host:port`
|
||||
on the LAN. Mesh peers connect to the gateway's mesh address on the
|
||||
listen port; the gateway terminates the connection and forwards the
|
||||
payload to the LAN target. This is the inverse of the outbound mode:
|
||||
the LAN service is exposed to the mesh, not the other way around. See
|
||||
[../how-to/deploy-gateway.md](../how-to/deploy-gateway.md) for the
|
||||
operator recipe.
|
||||
|
||||
| Parameter | Type | Default | Description |
|
||||
|-----------|------|---------|-------------|
|
||||
| `gateway.port_forwards[].listen_port` | u16 | *(required)* | Port on `fips0` that mesh peers connect to. Must be non-zero. The `(listen_port, proto)` pair must be unique across the list. |
|
||||
| `gateway.port_forwards[].proto` | string | *(required)* | Transport protocol: `tcp` or `udp`. |
|
||||
| `gateway.port_forwards[].target` | string | *(required)* | LAN destination as IPv6 `[addr]:port` (e.g., `"[fd12:3456::10]:80"`). IPv4 targets are rejected at config-load time. |
|
||||
|
||||
### Gateway Example
|
||||
|
||||
A typical gateway with both outbound (LAN-to-mesh) and inbound
|
||||
(mesh-to-LAN) modes enabled:
|
||||
|
||||
```yaml
|
||||
gateway:
|
||||
enabled: true
|
||||
pool: "fd01::/112"
|
||||
lan_interface: "enp3s0"
|
||||
dns:
|
||||
listen: "[::1]:5353"
|
||||
upstream: "[::1]:5354"
|
||||
ttl: 60
|
||||
pool_grace_period: 60
|
||||
conntrack:
|
||||
tcp_established: 432000
|
||||
udp_assured: 180
|
||||
port_forwards:
|
||||
- listen_port: 8080
|
||||
proto: tcp
|
||||
target: "[fd12:3456::10]:80"
|
||||
- listen_port: 5353
|
||||
proto: udp
|
||||
target: "[fd12:3456::10]:53"
|
||||
```
|
||||
|
||||
## Minimal Example
|
||||
|
||||
@@ -535,7 +840,7 @@ peers:
|
||||
### Mixed UDP + Ethernet Example
|
||||
|
||||
A node bridging internet peers (UDP) and a local Ethernet segment with
|
||||
beacon discovery:
|
||||
neighbor beacons:
|
||||
|
||||
```yaml
|
||||
node:
|
||||
@@ -551,7 +856,7 @@ transports:
|
||||
mtu: 1472
|
||||
ethernet:
|
||||
interface: "eth0"
|
||||
discovery: true
|
||||
listen: true
|
||||
announce: true
|
||||
auto_connect: true
|
||||
accept_connections: true
|
||||
@@ -621,13 +926,14 @@ node:
|
||||
identity_size: 10000
|
||||
discovery:
|
||||
ttl: 64
|
||||
timeout_secs: 10
|
||||
attempt_timeouts_secs: [1, 2, 4, 8]
|
||||
recent_expiry_secs: 10
|
||||
retry_interval_secs: 5
|
||||
max_attempts: 2
|
||||
backoff_base_secs: 30
|
||||
backoff_max_secs: 300
|
||||
backoff_base_secs: 0
|
||||
backoff_max_secs: 0
|
||||
forward_min_interval_secs: 2
|
||||
# lan: # uncomment to enable mDNS LAN discovery
|
||||
# enabled: true # opt-in, default false
|
||||
# scope: "my-mesh" # optional per-network scope filter
|
||||
tree:
|
||||
announce_min_interval_ms: 500
|
||||
parent_hysteresis: 0.2 # cost improvement fraction for parent switch
|
||||
@@ -638,6 +944,7 @@ node:
|
||||
flap_dampening_secs: 120 # extended hold-down on flap
|
||||
bloom:
|
||||
update_debounce_ms: 500
|
||||
max_inbound_fpr: 0.20 # antipoison cap on inbound FilterAnnounce FPR
|
||||
session:
|
||||
default_ttl: 64
|
||||
pending_packets_per_dest: 16
|
||||
@@ -676,7 +983,7 @@ tun:
|
||||
|
||||
dns:
|
||||
enabled: true
|
||||
bind_addr: "127.0.0.1"
|
||||
bind_addr: "::1"
|
||||
port: 5354
|
||||
ttl: 300
|
||||
|
||||
@@ -692,7 +999,7 @@ transports:
|
||||
# mtu: null # null = interface MTU - 3 (typically 1497)
|
||||
# recv_buf_size: 2097152 # 2 MB
|
||||
# send_buf_size: 2097152 # 2 MB
|
||||
# discovery: true # listen for beacons
|
||||
# listen: true # listen for beacons
|
||||
# announce: false # broadcast beacons
|
||||
# auto_connect: false # connect to discovered peers
|
||||
# accept_connections: false # accept inbound handshakes
|
||||
@@ -717,9 +1024,26 @@ transports:
|
||||
# # cookie_path: "/var/run/tor/control.authcookie"
|
||||
# # directory mode (inbound via Tor-managed onion service):
|
||||
# # directory_service:
|
||||
# # hostname_file: "/var/lib/tor/fips/hostname"
|
||||
# # bind_addr: "127.0.0.1:8444"
|
||||
# # hostname_file: "/var/lib/tor/fips_onion_service/hostname"
|
||||
# # bind_addr: "127.0.0.1:8443"
|
||||
# # max_inbound_connections: 64
|
||||
# # advertised_port: 443 # public-facing onion port for Nostr adverts
|
||||
# nym: # uncomment to enable Nym mixnet transport (outbound-only)
|
||||
# socks5_addr: "127.0.0.1:1080" # nym-socks5-client SOCKS5 proxy address
|
||||
# connect_timeout_ms: 300000 # connect timeout (300s for mixnet)
|
||||
# mtu: 1400 # default MTU
|
||||
# startup_timeout_secs: 120 # wait for nym-socks5-client to be ready
|
||||
# ble: # uncomment to enable BLE transport (Linux only, requires BlueZ)
|
||||
# adapter: "hci0" # HCI adapter name
|
||||
# psm: 0x0085 # L2CAP PSM (133)
|
||||
# mtu: 2048 # default MTU (negotiated per-link)
|
||||
# max_connections: 7 # max concurrent BLE connections
|
||||
# connect_timeout_ms: 10000 # outbound connect timeout
|
||||
# advertise: true # broadcast BLE beacons
|
||||
# scan: true # listen for BLE beacons
|
||||
# auto_connect: false # connect to discovered peers
|
||||
# accept_connections: true # accept incoming L2CAP connections
|
||||
# probe_cooldown_secs: 30 # cooldown before re-probing same address
|
||||
|
||||
peers: # static peer list
|
||||
# - npub: "npub1..."
|
||||
192
docs/reference/control-socket.md
Normal file
@@ -0,0 +1,192 @@
|
||||
# Control Socket Protocol
|
||||
|
||||
The FIPS daemon and `fips-gateway` each expose a local control socket
|
||||
that accepts line-delimited JSON requests and returns line-delimited
|
||||
JSON responses. `fipsctl` and `fipstop` are clients of this protocol;
|
||||
operators can also drive it directly with any tool that can speak
|
||||
length-bounded JSON over a stream socket.
|
||||
|
||||
## Connection
|
||||
|
||||
### Linux / macOS
|
||||
|
||||
A Unix domain socket. The default path is resolved in this order:
|
||||
|
||||
1. `/run/fips/control.sock` (or `/run/fips/gateway.sock` for the
|
||||
gateway), if `/run/fips` exists. This is what the `fips.service`
|
||||
systemd unit creates.
|
||||
2. `$XDG_RUNTIME_DIR/fips/control.sock` otherwise.
|
||||
3. `/tmp/fips-control.sock` if neither of the above is available.
|
||||
|
||||
The daemon `chown`s the socket file and its parent directory to the
|
||||
`fips` group at bind time and sets mode `0770`. Members of the `fips`
|
||||
group can therefore connect without root. Add a user with
|
||||
`sudo usermod -aG fips $USER` (re-login required).
|
||||
|
||||
The path can be overridden at the daemon side via
|
||||
`node.control.socket_path` in the YAML config, and at the client side
|
||||
via `fipsctl -s PATH` or `fipstop -s PATH`.
|
||||
|
||||
### Windows
|
||||
|
||||
A TCP listener bound to `127.0.0.1`. The daemon's port is `21210` by
|
||||
default; the gateway's is `21211`. Only loopback connections are
|
||||
accepted. Override via `node.control.socket_path` (which takes a port
|
||||
number string on Windows).
|
||||
|
||||
Windows TCP does not provide filesystem-level ACLs — any local user
|
||||
can connect. See the security note in
|
||||
[configuration.md](configuration.md#control-socket-nodecontrol).
|
||||
|
||||
## Request Format
|
||||
|
||||
One JSON object per line, terminated by `\n`. Maximum request size is
|
||||
4096 bytes; longer requests are dropped with `request too large`.
|
||||
|
||||
```json
|
||||
{"command": "<name>", "params": {<object>}}
|
||||
```
|
||||
|
||||
| Field | Type | Required | Description |
|
||||
| ----- | ---- | -------- | ----------- |
|
||||
| `command` | string | yes | Command name. See [Daemon command catalog](#daemon-command-catalog) and [Gateway command catalog](#gateway-command-catalog). |
|
||||
| `params` | object | only for commands that take parameters | Parameter object. Unknown fields are ignored; missing required fields produce an error response. |
|
||||
|
||||
Unknown top-level fields in the request are silently ignored.
|
||||
|
||||
## Response Format
|
||||
|
||||
One JSON object per line.
|
||||
|
||||
```json
|
||||
{"status": "ok", "data": {<object>}}
|
||||
{"status": "error", "message": "<reason>"}
|
||||
```
|
||||
|
||||
| Field | Type | When present |
|
||||
| ----- | ---- | ------------ |
|
||||
| `status` | string | always; one of `"ok"` or `"error"`. |
|
||||
| `data` | object | on `ok` responses. |
|
||||
| `message` | string | on `error` responses. |
|
||||
|
||||
### I/O timeouts
|
||||
|
||||
The daemon enforces a 5-second timeout for both the request read and
|
||||
the response write. If the connection idles longer than that, the
|
||||
daemon closes it with no response.
|
||||
|
||||
### Common error messages
|
||||
|
||||
| Message | Cause |
|
||||
| ------- | ----- |
|
||||
| `empty request` | Connection closed before a newline was received. |
|
||||
| `invalid request: <serde error>` | Malformed JSON or missing `command`. |
|
||||
| `request too large` | Request exceeded 4096 bytes. |
|
||||
| `read timeout` / `read error: ...` | Slow client or transport failure. |
|
||||
| `unknown command: <name>` | Command not registered with this daemon. |
|
||||
| `missing params for <name>` | Command requires `params` but none were provided. |
|
||||
| `missing '<field>' parameter` | Required parameter missing. |
|
||||
| `query timeout` | Internal handler did not respond within 5 seconds. |
|
||||
| `node shutting down` | Daemon is exiting. |
|
||||
| `gateway not yet initialized` | (Gateway socket only) snapshot has not been published yet. |
|
||||
|
||||
## Daemon Command Catalog
|
||||
|
||||
Read-only queries are dispatched in `src/control/queries.rs`;
|
||||
mutating commands are dispatched in `src/control/commands.rs`. The
|
||||
table below lists every command currently registered.
|
||||
|
||||
### Read-only queries
|
||||
|
||||
| Command | Params | `data` shape (top-level keys) |
|
||||
| ------- | ------ | ----------------------------- |
|
||||
| `show_status` | — | `version`, `npub`, `node_addr`, `ipv6_addr`, `state`, `is_leaf_only`, `is_root` (bool — this node is the spanning-tree root), `root` (hex node-addr of the current tree root), `persistent` (bool — identity is persisted, i.e. `persistent` set or an `nsec` configured), `peer_count`, `session_count`, `link_count`, `transport_count`, `connection_count`, `transport_peer_counts` (object mapping transport-type name to its connected-peer count; configured transports appear with `0`), `tun_state`, `tun_name`, `effective_ipv6_mtu`, `control_socket`, `pid`, `exe_path`, `uptime_secs`, `estimated_mesh_size`, `forwarding`, `sparklines`. |
|
||||
| `show_acl` | — | `allow_file`, `deny_file`, `enforcement_active`, `effective_mode`, `default_decision`, `allow_all`, `deny_all`, `allow_file_entries`, `deny_file_entries`, `allow_entries`, `deny_entries`. |
|
||||
| `show_peers` | — | `peers[]` — per-peer object: `node_addr`, `npub`, `display_name`, `ipv6_addr`, `connectivity`, `link_id`, `direction`, `transport_addr`, `transport_type`, `is_parent`, `is_child`, `tree_depth`, `effective_depth` (`tree_depth + link_cost` — the metric `evaluate_parent` ranks on; `null` when the peer has no coords, or is unmeasured while another peer has an SRTT sample, per the cold-start gate), `stats`, `noise`, `current_k_bit`, `mmp`, plus optional `nostr_traversal`, `rekey_in_progress`, `rekey_draining`. |
|
||||
| `show_links` | — | `links[]` — `link_id`, `transport_id`, `remote_addr`, `direction`, `state`, `created_at_ms`, `stats`. |
|
||||
| `show_tree` | — | `my_node_addr`, `root`, `root_npub` (bech32 npub of the current tree root), `is_root`, `depth`, `my_coords[]`, `parent`, `parent_display_name`, `declaration_sequence`, `declaration_signed`, `peer_tree_count`, `peers[]`, `stats`. |
|
||||
| `show_sessions` | — | `sessions[]` — `remote_addr`, `npub`, `display_name`, `state` (`established`, `initiating`, `awaiting_msg3`, `unknown`), `is_initiator`, `last_activity_ms`, `stats`, optional `mmp`, `current_k_bit`, `is_draining`. |
|
||||
| `show_bloom` | — | `own_node_addr`, `is_leaf_only`, `sequence`, `leaf_dependent_count`, `leaf_dependents[]`, `peer_filters[]`, `uptree_fill_ratio` (fill ratio of the last filter actually sent to the tree parent), `uptree_estimated_count` (cardinality estimate of that uptree filter — this node's whole subtree under split-horizon, not the mesh; both are `null` for a root node or before the first announce), `stats`. |
|
||||
| `show_mmp` | — | `peers[]` (link-layer per peer), `sessions[]` (session-layer per session). Each entry includes loss/RTT/ETX/goodput, smoothed values, trends. |
|
||||
| `show_cache` | — | `count`, `max_entries`, `fill_ratio`, `default_ttl_ms`, `expired`, `avg_age_ms`, `entries[]` — per-destination coords, depth, age, last-used, optional `path_mtu`. |
|
||||
| `show_connections` | — | `connections[]` — pending handshakes: `link_id`, `direction`, `handshake_state`, `started_at_ms`, `idle_ms`, `resend_count`, optional `expected_peer`. |
|
||||
| `show_transports` | — | `transports[]` — `transport_id`, `type`, `state`, `mtu`, `name`, `local_addr`, optional `tor_mode`, `onion_address`, `tor_monitoring`, `stats`. |
|
||||
| `show_routing` | — | `coord_cache_entries`, `identity_cache_entries`, `pending_lookups[]`, `pending_tun_destinations`, `pending_tun_packets`, `recent_requests`, `retries[]`, `forwarding`, `discovery` (request/response sub-counters; includes `req_deduplicated` — requests suppressed as recent duplicates — and `req_dedup_cache_full` — requests admitted because the dedup cache was full), `error_signals`, `congestion`. |
|
||||
| `show_identity_cache` | — | `entries[]`, `count`, `max_entries`. Each entry: `node_addr`, `npub`, `display_name`, `ipv6_addr`, `last_seen_ms`, `age_ms`. |
|
||||
| `show_listening_sockets` | — | `fips0_addr`, `firewall_active` (bool — `inet fips` table loaded), `sockets[]`. Each entry: `proto` (`tcp` / `udp`), `local_addr` (`::` or the node's fd00::/8 address), `port`, `pid` (nullable), `process` (nullable), `wildcard_bind` (bool — `local_addr == ::`), `filter` (`accept` / `drop` / `unknown` / `no_firewall`). Linux-only; returns an empty `sockets[]` on other platforms. |
|
||||
| `show_stats_list` | — | `metrics[]` (each with `name`, `unit`, `scope`), `fast_ring_seconds`, `slow_ring_minutes`, `peer_retention_seconds`. |
|
||||
| `show_metrics` | — | Flat snapshot of every counter family in the metrics registry: `forwarding`, `discovery`, `tree`, `bloom`, `congestion`, `errors`. Each value is that family's counter snapshot object. Counter-only — gauges/histograms that need the live node are excluded. Served off the main loop. Silent-rejection sites classify their reason as a typed `RejectReason` and increment the matching per-family counter exposed here — see [Rejection reasons](#rejection-reasons). |
|
||||
| `show_stats_history` | `metric` (req), `peer` (req for per-peer metrics), `window` (`<N>s` / `<N>m` / `<N>h`, default `10m`), `granularity` (`1s` / `1m`, default `1s`) | A single `Series`: `metric`, `unit`, `granularity_seconds`, `values[]`. |
|
||||
| `show_stats_all_history` | `peer` (optional npub), `window`, `granularity` | `granularity_seconds`, `window_seconds`, `peer`, `series[]` (one per metric). |
|
||||
| `show_stats_peers` | — | `peers[]`, `count`. Each entry: `npub`, `node_addr`, `display_name`, `is_active`, `first_seen_secs_ago`, `last_contact_secs_ago`. |
|
||||
| `show_stats_history_all_peers` | `metric` (req per-peer name), `window`, `granularity` | `metric`, `unit`, `granularity_seconds`, `window_seconds`, `peers[]` (each with `node_addr`, `display_name`, `is_active`, `values[]`). |
|
||||
|
||||
The schema of each query response is pinned by snapshot tests in
|
||||
`src/control/snapshots/`; intentional schema changes regenerate those
|
||||
fixtures.
|
||||
|
||||
### Rejection reasons
|
||||
|
||||
Silent-rejection paths across the node classify why a message was
|
||||
dropped via a typed `RejectReason` rather than only logging it, so the
|
||||
*what* of a rejection is visible in the counter snapshots above. The
|
||||
top-level reason set has eight families, mirroring the protocol-layer /
|
||||
subsystem split of the metrics:
|
||||
|
||||
- **Tree** — spanning-tree `TreeAnnounce` processing rejections.
|
||||
- **Bloom** — bloom-filter `FilterAnnounce` processing rejections.
|
||||
- **Discovery** — discovery request / response processing rejections.
|
||||
- **Handshake** — Noise handshake state-machine rejections.
|
||||
- **Session** — FSP session state-machine rejections.
|
||||
- **Mmp** — MMP link-layer rejections.
|
||||
- **Forwarding** — forwarding-path rejections (no-route, TTL, MTU).
|
||||
- **Transport** — transport-layer rejections (admission caps, framing).
|
||||
|
||||
Each rejection increments the corresponding counter in its family's
|
||||
stats, surfaced through `show_metrics` (the `tree`, `bloom`,
|
||||
`discovery`, and `forwarding` families carry their own counters; the
|
||||
`errors` family and the remaining subsystem counters carry the rest).
|
||||
The full per-family variant list lives in `src/node/reject.rs`; it is
|
||||
not reproduced here to avoid duplicating the source.
|
||||
|
||||
### Mutating commands
|
||||
|
||||
| Command | Required params | Behaviour |
|
||||
| ------- | --------------- | --------- |
|
||||
| `connect` | `npub` (bech32), `address` (transport endpoint), `transport` (`udp`, `tcp`, `tor`, `nym`, `ethernet`) | Asks the node to dial the peer over the named transport. The named transport must be configured and running. Returns the API result on success or an error string on failure. |
|
||||
| `disconnect` | `npub` (bech32) | Asks the node to drop the link to the named peer. |
|
||||
|
||||
Both commands run on the daemon's main task and may block briefly
|
||||
while the node mutates its state.
|
||||
|
||||
## Gateway Command Catalog
|
||||
|
||||
`fips-gateway` exposes a separate control socket with its own command
|
||||
set. Dispatch lives in `src/gateway/control.rs`.
|
||||
|
||||
| Command | Params | `data` shape |
|
||||
| ------- | ------ | ------------ |
|
||||
| `show_gateway` | — | `pool_total`, `pool_allocated`, `pool_active`, `pool_draining`, `pool_free`, `nat_mappings`, `dns_listen`, `uptime_secs`, `pool_cidr`, `lan_interface`, `dns_upstream`, `dns_ttl`, `pool_grace_period`. |
|
||||
| `show_mappings` | — | `mappings[]` — `virtual_ip`, `mesh_addr`, `node_addr`, `dns_name`, `state` (`Allocated`, `Active`, `Draining`), `sessions`, `age_secs`, `last_ref_secs`. |
|
||||
|
||||
Until the first snapshot has been published (very early in startup),
|
||||
both commands return `gateway not yet initialized`.
|
||||
|
||||
## Driving the Socket Directly
|
||||
|
||||
```sh
|
||||
# Linux / macOS
|
||||
echo '{"command":"show_status"}' | sudo nc -U /run/fips/control.sock
|
||||
|
||||
# Windows (PowerShell with a TCP-capable tool of your choice)
|
||||
```
|
||||
|
||||
The newline at the end of the request is required: the daemon reads
|
||||
one line per connection. The connection is closed after the single
|
||||
response is written.
|
||||
|
||||
## See also
|
||||
|
||||
- [`fipsctl`](cli-fipsctl.md) — full-featured client.
|
||||
- [`fipstop`](cli-fipstop.md) — read-only TUI.
|
||||
- [configuration.md](configuration.md) — `node.control.*` keys.
|
||||
|
Before Width: | Height: | Size: 1.6 KiB After Width: | Height: | Size: 1.6 KiB |
|
Before Width: | Height: | Size: 2.3 KiB After Width: | Height: | Size: 2.3 KiB |
|
Before Width: | Height: | Size: 2.0 KiB After Width: | Height: | Size: 2.0 KiB |
|
Before Width: | Height: | Size: 1.0 KiB After Width: | Height: | Size: 1.0 KiB |
|
Before Width: | Height: | Size: 5.6 KiB After Width: | Height: | Size: 5.6 KiB |
|
Before Width: | Height: | Size: 1.8 KiB After Width: | Height: | Size: 1.8 KiB |
|
Before Width: | Height: | Size: 3.2 KiB After Width: | Height: | Size: 3.2 KiB |
|
Before Width: | Height: | Size: 2.6 KiB After Width: | Height: | Size: 2.6 KiB |
|
Before Width: | Height: | Size: 6.0 KiB After Width: | Height: | Size: 6.0 KiB |
|
Before Width: | Height: | Size: 4.2 KiB After Width: | Height: | Size: 4.2 KiB |
|
Before Width: | Height: | Size: 3.8 KiB After Width: | Height: | Size: 3.8 KiB |
|
Before Width: | Height: | Size: 3.4 KiB After Width: | Height: | Size: 3.4 KiB |
|
Before Width: | Height: | Size: 2.5 KiB After Width: | Height: | Size: 2.5 KiB |
|
Before Width: | Height: | Size: 4.3 KiB After Width: | Height: | Size: 4.3 KiB |
|
Before Width: | Height: | Size: 4.1 KiB After Width: | Height: | Size: 4.1 KiB |
|
Before Width: | Height: | Size: 2.6 KiB After Width: | Height: | Size: 2.6 KiB |
|
Before Width: | Height: | Size: 1000 B After Width: | Height: | Size: 1000 B |
|
Before Width: | Height: | Size: 6.8 KiB After Width: | Height: | Size: 6.8 KiB |
|
Before Width: | Height: | Size: 4.1 KiB After Width: | Height: | Size: 4.1 KiB |
|
Before Width: | Height: | Size: 2.9 KiB After Width: | Height: | Size: 2.9 KiB |
|
Before Width: | Height: | Size: 2.9 KiB After Width: | Height: | Size: 2.9 KiB |
|
Before Width: | Height: | Size: 2.2 KiB After Width: | Height: | Size: 2.2 KiB |
|
Before Width: | Height: | Size: 2.9 KiB After Width: | Height: | Size: 2.9 KiB |
|
Before Width: | Height: | Size: 3.5 KiB After Width: | Height: | Size: 3.5 KiB |
197
docs/reference/nostr-events.md
Normal file
@@ -0,0 +1,197 @@
|
||||
# Nostr Event Reference
|
||||
|
||||
The Nostr-protocol surface FIPS uses for discovery and signaling. For
|
||||
the design of the discovery runtime and the rationale behind these
|
||||
event shapes, see
|
||||
[../design/fips-nostr-discovery.md](../design/fips-nostr-discovery.md).
|
||||
For operator activation recipes, see
|
||||
[../how-to/enable-nostr-discovery.md](../how-to/enable-nostr-discovery.md).
|
||||
|
||||
FIPS uses three Nostr event kinds:
|
||||
|
||||
| Kind | Name | Encryption | Storage | Purpose |
|
||||
| ---- | ---- | ---------- | ------- | ------- |
|
||||
| 37195 | Overlay advert | None (signed only) | Replaceable | Publish reachable transport endpoints |
|
||||
| 21059 | Traversal signaling | NIP-44 inside NIP-59 gift wrap | Ephemeral | Carry `TraversalOffer`/`TraversalAnswer` payloads |
|
||||
| 10050 | NIP-17 inbox relay list | None (signed only) | Replaceable | Tell dialers where to publish offers |
|
||||
|
||||
All three are signed with the node's FIPS identity key (the same
|
||||
secp256k1 keypair Nostr uses); there is no separate Nostr key.
|
||||
|
||||
## Kind 37195 — Overlay Advert
|
||||
|
||||
A parameterized replaceable event in the application-defined
|
||||
replaceable range `30000–39999` (the digits visually spell `FIPS`:
|
||||
7=F, 1=I, 9=P, 5=S). Each node has a single in-place-updatable advert
|
||||
under its identity.
|
||||
|
||||
### Tags
|
||||
|
||||
- `d` — fixed to the literal `fips-overlay-v1` (the application
|
||||
identifier baked into the binary). Together with `pubkey`, this
|
||||
identifies the unique replaceable event slot.
|
||||
- `protocol` — the configured `node.discovery.nostr.app` value
|
||||
(default `fips-overlay-v1`). Distinct from the `d` tag so the
|
||||
application string can evolve without breaking the replaceable
|
||||
event slot.
|
||||
- `version` — protocol version string (currently `"1"`).
|
||||
- `expiration` — NIP-40 expiration timestamp set to now +
|
||||
`node.discovery.nostr.advert_ttl_secs` (default 3600 seconds).
|
||||
Conforming relays stop serving the event after this time.
|
||||
|
||||
### Content
|
||||
|
||||
The event content is a JSON document shaped as `OverlayAdvert`:
|
||||
|
||||
```json
|
||||
{
|
||||
"identifier": "fips-overlay-v1",
|
||||
"version": 1,
|
||||
"endpoints": [
|
||||
{"transport": "udp", "addr": "203.0.113.45:2121"},
|
||||
{"transport": "tor", "addr": "xxxxx.onion:8443"},
|
||||
{"transport": "udp", "addr": "nat"}
|
||||
],
|
||||
"signalRelays": ["wss://relay.damus.io", "wss://nos.lol"],
|
||||
"stunServers": ["stun:stun.l.google.com:19302"]
|
||||
}
|
||||
```
|
||||
|
||||
Field semantics:
|
||||
|
||||
| Field | Type | Description |
|
||||
| ----- | ---- | ----------- |
|
||||
| `identifier` | string | Application namespace; must match the `d` tag. |
|
||||
| `version` | integer | Advert schema version (currently 1). |
|
||||
| `endpoints` | array | List of transport endpoints. Each is `{transport, addr}` where `transport` is `"udp"`, `"tcp"`, or `"tor"`, and `addr` is `"host:port"`, `".onion:port"`, or the literal `"nat"` (for UDP NAT-punch). |
|
||||
| `signalRelays` | array? | Optional. Relays the publisher prefers for offer/answer signaling. Present only when at least one endpoint is `udp:nat`. |
|
||||
| `stunServers` | array? | Optional. STUN servers the publisher uses for reflexive discovery. Present only when at least one endpoint is `udp:nat`. Informational — peers do not use these to choose their own STUN targets. |
|
||||
|
||||
### Signature scope
|
||||
|
||||
The Nostr event signature covers the standard Nostr event ID
|
||||
(serialized `[0, pubkey, created_at, kind, tags, content]`), so the
|
||||
content JSON, tags, kind, and timestamp are all bound to the signing
|
||||
identity.
|
||||
|
||||
### Replacement and deletion
|
||||
|
||||
Because kind 37195 is replaceable, publishing a new advert replaces
|
||||
the prior one in the same `(pubkey, d-tag)` slot. To withdraw an
|
||||
advert without publishing a successor, the node publishes a NIP-9
|
||||
kind 5 delete event referencing the prior advert.
|
||||
|
||||
## Kind 21059 — Traversal Signaling
|
||||
|
||||
An ephemeral event (kinds in the 20000–29999 range are not stored by
|
||||
conforming relays). Used to deliver gift-wrapped, NIP-44-encrypted
|
||||
`TraversalOffer` and `TraversalAnswer` payloads between dialer and
|
||||
responder during a UDP NAT hole-punch.
|
||||
|
||||
### Encryption envelope
|
||||
|
||||
The wire shape is the standard NIP-59 gift wrap:
|
||||
|
||||
1. **Rumor** — the unsigned `TraversalOffer`/`TraversalAnswer`
|
||||
payload (JSON), authored by the actual sender's identity.
|
||||
2. **Seal** — a kind 13 event whose content is the rumor
|
||||
NIP-44-encrypted to the recipient's pubkey, signed by the sender.
|
||||
3. **Gift wrap** — a kind 21059 event whose content is the seal
|
||||
NIP-44-encrypted to the recipient under an ephemeral key, signed
|
||||
by that ephemeral key. The outer `pubkey` of the kind 21059 event
|
||||
is the ephemeral identity, not the sender's real identity.
|
||||
|
||||
Only the intended recipient can decrypt the wrap to recover the seal,
|
||||
and only the recipient can decrypt the seal to recover the rumor.
|
||||
|
||||
### Wrapped payloads
|
||||
|
||||
The `TraversalOffer` carries:
|
||||
|
||||
- `type` — message-type tag.
|
||||
- `sessionId` — unique identifier correlating offer and answer.
|
||||
- `senderNpub` / `recipientNpub` — bech32-encoded pubkeys, repeated
|
||||
inside the encrypted payload (the outer wrap pubkey is ephemeral).
|
||||
- `issuedAt` / `expiresAt` — Unix-ms timestamps; `expiresAt` is
|
||||
`issuedAt + signal_ttl_secs * 1000`.
|
||||
- `nonce` — random per-offer value.
|
||||
- `reflexiveAddress` — `{protocol, ip, port}` observed via STUN, or
|
||||
`null` if STUN failed or returned no usable address.
|
||||
- `localAddresses` — array of `{protocol, ip, port}` private
|
||||
candidates, populated when `share_local_candidates` is enabled.
|
||||
- `stunServer` — the STUN server actually used (informational).
|
||||
|
||||
The `TraversalAnswer` echoes `sessionId` and carries:
|
||||
|
||||
- `type`, `senderNpub`, `recipientNpub`, `issuedAt`, `expiresAt`,
|
||||
`nonce` — same shape as the offer.
|
||||
- `inReplyTo` — the offer's event id.
|
||||
- `accepted` — boolean; false when the responder has no usable
|
||||
addresses.
|
||||
- `reflexiveAddress` and `localAddresses` — the responder's
|
||||
candidates, in the same shape as the offer.
|
||||
- `stunServer` — informational.
|
||||
- `punch` — a `PunchHint { startAtMs, intervalMs, durationMs }`
|
||||
telling both sides when to begin probing and how aggressively.
|
||||
Absent on rejected offers.
|
||||
- `reason` — optional rejection string when `accepted` is false.
|
||||
- `offerReceivedAt` — optional responder wall-clock (Unix ms) at
|
||||
the moment it received the offer; the initiator uses this to
|
||||
derive a clock-skew estimate.
|
||||
|
||||
### Relay selection
|
||||
|
||||
Dialer publishes offers to the recipient's NIP-17 inbox relays (kind
|
||||
10050) when available; otherwise to the local
|
||||
`node.discovery.nostr.dm_relays` list. The responder publishes the
|
||||
answer back through the same relay channel.
|
||||
|
||||
## Kind 10050 — NIP-17 Inbox Relay List
|
||||
|
||||
A standard NIP-17 event used by FIPS to advertise which relays this
|
||||
node prefers for receiving direct-message-style signaling — for FIPS,
|
||||
the gift-wrapped traversal offers (kind 21059).
|
||||
|
||||
This is **NIP-17** (`kind 10050`, inbox relays for DM delivery), not
|
||||
NIP-65 (`kind 10002`, general read/write relay list). The two serve
|
||||
different purposes:
|
||||
|
||||
- Kind 10002 (NIP-65) — general read/write relays for ordinary event
|
||||
publication and subscription.
|
||||
- Kind 10050 (NIP-17) — relays the recipient prefers for receiving
|
||||
DM-shaped (NIP-59 wrapped) events.
|
||||
|
||||
FIPS publishes its own kind 10050 on startup so dialers can discover
|
||||
where to send traversal offers. When dialing a peer, FIPS first
|
||||
fetches the peer's kind 10050 from the peer's `advert_relays`; on
|
||||
fetch failure it falls back to the local `dm_relays` list.
|
||||
|
||||
### Tags
|
||||
|
||||
Standard NIP-17 form: each relay is encoded as an `r` tag whose
|
||||
single value is the relay URL.
|
||||
|
||||
```text
|
||||
["r", "wss://relay.damus.io"]
|
||||
["r", "wss://nos.lol"]
|
||||
```
|
||||
|
||||
### Content
|
||||
|
||||
Empty per NIP-17.
|
||||
|
||||
## See also
|
||||
|
||||
- [../design/fips-nostr-discovery.md](../design/fips-nostr-discovery.md)
|
||||
— discovery runtime design and the five activation scenarios
|
||||
- [../how-to/enable-nostr-discovery.md](../how-to/enable-nostr-discovery.md)
|
||||
— operator recipes
|
||||
- [../tutorials/resolve-peers-via-nostr.md](../tutorials/resolve-peers-via-nostr.md),
|
||||
[../tutorials/advertise-your-node.md](../tutorials/advertise-your-node.md),
|
||||
[../tutorials/open-discovery.md](../tutorials/open-discovery.md)
|
||||
— hand-held tutorial walkthroughs of the three capabilities
|
||||
- [../design/port-advertisement-and-nat-traversal.md](../design/port-advertisement-and-nat-traversal.md)
|
||||
— generic protocol reference (event tags, NIP usage, on-the-wire
|
||||
offer/answer schema), with FIPS values as worked examples
|
||||
- [security.md](security.md) — how the FIPS identity key signs both
|
||||
adverts and Noise handshakes
|
||||
235
docs/reference/security.md
Normal file
@@ -0,0 +1,235 @@
|
||||
# Security Reference
|
||||
|
||||
Consolidated security reference covering the nftables baseline, peer
|
||||
ACL file format, cryptographic primitives, rekey defaults, replay
|
||||
window, filesystem permissions, threat-resistance matrix, and default
|
||||
network exposures per transport. For the threat-model design and
|
||||
rationale, see [../design/fips-security.md](../design/fips-security.md).
|
||||
For the operator activation steps and drop-in recipes, see
|
||||
[../how-to/enable-mesh-firewall.md](../how-to/enable-mesh-firewall.md).
|
||||
|
||||
## nftables Baseline
|
||||
|
||||
The shipped baseline is `/etc/fips/fips.nft`. It defines a single
|
||||
nftables table `inet fips` with one chain hooked at `input`, structured
|
||||
as follows:
|
||||
|
||||
| Step | Rule | Effect |
|
||||
| ---- | ---- | ------ |
|
||||
| 1 | `iifname != "fips0" return` | Match only traffic arriving on `fips0`; everything else short-circuits. |
|
||||
| 2 | `ct state established,related accept` | Allow conntrack replies and related ICMPv6 errors. |
|
||||
| 3 | `icmpv6 type echo-request accept` | Allow IPv6 echo (ping6 reachability). |
|
||||
| 4 | `include "/etc/fips/fips.d/*.nft"` | Splice in operator drop-ins (empty matches nothing). |
|
||||
| 5 | `counter drop` | Default-deny everything else; counter increments on every drop. |
|
||||
|
||||
Outbound from `fips0` is unrestricted. The baseline is a documented
|
||||
dpkg conffile — operator edits to `/etc/fips/fips.nft` are preserved
|
||||
across upgrades.
|
||||
|
||||
The systemd unit is `fips-firewall.service` (oneshot). It is **not**
|
||||
enabled by default; activation is an explicit operator gesture
|
||||
documented in
|
||||
[../how-to/enable-mesh-firewall.md](../how-to/enable-mesh-firewall.md).
|
||||
|
||||
## Drop-In File Format
|
||||
|
||||
Operator extensions live under `/etc/fips/fips.d/` with the `.nft`
|
||||
suffix. Each file is included inline into the `inbound` chain at the
|
||||
marked point and may contain any nftables rule lines valid in that
|
||||
context.
|
||||
|
||||
Naming convention: `<purpose>-from-<source>.nft` keeps drop-ins easy
|
||||
to scan. Examples shipped in the design discussion:
|
||||
|
||||
- `ssh-from-bastion.nft` — accept TCP/22 from a single mesh-node address
|
||||
- `http-from-cluster.nft` — accept TCP/80 from a `/64` mesh-address prefix
|
||||
- `dns-public.nft` — accept UDP/53 and TCP/53 from any mesh node
|
||||
- `git-from-trusted.nft` — accept TCP/9418 from a set of mesh-node addresses
|
||||
|
||||
After editing, reload via
|
||||
`sudo systemctl reload-or-restart fips-firewall.service` (or
|
||||
equivalently `sudo nft -f /etc/fips/fips.nft` since the file is
|
||||
idempotent).
|
||||
|
||||
## Cryptographic Primitives
|
||||
|
||||
| Component | Choice | Where Used |
|
||||
| --------- | ------ | ---------- |
|
||||
| Curve | secp256k1 | FMP IK, FSP XK, Schnorr signatures |
|
||||
| Diffie-Hellman | ECDH on secp256k1 (x-only normalized) | Noise IK, Noise XK |
|
||||
| AEAD | ChaCha20-Poly1305 | FMP link encryption, FSP session encryption |
|
||||
| Hash | SHA-256 | NodeAddr derivation, Noise transcript |
|
||||
| Key derivation | HKDF-SHA256 | Noise key schedule |
|
||||
| Signatures | secp256k1 Schnorr | TreeAnnounce, LookupResponse proof, Nostr adverts |
|
||||
| Noise pattern (link) | `Noise_IK_secp256k1_ChaChaPoly_SHA256` | FMP link layer (IK with epoch payload) |
|
||||
| Noise pattern (session) | `Noise_XK_secp256k1_ChaChaPoly_SHA256` | FSP session layer (XK with epoch payload) |
|
||||
|
||||
These choices align with the Nostr cryptographic stack
|
||||
(secp256k1 + ChaCha20-Poly1305 + SHA-256) and the NIP-44 encrypted
|
||||
messaging standard.
|
||||
|
||||
## Rekey Defaults
|
||||
|
||||
Both link-layer and session-layer Noise sessions rekey under one of
|
||||
two triggers, configurable under `node.rekey.*`:
|
||||
|
||||
| Parameter | Default | Description |
|
||||
| --------- | ------- | ----------- |
|
||||
| `enabled` | `true` | Master switch. |
|
||||
| `after_secs` | `120` | Time-based rekey threshold. |
|
||||
| `after_messages` | `65536` | Message-count rekey threshold. |
|
||||
|
||||
In addition to the configurable triggers, the daemon retains the old
|
||||
session keys for a fixed **10-second drain window** after each
|
||||
cutover (compile-time constant `DRAIN_WINDOW_SECS` in
|
||||
`src/node/handlers/rekey.rs`). Rekey rotates the Noise key schedule
|
||||
and the session indices; old session keys are kept in
|
||||
`previous_session` for the drain window so in-flight packets
|
||||
encrypted under the old keys still decrypt.
|
||||
|
||||
## Replay Window
|
||||
|
||||
Both layers use explicit per-packet counters with a sliding bitmap
|
||||
window for replay protection. The bitmap is **2048 entries** at both
|
||||
layers — large enough to accommodate UDP reordering and packet loss
|
||||
without false-positive replay rejection. Counters older than the
|
||||
window are rejected. The same `ReplayWindow` and
|
||||
`decrypt_with_replay_check()` implementation is used at both the FMP
|
||||
and FSP layers.
|
||||
|
||||
## Peer ACL
|
||||
|
||||
Mesh-level ACL files at `/etc/fips/peers.allow` and
|
||||
`/etc/fips/peers.deny` give the operator allowlist/blocklist control
|
||||
over which npubs may complete the FMP Noise IK link handshake.
|
||||
|
||||
File format:
|
||||
|
||||
- One entry per line. An entry is either a bech32 `npub1...`,
|
||||
an alias defined in `/etc/fips/hosts`, or the literal `ALL`
|
||||
wildcard (case-insensitive).
|
||||
- Lines beginning with `#` are comments.
|
||||
- Blank lines are ignored.
|
||||
|
||||
Evaluation order (first match wins, default-allow on no match):
|
||||
|
||||
1. `peers.allow` — if the peer matches an entry here (or `ALL` is
|
||||
in `peers.allow`), the handshake is admitted, regardless of any
|
||||
`peers.deny` entry.
|
||||
2. `peers.deny` — if the peer matches an entry here (or `ALL` is
|
||||
in `peers.deny`), the handshake is refused.
|
||||
3. Otherwise the peer is admitted.
|
||||
|
||||
`peers.allow` is **not** an exclusive gate on its own: an unlisted
|
||||
peer falls through to step 3 and is admitted unless it appears in
|
||||
`peers.deny`. To turn `peers.allow` into a strict allowlist, place
|
||||
`ALL` in `peers.deny` so every unlisted peer is rejected at step 2.
|
||||
|
||||
The `ALL` wildcard makes the operator's posture explicit:
|
||||
|
||||
- `ALL` in `peers.allow` admits every peer (same effect as the
|
||||
default-allow behavior, but documented in the file).
|
||||
- `ALL` in `peers.deny` blocks every peer except those listed in
|
||||
`peers.allow` — the "allowlist-strict" posture.
|
||||
|
||||
In practice this collapses to a few common postures:
|
||||
|
||||
- **Default-allow with denylist**: leave `peers.allow` empty;
|
||||
populate `peers.deny`. All npubs may peer except those listed.
|
||||
- **Allowlist-strict**: populate `peers.allow` and put `ALL`
|
||||
in `peers.deny`. Only the listed npubs may peer; everyone else
|
||||
is rejected at step 2.
|
||||
|
||||
A populated `peers.allow` with an empty `peers.deny` is not a
|
||||
strict allowlist — it is equivalent to default-allow plus an
|
||||
explicit "always-admit" set. The strict variant requires `ALL`
|
||||
in `peers.deny`.
|
||||
|
||||
Aliases are resolved through `/etc/fips/hosts` at file-load
|
||||
time. If `peers.allow` lists `core-vm` and `/etc/fips/hosts`
|
||||
maps `core-vm` to a specific npub, that npub is admitted. If
|
||||
`core-vm` is later remapped to a different npub, the ACL
|
||||
re-resolves on the next mtime change. Operators should be aware
|
||||
that ACL semantics follow the `hosts`-file aliasing, not just
|
||||
the literal npubs visible in the file.
|
||||
|
||||
Both files are reloaded automatically when their mtime changes
|
||||
— no daemon restart or signal is needed. ACL evaluation runs
|
||||
after msg1 decryption but before any further peer-state
|
||||
mutation; rate-limited msg1s never reach the ACL.
|
||||
|
||||
## Filesystem Permissions
|
||||
|
||||
| Path | Owner | Mode | Purpose |
|
||||
| ---- | ----- | ---- | ------- |
|
||||
| `/etc/fips/fips.key` | root:root | `0600` | Persistent identity private key (sensitive). |
|
||||
| `/etc/fips/fips.pub` | root:root | `0644` | Public key (npub). |
|
||||
| `/etc/fips/fips.yaml` | root:root | `0644` | Daemon configuration (dpkg conffile). |
|
||||
| `/etc/fips/fips.nft` | root:root | `0644` | nftables baseline (dpkg conffile). |
|
||||
| `/etc/fips/fips.d/` | root:root | `0755` | Operator drop-in directory. |
|
||||
| `/etc/fips/hosts` | root:root | `0644` | Optional hostname → npub map (dpkg conffile). |
|
||||
| `/etc/fips/peers.allow` | root:root | `0644` | Optional peer allowlist. |
|
||||
| `/etc/fips/peers.deny` | root:root | `0644` | Optional peer denylist. |
|
||||
| `/run/fips/control.sock` | root:fips | `0770` | Control socket (members of `fips` group can use `fipsctl`). |
|
||||
| `/run/fips/` | root:fips | `0750` | Control socket parent directory. |
|
||||
|
||||
Adding a user to the `fips` group grants `fipsctl` access without
|
||||
requiring root. The daemon `chown`s the control socket and its parent
|
||||
directory at bind time.
|
||||
|
||||
## Threat-Resistance Matrix
|
||||
|
||||
The link layer's threat-resistance matrix is consolidated here from
|
||||
the FMP design document:
|
||||
|
||||
| Threat | Mitigation |
|
||||
| ------ | ---------- |
|
||||
| Connection exhaustion | Token-bucket rate limit + connection count limit |
|
||||
| CPU exhaustion (msg1 flood) | Rate limit before crypto operations |
|
||||
| Replay attacks | Counter-based nonces with sliding window (2048 entries) |
|
||||
| State confusion | Strict handshake state machine validation |
|
||||
| Spoofed encrypted packets | Index lookup + AEAD verification |
|
||||
| Spoofed msg2 | Index lookup + Noise ephemeral key binding |
|
||||
| Address spoofing | Cryptographic authority, not address-based |
|
||||
| Session correlation | Index rotation on rekey |
|
||||
| Inbound exposure on `fips0` | Default-deny nftables baseline (operator opt-in) |
|
||||
| Sybil identities | Discretionary peering + handshake rate limiting + optional peer ACL |
|
||||
| Eclipse attack | Diverse peering across independent operators and transports |
|
||||
| Unauthorized peer admission | Optional `peers.allow` allowlist consulted before handshake |
|
||||
|
||||
See [../design/fips-mesh-layer.md](../design/fips-mesh-layer.md) for
|
||||
the unauthenticated-attack-surface analysis (only handshake msg1 is
|
||||
reachable by unauthenticated parties), and
|
||||
[../design/fips-mesh-operation.md](../design/fips-mesh-operation.md#privacy-considerations)
|
||||
for the metadata-privacy model and the rejection of onion routing.
|
||||
|
||||
## Default Network Exposures by Transport
|
||||
|
||||
| Transport | Default Inbound | Default Bind | Opt-in |
|
||||
| --------- | --------------- | ------------ | ------ |
|
||||
| UDP | None until `bind_addr` set | `0.0.0.0:2121` typical | Operator sets `transports.udp.bind_addr` |
|
||||
| TCP | None until `bind_addr` set | None — outbound-only without bind | Operator sets `transports.tcp.bind_addr` |
|
||||
| Ethernet | Listens on configured interface (raw `AF_PACKET`) | EtherType 0x2121 on selected interface | Per-flag `listen`, `announce`, `auto_connect`, `accept_connections` |
|
||||
| Tor | None until `directory_service` configured | `127.0.0.1:8443` (loopback only) | Operator sets `transports.tor.directory_service` and configures `HiddenServiceDir` in `torrc` |
|
||||
| BLE | Off by default | n/a | Operator enables `transports.ble.*` |
|
||||
| Nostr discovery | Off by default | n/a (relay client, not a listener) | Operator sets `node.discovery.nostr.enabled: true` |
|
||||
|
||||
The mesh-layer `fips0` interface is reachable from any mesh node that
|
||||
can route to you, not only direct peers — your direct peers forward
|
||||
traffic from any reachable mesh node onto your `fips0`. The
|
||||
default-deny nftables baseline (operator opt-in) is the recommended
|
||||
way to restrict inbound traffic on `fips0`. See
|
||||
[../how-to/enable-mesh-firewall.md](../how-to/enable-mesh-firewall.md).
|
||||
|
||||
## See also
|
||||
|
||||
- [../design/fips-security.md](../design/fips-security.md) — threat
|
||||
model and design rationale for the `fips0` baseline
|
||||
- [../design/fips-mesh-layer.md](../design/fips-mesh-layer.md) — FMP
|
||||
link encryption, replay protection, rate limiting
|
||||
- [../design/fips-session-layer.md](../design/fips-session-layer.md)
|
||||
— FSP end-to-end encryption, Noise XK, replay window
|
||||
- [../how-to/enable-mesh-firewall.md](../how-to/enable-mesh-firewall.md)
|
||||
— operator activation and drop-in recipes
|
||||
- [configuration.md](configuration.md) — full `node.rekey.*`,
|
||||
`node.rate_limit.*` parameter tables
|
||||
111
docs/reference/transports.md
Normal file
@@ -0,0 +1,111 @@
|
||||
# Transport Statistics Reference
|
||||
|
||||
Per-transport statistics counter inventories. Counters are exposed
|
||||
through the daemon control socket (`fipsctl show transports`) and the
|
||||
`fipstop` operator UI. For the transport-layer design (services
|
||||
provided to FMP, transport categories, the trait surface, connection
|
||||
model), see
|
||||
[../design/fips-transport-layer.md](../design/fips-transport-layer.md).
|
||||
|
||||
All transports report counters via `fipsctl show transports`; the
|
||||
tables below are source-extracted from each transport's `stats.rs`
|
||||
module.
|
||||
|
||||
## UDP
|
||||
|
||||
| Counter | Description |
|
||||
| ------- | ----------- |
|
||||
| `packets_sent` / `bytes_sent` | Successful sends |
|
||||
| `packets_recv` / `bytes_recv` | Successful receives |
|
||||
| `send_errors` / `recv_errors` | Send/receive failures |
|
||||
| `mtu_exceeded` | Packets rejected for MTU violation |
|
||||
| `kernel_drops` | Kernel `SO_RXQ_OVFL` drop count (feeds ECN congestion detection) |
|
||||
|
||||
## TCP
|
||||
|
||||
| Counter | Description |
|
||||
| ------- | ----------- |
|
||||
| `packets_sent` / `bytes_sent` | Successful sends |
|
||||
| `packets_recv` / `bytes_recv` | Successful receives |
|
||||
| `send_errors` / `recv_errors` | Send/receive failures |
|
||||
| `mtu_exceeded` | Packets rejected for MTU violation |
|
||||
| `connections_established` | Successful outbound connections |
|
||||
| `connections_accepted` | Accepted inbound connections |
|
||||
| `connections_rejected` | Rejected inbound connections (limit exceeded) |
|
||||
| `connect_timeouts` | Connection timeout count |
|
||||
| `connect_refused` | Connection refused count |
|
||||
| `pool_inbound` | Current inbound connections held in the connection pool (gauge) |
|
||||
| `pool_outbound` | Current outbound connections held in the connection pool (gauge) |
|
||||
|
||||
## Ethernet
|
||||
|
||||
| Counter | Description |
|
||||
| ------- | ----------- |
|
||||
| `frames_sent` / `frames_recv` | Successful frame send/receive |
|
||||
| `bytes_sent` / `bytes_recv` | Byte counters |
|
||||
| `send_errors` / `recv_errors` | Send/receive failures |
|
||||
| `beacons_sent` / `beacons_recv` | Peer-discovery beacon traffic |
|
||||
| `frames_too_short` | Frames below minimum length, dropped |
|
||||
| `frames_too_long` | Frames above transport MTU, dropped |
|
||||
|
||||
## Tor
|
||||
|
||||
| Counter | Description |
|
||||
| ------- | ----------- |
|
||||
| `packets_sent` / `bytes_sent` | Successful sends |
|
||||
| `packets_recv` / `bytes_recv` | Successful receives |
|
||||
| `send_errors` / `recv_errors` | Send/receive failures |
|
||||
| `connections_established` | Successful SOCKS5 connections |
|
||||
| `connect_timeouts` | Connection timeout count |
|
||||
| `connect_refused` | Connection refused count |
|
||||
| `socks5_errors` | SOCKS5 protocol errors |
|
||||
| `mtu_exceeded` | Packets rejected for MTU violation |
|
||||
| `connections_accepted` | Accepted inbound connections via onion service |
|
||||
| `connections_rejected` | Rejected inbound connections (limit exceeded) |
|
||||
| `control_errors` | Tor control port errors |
|
||||
| `pool_inbound` | Current inbound connections held in the connection pool (gauge) |
|
||||
| `pool_outbound` | Current outbound connections held in the connection pool (gauge) |
|
||||
|
||||
## Nym
|
||||
|
||||
| Counter | Description |
|
||||
| ------- | ----------- |
|
||||
| `packets_sent` / `bytes_sent` | Successful sends |
|
||||
| `packets_recv` / `bytes_recv` | Successful receives |
|
||||
| `send_errors` / `recv_errors` | Send/receive failures |
|
||||
| `mtu_exceeded` | Packets rejected for MTU violation |
|
||||
| `connections_established` | Successful SOCKS5 connections through `nym-socks5-client` |
|
||||
| `connect_timeouts` | Connection timeout count |
|
||||
| `socks5_errors` | SOCKS5 protocol errors |
|
||||
|
||||
Nym is outbound-only (no inbound listener), so there are no
|
||||
`connections_accepted` / `connections_rejected` counters.
|
||||
|
||||
## Bluetooth
|
||||
|
||||
| Counter | Description |
|
||||
| ------- | ----------- |
|
||||
| `packets_sent` / `bytes_sent` | Successful L2CAP CoC sends |
|
||||
| `packets_recv` / `bytes_recv` | Successful L2CAP CoC receives |
|
||||
| `send_errors` / `recv_errors` | Send/receive failures |
|
||||
| `mtu_exceeded` | Packets rejected for MTU violation |
|
||||
| `connections_established` | Successful outbound L2CAP connections |
|
||||
| `connections_accepted` | Accepted inbound L2CAP connections |
|
||||
| `connections_rejected` | Rejected inbound (limit exceeded) |
|
||||
| `connect_timeouts` | Connection timeout count |
|
||||
| `pool_evictions` | Connection-pool entries evicted |
|
||||
| `advertisements_sent` | BLE advertisements emitted |
|
||||
| `scan_results` | BLE scan results observed |
|
||||
|
||||
## See also
|
||||
|
||||
- [../design/fips-transport-layer.md](../design/fips-transport-layer.md)
|
||||
— transport-layer design, trait surface, per-transport sections
|
||||
- [configuration.md](configuration.md) — `transports.*` configuration
|
||||
blocks
|
||||
- [../how-to/tune-udp-buffers.md](../how-to/tune-udp-buffers.md) —
|
||||
host-side `net.core.rmem_max` / `net.core.wmem_max` setup for UDP
|
||||
- [../how-to/deploy-tor-onion.md](../how-to/deploy-tor-onion.md) —
|
||||
Tor `directory` mode operator setup
|
||||
- [../how-to/set-up-bluetooth-peer.md](../how-to/set-up-bluetooth-peer.md)
|
||||
— Linux BLE peer config
|
||||
@@ -5,6 +5,45 @@ protocol layers. It covers transport framing, link-layer message formats,
|
||||
and session-layer message formats, with an encapsulation walkthrough showing
|
||||
how application data is wrapped through each layer.
|
||||
|
||||
## FMP Message Type Catalog
|
||||
|
||||
The FMP link layer defines the following message types, dispatched by the
|
||||
`msg_type` byte in the encrypted inner header:
|
||||
|
||||
| Type | Name | Forwarding |
|
||||
| ---- | ---- | ---------- |
|
||||
| 0x00 | SessionDatagram | Routed hop-by-hop toward the destination |
|
||||
| 0x01 | SenderReport | Peer-to-peer (MMP, link-layer instance) |
|
||||
| 0x02 | ReceiverReport | Peer-to-peer (MMP, link-layer instance) |
|
||||
| 0x10 | TreeAnnounce | Peer-to-peer (spanning-tree gossip) |
|
||||
| 0x20 | FilterAnnounce | Peer-to-peer (bloom-filter gossip) |
|
||||
| 0x30 | LookupRequest | Forwarded — bloom-guided through tree peers |
|
||||
| 0x31 | LookupResponse | Forwarded — reverse-path via `recent_requests` |
|
||||
| 0x50 | Disconnect | Peer-to-peer (orderly link teardown) |
|
||||
| 0x51 | Heartbeat | Peer-to-peer (link liveness) |
|
||||
|
||||
Handshake messages travel before encryption is established and are identified
|
||||
by the FMP common-prefix `phase` field rather than a `msg_type` byte
|
||||
(phase 0x1 = Noise IK msg1, phase 0x2 = Noise IK msg2).
|
||||
|
||||
## Packet Type Summary
|
||||
|
||||
A higher-level summary that includes typical sizes and forwarding category:
|
||||
|
||||
| Message | Typical Size | When | Forwarded? |
|
||||
| ------- | ------------ | ---- | ---------- |
|
||||
| TreeAnnounce | Variable (depth-dependent) | Topology changes | No (peer-to-peer) |
|
||||
| FilterAnnounce | ~1 KB | Topology changes | No (peer-to-peer) |
|
||||
| LookupRequest | ~300 bytes | First contact, recovery | Yes (bloom-guided tree) |
|
||||
| LookupResponse | ~400 bytes | Response to discovery | Yes (reverse-path) |
|
||||
| SessionDatagram + SessionSetup | ~232–402 bytes | Session establishment | Yes (routed) |
|
||||
| SessionDatagram + SessionAck | ~170 bytes | Session confirmation | Yes (routed) |
|
||||
| SessionDatagram + Data (minimal) | 77 bytes + IPv6 payload | Bulk IPv6 traffic (compressed) | Yes (routed) |
|
||||
| SessionDatagram + Data (with CP) | 77 + coords + IPv6 payload | Warmup/recovery (compressed) | Yes (routed) |
|
||||
| SessionDatagram + CoordsRequired | 70 bytes | Cache miss error | Yes (routed) |
|
||||
| SessionDatagram + PathBroken | 70+ bytes | Dead-end error | Yes (routed) |
|
||||
| Disconnect | 2 bytes | Link teardown | No (peer-to-peer) |
|
||||
|
||||
## Encoding Rules
|
||||
|
||||
- All multi-byte integers are **little-endian** (LE)
|
||||
@@ -19,10 +58,10 @@ how application data is wrapped through each layer.
|
||||
|
||||
Datagram-oriented transports (UDP, raw Ethernet, radio) preserve natural
|
||||
packet boundaries and require no additional framing. Stream-oriented
|
||||
transports (TCP, WebSocket, Tor) must delineate FIPS packets within the
|
||||
byte stream; the common prefix `payload_len` field provides this
|
||||
framing directly. TCP and Tor share a common stream reader
|
||||
(`tcp/stream.rs`) that implements this framing.
|
||||
transports (TCP, Tor) must delineate FIPS packets within the byte
|
||||
stream; the common prefix `payload_len` field provides this framing
|
||||
directly. TCP and Tor share a common stream reader (`tcp/stream.rs`)
|
||||
that implements this framing.
|
||||
|
||||
**Ethernet data frame header.** The Ethernet transport prepends a 3-byte
|
||||
header before the FMP payload on data frames: a 1-byte frame type
|
||||
@@ -505,10 +544,21 @@ Message types 0x10-0x14 are carried inside the AEAD ciphertext (dispatched
|
||||
by the `msg_type` field in the encrypted inner header). Types 0x20-0x22 are
|
||||
plaintext error signals (U flag set, no encryption).
|
||||
|
||||
Session-layer SenderReport (0x11) and ReceiverReport (0x12) use the same
|
||||
body format as their link-layer counterparts (0x01 and 0x02). The msg_type
|
||||
byte in the body matches the link-layer value; dispatch to the correct layer
|
||||
happens at the session level based on the FSP message type.
|
||||
Session-layer SenderReport (0x11) and ReceiverReport (0x12) carry the same
|
||||
metric fields as their link-layer counterparts (0x01 and 0x02), but the
|
||||
body framing differs because the FSP encrypted inner header already
|
||||
carries the message-type byte. The session body therefore omits the
|
||||
msg_type byte and uses 2 reserved bytes (not 3) before the fields:
|
||||
|
||||
| Layer | Wire size | Header inside body |
|
||||
| ----- | --------- | ------------------ |
|
||||
| Link SenderReport (0x01) | 48 bytes | msg_type(1) + reserved(3) + fields(44) |
|
||||
| Session SenderReport (0x11) | 46 bytes | reserved(2) + fields(44) |
|
||||
| Link ReceiverReport (0x02) | 68 bytes | msg_type(1) + reserved(3) + fields(64) |
|
||||
| Session ReceiverReport (0x12) | 66 bytes | reserved(2) + fields(64) |
|
||||
|
||||
Dispatch happens at the session level via the `msg_type` byte in the FSP
|
||||
encrypted inner header.
|
||||
|
||||
### SessionSetup (phase 0x1)
|
||||
|
||||
@@ -844,7 +894,7 @@ endpoint session keys).
|
||||
| ------- | ---- | ----- |
|
||||
| TreeAnnounce | 100 + 32n bytes | n = depth + 1 |
|
||||
| FilterAnnounce | 1,035 bytes | v1 (1KB filter) |
|
||||
| LookupRequest | 303 + 16n bytes | n = origin depth + 1 |
|
||||
| LookupRequest | 46 + 16n bytes | n = origin depth + 1 |
|
||||
| LookupResponse | 93 + 16n bytes | n = target depth + 1 |
|
||||
| SessionDatagram | 36 + payload bytes | Fixed 36-byte header |
|
||||
| Disconnect | 2 bytes | |
|
||||
@@ -879,8 +929,10 @@ endpoint session keys).
|
||||
|
||||
## References
|
||||
|
||||
- [fips-mesh-layer.md](fips-mesh-layer.md) — FMP behavioral specification
|
||||
- [fips-session-layer.md](fips-session-layer.md) — FSP behavioral specification
|
||||
- [fips-transport-layer.md](fips-transport-layer.md) — Transport framing
|
||||
- [fips-mesh-operation.md](fips-mesh-operation.md) — How messages work together
|
||||
- [fips-ipv6-adapter.md](fips-ipv6-adapter.md) — MTU enforcement
|
||||
- [../design/fips-mesh-layer.md](../design/fips-mesh-layer.md) — FMP behavioral specification
|
||||
- [../design/fips-session-layer.md](../design/fips-session-layer.md) — FSP behavioral specification
|
||||
- [../design/fips-transport-layer.md](../design/fips-transport-layer.md) — Transport framing
|
||||
- [../design/fips-mesh-operation.md](../design/fips-mesh-operation.md) — How messages work together
|
||||
- [../design/fips-ipv6-adapter.md](../design/fips-ipv6-adapter.md) — MTU enforcement
|
||||
- [../design/fips-bloom-filters.md](../design/fips-bloom-filters.md) — FilterAnnounce parameters and FPR analysis
|
||||
- [../design/fips-mtu.md](../design/fips-mtu.md) — How `path_mtu` and MtuExceeded fit together
|
||||
764
docs/releases/release-notes-v0.3.0.md
Normal file
@@ -0,0 +1,764 @@
|
||||
# FIPS v0.3.0
|
||||
|
||||
**Released**: 2026-05-11
|
||||
|
||||
v0.3.0 is the testing-and-polishing release on the v0.2.x wire format.
|
||||
It widens the platform reach of FIPS from Linux-only to Linux, macOS,
|
||||
Windows, and OpenWrt; adds two large new mesh capabilities (Nostr-mediated
|
||||
peer discovery with UDP NAT traversal, and the `fips-gateway` LAN bridge);
|
||||
ships a default-deny security baseline for the mesh interface; introduces
|
||||
mesh-peer access control; substantially speeds up session-layer crypto and
|
||||
the Linux receive path; and tightens packaging across every supported
|
||||
distribution channel.
|
||||
|
||||
v0.3.0 is wire-compatible with v0.2.x. Mixed meshes interoperate; there
|
||||
is no flag-day upgrade.
|
||||
|
||||
v0.3.0 also rolls forward all changes from the v0.2.1 maintenance
|
||||
release. The sections below cover the cumulative v0.2.0 → v0.3.0
|
||||
delta; the per-section intros call out which entries first shipped
|
||||
in v0.2.1.
|
||||
|
||||
## At a glance
|
||||
|
||||
- 123 commits since v0.2.0 (109 non-merge), spanning 307 files with
|
||||
+44,186 / -4,078 lines.
|
||||
- 10 committers plus 3 issue reporters across feature work, fixes,
|
||||
packaging, and reviews.
|
||||
- 5 new GitHub Actions CI workflows (Linux Package, macOS Package,
|
||||
Windows Package, OpenWrt Package, AUR Publish) plus expanded
|
||||
integration matrices (gateway, NAT-cone, NAT-symmetric, NAT-LAN,
|
||||
rekey-accept-off, `.deb` install across Debian 12/13 + Ubuntu
|
||||
22/24/26, multi-backend `.fips` DNS resolver across the same five
|
||||
distros).
|
||||
- The long-standing systemd-resolved DNS-responder silent-drop is
|
||||
closed end-to-end.
|
||||
- Pre-1.0 control-socket JSON schema change for two query fields;
|
||||
see [Upgrade notes](#upgrade-notes).
|
||||
|
||||
## What's new
|
||||
|
||||
### Mesh discovery and NAT traversal
|
||||
|
||||
Previously, two FIPS nodes could only become peers if they had a way
|
||||
to find each other beforehand: a configured address, a shared LAN
|
||||
segment, or a Bluetooth radio range. v0.3.0 introduces a Nostr-based
|
||||
overlay-discovery channel that lets nodes find each other through any
|
||||
public Nostr relay set, plus a STUN-assisted UDP hole-punching path
|
||||
that connects peers across most consumer NATs.
|
||||
|
||||
Each participating node publishes a signed overlay advert as a Nostr
|
||||
**Kind 37195** parameterized replaceable event. (The kind sits in the
|
||||
application-defined replaceable range and the digits visually spell
|
||||
*FIPS*: 7=F, 1=I, 9=P, 5=S.) The advert lists reachable transport
|
||||
endpoints (UDP, TCP, Tor) and is consumed by other nodes to populate
|
||||
fallback addresses for `via_nostr` peers. Under `policy: open`, the
|
||||
advert cache is also dialed for non-configured peers within a budget
|
||||
cap.
|
||||
|
||||
When both peers are behind NAT, the daemon coordinates a UDP hole
|
||||
punch using NIP-59 gift-wrap signaling for the offer/answer exchange
|
||||
and STUN for reflexive address discovery. A candidate-pair punch
|
||||
planner attempts LAN-private and reflexive paths in parallel; on
|
||||
success the live socket is handed into the standard FIPS UDP transport
|
||||
via a bootstrap-handoff API.
|
||||
|
||||
Operators turn this on with `node.discovery.nostr.enabled: true` and
|
||||
the configured relay set. `policy: open` adds best-effort dialing of
|
||||
non-configured peers seen on the relays. New `peers[].via_nostr` and
|
||||
per-transport `advertise_on_nostr` / `public` flags control what each
|
||||
endpoint contributes to the published advert. Cross-field validation
|
||||
runs at startup to catch mis-configured combinations early.
|
||||
|
||||
A Docker NAT lab covering cone, symmetric, and LAN scenarios is wired
|
||||
into the integration CI matrix. A daemon-side failure-suppression
|
||||
layer (per-npub cooldown after consecutive failures, ±60s clock-skew
|
||||
tolerance, rate-limited WARN logs) keeps relay traffic well-mannered
|
||||
when peers come and go from the open discovery cache. A separate
|
||||
structural cooldown (`protocol_mismatch_cooldown_secs`, default 24h)
|
||||
suppresses retraversal when a punched peer turns out to be running an
|
||||
FMP version this daemon cannot handshake with: the punch completes at
|
||||
the UDP layer, the rx loop spots the version-mismatched packet,
|
||||
reverse-maps to the originating npub, and removes the peer from the
|
||||
next sweep until either side upgrades.
|
||||
|
||||
The auto-connect retry loop pins itself to relay ground truth. Each
|
||||
retry attempt refetches the cached overlay advert against the
|
||||
configured `advert_relays` (one filter query, 2s timeout) before
|
||||
dialing, so a peer whose NAT rebound to a fresh endpoint is recovered
|
||||
on the next retry rather than looping on a stale cached address.
|
||||
`NoTransportForType` triggers a fire-and-forget re-fetch that either
|
||||
replaces or evicts the cache entry. A startup peer-init failure (no
|
||||
operational transport, all addresses unreachable) now schedules a
|
||||
retry instead of leaving the peer in a dead state until the daemon is
|
||||
restarted. Adopted NAT-traversed UDP transports inherit the operator's
|
||||
primary `[transports.udp]` listener config (MTU, recv/send buffer
|
||||
sizes) instead of falling back to the 1280 IPv6-minimum default.
|
||||
|
||||
### Cross-platform reach
|
||||
|
||||
FIPS now ships first-class binaries for **Linux, macOS, Windows, and
|
||||
OpenWrt**.
|
||||
|
||||
- **macOS** support uses the native `utun` TUN interface, raw
|
||||
Ethernet via BPF, a `.pkg` installer with a launchd plist and
|
||||
uninstall script, and an x86_64 cross-compile from arm64 build
|
||||
hosts. A new CI matrix entry runs build and unit-test jobs on
|
||||
macOS hosts.
|
||||
- **Windows** support uses [wintun](https://www.wintun.net/) for the
|
||||
TUN device, a TCP control socket on `localhost:21210` (replacing
|
||||
the Unix domain socket Linux and macOS use), Windows Service
|
||||
lifecycle (`fips.exe --install-service`, `--uninstall-service`,
|
||||
`--service`), and a ZIP package with PowerShell install/uninstall
|
||||
scripts.
|
||||
- **MIPS** atomic-ABI portability lets the daemon build for 32-bit
|
||||
MIPS targets (`mips`, `mipsel`, MIPS32r2) by routing through
|
||||
`portable_atomic`. This unblocks OpenWrt deployments on
|
||||
consumer-grade MIPS routers.
|
||||
- **OpenWrt** packaging gets a procd init with dnsmasq forwarding,
|
||||
proxy NDP, RA route advertisements, and IPv6 forwarding sysctls.
|
||||
The `fips-gateway` is enabled by default in the OpenWrt build.
|
||||
|
||||
### FIPS gateway
|
||||
|
||||
The new `fips-gateway` binary lets unmodified LAN hosts reach FIPS
|
||||
mesh destinations without running the FIPS daemon themselves. Two
|
||||
flows ship together:
|
||||
|
||||
- **Outbound (LAN -> mesh)**: a virtual-IP pool (default
|
||||
`fd01::/112`) is allocated on demand from `.fips`-name DNS lookups.
|
||||
A state-machine lifecycle, conntrack-backed session tracking, proxy
|
||||
NDP on the LAN interface, and TTL-based reclamation handle the
|
||||
bookkeeping. A LAN host that resolves `peer.fips` gets a virtual
|
||||
address it can reach over IP, and the gateway translates the flow
|
||||
to the mesh.
|
||||
- **Inbound (mesh -> LAN)**: new `gateway.port_forwards` config
|
||||
installs prerouting DNAT rules so mesh peers can reach a configured
|
||||
`host:port` on the gateway's LAN. A LAN-side masquerade is added
|
||||
automatically when any forwards are configured, so replies flow
|
||||
back through conntrack.
|
||||
|
||||
A dedicated control socket at `/run/fips/gateway.sock` exposes
|
||||
`show_gateway` and `show_mappings`. `fipstop` adds a Gateway tab with
|
||||
a pool gauge and mappings table.
|
||||
|
||||
The gateway's `dns.listen` source default is now `[::1]:5353`,
|
||||
matching the canonical deployment model: the gateway sits on a host
|
||||
already serving DHCP and DNS to a LAN segment (an OpenWrt AP, a Linux
|
||||
router), port 53 there is taken by the existing resolver, and `.fips`
|
||||
queries are forwarded to the gateway over loopback. The OpenWrt ipk
|
||||
previously overrode the prior `[::]:53` source default in its packaged
|
||||
config; that override is now redundant and has been dropped.
|
||||
Operators on a host without a pre-existing resolver on port 53 can
|
||||
opt back into the wildcard bind by setting `dns.listen: "[::]:53"`
|
||||
explicitly. The new default binds IPv6 loopback only, so forwarders
|
||||
that reach the gateway over IPv4 loopback need an explicit IPv4
|
||||
listen address.
|
||||
|
||||
The cold-boot startup race between `fips.service` and
|
||||
`fips-gateway.service` is handled by a systemd `After=fips.service`
|
||||
ordering, an `ExecStartPre` poll loop that waits up to 30 seconds for
|
||||
the `fips0` interface to appear, and a DNS upstream probe in the
|
||||
gateway itself that retries up to 5 times with 1-second backoff.
|
||||
|
||||
Packaging covers systemd, Debian, AUR, and OpenWrt. The full design
|
||||
is in [`docs/design/fips-gateway.md`](../design/fips-gateway.md).
|
||||
|
||||
### Mesh-interface security baseline
|
||||
|
||||
The FIPS mesh is a flat layer-3 segment. Every authenticated peer can
|
||||
route packets to every other peer's `fips0` address. Peer identity is
|
||||
authenticated end-to-end by the FMP and FSP Noise handshakes, but
|
||||
identity is not authorization. A service on a mesh host that binds to
|
||||
a wildcard address is, by default, reachable from every peer in the
|
||||
mesh.
|
||||
|
||||
v0.3.0 ships an opt-in default-deny baseline that closes this gap on
|
||||
Linux:
|
||||
|
||||
- **`/etc/fips/fips.nft`** is installed as a documented operator
|
||||
conffile. It defines a single `inet fips` nftables table with one
|
||||
chain hooked at `input`, default-denies inbound traffic on
|
||||
`fips0`, and is a no-op for every other interface.
|
||||
- **`fips-firewall.service`** loads it. The unit ships **disabled by
|
||||
default**; activation is an explicit
|
||||
`systemctl enable --now fips-firewall.service`.
|
||||
- Per-service allowances live in **`/etc/fips/fips.d/*.nft`**
|
||||
drop-ins that the baseline includes.
|
||||
|
||||
Choosing opt-in keeps the mesh quick to bring up for evaluation while
|
||||
giving operators a documented, packaged path to lock it down for
|
||||
production. The full design (threat model, rule layout, conntrack
|
||||
handling, drop-in mechanism, and the rationale for a conffile rather
|
||||
than an auto-loaded package side-effect) is in
|
||||
[`docs/design/fips-security.md`](../design/fips-security.md).
|
||||
|
||||
`fipstop`'s Node tab gains a **"Listening on fips0" panel** that
|
||||
surfaces the answer to the operational question "what services on
|
||||
this host are reachable from the mesh, and what does the firewall
|
||||
currently say about each of them?" The panel lists every IPv6
|
||||
listening socket bound to either the wildcard address or this node's
|
||||
`fd00::/8` address, paired with its classification against the
|
||||
running `inet fips` baseline chain: `OPEN` (canonical accept rule),
|
||||
`filt` (falls through to drop), or `filt?` (referenced with matchers
|
||||
the panel cannot fully decompose, e.g. saddr filters or jumps). When
|
||||
`fips-firewall.service` is inactive, a yellow banner above the table
|
||||
reminds the operator that every listener is mesh-exposed; wildcard
|
||||
binds carry a trailing `*` in the Process column. The classifier is
|
||||
built on a new `show_listening_sockets` control query (Linux-only),
|
||||
which is also useful from `fipsctl` for scripting.
|
||||
|
||||
### Peer access control
|
||||
|
||||
Operators can now restrict which mesh peers a node will form direct
|
||||
links with. Optional `/etc/fips/peers.allow` and `/etc/fips/peers.deny`
|
||||
files (TCP-Wrappers style) match against npub, hex pubkey, host
|
||||
alias, or `ALL`. Enforcement runs at three points:
|
||||
|
||||
1. Outbound connect (before dialing).
|
||||
2. Inbound msg1 (the first FMP handshake message from a new peer).
|
||||
3. Outbound msg2 (the response).
|
||||
|
||||
Files reload automatically on mtime change; a new `fipsctl acl show`
|
||||
query reports the effective rule set. A six-node Docker integration
|
||||
harness (`testing/acl/`) exercises allowlist and denylist patterns
|
||||
end-to-end.
|
||||
|
||||
**Important scope distinction**: peer ACLs are an FMP-layer
|
||||
restriction. They control who can establish a *direct link* with this
|
||||
node. They do **not** control session-layer (FSP) reachability through
|
||||
the mesh. A node that denies peer X with an ACL can still receive FSP
|
||||
traffic from X relayed via other peers.
|
||||
|
||||
### Bluetooth Low Energy transport (experimental, Linux)
|
||||
|
||||
A new BLE L2CAP Connection-Oriented Channel transport lets FIPS nodes
|
||||
peer over Bluetooth Low Energy without any IP infrastructure in
|
||||
between. The transport handles per-link MTU negotiation, continuous
|
||||
scan/probe peer discovery with cooldown-based deduplication,
|
||||
continuous advertising, deterministic NodeAddr cross-probe
|
||||
tie-breaker, and a configurable connection pool with eviction.
|
||||
|
||||
This transport is **experimental in v0.3.0**. It is implemented and
|
||||
functional on Linux (BlueZ via `bluer`), but the reliability follow-up
|
||||
logic (probe cooldown, cross-probe tie-breaker, pubkey timeout,
|
||||
continuous advertising semantics, probe-promotion, fail-fast send) is
|
||||
not yet behaviorally tested in CI. Its maturity path is field-driven;
|
||||
please file issues with field reports. macOS BLE support is in
|
||||
development as a separate track and is not part of v0.3.0.
|
||||
|
||||
### UDP transport profiles
|
||||
|
||||
The UDP transport gains posture flags organized around deployment
|
||||
patterns:
|
||||
|
||||
- **Public-facing inbound nodes**: `bind_addr: "0.0.0.0:2121"`,
|
||||
`accept_connections: true` (default), `public: true` for advert
|
||||
publication. v0.3.0 adds STUN-based public-IP autodiscovery so
|
||||
cloud nodes (AWS EIP, GCP, Azure 1:1 NAT) advertise the right
|
||||
address even when the public IP isn't on a host interface.
|
||||
- **Ephemeral leaf nodes**: `outbound_only: true` binds an ephemeral
|
||||
port (`0.0.0.0:0`), refuses inbound msg1, and is never advertised
|
||||
on Nostr regardless of `advertise_on_nostr`. Use this for client
|
||||
postures that should connect outbound only, without exposing an
|
||||
inbound listener on a known port.
|
||||
- **General-purpose nodes**: `accept_connections: false` mirrors the
|
||||
Ethernet/BLE knob without changing the bind address. The Node-level
|
||||
handshake gate carves out msg1 from peers already established on
|
||||
this transport so rekey continues to work.
|
||||
|
||||
Startup validation now rejects `bind_addr` set to a loopback address
|
||||
when at least one peer has a non-loopback UDP address, closing a
|
||||
silent-failure trap from v0.2.0 where Linux's source-address routing
|
||||
check would drop outbound flows from the loopback-bound socket.
|
||||
|
||||
A new `external_addr` field on `transports.udp.*` and
|
||||
`transports.tcp.*` lets operators specify the advertise-as address
|
||||
explicitly. This is useful for UDP as a deterministic alternative to
|
||||
STUN, and required for TCP on cloud-NAT setups (where binding to the
|
||||
public IP fails with `EADDRNOTAVAIL` because the IP isn't on a host
|
||||
interface).
|
||||
|
||||
### `.fips` DNS resolver overhaul
|
||||
|
||||
The IPv6 adapter's `.fips` name resolution has been rebuilt around
|
||||
the constraints of contemporary systemd-based hosts. The default
|
||||
`dns.bind_addr` is now `::1` (IPv6 loopback), and a setup script
|
||||
picks one of five backends in priority order:
|
||||
|
||||
1. systemd-resolved global drop-in
|
||||
(`/etc/systemd/resolved.conf.d/fips.conf`).
|
||||
2. systemd dns-delegate (per-link configuration handed off to
|
||||
systemd-resolved).
|
||||
3. `resolvectl` per-link configuration.
|
||||
4. Standalone `dnsmasq`.
|
||||
5. NetworkManager's dnsmasq plugin.
|
||||
|
||||
Teardown reverses only what setup applied, recorded in a state file
|
||||
at `/run/fips/dns-backend`. A new `testing/dns-resolver/` harness
|
||||
exercises every backend across Debian 12, Debian 13, Ubuntu 22.04,
|
||||
Ubuntu 24.04, and Ubuntu 26.04, so a regression in any of the five
|
||||
backends shows up in CI rather than in the field.
|
||||
|
||||
This overhaul resolves the long-standing silent-drop case where the
|
||||
`resolvectl dns fips0 [<fips0_addr>]:5354` target collided with the
|
||||
daemon's mesh-interface filter on certain systemd-resolved
|
||||
deployments (typically Ubuntu 22 with systemd 249's interface-scoped
|
||||
routing).
|
||||
|
||||
### Operator tooling additions
|
||||
|
||||
A handful of additions land in `fipsctl`, `fipstop`, and the daemon's
|
||||
configuration surface:
|
||||
|
||||
- **`node.log_level`** config field replaces the hardcoded
|
||||
`RUST_LOG=info` previously baked into systemd units and the
|
||||
OpenWrt procd init. The daemon now loads config before
|
||||
initializing tracing so the configured level takes effect.
|
||||
`RUST_LOG` still overrides when set.
|
||||
- **`fipsctl show identity-cache`** is a new query that lists every
|
||||
cached node identity (npub, IPv6 address, display name, LRU age)
|
||||
alongside the configured cache capacity.
|
||||
- **`fipsctl show peers / sessions / cache / routing`** are
|
||||
substantially extended: per-peer security signals (replay
|
||||
suppression count, consecutive decrypt failures), Noise session
|
||||
counters, session indices, rekey lifecycle state, handshake resend
|
||||
counts, K-bit epoch, coords-warmup remaining, drain state, per-peer
|
||||
retry state, per-target lookup detail (attempt, age, last sent),
|
||||
and pending TUN packet queue depth.
|
||||
- **Historical statistics**: in-memory time-series rings on the
|
||||
daemon (1-second × 3600 fast, 1-minute × 1440 slow) cover per-node
|
||||
and per-peer metrics. New `show_stats_*` control-socket queries, a
|
||||
`fipsctl stats list / peers / history` subcommand with Unicode
|
||||
sparkline rendering, and a `fipstop` Graphs tab with btop-style
|
||||
sparklines surface them to the operator.
|
||||
|
||||
### Performance
|
||||
|
||||
Two independent perf threads land in v0.3.0: a session-layer crypto
|
||||
backend swap, and a Linux receive-path overhaul.
|
||||
|
||||
**Session-layer crypto backend.** The ChaCha20-Poly1305 backend used
|
||||
by every FIPS Noise session (end-to-end FSP traffic and link-layer
|
||||
FMP traffic alike) has been swapped from RustCrypto's
|
||||
`chacha20poly1305` crate to `ring 0.17`. ring wraps BoringSSL's
|
||||
hand-tuned ChaCha20-Poly1305 implementation, which dispatches to NEON
|
||||
on aarch64 and AVX2 / AVX-512 on x86_64. Typical throughput is in the
|
||||
3-5 GB/s/core range, versus the ~600-800 MB/s/core RustCrypto soft
|
||||
path on the same hardware.
|
||||
|
||||
Wire format is unchanged. ChaCha20-Poly1305 is byte-deterministic for
|
||||
a given `(key, nonce, plaintext, aad)`, so any correct AEAD
|
||||
implementation produces identical ciphertext. A mixed mesh with some
|
||||
nodes pre-swap and some post-swap interoperates without protocol
|
||||
awareness; v0.3.0 can roll out across a mesh in any order.
|
||||
|
||||
Measurements on an aarch64 Apple Silicon docker target:
|
||||
|
||||
- Two-node TCP single-stream: 437 -> 1097 Mbps (about 2.5×).
|
||||
- Two-node UDP at 1000 Mbit: 599 Mbps with 40% loss -> lossless at
|
||||
line rate.
|
||||
- Three-node ping under bulk-traffic load: 7.68 ms avg / 215 ms max
|
||||
-> 0.72 ms / 3.6 ms max as the relay path stops being crypto-bound.
|
||||
|
||||
No operator-visible action is required; the swap is internal to the
|
||||
session layer.
|
||||
|
||||
**Linux UDP receive path.** The Linux UDP receive path now uses
|
||||
`recvmmsg(2)` with a 32-packet batch in place of single-packet
|
||||
`recvmsg(2)`. A single `readable()` wakeup drains up to 32 datagrams
|
||||
in one syscall before yielding back to the reactor, eliminating the
|
||||
per-packet scheduler-hop and futex cost that previously capped
|
||||
inbound rate at one event per scheduler quantum independent of CPU.
|
||||
`SO_RXQ_OVFL` is sampled once per batch and surfaced through
|
||||
`AsyncUdpSocket::recv_batch` so the existing 1Hz transport-congestion
|
||||
detector continues to feed the per-transport `dropping` flag. macOS
|
||||
and Windows fall through to the per-packet path; `recvmmsg` is
|
||||
Linux-specific.
|
||||
|
||||
**Inner rx-loop drain batching.** `Node::run_rx_loop` drains up to
|
||||
256 additional ready items via `try_recv()` after each
|
||||
`tokio::select!` await fires on the packet and TUN-outbound branches,
|
||||
in a tight inner loop before yielding. Previously the select cost a
|
||||
full scheduler hop and futex per packet, capping throughput at one
|
||||
event per scheduler quantum with the worker near-idle. `biased`
|
||||
ordering keeps data-plane branches priority over tick / control / DNS
|
||||
under sustained load; the 256 cap keeps the worker on a busy stream
|
||||
between yield points (about 400 KB of contiguous traffic) while still
|
||||
bounding the inner loop so a flood on one branch cannot starve the
|
||||
periodic tick or control socket.
|
||||
|
||||
**Eager `pubkey_full` precompute.** `PeerIdentity::pubkey_full()`
|
||||
precomputes the parity-aware full secp256k1 public key at
|
||||
construction in `from_pubkey`. Previously the method fell through to
|
||||
an EC point parse on every call when the full key wasn't passed at
|
||||
construction (i.e. for every peer constructed from an npub or x-only
|
||||
key), about 6% of per-packet CPU on the bulk-data send path for a
|
||||
value that never changed after construction. The same parse already
|
||||
runs at construction inside `NodeAddr::from_pubkey`, so the cost is
|
||||
paid once where it would be paid anyway.
|
||||
|
||||
These three changes are a coordinated set: the syscall batching
|
||||
removes the per-packet kernel cost, the inner-loop drain removes the
|
||||
per-packet scheduler cost, and the pubkey-cache change removes the
|
||||
per-packet crypto-derivation cost. Like the AEAD swap, they are all
|
||||
internal and require no operator action.
|
||||
|
||||
### Examples
|
||||
|
||||
- **macOS WireGuard companion** ([#51](https://github.com/jmcorgan/fips/pull/51)):
|
||||
run FIPS in a local Docker container and route `.fips` traffic
|
||||
from the macOS host through a WireGuard tunnel to the container's
|
||||
`fips0`. Only traffic destined for `fd00::/8` transits the
|
||||
companion; regular internet traffic continues to use the host
|
||||
network. Persistent FIPS and WireGuard key material is generated
|
||||
on first run.
|
||||
|
||||
### Documentation
|
||||
|
||||
- **`docs/design/port-advertisement-and-nat-traversal.md`**
|
||||
documents how nodes find each other through Nostr relays and the
|
||||
STUN-assisted UDP hole punch.
|
||||
- **`docs/design/fips-gateway.md`** documents the gateway's virtual
|
||||
IP pool, lifecycle, control surface, and packaging.
|
||||
- **`docs/design/fips-security.md`** documents the mesh-interface
|
||||
security posture, threat model, default-deny baseline, and drop-in
|
||||
workflow.
|
||||
- **`CONTRIBUTING.md`** has been expanded with build prerequisites,
|
||||
Rust toolchain setup, and first-build steps.
|
||||
|
||||
The `docs/` tree has been reorganized end-to-end into four sections
|
||||
(*tutorials / how-to / reference / design*) with a new
|
||||
[`docs/getting-started.md`](../getting-started.md) and per-section
|
||||
landing pages. Content was reconciled against current source:
|
||||
protocol-layer details, wire-format diagrams, configuration knobs,
|
||||
and CLI references were brought back into agreement with the
|
||||
implementation. See [Documentation pointers](#documentation-pointers)
|
||||
below for entry points by reader intent.
|
||||
|
||||
## Behavior changes worth flagging
|
||||
|
||||
These default-config changes affect every operator on upgrade, even
|
||||
those with no explicit configuration. Two items below — bloom-filter
|
||||
fill-ratio validation and TreeAnnounce ancestry validation — first
|
||||
shipped in v0.2.1 and roll forward into v0.3.0; the rest are
|
||||
v0.3.0-net-new.
|
||||
|
||||
- **Discovery rate-limiting** has been retuned to be less aggressive
|
||||
at cold start. v0.2.0 used a single-lookup-with-internal-retry
|
||||
model where a timed-out lookup during bloom-filter propagation
|
||||
could suppress retries for 30 seconds while none of the reset
|
||||
triggers fired on a stable post-handshake topology. v0.3.0
|
||||
replaces this with a per-attempt timeout sequence
|
||||
(`node.discovery.attempt_timeouts_secs`, default `[1, 2, 4, 8]`,
|
||||
15s total). Each attempt sends a fresh `LookupRequest` with a new
|
||||
`request_id`, letting successive attempts take different
|
||||
forwarding paths as the bloom and tree state evolve. Post-failure
|
||||
suppression is **off by default**; operators with chatty
|
||||
applications can opt back in via `backoff_base_secs` /
|
||||
`backoff_max_secs`.
|
||||
- **MMP report intervals** are retuned for constrained transports.
|
||||
The steady-state floor moves from 100ms to 1000ms, the ceiling
|
||||
from 2000ms to 5000ms, with a cold-start phase running 200ms for
|
||||
the first 5 SRTT samples. This reduces BLE overhead by roughly
|
||||
10× while keeping reports well above the EWMA convergence
|
||||
threshold. Session-layer MMP intervals are unchanged.
|
||||
- **Bloom filter fill-ratio validation** runs on every inbound
|
||||
`FilterAnnounce`. Filters whose derived false-positive rate
|
||||
exceeds `node.bloom.max_inbound_fpr` (default 0.05) are rejected
|
||||
silently on the wire, logged at WARN, and counted in a new
|
||||
`bloom.fill_exceeded` counter. A rate-limited WARN also fires
|
||||
when the local outgoing filter exceeds the cap.
|
||||
- **TreeAnnounce ancestry validation** is now run before tree-state
|
||||
mutation, enforcing ancestry-self-match, root-single-entry,
|
||||
parent-second-entry, and root-is-minimum-NodeAddr. Non-conforming
|
||||
announces are rejected with a WARN. Mixed v0.2.0 / v0.2.1 / v0.3.0
|
||||
meshes may produce WARN log lines on the v0.2.1+ side until all
|
||||
peers upgrade; behavior is correct, log noise only.
|
||||
- **Log noise reduction**: 35 info-level log messages have been
|
||||
demoted to debug (handshake cross-connection mechanics, periodic
|
||||
MMP telemetry, TUN/transport shutdown, retry scheduling). The
|
||||
default `RUST_LOG` in systemd units is now `info`, where it
|
||||
previously ran at `debug`. Operator-visible info output now
|
||||
focuses on lifecycle events, peer promotions, session
|
||||
establishment, parent switches, and transport start/stop.
|
||||
|
||||
## Notable bug fixes
|
||||
|
||||
These pre-existing v0.2.0 bugs are worth singling out because they
|
||||
either affected real-world deployments or produced misleading
|
||||
operator experiences. The CHANGELOG has the exhaustive list; this is
|
||||
the operator-relevant subset. Four items below first shipped in
|
||||
v0.2.1 and roll forward into v0.3.0: auto-connect Disconnect-reconnect,
|
||||
`fipsctl connect` mesh-address rejection, `fd00::/8` routing
|
||||
protection from Tailscale interception, and bloom-filter routing
|
||||
greedy-tree fallback. The control-socket path-detection fix landed
|
||||
in v0.2.1 as well, and the unified resolver below is the v0.3.0
|
||||
refactor that builds on it.
|
||||
|
||||
- **DNS responder silent-drop on systemd-resolved** is fixed: the
|
||||
responder no longer drops queries on Ubuntu 22 / Debian 13 and
|
||||
similar deployments where systemd applies interface-scoped
|
||||
routing. Default bind moves to `::1`; new global drop-in backend
|
||||
available ([#52](https://github.com/jmcorgan/fips/issues/52),
|
||||
[#77](https://github.com/jmcorgan/fips/issues/77)).
|
||||
- **Auto-connect peers reconnect after a graceful Disconnect.**
|
||||
Previously, a clean upstream shutdown left the auto-connect peer
|
||||
orphaned; only the link-dead, decrypt-fail, and peer-restart
|
||||
paths scheduled a reconnect
|
||||
([#60](https://github.com/jmcorgan/fips/issues/60), reported by
|
||||
[@SwapMarket](https://github.com/SwapMarket)).
|
||||
- **`fipsctl connect` rejects FIPS mesh addresses** (`fd00::/8`)
|
||||
for `udp`, `tcp`, and `ethernet` transports with a clear error
|
||||
message, instead of echoing success while the daemon silently
|
||||
failed the bind with `EAFNOSUPPORT`
|
||||
([#61](https://github.com/jmcorgan/fips/issues/61), reported by
|
||||
[@SwapMarket](https://github.com/SwapMarket)).
|
||||
- **Default control-socket path resolution unified.** Daemon and
|
||||
client tools now share a single resolver, eliminating a divergence
|
||||
where `fipsctl` / `fipstop` could connect to a socket the daemon
|
||||
never bound (notably on dev runs with `XDG_RUNTIME_DIR` set, or
|
||||
after a prior packaged install left a root-owned `/run/fips`
|
||||
behind). Canonical order is
|
||||
`/run/fips` -> `$XDG_RUNTIME_DIR/fips/` -> `/tmp/fips-<name>`. The
|
||||
`/run/fips` arm is selected by directory existence; the kernel
|
||||
enforces actual access at `connect(2)` time, so users not yet in
|
||||
the `fips` group get a clear `EACCES` rather than a silent path
|
||||
mismatch and a misleading `No such file` fallback to
|
||||
`$XDG_RUNTIME_DIR`. `XDG_RUNTIME_DIR` is validated as an existing
|
||||
directory before being used so stale post-logout values are
|
||||
treated as missing. The deployed fleet is unaffected: packaged
|
||||
configs set `node.control.socket_path` explicitly
|
||||
([#30](https://github.com/jmcorgan/fips/issues/30), reported by
|
||||
[@Sebastix](https://github.com/Sebastix)).
|
||||
- **`fd00::/8` routing protected from Tailscale interception.** The
|
||||
daemon installs an IPv6 routing-policy rule
|
||||
(`ip -6 rule to fd00::/8 lookup main priority 5265`) at TUN
|
||||
setup, so Tailscale's table 52 default route can no longer divert
|
||||
mesh traffic.
|
||||
- **TCP-over-FIPS reliability on mixed-MTU paths** is markedly
|
||||
improved. Four interlocking changes ship together:
|
||||
`Node::transport_mtu()` is now deterministic across daemon
|
||||
restarts (min across operational transports rather than
|
||||
insertion-order-dependent); the TCP MSS clamp at the TUN boundary
|
||||
reads per-destination path MTU instead of a single global ceiling;
|
||||
reactive `MtuExceeded` from forwarders is mirrored back into the
|
||||
TUN-side `path_mtu_lookup` so later flows pick up forward-path
|
||||
bottlenecks without re-discovery; and the proactive end-to-end
|
||||
`PathMtuNotification` echoed by the destination is mirrored into
|
||||
the same TUN-side store. Without that fourth piece, on long-lived
|
||||
stable paths where the destination's echo had tightened the
|
||||
session MTU but no transit router had emitted a fresh
|
||||
`MtuExceeded`, new TCP flows opened in that window were clamped by
|
||||
the staler discovery-time value. The proactive mirror uses the
|
||||
same tighter-only semantics as the reactive mirror, so it never
|
||||
loosens the clamp. The Windows TUN reader receives the same
|
||||
per-destination plumbing.
|
||||
- **Bloom filter routing greedy-tree fallback.** `find_next_hop` no
|
||||
longer returns `NoRoute` when the bloom candidate set is non-empty
|
||||
but no candidate is strictly closer than the current node; it
|
||||
falls through to greedy tree routing instead. Previously, this
|
||||
caused dropped packets in topologies where the tree parent was
|
||||
closer but not a bloom candidate.
|
||||
- **`fipstop` graceful tty-init failure.** `ratatui::try_init()`
|
||||
produces a clean error message instead of a hard crash when
|
||||
terminal initialization fails (Docker on macOS Sequoia, ttyless
|
||||
environments).
|
||||
- **TreeAnnounce ancestry on self-root transitions.** When a node
|
||||
had no smaller-NodeAddr peer to use as a parent, the spanning-tree
|
||||
state correctly promoted it to root, but the ancestry advertised
|
||||
on the next `TreeAnnounce` still referenced its previous parent's
|
||||
path. Receiving peers rejected the announce as
|
||||
`invalid ancestry: advertised root X is not the minimum path entry
|
||||
Y`, blocking mesh transit on any path that needed to traverse the
|
||||
node. The self-root transition is now detected explicitly in
|
||||
`TreeState::become_root` and the advertised ancestry rebuilt to
|
||||
start from self; the MMP receive handler corrects stale ancestry
|
||||
inherited across reconnect eagerly rather than waiting for the
|
||||
next observation tick.
|
||||
- **Spanning-tree internal-path updates** that change only the
|
||||
internal path between root and leaf (without changing the root or
|
||||
the depth) now propagate to leaves correctly. Previously, a leaf
|
||||
could continue routing against a stale internal path until the
|
||||
parent or depth also changed.
|
||||
|
||||
## Upgrade notes
|
||||
|
||||
Operator-actionable items when moving from v0.2.x to v0.3.0:
|
||||
|
||||
- **Control socket JSON schema (breaking, pre-1.0).**
|
||||
- `show_cache` response field `entries` has changed type from a
|
||||
`u64` count to an array of entry objects. The previous scalar
|
||||
value is now in a new `count` field.
|
||||
- `show_routing` response field `pending_lookups` has changed
|
||||
type from a `u64` count to an array of per-target lookup
|
||||
objects.
|
||||
- External tooling parsing these fields as numbers must be
|
||||
updated. In-tree `fipstop` is adjusted to the new schema. The
|
||||
control-socket interface remains pre-1.0 and is not covered by
|
||||
stability guarantees.
|
||||
|
||||
- **Cargo feature flags removed.** `tui`, `ble`, `gateway`, and
|
||||
`nostr-discovery` are gone. Subsystem inclusion is now driven by
|
||||
platform `cfg` gates, so plain `cargo build` compiles everything
|
||||
available on the target without `--features` invocations.
|
||||
Source-build tooling that passed any of these features should be
|
||||
updated to omit them.
|
||||
|
||||
- **Discovery rate-limiting defaults changed.** Post-failure
|
||||
suppression is **off by default**
|
||||
(`node.discovery.backoff_base_secs: 0`, `backoff_max_secs: 0`).
|
||||
Operators relying on the prior 30s base / 300s cap behavior must
|
||||
set those fields explicitly. The per-attempt sequence
|
||||
(`attempt_timeouts_secs`, default `[1, 2, 4, 8]`) now governs
|
||||
cold-start lookup behavior.
|
||||
|
||||
- **`.fips` DNS bind address default changed.** The default
|
||||
`dns.bind_addr` is now `::1`. Operators with explicit overrides
|
||||
of this field should review them; many existing overrides were
|
||||
workarounds for the silent-drop bug that this release fixes
|
||||
properly.
|
||||
|
||||
- **Gateway `dns.listen` source default changed.** The
|
||||
`fips-gateway` `dns.listen` default is now `[::1]:5353` (was
|
||||
`[::]:53`), matching the canonical deployment model where a
|
||||
pre-existing resolver on the host already owns port 53. The
|
||||
OpenWrt ipk previously overrode this in its packaged config; the
|
||||
override is now redundant and has been dropped. Operators on a
|
||||
host without a pre-existing resolver on port 53 can opt back into
|
||||
the wildcard bind by setting `dns.listen: "[::]:53"` explicitly.
|
||||
The new default binds IPv6 loopback only, so forwarders that
|
||||
reach the gateway over IPv4 loopback need an explicit IPv4 listen
|
||||
address.
|
||||
|
||||
- **systemd unit log level.** The shipped systemd units no longer
|
||||
hardcode `RUST_LOG=info`; the daemon's effective log level is
|
||||
driven by `node.log_level` (default `info`). `RUST_LOG`, when
|
||||
set, still overrides.
|
||||
|
||||
- **UDP transport `bind_addr` validation.** Startup now rejects a
|
||||
`bind_addr` set to a loopback address when at least one peer has
|
||||
a non-loopback UDP address. Operators who configured a loopback
|
||||
UDP bind as a workaround should switch to `outbound_only: true`
|
||||
for the same effect, plus the correct semantics (kernel-assigned
|
||||
ephemeral port, refuses inbound, never advertised).
|
||||
|
||||
- **Tor advert port.** If the Tor `HiddenServicePort` virtual port
|
||||
isn't 443, set `transports.tor.advertised_port` to match. The
|
||||
default is 443 and matches the conventional virtual-port choice.
|
||||
|
||||
## Documentation pointers
|
||||
|
||||
v0.3.0 ships a `docs/` tree reorganized into four sections
|
||||
(*tutorials / how-to / reference / design*). A new top-level
|
||||
[`docs/getting-started.md`](../getting-started.md) and per-section
|
||||
landing pages anchor the entry points.
|
||||
|
||||
Entry points by reader intent:
|
||||
|
||||
- **New users**: [`docs/getting-started.md`](../getting-started.md)
|
||||
and [`docs/tutorials/`](../tutorials/) cover guided introductions
|
||||
for bringing up your first node, joining the test mesh,
|
||||
advertising a node over Nostr, hosting a service, deploying a
|
||||
gateway, walking through the IPv6 adapter, and resolving peers
|
||||
via Nostr.
|
||||
- **Operators with a specific task**:
|
||||
[`docs/how-to/`](../how-to/) holds task-driven guides for enabling
|
||||
Nostr discovery, deploying the gateway, troubleshooting the
|
||||
gateway, deploying a Tor onion, hosting aliases, persistent
|
||||
identity, running unprivileged, setting up a Bluetooth peer,
|
||||
enabling the mesh firewall, tuning UDP buffers, and diagnosing
|
||||
MTU issues.
|
||||
- **Reference lookups**: [`docs/reference/`](../reference/) holds
|
||||
the config field reference, control-socket query reference, the
|
||||
`fips`, `fipsctl`, `fipstop`, and `fips-gateway` CLI references,
|
||||
and the protocol diagram set.
|
||||
- **Architectural background**: [`docs/design/`](../design/) holds
|
||||
design rationale for FIPS as a whole, FMP and FSP, the spanning
|
||||
tree, bloom-filter discovery, transports, the IPv6 adapter, the
|
||||
Nostr discovery layer, and the gateway.
|
||||
- **Security**: [`docs/design/fips-security.md`](../design/fips-security.md)
|
||||
documents the mesh-interface security baseline, threat model, and
|
||||
drop-in workflow.
|
||||
|
||||
## Getting v0.3.0
|
||||
|
||||
- **Linux x86_64 / aarch64**: `.deb` and tarball at the
|
||||
[v0.3.0 release page](https://github.com/jmcorgan/fips/releases/tag/v0.3.0).
|
||||
- **Arch Linux**: `fips` from the AUR.
|
||||
- **macOS**: `.pkg` at the v0.3.0 release page.
|
||||
- **Windows**: ZIP at the v0.3.0 release page.
|
||||
- **OpenWrt**: `.ipk` at the v0.3.0 release page.
|
||||
- **From source**: `cargo build --release` from a checkout of the
|
||||
v0.3.0 tag.
|
||||
|
||||
The full per-commit changelog lives in
|
||||
[`CHANGELOG.md`](../../CHANGELOG.md). Issues and discussion at
|
||||
[github.com/jmcorgan/fips](https://github.com/jmcorgan/fips).
|
||||
|
||||
## Contributors
|
||||
|
||||
Thanks to everyone who contributed code, packaging work, bug reports,
|
||||
or reviews to this release.
|
||||
|
||||
**Code and packaging**:
|
||||
|
||||
- [@jcorgan](https://github.com/jmcorgan): release shepherd, Nostr
|
||||
discovery / NAT traversal, `fips-gateway`, ACL infrastructure,
|
||||
packaging, security baseline, BLE follow-ups.
|
||||
- [@Origami74](https://github.com/Origami74): macOS platform support,
|
||||
from-source Docker companion build and `fipstop` terminal-init
|
||||
handling, gateway co-development, OpenWrt BLE-feature build fix,
|
||||
AUR-workflow follow-ups.
|
||||
- [@jodobear](https://github.com/jodobear): Linux release-artifact
|
||||
workflow and target-aware build scripts, CONTRIBUTING.md
|
||||
expansion, rekey integration-test stabilization.
|
||||
- [@tidley](https://github.com/tidley): Nostr-mediated overlay
|
||||
discovery and UDP NAT traversal
|
||||
([#53](https://github.com/jmcorgan/fips/pull/53)).
|
||||
- [@alexxie16](https://github.com/alexxie16): peer ACL enforcement
|
||||
([#50](https://github.com/jmcorgan/fips/pull/50)),
|
||||
macOS WireGuard companion example
|
||||
([#51](https://github.com/jmcorgan/fips/pull/51)),
|
||||
follow-up ([#67](https://github.com/jmcorgan/fips/pull/67)).
|
||||
- [@osh](https://github.com/osh): diagnostic queries for security
|
||||
validation and mesh debugging
|
||||
([#42](https://github.com/jmcorgan/fips/pull/42)).
|
||||
- [@OceanSlim](https://github.com/0ceanSlim): Windows platform
|
||||
support ([#45](https://github.com/jmcorgan/fips/pull/45)).
|
||||
- [@mmalmi](https://github.com/mmalmi): ring AEAD backend
|
||||
([#80](https://github.com/jmcorgan/fips/pull/80)),
|
||||
hot-path drain batching + recvmmsg + eager pubkey_full
|
||||
([#81](https://github.com/jmcorgan/fips/pull/81)),
|
||||
TreeAnnounce self-root ancestry + overlay-advert retry hygiene
|
||||
([#82](https://github.com/jmcorgan/fips/pull/82)),
|
||||
NAT-traversal MTU inheritance
|
||||
([#83](https://github.com/jmcorgan/fips/pull/83)).
|
||||
- [@dskvr](https://github.com/dskvr): initial Arch Linux AUR
|
||||
packaging ([#21](https://github.com/jmcorgan/fips/pull/21)) and
|
||||
the AUR publish workflow.
|
||||
- [@SatsAndSports](https://github.com/SatsAndSports): rekey
|
||||
message-1 admit fix on non-accepting transports
|
||||
([#49](https://github.com/jmcorgan/fips/pull/49)),
|
||||
TreeAnnounce semantic validation, gateway test image fix
|
||||
([#69](https://github.com/jmcorgan/fips/pull/69)).
|
||||
- [@andrewheadricke](https://github.com/andrewheadricke): MIPS
|
||||
atomic-ABI portability via `portable_atomic`
|
||||
([#62](https://github.com/jmcorgan/fips/pull/62)).
|
||||
- [@sh1ftred](https://github.com/sh1ftred): Arch packaging namcap
|
||||
fixes ([#63](https://github.com/jmcorgan/fips/pull/63)).
|
||||
- [@oleksky](https://github.com/oleksky): macOS WireGuard companion
|
||||
collaboration on [#51](https://github.com/jmcorgan/fips/pull/51).
|
||||
|
||||
**Issue reports that drove fixes in this release**:
|
||||
|
||||
- [@deavmi](https://github.com/deavmi): MIPS daemon build support
|
||||
([#26](https://github.com/jmcorgan/fips/issues/26)).
|
||||
- [@Sebastix](https://github.com/Sebastix): fipsctl/fipstop
|
||||
control-socket path detection
|
||||
([#30](https://github.com/jmcorgan/fips/issues/30)).
|
||||
- [@SwapMarket](https://github.com/SwapMarket): auto-connect
|
||||
reconnect after graceful disconnect
|
||||
([#60](https://github.com/jmcorgan/fips/issues/60)) and
|
||||
fipsctl mesh-address rejection
|
||||
([#61](https://github.com/jmcorgan/fips/issues/61)).
|
||||
321
docs/releases/release-notes-v0.4.0.md
Normal file
@@ -0,0 +1,321 @@
|
||||
# FIPS v0.4.0
|
||||
|
||||
**Released**: 2026-06-27
|
||||
|
||||
v0.4.0 is the throughput-and-observability release on the v0.3.x wire
|
||||
format. It adds two new ways for nodes to find and reach each other (the
|
||||
Nym mixnet transport and opt-in mDNS LAN discovery), overhauls the data
|
||||
plane for higher single-node throughput and lower per-packet CPU, moves
|
||||
the entire operator read surface off the data-plane hot path so
|
||||
observability stays responsive under load, ships a reworked `fipstop`
|
||||
TUI, and hardens FMP and FSP rekey to be hitless under packet loss in
|
||||
both directions. It also folds in the accumulated mesh-convergence,
|
||||
admission-control, and packaging fixes from the maintenance line.
|
||||
|
||||
v0.4.0 is wire-compatible with v0.3.0. Mixed meshes interoperate; there
|
||||
is no flag-day upgrade. A deployed v0.3.0 node and an upgraded v0.4.0
|
||||
node peer, rekey, and route normally, so you can roll the upgrade out
|
||||
across a mesh in any order.
|
||||
|
||||
## At a glance
|
||||
|
||||
- New outbound Nym mixnet transport with a single-container demo and a
|
||||
new mixnet-relay example.
|
||||
- Opt-in mDNS / DNS-SD discovery on the local link.
|
||||
- Data-plane overhaul: off-task encrypt and decrypt worker pools, GSO,
|
||||
connected-UDP send path, copy-avoidance on receive, batched macOS
|
||||
receive.
|
||||
- The full `show_*` read surface now serves off the receive loop, so
|
||||
`fipsctl` and `fipstop` stay responsive on loaded nodes; a new
|
||||
counter-only `show_metrics` query enables a Prometheus scraper at no
|
||||
hot-path cost.
|
||||
- Reworked `fipstop` TUI on a machine-verified render-snapshot base.
|
||||
- Rekey is now hitless under loss and reordering in both directions.
|
||||
- New packaging targets: an OpenWrt `.apk` for OpenWrt 25+ and a Nix
|
||||
flake for reproducible from-source builds on Nix/NixOS.
|
||||
- Six route-class transit counters partition forwarded traffic by its
|
||||
tree relationship to the next hop, visible via `show_routing` and
|
||||
`show_status`.
|
||||
|
||||
## What's new
|
||||
|
||||
### Nym mixnet transport
|
||||
|
||||
FIPS can now peer over the [Nym](https://nymtech.net/) mixnet for
|
||||
metadata-resistant connectivity. The new `transports.nym` transport
|
||||
makes outbound connections through a `nym-socks5-client` SOCKS5 proxy
|
||||
that you run alongside the daemon (for example as a service running
|
||||
alongside the fips daemon, or as a sidecar container). The transport
|
||||
waits at startup for the nym-socks5-client to become ready before giving
|
||||
up.
|
||||
|
||||
This is a privacy and anonymity deployment mode chosen for its own
|
||||
properties. It mixes your FIPS traffic into the Nym cover-traffic
|
||||
network so that link-level observers cannot correlate which mesh peers
|
||||
are talking. A new `examples/sidecar-nostr-mixnet-relay/` demonstrates a
|
||||
FIPS-reachable Nostr relay peered across the mixnet end to end, and a
|
||||
single-container demo ships with the transport.
|
||||
|
||||
Enable it by adding a `transports.nym` instance and pointing it at your
|
||||
running nym-socks5-client. See the transports reference for the field
|
||||
set.
|
||||
|
||||
### mDNS LAN discovery
|
||||
|
||||
Nodes on a shared local link can now find each other with zero address
|
||||
configuration. The opt-in `node.discovery.lan` path runs an mDNS /
|
||||
DNS-SD responder and browser: each node advertises a FIPS service record
|
||||
on the link and adopts the peers it discovers. This complements the
|
||||
existing Nostr-mediated overlay discovery for the common case where the
|
||||
peers are simply on the same LAN.
|
||||
|
||||
Turn it on with `node.discovery.lan.enabled: true`. `service_type` and
|
||||
`scope` tune the advertised service record and which interfaces
|
||||
participate. Discovery on the local link needs no relay and no STUN.
|
||||
|
||||
### Data-plane throughput overhaul
|
||||
|
||||
The receive and send paths were reworked for higher single-node
|
||||
throughput and lower per-packet CPU, building on the v0.3.0
|
||||
crypto-backend swap:
|
||||
|
||||
- **Off-task encrypt and decrypt.** Per-peer encrypt and decrypt now run
|
||||
on dedicated worker tasks rather than inline on the receive loop, so a
|
||||
single busy peer no longer serializes the whole node's crypto.
|
||||
- **GSO and connected-UDP send.** The Linux send path uses generic
|
||||
segmentation offload and a connected-UDP socket where available,
|
||||
cutting syscall overhead on bulk flows.
|
||||
- **Copy-avoidance on receive.** The receive hot path avoids buffer
|
||||
copies it previously made per packet.
|
||||
- **Batched macOS receive.** macOS gains a `recvmsg_x` batched receive,
|
||||
mirroring the Linux `recvmmsg` batching from v0.3.0.
|
||||
- **Shared immutable-state context and an atomic metric registry.**
|
||||
Immutable per-node state moved into a single shared context, and
|
||||
counters live in an atomic metric registry that the new `show_metrics`
|
||||
query reads without touching the hot path.
|
||||
|
||||
These are all internal to the data plane and require no operator action.
|
||||
|
||||
### Observability off the hot path
|
||||
|
||||
Every read-only control query now renders from a snapshot published once
|
||||
per tick into a lock-free `ArcSwap`, served from the control accept task
|
||||
instead of round-tripping the data-plane receive loop. This covers
|
||||
`show_status`, `show_stats_*`, `show_peers`, `show_sessions`,
|
||||
`show_links`, `show_connections`, `show_transports`, `show_mmp`,
|
||||
`show_tree`, `show_bloom`, `show_cache`, `show_routing`,
|
||||
`show_identity_cache`, `show_acl`, `show_listening_sockets`, and the new
|
||||
`show_metrics`. Only the mutating `connect` and `disconnect` commands
|
||||
still reach the loop.
|
||||
|
||||
The practical effect: on a loaded node where the receive loop was busy,
|
||||
`fipsctl` and `fipstop` queries previously stalled or timed out (the
|
||||
five-second query pattern operators saw). They now answer promptly
|
||||
regardless of data-plane load. Per-entity snapshots reuse unchanged rows
|
||||
by pointer, so the per-tick publish cost stays bounded as peer and
|
||||
session counts grow.
|
||||
|
||||
A new **`show_metrics`** query (surfaced as `fipsctl stats metrics`)
|
||||
returns a counter-only snapshot of every metric family. It is the
|
||||
enabler for a Prometheus scraper that pulls node counters at no hot-path
|
||||
cost.
|
||||
|
||||
Six **route-class transit counters** partition transit-forwarded packets
|
||||
by their tree relationship to the chosen next hop — tree-up, tree-down,
|
||||
tree-down-cross, cross-link descend, cross-link ascend, and direct-peer
|
||||
— and the six classes sum to `forwarded_packets`. They surface through
|
||||
`show_routing` and `show_status`, and the `fipstop` routing tab is
|
||||
reorganized so its two columns separate own/endpoint traffic from
|
||||
forwarded/transit traffic with the tree-down-cross line visually flagged.
|
||||
|
||||
### Reworked fipstop TUI
|
||||
|
||||
`fipstop` gets a rendering, navigation, and read-surface overhaul on a
|
||||
machine-verified base: a render-snapshot harness asserts the exact text
|
||||
grid and per-cell style of every view against canned control-socket
|
||||
output. New daemon-resolved fields surface through the snapshots,
|
||||
including effective persistence, root and is-root state, a
|
||||
per-transport-type peer-count map, per-peer effective depth, the root
|
||||
npub, and the last-sent uptree filter fill ratio with the subtree size
|
||||
estimate.
|
||||
|
||||
A separate fix clears a garbled-screen problem on startup and stray
|
||||
bytes on quit, most visible over SSH and inside tmux: startup now forces
|
||||
a full repaint before the first draw, and quit stops and joins the
|
||||
stdin-poll thread before restoring the terminal, so post-raw-mode
|
||||
keystrokes no longer echo onto the restored screen.
|
||||
|
||||
### Rekey reliability
|
||||
|
||||
FMP and FSP session rekey are now hitless under packet loss and
|
||||
reordering in both directions:
|
||||
|
||||
- Inbound frames are authenticated against the pending session before
|
||||
the K-bit cutover promotes it, so a spoofed or stale frame cannot
|
||||
derail a rekey in progress.
|
||||
- Rekey message-1 retransmission is bounded, and the link-dead heartbeat
|
||||
is rekey-aware so an in-flight rekey is not mistaken for a dead link.
|
||||
- FSP session rekey holds connectivity across the rekey window under
|
||||
loss and reordering.
|
||||
- Dual-initiation races (both peers starting a rekey at once on a
|
||||
high-latency link) are desynchronized with symmetric jitter so the two
|
||||
sides converge on one session rather than fighting.
|
||||
- An exhausted retransmission-budget abort, an expected and self-limiting
|
||||
outcome on lossy or high-latency links, is logged at debug rather than
|
||||
warn.
|
||||
|
||||
The net operator takeaway: rekey completes cleanly without dropping
|
||||
traffic, even on lossy or high-latency links, and the log no longer
|
||||
cries wolf when a rekey gives up and retries.
|
||||
|
||||
### New packaging targets
|
||||
|
||||
- **OpenWrt `.apk`.** A new `.apk` package targets OpenWrt 25+, where
|
||||
apk-tools is the mandatory package manager; the existing `.ipk`
|
||||
continues to cover OpenWrt 24.x and earlier. It is built SDK-free,
|
||||
reusing the `.ipk` cross-compile and installed-filesystem payload, and
|
||||
releases publish `.apk` artifacts and checksums alongside `.ipk`. Like
|
||||
the `.ipk`, the package is unsigned and installed with
|
||||
`apk add --allow-untrusted`.
|
||||
- **Nix flake.** A `flake.nix` at the project root builds all four
|
||||
binaries (`fips`, `fipsctl`, `fips-gateway`, `fipstop`) from source on
|
||||
Nix/NixOS, pinning the exact toolchain and wiring the native build
|
||||
dependencies so no host setup is needed beyond Nix with flakes
|
||||
enabled. It exposes `nix build`, `nix run`, a `nix develop` dev shell,
|
||||
and `nix flake check`, with `flake.lock` committed for reproducibility.
|
||||
|
||||
## Behavior changes worth flagging
|
||||
|
||||
These affect operators on upgrade.
|
||||
|
||||
- **Bloom filter antipoison cap raised.** `node.bloom.max_inbound_fpr`
|
||||
moves from 0.05 to 0.10, accepting filters with a higher derived
|
||||
false-positive rate before rejecting them. This reduces spurious
|
||||
filter rejections on larger meshes while keeping the antipoison
|
||||
protection in place.
|
||||
- **TCP inbound cap honors `max_connections`.** The TCP inbound accept
|
||||
ceiling now resolves from explicit per-transport
|
||||
`max_inbound_connections`, then node-wide
|
||||
`node.limits.max_connections`, then the built-in default of 256.
|
||||
Previously the TCP inbound ceiling was hardwired to 256 and ignored
|
||||
`max_connections`, so raising it had no effect on inbound TCP.
|
||||
- **Static host aliases hot-reload.** `/etc/fips/hosts` now reloads on
|
||||
mtime change once per tick rather than only at startup, so display
|
||||
names in `fipsctl` and `fipstop` reflect edits without a daemon
|
||||
restart. The peer ACL reloads through the same lock-free snapshot
|
||||
mechanism.
|
||||
- **Quieter logs on busy public-mesh nodes.** Routine per-peer
|
||||
connection-lifecycle and capacity-cap events, no-route session-datagram
|
||||
drops, and exhausted rekey-budget aborts are demoted to debug, so
|
||||
genuinely notable info and warn lines are no longer drowned out.
|
||||
- **More visible drops.** Receive-path silent rejections now flow
|
||||
through typed reject-reason counters, and discovery counts requests
|
||||
dropped when the dedup cache is full (`req_dedup_cache_full`, visible
|
||||
via `show_routing`). Drops that were previously silent are now
|
||||
countable.
|
||||
- **Tor connect-refused accounting.** The Tor transport increments its
|
||||
`connect_refused` statistic (the "Refused" line in `fipstop`) on an
|
||||
actively-refused SOCKS5 connect, instead of recording every connect
|
||||
failure as a generic SOCKS5 error.
|
||||
|
||||
## Notable bug fixes
|
||||
|
||||
The CHANGELOG has the exhaustive list. This is the operator-relevant
|
||||
subset of fixes for behavior that shipped in v0.3.0.
|
||||
|
||||
- **Symmetric peer teardown on manual disconnect.** A manual
|
||||
`fipsctl disconnect` now sends the peer a scoped Disconnect so both
|
||||
ends tear down and re-handshake cleanly. Previously a manual
|
||||
disconnect tore down only the local side, leaving the peer with a
|
||||
stale session that was never re-adopted as a child and whose bloom
|
||||
filter was never re-recorded.
|
||||
- **Gateway holds long-lived and DNS-cached mappings.** `fips-gateway`
|
||||
no longer drops a virtual-IP mapping while traffic is still flowing.
|
||||
The mapping TTL clock previously advanced only on DNS re-query, so a
|
||||
busy long-lived or DNS-cached client could have its mapping reclaimed
|
||||
mid-flow. The tick now refreshes the mapping whenever conntrack reports
|
||||
active sessions and recovers a draining mapping to active when traffic
|
||||
resumes; only genuinely idle mappings drain.
|
||||
- **Accurate mesh-size estimate under filter overlap.** The mesh-size
|
||||
estimator now estimates the cardinality of the OR-union of self plus
|
||||
every connected peer's inbound filter, instead of summing per-filter
|
||||
cardinalities of tree peers. Summing assumed the filters were disjoint,
|
||||
so a stale or oversized parent filter or a routing loop inflated the
|
||||
reported mesh size and a tree rebalance flapped the count. OR-union
|
||||
deduplicates overlap, equals the old result in the disjoint case, and
|
||||
removes the estimate's dependence on tree-declaration cache freshness.
|
||||
- **Single-uplink node reattaches within a round-trip.** A node with one
|
||||
tree peer, which has periodic parent re-evaluation disabled, was left
|
||||
self-rooted and unreachable if its one-shot attaching TreeAnnounce was
|
||||
lost, until the next periodic re-broadcast. Tree-position exchange is
|
||||
now self-healing on the receive path: a node that hears an announce
|
||||
advertising a strictly worse root echoes its own declaration back,
|
||||
provoking the better-rooted peer to re-push its real position
|
||||
immediately.
|
||||
- **macOS self-connections work end to end (#117).** Traffic a macOS
|
||||
node sends to its own `<npub>.fips` address is now delivered locally
|
||||
for full TCP/UDP, not just `ping6`. The point-to-point `utun` egresses
|
||||
self-addressed packets into the daemon with an unfinished transport
|
||||
checksum (macOS offloads it on the `lo0` loopback route), so
|
||||
re-injecting them verbatim made the local stack drop every segment the
|
||||
MSS-clamp rewrite did not happen to fix and self-connections
|
||||
half-opened and hung. The hairpin path now recomputes the TCP/UDP
|
||||
checksum before re-injection. Linux was unaffected.
|
||||
|
||||
## Upgrade notes
|
||||
|
||||
Operator-actionable items moving from v0.3.0 to v0.4.0:
|
||||
|
||||
- **Wire-compatible, no flag day.** v0.4.0 peers with v0.3.0. Upgrade
|
||||
nodes in any order. During a rolling upgrade you may see some log lines
|
||||
on the upgraded side as it interacts with not-yet-upgraded peers;
|
||||
behavior is correct, log noise only.
|
||||
- **Bloom antipoison cap default changed.** `node.bloom.max_inbound_fpr`
|
||||
now defaults to 0.10 (was 0.05). If you set this explicitly, review
|
||||
whether you still want the old value.
|
||||
- **New optional config surfaces.** `transports.nym` (outbound Nym
|
||||
mixnet) and `node.discovery.lan` (mDNS LAN discovery) are both opt-in
|
||||
and off by default. Adding them is the only way to turn the new paths
|
||||
on.
|
||||
- **TCP inbound cap.** If you relied on the old hardwired 256 inbound-TCP
|
||||
ceiling, note it now honors `max_inbound_connections` then
|
||||
`node.limits.max_connections` then 256.
|
||||
- **New observability query.** `fipsctl stats metrics` (the
|
||||
`show_metrics` control query) returns a counter-only snapshot suitable
|
||||
for a scraper.
|
||||
|
||||
## Getting v0.4.0
|
||||
|
||||
- **Linux x86_64 / aarch64**: `.deb` and tarball at the
|
||||
[v0.4.0 release page](https://github.com/jmcorgan/fips/releases/tag/v0.4.0).
|
||||
- **Arch Linux**: `fips` from the AUR.
|
||||
- **macOS**: `.pkg` at the v0.4.0 release page.
|
||||
- **Windows**: ZIP at the v0.4.0 release page.
|
||||
- **OpenWrt**: `.ipk` (OpenWrt 24.x and earlier) or `.apk` (OpenWrt 25+)
|
||||
at the v0.4.0 release page.
|
||||
- **From source**: `cargo build --release` from a checkout of the v0.4.0
|
||||
tag (Rust 1.94.1 per `rust-toolchain.toml`; `libclang-dev` is a
|
||||
required Linux build prerequisite).
|
||||
- **Nix / NixOS**: `nix build .#fips` from a checkout of the v0.4.0 tag
|
||||
builds the binaries from source with the pinned toolchain and no manual
|
||||
prerequisites (see the Nix section of `packaging/README.md`).
|
||||
|
||||
The full per-commit changelog lives in
|
||||
[`CHANGELOG.md`](../../CHANGELOG.md). Issues and discussion at
|
||||
[github.com/jmcorgan/fips](https://github.com/jmcorgan/fips).
|
||||
|
||||
## Contributors
|
||||
|
||||
Thanks to everyone who contributed code, packaging work, bug reports, or
|
||||
reviews to this release.
|
||||
|
||||
- [@jcorgan](https://github.com/jmcorgan): release shepherd, high-level
|
||||
design, control read plane, rekey hardening, admission, bug fixes,
|
||||
testing, packaging, PR coordination, and issue resolution.
|
||||
- [@mmalmi](https://github.com/mmalmi): opt-in mDNS LAN discovery and
|
||||
data-plane performance work.
|
||||
- [@Origami74](https://github.com/Origami74): macOS packaging and
|
||||
website coordination.
|
||||
- [@dskvr](https://github.com/dskvr): AUR packaging.
|
||||
- [@oleksky](https://github.com/oleksky): Nym mixnet transport and the
|
||||
single-container mixnet demo.
|
||||
146
docs/releases/release-notes-v0.4.1.md
Normal file
@@ -0,0 +1,146 @@
|
||||
# FIPS v0.4.1
|
||||
|
||||
**Released**: 2026-07-19
|
||||
|
||||
v0.4.1 is a maintenance release on the v0.4.x line. It raises the default
|
||||
antipoison cap on inbound bloom filter announcements, removes a redundant
|
||||
spanning-tree metric counter, fixes two convergence and path-MTU bugs, and
|
||||
cuts per-packet CPU in the bloom and identity paths. There is no wire
|
||||
format change and no new feature surface.
|
||||
|
||||
v0.4.1 is wire-compatible with v0.4.0. Nodes can be upgraded one at a time
|
||||
with no coordinated restart, though one behavior change below is worth
|
||||
reading before you start a rolling upgrade.
|
||||
|
||||
## At a glance
|
||||
|
||||
- `node.bloom.max_inbound_fpr` default moves from `0.10` to `0.20`.
|
||||
- The `parent_switched` metric counter is gone. Use `parent_switches`.
|
||||
- Spanning tree no longer serves stale coordinates after a parent link is
|
||||
lost through peer removal.
|
||||
- Discovery no longer loosens a path MTU clamp it had correctly tightened.
|
||||
- Bloom probing and identity operations do measurably less work per call,
|
||||
with identical results.
|
||||
|
||||
## Behavior changes worth flagging
|
||||
|
||||
### The inbound filter FPR cap default doubles again
|
||||
|
||||
`node.bloom.max_inbound_fpr` goes from `0.10` to `0.20`. The cap rejects
|
||||
inbound `FilterAnnounce` frames whose advertised false positive rate
|
||||
exceeds it. On the fixed 1 KB, k=5 filter, `0.10` corresponds to a fill of
|
||||
0.631 and roughly 1,630 reachable entries, and the busiest nodes'
|
||||
aggregates had started reaching that ceiling as the mesh grew. `0.20`
|
||||
corresponds to a fill of 0.7248 and roughly 2,114 entries.
|
||||
|
||||
Be aware that this is the second time in two releases that this default
|
||||
has doubled, for the same reason both times. That is worth stating plainly
|
||||
rather than repeating the previous release's framing: raising the cap buys
|
||||
headroom, it does not fix anything. The real constraint is the fixed 1 KB
|
||||
filter size, which is a protocol constant. The structural remedy is the v2
|
||||
filter work, where filter capacity scales with the mesh instead of being
|
||||
pinned. This release is an interim step to keep legitimate aggregates from
|
||||
being rejected until that lands. It is not the start of a pattern of
|
||||
raising the cap once per release, and if you are sizing capacity planning
|
||||
around this number, plan against the v2 work rather than against a third
|
||||
raise.
|
||||
|
||||
The antipoison property the cap exists for is preserved. A saturated or
|
||||
deliberately poisoned filter still presents an FPR near 100% and is still
|
||||
rejected.
|
||||
|
||||
**This matters during a rolling upgrade.** A v0.4.1 node accepts a
|
||||
`FilterAnnounce` with a derived FPR between 0.10 and 0.20; a v0.4.0 node
|
||||
drops the same frame, and the drop is silent on the wire with no NACK. The
|
||||
cap also gates the mesh size estimator, which declines to produce a value
|
||||
when any contributing filter is over the cap. So while a mesh is partly
|
||||
upgraded, upgraded and not-yet-upgraded nodes can legitimately report
|
||||
different mesh sizes, or one can report a size while the other reports
|
||||
unknown. This resolves once every node is on v0.4.1. If you want to avoid
|
||||
the window entirely, set `node.bloom.max_inbound_fpr: 0.10` explicitly in
|
||||
your config before upgrading and remove it after the last node is done.
|
||||
|
||||
### The `parent_switched` counter is removed
|
||||
|
||||
`parent_switched` was incremented on the line immediately before
|
||||
`parent_switches` at every site and never independently, so the two
|
||||
counters always held the same value. `parent_switched` is now gone from
|
||||
the tree metrics, the control socket snapshot, and the `fipstop` tree
|
||||
view. `parent_switches` remains and is unchanged.
|
||||
|
||||
If you scrape the control socket, or have dashboards or alerts referencing
|
||||
`parent_switched`, point them at `parent_switches`. Anything still asking
|
||||
for `parent_switched` will find nothing rather than a zero.
|
||||
|
||||
## Notable bug fixes
|
||||
|
||||
### Stale coordinates after losing a parent through peer removal
|
||||
|
||||
When a node's parent link dropped via peer removal, the node correctly
|
||||
reparented or self-rooted, but skipped the coordinate cache invalidation
|
||||
that every other position-change path performs. Cached entries for
|
||||
downstream destinations kept the node's old coordinate prefix. This did
|
||||
not self-correct the way a stale cache entry normally would: routing
|
||||
access refreshes an entry's TTL, so an entry that was actively being
|
||||
routed through never expired, and was only fixed by an unrelated fresh
|
||||
insert. Both invalidation classes now run on this path, matching the
|
||||
loop-detection branch.
|
||||
|
||||
### Discovery could loosen a tightened path MTU clamp
|
||||
|
||||
An originator handling a `LookupResponse` overwrote its cached path MTU
|
||||
unconditionally. If a reactive `MtuExceeded` or `PathMtuNotification` had
|
||||
already taught it a tighter value, a later, looser discovery estimate
|
||||
would clobber that and re-loosen the clamp, risking a return to dropped
|
||||
oversized packets. The cached and received values are now compared and the
|
||||
tighter one is kept.
|
||||
|
||||
## Upgrade notes
|
||||
|
||||
This is a drop-in upgrade from v0.4.0 with no wire format change, no
|
||||
config migration, and no coordinated restart. Upgrade nodes in whatever
|
||||
order you like.
|
||||
|
||||
Two things to do rather than assume:
|
||||
|
||||
1. If you monitor `parent_switched`, move to `parent_switches` before
|
||||
upgrading, or your dashboards will go blank rather than error.
|
||||
2. During the rolling window, expect upgraded and not-yet-upgraded nodes
|
||||
to potentially disagree about mesh size, per the FPR cap section above.
|
||||
This is expected and self-resolves. Do not chase it as a bug unless it
|
||||
persists after every node reports `0.4.1`.
|
||||
|
||||
If you have pinned `node.bloom.max_inbound_fpr` explicitly in your config,
|
||||
your setting is honored and nothing changes for you. The change only
|
||||
affects nodes taking the default.
|
||||
|
||||
Downgrading to v0.4.0 is supported and needs no special handling.
|
||||
|
||||
## Getting v0.4.1
|
||||
|
||||
- **Linux x86_64 / aarch64**: `.deb` and tarball at the
|
||||
[v0.4.1 release page](https://github.com/jmcorgan/fips/releases/tag/v0.4.1).
|
||||
- **Arch Linux**: `fips` from the AUR.
|
||||
- **macOS**: `.pkg` at the v0.4.1 release page.
|
||||
- **Windows**: ZIP at the v0.4.1 release page.
|
||||
- **OpenWrt**: `.ipk` (OpenWrt 24.x and earlier) or `.apk` (OpenWrt 25+)
|
||||
at the v0.4.1 release page.
|
||||
- **From source**: `cargo build --release` from a checkout of the v0.4.1
|
||||
tag (Rust 1.94.1 per `rust-toolchain.toml`; `libclang-dev` is a
|
||||
required Linux build prerequisite).
|
||||
- **Nix / NixOS**: `nix build .#fips` from a checkout of the v0.4.1 tag
|
||||
builds the binaries from source with the pinned toolchain and no manual
|
||||
prerequisites (see the Nix section of `packaging/README.md`).
|
||||
|
||||
The full per-commit changelog lives in
|
||||
[`CHANGELOG.md`](../../CHANGELOG.md). Issues and discussion at
|
||||
[github.com/jmcorgan/fips](https://github.com/jmcorgan/fips).
|
||||
|
||||
## Contributors
|
||||
|
||||
Thanks to everyone who contributed code, packaging work, bug reports, or
|
||||
reviews to this release.
|
||||
|
||||
- [@jcorgan](https://github.com/jmcorgan): release shepherd, spanning-tree
|
||||
and discovery fixes, bloom and identity performance work, antipoison cap
|
||||
change, and testing.
|
||||