Files
fips/docs/design/fips-mtu.md
T
Johnathan Corgan 6a564e26ac Prepare the v0.5.0 release content
Everything the release needs except the version number, which stays at
0.5.0-dev until the tag.

The changelog entry covers only the work that is new on this line. The
point release's forty-six entries arrived under their own heading with the
forward merge and are left alone; the twenty that remained are regrouped by
topic and eight more added for changes no entry covered. Three of those
eight matter to someone upgrading. Five root modules and four re-exports
left the public library surface and Node::connections narrowed, none of it
recorded anywhere; the entry names what to use instead and distinguishes
the removed connection-phase enum from the Noise type of the same name,
which is a different type that still exists. Tracing targets moved, so an
existing RUST_LOG filter stops matching rather than erroring. And the
handshake resend interval key no longer governs the first resend, which is
now a constant, though it still governs later ones.

Seven more entries cover the work that landed after the first content pass
was written: the experimental native datagram API, the fipsctl probe
diagnostic, per-instance transport addressing, the app-owned UDP socket
seam, and the connect, disconnect and path-MTU fixes. The four bug fixes
among them all reach the deployed line, so the release notes no longer
claim this release carries exactly one fix for a shipped bug; it carries
four.

There is no security section, because after the split every security entry
belongs to the point release. The release notes say so plainly rather than
leaving a reader upgrading across both releases to conclude this one
carries no security work.

The notes are organized by audience, since the release spans OpenWrt
routers, embedders, FreeBSD, and the existing platforms, and a single list
serves none of them. The native datagram API is given a section of its own
rather than folded into the embedding seam: it is a client-facing API
rather than a way to host a node, and its one rule with no Berkeley-socket
counterpart, that the v1 wire carries no half-close, needs to be somewhere
a client author will read it. FreeBSD is advertised as supported on x86_64
only, stated wherever the platform appears. Android is advertised as an
embedding seam and not as a supported platform: a compile-gated library
surface with no artifact and no host application guide.

The configuration table rename is carried through every shipped file that
taught the old spelling: nine documentation files, the OpenWrt sample
config and a test generator, twenty-two sites in all. Guides written this
same cycle were among them, which is how the omission was found. The
documentation that arrived with the native API was checked for the same
omission and was already clean. The compatibility tests keep the old
spelling deliberately, since they exist to test the fold.

The changelog section is the fold of master's [Unreleased], not a snapshot
of it. An earlier version of this commit took a copy that then drifted, so
each section ended up holding a bullet the other did not and re-folding
them would have picked a winner silently. Both causes were fixed on master
instead — the NixOS module had never been recorded there, and the
pre-release batch of fixes was new — so [Unreleased] is a strict superset
and this is a copy rather than a merge. [0.5.0] carries all forty-six
bullets byte for byte, [Unreleased] is empty, and [0.4.2] is untouched,
checked by hashing it against master's copy.

The BLE work landed after the content pass and gets one summary entry in
the changelog and one section in the release notes rather than nine
bullets: the ble_available gate replacing target_os = "linux",
packet-boundary recovery for stream-oriented backends, peer recognition by
node identity instead of a rotating link address, the L2CAP PSM moving
into the backend seam and onto the advertisement, the embedder-supplied
Android radio, bounded probe retry, and inbound handshakes moved off the
accept loop.

The two release-notes copies no longer share their link paths. Relative
links resolve from one directory only, so the seven written for
docs/releases/ all 404ed from the root copy. The root copy now uses paths
from the repository root and the versioned copy keeps the ../ form; both
sets were resolved against the tree. The same two links are broken the
same way in the v0.4.0 through v0.4.2 notes, left as shipped history.

The contributor tallies are re-derived against maint..HEAD rather than
adjusted: twenty commits from outside the project and 171 from me, with
Arjen at fifteen and fr34aky at two. An earlier count of twelve and 138
was carried from a measurement taken three days before this content was
written, and the BLE branch widened the gap after it. Arjen's NixOS flake
module, the UDP sin6_scope_id fix and most of the BLE rework were
uncredited, as was fr34aky's L2CAP PSM seam. They want one last re-derive
at tag time if anything lands before the tag.

A sweep of all 99 tracked markdown files against the tree corrected
fifty-three of them. Four told the reader to run a build.sh that does not
exist; the only harness builder is testing/scripts/build.sh. The BLE build
prerequisites were described as optional on the strength of a probe that
build.rs does not perform, and bluez was named a build prerequisite when
libdbus-sys asks only for libdbus-1-dev and pkg-config and bluez is the
runtime daemon. Link cost is the primary sort key in next-hop ranking, not
reserved for future use; Ethernet runs on macOS as well as Linux; the BLE
MTU is the L2CAP CoC MTU rather than a negotiated ATT_MTU; effective
Ethernet MTU is 1497; the LAN discovery subsystem is src/mdns and eight
citations still named a src/discovery that never existed here. The
connectivity states in three tutorials were invented, and their jq filters
matched nothing including healthy peers. One command filtered on a literal
fd97: address prefix, which only the first byte of fixes, so it returned
empty for all but one reader in 256 and every later step using the
variable failed silently. transports.tor.advertise_on_nostr was
undocumented despite being validated against node.rendezvous.nostr.enabled.

The transport design document gains the BLE section it never had, written
from the source: the backend cascade and its compile_error tripwire, the
platform gate, the PSM advertisement wire layout and the byte budget that
forces a 16-bit service-data key, and the probe and admission bounds.

Three source files carried the same class of staleness and are corrected
with the documentation: the OpenWrt ipk usage line and Makefile error text
both named a packaging/openwrt that does not exist, and chaos.sh parsed
--subnet without listing it.

Folded in with the content commit, having been prepared alongside it:

The three GitHub Action pins that had gone stale. Every third-party
action is pinned to a commit SHA, nothing reports that a pin has aged,
and re-resolving all ten against their tags found dorny/test-reporter@v2,
taiki-e/install-action@v2 and vmactions/freebsd-vm@v1 had moved. The
three install-action@nextest references stay unpinned, since that action
reads the tool to install from the ref name. check-action-pins.sh passes
at 75 references and all nine workflow files parse.

The lockfile refresh, which is the mutating half of the dependency sweep.
Thirty-six packages move to their latest semver-compatible versions and
every one is transitive; nothing declared in Cargo.toml changes version.
No advisory forces any of them. It was taken before the validation
battery, because a gate run against a lockfile that later moves proves
nothing about what ships.

The sha2 0.10 to 0.11, hkdf 0.12 to 0.13 and bech32 0.11 to 0.12 majors,
three of the four deferred at v0.4.0 for change surface rather than
security. All three land with no source change. sha2 and hkdf must move
together, since both depend on digest 0.11, and neither changes an
algorithm. That matters because the chaining-key KDF in the Noise
handshake is built on Hkdf::<Sha256>, where an output change would be a
wire break rather than a compile error; no known-answer vectors exist for
that path, so the wire-compatibility gate is what covers it. secp256k1
0.31 is deliberately absent, since nostr's own requirement would leave
two copies of the ECC library in the tree.

The README support matrix, rebuilt as one feature table broken out by
Linux variety. A single Linux column hid that Debian, Ubuntu, Arch and
NixOS are one glibc build differing in packaging, that OpenWrt is musl
and drops BLE, and that Android is not a daemon platform. Transport rows
sort by how many platforms carry them. A Native API row reads its
platform set from the cfg gates. The installer row becomes a package
format row naming the artifact, and only the .deb is exercised per
release.

Four changelog and release-note gaps the BLE re-walk found: a Bluetooth
LE bullet stranded inside the released 0.4.2 section, a missing Fixed
entry for the scan and probe loop counting a pool-refused connection as
an established link, the unnamed embedder call that installs an
application-owned radio, and the fact that stopping the transport now
stops scanning as well as advertising.

Three release-document gaps found walking the unsurveyed commits: the UDP
reuse-flag fix stated in the direction opposite to the one it was made,
with the silent second-daemon bind it prevents left unsaid; the corrected
native-API socket paragraph carried into both release-note copies, which
still named SOCK_SEQPACKET on FreeBSD and two kernels where three are
handled; and the coordinate-cache hardening, which shipped with no text
anywhere despite adding four operator-visible status fields. That last
entry states plainly that the checks are mitigations and not a closure,
since the coordinate is still not authenticated.

Also folded in, the documentation pass that followed the content commit:

A stage-pipeline diagram for the probe, embedded in the fipsctl
reference under the five-stage list. It draws the five stages left to
right with each stage's failure reasons below it, and the bypass that
skips both lookup stages when the coordinates are cached or the target
is a direct peer. Its branches come from the probe state machine rather
than from the report, so the path stage is drawn as the one failure that
does not stop the probe.

A rewrite of the README's "What FIPS does" section. It now opens with
what a machine running FIPS gets, rather than with the two deployment
modes, and gives the self-organizing and permissionless property its own
paragraph since it holds for both modes.

A regrouping of the README's feature list into the mesh, getting traffic
onto it, and running a node, with a bullet added for the native datagram
API, which had none despite sitting in the support matrix. The Quick
start now leads with the released packages rather than a source build.
It also fixes a real defect: the package enables fips.service and
fips-dns.service and starts neither on a fresh install, so .fips name
resolution was silently dead until the next reboot and neither page said
to start the service.

A rewrite of the release notes. They opened with seven subsections of
upgrade caveats and reached the first feature two hundred lines in; they
now open with a summary of the release and elaborate below it in the
same order. Android is stated as supported through an embedded crate
rather than as a standalone daemon, consistently across all three
documents. The OpenWrt pair is corrected: it is 802.11s between routers
with FIPS supplying encryption, authentication and routing, plus a
convention of an open !FIPS SSID a client joins over WiFi, not meshing
over a router's own radios. The probe's path output is described as the
least-common-ancestor walk, which is the worst-case fallback route
rather than the route a packet takes. Detail that did not change what a
reader does was cut from the notes and kept in the changelog.
2026-08-30 10:42:59 +00:00

14 KiB
Raw Blame History

FIPS Path MTU and Encapsulation Overhead

MTU is a cross-cutting concern in FIPS. No single layer owns it: the transport reports per-link MTU, FMP propagates path_mtu along forward and reverse paths, FSP echoes the observed path MTU end-to-end back to the source, and the IPv6 adapter enforces the resulting effective MTU at the TUN interface. This document is the canonical home for the unified MTU model.

For operator-facing diagnostic recipes (interpreting MtuExceeded counters, tuning IPv6 application MSS, troubleshooting cold-flow oversize), see the relevant how-to under docs/how-to/.

The MTU Problem in FIPS

A FIPS path can traverse heterogeneous link types — UDP/IP (1280 default, IPv6 minimum), Ethernet (interface MTU 3, typically 1497), BLE (per-connection L2CAP CoC MTU, 2048 default), Tor stream (1400 default), radio (51222) — within a single end-to-end session. The minimum MTU along the path determines the largest datagram a session can deliver. Several properties make this harder than in classic IP networks:

  • No fragmentation. FIPS does not fragment at transit nodes (see No fragmentation policy). A datagram that exceeds the next-hop link MTU is dropped, and the source is signaled.
  • Forward/reverse path asymmetry. After tree reconvergence the return path may diverge from the forward path, so the bottleneck on each direction can differ.
  • First-flow race. The very first SessionDatagram races destination discovery — the source has not yet learned the path MTU but must pick a payload size for the queued packet.
  • Variable per-link MTU. Some transports (BLE, TCP via TCP_MAXSEG) report different MTUs for different links rather than a single transport-wide value.

The unified MTU model below combines proactive and reactive mechanisms to converge on a working effective MTU within the first few packets of a session, then maintain it across topology changes.

Encapsulation Overhead

The byte budget for a FIPS-encapsulated packet:

Layer Overhead Purpose
Link encryption 37 bytes 16-byte outer header + 5-byte inner header (timestamp + msg_type) + 16-byte AEAD tag
SessionDatagram body 35 bytes ttl + path_mtu + src_addr + dest_addr (msg_type counted in inner header)
FSP header 12 bytes 4-byte prefix + 8-byte counter (used as AEAD AAD)
FSP inner header 6 bytes 4-byte timestamp + 1-byte msg_type + 1-byte inner_flags (inside AEAD)
Session AEAD tag 16 bytes ChaCha20-Poly1305 tag on session-encrypted payload
Protocol envelope 106 bytes FIPS_OVERHEAD constant — the base payload budget for any service

FIPS_OVERHEAD = 106 is the constant the rest of the system reasons about. Coordinate piggybacking via the CP flag adds variable extra overhead — 2 + entries × 16 bytes per coordinate, with both source and destination coordinates carried — and the send path skips the CP flag if adding coords would exceed the transport MTU.

Service-specific overheads layer on top of FIPS_OVERHEAD:

Service Overhead Note
DataPacket port header +4 bytes Always present for port-multiplexed services
IPv6 compression 33 bytes 40-byte IPv6 header → 7-byte format + residual
IPv6 effective overhead 77 bytes FIPS_IPV6_OVERHEAD constant

See fips-ipv6-adapter.md for the IPv6 compression scheme that lets the adapter reach FIPS_IPV6_OVERHEAD.

Each transport implements two MTU methods on its trait:

  • mtu() -> u16 — Transport-wide default MTU.
  • link_mtu(addr: &TransportAddr) -> u16 — Per-link MTU for a specific remote address. The default implementation falls back to mtu(), so transports with uniform MTU (UDP, raw Ethernet) need not override it.

FMP uses link_mtu() when it needs to reason about a specific outbound link — typically for path_mtu annotation in SessionDatagram and LookupResponse. Per-transport defaults:

Transport Default MTU Per-link MTU source
UDP 1280 (IPv6 minimum) uniform (mtu() fallback)
Ethernet interface MTU 3 (typically 1497) uniform
TCP 1400 derived from TCP_MAXSEG per connection
Tor 1400 uniform
BLE 2048 default; per-connection L2CAP CoC MTU per-link (overrides mtu())

For TCP, the per-connection TCP_MAXSEG query lets FMP discover the actual MSS the kernel negotiated for each connection, rather than assuming a single value across all TCP peers.

Proactive PMTUD: SessionDatagram path_mtu

Every SessionDatagram and LookupResponse carries a 2-byte path_mtu field. The source initializes it to its outbound link MTU; each transit node applies min(current, link_mtu(next_hop)) before forwarding. The destination receives the forward-path minimum.

For SessionDatagram, the receiver of the forward-path minimum is the session-layer destination, which then echoes the value back to the source via PathMtuNotification (see End-to-end echo).

For LookupResponse, the receiver is the original requester, and the annotation is reverse-path-only: the LookupResponse path is the return path of the lookup, so the annotated path_mtu reflects what the requester can use to reach the discovered destination over the discovered path.

Because the field is initialized by the source and mins as it travels, it converges to the bottleneck without any additional probing. The first SessionDatagram on a fresh session may carry an over-estimate (the source has not yet been told a smaller min), which is what makes the reactive MtuExceeded path necessary.

Reactive PMTUD: MtuExceeded

When a transit node receives a SessionDatagram whose total wire size exceeds the next-hop link_mtu, it cannot forward without fragmentation. Instead:

  1. The transit node generates a SessionDatagram addressed back to the source carrying an MtuExceeded payload (msg_type 0x22). The payload identifies the destination, the reporting router, and the bottleneck MTU.
  2. The error is routed via find_next_hop(src_addr). If the source is also unreachable, the error is dropped silently (no cascading errors).
  3. The original oversized packet is dropped.

The source's FSP layer applies the reported bottleneck immediately — unlike the increase case (see hysteresis below), decrease is always take-the-lower-value because the original packet has already been dropped. The source can then reduce payload sizes on subsequent SessionDatagrams.

MtuExceeded is the reactive complement to the proactive path_mtu field. The proactive field tracks the minimum along the forward path under steady-state convergence; MtuExceeded handles the in-flight gap when an oversized packet hits a new bottleneck (forward path shifted, peer's outbound MTU dropped, BLE renegotiated) before the source has adapted.

Error generation is rate-limited at 100ms per destination at the transit node to prevent storms during topology changes.

End-to-End Echo: PathMtuNotification

PathMtuNotification (msg_type 0x13, session-layer) provides end-to-end path MTU feedback, adapting RFC 1191 Path MTU Discovery for overlay networks — the transit-node min() propagation replaces ICMP Packet Too Big.

Mechanism:

  1. The source sets path_mtu in each SessionDatagram envelope to its outbound link MTU.
  2. Each transit node applies min(current, transport.link_mtu(addr)) before forwarding.
  3. The destination receives the forward-path minimum and sends a PathMtuNotification (2-byte body: u16 LE path_mtu) back to the source.
  4. The source applies the notification with hysteresis:
    • Decrease: immediate (take lower value).
    • Increase: requires 3 consecutive higher-value notifications spanning at least 2 × notification interval.
  5. Notifications are sent on first measurement, on any decrease, and periodically at max(10s, 5 × SRTT).

The hysteresis on increase prevents oscillation when the path MTU fluctuates around a boundary; the immediate decrease prevents delivering oversized packets after a path has narrowed.

PathMtuNotification is wrapped in a session-layer encrypted message and travels back to the source via the session's normal forwarding path. It is part of the session-layer MMP report stream's traffic budget and (along with SenderReport and ReceiverReport) does not reset the session idle timer.

Per-Destination MTU Storage

Two storage locations track per-destination MTU, serving different consumers:

  • Session-canonical (MmpSessionState.path_mtu, type PathMtuState). Holds the running end-to-end path MTU for an established FSP session. Updated by both PathMtuNotification (proactive, end-to-end echo) and reactive MtuExceeded from transit routers. Read by the session layer when constructing outbound SessionDatagram envelopes.

  • TCP-clamp mirror (path_mtu_lookup, a HashMap<FipsAddress, u16> on the Node). Read by the TUN-side TCP MSS clamp (per_flow_max_mss in src/upper/tun.rs) at first-SYN time so outbound TCP flows are clamped to the per-destination MTU rather than a generic ceiling. Written from four sites, all using tighter-only semantics — the clamp is never loosened:

    • Discovery's LookupResponse handler — reverse-path annotated value carried back by the discovery target.
    • seed_path_mtu_for_link_peer when a peer is promoted to an active link, seeding with the new link's link_mtu so traffic to that peer immediately uses the per-link value rather than a generic default.
    • The reactive MtuExceeded handler, mirroring the bottleneck reported by a transit router.
    • The proactive PathMtuNotification handler, mirroring the new effective end-to-end value so a fresh TCP flow benefits immediately from PMTU knowledge the session has already acquired.

All four writers apply the same tighter-only rule, so the mirror converges to the smallest MTU any signal has reported for that destination and a subsequent looser observation cannot widen it.

TCP MSS Clamping

The IPv6 adapter intercepts TCP SYN and SYN-ACK packets at the TUN interface and clamps the Maximum Segment Size (MSS) option to:

clamped_mss = effective_ipv6_mtu - 40 (IPv6 header) - 20 (TCP header)

Clamping is applied in two places:

  • TUN reader (outbound): clamps MSS on outbound SYN packets
  • TUN writer (inbound): clamps MSS on inbound SYN-ACK packets

Together these ensure both directions of a TCP connection use appropriately-sized segments from the start, avoiding the initial oversized-packet loss that would occur if the adapter relied on ICMP Packet Too Big alone.

Clamping is conditional: when per_flow_max_mss already has an entry for the flow, that entry is used; otherwise the clamp falls back to a ceiling derived from the most pessimistic effective IPv6 MTU the adapter knows about (1143 with the typical 1280 transport floor). The fallback handles cold-flow first-SYN traffic — the very first SYN of a flow may arrive before the MMP path-MTU echo and any per-flow lookup has been populated, so the conservative ceiling prevents the SYN-ACK chain from negotiating a too-large MSS that would later drop.

The adapter integrates with the MTU subsystem rather than owning it. The "why we clamp and what max_mss means" lives here in the MTU design; the "how the clamp is implemented at the TUN" lives in the IPv6 adapter doc.

ICMP Packet Too Big

When an outbound packet at the TUN exceeds the effective IPv6 MTU, the adapter generates an ICMPv6 Packet Too Big message and delivers it back to the application via the TUN. This triggers the kernel's Path MTU Discovery mechanism for non-TCP traffic and for any TCP flow where MSS clamping was insufficient.

ICMPv6 Packet Too Big generation is rate-limited per source address (100ms interval) to prevent storms from applications sending many oversized packets. The ICMP response is delivered locally back through the TUN; no network traversal is needed, so delivery is reliable.

No Fragmentation Policy

FIPS does not perform fragmentation at transit nodes:

  • Why no transit fragmentation. Session-layer encryption is end-to-end — the AEAD tag authenticates the entire plaintext. Fragmenting an encrypted SessionDatagram would require either exposing plaintext structure to transit nodes (unacceptable) or reassembling before decryption (opens an attack surface — a transit node could replay or withhold fragments to influence reassembly).
  • Why no source-side fragmentation. The source doesn't need fragmentation because the proactive path_mtu field plus the reactive MtuExceeded signal converge on a working size within the first few packets. Applications that need oversized payloads run TCP over the IPv6 adapter, which has its own segmentation under MSS clamping.

Some transports may perform fragmentation and reassembly internally (e.g., BLE L2CAP) and can advertise a larger virtual MTU than the physical medium supports — this is transparent to FIPS.

Operational Considerations

Diagnosing MTU-related symptoms (handshakes succeed but bulk transfers stall, ssh hangs after Welcome banner, sporadic MtuExceeded spikes during topology changes) requires inspecting per-link MTU, per-session MTU, and the per-destination path_mtu_lookup table. See ../how-to/diagnose-mtu-issues.md for the operator recipes. The relevant control-socket queries are fipsctl show sessions (per-session MTU), fipsctl show transports (per-link MTU), and fipsctl show identity-cache (with adapter MTU context).

See also