multiport: make a lane a session on the tunnel, not a tunnel

Lanes used to be separate HostInfos, each with its own handshake, its own
half-established states, its own lifetime and its own slot bookkeeping. That
bought nothing: a lane is the same tunnel over a different underlay 5-tuple.

Derive lane sessions instead. Noise leaves us with A.eKey == B.dKey, so both
sides HKDF-expand the same two keys with the same per-lane label and land on a
matched pair without exchanging anything. Lanes now cost no handshake, have no
half-established state, and die exactly when their base tunnel does. Which lane
a packet belongs to rides the low byte of the nebula header's Reserved field,
inside the AEAD's associated data.

Receiving on a lane needs no permission, since the session exists the moment the
base handshake completes. Sending on one needs proof the new 5-tuple works, so a
lane stays down until a probe on it is acked and falls back to the base tunnel
the moment it stops being acked. Probing is demand-driven off the connection
manager's per-tunnel traffic tick: a peer we exchange a trickle with never costs
more than its base tunnel, however many lanes are configured. The ack rides the
base session on purpose, so a broken reverse lane can't fail a working one.

Because the data now rides lane counters, the rehandshake, exhaustion and
swap-primary checks take the max counter across the base session and its lanes;
otherwise the base counter would sit near zero while a lane ran its keys past
the nonce ceiling.

Removes OutboundLaneTimer, EnsureLanes, startLaneHandshake, completeLane,
completeLaneResponder, makeLaneTrafficDecision and the lane fields on HostInfo.
Handshake payload field 3 (the per-lane handshake index) is permanently
reserved; peers advertise a TxLanes count instead.
This commit is contained in:
Wade Simmons
2026-09-02 14:48:17 -04:00
parent 8b914e67f4
commit fb20de39b2
19 changed files with 1223 additions and 1273 deletions
+12 -9
View File
@@ -452,10 +452,12 @@ func (f *Interface) pinThisThread(i int) {
// txQueue is the per-routine TX state owned by one listenIn goroutine.
// laneSlot is the lane this routine's traffic rides (laneSlotFor); lane is
// bound to that slot's socket and carries lane-tunnel data; base is bound to
// socket 0 and carries base-tunnel and relay data, which must keep the base
// source port (a vanilla peer would otherwise see per-routine source ports
// and roam-thrash). The two alias when multiport is off or laneSlot is 0.
// bound to that slot's socket and carries traffic encrypted with that lane's
// session; base is bound to socket 0 and carries base-session and relay data,
// which must keep the base source port (a vanilla peer would otherwise see
// per-routine source ports and roam-thrash). A routine uses base whenever its
// lane is down, so both batches stay live for the life of the routine. The two
// alias when multiport is off or laneSlot is 0.
// Concurrent sendmmsg on a shared fd (socket 0, or a lane socket shared by
// overflow routines) is safe: a flow is pinned to one routine by tun
// steering, so per-flow wire order still holds.
@@ -473,8 +475,8 @@ func (tx *txQueue) full() bool {
}
// flush drains base before lane so that when a flow moves from the base
// tunnel onto a freshly established lane mid-window, its packets still leave
// this host in encryption order.
// session onto a freshly promoted lane mid-window, its packets still leave this
// host in encryption order.
func (tx *txQueue) flush(f *Interface) {
if tx.base != tx.lane {
f.flushSendBatch(tx.base, 0)
@@ -485,9 +487,10 @@ func (tx *txQueue) flush(f *Interface) {
// laneSlotFor maps a routine index to the lane its traffic rides. When
// multiport.lanes is below routines, overflow routines share the configured
// lanes round-robin instead of all falling back onto the base tunnel's
// single underlay flow. Sharers use the lane's own socket and 4-tuple, so
// this is vanilla-style same-flow sharing: no cross-path replay skew, and
// per-flow ordering still holds (a flow stays pinned to one routine).
// single underlay flow. Sharers use the lane's own socket, 4-tuple and
// session, so this is vanilla-style same-flow sharing: no cross-path replay
// skew, and per-flow ordering still holds (a flow stays pinned to one
// routine).
func (f *Interface) laneSlotFor(i int) int {
if f.multiport && f.laneCount > 0 {
return i % f.laneCount