Describe what a lane is, the grouped socket layout, negotiation and port
pairing, how a lane comes up, how a flow stays on one lane and one routine
(including the kernel tun queue feedback loop that forces the lane to come
from the flow hash rather than the routine index), and every level at which
multiport falls back to the old single-port behavior.
The lane byte was described in prose wedged between the first and second
rows of the table, which both broke up the diagram and left the wire layout
still claiming a single Reserved (uint16). Make it a field: widen the table
to fit Reserved (uint8) and Lane (uint8) with their types, and move the
explanation below the whole diagram.
The prose now covers what the picture can't: lane 0 is what a lanes-unaware
sender and every non-lane packet carry, so the field is compatible in both
directions; the remaining 8 bits stay reserved and zero; the H struct still
stores the pair as one Reserved field, so callers reach the low byte through
H.Lane and EncodeLane; and lane is part of the AEAD's associated data, so it
cannot be altered in flight.
Under multiport each port had exactly one socket, so each port had exactly
one core. That hurt worst on the base port, which carries every handshake,
every lighthouse and punch packet, every peer without multiport, and every
tunnel whose lanes are down or firewalled.
Bind multiport.ports consecutive ports and give each one a full group of
`routines` sockets sharing it through SO_REUSEPORT, so the kernel hashes
each arriving 4-tuple onto one of the group. `routines` is now per port and
the total worker count is routines * multiport.ports, computed for you: with
`routines: 8` and `multiport.ports: 4` you get 8 routines per port, 32
total. A routine still owns exactly one socket, which is what lets the read
path own its state without locking.
Sockets are laid out port-major, writers[s*routinesPerPort+r] being the r'th
socket on port listen.port+s, so socket selection is a pure function of
(queue, lane) with no borrowing: laneSock(q,s) = s*R + q%R, and base traffic
takes laneSock(q,0). Sibling routines land on different sockets of the same
port, so a lane's port is served by its whole group.
multiport.ports is a new key and must be > 1; there is no default, since
under these semantics defaulting it would silently multiply the worker
count. The multiport decision moves above tun creation because tun queues
are sized by the total.
Also drop the base-port fallback for unanswered lane probes. A lane now
aims only at its own paired peer port: sharing the base port's destination
buys the lane a source port of its own but costs the peer the receive
spread lanes exist to create, so a lane that cannot reach its port stays
down and its flows ride the base tunnel.
Lanes were collapsing onto lane 0 and staying there. Nebula writes a
peer's inbound packets to the tun queue matching the socket they arrived
on, and tun_flow_update teaches the kernel to steer that flow's outbound
packets to the same queue, preferring what it learned over the hash. So
while a tunnel's lanes were still down -- which is every tunnel, for its
first moments -- all traffic arrived on socket 0, pinning every flow to
queue 0 on both hosts for as long as it stayed busy. With the lane
following the routine, that meant lane 0 forever.
Two changes break the loop:
Choose the lane from the packet's own 5-tuple hash, so lane spread no
longer depends on tun steering at all. The hash is symmetric, and the
high-addressed side rotates its choice by the low side's port offset, so
a flow's two directions land on partner lanes with exact reverse
4-tuples and each arrives through the conntrack entry the other's probe
opened.
Seed every lane as demanded at lane-set creation, so the first traffic
tick probes them all at once rather than waiting for a flow to hash onto
each. Lanes have to be up before the flows are, not after.
A routine may now send on any lane, so it holds a SendBatch per lane,
built lazily and all borrowing one shared arena -- the slab is the
expensive part, and a batch per (routine, lane) would otherwise cost
hundreds of megabytes. full() counts across the batches so the total
outstanding stays bounded as before.
That also means several routines can write to one socket, which was
already true for base traffic under multiport: the linux batchWriter's
sendmmsg scratch had no lock. Take a mutex there, once per flush.
A lane coming up or going down is per-lane, per-tunnel, and repeats on the
keepalive: a node with many tunnels would bury its actual events under it. The
gauges are the place to watch lanes in aggregate, so drop these two to debug for
when you are looking at one tunnel. Probe failures stay at error.
Every multiport log line so far was a reason it turned itself off, so a node
running it looked identical to one that never tried. Add the other half:
- an Info at startup with the lane count and port range
- the negotiated lanes on both "Handshake message received" lines, which
required allocating the lane set before the log rather than after; it still
lands before CheckAndComplete/Complete, which was the ordering requirement
- multiport.lanes.up and multiport.lanes.tunnels gauges
The lane field is a slog.Attr so it vanishes on a node without multiport, and
reads zeros when we offered lanes and the peer had none - the case an operator
is actually looking for.
The gauges walk the hostmap instead of keeping a counter at promote/demote: a
tunnel torn down while its lanes are up never demotes them, so a counter would
drift upward forever. Only a multiport node pays for the walk.
Two routines can race on a lane's first packet and derive a session each. Only
one gets installed, and the loser was dropped along with the replay-window entry
for the packet it had just accepted, leaving that one counter replayable.
The two sessions hold identical keys, so the winner's window is the same window
in every respect that matters: mark the counter seen on it instead. Closes the
gap without putting a lock on the decrypt path.
How many lanes a tunnel has is partly the peer's call: it advertises how many it
sends on and we have to be able to receive all of them, up to the header's 256.
Deriving them all when the handshake completes meant a peer advertising a large
count cost us a replay window and two cipher states per lane, per tunnel, for
lanes it may never send on.
Derive each session on the first packet that needs it instead. The TX side asks
through laneSet.session, which installs on the spot — we only ask for lanes we
chose to send on. The RX side can't do that: anyone who can spoof a tunnel's
local index can name any lane, and installing on sight would hand them the same
allocation for free. So laneSession hands back a session without publishing it
and reports that it did; outside.go installs it only once the packet has
decrypted, which is the first moment the lane is known to be real. A spoofer
gets an HKDF per packet and nothing retained.
Only the session table is sized by the peer's advert now. txAddr, demand and
probe are sized by the lanes we will actually send on, so the peer can no longer
size our per-lane tx state either.
Also refuse relayed lane packets before the session lookup rather than after,
so a junk relay packet can't reach the derivation path at all.
Lanes used to be separate HostInfos, each with its own handshake, its own
half-established states, its own lifetime and its own slot bookkeeping. That
bought nothing: a lane is the same tunnel over a different underlay 5-tuple.
Derive lane sessions instead. Noise leaves us with A.eKey == B.dKey, so both
sides HKDF-expand the same two keys with the same per-lane label and land on a
matched pair without exchanging anything. Lanes now cost no handshake, have no
half-established state, and die exactly when their base tunnel does. Which lane
a packet belongs to rides the low byte of the nebula header's Reserved field,
inside the AEAD's associated data.
Receiving on a lane needs no permission, since the session exists the moment the
base handshake completes. Sending on one needs proof the new 5-tuple works, so a
lane stays down until a probe on it is acked and falls back to the base tunnel
the moment it stops being acked. Probing is demand-driven off the connection
manager's per-tunnel traffic tick: a peer we exchange a trickle with never costs
more than its base tunnel, however many lanes are configured. The ack rides the
base session on purpose, so a broken reverse lane can't fail a working one.
Because the data now rides lane counters, the rehandshake, exhaustion and
swap-primary checks take the max counter across the base session and its lanes;
otherwise the base counter would sit near zero while a lane ran its keys past
the nonce ceiling.
Removes OutboundLaneTimer, EnsureLanes, startLaneHandshake, completeLane,
completeLaneResponder, makeLaneTrafficDecision and the lane fields on HostInfo.
Handshake payload field 3 (the per-lane handshake index) is permanently
reserved; peers advertise a TxLanes count instead.
Fields 6 and 7 are reserved for work in progress, so the multiport lane
adverts move up to 9 and 10. Both are still single-byte tags, so the
encoded payload size does not change.
Also emit the lane fields after CertVersion so MarshalPayload stays in
ascending field-number order, matching what protoc-gen-go would produce
and keeping the hand-built expectations in payload_test.go
straightforward.
Field numbers are the only thing that identifies a lane advert, so a node
on this build and one on the previous numbering will each skip the other's
adverts as unknown fields and fall back to a single vanilla tunnel.
Lanes were established for every slot as soon as the base tunnel came up,
so a host with `routines: N` paid N-1 extra Noise handshakes, sessions and
keepalives for every peer it talked to, including peers it exchanges a
trickle with. Only routines that actually carry traffic to a peer need a
lane.
Add a per-slot demand flag to laneState. sendInsideMessage already loads
txLanes[laneSlot] on every direct-path packet; that load missing is now
the signal, and EnsureLanes only starts slots the data plane asked for.
The miss path is unchanged otherwise: the packet rides the base tunnel,
the same fallback used while a lane is down.
noteLaneDemand load-guards its store so repeated misses on a lane-less
slot are plain reads rather than a cache line ping-ponging between the
routines sharing it. The demand check is EnsureLanes' last condition, so
a slot that is pending or inside its backoff keeps the flag for the tick
that can act on it. Consuming the flag also stops a lane whose routine
went quiet from being rebuilt forever after it dies.
The wire format, lane negotiation and port-offset scheme are untouched,
so a lazy node interoperates with an eager one in both directions, and
inbound lane handshakes from a busy peer are still accepted regardless of
our own demand.
Add support for the "fips140" mode of Go:
- https://go.dev/doc/security/fips140
- https://csrc.nist.gov/projects/cryptographic-module-validation-program/certificate/5247
You can build with `make fips140`, see the README changes for more info.
Some differences from the boringcrypto builds:
- We switch to using `go:linkname crypto/tls.aeadAESGCMTLS13`, which gives us the fips implementation for both `boringcrypto` and `fips140` modes. This means we also no longer need `-checklinkname=0`
- Go native `fips140` doesn't need CGO_ENABLED=1
- We decide if we should use the fips140 GCM at runtime, if `fips140.Enabled()` is true. If you use the `make release-fips140`, we build with build tag `fips140enforce` which ensures the binary is running with fips140 enabled and that only P256 / AES-GCM is being used. If you don't want this enforce mode, you can build without the build tag.
The go-metrics-graphite package has been unmaintained for 10+ years,
which is a packaging and supply-chain concern for downstream
distributors (see #1831). Nebula only used its Config struct and Once()
entrypoint, so inline just those (~65 lines) into a local graphite.go,
preserving the upstream BSD-2-Clause copyright notice, and remove the
dependency.
Fixes#1831
Co-authored-by: Claude <svc-devxp-claude@slack-corp.com>
* use stdlib maps.Values
This has been in stdlib since go1.23, no need to use golang.org/x/exp
anymore just for this:
- https://pkg.go.dev/maps#Values
* use slices.AppendSeq to match old behavior
handleOutsideRelayPacket filled ViaSender.remoteIdx with relay.RemoteIndex,
an index from the relay peer's index space, but the rescue in
sendHandshakeResponse looks that value up in relayForByIdx, which is keyed
by local index. The lookup could never hit, so a terminal relay entry left
Disestablished by a one-sided teardown stayed Disestablished even after a
valid handshake arrived over it. The responder's first transmit then failed
to find an Established relay, deleted its only relay entry, and every
subsequent send was silently dropped until dead-tunnel detection forced a
re-handshake.