Under multiport each port had exactly one socket, so each port had exactly
one core. That hurt worst on the base port, which carries every handshake,
every lighthouse and punch packet, every peer without multiport, and every
tunnel whose lanes are down or firewalled.
Bind multiport.ports consecutive ports and give each one a full group of
`routines` sockets sharing it through SO_REUSEPORT, so the kernel hashes
each arriving 4-tuple onto one of the group. `routines` is now per port and
the total worker count is routines * multiport.ports, computed for you: with
`routines: 8` and `multiport.ports: 4` you get 8 routines per port, 32
total. A routine still owns exactly one socket, which is what lets the read
path own its state without locking.
Sockets are laid out port-major, writers[s*routinesPerPort+r] being the r'th
socket on port listen.port+s, so socket selection is a pure function of
(queue, lane) with no borrowing: laneSock(q,s) = s*R + q%R, and base traffic
takes laneSock(q,0). Sibling routines land on different sockets of the same
port, so a lane's port is served by its whole group.
multiport.ports is a new key and must be > 1; there is no default, since
under these semantics defaulting it would silently multiply the worker
count. The multiport decision moves above tun creation because tun queues
are sized by the total.
Also drop the base-port fallback for unanswered lane probes. A lane now
aims only at its own paired peer port: sharing the base port's destination
buys the lane a source port of its own but costs the peer the receive
spread lanes exist to create, so a lane that cannot reach its port stays
down and its flows ride the base tunnel.
Lanes were collapsing onto lane 0 and staying there. Nebula writes a
peer's inbound packets to the tun queue matching the socket they arrived
on, and tun_flow_update teaches the kernel to steer that flow's outbound
packets to the same queue, preferring what it learned over the hash. So
while a tunnel's lanes were still down -- which is every tunnel, for its
first moments -- all traffic arrived on socket 0, pinning every flow to
queue 0 on both hosts for as long as it stayed busy. With the lane
following the routine, that meant lane 0 forever.
Two changes break the loop:
Choose the lane from the packet's own 5-tuple hash, so lane spread no
longer depends on tun steering at all. The hash is symmetric, and the
high-addressed side rotates its choice by the low side's port offset, so
a flow's two directions land on partner lanes with exact reverse
4-tuples and each arrives through the conntrack entry the other's probe
opened.
Seed every lane as demanded at lane-set creation, so the first traffic
tick probes them all at once rather than waiting for a flow to hash onto
each. Lanes have to be up before the flows are, not after.
A routine may now send on any lane, so it holds a SendBatch per lane,
built lazily and all borrowing one shared arena -- the slab is the
expensive part, and a batch per (routine, lane) would otherwise cost
hundreds of megabytes. full() counts across the batches so the total
outstanding stays bounded as before.
That also means several routines can write to one socket, which was
already true for base traffic under multiport: the linux batchWriter's
sendmmsg scratch had no lock. Take a mutex there, once per flush.
Every multiport log line so far was a reason it turned itself off, so a node
running it looked identical to one that never tried. Add the other half:
- an Info at startup with the lane count and port range
- the negotiated lanes on both "Handshake message received" lines, which
required allocating the lane set before the log rather than after; it still
lands before CheckAndComplete/Complete, which was the ordering requirement
- multiport.lanes.up and multiport.lanes.tunnels gauges
The lane field is a slog.Attr so it vanishes on a node without multiport, and
reads zeros when we offered lanes and the peer had none - the case an operator
is actually looking for.
The gauges walk the hostmap instead of keeping a counter at promote/demote: a
tunnel torn down while its lanes are up never demotes them, so a counter would
drift upward forever. Only a multiport node pays for the walk.
Lanes used to be separate HostInfos, each with its own handshake, its own
half-established states, its own lifetime and its own slot bookkeeping. That
bought nothing: a lane is the same tunnel over a different underlay 5-tuple.
Derive lane sessions instead. Noise leaves us with A.eKey == B.dKey, so both
sides HKDF-expand the same two keys with the same per-lane label and land on a
matched pair without exchanging anything. Lanes now cost no handshake, have no
half-established state, and die exactly when their base tunnel does. Which lane
a packet belongs to rides the low byte of the nebula header's Reserved field,
inside the AEAD's associated data.
Receiving on a lane needs no permission, since the session exists the moment the
base handshake completes. Sending on one needs proof the new 5-tuple works, so a
lane stays down until a probe on it is acked and falls back to the base tunnel
the moment it stops being acked. Probing is demand-driven off the connection
manager's per-tunnel traffic tick: a peer we exchange a trickle with never costs
more than its base tunnel, however many lanes are configured. The ack rides the
base session on purpose, so a broken reverse lane can't fail a working one.
Because the data now rides lane counters, the rehandshake, exhaustion and
swap-primary checks take the max counter across the base session and its lanes;
otherwise the base counter would sit near zero while a lane ran its keys past
the nonce ceiling.
Removes OutboundLaneTimer, EnsureLanes, startLaneHandshake, completeLane,
completeLaneResponder, makeLaneTrafficDecision and the lane fields on HostInfo.
Handshake payload field 3 (the per-lane handshake index) is permanently
reserved; peers advertise a TxLanes count instead.
Add support for the "fips140" mode of Go:
- https://go.dev/doc/security/fips140
- https://csrc.nist.gov/projects/cryptographic-module-validation-program/certificate/5247
You can build with `make fips140`, see the README changes for more info.
Some differences from the boringcrypto builds:
- We switch to using `go:linkname crypto/tls.aeadAESGCMTLS13`, which gives us the fips implementation for both `boringcrypto` and `fips140` modes. This means we also no longer need `-checklinkname=0`
- Go native `fips140` doesn't need CGO_ENABLED=1
- We decide if we should use the fips140 GCM at runtime, if `fips140.Enabled()` is true. If you use the `make release-fips140`, we build with build tag `fips140enforce` which ensures the binary is running with fips140 enabled and that only P256 / AES-GCM is being used. If you don't want this enforce mode, you can build without the build tag.
remove runtime.LockOSThread() because it makes things worse now
remove the "custom" Write() method from tun_linux.go, the stdlib path via os.File performs better
We should change our guidance around number of routines, ~2 per thread (that you wish to use for Nebula) seems to be about right now
* Added firewall.rules.hash metric
Added a FNV-1 hash of the firewall rules as a Prometheus value.
* Switch FNV has to int64, include both hashes in log messages
* Use a uint32 for the FNV hash
Let go-metrics cast the uint32 to a int64, so it won't be lossy
when it eventually emits a float64 Prometheus metric.
This adds a few build targets to compile with `GOEXPERIMENT=boringcrypto`:
- `bin-boringcrypto`
- `release-boringcrypto`
It also adds a field to the intial start up log indicating if
boringcrypto is enabled in the binary.
These new helpers make the code a lot cleaner. I confirmed that the
simple helpers like `atomic.Int64` don't add any extra overhead as they
get inlined by the compiler. `atomic.Pointer` adds an extra method call
as it no longer gets inlined, but we aren't using these on the hot path
so it is probably okay.
By default, Nebula replies to packets it has no tunnel for with a `recv_error` packet. This packet helps speed up re-connection
in the case that Nebula on either side did not shut down cleanly. This response can be abused as a way to discover if Nebula is running
on a host though. This option lets you configure if you want to send `recv_error` packets always, never, or only to private network remotes.
valid values: always, never, private
This setting is reloadable with SIGHUP.
* Add more metrics
This change adds the following counter metrics:
Metrics to track packets dropped at the firewall:
firewall.dropped.local_ip
firewall.dropped.remote_ip
firewall.dropped.no_rule
Metrics to track handshakes attempts that have been initiated and ones
that have timed out (ones that have completed are tracked by the
existing "handshakes" histogram).
handshake_manager.initiated
handshake_manager.timed_out
Metrics to track when cached_packets are dropped because we run out of
buffer space, and how many are sent once the handshake completes.
hostinfo.cached_packets.dropped
hostinfo.cached_packets.sent
This change also notes how many cached packets we have when we log the
final "Handshake received" message for either stage1 for stage2.
* separate incoming/outgoing metrics
* remove "allowed" firewall metrics
We don't need this on the hotpath, they aren't worh it.
* don't need pointers here