Under multiport each port had exactly one socket, so each port had exactly
one core. That hurt worst on the base port, which carries every handshake,
every lighthouse and punch packet, every peer without multiport, and every
tunnel whose lanes are down or firewalled.
Bind multiport.ports consecutive ports and give each one a full group of
`routines` sockets sharing it through SO_REUSEPORT, so the kernel hashes
each arriving 4-tuple onto one of the group. `routines` is now per port and
the total worker count is routines * multiport.ports, computed for you: with
`routines: 8` and `multiport.ports: 4` you get 8 routines per port, 32
total. A routine still owns exactly one socket, which is what lets the read
path own its state without locking.
Sockets are laid out port-major, writers[s*routinesPerPort+r] being the r'th
socket on port listen.port+s, so socket selection is a pure function of
(queue, lane) with no borrowing: laneSock(q,s) = s*R + q%R, and base traffic
takes laneSock(q,0). Sibling routines land on different sockets of the same
port, so a lane's port is served by its whole group.
multiport.ports is a new key and must be > 1; there is no default, since
under these semantics defaulting it would silently multiply the worker
count. The multiport decision moves above tun creation because tun queues
are sized by the total.
Also drop the base-port fallback for unanswered lane probes. A lane now
aims only at its own paired peer port: sharing the base port's destination
buys the lane a source port of its own but costs the peer the receive
spread lanes exist to create, so a lane that cannot reach its port stays
down and its flows ride the base tunnel.
Lanes were established for every slot as soon as the base tunnel came up,
so a host with `routines: N` paid N-1 extra Noise handshakes, sessions and
keepalives for every peer it talked to, including peers it exchanges a
trickle with. Only routines that actually carry traffic to a peer need a
lane.
Add a per-slot demand flag to laneState. sendInsideMessage already loads
txLanes[laneSlot] on every direct-path packet; that load missing is now
the signal, and EnsureLanes only starts slots the data plane asked for.
The miss path is unchanged otherwise: the packet rides the base tunnel,
the same fallback used while a lane is down.
noteLaneDemand load-guards its store so repeated misses on a lane-less
slot are plain reads rather than a cache line ping-ponging between the
routines sharing it. The demand check is EnsureLanes' last condition, so
a slot that is pending or inside its backoff keeps the flag for the tick
that can act on it. Consuming the flag also stops a lane whose routine
went quiet from being rebuilt forever after it dies.
The wire format, lane negotiation and port-offset scheme are untouched,
so a lazy node interoperates with an eager one in both directions, and
inbound lane handshakes from a busy peer are still accepted regardless of
our own demand.
* add IPv6 reject packet generation (ICMPv6 Destination Unreachable and TCP RST)
* use ICMPv6 code 1 (administratively prohibited) and cap body at 1000 bytes
* cleanup, use ICMP error code 13 for ipv4
* better docs
* cleanup
* add sshd.sandbox_dir config option
Sanitize SSH profile paths (ssh.go:514,683,719) — restrict os.Create(a[0]) to a safe directory.
Add a config option in the config file to specify the sandbox directory. For backwards compatibility, if the config is not specified, keep the current behavior.
* update default and example
* use os.TempDir() for sshd.sandbox_dir default
* split sandbox path validation into separate conditionals
Separate the combined && check in sshSanitizeFilePath into two distinct
conditionals with specific error messages: one for paths resolving to the
sandbox directory itself, and one for paths outside the sandbox.
Co-Authored-By: Claude <svc-devxp-claude@slack-corp.com>
* fix: trim leading zeros from p256 signature swap result
bigmod.Nat.Bytes() returns fixed-size 32-byte slices, but ASN.1 INTEGER
parsing strips leading zeros. This caused a flaky test failure (~1/256
chance) when the S value's high byte was zero.
Co-Authored-By: Claude <svc-devxp-claude@slack-corp.com>
---------
Co-authored-by: Claude <svc-devxp-claude@slack-corp.com>
- Accept certs signed by trusted CAs
- Username must match the cert principal if set
- Any username can be used if cert principal is empty
- Don't allow removed pubkeys/CAs to be used after reload
* add calculated_remotes
This setting allows us to "guess" what the remote might be for a host
while we wait for the lighthouse response. For networks that hard
designed with in mind, it can help speed up handshake performance, as well as
improve resiliency in the case that all lighthouses are down.
Example:
lighthouse:
# ...
calculated_remotes:
# For any Nebula IPs in 10.0.10.0/24, this will apply the mask and add
# the calculated IP as an initial remote (while we wait for the response
# from the lighthouse). Both CIDRs must have the same mask size.
# For example, Nebula IP 10.0.10.123 will have a calculated remote of
# 192.168.1.123
10.0.10.0/24:
- mask: 192.168.1.0/24
port: 4242
* figure out what is up with this test
* add test
* better logic for sending handshakes
Keep track of the last light of hosts we sent handshakes to. Only log
handshake sent messages if the list has changed.
Remove the test Test_NewHandshakeManagerTrigger because it is faulty and
makes no sense. It relys on the fact that no handshake packets actually
get sent, but with these changes we would send packets now (which it
should!)
* use atomic.Pointer
* cleanup to make it clearer
* fix typo in example
* firewall: add option to send REJECT replies
This change allows you to configure the firewall to send REJECT packets
when a packet is denied.
firewall:
# Action to take when a packet is not allowed by the firewall rules.
# Can be one of:
# `drop` (default): silently drop the packet.
# `reject`: send a reject reply.
# - For TCP, this will be a RST "Connection Reset" packet.
# - For other protocols, this will be an ICMP port unreachable packet.
outbound_action: drop
inbound_action: drop
These packets are only sent to established tunnels, and only on the
overlay network (currently IPv4 only).
$ ping -c1 192.168.100.3
PING 192.168.100.3 (192.168.100.3) 56(84) bytes of data.
From 192.168.100.3 icmp_seq=2 Destination Port Unreachable
--- 192.168.100.3 ping statistics ---
2 packets transmitted, 0 received, +1 errors, 100% packet loss, time 31ms
$ nc -nzv 192.168.100.3 22
(UNKNOWN) [192.168.100.3] 22 (?) : Connection refused
This change also modifies the smoke test to capture tcpdump pcaps from
both the inside and outside to inspect what is going on over the wire.
It also now does TCP and UDP packet tests using the Nmap version of
ncat.
* calculate seq and ack the same was as the kernel
The logic a bit confusing, so we copy it straight from how the kernel
does iptables `--reject-with tcp-reset`:
- https://github.com/torvalds/linux/blob/v5.19/net/ipv4/netfilter/nf_reject_ipv4.c#L193-L221
* cleanup