234 lines
13 KiB
Markdown
234 lines
13 KiB
Markdown
# Technical reference: packet sniffing modes
|
||
|
||
This document specifies the implemented capture behavior in `network_sniffer.py`,
|
||
`bridge_telemetry.py`, `ebpf_bridge_events.py`, and `packet_tracker.py`. It makes a
|
||
deliberate distinction between observed facts and inferred forwarding results.
|
||
|
||
## Session model and mode selection
|
||
|
||
A capture session has a UUID and targets exactly one interface or exactly one
|
||
bridge. A bridge is expanded once with `get_bridge_ports_once`; its member list is
|
||
a creation-time snapshot. Later bridge membership changes are not added to the
|
||
existing session. Session state includes target label/type, effective mode, benchmark
|
||
flag, port snapshot, AF_PACKET sockets, optional reader thread, stop event, and
|
||
benchmark counters.
|
||
|
||
| Request | Effective mode | Capture source | Path/outcome evidence |
|
||
| --- | --- | --- | --- |
|
||
| Interface, any requested mode | `af_packet` | One raw socket on the interface | Packet-socket type labels an outgoing copy as egress; no kernel verdict telemetry. |
|
||
| Bridge, `af_packet` | `af_packet` | One raw socket per snapshot bridge port | Two matching port observations can infer forwarding. |
|
||
| Bridge, `tc_ebpf` (default) | `tc_ebpf` | One tc/eBPF helper across snapshot bridge ports | TC ingress/egress and skb-free/drop events, normally matched by skb mark. |
|
||
| Any mode with benchmark enabled | Same hook/socket setup | Counted but not processed | No parsing, enrichment, DB write, or WebSocket event. |
|
||
|
||
Interface targets always use AF_PACKET. The TC/eBPF mode is only selected for a
|
||
bridge target. A stop by session ID is the safest selector. Stopping a session also
|
||
discards pending tracker entries relating to its interfaces, so unpersisted data can
|
||
be lost deliberately at shutdown.
|
||
|
||
Multiple sessions may overlap on an interface. This is not an independent-capture
|
||
guarantee: the TC manager maps an interface to several sessions but assigns raw
|
||
ingress parsing to the first sorted session ID.
|
||
|
||
## AF_PACKET capture
|
||
|
||
### Socket behavior
|
||
|
||
For each capture interface the service opens `AF_PACKET` / `SOCK_RAW` with protocol
|
||
`htons(0x0003)` (`ETH_P_ALL`), requests the configured receive buffer (default
|
||
4 MiB), best-effort requests `TPACKET_V3`, binds to `(ifname, 0)`, and makes the
|
||
socket non-blocking. Failure to set the buffer or TPACKET version is non-fatal.
|
||
Failure to create or bind leaves the interface uncaptured; session creation can still
|
||
complete. This requires raw-socket privilege, commonly `CAP_NET_RAW`.
|
||
|
||
One daemon reader thread is started only when the session has sockets. It uses a
|
||
selector, receives at most `BACKEND_SNIFFER_RECV_BYTES` bytes per event (default
|
||
65,536), stamps the frame with userspace UTC receive time, creates Scapy `Ether`,
|
||
and calls the common parser. `ENODEV`, `ENETDOWN`, and `EBADF` close the affected
|
||
socket; it is not reopened in that session. The thread periodically attempts a
|
||
pre-DB-buffer drain during selector idle time.
|
||
|
||
### Direction and bridge inference
|
||
|
||
Packet-socket address metadata is used only as follows: `PACKET_OUTGOING` (normally
|
||
4) becomes `path_role: egress`; every other packet type becomes `path_role:
|
||
ingress`. This is a packet-socket perspective, not proof of a Linux bridge decision.
|
||
|
||
For a bridge session, the tracker groups AF_PACKET observations by session ID. It
|
||
uses an explicit ingress observation if available, otherwise the earliest one. It
|
||
prefers an explicit egress observation on a different port, otherwise a later
|
||
different-port observation. If one correlated packet is seen on at least two ports,
|
||
the tracker records:
|
||
|
||
```text
|
||
verdict = accept
|
||
verdict_reason = bridge-af_packet-forwarded-observed
|
||
verdict_confidence = medium
|
||
```
|
||
|
||
This means matching evidence was observed on two bridge ports. It does not prove a
|
||
particular kernel forwarding verdict and can be affected by duplicate copies, loops,
|
||
or fallback-identity collisions. A single-interface AF_PACKET record has no terminal
|
||
verdict from AF_PACKET itself.
|
||
|
||
### AF_PACKET implications
|
||
|
||
AF_PACKET provides full observed frame bytes without BCC or tc changes and is the
|
||
only interface-capture mode. It neither alters packets nor controls forwarding. It
|
||
also has no definitive drop visibility, reports userspace rather than kernel event
|
||
time, can observe local/outgoing copies, and can lose traffic under socket/userspace
|
||
load. The full-frame Scapy parse, tshark lookup, tracking and persistence path makes
|
||
it more expensive than sampled telemetry.
|
||
|
||
## TC/eBPF bridge capture
|
||
|
||
### Collector lifecycle and destructive qdisc behavior
|
||
|
||
Bridge sessions in `tc_ebpf` mode are aggregated into one helper process. Any change
|
||
to the active *interface set* stops the helper and recreates it for the new set;
|
||
there is a capture gap during that restart. The helper attaches direct-action
|
||
`BPF.SCHED_CLS` programs at TC ingress (`ffff:fff2`, handle `:20`) and egress
|
||
(`ffff:fff3`, handle `:30`) to every bridge **member interface**, not the bridge
|
||
device itself.
|
||
|
||
Before attachment the helper runs `tc qdisc del dev <iface> clsact` (ignoring its
|
||
result), then `tc qdisc add dev <iface> clsact`. It deletes `clsact` again for every
|
||
instrumented interface at helper shutdown and after an attachment failure.
|
||
|
||
> Starting, restarting, failing, or stopping TC/eBPF capture can remove pre-existing
|
||
> clsact qdiscs and their filters. Do not use it on interfaces with unrelated TC
|
||
> configuration unless coexistence and recovery are explicitly managed.
|
||
|
||
The BCC Python runtime, a compatible kernel, BPF/tracepoint access, TC and netlink
|
||
privileges are required. Session creation does not wait for a collector health
|
||
acknowledgement, so a successful start response is not proof that BPF attached.
|
||
|
||
### Kernel event generation
|
||
|
||
The helper opens a raw-ingress perf buffer and a metadata perf buffer. The ingress
|
||
TC program creates an skb mark only when it is zero, using the low 28 bits of
|
||
`bpf_ktime_get_ns()` and replacing zero with one. It preserves any existing nonzero
|
||
mark. It extracts Ethernet addresses, EtherType, a single 802.1Q/802.1AD VLAN ID,
|
||
ARP IPv4 addresses, and IPv4/IPv6 addresses with TCP/UDP ports. IPv6 extension
|
||
headers are not traversed; the base next-header is used as protocol.
|
||
|
||
The egress TC program never creates a mark. It exports metadata only for marked
|
||
packets. The `skb:kfree_skb` tracepoint reads the linear skb representation and
|
||
exports a drop event only for marked skbs whose device is a selected interface.
|
||
The emitted payload contains userspace and kernel-monotonic timestamps, interface,
|
||
mark, length, parsed L2–L4 fields, and event type. Drop events add a numerical
|
||
reason and `skb_drop_reason_<n>` label. Ingress events may contain `raw_b64`; egress
|
||
and drop events do not.
|
||
|
||
A kfree_skb event is evidence that a marked skb was freed in the kernel context. It
|
||
is not automatically evidence that nftables caused the outcome; interpret the
|
||
reason code in the context of kernel behavior and other instrumentation.
|
||
|
||
### Sampling
|
||
|
||
Raw and metadata sampling are independent settings.
|
||
|
||
| Value | Effect |
|
||
| --- | --- |
|
||
| `0` | Never emits that sample category. |
|
||
| `1` | Emits every marked packet in that category. |
|
||
| `N > 1` | Emits when `skb_mark % N == 0`. |
|
||
|
||
At ingress, a raw-selected packet emits a raw event; only a packet not chosen for
|
||
raw can emit an ingress metadata event. Egress and drop use metadata sampling only.
|
||
Therefore raw-enabled/meta-disabled capture stores sampled ingress frame records
|
||
without egress/drop visibility; raw-disabled/meta-enabled capture produces
|
||
metadata-only rows without raw bytes. Both enabled does not make raw and metadata
|
||
populations identical.
|
||
|
||
Sampling uses the entire existing skb mark. The documented mark layout reserves
|
||
bits 0–27 for packet ID and upper bits for drop/reject hints. This capture program
|
||
creates only the low-28-bit value for previously zero marks; it does not set verdict
|
||
hints. Any other mark-using subsystem must coordinate its mark semantics, because
|
||
it can change both sampling and correlation.
|
||
|
||
### Userspace event handling and loss
|
||
|
||
The manager reads JSON helper output into a bounded queue. For a non-benchmark
|
||
ingress event with `raw_b64`, it decodes the frame and sends it into the common Scapy
|
||
parser as source `tc_ingress_raw`, with `packet_id`, `skb_mark`, and capture mode
|
||
`tc_ingress`. It then sends every non-benchmark ingress/egress/drop event to the
|
||
tracker. A sampled raw ingress packet usually therefore has both a parsed capture
|
||
observation and a telemetry observation under the same mark-derived key.
|
||
|
||
When the telemetry queue is full, the manager drops oldest queued events down to
|
||
`BACKEND_BRIDGE_TELEMETRY_QUEUE_RECOVERY_SIZE`, attempts to keep the new event, and
|
||
counts dropped events and raw payloads. Perf buffers can also lose samples before
|
||
userspace. Neither loss mechanism is recovered. `/api/sniffer/debug` reports queue
|
||
size, queue drops, benchmark counts, collector interfaces, and tracker statistics.
|
||
|
||
### TC/eBPF implications
|
||
|
||
This mode yields better within-host correlation and explicit TC egress evidence. A
|
||
matching egress produces `accept`, `egress-observed`, confidence `high`; a matching
|
||
drop produces `drop` (or mark hint), confidence `high`. Absence of egress is not
|
||
proof of a drop: sampling, perf loss, queue loss, an uninstrumented path, teardown,
|
||
or collector failure can all explain it. Raw bytes are ingress-only and sampled.
|
||
|
||
## Common parsing, enrichment, and identity
|
||
|
||
Both modes use `parse_packet` / `parse_packet_bytes`. The parser records Ethernet
|
||
addresses, EtherType and VLAN, ARP operation/addressing, IPv4 ID or IPv6 base
|
||
header, TCP sequence/acknowledgement/flags, UDP ports, ICMP/ICMPv6 type/code, and
|
||
an embedded IPv4 tuple from eligible ICMP errors. It stores full raw frame bytes
|
||
when supplied by AF_PACKET or sampled TC ingress.
|
||
|
||
tshark workers are enabled for non-benchmark capture interfaces. They may add
|
||
application protocol, category, confidence, hostname, encryption/risk, and flow/DPI
|
||
metadata. They are optional and asynchronous; a failure or late match does not
|
||
discard underlying capture, and later backfill can enrich stored records.
|
||
|
||
The preferred identity is `pid:<packet-id>`, where the ID is bits 0–27 of skb mark.
|
||
Without it, a SHA-1 `uid` is calculated. The Scapy fallback includes L2–L4 fields,
|
||
IPv4 ID, ARP/ICMP fields and TCP sequence/ack/flags; eBPF metadata's fallback uses
|
||
only the smaller L2–L4 tuple and length. Hash-only correlation is consequently a
|
||
best-effort fallback, weaker for repeated/identical/fragmented traffic.
|
||
|
||
## Tracker outcomes and persistence
|
||
|
||
The tracker deduplicates observations, merges available fields, retains the earliest
|
||
timestamp, and waits the configured finalization delay (default 250 ms).
|
||
|
||
| Evidence | Verdict | Confidence |
|
||
| --- | --- | --- |
|
||
| TC egress telemetry | `accept`; `egress-observed` | high |
|
||
| TC drop telemetry | mark hint or `drop`; kernel reason / `kfree_skb` | high |
|
||
| Matching recent TCP RST or ICMP unreachable after drop | `reject` | medium |
|
||
| Same AF_PACKET bridge record on two ports | `accept`; forwarding observed | medium |
|
||
| No terminal evidence before delay expires | `unknown`; `timeout` | low |
|
||
|
||
It asynchronously upserts batches to PostgreSQL. Entry-cap pressure, persistence
|
||
failure/retry limits, dirty-age expiry, collector queue loss, socket loss, and
|
||
shutdown can all cause incompleteness. A packet history or WebSocket feed is never
|
||
a proof of lossless capture. Batch upserts also do not individually publish packet
|
||
updates, so realtime consumers must use history reconciliation.
|
||
|
||
## Benchmark mode
|
||
|
||
Benchmark mode still creates sockets or TC hooks but bypasses normal processing.
|
||
AF_PACKET increments received frame and byte counters. TC/eBPF increments helper
|
||
event counters and raw-payload-event counters. It does not parse Scapy, invoke
|
||
tshark, call the tracker, persist rows, or publish updates. AF_PACKET counters count
|
||
socket frames; TC counters count emitted sampled events. They are not comparable as
|
||
equal packet totals without accounting for sampling and multiple event types.
|
||
|
||
## Selection guidance
|
||
|
||
| Need | Mode | Important caveat |
|
||
| --- | --- | --- |
|
||
| Full raw visibility for a single interface | AF_PACKET | No definitive kernel egress/drop verdict. |
|
||
| Full raw frames across bridge ports | Bridge AF_PACKET | High userspace work; bridge forwarding is inferred. |
|
||
| Ingress/egress/drop evidence on a controlled bridge | TC/eBPF | Requires BPF/TC privileges and resets clsact. |
|
||
| Reduced overhead / sampled observability | TC/eBPF sampling | Data is intentionally incomplete. |
|
||
| Hook-overhead measurement | Benchmark mode | Counts differ between AF_PACKET and TC. |
|
||
|
||
Before TC/eBPF capture, inspect `tc qdisc` and filters for every target port,
|
||
coordinate skb-mark ownership, verify BCC/kernel support, and plan recovery of the
|
||
TC configuration. For every mode, monitor sniffer debug counters, system logs,
|
||
capture/process health, DB persistence failures, and expected traffic rate before
|
||
making operational or security conclusions.
|