documentation md files
This commit is contained in:
233
documentation/backend/sniffing.md
Normal file
233
documentation/backend/sniffing.md
Normal file
@@ -0,0 +1,233 @@
|
||||
# Technical reference: packet sniffing modes
|
||||
|
||||
This document specifies the implemented capture behavior in `network_sniffer.py`,
|
||||
`bridge_telemetry.py`, `ebpf_bridge_events.py`, and `packet_tracker.py`. It makes a
|
||||
deliberate distinction between observed facts and inferred forwarding results.
|
||||
|
||||
## Session model and mode selection
|
||||
|
||||
A capture session has a UUID and targets exactly one interface or exactly one
|
||||
bridge. A bridge is expanded once with `get_bridge_ports_once`; its member list is
|
||||
a creation-time snapshot. Later bridge membership changes are not added to the
|
||||
existing session. Session state includes target label/type, effective mode, benchmark
|
||||
flag, port snapshot, AF_PACKET sockets, optional reader thread, stop event, and
|
||||
benchmark counters.
|
||||
|
||||
| Request | Effective mode | Capture source | Path/outcome evidence |
|
||||
| --- | --- | --- | --- |
|
||||
| Interface, any requested mode | `af_packet` | One raw socket on the interface | Packet-socket type labels an outgoing copy as egress; no kernel verdict telemetry. |
|
||||
| Bridge, `af_packet` | `af_packet` | One raw socket per snapshot bridge port | Two matching port observations can infer forwarding. |
|
||||
| Bridge, `tc_ebpf` (default) | `tc_ebpf` | One tc/eBPF helper across snapshot bridge ports | TC ingress/egress and skb-free/drop events, normally matched by skb mark. |
|
||||
| Any mode with benchmark enabled | Same hook/socket setup | Counted but not processed | No parsing, enrichment, DB write, or WebSocket event. |
|
||||
|
||||
Interface targets always use AF_PACKET. The TC/eBPF mode is only selected for a
|
||||
bridge target. A stop by session ID is the safest selector. Stopping a session also
|
||||
discards pending tracker entries relating to its interfaces, so unpersisted data can
|
||||
be lost deliberately at shutdown.
|
||||
|
||||
Multiple sessions may overlap on an interface. This is not an independent-capture
|
||||
guarantee: the TC manager maps an interface to several sessions but assigns raw
|
||||
ingress parsing to the first sorted session ID.
|
||||
|
||||
## AF_PACKET capture
|
||||
|
||||
### Socket behavior
|
||||
|
||||
For each capture interface the service opens `AF_PACKET` / `SOCK_RAW` with protocol
|
||||
`htons(0x0003)` (`ETH_P_ALL`), requests the configured receive buffer (default
|
||||
4 MiB), best-effort requests `TPACKET_V3`, binds to `(ifname, 0)`, and makes the
|
||||
socket non-blocking. Failure to set the buffer or TPACKET version is non-fatal.
|
||||
Failure to create or bind leaves the interface uncaptured; session creation can still
|
||||
complete. This requires raw-socket privilege, commonly `CAP_NET_RAW`.
|
||||
|
||||
One daemon reader thread is started only when the session has sockets. It uses a
|
||||
selector, receives at most `BACKEND_SNIFFER_RECV_BYTES` bytes per event (default
|
||||
65,536), stamps the frame with userspace UTC receive time, creates Scapy `Ether`,
|
||||
and calls the common parser. `ENODEV`, `ENETDOWN`, and `EBADF` close the affected
|
||||
socket; it is not reopened in that session. The thread periodically attempts a
|
||||
pre-DB-buffer drain during selector idle time.
|
||||
|
||||
### Direction and bridge inference
|
||||
|
||||
Packet-socket address metadata is used only as follows: `PACKET_OUTGOING` (normally
|
||||
4) becomes `path_role: egress`; every other packet type becomes `path_role:
|
||||
ingress`. This is a packet-socket perspective, not proof of a Linux bridge decision.
|
||||
|
||||
For a bridge session, the tracker groups AF_PACKET observations by session ID. It
|
||||
uses an explicit ingress observation if available, otherwise the earliest one. It
|
||||
prefers an explicit egress observation on a different port, otherwise a later
|
||||
different-port observation. If one correlated packet is seen on at least two ports,
|
||||
the tracker records:
|
||||
|
||||
```text
|
||||
verdict = accept
|
||||
verdict_reason = bridge-af_packet-forwarded-observed
|
||||
verdict_confidence = medium
|
||||
```
|
||||
|
||||
This means matching evidence was observed on two bridge ports. It does not prove a
|
||||
particular kernel forwarding verdict and can be affected by duplicate copies, loops,
|
||||
or fallback-identity collisions. A single-interface AF_PACKET record has no terminal
|
||||
verdict from AF_PACKET itself.
|
||||
|
||||
### AF_PACKET implications
|
||||
|
||||
AF_PACKET provides full observed frame bytes without BCC or tc changes and is the
|
||||
only interface-capture mode. It neither alters packets nor controls forwarding. It
|
||||
also has no definitive drop visibility, reports userspace rather than kernel event
|
||||
time, can observe local/outgoing copies, and can lose traffic under socket/userspace
|
||||
load. The full-frame Scapy parse, tshark lookup, tracking and persistence path makes
|
||||
it more expensive than sampled telemetry.
|
||||
|
||||
## TC/eBPF bridge capture
|
||||
|
||||
### Collector lifecycle and destructive qdisc behavior
|
||||
|
||||
Bridge sessions in `tc_ebpf` mode are aggregated into one helper process. Any change
|
||||
to the active *interface set* stops the helper and recreates it for the new set;
|
||||
there is a capture gap during that restart. The helper attaches direct-action
|
||||
`BPF.SCHED_CLS` programs at TC ingress (`ffff:fff2`, handle `:20`) and egress
|
||||
(`ffff:fff3`, handle `:30`) to every bridge **member interface**, not the bridge
|
||||
device itself.
|
||||
|
||||
Before attachment the helper runs `tc qdisc del dev <iface> clsact` (ignoring its
|
||||
result), then `tc qdisc add dev <iface> clsact`. It deletes `clsact` again for every
|
||||
instrumented interface at helper shutdown and after an attachment failure.
|
||||
|
||||
> Starting, restarting, failing, or stopping TC/eBPF capture can remove pre-existing
|
||||
> clsact qdiscs and their filters. Do not use it on interfaces with unrelated TC
|
||||
> configuration unless coexistence and recovery are explicitly managed.
|
||||
|
||||
The BCC Python runtime, a compatible kernel, BPF/tracepoint access, TC and netlink
|
||||
privileges are required. Session creation does not wait for a collector health
|
||||
acknowledgement, so a successful start response is not proof that BPF attached.
|
||||
|
||||
### Kernel event generation
|
||||
|
||||
The helper opens a raw-ingress perf buffer and a metadata perf buffer. The ingress
|
||||
TC program creates an skb mark only when it is zero, using the low 28 bits of
|
||||
`bpf_ktime_get_ns()` and replacing zero with one. It preserves any existing nonzero
|
||||
mark. It extracts Ethernet addresses, EtherType, a single 802.1Q/802.1AD VLAN ID,
|
||||
ARP IPv4 addresses, and IPv4/IPv6 addresses with TCP/UDP ports. IPv6 extension
|
||||
headers are not traversed; the base next-header is used as protocol.
|
||||
|
||||
The egress TC program never creates a mark. It exports metadata only for marked
|
||||
packets. The `skb:kfree_skb` tracepoint reads the linear skb representation and
|
||||
exports a drop event only for marked skbs whose device is a selected interface.
|
||||
The emitted payload contains userspace and kernel-monotonic timestamps, interface,
|
||||
mark, length, parsed L2–L4 fields, and event type. Drop events add a numerical
|
||||
reason and `skb_drop_reason_<n>` label. Ingress events may contain `raw_b64`; egress
|
||||
and drop events do not.
|
||||
|
||||
A kfree_skb event is evidence that a marked skb was freed in the kernel context. It
|
||||
is not automatically evidence that nftables caused the outcome; interpret the
|
||||
reason code in the context of kernel behavior and other instrumentation.
|
||||
|
||||
### Sampling
|
||||
|
||||
Raw and metadata sampling are independent settings.
|
||||
|
||||
| Value | Effect |
|
||||
| --- | --- |
|
||||
| `0` | Never emits that sample category. |
|
||||
| `1` | Emits every marked packet in that category. |
|
||||
| `N > 1` | Emits when `skb_mark % N == 0`. |
|
||||
|
||||
At ingress, a raw-selected packet emits a raw event; only a packet not chosen for
|
||||
raw can emit an ingress metadata event. Egress and drop use metadata sampling only.
|
||||
Therefore raw-enabled/meta-disabled capture stores sampled ingress frame records
|
||||
without egress/drop visibility; raw-disabled/meta-enabled capture produces
|
||||
metadata-only rows without raw bytes. Both enabled does not make raw and metadata
|
||||
populations identical.
|
||||
|
||||
Sampling uses the entire existing skb mark. The documented mark layout reserves
|
||||
bits 0–27 for packet ID and upper bits for drop/reject hints. This capture program
|
||||
creates only the low-28-bit value for previously zero marks; it does not set verdict
|
||||
hints. Any other mark-using subsystem must coordinate its mark semantics, because
|
||||
it can change both sampling and correlation.
|
||||
|
||||
### Userspace event handling and loss
|
||||
|
||||
The manager reads JSON helper output into a bounded queue. For a non-benchmark
|
||||
ingress event with `raw_b64`, it decodes the frame and sends it into the common Scapy
|
||||
parser as source `tc_ingress_raw`, with `packet_id`, `skb_mark`, and capture mode
|
||||
`tc_ingress`. It then sends every non-benchmark ingress/egress/drop event to the
|
||||
tracker. A sampled raw ingress packet usually therefore has both a parsed capture
|
||||
observation and a telemetry observation under the same mark-derived key.
|
||||
|
||||
When the telemetry queue is full, the manager drops oldest queued events down to
|
||||
`BACKEND_BRIDGE_TELEMETRY_QUEUE_RECOVERY_SIZE`, attempts to keep the new event, and
|
||||
counts dropped events and raw payloads. Perf buffers can also lose samples before
|
||||
userspace. Neither loss mechanism is recovered. `/api/sniffer/debug` reports queue
|
||||
size, queue drops, benchmark counts, collector interfaces, and tracker statistics.
|
||||
|
||||
### TC/eBPF implications
|
||||
|
||||
This mode yields better within-host correlation and explicit TC egress evidence. A
|
||||
matching egress produces `accept`, `egress-observed`, confidence `high`; a matching
|
||||
drop produces `drop` (or mark hint), confidence `high`. Absence of egress is not
|
||||
proof of a drop: sampling, perf loss, queue loss, an uninstrumented path, teardown,
|
||||
or collector failure can all explain it. Raw bytes are ingress-only and sampled.
|
||||
|
||||
## Common parsing, enrichment, and identity
|
||||
|
||||
Both modes use `parse_packet` / `parse_packet_bytes`. The parser records Ethernet
|
||||
addresses, EtherType and VLAN, ARP operation/addressing, IPv4 ID or IPv6 base
|
||||
header, TCP sequence/acknowledgement/flags, UDP ports, ICMP/ICMPv6 type/code, and
|
||||
an embedded IPv4 tuple from eligible ICMP errors. It stores full raw frame bytes
|
||||
when supplied by AF_PACKET or sampled TC ingress.
|
||||
|
||||
tshark workers are enabled for non-benchmark capture interfaces. They may add
|
||||
application protocol, category, confidence, hostname, encryption/risk, and flow/DPI
|
||||
metadata. They are optional and asynchronous; a failure or late match does not
|
||||
discard underlying capture, and later backfill can enrich stored records.
|
||||
|
||||
The preferred identity is `pid:<packet-id>`, where the ID is bits 0–27 of skb mark.
|
||||
Without it, a SHA-1 `uid` is calculated. The Scapy fallback includes L2–L4 fields,
|
||||
IPv4 ID, ARP/ICMP fields and TCP sequence/ack/flags; eBPF metadata's fallback uses
|
||||
only the smaller L2–L4 tuple and length. Hash-only correlation is consequently a
|
||||
best-effort fallback, weaker for repeated/identical/fragmented traffic.
|
||||
|
||||
## Tracker outcomes and persistence
|
||||
|
||||
The tracker deduplicates observations, merges available fields, retains the earliest
|
||||
timestamp, and waits the configured finalization delay (default 250 ms).
|
||||
|
||||
| Evidence | Verdict | Confidence |
|
||||
| --- | --- | --- |
|
||||
| TC egress telemetry | `accept`; `egress-observed` | high |
|
||||
| TC drop telemetry | mark hint or `drop`; kernel reason / `kfree_skb` | high |
|
||||
| Matching recent TCP RST or ICMP unreachable after drop | `reject` | medium |
|
||||
| Same AF_PACKET bridge record on two ports | `accept`; forwarding observed | medium |
|
||||
| No terminal evidence before delay expires | `unknown`; `timeout` | low |
|
||||
|
||||
It asynchronously upserts batches to PostgreSQL. Entry-cap pressure, persistence
|
||||
failure/retry limits, dirty-age expiry, collector queue loss, socket loss, and
|
||||
shutdown can all cause incompleteness. A packet history or WebSocket feed is never
|
||||
a proof of lossless capture. Batch upserts also do not individually publish packet
|
||||
updates, so realtime consumers must use history reconciliation.
|
||||
|
||||
## Benchmark mode
|
||||
|
||||
Benchmark mode still creates sockets or TC hooks but bypasses normal processing.
|
||||
AF_PACKET increments received frame and byte counters. TC/eBPF increments helper
|
||||
event counters and raw-payload-event counters. It does not parse Scapy, invoke
|
||||
tshark, call the tracker, persist rows, or publish updates. AF_PACKET counters count
|
||||
socket frames; TC counters count emitted sampled events. They are not comparable as
|
||||
equal packet totals without accounting for sampling and multiple event types.
|
||||
|
||||
## Selection guidance
|
||||
|
||||
| Need | Mode | Important caveat |
|
||||
| --- | --- | --- |
|
||||
| Full raw visibility for a single interface | AF_PACKET | No definitive kernel egress/drop verdict. |
|
||||
| Full raw frames across bridge ports | Bridge AF_PACKET | High userspace work; bridge forwarding is inferred. |
|
||||
| Ingress/egress/drop evidence on a controlled bridge | TC/eBPF | Requires BPF/TC privileges and resets clsact. |
|
||||
| Reduced overhead / sampled observability | TC/eBPF sampling | Data is intentionally incomplete. |
|
||||
| Hook-overhead measurement | Benchmark mode | Counts differ between AF_PACKET and TC. |
|
||||
|
||||
Before TC/eBPF capture, inspect `tc qdisc` and filters for every target port,
|
||||
coordinate skb-mark ownership, verify BCC/kernel support, and plan recovery of the
|
||||
TC configuration. For every mode, monitor sniffer debug counters, system logs,
|
||||
capture/process health, DB persistence failures, and expected traffic rate before
|
||||
making operational or security conclusions.
|
||||
Reference in New Issue
Block a user