# Packet capture and correlation pipeline ## Capture modes A sniffer session targets exactly one interface or bridge. - **Interface target:** an `AF_PACKET` raw socket is opened on that interface; its effective mode is always `af_packet`. - **Bridge target with `af_packet`:** the bridge's member interfaces are captured with raw sockets. - **Bridge target with `tc_ebpf` (default):** no raw socket is opened for bridge ports. `BridgeTelemetryManager` manages an eBPF/tc helper that exports ingress raw data and egress/drop verdict-related events. - **Benchmark mode:** preserves session accounting but skips the normal userspace packet processing path, allowing capture-overhead measurements. Sessions have UUIDs and record their label, target type, mode, snapshot of bridge ports, capture interfaces, socket map, thread, and stop event. Stopping by session ID is preferred. A target-specific stop finds matching sessions; an unqualified stop stops every session. ## End-to-end lifecycle ```mermaid sequenceDiagram participant C as Capture socket or tc/eBPF participant N as network_sniffer participant T as tshark manager participant P as PacketTracker participant D as DatabasePool participant W as packet WebSocket C->>N: frame / telemetry event N->>N: parse headers, identity, observation metadata N->>T: lookup or schedule enrichment N->>P: capture observation C->>P: ingress/egress/verdict telemetry P->>P: correlate, merge and finalize record P->>D: upsert packet D->>W: publish normalized row ``` `network_sniffer.parse_packet` parses Scapy packet objects, while `parse_packet_bytes` supports raw data. It extracts link, network, and transport fields; adds capture session/observation data; calculates or obtains correlation identifiers; and hands observations to `PacketTracker`. If the shared database or event loop is not yet usable, records are retained in a bounded in-memory buffer; `drain_buffer_to_shared_db` flushes it at application startup. ## Identity and merging `packet_identity.build_packet_uid` makes a stable hash-based fallback identity from normalized packet fields. `packet_mark` decodes the shared skb-mark layout: it normalizes an observed mark, extracts a packet ID, and extracts a verdict hint. Kernel-mark identity is preferred when present; the hash fallback keeps capture and telemetry correlation possible when it is not. `PacketTracker` aggregates observations in a bounded dictionary. It deduplicates observations, keeps capture and telemetry provenance, combines ingress/egress and verdict timing, and delays finalization briefly so companion events can arrive. It writes finalized or aged dirty records in batches, retries failed persistence with bounded exponential backoff, discards pending records for stopped interfaces, and exposes a debug snapshot. Its limits and timings are all configured through `BACKEND_PACKET_TRACKER_*` settings. ## tshark enrichment `TsharkManager` starts one long-lived `tshark` process per enabled interface. It reads JSON output in a thread, derives protocol stacks, HTTP/TLS/DNS information, TCP flags, and stream context, then caches matching data for a configurable time window. The capture parser can use a heuristic immediately and the manager can backfill metadata or stream context into already stored rows. tshark is optional in the logical pipeline but enabled by default; a missing executable or worker failure is logged and does not stop capture. ## Realtime delivery `PacketBroadcaster` maintains a bounded asyncio queue per subscriber. A successful single-row database upsert serializes the row and publishes it to subscribers. Slow consumers lose queued messages when their individual queue is full rather than blocking the capture or database path. `/api/packets/ws/packets` is therefore a live-update channel, not a lossless event log; clients should retrieve history over REST and use the WebSocket for incremental updates.