documentation md files
This commit is contained in:
82
documentation/backend/capture-pipeline.md
Normal file
82
documentation/backend/capture-pipeline.md
Normal file
@@ -0,0 +1,82 @@
|
||||
# Packet capture and correlation pipeline
|
||||
|
||||
## Capture modes
|
||||
|
||||
A sniffer session targets exactly one interface or bridge.
|
||||
|
||||
- **Interface target:** an `AF_PACKET` raw socket is opened on that interface;
|
||||
its effective mode is always `af_packet`.
|
||||
- **Bridge target with `af_packet`:** the bridge's member interfaces are captured
|
||||
with raw sockets.
|
||||
- **Bridge target with `tc_ebpf` (default):** no raw socket is opened for bridge
|
||||
ports. `BridgeTelemetryManager` manages an eBPF/tc helper that exports ingress
|
||||
raw data and egress/drop verdict-related events.
|
||||
- **Benchmark mode:** preserves session accounting but skips the normal userspace
|
||||
packet processing path, allowing capture-overhead measurements.
|
||||
|
||||
Sessions have UUIDs and record their label, target type, mode, snapshot of bridge
|
||||
ports, capture interfaces, socket map, thread, and stop event. Stopping by session
|
||||
ID is preferred. A target-specific stop finds matching sessions; an unqualified
|
||||
stop stops every session.
|
||||
|
||||
## End-to-end lifecycle
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
participant C as Capture socket or tc/eBPF
|
||||
participant N as network_sniffer
|
||||
participant T as tshark manager
|
||||
participant P as PacketTracker
|
||||
participant D as DatabasePool
|
||||
participant W as packet WebSocket
|
||||
C->>N: frame / telemetry event
|
||||
N->>N: parse headers, identity, observation metadata
|
||||
N->>T: lookup or schedule enrichment
|
||||
N->>P: capture observation
|
||||
C->>P: ingress/egress/verdict telemetry
|
||||
P->>P: correlate, merge and finalize record
|
||||
P->>D: upsert packet
|
||||
D->>W: publish normalized row
|
||||
```
|
||||
|
||||
`network_sniffer.parse_packet` parses Scapy packet objects, while
|
||||
`parse_packet_bytes` supports raw data. It extracts link, network, and transport
|
||||
fields; adds capture session/observation data; calculates or obtains correlation
|
||||
identifiers; and hands observations to `PacketTracker`. If the shared database or
|
||||
event loop is not yet usable, records are retained in a bounded in-memory buffer;
|
||||
`drain_buffer_to_shared_db` flushes it at application startup.
|
||||
|
||||
## Identity and merging
|
||||
|
||||
`packet_identity.build_packet_uid` makes a stable hash-based fallback identity
|
||||
from normalized packet fields. `packet_mark` decodes the shared skb-mark layout:
|
||||
it normalizes an observed mark, extracts a packet ID, and extracts a verdict hint.
|
||||
Kernel-mark identity is preferred when present; the hash fallback keeps capture and
|
||||
telemetry correlation possible when it is not.
|
||||
|
||||
`PacketTracker` aggregates observations in a bounded dictionary. It deduplicates
|
||||
observations, keeps capture and telemetry provenance, combines ingress/egress and
|
||||
verdict timing, and delays finalization briefly so companion events can arrive.
|
||||
It writes finalized or aged dirty records in batches, retries failed persistence
|
||||
with bounded exponential backoff, discards pending records for stopped interfaces,
|
||||
and exposes a debug snapshot. Its limits and timings are all configured through
|
||||
`BACKEND_PACKET_TRACKER_*` settings.
|
||||
|
||||
## tshark enrichment
|
||||
|
||||
`TsharkManager` starts one long-lived `tshark` process per enabled interface. It
|
||||
reads JSON output in a thread, derives protocol stacks, HTTP/TLS/DNS information,
|
||||
TCP flags, and stream context, then caches matching data for a configurable time
|
||||
window. The capture parser can use a heuristic immediately and the manager can
|
||||
backfill metadata or stream context into already stored rows. tshark is optional in
|
||||
the logical pipeline but enabled by default; a missing executable or worker failure
|
||||
is logged and does not stop capture.
|
||||
|
||||
## Realtime delivery
|
||||
|
||||
`PacketBroadcaster` maintains a bounded asyncio queue per subscriber. A successful
|
||||
single-row database upsert serializes the row and publishes it to subscribers.
|
||||
Slow consumers lose queued messages when their individual queue is full rather than
|
||||
blocking the capture or database path. `/api/packets/ws/packets` is therefore a
|
||||
live-update channel, not a lossless event log; clients should retrieve history over
|
||||
REST and use the WebSocket for incremental updates.
|
||||
Reference in New Issue
Block a user