documentation md files
This commit is contained in:
39
documentation/backend/README.md
Normal file
39
documentation/backend/README.md
Normal file
@@ -0,0 +1,39 @@
|
|||||||
|
# Backend documentation
|
||||||
|
|
||||||
|
This directory documents the Python service in `backend/src`. It is written for
|
||||||
|
developers and operators of the inline MITM test system. The source code remains
|
||||||
|
the implementation authority; this documentation records the externally useful
|
||||||
|
contracts, lifecycle, Linux integration, and data semantics that are easy to lose
|
||||||
|
when reading individual modules.
|
||||||
|
|
||||||
|
## Reading order
|
||||||
|
|
||||||
|
1. [Architecture](architecture.md) explains the process, responsibilities, and
|
||||||
|
lifecycle.
|
||||||
|
2. [Sniffing modes](sniffing.md) gives the complete technical behavior and
|
||||||
|
implications of AF_PACKET and TC/eBPF capture.
|
||||||
|
3. [Capture pipeline](capture-pipeline.md) follows a packet from observation to
|
||||||
|
persistence and realtime delivery.
|
||||||
|
4. [HTTP and WebSocket API](api.md) lists every router mounted by the application.
|
||||||
|
5. [Data and analysis](data-and-analysis.md) describes the packet record, database
|
||||||
|
operations, and derived analysis views.
|
||||||
|
6. [Host integration](host-integration.md) covers network, eBPF, nftables, tshark,
|
||||||
|
and systemd side effects.
|
||||||
|
7. [Configuration and deployment](configuration.md) records dependencies and all
|
||||||
|
`BACKEND_*` settings.
|
||||||
|
8. [Source reference](source-reference.md) documents every backend source module,
|
||||||
|
including modules not mounted by the current application.
|
||||||
|
|
||||||
|
## Scope and conventions
|
||||||
|
|
||||||
|
All HTTP paths below include the FastAPI `root_path`, `/api`. The interactive
|
||||||
|
schema is available at `/api/docs`, the alternative reference UI at `/api/redoc`,
|
||||||
|
and the machine-readable contract at `/api/openapi.json`.
|
||||||
|
|
||||||
|
"Live" means a router is included by `src.main`. `nft_api.py` and
|
||||||
|
`nftables_api.py` contain independent routers but are not included by the current
|
||||||
|
entrypoint; they are documented as available-but-unmounted implementation paths.
|
||||||
|
|
||||||
|
Packet capture, firewall changes, bridge changes, and script deployment alter the
|
||||||
|
host system. They must be used only in a controlled environment with explicit
|
||||||
|
operator authorization.
|
||||||
112
documentation/backend/api.md
Normal file
112
documentation/backend/api.md
Normal file
@@ -0,0 +1,112 @@
|
|||||||
|
# HTTP and WebSocket API
|
||||||
|
|
||||||
|
The application is served below `/api`. FastAPI validates request models and
|
||||||
|
publishes the complete JSON Schema at `/api/openapi.json`; use it for exact field
|
||||||
|
types and the current response schema. This page documents semantics and all
|
||||||
|
mounted operations.
|
||||||
|
|
||||||
|
## General endpoints
|
||||||
|
|
||||||
|
| Method/path | Meaning |
|
||||||
|
| --- | --- |
|
||||||
|
| `GET /api/hello` | Simple application health response. |
|
||||||
|
| `GET /api/versions` | Returns the Python runtime version. |
|
||||||
|
|
||||||
|
## Network: `/api/network`
|
||||||
|
|
||||||
|
| Method/path | Parameters/body | Behaviour |
|
||||||
|
| --- | --- | --- |
|
||||||
|
| `GET /interfaces` | none | Lists interfaces, addresses, flags, MTU, MAC, state and Ethernet profile. |
|
||||||
|
| `GET /routes` | none | Lists kernel route entries and resolved output-interface names. |
|
||||||
|
| `GET /links` | none | Lists raw link information. |
|
||||||
|
| `GET /bridges` | none | Lists Linux bridges, STP state and current member details. |
|
||||||
|
| `GET /full-state` | none | Combines interfaces, routes, links and bridges into one snapshot. |
|
||||||
|
| `POST /interfaces/reset-defaults` | `{ interfaces: string[] }` | Resets each requested interface to MTU 1500 and attempts to restore an automatic Ethernet profile through `ethtool`. |
|
||||||
|
| `POST /bridge/create` | `{ name, interfaces }` | Creates a Linux bridge and attaches listed interfaces. |
|
||||||
|
| `POST /bridge/remove` | `{ name }` | Removes an existing bridge. |
|
||||||
|
| `GET /bridge/link-state-watchers` | none | Returns all watcher states. |
|
||||||
|
| `GET /bridge/{bridge_name}/link-state-watcher` | path name | Returns one bridge watcher state. |
|
||||||
|
| `POST /bridge/{bridge_name}/link-state-watcher/enable` | optional recovery holdoff | Enables member failure/recovery propagation. |
|
||||||
|
| `POST /bridge/{bridge_name}/link-state-watcher/disable` | path name | Stops and removes that watcher. |
|
||||||
|
| `WS /ws/state` | none | Receives full network-state update payloads after network mutations. |
|
||||||
|
|
||||||
|
An interface object includes its kernel index, name, state, MAC, MTU, decoded flags,
|
||||||
|
assigned IPv4/IPv6 addresses, and, where available, speed/duplex/autoneg data.
|
||||||
|
|
||||||
|
## Sniffer: `/api/sniffer`
|
||||||
|
|
||||||
|
| Method/path | Parameters/body | Behaviour |
|
||||||
|
| --- | --- | --- |
|
||||||
|
| `POST /start` | exactly one of `bridge` or `interface`; optional `bridge_capture_mode`, `benchmark_mode` | Creates a capture session. Bridge modes are `tc_ebpf` and `af_packet`. |
|
||||||
|
| `POST /stop` | optional session ID or bridge/interface selector | Stops an identified session, target sessions, or all sessions according to the request. |
|
||||||
|
| `GET /status` | none | Returns status keyed by captured interface: running/existing/up state, owner session, mode, and benchmark mode. |
|
||||||
|
| `GET /debug` | none | Returns internal session, buffered-record, tshark, telemetry, and tracker state. Treat as diagnostic output, not a stable client contract. |
|
||||||
|
|
||||||
|
The start endpoint rejects requests containing both a bridge and an interface, or
|
||||||
|
neither. A bridge defaults to `tc_ebpf`; an interface always captures using
|
||||||
|
AF_PACKET.
|
||||||
|
|
||||||
|
## Packets: `/api/packets`
|
||||||
|
|
||||||
|
| Method/path | Parameters/body | Behaviour |
|
||||||
|
| --- | --- | --- |
|
||||||
|
| `GET /packets?limit=100` | `limit` 1–10,000 | Fetches most-recent normalized packet rows. |
|
||||||
|
| `DELETE /packets?reset_id=true` | optional boolean | Clears packet history; can reset database identity state. |
|
||||||
|
| `WS /ws/packets` | none | Receives packet updates from the in-process broadcaster. |
|
||||||
|
|
||||||
|
REST history is authoritative. WebSocket clients must expect connection loss and
|
||||||
|
dropped messages for a slow subscriber, then refill missed state with `GET`.
|
||||||
|
|
||||||
|
## Analysis: `/api/analysis`
|
||||||
|
|
||||||
|
Every analysis endpoint accepts `since_minutes` when shown; its valid range is
|
||||||
|
1 minute to 30 days. Results are derived from the stored packet history and do not
|
||||||
|
claim ground truth about a physical topology or attack.
|
||||||
|
|
||||||
|
| Method/path | Main query controls | Result |
|
||||||
|
| --- | --- | --- |
|
||||||
|
| `GET /interface-hosts` | `since_minutes`, `limit_per_interface` | Likely hosts attached to each MITM-side interface. |
|
||||||
|
| `GET /interface-host-protocols` | plus `limit_protocols_per_host` | Attachment inference with per-host protocol evidence. |
|
||||||
|
| `GET /interface-protocol-paths` | `limit_paths` | Directional aggregated paths for a Sankey-style view. |
|
||||||
|
| `GET /conversations` | `limit` | Aggregated directional endpoint conversations. |
|
||||||
|
| `GET /conversation-flow-detail` | `flow_id` or directional endpoint/port fields; `protocol`, `limit_packets` | Ordered packets, subflows, and derived request/response events. |
|
||||||
|
| `GET /host-intelligence` | `limit_hosts` | Host-centric peers, service and hostname hints. |
|
||||||
|
| `GET /discovery` | `limit` | Discovery, naming and service-advertisement activity. |
|
||||||
|
| `GET /anomalies` | `limit` | Heuristic scan, beacon, rare service, reset-heavy and drop-heavy candidates. |
|
||||||
|
|
||||||
|
`conversation-flow-detail` requires a `flow_id` or enough directional fields to
|
||||||
|
identify a conversation. All analysis endpoints return 503 while the database is
|
||||||
|
unavailable and 500 when their underlying query fails.
|
||||||
|
|
||||||
|
## Firewall: `/api/firewall`
|
||||||
|
|
||||||
|
| Method/path | Body/query | Behaviour |
|
||||||
|
| --- | --- | --- |
|
||||||
|
| `GET /rules` | none | Lists nftables ruleset in a predictable structured representation, enriched with textual rule data where possible. |
|
||||||
|
| `DELETE /rules/{handle}` | optional family/table/chain defaults | Deletes the rule identified by its nft handle. |
|
||||||
|
| `POST /raw` | `{ cmd: string }` | Executes an arbitrary textual nft command and returns stdout/stderr/return code. |
|
||||||
|
|
||||||
|
The raw endpoint is intentionally powerful and must not be exposed to untrusted
|
||||||
|
clients. It changes the host firewall, not an application-local simulation.
|
||||||
|
|
||||||
|
## Scripts: `/api/scripts/scripts`
|
||||||
|
|
||||||
|
The doubled path is produced by the current combination of router and application
|
||||||
|
prefixes. Scripts are Python NFQUEUE workers installed under `/srv/fw-scripts` and
|
||||||
|
can have systemd units and isolated virtual environments.
|
||||||
|
|
||||||
|
| Method/path | Behaviour |
|
||||||
|
| --- | --- |
|
||||||
|
| `GET /` | Lists scripts and their unit mappings/status. |
|
||||||
|
| `POST /` | Uploads a script multipart payload; accepts a script, optional requirements file and required name form field. |
|
||||||
|
| `GET /{name}` | Downloads script source. |
|
||||||
|
| `GET /{name}/requirements` | Downloads its requirements file. |
|
||||||
|
| `PUT /{name}/requirements` | Replaces requirements and runs pip install in the script venv. |
|
||||||
|
| `DELETE /{name}/requirements` | Deletes requirements and removes the venv. |
|
||||||
|
| `POST /{name}/enable` | Creates/starts an NFQUEUE systemd service for a requested queue number. |
|
||||||
|
| `POST /{name}/disable` | Stops/disables the service for a queue number. |
|
||||||
|
| `DELETE /{name}` | Removes all, or one requested queue-number unit, then cleans script-related files as appropriate. |
|
||||||
|
|
||||||
|
Names allow letters, digits, `.`, `_`, and `-`; `.` and `..` are prohibited.
|
||||||
|
Repository example scripts are protected from API modification. Enabling/uploading
|
||||||
|
requirements has code-execution and host-service consequences.
|
||||||
70
documentation/backend/architecture.md
Normal file
70
documentation/backend/architecture.md
Normal file
@@ -0,0 +1,70 @@
|
|||||||
|
# Backend architecture
|
||||||
|
|
||||||
|
## Process model
|
||||||
|
|
||||||
|
`src.main` constructs one FastAPI application with `root_path="/api"`. During
|
||||||
|
startup it stores the running asyncio loop in `src.shared_objects`, creates an
|
||||||
|
asyncpg `DatabasePool`, attaches a packet broadcaster to it, creates a second
|
||||||
|
network-state broadcaster, and drains any capture records buffered before the DB
|
||||||
|
became available. Shutdown stops capture, network resources, telemetry, tshark,
|
||||||
|
and the packet tracker; then closes WebSocket broadcasters and the DB pool.
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
flowchart LR
|
||||||
|
UI[Frontend/client] --> API[FastAPI /api]
|
||||||
|
API --> NET[Network and bridge API]
|
||||||
|
API --> CAP[Sniffer API]
|
||||||
|
API --> FW[Firewall API]
|
||||||
|
API --> SCR[Script API]
|
||||||
|
CAP --> NS[network_sniffer]
|
||||||
|
NS --> PT[PacketTracker]
|
||||||
|
EBPF[tc/eBPF telemetry process] --> PT
|
||||||
|
NS <--> TS[tshark workers]
|
||||||
|
PT --> DB[(PostgreSQL packets)]
|
||||||
|
DB --> PB[PacketBroadcaster]
|
||||||
|
PB --> WS1[Packet WebSocket]
|
||||||
|
NET --> NB[Network broadcaster]
|
||||||
|
NB --> WS2[Network WebSocket]
|
||||||
|
API --> DB
|
||||||
|
```
|
||||||
|
|
||||||
|
## Component boundaries
|
||||||
|
|
||||||
|
| Component | Responsibility | Persistent state | Important side effects |
|
||||||
|
| --- | --- | --- | --- |
|
||||||
|
| `main.py` | app construction and lifecycle wiring | shared object references | starts/stops resources |
|
||||||
|
| `api/` | validates requests and presents HTTP/WebSocket contracts | none by default | may alter Linux networking, nftables, or services |
|
||||||
|
| `network_sniffer.py` | owns capture sessions and AF_PACKET sockets | in-process session map and pre-DB buffer | raw sockets, reader threads |
|
||||||
|
| `packet_tracker.py` | merges capture and telemetry observations | bounded in-memory pending entries | asynchronous database persistence |
|
||||||
|
| `database.py` | packet upsert/retrieval and SQL analysis | PostgreSQL `packets` table | WebSocket publication after single-row upserts |
|
||||||
|
| `tshark_manager.py` | optional application-protocol enrichment | worker and metadata caches | `tshark` subprocesses/threads |
|
||||||
|
| `bridge_telemetry.py` and `ebpf_bridge_events.py` | bridge tc/eBPF event collection | subprocess state and event queue | compiles/attaches tc programs |
|
||||||
|
| `bridge_link_state_manager.py` | optionally propagates member failure/recovery state | watcher registry | link and Ethernet-profile changes |
|
||||||
|
|
||||||
|
## Shared runtime state
|
||||||
|
|
||||||
|
`shared_objects.py` intentionally holds process-wide references rather than using
|
||||||
|
request-scoped dependency injection:
|
||||||
|
|
||||||
|
- `db`: initialized `DatabasePool`, or `None` after shutdown.
|
||||||
|
- `web_loop`: FastAPI event loop used when worker threads need to schedule work.
|
||||||
|
- `broadcaster`: packet update broadcaster.
|
||||||
|
- `network_broadcaster`: network-state update broadcaster.
|
||||||
|
|
||||||
|
Endpoints that require the database return HTTP 503 when `shared_objects.db` is
|
||||||
|
unavailable. Worker components should tolerate the DB not being ready by buffering
|
||||||
|
or logging failure, rather than assuming the application has fully started.
|
||||||
|
|
||||||
|
## Router mounting
|
||||||
|
|
||||||
|
| Router module | Prefix added by `main.py` | Router-local prefix | Result |
|
||||||
|
| --- | --- | --- | --- |
|
||||||
|
| `network_api` | `/network` | none | `/api/network/...` |
|
||||||
|
| `sniffer_api` | `/sniffer` | none | `/api/sniffer/...` |
|
||||||
|
| `packet_api` | `/packets` | none | `/api/packets/...` |
|
||||||
|
| `analysis_api` | `/analysis` | none | `/api/analysis/...` |
|
||||||
|
| `nft_manager` | none | `/firewall` | `/api/firewall/...` |
|
||||||
|
| `packet_scripting_api` | `/scripts` | `/scripts` | `/api/scripts/scripts/...` |
|
||||||
|
|
||||||
|
The last row reflects the current code exactly. It is worth preserving this fact in
|
||||||
|
examples until the duplicated prefix is deliberately changed.
|
||||||
82
documentation/backend/capture-pipeline.md
Normal file
82
documentation/backend/capture-pipeline.md
Normal file
@@ -0,0 +1,82 @@
|
|||||||
|
# Packet capture and correlation pipeline
|
||||||
|
|
||||||
|
## Capture modes
|
||||||
|
|
||||||
|
A sniffer session targets exactly one interface or bridge.
|
||||||
|
|
||||||
|
- **Interface target:** an `AF_PACKET` raw socket is opened on that interface;
|
||||||
|
its effective mode is always `af_packet`.
|
||||||
|
- **Bridge target with `af_packet`:** the bridge's member interfaces are captured
|
||||||
|
with raw sockets.
|
||||||
|
- **Bridge target with `tc_ebpf` (default):** no raw socket is opened for bridge
|
||||||
|
ports. `BridgeTelemetryManager` manages an eBPF/tc helper that exports ingress
|
||||||
|
raw data and egress/drop verdict-related events.
|
||||||
|
- **Benchmark mode:** preserves session accounting but skips the normal userspace
|
||||||
|
packet processing path, allowing capture-overhead measurements.
|
||||||
|
|
||||||
|
Sessions have UUIDs and record their label, target type, mode, snapshot of bridge
|
||||||
|
ports, capture interfaces, socket map, thread, and stop event. Stopping by session
|
||||||
|
ID is preferred. A target-specific stop finds matching sessions; an unqualified
|
||||||
|
stop stops every session.
|
||||||
|
|
||||||
|
## End-to-end lifecycle
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
sequenceDiagram
|
||||||
|
participant C as Capture socket or tc/eBPF
|
||||||
|
participant N as network_sniffer
|
||||||
|
participant T as tshark manager
|
||||||
|
participant P as PacketTracker
|
||||||
|
participant D as DatabasePool
|
||||||
|
participant W as packet WebSocket
|
||||||
|
C->>N: frame / telemetry event
|
||||||
|
N->>N: parse headers, identity, observation metadata
|
||||||
|
N->>T: lookup or schedule enrichment
|
||||||
|
N->>P: capture observation
|
||||||
|
C->>P: ingress/egress/verdict telemetry
|
||||||
|
P->>P: correlate, merge and finalize record
|
||||||
|
P->>D: upsert packet
|
||||||
|
D->>W: publish normalized row
|
||||||
|
```
|
||||||
|
|
||||||
|
`network_sniffer.parse_packet` parses Scapy packet objects, while
|
||||||
|
`parse_packet_bytes` supports raw data. It extracts link, network, and transport
|
||||||
|
fields; adds capture session/observation data; calculates or obtains correlation
|
||||||
|
identifiers; and hands observations to `PacketTracker`. If the shared database or
|
||||||
|
event loop is not yet usable, records are retained in a bounded in-memory buffer;
|
||||||
|
`drain_buffer_to_shared_db` flushes it at application startup.
|
||||||
|
|
||||||
|
## Identity and merging
|
||||||
|
|
||||||
|
`packet_identity.build_packet_uid` makes a stable hash-based fallback identity
|
||||||
|
from normalized packet fields. `packet_mark` decodes the shared skb-mark layout:
|
||||||
|
it normalizes an observed mark, extracts a packet ID, and extracts a verdict hint.
|
||||||
|
Kernel-mark identity is preferred when present; the hash fallback keeps capture and
|
||||||
|
telemetry correlation possible when it is not.
|
||||||
|
|
||||||
|
`PacketTracker` aggregates observations in a bounded dictionary. It deduplicates
|
||||||
|
observations, keeps capture and telemetry provenance, combines ingress/egress and
|
||||||
|
verdict timing, and delays finalization briefly so companion events can arrive.
|
||||||
|
It writes finalized or aged dirty records in batches, retries failed persistence
|
||||||
|
with bounded exponential backoff, discards pending records for stopped interfaces,
|
||||||
|
and exposes a debug snapshot. Its limits and timings are all configured through
|
||||||
|
`BACKEND_PACKET_TRACKER_*` settings.
|
||||||
|
|
||||||
|
## tshark enrichment
|
||||||
|
|
||||||
|
`TsharkManager` starts one long-lived `tshark` process per enabled interface. It
|
||||||
|
reads JSON output in a thread, derives protocol stacks, HTTP/TLS/DNS information,
|
||||||
|
TCP flags, and stream context, then caches matching data for a configurable time
|
||||||
|
window. The capture parser can use a heuristic immediately and the manager can
|
||||||
|
backfill metadata or stream context into already stored rows. tshark is optional in
|
||||||
|
the logical pipeline but enabled by default; a missing executable or worker failure
|
||||||
|
is logged and does not stop capture.
|
||||||
|
|
||||||
|
## Realtime delivery
|
||||||
|
|
||||||
|
`PacketBroadcaster` maintains a bounded asyncio queue per subscriber. A successful
|
||||||
|
single-row database upsert serializes the row and publishes it to subscribers.
|
||||||
|
Slow consumers lose queued messages when their individual queue is full rather than
|
||||||
|
blocking the capture or database path. `/api/packets/ws/packets` is therefore a
|
||||||
|
live-update channel, not a lossless event log; clients should retrieve history over
|
||||||
|
REST and use the WebSocket for incremental updates.
|
||||||
73
documentation/backend/configuration.md
Normal file
73
documentation/backend/configuration.md
Normal file
@@ -0,0 +1,73 @@
|
|||||||
|
# Configuration and deployment
|
||||||
|
|
||||||
|
## Runtime dependencies
|
||||||
|
|
||||||
|
The service runs with Python 3.11 in the supplied Dockerfile and starts Uvicorn as
|
||||||
|
`src.main:app` on port 8000 with reload enabled. Python dependencies include
|
||||||
|
FastAPI/Pydantic, asyncpg, pyroute2, Scapy, pip-nftables, multipart handling, and
|
||||||
|
WebSocket support. The image installs build tools, libpcap development headers,
|
||||||
|
pkg-config, and `tshark`.
|
||||||
|
|
||||||
|
The host also needs facilities that a minimal application container normally does
|
||||||
|
not have: a reachable PostgreSQL database, Linux network namespace permissions,
|
||||||
|
raw-socket capability, access to `ip`/pyroute2 netlink operations, nftables and
|
||||||
|
appropriate capability, `ethtool` where profile/reset functions are used,
|
||||||
|
systemd/systemctl for scripts, and BCC/eBPF/tc tooling for `tc_ebpf` capture.
|
||||||
|
|
||||||
|
## Environment variables
|
||||||
|
|
||||||
|
All settings are loaded once by `src.config.load_settings`. Empty values use their
|
||||||
|
default. Boolean true values are `1`, `true`, `yes`, or `on` (case-insensitive).
|
||||||
|
|
||||||
|
| Variable | Default | Purpose |
|
||||||
|
| --- | --- | --- |
|
||||||
|
| `BACKEND_DB_DSN` | `postgresql://mitm_user:mitm_password@localhost:5432/mitm_db` | PostgreSQL connection string. |
|
||||||
|
| `BACKEND_LOG_LEVEL` | `DEBUG` | Python logging level. |
|
||||||
|
| `BACKEND_DB_POOL_MIN_SIZE` / `MAX_SIZE` | `1` / `5` | asyncpg pool bounds. |
|
||||||
|
| `BACKEND_BROADCAST_QUEUE_MAXSIZE` | `1024` | Per-WebSocket broadcast queue size. |
|
||||||
|
| `BACKEND_PACKET_TRACKER_FINALIZE_DELAY_SECONDS` | `0.25` | Wait for related observations before finalizing. |
|
||||||
|
| `BACKEND_PACKET_TRACKER_RETENTION_SECONDS` | `10.0` | Pending-entry retention. |
|
||||||
|
| `BACKEND_PACKET_TRACKER_MIN_FLUSH_INTERVAL_SECONDS` | `0.05` | Minimum persistence flush interval. |
|
||||||
|
| `BACKEND_PACKET_TRACKER_PERSIST_TIMEOUT_SECONDS` | `2.0` | One persistence attempt timeout. |
|
||||||
|
| `BACKEND_PACKET_TRACKER_BATCH_PERSIST_TIMEOUT_SECONDS` | `10.0` | Batch persistence timeout. |
|
||||||
|
| `BACKEND_PACKET_TRACKER_PERSIST_RETRY_BACKOFF_SECONDS` / `MAX_SECONDS` | `0.25` / `5.0` | Retry backoff bounds. |
|
||||||
|
| `BACKEND_PACKET_TRACKER_ERROR_LOG_INTERVAL_SECONDS` | `5.0` | Failure-log throttling interval. |
|
||||||
|
| `BACKEND_PACKET_TRACKER_FLUSH_BATCH_SIZE` | `500` | Maximum batch size; clamped to at least 1. |
|
||||||
|
| `BACKEND_PACKET_TRACKER_MAX_ENTRIES` | `50000` | Bounded in-memory correlation capacity; clamped to at least 1. |
|
||||||
|
| `BACKEND_PACKET_TRACKER_MAX_PERSIST_FAILURES` | `3` | Failure threshold; clamped to at least 1. |
|
||||||
|
| `BACKEND_PACKET_TRACKER_MAX_DIRTY_AGE_SECONDS` | `60.0` | Maximum age before dirty data must be flushed. |
|
||||||
|
| `BACKEND_PACKET_TRACKER_STOP_JOIN_TIMEOUT_SECONDS` | `2.0` | Tracker thread join timeout. |
|
||||||
|
| `BACKEND_PACKET_TRACKER_REJECT_CORRELATION_WINDOW_SECONDS` | `1.0` | Rejection-event matching window. |
|
||||||
|
| `BACKEND_SNIFFER_BUFFER_CAPACITY` | `20000` | Pre-DB capture buffer capacity. |
|
||||||
|
| `BACKEND_SNIFFER_SOCKET_RCVBUF_BYTES` | `4194304` | Requested raw-socket receive buffer. |
|
||||||
|
| `BACKEND_SNIFFER_SELECTOR_TIMEOUT_SECONDS` | `1.0` | Reader select timeout. |
|
||||||
|
| `BACKEND_SNIFFER_RECV_BYTES` | `65536` | Maximum raw receive length. |
|
||||||
|
| `BACKEND_SNIFFER_BUFFER_DRAIN_INTERVAL_SECONDS` | `5.0` | Buffered-record drain frequency. |
|
||||||
|
| `BACKEND_SNIFFER_THREAD_JOIN_TIMEOUT_SECONDS` | `2.0` | Session reader join timeout. |
|
||||||
|
| `BACKEND_BRIDGE_BPF_BUILD_DIR` | `/tmp/mitm-bpf` | eBPF build artifacts directory. |
|
||||||
|
| `BACKEND_BRIDGE_TELEMETRY_RAW_SAMPLE_EVERY` / `META_SAMPLE_EVERY` | `1` / `1` | Raw/meta sampling rates; zero is allowed. |
|
||||||
|
| `BACKEND_BRIDGE_TELEMETRY_INGRESS_PERF_PAGES` / `META_PERF_PAGES` | `256` / `128` | eBPF perf-buffer page counts. |
|
||||||
|
| `BACKEND_BRIDGE_TELEMETRY_EVENT_QUEUE_MAXSIZE` | `20000` | Telemetry event queue cap. |
|
||||||
|
| `BACKEND_BRIDGE_TELEMETRY_QUEUE_RECOVERY_SIZE` | `1000` | Queue recovery threshold. |
|
||||||
|
| `BACKEND_BRIDGE_TELEMETRY_DROP_LOG_INTERVAL_SECONDS` | `5.0` | Telemetry-drop log throttling. |
|
||||||
|
| `BACKEND_BRIDGE_LINK_STATE_THREAD_JOIN_TIMEOUT_SECONDS` | `2.0` | Link watcher join timeout. |
|
||||||
|
| `BACKEND_BRIDGE_LINK_STATE_FAILURE_HOLDOFF_SECONDS` / `RECOVERY_HOLDOFF_SECONDS` | `0.75` / `1.0` | Delay before propagating failure/recovery. |
|
||||||
|
| `BACKEND_BRIDGE_LINK_STATE_DEGRADED_RECHECK_SECONDS` | `0.5` | Degraded-link polling period. |
|
||||||
|
| `BACKEND_TELEMETRY_PROCESS_STOP_TIMEOUT_SECONDS` / `READER_JOIN_TIMEOUT_SECONDS` | `3.0` / `2.0` | Telemetry subprocess shutdown limits. |
|
||||||
|
| `BACKEND_TSHARK_ENABLED` | `true` | Enables tshark worker management. |
|
||||||
|
| `BACKEND_TSHARK_DISPLAY_FILTER` | empty | Optional tshark display filter. |
|
||||||
|
| `BACKEND_TSHARK_TRY_HEURISTIC_FIRST` | `true` | Applies local heuristic before tshark match. |
|
||||||
|
| `BACKEND_TSHARK_CACHE_TTL_SECONDS` | `5.0` | Enrichment cache lifetime. |
|
||||||
|
| `BACKEND_TSHARK_MATCH_WINDOW_MS` | `5000` | Capture-to-tshark matching window. |
|
||||||
|
| `BACKEND_TSHARK_READER_JOIN_TIMEOUT_SECONDS` / `PROCESS_STOP_TIMEOUT_SECONDS` | `2.0` / `3.0` | tshark shutdown limits. |
|
||||||
|
|
||||||
|
## Operational safeguards
|
||||||
|
|
||||||
|
Run the API behind an authenticated, access-controlled boundary. The configured
|
||||||
|
CORS policy currently permits every origin, method, and header; it is convenient
|
||||||
|
for development but should not be treated as an authorization control. Keep DB
|
||||||
|
credentials out of version control and use a production-specific DSN.
|
||||||
|
|
||||||
|
Before starting capture, verify target interface/bridge names and ensure recovery
|
||||||
|
access to the host. Before using firewall or script endpoints, snapshot the nft
|
||||||
|
ruleset and understand which systemd units and filesystem paths are in scope.
|
||||||
65
documentation/backend/data-and-analysis.md
Normal file
65
documentation/backend/data-and-analysis.md
Normal file
@@ -0,0 +1,65 @@
|
|||||||
|
# Packet data, persistence, and analysis
|
||||||
|
|
||||||
|
## Packet record
|
||||||
|
|
||||||
|
`Models/packets.py` defines the normalized `PacketDBModel` returned by packet
|
||||||
|
history APIs. It represents one correlated packet record, not necessarily one raw
|
||||||
|
capture callback. A record may combine several observations.
|
||||||
|
|
||||||
|
| Field group | Fields | Meaning |
|
||||||
|
| --- | --- | --- |
|
||||||
|
| Identity | `id`, `timestamp`, `updated_at`, `correlation_key`, `correlation_source`, `packet_id`, `packet_uid`, `skb_mark` | Database identity and the evidence used to correlate capture/telemetry data. |
|
||||||
|
| Path | `capture_iface`, `ingress_if`, `egress_if`, `capture_session_id`, `capture_sources` | Where and how it was observed. |
|
||||||
|
| Link/network/transport | MACs, EtherType, IP protocol, IPs, ports, VLAN, length | Parsed packet headers. Raw numeric values are retained beside human-readable names. |
|
||||||
|
| Application enrichment | `flow_id`, app protocol/master protocol/category/confidence/hostname/encryption/risk, `dpi_metadata` | tshark-derived context when available. |
|
||||||
|
| Evidence | `raw_b64`, `raw_present`, capture/telemetry metadata and `capture_observations` | Raw bytes and provenance; may be absent. |
|
||||||
|
| Outcome | `verdict`, reason/confidence, ingress/egress/verdict timestamps | Forwarding outcome inferred from telemetry. |
|
||||||
|
|
||||||
|
Each `PacketObservationModel` identifies whether its contribution was `capture` or
|
||||||
|
`telemetry`, the source, interface, timestamp, event type, capture mode, session,
|
||||||
|
and optional reason. Consumers should not assume that every optional field exists:
|
||||||
|
AF_PACKET, tc/eBPF, and enrichment sources provide different evidence.
|
||||||
|
|
||||||
|
## DatabasePool
|
||||||
|
|
||||||
|
`DatabasePool` is an asyncpg wrapper initialized from `BACKEND_DB_DSN`. Startup
|
||||||
|
makes compatibility changes to an existing `packets` table: fills missing
|
||||||
|
timestamps, sets timestamp defaults/non-null constraints, adds
|
||||||
|
`capture_observations` JSONB if needed, and creates an index on `packet_uid`.
|
||||||
|
|
||||||
|
Writes use an upsert keyed by `correlation_key`. `upsert_packet` returns a row and
|
||||||
|
publishes it to the packet broadcaster; `upsert_packets` is a batched performance
|
||||||
|
path and does not individually publish rows. Before writing, it normalizes JSON
|
||||||
|
fields and attaches derived protocol, flow, and analysis fields.
|
||||||
|
|
||||||
|
Read/analysis methods are:
|
||||||
|
|
||||||
|
- `fetch_latest(limit)` for history.
|
||||||
|
- `backfill_packet_metadata` and `backfill_stream_metadata` for late tshark data.
|
||||||
|
- `infer_interface_hosts`, `infer_interface_host_protocols`, and
|
||||||
|
`infer_interface_protocol_paths` for topology/protocol views.
|
||||||
|
- `analyze_conversations` and `fetch_conversation_flow_detail` for directional
|
||||||
|
communication views.
|
||||||
|
- `analyze_host_intelligence`, `analyze_discovery_activity`, and
|
||||||
|
`analyze_anomalies` for investigation aids.
|
||||||
|
- `clear_all_packets(reset_identity)` for destructive history cleanup.
|
||||||
|
|
||||||
|
## Analysis interpretation
|
||||||
|
|
||||||
|
The analysis API runs SQL aggregations over what the system captured. It infers
|
||||||
|
attachment from traffic evidence, groups protocol paths and conversations, and
|
||||||
|
derives host/service/hostname hints. Its anomaly queries rank plausible scan,
|
||||||
|
beacon, rare-service, TCP-reset, and drop-heavy patterns. These are leads for an
|
||||||
|
operator—not assertions of a network's real topology, attribution, or malicious
|
||||||
|
intent. Missing capture events, encrypted traffic, NAT, asymmetric paths, and
|
||||||
|
limits change the output.
|
||||||
|
|
||||||
|
## Enumerations and configuration models
|
||||||
|
|
||||||
|
`Models/etherType.py` provides a string-valued `EtherTypeEnum` and
|
||||||
|
`ethertype_from_int`; `Models/ip_protocol.py` provides `IPProtocolEnum` and
|
||||||
|
`protocol_from_number`. They turn numeric protocol fields into readable labels
|
||||||
|
while keeping raw values. `Models/netplan.py` provides Pydantic schemas for
|
||||||
|
nameservers, Ethernet settings, bridge settings, and a full Netplan-style network
|
||||||
|
configuration. These models are reusable schemas; they are not a substitute for
|
||||||
|
applying a Netplan configuration in the currently mounted API.
|
||||||
90
documentation/backend/host-integration.md
Normal file
90
documentation/backend/host-integration.md
Normal file
@@ -0,0 +1,90 @@
|
|||||||
|
# Linux host integration and side effects
|
||||||
|
|
||||||
|
## Network and bridge control
|
||||||
|
|
||||||
|
`api/network_api.py` retains process-wide pyroute2 `IPRoute` and `NDB` objects.
|
||||||
|
It reads addresses, link flags, routes and bridge membership through netlink, and
|
||||||
|
uses NDB/pyroute2 to create or remove bridges. Resetting interfaces executes
|
||||||
|
`ethtool` and changes MTU/profile values. These operations affect the host's live
|
||||||
|
connectivity; API errors must be treated as operational failures, not merely input
|
||||||
|
validation failures.
|
||||||
|
|
||||||
|
`utilities/interface_bridge_helpers.py` is the low-level read layer. It checks
|
||||||
|
interface presence/up state, reads sysfs operational/carrier/admin/MTU values,
|
||||||
|
obtains Ethernet profile data using `ethtool`, caches profile data, and reads bridge
|
||||||
|
members from sysfs. It deliberately supplies best-effort information when a driver
|
||||||
|
or platform cannot report every property.
|
||||||
|
|
||||||
|
`bridge_link_state_manager.py` owns optional event-driven bridge watchers. Each
|
||||||
|
watcher tracks Ethernet profile and member readiness, uses failure and recovery
|
||||||
|
holdoffs to avoid flapping, and adjusts selected peer state so an inline bridge
|
||||||
|
reacts coherently to member link loss. `BridgeLinkStateManager` indexes watchers,
|
||||||
|
enables/disables them, reports individual/all status, and stops all during shutdown.
|
||||||
|
|
||||||
|
## eBPF/tc telemetry
|
||||||
|
|
||||||
|
`bridge_telemetry.py` manages the lifecycle of the telemetry helper. Its
|
||||||
|
`update_sessions` method reconciles currently requested bridge ports with the
|
||||||
|
subprocess; `stop` terminates it and `get_debug_snapshot` provides operator
|
||||||
|
diagnostics. It does not itself parse kernel events.
|
||||||
|
|
||||||
|
`ebpf_bridge_events.py` is the helper process. It builds BPF source, attaches tc
|
||||||
|
programs to requested interfaces, reads perf events, and writes JSON-safe event
|
||||||
|
payloads. Events cover ingress raw capture plus egress/drop metadata, including
|
||||||
|
interfaces, MAC/IP information, packet identity, event/reason names, and timing.
|
||||||
|
It cleans existing clsact qdiscs/program attachment as part of setup/cleanup. This
|
||||||
|
requires an appropriate kernel, BCC Python bindings/toolchain, tc, and privileges.
|
||||||
|
|
||||||
|
`tools/ebpf/mark_packet_id.c` is related kernel-side support for packet marking;
|
||||||
|
the mark is decoded by `utilities/packet_mark.py` and used in tracker correlation.
|
||||||
|
|
||||||
|
## nftables
|
||||||
|
|
||||||
|
The mounted `api/nft_manager.py` uses `pip-nftables` to list JSON/text rulesets,
|
||||||
|
normalize them into stable table/chain/rule models, parse rule priorities/text, and
|
||||||
|
delete a rule by handle. Its raw-command endpoint forwards textual nft commands.
|
||||||
|
It therefore needs capability to inspect and change the host nftables ruleset.
|
||||||
|
|
||||||
|
Two alternative implementations exist but are currently unmounted:
|
||||||
|
|
||||||
|
- `api/nft_api.py` is bridge-family oriented. It models meta, Ethernet, IP, port,
|
||||||
|
conntrack, verdict, reject, log, and raw expressions; can generate previews,
|
||||||
|
list rules with authoritative handles, add/delete/update rules, and uses the
|
||||||
|
`nft` CLI.
|
||||||
|
- `api/nftables_api.py` is a stateless typed replacement API. It models matches and
|
||||||
|
actions, chooses a pyroute2 binding when viable or a CLI wrapper otherwise,
|
||||||
|
ensures table/chain presence, reconstructs readable rules, and replaces a chain's
|
||||||
|
ruleset. Its own source warns that a running asyncio loop may force CLI fallback.
|
||||||
|
|
||||||
|
Do not mount more than one firewall router without an explicit API versioning and
|
||||||
|
conflict review: all manipulate shared kernel state and have overlapping concepts.
|
||||||
|
|
||||||
|
## NFQUEUE script services
|
||||||
|
|
||||||
|
`api/packet_scripting_api.py` manages executable Python scripts. It makes these
|
||||||
|
directories at import time: `/srv/fw-scripts`, `/srv/fw-scripts/venvs`, and the
|
||||||
|
repository's `backend/example_scripts`. Scripts are named `<name>.py`; requirements
|
||||||
|
are `<name>-requirements.txt`; virtual environments are per-script. Units use the
|
||||||
|
deterministic name `fw-script-<name>-q<queue>.service` and are written under
|
||||||
|
`/etc/systemd/system`.
|
||||||
|
|
||||||
|
The module discovers services through `systemctl`, writes/parses unit `ExecStart`,
|
||||||
|
runs `daemon-reload`, starts/stops/enables/disables units, creates virtualenvs, and
|
||||||
|
uses pip to install user-provided requirements. Startup can copy protected example
|
||||||
|
scripts and optionally deploy them from `<name>.deploy.json`. This API is a remote
|
||||||
|
code/service-management surface and requires strict authentication plus host-level
|
||||||
|
least privilege.
|
||||||
|
|
||||||
|
## External subprocesses
|
||||||
|
|
||||||
|
| Integration | Commands/facility | Used by |
|
||||||
|
| --- | --- | --- |
|
||||||
|
| tshark | long-lived `tshark` subprocesses | DPI enrichment |
|
||||||
|
| nftables | `nft` CLI and/or pip-nftables bindings | firewall APIs |
|
||||||
|
| Ethernet control | `ethtool` | interface profile/reset |
|
||||||
|
| system services | `systemctl`, virtualenv, pip | script lifecycle |
|
||||||
|
| BPF/tc | BCC, tc, qdisc/program attachment | bridge telemetry |
|
||||||
|
|
||||||
|
Failures are generally logged and translated to endpoint failures or degraded
|
||||||
|
capture. Operators should collect `/api/sniffer/debug`, service logs, nftables
|
||||||
|
state, and interface state when investigating a problem.
|
||||||
233
documentation/backend/sniffing.md
Normal file
233
documentation/backend/sniffing.md
Normal file
@@ -0,0 +1,233 @@
|
|||||||
|
# Technical reference: packet sniffing modes
|
||||||
|
|
||||||
|
This document specifies the implemented capture behavior in `network_sniffer.py`,
|
||||||
|
`bridge_telemetry.py`, `ebpf_bridge_events.py`, and `packet_tracker.py`. It makes a
|
||||||
|
deliberate distinction between observed facts and inferred forwarding results.
|
||||||
|
|
||||||
|
## Session model and mode selection
|
||||||
|
|
||||||
|
A capture session has a UUID and targets exactly one interface or exactly one
|
||||||
|
bridge. A bridge is expanded once with `get_bridge_ports_once`; its member list is
|
||||||
|
a creation-time snapshot. Later bridge membership changes are not added to the
|
||||||
|
existing session. Session state includes target label/type, effective mode, benchmark
|
||||||
|
flag, port snapshot, AF_PACKET sockets, optional reader thread, stop event, and
|
||||||
|
benchmark counters.
|
||||||
|
|
||||||
|
| Request | Effective mode | Capture source | Path/outcome evidence |
|
||||||
|
| --- | --- | --- | --- |
|
||||||
|
| Interface, any requested mode | `af_packet` | One raw socket on the interface | Packet-socket type labels an outgoing copy as egress; no kernel verdict telemetry. |
|
||||||
|
| Bridge, `af_packet` | `af_packet` | One raw socket per snapshot bridge port | Two matching port observations can infer forwarding. |
|
||||||
|
| Bridge, `tc_ebpf` (default) | `tc_ebpf` | One tc/eBPF helper across snapshot bridge ports | TC ingress/egress and skb-free/drop events, normally matched by skb mark. |
|
||||||
|
| Any mode with benchmark enabled | Same hook/socket setup | Counted but not processed | No parsing, enrichment, DB write, or WebSocket event. |
|
||||||
|
|
||||||
|
Interface targets always use AF_PACKET. The TC/eBPF mode is only selected for a
|
||||||
|
bridge target. A stop by session ID is the safest selector. Stopping a session also
|
||||||
|
discards pending tracker entries relating to its interfaces, so unpersisted data can
|
||||||
|
be lost deliberately at shutdown.
|
||||||
|
|
||||||
|
Multiple sessions may overlap on an interface. This is not an independent-capture
|
||||||
|
guarantee: the TC manager maps an interface to several sessions but assigns raw
|
||||||
|
ingress parsing to the first sorted session ID.
|
||||||
|
|
||||||
|
## AF_PACKET capture
|
||||||
|
|
||||||
|
### Socket behavior
|
||||||
|
|
||||||
|
For each capture interface the service opens `AF_PACKET` / `SOCK_RAW` with protocol
|
||||||
|
`htons(0x0003)` (`ETH_P_ALL`), requests the configured receive buffer (default
|
||||||
|
4 MiB), best-effort requests `TPACKET_V3`, binds to `(ifname, 0)`, and makes the
|
||||||
|
socket non-blocking. Failure to set the buffer or TPACKET version is non-fatal.
|
||||||
|
Failure to create or bind leaves the interface uncaptured; session creation can still
|
||||||
|
complete. This requires raw-socket privilege, commonly `CAP_NET_RAW`.
|
||||||
|
|
||||||
|
One daemon reader thread is started only when the session has sockets. It uses a
|
||||||
|
selector, receives at most `BACKEND_SNIFFER_RECV_BYTES` bytes per event (default
|
||||||
|
65,536), stamps the frame with userspace UTC receive time, creates Scapy `Ether`,
|
||||||
|
and calls the common parser. `ENODEV`, `ENETDOWN`, and `EBADF` close the affected
|
||||||
|
socket; it is not reopened in that session. The thread periodically attempts a
|
||||||
|
pre-DB-buffer drain during selector idle time.
|
||||||
|
|
||||||
|
### Direction and bridge inference
|
||||||
|
|
||||||
|
Packet-socket address metadata is used only as follows: `PACKET_OUTGOING` (normally
|
||||||
|
4) becomes `path_role: egress`; every other packet type becomes `path_role:
|
||||||
|
ingress`. This is a packet-socket perspective, not proof of a Linux bridge decision.
|
||||||
|
|
||||||
|
For a bridge session, the tracker groups AF_PACKET observations by session ID. It
|
||||||
|
uses an explicit ingress observation if available, otherwise the earliest one. It
|
||||||
|
prefers an explicit egress observation on a different port, otherwise a later
|
||||||
|
different-port observation. If one correlated packet is seen on at least two ports,
|
||||||
|
the tracker records:
|
||||||
|
|
||||||
|
```text
|
||||||
|
verdict = accept
|
||||||
|
verdict_reason = bridge-af_packet-forwarded-observed
|
||||||
|
verdict_confidence = medium
|
||||||
|
```
|
||||||
|
|
||||||
|
This means matching evidence was observed on two bridge ports. It does not prove a
|
||||||
|
particular kernel forwarding verdict and can be affected by duplicate copies, loops,
|
||||||
|
or fallback-identity collisions. A single-interface AF_PACKET record has no terminal
|
||||||
|
verdict from AF_PACKET itself.
|
||||||
|
|
||||||
|
### AF_PACKET implications
|
||||||
|
|
||||||
|
AF_PACKET provides full observed frame bytes without BCC or tc changes and is the
|
||||||
|
only interface-capture mode. It neither alters packets nor controls forwarding. It
|
||||||
|
also has no definitive drop visibility, reports userspace rather than kernel event
|
||||||
|
time, can observe local/outgoing copies, and can lose traffic under socket/userspace
|
||||||
|
load. The full-frame Scapy parse, tshark lookup, tracking and persistence path makes
|
||||||
|
it more expensive than sampled telemetry.
|
||||||
|
|
||||||
|
## TC/eBPF bridge capture
|
||||||
|
|
||||||
|
### Collector lifecycle and destructive qdisc behavior
|
||||||
|
|
||||||
|
Bridge sessions in `tc_ebpf` mode are aggregated into one helper process. Any change
|
||||||
|
to the active *interface set* stops the helper and recreates it for the new set;
|
||||||
|
there is a capture gap during that restart. The helper attaches direct-action
|
||||||
|
`BPF.SCHED_CLS` programs at TC ingress (`ffff:fff2`, handle `:20`) and egress
|
||||||
|
(`ffff:fff3`, handle `:30`) to every bridge **member interface**, not the bridge
|
||||||
|
device itself.
|
||||||
|
|
||||||
|
Before attachment the helper runs `tc qdisc del dev <iface> clsact` (ignoring its
|
||||||
|
result), then `tc qdisc add dev <iface> clsact`. It deletes `clsact` again for every
|
||||||
|
instrumented interface at helper shutdown and after an attachment failure.
|
||||||
|
|
||||||
|
> Starting, restarting, failing, or stopping TC/eBPF capture can remove pre-existing
|
||||||
|
> clsact qdiscs and their filters. Do not use it on interfaces with unrelated TC
|
||||||
|
> configuration unless coexistence and recovery are explicitly managed.
|
||||||
|
|
||||||
|
The BCC Python runtime, a compatible kernel, BPF/tracepoint access, TC and netlink
|
||||||
|
privileges are required. Session creation does not wait for a collector health
|
||||||
|
acknowledgement, so a successful start response is not proof that BPF attached.
|
||||||
|
|
||||||
|
### Kernel event generation
|
||||||
|
|
||||||
|
The helper opens a raw-ingress perf buffer and a metadata perf buffer. The ingress
|
||||||
|
TC program creates an skb mark only when it is zero, using the low 28 bits of
|
||||||
|
`bpf_ktime_get_ns()` and replacing zero with one. It preserves any existing nonzero
|
||||||
|
mark. It extracts Ethernet addresses, EtherType, a single 802.1Q/802.1AD VLAN ID,
|
||||||
|
ARP IPv4 addresses, and IPv4/IPv6 addresses with TCP/UDP ports. IPv6 extension
|
||||||
|
headers are not traversed; the base next-header is used as protocol.
|
||||||
|
|
||||||
|
The egress TC program never creates a mark. It exports metadata only for marked
|
||||||
|
packets. The `skb:kfree_skb` tracepoint reads the linear skb representation and
|
||||||
|
exports a drop event only for marked skbs whose device is a selected interface.
|
||||||
|
The emitted payload contains userspace and kernel-monotonic timestamps, interface,
|
||||||
|
mark, length, parsed L2–L4 fields, and event type. Drop events add a numerical
|
||||||
|
reason and `skb_drop_reason_<n>` label. Ingress events may contain `raw_b64`; egress
|
||||||
|
and drop events do not.
|
||||||
|
|
||||||
|
A kfree_skb event is evidence that a marked skb was freed in the kernel context. It
|
||||||
|
is not automatically evidence that nftables caused the outcome; interpret the
|
||||||
|
reason code in the context of kernel behavior and other instrumentation.
|
||||||
|
|
||||||
|
### Sampling
|
||||||
|
|
||||||
|
Raw and metadata sampling are independent settings.
|
||||||
|
|
||||||
|
| Value | Effect |
|
||||||
|
| --- | --- |
|
||||||
|
| `0` | Never emits that sample category. |
|
||||||
|
| `1` | Emits every marked packet in that category. |
|
||||||
|
| `N > 1` | Emits when `skb_mark % N == 0`. |
|
||||||
|
|
||||||
|
At ingress, a raw-selected packet emits a raw event; only a packet not chosen for
|
||||||
|
raw can emit an ingress metadata event. Egress and drop use metadata sampling only.
|
||||||
|
Therefore raw-enabled/meta-disabled capture stores sampled ingress frame records
|
||||||
|
without egress/drop visibility; raw-disabled/meta-enabled capture produces
|
||||||
|
metadata-only rows without raw bytes. Both enabled does not make raw and metadata
|
||||||
|
populations identical.
|
||||||
|
|
||||||
|
Sampling uses the entire existing skb mark. The documented mark layout reserves
|
||||||
|
bits 0–27 for packet ID and upper bits for drop/reject hints. This capture program
|
||||||
|
creates only the low-28-bit value for previously zero marks; it does not set verdict
|
||||||
|
hints. Any other mark-using subsystem must coordinate its mark semantics, because
|
||||||
|
it can change both sampling and correlation.
|
||||||
|
|
||||||
|
### Userspace event handling and loss
|
||||||
|
|
||||||
|
The manager reads JSON helper output into a bounded queue. For a non-benchmark
|
||||||
|
ingress event with `raw_b64`, it decodes the frame and sends it into the common Scapy
|
||||||
|
parser as source `tc_ingress_raw`, with `packet_id`, `skb_mark`, and capture mode
|
||||||
|
`tc_ingress`. It then sends every non-benchmark ingress/egress/drop event to the
|
||||||
|
tracker. A sampled raw ingress packet usually therefore has both a parsed capture
|
||||||
|
observation and a telemetry observation under the same mark-derived key.
|
||||||
|
|
||||||
|
When the telemetry queue is full, the manager drops oldest queued events down to
|
||||||
|
`BACKEND_BRIDGE_TELEMETRY_QUEUE_RECOVERY_SIZE`, attempts to keep the new event, and
|
||||||
|
counts dropped events and raw payloads. Perf buffers can also lose samples before
|
||||||
|
userspace. Neither loss mechanism is recovered. `/api/sniffer/debug` reports queue
|
||||||
|
size, queue drops, benchmark counts, collector interfaces, and tracker statistics.
|
||||||
|
|
||||||
|
### TC/eBPF implications
|
||||||
|
|
||||||
|
This mode yields better within-host correlation and explicit TC egress evidence. A
|
||||||
|
matching egress produces `accept`, `egress-observed`, confidence `high`; a matching
|
||||||
|
drop produces `drop` (or mark hint), confidence `high`. Absence of egress is not
|
||||||
|
proof of a drop: sampling, perf loss, queue loss, an uninstrumented path, teardown,
|
||||||
|
or collector failure can all explain it. Raw bytes are ingress-only and sampled.
|
||||||
|
|
||||||
|
## Common parsing, enrichment, and identity
|
||||||
|
|
||||||
|
Both modes use `parse_packet` / `parse_packet_bytes`. The parser records Ethernet
|
||||||
|
addresses, EtherType and VLAN, ARP operation/addressing, IPv4 ID or IPv6 base
|
||||||
|
header, TCP sequence/acknowledgement/flags, UDP ports, ICMP/ICMPv6 type/code, and
|
||||||
|
an embedded IPv4 tuple from eligible ICMP errors. It stores full raw frame bytes
|
||||||
|
when supplied by AF_PACKET or sampled TC ingress.
|
||||||
|
|
||||||
|
tshark workers are enabled for non-benchmark capture interfaces. They may add
|
||||||
|
application protocol, category, confidence, hostname, encryption/risk, and flow/DPI
|
||||||
|
metadata. They are optional and asynchronous; a failure or late match does not
|
||||||
|
discard underlying capture, and later backfill can enrich stored records.
|
||||||
|
|
||||||
|
The preferred identity is `pid:<packet-id>`, where the ID is bits 0–27 of skb mark.
|
||||||
|
Without it, a SHA-1 `uid` is calculated. The Scapy fallback includes L2–L4 fields,
|
||||||
|
IPv4 ID, ARP/ICMP fields and TCP sequence/ack/flags; eBPF metadata's fallback uses
|
||||||
|
only the smaller L2–L4 tuple and length. Hash-only correlation is consequently a
|
||||||
|
best-effort fallback, weaker for repeated/identical/fragmented traffic.
|
||||||
|
|
||||||
|
## Tracker outcomes and persistence
|
||||||
|
|
||||||
|
The tracker deduplicates observations, merges available fields, retains the earliest
|
||||||
|
timestamp, and waits the configured finalization delay (default 250 ms).
|
||||||
|
|
||||||
|
| Evidence | Verdict | Confidence |
|
||||||
|
| --- | --- | --- |
|
||||||
|
| TC egress telemetry | `accept`; `egress-observed` | high |
|
||||||
|
| TC drop telemetry | mark hint or `drop`; kernel reason / `kfree_skb` | high |
|
||||||
|
| Matching recent TCP RST or ICMP unreachable after drop | `reject` | medium |
|
||||||
|
| Same AF_PACKET bridge record on two ports | `accept`; forwarding observed | medium |
|
||||||
|
| No terminal evidence before delay expires | `unknown`; `timeout` | low |
|
||||||
|
|
||||||
|
It asynchronously upserts batches to PostgreSQL. Entry-cap pressure, persistence
|
||||||
|
failure/retry limits, dirty-age expiry, collector queue loss, socket loss, and
|
||||||
|
shutdown can all cause incompleteness. A packet history or WebSocket feed is never
|
||||||
|
a proof of lossless capture. Batch upserts also do not individually publish packet
|
||||||
|
updates, so realtime consumers must use history reconciliation.
|
||||||
|
|
||||||
|
## Benchmark mode
|
||||||
|
|
||||||
|
Benchmark mode still creates sockets or TC hooks but bypasses normal processing.
|
||||||
|
AF_PACKET increments received frame and byte counters. TC/eBPF increments helper
|
||||||
|
event counters and raw-payload-event counters. It does not parse Scapy, invoke
|
||||||
|
tshark, call the tracker, persist rows, or publish updates. AF_PACKET counters count
|
||||||
|
socket frames; TC counters count emitted sampled events. They are not comparable as
|
||||||
|
equal packet totals without accounting for sampling and multiple event types.
|
||||||
|
|
||||||
|
## Selection guidance
|
||||||
|
|
||||||
|
| Need | Mode | Important caveat |
|
||||||
|
| --- | --- | --- |
|
||||||
|
| Full raw visibility for a single interface | AF_PACKET | No definitive kernel egress/drop verdict. |
|
||||||
|
| Full raw frames across bridge ports | Bridge AF_PACKET | High userspace work; bridge forwarding is inferred. |
|
||||||
|
| Ingress/egress/drop evidence on a controlled bridge | TC/eBPF | Requires BPF/TC privileges and resets clsact. |
|
||||||
|
| Reduced overhead / sampled observability | TC/eBPF sampling | Data is intentionally incomplete. |
|
||||||
|
| Hook-overhead measurement | Benchmark mode | Counts differ between AF_PACKET and TC. |
|
||||||
|
|
||||||
|
Before TC/eBPF capture, inspect `tc qdisc` and filters for every target port,
|
||||||
|
coordinate skb-mark ownership, verify BCC/kernel support, and plan recovery of the
|
||||||
|
TC configuration. For every mode, monitor sniffer debug counters, system logs,
|
||||||
|
capture/process health, DB persistence failures, and expected traffic rate before
|
||||||
|
making operational or security conclusions.
|
||||||
70
documentation/backend/source-reference.md
Normal file
70
documentation/backend/source-reference.md
Normal file
@@ -0,0 +1,70 @@
|
|||||||
|
# Source reference
|
||||||
|
|
||||||
|
This index covers every Python module under `backend/src`, including helper and
|
||||||
|
unmounted-router code. Function names prefixed with `_` are private implementation
|
||||||
|
details; they are described by their owning module's responsibility rather than as
|
||||||
|
separate public contracts.
|
||||||
|
|
||||||
|
## Application and configuration
|
||||||
|
|
||||||
|
| Module | Public surface and role |
|
||||||
|
| --- | --- |
|
||||||
|
| `main.py` | Creates FastAPI, enables permissive CORS, registers startup/shutdown handlers, provides `/hello` and `/versions`, includes live routers, and registers script lifecycle hooks. |
|
||||||
|
| `config.py` | Parses environment strings/integers/floats/booleans; immutable `BackendSettings`; `load_settings`; module-global `settings`. See [configuration](configuration.md). |
|
||||||
|
| `shared_objects.py` | Process-global `db`, `web_loop`, packet broadcaster, and network broadcaster references initialized by `main`. |
|
||||||
|
|
||||||
|
## Models
|
||||||
|
|
||||||
|
| Module | Public surface and role |
|
||||||
|
| --- | --- |
|
||||||
|
| `Models/packets.py` | `PacketObservationModel` and `PacketDBModel`, the normalized persisted/API packet schemas. |
|
||||||
|
| `Models/ip_protocol.py` | `IPProtocolEnum` and `protocol_from_number`, translating IANA protocol numbers to labels. |
|
||||||
|
| `Models/etherType.py` | `EtherTypeEnum` and `ethertype_from_int`, translating Ethernet type values to labels. |
|
||||||
|
| `Models/netplan.py` | `Nameservers`, `EthernetConfig`, `BridgeConfig`, and `NetworkConfig` Pydantic schemas for Netplan-shaped network data. |
|
||||||
|
|
||||||
|
## API routers
|
||||||
|
|
||||||
|
| Module | Public surface and role |
|
||||||
|
| --- | --- |
|
||||||
|
| `api/network_api.py` | Network inspection, bridge create/remove, default reset, link-state watcher control, and network-state WebSocket. Holds shared `IPRoute`/`NDB`; converts netlink messages to Pydantic interface/route/bridge models; publishes state after mutations. |
|
||||||
|
| `api/sniffer_api.py` | Pydantic start/stop/status models and endpoints. Validates one capture target and calls the capture-session API. |
|
||||||
|
| `api/packet_api.py` | Latest packet retrieval, packet-history deletion, and packet-update WebSocket. Serialization handles database records and Pydantic values safely for JSON. |
|
||||||
|
| `api/analysis_api.py` | Pydantic evidence/response models for attachment, protocols, paths, conversations, flow detail, hosts, discovery and anomaly views; delegates each endpoint to `DatabasePool`. |
|
||||||
|
| `api/nft_manager.py` | **Mounted.** `NftManager` wrapper, normalized ruleset models and functions to list rules, delete by handle, and run textual nft. It parses JSON and textual output to enrich rule data. |
|
||||||
|
| `api/packet_scripting_api.py` | **Mounted with doubled prefix.** Name/path validation, example deployment, systemd unit management, venv/pip operations, script status models, and upload/download/enable/disable/delete endpoints. |
|
||||||
|
| `api/nft_api.py` | **Not mounted.** Bridge nftables typed expression model, command generator, handle mapping, and CRUD/preview endpoint functions. `RuleModel.only_bridge` rejects other families. |
|
||||||
|
| `api/nftables_api.py` | **Not mounted.** Generic typed match/action models, resilient binding/CLI wrapper selection, rule reconstruction, and list/replace endpoint functions. |
|
||||||
|
|
||||||
|
## Capture, telemetry, and broadcasting utilities
|
||||||
|
|
||||||
|
| Module | Public surface and role |
|
||||||
|
| --- | --- |
|
||||||
|
| `network_sniffer.py` | Defines flexible `PacketInfo`; parses packet objects/bytes; opens/closes AF_PACKET sockets; owns session reader loops; coordinates telemetry; exposes `start_capture_session`, `stop_capture_session`, status and debug accessors. Legacy `*_afpacket_sniffer` functions delegate to current session functions. |
|
||||||
|
| `utilities/packet_tracker.py` | `PacketTracker` observes capture or telemetry events, aggregates observations, schedules persistence, stops/discards state, and exposes diagnostics. The module-global tracker is the correlation entrypoint. |
|
||||||
|
| `utilities/packet_identity.py` | Builds deterministic fallback packet UID and the minimum fields used to calculate it. |
|
||||||
|
| `utilities/packet_mark.py` | Decodes numeric skb marks into a normalized mark, packet ID, and verdict hint according to the shared mark layout. |
|
||||||
|
| `utilities/tshark_manager.py` | `TsharkManager` owns optional worker processes and caches. Parsing helpers safely coerce nested tshark JSON, extract protocol/HTTP/TLS/DNS/TCP data, derive stream context, and merge enrichment. Module-global `tshark_manager` is used by capture. |
|
||||||
|
| `utilities/bridge_telemetry.py` | `BridgeTelemetryManager` starts/reconciles/stops the eBPF helper and reports subprocess/queue state. Module-global manager is invoked by sniffer lifecycle. |
|
||||||
|
| `utilities/ebpf_bridge_events.py` | Standalone helper program: ctypes event format, BPF-source construction, tc attach/cleanup, perf callbacks, JSON output, signal handling, and `main`. |
|
||||||
|
| `utilities/packet_broadcaster.py` | `PacketBroadcaster` manages subscriber queues. `subscribe`/`unsubscribe`, async `publish`, cross-thread `sync_publish`, and async `close` provide the WebSocket transport primitive. |
|
||||||
|
|
||||||
|
## Network and persistence utilities
|
||||||
|
|
||||||
|
| Module | Public surface and role |
|
||||||
|
| --- | --- |
|
||||||
|
| `utilities/interface_bridge_helpers.py` | Interface existence/up tests; sysfs readers for operational/carrier/admin/MTU state; Ethernet profile retrieval/cache; bridge-port discovery. |
|
||||||
|
| `utilities/bridge_link_state_manager.py` | `EthernetProfile` and `MemberLinkState` data objects; `BridgeLinkStateWatcher` start/stop/status; `BridgeLinkStateManager` enable/disable/query/stop. It embodies debounce, failure, recovery, and profile propagation logic. |
|
||||||
|
| `utilities/database.py` | `DatabasePool` initialization/closure, upsert/batch-upsert, enrichment backfills, latest-packet query, all analysis SQL, and history clearing. Internal helpers normalize values, derive protocol/flow/analysis fields, serialize outgoing rows, and classify discovery activity. |
|
||||||
|
|
||||||
|
## Extension points and maintenance notes
|
||||||
|
|
||||||
|
- New API functionality should live in an `APIRouter`, use Pydantic request and
|
||||||
|
response models, and be explicitly included from `main.py`; otherwise it is not
|
||||||
|
live.
|
||||||
|
- New capture fields must be updated consistently in packet parsing, tracker merge,
|
||||||
|
database upsert SQL, `PacketDBModel`, broadcaster serialization, and analysis
|
||||||
|
queries where relevant.
|
||||||
|
- Any new Linux side effect belongs in [host integration](host-integration.md),
|
||||||
|
including required binary/capability, rollback behavior, and its API exposure.
|
||||||
|
- If an unmounted nft router is adopted, document the migration and remove or
|
||||||
|
version conflicting endpoints instead of silently mounting another implementation.
|
||||||
Reference in New Issue
Block a user