documentation md files
This commit is contained in:
39
documentation/backend/README.md
Normal file
39
documentation/backend/README.md
Normal file
@@ -0,0 +1,39 @@
|
||||
# Backend documentation
|
||||
|
||||
This directory documents the Python service in `backend/src`. It is written for
|
||||
developers and operators of the inline MITM test system. The source code remains
|
||||
the implementation authority; this documentation records the externally useful
|
||||
contracts, lifecycle, Linux integration, and data semantics that are easy to lose
|
||||
when reading individual modules.
|
||||
|
||||
## Reading order
|
||||
|
||||
1. [Architecture](architecture.md) explains the process, responsibilities, and
|
||||
lifecycle.
|
||||
2. [Sniffing modes](sniffing.md) gives the complete technical behavior and
|
||||
implications of AF_PACKET and TC/eBPF capture.
|
||||
3. [Capture pipeline](capture-pipeline.md) follows a packet from observation to
|
||||
persistence and realtime delivery.
|
||||
4. [HTTP and WebSocket API](api.md) lists every router mounted by the application.
|
||||
5. [Data and analysis](data-and-analysis.md) describes the packet record, database
|
||||
operations, and derived analysis views.
|
||||
6. [Host integration](host-integration.md) covers network, eBPF, nftables, tshark,
|
||||
and systemd side effects.
|
||||
7. [Configuration and deployment](configuration.md) records dependencies and all
|
||||
`BACKEND_*` settings.
|
||||
8. [Source reference](source-reference.md) documents every backend source module,
|
||||
including modules not mounted by the current application.
|
||||
|
||||
## Scope and conventions
|
||||
|
||||
All HTTP paths below include the FastAPI `root_path`, `/api`. The interactive
|
||||
schema is available at `/api/docs`, the alternative reference UI at `/api/redoc`,
|
||||
and the machine-readable contract at `/api/openapi.json`.
|
||||
|
||||
"Live" means a router is included by `src.main`. `nft_api.py` and
|
||||
`nftables_api.py` contain independent routers but are not included by the current
|
||||
entrypoint; they are documented as available-but-unmounted implementation paths.
|
||||
|
||||
Packet capture, firewall changes, bridge changes, and script deployment alter the
|
||||
host system. They must be used only in a controlled environment with explicit
|
||||
operator authorization.
|
||||
112
documentation/backend/api.md
Normal file
112
documentation/backend/api.md
Normal file
@@ -0,0 +1,112 @@
|
||||
# HTTP and WebSocket API
|
||||
|
||||
The application is served below `/api`. FastAPI validates request models and
|
||||
publishes the complete JSON Schema at `/api/openapi.json`; use it for exact field
|
||||
types and the current response schema. This page documents semantics and all
|
||||
mounted operations.
|
||||
|
||||
## General endpoints
|
||||
|
||||
| Method/path | Meaning |
|
||||
| --- | --- |
|
||||
| `GET /api/hello` | Simple application health response. |
|
||||
| `GET /api/versions` | Returns the Python runtime version. |
|
||||
|
||||
## Network: `/api/network`
|
||||
|
||||
| Method/path | Parameters/body | Behaviour |
|
||||
| --- | --- | --- |
|
||||
| `GET /interfaces` | none | Lists interfaces, addresses, flags, MTU, MAC, state and Ethernet profile. |
|
||||
| `GET /routes` | none | Lists kernel route entries and resolved output-interface names. |
|
||||
| `GET /links` | none | Lists raw link information. |
|
||||
| `GET /bridges` | none | Lists Linux bridges, STP state and current member details. |
|
||||
| `GET /full-state` | none | Combines interfaces, routes, links and bridges into one snapshot. |
|
||||
| `POST /interfaces/reset-defaults` | `{ interfaces: string[] }` | Resets each requested interface to MTU 1500 and attempts to restore an automatic Ethernet profile through `ethtool`. |
|
||||
| `POST /bridge/create` | `{ name, interfaces }` | Creates a Linux bridge and attaches listed interfaces. |
|
||||
| `POST /bridge/remove` | `{ name }` | Removes an existing bridge. |
|
||||
| `GET /bridge/link-state-watchers` | none | Returns all watcher states. |
|
||||
| `GET /bridge/{bridge_name}/link-state-watcher` | path name | Returns one bridge watcher state. |
|
||||
| `POST /bridge/{bridge_name}/link-state-watcher/enable` | optional recovery holdoff | Enables member failure/recovery propagation. |
|
||||
| `POST /bridge/{bridge_name}/link-state-watcher/disable` | path name | Stops and removes that watcher. |
|
||||
| `WS /ws/state` | none | Receives full network-state update payloads after network mutations. |
|
||||
|
||||
An interface object includes its kernel index, name, state, MAC, MTU, decoded flags,
|
||||
assigned IPv4/IPv6 addresses, and, where available, speed/duplex/autoneg data.
|
||||
|
||||
## Sniffer: `/api/sniffer`
|
||||
|
||||
| Method/path | Parameters/body | Behaviour |
|
||||
| --- | --- | --- |
|
||||
| `POST /start` | exactly one of `bridge` or `interface`; optional `bridge_capture_mode`, `benchmark_mode` | Creates a capture session. Bridge modes are `tc_ebpf` and `af_packet`. |
|
||||
| `POST /stop` | optional session ID or bridge/interface selector | Stops an identified session, target sessions, or all sessions according to the request. |
|
||||
| `GET /status` | none | Returns status keyed by captured interface: running/existing/up state, owner session, mode, and benchmark mode. |
|
||||
| `GET /debug` | none | Returns internal session, buffered-record, tshark, telemetry, and tracker state. Treat as diagnostic output, not a stable client contract. |
|
||||
|
||||
The start endpoint rejects requests containing both a bridge and an interface, or
|
||||
neither. A bridge defaults to `tc_ebpf`; an interface always captures using
|
||||
AF_PACKET.
|
||||
|
||||
## Packets: `/api/packets`
|
||||
|
||||
| Method/path | Parameters/body | Behaviour |
|
||||
| --- | --- | --- |
|
||||
| `GET /packets?limit=100` | `limit` 1–10,000 | Fetches most-recent normalized packet rows. |
|
||||
| `DELETE /packets?reset_id=true` | optional boolean | Clears packet history; can reset database identity state. |
|
||||
| `WS /ws/packets` | none | Receives packet updates from the in-process broadcaster. |
|
||||
|
||||
REST history is authoritative. WebSocket clients must expect connection loss and
|
||||
dropped messages for a slow subscriber, then refill missed state with `GET`.
|
||||
|
||||
## Analysis: `/api/analysis`
|
||||
|
||||
Every analysis endpoint accepts `since_minutes` when shown; its valid range is
|
||||
1 minute to 30 days. Results are derived from the stored packet history and do not
|
||||
claim ground truth about a physical topology or attack.
|
||||
|
||||
| Method/path | Main query controls | Result |
|
||||
| --- | --- | --- |
|
||||
| `GET /interface-hosts` | `since_minutes`, `limit_per_interface` | Likely hosts attached to each MITM-side interface. |
|
||||
| `GET /interface-host-protocols` | plus `limit_protocols_per_host` | Attachment inference with per-host protocol evidence. |
|
||||
| `GET /interface-protocol-paths` | `limit_paths` | Directional aggregated paths for a Sankey-style view. |
|
||||
| `GET /conversations` | `limit` | Aggregated directional endpoint conversations. |
|
||||
| `GET /conversation-flow-detail` | `flow_id` or directional endpoint/port fields; `protocol`, `limit_packets` | Ordered packets, subflows, and derived request/response events. |
|
||||
| `GET /host-intelligence` | `limit_hosts` | Host-centric peers, service and hostname hints. |
|
||||
| `GET /discovery` | `limit` | Discovery, naming and service-advertisement activity. |
|
||||
| `GET /anomalies` | `limit` | Heuristic scan, beacon, rare service, reset-heavy and drop-heavy candidates. |
|
||||
|
||||
`conversation-flow-detail` requires a `flow_id` or enough directional fields to
|
||||
identify a conversation. All analysis endpoints return 503 while the database is
|
||||
unavailable and 500 when their underlying query fails.
|
||||
|
||||
## Firewall: `/api/firewall`
|
||||
|
||||
| Method/path | Body/query | Behaviour |
|
||||
| --- | --- | --- |
|
||||
| `GET /rules` | none | Lists nftables ruleset in a predictable structured representation, enriched with textual rule data where possible. |
|
||||
| `DELETE /rules/{handle}` | optional family/table/chain defaults | Deletes the rule identified by its nft handle. |
|
||||
| `POST /raw` | `{ cmd: string }` | Executes an arbitrary textual nft command and returns stdout/stderr/return code. |
|
||||
|
||||
The raw endpoint is intentionally powerful and must not be exposed to untrusted
|
||||
clients. It changes the host firewall, not an application-local simulation.
|
||||
|
||||
## Scripts: `/api/scripts/scripts`
|
||||
|
||||
The doubled path is produced by the current combination of router and application
|
||||
prefixes. Scripts are Python NFQUEUE workers installed under `/srv/fw-scripts` and
|
||||
can have systemd units and isolated virtual environments.
|
||||
|
||||
| Method/path | Behaviour |
|
||||
| --- | --- |
|
||||
| `GET /` | Lists scripts and their unit mappings/status. |
|
||||
| `POST /` | Uploads a script multipart payload; accepts a script, optional requirements file and required name form field. |
|
||||
| `GET /{name}` | Downloads script source. |
|
||||
| `GET /{name}/requirements` | Downloads its requirements file. |
|
||||
| `PUT /{name}/requirements` | Replaces requirements and runs pip install in the script venv. |
|
||||
| `DELETE /{name}/requirements` | Deletes requirements and removes the venv. |
|
||||
| `POST /{name}/enable` | Creates/starts an NFQUEUE systemd service for a requested queue number. |
|
||||
| `POST /{name}/disable` | Stops/disables the service for a queue number. |
|
||||
| `DELETE /{name}` | Removes all, or one requested queue-number unit, then cleans script-related files as appropriate. |
|
||||
|
||||
Names allow letters, digits, `.`, `_`, and `-`; `.` and `..` are prohibited.
|
||||
Repository example scripts are protected from API modification. Enabling/uploading
|
||||
requirements has code-execution and host-service consequences.
|
||||
70
documentation/backend/architecture.md
Normal file
70
documentation/backend/architecture.md
Normal file
@@ -0,0 +1,70 @@
|
||||
# Backend architecture
|
||||
|
||||
## Process model
|
||||
|
||||
`src.main` constructs one FastAPI application with `root_path="/api"`. During
|
||||
startup it stores the running asyncio loop in `src.shared_objects`, creates an
|
||||
asyncpg `DatabasePool`, attaches a packet broadcaster to it, creates a second
|
||||
network-state broadcaster, and drains any capture records buffered before the DB
|
||||
became available. Shutdown stops capture, network resources, telemetry, tshark,
|
||||
and the packet tracker; then closes WebSocket broadcasters and the DB pool.
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
UI[Frontend/client] --> API[FastAPI /api]
|
||||
API --> NET[Network and bridge API]
|
||||
API --> CAP[Sniffer API]
|
||||
API --> FW[Firewall API]
|
||||
API --> SCR[Script API]
|
||||
CAP --> NS[network_sniffer]
|
||||
NS --> PT[PacketTracker]
|
||||
EBPF[tc/eBPF telemetry process] --> PT
|
||||
NS <--> TS[tshark workers]
|
||||
PT --> DB[(PostgreSQL packets)]
|
||||
DB --> PB[PacketBroadcaster]
|
||||
PB --> WS1[Packet WebSocket]
|
||||
NET --> NB[Network broadcaster]
|
||||
NB --> WS2[Network WebSocket]
|
||||
API --> DB
|
||||
```
|
||||
|
||||
## Component boundaries
|
||||
|
||||
| Component | Responsibility | Persistent state | Important side effects |
|
||||
| --- | --- | --- | --- |
|
||||
| `main.py` | app construction and lifecycle wiring | shared object references | starts/stops resources |
|
||||
| `api/` | validates requests and presents HTTP/WebSocket contracts | none by default | may alter Linux networking, nftables, or services |
|
||||
| `network_sniffer.py` | owns capture sessions and AF_PACKET sockets | in-process session map and pre-DB buffer | raw sockets, reader threads |
|
||||
| `packet_tracker.py` | merges capture and telemetry observations | bounded in-memory pending entries | asynchronous database persistence |
|
||||
| `database.py` | packet upsert/retrieval and SQL analysis | PostgreSQL `packets` table | WebSocket publication after single-row upserts |
|
||||
| `tshark_manager.py` | optional application-protocol enrichment | worker and metadata caches | `tshark` subprocesses/threads |
|
||||
| `bridge_telemetry.py` and `ebpf_bridge_events.py` | bridge tc/eBPF event collection | subprocess state and event queue | compiles/attaches tc programs |
|
||||
| `bridge_link_state_manager.py` | optionally propagates member failure/recovery state | watcher registry | link and Ethernet-profile changes |
|
||||
|
||||
## Shared runtime state
|
||||
|
||||
`shared_objects.py` intentionally holds process-wide references rather than using
|
||||
request-scoped dependency injection:
|
||||
|
||||
- `db`: initialized `DatabasePool`, or `None` after shutdown.
|
||||
- `web_loop`: FastAPI event loop used when worker threads need to schedule work.
|
||||
- `broadcaster`: packet update broadcaster.
|
||||
- `network_broadcaster`: network-state update broadcaster.
|
||||
|
||||
Endpoints that require the database return HTTP 503 when `shared_objects.db` is
|
||||
unavailable. Worker components should tolerate the DB not being ready by buffering
|
||||
or logging failure, rather than assuming the application has fully started.
|
||||
|
||||
## Router mounting
|
||||
|
||||
| Router module | Prefix added by `main.py` | Router-local prefix | Result |
|
||||
| --- | --- | --- | --- |
|
||||
| `network_api` | `/network` | none | `/api/network/...` |
|
||||
| `sniffer_api` | `/sniffer` | none | `/api/sniffer/...` |
|
||||
| `packet_api` | `/packets` | none | `/api/packets/...` |
|
||||
| `analysis_api` | `/analysis` | none | `/api/analysis/...` |
|
||||
| `nft_manager` | none | `/firewall` | `/api/firewall/...` |
|
||||
| `packet_scripting_api` | `/scripts` | `/scripts` | `/api/scripts/scripts/...` |
|
||||
|
||||
The last row reflects the current code exactly. It is worth preserving this fact in
|
||||
examples until the duplicated prefix is deliberately changed.
|
||||
82
documentation/backend/capture-pipeline.md
Normal file
82
documentation/backend/capture-pipeline.md
Normal file
@@ -0,0 +1,82 @@
|
||||
# Packet capture and correlation pipeline
|
||||
|
||||
## Capture modes
|
||||
|
||||
A sniffer session targets exactly one interface or bridge.
|
||||
|
||||
- **Interface target:** an `AF_PACKET` raw socket is opened on that interface;
|
||||
its effective mode is always `af_packet`.
|
||||
- **Bridge target with `af_packet`:** the bridge's member interfaces are captured
|
||||
with raw sockets.
|
||||
- **Bridge target with `tc_ebpf` (default):** no raw socket is opened for bridge
|
||||
ports. `BridgeTelemetryManager` manages an eBPF/tc helper that exports ingress
|
||||
raw data and egress/drop verdict-related events.
|
||||
- **Benchmark mode:** preserves session accounting but skips the normal userspace
|
||||
packet processing path, allowing capture-overhead measurements.
|
||||
|
||||
Sessions have UUIDs and record their label, target type, mode, snapshot of bridge
|
||||
ports, capture interfaces, socket map, thread, and stop event. Stopping by session
|
||||
ID is preferred. A target-specific stop finds matching sessions; an unqualified
|
||||
stop stops every session.
|
||||
|
||||
## End-to-end lifecycle
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
participant C as Capture socket or tc/eBPF
|
||||
participant N as network_sniffer
|
||||
participant T as tshark manager
|
||||
participant P as PacketTracker
|
||||
participant D as DatabasePool
|
||||
participant W as packet WebSocket
|
||||
C->>N: frame / telemetry event
|
||||
N->>N: parse headers, identity, observation metadata
|
||||
N->>T: lookup or schedule enrichment
|
||||
N->>P: capture observation
|
||||
C->>P: ingress/egress/verdict telemetry
|
||||
P->>P: correlate, merge and finalize record
|
||||
P->>D: upsert packet
|
||||
D->>W: publish normalized row
|
||||
```
|
||||
|
||||
`network_sniffer.parse_packet` parses Scapy packet objects, while
|
||||
`parse_packet_bytes` supports raw data. It extracts link, network, and transport
|
||||
fields; adds capture session/observation data; calculates or obtains correlation
|
||||
identifiers; and hands observations to `PacketTracker`. If the shared database or
|
||||
event loop is not yet usable, records are retained in a bounded in-memory buffer;
|
||||
`drain_buffer_to_shared_db` flushes it at application startup.
|
||||
|
||||
## Identity and merging
|
||||
|
||||
`packet_identity.build_packet_uid` makes a stable hash-based fallback identity
|
||||
from normalized packet fields. `packet_mark` decodes the shared skb-mark layout:
|
||||
it normalizes an observed mark, extracts a packet ID, and extracts a verdict hint.
|
||||
Kernel-mark identity is preferred when present; the hash fallback keeps capture and
|
||||
telemetry correlation possible when it is not.
|
||||
|
||||
`PacketTracker` aggregates observations in a bounded dictionary. It deduplicates
|
||||
observations, keeps capture and telemetry provenance, combines ingress/egress and
|
||||
verdict timing, and delays finalization briefly so companion events can arrive.
|
||||
It writes finalized or aged dirty records in batches, retries failed persistence
|
||||
with bounded exponential backoff, discards pending records for stopped interfaces,
|
||||
and exposes a debug snapshot. Its limits and timings are all configured through
|
||||
`BACKEND_PACKET_TRACKER_*` settings.
|
||||
|
||||
## tshark enrichment
|
||||
|
||||
`TsharkManager` starts one long-lived `tshark` process per enabled interface. It
|
||||
reads JSON output in a thread, derives protocol stacks, HTTP/TLS/DNS information,
|
||||
TCP flags, and stream context, then caches matching data for a configurable time
|
||||
window. The capture parser can use a heuristic immediately and the manager can
|
||||
backfill metadata or stream context into already stored rows. tshark is optional in
|
||||
the logical pipeline but enabled by default; a missing executable or worker failure
|
||||
is logged and does not stop capture.
|
||||
|
||||
## Realtime delivery
|
||||
|
||||
`PacketBroadcaster` maintains a bounded asyncio queue per subscriber. A successful
|
||||
single-row database upsert serializes the row and publishes it to subscribers.
|
||||
Slow consumers lose queued messages when their individual queue is full rather than
|
||||
blocking the capture or database path. `/api/packets/ws/packets` is therefore a
|
||||
live-update channel, not a lossless event log; clients should retrieve history over
|
||||
REST and use the WebSocket for incremental updates.
|
||||
73
documentation/backend/configuration.md
Normal file
73
documentation/backend/configuration.md
Normal file
@@ -0,0 +1,73 @@
|
||||
# Configuration and deployment
|
||||
|
||||
## Runtime dependencies
|
||||
|
||||
The service runs with Python 3.11 in the supplied Dockerfile and starts Uvicorn as
|
||||
`src.main:app` on port 8000 with reload enabled. Python dependencies include
|
||||
FastAPI/Pydantic, asyncpg, pyroute2, Scapy, pip-nftables, multipart handling, and
|
||||
WebSocket support. The image installs build tools, libpcap development headers,
|
||||
pkg-config, and `tshark`.
|
||||
|
||||
The host also needs facilities that a minimal application container normally does
|
||||
not have: a reachable PostgreSQL database, Linux network namespace permissions,
|
||||
raw-socket capability, access to `ip`/pyroute2 netlink operations, nftables and
|
||||
appropriate capability, `ethtool` where profile/reset functions are used,
|
||||
systemd/systemctl for scripts, and BCC/eBPF/tc tooling for `tc_ebpf` capture.
|
||||
|
||||
## Environment variables
|
||||
|
||||
All settings are loaded once by `src.config.load_settings`. Empty values use their
|
||||
default. Boolean true values are `1`, `true`, `yes`, or `on` (case-insensitive).
|
||||
|
||||
| Variable | Default | Purpose |
|
||||
| --- | --- | --- |
|
||||
| `BACKEND_DB_DSN` | `postgresql://mitm_user:mitm_password@localhost:5432/mitm_db` | PostgreSQL connection string. |
|
||||
| `BACKEND_LOG_LEVEL` | `DEBUG` | Python logging level. |
|
||||
| `BACKEND_DB_POOL_MIN_SIZE` / `MAX_SIZE` | `1` / `5` | asyncpg pool bounds. |
|
||||
| `BACKEND_BROADCAST_QUEUE_MAXSIZE` | `1024` | Per-WebSocket broadcast queue size. |
|
||||
| `BACKEND_PACKET_TRACKER_FINALIZE_DELAY_SECONDS` | `0.25` | Wait for related observations before finalizing. |
|
||||
| `BACKEND_PACKET_TRACKER_RETENTION_SECONDS` | `10.0` | Pending-entry retention. |
|
||||
| `BACKEND_PACKET_TRACKER_MIN_FLUSH_INTERVAL_SECONDS` | `0.05` | Minimum persistence flush interval. |
|
||||
| `BACKEND_PACKET_TRACKER_PERSIST_TIMEOUT_SECONDS` | `2.0` | One persistence attempt timeout. |
|
||||
| `BACKEND_PACKET_TRACKER_BATCH_PERSIST_TIMEOUT_SECONDS` | `10.0` | Batch persistence timeout. |
|
||||
| `BACKEND_PACKET_TRACKER_PERSIST_RETRY_BACKOFF_SECONDS` / `MAX_SECONDS` | `0.25` / `5.0` | Retry backoff bounds. |
|
||||
| `BACKEND_PACKET_TRACKER_ERROR_LOG_INTERVAL_SECONDS` | `5.0` | Failure-log throttling interval. |
|
||||
| `BACKEND_PACKET_TRACKER_FLUSH_BATCH_SIZE` | `500` | Maximum batch size; clamped to at least 1. |
|
||||
| `BACKEND_PACKET_TRACKER_MAX_ENTRIES` | `50000` | Bounded in-memory correlation capacity; clamped to at least 1. |
|
||||
| `BACKEND_PACKET_TRACKER_MAX_PERSIST_FAILURES` | `3` | Failure threshold; clamped to at least 1. |
|
||||
| `BACKEND_PACKET_TRACKER_MAX_DIRTY_AGE_SECONDS` | `60.0` | Maximum age before dirty data must be flushed. |
|
||||
| `BACKEND_PACKET_TRACKER_STOP_JOIN_TIMEOUT_SECONDS` | `2.0` | Tracker thread join timeout. |
|
||||
| `BACKEND_PACKET_TRACKER_REJECT_CORRELATION_WINDOW_SECONDS` | `1.0` | Rejection-event matching window. |
|
||||
| `BACKEND_SNIFFER_BUFFER_CAPACITY` | `20000` | Pre-DB capture buffer capacity. |
|
||||
| `BACKEND_SNIFFER_SOCKET_RCVBUF_BYTES` | `4194304` | Requested raw-socket receive buffer. |
|
||||
| `BACKEND_SNIFFER_SELECTOR_TIMEOUT_SECONDS` | `1.0` | Reader select timeout. |
|
||||
| `BACKEND_SNIFFER_RECV_BYTES` | `65536` | Maximum raw receive length. |
|
||||
| `BACKEND_SNIFFER_BUFFER_DRAIN_INTERVAL_SECONDS` | `5.0` | Buffered-record drain frequency. |
|
||||
| `BACKEND_SNIFFER_THREAD_JOIN_TIMEOUT_SECONDS` | `2.0` | Session reader join timeout. |
|
||||
| `BACKEND_BRIDGE_BPF_BUILD_DIR` | `/tmp/mitm-bpf` | eBPF build artifacts directory. |
|
||||
| `BACKEND_BRIDGE_TELEMETRY_RAW_SAMPLE_EVERY` / `META_SAMPLE_EVERY` | `1` / `1` | Raw/meta sampling rates; zero is allowed. |
|
||||
| `BACKEND_BRIDGE_TELEMETRY_INGRESS_PERF_PAGES` / `META_PERF_PAGES` | `256` / `128` | eBPF perf-buffer page counts. |
|
||||
| `BACKEND_BRIDGE_TELEMETRY_EVENT_QUEUE_MAXSIZE` | `20000` | Telemetry event queue cap. |
|
||||
| `BACKEND_BRIDGE_TELEMETRY_QUEUE_RECOVERY_SIZE` | `1000` | Queue recovery threshold. |
|
||||
| `BACKEND_BRIDGE_TELEMETRY_DROP_LOG_INTERVAL_SECONDS` | `5.0` | Telemetry-drop log throttling. |
|
||||
| `BACKEND_BRIDGE_LINK_STATE_THREAD_JOIN_TIMEOUT_SECONDS` | `2.0` | Link watcher join timeout. |
|
||||
| `BACKEND_BRIDGE_LINK_STATE_FAILURE_HOLDOFF_SECONDS` / `RECOVERY_HOLDOFF_SECONDS` | `0.75` / `1.0` | Delay before propagating failure/recovery. |
|
||||
| `BACKEND_BRIDGE_LINK_STATE_DEGRADED_RECHECK_SECONDS` | `0.5` | Degraded-link polling period. |
|
||||
| `BACKEND_TELEMETRY_PROCESS_STOP_TIMEOUT_SECONDS` / `READER_JOIN_TIMEOUT_SECONDS` | `3.0` / `2.0` | Telemetry subprocess shutdown limits. |
|
||||
| `BACKEND_TSHARK_ENABLED` | `true` | Enables tshark worker management. |
|
||||
| `BACKEND_TSHARK_DISPLAY_FILTER` | empty | Optional tshark display filter. |
|
||||
| `BACKEND_TSHARK_TRY_HEURISTIC_FIRST` | `true` | Applies local heuristic before tshark match. |
|
||||
| `BACKEND_TSHARK_CACHE_TTL_SECONDS` | `5.0` | Enrichment cache lifetime. |
|
||||
| `BACKEND_TSHARK_MATCH_WINDOW_MS` | `5000` | Capture-to-tshark matching window. |
|
||||
| `BACKEND_TSHARK_READER_JOIN_TIMEOUT_SECONDS` / `PROCESS_STOP_TIMEOUT_SECONDS` | `2.0` / `3.0` | tshark shutdown limits. |
|
||||
|
||||
## Operational safeguards
|
||||
|
||||
Run the API behind an authenticated, access-controlled boundary. The configured
|
||||
CORS policy currently permits every origin, method, and header; it is convenient
|
||||
for development but should not be treated as an authorization control. Keep DB
|
||||
credentials out of version control and use a production-specific DSN.
|
||||
|
||||
Before starting capture, verify target interface/bridge names and ensure recovery
|
||||
access to the host. Before using firewall or script endpoints, snapshot the nft
|
||||
ruleset and understand which systemd units and filesystem paths are in scope.
|
||||
65
documentation/backend/data-and-analysis.md
Normal file
65
documentation/backend/data-and-analysis.md
Normal file
@@ -0,0 +1,65 @@
|
||||
# Packet data, persistence, and analysis
|
||||
|
||||
## Packet record
|
||||
|
||||
`Models/packets.py` defines the normalized `PacketDBModel` returned by packet
|
||||
history APIs. It represents one correlated packet record, not necessarily one raw
|
||||
capture callback. A record may combine several observations.
|
||||
|
||||
| Field group | Fields | Meaning |
|
||||
| --- | --- | --- |
|
||||
| Identity | `id`, `timestamp`, `updated_at`, `correlation_key`, `correlation_source`, `packet_id`, `packet_uid`, `skb_mark` | Database identity and the evidence used to correlate capture/telemetry data. |
|
||||
| Path | `capture_iface`, `ingress_if`, `egress_if`, `capture_session_id`, `capture_sources` | Where and how it was observed. |
|
||||
| Link/network/transport | MACs, EtherType, IP protocol, IPs, ports, VLAN, length | Parsed packet headers. Raw numeric values are retained beside human-readable names. |
|
||||
| Application enrichment | `flow_id`, app protocol/master protocol/category/confidence/hostname/encryption/risk, `dpi_metadata` | tshark-derived context when available. |
|
||||
| Evidence | `raw_b64`, `raw_present`, capture/telemetry metadata and `capture_observations` | Raw bytes and provenance; may be absent. |
|
||||
| Outcome | `verdict`, reason/confidence, ingress/egress/verdict timestamps | Forwarding outcome inferred from telemetry. |
|
||||
|
||||
Each `PacketObservationModel` identifies whether its contribution was `capture` or
|
||||
`telemetry`, the source, interface, timestamp, event type, capture mode, session,
|
||||
and optional reason. Consumers should not assume that every optional field exists:
|
||||
AF_PACKET, tc/eBPF, and enrichment sources provide different evidence.
|
||||
|
||||
## DatabasePool
|
||||
|
||||
`DatabasePool` is an asyncpg wrapper initialized from `BACKEND_DB_DSN`. Startup
|
||||
makes compatibility changes to an existing `packets` table: fills missing
|
||||
timestamps, sets timestamp defaults/non-null constraints, adds
|
||||
`capture_observations` JSONB if needed, and creates an index on `packet_uid`.
|
||||
|
||||
Writes use an upsert keyed by `correlation_key`. `upsert_packet` returns a row and
|
||||
publishes it to the packet broadcaster; `upsert_packets` is a batched performance
|
||||
path and does not individually publish rows. Before writing, it normalizes JSON
|
||||
fields and attaches derived protocol, flow, and analysis fields.
|
||||
|
||||
Read/analysis methods are:
|
||||
|
||||
- `fetch_latest(limit)` for history.
|
||||
- `backfill_packet_metadata` and `backfill_stream_metadata` for late tshark data.
|
||||
- `infer_interface_hosts`, `infer_interface_host_protocols`, and
|
||||
`infer_interface_protocol_paths` for topology/protocol views.
|
||||
- `analyze_conversations` and `fetch_conversation_flow_detail` for directional
|
||||
communication views.
|
||||
- `analyze_host_intelligence`, `analyze_discovery_activity`, and
|
||||
`analyze_anomalies` for investigation aids.
|
||||
- `clear_all_packets(reset_identity)` for destructive history cleanup.
|
||||
|
||||
## Analysis interpretation
|
||||
|
||||
The analysis API runs SQL aggregations over what the system captured. It infers
|
||||
attachment from traffic evidence, groups protocol paths and conversations, and
|
||||
derives host/service/hostname hints. Its anomaly queries rank plausible scan,
|
||||
beacon, rare-service, TCP-reset, and drop-heavy patterns. These are leads for an
|
||||
operator—not assertions of a network's real topology, attribution, or malicious
|
||||
intent. Missing capture events, encrypted traffic, NAT, asymmetric paths, and
|
||||
limits change the output.
|
||||
|
||||
## Enumerations and configuration models
|
||||
|
||||
`Models/etherType.py` provides a string-valued `EtherTypeEnum` and
|
||||
`ethertype_from_int`; `Models/ip_protocol.py` provides `IPProtocolEnum` and
|
||||
`protocol_from_number`. They turn numeric protocol fields into readable labels
|
||||
while keeping raw values. `Models/netplan.py` provides Pydantic schemas for
|
||||
nameservers, Ethernet settings, bridge settings, and a full Netplan-style network
|
||||
configuration. These models are reusable schemas; they are not a substitute for
|
||||
applying a Netplan configuration in the currently mounted API.
|
||||
90
documentation/backend/host-integration.md
Normal file
90
documentation/backend/host-integration.md
Normal file
@@ -0,0 +1,90 @@
|
||||
# Linux host integration and side effects
|
||||
|
||||
## Network and bridge control
|
||||
|
||||
`api/network_api.py` retains process-wide pyroute2 `IPRoute` and `NDB` objects.
|
||||
It reads addresses, link flags, routes and bridge membership through netlink, and
|
||||
uses NDB/pyroute2 to create or remove bridges. Resetting interfaces executes
|
||||
`ethtool` and changes MTU/profile values. These operations affect the host's live
|
||||
connectivity; API errors must be treated as operational failures, not merely input
|
||||
validation failures.
|
||||
|
||||
`utilities/interface_bridge_helpers.py` is the low-level read layer. It checks
|
||||
interface presence/up state, reads sysfs operational/carrier/admin/MTU values,
|
||||
obtains Ethernet profile data using `ethtool`, caches profile data, and reads bridge
|
||||
members from sysfs. It deliberately supplies best-effort information when a driver
|
||||
or platform cannot report every property.
|
||||
|
||||
`bridge_link_state_manager.py` owns optional event-driven bridge watchers. Each
|
||||
watcher tracks Ethernet profile and member readiness, uses failure and recovery
|
||||
holdoffs to avoid flapping, and adjusts selected peer state so an inline bridge
|
||||
reacts coherently to member link loss. `BridgeLinkStateManager` indexes watchers,
|
||||
enables/disables them, reports individual/all status, and stops all during shutdown.
|
||||
|
||||
## eBPF/tc telemetry
|
||||
|
||||
`bridge_telemetry.py` manages the lifecycle of the telemetry helper. Its
|
||||
`update_sessions` method reconciles currently requested bridge ports with the
|
||||
subprocess; `stop` terminates it and `get_debug_snapshot` provides operator
|
||||
diagnostics. It does not itself parse kernel events.
|
||||
|
||||
`ebpf_bridge_events.py` is the helper process. It builds BPF source, attaches tc
|
||||
programs to requested interfaces, reads perf events, and writes JSON-safe event
|
||||
payloads. Events cover ingress raw capture plus egress/drop metadata, including
|
||||
interfaces, MAC/IP information, packet identity, event/reason names, and timing.
|
||||
It cleans existing clsact qdiscs/program attachment as part of setup/cleanup. This
|
||||
requires an appropriate kernel, BCC Python bindings/toolchain, tc, and privileges.
|
||||
|
||||
`tools/ebpf/mark_packet_id.c` is related kernel-side support for packet marking;
|
||||
the mark is decoded by `utilities/packet_mark.py` and used in tracker correlation.
|
||||
|
||||
## nftables
|
||||
|
||||
The mounted `api/nft_manager.py` uses `pip-nftables` to list JSON/text rulesets,
|
||||
normalize them into stable table/chain/rule models, parse rule priorities/text, and
|
||||
delete a rule by handle. Its raw-command endpoint forwards textual nft commands.
|
||||
It therefore needs capability to inspect and change the host nftables ruleset.
|
||||
|
||||
Two alternative implementations exist but are currently unmounted:
|
||||
|
||||
- `api/nft_api.py` is bridge-family oriented. It models meta, Ethernet, IP, port,
|
||||
conntrack, verdict, reject, log, and raw expressions; can generate previews,
|
||||
list rules with authoritative handles, add/delete/update rules, and uses the
|
||||
`nft` CLI.
|
||||
- `api/nftables_api.py` is a stateless typed replacement API. It models matches and
|
||||
actions, chooses a pyroute2 binding when viable or a CLI wrapper otherwise,
|
||||
ensures table/chain presence, reconstructs readable rules, and replaces a chain's
|
||||
ruleset. Its own source warns that a running asyncio loop may force CLI fallback.
|
||||
|
||||
Do not mount more than one firewall router without an explicit API versioning and
|
||||
conflict review: all manipulate shared kernel state and have overlapping concepts.
|
||||
|
||||
## NFQUEUE script services
|
||||
|
||||
`api/packet_scripting_api.py` manages executable Python scripts. It makes these
|
||||
directories at import time: `/srv/fw-scripts`, `/srv/fw-scripts/venvs`, and the
|
||||
repository's `backend/example_scripts`. Scripts are named `<name>.py`; requirements
|
||||
are `<name>-requirements.txt`; virtual environments are per-script. Units use the
|
||||
deterministic name `fw-script-<name>-q<queue>.service` and are written under
|
||||
`/etc/systemd/system`.
|
||||
|
||||
The module discovers services through `systemctl`, writes/parses unit `ExecStart`,
|
||||
runs `daemon-reload`, starts/stops/enables/disables units, creates virtualenvs, and
|
||||
uses pip to install user-provided requirements. Startup can copy protected example
|
||||
scripts and optionally deploy them from `<name>.deploy.json`. This API is a remote
|
||||
code/service-management surface and requires strict authentication plus host-level
|
||||
least privilege.
|
||||
|
||||
## External subprocesses
|
||||
|
||||
| Integration | Commands/facility | Used by |
|
||||
| --- | --- | --- |
|
||||
| tshark | long-lived `tshark` subprocesses | DPI enrichment |
|
||||
| nftables | `nft` CLI and/or pip-nftables bindings | firewall APIs |
|
||||
| Ethernet control | `ethtool` | interface profile/reset |
|
||||
| system services | `systemctl`, virtualenv, pip | script lifecycle |
|
||||
| BPF/tc | BCC, tc, qdisc/program attachment | bridge telemetry |
|
||||
|
||||
Failures are generally logged and translated to endpoint failures or degraded
|
||||
capture. Operators should collect `/api/sniffer/debug`, service logs, nftables
|
||||
state, and interface state when investigating a problem.
|
||||
233
documentation/backend/sniffing.md
Normal file
233
documentation/backend/sniffing.md
Normal file
@@ -0,0 +1,233 @@
|
||||
# Technical reference: packet sniffing modes
|
||||
|
||||
This document specifies the implemented capture behavior in `network_sniffer.py`,
|
||||
`bridge_telemetry.py`, `ebpf_bridge_events.py`, and `packet_tracker.py`. It makes a
|
||||
deliberate distinction between observed facts and inferred forwarding results.
|
||||
|
||||
## Session model and mode selection
|
||||
|
||||
A capture session has a UUID and targets exactly one interface or exactly one
|
||||
bridge. A bridge is expanded once with `get_bridge_ports_once`; its member list is
|
||||
a creation-time snapshot. Later bridge membership changes are not added to the
|
||||
existing session. Session state includes target label/type, effective mode, benchmark
|
||||
flag, port snapshot, AF_PACKET sockets, optional reader thread, stop event, and
|
||||
benchmark counters.
|
||||
|
||||
| Request | Effective mode | Capture source | Path/outcome evidence |
|
||||
| --- | --- | --- | --- |
|
||||
| Interface, any requested mode | `af_packet` | One raw socket on the interface | Packet-socket type labels an outgoing copy as egress; no kernel verdict telemetry. |
|
||||
| Bridge, `af_packet` | `af_packet` | One raw socket per snapshot bridge port | Two matching port observations can infer forwarding. |
|
||||
| Bridge, `tc_ebpf` (default) | `tc_ebpf` | One tc/eBPF helper across snapshot bridge ports | TC ingress/egress and skb-free/drop events, normally matched by skb mark. |
|
||||
| Any mode with benchmark enabled | Same hook/socket setup | Counted but not processed | No parsing, enrichment, DB write, or WebSocket event. |
|
||||
|
||||
Interface targets always use AF_PACKET. The TC/eBPF mode is only selected for a
|
||||
bridge target. A stop by session ID is the safest selector. Stopping a session also
|
||||
discards pending tracker entries relating to its interfaces, so unpersisted data can
|
||||
be lost deliberately at shutdown.
|
||||
|
||||
Multiple sessions may overlap on an interface. This is not an independent-capture
|
||||
guarantee: the TC manager maps an interface to several sessions but assigns raw
|
||||
ingress parsing to the first sorted session ID.
|
||||
|
||||
## AF_PACKET capture
|
||||
|
||||
### Socket behavior
|
||||
|
||||
For each capture interface the service opens `AF_PACKET` / `SOCK_RAW` with protocol
|
||||
`htons(0x0003)` (`ETH_P_ALL`), requests the configured receive buffer (default
|
||||
4 MiB), best-effort requests `TPACKET_V3`, binds to `(ifname, 0)`, and makes the
|
||||
socket non-blocking. Failure to set the buffer or TPACKET version is non-fatal.
|
||||
Failure to create or bind leaves the interface uncaptured; session creation can still
|
||||
complete. This requires raw-socket privilege, commonly `CAP_NET_RAW`.
|
||||
|
||||
One daemon reader thread is started only when the session has sockets. It uses a
|
||||
selector, receives at most `BACKEND_SNIFFER_RECV_BYTES` bytes per event (default
|
||||
65,536), stamps the frame with userspace UTC receive time, creates Scapy `Ether`,
|
||||
and calls the common parser. `ENODEV`, `ENETDOWN`, and `EBADF` close the affected
|
||||
socket; it is not reopened in that session. The thread periodically attempts a
|
||||
pre-DB-buffer drain during selector idle time.
|
||||
|
||||
### Direction and bridge inference
|
||||
|
||||
Packet-socket address metadata is used only as follows: `PACKET_OUTGOING` (normally
|
||||
4) becomes `path_role: egress`; every other packet type becomes `path_role:
|
||||
ingress`. This is a packet-socket perspective, not proof of a Linux bridge decision.
|
||||
|
||||
For a bridge session, the tracker groups AF_PACKET observations by session ID. It
|
||||
uses an explicit ingress observation if available, otherwise the earliest one. It
|
||||
prefers an explicit egress observation on a different port, otherwise a later
|
||||
different-port observation. If one correlated packet is seen on at least two ports,
|
||||
the tracker records:
|
||||
|
||||
```text
|
||||
verdict = accept
|
||||
verdict_reason = bridge-af_packet-forwarded-observed
|
||||
verdict_confidence = medium
|
||||
```
|
||||
|
||||
This means matching evidence was observed on two bridge ports. It does not prove a
|
||||
particular kernel forwarding verdict and can be affected by duplicate copies, loops,
|
||||
or fallback-identity collisions. A single-interface AF_PACKET record has no terminal
|
||||
verdict from AF_PACKET itself.
|
||||
|
||||
### AF_PACKET implications
|
||||
|
||||
AF_PACKET provides full observed frame bytes without BCC or tc changes and is the
|
||||
only interface-capture mode. It neither alters packets nor controls forwarding. It
|
||||
also has no definitive drop visibility, reports userspace rather than kernel event
|
||||
time, can observe local/outgoing copies, and can lose traffic under socket/userspace
|
||||
load. The full-frame Scapy parse, tshark lookup, tracking and persistence path makes
|
||||
it more expensive than sampled telemetry.
|
||||
|
||||
## TC/eBPF bridge capture
|
||||
|
||||
### Collector lifecycle and destructive qdisc behavior
|
||||
|
||||
Bridge sessions in `tc_ebpf` mode are aggregated into one helper process. Any change
|
||||
to the active *interface set* stops the helper and recreates it for the new set;
|
||||
there is a capture gap during that restart. The helper attaches direct-action
|
||||
`BPF.SCHED_CLS` programs at TC ingress (`ffff:fff2`, handle `:20`) and egress
|
||||
(`ffff:fff3`, handle `:30`) to every bridge **member interface**, not the bridge
|
||||
device itself.
|
||||
|
||||
Before attachment the helper runs `tc qdisc del dev <iface> clsact` (ignoring its
|
||||
result), then `tc qdisc add dev <iface> clsact`. It deletes `clsact` again for every
|
||||
instrumented interface at helper shutdown and after an attachment failure.
|
||||
|
||||
> Starting, restarting, failing, or stopping TC/eBPF capture can remove pre-existing
|
||||
> clsact qdiscs and their filters. Do not use it on interfaces with unrelated TC
|
||||
> configuration unless coexistence and recovery are explicitly managed.
|
||||
|
||||
The BCC Python runtime, a compatible kernel, BPF/tracepoint access, TC and netlink
|
||||
privileges are required. Session creation does not wait for a collector health
|
||||
acknowledgement, so a successful start response is not proof that BPF attached.
|
||||
|
||||
### Kernel event generation
|
||||
|
||||
The helper opens a raw-ingress perf buffer and a metadata perf buffer. The ingress
|
||||
TC program creates an skb mark only when it is zero, using the low 28 bits of
|
||||
`bpf_ktime_get_ns()` and replacing zero with one. It preserves any existing nonzero
|
||||
mark. It extracts Ethernet addresses, EtherType, a single 802.1Q/802.1AD VLAN ID,
|
||||
ARP IPv4 addresses, and IPv4/IPv6 addresses with TCP/UDP ports. IPv6 extension
|
||||
headers are not traversed; the base next-header is used as protocol.
|
||||
|
||||
The egress TC program never creates a mark. It exports metadata only for marked
|
||||
packets. The `skb:kfree_skb` tracepoint reads the linear skb representation and
|
||||
exports a drop event only for marked skbs whose device is a selected interface.
|
||||
The emitted payload contains userspace and kernel-monotonic timestamps, interface,
|
||||
mark, length, parsed L2–L4 fields, and event type. Drop events add a numerical
|
||||
reason and `skb_drop_reason_<n>` label. Ingress events may contain `raw_b64`; egress
|
||||
and drop events do not.
|
||||
|
||||
A kfree_skb event is evidence that a marked skb was freed in the kernel context. It
|
||||
is not automatically evidence that nftables caused the outcome; interpret the
|
||||
reason code in the context of kernel behavior and other instrumentation.
|
||||
|
||||
### Sampling
|
||||
|
||||
Raw and metadata sampling are independent settings.
|
||||
|
||||
| Value | Effect |
|
||||
| --- | --- |
|
||||
| `0` | Never emits that sample category. |
|
||||
| `1` | Emits every marked packet in that category. |
|
||||
| `N > 1` | Emits when `skb_mark % N == 0`. |
|
||||
|
||||
At ingress, a raw-selected packet emits a raw event; only a packet not chosen for
|
||||
raw can emit an ingress metadata event. Egress and drop use metadata sampling only.
|
||||
Therefore raw-enabled/meta-disabled capture stores sampled ingress frame records
|
||||
without egress/drop visibility; raw-disabled/meta-enabled capture produces
|
||||
metadata-only rows without raw bytes. Both enabled does not make raw and metadata
|
||||
populations identical.
|
||||
|
||||
Sampling uses the entire existing skb mark. The documented mark layout reserves
|
||||
bits 0–27 for packet ID and upper bits for drop/reject hints. This capture program
|
||||
creates only the low-28-bit value for previously zero marks; it does not set verdict
|
||||
hints. Any other mark-using subsystem must coordinate its mark semantics, because
|
||||
it can change both sampling and correlation.
|
||||
|
||||
### Userspace event handling and loss
|
||||
|
||||
The manager reads JSON helper output into a bounded queue. For a non-benchmark
|
||||
ingress event with `raw_b64`, it decodes the frame and sends it into the common Scapy
|
||||
parser as source `tc_ingress_raw`, with `packet_id`, `skb_mark`, and capture mode
|
||||
`tc_ingress`. It then sends every non-benchmark ingress/egress/drop event to the
|
||||
tracker. A sampled raw ingress packet usually therefore has both a parsed capture
|
||||
observation and a telemetry observation under the same mark-derived key.
|
||||
|
||||
When the telemetry queue is full, the manager drops oldest queued events down to
|
||||
`BACKEND_BRIDGE_TELEMETRY_QUEUE_RECOVERY_SIZE`, attempts to keep the new event, and
|
||||
counts dropped events and raw payloads. Perf buffers can also lose samples before
|
||||
userspace. Neither loss mechanism is recovered. `/api/sniffer/debug` reports queue
|
||||
size, queue drops, benchmark counts, collector interfaces, and tracker statistics.
|
||||
|
||||
### TC/eBPF implications
|
||||
|
||||
This mode yields better within-host correlation and explicit TC egress evidence. A
|
||||
matching egress produces `accept`, `egress-observed`, confidence `high`; a matching
|
||||
drop produces `drop` (or mark hint), confidence `high`. Absence of egress is not
|
||||
proof of a drop: sampling, perf loss, queue loss, an uninstrumented path, teardown,
|
||||
or collector failure can all explain it. Raw bytes are ingress-only and sampled.
|
||||
|
||||
## Common parsing, enrichment, and identity
|
||||
|
||||
Both modes use `parse_packet` / `parse_packet_bytes`. The parser records Ethernet
|
||||
addresses, EtherType and VLAN, ARP operation/addressing, IPv4 ID or IPv6 base
|
||||
header, TCP sequence/acknowledgement/flags, UDP ports, ICMP/ICMPv6 type/code, and
|
||||
an embedded IPv4 tuple from eligible ICMP errors. It stores full raw frame bytes
|
||||
when supplied by AF_PACKET or sampled TC ingress.
|
||||
|
||||
tshark workers are enabled for non-benchmark capture interfaces. They may add
|
||||
application protocol, category, confidence, hostname, encryption/risk, and flow/DPI
|
||||
metadata. They are optional and asynchronous; a failure or late match does not
|
||||
discard underlying capture, and later backfill can enrich stored records.
|
||||
|
||||
The preferred identity is `pid:<packet-id>`, where the ID is bits 0–27 of skb mark.
|
||||
Without it, a SHA-1 `uid` is calculated. The Scapy fallback includes L2–L4 fields,
|
||||
IPv4 ID, ARP/ICMP fields and TCP sequence/ack/flags; eBPF metadata's fallback uses
|
||||
only the smaller L2–L4 tuple and length. Hash-only correlation is consequently a
|
||||
best-effort fallback, weaker for repeated/identical/fragmented traffic.
|
||||
|
||||
## Tracker outcomes and persistence
|
||||
|
||||
The tracker deduplicates observations, merges available fields, retains the earliest
|
||||
timestamp, and waits the configured finalization delay (default 250 ms).
|
||||
|
||||
| Evidence | Verdict | Confidence |
|
||||
| --- | --- | --- |
|
||||
| TC egress telemetry | `accept`; `egress-observed` | high |
|
||||
| TC drop telemetry | mark hint or `drop`; kernel reason / `kfree_skb` | high |
|
||||
| Matching recent TCP RST or ICMP unreachable after drop | `reject` | medium |
|
||||
| Same AF_PACKET bridge record on two ports | `accept`; forwarding observed | medium |
|
||||
| No terminal evidence before delay expires | `unknown`; `timeout` | low |
|
||||
|
||||
It asynchronously upserts batches to PostgreSQL. Entry-cap pressure, persistence
|
||||
failure/retry limits, dirty-age expiry, collector queue loss, socket loss, and
|
||||
shutdown can all cause incompleteness. A packet history or WebSocket feed is never
|
||||
a proof of lossless capture. Batch upserts also do not individually publish packet
|
||||
updates, so realtime consumers must use history reconciliation.
|
||||
|
||||
## Benchmark mode
|
||||
|
||||
Benchmark mode still creates sockets or TC hooks but bypasses normal processing.
|
||||
AF_PACKET increments received frame and byte counters. TC/eBPF increments helper
|
||||
event counters and raw-payload-event counters. It does not parse Scapy, invoke
|
||||
tshark, call the tracker, persist rows, or publish updates. AF_PACKET counters count
|
||||
socket frames; TC counters count emitted sampled events. They are not comparable as
|
||||
equal packet totals without accounting for sampling and multiple event types.
|
||||
|
||||
## Selection guidance
|
||||
|
||||
| Need | Mode | Important caveat |
|
||||
| --- | --- | --- |
|
||||
| Full raw visibility for a single interface | AF_PACKET | No definitive kernel egress/drop verdict. |
|
||||
| Full raw frames across bridge ports | Bridge AF_PACKET | High userspace work; bridge forwarding is inferred. |
|
||||
| Ingress/egress/drop evidence on a controlled bridge | TC/eBPF | Requires BPF/TC privileges and resets clsact. |
|
||||
| Reduced overhead / sampled observability | TC/eBPF sampling | Data is intentionally incomplete. |
|
||||
| Hook-overhead measurement | Benchmark mode | Counts differ between AF_PACKET and TC. |
|
||||
|
||||
Before TC/eBPF capture, inspect `tc qdisc` and filters for every target port,
|
||||
coordinate skb-mark ownership, verify BCC/kernel support, and plan recovery of the
|
||||
TC configuration. For every mode, monitor sniffer debug counters, system logs,
|
||||
capture/process health, DB persistence failures, and expected traffic rate before
|
||||
making operational or security conclusions.
|
||||
70
documentation/backend/source-reference.md
Normal file
70
documentation/backend/source-reference.md
Normal file
@@ -0,0 +1,70 @@
|
||||
# Source reference
|
||||
|
||||
This index covers every Python module under `backend/src`, including helper and
|
||||
unmounted-router code. Function names prefixed with `_` are private implementation
|
||||
details; they are described by their owning module's responsibility rather than as
|
||||
separate public contracts.
|
||||
|
||||
## Application and configuration
|
||||
|
||||
| Module | Public surface and role |
|
||||
| --- | --- |
|
||||
| `main.py` | Creates FastAPI, enables permissive CORS, registers startup/shutdown handlers, provides `/hello` and `/versions`, includes live routers, and registers script lifecycle hooks. |
|
||||
| `config.py` | Parses environment strings/integers/floats/booleans; immutable `BackendSettings`; `load_settings`; module-global `settings`. See [configuration](configuration.md). |
|
||||
| `shared_objects.py` | Process-global `db`, `web_loop`, packet broadcaster, and network broadcaster references initialized by `main`. |
|
||||
|
||||
## Models
|
||||
|
||||
| Module | Public surface and role |
|
||||
| --- | --- |
|
||||
| `Models/packets.py` | `PacketObservationModel` and `PacketDBModel`, the normalized persisted/API packet schemas. |
|
||||
| `Models/ip_protocol.py` | `IPProtocolEnum` and `protocol_from_number`, translating IANA protocol numbers to labels. |
|
||||
| `Models/etherType.py` | `EtherTypeEnum` and `ethertype_from_int`, translating Ethernet type values to labels. |
|
||||
| `Models/netplan.py` | `Nameservers`, `EthernetConfig`, `BridgeConfig`, and `NetworkConfig` Pydantic schemas for Netplan-shaped network data. |
|
||||
|
||||
## API routers
|
||||
|
||||
| Module | Public surface and role |
|
||||
| --- | --- |
|
||||
| `api/network_api.py` | Network inspection, bridge create/remove, default reset, link-state watcher control, and network-state WebSocket. Holds shared `IPRoute`/`NDB`; converts netlink messages to Pydantic interface/route/bridge models; publishes state after mutations. |
|
||||
| `api/sniffer_api.py` | Pydantic start/stop/status models and endpoints. Validates one capture target and calls the capture-session API. |
|
||||
| `api/packet_api.py` | Latest packet retrieval, packet-history deletion, and packet-update WebSocket. Serialization handles database records and Pydantic values safely for JSON. |
|
||||
| `api/analysis_api.py` | Pydantic evidence/response models for attachment, protocols, paths, conversations, flow detail, hosts, discovery and anomaly views; delegates each endpoint to `DatabasePool`. |
|
||||
| `api/nft_manager.py` | **Mounted.** `NftManager` wrapper, normalized ruleset models and functions to list rules, delete by handle, and run textual nft. It parses JSON and textual output to enrich rule data. |
|
||||
| `api/packet_scripting_api.py` | **Mounted with doubled prefix.** Name/path validation, example deployment, systemd unit management, venv/pip operations, script status models, and upload/download/enable/disable/delete endpoints. |
|
||||
| `api/nft_api.py` | **Not mounted.** Bridge nftables typed expression model, command generator, handle mapping, and CRUD/preview endpoint functions. `RuleModel.only_bridge` rejects other families. |
|
||||
| `api/nftables_api.py` | **Not mounted.** Generic typed match/action models, resilient binding/CLI wrapper selection, rule reconstruction, and list/replace endpoint functions. |
|
||||
|
||||
## Capture, telemetry, and broadcasting utilities
|
||||
|
||||
| Module | Public surface and role |
|
||||
| --- | --- |
|
||||
| `network_sniffer.py` | Defines flexible `PacketInfo`; parses packet objects/bytes; opens/closes AF_PACKET sockets; owns session reader loops; coordinates telemetry; exposes `start_capture_session`, `stop_capture_session`, status and debug accessors. Legacy `*_afpacket_sniffer` functions delegate to current session functions. |
|
||||
| `utilities/packet_tracker.py` | `PacketTracker` observes capture or telemetry events, aggregates observations, schedules persistence, stops/discards state, and exposes diagnostics. The module-global tracker is the correlation entrypoint. |
|
||||
| `utilities/packet_identity.py` | Builds deterministic fallback packet UID and the minimum fields used to calculate it. |
|
||||
| `utilities/packet_mark.py` | Decodes numeric skb marks into a normalized mark, packet ID, and verdict hint according to the shared mark layout. |
|
||||
| `utilities/tshark_manager.py` | `TsharkManager` owns optional worker processes and caches. Parsing helpers safely coerce nested tshark JSON, extract protocol/HTTP/TLS/DNS/TCP data, derive stream context, and merge enrichment. Module-global `tshark_manager` is used by capture. |
|
||||
| `utilities/bridge_telemetry.py` | `BridgeTelemetryManager` starts/reconciles/stops the eBPF helper and reports subprocess/queue state. Module-global manager is invoked by sniffer lifecycle. |
|
||||
| `utilities/ebpf_bridge_events.py` | Standalone helper program: ctypes event format, BPF-source construction, tc attach/cleanup, perf callbacks, JSON output, signal handling, and `main`. |
|
||||
| `utilities/packet_broadcaster.py` | `PacketBroadcaster` manages subscriber queues. `subscribe`/`unsubscribe`, async `publish`, cross-thread `sync_publish`, and async `close` provide the WebSocket transport primitive. |
|
||||
|
||||
## Network and persistence utilities
|
||||
|
||||
| Module | Public surface and role |
|
||||
| --- | --- |
|
||||
| `utilities/interface_bridge_helpers.py` | Interface existence/up tests; sysfs readers for operational/carrier/admin/MTU state; Ethernet profile retrieval/cache; bridge-port discovery. |
|
||||
| `utilities/bridge_link_state_manager.py` | `EthernetProfile` and `MemberLinkState` data objects; `BridgeLinkStateWatcher` start/stop/status; `BridgeLinkStateManager` enable/disable/query/stop. It embodies debounce, failure, recovery, and profile propagation logic. |
|
||||
| `utilities/database.py` | `DatabasePool` initialization/closure, upsert/batch-upsert, enrichment backfills, latest-packet query, all analysis SQL, and history clearing. Internal helpers normalize values, derive protocol/flow/analysis fields, serialize outgoing rows, and classify discovery activity. |
|
||||
|
||||
## Extension points and maintenance notes
|
||||
|
||||
- New API functionality should live in an `APIRouter`, use Pydantic request and
|
||||
response models, and be explicitly included from `main.py`; otherwise it is not
|
||||
live.
|
||||
- New capture fields must be updated consistently in packet parsing, tracker merge,
|
||||
database upsert SQL, `PacketDBModel`, broadcaster serialization, and analysis
|
||||
queries where relevant.
|
||||
- Any new Linux side effect belongs in [host integration](host-integration.md),
|
||||
including required binary/capability, rollback behavior, and its API exposure.
|
||||
- If an unmounted nft router is adopted, document the migration and remove or
|
||||
version conflicting endpoints instead of silently mounting another implementation.
|
||||
Reference in New Issue
Block a user