documentation md files
Some checks failed
Build and Deploy MITM Webserver / build (push) Has been cancelled
Build and Deploy MITM Webserver / traffic_target (push) Has been cancelled

This commit is contained in:
Marcus Almert
2026-08-30 17:31:46 +02:00
parent 68827ed7e3
commit 9106cac2a2
9 changed files with 834 additions and 0 deletions

View File

@@ -0,0 +1,39 @@
# Backend documentation
This directory documents the Python service in `backend/src`. It is written for
developers and operators of the inline MITM test system. The source code remains
the implementation authority; this documentation records the externally useful
contracts, lifecycle, Linux integration, and data semantics that are easy to lose
when reading individual modules.
## Reading order
1. [Architecture](architecture.md) explains the process, responsibilities, and
lifecycle.
2. [Sniffing modes](sniffing.md) gives the complete technical behavior and
implications of AF_PACKET and TC/eBPF capture.
3. [Capture pipeline](capture-pipeline.md) follows a packet from observation to
persistence and realtime delivery.
4. [HTTP and WebSocket API](api.md) lists every router mounted by the application.
5. [Data and analysis](data-and-analysis.md) describes the packet record, database
operations, and derived analysis views.
6. [Host integration](host-integration.md) covers network, eBPF, nftables, tshark,
and systemd side effects.
7. [Configuration and deployment](configuration.md) records dependencies and all
`BACKEND_*` settings.
8. [Source reference](source-reference.md) documents every backend source module,
including modules not mounted by the current application.
## Scope and conventions
All HTTP paths below include the FastAPI `root_path`, `/api`. The interactive
schema is available at `/api/docs`, the alternative reference UI at `/api/redoc`,
and the machine-readable contract at `/api/openapi.json`.
"Live" means a router is included by `src.main`. `nft_api.py` and
`nftables_api.py` contain independent routers but are not included by the current
entrypoint; they are documented as available-but-unmounted implementation paths.
Packet capture, firewall changes, bridge changes, and script deployment alter the
host system. They must be used only in a controlled environment with explicit
operator authorization.

View File

@@ -0,0 +1,112 @@
# HTTP and WebSocket API
The application is served below `/api`. FastAPI validates request models and
publishes the complete JSON Schema at `/api/openapi.json`; use it for exact field
types and the current response schema. This page documents semantics and all
mounted operations.
## General endpoints
| Method/path | Meaning |
| --- | --- |
| `GET /api/hello` | Simple application health response. |
| `GET /api/versions` | Returns the Python runtime version. |
## Network: `/api/network`
| Method/path | Parameters/body | Behaviour |
| --- | --- | --- |
| `GET /interfaces` | none | Lists interfaces, addresses, flags, MTU, MAC, state and Ethernet profile. |
| `GET /routes` | none | Lists kernel route entries and resolved output-interface names. |
| `GET /links` | none | Lists raw link information. |
| `GET /bridges` | none | Lists Linux bridges, STP state and current member details. |
| `GET /full-state` | none | Combines interfaces, routes, links and bridges into one snapshot. |
| `POST /interfaces/reset-defaults` | `{ interfaces: string[] }` | Resets each requested interface to MTU 1500 and attempts to restore an automatic Ethernet profile through `ethtool`. |
| `POST /bridge/create` | `{ name, interfaces }` | Creates a Linux bridge and attaches listed interfaces. |
| `POST /bridge/remove` | `{ name }` | Removes an existing bridge. |
| `GET /bridge/link-state-watchers` | none | Returns all watcher states. |
| `GET /bridge/{bridge_name}/link-state-watcher` | path name | Returns one bridge watcher state. |
| `POST /bridge/{bridge_name}/link-state-watcher/enable` | optional recovery holdoff | Enables member failure/recovery propagation. |
| `POST /bridge/{bridge_name}/link-state-watcher/disable` | path name | Stops and removes that watcher. |
| `WS /ws/state` | none | Receives full network-state update payloads after network mutations. |
An interface object includes its kernel index, name, state, MAC, MTU, decoded flags,
assigned IPv4/IPv6 addresses, and, where available, speed/duplex/autoneg data.
## Sniffer: `/api/sniffer`
| Method/path | Parameters/body | Behaviour |
| --- | --- | --- |
| `POST /start` | exactly one of `bridge` or `interface`; optional `bridge_capture_mode`, `benchmark_mode` | Creates a capture session. Bridge modes are `tc_ebpf` and `af_packet`. |
| `POST /stop` | optional session ID or bridge/interface selector | Stops an identified session, target sessions, or all sessions according to the request. |
| `GET /status` | none | Returns status keyed by captured interface: running/existing/up state, owner session, mode, and benchmark mode. |
| `GET /debug` | none | Returns internal session, buffered-record, tshark, telemetry, and tracker state. Treat as diagnostic output, not a stable client contract. |
The start endpoint rejects requests containing both a bridge and an interface, or
neither. A bridge defaults to `tc_ebpf`; an interface always captures using
AF_PACKET.
## Packets: `/api/packets`
| Method/path | Parameters/body | Behaviour |
| --- | --- | --- |
| `GET /packets?limit=100` | `limit` 1–10,000 | Fetches most-recent normalized packet rows. |
| `DELETE /packets?reset_id=true` | optional boolean | Clears packet history; can reset database identity state. |
| `WS /ws/packets` | none | Receives packet updates from the in-process broadcaster. |
REST history is authoritative. WebSocket clients must expect connection loss and
dropped messages for a slow subscriber, then refill missed state with `GET`.
## Analysis: `/api/analysis`
Every analysis endpoint accepts `since_minutes` when shown; its valid range is
1 minute to 30 days. Results are derived from the stored packet history and do not
claim ground truth about a physical topology or attack.
| Method/path | Main query controls | Result |
| --- | --- | --- |
| `GET /interface-hosts` | `since_minutes`, `limit_per_interface` | Likely hosts attached to each MITM-side interface. |
| `GET /interface-host-protocols` | plus `limit_protocols_per_host` | Attachment inference with per-host protocol evidence. |
| `GET /interface-protocol-paths` | `limit_paths` | Directional aggregated paths for a Sankey-style view. |
| `GET /conversations` | `limit` | Aggregated directional endpoint conversations. |
| `GET /conversation-flow-detail` | `flow_id` or directional endpoint/port fields; `protocol`, `limit_packets` | Ordered packets, subflows, and derived request/response events. |
| `GET /host-intelligence` | `limit_hosts` | Host-centric peers, service and hostname hints. |
| `GET /discovery` | `limit` | Discovery, naming and service-advertisement activity. |
| `GET /anomalies` | `limit` | Heuristic scan, beacon, rare service, reset-heavy and drop-heavy candidates. |
`conversation-flow-detail` requires a `flow_id` or enough directional fields to
identify a conversation. All analysis endpoints return 503 while the database is
unavailable and 500 when their underlying query fails.
## Firewall: `/api/firewall`
| Method/path | Body/query | Behaviour |
| --- | --- | --- |
| `GET /rules` | none | Lists nftables ruleset in a predictable structured representation, enriched with textual rule data where possible. |
| `DELETE /rules/{handle}` | optional family/table/chain defaults | Deletes the rule identified by its nft handle. |
| `POST /raw` | `{ cmd: string }` | Executes an arbitrary textual nft command and returns stdout/stderr/return code. |
The raw endpoint is intentionally powerful and must not be exposed to untrusted
clients. It changes the host firewall, not an application-local simulation.
## Scripts: `/api/scripts/scripts`
The doubled path is produced by the current combination of router and application
prefixes. Scripts are Python NFQUEUE workers installed under `/srv/fw-scripts` and
can have systemd units and isolated virtual environments.
| Method/path | Behaviour |
| --- | --- |
| `GET /` | Lists scripts and their unit mappings/status. |
| `POST /` | Uploads a script multipart payload; accepts a script, optional requirements file and required name form field. |
| `GET /{name}` | Downloads script source. |
| `GET /{name}/requirements` | Downloads its requirements file. |
| `PUT /{name}/requirements` | Replaces requirements and runs pip install in the script venv. |
| `DELETE /{name}/requirements` | Deletes requirements and removes the venv. |
| `POST /{name}/enable` | Creates/starts an NFQUEUE systemd service for a requested queue number. |
| `POST /{name}/disable` | Stops/disables the service for a queue number. |
| `DELETE /{name}` | Removes all, or one requested queue-number unit, then cleans script-related files as appropriate. |
Names allow letters, digits, `.`, `_`, and `-`; `.` and `..` are prohibited.
Repository example scripts are protected from API modification. Enabling/uploading
requirements has code-execution and host-service consequences.

View File

@@ -0,0 +1,70 @@
# Backend architecture
## Process model
`src.main` constructs one FastAPI application with `root_path="/api"`. During
startup it stores the running asyncio loop in `src.shared_objects`, creates an
asyncpg `DatabasePool`, attaches a packet broadcaster to it, creates a second
network-state broadcaster, and drains any capture records buffered before the DB
became available. Shutdown stops capture, network resources, telemetry, tshark,
and the packet tracker; then closes WebSocket broadcasters and the DB pool.
```mermaid
flowchart LR
UI[Frontend/client] --> API[FastAPI /api]
API --> NET[Network and bridge API]
API --> CAP[Sniffer API]
API --> FW[Firewall API]
API --> SCR[Script API]
CAP --> NS[network_sniffer]
NS --> PT[PacketTracker]
EBPF[tc/eBPF telemetry process] --> PT
NS <--> TS[tshark workers]
PT --> DB[(PostgreSQL packets)]
DB --> PB[PacketBroadcaster]
PB --> WS1[Packet WebSocket]
NET --> NB[Network broadcaster]
NB --> WS2[Network WebSocket]
API --> DB
```
## Component boundaries
| Component | Responsibility | Persistent state | Important side effects |
| --- | --- | --- | --- |
| `main.py` | app construction and lifecycle wiring | shared object references | starts/stops resources |
| `api/` | validates requests and presents HTTP/WebSocket contracts | none by default | may alter Linux networking, nftables, or services |
| `network_sniffer.py` | owns capture sessions and AF_PACKET sockets | in-process session map and pre-DB buffer | raw sockets, reader threads |
| `packet_tracker.py` | merges capture and telemetry observations | bounded in-memory pending entries | asynchronous database persistence |
| `database.py` | packet upsert/retrieval and SQL analysis | PostgreSQL `packets` table | WebSocket publication after single-row upserts |
| `tshark_manager.py` | optional application-protocol enrichment | worker and metadata caches | `tshark` subprocesses/threads |
| `bridge_telemetry.py` and `ebpf_bridge_events.py` | bridge tc/eBPF event collection | subprocess state and event queue | compiles/attaches tc programs |
| `bridge_link_state_manager.py` | optionally propagates member failure/recovery state | watcher registry | link and Ethernet-profile changes |
## Shared runtime state
`shared_objects.py` intentionally holds process-wide references rather than using
request-scoped dependency injection:
- `db`: initialized `DatabasePool`, or `None` after shutdown.
- `web_loop`: FastAPI event loop used when worker threads need to schedule work.
- `broadcaster`: packet update broadcaster.
- `network_broadcaster`: network-state update broadcaster.
Endpoints that require the database return HTTP 503 when `shared_objects.db` is
unavailable. Worker components should tolerate the DB not being ready by buffering
or logging failure, rather than assuming the application has fully started.
## Router mounting
| Router module | Prefix added by `main.py` | Router-local prefix | Result |
| --- | --- | --- | --- |
| `network_api` | `/network` | none | `/api/network/...` |
| `sniffer_api` | `/sniffer` | none | `/api/sniffer/...` |
| `packet_api` | `/packets` | none | `/api/packets/...` |
| `analysis_api` | `/analysis` | none | `/api/analysis/...` |
| `nft_manager` | none | `/firewall` | `/api/firewall/...` |
| `packet_scripting_api` | `/scripts` | `/scripts` | `/api/scripts/scripts/...` |
The last row reflects the current code exactly. It is worth preserving this fact in
examples until the duplicated prefix is deliberately changed.

View File

@@ -0,0 +1,82 @@
# Packet capture and correlation pipeline
## Capture modes
A sniffer session targets exactly one interface or bridge.
- **Interface target:** an `AF_PACKET` raw socket is opened on that interface;
its effective mode is always `af_packet`.
- **Bridge target with `af_packet`:** the bridge's member interfaces are captured
with raw sockets.
- **Bridge target with `tc_ebpf` (default):** no raw socket is opened for bridge
ports. `BridgeTelemetryManager` manages an eBPF/tc helper that exports ingress
raw data and egress/drop verdict-related events.
- **Benchmark mode:** preserves session accounting but skips the normal userspace
packet processing path, allowing capture-overhead measurements.
Sessions have UUIDs and record their label, target type, mode, snapshot of bridge
ports, capture interfaces, socket map, thread, and stop event. Stopping by session
ID is preferred. A target-specific stop finds matching sessions; an unqualified
stop stops every session.
## End-to-end lifecycle
```mermaid
sequenceDiagram
participant C as Capture socket or tc/eBPF
participant N as network_sniffer
participant T as tshark manager
participant P as PacketTracker
participant D as DatabasePool
participant W as packet WebSocket
C->>N: frame / telemetry event
N->>N: parse headers, identity, observation metadata
N->>T: lookup or schedule enrichment
N->>P: capture observation
C->>P: ingress/egress/verdict telemetry
P->>P: correlate, merge and finalize record
P->>D: upsert packet
D->>W: publish normalized row
```
`network_sniffer.parse_packet` parses Scapy packet objects, while
`parse_packet_bytes` supports raw data. It extracts link, network, and transport
fields; adds capture session/observation data; calculates or obtains correlation
identifiers; and hands observations to `PacketTracker`. If the shared database or
event loop is not yet usable, records are retained in a bounded in-memory buffer;
`drain_buffer_to_shared_db` flushes it at application startup.
## Identity and merging
`packet_identity.build_packet_uid` makes a stable hash-based fallback identity
from normalized packet fields. `packet_mark` decodes the shared skb-mark layout:
it normalizes an observed mark, extracts a packet ID, and extracts a verdict hint.
Kernel-mark identity is preferred when present; the hash fallback keeps capture and
telemetry correlation possible when it is not.
`PacketTracker` aggregates observations in a bounded dictionary. It deduplicates
observations, keeps capture and telemetry provenance, combines ingress/egress and
verdict timing, and delays finalization briefly so companion events can arrive.
It writes finalized or aged dirty records in batches, retries failed persistence
with bounded exponential backoff, discards pending records for stopped interfaces,
and exposes a debug snapshot. Its limits and timings are all configured through
`BACKEND_PACKET_TRACKER_*` settings.
## tshark enrichment
`TsharkManager` starts one long-lived `tshark` process per enabled interface. It
reads JSON output in a thread, derives protocol stacks, HTTP/TLS/DNS information,
TCP flags, and stream context, then caches matching data for a configurable time
window. The capture parser can use a heuristic immediately and the manager can
backfill metadata or stream context into already stored rows. tshark is optional in
the logical pipeline but enabled by default; a missing executable or worker failure
is logged and does not stop capture.
## Realtime delivery
`PacketBroadcaster` maintains a bounded asyncio queue per subscriber. A successful
single-row database upsert serializes the row and publishes it to subscribers.
Slow consumers lose queued messages when their individual queue is full rather than
blocking the capture or database path. `/api/packets/ws/packets` is therefore a
live-update channel, not a lossless event log; clients should retrieve history over
REST and use the WebSocket for incremental updates.

View File

@@ -0,0 +1,73 @@
# Configuration and deployment
## Runtime dependencies
The service runs with Python 3.11 in the supplied Dockerfile and starts Uvicorn as
`src.main:app` on port 8000 with reload enabled. Python dependencies include
FastAPI/Pydantic, asyncpg, pyroute2, Scapy, pip-nftables, multipart handling, and
WebSocket support. The image installs build tools, libpcap development headers,
pkg-config, and `tshark`.
The host also needs facilities that a minimal application container normally does
not have: a reachable PostgreSQL database, Linux network namespace permissions,
raw-socket capability, access to `ip`/pyroute2 netlink operations, nftables and
appropriate capability, `ethtool` where profile/reset functions are used,
systemd/systemctl for scripts, and BCC/eBPF/tc tooling for `tc_ebpf` capture.
## Environment variables
All settings are loaded once by `src.config.load_settings`. Empty values use their
default. Boolean true values are `1`, `true`, `yes`, or `on` (case-insensitive).
| Variable | Default | Purpose |
| --- | --- | --- |
| `BACKEND_DB_DSN` | `postgresql://mitm_user:mitm_password@localhost:5432/mitm_db` | PostgreSQL connection string. |
| `BACKEND_LOG_LEVEL` | `DEBUG` | Python logging level. |
| `BACKEND_DB_POOL_MIN_SIZE` / `MAX_SIZE` | `1` / `5` | asyncpg pool bounds. |
| `BACKEND_BROADCAST_QUEUE_MAXSIZE` | `1024` | Per-WebSocket broadcast queue size. |
| `BACKEND_PACKET_TRACKER_FINALIZE_DELAY_SECONDS` | `0.25` | Wait for related observations before finalizing. |
| `BACKEND_PACKET_TRACKER_RETENTION_SECONDS` | `10.0` | Pending-entry retention. |
| `BACKEND_PACKET_TRACKER_MIN_FLUSH_INTERVAL_SECONDS` | `0.05` | Minimum persistence flush interval. |
| `BACKEND_PACKET_TRACKER_PERSIST_TIMEOUT_SECONDS` | `2.0` | One persistence attempt timeout. |
| `BACKEND_PACKET_TRACKER_BATCH_PERSIST_TIMEOUT_SECONDS` | `10.0` | Batch persistence timeout. |
| `BACKEND_PACKET_TRACKER_PERSIST_RETRY_BACKOFF_SECONDS` / `MAX_SECONDS` | `0.25` / `5.0` | Retry backoff bounds. |
| `BACKEND_PACKET_TRACKER_ERROR_LOG_INTERVAL_SECONDS` | `5.0` | Failure-log throttling interval. |
| `BACKEND_PACKET_TRACKER_FLUSH_BATCH_SIZE` | `500` | Maximum batch size; clamped to at least 1. |
| `BACKEND_PACKET_TRACKER_MAX_ENTRIES` | `50000` | Bounded in-memory correlation capacity; clamped to at least 1. |
| `BACKEND_PACKET_TRACKER_MAX_PERSIST_FAILURES` | `3` | Failure threshold; clamped to at least 1. |
| `BACKEND_PACKET_TRACKER_MAX_DIRTY_AGE_SECONDS` | `60.0` | Maximum age before dirty data must be flushed. |
| `BACKEND_PACKET_TRACKER_STOP_JOIN_TIMEOUT_SECONDS` | `2.0` | Tracker thread join timeout. |
| `BACKEND_PACKET_TRACKER_REJECT_CORRELATION_WINDOW_SECONDS` | `1.0` | Rejection-event matching window. |
| `BACKEND_SNIFFER_BUFFER_CAPACITY` | `20000` | Pre-DB capture buffer capacity. |
| `BACKEND_SNIFFER_SOCKET_RCVBUF_BYTES` | `4194304` | Requested raw-socket receive buffer. |
| `BACKEND_SNIFFER_SELECTOR_TIMEOUT_SECONDS` | `1.0` | Reader select timeout. |
| `BACKEND_SNIFFER_RECV_BYTES` | `65536` | Maximum raw receive length. |
| `BACKEND_SNIFFER_BUFFER_DRAIN_INTERVAL_SECONDS` | `5.0` | Buffered-record drain frequency. |
| `BACKEND_SNIFFER_THREAD_JOIN_TIMEOUT_SECONDS` | `2.0` | Session reader join timeout. |
| `BACKEND_BRIDGE_BPF_BUILD_DIR` | `/tmp/mitm-bpf` | eBPF build artifacts directory. |
| `BACKEND_BRIDGE_TELEMETRY_RAW_SAMPLE_EVERY` / `META_SAMPLE_EVERY` | `1` / `1` | Raw/meta sampling rates; zero is allowed. |
| `BACKEND_BRIDGE_TELEMETRY_INGRESS_PERF_PAGES` / `META_PERF_PAGES` | `256` / `128` | eBPF perf-buffer page counts. |
| `BACKEND_BRIDGE_TELEMETRY_EVENT_QUEUE_MAXSIZE` | `20000` | Telemetry event queue cap. |
| `BACKEND_BRIDGE_TELEMETRY_QUEUE_RECOVERY_SIZE` | `1000` | Queue recovery threshold. |
| `BACKEND_BRIDGE_TELEMETRY_DROP_LOG_INTERVAL_SECONDS` | `5.0` | Telemetry-drop log throttling. |
| `BACKEND_BRIDGE_LINK_STATE_THREAD_JOIN_TIMEOUT_SECONDS` | `2.0` | Link watcher join timeout. |
| `BACKEND_BRIDGE_LINK_STATE_FAILURE_HOLDOFF_SECONDS` / `RECOVERY_HOLDOFF_SECONDS` | `0.75` / `1.0` | Delay before propagating failure/recovery. |
| `BACKEND_BRIDGE_LINK_STATE_DEGRADED_RECHECK_SECONDS` | `0.5` | Degraded-link polling period. |
| `BACKEND_TELEMETRY_PROCESS_STOP_TIMEOUT_SECONDS` / `READER_JOIN_TIMEOUT_SECONDS` | `3.0` / `2.0` | Telemetry subprocess shutdown limits. |
| `BACKEND_TSHARK_ENABLED` | `true` | Enables tshark worker management. |
| `BACKEND_TSHARK_DISPLAY_FILTER` | empty | Optional tshark display filter. |
| `BACKEND_TSHARK_TRY_HEURISTIC_FIRST` | `true` | Applies local heuristic before tshark match. |
| `BACKEND_TSHARK_CACHE_TTL_SECONDS` | `5.0` | Enrichment cache lifetime. |
| `BACKEND_TSHARK_MATCH_WINDOW_MS` | `5000` | Capture-to-tshark matching window. |
| `BACKEND_TSHARK_READER_JOIN_TIMEOUT_SECONDS` / `PROCESS_STOP_TIMEOUT_SECONDS` | `2.0` / `3.0` | tshark shutdown limits. |
## Operational safeguards
Run the API behind an authenticated, access-controlled boundary. The configured
CORS policy currently permits every origin, method, and header; it is convenient
for development but should not be treated as an authorization control. Keep DB
credentials out of version control and use a production-specific DSN.
Before starting capture, verify target interface/bridge names and ensure recovery
access to the host. Before using firewall or script endpoints, snapshot the nft
ruleset and understand which systemd units and filesystem paths are in scope.

View File

@@ -0,0 +1,65 @@
# Packet data, persistence, and analysis
## Packet record
`Models/packets.py` defines the normalized `PacketDBModel` returned by packet
history APIs. It represents one correlated packet record, not necessarily one raw
capture callback. A record may combine several observations.
| Field group | Fields | Meaning |
| --- | --- | --- |
| Identity | `id`, `timestamp`, `updated_at`, `correlation_key`, `correlation_source`, `packet_id`, `packet_uid`, `skb_mark` | Database identity and the evidence used to correlate capture/telemetry data. |
| Path | `capture_iface`, `ingress_if`, `egress_if`, `capture_session_id`, `capture_sources` | Where and how it was observed. |
| Link/network/transport | MACs, EtherType, IP protocol, IPs, ports, VLAN, length | Parsed packet headers. Raw numeric values are retained beside human-readable names. |
| Application enrichment | `flow_id`, app protocol/master protocol/category/confidence/hostname/encryption/risk, `dpi_metadata` | tshark-derived context when available. |
| Evidence | `raw_b64`, `raw_present`, capture/telemetry metadata and `capture_observations` | Raw bytes and provenance; may be absent. |
| Outcome | `verdict`, reason/confidence, ingress/egress/verdict timestamps | Forwarding outcome inferred from telemetry. |
Each `PacketObservationModel` identifies whether its contribution was `capture` or
`telemetry`, the source, interface, timestamp, event type, capture mode, session,
and optional reason. Consumers should not assume that every optional field exists:
AF_PACKET, tc/eBPF, and enrichment sources provide different evidence.
## DatabasePool
`DatabasePool` is an asyncpg wrapper initialized from `BACKEND_DB_DSN`. Startup
makes compatibility changes to an existing `packets` table: fills missing
timestamps, sets timestamp defaults/non-null constraints, adds
`capture_observations` JSONB if needed, and creates an index on `packet_uid`.
Writes use an upsert keyed by `correlation_key`. `upsert_packet` returns a row and
publishes it to the packet broadcaster; `upsert_packets` is a batched performance
path and does not individually publish rows. Before writing, it normalizes JSON
fields and attaches derived protocol, flow, and analysis fields.
Read/analysis methods are:
- `fetch_latest(limit)` for history.
- `backfill_packet_metadata` and `backfill_stream_metadata` for late tshark data.
- `infer_interface_hosts`, `infer_interface_host_protocols`, and
`infer_interface_protocol_paths` for topology/protocol views.
- `analyze_conversations` and `fetch_conversation_flow_detail` for directional
communication views.
- `analyze_host_intelligence`, `analyze_discovery_activity`, and
`analyze_anomalies` for investigation aids.
- `clear_all_packets(reset_identity)` for destructive history cleanup.
## Analysis interpretation
The analysis API runs SQL aggregations over what the system captured. It infers
attachment from traffic evidence, groups protocol paths and conversations, and
derives host/service/hostname hints. Its anomaly queries rank plausible scan,
beacon, rare-service, TCP-reset, and drop-heavy patterns. These are leads for an
operator—not assertions of a network's real topology, attribution, or malicious
intent. Missing capture events, encrypted traffic, NAT, asymmetric paths, and
limits change the output.
## Enumerations and configuration models
`Models/etherType.py` provides a string-valued `EtherTypeEnum` and
`ethertype_from_int`; `Models/ip_protocol.py` provides `IPProtocolEnum` and
`protocol_from_number`. They turn numeric protocol fields into readable labels
while keeping raw values. `Models/netplan.py` provides Pydantic schemas for
nameservers, Ethernet settings, bridge settings, and a full Netplan-style network
configuration. These models are reusable schemas; they are not a substitute for
applying a Netplan configuration in the currently mounted API.

View File

@@ -0,0 +1,90 @@
# Linux host integration and side effects
## Network and bridge control
`api/network_api.py` retains process-wide pyroute2 `IPRoute` and `NDB` objects.
It reads addresses, link flags, routes and bridge membership through netlink, and
uses NDB/pyroute2 to create or remove bridges. Resetting interfaces executes
`ethtool` and changes MTU/profile values. These operations affect the host's live
connectivity; API errors must be treated as operational failures, not merely input
validation failures.
`utilities/interface_bridge_helpers.py` is the low-level read layer. It checks
interface presence/up state, reads sysfs operational/carrier/admin/MTU values,
obtains Ethernet profile data using `ethtool`, caches profile data, and reads bridge
members from sysfs. It deliberately supplies best-effort information when a driver
or platform cannot report every property.
`bridge_link_state_manager.py` owns optional event-driven bridge watchers. Each
watcher tracks Ethernet profile and member readiness, uses failure and recovery
holdoffs to avoid flapping, and adjusts selected peer state so an inline bridge
reacts coherently to member link loss. `BridgeLinkStateManager` indexes watchers,
enables/disables them, reports individual/all status, and stops all during shutdown.
## eBPF/tc telemetry
`bridge_telemetry.py` manages the lifecycle of the telemetry helper. Its
`update_sessions` method reconciles currently requested bridge ports with the
subprocess; `stop` terminates it and `get_debug_snapshot` provides operator
diagnostics. It does not itself parse kernel events.
`ebpf_bridge_events.py` is the helper process. It builds BPF source, attaches tc
programs to requested interfaces, reads perf events, and writes JSON-safe event
payloads. Events cover ingress raw capture plus egress/drop metadata, including
interfaces, MAC/IP information, packet identity, event/reason names, and timing.
It cleans existing clsact qdiscs/program attachment as part of setup/cleanup. This
requires an appropriate kernel, BCC Python bindings/toolchain, tc, and privileges.
`tools/ebpf/mark_packet_id.c` is related kernel-side support for packet marking;
the mark is decoded by `utilities/packet_mark.py` and used in tracker correlation.
## nftables
The mounted `api/nft_manager.py` uses `pip-nftables` to list JSON/text rulesets,
normalize them into stable table/chain/rule models, parse rule priorities/text, and
delete a rule by handle. Its raw-command endpoint forwards textual nft commands.
It therefore needs capability to inspect and change the host nftables ruleset.
Two alternative implementations exist but are currently unmounted:
- `api/nft_api.py` is bridge-family oriented. It models meta, Ethernet, IP, port,
conntrack, verdict, reject, log, and raw expressions; can generate previews,
list rules with authoritative handles, add/delete/update rules, and uses the
`nft` CLI.
- `api/nftables_api.py` is a stateless typed replacement API. It models matches and
actions, chooses a pyroute2 binding when viable or a CLI wrapper otherwise,
ensures table/chain presence, reconstructs readable rules, and replaces a chain's
ruleset. Its own source warns that a running asyncio loop may force CLI fallback.
Do not mount more than one firewall router without an explicit API versioning and
conflict review: all manipulate shared kernel state and have overlapping concepts.
## NFQUEUE script services
`api/packet_scripting_api.py` manages executable Python scripts. It makes these
directories at import time: `/srv/fw-scripts`, `/srv/fw-scripts/venvs`, and the
repository's `backend/example_scripts`. Scripts are named `<name>.py`; requirements
are `<name>-requirements.txt`; virtual environments are per-script. Units use the
deterministic name `fw-script-<name>-q<queue>.service` and are written under
`/etc/systemd/system`.
The module discovers services through `systemctl`, writes/parses unit `ExecStart`,
runs `daemon-reload`, starts/stops/enables/disables units, creates virtualenvs, and
uses pip to install user-provided requirements. Startup can copy protected example
scripts and optionally deploy them from `<name>.deploy.json`. This API is a remote
code/service-management surface and requires strict authentication plus host-level
least privilege.
## External subprocesses
| Integration | Commands/facility | Used by |
| --- | --- | --- |
| tshark | long-lived `tshark` subprocesses | DPI enrichment |
| nftables | `nft` CLI and/or pip-nftables bindings | firewall APIs |
| Ethernet control | `ethtool` | interface profile/reset |
| system services | `systemctl`, virtualenv, pip | script lifecycle |
| BPF/tc | BCC, tc, qdisc/program attachment | bridge telemetry |
Failures are generally logged and translated to endpoint failures or degraded
capture. Operators should collect `/api/sniffer/debug`, service logs, nftables
state, and interface state when investigating a problem.

View File

@@ -0,0 +1,233 @@
# Technical reference: packet sniffing modes
This document specifies the implemented capture behavior in `network_sniffer.py`,
`bridge_telemetry.py`, `ebpf_bridge_events.py`, and `packet_tracker.py`. It makes a
deliberate distinction between observed facts and inferred forwarding results.
## Session model and mode selection
A capture session has a UUID and targets exactly one interface or exactly one
bridge. A bridge is expanded once with `get_bridge_ports_once`; its member list is
a creation-time snapshot. Later bridge membership changes are not added to the
existing session. Session state includes target label/type, effective mode, benchmark
flag, port snapshot, AF_PACKET sockets, optional reader thread, stop event, and
benchmark counters.
| Request | Effective mode | Capture source | Path/outcome evidence |
| --- | --- | --- | --- |
| Interface, any requested mode | `af_packet` | One raw socket on the interface | Packet-socket type labels an outgoing copy as egress; no kernel verdict telemetry. |
| Bridge, `af_packet` | `af_packet` | One raw socket per snapshot bridge port | Two matching port observations can infer forwarding. |
| Bridge, `tc_ebpf` (default) | `tc_ebpf` | One tc/eBPF helper across snapshot bridge ports | TC ingress/egress and skb-free/drop events, normally matched by skb mark. |
| Any mode with benchmark enabled | Same hook/socket setup | Counted but not processed | No parsing, enrichment, DB write, or WebSocket event. |
Interface targets always use AF_PACKET. The TC/eBPF mode is only selected for a
bridge target. A stop by session ID is the safest selector. Stopping a session also
discards pending tracker entries relating to its interfaces, so unpersisted data can
be lost deliberately at shutdown.
Multiple sessions may overlap on an interface. This is not an independent-capture
guarantee: the TC manager maps an interface to several sessions but assigns raw
ingress parsing to the first sorted session ID.
## AF_PACKET capture
### Socket behavior
For each capture interface the service opens `AF_PACKET` / `SOCK_RAW` with protocol
`htons(0x0003)` (`ETH_P_ALL`), requests the configured receive buffer (default
4 MiB), best-effort requests `TPACKET_V3`, binds to `(ifname, 0)`, and makes the
socket non-blocking. Failure to set the buffer or TPACKET version is non-fatal.
Failure to create or bind leaves the interface uncaptured; session creation can still
complete. This requires raw-socket privilege, commonly `CAP_NET_RAW`.
One daemon reader thread is started only when the session has sockets. It uses a
selector, receives at most `BACKEND_SNIFFER_RECV_BYTES` bytes per event (default
65,536), stamps the frame with userspace UTC receive time, creates Scapy `Ether`,
and calls the common parser. `ENODEV`, `ENETDOWN`, and `EBADF` close the affected
socket; it is not reopened in that session. The thread periodically attempts a
pre-DB-buffer drain during selector idle time.
### Direction and bridge inference
Packet-socket address metadata is used only as follows: `PACKET_OUTGOING` (normally
4) becomes `path_role: egress`; every other packet type becomes `path_role:
ingress`. This is a packet-socket perspective, not proof of a Linux bridge decision.
For a bridge session, the tracker groups AF_PACKET observations by session ID. It
uses an explicit ingress observation if available, otherwise the earliest one. It
prefers an explicit egress observation on a different port, otherwise a later
different-port observation. If one correlated packet is seen on at least two ports,
the tracker records:
```text
verdict = accept
verdict_reason = bridge-af_packet-forwarded-observed
verdict_confidence = medium
```
This means matching evidence was observed on two bridge ports. It does not prove a
particular kernel forwarding verdict and can be affected by duplicate copies, loops,
or fallback-identity collisions. A single-interface AF_PACKET record has no terminal
verdict from AF_PACKET itself.
### AF_PACKET implications
AF_PACKET provides full observed frame bytes without BCC or tc changes and is the
only interface-capture mode. It neither alters packets nor controls forwarding. It
also has no definitive drop visibility, reports userspace rather than kernel event
time, can observe local/outgoing copies, and can lose traffic under socket/userspace
load. The full-frame Scapy parse, tshark lookup, tracking and persistence path makes
it more expensive than sampled telemetry.
## TC/eBPF bridge capture
### Collector lifecycle and destructive qdisc behavior
Bridge sessions in `tc_ebpf` mode are aggregated into one helper process. Any change
to the active *interface set* stops the helper and recreates it for the new set;
there is a capture gap during that restart. The helper attaches direct-action
`BPF.SCHED_CLS` programs at TC ingress (`ffff:fff2`, handle `:20`) and egress
(`ffff:fff3`, handle `:30`) to every bridge **member interface**, not the bridge
device itself.
Before attachment the helper runs `tc qdisc del dev <iface> clsact` (ignoring its
result), then `tc qdisc add dev <iface> clsact`. It deletes `clsact` again for every
instrumented interface at helper shutdown and after an attachment failure.
> Starting, restarting, failing, or stopping TC/eBPF capture can remove pre-existing
> clsact qdiscs and their filters. Do not use it on interfaces with unrelated TC
> configuration unless coexistence and recovery are explicitly managed.
The BCC Python runtime, a compatible kernel, BPF/tracepoint access, TC and netlink
privileges are required. Session creation does not wait for a collector health
acknowledgement, so a successful start response is not proof that BPF attached.
### Kernel event generation
The helper opens a raw-ingress perf buffer and a metadata perf buffer. The ingress
TC program creates an skb mark only when it is zero, using the low 28 bits of
`bpf_ktime_get_ns()` and replacing zero with one. It preserves any existing nonzero
mark. It extracts Ethernet addresses, EtherType, a single 802.1Q/802.1AD VLAN ID,
ARP IPv4 addresses, and IPv4/IPv6 addresses with TCP/UDP ports. IPv6 extension
headers are not traversed; the base next-header is used as protocol.
The egress TC program never creates a mark. It exports metadata only for marked
packets. The `skb:kfree_skb` tracepoint reads the linear skb representation and
exports a drop event only for marked skbs whose device is a selected interface.
The emitted payload contains userspace and kernel-monotonic timestamps, interface,
mark, length, parsed L2–L4 fields, and event type. Drop events add a numerical
reason and `skb_drop_reason_<n>` label. Ingress events may contain `raw_b64`; egress
and drop events do not.
A kfree_skb event is evidence that a marked skb was freed in the kernel context. It
is not automatically evidence that nftables caused the outcome; interpret the
reason code in the context of kernel behavior and other instrumentation.
### Sampling
Raw and metadata sampling are independent settings.
| Value | Effect |
| --- | --- |
| `0` | Never emits that sample category. |
| `1` | Emits every marked packet in that category. |
| `N > 1` | Emits when `skb_mark % N == 0`. |
At ingress, a raw-selected packet emits a raw event; only a packet not chosen for
raw can emit an ingress metadata event. Egress and drop use metadata sampling only.
Therefore raw-enabled/meta-disabled capture stores sampled ingress frame records
without egress/drop visibility; raw-disabled/meta-enabled capture produces
metadata-only rows without raw bytes. Both enabled does not make raw and metadata
populations identical.
Sampling uses the entire existing skb mark. The documented mark layout reserves
bits 0–27 for packet ID and upper bits for drop/reject hints. This capture program
creates only the low-28-bit value for previously zero marks; it does not set verdict
hints. Any other mark-using subsystem must coordinate its mark semantics, because
it can change both sampling and correlation.
### Userspace event handling and loss
The manager reads JSON helper output into a bounded queue. For a non-benchmark
ingress event with `raw_b64`, it decodes the frame and sends it into the common Scapy
parser as source `tc_ingress_raw`, with `packet_id`, `skb_mark`, and capture mode
`tc_ingress`. It then sends every non-benchmark ingress/egress/drop event to the
tracker. A sampled raw ingress packet usually therefore has both a parsed capture
observation and a telemetry observation under the same mark-derived key.
When the telemetry queue is full, the manager drops oldest queued events down to
`BACKEND_BRIDGE_TELEMETRY_QUEUE_RECOVERY_SIZE`, attempts to keep the new event, and
counts dropped events and raw payloads. Perf buffers can also lose samples before
userspace. Neither loss mechanism is recovered. `/api/sniffer/debug` reports queue
size, queue drops, benchmark counts, collector interfaces, and tracker statistics.
### TC/eBPF implications
This mode yields better within-host correlation and explicit TC egress evidence. A
matching egress produces `accept`, `egress-observed`, confidence `high`; a matching
drop produces `drop` (or mark hint), confidence `high`. Absence of egress is not
proof of a drop: sampling, perf loss, queue loss, an uninstrumented path, teardown,
or collector failure can all explain it. Raw bytes are ingress-only and sampled.
## Common parsing, enrichment, and identity
Both modes use `parse_packet` / `parse_packet_bytes`. The parser records Ethernet
addresses, EtherType and VLAN, ARP operation/addressing, IPv4 ID or IPv6 base
header, TCP sequence/acknowledgement/flags, UDP ports, ICMP/ICMPv6 type/code, and
an embedded IPv4 tuple from eligible ICMP errors. It stores full raw frame bytes
when supplied by AF_PACKET or sampled TC ingress.
tshark workers are enabled for non-benchmark capture interfaces. They may add
application protocol, category, confidence, hostname, encryption/risk, and flow/DPI
metadata. They are optional and asynchronous; a failure or late match does not
discard underlying capture, and later backfill can enrich stored records.
The preferred identity is `pid:<packet-id>`, where the ID is bits 0–27 of skb mark.
Without it, a SHA-1 `uid` is calculated. The Scapy fallback includes L2–L4 fields,
IPv4 ID, ARP/ICMP fields and TCP sequence/ack/flags; eBPF metadata's fallback uses
only the smaller L2–L4 tuple and length. Hash-only correlation is consequently a
best-effort fallback, weaker for repeated/identical/fragmented traffic.
## Tracker outcomes and persistence
The tracker deduplicates observations, merges available fields, retains the earliest
timestamp, and waits the configured finalization delay (default 250 ms).
| Evidence | Verdict | Confidence |
| --- | --- | --- |
| TC egress telemetry | `accept`; `egress-observed` | high |
| TC drop telemetry | mark hint or `drop`; kernel reason / `kfree_skb` | high |
| Matching recent TCP RST or ICMP unreachable after drop | `reject` | medium |
| Same AF_PACKET bridge record on two ports | `accept`; forwarding observed | medium |
| No terminal evidence before delay expires | `unknown`; `timeout` | low |
It asynchronously upserts batches to PostgreSQL. Entry-cap pressure, persistence
failure/retry limits, dirty-age expiry, collector queue loss, socket loss, and
shutdown can all cause incompleteness. A packet history or WebSocket feed is never
a proof of lossless capture. Batch upserts also do not individually publish packet
updates, so realtime consumers must use history reconciliation.
## Benchmark mode
Benchmark mode still creates sockets or TC hooks but bypasses normal processing.
AF_PACKET increments received frame and byte counters. TC/eBPF increments helper
event counters and raw-payload-event counters. It does not parse Scapy, invoke
tshark, call the tracker, persist rows, or publish updates. AF_PACKET counters count
socket frames; TC counters count emitted sampled events. They are not comparable as
equal packet totals without accounting for sampling and multiple event types.
## Selection guidance
| Need | Mode | Important caveat |
| --- | --- | --- |
| Full raw visibility for a single interface | AF_PACKET | No definitive kernel egress/drop verdict. |
| Full raw frames across bridge ports | Bridge AF_PACKET | High userspace work; bridge forwarding is inferred. |
| Ingress/egress/drop evidence on a controlled bridge | TC/eBPF | Requires BPF/TC privileges and resets clsact. |
| Reduced overhead / sampled observability | TC/eBPF sampling | Data is intentionally incomplete. |
| Hook-overhead measurement | Benchmark mode | Counts differ between AF_PACKET and TC. |
Before TC/eBPF capture, inspect `tc qdisc` and filters for every target port,
coordinate skb-mark ownership, verify BCC/kernel support, and plan recovery of the
TC configuration. For every mode, monitor sniffer debug counters, system logs,
capture/process health, DB persistence failures, and expected traffic rate before
making operational or security conclusions.

View File

@@ -0,0 +1,70 @@
# Source reference
This index covers every Python module under `backend/src`, including helper and
unmounted-router code. Function names prefixed with `_` are private implementation
details; they are described by their owning module's responsibility rather than as
separate public contracts.
## Application and configuration
| Module | Public surface and role |
| --- | --- |
| `main.py` | Creates FastAPI, enables permissive CORS, registers startup/shutdown handlers, provides `/hello` and `/versions`, includes live routers, and registers script lifecycle hooks. |
| `config.py` | Parses environment strings/integers/floats/booleans; immutable `BackendSettings`; `load_settings`; module-global `settings`. See [configuration](configuration.md). |
| `shared_objects.py` | Process-global `db`, `web_loop`, packet broadcaster, and network broadcaster references initialized by `main`. |
## Models
| Module | Public surface and role |
| --- | --- |
| `Models/packets.py` | `PacketObservationModel` and `PacketDBModel`, the normalized persisted/API packet schemas. |
| `Models/ip_protocol.py` | `IPProtocolEnum` and `protocol_from_number`, translating IANA protocol numbers to labels. |
| `Models/etherType.py` | `EtherTypeEnum` and `ethertype_from_int`, translating Ethernet type values to labels. |
| `Models/netplan.py` | `Nameservers`, `EthernetConfig`, `BridgeConfig`, and `NetworkConfig` Pydantic schemas for Netplan-shaped network data. |
## API routers
| Module | Public surface and role |
| --- | --- |
| `api/network_api.py` | Network inspection, bridge create/remove, default reset, link-state watcher control, and network-state WebSocket. Holds shared `IPRoute`/`NDB`; converts netlink messages to Pydantic interface/route/bridge models; publishes state after mutations. |
| `api/sniffer_api.py` | Pydantic start/stop/status models and endpoints. Validates one capture target and calls the capture-session API. |
| `api/packet_api.py` | Latest packet retrieval, packet-history deletion, and packet-update WebSocket. Serialization handles database records and Pydantic values safely for JSON. |
| `api/analysis_api.py` | Pydantic evidence/response models for attachment, protocols, paths, conversations, flow detail, hosts, discovery and anomaly views; delegates each endpoint to `DatabasePool`. |
| `api/nft_manager.py` | **Mounted.** `NftManager` wrapper, normalized ruleset models and functions to list rules, delete by handle, and run textual nft. It parses JSON and textual output to enrich rule data. |
| `api/packet_scripting_api.py` | **Mounted with doubled prefix.** Name/path validation, example deployment, systemd unit management, venv/pip operations, script status models, and upload/download/enable/disable/delete endpoints. |
| `api/nft_api.py` | **Not mounted.** Bridge nftables typed expression model, command generator, handle mapping, and CRUD/preview endpoint functions. `RuleModel.only_bridge` rejects other families. |
| `api/nftables_api.py` | **Not mounted.** Generic typed match/action models, resilient binding/CLI wrapper selection, rule reconstruction, and list/replace endpoint functions. |
## Capture, telemetry, and broadcasting utilities
| Module | Public surface and role |
| --- | --- |
| `network_sniffer.py` | Defines flexible `PacketInfo`; parses packet objects/bytes; opens/closes AF_PACKET sockets; owns session reader loops; coordinates telemetry; exposes `start_capture_session`, `stop_capture_session`, status and debug accessors. Legacy `*_afpacket_sniffer` functions delegate to current session functions. |
| `utilities/packet_tracker.py` | `PacketTracker` observes capture or telemetry events, aggregates observations, schedules persistence, stops/discards state, and exposes diagnostics. The module-global tracker is the correlation entrypoint. |
| `utilities/packet_identity.py` | Builds deterministic fallback packet UID and the minimum fields used to calculate it. |
| `utilities/packet_mark.py` | Decodes numeric skb marks into a normalized mark, packet ID, and verdict hint according to the shared mark layout. |
| `utilities/tshark_manager.py` | `TsharkManager` owns optional worker processes and caches. Parsing helpers safely coerce nested tshark JSON, extract protocol/HTTP/TLS/DNS/TCP data, derive stream context, and merge enrichment. Module-global `tshark_manager` is used by capture. |
| `utilities/bridge_telemetry.py` | `BridgeTelemetryManager` starts/reconciles/stops the eBPF helper and reports subprocess/queue state. Module-global manager is invoked by sniffer lifecycle. |
| `utilities/ebpf_bridge_events.py` | Standalone helper program: ctypes event format, BPF-source construction, tc attach/cleanup, perf callbacks, JSON output, signal handling, and `main`. |
| `utilities/packet_broadcaster.py` | `PacketBroadcaster` manages subscriber queues. `subscribe`/`unsubscribe`, async `publish`, cross-thread `sync_publish`, and async `close` provide the WebSocket transport primitive. |
## Network and persistence utilities
| Module | Public surface and role |
| --- | --- |
| `utilities/interface_bridge_helpers.py` | Interface existence/up tests; sysfs readers for operational/carrier/admin/MTU state; Ethernet profile retrieval/cache; bridge-port discovery. |
| `utilities/bridge_link_state_manager.py` | `EthernetProfile` and `MemberLinkState` data objects; `BridgeLinkStateWatcher` start/stop/status; `BridgeLinkStateManager` enable/disable/query/stop. It embodies debounce, failure, recovery, and profile propagation logic. |
| `utilities/database.py` | `DatabasePool` initialization/closure, upsert/batch-upsert, enrichment backfills, latest-packet query, all analysis SQL, and history clearing. Internal helpers normalize values, derive protocol/flow/analysis fields, serialize outgoing rows, and classify discovery activity. |
## Extension points and maintenance notes
- New API functionality should live in an `APIRouter`, use Pydantic request and
response models, and be explicitly included from `main.py`; otherwise it is not
live.
- New capture fields must be updated consistently in packet parsing, tracker merge,
database upsert SQL, `PacketDBModel`, broadcaster serialization, and analysis
queries where relevant.
- Any new Linux side effect belongs in [host integration](host-integration.md),
including required binary/capability, rollback behavior, and its API exposure.
- If an unmounted nft router is adopted, document the migration and remove or
version conflicting endpoints instead of silently mounting another implementation.