Files
mitm-webserver/documentation/backend/data-and-analysis.md
Marcus Almert 9106cac2a2
Some checks failed
Build and Deploy MITM Webserver / build (push) Has been cancelled
Build and Deploy MITM Webserver / traffic_target (push) Has been cancelled
documentation md files
2026-08-30 17:31:46 +02:00

3.8 KiB

Packet data, persistence, and analysis

Packet record

Models/packets.py defines the normalized PacketDBModel returned by packet history APIs. It represents one correlated packet record, not necessarily one raw capture callback. A record may combine several observations.

Field group Fields Meaning
Identity id, timestamp, updated_at, correlation_key, correlation_source, packet_id, packet_uid, skb_mark Database identity and the evidence used to correlate capture/telemetry data.
Path capture_iface, ingress_if, egress_if, capture_session_id, capture_sources Where and how it was observed.
Link/network/transport MACs, EtherType, IP protocol, IPs, ports, VLAN, length Parsed packet headers. Raw numeric values are retained beside human-readable names.
Application enrichment flow_id, app protocol/master protocol/category/confidence/hostname/encryption/risk, dpi_metadata tshark-derived context when available.
Evidence raw_b64, raw_present, capture/telemetry metadata and capture_observations Raw bytes and provenance; may be absent.
Outcome verdict, reason/confidence, ingress/egress/verdict timestamps Forwarding outcome inferred from telemetry.

Each PacketObservationModel identifies whether its contribution was capture or telemetry, the source, interface, timestamp, event type, capture mode, session, and optional reason. Consumers should not assume that every optional field exists: AF_PACKET, tc/eBPF, and enrichment sources provide different evidence.

DatabasePool

DatabasePool is an asyncpg wrapper initialized from BACKEND_DB_DSN. Startup makes compatibility changes to an existing packets table: fills missing timestamps, sets timestamp defaults/non-null constraints, adds capture_observations JSONB if needed, and creates an index on packet_uid.

Writes use an upsert keyed by correlation_key. upsert_packet returns a row and publishes it to the packet broadcaster; upsert_packets is a batched performance path and does not individually publish rows. Before writing, it normalizes JSON fields and attaches derived protocol, flow, and analysis fields.

Read/analysis methods are:

  • fetch_latest(limit) for history.
  • backfill_packet_metadata and backfill_stream_metadata for late tshark data.
  • infer_interface_hosts, infer_interface_host_protocols, and infer_interface_protocol_paths for topology/protocol views.
  • analyze_conversations and fetch_conversation_flow_detail for directional communication views.
  • analyze_host_intelligence, analyze_discovery_activity, and analyze_anomalies for investigation aids.
  • clear_all_packets(reset_identity) for destructive history cleanup.

Analysis interpretation

The analysis API runs SQL aggregations over what the system captured. It infers attachment from traffic evidence, groups protocol paths and conversations, and derives host/service/hostname hints. Its anomaly queries rank plausible scan, beacon, rare-service, TCP-reset, and drop-heavy patterns. These are leads for an operator—not assertions of a network's real topology, attribution, or malicious intent. Missing capture events, encrypted traffic, NAT, asymmetric paths, and limits change the output.

Enumerations and configuration models

Models/etherType.py provides a string-valued EtherTypeEnum and ethertype_from_int; Models/ip_protocol.py provides IPProtocolEnum and protocol_from_number. They turn numeric protocol fields into readable labels while keeping raw values. Models/netplan.py provides Pydantic schemas for nameservers, Ethernet settings, bridge settings, and a full Netplan-style network configuration. These models are reusable schemas; they are not a substitute for applying a Netplan configuration in the currently mounted API.