# Packet data, persistence, and analysis ## Packet record `Models/packets.py` defines the normalized `PacketDBModel` returned by packet history APIs. It represents one correlated packet record, not necessarily one raw capture callback. A record may combine several observations. | Field group | Fields | Meaning | | --- | --- | --- | | Identity | `id`, `timestamp`, `updated_at`, `correlation_key`, `correlation_source`, `packet_id`, `packet_uid`, `skb_mark` | Database identity and the evidence used to correlate capture/telemetry data. | | Path | `capture_iface`, `ingress_if`, `egress_if`, `capture_session_id`, `capture_sources` | Where and how it was observed. | | Link/network/transport | MACs, EtherType, IP protocol, IPs, ports, VLAN, length | Parsed packet headers. Raw numeric values are retained beside human-readable names. | | Application enrichment | `flow_id`, app protocol/master protocol/category/confidence/hostname/encryption/risk, `dpi_metadata` | tshark-derived context when available. | | Evidence | `raw_b64`, `raw_present`, capture/telemetry metadata and `capture_observations` | Raw bytes and provenance; may be absent. | | Outcome | `verdict`, reason/confidence, ingress/egress/verdict timestamps | Forwarding outcome inferred from telemetry. | Each `PacketObservationModel` identifies whether its contribution was `capture` or `telemetry`, the source, interface, timestamp, event type, capture mode, session, and optional reason. Consumers should not assume that every optional field exists: AF_PACKET, tc/eBPF, and enrichment sources provide different evidence. ## DatabasePool `DatabasePool` is an asyncpg wrapper initialized from `BACKEND_DB_DSN`. Startup makes compatibility changes to an existing `packets` table: fills missing timestamps, sets timestamp defaults/non-null constraints, adds `capture_observations` JSONB if needed, and creates an index on `packet_uid`. Writes use an upsert keyed by `correlation_key`. `upsert_packet` returns a row and publishes it to the packet broadcaster; `upsert_packets` is a batched performance path and does not individually publish rows. Before writing, it normalizes JSON fields and attaches derived protocol, flow, and analysis fields. Read/analysis methods are: - `fetch_latest(limit)` for history. - `backfill_packet_metadata` and `backfill_stream_metadata` for late tshark data. - `infer_interface_hosts`, `infer_interface_host_protocols`, and `infer_interface_protocol_paths` for topology/protocol views. - `analyze_conversations` and `fetch_conversation_flow_detail` for directional communication views. - `analyze_host_intelligence`, `analyze_discovery_activity`, and `analyze_anomalies` for investigation aids. - `clear_all_packets(reset_identity)` for destructive history cleanup. ## Analysis interpretation The analysis API runs SQL aggregations over what the system captured. It infers attachment from traffic evidence, groups protocol paths and conversations, and derives host/service/hostname hints. Its anomaly queries rank plausible scan, beacon, rare-service, TCP-reset, and drop-heavy patterns. These are leads for an operator—not assertions of a network's real topology, attribution, or malicious intent. Missing capture events, encrypted traffic, NAT, asymmetric paths, and limits change the output. ## Enumerations and configuration models `Models/etherType.py` provides a string-valued `EtherTypeEnum` and `ethertype_from_int`; `Models/ip_protocol.py` provides `IPProtocolEnum` and `protocol_from_number`. They turn numeric protocol fields into readable labels while keeping raw values. `Models/netplan.py` provides Pydantic schemas for nameservers, Ethernet settings, bridge settings, and a full Netplan-style network configuration. These models are reusable schemas; they are not a substitute for applying a Netplan configuration in the currently mounted API.