Files
mitm-webserver/documentation/thesis/notes/preliminaries ideas.md
malmert 945b259ebb
All checks were successful
Build and Deploy MITM Webserver / traffic_target (push) Successful in 0s
Build and Deploy MITM Webserver / build (push) Successful in 12s
scripts and texts
2026-05-03 16:38:37 +02:00

5.6 KiB

Your project is already much richer than a generic “MITM tool.” From the code, it is really a transparent inline Layer-2 observation and manipulation platform: it creates a Linux bridge with STP disabled, manages bridge member behavior, captures traffic either via AF_PACKET or tc/eBPF, correlates packet observations with kernel telemetry, applies nftables/NFQUEUE manipulation, and builds higher-level traffic intelligence on top of that. You can see those pillars in network_api.py, bridge_link_state_manager.py, network_sniffer.py, bridge_telemetry.py, packet_tracker.py, nftables_api.py, packet_scripting_api.py, and analysis_api.py. Compared with your current 02-preliminaries.tex, the thesis would benefit from moving beyond a mainly OSI-focused introduction.

What I would definitely add to the preliminaries

  • Transparent Layer-2 MITM / inline bridge systems: difference between routed MITM, proxying, TAP/SPAN capture, and transparent bridging.
  • Ethernet switching and Linux bridge internals: MAC learning, forwarding database, flooding, broadcast domains, unknown unicast, VLAN awareness, STP/RSTP, and why disabling STP matters for your setup.
  • Stealth / transparency criteria: what “hidden” means technically in your thesis. For example: no IP hop added, no TTL change, minimal forwarding delay, preserved link properties, no obvious protocol artifacts.
  • Link-state propagation and fail behavior: your code actively mirrors link failures and synchronizes MTU / speed / duplex / autoneg, which is unusually relevant for an inline appliance and worth explaining conceptually.
  • Linux packet-processing path: NIC, driver, sk_buff, bridge forwarding path, netfilter hooks, tc ingress/egress, and where capture/manipulation can be attached.
  • AF_PACKET raw sockets: why they are suitable for passive L2 capture, and their trade-offs.
  • nftables and the bridge family: tables, chains, hooks, priorities, verdicts, and why bridge-family filtering is important in a transparent bridge scenario.
  • NFQUEUE: how packets are punted to user space, latency/performance implications, and the difference between passive observation and inline modification.
  • eBPF at tc: attach points, maps, helpers, packet metadata access, and why eBPF is useful for low-overhead telemetry and packet correlation.
  • Packet marking and correlation: your project uses skb->mark-based packet IDs and verdict bits, which is a very strong thesis concept because it ties kernel events to captured packets (mark_packet_id.c).
  • Protocol parsing and enrichment: Ethernet, ARP, IPv4/IPv6, TCP/UDP/ICMP, plus DPI/enrichment with Scapy and tshark (tshark_manager.py).
  • Flow/conversation reconstruction: packet identity, deduplication, ingress/egress inference, flow IDs, conversations, and discovery traffic classification.
  • Threat model and limitations: what kinds of traffic can be observed/manipulated, how encryption limits analysis, and where the bridge can still become detectable.

Very thesis-relevant concepts that are specific to your implementation

  • Bridge transparency vs detectability
  • Bridge member synchronization
  • Event-driven network control via netlink / pyroute2
  • Hybrid observation pipeline: raw capture + kernel telemetry + DPI enrichment
  • Correlation of data-plane and control-plane evidence
  • Programmable packet handling with nftables + NFQUEUE scripts
  • Traffic-intelligence extraction from passive observations
  • Discovery protocol analysis: ARP, DHCP, mDNS, SSDP, LLMNR, NBNS, ICMPv6 discovery

What I would keep short or move to implementation

  • FastAPI, React, WebSockets, Docker, and general UI architecture
  • PostgreSQL schema details
  • systemd service deployment details for scripts

Those matter, but they feel more like implementation chapter material than preliminaries unless your thesis is explicitly about the full software platform architecture.

A strong chapter structure could be

  1. Communication models: brief OSI and TCP/IP mapping
  2. Ethernet and transparent bridging
  3. Linux bridge architecture and link-state behavior
  4. Linux packet path: raw sockets, netfilter, tc, and sk_buff
  5. nftables, bridge-family filtering, and NFQUEUE
  6. eBPF for packet telemetry and correlation
  7. Packet parsing, DPI, and flow reconstruction
  8. Stealth, detectability, and operational limitations
  9. Ethical and legal boundaries of MITM experimentation

If you want, I can turn this directly into a thesis-ready rewrite for 02-preliminaries.tex with subsection titles and short starter paragraphs.