GitHub

Detection & thresholds

Kapkan detects volumetric attacks — floods that try to saturate a host with raw packet or bit volume, as opposed to application-layer attacks — per destination host over a sliding time window. Rates are evaluated as per-second averages over that window (5 seconds by default), so a sub-second spike that does not sustain across the window will not, by itself, open an attack. For every protected address it keeps windowed rate counters, and the moment any configured threshold is crossed it opens an attack, attaches a flow sample, classifies the vector, and hands the event to mitigation, notifications and the API.

Every rate is evaluated in real, unsampled units. Flow exporters sample — they report one record per N packets — so Kapkan multiplies each observed rate by the exporter's sampling rate before comparing it to a threshold. You configure thresholds in the traffic the attacker is actually sending, not in sampled records.

iDetection is read-only

Detecting an attack never announces a route on its own. Detection produces an attack record and an event; mitigation is a separate, dry-run-by-default step. See Mitigation and the Safety model.

Core thresholds

The top-level thresholds block defines the baseline limits applied to each destination host inside your protected networks. Three core metrics are always evaluated:

MetricUnitMeaning
ppspackets/secTotal inbound packet rate to the host.
mbpsmegabits/secTotal inbound bit rate to the host.
flows_per_secflows/secDistinct flow records/sec toward the host. NetFlow/IPFIX only — see the note below.
thresholds:
  pps: 80000
  mbps: 1000
  flows_per_sec: 35000

All three core thresholds must be greater than 0. A host crosses into "under attack" as soon as any one of them is exceeded over the sliding window.

!flows_per_sec applies to NetFlow/IPFIX only

flows_per_sec counts aggregated flow records. sFlow does not aggregate — it exports one sample per packet — so Kapkan counts 0 flows for sFlow-sourced traffic, and the flows_per_sec threshold never fires for sFlow hosts (their flows_per_sec reads 0 in the API and dashboard). Counting each sample as a flow would just duplicate pps. On sFlow deployments, rely on pps, tcp_syn_pps and udp_pps to catch packet and connection floods. The value is still required (> 0) because NetFlow/IPFIX exporters can feed the same engine.

iPer-group overrides

These global thresholds are the fallback. You can give specific prefixes tighter or looser limits, or evaluate a pool's summed traffic, with hostgroups. Kapkan can also learn each host's normal level and tighten the effective threshold automatically — see Baselines.

Sampling correction

Sampled flow telemetry undercounts traffic by design: an exporter sampling 1-in-1000 reports roughly one record for every thousand packets. Kapkan corrects for this so your thresholds stay expressed in real traffic.

Every rate is multiplied by the exporter's sampling rate:

corrected_rate = observed_rate * sampling_rate

The sampling rate comes from the flow packet itself when the exporter reports it; otherwise Kapkan falls back to the configured sampling.default_rate:

sampling:
  default_rate: 1000   # used only when a packet carries no sampling rate; must be >= 1

Because correction happens before threshold comparison, a pps: 80000 threshold means 80,000 real packets per second regardless of how aggressively your routers sample.

!Set the right default rate

If your exporters do not advertise their sampling rate, sampling.default_rate must match what they actually sample. A default that is too low under-counts traffic and delays detection; one that is too high over-counts and trips early.

Correction assumes each packet is sampled once. If the same packet is observed at several vantage points — redundant exporters, ingress and egress sampling on one device, or transit links — Kapkan would sum every copy and over-count. Interface-boundary counting deduplicates this by counting a flow only where it crosses an external (uplink/border) interface; see interface-boundary counting for how to configure it.

Per-protocol thresholds

Beyond the core totals, you can set optional per-protocol limits. These let you catch an attack that is large for one protocol but invisible in the aggregate — for example a SYN flood whose packet rate is high but whose bit rate stays under the global mbps.

Each protocol has a pps and an mbps variant. The metric names are exactly:

Protocolpps metricmbps metricCounts
TCPtcp_ppstcp_mbpsAll TCP packets.
UDPudp_ppsudp_mbpsAll UDP packets.
ICMPicmp_ppsicmp_mbpsAll ICMP packets.
TCP SYNtcp_syn_ppstcp_syn_mbpsPure SYNs only (SYN flag set, ACK clear).
Fragmentsfrag_ppsfrag_mbpsNon-first IP fragments.
thresholds:
  pps: 80000
  mbps: 1000
  flows_per_sec: 35000
  tcp_syn_pps: 50000
  udp_mbps: 800
  frag_pps: 30000

A few rules govern how these behave:

  • OR semantics. Any single crossed threshold triggers detection. The core metrics and every configured per-protocol metric are evaluated independently; the first to exceed its limit opens the attack, and the triggering metric is recorded on the event.
  • tcp_syn is SYN-only. It counts packets with the SYN flag set and the ACK flag clear — the half-open packets of a SYN flood — not all TCP packets.
  • frag is non-first fragments. It counts the trailing fragments of fragmented datagrams, the hallmark of a fragmentation flood.
  • 0 or absent disables. A per-protocol metric left out of the config, or set to 0, is simply not evaluated and costs nothing.

Scope

Detection is scoped to your protected prefixes (the IP ranges in CIDR form, e.g. 203.0.113.0/24, that you protect). Only destinations inside the configured networks are ever acted on. Traffic to any other destination is still counted in Prometheus metrics for visibility, but it can never open an attack or trigger a ban.

networks:
  - 203.0.113.0/24
  - 2001:db8::/32

This is one of the enforced safety rules: a destination outside networks is out of scope for mitigation entirely. Protected prefixes must not overlap. See the Safety model for the full list of guarantees.

Carpet-bombing detection

Per-host detection is blind to a carpet bomb (subnet-spread attack): volume spread thinly across every address in a prefix, so no single host crosses its threshold even though the aggregate is huge. The optional carpet block catches it. Each monitored host's incoming rates are folded into its aggregation prefix (a /24 for IPv4, /48 for IPv6 by default), and a prefix-scoped attack is raised when the aggregate crosses carpet.thresholds and the traffic is spread across at least min_hosts distinct hosts — the fan-out gate that separates a real subnet-spread flood from one heavy host (already caught per-host).

carpet:
  aggregation_prefix_v4: 24  # fold per /24 (default)
  aggregation_prefix_v6: 48  # and per /48 (default)
  min_hosts: 10              # fan-out gate (minimum 2)
  thresholds:                # AGGREGATE volume over the prefix
    pps: 2000000
    mbps: 20000

Set the aggregate thresholds well above the per-host ones — they sum the whole prefix. A carpet attack carries scope: prefix, the aggregation CIDR in prefix, and its fan-out in hosts, alongside the usual sample and classification.

iMitigation is opt-in

Carpet attacks are alert-only by default. Set carpet.mitigation to auto-mitigate the aggregation prefix:

  • flowspec — announces filter rules matching the attack's traffic pattern across the whole prefix (surgical: legitimate traffic to those hosts keeps flowing). Needs upstreams that honour FlowSpec.
  • dataplane — installs that same vector-narrowed rule into this box's XDP data plane, so the drop happens locally and needs no router cooperation. It only sees traffic that reaches this machine's NIC, so it does not help once the uplink itself is saturated. Requires a dataplane block with enabled: true; the engine refuses the configuration otherwise rather than silently falling back.
  • blackhole — RTBHs the entire prefix (blunt: every host in it goes dark).

Every method refuses a prefix that contains a protected_whitelist address outright: a prefix mitigation cannot exempt one member and the whitelist guarantee is absolute. If a reload adds a protected address to a prefix that is already mitigated, the ban is withdrawn on the next sweep — and with dataplane, the kernel protects that host immediately, without waiting for the withdrawal, because the protected-destination check outranks every installed rule.

A prefix ban also counts its full address span (256 addresses for a /24) toward ban.max_banned_fraction and has its own carpet.max_active_prefix_bans cap. See Mitigation.

Outgoing detection

By default Kapkan watches traffic arriving at your protected hosts (direction: incoming). Add a thresholds_outgoing block — globally or per hostgroup — and it also watches traffic leaving those hosts, reporting direction: outgoing attacks. This is the signature of a compromised machine inside your network being used to attack someone else.

thresholds_outgoing:
  pps: 50000
  udp_pps: 20000

thresholds_outgoing takes the same keys as thresholds; at least one must be set. A few details matter:

  • Direction values. Attack events carry direction: incoming or direction: outgoing. Outgoing means the target host is the source of the flood.
  • Shared RTBH route. A host that is being attacked and attacking at the same time holds two independent attack records — one per direction — but shares a single RTBH route. The route is withdrawn only when the last of the two attacks ends, so clearing one direction never prematurely un-blackholes a host still active in the other.
  • Zero cost when absent. Without a thresholds_outgoing block, outgoing traffic is not even counted — there is no hot-path overhead.

!RTBH is destination-based

RTBH is destination-based: it drops traffic to the host. For an outgoing attack that takes the compromised host offline, which usually stops the abuse — but it does not, by itself, drop the outbound packets. It only stops the outbound flood if your edge separately drops packets whose source is in a blackholed prefix — e.g. with uRPF (unicast Reverse Path Forwarding, a router feature that discards packets arriving from an address it would not route back to). Set ban: false on the hostgroup if you want the alert without the route. See Mitigation.

Attack samples

When an attack opens, Kapkan attaches a flow sample to it immediately — the top sources, source ports, destination ports and protocols driving the traffic. This is possible because Kapkan keeps a continuous traffic buffer: recent flows are buffered as they arrive, so the window leading up to the threshold trip is already on hand the instant detection fires. There is no post-detection capture delay and no risk of missing the start of the attack.

samples:
  enabled: true        # default on
  buffer_flows: 65536  # rolling buffer size
  flows_per_attack: 20 # raw flow records kept per attack

The sample's counter totals (top sources, ports and protocols) are sampling-corrected; the raw flow records carry the exporter's pre-correction numbers plus their sampling_rate. The total_packets field is the untruncated corrected packet total used as the denominator for shares, since the top-K lists drop the lighter keys. Samples appear on the attack_started event, in notifications, and in the API. Buffer sizing changes require a restart.

Classification

Each attack is labelled with an inferred vector at detection time, computed from the windowed per-protocol rates and the flow sample. The classification carries a confidence — the share (0..1) of the attack traffic matching the winning signature — and, for amplification vectors, the reflected service src_port.

The classifier recognizes these type values:

TypeVector
ntp_amplificationReflected NTP responses (src_port 123).
dns_amplificationReflected DNS responses (src_port 53).
cldap_amplificationReflected CLDAP responses (src_port 389).
memcached_amplificationReflected memcached responses (src_port 11211).
ssdp_amplificationReflected SSDP responses (src_port 1900).
chargen_amplificationReflected chargen responses (src_port 19).
syn_floodTCP SYN-dominated flood.
fragment_floodNon-first IP fragment flood.
icmp_floodICMP-dominated flood.
udp_floodUDP-dominated flood with no amplification signature.
tcp_floodTCP-dominated flood.
mixedNo single vector dominates.

How a vector is chosen:

  • A signature needs a dominant share — at least half (>= 0.5) the windowed traffic — to win. When UDP is the dominant protocol, Kapkan first checks for an amplification signature (most specific). Otherwise — or if no amplification port qualifies — it falls through to pure SYN, then fragments, then ICMP, UDP and TCP volumetric floods.
  • An amplification vector also requires response-sized packets from the reflected service port: the average packet size from that source port must be at least ~200 bytes to read as a reflected response (request-sized packets from a service port classify as a plain flood instead), and that single source port must itself carry at least half the sampled packets.
  • When no signature reaches the dominant share, the attack is labelled mixed and confidence is 0.

iConfidence is a share, not a probability

confidence is the fraction of the attack traffic that matched the winning signature, in the range 0..1. A value of 0 means the attack is mixed — no signature matched at all.

Why a detection fired

Every attack carries a reason — a small record, captured the instant the threshold crosses, that explains the detection without making you re-derive the math. It lets an operator tell a real attack from a sampling or threshold artifact at a glance. It is built once per attack start from inputs the engine already has, so it costs nothing on the hot path, and it rides along on the attack_started event, in storage, and on the API's reason field.

It answers two questions:

  • Which threshold was crossed, and where did it come from? threshold_source is either static (the limit came straight from your config) or baseline (a warmed-up learned baseline set the limit). The base trio (pps, mbps, flows_per_sec) becomes baseline-derived once the baseline warms up; per-protocol metrics are always static. For a baseline threshold the reason carries the full computation — the learned normal, the factor, the floor, and the static ceiling — so you can see the effective threshold as min(ceiling, max(floor, normal × factor)).
  • Was the baseline ready? If a baseline is configured but still warming up, threshold_source stays static, warming_up is true, and warmup_remaining_seconds tells you how long until the learned threshold takes over. A detection during warm-up is therefore expected to use the static limit.

It also restates the classification inputs: shares is the per-protocol fraction of total PPS (the same ratios the classifier uses), and dominant_share_gate is the share one protocol needs to win a vector — 0.5. If no protocol reaches it, the attack is mixed. Seeing, say, a UDP share of 0.98 next to an ntp_amplification classification confirms at a glance why that vector won.

iPer-protocol metrics are always static

Baselines learn only the base trio. When a per-protocol metric (tcp_syn_pps, udp_mbps, …) trips the attack, threshold_source is static even if the group has a warmed-up baseline — there is no learned normal for that metric to compare against.

  • Hostgroups — per-prefix thresholds, total vs. per-host calculation, and per-group policy.
  • Baselines — learned per-host thresholds that tighten the static limits automatically.
  • Mitigation — how a detected attack becomes an RTBH route.
  • REST API — read active and recent attacks with their samples and classification.