GitHub

RTBH mitigation

BGP (Border Gateway Protocol — how routers exchange routes) is what all of Kapkan's mitigation rides on: Kapkan acts as a BGP speaker that tells your routers which routes to install. Kapkan tracks each mitigation decision as a ban; for the default method a ban is a blackhole route (RTBH).

When the engine reports an attack, Kapkan mitigates it by announcing a remotely-triggered blackhole (RTBH) route for the targeted host: a /32 (IPv4) or /128 (IPv6) prefix carried over an embedded GoBGP speaker, tagged with your RTBH community and pointed at a discard next-hop — an address your routers send blackholed traffic to and then drop (a null route). Your edge routers match the community and drop all traffic to that host, dropping the flood before it reaches your network's core.

Mitigation is destination-based: the blackhole drops traffic to the target. For an incoming flood this stops the attack. For an outgoing attack from a compromised host it takes that host offline, which usually stops the abuse — but it does not directly drop the outbound packets unless your edge also filters packets whose source is in a blackholed prefix, e.g. with uRPF (unicast Reverse Path Forwarding, a router feature that drops spoofed/return-path-mismatched traffic). To get the alert without the route, set ban: false on the hostgroup.

!Dry-run by default

Until you explicitly set dry_run: false, every would-be blackhole is computed, logged and exposed through the API — but never announced to your routers. Validate detection and BGP peering against production telemetry before any route can be sent. See Going live.

Mitigation methods

blackhole (RTBH) is the default method and the subject of the rest of this page. Other methods are available via the mitigation key, each overridable per hostgroup:

  • FlowSpec (mitigation: flowspec) — drop only the attack vector with BGP FlowSpec rules (RFC 8955/8956) instead of blackholing the whole victim.
  • Traffic diversion (mitigation: divert) — announce the victim toward a scrubbing center so its traffic is cleaned and reinjected rather than dropped.
  • Escalation ladders (escalation:) — step the response up the longer an attack persists, from alert to FlowSpec to divert to blackhole.

All share the same ban lifecycle, TTL, unban hysteresis and max_active_bans cap described below.

BGP configuration

The bgp block defines Kapkan's BGP identity, the blackhole next-hops, the RTBH community, and the eBGP peers it announces to.

bgp:
  local_asn: 65000               # Kapkan's local ASN
  router_id: "198.51.100.1"      # must be a valid IPv4 address
  next_hop: "192.0.2.1"          # IPv4 discard next-hop
  next_hop6: "100::1"            # IPv6 discard next-hop (optional)
  community: "65000:666"         # RTBH community, ASN:value
  # communities: ["65000:666", "65000:777"]  # full set; overrides `community`
  # local_pref: 100              # optional LOCAL_PREF for iBGP peers; 0 = omit
  neighbors:
    - address: "198.51.100.254"
      remote_asn: 65001
      port: 179                  # optional; defaults to 179, override for testing
KeyMeaning
local_asnKapkan's local ASN for the eBGP sessions. Required.
router_idThe BGP router ID. Must be a valid IPv4 address.
next_hopThe IPv4 discard next-hop announced with each /32 blackhole.
next_hop6The IPv6 discard next-hop announced with each /128 blackhole. Optional; when unset, Kapkan falls back to the RFC 6666 discard address 100::1 (from the 100::/64 discard prefix).
communityThe RTBH community in ASN:value form (for example 65000:666). Parsed once at load into the wire value sent with every route.
communitiesOptional list of communities that overrides community when set, for upstreams that expect a full set.
local_prefOptional LOCAL_PREF path attribute attached to blackholes. It only travels within your own network (iBGP — internal BGP), not to external peers (eBGP — external BGP), so it is meaningful only to your internal neighbors. Default 0 omits it.
neighbors[]The BGP peers Kapkan announces to, each either eBGP or iBGP. Each entry takes an address, a remote_asn, and an optional port (default 179, intended for testing).

community, communities, next_hop, next_hop6 and local_pref can all be overridden per hostgroup — see Per-hostgroup BGP attributes.

The blackhole route Kapkan builds for a target reads roughly as blackhole <prefix> next-hop <next_hop> community <community> — the string surfaced in the route field of the ban object. The community segment is dropped if no community is set, and a local-pref <n> segment is appended when local_pref is configured. IPv6 targets use next_hop6 (or the 100::1 fallback) in place of next_hop.

iPeering happens in dry-run

The BGP speaker peers with your neighbors even while dry_run is true — only route announcements are gated on the flag. This lets you confirm sessions reach ESTABLISHED (logged as bgp peer state) and validate connectivity before you announce a single route.

Per-hostgroup BGP attributes

The global bgp block sets the default blackhole next-hops and RTBH community. A hostgroup can override any of them, so different customers signal their own upstreams:

bgp:
  next_hop: "192.0.2.1"
  community: "65000:666"                          # the default RTBH community

hostgroups:
  - name: customer-a
    networks: ["203.0.113.64/26"]
    bgp:
      communities: ["65000:100", "65001:200"]     # customer-A's own blackhole signal
      next_hop: "192.0.2.50"                       # and discard next-hop
      local_pref: 250

Each field left unset inherits the global bgp value, so a group can override just its community while sharing the global next-hop. The resolved attributes are frozen on each ban when it is created: a config reload changes only future bans, never the route a live ban already announced. The per-ban next_hop, community and local_pref are visible in /api/v1/bans and in the route field.

The ban lifecycle

A ban is the unit Kapkan tracks for each blackhole decision. Its lifecycle is bounded at every step by the safety model — there is no path to a permanent or runaway ban.

  1. Announce on detection. When the engine reports a new per-host attack and the host's policy permits banning, Kapkan announces the blackhole route (or, in dry-run, records the would-be route) and the ban enters the active state. A second attack — for example an incoming and an outgoing flood on the same host at once — shares the one RTBH route rather than announcing twice.
  2. Auto-withdraw on TTL. Every announcement carries a TTL. While the attack is still detected, each detection window refreshes the ban's TTL (never beyond a fresh ban.ttl_seconds from now), so the route stays up for the whole attack. Once detection stops, the route is auto-withdrawn within one TTL. There are no permanent bans.
  3. Unban with hysteresis. A ban driven by an attack is withdrawn only after traffic stays below threshold for ban.unban_hysteresis_seconds. This anti-flap delay prevents a borderline attack from rapidly announcing and withdrawing the same route.
  4. Hard cap. ban.max_active_bans is an absolute ceiling on simultaneous active bans. Past the cap, new bans are refused (returned as a ban in the rejected state with reason max_active_bans reached) and operators are alerted — so Kapkan can never blackhole half your network in a broad attack.
ban:
  ttl_seconds: 1800            # every announcement auto-withdraws after this
  unban_hysteresis_seconds: 60 # traffic must stay below threshold this long before withdraw
  max_active_bans: 100         # hard cap on simultaneous bans

A ban is also withdrawn early if a config reload removes its target from the protected networks — Kapkan will not keep a route up for address space it no longer protects.

Persistence across restarts

A restart — a kernel upgrade, a new Kapkan version — would normally drop every active route and leave the victims exposed until the engine re-detects each attack. Two complementary, optional mechanisms keep mitigation up across that gap:

  • Ban persistence (ban.state_file). Point it at a writable path (for example /var/lib/kapkan/bans.json) and Kapkan writes its active bans to disk. On the next start it rehydrates them and re-announces each route immediately, instead of waiting for the attack to trip again. Empty (the default) disables persistence. The systemd unit provides the directory via StateDirectory=kapkan.
  • BGP Graceful Restart (bgp.graceful_restart). On by default. A peer that supports Graceful Restart (RFC 4724) retains Kapkan's blackhole routes as stale while the session is down, rather than flushing them the instant the BGP connection drops — so the blackhole stays in effect during the brief restart window. Advertising the capability is harmless to a peer that does not support it.
ban:
  state_file: /var/lib/kapkan/bans.json   # persist & re-announce active bans across a restart

bgp:
  graceful_restart:
    enabled: true              # default; set false to opt out
    restart_seconds: 120       # how long the peer holds routes as stale (0..4095; 0 = use the default 120)
    # long_lived: false        # LLGR (Long-Lived Graceful Restart): extend retention beyond restart_seconds
    # long_lived_stale_seconds: 3600  # LLGR stale time when long_lived is set (<= 86400)

Graceful Restart bridges the session gap, but on its own a peer purges the stale routes once Kapkan reconnects and signals End-of-RIB (the marker that says "I've finished re-sending my routes"). Pair it with state_file so the restarted Kapkan re-announces the bans before that happens: used together, the bans survive the process restart (state_file) and the routes survive the session gap (graceful restart), so an upgrade no longer un-mitigates an in-progress attack.

Fallback when a method is rejected

Upstreams vary in what they honor — a peer may reject a FlowSpec or divert announcement outright. By default Kapkan does not leave the victim undefended in that case: ban.fallback is blackhole, so a rejected flowspec/divert announce degrades to an RTBH route rather than failing the ban. Set ban.fallback: none to disable that and reject the ban instead.

ban:
  fallback: blackhole   # degrade a rejected flowspec/divert announce to RTBH (default); none disables

When a fallback happens the ban records fell_back_from (the method the peer rejected) and the kapkan_mitigate_fallback_total metric increments — a non-zero from="flowspec" series flags upstreams that do not honor FlowSpec. A total BGP outage fails both methods, so the ban is still rejected as before; fallback only changes the case where the surgical method specifically is unsupported.

Dry-run

In dry-run mode (dry_run: true, the default) Kapkan does everything except send the route. The target prefix, next-hop, community and the full route string are computed, the decision is logged as DRY-RUN: would announce mitigation (not sent) (and the matching withdraw line DRY-RUN: would withdraw mitigation (not sent)), and the ban is tracked in the API with dry_run: true and the live active/withdrawn states. TTL expiry and hysteresis run exactly as they would in production, so the entire lifecycle is observable before any route reaches a router.

Nothing is announced until you set dry_run: false and reload (SIGHUP or POST /api/v1/config/reload). Because peering happens in dry-run, the only behavior that changes when you flip the flag is the announce/withdraw itself.

iValidate first

Run against production telemetry in dry-run, confirm detections fire on the right prefixes and the route fields are correct, then turn off dry-run. The full procedure is in Going live.

Manual bans

Operators can blackhole or release a host directly through the REST API. If you are mid-attack right now, start with the under-attack runbook, which walks the ban-a-host sequence step by step.

POST /api/v1/ban
Content-Type: application/json

{"ip": "203.0.113.66"}
POST /api/v1/unban
Content-Type: application/json

{"ip": "203.0.113.66"}

A manual ban is created with manual: true and follows the same TTL (so it still auto-expires — there are no permanent bans), but it is released only by an explicit unban or TTL expiry, never by traffic dropping below threshold. It honors every safety rule with no exceptions: a whitelisted target is refused (returns 409, reason whitelisted), a target outside the configured networks is refused (returns 409, reason outside configured networks), and any blast-radius cap — max_active_bans, max_banned_fraction or max_bans_per_window — is refused (returns 409). The reason field names which rule fired; see the safety model. Both POST endpoints require the application/json content type. See the API reference for the full request and response shapes.

The ban object

The API returns each ban as a JSON object. These are the fields exactly as serialized by the mitigator:

FieldTypeMeaning
targetstringThe blackholed host address (IPv4 or IPv6).
prefixstringThe host prefix announced: /32 for IPv4, /128 for IPv6.
metricstringThe metric that triggered the attack-driven ban (omitted for manual bans).
ratenumberThe observed rate at ban time, in the metric's units (omitted when zero).
thresholdnumberThe threshold the rate crossed (omitted when zero).
next_hopstringThe discard next-hop announced with the route.
communitystringThe community set attached to the route (the RTBH community for a blackhole, the divert community for a divert ban), space-joined when more than one.
local_prefnumberThe LOCAL_PREF attached to the route (omitted when zero).
routestringThe full route as blackhole <prefix> next-hop <next_hop> community <community> (or divert ... for divert bans), or a flowspec: ... summary for FlowSpec bans.
statestringOne of active, withdrawn, or rejected.
dry_runbooleanWhether the ban was made in dry-run mode (route not actually sent).
manualbooleantrue for operator-requested bans, false for automatic ones.
started_attimestampWhen the ban was created.
expires_attimestampWhen the TTL auto-withdraws the route.
withdrawn_attimestampWhen the route was withdrawn (omitted while active).
reasonstringWhy the ban was withdrawn or rejected (omitted while active).
methodstringThe mitigation method that produced the ban: blackhole, flowspec, or divert.
flowspecarrayThe generated FlowSpec rules, when the method is flowspec.
escalationarrayThe configured escalation ladder, when one is set.
escalation_stepnumberThe index of the ladder's current rung.
fell_back_fromstringThe method whose announce the peer rejected, causing the ban to degrade to its fallback (omitted when no fallback occurred).
fell_back_reasonstringThe human-readable cause of the fallback (omitted when no fallback occurred).

A rejected ban is never announced — it records a refusal (whitelisted, outside networks, or cap reached) so the decision is auditable in the API and notifications.

  • FlowSpec mitigation — surgical drops that spare the victim's other traffic.
  • Traffic diversion — divert to a scrubbing center instead of dropping.
  • Escalation ladders — step the response up the longer an attack persists.
  • Safety model — the dry-run, TTL, hysteresis, cap and whitelist rules enforced in code.
  • Going live — validate detection and peering, then turn off dry-run.
  • REST API — the ban endpoints, request shapes and response codes.
  • Under attack now — the step-by-step runbook for an attack in progress.
  • Troubleshooting — symptom-keyed fixes when something is not working.
  • Glossary — BGP, RTBH, uRPF and the other terms used on this page.