Skip to content
Hogin Hogin
Go back

Cilium Tetragon: eBPF runtime security that blocks, not just logs

10 мин чтения

Falco, Tracee, and most eBPF security monitoring tools can do one thing: notice that something happened and fire an alert — after the syscall has already completed and the data may already be off the host. Tetragon, Cilium’s eBPF runtime security engine, can do more: stop a process synchronously, in the kernel, the moment a kprobe fires — before the syscall even gets to return control to the calling code. Let’s break down what a TracingPolicy is made of, how in-kernel blocking differs from a postmortem alert, and why flipping on Sigkill on day one is a way to take down production, not protect it.

Table of contents

Open Table of contents

From alert to Sigkill: why yet another eBPF security tool

Detection-only isn’t a flaw in any specific tool — it’s an architectural ceiling for an entire class of solutions. Falco hooks kprobes and tracepoints, collects events in userspace, runs them through rules, and on a match sends a notification — to Slack, to a SIEM, wherever. By the time anyone (or even an automation) reads that notification, the original syscall is long done: the file has been read, the shell has launched, the connection is established. Detection here works like a security camera — a great source for investigating an incident afterward, but not a physical barrier at the moment the violation happens.

Tetragon isn’t built differently because it uses some other eBPF hook — it’s built differently because an action can be attached directly to the interception point: a TracingPolicy with a Sigkill action terminates the process synchronously, from the kernel, before control returns from the intercepted function back to the caller. The difference isn’t reaction speed — this isn’t milliseconds-of-alerting versus milliseconds-of-blocking — it’s that by the time the event gets logged, Tetragon has already irreversibly prevented it, not merely recorded it.

Postmortem detection vs Sigkill in the kernel before the syscall completes

Why this is worth covering now, rather than back when Tetragon was first announced: over 2026 the project went from version 1.4 (February), which meaningfully smoothed out writing and debugging TracingPolicy — before it, policy authors had to guess a lot more about why a selector wasn’t matching — to 1.7 (April), with noticeably more mature tooling around enforcement. This isn’t a “feature that shipped in beta last week” — it’s a tool that, over a couple of releases, reached a point where turning on blocking in production stopped being a gamble, provided you follow the right rollout pattern (more on that below).

Anatomy of a TracingPolicy: what a rule is made of

A TracingPolicy is a custom resource (apiVersion: cilium.io/v1alpha1, kind: TracingPolicy) that describes three things: where to intercept, what to compare, and what to do on a match.

TracingPolicy anatomy: hook → selectors → matchActions

The hook point is one of several hook types: kprobes (kernel functions and syscalls), tracepoints (predefined kernel events, more stable across kernel versions than arbitrary symbols), uprobes (userspace binary functions), and the rarer fentries/lsmhooks/usdts. In practice, kprobes matter most — almost every runtime security policy is built on them: hooking execve to control process launches, hooking file descriptor operations to control file access.

Selectors filter within a hook — without them, you’d end up catching every single call to that function across the entire cluster:

matchActions is what happens once every condition in the selector matches. Post simply generates an event (this is the observation mode); Sigkill synchronously kills the process from the kernel; Override replaces the call’s return value (requires CONFIG_BPF_KPROBE_OVERRIDE kernel support) and hands control back as if the function had returned, say, -EPERM; Signal sends an arbitrary signal, but asynchronously — the operation can complete before the process receives and handles it, so Signal isn’t something to rely on for hard blocking; it’s closer to a soft nudge to the process than a guarantee it stops.

Runtime vs deploy-time: where Tetragon sits next to admission control

It’s easy to mistake Tetragon for admission control like Kyverno or ValidatingAdmissionPolicy — both ultimately “block something bad in Kubernetes.” The difference is timing, not strictness.

Deploy-time admission control vs runtime enforcement on the node

Admission control is deploy-time: kube-apiserver validates the object at kubectl apply time, before the pod is even scheduled onto a node. ValidatingAdmissionPolicy or Kyverno will happily catch a pod requesting a hostPath, running as root, or pulling an image tagged :latest — but they know nothing about what happens inside an already-approved, already-running container ten minutes after it starts. If a legitimate image’s dependency gets a supply-chain compromise that looks harmless at startup and then, mid-shift, tries to spawn /bin/sh and exfiltrate environment variables — admission control simply never sees that event: the pod already passed its check and is living its own life.

Tetragon operates at a different layer — runtime, at the node level, across the container’s whole lifecycle, not just the moment it’s created. It doesn’t know and doesn’t need to know whether a manifest matches the organization’s security policy — but it sees every execve, every file descriptor operation, every intercepted syscall inside an already-running container. That’s exactly the spot admission control structurally can’t reach — not because Kyverno or VAP are poorly written, but because they don’t re-run on every syscall inside an already-live pod, and shouldn’t, or admission would become a bottleneck for the whole cluster’s runtime. The right model isn’t “pick one” — it’s two independent layers of defense: one guards the entry point into the cluster, the other guards what happens inside after entry.

Comparison table

DimensionFalco / detection-onlyTetragon Sigkill/OverrideAdmission control (Kyverno/VAP)
When it actsafter the syscall completesbefore returning from the intercepted functionbefore the object is created in the API
What it seesprocesses, files, network on the nodethe same, plus the ability to actthe resource manifest
Reaction to a matchuserspace alertin-kernel blockAPI request rejected
False-positive riska noisy alerta killed legitimate processa rejected deploy
Protects againstknown activity, after the factexploitation already inside the containeran unsafe manifest at the door
Requires an eBPF-capable kernelyesyesno

What it takes to roll out a policy the right way

The goal: block /bin/sh and /bin/bash from launching inside containers labeled security.hogin.pro/immutable: "true" — a typical case for services that have no legitimate need for an interactive shell, ever.

Start with a policy that has no enforcement — observation only:

apiVersion: cilium.io/v1alpha1
kind: TracingPolicy
metadata:
  name: "block-shell-in-immutable"
spec:
  podSelector:
    matchLabels:
      security.hogin.pro/immutable: "true"
  kprobes:
  - call: "sys_execve"
    syscall: true
    args:
    - index: 0
      type: "string"
    selectors:
    - matchArgs:
      - index: 0
        operator: "Prefix"
        values:
        - "/bin/sh"
        - "/bin/bash"
      matchActions:
      - action: "Post"

matchActions: [Post] is the audit mode: the selector is already fully configured and matches exactly the events it will later block, but for now it only generates an event without stopping anything. Apply it and watch the stream of matches through tetra:

kubectl apply -f block-shell-in-immutable.yaml

tetra getevents -o compact --namespace prod \
  --pod-labels security.hogin.pro/immutable=true
# 🚀 process   prod/api-6f9. /bin/sh -c "healthcheck.sh"
# 💥 exit      prod/api-6f9. /bin/sh     0

This is exactly the step where the whole point of audit mode tends to surface: a healthcheck script, a sidecar, or a CI runner that legitimately calls /bin/sh and that nobody remembered when designing the policy. There are two honest options here: narrow the podSelector/containerSelector to exclude that container from the policy, or admit that the immutable label was applied to that workload too early. Either beats discovering that /bin/sh call the moment Sigkill has already killed the healthcheck and the pod is spinning into CrashLoopBackOff.

Once the event stream over an observation window (realistically, days to a couple of weeks under real load, not five minutes in staging) shows nothing but confirmed unwanted activity, flip matchActions in that same policy:

      matchActions:
      - action: "Sigkill"
kubectl apply -f block-shell-in-immutable.yaml

The policy’s structure, its selectors, its podSelector — all of that stays exactly the same. One action, in one place, changed. That’s the whole point of the pattern: audit and enforcement aren’t two different resources or two different Tetragon modes — they’re the same policy with a different matchActions, which makes the switch a reversible, local change rather than a migration to a different system.

How to verify it works

kubectl exec -n prod api-6f9xy -- /bin/sh -c "echo test"
# command terminated with exit code 137

tetra getevents -o compact --namespace prod
# 🚀 process   prod/api-6f9. /bin/sh -c "echo test"
# 💀 kill      prod/api-6f9. /bin/sh                    SIGKILL

Exit code 137 (128 + signal 9) plus a kill event with SIGKILL in the Tetragon stream confirm the block fired synchronously from the kernel, not through some delayed external controller. Legitimate traffic from the same pod — HTTP requests, a healthy execve of allowed binaries — keeps working as usual: the policy only matches sys_execve with an argument starting with /bin/sh or /bin/bash, and touches nothing outside that condition.

Bottom line

eBPF runtime security stops being a nice source of logs the moment matchActions in a TracingPolicy flips from Post to Sigkill — but that switch has to be earned with an observation period against real traffic, not flipped on day one out of enthusiasm. Tetragon doesn’t replace admission control like Kyverno or VAP, and isn’t replaced by it either — they protect different moments in a workload’s life, the entry into the cluster and what happens after, and ideally run together. The cost of kernel-level enforcement is real: a badly calibrated selector kills your own healthcheck, not an attacker. What you get in return is the ability to stop exploitation before it gets to do anything, rather than write a postmortem about it.


Share this post:

Next Post
containerd 1.x is done: what breaks moving to containerd 2.0 before Kubernetes 1.36