Runtime Security Monitoring for Kubernetes Workloads

Kernel-level visibility catches container threats that detection at deployment time cannot.

Reporter · · 10 min read
Cover illustration for “Runtime Security Monitoring for Kubernetes Workloads”
Security Posture · October 4, 2026 · 10 min read · 2,255 words

Runtime security monitoring for Kubernetes workloads answers a question that image scanning and admission control were never built to answer: what happens after a container starts running. Static checks and policy gates do their job at the front door, but a workload that gets through cleanly can still turn hostile once it's live, and closing that gap is now a matter of kernel-level visibility rather than better front-door rules.

Where admission control and image scanning leave a gap

The standard Kubernetes security stack, built from image scanning, admission control, and RBAC hardening, is designed to evaluate what enters a cluster. None of it is built to watch what happens once a workload is inside. That's a structural boundary, and it's the root cause of what gets called the post-deployment detection gap.

Image scanning checks a static artifact against known vulnerabilities, but only at one point in time. It can tell you a container image has a patched or unpatched CVE. It has nothing to say about what that container's process actually does once it starts executing.

Admission control works at the API server, so it enforces policy on create and update events. Once a pod clears admission and gets scheduled, the admission controller's job is done. A workload that passes every check and then behaves maliciously, whether through a logic bug, a compromised dependency, or a supply-chain backdoor planted upstream, is simply invisible to it from that point forward.

The division is clean once you see it: pre-runtime controls protect against threats that are detectable before a container runs, and runtime monitoring protects against what happens after. No amount of tightening on the pre-runtime side substitutes for the second half. A workload can pass every gate without incident and still spawn a reverse shell, write to filesystem paths it has no business touching, escalate its own privileges, or open network connections to places it was never supposed to reach, all after it's already live. None of that is a failure of image scanning or admission control. It's simply outside the range of what those tools were built to see.

How eBPF moves runtime observation to the kernel

eBPF closes that gap by moving the observation point down to the kernel itself. It works by attaching small programs to tracepoints, kprobes, and LSM hooks, which gives it a direct view of system call activity, network connections, file access patterns, and process execution as they happen, not after the fact.

The structural advantage is what matters most here. Because eBPF sensors run on the host kernel, they sit below the container abstraction layer. A process that manages to escape its container boundary can't evade a sensor that was never relying on the container boundary to begin with, because the observation happens at a layer the container runtime doesn't control.

Compare that to userspace monitoring, which typically watches what a container reports about its own state. An attacker who breaks out of a container can silence or falsify those reports from the inside. An eBPF sensor watching from the host kernel observes the escape itself as it happens, independent of anything the compromised container claims is going on.

System calls like ptrace, mount, unshare, and keyctl appear repeatedly in container escape exploits because they touch namespace and capability boundaries directly. At the kernel level, these calls are visible regardless of what the container runtime believes is happening inside its own boundary.

eBPF maturity in 2026

Eyeing eBPF-based runtime security as a cutting-edge bet made sense a few years ago, but it no longer does.

Production adoption backs this up. GitLab, Shopify, and Skyscanner all run Falco, the leading open-source cloud native runtime security project, in production. These are organizations operating at real scale, with diverse workloads and high-availability requirements that leave little room for tooling that isn't dependable.

The clearest evidence of the shift is the CVE-to-mitigation pattern itself. Disclosures that showed admission-time controls missing post-deployment attacks pushed teams to treat kernel-level runtime visibility as a requirement, not an enhancement. That's the kind of pressure that turns a promising technology into a default one.

Falco's detection model: rules, syscall events, and MITRE mapping in practice

Falco takes the raw syscall visibility eBPF provides and turns it into something a security team can actually act on. It matches a continuous stream of runtime events against a rule engine, and that engine maps suspicious behaviors, things like reverse shells, unexpected privilege changes, and unauthorized file writes, to named threat patterns.

Rule coverage maps to the MITRE ATT&CK framework. That matters because it changes what an alert tells a responder. Instead of a bare notice that something happened, the alert arrives with context: this activity matches lateral movement, or this matches privilege escalation, or this matches defense evasion.

Falco's detection scope across containers and Kubernetes is broad and genuinely strong. It is not the same thing as a full commercial platform, though. Commercial offerings add response automation, a user interface, alert correlation across many signals, and in-kernel enforcement that Falco doesn't provide on its own. What Falco doesn't ship with natively is the operations layer: the dashboards, the compliance reporting, the correlation logic, and the automated response workflows that commercial products build on top of the same underlying detection engine.

The honest cost framing matters here too. Falco is free to use. Running Falco in production is a different kind of expense. The real investment goes into tuning rules so alert volume stays manageable, and into wiring Falco into a SIEM and staffing the team that responds when an alert fires.

Why detection alone is not enforcement

The strongest objection to stopping at detection is simple: an alert generated after a malicious action has already completed hasn't prevented anything. Detection tells a team what happened. It doesn't close the window between the action and the response.

Tetragon takes a different architectural position on this exact question. It loads security logic directly into the Linux kernel via eBPF, using hook points that include LSM hooks, kprobes, and tracepoints, and that placement lets it kill a process, block a syscall, or drop a network packet before the action completes, rather than logging it after the fact.

That's a genuine architectural difference from Falco, not a marketing distinction. Falco observes and alerts. Tetragon can enforce in-kernel, which makes the runtime security posture active instead of reactive.

That capability carries a real trade-off. In-kernel enforcement is a more demanding operational posture to run, because an incorrectly written enforcement rule can block legitimate workload behavior and take down something that was never actually a threat. Teams adopting Tetragon typically start in detection mode and promote individual rules to enforcement only once they've proven out against real traffic.

Choosing between a detection-first tool like Falco and an enforcement-capable tool like Tetragon isn't about picking a winner. It comes down to the threat model a team is defending against, how mature the team's operational practices are, and how much tolerance there is for an enforcement rule misfiring while it's still being tuned.

How commercial platforms layer on top of the open-source detection core

Commercial Kubernetes runtime security platforms don't replace the eBPF detection core that Falco and Tetragon represent. They build the operations layer most teams can't cost-effectively build and staff on their own: the UI, the alert correlation across signals, the compliance reporting, and the automated response workflows.

Aqua Security takes a full lifecycle approach, connecting pre-deployment image scanning through Trivy with drift prevention, supply-chain controls, admission policies, and runtime protection inside a single platform. That matters most to teams that want one vendor spanning the whole build-to-run boundary, instead of stitching several tools together themselves.

Red Hat Advanced Cluster Security, built on StackRox heritage, works deeply with Kubernetes objects directly: RBAC, network policies, admission controls, and runtime behavior all inside one model. That depth of integration is especially relevant for organizations running OpenShift.

The cost crossover tends to follow team size and cluster count. A small security team running a handful of clusters often comes out ahead building on open-source tooling plus operational discipline. A larger team running many clusters at scale typically comes out ahead paying for commercial tooling, even at a meaningful license cost, because the labor the platform replaces becomes the larger expense once headcount and cluster count both grow.

The expanded attack surface GPU and AI workloads introduce at runtime

GPU-backed inference and training workloads present a different runtime security surface than a typical stateless service. These workloads run with high privilege, hold GPU access for long stretches of time, and sit on hardware expensive enough to make GPU nodes a prime target for cryptomining and data exfiltration.

Microsoft documented a campaign that abused Kubeflow ML pipelines and legitimate TensorFlow images, including tensorflow/tensorflow:latest and tensorflow/tensorflow:latest-gpu, to deploy Monero and Ethereum miners across Kubernetes clusters. That incident shows AI infrastructure is being exploited right now, not some hypothetical risk you can plan around someday.

At the hardware isolation layer, NVIDIA MIG, available on supported GPU generations, splits a single GPU into isolated instances, each with its own dedicated compute and memory. That isolation cuts the risk of one tenant snooping another tenant's GPU memory, compared to software-only sharing approaches like time-slicing.

AI workloads also tend to cross compliance boundaries in ways other workloads don't. Model training often runs in the cloud to get GPU availability, while data processing stays on-premises because of compliance reasons. So runtime security has to span multiple environments at once, and a single-cluster monitoring deployment simply can't cover that.

The operational shape of these workloads is changing too. Robinhood adopted KubeRay with an ephemeral cluster-per-job pattern: it spins up a distributed Ray cluster, runs the job, then tears the cluster down. Traditional persistent-DaemonSet monitoring assumes nodes that stick around. Ephemeral GPU clusters need monitoring that provisions and deprovisions on the same schedule as the workload itself, which is a meaningfully different operational demand than monitoring a long-lived service.

Where supply chain integrity ends and runtime monitoring begins

Runtime monitoring and supply chain integrity sit at different points on the attack timeline, and they reinforce each other rather than duplicate each other's work. Supply chain controls constrain what's allowed to enter a cluster. Runtime monitoring catches what those controls miss, or what changes after the workload has already been admitted.

The SLSA framework (Supply-chain Levels for Software Artifacts), paired with attestations signed through Sigstore or a platform-controlled signing key, gives cryptographic proof of build provenance: where an artifact was built, from what source, and by which build system. That proof lets an admission controller enforce a policy as specific as accepting only images built by a team's own CI/CD system from its own trusted repositories.

Even a well-built provenance chain has a limit. A backdoored image can pass every supply chain check if it was legitimately built from a source repository that was already compromised upstream. That image still has to run eventually, and once it does, it tends to show anomalous behavior: unexpected process spawning, unauthorized outbound connections, file writes that don't match the workload's normal pattern. That's the behavior you catch with eBPF-based detection after admission, even when the artifact itself looked clean going in.

The relationship runs in both directions. Vulnerabilities and misconfigurations caught in production runtime environments can feed back into CI/CD pipelines as updated checks, preventing the same class of issue from shipping again. Runtime detection becomes an input to how the build process evolves.

Continuous compliance and the role OPA Gatekeeper v4 plays in closing configuration drift

Admission-time policy enforcement creates a kind of false confidence. It validates a resource the moment it's created, but it has nothing to say about what happens afterward: manual kubectl patches, controller reconciliation loops, or policy changes that quietly make a previously-compliant resource non-compliant with no new admission event to catch it.

OPA Gatekeeper v4 reached general availability in February 2026, and it introduces continuous validation: the ability to periodically re-evaluate every existing resource against the current policy set, not just resources that are newly created or updated. That closes the drift gap that admission-only enforcement leaves open.

For regulated workloads, the highest-leverage move is often narrower than it sounds: limiting PHI to clearly bounded namespaces or clusters. That shrinks the compliance surface instead of trying to apply uniform controls across an entire cluster, which is harder to sustain and harder to audit.

Regulatory expectations in 2026 have shifted toward continuous, automated compliance evidence rather than periodic manual assessments. That shift is what makes a tool like Gatekeeper v4 directly relevant to audit readiness, not simply a convenience for the engineering team running the cluster day to day.

A practical runtime security stack for a small engineering team

A small engineering team doesn't need to build a comprehensive security engineering project to get real protection. A layered posture with low operational overhead covers most of the realistic risk a small team actually faces: posture scanning to catch misconfigurations before they ship, runtime monitoring to catch what happens after deployment, and RBAC controls sized sensibly to the team's actual access needs.

Each layer does a distinct job, and none of them substitutes for the others. Posture scanning catches what's wrong before a workload runs. Runtime monitoring catches what posture scanning structurally cannot see. RBAC limits what any single compromised identity or workload can reach, regardless of what else goes wrong. Put together, at a scale a small team can actually tune, staff, and maintain, that combination handles the overwhelming majority of risk a team running production Kubernetes will encounter.

Sources

  1. 10 Open Source Kubernetes Security Tools for 2026
  2. What Is eBPF Observability and How Does It Work? · Dash0
  3. Kubernetes Runtime Security : The Guide
  4. Tetragon - eBPF-based Security Observability and Runtime Enforcement
  5. Falco
Filed underSecurity Posture

More in Security Posture