Container Image Hardening for Production Workloads

Reduce image attack surface with minimal base images and runtime controls.

Staff Writer · · 11 min read
Cover illustration for “Container Image Hardening for Production Workloads”
Security Posture · October 5, 2026 · 11 min read · 2,436 words

A container image is an inherited environment, and every workload built on top of it trusts that inheritance by default. That's what makes image security different from application security: a flaw in a base image doesn't wait for a developer to write a bad line of code. It's already there, baked into every layer, before a single feature gets built.

The reuse pattern makes this worse. Teams don't build one image for one service. They pull a base image once and build dozens of services on top of it, so an unpatched vulnerability doesn't scale with how dangerous it looks in isolation. It scales with how many services quietly depend on it. If a flaw seems minor in a security scanner's output, it can still sit underneath an entire company's infrastructure.

So fixing an image problem looks nothing like fixing a code bug. Patching one service and shipping it doesn't close anything. The base itself has to change, and every service built on it has to rebuild and redeploy. Image hardening is a standing discipline, applied at build time, ship time, and runtime, continuously, not a box checked once before launch.

Base image choice and its ceiling on everything that follows

Before a team writes an application, the base image has already decided how many known vulnerabilities ship by default, how much an attacker can do after breaking in, and how much work compliance reviews will take later. Standard public images often carry a large number of known CVEs out of the box. Minimal images cut that number dramatically. That gap shapes everything downstream: a scanner, a patching cadence, an incident response plan, all of it inherits whatever the base image started with.

A few paths cover most real situations, and none of them is universally right.

Google Distroless images run between 2 and 230 MB, and they include only the application and the runtime dependencies it needs to execute, nothing else. No shell. No package manager. An attacker who breaks into a Distroless container can't spawn a shell, can't install new tools, and has a hard time getting data out. For production Java, Python, Node.js, and Go workloads, that makes Distroless a strong default.

Chainguard images take a different approach built around freshness and proof. They're rebuilt daily with the latest security patches, they ship with a software bill of materials (SBOM) attached to every image, and they meet FIPS compliance requirements. That combination matters most in regulated environments, where auditors want documented evidence of what's in an image and when it was last patched, not just a low CVE count.

Teams run into trouble with a specific set of choices: the latest tag, which is mutable and can point to a different image tomorrow than it did today; images from publishers nobody can identify; and images built on operating systems that have reached end of life and stopped receiving patches. None of these belong in production, regardless of how convenient they seem during development.

The common objection is that an application needs a full-featured operating system environment, with real package managers and system libraries, to even build correctly. That's a fair concern, but it has a structural answer rather than a compromise on the runtime image: use the full-featured environment only to build the application, then carry just the finished artifact into a minimal image for production. That split is what multi-stage builds are for.

What multi-stage builds remove from production images

Base image choice sets what you start with. Multi-stage builds decide what makes it to production, and that's just as consequential. A multi-stage build uses one image, often something like the Go SDK or a full Node build environment, to compile and assemble the application. Then a second, separate stage copies only the finished binary or bundle into a minimal runtime image. Everything used to build the application, compilers, package managers, test frameworks, shell utilities, stays behind in the first stage and never reaches production.

A Go service shows the effect clearly. Go compiles to a single statically linked binary. Copy that binary into a Distroless image, and the resulting production image has no shell, no package manager, and no build tools of any kind. An attacker who manages to exploit the running service has almost nothing to work with afterward. Post-exploitation capability drops to nearly zero because the environment it would run in has been stripped bare, not because the vulnerability was prevented.

Two smaller choices inside the build configuration reinforce this. Use COPY instead of ADD when moving files into an image, since ADD can pull from remote URLs and automatically unpack archives, letting files enter the build from outside anyone's direct control.

The second choice concerns how images are referenced once they reach production: by digest, not by tag. A tag like nginx:1.25 can point to a completely different image tomorrow than it does today, because tags are mutable labels, not fixed identifiers. A SHA256 digest is immutable. Pinning to a digest guarantees that the exact image tested in staging is the exact image running in production, with nothing substituted in between.

The runtime controls that limit what a compromise can do

Everything up to this point reduces what ships inside an image. Runtime controls take over from there, by assuming that some exploit, somewhere, eventually succeeds, and asking what it can actually reach once it does. That's the shift from prevention to containment, and it's the first real application of defense in depth in the hardening process: you've already reduced what's in the image, now reduce what it's allowed to do.

Running as root is the single most common and most damaging misconfiguration in production containers. A rooted container that manages to escape its namespace carries host-level access with it, turning a contained compromise into a host compromise. Creating a dedicated non-root user for the application, and enforcing it in the image and in the deployment configuration, closes that path directly.

A read-only root filesystem closes a related path. If the container filesystem can't be written to, an attacker can't drop new tools, rewrite existing binaries, or set up anything persistent inside the container. Applications that genuinely need to write data, logs, temp files, caches, get a specific writable volume mounted for that purpose, and nothing else stays writable.

Seccomp, AppArmor, and SELinux profiles filter which system calls a container process can make to the kernel. Kubernetes includes a RuntimeDefault seccomp profile that blocks a wide range of dangerous syscalls, but it isn't automatic. It has to be turned on explicitly, either with the --seccomp-default kubelet flag at the node level or set directly in a pod's security context. Teams that assume this protection exists by default are often running completely unconfined.

Port binding deserves specific attention because the default behavior surprises people. Publishing a container port binds it to every network interface unless told otherwise. A database exposed this way becomes reachable from the public internet with no change at all to cloud security group rules, because the exposure happened inside the container's own configuration. Binding to loopback unless a service genuinely needs to be public closes that gap quietly and permanently.

The common pushback here is that an application needs elevated capabilities to function. For most stateless API services, that's rarely actually true. The right test is to drop all capabilities, run the application, and add back only the specific ones that break something, rather than keeping the defaults because removing them hasn't been tried.

Secrets in images: a permanent problem, not a deployment mistake

A secret baked into an image layer doesn't go away when a later step deletes the file that held it. Docker's layer caching means that a credential passed through a --build-arg or copied in as part of a .env file gets recorded permanently in the image's build history, even after a subsequent layer removes the file from the final filesystem. Anyone with access to the image can run docker history and pull that secret back out, long after the team assumes it's gone. That permanence is what separates this from an ordinary mistake. It's a structural property of how image layers work. The fix has to happen at the build mechanism, not in a developer's checklist.

BuildKit secret mounts solve the problem at its source. A secret gets injected into the build environment only for the duration of a single RUN command and is never written into any image layer. The credential does its job during the build and then disappears, leaving nothing recoverable in the image.

For secrets a running application needs, dedicated secret managers do the job that environment variables can't. HashiCorp Vault, AWS Secrets Manager, Azure Key Vault, and Google Secret Manager all provide encryption at rest and in transit, access control, audit logging, and automatic rotation. Vault goes a step further with dynamic secrets that expire automatically after use, a capability the cloud-native managers don't offer.

Environment variables sit in a middle ground that teams often treat as safe when it isn't. Any process running inside the container can read them, many logging frameworks capture them without being told to, and anything with the right access can see them through /proc. They're fine for configuration values that carry no sensitivity. They're the wrong place for credentials.

Catching a secret earlier, before production, is what prevents it from being recoverable from the image later. Secret scanning belongs in CI, not as a cleanup step after something leaks. Tools like Trivy, run with its --scanners secret mode, catch hardcoded credentials before an image ever reaches the registry.

The software supply chain as the largest source of image risk

Most of what sits inside a production image wasn't written by the team running it. It was written upstream, built somewhere else, and pulled in automatically as a dependency. So the integrity of an image depends on a chain of custody the team deploying it never directly controlled, across every package, library, and base layer involved.

This isn't a hypothetical risk. Public package registries, npm, PyPI, Maven Central, NuGet, and Hugging Face among them, have all seen a growing volume of malicious packages uploaded and distributed to unsuspecting developers. CI/CD systems add a second exposure point of their own. A compromised workflow or action with access to production deployment credentials can inject malicious behavior directly into a pipeline. The GhostAction compromise of 2025 showed how this plays out at scale: attackers compromised maintainer accounts across hundreds of repositories, stealing credentials from CI/CD pipelines over a window of several days before the activity was caught and remediated.

Automated dependency updates cut both ways. Pulling in new package versions quickly keeps software current, but it also means a malicious or compromised version can reach a production build before anyone on the team has reviewed it.

Provenance verification is the direct response to this exposure. If you cryptographically sign every image and bind that signature to a specific build workflow and source commit, you can catch a tampered artifact before it ever runs. Keyless signing with Cosign makes this practical at scale: each signing event gets tied to a build identity and logged in a public transparency log, so there's no long-lived signing key for an attacker to steal or misuse.

The SLSA framework as a build integrity standard teams can adopt

Provenance as a concept is easy to agree with and hard to operationalize without a concrete target. SLSA (Supply-chain Levels for Software Artifacts) gives teams a graduated framework with specific, buildable requirements at each level.

At Level 1, the build platform generates provenance on its own, describing how an artifact was built. That's a modest requirement, but it rules out the riskiest pattern of all: ad-hoc local builds pushed straight to production with no record of how they were made.

At Level 3, the build environment itself is hardened, isolated, and ephemeral. An attacker who manages to reach the CI system at this level can't slip a manipulated artifact through undetected, because the build environment is designed to catch exactly that kind of tampering.

Level 3 is the realistic target for most teams, because it's the point where compromising CI no longer automatically means compromising production. Getting to Level 2 doesn't require custom infrastructure. Standard hosted CI providers like GitHub Actions and GitLab CI, combined with platform-native provenance tools such as the slsa-github-generator reusable workflows or GitHub artifact attestations, get a team there. Level 3 takes more deliberate work: you have to isolate the build environment so one run can't influence another.

None of this matters downstream without an SBOM attached to the image. The SBOM is what turns a SLSA attestation into something usable: compliance reviewers, admission controllers, and downstream teams can check an image's provenance and contents without rebuilding it themselves, as long as the SBOM and scan results are attached as attestations alongside it.

Embedding scanning and signing into CI/CD so hardening is not a manual gate

Every control described so far works only if it happens the same way every time. It can't depend on a developer remembering to run it under deadline pressure. The fix is to build these steps directly into the pipeline, so the secure path is the only path an image can take to reach production.

A production pipeline that does this starts with reproducible, hermetic builds run entirely in CI, and no local builds get pushed directly to production. Once an image is built, it gets scanned for vulnerabilities with a tool like Trivy or Grype before it's pushed anywhere, and the pipeline fails outright on critical or high-severity findings. A passing image is then signed using Cosign's keyless signing, which binds it to the build identity that produced it. Its SBOM and scan results get attached to the image as attestations once it lands in the registry, so anyone downstream can check its provenance without rerunning anything. At deploy time, an admission controller such as Kyverno or OPA Gatekeeper (paired with an external data provider like Ratify) rejects any image that arrives unsigned or unattested, and that rejection happens at the cluster level, not as a step someone has to remember on their own workstation. Finally, the work doesn't stop once an image is running: continuous scanning keeps checking deployed images as new CVEs get published, catching vulnerabilities that didn't exist, or weren't yet known, at the moment the image was built.

That sequence, build, scan, sign, attest, verify, monitor, is what turns hardening from a set of good intentions into something that holds up under real deadline pressure, because none of it depends on anyone choosing to do it right that day.

Sources

  1. Best Practices for Securing and Hardening Container Images
  2. Container Vulnerability Management: Securing in 2026
  3. 10 Container Security Best Practices in 2026
Filed underSecurity Posture

More in Security Posture