Chaos Engineering Without a Dedicated SRE Team

Product teams can adopt chaos engineering without dedicated SREs.

Contributing Editor · · 11 min read
Cover illustration for “Chaos Engineering Without a Dedicated SRE Team”
Reliability & Uptime · September 29, 2026 · 11 min read · 2,414 words

Chaos engineering doesn't need a dedicated SRE team to exist inside a company. It needs a deliberate, lightweight process that a product engineering team can own outright, starting small and building into a habit.

Chaos engineering became an SRE specialty for the wrong reason

Netflix's Chaos Monkey gave the industry its origin story for this practice, and that story left behind a mental model that doesn't hold up: chaos engineering equals hundreds of SREs and a budget with no ceiling. SRE itself was coined at Google back in 2003, and it was never meant to describe a headcount spendark.com. It describes a set of practices: service level objectives, error budgets, toil management, chaos experiments, that any team can pick up regardless of size spendark.com. The trouble is that most of the guidance floating around for engineering leaders comes from Google's own 500-page book, Netflix's chaos engineering blog posts, and Spotify's squad model, all written by and for organizations running thousands of engineers. A ten-person startup reading that material comes away thinking chaos engineering is out of reach spendark.com.

It is not, since deliberately introducing controlled failure to find hidden failure modes before a customer finds them first is precisely what chaos engineering means. Chaos engineering just means deliberately introducing controlled failure to find the hidden failure modes before a customer finds them first, and the blast radius gets bounded by design, not by how many engineers are on staff. Mordor Intelligence sized the chaos engineering market at $2.36 billion in 2025, growing at an 8.28% compound annual rate toward $3.51 billion by 2030 Why Chaos Engineering Powers Modern SRE | HCLTech spendark.com coderio.com. That's not a research curiosity confined to a few unicorns. That's a commercial ecosystem with vendors competing for a small team's business too.

So the question a team should be asking isn't "do we have an SRE team." It's "do we have a deliberate process." This reframe changes what comes next from an org-chart problem into a workflow problem, and workflow problems are solvable at any size.

The reliability wall small teams hit before they think about chaos

A startup scales engineering headcount without scaling its operational practices, and it hits a wall, a pattern that occurs the same way almost everywhere spendark.com. Shipping more services follows from more engineers. More services produce more incidents. More incidents mean more manual firefighting, and none of it self-corrects just because the team got bigger spendark.com. Toil, the manual, repetitive ops work that doesn't scale with automation, rose 30% in 2026, the first increase in five years SRE Practices for Startups | The Good Shell spendark.com coderio.com. That's not a story about one struggling company. That's an industry-wide signal that the problem is getting worse, not better.

The trigger is usually specific and recognizable. The scrappy DevOps setup that shipped fast at ten engineers starts throwing cascading alerts at sixty, and around the same time, a prospective enterprise customer starts asking about the SLA before they'll sign the contract spendark.com. Reliability debt behaves the same way technical debt does: retrofitting observability, ownership clarity, and incident process into a system that was built without any of it costs a lot more than building it in gradually from the start.

Chaos engineering isn't the first SRE practice worth adopting. But it actually tests whether the foundation a team thinks it has, ownership, observability, runbooks, holds up under a realistic failure, rather than just sounding good in a wiki page. Before any of that testing makes sense, explicit service ownership, a baseline of observability, and runbooks for the incidents that keep recurring need to exist, the scaffolding the rest of this depends on. That's the scaffolding the rest of this depends on.

The foundation a team must have before running any chaos experiment

The organizing rule here is simple: implement whatever eliminates the most recurring pain right now, and nothing more. Three things need to be in place before a chaos experiment produces anything useful.

First, ownership has to be explicit. A shared document, a Notion page, a YAML file, doesn't matter which, needs to record the service name, the owning team, and an on-call contact, and it needs to be findable in under 30 seconds spendark.com coderio.com. Skip this and a chaos experiment reveals a failure that nobody actually knows how to route.

Second, there needs to be a minimal observability baseline. That means the four golden signals, latency, traffic, errors, saturation, instrumented on every service that touches a user. Prometheus and Grafana get a team under 50 engineers this far at close to zero cost, so budget isn't really the excuse it sometimes gets treated as.

Third, write runbooks for the five incident types that come up most often SRE Practices for Startups | The Good Shell spendark.com coderio.com. Specific commands, specific dashboards, specific escalation paths, not vague guidance. Without that, a chaos experiment that uncovers a real failure mode produces panic instead of a structured response, and panic teaches a team nothing.

Cap manual ops work at half of an engineer's time, and force the rest toward automation spendark.com. Chaos engineering only pays off if the team actually has room left to act on what the experiments find. This section is not recommending implementing SLOs, error budgets, a service catalog, observability-as-code, and chaos engineering all at once; doing so is a recipe for doing none of them well. Sequence it.

Once those three foundations exist, ownership, observability, runbooks, the first chaos experiment stops being a research project. It becomes a structured check that the system behaves the way its documentation claims it does.

How to design a first chaos experiment with a low blast radius

Three properties define a real chaos experiment, and none of them are optional: small scope with a bounded blast radius, full observability during the run, and repeatability spendark.com. Drop any one of those and what's left is just an outage that happens to be intentional.

A practical first month looks like this. In week one, define steady state for the three services that matter most, write down three to five metrics per service that indicate health, and wire those into whatever monitoring already exists. In week two, run the first experiment by hand: terminate one instance of one service, during business hours, and write down exactly what happens. In week three, fix whatever broke, because chaos engineering without remediation is just theater dressed up as engineering discipline.

A handful of blast-radius controls make that first experiment survivable. Pick a non-critical service, or something staging-adjacent, rather than the payments system. Run it while the whole team is around to watch and respond, never overnight when nobody's looking. Set a stop condition before starting: if metric X crosses threshold Y, the experiment halts immediately, no debate. And test exactly one failure mode at a time. Combining failure injections before single-fault behavior is well understood just muddies the read on what actually caused what.

On what to break first: network latency, instance termination, resource exhaustion, roughly in that order of risk, with instance termination serving as the standard "first real experiment" most teams run. What that first experiment actually tests is whether the observability, the runbooks, and the ownership structure work the way they're supposed to. It's whether the observability, the runbooks, and the ownership structure work the way they're supposed to. A failure that gets caught cleanly counts as a win.

Choosing a chaos tool when there is no platform team to maintain one

For a small team, the variable that matters is cost of ownership, not the length of a feature list. Open-source tools skip the license fee but eat engineering time; commercial tools charge by the host or by the minute, and that adds up faster than it looks like it will.

Some real numbers to anchor expectations. Gremlin runs $50 per host per month on an annual contract, so a 100-node Kubernetes cluster costs $5,000 a month, or $60,000 a year, before any volume discount kicks in cubeapm.com. Azure Chaos Studio charges $0.20 per experiment minute cubeapm.com spendark.com.

LitmusChaos is the open-source default for teams watching their budget and living on Kubernetes. It's free, it installs through the official Helm Chart, it ships with a native web interface called ChaosCenter, and there's a public experiment repository, ChaosHub, to pull from. The fault library already covers containers, hosts, Amazon EC2, Apache Kafka, and Azure. Harness owns the project and sells a fully managed SaaS version, Harness Chaos Engineering, at $200 per service per month for teams that want the support without running the infrastructure themselves speedscale.com spendark.com. It fits best for Kubernetes-native teams that have engineers willing to own setup and upkeep.

Gremlin sits at the other end: the first commercial chaos engineering tool to hit the market, founded in 2016, covering AWS, Azure, GCP, Linux, Windows, Kubernetes, and on-prem through Gremlin Private Edition, and it's SOC II compliant spendark.com. It's closed-source with limited room for customization, and a five-person team running 20 nodes is probably not who this is built for, given the cost math above spendark.com.

Steadybit, founded in 2019, deserves a mention for teams that want commercial backing without giving up flexibility spendark.com. It's built around an open-source extension framework so teams can add their own attacks and integrations, and its Reliability Hub is an open-source library of experiment templates and actions spendark.com. And for teams already living inside a CI/CD workflow, Harness Chaos Engineering integrates directly with CI/CD, GitOps, and existing observability tooling, and Harness describes it as more pipeline-aware than competing options, which in practice means it's usable by engineers who've never held an SRE title.

The heuristic to walk away with: Kubernetes-native team with two or three engineers willing to own the tooling, start with LitmusChaos. Need turn-key experiments with no setup overhead and the budget covers it, Gremlin removes the friction cubeapm.com. CI/CD integration is the priority, Harness CE slots into an existing pipeline more naturally than anything else on this list. AWS FIS charges $0.10 per minute per action, so a 10-minute experiment terminating 20 EC2 instances costs $20.00, and CubeAPM reports that daily tests across multiple AWS accounts accumulate $600/month minimum, while Gremlin's own comparison shows AWS FIS remains limited to AWS only with no multi-cloud or hybrid support cubeapm.com spendark.com.

Embedding chaos experiments into a CI/CD pipeline without a dedicated operator

The real sign of a maturing practice is when chaos experiments become part of the release checklist itself: small, frequent tests with remediation that's visible to the whole team, instead of one-off events run by whoever happened to have time that week.

The pattern for wiring this in isn't complicated. Define the experiment as code, an experiment YAML file or equivalent, and version it right alongside the application it tests. Gate the run on a passing observability baseline: if the four golden signals aren't healthy before the experiment starts, it doesn't fire. Run a scoped fault, latency injection or a single pod termination, as a post-deploy step in staging, before anything gets promoted to production. Fail the pipeline outright if the system doesn't return to steady state inside a defined recovery window.

Harness CE's CI/CD integration is the cleanest existing implementation of this pattern, and LitmusChaos supports GitOps-driven scheduling too, for teams already running Argo CD or Flux. For a team with zero chaos engineering background, this can look as simple as one engineer owning the experiment definitions as a code artifact, experiments firing automatically on every release to staging, and production experiments getting scheduled weekly, during business hours, once the staging pattern has proven stable.

Buying the tooling without changing the culture, specifically without blameless post-mortems and action items that actually get tracked, means the pipeline produces findings that sit there and get ignored. Infrastructure that auto-patches CVEs and manages cluster upgrades on its own removes a whole category of platform-layer failure from the experiment surface, which keeps the testing focused on how the application itself handles failure rather than on infrastructure drift.

Running blameless post-mortems when there is no SRE team to own them

Blameless post-mortems with tracked action items are a culture shift, not a process document, and the tooling that supports them matters less than that culture shift. Chaos engineering without a post-mortem to follow it produces nothing but stress; without remediation, the whole exercise is theater, no matter how well the experiment itself was designed.

Ownership doesn't require a dedicated team to work. The engineer who designed the experiment owns the post-mortem document. The team lead owns tracking the action items. And the post-mortem gets shared openly across engineering rather than buried in a private incident channel where nobody but the on-call rotation ever reads it. Two metrics tell a small team whether any of this is actually working: mean time to recovery going down, and alert accuracy going up. Those are the leading indicators that the chaos program is producing real gains, not just a folder of interesting findings nobody acted on.

What separates a team that keeps this practice alive from one that quietly drops it after three experiments comes down to the post-mortem. It's the mechanism that turns an experimental finding into an actual engineering priority, and that's what gives chaos work a business justification instead of making it look like deliberate self-sabotage.

Scaling the chaos practice as the team and system grow

The same staged logic that applies to the foundation work applies here too: implement whatever eliminates the most pain at the current scale, then expand, not the other way around.

Under 20 engineers, roughly seed to Series A, the process stays manual, limited to business hours, one service at a time, with findings tracked in a shared doc spendark.com. That's the entire process this article has walked through spendark.com. Past 80 engineers, GameDays become worth running, full-team exercises scheduled around complex failure scenarios, and it's common at this stage for one engineer to take on reliability work part-time, functioning as something close to an embedded SRE without the title or the dedicated headcount spendark.com.

A dedicated SRE team earns its place at scale, not before it. Before that point, engineers build SRE instincts inside a DevOps culture, one deliberate experiment at a time spendark.com. At 20–80 engineers (Series A to Series B), The Good Shell reports that experiments graduate into the CI/CD pipeline, a dedicated experiment library is maintained as code, and SLOs are formally defined with chaos experiments designed to test SLO breach conditions specifically spendark.com.

Sources

  1. Why Chaos Engineering Powers Modern SRE | HCLTech
  2. SRE Practices for Startups: The Essential Guide to Reliability Without a Dedicated SRE Team - The Good Shell
  3. SRE for Scale-Ups: Reliability Engineering Guide

More in Reliability & Uptime