Designing for Failure: Chaos Engineering Basics for Non-Unicorns
Chaos engineering has a reputation problem: it sounds like something only companies with an SRE org and a platform team can afford. It isn't.
Most of what gets written about chaos engineering opens with Netflix, Chaos Monkey, and a platform team large enough to staff a mid-size startup on its own. That framing is accurate for Netflix and useless for almost everyone else, because it quietly tells every other engineering leader that this discipline is out of reach until they have a dedicated SRE org. It isn't. Chaos engineering is not a tool, a team, or a budget line. It is a hypothesis-testing habit: pick a failure you are already afraid of, inject it on purpose in a controlled way, and see whether your system behaves the way you assumed. You can start that habit with three prerequisites, five experiments that fit on an afternoon, and no new headcount.
The honest reason most mid-size and regulated-industry teams never start is not technical, it is organizational: nobody owns resilience testing, so it never gets scheduled. Teams without the in-house bandwidth to build that discipline themselves often end up leaning on an outside partner instead, which is really the same conversation from a different angle. A custom software development team at YuSMP that designs for failure from day one is, in effect, running the chaos engineering mindset upstream of your first incident rather than retrofitting it after one. Either way, the discipline itself does not require a platform team to begin, just a willingness to break something on purpose before reality does it for you.
What Chaos Engineering Actually Is (Minus the Netflix Mythology)
Strip away the tooling and the war stories, and chaos engineering is deliberate, controlled failure injection used to test whether your assumptions about system behavior actually hold. That is a meaningfully different activity from traditional testing. A unit test or an integration test confirms known behavior: given this input, does the system produce the expected output? Chaos engineering probes unknown failure modes: if this dependency times out, does the rest of the system degrade gracefully, or does it fall over in a way nobody predicted? You are not trying to break things for fun. You are trying to find out, on your own schedule, what you would otherwise find out during an incident at 2 a.m.
Why "We're Not Big Enough for This" Is the Wrong Instinct
Small and mid-size systems fail in the same categories as hyperscale ones: dependency timeouts, resource exhaustion, partial outages, a third-party API that silently starts returning errors. The difference is not the category of failure, it is predictability. A company running thousands of instances has enough statistical volume that rare failure modes show up constantly, which forces the discipline into existence. A company running a few dozen services has the same failure modes, they just surface rarely and unpredictably, usually during a product launch or a traffic spike, which is the worst possible time to discover them. You do not need Netflix's scale to have unknown failure modes in production. You need customers who notice when something goes down, and most companies already have those.
Three Things You Need Before Your First Experiment
Three prerequisites, and none of them require new tooling or a reorg:
- A steady-state definition. What does "normal" look like, measurably? Response time, error rate, order completion rate, whatever the system's health actually means in business terms.
- A specific hypothesis. Not "let's see what breaks," but "if the payments API times out, checkout should fall back to a queued retry instead of failing the order."
- A blast radius and abort criteria. Run it in staging first, then on the smallest reversible slice of production you can justify, with a clear line for when you stop the experiment and roll back.
Skip any one of these and you are not doing chaos engineering, you are just causing an outage with extra steps.
Five Starter Experiments That Don't Require a Platform Team
None of these need Gremlin, LitmusChaos, or a dedicated fault-injection platform to start, though tools like those earn their keep once you are running experiments regularly. A few lines of kill -9, tc netem, or iptables, or your cloud provider's built-in fault-injection service, is enough for a first pass:
- Kill a non-critical background process and watch whether it recovers on its own or needs a human.
- Inject artificial latency into one dependency call and see what happens to the request that depends on it.
- Exhaust memory or CPU on a single non-critical instance and confirm the rest of the fleet absorbs the load.
- Simulate a failed third-party API call, a timeout or a 500, and check whether your error handling degrades gracefully or cascades.
- Pull a database replica out of rotation and confirm failover actually happens the way the runbook says it does.
Each of these is runnable by one engineer in an afternoon, with a hypothesis written down beforehand and a rollback plan ready before you start.
What Changes When You're in a Regulated Industry
In fintech, healthcare, aerospace, and similar sectors, untested recovery behavior is not just a reliability gap, it is an audit-trail gap. If an auditor asks what happens when a core dependency fails, and the honest answer is "we think it fails over, we've never actually watched it happen," that is a finding waiting to be written up. The fix is not more documentation written from memory. It is running the experiment, logging what actually happened, and keeping that output as evidence. A chaos experiment that produces a timestamped log of "dependency failed, system degraded as designed, recovery completed in 40 seconds" is worth more to a compliance review than a diagram claiming the same thing never tested.
How Do You Run a Game Day Without Scaring Everyone?
Schedule it like any other planned work: pick a date, scope the blast radius in writing, tell the on-call team it's happening, and keep a kill switch within arm's reach the entire time. Treat the first few game days as learning exercises for the team, not pass/fail audits of anyone's competence. The goal of the first game day is not a clean result, it is a team that is no longer afraid of the exercise, which makes the second and third ones far more useful.
Tools: What You Actually Need vs. What the Vendors Will Sell You
You do not need a commercial platform to start. Manual scripts and your cloud provider's native fault-injection options cover the first several months of experiments for most teams. Commercial tools like Gremlin or open-source options like LitmusChaos earn their place later, once you are running experiments often enough that scheduling, reporting, and blast-radius controls become worth automating. Buy the tool when the manual process becomes the bottleneck, not before.
You don't need a platform team to start chaos engineering. You need a hypothesis, a blast radius, and the discipline to run the experiment anyway.
FAQ
Is chaos engineering safe to run in production?
Eventually, yes, on a narrow and reversible slice, with abort criteria defined before you start. Your first several experiments should run in staging, not production.
Do we need microservices or Kubernetes to do this?
No. A monolith with external dependencies has the same categories of failure: a slow database, a flaky third-party API, a resource limit. The experiments described here apply regardless of architecture.
How is this different from normal QA or load testing?
QA confirms the system does what you expect under expected conditions. Load testing confirms it holds up under expected traffic. Chaos engineering specifically tests what happens when something you depend on fails, which neither of the other two is designed to probe.
What's the smallest possible first experiment?
Kill a single non-critical background process in staging and watch whether it recovers without a human. It takes under an hour and tells you something real about your assumptions.
How do we convince leadership this is worth the risk?
Frame it as controlled testing on your schedule versus uncontrolled testing on an outage's schedule. The second one already happens whether you plan for it or not; this just moves it to a Tuesday afternoon instead of a Saturday night.