Named after the electrical device that trips to stop a current before it burns down the wiring, a software circuit breaker trips to stop calls to a dependency that’s already failing, before those calls burn down the rest of your system too.
The Problem: Failures That Spread
Without a circuit breaker, a struggling downstream service tends to make things worse for everyone calling it. Requests pile up waiting on a slow or failing dependency, threads or connections get held open, retries add even more load on top of a service that’s already struggling, and the calling service itself starts timing out — which then cascades to whatever’s calling it. A single failing dependency, left unmanaged, can take down services three hops away that never directly touch it.
Three States
A circuit breaker wraps calls to a dependency and tracks recent failures, moving between three states based on what it sees.
| State | Behavior | Transitions to |
|---|---|---|
| Closed | Calls pass through normally, failures are counted | Open, once the failure threshold is crossed |
| Open | Calls fail immediately without touching the dependency | Half-open, after a cooldown period |
| Half-open | A limited number of test calls are allowed through | Closed if they succeed, back to Open if they fail |
The open state is the important one: instead of letting every caller individually discover the dependency is down (each paying a full timeout to find out), the breaker fails fast, immediately, without even attempting the call. That protects both sides — callers stop wasting time and threads on doomed requests, and the struggling dependency stops receiving load it can’t handle, giving it room to actually recover.
A Basic Implementation
type State = 'closed' | 'open' | 'half-open';
class CircuitBreaker {
private state: State = 'closed';
private failureCount = 0;
private lastFailureAt = 0;
constructor(
private readonly failureThreshold = 5,
private readonly cooldownMs = 30_000,
) {}
async call<T>(fn: () => Promise<T>): Promise<T> {
if (this.state === 'open') {
if (Date.now() - this.lastFailureAt < this.cooldownMs) {
throw new Error('Circuit is open, failing fast');
}
this.state = 'half-open';
}
try {
const result = await fn();
this.onSuccess();
return result;
} catch (err) {
this.onFailure();
throw err;
}
}
private onSuccess() {
this.failureCount = 0;
this.state = 'closed';
}
private onFailure() {
this.failureCount++;
this.lastFailureAt = Date.now();
if (this.failureCount >= this.failureThreshold || this.state === 'half-open') {
this.state = 'open';
}
}
}
const paymentBreaker = new CircuitBreaker(5, 30_000);
async function chargeCard(payload: ChargeRequest) {
return paymentBreaker.call(() => paymentProvider.charge(payload));
}
Real implementations usually track a rolling window of recent call outcomes (a failure rate over the last N requests or last N seconds) rather than a simple consecutive-failure counter, which avoids tripping on a couple of unlucky failures mixed into mostly-healthy traffic.
Choosing Thresholds
Getting the thresholds right matters more than getting the pattern right. Too sensitive, and normal transient blips trip the breaker constantly, adding latency and failed requests that a couple of retries would have absorbed just fine. Too lax, and the breaker doesn’t trip until the dependency and everything calling it are already in serious trouble.
Circuit Breaker, Retry, and Timeout Are Different Tools
These three get bundled together constantly, but they solve different problems and are meant to be used together, not as substitutes for each other.
| Pattern | Question it answers | Protects |
|---|---|---|
| Timeout | How long will I wait for one call? | The calling thread/request from hanging forever |
| Retry | Should I try this specific call again? | Recovery from a single transient blip |
| Circuit breaker | Should I even attempt calls to this dependency right now? | Both sides from sustained, ongoing failure |
A sensible stack applies all three: a timeout bounds each individual call, a small number of retries with backoff handle transient blips, and the circuit breaker sits above both, tracking the aggregate failure rate and cutting off the retries-plus-timeouts machinery entirely once the dependency is clearly down rather than just having a bad moment.
Circuit Breakers Interact With Idempotency
Takeaway
A circuit breaker protects a system from a failing dependency by tracking recent failures and, past a threshold, failing fast instead of continuing to hammer something that’s already struggling — then cautiously testing recovery through a half-open state before fully reopening the gate. It’s not a substitute for timeouts or retries, it’s the layer that decides when those mechanisms should stop being tried at all, and it only works safely alongside idempotent calls that can tolerate the retries a recovering circuit will generate.