Fault Tolerance and Redundant Systems

A progressive guide to designing, operating, and testing systems that continue useful service despite component faults, overload, data loss, and recovery challenges.

Foundations of

A dependable system is designed on the assumption that components will eventually malfunction. A server may lose power, a network link may drop packets, a process may deadlock, or a database replica may become unreachable. focuses on continuing correct or acceptable service despite such events, while also includes absorbing disruption, recovering, and adapting.

The central design principle is containment: one component's problem should not become the entire system's outage. This requires more than making individual components reliable. Dependencies, capacity, detection, data protection, recovery procedures, and testing must work together.

Takeaway: Reliability is a system property. It comes from preventing, containing, and recovering from faults rather than assuming that every component will remain perfect.

Faults, Errors, Failures, and Performance

These terms describe a progression from an internal problem to an externally visible consequence:

  1. A is an abnormal condition or defect, such as a damaged disk or software bug.

  2. An is the incorrect internal state produced by that , such as corrupted memory or an invalid routing entry.

  3. A occurs when externally visible behavior no longer satisfies the specification.

  4. A performance problem occurs when results remain correct but timing or capacity expectations are violated.

A performance problem can become a . For example, a slow database can cause requests to miss deadlines. Clients may retry, increasing load and creating a cascading . A service that eventually returns a correct response after two minutes may be technically available but still fail a user's performance requirement.

This chain helps teams choose the right response: repair the , prevent the from spreading, detect the , and protect the system from overload caused by degraded performance.

Takeaway: Correctness, availability, latency, and capacity are related but distinct service requirements.

and Domains

provides more than one instance of a resource so that loss of one instance does not necessarily interrupt service. It can be applied to compute, networks, storage, power, facilities, and software services.

A is a dependency whose loss can stop the system. Removing one such dependency may expose another: several application servers provide little protection if they all depend on one unreplicated database or one power circuit.

Copies should be distributed across independent domains, such as separate racks, availability zones, data centers, or regions. Apparent is weak when all copies share a common software bug, faulty deployment, misconfiguration, regional outage, or physical dependency.

For two components that are both required, availability in a simplified independent model is:

0.99×0.99=0.98010.99 \times 0.99 = 0.9801

If either of two equivalent components can serve a request, availability is approximately:

1−(1−0.99)2=0.99991 - (1 - 0.99)^2 = 0.9999

The second arrangement is a parallel system, whereas the first is a series system. Real systems achieve less than the simplified result when failures are correlated.

Takeaway: Add across the dependency chain and domains, not merely within one service tier.

, Backups, and Recovery Objectives

When is applied to data or system state, it becomes . An active-passive arrangement uses one primary and a standby that takes over during . An active-active arrangement allows multiple replicas to serve traffic but may require conflict resolution or coordination for concurrent writes.

Synchronous waits for required replicas to acknowledge a write. This reduces potential data loss but can increase latency and make availability sensitive to network partitions. Asynchronous acknowledges before every replica receives the write. This often improves latency and availability but creates a window in which the newest data may be lost.

is not the same as a . Live replicas help maintain service during ordinary component failures, but accidental deletions, corruption, and malicious changes may be copied to all replicas. A retained historical copy is needed for restoration.

Two objectives make recovery requirements measurable:

  • : the maximum acceptable time to restore service.

  • : the maximum acceptable data loss, measured in time or transactions.

For example, an RTO of 15 minutes and an RPO of 5 minutes means service should return within 15 minutes and lose no more than approximately five minutes of recently committed data after a disaster.

Distributed systems may use so nodes can agree on decisions despite some node failures. A five-node cluster can generally continue making progress when two nodes fail because a majority of three remains available.

Takeaway: Choose and strategies according to both availability needs and the consequences of data loss or corruption.

and Capacity Planning

redirects work from a failed or unhealthy component to a surviving one. It may be automatic, manual, planned before maintenance, or unplanned after an unexpected .

A complete design must answer these questions:

  1. How is detected through health checks, timeouts, heartbeats, or application probes?

  2. What does unhealthy mean? A running process may still be unable to serve correct requests.

  3. Where does traffic go, and does the destination have sufficient capacity and compatible state?

  4. How is split-brain prevented so two nodes do not both accept conflicting writes as the sole primary?

  5. How are data repair, synchronization, cache warming, and controlled traffic restoration handled?

  6. How is failback treated as a separate change that is tested carefully?

does not create capacity. If one of two servers fails, the survivor receives all traffic and may become overloaded. Systems therefore need spare capacity, load shedding, or normal-operation scaling so surviving components can handle the load.

Takeaway: A path is a complete operating process, not merely a routing rule or replica-promotion command.

Load Balancing and Isolation

A load balancer distributes requests or tasks among multiple instances. Common policies include round robin, weighted routing, least connections, locality-based routing, and health-aware routing. Health checks should distinguish a process that is alive from an instance that is ready to accept work.

isolation limits the blast radius of a . Useful controls include:

  • Bulkheads: separate resources into pools so one tenant, request class, or dependency cannot consume everything.

  • Timeouts and deadlines: stop waiting for work that is unlikely to complete.

  • Circuit breakers: temporarily stop requests to a failing dependency.

  • Rate limiting and load shedding: reject or defer excess work before resources are exhausted.

  • Queues: buffer bursts and process work at a controlled rate.

  • Partitions or shards: limit the affected data and traffic to part of the system.

Retries need safeguards. Bounded attempts, timeouts, exponential backoff, random jitter, and a distinction between retryable and permanent errors reduce retry storms. Uncontrolled or synchronized retries can turn a small transient into a cascading .

Takeaway: Routing and isolation controls should protect remaining capacity instead of moving overload from one failed pool to another.

Degradation, Health Checks, and

preserves essential functionality when full functionality is unavailable. Examples include serving cached product information when recommendations fail, disabling image processing while preserving text access, returning basic search results when ranking is overloaded, or serving read-only data while writes are repaired.

This approach requires explicit priorities. The system must identify which functions are essential, which can be delayed, and which can be omitted. Optional dependencies should not be treated as hard dependencies.

The three main signals are:

  • Metrics: numerical measurements such as request rate, rate, queue depth, CPU utilization, and latency percentiles.

  • Logs: timestamped records of events, warnings, errors, and state changes.

  • Traces: records of a request's path across services, including timing and dependency relationships.

Health checks should match their purpose. A liveness check asks whether a process is functioning or making progress and may trigger a restart. A readiness check asks whether an instance should receive traffic and normally removes it from routing without restarting it. A startup check allows a slow-starting application time to initialize before liveness checks begin.

Monitoring should include user-visible outcomes such as availability, rate, latency, data freshness, and successful completion of user operations. Correlating metrics, logs, and traces helps distinguish causes such as database saturation, network partitions, expired certificates, configuration errors, and retry storms.

Takeaway: Good detects both infrastructure symptoms and the user-visible effects that determine whether the service is actually useful.

Recovery and -Tolerance Testing

Recovery is a controlled sequence rather than a simple server restart:

  1. Detect and classify the incident.

  2. Contain the and protect remaining capacity.

  3. Redirect traffic or enter a degraded mode.

  4. Repair or replace failed components.

  5. Rebuild or resynchronize replicas.

  6. Validate data integrity and application behavior.

  7. Restore normal traffic gradually.

  8. Review the incident and improve the design.

Recovery procedures should be tested under realistic conditions. Load tests, drills, -restore tests, dependency injection, and controlled chaos engineering experiments can reveal insufficient post- capacity, stale replicas, unsafe health checks, or recovery steps that depend on the failed system itself.

A design is not tolerant merely because it contains replicas. Its path, monitoring, configuration, data consistency, operator procedures, and recovery capacity must work together. Testing also exposes correlated failures that simplified independent-availability calculations do not capture.

Takeaway: is demonstrated by successful behavior during disruption and recovery, not by architecture diagrams alone.