A service continues returning correct results, but its latency is so high that it violates the application's timing requirement. Which term best describes this condition?
Fault Tolerance and Redundant Systems Online Quiz Questions
Use this free practice quiz with 20 questions to review Fault Tolerance and Redundant Systems, test your knowledge, and prepare for your next test or exam.
A service has three application servers, but all of them depend on one unreplicated database. What is the unreplicated database in this design?
- A
A failure domain
- B
A replica
- C
A single point of failure
- D
A bulkhead
An application process is running, but it cannot currently serve requests because a required dependency is unavailable. Which health check should normally prevent the instance from receiving traffic without necessarily restarting it?
- A
Readiness check
- B
Liveness check
- C
Startup check
- D
Backup-restore test
Which two statements correctly describe synchronous replication? Select all that apply.
- A
It can reduce the amount of recently committed data lost after a failure.
- B
It always improves write latency compared with asynchronous replication.
- C
It can make availability more sensitive to network partitions.
- D
It acknowledges every write before any replica receives it.
Which two patterns directly help contain the blast radius of a failing dependency or overloaded tenant? Select all that apply.
- A
Unbounded retries
- B
Bulkheads
- C
Removing all timeouts
- D
Circuit breakers
True or false: Replication alone is sufficient protection against accidental deletion or corruption because every replica contains a copy of the data.
- A
True
- B
False
True or false: When one of two servers fails, failover automatically gives the surviving server enough capacity to handle all the original traffic safely.
- A
True
- B
False
What is the standard term for the maximum acceptable amount of data loss after a disaster?
What failure-isolation pattern temporarily stops requests to a dependency that is repeatedly failing?
Complete the causal sequence: An abnormal component condition is a ; the incorrect internal state it produces is an ; and externally visible behavior that violates the specification is a .
Complete the observability description: Numerical measurements are , timestamped records of events are , and records of a request's path across services are .
A consensus cluster has five nodes, and two nodes fail. Explain why the cluster can generally continue making progress, and state the key assumption that makes this possible.
After repairing and resynchronizing a failed replica, which next action best reduces the risk of returning an unsafe system to full production load?
- A
Disable all monitoring during recovery
- B
Delete the old replica immediately
- C
Validate data integrity and application behavior before gradually restoring traffic
- D
Route all traffic back as soon as a process starts
A service process is still running, but it cannot currently serve requests because a required dependency is unavailable. Which health check should normally prevent this instance from receiving new traffic without necessarily restarting it?
- A
A liveness check
- B
A readiness check
- C
A startup check
- D
A backup-restore test
A database must remain available during an ordinary node failure and also recover from accidental deletion that is replicated to every live copy. Which design best addresses both requirements?
- A
Use only active-active replication
- B
Use replication for availability plus a retained backup that can restore an earlier state
- C
Increase the number of live replicas without changing anything else
- D
Use round-robin load balancing
True or false: A performance problem can never become a system failure if the system eventually returns correct results.
- A
True
- B
False
A disaster-recovery requirement says that service must be restored within 15 minutes and that no more than approximately 5 minutes of recently committed data may be lost. What is the RTO, expressed in minutes?
What is the exact health-check term for a check that determines whether an instance should receive traffic?
Two independent components are both required for a request to succeed. Each has availability 0.99. What is the approximate combined availability of the series system?
- A
98.01%
- B
99.00%
- C
99.99%
- D
100.00%
An API repeatedly calls a dependency that is timing out. The system should temporarily stop sending calls to that dependency and resume them only after it appears healthy. Which failure-isolation pattern is most appropriate?
- A
A queue
- B
A bulkhead
- C
A circuit breaker
- D
A startup check