Using Hedged Requests to Reduce Tail Latency
Most calls to a dependency may finish quickly while a small fraction take much longer. A request that depends on one of those slow calls inherits the delay even when another healthy instance could have answered sooner. Increasing the timeout does not solve this problem. Retrying only after the timeout may also be too late: by then, the caller has already spent most of its latency budget. A hedged request is a deliberately delayed duplicate of an operation that is still in progress. The original request starts normally. If it has not completed after a chosen delay, the caller sends one additional equivalent request, usually to another eligible instance. The first acceptable result wins, and the remaining work is cancelled or ignored.