7166071fca
The compression/aux fallback ladder decides "payment error" via _is_payment_error(), which only knew the spaced "resource exhausted". NVIDIA NIM and gRPC-style wrappers serialize the same quota signal as ResourceExhausted / RESOURCE_EXHAUSTED / resource-exhausted, so a 403/429/status-less body carrying it re-raised instead of walking the configured fallback_chain, and compression fell to the lossy emergency path. Add the three separator variants to _PAYMENT_KEYWORDS (same status gate, no broader "exhausted" matching). evals/auxiliary_resource_exhausted.py drives the real call_llm/async_call_llm against two local OpenAI-compatible listeners (nvidia profile primary, named custom fallback) for the before/after check. #85649