Experimenting with TPUs, GKE Managed DRANET, and Multi-cluster Inference Gateway | Google Cloud Blog
… Apply health checks and backend policies to the pool to ensure load balancing relies on your custom hardware metrics. Configure an InferenceObjective to instruct the gateway to route prompts to the region with the highest availability, avoiding overloaded TPUs. …