Skip to content
TaeyoungKim.dev

Azure Traffic Manager priority failover: Why users switch at different times

LinuxWritten 3 min readTaeyoungKim
LinkedInX

You configure Azure Traffic Manager Priority routing to send users to a backup service if the primary fails. Yet one user already sees the backup while another continues toward the primary. Before calling failover broken, separate two questions: which endpoint Traffic Manager considers healthy, and when each user's DNS result is refreshed.

Which endpoint does Priority choose?

Traffic Manager chooses the endpoint for new DNS answers from priority and monitored health. Expiration of DNS entries already cached by clients or resolvers is a separate step, so the change is not visible to everyone instantly.

Priority routing suits a primary-and-backup arrangement: choose the highest-priority eligible healthy endpoint. For illustration, assign the primary priority 1 and the backup priority 2. When both are healthy, select the primary; after monitoring marks it unhealthy, select the backup. These numbers illustrate the concept rather than reproduce a particular production configuration.

EndpointPriorityMonitored stateExpected selection
Primary1HealthyPrimary
Backup2HealthyPrimary
Primary1UnhealthyBackup
Backup2HealthyBackup

This simplified table explains the order. Actual DNS answers also depend on whether endpoints are enabled and how monitoring is configured. Assigning the backup a lower priority does not make it a viable recovery target if the backup application is not ready.

Health detection and DNS refresh happen at different times

Traffic Manager is a DNS-based routing service, not a proxy that forwards every application request. It monitors endpoint health and returns a suitable target in DNS answers. It cannot forcibly move every existing client connection at the instant a service fails. Microsoft's monitoring guidance describes the endpoint-health side of that decision.

First check whether Traffic Manager has marked the primary unhealthy. Then inspect which target a new DNS lookup receives, and whether a user's resolver or client still has a cached answer for the old target. Two users seeing different services is not, by itself, proof of broken monitoring. Existing connections and client-side caching also affect observations.

An extremely short DNS TTL may sound like a guarantee of instant change, but it cannot guarantee every client's behavior. Measure monitoring interval, failure-detection settings, DNS TTL, and the backup application's actual readiness together. A priority-routing concept alone cannot promise failover within a fixed number of seconds.

What must be ready after DNS selects the backup?

The backup can receive traffic yet still fail users if it lacks the required data or dependencies. Decide how both sites access data, how login sessions continue, and where writes go. Traffic Manager's DNS decision does not replicate application state.

Begin a controlled test by confirming that the healthy primary is selected. Then record when monitoring marks it unhealthy, when a new DNS lookup points to the backup, and whether the backup handles critical application requests. Finally, verify what happens when the primary recovers. Plan the potential service impact and a recovery procedure before deliberately testing failure.

Key takeaways: separate routing change from service recovery

Priority routing selects the highest-priority eligible healthy endpoint. Health detection, DNS cache refresh, and backup application readiness are distinct. If failover looks slow, inspect monitoring state → new DNS answer → client connection and application behavior in that order.

Author

TaeyoungKim

Connecting technical foundations with implementation, verification, and production decisions.

#Azure#Traffic Manager#Priority#DNS#Failover

Read next