# Keep an application online when an origin fails

An SRE team runs an application on a primary origin and keeps a standby copy of it in a second data center, cloud, or region. Users must keep being served when the primary fails, and the switch must happen on Azion's side, with no DNS change and no logic in the client. This page groups both origins in one connector with Load Balancer, so the standby receives requests only when the primary fails, and replaces Azion's default error page with your own when neither origin can answer. The result is measured by the share of requests that still succeed during an origin incident and by the error rate users see while traffic moves to the standby.

This use case does not cover data replication between the origins, or the performance of a single origin. For a single origin, refer to [Accelerate websites and APIs with a CDN](/en/documentation/use-cases/improve-performance-and-reliability/accelerate-websites-and-apis-with-a-cdn/).

## Prerequisites

- An application that serves your site through a workload, with a rule whose **Set Connector** behavior sends its requests to a connector of your primary origin. To create them, refer to [Applications quickstart](/en/documentation/platform/applications/quickstart/).
- A standby origin that holds the same content as the primary and is set up the same way for the application.
- A second connector that serves your error page at `/errors/503.html` and reaches neither origin, such as a connector of type `storage` that reads an [Object Storage](/en/documentation/platform/object-storage/) bucket. For its fields, refer to [Connector settings](/en/documentation/platform/connectors/settings/#storage).
- A personal token, for the API steps. To create one, refer to [Personal tokens](/en/documentation/guides/platform/account-and-billing/personal-tokens/).
- The names of your origins and your domain. This page uses `app-origin` for the connector of the primary origin, `primary.example.com` for the primary origin, `standby.example.com` for the standby, `error-pages` for the connector of the error page, and `www.example.com` for the domain. Each origin adds an `X-Origin-Site` response header with the value `primary` or `standby`, so a response names the origin that answered. Replace each value with yours in every step.

---

## Required products

| The application needs                                          | Which means                                                                                                          | Product           | Documented in                                                                                                                                         |
| -------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------- | ----------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------- |
| A standby origin that takes the traffic when the primary fails | Load Balancer on the connector, with the primary origin as a *Primary* address and the standby as a *Backup* address | Load Balancer     | [Add a backup origin to a connector](/en/documentation/guides/application-performance/availability/add-a-backup-origin-to-a-connector/)               |
| A page of your own when neither origin can answer              | A custom page set for `502`, `503`, and `504`, assigned in the workload's deployment                                 | Workloads         | [Show your own page when no origin answers](/en/documentation/guides/application-performance/availability/show-your-own-page-when-no-origin-answers/) |
| Cached content that keeps answering while the origins change   | The application's cache settings, and stale cache for expired copies                                                 | Cache             | [Expiration and freshness](/en/documentation/platform/applications/cache/expiration-and-freshness/#stale-cache)                                       |
| Errors and traffic per origin, during and after an incident    | The **Status Codes** dashboard and the upstream address of each request                                              | Real-Time Metrics | [Build dashboards](/en/documentation/platform/real-time-metrics/build-dashboards/#status-codes)                                                       |

---

## Reference architecture

This page builds the *Active-passive origin pool*: one connector that groups a primary origin and a standby, with the standby serving only when the primary fails.

```mermaid
%%{init: {"layout": "dagre", "themeVariables": {"fontSize": "13px"}, "flowchart": {"nodeSpacing": 12, "rankSpacing": 12, "padding": 6, "wrappingWidth": 70, "minNodeWidth": 40, "useMaxWidth": true}}}%%
flowchart TD
  User["User"] -->|"HTTPS request"| App["application"]
  App -->|"cached copy"| Cache["Cache"]
  App -->|"Set Connector"| Pool["app-origin connector with Load Balancer"]
  Pool -->|"Primary role: every request"| Primary["primary.example.com"]
  Pool -->|"Backup role: only when every primary fails"| Standby["standby.example.com"]
  Pool -->|"no origin answers: 502, 503, or 504"| Pages["custom page set"]
  Pages -->|"fetches the page"| Errors["error-pages connector"]
```

Read the diagram from the application. Cache answers what it already holds, and every other request reaches one connector, which holds both origins. Inside the connector, the role of each address decides who serves: the primary in normal operation, the standby only after the primary fails. The custom page set sits at the end of the failure path, for the moment when neither origin can answer.

### Dataflow

1. A user's request reaches the application, and a valid copy that Cache holds answers it without reaching either origin.
2. Any other request goes to the `app-origin` connector through the application's **Set Connector** rule.
3. Load Balancer sends the request to `primary.example.com`, the only *Primary* address. The standby receives nothing while the primary answers, so only one origin writes data at a time.
4. When every *Primary* address fails, as detected from failed and timed-out connections, Load Balancer sends the request to `standby.example.com`, the *Backup* address. The switch to the standby needs no DNS change.
5. When no origin can be selected or none answers in time, Azion returns `502`, `503`, or `504`, and the custom page set replaces that response with the page the `error-pages` connector serves, outside the pool.
6. Real-Time Metrics counts the status codes of the domain and, per request, the origin that answered, while traffic moves to the standby and back.

### Components

- **connector**: the Platform Resource that groups the origins. One connector, `app-origin`, holds the primary and the standby, so the rule that names it needs no change when traffic moves between them.
- **Load Balancer**: assigns each address the *Primary* or the *Backup* role. *Primary* addresses are always preferred, and a *Backup* address takes requests only when every *Primary* address fails. **Max Retries**, **Connection Timeout**, and **Read/Write Timeout** decide how long a failing connection holds a request. The *IP Hash* method refuses *Backup* addresses.
- **application**: the Platform Resource that routes requests to `app-origin` with a **Set Connector** rule, and applies the cache settings.
- **Cache**: keeps answering the content it holds while the origins change. With **Stale cache** on, it also serves an expired copy when the origin returns a `5xx` error or times out.
- **Custom Pages**: the Platform Resource that replaces Azion's error responses with a page of your own. Each page is fetched from a connector, so the page lives on `error-pages`, a connector that reaches neither origin.
- **Real-Time Metrics**: shows the traffic per origin and the errors users see, through the status codes and the upstream address of each request.

### Other designs for this use case

- *Active-active origin pool across clouds*: for teams that serve from two or more origins at once, to spread load or cost. Every origin serves live traffic, so the design needs replicated or shared state and a session-affinity decision, and a failure removes capacity instead of switching sides.
- *Single origin with stale-cache fallback*: for applications with one origin whose content tolerates being briefly out of date. There is no second origin, so the failure flow serves expired content instead of switching origins, and only content already in cache survives the outage.

---

## Configure the origin pool

The origin pool is the `app-origin` connector with Load Balancer on and two addresses. The primary origin takes the *Primary* role, so it receives every request while it answers. The standby takes the *Backup* role, so it stands by out of daily traffic and receives requests only when every *Primary* address fails.

The pool is the procedure that [Add a backup origin to a connector](/en/documentation/guides/application-performance/availability/add-a-backup-origin-to-a-connector/) describes, run on `app-origin` with these values:

- **Addresses** `primary.example.com` with **Server Role** *Primary*, and `standby.example.com` with *Backup* and **Active** on.
- **Method** *Round Robin*. With one *Primary* address, there is nothing to rotate, and *IP Hash* refuses *Backup* addresses with `28005`.
- **Max Retries** `1`. A connection to the origin that fails is retried once. The user waits through every retry before receiving an answer, so one retry bounds that wait.
- **Connection Timeout** `10` seconds. An origin that accepts no connection within 10 seconds fails the connection, instead of holding the user for the 60-second API default.
- **Read/Write Timeout** `60` seconds, the value the Console fills in. It bounds the wait for data on an open connection to an origin that has stopped answering.

Load Balancer has no health check: no probe and no interval, so nothing tests either origin between requests. The API default of **Max Retries** is `0`, so an API or CLI body sets every value explicitly, as in this `config`:

```json
{"method": "round_robin", "max_retries": 1, "connection_timeout": 10, "read_write_timeout": 60}
```

The connector holds the primary origin and the standby, and the rule that names `app-origin` needs no change. A connector change reaches Azion's distributed infrastructure over several minutes, and data centers apply it at different times.

Both addresses receive the same `Host` header, path prefix, and protocol from the connector. When the standby answers under another name than the primary, set the connector's **Host** to `${host}`, which sends the host the user requested. For the connection options, refer to [Connector settings](/en/documentation/platform/connectors/settings/#connection-options).

---

## Configure the error page when no origin answers

When no origin can answer, Azion returns a status code of its own: `502` when no origin server can be selected or the origin returns an invalid response, and `504` when the origin does not answer in time. Without a custom page, the user receives the page titled `Azion - Default error page`. A custom page set replaces those responses with your page.

The set in this section binds `502`, `503`, and `504` to `/errors/503.html` on the `error-pages` connector, and answers each one with `503`. The `503` status marks a temporary condition, which is what an origin outage is. The page's **Response TTL** is `86400` seconds, one day: an error page is static and rarely changes, and a long TTL keeps requests away from the `error-pages` connector. The page comes from a connector that reaches neither origin, because a page fetched from the failing pool would fail with it.

The set is the one that [Show your own page when no origin answers](/en/documentation/guides/application-performance/availability/show-your-own-page-when-no-origin-answers/) creates and assigns, with these values:

| Field                             | Value                                                        |
| --------------------------------- | ------------------------------------------------------------ |
| **Name**                          | `origin-down`                                                |
| **Page Code**                     | `502`, `503`, and `504`, one page each                       |
| **Connector**                     | `error-pages`                                                |
| **Page Path (URI)**               | `/errors/503.html`                                           |
| **Response TTL**                  | `86400`                                                      |
| **Response Custom Status Code**   | `503`                                                        |
| **Custom Page** of the deployment | `origin-down`, on the workload that serves `www.example.com` |

The workload's deployment names the `origin-down` set. A deployment change takes several minutes to reach Azion's distributed infrastructure, and requests can receive Azion's default page until it does.

---

## Verify the setup

Each check reads the `X-Origin-Site` header your origins set. Run the failure checks in a maintenance window, because they take an origin out of service.

- **The primary serves while it answers.** Send several requests:

  ```bash
  curl -s -D - -o /dev/null https://www.example.com/
  ```

  Every response carries `x-origin-site: primary`. Under HTTP/2, header names arrive in lower case.

- **The standby takes over when the primary fails.** Stop the web server on `primary.example.com`, or block the connections from Azion at its firewall. Send the request again, several times. The responses carry `x-origin-site: standby`, with no DNS change on your side.

- **The primary takes the traffic back.** Start the web server on `primary.example.com` again, and send the request several times. The responses carry `x-origin-site: primary` again, because *Primary* addresses are always preferred over *Backup* addresses.

- **Users get your page when no origin answers.** Stop both origins and send the request. The response carries status `503` and the body of `/errors/503.html`, not the page titled `Azion - Default error page`. Start both origins again.

A response from a data center that has not yet received a connector or deployment change can still show the previous behavior. Repeat the request until the answers agree.

---

## Measuring results

| Metric                                                   | Where to read it                                                                                                                                                                                                                                                                         | What working looks like                                                            |
| -------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------- |
| Share of requests that succeed during an origin incident | The **HTTP Status Codes 2XX** and **HTTP Status Codes 5XX** charts of Real-Time Metrics, filtered to the host, over the incident                                                                                                                                                         | 2XX responses continue through the incident, carried by the standby                |
| Error rate users see while traffic moves                 | **Requests by Status and Upstream Status**, on the **Status Codes** dashboard. Upstream status `502` marks a request for which no origin server could be selected. Refer to [Build dashboards](/en/documentation/platform/real-time-metrics/build-dashboards/#status-codes)              | 5XX responses rise only while the primary fails, and fall once the standby answers |
| Traffic per origin                                       | The `upstreamAddr` field of the `workloadBreakdownMetrics` dataset, which holds the address and port of the origin that answered, summed by hour. Refer to [Real-Time Metrics GraphQL fields](/en/documentation/devtools/graphql/gql-real-time-metrics-fields/#workloadbreakdownmetrics) | Only the primary's address outside incidents; the standby's address during one     |

---

## Best practices

- **Test the standby in every maintenance window.** A *Backup* address carries no traffic while the primary answers, so nothing exercises it. Run the failure check of this page in each maintenance window, so a standby that drifted from the primary is found before an incident.
- **Set Max Retries and the timeouts in every API call.** A connector whose Load Balancer configuration is created through the API takes `max_retries` `0`, `connection_timeout` `60`, and `read_write_timeout` `120` for every key the body leaves out. A pool without retries fails a request on the first connection failure.
- **Never put the error page behind the pool it covers.** A custom page fetches its content from a connector. A page served by `app-origin` fails in the same incident that should show it.
- **Let Cache answer what it can during the switch.** A request that a valid cached copy answers never reaches either origin. With **Stale cache** on in a cache setting, Azion also serves an expired copy when the origin returns a `5xx` error or times out, for 300 seconds under *Override cache behavior*, and the response reports `STALE`. For when it applies, refer to [Expiration and freshness](/en/documentation/platform/applications/cache/expiration-and-freshness/#stale-cache).
- **Do not use IP Hash for this pool.** *IP Hash* keeps a client on one address, and it refuses *Backup* addresses with `28005`. A pool that needs session affinity across several live origins is the *Active-active origin pool across clouds* design, not this one.

---

## Guides in this use case

- [Add a backup origin to a connector](/en/documentation/guides/application-performance/availability/add-a-backup-origin-to-a-connector.md): Turns on Load Balancer on app-origin, with the primary as a Primary address and the standby as a Backup address.
- [Show your own page when no origin answers](/en/documentation/guides/application-performance/availability/show-your-own-page-when-no-origin-answers.md): Creates the origin-down custom page set and assigns it in the workload's deployment.
