Route model calls through a gateway function
Put one function in front of AI Inference and a third-party provider that authenticates teams, routes by alias, falls back, caches, and logs each call.
You route the model calls of several teams through one function: it reads each team’s record from KV Store, calls the routes of a model alias in order, a model on AI Inference first and a third-party provider on failure, caches the responses a caller allows, and writes one log line per call. To call one model from a function first, refer to Call a model on AI Inference from a function.
Prerequisites
- An application and a workload that serve the gateway’s domain, with Application Accelerator turned on, which the Run Function behavior requires. To create them, refer to Applications quickstart.
- KV Store enabled on the account. The product is in Preview and is not enabled by default, so request access through Technical Support.
- A personal token, for the KV Store call. To create one, refer to Personal tokens.
- The Azion CLI, installed and authorized, to store the environment variables.
- A third-party provider whose chat endpoint accepts the OpenAI chat completions format, with its URL, a model name, and an API key.
The Cache API is not defined under azion dev, so test the function once it is deployed.
The examples use ai-gateway for the KV Store namespace, the function, and the cache, general and long-context for the two model aliases the gateway offers, checkout-team for one team, and gateway.example.com for the domain. Replace them with your values.
Create the namespace
The function opens the ai-gateway namespace on every request, so the namespace exists before the function runs.
To create the namespace, send its name to the KV Store API:
The API answers 201 with the namespace. A namespace cannot be renamed or deleted, so check the name before you send it:
The account holds the empty ai-gateway namespace.
The Govern access to multiple AI models use case uses the values of this example.
Store the values the function reads
The provider key and the admin secret stay out of the code, as environment variables.
To store the values the function reads, run these commands with the Azion CLI. A key that contains key or secret is stored as a secret by default:
The account holds the four variables the gateway function reads with Azion.env.get().
The Govern access to multiple AI models use case uses the values of this example.
Create the gateway function
The function answers /v1/chat/completions for the teams and /admin/teams for the operator, and every other path answers 404. The fallback-test alias names a model id that does not exist, so every call to it falls back to the provider.
Create a function named ai-gateway with this code. The provider call sends the key as Authorization: Bearer; change that header to the one your provider requires:
To create the function and its instance, follow Functions quickstart with the name ai-gateway, and name the instance ai-gateway, with no Args.
The application carries an ai-gateway instance that authenticates a team, routes its request by alias, and logs the call.
The Govern access to multiple AI models use case uses the values of this example.
Run the function on every path
Create a Request Phase rule as Add the rule that runs the function shows, with the name gateway - all paths, the criterion ${uri} starts with /, and the Run Function behavior selecting the ai-gateway instance.
The gateway answers on /v1/chat/completions and /admin/teams, and every other path answers 404. A new rule takes a few minutes to propagate.
Add a team record
Keys are written from a function rather than through the API, so the gateway’s /admin/teams path writes each record. To add a team that may call general and fallback-test, generate a random team key, then send it with the admin secret:
The gateway answers with the record it stored, without the key:
Hand the team key to the team once, because the gateway keeps only its hash. To block the team, send the same request with "status":"blocked", and with "status":"active" to admit it again.
The ai-gateway namespace holds the checkout-team record under the hash of its key.
The Govern access to multiple AI models use case uses the values of this example.
Confirm the gateway routes, falls back, and refuses
Each check sends a chat request with the team key of Add a team record.
To send a request to general:
The response carries x-gateway-route: Qwen/Qwen3-30B-A3B-Instruct-2507-FP8 and a chat.completion object whose generated text sits at choices[0].message.content.
Send the same request with "model":"fallback-test". The response carries x-gateway-route: provider, and the log lines for that request hold one route_failed line for no-such-model, then a model_call line with "fallback":true.
Send the general request twice with the header x-gateway-cache: allow. The second response carries x-gateway-cache: hit.
A request with no Authorization header answers 401. A request for long-context, which checkout-team may not call, answers 403. After you set the team’s status to blocked, any request with its key answers 429.
The gateway routes each team’s request by alias, falls back when a route fails, and refuses what the team record does not allow. Remove the fallback-test alias from ROUTES and from the team record once the fallback check passes.
These checks confirm the Govern access to multiple AI models use case.