Govern access to multiple AI models
Put one gateway function in front of AI Inference and a third-party provider to authenticate teams, route by policy, fall back, cache, and log every call.
A platform team supports product teams that call AI models from several providers and from AI Inference. Each team holds its own provider keys and its own fallback code, so nobody can route a request to another model when a provider fails, stop a team that spends past its budget, or say which model answered a given request. This page puts a gateway built on Functions in front of every model. Applications call one endpoint with a team key, and the gateway authenticates the team, picks a model by policy, falls back to the next model when a call fails, caches responses the caller allows, and logs every call. The result is measured by requests answered despite a provider failure, spend per team within budget, the share of requests answered from cache, and audit coverage of model calls.
This use case does not cover building the applications or agents that call the models. For that, refer to Build AI agents.
Prerequisites
- An application and a workload that serve the gateway’s domain, with Application Accelerator turned on, which the Run Function behavior requires. To create them, refer to Applications quickstart.
- KV Store enabled on the account. The product is in Preview and is not enabled by default, so request access through Technical Support.
- A personal token, for the KV Store call. To create one, refer to Personal tokens.
- The Azion CLI, installed and authorized, to store the environment variables.
- A third-party provider whose chat endpoint accepts the OpenAI chat completions format, with its URL, a model name, and an API key.
- An HTTPS endpoint of your analytics or log platform that accepts
POSTrequests, for the usage and audit logs. - The values of your own setup. This page uses
ai-gatewayfor the KV Store namespace, the function, and the cache,generalandlong-contextfor the two model aliases the gateway offers,checkout-teamfor one team, andgateway.example.comfor the domain. Replace each value with yours in every step.
Required products
| The gateway needs | Which means | Product | Documented in |
|---|---|---|---|
| One endpoint that holds the routing, fallback, and authentication logic | A function on the gateway’s application, run on every path by one rule | Functions | Functions quickstart |
| Platform-hosted models among the routes | Azion.AI.run with a model id from the AI Inference model catalog | AI Inference | Call a model on AI Inference from a function |
| Team keys, the models each team may call, and whether a team is still within budget | One record per team, read by key on every request | KV Store | KV Store API |
| Repeated prompts answered without a model call, where the caller allows it | Responses stored and matched with the Cache API | Cache | Cache a function’s response with the Cache API |
| A usage and audit record of every model call | A log line per call, streamed from the Functions data source | Data Stream | Send logs to an HTTP endpoint |
| One request investigated after the fact | The function’s log lines for that request | Real-Time Events | Real-Time Events data sources |
| A rule that runs the function | The Run Function behavior, which requires Application Accelerator on the application | Application Accelerator | How Functions works |
Reference architecture
This page builds the Multi-model AI gateway: one function in front of every model, with the policy, the keys, and the log in one place.
Read the diagram from the function outward. Every model call passes through it, so it is the one place where a team is identified, a model is chosen, and a call is recorded. The arrows to the models are ordered: the policy names a first route and the routes after it, and the function moves to the next one only when a call fails. The arrows to KV Store and the Cache API are reads that come before any model call, so a refused team or a stored answer costs no model call at all.
Dataflow
- A team application sends an OpenAI-compatible request to
/v1/chat/completions, with its team key as a bearer token and a model alias inmodelrather than a model id. - The gateway hashes the key and reads the team’s record from KV Store. An unknown key answers
401, a team over budget answers429, and an alias the team may not call answers403. - When the caller allows caching, the gateway looks the request up in the cache and returns a stored response on a match.
- Otherwise the gateway calls the routes of the alias in order: a model on AI Inference first, then the third-party provider when the first call fails.
- The gateway writes one log line with the team, the alias, the model that answered, whether it fell back, the cache result, and the tokens used. Data Stream sends it to your analytics platform, and Real-Time Events holds the lines of each request for investigation.
- The gateway returns the model’s response, and stores it in the cache when the caller allowed it.
Components
- Functions: runs authentication, routing, and fallback. The policy is code in the
ai-gatewayfunction, so a change to a route reaches every team at once, and no team holds a provider key or fallback logic of its own. - AI Inference: runs the platform-hosted models, called with
Azion.AI.runand a model id. A call to it stays inside Azion and needs no provider key. - third-party LLM providers: the integrations that run external models. The function calls them with keys it reads from environment variables, so their latency and failures enter only the routes that name them.
- Cache: holds responses to repeated prompts, through the Cache API of the function. A stored response is returned without a model call, only for requests whose caller allows it.
- KV Store: holds the keys, quotas, and budgets: one record per team in the
ai-gatewaynamespace, keyed by a hash of the team key, with the models the team may call and whether it is within budget. The function reads it on every request. - Data Stream: sends the usage and audit log lines the function writes, from the Functions data source, to your analytics platform, where spend per team is summed.
- Real-Time Events: holds the function’s log lines grouped by request, for investigating one call after the fact.
- application: the Platform Resource that is the gateway endpoint. A rule runs the function on its paths, so every application points at one domain.
Configure the gateway function
The gateway is one function. Every decision it applies is in the code, so a policy change is a code change that every team gets at once.
- Teams send an alias, not a model id.
generalroutes toQwen/Qwen3-30B-A3B-Instruct-2507-FP8on AI Inference, andlong-contexttogpt-oss-20b, whose 131k-token context is the longest among the models AI Inference runs. Both fall back to the provider. An alias lets the platform team change the model behind it without a change in any team’s code. The gateway calls an AI Inference route as Call a model on AI Inference from a function describes, with these values: the route’s model id, and the team’s request body withstreamset tofalse. - The key is stored as a hash. The record key is
team:followed by the SHA-256 of the team key, so KV Store never holds a usable key. SHA-256 is one of the digestscrypto.subtle.digestsupports. - A failed call moves to the next route. A call that throws, returns no
choices, answers with an error status, or runs past 30 seconds counts as failed. Thirty seconds bounds how long a team waits on one route before the next one runs, inside the 5-minute wall-clock limit of a function. - Caching is the caller’s decision. The gateway caches only requests that carry
x-gateway-cache: allow, because only the team knows whether one answer can stand for every identical prompt. The gateway caches as Cache a function’s response with the Cache API describes, with the cacheai-gateway, the keyhttps://ai-gateway.cache/<sha-256 of the alias and the request body>, andmax-age=3600, one hour. - Every call writes one log line. The line is the usage and audit record: a team’s spend is the sum of its
total_tokens, and the line names the model that answered. - The
fallback-testalias proves the fallback. Its first route names a model id that does not exist, so every call to it falls back. Remove it after Verify the setup.
To store the values the function reads, run these commands with the Azion CLI. A key that contains key or secret is stored as a secret by default:
Create a function named ai-gateway with this code. The provider call sends the key as Authorization: Bearer; change that header to the one your provider requires:
Run the function with these values, following Functions quickstart:
- Function instance:
ai-gateway, with no Args. - Rule: a Request Phase rule named
gateway - all paths, with the criterion${uri}starts with/and the Run Function behavior selecting theai-gatewayinstance.
The gateway answers on /v1/chat/completions and /admin/teams, and every other path answers 404. The Cache API is not defined under azion dev, so test the function once it is deployed.
Configure the team records
A team record is a JSON object under team:<sha256 of the team key> in the ai-gateway namespace: the team’s name, the aliases it may call, and its status. The gateway reads the record on every request, and admits a request only while status is active.
Budget enforcement reads that status rather than counting each call. KV Store accepts one write per second to the same key and offers no atomic increment, so a counter written on every request would lose updates under concurrent traffic. Spend is computed instead in your analytics platform from the total_tokens of each team’s log lines. When a team reaches its budget, an operator sets its status to blocked, and the gateway answers 429 to that team from then on.
To create the namespace, send its name to the KV Store API:
The API answers 201 with the namespace. A namespace cannot be renamed or deleted, so check the name before you send it:
Keys are written from a function rather than through the API, so the gateway’s /admin/teams path writes each record. To add a team that may call general and fallback-test, generate a random team key, then send it with the admin secret:
The gateway answers with the record it stored, without the key:
Hand the team key to the team once, because the gateway keeps only its hash. To block the team, send the same request with "status":"blocked", and with "status":"active" to admit it again.
Configure the usage and audit stream
Every model_call line is a usage and audit record, and Data Stream delivers the lines from the Functions data source to your analytics platform, where spend and fallbacks are summed per team.
Create the stream as Send logs to an HTTP endpoint shows, with these values:
- Data Source: Functions. In the API,
functions_console. - Template: Functions Event Collector, whose
$log_messagevariable carries each line the function writes. - Option: Filter Workloads, with the gateway’s workload as the only chosen workload. A filter keeps this stream from deactivating the other streams on the account, which saving an active stream with sampling does.
- Connector: Standard HTTP/HTTPS POST, with your analytics platform’s URL and the header it requires to accept the request.
The stream becomes active within one to two minutes of saving it. Your analytics platform then receives one model_call line per request, and one route_failed line per route that failed.
Verify the setup
Each check sends a chat request with the team key from the team records step.
-
A team reaches a model through an alias. Send a request to
general:ShellThe response carries
x-gateway-route: Qwen/Qwen3-30B-A3B-Instruct-2507-FP8and achat.completionobject whose generated text sits atchoices[0].message.content. -
A failed route falls back. Send the same request with
"model":"fallback-test". The response carriesx-gateway-route: provider, and the log lines for that request hold oneroute_failedline forno-such-model, then amodel_callline with"fallback":true. -
A repeated prompt the caller allows is answered from cache. Send the
generalrequest twice with the headerx-gateway-cache: allow. The second response carriesx-gateway-cache: hit. -
The gateway refuses what the policy refuses. A request with no
Authorizationheader answers401. A request forlong-context, whichcheckout-teammay not call, answers403. After you set the team’s status toblocked, any request with its key answers429. -
Every call is logged. Your analytics platform receives a
model_callline for each request above, with"team":"checkout-team". To read the lines of one request, open the Functions Console data source in Real-Time Events, where theIDvariable groups the lines of a single request.
Remove the fallback-test alias from ROUTES and from the team record once the fallback check passes. A new rule takes a few minutes to propagate.
Measuring results
| Metric | Where to read it | What working looks like |
|---|---|---|
| Requests answered despite a provider failure | The model_call lines with "ok":true and "fallback":true, in your analytics platform | Every route failure in a route_failed line is followed by a successful model_call for the same request, unless every route failed |
| Spend per team within budget | The sum of total_tokens per team over the budget period, in your analytics platform | Each team’s sum stays under its budget, and a team that reaches it has status set to blocked |
| Share of requests answered from cache | The model_call lines with "cache":"hit" over the lines with "cache":"hit" or "cache":"miss" | Rises for teams that allow caching on prompts that repeat |
| Audit coverage of model calls | The count of model_call lines against the gateway’s invocations in the Functions tab of Real-Time Metrics, less the /admin/teams calls and the refused requests | The two counts match, so every model call has a record |
Best practices
- Give each team its own key, and rotate it by replacing the record. The record key is the hash of the team key, so a new key is a new record. Set the old record to
blockedonce the team switches, because a KV Store key cannot be listed and a forgotten record stays readable. - Keep the provider key in an environment variable. The function reads
PROVIDER_API_KEYwithAzion.env.get(), so the key never appears in the code, and teams never hold it. A changed variable reaches the function only after it is redeployed. - Let teams opt in to caching per request. A cached answer is returned for every identical prompt, whichever team sends it. Caching a prompt whose answer must change, such as one that asks about the current state of an account, returns a stale answer for an hour.
- Account for KV Store reads at the gateway’s request rate. Every request reads one team record, and KV Store includes 100,000 keys read per day before a charge applies. For the included amounts and the rate past them, refer to KV Store limits.