AI Inference in the application request path
Serve an OpenAI-compatible inference endpoint from your own application domain, with a function that calls the model with Azion.AI.run.
A client that calls a model needs one address to post to and one response shape to read. This design puts that address on an application you deploy: the application is the inference endpoint, and the model runs behind it. Use it when your team already serves an HTTP API and wants inference to be one more path on it. You own the domain, and no inference server stays running.
How AI Inference works covers the hand-offs inside a single call. This page describes the shape you deploy around them.
Architecture diagram
The diagram traces one request through the deployed shape, from the client to the model and back:
Read the diagram from the application domain outward. To its left is a client that knows nothing about Azion. It posts JSON to a URL and reads JSON back, as it would against any OpenAI-compatible service. Behind the domain, one execution does the rest: the application calls the function, and the function holds the request open until the model responds. No arrow leaves for an inference host, because the function configures no base URL and names no host. It passes a model id, and Azion places the call.
Dataflow
A request moves through the design in this order:
- A client posts an OpenAI-compatible request to
/v1/chat/completionson the application domain. - The nearest data center receives the request and hands it to the application.
- The application applies its policies, checks the authentication you configured, and calls the function.
- The function preprocesses the input and calls the model with
Azion.AI.run, which reaches AI Inference through the Azion Cells API. - AI Inference runs the model and returns the generated response to the function.
- The function shapes that response and returns it to the client along the same path.
Components
- Applications: receives the request on your domain and applies its policies before any of your code runs. It is what makes the endpoint yours, because the domain, the path, and the authentication in front of the function are all configured here. Azion issues no credential for a model call, so an endpoint with no check of its own answers anyone who finds it.
- Functions: holds the inference logic and issues the model call. Your code runs nowhere else in the design, so the prompt, the model id, the context retrieved from a store, and any external service the call needs all sit in one execution.
- AI Inference: runs the model, which is what removes the inference server from the design. You load no model and size no cluster, and the id from AI models is the only address the function needs.
- Orchestrator: manages the requests for Azion Cells execution.
Azion.AI.runreaches a model through the Azion Cells API, which is the part of Orchestrator this design depends on. - Azion’s distributed infrastructure: runs every part above. It is why the design has no region to choose. The data center nearest the client answers the request, and you select none of them.
- Real-Time Metrics: reports the traffic and performance of the application that serves the endpoint. A design with no inference server of your own still needs a view of what the endpoint carried.
- Real-Time Events: tracks individual requests, raw data, and logs. When a response is wrong, the request that produced it is what you read here.
Implementation
- AI Inference quickstart - deploys the design and posts a first request to a model on the domain it assigns.
- How to deploy the AI Inference Starter Kit template - creates the application, the function, and the domain this page describes, in one deployment.
- Model invocation - carries the request body both interfaces accept, and the
curlform for the HTTP endpoint.