AI agents on third-party LLM providers
Azion hosts the agent, the chat state, and the frontend, while an external provider such as OpenAI or Anthropic runs the model.
An agent answers inside a conversation, so every turn pays for the model call plus everything the agent does around it. This design puts everything around the call on Azion: the interface, the agent logic, and the chat state all run here, close to the user. The model is not one of them. It runs at a third-party provider such as OpenAI or Anthropic, and the agent reaches it over that provider’s API with your key.
Use this design when the provider is already chosen. What you want near the user is the agent itself, its history, and the documents it reads. AI Inference describes the design where Azion runs the model instead.
Architecture diagram
The diagram traces one turn of a conversation, from the browser to the provider and back:
Read the diagram from the function outward. Every node on it runs on Azion except one: the third-party LLM sits outside, and the arrow reaching it is the only one that leaves. The function is where the design converges, because it holds the agent, owns the reads and writes of chat state, and places the model call. Object Storage stays off that path. It serves the static files the interface is built from, and the documents the agent later answers from.
Dataflow
A request moves through the design in this order:
- A request arrives at Azion.
- The frontend application serves the user interface, with the static files coming from Object Storage.
- The frontend application sends an API request to the backend function.
- The backend function reads and writes chat state in SQL Database for persistence and context. It then runs the LangGraph agent and streams the response to the client.
- The agent calls the third-party LLM, uses the data in the database to form a response, and sends it back along the same path.
- The frontend application displays the response to the user.
Components
- Applications: hosts the AI agent on Azion. It is the address the browser talks to, and it routes the interface and the API request to the right place behind one domain.
- Functions: contains the AI agent logic. Your code runs nowhere else in the design, so the LangGraph graph, the provider call, and the response stream all sit in one execution.
- Object Storage: stores the data the agent uses to answer questions, and the static files the interface loads. Neither one is fetched from a server you keep running.
- SQL Database: stores chat state and documents. It is what makes a turn contextual, because the agent reads the earlier messages and the matching documents from here before it writes a prompt.
- Third-party LLM: the external service that runs the model, such as OpenAI or Anthropic. It is the one part Azion does not operate, so its availability, its rate limits, and its billing stay with the provider.
Implementation
- How to deploy the LangGraph AI Agent Boilerplate - creates the frontend, the function, and the database this page describes, and carries the document loader that fills them.
- Azion GitHub App - connects the repository that deployment pushes to, which the guide requires before it starts.