Build and run customer support AI assistants
Answer customer questions from your own documentation, with passages retrieved by vector search and an answer generated by a model on Azion.
A support or product team wants an assistant that answers customers from the company’s own documentation, help center, or product data, instead of from a model’s general knowledge. A model alone answers from what it learned in training, so it cannot know a return policy or a plan limit that only your documents state. This page sets up the whole assistant on Azion: a function splits the documents stored in Object Storage, embeds each passage with a model on AI Inference, and stores the passages in SQL Database for vector search. A second function embeds each question, retrieves the closest passages, and asks a model on AI Inference to answer from them and cite them. The result is measured by answer latency, the share of answers that cite retrieved sources, and answer accuracy on a test set.
This use case does not cover AI features that are not conversational. For those, refer to Add AI features to existing applications.
Prerequisites
- An application and a workload that serve your domain, with Application Accelerator turned on, which the Run Function behavior requires. To create them, refer to Applications quickstart.
- SQL Database enabled on the account. The product is in Preview and is not enabled by default, so request access through Technical Support.
- An Object Storage bucket with
workloads_accessset toread_only, holding the documents as UTF-8 text or Markdown files. To create the bucket and upload the files, refer to Object Storage quickstart. - A personal token with the Edit SQL Database permission, for the database calls. To create one, refer to Personal tokens.
- The Azion CLI, installed and authorized, to store the environment variables.
- The values of your own setup. This page uses
support-docsfor the bucket,returns-policy.mdfor one document in it,support-kbfor the database,/admin/ingestand/api/askfor the two paths, andwww.example.comfor the domain. Replace each value with yours in every step.
Required products
| The assistant needs | Which means | Product | Documented in |
|---|---|---|---|
| The company’s documents in one place the assistant reads from | A bucket that the ingestion function reads with the azion:storage module | Object Storage | Object Storage API |
| Passages that can be found by meaning, not by keyword | A table with a vector column and a vector index, queried with vector_top_k | SQL Database | Vector search |
| A vector for every passage and every question, and an answer written from the passages | An embedding model and a chat model called with Azion.AI.run | AI Inference | Embed documents into a vector table with AI Inference |
| Code that splits, embeds, retrieves, and assembles the prompt | Two functions, each run by a rule on its own path | Functions | Functions quickstart |
| Rules that run a function on a path | The Run Function behavior, which requires Application Accelerator on the application | Application Accelerator | How Functions works |
Reference architecture
This page builds the Retrieval-augmented assistant on platform-hosted models: the documents, the vectors, the retrieval, and both models stay on Azion.
The diagram carries two flows that meet at the database. The ingestion flow, at the top, runs when a document changes: it turns a document into passages and vectors, and writes them to the passages table. The request flow runs on every question: the assistant function embeds the question, retrieves passages by vector distance, and asks the chat model to answer from them. Both models run on AI Inference, so neither flow leaves Azion. A failed model call ends inside the function, and the function decides what the customer receives.
Dataflow
- An operator sends
POST /admin/ingestwith a document key, and the ingestion function reads that document from thesupport-docsbucket in Object Storage. - The function splits the document into passages, embeds all of them in one call to the embedding model on AI Inference, and writes them to the
passagestable through the Azion API, because the runtime connection to a database is read-only. - A customer sends a question to
POST /api/ask. The application’s rule runs the assistant function, which embeds the question with the same embedding model. - The assistant function asks the vector index of
passagesfor the four passages nearest to the question, through a read replica of the database. - The function sends the passages, the earlier turns of the conversation, and the question to the chat model on AI Inference, which answers and cites the passages it used.
- The function returns the answer and the list of its sources to the customer. A model call that fails ends inside the function, so the failure flow stays inside Azion.
Components
- Functions: runs ingestion and retrieval. The
support-ingestfunction splits and embeds documents, and thesupport-assistantfunction embeds the question, retrieves passages, assembles the prompt, and calls the chat model. Both models are reached withAzion.AI.run, so the code names no host and holds no model credential. - AI Inference: runs the embedding model,
Qwen/Qwen3-Embedding-4B, that turns passages and questions into vectors, and the chat model,Qwen/Qwen3-30B-A3B-Instruct-2507-FP8, that writes the answer. Both run on Azion’s infrastructure, which keeps the request flow and its failures inside Azion. - SQL Database: stores the passages, their source documents, and their vectors in one table,
passages, in thesupport-kbdatabase. The assistant reads them through a read replica, and ingestion writes them through the Azion API. - Vector Search: the Feature of SQL Database that ranks passages by vector distance. The
passages_idxvector index answersvector_top_kwithout comparing the question with every row, so retrieval reads a fixed number of passages however large the table grows. - Object Storage: holds the source documents in the
support-docsbucket. The ingestion function reads them by key, so the documents stay the single copy that the passages are derived from. - application: the Platform Resource that receives the questions and serves the front end. Its rules decide which path runs which function, and they run before the function does.
Other designs for this use case
- Retrieval-augmented assistant over third-party LLMs: for teams committed to a model provider such as OpenAI or Anthropic. A function on Azion still retrieves the passages from SQL Database, but it sends them with the question to the provider’s API and keeps conversation state between turns, so generation crosses to an external provider and adds provider keys, provider latency, and fallback decisions to the request flow.
Configure the passages table
The passages table holds each passage, the document it came from, and its vector. The vector column declares F32_BLOB(1024) because the functions request 1,024-dimension vectors from Qwen/Qwen3-Embedding-4B, one of the five widths the model returns. The column and the request must state the same number, and a narrower vector stores less per passage. The index uses the cosine metric, and the table carries INTEGER PRIMARY KEY AUTOINCREMENT, because a vector index requires a ROWID or a single-column primary key.
To create the database, send its name to the SQL Database API:
The API answers 202 with the new database. Keep its id, which the next call and the ingestion function use:
Provisioning takes about 15 seconds. Send GET /v4/workspace/sql/databases/<database-id> until status reads created. Then create the table and its index:
The API answers 200 with "state": "executed" and one entry in data per statement. A statement that fails still answers 200, with error in place of results in its entry, so read both entries before you continue.
The support-kb database holds an empty passages table and the passages_idx vector index. The index also adds a shadow table, passages_idx_shadow, that a table listing shows beside passages.
Configure the ingestion function
The ingestion function turns one document into rows of passages. It reads, splits, and embeds the document as Embed documents into a vector table with AI Inference describes, and writes the rows as Write rows to SQL Database from a function describes, with the assistant’s values:
- Path:
POST /admin/ingest?key=<object-key>, one document per request. The route writes to the database, so it refuses a request without the ingestion secret. - Bucket:
support-docs, read fromDOCS_BUCKET. - Passages: runs of paragraphs up to 1,500 characters. A short passage points the answer at one section of a document, and four of them still fit the prompt with room for the conversation. A single paragraph longer than that stays one passage, which the embedding model’s 32k-token context reads whole.
- Passages per document: at most 99, so the
DELETEfor the document’s old rows and oneINSERTper passage fit one call of 100 statements. A longer document is refused with413, so split it into two files. - Vectors:
Qwen/Qwen3-Embedding-4Bat 1,024 dimensions, the width of theembeddingcolumn. - Write credentials:
SQL_DATABASE_IDandAZION_TOKEN.
To store the four values the function reads, run these commands with the Azion CLI. A key that contains token or secret is stored as a secret by default:
Create the variables before the function: a function that is already running does not read a changed value until it is redeployed.
Create a function named support-ingest with this code:
Run the function on its path with these values, following Functions quickstart:
- Function instance:
support-ingest, with no Args. - Rule: a Request Phase rule named
support - ingest, with the criterion${uri}starts with/admin/ingestand the Run Function behavior selecting thesupport-ingestinstance.
A POST to /admin/ingest with the secret and a document key replaces that document’s passages in passages, and answers with the number it stored. The DELETE makes a second ingestion of the same key replace its rows instead of duplicating them.
Configure the assistant function
The assistant function answers one question per request. It embeds the question with the same model and width as the passages, because a query vector is comparable only with vectors that model produced. It retrieves the four nearest passages: four passages of up to 1,500 characters ground the answer in more than one section, and keep the prompt far below the 64k-token context of the chat model.
Three more decisions are in the code:
- The query vector is written into the SQL text. The runtime refuses a JavaScript string as a parameter value, so the vector goes inside
vector('[...]'), as the vector search reference does. It holds only numbers the embedding model returned. - The client carries the conversation. The request body holds
history, the earlieruserandassistantmessages, and the function keeps the last six, which is three turns. The function stores nothing between requests, and a bounded history keeps the prompt from growing with each turn. max_tokensis800. The function waits while the model generates, so a cap on the answer’s length is also a cap on the time a customer waits.
The function writes one log line per answer, recording whether the answer cites a passage. That line is the source of the citation metric in Measuring results.
To store the database name the function opens, run:
Create a function named support-assistant with this code:
Run the function on its path with these values, following Functions quickstart:
- Function instance:
support-assistant, with no Args. - Rule: a Request Phase rule named
support - ask, with the criterion${uri}is equal/api/askand the Run Function behavior selecting thesupport-assistantinstance.
A POST to /api/ask with a question answers with the model’s answer, which cites passages as [n], and with the document each cited passage came from. A new rule takes a few minutes to propagate.
Verify the setup
-
A document becomes passages. Ingest the example document:
ShellThe function answers with the key and the number of passages it stored, such as
{"key":"returns-policy.md","passages":3}. A request without theAuthorizationheader answers401. -
The passages carry vectors. Count the rows of the document through the SQL Database API:
ShellThe single entry in
datacarriesresults, whose one row holds the same number the ingestion returned. -
An answer comes from the documents and cites them. Ask a question the document answers:
ShellThe response carries
answer, with at least one[n]citation, andsources, whose entries namereturns-policy.md. -
A question outside the documents is not answered from training. Ask a question no document covers, such as
What is the capital of France?. The answer says that the assistant does not know, and it carries no citation.
When a request answers 404 or the default page of the application, the rule may still be propagating. When it persists after a few minutes, read the function’s log lines in Real-Time Events, under the Functions Console data source.
Measuring results
| Metric | Where to read it | What working looks like |
|---|---|---|
| Answer latency | The Request Time of requests to /api/ask, in the HTTP Requests data source of Real-Time Events | Stable as the number of documents grows, because each answer reads four passages whatever the size of the table |
| Share of answers that cite retrieved sources | The "event":"answer" lines the assistant function writes, in the Functions Console data source of Real-Time Events: the lines with "cited":true over all of them | Close to every answer. A drop points at a question the documents do not cover, or at passages that no longer match the documents |
| Answer accuracy on a test set | A fixed list of questions with known answers from your documents, sent to /api/ask after every change to the documents, the prompt, or the models | The share of correct answers holds or rises from one run to the next |
Best practices
- Embed questions and passages with the same model and width. Vector distance compares two vectors only when one model produced both at one width. Changing the model or
dimensionsmeans a new column declaration and a full re-ingestion, so keep both values in one constant shared by the two functions. - Re-ingest a document when it changes. The
DELETEbefore theINSERTstatements replaces a document’s passages, so the table never answers from a superseded version. A retired document needs its ownDELETE FROM passages WHERE source = '<key>';. - Check every statement entry as well as the status. The SQL Database API answers
200even when a statement fails, witherrorin that statement’s entry. The ingestion function reads every entry, and any script that writes passages should do the same. For the error shapes, refer to SQL Database best practices. - Read the model’s text with optional chaining. The function reads
modelResponse?.choices?.[0]?.message?.content, so a response without one of those levels yields an empty answer instead of an exception. For the reasoning, refer to AI Inference best practices.