# Build and run customer support AI assistants

A support or product team wants an assistant that answers customers from the company's own documentation, help center, or product data, instead of from a model's general knowledge. A model alone answers from what it learned in training, so it cannot know a return policy or a plan limit that only your documents state. This page sets up the whole assistant on Azion: a function splits the documents stored in Object Storage, embeds each passage with a model on AI Inference, and stores the passages in SQL Database for vector search. A second function embeds each question, retrieves the closest passages, and asks a model on AI Inference to answer from them and cite them. The result is measured by answer latency, the share of answers that cite retrieved sources, and answer accuracy on a test set.

This use case does not cover AI features that are not conversational. For those, refer to [Add AI features to existing applications](/en/documentation/use-cases/build-and-run-ai-workloads/add-ai-features-to-existing-applications/).

## Prerequisites

- An application and a workload that serve your domain, with **Application Accelerator** turned on, which the **Run Function** behavior requires. To create them, refer to [Applications quickstart](/en/documentation/platform/applications/quickstart/).
- SQL Database enabled on the account. The product is in Preview and is not enabled by default, so request access through [Technical Support](/en/documentation/support/).
- An Object Storage bucket with `workloads_access` set to `read_only`, holding the documents as UTF-8 text or Markdown files. To create the bucket and upload the files, refer to [Object Storage quickstart](/en/documentation/platform/object-storage/quickstart/).
- A personal token with the **Edit SQL Database** permission, for the database calls. To create one, refer to [Personal tokens](/en/documentation/guides/platform/account-and-billing/personal-tokens/).
- The [Azion CLI](/en/documentation/devtools/cli/), installed and authorized, to store the environment variables.
- The values of your own setup. This page uses `support-docs` for the bucket, `returns-policy.md` for one document in it, `support-kb` for the database, `/admin/ingest` and `/api/ask` for the two paths, and `www.example.com` for the domain. Replace each value with yours in every step.

---

## Required products

| The assistant needs                                                                    | Which means                                                                              | Product                 | Documented in                                                                                                                            |
| -------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------- | ----------------------- | ---------------------------------------------------------------------------------------------------------------------------------------- |
| The company's documents in one place the assistant reads from                          | A bucket that the ingestion function reads with the `azion:storage` module               | Object Storage          | [Object Storage API](/en/documentation/devtools/runtime/api-reference/storage/)                                                          |
| Passages that can be found by meaning, not by keyword                                  | A table with a vector column and a vector index, queried with `vector_top_k`             | SQL Database            | [Vector search](/en/documentation/platform/sql-database/vector-search/)                                                                  |
| A vector for every passage and every question, and an answer written from the passages | An embedding model and a chat model called with `Azion.AI.run`                           | AI Inference            | [Embed documents into a vector table with AI Inference](/en/documentation/guides/ai/agents-and-rag/embed-documents-into-a-vector-table/) |
| Code that splits, embeds, retrieves, and assembles the prompt                          | Two functions, each run by a rule on its own path                                        | Functions               | [Functions quickstart](/en/documentation/platform/functions/quickstart/)                                                                 |
| Rules that run a function on a path                                                    | The **Run Function** behavior, which requires Application Accelerator on the application | Application Accelerator | [How Functions works](/en/documentation/platform/functions/how-it-works/#execution-phases-and-criteria)                                  |

---

## Reference architecture

This page builds the *Retrieval-augmented assistant on platform-hosted models*: the documents, the vectors, the retrieval, and both models stay on Azion.

```mermaid
%%{init: {"layout": "dagre", "themeVariables": {"fontSize": "13px"}, "flowchart": {"nodeSpacing": 12, "rankSpacing": 12, "padding": 6, "wrappingWidth": 70, "minNodeWidth": 40, "useMaxWidth": true}}}%%
flowchart TD
  Bucket["Object Storage bucket: support-docs"] -->|"document text"| Ingest["ingestion function on /admin/ingest"]
  Ingest -->|"passages"| Embed["AI Inference: Qwen3 Embedding 4B"]
  Embed -->|"vectors"| Ingest
  Ingest -->|"INSERT through the Azion API"| DB["SQL Database: passages table and vector index"]
  Customer["Customer"] -->|"POST /api/ask"| Assist["assistant function on /api/ask"]
  Assist -->|"question"| Embed
  Assist -->|"vector_top_k, read replica"| DB
  Assist -->|"passages and question"| Chat["AI Inference: Qwen3 30B A3B Instruct"]
  Chat -->|"answer with citations"| Assist
  Assist -->|"answer and sources"| Customer
```

The diagram carries two flows that meet at the database. The ingestion flow, at the top, runs when a document changes: it turns a document into passages and vectors, and writes them to the `passages` table. The request flow runs on every question: the assistant function embeds the question, retrieves passages by vector distance, and asks the chat model to answer from them. Both models run on AI Inference, so neither flow leaves Azion. A failed model call ends inside the function, and the function decides what the customer receives.

### Dataflow

1. An operator sends `POST /admin/ingest` with a document key, and the ingestion function reads that document from the `support-docs` bucket in Object Storage.
2. The function splits the document into passages, embeds all of them in one call to the embedding model on AI Inference, and writes them to the `passages` table through the Azion API, because the runtime connection to a database is read-only.
3. A customer sends a question to `POST /api/ask`. The application's rule runs the assistant function, which embeds the question with the same embedding model.
4. The assistant function asks the vector index of `passages` for the four passages nearest to the question, through a read replica of the database.
5. The function sends the passages, the earlier turns of the conversation, and the question to the chat model on AI Inference, which answers and cites the passages it used.
6. The function returns the answer and the list of its sources to the customer. A model call that fails ends inside the function, so the failure flow stays inside Azion.

### Components

- **Functions**: runs ingestion and retrieval. The `support-ingest` function splits and embeds documents, and the `support-assistant` function embeds the question, retrieves passages, assembles the prompt, and calls the chat model. Both models are reached with `Azion.AI.run`, so the code names no host and holds no model credential.
- **AI Inference**: runs the embedding model, `Qwen/Qwen3-Embedding-4B`, that turns passages and questions into vectors, and the chat model, `Qwen/Qwen3-30B-A3B-Instruct-2507-FP8`, that writes the answer. Both run on Azion's infrastructure, which keeps the request flow and its failures inside Azion.
- **SQL Database**: stores the passages, their source documents, and their vectors in one table, `passages`, in the `support-kb` database. The assistant reads them through a read replica, and ingestion writes them through the Azion API.
- **Vector Search**: the Feature of SQL Database that ranks passages by vector distance. The `passages_idx` vector index answers `vector_top_k` without comparing the question with every row, so retrieval reads a fixed number of passages however large the table grows.
- **Object Storage**: holds the source documents in the `support-docs` bucket. The ingestion function reads them by key, so the documents stay the single copy that the passages are derived from.
- **application**: the Platform Resource that receives the questions and serves the front end. Its rules decide which path runs which function, and they run before the function does.

### Other designs for this use case

- *Retrieval-augmented assistant over third-party LLMs*: for teams committed to a model provider such as OpenAI or Anthropic. A function on Azion still retrieves the passages from SQL Database, but it sends them with the question to the provider's API and keeps conversation state between turns, so generation crosses to an external provider and adds provider keys, provider latency, and fallback decisions to the request flow.

---

## Configure the passages table

The passages table holds each passage, the document it came from, and its vector. The vector column declares `F32_BLOB(1024)` because the functions request 1,024-dimension vectors from `Qwen/Qwen3-Embedding-4B`, one of the five widths the model returns. The column and the request must state the same number, and a narrower vector stores less per passage. The index uses the cosine metric, and the table carries `INTEGER PRIMARY KEY AUTOINCREMENT`, because a vector index requires a `ROWID` or a single-column primary key.

To create the database, send its name to the SQL Database API:

```bash
curl --request POST \
  --url https://api.azion.com/v4/workspace/sql/databases \
  --header 'Accept: application/json' \
  --header 'Authorization: Token [TOKEN VALUE]' \
  --header 'Content-Type: application/json' \
  --data '{"name":"support-kb"}'
```

The API answers `202` with the new database. Keep its `id`, which the next call and the ingestion function use:

```json
{
  "state": "pending",
  "data": {
    "id": <database-id>,
    "name": "support-kb",
    "status": "creating",
    "active": true,
    ...
  }
}
```

Provisioning takes about 15 seconds. Send `GET /v4/workspace/sql/databases/<database-id>` until `status` reads `created`. Then create the table and its index:

```bash
curl --request POST \
  --url https://api.azion.com/v4/workspace/sql/databases/<database-id>/query \
  --header 'Accept: application/json' \
  --header 'Authorization: Token [TOKEN VALUE]' \
  --header 'Content-Type: application/json' \
  --data '{"statements":[
    "CREATE TABLE passages (id INTEGER PRIMARY KEY AUTOINCREMENT, source TEXT NOT NULL, content TEXT NOT NULL, embedding F32_BLOB(1024));",
    "CREATE INDEX passages_idx ON passages (libsql_vector_idx(embedding, '\''metric=cosine'\''));"
  ]}'
```

The API answers `200` with `"state": "executed"` and one entry in `data` per statement. A statement that fails still answers `200`, with `error` in place of `results` in its entry, so read both entries before you continue.

The `support-kb` database holds an empty `passages` table and the `passages_idx` vector index. The index also adds a shadow table, `passages_idx_shadow`, that a table listing shows beside `passages`.

---

## Configure the ingestion function

The ingestion function turns one document into rows of `passages`. It reads, splits, and embeds the document as [Embed documents into a vector table with AI Inference](/en/documentation/guides/ai/agents-and-rag/embed-documents-into-a-vector-table/) describes, and writes the rows as [Write rows to SQL Database from a function](/en/documentation/guides/application-development/data/write-sql-database-rows-from-a-function/) describes, with the assistant's values:

- **Path**: `POST /admin/ingest?key=<object-key>`, one document per request. The route writes to the database, so it refuses a request without the ingestion secret.
- **Bucket**: `support-docs`, read from `DOCS_BUCKET`.
- **Passages**: runs of paragraphs up to 1,500 characters. A short passage points the answer at one section of a document, and four of them still fit the prompt with room for the conversation. A single paragraph longer than that stays one passage, which the embedding model's 32k-token context reads whole.
- **Passages per document**: at most 99, so the `DELETE` for the document's old rows and one `INSERT` per passage fit one call of 100 statements. A longer document is refused with `413`, so split it into two files.
- **Vectors**: `Qwen/Qwen3-Embedding-4B` at 1,024 dimensions, the width of the `embedding` column.
- **Write credentials**: `SQL_DATABASE_ID` and `AZION_TOKEN`.

To store the four values the function reads, run these commands with the Azion CLI. A key that contains `token` or `secret` is stored as a secret by default:

```bash
azion create variables --key "AZION_TOKEN" --value "[TOKEN VALUE]"
azion create variables --key "INGEST_SECRET" --value "<ingest-secret>"
azion create variables --key "SQL_DATABASE_ID" --value "<database-id>" --secret false
azion create variables --key "DOCS_BUCKET" --value "support-docs" --secret false
```

Create the variables before the function: a function that is already running does not read a changed value until it is redeployed.

Create a function named `support-ingest` with this code:

```javascript
import Storage from "azion:storage";

const EMBEDDING_MODEL = "Qwen/Qwen3-Embedding-4B";
const DIMENSIONS = 1024;
const MAX_PASSAGE = 1500;
const MAX_PASSAGES = 99;

function split(text) {
  const passages = [];
  let current = "";
  for (const paragraph of text.split(/\n\s*\n/)) {
    const p = paragraph.trim();
    if (!p) continue;
    if (current && current.length + p.length > MAX_PASSAGE) {
      passages.push(current);
      current = "";
    }
    current = current ? current + "\n\n" + p : p;
  }
  if (current) passages.push(current);
  return passages;
}

const quote = (value) => "'" + value.replaceAll("'", "''") + "'";

export default {
  async fetch(request, env, ctx) {
    if (request.method !== "POST") {
      return new Response("Method not allowed", { status: 405 });
    }
    if (request.headers.get("Authorization") !== `Bearer ${Azion.env.get("INGEST_SECRET")}`) {
      return new Response("Unauthorized", { status: 401 });
    }
    const key = new URL(request.url).searchParams.get("key");
    if (!key) {
      return Response.json({ error: "key is required" }, { status: 400 });
    }

    const object = await new Storage(Azion.env.get("DOCS_BUCKET")).get(key);
    const passages = split(new TextDecoder().decode(await object.arrayBuffer()));
    if (passages.length === 0 || passages.length > MAX_PASSAGES) {
      return Response.json({ error: `${passages.length} passages; split the document` }, { status: 413 });
    }

    const embedded = await Azion.AI.run(EMBEDDING_MODEL, {
      "input": passages,
      "encoding_format": "float",
      "dimensions": DIMENSIONS
    });

    const statements = [`DELETE FROM passages WHERE source = ${quote(key)};`];
    for (const item of embedded.data) {
      statements.push(
        `INSERT INTO passages (source, content, embedding) VALUES (${quote(key)}, ${quote(passages[item.index])}, vector('[${item.embedding.join(",")}]'));`
      );
    }

    const result = await fetch(
      `https://api.azion.com/v4/workspace/sql/databases/${Azion.env.get("SQL_DATABASE_ID")}/query`,
      {
        method: "POST",
        headers: {
          "Accept": "application/json",
          "Authorization": `Token ${Azion.env.get("AZION_TOKEN")}`,
          "Content-Type": "application/json"
        },
        body: JSON.stringify({ statements })
      }
    );
    const body = await result.json();
    const failed = (body.data ?? []).find((entry) => entry.error);
    if (!result.ok || failed) {
      return Response.json({ error: failed?.error ?? `SQL API answered ${result.status}` }, { status: 502 });
    }
    return Response.json({ key, passages: passages.length });
  },
};
```

Run the function on its path with these values, following [Functions quickstart](/en/documentation/platform/functions/quickstart/):

- **Function instance**: `support-ingest`, with no Args.
- **Rule**: a Request Phase rule named `support - ingest`, with the criterion `${uri}` *starts with* `/admin/ingest` and the **Run Function** behavior selecting the `support-ingest` instance.

A `POST` to `/admin/ingest` with the secret and a document key replaces that document's passages in `passages`, and answers with the number it stored. The `DELETE` makes a second ingestion of the same key replace its rows instead of duplicating them.

---

## Configure the assistant function

The assistant function answers one question per request. It embeds the question with the same model and width as the passages, because a query vector is comparable only with vectors that model produced. It retrieves the four nearest passages: four passages of up to 1,500 characters ground the answer in more than one section, and keep the prompt far below the 64k-token context of the chat model.

Three more decisions are in the code:

- **The query vector is written into the SQL text.** The runtime refuses a JavaScript string as a parameter value, so the vector goes inside `vector('[...]')`, as the vector search reference does. It holds only numbers the embedding model returned.
- **The client carries the conversation.** The request body holds `history`, the earlier `user` and `assistant` messages, and the function keeps the last six, which is three turns. The function stores nothing between requests, and a bounded history keeps the prompt from growing with each turn.
- **`max_tokens` is `800`.** The function waits while the model generates, so a cap on the answer's length is also a cap on the time a customer waits.

The function writes one log line per answer, recording whether the answer cites a passage. That line is the source of the citation metric in [Measuring results](#measuring-results).

To store the database name the function opens, run:

```bash
azion create variables --key "SQL_DATABASE_NAME" --value "support-kb" --secret false
```

Create a function named `support-assistant` with this code:

```javascript
const EMBEDDING_MODEL = "Qwen/Qwen3-Embedding-4B";
const CHAT_MODEL = "Qwen/Qwen3-30B-A3B-Instruct-2507-FP8";
const DIMENSIONS = 1024;
const TOP_K = 4;

const SYSTEM_PROMPT =
  "You answer customer questions using only the numbered passages in the user message. " +
  "Cite every passage you use as [n]. If the passages do not contain the answer, " +
  "say that you do not know and suggest contacting support.";

async function retrieve(question) {
  const embedded = await Azion.AI.run(EMBEDDING_MODEL, {
    "input": question,
    "encoding_format": "float",
    "dimensions": DIMENSIONS
  });
  const vector = embedded.data[0].embedding.join(",");

  const { Database } = Azion.Sql;
  const connection = await Database.open(Azion.env.get("SQL_DATABASE_NAME"));
  const rows = await connection.query(
    `SELECT passages.source, passages.content FROM vector_top_k('passages_idx', vector('[${vector}]'), ${TOP_K}) JOIN passages ON passages.rowid = id;`
  );

  const passages = [];
  let row = await rows.next();
  while (row) {
    passages.push({ source: row.getString(0), content: row.getString(1) });
    row = await rows.next();
  }
  return passages;
}

export default {
  async fetch(request, env, ctx) {
    if (request.method !== "POST") {
      return new Response("Method not allowed", { status: 405 });
    }
    const { question, history = [] } = await request.json();
    if (typeof question !== "string" || question.trim() === "") {
      return Response.json({ error: "question is required" }, { status: 400 });
    }

    const passages = await retrieve(question);
    const context = passages
      .map((p, i) => `[${i + 1}] (${p.source})\n${p.content}`)
      .join("\n\n");

    const modelResponse = await Azion.AI.run(CHAT_MODEL, {
      "stream": false,
      "max_tokens": 800,
      "messages": [
        { "role": "system", "content": SYSTEM_PROMPT },
        ...history.slice(-6),
        { "role": "user", "content": `Passages:\n${context}\n\nQuestion: ${question}` }
      ]
    });

    const answer = modelResponse?.choices?.[0]?.message?.content ?? "";
    const cited = /\[\d+\]/.test(answer);
    console.log(JSON.stringify({ event: "answer", cited, passages: passages.length }));

    return Response.json({
      answer,
      sources: passages.map((p, i) => ({ n: i + 1, source: p.source }))
    });
  },
};
```

Run the function on its path with these values, following [Functions quickstart](/en/documentation/platform/functions/quickstart/):

- **Function instance**: `support-assistant`, with no Args.
- **Rule**: a Request Phase rule named `support - ask`, with the criterion `${uri}` *is equal* `/api/ask` and the **Run Function** behavior selecting the `support-assistant` instance.

A `POST` to `/api/ask` with a question answers with the model's answer, which cites passages as `[n]`, and with the document each cited passage came from. A new rule takes a few minutes to propagate.

---

## Verify the setup

- **A document becomes passages.** Ingest the example document:

  ```bash
  curl -X POST 'https://www.example.com/admin/ingest?key=returns-policy.md' \
    -H 'Authorization: Bearer <ingest-secret>'
  ```

  The function answers with the key and the number of passages it stored, such as `{"key":"returns-policy.md","passages":3}`. A request without the `Authorization` header answers `401`.

- **The passages carry vectors.** Count the rows of the document through the SQL Database API:

  ```bash
  curl --request POST \
    --url https://api.azion.com/v4/workspace/sql/databases/<database-id>/query \
    --header 'Accept: application/json' \
    --header 'Authorization: Token [TOKEN VALUE]' \
    --header 'Content-Type: application/json' \
    --data '{"statements":["SELECT COUNT(*) FROM passages WHERE source = '\''returns-policy.md'\'' AND embedding IS NOT NULL;"]}'
  ```

  The single entry in `data` carries `results`, whose one row holds the same number the ingestion returned.

- **An answer comes from the documents and cites them.** Ask a question the document answers:

  ```bash
  curl -X POST https://www.example.com/api/ask \
    -H 'Content-Type: application/json' \
    -d '{"question":"How many days do I have to return an item?"}'
  ```

  The response carries `answer`, with at least one `[n]` citation, and `sources`, whose entries name `returns-policy.md`.

- **A question outside the documents is not answered from training.** Ask a question no document covers, such as `What is the capital of France?`. The answer says that the assistant does not know, and it carries no citation.

When a request answers `404` or the default page of the application, the rule may still be propagating. When it persists after a few minutes, read the function's log lines in [Real-Time Events](/en/documentation/platform/real-time-events/quickstart/), under the **Functions Console** data source.

---

## Measuring results

| Metric                                       | Where to read it                                                                                                                                                                                                                                        | What working looks like                                                                                                          |
| -------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------- |
| Answer latency                               | The **Request Time** of requests to `/api/ask`, in the **HTTP Requests** data source of [Real-Time Events](/en/documentation/platform/real-time-events/data-sources/#http-requests)                                                                     | Stable as the number of documents grows, because each answer reads four passages whatever the size of the table                  |
| Share of answers that cite retrieved sources | The `"event":"answer"` lines the assistant function writes, in the **Functions Console** data source of [Real-Time Events](/en/documentation/platform/real-time-events/data-sources/#functions-console): the lines with `"cited":true` over all of them | Close to every answer. A drop points at a question the documents do not cover, or at passages that no longer match the documents |
| Answer accuracy on a test set                | A fixed list of questions with known answers from your documents, sent to `/api/ask` after every change to the documents, the prompt, or the models                                                                                                     | The share of correct answers holds or rises from one run to the next                                                             |

---

## Best practices

- **Embed questions and passages with the same model and width.** Vector distance compares two vectors only when one model produced both at one width. Changing the model or `dimensions` means a new column declaration and a full re-ingestion, so keep both values in one constant shared by the two functions.
- **Re-ingest a document when it changes.** The `DELETE` before the `INSERT` statements replaces a document's passages, so the table never answers from a superseded version. A retired document needs its own `DELETE FROM passages WHERE source = '<key>';`.
- **Check every statement entry as well as the status.** The SQL Database API answers `200` even when a statement fails, with `error` in that statement's entry. The ingestion function reads every entry, and any script that writes passages should do the same. For the error shapes, refer to [SQL Database best practices](/en/documentation/platform/sql-database/best-practices/).
- **Read the model's text with optional chaining.** The function reads `modelResponse?.choices?.[0]?.message?.content`, so a response without one of those levels yields an empty answer instead of an exception. For the reasoning, refer to [AI Inference best practices](/en/documentation/platform/ai-inference/best-practices/).

---

## Guides in this use case

- [Embed documents into a vector table with AI Inference](/en/documentation/guides/ai/agents-and-rag/embed-documents-into-a-vector-table.md): Reads, splits, and embeds each document, the pattern the ingestion function follows.
- [Write rows to SQL Database from a function](/en/documentation/guides/application-development/data/write-sql-database-rows-from-a-function.md): Writes the passages through the Azion API and checks every statement entry.
