Answer questions from a vector table
Answer a question from a function with the nearest passages of a SQL Database vector table and a chat model on AI Inference that cites them.
You answer a question from a function: it embeds the question with a model on AI Inference, retrieves the nearest passages of a SQL Database vector table with vector_top_k, and asks a chat model to answer from them and cite them. The passages come from documents in an Object Storage bucket, which a second function on the same application ingests. For the general ingestion pattern, refer to Embed documents into a vector table with AI Inference.
Prerequisites
- An application and a workload that serve your domain, with Application Accelerator turned on, which the Run Function behavior requires. To create them, refer to Applications quickstart.
- SQL Database enabled on the account. The product is in Preview and is not enabled by default, so request access through Technical Support.
- A
passagestable with apassages_idxvector index for 1,024-dimension vectors. To create them, refer to Create the vector table. - An Object Storage bucket with
workloads_accessset toread_only, holding the documents as UTF-8 text or Markdown files. To create the bucket and upload the files, refer to Object Storage quickstart. - A personal token with the Edit SQL Database permission, for the database calls. To create one, refer to Personal tokens.
- The Azion CLI, installed and authorized, to store the environment variables.
The examples use support-docs for the bucket, support-kb for the database, returns-policy.md for one document in it, /admin/ingest and /api/ask for the two paths, and www.example.com for the domain. Replace them with your values.
Ingest the documents from a bucket
The ingestion function turns one document into rows of passages: it reads the document from the bucket, splits it into passages of up to 1,500 characters, embeds all of them in one call, and writes them through the Azion API, because the runtime connection to a database is read-only. The route writes to the database, so it refuses a request without the ingestion secret.
To store the four values the function reads, run these commands with the Azion CLI. A key that contains token or secret is stored as a secret by default:
Create the variables before the function: a function that is already running does not read a changed value until it is redeployed.
Create a function named support-ingest with this code:
Run the function on its path with these values, following Functions quickstart:
- Function instance:
support-ingest, with no Args. - Rule: a Request Phase rule named
support - ingest, with the criterion${uri}starts with/admin/ingestand the Run Function behavior selecting thesupport-ingestinstance.
To create the rule in Azion Console or with the API, refer to Add the rule that runs the function.
A POST to /admin/ingest with the secret and a document key replaces that document’s passages in passages, and answers with the number it stored. The DELETE makes a second ingestion of the same key replace its rows instead of duplicating them.
To ingest the example document:
The function answers with the key and the number of passages it stored, such as {"key":"returns-policy.md","passages":3}. A request without the Authorization header answers 401. To count the rows of the document through the SQL Database API:
The single entry in data carries results, whose one row holds the same number the ingestion returned.
The Build and run customer support AI assistants use case uses the values of this example.
These checks confirm the Build and run customer support AI assistants use case.
Store the database name
The function opens the database by name, through a read replica.
To store the database name the function opens, run:
The account holds the SQL_DATABASE_NAME variable that the function reads with Azion.env.get().
The Build and run customer support AI assistants use case uses the values of this example.
Create the assistant function
The function embeds the question with the same model and width as the passages, retrieves the four nearest passages, and sends them with the last six messages of history and the question to the chat model. It writes one "event":"answer" log line per answer, recording whether the answer cites a passage.
Create a function named support-assistant with this code:
To create the function and its instance, follow Functions quickstart with the name support-assistant, and name the instance support-assistant, with no Args.
The application carries a support-assistant instance that answers a question from the passages nearest to it.
The Build and run customer support AI assistants use case uses the values of this example.
Run the function on the path
Create a Request Phase rule as Add the rule that runs the function shows, with the name support - ask, the criterion ${uri} is equal /api/ask, and the Run Function behavior selecting the support-assistant instance.
A POST to /api/ask with a question answers with the model’s answer, which cites passages as [n], and with the document each cited passage came from. A new rule takes a few minutes to propagate.
Confirm the answers cite the documents
To ask a question the document answers:
The response carries answer, with at least one [n] citation, and sources, whose entries name returns-policy.md.
Ask a question no document covers, such as What is the capital of France?. The answer says that the assistant does not know, and it carries no citation.
The function answers from the passages of the table, and only from them.
These checks confirm the Build and run customer support AI assistants use case.