Embed documents into a vector table with AI Inference
Turn the documents in a bucket into passages with vectors from an AI Inference embedding model, stored in a SQL Database vector table.
You turn a document stored in Object Storage into passages, embed each passage with a model on AI Inference, and store the passages and their vectors in a SQL Database table, from a function and the Azion API. To query the table by meaning once it holds vectors, refer to Vector search.
A retrieval step finds passages by comparing vectors, so every passage needs one, produced by the same model and at the same width as the vector of each later question. The function reads one document per request, embeds all of its passages in one model call, and writes them in one call to the Azion API.
- A request names the object key of one document, and the function reads that document from the bucket.
- The function splits the text into passages.
- One call to the embedding model returns a vector for every passage.
- The function builds a
DELETEfor the document’s earlier rows and oneINSERTper passage, and sends them to the query endpoint of the Azion API. - The rows land in the vector table, and the vector index covers them.
Prerequisites
- SQL Database enabled on your account. The product is in Preview and is not enabled by default, so request access through Technical Support.
- A database, its identifier, and a personal token, stored as the
SQL_DATABASE_IDandSQL_TOKENenvironment variables, as Write rows to SQL Database from a function describes. A function writes a database only through the Azion API, because its own connection is read-only. - A bucket that holds the documents as UTF-8 text or Markdown files. To create one and upload the files, refer to Object Storage quickstart.
- The Azion CLI installed and authorized, and a function project to add the code to. To create and deploy one, refer to Deploy a function with Azion CLI.
The examples read the bucket my-bucket, embed with Qwen/Qwen3-Embedding-4B at 1,024 dimensions, write to a table named passages, and answer on www.example.com. Replace them, <database-id>, <personal-token>, and <object-key> with your own values.
Create the vector table
The vector column declares the width of the vectors it holds, and the embedding request asks for the same width. Qwen/Qwen3-Embedding-4B returns one of five widths, 256, 512, 1024, 2048, or 4096, through the dimensions field, so a F32_BLOB(1024) column matches a request for 1024. A vector index needs a table with a ROWID or a single-column primary key, and its second argument sets the distance metric.
To create the table and its index, send both statements to the query endpoint:
The API answers 200 with "state": "executed" and one entry per statement. A statement that fails still answers 200, with error in place of results in its entry, so read both entries:
The database holds an empty passages table and the passages_idx index. The index adds a shadow table, passages_idx_shadow, that a table listing shows beside passages.
Embed a document into the table
The function reads the document with the Storage class of the azion:storage module, which takes the bucket name and no token. It sends every passage in one input array, and the model answers with one entry per input in data, where data[].index is the position of the passage and data[].embedding its vector. Each vector goes into its statement as text inside vector('[...]').
The route writes to the database, so it refuses a request that does not carry a secret. To store that secret with the Azion CLI:
The command prints the UUID of the variable it created:
To embed one document per request, use this code as the function’s entrypoint:
The code applies four decisions:
- One document per request. One read, one model call, and one write bound the work of each invocation. A function may use 2 seconds of CPU time and 50 outbound
fetch()calls per invocation. - At most 99 passages per document. The
DELETEand 99INSERTstatements make 100 statements, and a call carrying 100 statements succeeds. A longer document answers413, so split it into two files. - The width is one constant.
DIMENSIONSsets the request, and it must equal theF32_BLOBwidth of the column. A question is comparable with a passage only when the same model produced both vectors at the same width. - A missing key answers
404.getthrowsStorageError: Object not foundfor a key that holds no object, andBucket not foundfor a bucket the account does not have.
Under azion dev, Azion.AI is undefined and azion:storage reads the local disk, so deploy the function, as Deploy a function with Azion CLI shows, and test it there. Then send one document:
The function answers with the key and the number of passages it stored, such as {"key":"<object-key>","passages":3}. A request without the Authorization header answers 401. The table holds one row per passage of the document, each with its vector.
Confirm the passages carry vectors
To count the rows of the document that hold a vector, send a SELECT to the query endpoint:
The single entry in data carries results, whose one row holds the number the function returned:
Every passage of the document carries a vector, and vector_top_k on passages_idx can return it. A query vector must come from Qwen/Qwen3-Embedding-4B at 1,024 dimensions too.