Add AI features to existing applications
Add a summarization API route to an application you already run, with a function that calls a catalog model on AI Inference and caches results.
A product team wants AI features such as summarization, classification, moderation, or image understanding in an application it already runs, without operating GPU infrastructure. The application already sits behind Azion, and the team wants one more API route that it can call from its code like any other. This page adds a summarization route to that application: a function validates the request, calls a model from the AI Inference catalog, caches the result for identical text, and returns the summary. The result is measured by inference latency per request, consumption per thousand requests, and task accuracy on a test set.
This use case does not cover conversational assistants grounded in company content, or training models from scratch. For assistants, refer to Build and run customer support AI assistants.
Prerequisites
- An application and a workload that already serve your application, with Application Accelerator turned on, which the Run Function behavior requires. To create them, refer to Applications quickstart.
- A personal token, for the API steps. To create one, refer to Personal tokens.
- The values of your own setup. This page uses
/api/summarizefor the route,ai-summarizefor the function and its instance, andwww.example.comfor the domain your application answers on. Replace each value with yours in every step.
Required products
| The feature needs | Which means | Product | Documented in |
|---|---|---|---|
| A model that summarizes, with no inference server to run | A catalog model called with Azion.AI.run | AI Inference | Call a model on AI Inference from a function |
| One API route that validates input and shapes output | A function that the application runs on /api/summarize | Functions | Functions quickstart |
| The same text summarized once, not on every request | Results stored and matched with the Cache API, keyed by a hash of the text | Cache | Cache a function’s response with the Cache API |
| A view of how often the route runs | The Invocations dashboard of the Functions tab | Real-Time Metrics | Build dashboards |
| A rule that runs the function on the route | The Run Function behavior, which requires Application Accelerator on the application | Application Accelerator | Run a function on an application |
Reference architecture
This page builds the Catalog-model inference API: the model is used as AI Inference publishes it, behind a function your application calls.
Read the diagram from the function outward. The arrows that reach the function come from a route of an application that already serves your traffic, so your application’s code reaches the model through one more path on a domain it already calls. The arrows that leave the function are one execution: the function holds the request open while the model generates, and no arrow leads to an inference server, because the function names a model id and Azion runs the model. The model is used as published, so the design has no training step and no model version to choose beyond the id.
Dataflow
- Your application’s code sends
POST /api/summarizewith the text to summarize, and the application’s rule runs theai-summarizefunction. Every other path keeps its own rules. - The function validates the body before any model call, and refuses a missing text with
400and a text over 20,000 characters with413. - The function hashes the text and looks the hash up in the cache. A match returns the stored summary with no model call.
- On a miss, the function sends the text to the catalog model with
Azion.AI.run, with an instruction to summarize it, and AI Inference runs the model and returns the summary. - The function stores the summary in the cache under the hash, for one day.
- The function returns the summary and whether it came from the cache.
Components
- application: the Platform Resource that routes
/api/summarizeto the function. The route is one rule on the application that already serves your traffic, so the domain, the TLS certificate, and any authentication in front of it are the ones the application already has. Azion issues no credential for a model call, so a route with no check of its own answers anyone who reaches it. - Functions: handles each request. The
ai-summarizefunction holds the prompt, the model id, the input bounds, and the shape of the response, so every decision about what the model receives is code your team owns. - AI Inference: runs the catalog model,
Qwen/Qwen3-30B-A3B-Instruct-2507-FP8. Your team loads no model and sizes no cluster, and the id from the model’s page is the only address the function needs. - Cache: stores results for repeated inputs, a design option. The function keys a result by a hash of its input through the Cache API, and uses it only for tasks whose output is the same for every identical input.
- Real-Time Metrics: reports the latency and volume of the route. A design with no inference server of its own still needs a view of what the route carried.
Other designs for this use case
- Fine-tuned inference API with LoRA adapters: for teams whose task needs a specialized model. The team trains a LoRA adapter on its own dataset with LoRA Fine-Tune, evaluates it, and serves the base model with the adapter on AI Inference behind the same kind of function API, which adds a training and evaluation control flow and a model-version decision before serving.
Configure the summarization function
The function is the whole feature: it decides what the model is asked, what it may receive, and what the caller gets back.
- The model is
Qwen/Qwen3-30B-A3B-Instruct-2507-FP8. Its page lists summarization among its uses, it serves a 64k-token context, and it is the model the AI Inference Starter Kit calls. The function calls it as Call a model on AI Inference from a function describes, withstreamset tofalseand the summarization instruction as thesystemmessage. - A text is at most 20,000 characters. The model reads every character it receives, and the function waits while it does, so a bound on the input is a bound on the wait. The response’s
usage.prompt_tokensshows what a text consumed, which is how to adjust the bound. max_tokensis300. A summary is short by definition, and the cap ends a generation that runs on.- A result is cached for one day under the SHA-256 of the text. The same text yields a summary the caller can reuse, so a repeated text costs no model call. The function caches as Cache a function’s response with the Cache API describes, with the cache
ai-summarize, the keyhttps://ai-summarize.cache/<sha-256 of the text>, andmax-age=86400. - The model’s text is read with optional chaining. A response missing a level yields an empty summary, which the function reports as
502, instead of an exception.
Create a function named ai-summarize with this code:
To create the function and its instance, follow Functions quickstart with the name ai-summarize, and name the instance on your application ai-summarize, with no Args. The Cache API is not defined under azion dev, so test the function once it runs on the application.
The application carries an ai-summarize instance that returns a summary for a text, from the cache when the same text was summarized in the last day.
Configure the API route
The route is a rule on the application that already serves your application. It matches /api/summarize exactly, so no other path of your application runs the function, and it runs in the Request Phase, because the function produces the whole response.
To create the rule:
Access Azion Console > Applications > your application.
Enter ai - summarize.
In the Criteria section, select the ${uri} variable, the is equal operator, and /api/summarize as the argument.
A POST to /api/summarize on your domain runs the function, and every other path keeps the rules it had. A new rule takes a few minutes to propagate.
Verify the setup
-
The route returns a summary. Send a text:
ShellThe response carries
x-summary-cache: missand a JSON body withsummary. -
The same text is answered from the cache. Send the same request again. The response carries
x-summary-cache: hitand the samesummary. -
Invalid input never reaches the model. Send
{"text":""}. The response is400withtext is required, which the function returns before the model call. -
The route runs only on its path. Request any other path of your application. It answers as it did before the rule existed.
When the route answers with your application’s own page instead of JSON, the rule may still be propagating. When it persists after a few minutes, read the function’s log lines under the Functions Console data source of Real-Time Events.
Measuring results
| Metric | Where to read it | What working looks like |
|---|---|---|
| Inference latency per request | The Request Time of requests to /api/summarize, in the HTTP Requests data source of Real-Time Events | Requests answered from the cache are faster than misses, and misses stay stable as volume grows |
| Request volume | Total Invocations on the Invocations dashboard of the Functions tab in Real-Time Metrics | Follows your application’s own traffic to the feature |
| Consumption per thousand requests | compute_time over invocations, times 1,000, from the consumption dataset in Query usage data from Functions. Functions and AI Inference are each charged at their own rates, in Pricing | Falls as the share of cached answers rises |
| Task accuracy on a test set | A fixed set of texts with reference summaries, sent to /api/summarize after every change to the prompt or the model | Holds or rises from one run to the next |
Best practices
- Validate in the function, before the model call. The function is the one place a request can be refused before a model runs, and every model call consumes Compute Time. For the reasoning, refer to AI Inference best practices.
- Cache only tasks whose output can be reused. A summary of identical text can stand for every request that sends it. A task whose answer must differ per user or per moment, such as one that reads the current time, must not go through the cache.
- Copy the model id from its page. Model ids do not share one form, so an id derived from the model name can be wrong for another model. To change the model, copy the id from AI models and change the cache name, so summaries from the old model are not returned for the new one.
- Add authentication when the route is public. The route answers anyone who reaches your domain, and each miss runs a model. When your application has a session or an API key, check it in the function before the cache lookup.