LoRA Fine-Tune
Find which AI Inference models support Low-Rank Adaptation, and how LoRA Fine-Tune extends AI Inference and is metered.
Low-Rank Adaptation (LoRA) adapts a model to a task by training a small set of additional weights instead of retraining the whole model. LoRA Fine-Tune applies that method to supported models on Azion’s infrastructure, as an extension of AI Inference. LoRA applies to some models and not to others, and each model page states which case it is.
Supported models
The table lists every model AI Inference runs, with what the model’s own page states about LoRA. Yes and No repeat that statement. A dash means the page states no value, which is not the same as a No.
For the category, the id, the context length, and the input types of a model in this table, refer to AI models.
Relationship to AI Inference
LoRA Fine-Tune is an Azion product of its own, and the product it extends is AI Inference. AI Inference supports LoRA through an add-on, and the adaptation runs on Azion’s infrastructure, with no infrastructure for you to provision or manage. You are responsible for the training data, the model inputs and outputs, and the third-party license of the model you adapt.
Billing metrics
LoRA Fine-Tune is billed on two metrics, Compute Time and Requests, and it is metered separately from AI Inference. Compute Time is the duration of active execution in hours multiplied by the memory allocated in gigabytes, stated in GB-hour. For the rate charged for each metric, refer to Pricing, which carries a section of its own for LoRA Fine-Tune.