Qwen3 30B A3B Instruct 2507 FP8
Invoke Qwen3 30B A3B Instruct 2507 FP8, a 30B-parameter language model with a 64k-token context length, tool calling, and LoRA support.
Qwen3 30B A3B Instruct 2507 FP8 is an instruction-tuned large language model with 30 billion parameters, in FP8 precision. It reads text and returns text, for chat and question answering, summarization, multilingual tasks, mathematics and science problems, coding, and workflows that call tools. AI Inference runs it under the id Qwen/Qwen3-30B-A3B-Instruct-2507-FP8.
Model details
The id is the string a request passes to select this model, and Azion.AI.run takes it as the first argument. The HuggingFace repository publishes the model card.
| Detail | Value |
|---|---|
| Model name | Qwen3 30B A3B Instruct 2507 FP8 |
| Version | 30B - FP8 |
| Model category | Large Language Model (LLM) |
| Model id | Qwen/Qwen3-30B-A3B-Instruct-2507-FP8 |
| Size | 30B parameters |
| HuggingFace model | Qwen/Qwen3-30B-A3B-Instruct-2507-FP8 |
| OpenAI-compatible endpoint | OpenAI Chat API |
| License | Apache 2.0 |
Capabilities
The values below are properties of the model, and a request body does not set them. For the fields a request body does carry, with their types, defaults, and bounds, refer to Model invocation.
| Capability | Value |
|---|---|
| Input data | Text |
| Context length | 64k tokens |
| Tool calling | Yes |
| Supports LoRA | Yes |
The model’s native context window upstream is larger, and 64k tokens is the context length AI Inference serves.
Usage
The examples below call the model from a function with Azion.AI.run. The binding takes the model id as its first argument and the request body as its second, so the body never repeats the id. To send the same body to the OpenAI-compatible HTTP endpoint instead, refer to Model invocation.
Chat completion
This request carries a system message that sets the role the model answers in, and a user message holding the prompt:
The call resolves to one response object, and the generated text sits at modelResponse?.choices?.[0]?.message?.content, the content of the first entry in choices.
Tool calling
The same call becomes a tool-calling request when the body carries a tools array. Each entry sets type to function and holds a function object naming one function the model may call, what it does, and the parameters it takes in JSON Schema form:
When the model selects a tool, the assistant message carries a tool_calls array instead of text, and each entry names the function and the arguments the model generated for it.