# GPT-OSS 20B

**GPT-OSS 20B** is an open-source large language model with 20 billion parameters, built for text generation, conversation, and other natural language processing tasks. It reads text, returns text, and accepts tool definitions in the request. [AI Inference](/en/documentation/platform/ai-inference/) runs it under the id `gpt-oss-20b`.

## Model details

The model id is the string that selects this model, and it is the first argument `Azion.AI.run` takes. The HuggingFace repository holds the model card.

| Detail                     | Value                                                                   |
| -------------------------- | ----------------------------------------------------------------------- |
| Model name                 | GPT-OSS 20B                                                             |
| Version                    | 20B                                                                     |
| Model category             | Large Language Model (LLM)                                              |
| Model id                   | `gpt-oss-20b`                                                           |
| Size                       | 20B parameters                                                          |
| HuggingFace model          | [openai/gpt-oss-20b](https://huggingface.co/openai/gpt-oss-20b)         |
| OpenAI-compatible endpoint | [OpenAI Chat API](https://developers.openai.com/api/reference/overview) |
| License                    | [Apache 2.0](https://choosealicense.com/licenses/apache-2.0/)           |

## Capabilities

These values belong to the model and are not part of the request body. For the fields every chat request accepts, with their types, defaults, and bounds, refer to [Model invocation](/en/documentation/platform/ai-inference/model-invocation/).

| Capability     | Value       |
| -------------- | ----------- |
| Input data     | Text        |
| Context length | 131k tokens |
| Tool calling   | Yes         |
| Supports LoRA  | No          |

---

## Usage

A [function](/en/documentation/platform/functions/) invokes the model with `Azion.AI.run`, passing the id as the first argument and an OpenAI-compatible request body as the second. The examples below use that binding. To send the same body to the OpenAI-compatible HTTP endpoint instead, refer to [Model invocation](/en/documentation/platform/ai-inference/model-invocation/).

### Chat completion

This request carries a system message, a user message, and the sampling fields that shape the output:

```ts
const modelResponse = await Azion.AI.run("gpt-oss-20b", {
  "stream": false,
  "max_tokens": 1024,
  "temperature": 0.7,
  "top_p": 0.9,
  "messages": [
    { "role": "system", "content": "You are a helpful assistant." },
    { "role": "user", "content": "Name the European capitals." }
  ]
})
```

The model answers with one entry in `choices`, and the generated text sits at `choices[0].message.content`:

```json
{
  "id": "chatcmpl-example",
  "object": "chat.completion",
  "created": 1767268800,
  "model": "gpt-oss-20b",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "reasoning_content": null,
        "content": "Sure! Here is a list of some European capitals...",
        "tool_calls": []
      },
      "logprobs": null,
      "finish_reason": "stop",
      "stop_reason": null
    }
  ],
  "usage": {
    "prompt_tokens": 9,
    "total_tokens": 527,
    "completion_tokens": 518,
    "prompt_tokens_details": null
  },
  "prompt_logprobs": null
}
```

### Tool calling

A tool-calling request adds a `tools` array to the body. Each entry sets `type` to `function` and carries a `function` object holding the name the model calls, a description of what it does, and its `parameters` in JSON Schema form:

```ts
const modelResponse = await Azion.AI.run("gpt-oss-20b", {
  "stream": false,
  "max_tokens": 1024,
  "messages": [
    { "role": "system", "content": "You are a helpful assistant with access to tools." },
    { "role": "user", "content": "What is the weather in London?" }
  ],
  "tools": [
    {
      "type": "function",
      "function": {
        "name": "get_weather",
        "description": "Get the current weather for a location",
        "parameters": {
          "type": "object",
          "properties": {
            "location": {
              "type": "string",
              "description": "The city and state"
            }
          },
          "required": ["location"]
        }
      }
    }
  ]
})
```

When the model selects a tool, it returns `content` as `null`, names the call in `tool_calls` with the arguments it chose, and sets `finish_reason` to `tool_calls`:

```json
{
  "id": "chatcmpl-tool-example",
  "object": "chat.completion",
  "created": 1746821866,
  "model": "gpt-oss-20b",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "reasoning_content": null,
        "content": null,
        "tool_calls": [
          {
            "id": "chatcmpl-tool-0123456789abcdef0123456789abcdef",
            "type": "function",
            "function": {
              "name": "get_weather",
              "arguments": "{\"location\": \"London\"}"
            }
          }
        ]
      },
      "logprobs": null,
      "finish_reason": "tool_calls",
      "stop_reason": null
    }
  ],
  "usage": {
    "prompt_tokens": 293,
    "total_tokens": 313,
    "completion_tokens": 20,
    "prompt_tokens_details": null
  },
  "prompt_logprobs": null
}
```

---

## Related resources

- [Model invocation](/en/documentation/platform/ai-inference/model-invocation.md): Every field a request body accepts, and the HTTP endpoint that takes the same body.
- [AI models](/en/documentation/platform/ai-inference/models.md): The other models AI Inference runs, and the id each one answers to.
- [AI Inference](/en/documentation/platform/ai-inference.md): The product that runs this model.
- [Functions](/en/documentation/platform/functions.md): The product whose runtime holds the binding the examples on this page call.
- [AI Inference limits](/en/documentation/platform/ai-inference/limits.md): The conditions under which Azion terminates or deprovisions a model.
- [Glossary](/en/documentation/platform/ai-inference/glossary.md): Where model id, context length, and the rest of the AI Inference vocabulary are defined.
