# GLM-5.3

## Best for

Long-running coding agents that need a million-token context

## When not to use

Image, audio, or video. Not GLM-5.3 Flash — that is the cheaper, natively visual sibling; this is the larger text-only model.

## Overview

GLM-5.3 is Z.AI's 753-billion-parameter mixture-of-experts flagship, with a sparse-attention design that keeps long prompts affordable. It always thinks, and reasoning effort can be set to low, high, or max.

Its 1,048,576-token context window suits coding agents that carry a large repository and long tool traces. Z.AI reports stronger agentic coding results than GLM-5.2 at a lower token cost per task.

| | |
| --- | --- |
| API identifier | `glm-5.3` |
| Developer | Z.ai |
| Context window | 1,048,576 tokens |
| Supports | text, reasoning, tool-use |
| Status | Serving |

## Rates

USD per 1M tokens, on-demand tier.

| Input | Cached input | Output |
| --- | --- | --- |
| $1.12 | $0.14 | $3.52 |

## Call it

Put `glm-5.3` in the `model` field. The same request works against the Anthropic Messages API — see [API Compatibility](https://inference.boundless.network/docs/api-compatibility).

```bash
curl https://api.inference.boundless.network/v1/chat/completions \
  -H "Authorization: Bearer $BOUNDLESS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "glm-5.3",
    "messages": [{
      "content": "Say hello",
      "role": "user"
    }]
  }'
```

Every model we serve is listed at [Models](https://inference.boundless.network/models); the rate card in full is on [Pricing](https://inference.boundless.network/docs/pricing).
