---
title: "Pricing"
description: "The models we serve and the live rate card"
canonical: "https://inference.boundless.network/docs/pricing"
last-updated: "2018-10-20T01:46:40.000Z"
---

# Pricing

The models we serve and what they cost.


> We serve a small set of open-weight models rather than a long catalog. If one you need is missing, [tell us](/support).

## Rate card

Per 1M tokens, on-demand tier. Cached input is charged at the cached rate when a request reuses a prefix the gateway already holds.

US dollars per 1M tokens.

| Model | Lab | Input | Cached input | Output |
| --- | --- | --- | --- | --- |
| `claude-opus-5` | Anthropic | $5.00 | $0.50 | $25.00 |
| `claude-fable-5.1` | Anthropic | $10.00 | $0.25 | $50.00 |
| `dsv4` | DeepSeek | $0.04 | $0.02 | $0.21 |
| `deepseek-v4-pro-0813` | DeepSeek | $1.04 | $0.0352 | $2.08 |
| `minimax-m3` | MiniMax | $0.28 | $0.056 | $1.10 |
| `minimax-m2.7` | MiniMax | $0.30 | $0.06 | $1.20 |
| `kimi-k2.5` | Moonshot AI | $0.45 | $0.07 | $2.25 |
| `kimi-k2.6` | Moonshot AI | $0.45 | $0.13 | $2.71 |
| `kimi-k3` | Moonshot AI | $2.30 | $0.23 | $11.40 |
| `nemotron3-super` | NVIDIA | $0.12 | $0.06 | $0.52 |
| `gpt-5.6-luna` | OpenAI | $0.20 | $0.02 | $1.20 |
| `gpt-5.6-sol` | OpenAI | $2.00 | $0.20 | $10.00 |
| `qwen3.6` | Qwen | $0.04 | $0.02 | $0.35 |
| `qwen3.8-flash` | Qwen | $0.15 | $0.016 | $0.47 |
| `qwen3-coder-next` | Qwen | $0.12 | $0.07 | $0.80 |
| `hy3` | Tencent | $0.132 | $0.033 | $0.528 |
| `hy4-preview` | Tencent | $0.834 | $0.042 | $2.501 |
| `mimo-v2.5` | Xiaomi | $0.119 | $0.00255 | $0.238 |
| `mimo-v2.5-pro` | Xiaomi | $0.35 | $0.0029 | $0.70 |
| `glm-5.3-flash` | Z.ai | $0.12 | $0.024 | $0.40 |
| `glm-5.2` | Z.ai | $0.34 | $0.17 | $2.30 |
| `glm-5.1` | Z.ai | $0.84 | $0.164 | $2.80 |
