# System One

Call Laya at POST /v1/systemone, and wire it into harnesses and agents as a fast gate or router.


A System One model reads a state and answers typed questions about it in one forward pass. It returns probabilities, never text. [Laya](/models/convaiinnovations/laya) is the System One model we serve, designed for triage, routing and guardrails.

System One models have their own endpoint, `POST /v1/systemone`, on the same base URL and with the same API key as chat. They do not answer chat completions, Responses or Messages requests, so a harness cannot run Laya as its model. You call Laya next to the model the harness runs: from a hook before a tool call, or from your own agent code before it picks a model.

## Send a request

```bash
curl https://api.inference.boundless.network/v1/systemone \
  -H "Authorization: Bearer $BOUNDLESS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "convaiinnovations/laya",
    "state": "I was charged twice for my October invoice. Please refund the duplicate and cancel my plan.",
    "questions": {
      "refund": { "type": "noul", "instructions": "Does the customer ask for a refund?" },
      "topic": {
        "type": "choice",
        "instructions": "What is the message about?",
        "criteria": {
          "billing": "Charges, invoices or refunds",
          "technical": "Errors, outages or bugs",
          "account": "Sign-in, profile or team settings"
        }
      },
      "urgency": { "type": "score", "instructions": "How urgent is this message?", "criteria": ["low", "medium", "high"] }
    }
  }'
```

Each question comes back under its own name in `answers`:

```json
{
  "model": "convaiinnovations/laya",
  "answers": {
    "refund": { "type": "noul", "noul": 0.8932 },
    "topic": {
      "type": "choice",
      "choice": "billing",
      "confidence": 0.7971,
      "probabilities": { "billing": 0.9528, "technical": 0.0214, "account": 0.0259 }
    },
    "urgency": {
      "type": "score",
      "score": 1.1619,
      "confidence": 0.0644,
      "legend": { "0": "low", "1": "medium", "2": "high" },
      "probabilities": { "0": 0.1796, "1": 0.479, "2": 0.3414 }
    }
  },
  "usage": { "input_tokens": 161, "output_tokens": 0 }
}
```

`model` is required. Use `convaiinnovations/laya` or its alias `laya`.

## Questions

| `type` | Asks | `criteria` | Answer |
| --- | --- | --- | --- |
| `noul` | Whether a statement about the state is true | Optional: `{ "true": "...", "false": "..." }` descriptions | `noul`, the probability that it is true |
| `choice` | Which named option fits | Required: an object of option names to descriptions | `choice`, `confidence`, and `probabilities` per option |
| `score` | Where the state sits on an ordered scale | Required: an array of levels, lowest first | `score`, `confidence`, `legend`, and `probabilities` per level index |

A `score` is the probability-weighted mean of the level indexes, so `1.16` on a three-level scale sits just above `medium`.

`state` can be a string, an object or an array. Send structured data, such as a tool call or a ticket, as JSON rather than flattening it into a sentence. Ask as many questions about one state as you need in a single request.

Decide on probabilities with a threshold you have tested, not on `choice` alone. `choice` is the likeliest option even when no option is likely, and a low `confidence` means the options are close. In one test, a `choice` between allow, ask and deny picked allow for `git push --force` at 0.46, while a `noul` asking whether the command could destroy data returned 0.90. Describing each option in `criteria` separates them better than a bare name.

## Context budget

Laya sends each request to its English or its multilingual checkpoint. The default budget is 512 tokens on English and 1,024 on multilingual. Set `max_len` in the request, up to 8,192, for longer input.

The budget covers the state, the question, its options and special tokens, for each question. Each `criteria` entry has its own limit of 48 tokens. Input over either limit is rejected with a 422 and never truncated.

`usage.input_tokens` counts the encoded input once per question. `output_tokens` is always 0. Laya is currently free to call. Its [model page](/models/convaiinnovations/laya) shows the live rate.

## Errors

Errors are JSON with `message`, `error_type` and `code`.

| Status | `code` | Meaning |
| --- | --- | --- |
| 400 | `invalid_json` | The body is not valid JSON. |
| 400 | `unsupported_model` | `model` is missing or is not a System One model. |
| 401 | `401` | The API key is missing or not valid. |
| 403 | `403` | The key is not allowed to call this model. |
| 422 | `invalid_request` | A field is malformed. `message` names it, as in `questions: expected at least one named question`. |
| 422 | `422` | The input is over its context budget or a criterion is too long. |
| 429 | `capacity_exceeded`, `rate_limit_exceeded` or a team limit | Retry after the `Retry-After` header. [Rate Limits](/docs/rate-limits) covers team limits. |
| 502 | `502` | The request did not complete. Retry it. |

Every response carries an `X-Request-Id`. Quote it to [support](/support). `Server-Timing` reports the request's validation, queue and inference time in milliseconds.

## Gate tool calls in Claude Code

A Claude Code `PreToolUse` hook runs before each tool call. This one sends the call to Laya and, when Laya rates it likely to delete data, discard work or rewrite shared history, asks you to approve it. It needs `curl`, `jq`, and `BOUNDLESS_API_KEY` in Claude Code's environment, which [one-command setup](/docs/coding-agents-and-harnesses#set-up-with-one-command) exports from your shell profile.

Save it as `.claude/hooks/laya-gate.sh` in your project and make it executable:

```bash
#!/usr/bin/env bash
input=$(cat)
answer=$(jq -c '{
  model: "convaiinnovations/laya",
  state: {tool_name, tool_input},
  questions: {destructive: {type: "noul", instructions: "Could this tool call delete data, discard work or rewrite shared history?"}}
}' <<<"$input" | curl -sf --max-time 5 https://api.inference.boundless.network/v1/systemone \
  -H "Authorization: Bearer $BOUNDLESS_API_KEY" -H "Content-Type: application/json" -d @-) || exit 0
jq -c 'select(.answers.destructive.noul >= 0.5) | {hookSpecificOutput: {
  hookEventName: "PreToolUse",
  permissionDecision: "ask",
  permissionDecisionReason: "Laya rates this \(.answers.destructive.noul * 100 | round)% likely to be destructive"
}}' <<<"$answer"
```

Register it in `.claude/settings.json`:

```json
{
  "hooks": {
    "PreToolUse": [
      {
        "matcher": "Bash",
        "hooks": [{ "type": "command", "command": "${CLAUDE_PROJECT_DIR}/.claude/hooks/laya-gate.sh", "timeout": 10 }]
      }
    ]
  }
}
```

Below the threshold the hook prints nothing, and Claude Code's normal permission rules decide. If Laya does not answer within five seconds, the hook also stays silent, so a failed call never approves anything. In `claude -p`, where nobody can answer the prompt, an `ask` blocks the call and the agent sees the reason.

In our tests, `git status`, `npm test` and `cat package.json` scored below 0.12, and `rm -rf build/`, `git reset --hard`, `git push --force` and `DROP TABLE` scored above 0.81. Check the threshold against the commands your agents run.

Any harness that can run a command or plugin before a tool call can make the same request. Send the tool name and its arguments as `state`.

## Route requests in your own agent

Laya can choose a model before your agent spends tokens. This sends lookups and rewrites to a fast model and code or reasoning to a stronger one, and falls back to the stronger model if Laya returns an error:

```ts
import OpenAI from "openai";

const apiKey = process.env.BOUNDLESS_API_KEY;
const client = new OpenAI({ apiKey, baseURL: "https://api.inference.boundless.network/v1" });

async function pickModel(prompt: string) {
  const response = await fetch("https://api.inference.boundless.network/v1/systemone", {
    body: JSON.stringify({
      model: "convaiinnovations/laya",
      questions: {
        tier: {
          criteria: {
            fast: "A lookup, translation, summary or reformatting",
            strong: "Code, debugging, proofs or multi-step reasoning",
          },
          instructions: "Which kind of model should answer this request?",
          type: "choice",
        },
      },
      state: prompt,
    }),
    headers: { Authorization: `Bearer ${apiKey}`, "Content-Type": "application/json" },
    method: "POST",
  });
  if (!response.ok) return "glm-5.3";
  const { answers } = await response.json();
  return answers.tier.probabilities.strong >= 0.5 ? "glm-5.3" : "glm-5.3-flash";
}

const prompt = "Translate thank you into Spanish";
const completion = await client.chat.completions.create({
  messages: [{ content: prompt, role: "user" }],
  model: await pickModel(prompt),
});
```

In our tests, `strong` scored 0.17 to 0.25 on a capital-city lookup, a translation and an email summary, and 0.80 to 0.85 on a refactor and a race-condition hunt. A proof that the square root of 2 is irrational scored 0.55, close to the line. Pick model IDs from [Models](/models).
