---
name: levanto-sage-decide
description: Call the Levanto Sage decision API (/decide): Yes/No, Choice, Scale, Sort and Tags decisions, with optional reasoning, web grounding and images (beta). Use when the user gives a Sage endpoint, asks to classify, route, score, rank or tag content, or wants /decide examples.
---

# Levanto Sage Decision Model

Sage turns content plus a question into a structured decision your code acts on: route, approve, block, score, rank, tag or escalate. Typical uses: agent workflows and guardrails, ops triage, moderation, risk and fraud.

When you write about Sage:

- Say "decision kind" in prose; keep JSON wire fields exactly as shown.
- When Sage isn't sure, say "the answer is `null`", never "abstains".
- Don't describe internals, architecture or training unless the user asks.

## Endpoint and key

Base URL: `https://sage.levanto.ai`. Use the user's URL if they give one; never invent one.

| Method | Path | Use |
|---|---|---|
| GET | `/ready` | `200` ready, `503` loading. No key. |
| POST | `/decide` | One decision |
| POST | `/decide/batch` | Many decisions |
| POST | `/usage/estimate` | Token counts without running |

`/decide` and `/decide/batch` need a Levanto API key, sent as `Authorization: Bearer lv_live_...`; without one they return `401`. If the user has no key, ask for one. Read it from an environment variable (e.g. `SAGE_API_KEY`); never hardcode, log, echo or commit it.

## Request

```json
{
  "content": "<what to judge>",
  "question": {"id": "<your result key>", "kind": "yesno | choice | scale | sort | tags", "instructions": "<the question>"}
}
```

`content` is a string, `{"kind": "text", "value": "..."}`, a list for Sort, or an image ([Images](#images-beta)). Optional top-level fields: `reasoning`, `grounding`. Every response is `{ id, kind, result, meta }`.

## Acting on answers

Yes/No, Choice and Tags return `null` when Sage isn't sure. Act on decisions; send `null` to a human:

```python
result = response["result"]
if result["answer"] is None:        # Choice: result["chosen"]; Tags: tag["applies"]
    escalate_to_human(response)
else:
    act_on(result["answer"])
```

Don't build review bands on `probability`; use it only when a workflow needs a stricter bar than Sage's answer. Scale and Sort return `confidence` instead; route on it as the workflow needs.

## Yes/No

```json
{
  "content": "Email draft promises \"guaranteed 40% returns\" and calls the product \"risk-free\".",
  "question": {"id": "needs_review", "kind": "yesno", "instructions": "Does this copy need compliance review before send?"}
}
```

```json
{
  "id": "needs_review",
  "kind": "yesno",
  "result": {
    "answer": "yes",
    "probability": 0.92
  },
  "meta": { "model": "levanto-sage-v1.3", "latency_ms": 97.4 }
}
```

- `answer`: `"yes"`, `"no"`, or `null` on a near-tie (e.g. `probability` 0.51).
- `probability`: calibrated P(yes), always returned.

## Choice

```json
{
  "content": "Vendor contract with automatic renewal and a 90-day termination clause.",
  "question": {
    "id": "contract", "kind": "choice", "instructions": "How should legal handle this contract?",
    "options": [
      {"option": "approve", "description": "Standard terms, no red flags"},
      {"option": "revise", "description": "Acceptable with clause changes"},
      {"option": "escalate", "description": "Unusual or high-risk terms"}
    ]
  }
}
```

```json
{
  "id": "contract",
  "kind": "choice",
  "result": {
    "chosen": "revise",
    "probability": 0.88,
    "probabilities": [
      {"option": "approve", "probability": 0.08},
      {"option": "revise", "probability": 0.88},
      {"option": "escalate", "probability": 0.04}
    ]
  },
  "meta": { "model": "levanto-sage-v1.3", "latency_ms": 216.8 }
}
```

- `chosen`: the winning option, or `null` when the top two are too close to call.
- `probability`: calibrated P(`chosen` is correct); `null` when `chosen` is.
- `probabilities`: every option in request order; they sum to 1.
- 2–120 options (20 with an image).
- Strong on rules-heavy decisions (several rules, exceptions, precedence), such as compliance and trust & safety. Put the rules in `content` or `instructions` and describe each option.

## Scale

Send 2–26 `levels`, distinct whole numbers of your choice (e.g. `0`–`4`, `1`–`5`); anything else returns `400`.

```json
{
  "content": "Candidate explained the roadmap clearly but struggled with database scaling tradeoffs.",
  "question": {
    "id": "interview", "kind": "scale", "instructions": "How strong was this technical interview?",
    "levels": [
      {"level": 0, "description": "Not qualified"},
      {"level": 1, "description": "Weak"},
      {"level": 2, "description": "Mixed"},
      {"level": 3, "description": "Solid"},
      {"level": 4, "description": "Strong hire"}
    ]
  }
}
```

```json
{
  "id": "interview",
  "kind": "scale",
  "result": {
    "expectation": 2.89,
    "confidence": 0.66,
    "probabilities": [
      {"level": 0, "probability": 0.0},
      {"level": 1, "probability": 0.01},
      {"level": 2, "probability": 0.12},
      {"level": 3, "probability": 0.84},
      {"level": 4, "probability": 0.03}
    ]
  }
}
```

`expectation` is the expected level, in your levels' units; `probabilities` gives each level's probability (they sum to 1); `confidence` is how sure Sage is of the score, independent of the score.

## Sort

`content` is a list of 2–120 items. The question takes only `kind`, `id` and `instructions`; other fields are rejected, and the server picks the ranking method.

```json
{
  "content": {"kind": "list", "value": [
    {"id": "typo", "content": "Spelling mistake on the FAQ page."},
    {"id": "db_down", "content": "Production database is unreachable."},
    {"id": "pricing", "content": "Prospect asking about Enterprise pricing."}
  ]},
  "question": {"id": "triage", "kind": "sort", "instructions": "Most urgent first."}
}
```

```json
{ "id": "triage", "kind": "sort", "result": { "sorted": ["db_down", "pricing", "typo"], "confidence": 0.92 } }
```

`confidence` (0–1, or `null`) is how trustworthy the whole order is.

## Tags

Decide which of 1–120 labels apply. Each tag returns `probability` and a verdict, `applies`: `true`, `false`, or `null` when too close to call. `null` means Sage isn't sure whether that tag applies: don't read it as `false`; send the item to a person or skip the action that depends on that tag.

- Write `instructions` as the rule for when a tag applies. Sage reads it with the whole tag list, so similar tags are judged against each other.
- Give each tag a `name`: a short label plus one line, e.g. `"promotion: advertises the poster's own product"`. The `id` is only the key you get back; without a `name`, Sage reads the `id`.
- When two tags are easy to confuse, say what doesn't count ("a single genuine offer is not spam").
- `{label}` or `{tag}` in instructions still work as a per-tag template; other braces are plain text.
- A tag's `threshold` is deprecated and ignored; apply a stricter bar to `probability` yourself.
- Tags answer "which apply", not "does anything apply". If none may apply, gate with Yes/No first. For exactly one of N, use Choice.

```json
{
  "content": "Forum post: \"I help indie artists grow their streams. First month half price, DM me.\"",
  "question": {
    "id": "moderate", "kind": "tags",
    "instructions": "Tag the issues. Members may promote their own work once.",
    "tags": [
      {"id": "spam", "name": "spam: bulk or repeated posting; a single genuine offer is not spam"},
      {"id": "promotion", "name": "promotion: advertises the poster's own product"},
      {"id": "scam", "name": "scam: deceptive offer, fake results or payment tricks"}
    ]
  }
}
```

```json
{
  "id": "moderate",
  "kind": "tags",
  "result": {
    "tags": [
      {"id": "spam", "probability": 0.03, "applies": false},
      {"id": "promotion", "probability": 1.0, "applies": true},
      {"id": "scam", "probability": 0.03, "applies": false}
    ]
  }
}
```

## Images (beta)

Yes/No, Choice, Scale and Tags accept an image as `content`, on `/decide` and per batch group.

```json
{
  "content": {"kind": "image", "media": "data:image/png;base64,iVBORw0KGgo...", "text": "Ticket: checkout button missing on mobile."},
  "question": {"id": "shows_bug", "kind": "yesno", "instructions": "Does the screenshot show the reported problem?"}
}
```

- `media`: base64 `data:` URI (PNG, JPEG or WebP), up to 4 MiB decoded, longest edge 8192 px, 4096² pixels. No remote URLs.
- `text` (optional): context to judge with the image.
- Returns `400` with Sort, inside a `list` item, with `grounding`, or with more than 20 Choice options.
- On images, Tags always decide (`applies` is never `null`).
- Image tokens are input tokens; `meta.usage` reports `image_count` and `image_tokens`.

## Reasoning

Top-level `reasoning` on `/decide` and `/decide/batch`:

- `off` (default): answer right away. Fastest and cheapest; right for most calls and tight latency, such as a gate in an agent loop.
- `auto`: reason only when Sage judges the question needs it (long policy text, multi-step rules).
- `on`: always reason, up to a few seconds. For rules-heavy decisions where accuracy beats speed.

With `auto` or `on`, set client timeouts above 10 seconds.

```json
{
  "content": "Refund request: $180 order, bought 41 days ago, box unopened. Policy: refunds within 30 days; store credit up to 60 days if unopened.",
  "question": {"id": "refund_ok", "kind": "yesno", "instructions": "Should we issue a cash refund?"},
  "reasoning": "auto"
}
```

`meta.reasoning` (omitted on kinds without reasoning):

```json
"reasoning": {"fired": true, "ran": true, "finished": true, "tokens": 312, "limited": null}
```

- `fired`: Sage judged reasoning needed. `ran`: it reasoned.
- `finished`: `false` if cut off, `null` if Sage didn't reason.
- `limited`: why it was cut off: `cap` (length), `timeout` (10 s) or `budget` (no time to start). Sage still answers.
- `forced` (only after `cap` or `timeout`): `true` answered from the partial reasoning, `false` gave the quick answer.
- `tokens`: thinking tokens, charged as output. Thinking also reads, charged as input.

## Grounding (optional)

Add `grounding` to a Yes/No, Choice, Scale or Tags request to search the web before deciding, for recent or factual questions. Not with images. Omit it if the server isn't set up for grounding.

```json
{
  "content": "Has there been a major Cloudflare incident this month?",
  "question": {"id": "incident", "kind": "yesno", "instructions": "Is this likely true?"},
  "grounding": {"trigger": "low_confidence", "confidence_floor": 0.8}
}
```

- `trigger`: `low_confidence` (default: search when Sage is less sure than `confidence_floor`), `always` or `never`.
- `confidence_floor` (default `0.80`): for Yes/No, Sage skips search only when `probability` ≤ 0.10 or ≥ 0.90.
- Tags search when any tag is unsure, then decide every tag again.
- Search runs before the `null` check, so it can turn `null` into an answer.
- The response adds `grounding_meta`: `triggered`, `queries`, `sources`, timing.
- Each search that runs is charged at your plan's search rate, plus tokens.

## Pricing

Each plan includes usage each month: a paid plan its price, Free a small amount, a custom package what was agreed. Calls are charged per input token, output token and search at your plan's rates:

| Plan | Price | Input / 1M tokens |
|------|------:|------------------:|
| Free | Free ($0.20 usage / month) | $0.05 |
| Developer | $14 / month | $0.05 |
| Starter | $49 / month | $0.046 |
| Pro | $99 / month | $0.042 |
| Growth | $249 / month | $0.038 |

All plans: output $10 / 1M tokens · search $0.10.

- **Input**: everything Sage reads, including instructions, options and images. Content counts once per question (ten questions on one document count it ten times), once per tag for Tags, once per item for Sort.
- **Output**: one token per answer, plus thinking when Sage reasons.
- **Search**: per grounding search that runs.
- `meta.usage` reports `input_tokens`, `output_tokens` and `image_tokens`; `POST /usage/estimate` predicts them (images excluded).
- Unused usage doesn't roll over. When it runs out, calls return `402` until the next period or an upgrade. No seats, no overage invoices.

## Batch

`requests` is a list of groups: one `content` and its `questions`. Content is sent once per group; add groups for other content. Batching saves round trips, not tokens. Top-level `reasoning` applies to every question; `grounding` goes on each question (not Sort).

```json
{
  "reasoning": "auto",
  "requests": [
    {"content": "Selling verified Instagram accounts, DM for bulk pricing.", "questions": [
      {"id": "violates", "kind": "yesno", "instructions": "Does this post violate marketplace policy?"},
      {"id": "action", "kind": "choice", "instructions": "What should moderation do?",
       "options": [{"option": "remove"}, {"option": "warn"}, {"option": "allow"}]}
    ]},
    {"content": "doc B", "questions": [{"id": "b1", "kind": "yesno", "instructions": "Urgent?"}]}
  ]
}
```

Response: `results[i]` matches `requests[i]`; `results[i].answers[j]` matches its `questions[j]` and holds `ok` plus `result` (a normal decision) or `error`. `meta.usage` counts tokens for the whole call.

```json
{
  "results": [
    {"answers": [
      {"ok": true, "result": {"id": "violates", "kind": "yesno", "result": {"answer": "yes", "probability": 0.94}}},
      {"ok": true, "result": {"id": "action", "kind": "choice", "result": {"chosen": "remove", "probability": 0.9, "probabilities": [...]}}}
    ]},
    {"answers": [{"ok": true, "result": {"id": "b1", "kind": "yesno", "result": {"answer": "no", "probability": 0.08}}}]}
  ],
  "meta": {"model": "levanto-sage-v1.3", "request_count": 2, "question_count": 3, "latency_ms": 150.0,
           "usage": {"input_tokens": 412, "output_tokens": 3}}
}
```

## Errors and limits

| HTTP | Meaning | Example body |
|---|---|---|
| `400` | Invalid request, a limit below, or an unsupported image combination | `{"detail": "content: Field required"}` |
| `401` | Missing or invalid API key | `{"detail": "API key required. Provide a valid Levanto API key."}` |
| `402` | This period's included usage is used up | |
| `503` | Loading or temporarily unavailable | `{"detail": "Service is still loading."}` |

| Kind | Field | Min | Max |
|---|---|---:|---:|
| choice | `options` | 2 | 120 (20 with an image) |
| scale | `levels` | 2 | 26 (distinct whole numbers) |
| sort | `content.value` items | 2 | 120 |
| tags | `tags` | 1 | 120 |

The whole request (content, instructions, options, levels, items, tags and any search results) must fit the plan's context: about 32K tokens, or 128K on Growth (extended context). Oversized requests fail, never truncate. Very long content can return `503` sooner, so stay well under the limit.

## Smoke Test

```bash
endpoint="https://sage.levanto.ai"
auth="Authorization: Bearer ${SAGE_API_KEY:?set SAGE_API_KEY to your lv_live_ key}"

# /ready needs no key: 200 ready, 503 loading.
curl -s -o /dev/null -w '%{http_code}\n' "$endpoint/ready" | grep -qx 200

curl -s -H "$auth" -H 'Content-Type: application/json' "$endpoint/decide" -d '{
  "content": "Email draft promises \"guaranteed 40% returns\" and calls the product \"risk-free\".",
  "question": {"id": "needs_review", "kind": "yesno", "instructions": "Does this copy need compliance review before send?"}
}' | jq -e '.kind == "yesno" and .result.probability != null and (.result | has("answer")) and (.result.answer == "yes" or .result.answer == "no" or .result.answer == null)'

curl -s -H "$auth" -H 'Content-Type: application/json' "$endpoint/decide/batch" -d '{
  "requests": [{"content": "same doc", "questions": [
    {"id": "a", "kind": "yesno", "instructions": "urgent?"},
    {"id": "b", "kind": "yesno", "instructions": "positive?"}
  ]}]
}' | jq -e '.meta.request_count == 1 and .meta.question_count == 2 and (.results | length) == 1 and (.results[0].answers | length) == 2'
```
