Are you an LLM? Read llms.txt for a summary of the docs, or llms-full.txt for the full context.
Skip to content

Decisions

POST /v1/decisions evaluates explicit questions over text, JSON or embedded images. Discover available models and their decision_capabilities through the model catalog. This endpoint returns one JSON response and does not support streaming.

curl https://api.lazu.ai/v1/decisions \
  -H "Authorization: Bearer $LAZU_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "jev-1.13",
    "state": {"ticket": "The checkout service is down."},
    "questions": {
      "urgent": {"type": "noul", "instructions": "Does this need immediate attention?"},
      "owner": {"type": "choice", "instructions": "Choose the responsible team.", "criteria": {"engineering": "Service reliability", "billing": "Payment reconciliation"}},
      "severity": {"type": "score", "instructions": "Rate operational impact.", "criteria": ["No impact", "Degraded service", "Service unavailable"]}
    }
  }'

state is the material to judge. questions defines what to judge: noul returns a probability from 0 to 1; choice returns an option ID and probabilities; score returns a score against your ordered rubric. Business JSON inside instructions and criteria is preserved. Supply 1–64 questions with unique ASCII IDs using letters, numbers, _, . or -. Choice allows 2–255 options; score allows 2–10 levels. Answers include confidence and preserve additional provider fields.

Images

clef and clef-flash accept images containing PNG, JPEG or WebP data URLs, either as strings or { "url": "data:image/png;base64,..." }. State must still be supplied; it may be empty when images are valid. Remote URLs, uploaded file IDs, audio and video are unsupported. Jev accepts text and JSON.

Each image may contain at most 4 MiB decoded data and 16 million pixels. At most 4 images and 8 MiB total decoded data are allowed; the entire encoded upstream request may not exceed 13 MiB. Clef currently applies a conservative 2000 UTF-8 byte limit to state, in addition to token estimates, to prevent silent upstream truncation. Check max_state_bytes and token limits in the catalog.

Optional question generation

Question generation is disabled by default. When your operator enables it, opt in per request or in your API Key settings:

{
  "model": "jev-1.13",
  "state": "Checkout is down. Customers cannot pay.",
  "question_generation": {"enabled": true, "model": "gpt-6-luna"}
}

You may omit the generation model to use the Key's model or the platform default. Specifying a model alone does not enable generation. Explicit questions always take precedence. Automatically inferred questions can change with the material; use explicit questions when stable business rules are required.

Generation uses an independently authorized Responses call. Your Key and temporary credential must allow both models. The gateway checks pricing, Key, project, member and account budgets before generation, with a separate 2 requests/minute generation limit per Key and account. There is no questions cache. Generation may take up to 15 seconds; native evaluation has a total 10-second budget and at most two provider attempts; both stages share a 25-second budget.

Usage and failures

Native usage contains only the decision model's input and output tokens. question_generation includes its own model, request ID, prompt version and usage; questions contains the generated questions. Each stage has a separate usage record and charge. Upstream output tokens remain visible even when configured output pricing is zero.

Unusable generated questions are refunded to the original wallet and Key with an auditable refund record. If valid generation is followed by a native failure, the generation charge is retained and error.details.questions contains reusable questions. Retry those explicitly to avoid another generation charge. Request details and Console logs link the parent and generation requests.

When the decision API is disabled it returns 404 and decision models are hidden from the catalog. Use error.code and Retry-After for machine handling; see errors.

Operator configuration

Select a decision dialect on each source: TypeSafe, OpenRouter, Cloudflare AI Gateway, or Cloudflare Workers AI. Cloudflare sources also require the account ID. Register Jev, Clef and Clef Flash as decision models, configure input pricing and set output pricing to zero where applicable. Set channel and shared account RPS/in-flight caps; a shared group uses the smallest positive configured member limit. Multiple API replicas require Redis for shared admission. With Redis unavailable, the gateway retains local in-flight protection and emits an operational error; shared multi-replica capacity is temporarily unavailable.

The generator model must declare structured-output support in the model catalog, plus vision support for image requests. Native and generation permissions are checked before the paid generation call.

Enable /v1/decisions and automatic generation separately in upstream settings. Both switches default to off. Use the source health test with endpoint type decisions before exposing a model.