Are you an LLM? Read llms.txt for a summary of the docs, or llms-full.txt for the full context.
Skip to content
Endpoints / OpenAI-compatible
POSThttps://api.lazu.ai/v1/responses

Responses API

Use Responses for reasoning models, multimodal inputs and Lazu file_id dereferencing. This is the recommended endpoint when a request needs uploaded files or newer OpenAI response features.

BillingUsage tokens at the selected model and lane price; other billable dimensions may apply. Pricing & lanes ↗

When to use it

Use Responses

Uploaded files, reasoning controls, document workflows, and multimodal inputs that should be normalized server-side.

Use Chat completions

Simple chat, SDK compatibility, tool calling, and existing apps already built around /v1/chat/completions.

Request body

modelstringrequired

Model ID from /api/models/catalog. Prefer models where supported_endpoints includes the path /v1/responses.

inputstring | object[]required

Text input or an array of response input messages.

instructionsstringnullable

System-level instructions for the response.

reasoningobjectnullable

Reasoning effort controls for supported models, for example {"effort":"medium"}.

toolsobject[]nullable

Tool definitions using the OpenAI-compatible Responses shape.

streambooleannullable

Streams response events when supported by the selected model.

max_output_tokensintegernullable

Upper bound for generated output tokens.

Content parts

input_textcontent part

Text content sent to the model.

input_imagecontent part

Image content. Lazu can dereference uploaded file_id values with purpose vision.

input_filecontent part

Uploaded file reference. Lazu dereferences file content server-side before forwarding the request upstream.

File dereferencing

Upload via Files, then reference the resulting

file_id. Lazu adds X-Lazu-File-Dereference: 1 when the request dereferenced files.

Limits:

  • Single file purpose limit still applies.
  • Total dereferenced files in one Responses call must stay under 64 MB.
  • Chat completions does not auto-dereference file_id.

Stateless compatibility bridges

Some catalog entries expose Responses through a lossless Chat bridge; check supported_endpoints[].mode. On such a route, omitted store is accepted and returns X-Lazu-Warning: stateless_bridge, while store: false is accepted without a warning. Explicit store: true, stateful fields, and unknown cross-protocol fields return 400 protocol_bridge_unsupported before an upstream call. The standard Responses body is not extended with Lazu-only fields.

Response

idstring

Response ID.

outputobject[]

Output messages, reasoning items, tool calls, or other response events.

usage.input_tokensintegernullable

Input token count when the upstream reports usage.

usage.output_tokensintegernullable

Output token count when the upstream reports usage.

usage.input_tokens_details.cached_tokensintegernullable

Cache read tokens for providers that expose response-level cache usage.

For full reconciliation, use

GET /api/usage/requests/{request_id} with the same API key.

See also

Complete tool round trip

Select a model with Responses and function tools from your token-scoped catalog, then set LAZU_API_KEY and LAZU_MODEL. Keep every output item when continuing a stateless conversation, including reasoning items; append the tool result with the same call_id. Encrypted or provider-specific reasoning state remains native-only when a bridge cannot preserve it.

import json
import os
from openai import OpenAI, APIStatusError

client = OpenAI(
    api_key=os.environ["LAZU_API_KEY"],
    base_url=os.environ.get("LAZU_BASE_URL", "https://api.lazu.ai/v1"),
    max_retries=0,
)
model = os.environ["LAZU_MODEL"]
tools = [{"type": "function", "name": "add", "description": "Add two integers",
          "parameters": {"type": "object", "properties": {
              "a": {"type": "integer"}, "b": {"type": "integer"}},
              "required": ["a", "b"], "additionalProperties": False}}]
history = [{"role": "user", "content": "Use add to calculate 2 + 3."}]
try:
    first = client.responses.create(model=model, input=history, tools=tools,
        tool_choice={"type": "function", "name": "add"}, store=False,
        max_output_tokens=1024)
    history.extend(item.model_dump(exclude_unset=True) for item in first.output)
    for item in first.output:
        if item.type == "function_call":
            if item.name != "add":
                raise ValueError("Unexpected tool")
            args = json.loads(item.arguments)
            history.append({"type": "function_call_output", "call_id": item.call_id,
                            "output": json.dumps({"result": args["a"] + args["b"]})})
    final = client.responses.create(model=model, input=history, tools=tools,
        tool_choice="none", store=False, max_output_tokens=1024)
    print(final.output_text)
except APIStatusError as exc:
    print(exc.status_code, exc.response.headers.get("x-lazu-request-id"), exc.body)
    raise

Do not replay automatically after output has arrived. For 413 request_body_too_large, reduce the request; for 503 gateway_overloaded, retry later with backoff. Local overload does not imply a provider failure. Use request details to reconcile partial usage.

Sent from your browser to api.lazu.ai · key stays in this tab
RequestPOST /v1/responses
curl https://api.lazu.ai/v1/responses \
  -H "Authorization: Bearer $LAZU_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-6-luna",
    "input": [{
      "role": "user",
      "content": [
        {"type": "input_text", "text": "Say hi"}
      ]
    }]
  }'
ResponseExample
{
  "id": "resp_01ABCDEF",
  "model": "gpt-6-luna",
  "output": [
    {
      "type": "message",
      "role": "assistant",
      "content": [{ "type": "output_text", "text": "Hi!" }]
    }
  ],
  "usage": { "input_tokens": 9, "output_tokens": 2 }
}
Request receiptIllustrative example
$0.00028200 · 842 ms
Request ID
req_demo_01
Model · Lane
gpt-6-luna · stable
Tokens · input / output
1,200 / 320
Key
demo-key

Example only — not a quote and not the result of your request. Send the request above to see your own receipt.

Read a real request receipt →

Best for

  • Reasoning models
  • Uploaded PDFs or images
  • Structured multimodal input