Are you an LLM? Read llms.txt for a summary of the docs, or llms-full.txt for the full context.
Skip to content
介面 / OpenAI-compatible
POSThttps://api.lazu.ai/v1/responses

Responses API

Responses 適合 reasoning 模型、多模態輸入和 Lazu file_id 解引用。上傳檔案或使用較新的 OpenAI response 特性時,優先使用這個 endpoint。

計費依所選模型與線路的價格計算 token 用量;可能包含其他計費維度。 價格與線路 ↗

何時使用

使用 Responses

上傳檔案、reasoning 控制、文件 workflow,以及需要服務端標準化的多模態輸入。

使用 Chat completions

簡單聊天、SDK 相容、tool calling,以及已經圍繞 /v1/chat/completions 建構的應用。

請求 Body

modelstringrequired

來自 /api/models/catalog 的模型 ID。優先選擇 supported_endpoints 包含 /v1/responses 路径 的模型。

inputstring | object[]required

文字輸入,或 response input messages 陣列。

instructionsstringnullable

response 的系統級 instructions。

reasoningobjectnullable

支援 reasoning 的模型可使用 effort 控制,例如 {"effort":"medium"}。

toolsobject[]nullable

OpenAI-compatible Responses shape 的工具定義。

streambooleannullable

當模型支援時,返回 response events stream。

max_output_tokensintegernullable

生成輸出 token 的上限。

File 解引用

先透過 Files 上傳檔案,再引用返回的

file_id。如果請求中發生了解引用,Lazu 會加入 X-Lazu-File-Dereference: 1。

限制:

  • 單檔案仍受 purpose 對應大小限制。
  • 單次 Responses 呼叫中解引用的檔案總大小必須低於 64 MB。
  • Chat completions 不會自動解引用 file_id。

無狀態相容橋接

部分 catalog 項目會透過無損 Chat 橋接提供 Responses;可查看 supported_endpoints[].mode。走這類路由時,省略 store 可以正常請求,並回傳 X-Lazu-Warning: stateless_bridge;明確指定 store: false 正常請求且不產生 warning。明確指定 store: true、無法保真的狀態欄位或未知跨協議欄位,會在呼叫上游前回傳 400 protocol_bridge_unsupported。標準 Responses body 不會增加 Lazu 私有欄位。

響應

idstring

Response ID。

outputobject[]

輸出訊息、reasoning items、tool calls 或其它 response events。

usage.input_tokensintegernullable

上游返回 usage 時的輸入 token 數。

usage.output_tokensintegernullable

上游返回 usage 時的輸出 token 數。

usage.input_tokens_details.cached_tokensintegernullable

支援 response-level cache usage 的 provider 返回的 cache read tokens。

完整對帳請使用同一把 API Key 呼叫:

GET /api/usage/requests/{request_id}

相關頁面

完整工具續輪範例

先從目前 Key 的目錄選擇支援 Responses 和函式工具的模型,再設定 LAZU_API_KEY、LAZU_MODEL。無狀態續輪要保留全部 output item(包括 reasoning),再用原 call_id 附加工具結果。橋接無法保留的加密或供應商私有 reasoning 狀態仍只能走原生路徑。

import json
import os
from openai import OpenAI, APIStatusError

client = OpenAI(
    api_key=os.environ["LAZU_API_KEY"],
    base_url=os.environ.get("LAZU_BASE_URL", "https://api.lazu.ai/v1"),
    max_retries=0,
)
model = os.environ["LAZU_MODEL"]
tools = [{"type": "function", "name": "add", "description": "Add two integers",
          "parameters": {"type": "object", "properties": {
              "a": {"type": "integer"}, "b": {"type": "integer"}},
              "required": ["a", "b"], "additionalProperties": False}}]
history = [{"role": "user", "content": "Use add to calculate 2 + 3."}]
try:
    first = client.responses.create(model=model, input=history, tools=tools,
        tool_choice={"type": "function", "name": "add"}, store=False,
        max_output_tokens=1024)
    history.extend(item.model_dump(exclude_unset=True) for item in first.output)
    for item in first.output:
        if item.type == "function_call":
            if item.name != "add":
                raise ValueError("Unexpected tool")
            args = json.loads(item.arguments)
            history.append({"type": "function_call_output", "call_id": item.call_id,
                            "output": json.dumps({"result": args["a"] + args["b"]})})
    final = client.responses.create(model=model, input=history, tools=tools,
        tool_choice="none", store=False, max_output_tokens=1024)
    print(final.output_text)
except APIStatusError as exc:
    print(exc.status_code, exc.response.headers.get("x-lazu-request-id"), exc.body)
    raise

收到輸出後不要自動重播。413 request_body_too_large 應縮小請求;503 gateway_overloaded 可稍後退避重試,它不表示供應商故障。部分用量透過請求詳情對帳。

Sent from your browser to api.lazu.ai · key stays in this tab
RequestPOST /v1/responses
curl https://api.lazu.ai/v1/responses \
  -H "Authorization: Bearer $LAZU_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-6-luna",
    "input": [{
      "role": "user",
      "content": [
        {"type": "input_text", "text": "Say hi"}
      ]
    }]
  }'
ResponseExample
{
  "id": "resp_01ABCDEF",
  "model": "gpt-6-luna",
  "output": [
    {
      "type": "message",
      "role": "assistant",
      "content": [{ "type": "output_text", "text": "Hi!" }]
    }
  ],
  "usage": { "input_tokens": 9, "output_tokens": 2 }
}
請求回執範例回執
$0.00028200 · 842 ms
Request ID
req_demo_01
模型 · 線路
gpt-6-luna · stable
Token · 輸入 / 輸出
1,200 / 320
Key
demo-key

僅為範例,不是報價,也不是你這次請求的結果。送出上方請求即可看到你自己的回執。

查看真實請求回執 →

適合場景

  • Reasoning 模型
  • 已上傳的 PDF 或圖片
  • 結構化多模態輸入