Google · Released 2026-02-19

Gemini 3.1 Pro Preview

Available
gemini-3.1-pro-preview

Call gemini-3.1-pro-preview on Lazu for $2.00 per million input tokens and $12.00 per million output tokens, the same as the Google official price. 1M context, up to 66k output tokens, with Tool calling, Vision, Reasoning and Structured output.

Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more efficient token usage across complex workflows. Building on the multimodal foundation...

Get an API key
Copy page
Context1M
Max output66k
Knowledge cutoff2025-01
InputText · Image · Audio · Video · PDF
OutputText
Endpoints/v1beta/models/gemini-3.1-pro-preview:generateContent · /v1/chat/completions

Prices and lanes

USD per million tokens. Availability covers the last 7 days, 2 hours per bar. A request whose input (cached tokens included) passes the threshold is billed entirely at the long-context prices.

Lane InputOutputCache readCache write 5m 7-day availability
Google official price $2.00 $12.00$0.20— —
Input over 200k $4.00 $18.00$0.40—
Stable Same as official $2.00 $12.00$0.20— 97.0%

Prices checked 2026-10-03 16:41 UTC · How to choose a lane · gemini-3.1-pro-preview.json · gemini-3.1-pro-preview.md

Capabilities

Tool callingVisionReasoningStructured outputPDF inputAudio input
Reasoning effort
low · medium · high

Example request

POST /v1beta/models/gemini-3.1-pro-preview:generateContent
curl -X POST https://api.lazu.ai/v1beta/models/gemini-3.1-pro-preview:generateContent \
  -H "Authorization: Bearer $LAZU_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "contents": [{ "parts": [{ "text": "Hello from Lazu" }] }]
  }'

Any OpenAI SDK works too: set base_url to https://api.lazu.ai/v1. Full integration docs

Questions

What is the difference between the stable and discount lanes?

The stable lane runs on each maker's official API, is billed at the list price, has the highest availability and suits every use. The discount lane is a reverse-engineered route: a third-party supplier serves the model through its official client, not the official API. It costs far less, but availability varies, a failed request does not fall back to the stable lane, and the supplier may add its own system prompt. So use it only inside the matching agent tool: Claude models in Claude Code, GPT models in Codex, Gemini models in Antigravity. For your own code, or anywhere the output must be controlled exactly, use the stable lane. Both lanes work with the same key, and each key picks the lane per model.

Can I call gemini-3.1-pro-preview with the OpenAI SDK?

Yes. Point the SDK's base URL at https://api.lazu.ai/v1 and use gemini-3.1-pro-preview as the model name. The native Gemini endpoint /v1beta/models/gemini-3.1-pro-preview:generateContent works as well.

How is it billed?

Per token actually used — input, output and cache — deducted from your wallet balance. There is no monthly fee, and every request shows up itemized in your usage log.

Same family

Get a key and call it now

Start free Read the docs