ZeroTwo home

Gemini 3.1 Flash-Lite

Default
A low-cost multimodal model in the Gemini 3.1 family, aimed at high-frequency, lightweight tasks. Google positions it for high-volume agentic work, simple data extraction, translation, transcription, and latency-sensitive applications where budget and speed are the primary constraints. Supports thinking for enhanced accuracy.
A low-cost multimodal model in the Gemini 3.1 family, aimed at high-frequency, lightweight tasks. Google positions it for high-volume agentic work, simple data extraction, translation, transcription, and latency-sensitive applications where budget and speed are the primary constraints. Supports thinking for enhanced accuracy.

Gemini 3.1 Flash-Lite is an AI model from Google, available on ZeroTwo. It costs $0.25 per million input tokens and $1.50 per million output tokens, has a 1,048,576-token context window, can return up to 65,536 output tokens, and has a knowledge cutoff of January 2025.

Intelligence
Fair
Speed
Fast
Price
$0.25 • $1.50
Input • Output
Input
Text, Image, Audio
Output
Text

A low-cost multimodal model in the Gemini 3.1 family, aimed at high-frequency, lightweight tasks. Google positions it for high-volume agentic work, simple data extraction, translation, transcription, and latency-sensitive applications where budget and speed are the primary constraints. Supports thinking for enhanced accuracy.

1,048,576 context window
65,536 max output tokens
January 2025 knowledge cutoff
Reasoning token support
Pricing
Pricing is based on the number of tokens used. For tool-specific models, like search and computer use, there's a fee per tool call. See details in the pricing page.

The $0.25 input and $0.025 cached input rates cover text, image and video. Audio input is billed separately at $0.50 per million tokens, and cached audio input at $0.05 per million tokens.

Text tokens
Per 1M tokens
∙
Batch API price
Input
$0.25
Cached input
$0.025
Output
$1.50
Quick comparison
Input
Cached input
Output
Gemini 3.1 Flash-Lite
$0.25
Modalities
Text
Input and output
Image
Input only
Audio
Input only
Endpoints
Chat Completions
v1/chat/completions
Responses
v1/responses
Batch
v1/batch
Realtime
v1/realtime
Assistants
v1/assistants
Fine-tuning
v1/fine-tuning
Embeddings
v1/embeddings
Image Generation
v1/images/generations
Image Edit
v1/images/edits
Speech Generation
v1/audio/speech
Transcription
v1/audio/transcriptions
Translation
v1/translations
Moderation
v1/moderations
Completions (legacy)
v1/completions
Features
Streaming
Supported
Function calling
Supported
Structured outputs
Supported
Fine-tuning
Not supported
Distillation
Not supported
Fast response
Not supported
Cost efficient
Not supported
Tools
Tools supported by this model when using the Responses API.
Web search
Supported
Code interpreter
Supported
File search
Supported
Image generation
Not supported
MCP
Not supported
Snapshots
Snapshots let you lock in a specific version of the model so that performance and behavior remain consistent. Below is a list of all available snapshots and aliases for Gemini 3.1 Flash-Lite.
gemini-3.1-flash-lite-preview

Questions

How much does Gemini 3.1 Flash-Lite cost?

Gemini 3.1 Flash-Lite is priced at $0.25 per million input tokens and $1.50 per million output tokens, with cached input at $0.025 per million. On ZeroTwo it is included in your plan's credits rather than billed per token.

What is Gemini 3.1 Flash-Lite's context window?

Gemini 3.1 Flash-Lite accepts up to 1,048,576 tokens of context and can return up to 65,536 output tokens.

What is Gemini 3.1 Flash-Lite's knowledge cutoff date?

Gemini 3.1 Flash-Lite has a training knowledge cutoff of January 2025. For anything later than that, use a model with web search enabled.

Can I use Gemini 3.1 Flash-Lite for free?

You can try Gemini 3.1 Flash-Lite on ZeroTwo's free plan, which includes a monthly credit allowance across every model in the catalogue. Heavier use moves to a paid plan rather than a per-token bill.