AI Models

Explore our collection of AI models

Compare pricing and context

AI models available on ZeroTwo, with input and output pricing per million tokens and context window size.
Claude Fable 5.1Anthropic$10.00$50.001,000,000
Claude Haiku 4.5Anthropic$1.00$5.00200,000
Claude Opus 4.6Anthropic$5.00$25.001,000,000
Claude Opus 4.7Anthropic$5.00$25.001,000,000
Claude Opus 4.8Anthropic$5.00$25.001,000,000
Claude Opus 5Anthropic$5.00$25.001,000,000
Claude Sonnet 4.5Anthropic$3.00$15.00200,000
Claude Sonnet 4.6Anthropic$3.00$15.001,000,000
Claude Sonnet 5Anthropic$2.00$10.001,000,000
Command ACohere$2.50$10.00256,000
Command A ReasoningCohere——256,000
Command R7BCohere$0.0375$0.15128,000
DeepSeek V4 FlashDeepSeek$0.22$0.661,000,000
DeepSeek V4 ProDeepSeek$0.66$1.981,000,000
DeepSeek V4.1 FlashDeepSeek$0.15$0.601,000,000
Flux ProFAL———
Flux Pro 2FAL———
Gemini 2.5 FlashGoogle$0.30$2.501,000,000
Gemini 2.5 Flash LiteGoogle$0.10$0.401,000,000
Gemini 2.5 ProGoogle$1.25$10.001,000,000
Gemini 3 FlashGoogle$0.50$3.001,000,000
Gemini 3.1 Flash-LiteGoogle$0.25$1.501,048,576
Gemini 3.1 ProGoogle$2.00$12.001,000,000
Gemini 3.5 FlashGoogle$1.50$9.001,000,000
Gemini 3.5 Flash-LiteGoogle$0.30$2.501,000,000
Gemini 3.6 FlashGoogle$0.75$3.751,000,000
Gemini 3.7 FlashGoogle$0.75$3.751,000,000
Gemini 3.8 FlashGoogle$0.75$3.751,000,000
GLM 4.7 Flash HereticZ.ai$0.07$0.40200,000
GLM 5Z.ai$1.00$3.20200,000
GLM 5.2Z.ai$1.40$4.401,000,000
GLM 5.3Z.ai$1.40$4.401,000,000
GPT 4.1OpenAI$2.00$8.001,000,000
GPT 4.1 MiniOpenAI$0.40$1.601,000,000
GPT 4.1 NanoOpenAI$0.10$0.401,000,000
GPT 4oOpenAI$2.50$10.00128,000
GPT 4o MiniOpenAI$0.15$0.60128,000
GPT 5OpenAI$1.25$10.00400,000
GPT 5 MiniOpenAI$0.25$2.00400,000
GPT 5 NanoOpenAI$0.05$0.40400,000
GPT 5.1OpenAI$1.25$10.00400,000
GPT 5.2OpenAI$1.75$14.00400,000
GPT 5.3 CodexOpenAI$1.75$14.00400,000
GPT 5.4OpenAI$2.50$15.001,050,000
GPT 5.4 MiniOpenAI$0.75$4.50400,000
GPT 5.4 NanoOpenAI$0.20$1.25400,000
GPT 5.5OpenAI$5.00$30.001,050,000
GPT 5.6 SolOpenAI$4.00$20.001,050,000
GPT 6 AstraOpenAI$10.00$50.001,050,000
GPT 6.1 SolOpenAI$2.00$10.001,050,000
GPT Image 1OpenAI$5.00$40.00N/A
GPT Image 1 MiniOpenAI$2.00$8.00N/A
GPT Image 1.5OpenAI$5.00$32.00N/A
Grok 4xAI——256,000
Grok 4 FastxAI——2,000,000
Grok 4 Fast ReasoningxAI——2,000,000
Grok 4.1 FastxAI——2,000,000
Grok 4.1 Fast ReasoningxAI——2,000,000
Grok 4.3xAI$1.25$2.501,000,000
Grok 4.5xAI$2.00$6.00500,000
Grok 4.6xAI$2.00$6.00500,000
Grok 4.7xAI$2.00$6.00500,000
Grok Build 0.1xAI$1.00$2.00256,000
Grok Code FastxAI——256,000
Grok Imagine Image 2.0xAI—$0.04/imageN/A
Kimi K2.5Kimi$0.60$3.00256,000
Kimi K2.6Kimi$0.95$4.00262,144
Kimi K2.7 CodeKimi$0.95$4.00262,144
Kimi K3Kimi$3.00$15.001,048,576
Mistral 3.1 24BVenice——128,000
Nano BananaGoogle$0.30$30.00N/A
Nano Banana ProGoogle$2.00$12.00N/A
o3OpenAI$2.00$8.00200,000
o4 MiniOpenAI$1.10$4.40200,000
QwenQwen——N/A
Qwen 3.5 PlusQwen$0.40$2.401,000,000
Qwen 3.6 FlashQwen$0.25$1.501,000,000
Qwen 3.6 PlusQwen$0.50$3.001,000,000
Qwen 3.7 MaxQwen$2.50$7.501,000,000
Qwen 3.7 PlusQwen$0.40$1.601,000,000
Qwen 3.8 FlashQwen$0.15$0.471,000,000
Qwen 3.8 MaxQwen$2.00$6.001,000,000
Qwen Coder PlusQwen$1.00$5.001,000,000
Qwen EditQwen——N/A
Qwen Plus CharacterQwen$0.50$1.4032,768
Qwen3 Next 80B InstructQwen$0.15$1.20131,072
Qwen3 Next 80B ThinkingQwen$0.15$1.20131,072
SonarPerplexity$1.00$1.00128,000
Sonar ProPerplexity$3.00$15.00200,000
OpenAI models
Latest models from OpenAI including GPT-5 and reasoning models.
GPT 6.1 Sol
OpenAI reasoning model for complex coding, computer use, and professional work. Tool calling uses the Responses API.
GPT 4.1
GPT-4.1 excels at instruction following and tool calling, with broad knowledge across domains.
GPT 4.1 Mini
GPT-4.1 mini excels at instruction following and tool calling. It features a 1M token context window, and low latency without a reasoning step.
GPT 4.1 Nano
GPT-4.1 nano excels at instruction following and tool calling. It features a 1M token context window, and low latency without a reasoning step.
GPT 4o
GPT-4o ("o" for "omni") is OpenAI's versatile multimodal model, accepting text and image input and returning text, with a 128,000-token context window.
GPT 4o Mini
GPT-4o mini is openai's fast, affordable small model for focused tasks. It features a 128k token context window, and low latency without a reasoning step.
GPT 5
GPT-5 is OpenAI's multimodal model with a 400,000-token context window, accepting text and image input, priced at $1.25 per million input tokens and $10.00 per million output tokens.
GPT 5 Mini
GPT-5 Mini is a cost-efficient variant of GPT-5 optimized for high-throughput tasks with fast responses and strong general capabilities.
GPT 5 Nano
GPT-5 Nano is OpenAI's lightest GPT-5 model for ultra-low-latency responses and production applications at minimal cost.
GPT 5.1
GPT-5.1 builds on GPT-5 with improved coding accuracy and instruction-following. It accepts text and image input and returns text.
GPT 5.2
GPT-5.2 is an OpenAI model for coding, math, and writing tasks, with a 400,000-token context window and text and image input.
o4 Mini
o4-mini is a small o-series reasoning model, optimized for fast, effective reasoning with efficient performance in coding and visual tasks.
o3
o3 is a reasoning model that works across domains, aimed at math, science, coding, and visual reasoning tasks. It also handles technical writing and instruction-following.
Anthropic models
Claude models for advanced reasoning and conversation.
Claude Sonnet 4.6
Anthropic's Claude Sonnet 4.6, now a legacy model superseded by Claude Sonnet 5. Sonnet 4.6 targets agents, coding, and computer use with adaptive thinking, a 1M token context window, and 128K token synchronous output.
Claude Sonnet 4.5
Anthropic's Claude Sonnet 4.5, now a legacy model superseded by Claude Sonnet 5. Sonnet 4.5 targets agents, coding, and computer use, with extended thinking, a 200K token context window, and 64K token output.
Claude Haiku 4.5
Anthropic's Claude Haiku 4.5 is a small, low-latency model with a 200K token context window and 64K token output. It delivers coding performance similar to Claude Sonnet 4 at one-third the cost and more than twice the speed, and scores 73.3% on SWE-bench Verified. Built for latency-sensitive experiences, free user experiences, and high-volume operations.
claude-3-5-sonnet-20241022
claude-3-5-haiku-20241022
c1-anthropic-claude-sonnet-4-v-20250815
Google models
Gemini models with multimodal capabilities.
gemini-3-pro-preview
Gemini 3.1 Pro
Refined Gemini 3 Pro with better thinking, improved token efficiency, and more grounded, factually consistent output. Optimized for software engineering and agentic workflows.
Gemini 3 Flash
A Gemini 3 series Flash model for multimodal understanding, agentic workflows, and coding, built on the Gemini 3 reasoning foundation. Accepts text, image, audio, video, and PDF input and returns text. Uses dynamic thinking by default, with a configurable thinking_level. Supports 1M token context window.
Gemini 2.5 Pro
A Gemini 2.5 thinking model that reasons over complex problems in code, math, and STEM. Suited to analyzing large datasets, codebases, and documents using long context. Accepts audio, image, video, text, and PDF input and returns text, with a 1M token context window.
Gemini 2.5 Flash
A Gemini 2.5 model balancing cost and capability for large scale processing, low-latency, high volume tasks that require thinking, and agentic use cases. Accepts text, image, video, and audio input and returns text, with a 1M token context window and adjustable thinking budgets.
Gemini 2.5 Flash Lite
Gemini 2.5 Flash Lite is a faster, cost-efficient version of Gemini 2.5 Flash. It's optimized for quick responses and cost-effective applications. Supports multimodal inputs including text, images, video, and audio.
Gemini 3.1 Flash-Lite
A low-cost multimodal model in the Gemini 3.1 family, aimed at high-frequency, lightweight tasks. Google positions it for high-volume agentic work, simple data extraction, translation, transcription, and latency-sensitive applications where budget and speed are the primary constraints. Supports thinking for enhanced accuracy.
Gemini 2.5 Flash
A Gemini 2.5 model balancing cost and capability for large scale processing, low-latency, high volume tasks that require thinking, and agentic use cases. Accepts text, image, video, and audio input and returns text, with a 1M token context window and adjustable thinking budgets.
gemini-2-0-flash
gemini-1-5-pro
X AI models
Grok models with real-time information access.
Grok 4.1 Fast
Grok 4.1 Fast in non-reasoning mode for instant responses, with a 2M token context window and tool calling for agentic workflows such as customer support and finance. xAI retired this slug on May 15, 2026; requests now route to Grok 4.3 in non-reasoning mode.
Grok 4.1 Fast Reasoning
Grok 4.1 Fast with reasoning enabled — an xAI tool-calling model for agentic tasks, with a 2M token context window. xAI retired this slug on May 15, 2026; requests now route to Grok 4.3 with low reasoning effort.
Grok 4 Fast Reasoning
Grok 4 Fast with reasoning enabled, using a unified architecture that blends reasoning and non-reasoning modes, with web and X search and a 2M token context window. xAI retired this slug on May 15, 2026; requests now route to Grok 4.3 with low reasoning effort.
Grok 4 Fast
Grok 4 Fast in non-reasoning mode for instant responses, with web and X search and a 2M token context window. xAI retired this slug on May 15, 2026; requests now route to Grok 4.3 in non-reasoning mode.
Grok 4
Grok 4 is xAI's July 2025 reasoning model, trained with scaled reinforcement learning on the Colossus cluster, with native tool use, real-time search, and text and image input. xAI retired the grok-4-0709 slug on May 15, 2026; requests now route to Grok 4.3 with low reasoning effort.
Grok Code Fast
An xAI coding model for agentic workflows, built on a new architecture with prompt caching and adept at TypeScript, Python, Java, Rust, C++, and Go. xAI retired this slug on May 15, 2026; requests now route to Grok Build 0.1.
Qwen models
Multilingual models from Alibaba with strong reasoning and coding.
qwen3-max
qwen-plus
Qwen Plus Character
The role-playing model of the Qwen series. This is a dynamically updated version, and notifications will be provided in advance for any model updates. It is suitable for anthropomorphic role-playing and has optimized capabilities in following character personalities and maintaining consistent role-play interactions.
qwen-flash
Qwen Coder Plus
Powered by Qwen3, this is a powerful Coding Agent that excels in tool calling and environment interaction to achieve autonomous programming. It combines outstanding coding proficiency with versatile general-purpose abilities.
Qwen3 Next 80B Instruct
A new generation of open-source, non-thinking mode model powered by Qwen3. This version demonstrates superior Chinese text understanding, augmented logical reasoning, and enhanced capabilities in text generation tasks over the previous iteration (Qwen3-235B-A22B-Instruct-2507).
Qwen3 Next 80B Thinking
A new generation of Qwen3-based open-source thinking mode models. This version offers improved instruction following and streamlined summary responses over the previous iteration (Qwen3-235B-A22B-Thinking-2507).
Qwen 3.5 Plus
Hosted Plus-tier model in the Qwen 3.5 series (February 2026) on Alibaba Cloud Model Studio. Native vision-language model combining linear attention with sparse MoE (Qwen3.5-397B-A17B class), with optional chain-of-thought via enable_thinking. Strong on coding, agents, tool use, and multimodal reasoning across text, images, and video, with a 1M-token context window by default.
Qwen 3.6 Plus
Alibaba’s April 2026 Plus-tier model on Model Studio: major upgrade in agentic coding (frontend through repo-level work), sharper multimodal perception and reasoning, and a 1M-token default context. Supports enable_thinking for hybrid reasoning and preserve_thinking for multi-turn agents so prior reasoning traces stay in context. Strong on coding-agent, tool-use, STEM, and long-context benchmarks.
qwen-max-latest
qwen-turbo-latest
qwen3-235b-a22b-instruct
qwen3-32b-instruct
Kimi models
Frontier mixture-of-experts models from Moonshot AI with tool use.
Groq models
High-speed inference models optimized for performance.
Thesys models
C1 models that generate interactive UI components and visualizations.
Community & Open models
Open-source and community-fine-tuned models with unique capabilities.