AI Models
Explore our collection of AI models
Compare pricing and context
| Claude Fable 5.1 | Anthropic | $10.00 | $50.00 | 1,000,000 |
| Claude Haiku 4.5 | Anthropic | $1.00 | $5.00 | 200,000 |
| Claude Opus 4.6 | Anthropic | $5.00 | $25.00 | 1,000,000 |
| Claude Opus 4.7 | Anthropic | $5.00 | $25.00 | 1,000,000 |
| Claude Opus 4.8 | Anthropic | $5.00 | $25.00 | 1,000,000 |
| Claude Opus 5 | Anthropic | $5.00 | $25.00 | 1,000,000 |
| Claude Sonnet 4.5 | Anthropic | $3.00 | $15.00 | 200,000 |
| Claude Sonnet 4.6 | Anthropic | $3.00 | $15.00 | 1,000,000 |
| Claude Sonnet 5 | Anthropic | $2.00 | $10.00 | 1,000,000 |
| Command A | Cohere | $2.50 | $10.00 | 256,000 |
| Command A Reasoning | Cohere | — | — | 256,000 |
| Command R7B | Cohere | $0.0375 | $0.15 | 128,000 |
| DeepSeek V4 Flash | DeepSeek | $0.22 | $0.66 | 1,000,000 |
| DeepSeek V4 Pro | DeepSeek | $0.66 | $1.98 | 1,000,000 |
| DeepSeek V4.1 Flash | DeepSeek | $0.15 | $0.60 | 1,000,000 |
| Flux Pro | FAL | — | — | — |
| Flux Pro 2 | FAL | — | — | — |
| Gemini 2.5 Flash | $0.30 | $2.50 | 1,000,000 | |
| Gemini 2.5 Flash Lite | $0.10 | $0.40 | 1,000,000 | |
| Gemini 2.5 Pro | $1.25 | $10.00 | 1,000,000 | |
| Gemini 3 Flash | $0.50 | $3.00 | 1,000,000 | |
| Gemini 3.1 Flash-Lite | $0.25 | $1.50 | 1,048,576 | |
| Gemini 3.1 Pro | $2.00 | $12.00 | 1,000,000 | |
| Gemini 3.5 Flash | $1.50 | $9.00 | 1,000,000 | |
| Gemini 3.5 Flash-Lite | $0.30 | $2.50 | 1,000,000 | |
| Gemini 3.6 Flash | $0.75 | $3.75 | 1,000,000 | |
| Gemini 3.7 Flash | $0.75 | $3.75 | 1,000,000 | |
| Gemini 3.8 Flash | $0.75 | $3.75 | 1,000,000 | |
| GLM 4.7 Flash Heretic | Z.ai | $0.07 | $0.40 | 200,000 |
| GLM 5 | Z.ai | $1.00 | $3.20 | 200,000 |
| GLM 5.2 | Z.ai | $1.40 | $4.40 | 1,000,000 |
| GLM 5.3 | Z.ai | $1.40 | $4.40 | 1,000,000 |
| GPT 4.1 | OpenAI | $2.00 | $8.00 | 1,000,000 |
| GPT 4.1 Mini | OpenAI | $0.40 | $1.60 | 1,000,000 |
| GPT 4.1 Nano | OpenAI | $0.10 | $0.40 | 1,000,000 |
| GPT 4o | OpenAI | $2.50 | $10.00 | 128,000 |
| GPT 4o Mini | OpenAI | $0.15 | $0.60 | 128,000 |
| GPT 5 | OpenAI | $1.25 | $10.00 | 400,000 |
| GPT 5 Mini | OpenAI | $0.25 | $2.00 | 400,000 |
| GPT 5 Nano | OpenAI | $0.05 | $0.40 | 400,000 |
| GPT 5.1 | OpenAI | $1.25 | $10.00 | 400,000 |
| GPT 5.2 | OpenAI | $1.75 | $14.00 | 400,000 |
| GPT 5.3 Codex | OpenAI | $1.75 | $14.00 | 400,000 |
| GPT 5.4 | OpenAI | $2.50 | $15.00 | 1,050,000 |
| GPT 5.4 Mini | OpenAI | $0.75 | $4.50 | 400,000 |
| GPT 5.4 Nano | OpenAI | $0.20 | $1.25 | 400,000 |
| GPT 5.5 | OpenAI | $5.00 | $30.00 | 1,050,000 |
| GPT 5.6 Sol | OpenAI | $4.00 | $20.00 | 1,050,000 |
| GPT 6 Astra | OpenAI | $10.00 | $50.00 | 1,050,000 |
| GPT 6.1 Sol | OpenAI | $2.00 | $10.00 | 1,050,000 |
| GPT Image 1 | OpenAI | $5.00 | $40.00 | N/A |
| GPT Image 1 Mini | OpenAI | $2.00 | $8.00 | N/A |
| GPT Image 1.5 | OpenAI | $5.00 | $32.00 | N/A |
| Grok 4 | xAI | — | — | 256,000 |
| Grok 4 Fast | xAI | — | — | 2,000,000 |
| Grok 4 Fast Reasoning | xAI | — | — | 2,000,000 |
| Grok 4.1 Fast | xAI | — | — | 2,000,000 |
| Grok 4.1 Fast Reasoning | xAI | — | — | 2,000,000 |
| Grok 4.3 | xAI | $1.25 | $2.50 | 1,000,000 |
| Grok 4.5 | xAI | $2.00 | $6.00 | 500,000 |
| Grok 4.6 | xAI | $2.00 | $6.00 | 500,000 |
| Grok 4.7 | xAI | $2.00 | $6.00 | 500,000 |
| Grok Build 0.1 | xAI | $1.00 | $2.00 | 256,000 |
| Grok Code Fast | xAI | — | — | 256,000 |
| Grok Imagine Image 2.0 | xAI | — | $0.04/image | N/A |
| Kimi K2.5 | Kimi | $0.60 | $3.00 | 256,000 |
| Kimi K2.6 | Kimi | $0.95 | $4.00 | 262,144 |
| Kimi K2.7 Code | Kimi | $0.95 | $4.00 | 262,144 |
| Kimi K3 | Kimi | $3.00 | $15.00 | 1,048,576 |
| Mistral 3.1 24B | Venice | — | — | 128,000 |
| Nano Banana | $0.30 | $30.00 | N/A | |
| Nano Banana Pro | $2.00 | $12.00 | N/A | |
| o3 | OpenAI | $2.00 | $8.00 | 200,000 |
| o4 Mini | OpenAI | $1.10 | $4.40 | 200,000 |
| Qwen | Qwen | — | — | N/A |
| Qwen 3.5 Plus | Qwen | $0.40 | $2.40 | 1,000,000 |
| Qwen 3.6 Flash | Qwen | $0.25 | $1.50 | 1,000,000 |
| Qwen 3.6 Plus | Qwen | $0.50 | $3.00 | 1,000,000 |
| Qwen 3.7 Max | Qwen | $2.50 | $7.50 | 1,000,000 |
| Qwen 3.7 Plus | Qwen | $0.40 | $1.60 | 1,000,000 |
| Qwen 3.8 Flash | Qwen | $0.15 | $0.47 | 1,000,000 |
| Qwen 3.8 Max | Qwen | $2.00 | $6.00 | 1,000,000 |
| Qwen Coder Plus | Qwen | $1.00 | $5.00 | 1,000,000 |
| Qwen Edit | Qwen | — | — | N/A |
| Qwen Plus Character | Qwen | $0.50 | $1.40 | 32,768 |
| Qwen3 Next 80B Instruct | Qwen | $0.15 | $1.20 | 131,072 |
| Qwen3 Next 80B Thinking | Qwen | $0.15 | $1.20 | 131,072 |
| Sonar | Perplexity | $1.00 | $1.00 | 128,000 |
| Sonar Pro | Perplexity | $3.00 | $15.00 | 200,000 |
Featured models
Our most advanced and popular models.
OpenAI models
Latest models from OpenAI including GPT-5 and reasoning models.
GPT 6.1 Sol
OpenAI reasoning model for complex coding, computer use, and professional work. Tool calling uses the Responses API.
GPT 4.1
GPT-4.1 excels at instruction following and tool calling, with broad knowledge across domains.
GPT 4.1 Mini
GPT-4.1 mini excels at instruction following and tool calling. It features a 1M token context window, and low latency without a reasoning step.
GPT 4.1 Nano
GPT-4.1 nano excels at instruction following and tool calling. It features a 1M token context window, and low latency without a reasoning step.
GPT 4o
GPT-4o ("o" for "omni") is OpenAI's versatile multimodal model, accepting text and image input and returning text, with a 128,000-token context window.
GPT 4o Mini
GPT-4o mini is openai's fast, affordable small model for focused tasks. It features a 128k token context window, and low latency without a reasoning step.
GPT 5
GPT-5 is OpenAI's multimodal model with a 400,000-token context window, accepting text and image input, priced at $1.25 per million input tokens and $10.00 per million output tokens.
GPT 5 Mini
GPT-5 Mini is a cost-efficient variant of GPT-5 optimized for high-throughput tasks with fast responses and strong general capabilities.
GPT 5 Nano
GPT-5 Nano is OpenAI's lightest GPT-5 model for ultra-low-latency responses and production applications at minimal cost.
GPT 5.1
GPT-5.1 builds on GPT-5 with improved coding accuracy and instruction-following. It accepts text and image input and returns text.
GPT 5.2
GPT-5.2 is an OpenAI model for coding, math, and writing tasks, with a 400,000-token context window and text and image input.
o4 Mini
o4-mini is a small o-series reasoning model, optimized for fast, effective reasoning with efficient performance in coding and visual tasks.
o3
o3 is a reasoning model that works across domains, aimed at math, science, coding, and visual reasoning tasks. It also handles technical writing and instruction-following.
Anthropic models
Claude models for advanced reasoning and conversation.
Claude Sonnet 4.6
Anthropic's Claude Sonnet 4.6, now a legacy model superseded by Claude Sonnet 5. Sonnet 4.6 targets agents, coding, and computer use with adaptive thinking, a 1M token context window, and 128K token synchronous output.
Claude Sonnet 4.5
Anthropic's Claude Sonnet 4.5, now a legacy model superseded by Claude Sonnet 5. Sonnet 4.5 targets agents, coding, and computer use, with extended thinking, a 200K token context window, and 64K token output.
Claude Haiku 4.5
Anthropic's Claude Haiku 4.5 is a small, low-latency model with a 200K token context window and 64K token output. It delivers coding performance similar to Claude Sonnet 4 at one-third the cost and more than twice the speed, and scores 73.3% on SWE-bench Verified. Built for latency-sensitive experiences, free user experiences, and high-volume operations.
claude-3-5-sonnet-20241022
claude-3-5-haiku-20241022
c1-anthropic-claude-sonnet-4-v-20250815
Google models
Gemini models with multimodal capabilities.
gemini-3-pro-preview
Gemini 3.1 Pro
Refined Gemini 3 Pro with better thinking, improved token efficiency, and more grounded, factually consistent output. Optimized for software engineering and agentic workflows.
Gemini 3 Flash
A Gemini 3 series Flash model for multimodal understanding, agentic workflows, and coding, built on the Gemini 3 reasoning foundation. Accepts text, image, audio, video, and PDF input and returns text. Uses dynamic thinking by default, with a configurable thinking_level. Supports 1M token context window.
Gemini 2.5 Pro
A Gemini 2.5 thinking model that reasons over complex problems in code, math, and STEM. Suited to analyzing large datasets, codebases, and documents using long context. Accepts audio, image, video, text, and PDF input and returns text, with a 1M token context window.
Gemini 2.5 Flash
A Gemini 2.5 model balancing cost and capability for large scale processing, low-latency, high volume tasks that require thinking, and agentic use cases. Accepts text, image, video, and audio input and returns text, with a 1M token context window and adjustable thinking budgets.
Gemini 2.5 Flash Lite
Gemini 2.5 Flash Lite is a faster, cost-efficient version of Gemini 2.5 Flash. It's optimized for quick responses and cost-effective applications. Supports multimodal inputs including text, images, video, and audio.
Gemini 3.1 Flash-Lite
A low-cost multimodal model in the Gemini 3.1 family, aimed at high-frequency, lightweight tasks. Google positions it for high-volume agentic work, simple data extraction, translation, transcription, and latency-sensitive applications where budget and speed are the primary constraints. Supports thinking for enhanced accuracy.
Gemini 2.5 Flash
A Gemini 2.5 model balancing cost and capability for large scale processing, low-latency, high volume tasks that require thinking, and agentic use cases. Accepts text, image, video, and audio input and returns text, with a 1M token context window and adjustable thinking budgets.
gemini-2-0-flash
gemini-1-5-pro
X AI models
Grok models with real-time information access.
Grok 4.1 Fast
Grok 4.1 Fast in non-reasoning mode for instant responses, with a 2M token context window and tool calling for agentic workflows such as customer support and finance. xAI retired this slug on May 15, 2026; requests now route to Grok 4.3 in non-reasoning mode.
Grok 4.1 Fast Reasoning
Grok 4.1 Fast with reasoning enabled — an xAI tool-calling model for agentic tasks, with a 2M token context window. xAI retired this slug on May 15, 2026; requests now route to Grok 4.3 with low reasoning effort.
Grok 4 Fast Reasoning
Grok 4 Fast with reasoning enabled, using a unified architecture that blends reasoning and non-reasoning modes, with web and X search and a 2M token context window. xAI retired this slug on May 15, 2026; requests now route to Grok 4.3 with low reasoning effort.
Grok 4 Fast
Grok 4 Fast in non-reasoning mode for instant responses, with web and X search and a 2M token context window. xAI retired this slug on May 15, 2026; requests now route to Grok 4.3 in non-reasoning mode.
Grok 4
Grok 4 is xAI's July 2025 reasoning model, trained with scaled reinforcement learning on the Colossus cluster, with native tool use, real-time search, and text and image input. xAI retired the grok-4-0709 slug on May 15, 2026; requests now route to Grok 4.3 with low reasoning effort.
Grok Code Fast
An xAI coding model for agentic workflows, built on a new architecture with prompt caching and adept at TypeScript, Python, Java, Rust, C++, and Go. xAI retired this slug on May 15, 2026; requests now route to Grok Build 0.1.
DeepSeek models
Advanced reasoning models with thinking capabilities.
deepseek-chat
deepseek-reasoner
deepseek-coder
DeepSeek V4.1 Flash
DeepSeek's official V4.1 Flash model (`deepseek-flash`) with a 1M-token context window, up to 384K output tokens, hybrid thinking mode (configurable reasoning effort), native Chat Completions vision (JPEG/PNG/GIF/WebP), function calling, and streaming. Served on the official DeepSeek API at off-peak list pricing.
Mistral models
Open-source and commercial models from Mistral AI.
Qwen models
Multilingual models from Alibaba with strong reasoning and coding.
qwen3-max
qwen-plus
Qwen Plus Character
The role-playing model of the Qwen series. This is a dynamically updated version, and notifications will be provided in advance for any model updates. It is suitable for anthropomorphic role-playing and has optimized capabilities in following character personalities and maintaining consistent role-play interactions.
qwen-flash
Qwen Coder Plus
Powered by Qwen3, this is a powerful Coding Agent that excels in tool calling and environment interaction to achieve autonomous programming. It combines outstanding coding proficiency with versatile general-purpose abilities.
Qwen3 Next 80B Instruct
A new generation of open-source, non-thinking mode model powered by Qwen3. This version demonstrates superior Chinese text understanding, augmented logical reasoning, and enhanced capabilities in text generation tasks over the previous iteration (Qwen3-235B-A22B-Instruct-2507).
Qwen3 Next 80B Thinking
A new generation of Qwen3-based open-source thinking mode models. This version offers improved instruction following and streamlined summary responses over the previous iteration (Qwen3-235B-A22B-Thinking-2507).
Qwen 3.5 Plus
Hosted Plus-tier model in the Qwen 3.5 series (February 2026) on Alibaba Cloud Model Studio. Native vision-language model combining linear attention with sparse MoE (Qwen3.5-397B-A17B class), with optional chain-of-thought via enable_thinking. Strong on coding, agents, tool use, and multimodal reasoning across text, images, and video, with a 1M-token context window by default.
Qwen 3.6 Plus
Alibaba’s April 2026 Plus-tier model on Model Studio: major upgrade in agentic coding (frontend through repo-level work), sharper multimodal perception and reasoning, and a 1M-token default context. Supports enable_thinking for hybrid reasoning and preserve_thinking for multi-turn agents so prior reasoning traces stay in context. Strong on coding-agent, tool-use, STEM, and long-context benchmarks.
qwen-max-latest
qwen-turbo-latest
qwen3-235b-a22b-instruct
qwen3-32b-instruct
Kimi models
Frontier mixture-of-experts models from Moonshot AI with tool use.
Kimi K3
Kimi K3 is Moonshot AI's 2.8T-parameter Mixture-of-Experts model for long-horizon coding, knowledge work, deep reasoning, visual understanding, and agentic tool use. It always reasons, supports native image and video input, and provides a 1,048,576-token context window.
Kimi K2.5
Kimi K2.5 is Moonshot AI's native multimodal Mixture-of-Experts model with 1T total and 32B activated parameters. Its MoonViT vision encoder supports text, image and video input, and the model runs in both thinking and non-thinking modes across dialogue and agent tasks. It decomposes complex tasks into parallel sub-tasks executed by a self-directed swarm of dynamically instantiated sub-agents.
kimi-k2-turbo-preview
moonshotai-kimi-k2-instruct
Z.AI models
GLM models from Z.AI designed for real-world development with strong coding capabilities.
glm-4.7
glm-4-32b-0414-128k
GLM 4.7 Flash Heretic
A community-abliterated, open-weights variant of Z.ai's GLM-4.7-Flash (30B-A3B mixture-of-experts), decensored with Heretic v1.1.0 to reduce refusals while preserving the base model's thinking mode. Text in and out, with a 200K-token context window on Venice.
Cohere models
Enterprise-focused Command models for business applications.
Command A
Cohere's Command A, a 111-billion-parameter model for enterprise tasks including tool use, retrieval-augmented generation, agents, and multilingual use cases, with a 256K context window.
Command A Reasoning
Command A Reasoning is Cohere's first reasoning model, producing a reasoning pass before its answer, and is aimed at enterprise tasks including tool use, retrieval augmented generation (RAG), agents, and multilingual use across 23 languages.
Command R7B
Cohere's Command R7B, the 7-billion-parameter entry in its R family of enterprise-focused large language models, built for high-volume workloads such as retrieval-augmented generation, tool use, and agents.
command-r-plus
command-r
Groq models
High-speed inference models optimized for performance.
Perplexity models
Search-augmented Sonar models with real-time web access.
Thesys models
C1 models that generate interactive UI components and visualizations.
Community & Open models
Open-source and community-fine-tuned models with unique capabilities.