Explore Models

Browse and compare available AI models, capabilities, and pricing.

AionLabs: Aion-1.0
aion-1.0

In: textOut: text
Context:131.1K
In:$4.8/1M
Out:$9.6/1M
AionLabs: Aion-1.0-Mini
aion-1.0-mini

In: textOut: text
Context:131.1K
In:$0.84/1M
Out:$1.68/1M
AionLabs: Aion-2.0
aion-2.0

Aion-2.0 is a variant of DeepSeek V3.2 optimized for immersive roleplaying and storytelling. It is particularly strong at introducing tension, crises, and conflict into stories, making narratives feel more engaging....

In: textOut: text
Context:131.1K
In:$0.96/1M
Out:$1.92/1M
AionLabs: Aion-3.0
aion-3.0

Aion-3.0 is a multi-model roleplaying and storytelling system from AionLabs, built on the GLM family of models. It uses a collaborative generation process in which multiple specialized models each contribute...

In: textOut: text
Context:131.1K
In:$3.6/1M
Out:$7.2/1M
AionLabs: Aion-3.0-Mini
aion-3.0-mini

Aion-3.0 Mini is a multi-model roleplaying and storytelling system from AionLabs, built on the DeepSeek family of models. It uses a collaborative generation process in which multiple specialized models each...

In: textOut: text
Context:131.1K
In:$0.84/1M
Out:$1.68/1M
AionLabs: Aion-RP 1.0 (8B)
aion-rp-llama-3.1-8b

Aion-RP-Llama-3.1-8B ranks the highest in the character evaluation portion of the RPBench-Auto benchmark, a roleplaying-specific variant of Arena-Hard-Auto, where LLMs evaluate each other’s responses. It is a fine-tuned base model...

In: textOut: text
Context:32.8K
In:$0.96/1M
Out:$1.92/1M
baichuan-m2-32b
baichuan-m2-32b

In: textOut: text
Context:131.1K
In:$0.084/1M
Out:$0.084/1M
BGE Reranker v2 M3
bge-reranker-v2-m3

In: textOut: text
Context:8.2K
In:$0.012/1M
claude-3-5-haiku-20241022
claude-3-5-haiku-20241022

In: textIn: imageIn: pdfOut: text
Context:200K
In:$0.96/1M
Out:$4.8/1M
Claude Sonnet 3.5 v2
claude-3-5-sonnet-20241022

In: textIn: imageIn: pdfOut: text
Context:200K
In:$3.6/1M
Out:$18/1M
Claude Sonnet 3.7
claude-3-7-sonnet

In: textIn: imageIn: pdfOut: text
Context:200K
In:$3.6/1M
Out:$18/1M
Claude Sonnet 3.7
claude-3-7-sonnet-20250219

In: textIn: imageIn: pdfOut: text
Context:200K
In:$3.6/1M
Out:$18/1M
Anthropic: Claude 3 Haiku
claude-3-haiku

Claude 3 Haiku is Anthropic's fastest and most compact model for near-instant responsiveness. Quick and accurate targeted performance. See the launch announcement and benchmark results [here](https://www.anthropic.com/news/claude-3-haiku) #multimodal

In: textIn: imageOut: text
Context:200K
In:$0.3/1M
Out:$1.5/1M
Anthropic: Claude 3 Haiku
claude-3-haiku-20240307

In: textIn: imageOut: text
Context:200K
In:$0.3/1M
Out:$1.5/1M
Claude Opus 3
claude-3-opus-20240229

In: textIn: imageIn: pdfOut: text
Context:200K
In:$18/1M
Out:$90/1M
Claude 3.5 Haiku
claude-3.5-haiku

In: textIn: imageOut: text
Context:200K
In:$0.96/1M
Out:$4.8/1M
Claude 3.5 Sonnet
claude-3.5-sonnet

In: textIn: imageOut: text
Context:200K
In:$3.6/1M
Out:$18/1M
Claude 3.7 Sonnet
claude-3.7-sonnet

In: textIn: imageOut: text
Context:200K
In:$3.6/1M
Out:$18/1M
Anthropic: Claude Fable 5
claude-fable-5

Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text output, with reasoning support and...

In: textIn: imageIn: pdfOut: text
Context:1M
In:$12/1M
Out:$60/1M
Claude Fable 5.1
claude-fable-5-1

In: textIn: imageIn: pdfOut: text
Context:1M
In:$12/1M
Out:$60/1M
Anthropic: Claude Fable 5.1
claude-fable-5.1

Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long code refactors, front-end and visual...

In: textIn: imageIn: pdfOut: text
Context:1M
In:$12/1M
Out:$60/1M
Anthropic: Claude Fable Latest
claude-fable-latest

This model always redirects to the latest model in the Claude Fable family.

In: textIn: imageIn: pdfOut: text
Context:1M
In:$12/1M
Out:$60/1M
Claude Haiku 4.5
claude-haiku-4-5-20251001

In: textIn: imageIn: pdfOut: text
Context:200K
In:$1.2/1M
Out:$6/1M
Claude Haiku 4.5 Thinking
claude-haiku-4-5-20251001-thinking

In: textIn: imageIn: pdfOut: text
Context:200K
In:$1.2/1M
Out:$6/1M
Anthropic: Claude Haiku 4.5
claude-haiku-4.5

Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance...

In: textIn: imageIn: pdfOut: text
Context:200K
In:$1.2/1M
Out:$6/1M
Anthropic Claude Haiku Latest
claude-haiku-latest

This model always redirects to the latest model in the Anthropic Claude Haiku family.

In: textIn: imageIn: pdfOut: text
Context:200K
In:$1.2/1M
Out:$6/1M
Claude Mythos 5
claude-mythos-5

In: textIn: imageIn: pdfOut: text
Context:1M
In:$12/1M
Out:$60/1M
Anthropic: Claude Opus 4
claude-opus-4

Claude Opus 4 is benchmarked as the world’s best coding model, at time of release, bringing sustained performance on complex, long-running tasks and agent workflows. It sets new benchmarks in...

In: imageIn: textIn: pdfOut: text
Context:200K
In:$18/1M
Out:$90/1M
claude-opus-4-1-20250805
claude-opus-4-1-20250805

In: textIn: imageIn: pdfOut: text
Context:200K
In:$18/1M
Out:$90/1M
claude-opus-4-1-20250805-thinking
claude-opus-4-1-20250805-thinking

In: textIn: imageOut: text
Context:200K
In:$18/1M
Out:$90/1M
claude-opus-4-20250514
claude-opus-4-20250514

In: textIn: imageIn: pdfOut: text
Context:200K
In:$18/1M
Out:$90/1M
Claude Opus 4.5
claude-opus-4-5-20251101

In: textIn: imageIn: pdfOut: text
Context:200K
In:$6/1M
Out:$30/1M
claude-opus-4-5-20251101-thinking
claude-opus-4-5-20251101-thinking

In: textIn: imageOut: text
Context:200K
In:$6/1M
Out:$30/1M
Claude Opus 4.6
claude-opus-4-6

In: textIn: imageIn: pdfOut: text
Context:1M
In:$6/1M
Out:$30/1M
Claude Opus 4.6 Fast
claude-opus-4-6-fast

In: textIn: imageOut: text
Context:1M
In:$43.2/1M
Out:$216/1M
claude-opus-4-6-thinking
claude-opus-4-6-thinking

In: textIn: imageIn: pdfOut: text
Context:1M
In:$6/1M
Out:$30/1M
Claude Opus 4.7
claude-opus-4-7

In: textIn: imageIn: pdfOut: text
Context:1M
In:$6/1M
Out:$30/1M
Claude Opus 4.7 Fast
claude-opus-4-7-fast

In: textIn: imageOut: text
Context:1M
In:$43.2/1M
Out:$216/1M
Claude Opus 4.8
claude-opus-4-8

In: textIn: imageIn: pdfOut: text
Context:1M
In:$6/1M
Out:$30/1M
Claude Opus 4.8 Fast
claude-opus-4-8-fast

In: textIn: imageOut: text
Context:1M
In:$14.4/1M
Out:$72/1M
Anthropic: Claude Opus 4.1
claude-opus-4.1

Claude Opus 4.1 is an updated version of Anthropic’s flagship model, offering improved performance in coding, reasoning, and agentic tasks. It achieves 74.5% on SWE-bench Verified and shows notable gains...

In: textIn: imageIn: pdfOut: text
Context:200K
In:$18/1M
Out:$90/1M
Anthropic: Claude Opus 4.5
claude-opus-4.5

Claude Opus 4.5 is Anthropic’s frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon computer use. It offers strong multimodal capabilities, competitive performance across real-world coding and...

In: textIn: imageIn: pdfOut: text
Context:200K
In:$6/1M
Out:$30/1M
Anthropic: Claude Opus 4.6
claude-opus-4.6

Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks. It is built for agents that operate across entire workflows rather than single prompts, making it especially effective...

In: textIn: imageIn: pdfOut: text
Context:1M
In:$6/1M
Out:$30/1M
Anthropic: Claude Opus 4.6 (Fast)
claude-opus-4.6-fast

In: imageIn: textOut: text
Context:1M
In:$36/1M
Out:$180/1M
Anthropic: Claude Opus 4.7
claude-opus-4.7

Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic strengths of Opus 4.6, it delivers stronger performance on...

In: textIn: imageIn: pdfOut: text
Context:1M
In:$6/1M
Out:$30/1M
Anthropic: Claude Opus 4.7 (Fast)
claude-opus-4.7-fast

Fast-mode variant of [Opus 4.7](/anthropic/claude-opus-4.7) - identical capabilities with higher output speed at premium 6x pricing. Learn more in Anthropic's docs: https://platform.claude.com/docs/en/build-with-claude/fast-mode

In: textIn: imageIn: pdfOut: text
Context:1M
In:$36/1M
Out:$180/1M
Anthropic: Claude Opus 4.8
claude-opus-4.8

Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token...

In: textIn: imageIn: pdfOut: text
Context:1M
In:$6/1M
Out:$30/1M
Claude Opus 4.8 (Fast)
claude-opus-4.8-fast

In: textIn: imageIn: pdfOut: text
Context:1M
In:$12/1M
Out:$60/1M
Claude Opus 5
claude-opus-5

Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software tasks, code review and bug finding, visual analysis...

In: textIn: imageIn: pdfOut: text
Context:1M
In:$6/1M
Out:$30/1M
Claude Opus 5 Fast
claude-opus-5-fast

In: textIn: imageOut: text
Context:1M
In:$14.4/1M
Out:$72/1M
Anthropic: Claude Opus Latest
claude-opus-latest

This model always redirects to the latest model in the Claude Opus family.

In: textIn: imageIn: pdfOut: text
Context:1M
In:$6/1M
Out:$30/1M
Anthropic: Claude Sonnet 4
claude-sonnet-4

Claude Sonnet 4 significantly enhances the capabilities of its predecessor, Sonnet 3.7, excelling in both coding and reasoning tasks with improved precision and controllability. Achieving state-of-the-art performance on SWE-bench (72.7%),...

In: imageIn: textIn: pdfOut: text
Context:1M
In:$3.6/1M
Out:$18/1M
claude-sonnet-4-20250514
claude-sonnet-4-20250514

In: textIn: imageIn: pdfOut: text
Context:200K
In:$3.6/1M
Out:$18/1M
Claude Sonnet 4.5
claude-sonnet-4-5-20250929

In: textIn: imageIn: pdfOut: text
Context:1M
In:$3.6/1M
Out:$18/1M
claude-sonnet-4-5-20250929-thinking
claude-sonnet-4-5-20250929-thinking

In: textIn: imageIn: pdfOut: text
Context:200K
In:$3.6/1M
Out:$18/1M
Claude Sonnet 4.6
claude-sonnet-4-6

In: textIn: imageIn: pdfOut: text
Context:1M
In:$3.6/1M
Out:$18/1M
Anthropic: Claude Sonnet 4.5
claude-sonnet-4.5

Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with...

In: textIn: imageIn: pdfOut: text
Context:1M
In:$3.6/1M
Out:$18/1M
Anthropic: Claude Sonnet 4.6
claude-sonnet-4.6

Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. It excels at iterative development, complex codebase navigation, end-to-end project management with...

In: textIn: imageIn: pdfOut: text
Context:1M
In:$3.6/1M
Out:$18/1M
Anthropic: Claude Sonnet 5
claude-sonnet-5

Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It supports adaptive thinking with selectable reasoning effort levels (low, medium, high, max,...

In: textIn: imageIn: pdfOut: text
Context:1M
In:$2.4/1M
Out:$12/1M
Anthropic Claude Sonnet Latest
claude-sonnet-latest

This model always redirects to the latest model in the Anthropic Claude Sonnet family.

In: textIn: imageIn: pdfOut: text
Context:1M
In:$2.4/1M
Out:$12/1M
AlfredPros: CodeLLaMa 7B Instruct Solidity
codellama-7b-instruct-solidity

In: textOut: text
Context:4.1K
In:$0.96/1M
Out:$1.44/1M
Arcee AI: Coder Large
coder-large

In: textOut: text
Context:32.8K
In:$0.6/1M
Out:$0.96/1M
Mistral: Codestral 2508
codestral-2508

Mistral's cutting-edge language model for coding released end of July 2025. Codestral specializes in low-latency, high-frequency tasks such as fill-in-the-middle (FIM), code correction and test generation. [Blog Post](https://mistral.ai/news/codestral-25-08)

In: textIn: pdfOut: text
Context:256K
In:$0.36/1M
Out:$1.08/1M
Codestral (latest)
codestral-latest

In: textOut: text
Context:256K
In:$0.36/1M
Out:$1.08/1M
Cogito v2.1 671B
cogito-v2-1-671b

In: textOut: text
Context:163.8K
In:$1.5/1M
Out:$1.5/1M
Deep Cogito: Cogito v2.1 671B
cogito-v2.1-671b

Cogito v2.1 671B MoE represents one of the strongest open models globally, matching performance of frontier closed and open models. This model is trained using self play with reinforcement learning...

In: textOut: text
Context:128K
In:$1.5/1M
Out:$1.5/1M
Cohere: Command A
command-a

Command A is an open-weights 111B parameter model with a 256k context window focused on delivering great performance across agentic, multilingual, and coding use cases. Compared to other leading proprietary...

In: textOut: text
Context:256K
In:$3/1M
Out:$12/1M
Cohere: Command R (08-2024)
command-r-08-2024

command-r-08-2024 is an update of the [Command R](/models/cohere/command-r) with improved performance for multilingual retrieval-augmented generation (RAG) and tool use. More broadly, it is better at math, code and reasoning and...

In: textOut: text
Context:128K
In:$0.18/1M
Out:$0.72/1M
Cohere: Command R+ (08-2024)
command-r-plus-08-2024

command-r-plus-08-2024 is an update of the [Command R+](/models/cohere/command-r-plus) with roughly 50% higher throughput and 25% lower latencies as compared to the previous Command R+ version, while keeping the hardware footprint...

In: textOut: text
Context:128K
In:$3/1M
Out:$12/1M
Cohere: Command R7B (12-2024)
command-r7b-12-2024

Command R7B (12-2024) is a small, fast update of the Command R+ model, delivered in December 2024. It excels at RAG, tool use, agents, and similar tasks requiring complex reasoning...

In: textOut: text
Context:128K
In:$0.045/1M
Out:$0.18/1M
TheDrummer: Cydonia 24B V4.1
cydonia-24b-v4.1

Uncensored and creative writing model based on Mistral Small 3.2 24B with good recall, prompt adherence, and intelligence.

In: textOut: text
Context:131.1K
In:$0.36/1M
Out:$0.6/1M
DALL-E-3
dall-e-3

In: textOut: image
Context:800
In:$0.036/1M
Out:$0.18/1M
DeepSeek: DeepSeek V3
deepseek-chat

DeepSeek-V3 is the latest model from the DeepSeek team, building upon the instruction following and coding abilities of the previous versions. Pre-trained on nearly 15 trillion tokens, the reported evaluations...

In: textOut: text
Context:163.8K
In:$0.384/1M
Out:$1.068/1M
DeepSeek: DeepSeek V3 0324
deepseek-chat-v3-0324

DeepSeek V3, a 685B-parameter, mixture-of-experts model, is the latest iteration of the flagship chat model family from the DeepSeek team. It succeeds the [DeepSeek V3](/deepseek/deepseek-chat-v3) model and performs really well...

In: textOut: text
Context:163.8K
In:$0.3/1M
Out:$1.2/1M
DeepSeek: DeepSeek V3.1
deepseek-chat-v3.1

DeepSeek-V3.1 is a large hybrid reasoning model (671B parameters, 37B active) that supports both thinking and non-thinking modes via prompt templates. It extends the DeepSeek-V3 base with a two-phase long-context...

In: textOut: text
Context:163.8K
In:$0.66/1M
Out:$1.98/1M
Deepseek Prover V2 671B
deepseek-prover-v2-671b

In: textOut: text
Context:160K
In:$0.84/1M
Out:$3/1M
DeepSeek: R1
deepseek-r1

DeepSeek R1 is here: Performance on par with [OpenAI o1](/openai/o1), but open-sourced and with fully open reasoning tokens. It's 671B parameters in size, with 37B active in an inference pass....

In: textOut: text
Context:64K
In:$0.84/1M
Out:$3/1M
DeepSeek: R1 0528
deepseek-r1-0528

May 28th update to the [original DeepSeek R1](/deepseek/deepseek-r1) Performance on par with [OpenAI o1](/openai/o1), but open-sourced and with fully open reasoning tokens. It's 671B parameters in size, with 37B active...

In: textOut: text
Context:163.8K
In:$0.6/1M
Out:$2.58/1M
DeepSeek: R1 Distill Llama 70B
deepseek-r1-distill-llama-70b

DeepSeek R1 Distill Llama 70B is a distilled large language model based on [Llama-3.3-70B-Instruct](/meta-llama/llama-3.3-70b-instruct), using outputs from [DeepSeek R1](/deepseek/deepseek-r1). The model combines advanced distillation techniques to achieve high performance across...

In: textOut: text
Context:8.2K
In:$0.96/1M
Out:$0.96/1M
DeepSeek R1 (Turbo)
deepseek-r1-turbo

In: textOut: text
Context:64K
In:$0.84/1M
Out:$3/1M
TNG: DeepSeek R1T2 Chimera
deepseek-r1t2-chimera

DeepSeek-TNG-R1T2-Chimera is the second-generation Chimera model from TNG Tech. It is a 671 B-parameter mixture-of-experts text-generation model assembled from DeepSeek-AI’s R1-0528, R1, and V3-0324 checkpoints with an Assembly-of-Experts merge. The...

In: textOut: text
Context:163.8K
In:$0.36/1M
Out:$1.32/1M
Deepseek-Reasoner
deepseek-reasoner

In: textOut: text
Context:128K
In:$0.348/1M
Out:$0.516/1M
DeepSeek-V3
deepseek-v3

In: textOut: text
Context:128K
In:$0.344/1M
Out:$1.376/1M
DeepSeek-V3-0324
deepseek-v3-0324

In: textOut: text
Context:128K
In:$0.336/1M
Out:$1.368/1M
DeepSeek-V3-0324 (Fast)
DeepSeek-V3-0324-fast

In: textOut: text
Context:128K
In:$0.9/1M
Out:$2.7/1M
DeepSeek V3.1
deepseek-v3-1

In: textOut: text
Context:131.1K
In:$0.689/1M
Out:$2.065/1M
DeepSeek V3.2
deepseek-v3-2

In: textOut: text
Context:128K
In:$0.684/1M
Out:$2.052/1M
DeepSeek V3.2 Exp
deepseek-v3-2-exp

In: textOut: text
Context:131.1K
In:$0.344/1M
Out:$0.517/1M
DeepSeek V3 (Turbo)
deepseek-v3-turbo

In: textOut: text
Context:64K
In:$0.48/1M
Out:$1.56/1M
DeepSeek-V3.1
deepseek-v3.1

In: textOut: text
Context:131.1K
In:$0.228/1M
Out:$0.852/1M
Nex AGI: DeepSeek V3.1 Nex N1
deepseek-v3.1-nex-n1

In: textOut: text
Context:131.1K
In:$0.324/1M
Out:$1.2/1M
DeepSeek: DeepSeek V3.1 Terminus
deepseek-v3.1-terminus

DeepSeek-V3.1 Terminus is an update to [DeepSeek V3.1](/deepseek/deepseek-chat-v3.1) that maintains the model's original capabilities while addressing issues reported by users, including language consistency and agent capabilities, further optimizing the model's...

In: textOut: text
Context:163.8K
In:$0.324/1M
Out:$1.2/1M
DeepSeek: DeepSeek V3.2
deepseek-v3.2

DeepSeek-V3.2 is a large language model designed to harmonize high computational efficiency with strong reasoning and agentic tool-use performance. It introduces DeepSeek Sparse Attention (DSA), a fine-grained sparse attention mechanism...

In: textOut: text
Context:163.8K
In:$0.323/1M
Out:$0.48/1M
DeepSeek: DeepSeek V3.2 Exp
deepseek-v3.2-exp

DeepSeek-V3.2-Exp is an experimental large language model released by DeepSeek as an intermediate step between V3.1 and future architectures. It introduces DeepSeek Sparse Attention (DSA), a fine-grained sparse attention mechanism...

In: textOut: text
Context:163.8K
In:$0.324/1M
Out:$0.492/1M
DeepSeek/DeepSeek-V3.2-Exp-Thinking
deepseek-v3.2-exp-thinking

In: textOut: text
Context:128K
In:$0.336/1M
Out:$0.504/1M
DeepSeek-V3.2-Fast
deepseek-v3.2-fast

In: textOut: text
Context:128K
In:$1.32/1M
Out:$3.948/1M
DeepSeek-V3.2-Speciale
deepseek-v3.2-speciale

In: textOut: text
Context:128K
In:$0.696/1M
Out:$2.016/1M
DeepSeek-V3.2-Thinking
deepseek-v3.2-thinking

In: textOut: text
Context:128K
In:$0.348/1M
Out:$0.516/1M
DeepSeek: DeepSeek V4 Flash 0423
deepseek-v4-flash

DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and...

In: textOut: text
Context:1M
In:$0.108/1M
Out:$0.216/1M
DeepSeek: DeepSeek V4 Flash 0731
deepseek-v4-flash-0731

DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows....

In: textOut: text
Context:1.3M
In:$0.168/1M
Out:$0.336/1M
DeepSeek V4 Flash Latest
deepseek-v4-flash-latest

This model always redirects to the latest model in the DeepSeek V4 Flash family.

In: textOut: text
Context:1.3M
In:$0.054/1M
Out:$0.108/1M
DeepSeek: DeepSeek V4 Flash Vision Exp
deepseek-v4-flash-vision-exp

DeepSeek V4 Flash Vision Exp is an experimental vision-enabled version of [DeepSeek V4 Flash 0731](https://openrouter.ai/deepseek/deepseek-v4-flash-0731) from DeepSeek, adding image understanding while matching the base model on text capabilities including agents,...

In: textIn: imageOut: text
Context:1M
In:$0.264/1M
Out:$0.792/1M
DeepSeek: DeepSeek V4 Pro 0423
deepseek-v4-pro

DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window. It is designed for advanced reasoning, coding,...

In: textOut: text
Context:1M
In:$1.242/1M
Out:$2.485/1M
DeepSeek: DeepSeek V4 Pro 0813
deepseek-v4-pro-0813

DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model from DeepSeek. This is the GA release of DeepSeek V4 Pro.

In: textOut: text
Context:1M
In:$1.345/1M
Out:$4.034/1M
Mistral: Devstral 2 2512
devstral-2512

Devstral 2 is a state-of-the-art open-source model by Mistral AI specializing in agentic coding. It is a 123B-parameter dense transformer model supporting a 256K context window. Devstral 2 supports exploring...

In: textIn: pdfOut: text
Context:262.1K
In:$0.48/1M
Out:$2.4/1M
devstral-latest
devstral-latest

In: textOut: text
Context:256K
In:$0.528/1M
Out:$2.64/1M
Mistral: Devstral Medium
devstral-medium

In: textOut: text
Context:131.1K
In:$0.48/1M
Out:$2.4/1M
Devstral Medium
devstral-medium-2507

In: textOut: text
Context:128K
In:$0.48/1M
Out:$2.4/1M
Mistral: Devstral Small 1.1
devstral-small

In: textOut: text
Context:131.1K
In:$0.12/1M
Out:$0.36/1M
Devstral Small
devstral-small-2507

In: textOut: text
Context:128K
In:$0.12/1M
Out:$0.36/1M
Venice: Uncensored
dolphin-mistral-24b-venice-edition

Venice Uncensored Dolphin Mistral 24B Venice Edition is a fine-tuned variant of Mistral-Small-24B-Instruct-2501, developed by dphn.ai in collaboration with Venice.ai. This model is designed as an “uncensored” instruct-tuned LLM, preserving...

In: textOut: text
Context:128K
In:$0.24/1M
Out:$1.08/1M
Seed 2.0 Code
doubao-seed-2-0-code-preview-260215

In: textIn: imageIn: videoOut: text
Context:262.1K
In:$0.57/1M
Out:$2.85/1M
Doubao Seed 2.0 Lite
doubao-seed-2-0-lite-260215

In: textOut: text
Context:256K
In:$0.175/1M
Out:$1.049/1M
Seed 2.0 Lite
doubao-seed-2-0-lite-260428

In: textIn: imageIn: videoOut: text
Context:256K
In:$0.107/1M
Out:$0.641/1M
Doubao Seed 2.0 Mini
doubao-seed-2-0-mini-260215

In: textOut: text
Context:256K
In:$0.059/1M
Out:$0.581/1M
Seed 2.0 Mini
doubao-seed-2-0-mini-260428

In: textIn: imageIn: videoOut: text
Context:256K
In:$0.036/1M
Out:$0.356/1M
Seed 2.0 Pro
doubao-seed-2-0-pro-260215

In: textIn: imageIn: videoOut: text
Context:256K
In:$0.57/1M
Out:$2.85/1M
Seed 2.1 Pro
doubao-seed-2-1-pro-260628

In: textIn: imageIn: videoOut: text
Context:256K
In:$1.069/1M
Out:$5.344/1M
Seed 2.1 Turbo
doubao-seed-2-1-turbo-260628

In: textIn: imageIn: videoOut: text
Context:256K
In:$0.534/1M
Out:$2.672/1M
Seed Evolving
doubao-seed-evolving

In: textIn: imageIn: videoOut: text
Context:256K
In:$1.069/1M
Out:$5.344/1M
Baidu: ERNIE 4.5 21B A3B
ernie-4.5-21b-a3b

In: textOut: text
Context:120K
In:$0.084/1M
Out:$0.336/1M
Baidu Ernie 4.5 21B A3B Thinking
ernie-4.5-21b-a3b-thinking

In: textOut: text
Context:128K
In:$0.084/1M
Out:$0.336/1M
Baidu: ERNIE 4.5 300B A47B
ernie-4.5-300b-a47b

In: textOut: text
Context:123K
In:$0.336/1M
Out:$1.32/1M
ERNIE 4.5 300B A47B
ernie-4.5-300b-a47b-paddle

In: textOut: text
Context:123K
In:$0.336/1M
Out:$1.32/1M
ERNIE 4.5 VL 28B A3B
ernie-4.5-vl-28b-a3b

In: textIn: imageOut: text
Context:30K
In:$0.168/1M
Out:$0.672/1M
Baidu: ERNIE 4.5 VL 424B A47B
ernie-4.5-vl-424b-a47b

ERNIE-4.5-VL-424B-A47B is a multimodal Mixture-of-Experts (MoE) model from Baidu’s ERNIE 4.5 series, featuring 424B total parameters with 47B active per token. It is trained jointly on text and image data...

In: imageIn: textOut: text
Context:123K
In:$0.504/1M
Out:$1.5/1M
Sakana: Fugu Ultra
fugu-ultra

Fugu Ultra is the higher-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent orchestration system: a language model trained to route...

In: textIn: imageOut: text
Context:1M
In:$6/1M
Out:$36/1M
Gemini 2.5 Pro
gemini-2-5-pro

In: textIn: imageIn: audioIn: videoIn: pdfOut: text
Context:1M
In:$1.5/1M
Out:$12/1M
Google: Gemini 2.0 Flash
gemini-2.0-flash-001

In: audioIn: imageIn: pdfIn: textIn: videoOut: text
Context:1M
In:$0.12/1M
Out:$0.48/1M
Google: Gemini 2.0 Flash Lite
gemini-2.0-flash-lite-001

In: audioIn: imageIn: pdfIn: textIn: videoOut: text
Context:1M
In:$0.09/1M
Out:$0.36/1M
Google: Gemini 2.5 Flash
gemini-2.5-flash

Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific tasks. It includes built-in "thinking" capabilities, enabling it to provide responses with greater...

In: textIn: imageIn: audioIn: videoIn: pdfOut: text
Context:1M
In:$0.36/1M
Out:$3/1M
Google: Nano Banana (Gemini 2.5 Flash Image)
gemini-2.5-flash-image

Gemini 2.5 Flash Image, a.k.a. "Nano Banana," is now generally available. It is a state of the art image generation model with contextual understanding. It is capable of image generation, edits, and multi-turn conversations. Aspect ratios can be controlled with the [image_config API Parameter](https://openrouter.ai/docs/features/multimodal/image-generation#image-aspect-ratio-configuration)

In: textIn: imageOut: textOut: image
Context:32.8K
In:$0.36/1M
Out:$36/1M
Google: Gemini 2.5 Flash Lite
gemini-2.5-flash-lite

Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency. It offers improved throughput, faster token generation, and better performance...

In: textIn: imageIn: audioIn: videoIn: pdfOut: text
Context:1M
In:$0.12/1M
Out:$0.48/1M
gemini-2.5-flash-lite-preview-09-2025
gemini-2.5-flash-lite-preview-09-2025

In: textIn: imageOut: text
Context:1M
In:$0.12/1M
Out:$0.48/1M
gemini-2.5-flash-preview-09-2025
gemini-2.5-flash-preview-09-2025

In: textIn: imageOut: text
Context:1M
In:$0.36/1M
Out:$3/1M
Google: Gemini 2.5 Pro
gemini-2.5-pro

Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy...

In: textIn: imageIn: audioIn: videoIn: pdfOut: text
Context:1M
In:$1.5/1M
Out:$12/1M
Google: Gemini 2.5 Pro Preview 06-05
gemini-2.5-pro-preview

Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy...

In: pdfIn: imageIn: textIn: audioOut: text
Context:1M
In:$1.5/1M
Out:$12/1M
Google: Gemini 2.5 Pro Preview 05-06
gemini-2.5-pro-preview-05-06

Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy...

In: textIn: imageIn: pdfIn: audioIn: videoOut: text
Context:1M
In:$1.5/1M
Out:$12/1M
Gemini 3.1 Flash Lite Preview
gemini-3-1-flash-lite

In: textIn: imageIn: videoIn: audioIn: pdfOut: text
Context:1M
In:$0.3/1M
Out:$1.8/1M
Gemini 3.1 Pro Preview
gemini-3-1-pro-preview

In: textIn: imageIn: audioIn: videoOut: text
Context:1M
In:$3/1M
Out:$18/1M
Gemini 3.5 Flash
gemini-3-5-flash

In: textIn: imageIn: audioIn: videoOut: text
Context:1M
In:$1.86/1M
Out:$11.34/1M
Gemini 3.5 Flash-Lite
gemini-3-5-flash-lite

In: textIn: imageIn: audioIn: videoOut: text
Context:1M
In:$0.45/1M
Out:$3.75/1M
Gemini 3.6 Flash
gemini-3-6-flash

In: textIn: imageIn: audioIn: videoOut: text
Context:1M
In:$1.125/1M
Out:$5.625/1M
Gemini 3.7 Flash
gemini-3-7-flash

In: textIn: imageIn: audioIn: videoOut: text
Context:1M
In:$1.125/1M
Out:$5.625/1M
Gemini 3.8 Flash
gemini-3-8-flash

In: textIn: imageIn: audioIn: videoOut: text
Context:1M
In:$1.125/1M
Out:$5.625/1M
Gemini 3 Flash
gemini-3-flash

In: textIn: imageIn: videoIn: audioIn: pdfOut: text
Context:1M
In:$0.6/1M
Out:$3.6/1M
Google: Gemini 3 Flash Preview
gemini-3-flash-preview

Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance. It delivers near Pro level reasoning and tool...

In: textIn: imageIn: videoIn: audioIn: pdfOut: text
Context:1M
In:$0.6/1M
Out:$3.6/1M
Google: Nano Banana Pro (Gemini 3 Pro Image)
gemini-3-pro-image

Nano Banana Pro is Google’s most advanced image-generation and editing model, built on Gemini 3 Pro. It extends the original Nano Banana with significantly improved multimodal reasoning, real-world grounding, and...

In: textIn: imageOut: textOut: image
Context:131.1K
In:$2.4/1M
Out:$14.4/1M
Google: Nano Banana Pro (Gemini 3 Pro Image Preview)
gemini-3-pro-image-preview

Nano Banana Pro is Google’s most advanced image-generation and editing model, built on Gemini 3 Pro. It extends the original Nano Banana with significantly improved multimodal reasoning, real-world grounding, and high-fidelity visual synthesis. The model generates context-rich graphics, from infographics and diagrams to cinematic composites, and can incorporate real-time information via Search grounding. It offers industry-leading text rendering in images (including long passages and multilingual layouts), consistent multi-image blending, and accurate identity preservation across up to five subjects. Nano Banana Pro adds fine-grained creative controls such as localized edits, lighting and focus adjustments, camera transformations, and support for 2K/4K outputs and flexible aspect ratios. It is designed for professional-grade design, product visualization, storyboarding, and complex multi-element compositions while remaining efficient for general image creation workflows.

In: textIn: imageOut: text
Context:1M
In:$2.4/1M
Out:$72/1M
gemini-3-pro-preview
gemini-3-pro-preview

In: textIn: imageOut: text
Context:1M
In:$2.4/1M
Out:$14.4/1M
Google: Nano Banana 2 (Gemini 3.1 Flash Image)
gemini-3.1-flash-image

Gemini 3.1 Flash Image, a.k.a. "Nano Banana 2," is Google’s latest state of the art image generation and editing model, delivering Pro-level visual quality at Flash speed. It combines advanced...

In: imageIn: textOut: textOut: image
Context:131.1K
In:$0.6/1M
Out:$3.6/1M
Google: Nano Banana 2 (Gemini 3.1 Flash Image Preview)
gemini-3.1-flash-image-preview

Gemini 3.1 Flash Image Preview, a.k.a. "Nano Banana 2," is Google’s latest state of the art image generation and editing model, delivering Pro-level visual quality at Flash speed. It combines advanced contextual understanding with fast, cost-efficient inference, making complex image generation and iterative edits significantly more accessible. Aspect ratios can be controlled with the [image_config API Parameter](https://openrouter.ai/docs/features/multimodal/image-generation#image-aspect-ratio-configuration)

In: textIn: imageIn: pdfOut: textOut: image
Context:131.1K
In:$0.6/1M
Out:$72/1M
Google: Gemini 3.1 Flash Lite
gemini-3.1-flash-lite

Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, image, video, audio, and PDF inputs, and is designed for lightweight agentic...

In: textIn: imageIn: videoIn: audioIn: pdfOut: text
Context:1M
In:$0.3/1M
Out:$1.8/1M
Google: Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)
gemini-3.1-flash-lite-image

Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) is Google's fastest, most cost-efficient Gemini image model, built for high-velocity developer pipelines and rapid-fire visual exploration. It delivers text-to-image generation...

In: textIn: imageOut: textOut: image
Context:65.5K
In:$0.3/1M
Out:$1.8/1M
Google: Gemini 3.1 Flash Lite Preview
gemini-3.1-flash-lite-preview

Gemini 3.1 Flash Lite Preview is Google's high-efficiency model optimized for high-volume use cases. It outperforms Gemini 2.5 Flash Lite on overall quality and approaches Gemini 2.5 Flash performance across...

In: textIn: imageIn: videoIn: audioIn: pdfOut: text
Context:1M
In:$0.3/1M
Out:$1.8/1M
Gemini 3.1 Flash TTS Preview
gemini-3.1-flash-tts-preview

In: textOut: audio
Context:8.2K
In:$1.2/1M
Out:$24/1M
Gemini 3.1 Pro Preview
gemini-3.1-pro

In: textIn: imageIn: videoIn: audioIn: pdfOut: text
Context:1M
In:$2.4/1M
Out:$14.4/1M
Google: Gemini 3.1 Pro Preview
gemini-3.1-pro-preview

Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more efficient token usage across complex workflows. Building on the multimodal foundation...

In: textIn: imageIn: videoIn: audioIn: pdfOut: text
Context:1M
In:$2.4/1M
Out:$14.4/1M
Google: Gemini 3.1 Pro Preview Custom Tools
gemini-3.1-pro-preview-customtools

Gemini 3.1 Pro Preview Custom Tools is a variant of Gemini 3.1 Pro that improves tool selection behavior by preventing overuse of a general bash tool when more efficient third-party...

In: textIn: imageIn: videoIn: audioIn: pdfOut: text
Context:1M
In:$2.4/1M
Out:$14.4/1M
Google: Gemini 3.5 Flash
gemini-3.5-flash

Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution...

In: textIn: imageIn: videoIn: audioIn: pdfOut: text
Context:1M
In:$1.8/1M
Out:$10.8/1M
Google: Gemini 3.5 Flash Lite
gemini-3.5-flash-lite

Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows.

In: textIn: imageIn: videoIn: audioIn: pdfOut: text
Context:1M
In:$0.36/1M
Out:$3/1M
Google: Gemini 3.6 Flash
gemini-3.6-flash

Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and...

In: textIn: imageIn: videoIn: audioIn: pdfOut: text
Context:1M
In:$0.9/1M
Out:$4.5/1M
Google: Gemini 3.7 Flash
gemini-3.7-flash

Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step...

In: textIn: imageIn: videoIn: audioIn: pdfOut: text
Context:1M
In:$0.9/1M
Out:$4.5/1M
Google: Gemini 3.8 Flash
gemini-3.8-flash

Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning.

In: textIn: imageIn: videoIn: audioIn: pdfOut: text
Context:1M
In:$0.9/1M
Out:$4.5/1M
Gemini Omni Flash Preview
gemini-omni-flash-preview

In: textIn: imageIn: videoOut: video
Context:131.1K
In:$1.8/1M
Out:$21/1M
Google: Gemma 2 27B
gemma-2-27b-it

Gemma 2 27B by Google is an open model built from the same research and technology used to create the [Gemini models](/models?q=gemini). Gemma models are well-suited for a variety of...

In: textOut: text
Context:8.2K
In:$0.78/1M
Out:$0.78/1M
Gemma 2 9B
gemma-2-9b-it

In: textOut: text
Context:8.2K
In:$0.036/1M
Out:$0.108/1M
Google: Gemma 3 12B
gemma-3-12b-it

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities,...

In: textIn: imageOut: text
Context:131.1K
In:$0.06/1M
Out:$0.18/1M
Google: Gemma 3 27B
gemma-3-27b-it

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities,...

In: textIn: imageOut: text
Context:131.1K
In:$0.096/1M
Out:$0.54/1M
Google: Gemma 3 4B
gemma-3-4b-it

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities,...

In: textIn: imageOut: text
Context:131.1K
In:$0.06/1M
Out:$0.12/1M
Google: Gemma 4 26B A4B
gemma-4-26b-a4b-it

Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at...

In: imageIn: textIn: videoOut: text
Context:262.1K
In:$0.084/1M
Out:$0.408/1M
Google: Gemma 4 31B
gemma-4-31b-it

Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...

In: imageIn: textIn: videoOut: text
Context:262.1K
In:$0.108/1M
Out:$0.408/1M
GLM-4
glm-4

In: textOut: text
Context:128K
In:$17.993/1M
Out:$17.993/1M
Z.ai: GLM 4 32B
glm-4-32b

In: textOut: text
Context:128K
In:$0.12/1M
Out:$0.12/1M
GLM-4 Air
glm-4-air

In: textOut: text
Context:128K
In:$0.241/1M
Out:$0.241/1M
GLM-4 AirX
glm-4-airx

In: textOut: text
Context:8K
In:$2.407/1M
Out:$2.407/1M
GLM-4 Flash
glm-4-flash

In: textOut: text
Context:128K
In:$0.12/1M
Out:$0.12/1M
GLM-4 Long
glm-4-long

In: textOut: text
Context:1M
In:$0.241/1M
Out:$0.241/1M
Z.ai: GLM 4.5
glm-4.5

GLM-4.5 is our latest flagship foundation model, purpose-built for agent-based applications. It leverages a Mixture-of-Experts (MoE) architecture and supports a context length of up to 128k tokens. GLM-4.5 delivers significantly...

In: textOut: text
Context:131.1K
In:$0.72/1M
Out:$2.64/1M
Z.ai: GLM 4.5 Air
glm-4.5-air

GLM-4.5-Air is the lightweight variant of our latest flagship model family, also purpose-built for agent-centric applications. Like GLM-4.5, it adopts the Mixture-of-Experts (MoE) architecture but with a more compact parameter...

In: textOut: text
Context:131.1K
In:$0.156/1M
Out:$1.02/1M
glm-4.5-airx
glm-4.5-airx

In: textOut: text
Context:128K
In:$0.686/1M
Out:$2.057/1M
glm-4.5-x
glm-4.5-x

In: textOut: text
Context:128K
In:$1.372/1M
Out:$2.748/1M
Z.ai: GLM 4.5V
glm-4.5v

GLM-4.5V is a vision-language foundation model for multimodal agent applications. Built on a Mixture-of-Experts (MoE) architecture with 106B parameters and 12B activated parameters, it achieves state-of-the-art results in video understanding,...

In: textIn: imageOut: text
Context:65.5K
In:$0.72/1M
Out:$2.16/1M
Z.ai: GLM 4.6
glm-4.6

Compared with GLM-4.5, this generation brings several key improvements: Longer context window: The context window has been expanded from 128K to 200K tokens, enabling the model to handle more complex...

In: textOut: text
Context:204.8K
In:$0.516/1M
Out:$2.1/1M
Z.ai: GLM 4.6V
glm-4.6v

GLM-4.6V is a large multimodal model designed for high-fidelity visual understanding and long-context reasoning across images, documents, and mixed media. It supports up to 128K tokens, processes complex page layouts...

In: textIn: imageIn: videoOut: text
Context:131.1K
In:$0.36/1M
Out:$1.08/1M
Z.ai: GLM 4.7
glm-4.7

GLM-4.7 is Z.ai’s latest flagship model, featuring upgrades in two key areas: enhanced programming capabilities and more stable multi-step reasoning/execution. It demonstrates significant improvements in executing complex agent tasks while...

In: textOut: text
Context:204.8K
In:$0.48/1M
Out:$2.1/1M
Z.ai: GLM 4.7 Flash
glm-4.7-flash

As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency. It is further optimized for agentic coding use cases, strengthening coding capabilities, long-horizon task planning,...

In: textOut: text
Context:202.8K
In:$0.072/1M
Out:$0.48/1M
Z.ai: GLM 5
glm-5

GLM-5 is Z.ai’s flagship open-source foundation model engineered for complex systems design and long-horizon agent workflows. Built for expert developers, it delivers production-grade performance on large-scale programming tasks, rivaling leading...

In: textOut: text
Context:204.8K
In:$0.72/1M
Out:$2.304/1M
GLM-5.2
glm-5-2

In: textOut: text
Context:1M
In:$1.68/1M
Out:$5.28/1M
GLM 5.3
glm-5-3

In: textOut: text
Context:1M
In:$1.68/1M
Out:$5.28/1M
GLM-5.3 Flash
glm-5-3-flash

In: textIn: imageIn: videoIn: pdfOut: text
Context:1M
In:$0.18/1M
Out:$0.6/1M
Z.ai: GLM 5 Turbo
glm-5-turbo

GLM-5 Turbo is a new model from Z.ai designed for fast inference and strong performance in agent-driven environments such as OpenClaw scenarios. It is deeply optimized for real-world agent workflows...

In: textOut: text
Context:202.8K
In:$1.44/1M
Out:$4.8/1M
Z.ai: GLM 5.1
glm-5.1

GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks. Unlike previous models built around minute-level interactions, GLM-5.1 can work independently and continuously on...

In: textOut: text
Context:204.8K
In:$1.159/1M
Out:$3.643/1M
Z.ai: GLM 5.2
glm-5.2

GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering,...

In: textOut: text
Context:1M
In:$1.159/1M
Out:$3.643/1M
GLM-5.2
glm-5.2-fast

In: textOut: text
Context:1M
In:$2.388/1M
Out:$7.392/1M
Z.ai: GLM 5.3
glm-5.3

GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves...

In: textOut: text
Context:1.3M
In:$1.68/1M
Out:$5.28/1M
Z.ai: GLM 5.3 Flash
glm-5.3-flash

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...

In: textIn: imageIn: videoOut: text
Context:1.3M
In:$0.09/1M
Out:$0.3/1M
Z.ai: GLM 5V Turbo
glm-5v-turbo

GLM-5V-Turbo is Z.ai’s first native multimodal agent foundation model, built for vision-based coding and agent-driven tasks. It natively handles image, video, and text inputs, excels at long-horizon planning, complex coding,...

In: imageIn: textIn: videoOut: text
Context:202.8K
In:$1.44/1M
Out:$4.8/1M
Z.ai: GLM Flash Latest
glm-flash-latest

This model always redirects to the latest model in the GLM Flash family.

In: textIn: imageIn: videoOut: text
Context:1.3M
In:$0.09/1M
Out:$0.3/1M
Z.ai: GLM Latest
glm-latest

This model always redirects to the latest GLM model from Z.ai.

In: textOut: text
Context:1.3M
In:$1.404/1M
Out:$4.752/1M
Goliath 120B
goliath-120b

A large LLM created by combining two fine-tuned Llama 70B models into one 120B model. Combines Xwin and Euryale. Credits to - [@chargoddard](https://huggingface.co/chargoddard) for developing the framework used to merge...

In: textOut: text
Context:6.1K
In:$4.5/1M
Out:$9/1M
OpenAI: GPT-3.5 Turbo
gpt-3.5-turbo

GPT-3.5 Turbo is OpenAI's fastest model. It can understand and generate natural language or code, and is optimized for chat and traditional completion tasks. Training data up to Sep 2021.

In: textOut: text
Context:16.4K
In:$0.6/1M
Out:$1.8/1M
GPT-3.5 Turbo 0301
gpt-3.5-turbo-0301

In: textOut: text
Context:4.1K
In:$1.8/1M
Out:$2.4/1M
OpenAI: GPT-3.5 Turbo (older v0613)
gpt-3.5-turbo-0613

GPT-3.5 Turbo is OpenAI's fastest model. It can understand and generate natural language or code, and is optimized for chat and traditional completion tasks. Training data up to Sep 2021.

In: textOut: text
Context:4.1K
In:$1.2/1M
Out:$2.4/1M
GPT-3.5 Turbo 1106
gpt-3.5-turbo-1106

In: textOut: text
Context:16.4K
In:$1.2/1M
Out:$2.4/1M
OpenAI: GPT-3.5 Turbo 16k
gpt-3.5-turbo-16k

This model offers four times the context length of gpt-3.5-turbo, allowing it to support approximately 20 pages of text in a single request at a higher cost. Training data: up...

In: textOut: text
Context:16.4K
In:$3.6/1M
Out:$4.8/1M
OpenAI: GPT-3.5 Turbo Instruct
gpt-3.5-turbo-instruct

This model is a variant of GPT-3.5 Turbo tuned for instructional prompts and omitting chat-related optimizations. Training data: up to Sep 2021.

In: textOut: text
Context:4.1K
In:$1.8/1M
Out:$2.4/1M
OpenAI: GPT-4
gpt-4

OpenAI's flagship model, GPT-4 is a large-scale multimodal language model capable of solving difficult problems with greater accuracy than previous models due to its broader general knowledge and advanced reasoning...

In: textOut: text
Context:8.2K
In:$36/1M
Out:$72/1M
OpenAI: GPT-4 (older v0314)
gpt-4-0314

In: textOut: text
Context:8.2K
In:$36/1M
Out:$72/1M
OpenAI: GPT-4 Turbo (older v1106)
gpt-4-1106-preview

In: textOut: text
Context:128K
In:$12/1M
Out:$36/1M
GPT-4 32K
gpt-4-32k

In: textOut: text
Context:32.8K
In:$72/1M
Out:$144/1M
OpenAI: GPT-4 Turbo
gpt-4-turbo

The latest GPT-4 Turbo model with vision capabilities. Vision requests can now use JSON mode and function calling. Training data: up to December 2023.

In: textIn: imageOut: text
Context:128K
In:$12/1M
Out:$36/1M
OpenAI: GPT-4 Turbo Preview
gpt-4-turbo-preview

The preview GPT-4 model with improved instruction following, JSON mode, reproducible outputs, parallel function calling, and more. Training data: up to Dec 2023. **Note:** heavily rate limited by OpenAI while...

In: textOut: text
Context:128K
In:$12/1M
Out:$36/1M
OpenAI: GPT-4.1
gpt-4.1

GPT-4.1 is a flagship large language model optimized for advanced instruction following, real-world software engineering, and long-context reasoning. It supports a 1 million token context window and outperforms GPT-4o and...

In: textIn: imageIn: pdfOut: text
Context:1M
In:$2.4/1M
Out:$9.6/1M
OpenAI: GPT-4.1 Mini
gpt-4.1-mini

GPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost. It retains a 1 million token context window and scores 45.1% on hard...

In: textIn: imageIn: pdfOut: text
Context:1M
In:$0.48/1M
Out:$1.92/1M
OpenAI GPT-4.1 Mini
gpt-4.1-mini-2025-04-14

In: textIn: imageOut: text
Context:1M
In:$0.48/1M
Out:$1.92/1M
OpenAI: GPT-4.1 Nano
gpt-4.1-nano

For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance at a small size with its 1 million...

In: imageIn: textIn: pdfOut: text
Context:1M
In:$0.12/1M
Out:$0.48/1M
OpenAI: GPT-4o
gpt-4o

GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. It maintains the intelligence level of [GPT-4 Turbo](/models/openai/gpt-4-turbo) while being twice as...

In: textIn: imageIn: pdfOut: text
Context:128K
In:$3/1M
Out:$12/1M
OpenAI: GPT-4o (2024-05-13)
gpt-4o-2024-05-13

GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. It maintains the intelligence level of [GPT-4 Turbo](/models/openai/gpt-4-turbo) while being twice as...

In: textIn: imageIn: pdfOut: text
Context:128K
In:$6/1M
Out:$18/1M
OpenAI: GPT-4o (2024-08-06)
gpt-4o-2024-08-06

The 2024-08-06 version of GPT-4o offers improved performance in structured outputs, with the ability to supply a JSON schema in the respone_format. Read more [here](https://openai.com/index/introducing-structured-outputs-in-the-api/). GPT-4o ("o" for "omni") is...

In: textIn: imageIn: pdfOut: text
Context:128K
In:$3/1M
Out:$12/1M
OpenAI: GPT-4o (2024-11-20)
gpt-4o-2024-11-20

The 2024-11-20 version of GPT-4o offers a leveled-up creative writing ability with more natural, engaging, and tailored writing to improve relevance & readability. It’s also better at working with uploaded...

In: textIn: imageIn: pdfOut: text
Context:128K
In:$3/1M
Out:$12/1M
OpenAI: GPT-4o Audio
gpt-4o-audio-preview

In: audioIn: textOut: audioOut: text
Context:128K
In:$3/1M
Out:$12/1M
OpenAI: GPT-4o-mini
gpt-4o-mini

GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs. As their most advanced small model, it is many multiples more affordable...

In: textIn: imageIn: pdfOut: text
Context:128K
In:$0.18/1M
Out:$0.72/1M
OpenAI: GPT-4o-mini (2024-07-18)
gpt-4o-mini-2024-07-18

GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs. As their most advanced small model, it is many multiples more affordable...

In: textIn: imageIn: pdfOut: text
Context:128K
In:$0.18/1M
Out:$0.72/1M
GPT 4o Mini Search Preview
gpt-4o-mini-search-preview

In: textOut: text
Context:128K
In:$0.18/1M
Out:$0.72/1M
GPT-4o Mini Transcribe
gpt-4o-mini-transcribe

In: textIn: audioOut: text
Context:16K
In:$1.5/1M
Out:$6/1M
OpenAI: GPT-4o Search Preview
gpt-4o-search-preview

In: textOut: text
Context:128K
In:$3/1M
Out:$12/1M
GPT-4o Transcribe
gpt-4o-transcribe

In: textIn: audioOut: text
Context:16K
In:$3/1M
Out:$12/1M
OpenAI: GPT-5
gpt-5

GPT-5 is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex tasks that require step-by-step reasoning, instruction following, and accuracy...

In: textIn: imageIn: pdfOut: text
Context:400K
In:$1.5/1M
Out:$12/1M
GPT-5.1
gpt-5-1

In: textIn: imageOut: textOut: image
Context:400K
In:$1.5/1M
Out:$12/1M
GPT-5.1 Codex Max
gpt-5-1-codex-max

In: textIn: imageOut: textOut: image
Context:400K
In:$1.5/1M
Out:$12/1M
GPT-5.1 Codex mini
gpt-5-1-codex-mini

In: textIn: imageOut: textOut: image
Context:400K
In:$0.3/1M
Out:$2.4/1M
GPT-5.2
gpt-5-2

In: textIn: imageOut: textOut: image
Context:400K
In:$2.1/1M
Out:$16.8/1M
GPT-5.2 Codex
gpt-5-2-codex

In: textIn: imageIn: pdfOut: textOut: image
Context:400K
In:$2.1/1M
Out:$16.8/1M
GPT-5.3 Codex
gpt-5-3-codex

In: textIn: imageIn: pdfOut: textOut: image
Context:400K
In:$2.1/1M
Out:$16.8/1M
GPT-5.4
gpt-5-4

In: textIn: imageIn: pdfOut: textOut: image
Context:1.1M
In:$3/1M
Out:$18/1M
GPT-5.4 mini
gpt-5-4-mini