AI Models
Configure available AI models and providers for your organization
OpenAI
GPT-6 Astra
gpt-6-astra
OpenAI's most intelligent and aligned model. State-of-the-art on computer use, browsing, software engineering, science and professional work. Saturates FrontierMath Tier 4 (98%) and ARC-AGI-3 (99.9%).
GPT-5.5 Pro
gpt-5-5-pro
Maximum-compute GPT-5.5 for the hardest reasoning tasks.
o5
o5-latest
Deep deliberate reasoning model for math, science and multi-step planning.
GPT-5.2
gpt-5-2-latest
OpenAI's flagship frontier model with adaptive reasoning and native agentic tool use.
GPT-5.6 Sol
gpt-5-6-sol
Flagship GPT-5.6 tier for frontier coding, reasoning and agentic work.
GPT-5.4 Pro
gpt-5-4-pro
Maximum-compute GPT-5.4 reasoning model.
GPT-5.5
gpt-5-5
Previous flagship generation; strong coding, reasoning and agentic performance.
GPT-5 pro
gpt-5-pro-latest
GPT-5 Pro is OpenAI's most advanced AI model, optimized for complex tasks requiring deep reasoning and precise coding capabilities. It features a 400,000-token context window, supports multimodal inputs, and offers high accuracy in various benchmarks.
o1 Preview
o1-preview-old
OpenAI's o1-preview is a cutting-edge AI model optimized for complex reasoning tasks, particularly in STEM fields, offering high performance and cost efficiency.
GPT-5 Thinking
gpt-5-thinking-latest
GPT-5 is OpenAI's advanced AI model designed for complex coding, reasoning, and multimodal tasks, offering a large context window and high accuracy across various benchmarks.
GPT-5.6 Terra
gpt-5-6-terra
Production default of the GPT-5.6 family — strong capability at mid-tier cost.
GPT-5.4
gpt-5-4
Balanced GPT-5.4 model with 1M context.
GPT-5.3 Codex
gpt-5-3-codex
Agentic coding specialist optimized for the Codex harness.
GPT-5
gpt-5-latest
GPT-5 is OpenAI's advanced AI model designed for complex reasoning, coding, and multimodal tasks, offering a large context window and competitive pricing.
GPT-5.6 Luna
gpt-5-6-luna
High-volume, cost-sensitive GPT-5.6 model for classification, extraction and routing.
gpt-oss-120b
gpt-oss-120b-120b-latest
gpt-oss-120b is OpenAI's most powerful open-weight model, designed for complex reasoning and coding tasks, offering high performance at a cost-effective price point.
GPT-5.2 Mini
gpt-5-2-mini-latest
Fast, low-cost workhorse for high-volume production workloads.
GPT 4o
gpt-4o-latest
GPT-4o is OpenAI's advanced AI model released in May 2024, excelling in voice, multilingual, and vision tasks, with a context window of 128,000 tokens and a knowledge cutoff in October 2023.
GPT 4o
gpt-4o-old
GPT-4o is OpenAI's advanced AI model released in May 2024, excelling in voice, multilingual, and vision tasks, with a context window of 128,000 tokens and a knowledge cutoff in October 2023.
GPT-5.4 Mini
gpt-5-4-mini
Fast, affordable GPT-5.4 mini model.
GPT-5 nano
gpt-5-nano-latest
GPT-5 nano is OpenAI's most compact and cost-effective variant in the GPT-5 family, designed for efficient processing of text and image inputs, making it ideal for applications where speed and affordability are paramount.
GPT 4.1 mini
gpt-4-1-mini-latest
GPT-4.1 mini is a mid-sized AI model by OpenAI, offering strong performance in coding and instruction following tasks, with a 1 million token context window and a knowledge cutoff in June 2024.
gpt-oss-20b
gpt-oss-20b-20b-latest
gpt-oss-20b is OpenAI's open-weight language model with 21 billion parameters, designed for efficient deployment on consumer hardware with just 16GB memory. Released under Apache 2.0 license, it demonstrates strong performance across benchmarks including 85.3% MMLU and exceptional AIME scores.
o3 mini high
o3-mini-high-latest
OpenAI's o3-mini is a cost-effective, high-speed reasoning model optimized for STEM tasks, offering advanced capabilities in coding, math, and science.
GPT-5 mini
gpt-5-mini-latest
GPT-5 mini is a cost-effective, high-performance AI model by OpenAI, designed for well-defined tasks requiring efficient processing and strong capabilities in coding and reasoning.
o1
o1-latest
OpenAI's o1 model is a cutting-edge AI designed for complex reasoning and problem-solving tasks, excelling in STEM-related applications.
o3 mini medium
o3-mini-medium-latest
OpenAI's o3-mini is a cost-efficient reasoning model optimized for STEM tasks, particularly excelling in science, math, and coding. It offers advanced reasoning capabilities with a 200,000-token context window and a knowledge cutoff in October 2023.
o1 Pro
o1-pro-latest
OpenAI's o1 Pro is a high-performance AI model designed for advanced reasoning and complex problem-solving tasks, offering a large context window and high accuracy in various benchmarks.
GPT 4.5
gpt-4-5-preview
GPT-4.5 is OpenAI's advanced language model, offering enhanced reasoning, coding capabilities, and a 128,000-token context window, suitable for complex tasks but at a higher cost.
o3 mini low
o3-mini-low-latest
OpenAI's o3-mini is a cost-efficient reasoning model optimized for STEM tasks, particularly excelling in science, math, and coding. It offers high intelligence at the same cost and latency targets as o1-mini, supporting key developer features like function calling, structured outputs, and batch API. Released on January 31, 2025, with a knowledge cutoff in October 2023, o3-mini provides a context window of 200,000 tokens and a maximum output of 100,000 tokens. It is priced at $1.10 per million input tokens and $4.40 per million output tokens. Benchmark performance includes 97.9% on MATH, 86.9% on MMLU, and 79.7% on GPQA. While it does not support vision capabilities, o3-mini is ideal for tasks requiring high intelligence at lower cost, such as landing page generation, policy analysis, text-to-SQL conversion, and graph entity extraction. It is available through OpenAI's API and supports text input and output.
GPT 4.1
gpt-4-1-latest
GPT-4.1 is a versatile AI model by OpenAI, excelling in instruction following, tool calling, and handling long-context tasks. It offers a 1 million token context window and low latency without a reasoning step.
GPT-5.4 Nano
gpt-5-4-nano
Lowest-cost GPT-5.4 model for classification and simple tasks.
GPT 4o mini
gpt-4o-mini-latest
GPT-4o mini is OpenAI's most cost-efficient small model, offering strong performance in coding and reasoning tasks, with a context window of 128,000 tokens and support for text and image inputs.
GPT 4.1 nano
gpt-4-1-nano-latest
GPT-4.1 nano is OpenAI's fastest and most cost-efficient model, designed for tasks requiring low latency, such as classification and autocompletion. It features a 1 million token context window and supports both text and image inputs.
GPT 4o
gpt-4o-old-2
GPT-4o is OpenAI's general-purpose language model designed for a wide range of language processing tasks, offering a balance between performance and cost efficiency.
o1 mini
o1-mini-latest
o1-mini is a cost-efficient reasoning model optimized for STEM tasks, offering strong performance in coding and math while maintaining affordability.
GPT 4o
gpt-4o-old-3
GPT-4o is a versatile AI model by OpenAI, known for its large context window and support for image inputs, making it suitable for tasks involving extensive text and visual data.
Anthropic
Claude Opus 4.6
claude-opus-4-6-latest
Anthropic's most capable model — best-in-class coding and long-horizon agent tasks.
Claude Sonnet 4.6
claude-sonnet-4-6-latest
The best balance of intelligence, speed and cost for enterprise workloads.
Claude Sonnet 4.5
claude-sonnet-4-5-latest
Claude Sonnet 4.5 is Anthropic's most advanced AI model, excelling in complex coding, agent-based workflows, and long-duration tasks.
Claude Opus 4.1
claude-opus-4-1-latest
Claude Opus 4.1 is Anthropic's advanced AI model, excelling in coding and reasoning tasks, with a 200K token context window and a knowledge cutoff in March 2025.
Claude 3 Opus
claude-3-opus-latest
Claude 3 Opus is Anthropic's most intelligent AI model, offering exceptional performance across a wide range of cognitive tasks, including advanced reasoning, general knowledge, and coding. It features a 200,000-token context window and supports both text and image inputs. Released on March 4, 2024, it has a knowledge cutoff in August 2023. While it is priced at $15 per million input tokens and $75 per million output tokens, its high accuracy and reliability make it a valuable tool for complex applications.
Claude 3.5 Sonnet
claude-3-5-sonnet-old
Claude 3.5 Sonnet is a large language model developed by Anthropic, known for its strong coding capabilities and extended context window, making it suitable for complex coding tasks and scenarios requiring nuanced understanding.
Claude 3.5 Sonnet
claude-3-5-sonnet-latest
Claude 3.5 Sonnet is a large language model developed by Anthropic, known for its strong coding capabilities and extended context window.
Claude 3.7 Sonnet
claude-3-7-sonnet-latest
Claude 3.7 Sonnet is a high-performance AI model developed by Anthropic, offering advanced capabilities in coding, reasoning, and extended thinking.
Claude 3.5 Haiku
claude-3-5-haiku-latest
Claude 3.5 Haiku is a language model developed by Anthropic, designed for high-throughput, simple tasks. It offers a balance between performance and cost, making it suitable for applications requiring rapid responses and efficient processing.
Claude 3 Haiku
claude-3-haiku-old
Claude 3 Haiku is Anthropic's fastest and most affordable model in its intelligence class, designed for enterprise applications requiring rapid analysis of large datasets.
Google AI
Gemini 3.1 Pro
gemini-3-1-pro-latest
Google's frontier multimodal model with 2M context and native video understanding.
Gemma 3
gemma-3-12b-latest
Gemma 3 is a versatile AI model developed by Google, designed for efficient processing of extensive documents and multimodal data, offering support for over 140 languages and advanced reasoning capabilities.
Gemma 3
gemma-3-4b-latest
Gemma 3 is a state-of-the-art AI model developed by Google, designed for efficient deployment on single GPUs or TPUs. It offers advanced text and visual reasoning capabilities, supports over 140 languages, and features a 128,000-token context window, enabling the processing of extensive documents.
Gemini 1.5 Pro 001
gemini-1-5-pro-001-old
Gemini 1.5 Pro is Google's advanced AI model designed for complex reasoning tasks, offering a 2 million token context window and support for various input modalities, including text, images, audio, and video.
Gemini 1.5 Pro 002
gemini-1-5-pro-002-old
Gemini 1.5 Pro 002 is a high-performance AI model developed by Google, designed for complex reasoning tasks and large-scale data analysis. It offers a 2 million token context window, supports multimodal inputs (text, images, audio, video), and demonstrates strong performance across various benchmarks. Released in September 2024, it provides cost-effective pricing for both input and output tokens.
Gemini Exp 1206
gemini-exp-1206-old
Gemini Exp 1206 is a cutting-edge AI model developed by Google, released on December 17, 2025. It boasts a substantial context window of 2,097,152 tokens, enabling it to process and understand extensive inputs effectively. The model demonstrates strong performance across various benchmarks, including a 90.04% accuracy on MMLU, 94.4% performance on GSM8K, and a 74.4% success rate on Python coding tasks (HumanEval).
Gemini 3 Flash
gemini-3-flash-latest
Ultra-fast multimodal model with 1M context — the volume default.
Gemma 3
gemma-3-27b-latest
Gemma 3 is a state-of-the-art AI model developed by Google, designed for efficient processing of extensive documents and multimodal data. It offers a large context window, supports over 140 languages, and is optimized for single-GPU deployment.
Gemini 2.0 Flash Exp
gemini-2-0-flash-exp-old
Gemini 2.0 Flash is Google's next-generation AI model designed for high-speed, cost-effective processing across multiple modalities, including text, images, video, and audio, with a 1 million token context window.
Gemini 2.0 Flash
gemini-2-0-flash-latest
Gemini 2.0 Flash is Google's advanced AI model designed for high-speed, cost-effective, and reliable performance across various tasks, including coding and reasoning.
Gemini 1.5 Flash 002
gemini-1-5-flash-002-old
Gemini 1.5 Flash 002 is a high-performance, cost-effective AI model developed by Google, designed for applications requiring large context processing and multimodal input handling.
Gemini 2.5 Flash Lite Thinking
gemini-2-5-flash-lite-thinking-latest
Gemini 2.5 Flash-Lite is Google's fastest and most cost-efficient AI model, optimized for high-throughput applications requiring adaptive thinking.
Gemini 2.5 Flash Lite
gemini-2-5-flash-lite-latest
Gemini 2.5 Flash-Lite is Google's fastest and most cost-efficient AI model, optimized for high-throughput tasks requiring rapid responses.
Gemini 2.0 Flash-Lite
gemini-2-0-flash-lite-latest
Gemini 2.0 Flash-Lite is a lightweight, cost-efficient AI model by Google, optimized for low-latency, high-throughput tasks, supporting text, image, audio, and video inputs with a 1 million token context window and a knowledge cutoff in August 2024.
Gemini 2.0 Pro
gemini-2-0-pro-preview
Gemini 2.0 Pro is Google's advanced AI model, released on December 11, 2024, offering enhanced capabilities in coding performance and complex prompt handling. It supports multimodal inputs and outputs, including text, images, video, and audio, with a context window of 2 million input tokens and 8,192 output tokens. The model has a knowledge cutoff in August 2024 and is available through Google AI Studio and Vertex AI. Pricing is approximately $0.10 per million input tokens and $0.40 per million output tokens.
Gemini 2.5 Deep Think
gemini-2-5-deep-think-latest
Gemini 2.5 Deep Think is Google's advanced AI model, offering a massive 1 million-token context window and robust multimodal capabilities, excelling in complex reasoning and coding tasks.
Gemini 2.5 Pro
gemini-2-5-pro-old-2
Gemini 2.5 Pro is Google's state-of-the-art AI model, excelling in complex reasoning and coding tasks with a vast context window and multimodal capabilities.
Gemini 2.5 Flash
gemini-2-5-flash-preview
Gemini 2.5 Flash is Google's most efficient AI model, designed for speed and low-cost, supporting text, images, video, and audio inputs with a 1 million token context window and adaptive thinking capabilities.
Gemini 2 Flash Thinking
gemini-2-flash-thinking-old
Gemini 2.5 Flash is Google's advanced AI model designed for large-scale, low-latency tasks that require reasoning and agentic capabilities. It supports multimodal inputs, including text, images, video, and audio, with a context window of 1 million tokens and a knowledge cutoff in January 2025. The model offers a cost-effective pricing structure, with input costs at $0.30 per million tokens and output costs at $2.50 per million tokens. Its strengths include high performance, adaptive thinking capabilities, and cost efficiency, making it suitable for applications requiring complex reasoning and multimodal processing.
Gemini Exp 1114
gemini-exp-1114-old
Gemini 3 Pro is Google's advanced AI model, released in November 2025, known for its large context window and multimodal processing capabilities.
Gemini Exp 1121
gemini-exp-1121-old
Google's Gemini Exp 1121 is an AI model released on November 21, 2024, featuring a 1 million-token context window and a knowledge cutoff in August 2024. It excels in reasoning tasks with a 59% average score on LiveBench, and offers cost-effective pricing at $1.25 per million input tokens and $10 per million output tokens. Its primary strength lies in handling extensive context for research and multimodal applications.
Gemini 2 Flash Thinking
gemini-2-flash-thinking-preview
Gemini 2.5 Flash is Google's advanced AI model designed for high-volume, low-latency tasks that require reasoning and multimodal processing. It offers a 1 million-token context window, supports text, images, video, and audio inputs, and provides text outputs. The model is optimized for cost-effectiveness and speed, making it suitable for applications needing rapid and reliable AI responses.
Gemini 1.5 Flash
gemini-1-5-flash-8b-old
Gemini 1.5 Flash is a fast and versatile multimodal AI model by Google, optimized for high-volume, real-time tasks, supporting text, images, audio, and video inputs with a large context window of 1,048,576 tokens.
Gemini 1.5 Flash 001
gemini-1-5-flash-001-old
Gemini 1.5 Flash is a fast and versatile multimodal AI model by Google, optimized for diverse tasks with a 1 million token context window and support for text, images, audio, and video inputs.
Gemma 2
gemma-2-2b-latest
Gemma 2 9B is a 9-billion parameter language model developed by Google, released on June 28, 2024. It offers a context window of 8,192 tokens and is priced at $0.03 per million input tokens and $0.09 per million output tokens. The model's knowledge cutoff is April 1, 2024. It excels in reasoning tasks and is suitable for educational item generation, adaptation to low-resource languages, and long-context deployments. As an open-source model, it provides flexibility for various applications.
Gemma 2
gemma-2-9b-latest
Gemma 2 is Google's efficient second-generation open-source language model, available in 9B and 27B parameter versions, designed for tasks requiring strong reasoning capabilities.
Gemma 2
gemma-2-27b-latest
Gemma 2 9B is a 9-billion parameter language model developed by Google, designed for efficient language understanding and generation tasks.
alibaba
Qwen 3 A22B Thinking
qwen-3-a22b-thinking-235b-latest
Qwen3-235B-A22B-Thinking-2507 is a large-scale AI model developed by Alibaba, excelling in complex reasoning and coding tasks with a maximum context window of 387,100 tokens.
Qwen 3
qwen-3-32b-latest
Qwen 3 is Alibaba's advanced AI model series, excelling in complex reasoning, coding tasks, and multilingual support.
Qwen 3
qwen-3-4b-latest
Qwen 3 is Alibaba's advanced AI model series, excelling in complex reasoning, coding, and general knowledge tasks.
Qwen 3.5 Max
qwen-3-5-max-latest
Alibaba's top open model — exceptional multilingual and coding ability.
Qwen 2.5 VL
qwen-2-5-vl-32b-latest
Qwen 2.5 VL is a vision-language AI model developed by Alibaba, designed for complex reasoning tasks, coding assistance, and visual processing.
Qwen 2.5 Coder
qwen-2-5-coder-0-5b-latest
Qwen 2.5 Coder is a high-performance AI model developed by Alibaba, optimized for coding and reasoning tasks across multiple programming languages. It offers a large context window, competitive benchmark scores, and cost-effective usage, making it suitable for various software development applications.
Qwen 2.5
qwen-2-5-7b-latest
Qwen 2.5 is an advanced AI language model developed by Alibaba, renowned for its exceptional long-context processing capabilities and competitive performance across various benchmarks.
Qwen 2.5 Coder
qwen-2-5-coder-32b-latest
Qwen 2.5 Coder is an advanced AI model developed by Alibaba, designed to excel in code generation, debugging, and understanding across multiple programming languages. It offers a large context window, high performance in coding benchmarks, and cost-effective usage, making it suitable for various software development tasks.
Qwen 2.5 Coder
qwen-2-5-coder-14b-latest
Qwen 2.5 Coder is an advanced AI model developed by Alibaba, designed to excel in code generation, debugging, and understanding across multiple programming languages. It offers a large context window, high performance in coding benchmarks, and cost-effective usage, making it suitable for various software development tasks.
Qwen 2.5 Coder
qwen-2-5-coder-7b-latest
Qwen 2.5 Coder is an advanced AI model developed by Alibaba, designed to excel in code generation, debugging, and understanding complex codebases. It offers a substantial context window of 1 million tokens, enabling it to process extensive code and documentation efficiently. With competitive benchmark scores and cost-effective usage, it serves as a valuable tool for developers and educators alike.
QWQ Max
qwq-max-preview
Qwen3-Max is Alibaba's flagship AI model, featuring over 1 trillion parameters and a 262,144-token context window, excelling in advanced reasoning and instruction-following tasks.
Qwen 2.5
qwen-2-5-14b-latest
Qwen 2.5 is a state-of-the-art AI language model developed by Alibaba, renowned for its exceptional performance in handling long-context tasks, coding, and reasoning. It offers cost-effective solutions for processing extensive documents and codebases.
Qwen 2.5 Coder
qwen-2-5-coder-3b-latest
Qwen 2.5 Coder is an advanced AI model developed by Alibaba, designed to excel in coding and reasoning tasks across multiple programming languages. It offers a large context window, high benchmark scores, and cost-effective usage, making it suitable for various software development applications.
Qwen 3 A22B
qwen-3-a22b-235b-old
Qwen 3 A22B is a large-scale AI model developed by Alibaba, excelling in advanced reasoning and coding tasks, suitable for enterprise applications and complex problem-solving scenarios.
Qwen 3 Coder
qwen-3-coder-480b-latest
Qwen 3 Coder is a state-of-the-art AI model developed by Alibaba, designed for advanced coding and reasoning tasks with a vast context window and multilingual support.
Qwen3 A3B Thinking
qwen3-a3b-thinking-30b-latest
Qwen3-Next-80B-A3B-Thinking is a reasoning-enhanced variant of Alibaba's Qwen3-Next architecture, designed for complex problem-solving with mixture-of-experts efficiency.
Qwen 2.5
qwen-2-5-32b-latest
Qwen 2.5 is an advanced AI language model developed by Alibaba, renowned for its exceptional long-context processing capabilities and competitive performance across various benchmarks.
Qwen 2.5
qwen-2-5-0-5b-latest
Qwen 2.5 is an advanced AI language model developed by Alibaba, renowned for its exceptional long-context processing capabilities and competitive performance across various benchmarks.
Qwen 2.5 Turbo
qwen-2-5-turbo-latest
Qwen2.5-Turbo is an AI language model developed by Alibaba, designed to handle extensive contexts efficiently and affordably, excelling in both coding and reasoning tasks.
Qwen 2.5
qwen-2-5-3b-latest
Qwen 2.5 is an advanced AI language model developed by Alibaba, renowned for its exceptional long-context processing capabilities and strong performance in coding and reasoning tasks.
Qwen 2.5 Coder
qwen-2-5-coder-1-5b-latest
Qwen 2.5 Coder is a versatile AI model developed by Alibaba, designed to excel in coding and reasoning tasks across multiple programming languages, with support for long-context processing up to 1 million tokens.
Qwen 2.5
qwen-2-5-1-5b-latest
Qwen 2.5 is an advanced AI language model developed by Alibaba, renowned for its exceptional long-context processing capabilities and competitive performance across various benchmarks.
QwQ
qwq-32b-latest
QwQ is an open-source AI model developed by Alibaba, designed to enhance reasoning and coding capabilities. It boasts a large context window and competitive performance across various benchmarks.
Qwen 3 A22B
qwen-3-a22b-235b-latest
Qwen 3 A22B is a large-scale AI model developed by Alibaba, excelling in complex reasoning and coding tasks with a substantial context window and competitive pricing.
Qwen 3 A3B
qwen-3-a3b-30b-old
Qwen 3 A3B is a reasoning-enhanced AI model by Alibaba, designed for complex problem-solving and efficient long-context processing.
Qwen 2.5
qwen-2-5-72b-latest
Qwen 2.5 is an advanced AI language model developed by Alibaba, renowned for its exceptional long-context processing capabilities and cost efficiency.
Qwen 3 A3B
qwen-3-a3b-30b-latest
Qwen 3 A3B is a reasoning-enhanced variant of Alibaba's Qwen 3 series, designed for complex problem-solving with mixture-of-experts efficiency.
Qwen 2.5 Max
qwen-2-5-max-latest
Qwen 2.5 Max is a versatile AI model developed by Alibaba, excelling in tasks that demand extensive context understanding, coding assistance, and general knowledge retrieval. It offers a substantial context window of 131,072 tokens, enabling it to process and comprehend lengthy documents effectively. The model demonstrates strong performance in various benchmarks, including a score of 89.4 in the Arena-Hard preference benchmark, surpassing competitors like DeepSeek V3 (85.5) and Claude 3.5 Sonnet (85.2). In coding tasks, it achieves a HumanEval score of 73.2, indicating its proficiency in code generation and problem-solving. Despite its capabilities, users should be aware of its cost structure, with input tokens priced at $1.6 per million and output tokens at $6.4 per million, which may impact cost-effectiveness for extensive usage. Overall, Qwen 2.5 Max is well-suited for applications requiring deep contextual understanding and complex reasoning, such as legal document analysis, technical content generation, and advanced coding assistance.
QwQ Preview
qwq-preview-32b-preview
QwQ-32B-Preview is an experimental AI model developed by Alibaba's Qwen team in 2024, focusing on enhancing AI reasoning capabilities, particularly in mathematics and programming.
deepseek
DeepSeek Coder 2
deepseek-coder-2-236b-latest
DeepSeek Coder 2 is an advanced AI model specializing in coding and reasoning tasks, offering high performance and cost efficiency.
DeepSeek R2
deepseek-r2-latest
Open reasoning model rivaling proprietary frontier on math and code.
DeepSeek V4
deepseek-v4-latest
Open-weights frontier MoE — near-flagship quality at a fraction of the cost.
DeepSeek V3.1
deepseek-v3-1-840b-latest
DeepSeek V3.1 is an advanced open-source AI model developed by DeepSeek, known for its strong performance in coding and reasoning tasks, large context window, and cost-effective operation.
DeepSeek 2.5
deepseek-2-5-236b-old
DeepSeek 2.5 is an AI model developed by DeepSeek, known for its strong performance in coding and reasoning tasks, cost efficiency, and extended context window capabilities.
Deepseek 2.5
deepseek-2-5-old
DeepSeek 2.5 is an advanced AI model developed by DeepSeek, known for its cost-effectiveness and strong performance in coding and reasoning tasks. It offers a context window of 128,000 tokens, enabling efficient processing of extensive inputs. The model is open-source, allowing for customization and integration into various applications.
R1 Distill Qwen
r1-distill-qwen-1-5b-latest
DeepSeek's R1 Distill Qwen is a distilled language model designed to enhance reasoning and coding capabilities while maintaining efficiency.
R1 Distill Qwen
r1-distill-qwen-7b-latest
DeepSeek's R1 Distill Qwen is a distilled language model designed to enhance reasoning and coding capabilities while maintaining cost efficiency.
R1 Distill Llama
r1-distill-llama-8b-latest
R1 Distill Llama 70B is a distilled large language model by DeepSeek, offering advanced reasoning and coding capabilities at a competitive cost.
DeepSeek V3.1 Reasoner
deepseek-v3-1-reasoner-840b-latest
DeepSeek V3.1 Reasoner is an advanced AI model by DeepSeek, featuring a hybrid reasoning architecture that integrates chat, reasoning, and coding capabilities into a unified model. It offers a 128K token context window, a knowledge cutoff in July 2025, and competitive pricing, making it suitable for complex problem-solving and large-scale document processing tasks.
DeepSeek R1
deepseek-r1-671b-latest
DeepSeek R1 is an advanced AI model developed by DeepSeek, known for its strong performance in reasoning and coding tasks, cost efficiency, and long-context processing capabilities.
DeepSeek V3
deepseek-v3-685b-latest
DeepSeek V3 is an advanced AI model developed by DeepSeek, known for its exceptional coding and reasoning capabilities, large context window, and cost-effective pricing.
R1 Distill Llama
r1-distill-llama-70b-latest
R1 Distill Llama 70B is a distilled large language model by DeepSeek, offering advanced reasoning and coding capabilities with a cost-effective pricing structure.
Deepseek V3
deepseek-v3-671b-old
DeepSeek V3 is an advanced AI model developed by DeepSeek, renowned for its exceptional coding and reasoning capabilities, extensive context window, and cost-effective performance.
R1 Distill Qwen
r1-distill-qwen-32b-latest
DeepSeek's R1 Distill Qwen is a distilled language model designed to enhance reasoning and coding capabilities while maintaining cost efficiency.
R1 Lite Preview
r1-lite-preview-preview
DeepSeek R1 is a large language model designed for advanced reasoning and coding tasks, offering a 128,000-token context window and competitive pricing.
R1 Distill Qwen
r1-distill-qwen-14b-latest
DeepSeek's R1 Distill Qwen is a series of distilled language models designed to enhance reasoning and coding capabilities while maintaining cost efficiency.
Meta
Llama 4 Behemoth
llama-4-behemoth-2t-preview
Llama 4 Behemoth is Meta's most powerful AI model, featuring 288 billion active parameters and a context window of up to 10 million tokens, excelling in multimodal reasoning and coding tasks.
Llama 4 Behemoth
llama-4-behemoth-latest
Meta's largest open-weights teacher model for distillation and research.
Llama 3.3
llama-3-3-70b-latest
Meta's Llama 3.3 is a 70-billion-parameter multilingual large language model optimized for dialogue, reasoning, and coding tasks, offering high performance at a cost-effective price point.
Llama 3.1
llama-3-1-70b-old
Llama 3.1 is Meta's latest open-weight large language model, offering enhanced performance, a vast context window, and cost-effective usage.
Llama 3.1
llama-3-1-405b-latest
Llama 3.1 is Meta's latest large language model, offering enhanced performance and efficiency across various tasks.
Llama 4 Maverick
llama-4-maverick-400b-latest
Llama 4 Maverick is Meta's advanced multimodal AI model, excelling in coding, reasoning, and vision tasks, offering high efficiency and a large context window.
Llama 3.1
llama-3-1-8b-latest
Meta's Llama 3.1 is a state-of-the-art language model known for its extensive context window and competitive performance across various benchmarks.
Llama 3.2
llama-3-2-3b-latest
Llama 3.2 is Meta's latest AI model, offering enhanced capabilities in natural language processing tasks.
Llama 3.2
llama-3-2-1b-latest
Llama 3.2 is Meta's latest AI model, offering enhanced capabilities in natural language processing tasks.
Llama 4 Scout
llama-4-scout-109b-latest
Llama 4 Scout is Meta's efficient multimodal AI model, featuring a 10M token context window and native support for text and image inputs.
Llama 3.2 (Vision)
llama-3-2-vision-11b-latest
Llama 3.2 (Vision) is a 90-billion-parameter multimodal AI model developed by Meta, designed for advanced visual reasoning and language tasks, including image captioning, visual question answering, and complex image-text comprehension.
Llama 3.2 (Vision)
llama-3-2-vision-90b-latest
Llama 3.2 (Vision) is a 90-billion-parameter multimodal AI model developed by Meta, designed for advanced visual reasoning and language tasks, including image captioning, visual question answering, and complex image-text comprehension.
xAI
Grok 4.1
grok-4-1-latest
xAI's frontier model with real-time knowledge and strong agentic reasoning.
Grok 3 mini
grok-3-mini-latest
Grok 3 Mini is a lightweight AI model developed by xAI, optimized for efficient performance in logic-based tasks and coding applications, offering transparent reasoning traces and cost-effective usage.
Grok 3 mini (Think)
grok-3-mini-think-preview
Grok 3 Mini is a cost-efficient AI model by xAI, optimized for advanced reasoning and problem-solving tasks, particularly in STEM fields.
Grok 3
grok-3-latest
Grok 3 is xAI's third-generation AI model, designed to excel in reasoning, coding, and processing large-scale inputs.
Grok 3 (Think)
grok-3-think-preview
Grok 3 is xAI's advanced AI model, known for its exceptional reasoning and coding capabilities, large context window, and cost-effective pricing.
Grok 2
grok-2-latest
Grok-2 is xAI's second-generation language model, offering enhanced reasoning, coding, and vision capabilities, outperforming models like GPT-4-Turbo and Claude 3.5 in various benchmarks.
Grok-2 mini
grok-2-mini-latest
Grok-2 mini is xAI's efficient language model designed for speed without sacrificing quality, offering strong performance in reasoning, coding, and conversational tasks.
Grok 4 Heavy
grok-4-heavy-latest
Grok 4 Heavy by xAI is a high-performance AI model designed for complex reasoning and coding tasks, offering a 256,000-token context window and competitive pricing.
Mistral
Mistral Large 3
mistral-large-3-latest
European frontier model — GDPR-native, strong function calling.
Codestral 25.01
codestral-25-01-latest
Codestral 25.01 is Mistral AI's advanced code generation model, offering enhanced performance and efficiency for software development tasks.
Devstral Medium
devstral-medium-latest
Devstral Medium is a high-performance language model developed by Mistral AI, specializing in code generation and agentic reasoning tasks. Released on July 10, 2025, it offers a 131,072-token context window, enabling efficient processing of large codebases. The model is accessible via API and supports enterprise deployment with optional fine-tuning capabilities.
Devstral Small 1.1
devstral-small-1-1-24b-latest
Devstral Small 1.1 is a 24-billion parameter language model developed by Mistral AI, optimized for software engineering tasks and agentic coding workflows.
Mistral Large
mistral-large-123b-old
Mistral Large is a high-performance AI model developed by Mistral, designed for enterprise applications requiring extensive context processing and multilingual support.
Mistral Large
mistral-large-123b-latest
Mistral Large is a 123-billion-parameter language model developed by Mistral AI, designed for enterprise applications requiring extensive multilingual support and long-context processing. It features a 32,000-token context window and is priced at $2.00 per million input tokens and $6.00 per million output tokens. Released on February 26, 2024, it has achieved an MMLU score of 81.2% in a 5-shot scenario and a HellaSwag score of 89.2% in a 10-shot scenario. The model is open-source, allowing for customization and fine-tuning to specific tasks. Its strengths include strong coding capabilities, multilingual support, and efficient token usage, making it suitable for enterprise deployments requiring extensive multilingual support and long-context processing.
Mistral Medium 3
mistral-medium-3-latest
Mistral Medium 3 is a high-performance, cost-effective AI model designed for enterprise applications, offering strong coding and reasoning capabilities with multimodal processing support.
Mistral Small 3
mistral-small-3-24b-latest
Mistral Small 3 is a 24-billion parameter language model optimized for speed and efficiency, featuring a 32,000-token context window and competitive pricing.
Pixtral Large
pixtral-large-123b-latest
Pixtral Large is a 124-billion parameter, open-weight, multimodal AI model developed by Mistral, designed to process and understand both text and images. Released on November 18, 2024, it offers a substantial context window of 131,072 tokens, enabling efficient handling of extensive inputs. The model is available under the Mistral Research License for research and educational use, and the Mistral Commercial License for commercial applications. Pixtral Large demonstrates strong performance in various benchmarks, particularly excelling in reliability and specific accuracy tasks. It achieved a 100% success rate across all benchmarks, indicating consistent and dependable outputs. In terms of accuracy, the model excels in areas requiring precise knowledge and ethical reasoning, achieving perfect 100% accuracy in both Hallucinations (Baseline) and Ethics (Baseline) benchmarks. It also performed well in General Knowledge (99.5% accuracy) and Mathematics (90.0% accuracy). While its Instruction Following (64.0% accuracy) and Reasoning (60.0% accuracy) scores are respectable, they represent areas with potential for further improvement compared to its top-tier performance in other categories. The model's performance in Email Classification (98.0% accuracy) is strong, though its duration for this benchmark was notably high. Coding accuracy (82.0%) is moderate. Overall, Pixtral Large is a highly reliable and accurate model, particularly strong in knowledge-based and ethical tasks, with competitive speed and moderate pricing. However, its higher cost per million tokens and slower processing speed compared to some alternatives may be considerations for specific use cases. The model supports image input, allowing it to process and understand documents, charts, and natural images. This capability makes it suitable for applications requiring multimodal understanding, such as document analysis, content creation, and complex query answering. Its large context window further enhances its suitability for tasks involving extensive information processing, such as long-context chat and document analysis, agent workflows with large memory windows, agentic systems with function or tool calling, workflow automation and API orchestration, multimodal applications requiring image or audio processing, and content analysis across multiple media types. In summary, Pixtral Large offers a robust set of features and capabilities, making it a strong candidate for various advanced AI applications, particularly those requiring multimodal understanding and extensive context processing.
Mistral Small 3.1
mistral-small-3-1-24b-latest
Mistral Small 3.1 is a 24-billion-parameter AI model developed by Mistral, featuring advanced multimodal capabilities and an extended context window of 128,000 tokens, enabling efficient processing of long documents and complex tasks.
Mistral Nemo
mistral-nemo-12b-latest
Mistral NeMo is a 12-billion parameter multilingual language model developed by Mistral in collaboration with NVIDIA, designed for efficient processing of large context windows and cost-effective deployment.
moonshotai

Kimi K2.5
kimi-k2-5-latest
Moonshot's agentic open MoE with elite long-context recall.
moonshot

Kimi K2
kimi-k2-latest
Kimi K2 is an open-source AI model developed by Moonshot AI, featuring a Mixture-of-Experts (MoE) architecture with 1 trillion total parameters and 32 billion active parameters per token. It offers a 128,000-token context window, enabling it to process and understand extensive documents and conversations. Kimi K2 excels in coding tasks, achieving a 65.8% pass rate on the SWE-bench Verified benchmark, and demonstrates strong reasoning capabilities, scoring 97.4% on the MATH-500 benchmark. It is competitively priced at $0.15 per million input tokens and $2.50 per million output tokens, making it a cost-effective choice for developers and researchers. The model's knowledge is up-to-date as of July 12, 2025, and it is licensed under a modified MIT license, allowing for broad usage and modification.
Cohere
Command A
command-a-111b-latest
Command A is Cohere's advanced AI model optimized for enterprise applications, offering high performance and multilingual support.
novasky

Sky T1 Preview
sky-t1-preview-32b-preview
Sky-T1-32B-Preview is an open-source AI model developed by NovaSky, offering high performance in reasoning and coding tasks at a fraction of traditional training costs.
Perplexity
R1 1776
r1-1776-latest
R1 1776 is a post-trained reasoning model by Perplexity AI, designed to provide unbiased and accurate information while maintaining high reasoning capabilities. It offers a 128K token context window and is available under the Apache 2.0 license, allowing for open-source deployment.
microsoft
Phi 4
phi-4-14b-latest
Phi-4 is Microsoft's fourth-generation language model, designed for complex reasoning and coding tasks, offering high performance and efficiency.
amazon
Nova Pro
nova-pro-latest
Amazon Nova Pro is a highly capable multimodal AI model designed for tasks requiring advanced understanding of text, images, and videos. It offers a large context window, competitive pricing, and strong performance across various benchmarks.
Nova Lite
nova-lite-latest
Amazon Nova Lite is a cost-effective, fast, and efficient multimodal AI model designed for real-time interactions, document analysis, and visual question answering. It supports text, image, and video inputs with a 300K token context window, making it suitable for applications requiring quick responses and processing of diverse data types.
Nova Micro
nova-micro-latest
Amazon's Nova Micro is a text-only AI model optimized for speed and cost, featuring a 128K token context window and strong performance in core language tasks.