OrchestrAIte
Home/AI Models

AI Models

Configure available AI models and providers for your organization

OpenAI

No API Key
No API Key
OpenAI

GPT-6 Astra

gpt-6-astra

OpenAI's most intelligent and aligned model. State-of-the-art on computer use, browsing, software engineering, science and professional work. Saturates FrontierMath Tier 4 (98%) and ARC-AGI-3 (99.9%).

frontier
agentic
computer-use
coding
reasoning
multimodal
🌍 Global
1050K ctx
premium
No API Key
OpenAI

GPT-5.5 Pro

gpt-5-5-pro

Maximum-compute GPT-5.5 for the hardest reasoning tasks.

reasoning
premium
multimodal
🌍 Global
1000K ctx
premium
No API Key
OpenAI

o5

o5-latest

Deep deliberate reasoning model for math, science and multi-step planning.

512K ctx
premium
No API Key
OpenAI

GPT-5.2

gpt-5-2-latest

OpenAI's flagship frontier model with adaptive reasoning and native agentic tool use.

coding
reasoning
multimodal
agentic
cost-effective
1000K ctx
premium
No API Key
OpenAI

GPT-5.6 Sol

gpt-5-6-sol

Flagship GPT-5.6 tier for frontier coding, reasoning and agentic work.

frontier
coding
reasoning
agentic
multimodal
🌍 Global
1050K ctx
premium
No API Key
OpenAI

GPT-5.4 Pro

gpt-5-4-pro

Maximum-compute GPT-5.4 reasoning model.

reasoning
premium
🌍 Global
1050K ctx
premium
No API Key
OpenAI

GPT-5.5

gpt-5-5

Previous flagship generation; strong coding, reasoning and agentic performance.

coding
reasoning
agentic
multimodal
🌍 Global
1000K ctx
premium
No API Key
OpenAI

GPT-5 pro

gpt-5-pro-latest

GPT-5 Pro is OpenAI's most advanced AI model, optimized for complex tasks requiring deep reasoning and precise coding capabilities. It features a 400,000-token context window, supports multimodal inputs, and offers high accuracy in various benchmarks.

coding
reasoning
multimodal
fast
cost-effective
🌍 Global
400K ctx
No API Key
OpenAI

o1 Preview

o1-preview-old

OpenAI's o1-preview is a cutting-edge AI model optimized for complex reasoning tasks, particularly in STEM fields, offering high performance and cost efficiency.

reasoning
coding
STEM
high-performance
cost-effective
🌍 Global
128K ctx
No API Key
OpenAI

GPT-5 Thinking

gpt-5-thinking-latest

GPT-5 is OpenAI's advanced AI model designed for complex coding, reasoning, and multimodal tasks, offering a large context window and high accuracy across various benchmarks.

coding
reasoning
multimodal
high-performance
cost-effective
🌍 Global
400K ctx
No API Key
OpenAI

GPT-5.6 Terra

gpt-5-6-terra

Production default of the GPT-5.6 family — strong capability at mid-tier cost.

coding
reasoning
production
multimodal
🌍 Global
1050K ctx
high
No API Key
OpenAI

GPT-5.4

gpt-5-4

Balanced GPT-5.4 model with 1M context.

coding
reasoning
multimodal
🌍 Global
1050K ctx
medium
No API Key
OpenAI

GPT-5.3 Codex

gpt-5-3-codex

Agentic coding specialist optimized for the Codex harness.

coding
agentic
🌍 Global
400K ctx
medium
No API Key
OpenAI

GPT-5

gpt-5-latest

GPT-5 is OpenAI's advanced AI model designed for complex reasoning, coding, and multimodal tasks, offering a large context window and competitive pricing.

coding
reasoning
multimodal
high-performance
cost-effective
🌍 Global
400K ctx
No API Key
OpenAI

GPT-5.6 Luna

gpt-5-6-luna

High-volume, cost-sensitive GPT-5.6 model for classification, extraction and routing.

fast
cost-effective
multimodal
🌍 Global
1050K ctx
low
No API Key
OpenAI

gpt-oss-120b

gpt-oss-120b-120b-latest

gpt-oss-120b is OpenAI's most powerful open-weight model, designed for complex reasoning and coding tasks, offering high performance at a cost-effective price point.

reasoning
coding
agentic
high-performance
cost-effective
🌍 Global
131K ctx
No API Key
OpenAI

GPT-5.2 Mini

gpt-5-2-mini-latest

Fast, low-cost workhorse for high-volume production workloads.

coding
reasoning
multimodal
professional
agentic
frontier
400K ctx
low
No API Key
OpenAI

GPT 4o

gpt-4o-latest

GPT-4o is OpenAI's advanced AI model released in May 2024, excelling in voice, multilingual, and vision tasks, with a context window of 128,000 tokens and a knowledge cutoff in October 2023.

multimodal
voice
multilingual
vision
cost-effective
🌍 Global
128K ctx
No API Key
OpenAI

GPT 4o

gpt-4o-old

GPT-4o is OpenAI's advanced AI model released in May 2024, excelling in voice, multilingual, and vision tasks, with a context window of 128,000 tokens and a knowledge cutoff in October 2023.

multimodal
voice
multilingual
vision
cost-effective
🌍 Global
128K ctx
No API Key
OpenAI

GPT-5.4 Mini

gpt-5-4-mini

Fast, affordable GPT-5.4 mini model.

fast
cost-effective
🌍 Global
400K ctx
low
No API Key
OpenAI

GPT-5 nano

gpt-5-nano-latest

GPT-5 nano is OpenAI's most compact and cost-effective variant in the GPT-5 family, designed for efficient processing of text and image inputs, making it ideal for applications where speed and affordability are paramount.

coding
reasoning
multimodal
fast
cost-effective
🌍 Global
400K ctx
No API Key
OpenAI

GPT 4.1 mini

gpt-4-1-mini-latest

GPT-4.1 mini is a mid-sized AI model by OpenAI, offering strong performance in coding and instruction following tasks, with a 1 million token context window and a knowledge cutoff in June 2024.

coding
reasoning
instruction following
tool calling
cost-effective
🌍 Global
1000K ctx
No API Key
OpenAI

gpt-oss-20b

gpt-oss-20b-20b-latest

gpt-oss-20b is OpenAI's open-weight language model with 21 billion parameters, designed for efficient deployment on consumer hardware with just 16GB memory. Released under Apache 2.0 license, it demonstrates strong performance across benchmarks including 85.3% MMLU and exceptional AIME scores.

coding
reasoning
cost-effective
fast
open-weight
🌍 Global
131K ctx
No API Key
OpenAI

o3 mini high

o3-mini-high-latest

OpenAI's o3-mini is a cost-effective, high-speed reasoning model optimized for STEM tasks, offering advanced capabilities in coding, math, and science.

coding
reasoning
cost-effective
fast
🌍 Global
200K ctx
No API Key
OpenAI

GPT-5 mini

gpt-5-mini-latest

GPT-5 mini is a cost-effective, high-performance AI model by OpenAI, designed for well-defined tasks requiring efficient processing and strong capabilities in coding and reasoning.

multimodal
cost-effective
fast
coding
reasoning
🌍 Global
400K ctx
No API Key
OpenAI

o1

o1-latest

OpenAI's o1 model is a cutting-edge AI designed for complex reasoning and problem-solving tasks, excelling in STEM-related applications.

reasoning
coding
multimodal
high-cost
large-context-window
🌍 Global
200K ctx
No API Key
OpenAI

o3 mini medium

o3-mini-medium-latest

OpenAI's o3-mini is a cost-efficient reasoning model optimized for STEM tasks, particularly excelling in science, math, and coding. It offers advanced reasoning capabilities with a 200,000-token context window and a knowledge cutoff in October 2023.

coding
reasoning
cost-effective
fast
🌍 Global
200K ctx
No API Key
OpenAI

o1 Pro

o1-pro-latest

OpenAI's o1 Pro is a high-performance AI model designed for advanced reasoning and complex problem-solving tasks, offering a large context window and high accuracy in various benchmarks.

reasoning
complex problem-solving
large context window
high cost
🌍 Global
200K ctx
No API Key
OpenAI

GPT 4.5

gpt-4-5-preview

GPT-4.5 is OpenAI's advanced language model, offering enhanced reasoning, coding capabilities, and a 128,000-token context window, suitable for complex tasks but at a higher cost.

multimodal
reasoning
coding
high-context
costly
🌍 Global
128K ctx
No API Key
OpenAI

o3 mini low

o3-mini-low-latest

OpenAI's o3-mini is a cost-efficient reasoning model optimized for STEM tasks, particularly excelling in science, math, and coding. It offers high intelligence at the same cost and latency targets as o1-mini, supporting key developer features like function calling, structured outputs, and batch API. Released on January 31, 2025, with a knowledge cutoff in October 2023, o3-mini provides a context window of 200,000 tokens and a maximum output of 100,000 tokens. It is priced at $1.10 per million input tokens and $4.40 per million output tokens. Benchmark performance includes 97.9% on MATH, 86.9% on MMLU, and 79.7% on GPQA. While it does not support vision capabilities, o3-mini is ideal for tasks requiring high intelligence at lower cost, such as landing page generation, policy analysis, text-to-SQL conversion, and graph entity extraction. It is available through OpenAI's API and supports text input and output.

coding
reasoning
cost-effective
fast
🌍 Global
200K ctx
No API Key
OpenAI

GPT 4.1

gpt-4-1-latest

GPT-4.1 is a versatile AI model by OpenAI, excelling in instruction following, tool calling, and handling long-context tasks. It offers a 1 million token context window and low latency without a reasoning step.

coding
reasoning
instruction following
long context
cost-effective
🌍 Global
1000K ctx
No API Key
OpenAI

GPT-5.4 Nano

gpt-5-4-nano

Lowest-cost GPT-5.4 model for classification and simple tasks.

fast
cost-effective
🌍 Global
400K ctx
low
No API Key
OpenAI

GPT 4o mini

gpt-4o-mini-latest

GPT-4o mini is OpenAI's most cost-efficient small model, offering strong performance in coding and reasoning tasks, with a context window of 128,000 tokens and support for text and image inputs.

coding
reasoning
cost-effective
multimodal
fast
🌍 Global
128K ctx
No API Key
OpenAI

GPT 4.1 nano

gpt-4-1-nano-latest

GPT-4.1 nano is OpenAI's fastest and most cost-efficient model, designed for tasks requiring low latency, such as classification and autocompletion. It features a 1 million token context window and supports both text and image inputs.

coding
reasoning
multimodal
fast
cost-effective
🌍 Global
1000K ctx
No API Key
OpenAI

GPT 4o

gpt-4o-old-2

GPT-4o is OpenAI's general-purpose language model designed for a wide range of language processing tasks, offering a balance between performance and cost efficiency.

coding
reasoning
multimodal
fast
cost-effective
🌍 Global
128K ctx
No API Key
OpenAI

o1 mini

o1-mini-latest

o1-mini is a cost-efficient reasoning model optimized for STEM tasks, offering strong performance in coding and math while maintaining affordability.

coding
reasoning
cost-effective
fast
🌍 Global
128K ctx
No API Key
OpenAI

GPT 4o

gpt-4o-old-3

GPT-4o is a versatile AI model by OpenAI, known for its large context window and support for image inputs, making it suitable for tasks involving extensive text and visual data.

multimodal
cost-effective
large-context
image-input
🌍 Global
128K ctx

Anthropic

No API Key
No API Key
Anthropic

Claude Opus 4.6

claude-opus-4-6-latest

Anthropic's most capable model — best-in-class coding and long-horizon agent tasks.

coding
reasoning
multimodal
agentic
long-context
1000K ctx
premium
No API Key
Anthropic

Claude Sonnet 4.6

claude-sonnet-4-6-latest

The best balance of intelligence, speed and cost for enterprise workloads.

coding
reasoning
multimodal
fast
cost-effective
agentic
1000K ctx
medium
No API Key
Anthropic

Claude Sonnet 4.5

claude-sonnet-4-5-latest

Claude Sonnet 4.5 is Anthropic's most advanced AI model, excelling in complex coding, agent-based workflows, and long-duration tasks.

coding
reasoning
multimodal
fast
cost-effective
🌍 Global
200K ctx
No API Key
Anthropic

Claude Opus 4.1

claude-opus-4-1-latest

Claude Opus 4.1 is Anthropic's advanced AI model, excelling in coding and reasoning tasks, with a 200K token context window and a knowledge cutoff in March 2025.

coding
reasoning
agentic
cost-effective
🌍 Global
200K ctx
No API Key
Anthropic

Claude 3 Opus

claude-3-opus-latest

Claude 3 Opus is Anthropic's most intelligent AI model, offering exceptional performance across a wide range of cognitive tasks, including advanced reasoning, general knowledge, and coding. It features a 200,000-token context window and supports both text and image inputs. Released on March 4, 2024, it has a knowledge cutoff in August 2023. While it is priced at $15 per million input tokens and $75 per million output tokens, its high accuracy and reliability make it a valuable tool for complex applications.

reasoning
general knowledge
coding
multimodal
high accuracy
cost-effective
🌍 Global
200K ctx
No API Key
Anthropic

Claude 3.5 Sonnet

claude-3-5-sonnet-old

Claude 3.5 Sonnet is a large language model developed by Anthropic, known for its strong coding capabilities and extended context window, making it suitable for complex coding tasks and scenarios requiring nuanced understanding.

coding
reasoning
multimodal
fast
cost-effective
🌍 Global
200K ctx
No API Key
Anthropic

Claude 3.5 Sonnet

claude-3-5-sonnet-latest

Claude 3.5 Sonnet is a large language model developed by Anthropic, known for its strong coding capabilities and extended context window.

coding
reasoning
multimodal
fast
cost-effective
🌍 Global
200K ctx
No API Key
Anthropic

Claude 3.7 Sonnet

claude-3-7-sonnet-latest

Claude 3.7 Sonnet is a high-performance AI model developed by Anthropic, offering advanced capabilities in coding, reasoning, and extended thinking.

coding
reasoning
multimodal
fast
cost-effective
🌍 Global
200K ctx
No API Key
Anthropic

Claude 3.5 Haiku

claude-3-5-haiku-latest

Claude 3.5 Haiku is a language model developed by Anthropic, designed for high-throughput, simple tasks. It offers a balance between performance and cost, making it suitable for applications requiring rapid responses and efficient processing.

fast
cost-effective
coding
reasoning
🌍 Global
200K ctx
No API Key
Anthropic

Claude 3 Haiku

claude-3-haiku-old

Claude 3 Haiku is Anthropic's fastest and most affordable model in its intelligence class, designed for enterprise applications requiring rapid analysis of large datasets.

fast
cost-effective
multimodal
🌍 Global
200K ctx

Google AI

No API Key
No API Key
Google AI

Gemini 3.1 Pro

gemini-3-1-pro-latest

Google's frontier multimodal model with 2M context and native video understanding.

coding
reasoning
multimodal
cost-effective
agentic
long-context
2000K ctx
high
No API Key
Google AI

Gemma 3

gemma-3-12b-latest

Gemma 3 is a versatile AI model developed by Google, designed for efficient processing of extensive documents and multimodal data, offering support for over 140 languages and advanced reasoning capabilities.

multimodal
coding
reasoning
cost-effective
high-context
🌍 Global
131K ctx
No API Key
Google AI

Gemma 3

gemma-3-4b-latest

Gemma 3 is a state-of-the-art AI model developed by Google, designed for efficient deployment on single GPUs or TPUs. It offers advanced text and visual reasoning capabilities, supports over 140 languages, and features a 128,000-token context window, enabling the processing of extensive documents.

multimodal
coding
reasoning
cost-effective
fast
🌍 Global
128K ctx
No API Key
Google AI

Gemini 1.5 Pro 001

gemini-1-5-pro-001-old

Gemini 1.5 Pro is Google's advanced AI model designed for complex reasoning tasks, offering a 2 million token context window and support for various input modalities, including text, images, audio, and video.

multimodal
long-context
high-performance
cost-effective
🌍 Global
2000K ctx
No API Key
Google AI

Gemini 1.5 Pro 002

gemini-1-5-pro-002-old

Gemini 1.5 Pro 002 is a high-performance AI model developed by Google, designed for complex reasoning tasks and large-scale data analysis. It offers a 2 million token context window, supports multimodal inputs (text, images, audio, video), and demonstrates strong performance across various benchmarks. Released in September 2024, it provides cost-effective pricing for both input and output tokens.

multimodal
long-context
high-performance
cost-effective
🌍 Global
2000K ctx
No API Key
Google AI

Gemini Exp 1206

gemini-exp-1206-old

Gemini Exp 1206 is a cutting-edge AI model developed by Google, released on December 17, 2025. It boasts a substantial context window of 2,097,152 tokens, enabling it to process and understand extensive inputs effectively. The model demonstrates strong performance across various benchmarks, including a 90.04% accuracy on MMLU, 94.4% performance on GSM8K, and a 74.4% success rate on Python coding tasks (HumanEval).

coding
reasoning
multimodal
fast
cost-effective
🌍 Global
2000K ctx
No API Key
Google AI

Gemini 3 Flash

gemini-3-flash-latest

Ultra-fast multimodal model with 1M context — the volume default.

coding
reasoning
multimodal
fast
cost-effective
agentic
1000K ctx
low
No API Key
Google AI

Gemma 3

gemma-3-27b-latest

Gemma 3 is a state-of-the-art AI model developed by Google, designed for efficient processing of extensive documents and multimodal data. It offers a large context window, supports over 140 languages, and is optimized for single-GPU deployment.

multimodal
large context window
multilingual
cost-effective
high performance
🌍 Global
131K ctx
No API Key
Google AI

Gemini 2.0 Flash Exp

gemini-2-0-flash-exp-old

Gemini 2.0 Flash is Google's next-generation AI model designed for high-speed, cost-effective processing across multiple modalities, including text, images, video, and audio, with a 1 million token context window.

coding
reasoning
multimodal
fast
cost-effective
🌍 Global
1000K ctx
No API Key
Google AI

Gemini 2.0 Flash

gemini-2-0-flash-latest

Gemini 2.0 Flash is Google's advanced AI model designed for high-speed, cost-effective, and reliable performance across various tasks, including coding and reasoning.

multimodal
fast
cost-effective
large context window
coding
reasoning
🌍 Global
1000K ctx
No API Key
Google AI

Gemini 1.5 Flash 002

gemini-1-5-flash-002-old

Gemini 1.5 Flash 002 is a high-performance, cost-effective AI model developed by Google, designed for applications requiring large context processing and multimodal input handling.

multimodal
large context window
cost-effective
high performance
🌍 Global
1000K ctx
No API Key
Google AI

Gemini 2.5 Flash Lite Thinking

gemini-2-5-flash-lite-thinking-latest

Gemini 2.5 Flash-Lite is Google's fastest and most cost-efficient AI model, optimized for high-throughput applications requiring adaptive thinking.

coding
reasoning
multimodal
fast
cost-effective
🌍 Global
1000K ctx
No API Key
Google AI

Gemini 2.5 Flash Lite

gemini-2-5-flash-lite-latest

Gemini 2.5 Flash-Lite is Google's fastest and most cost-efficient AI model, optimized for high-throughput tasks requiring rapid responses.

multimodal
fast
cost-effective
high-throughput
low-latency
🌍 Global
1000K ctx
No API Key
Google AI

Gemini 2.0 Flash-Lite

gemini-2-0-flash-lite-latest

Gemini 2.0 Flash-Lite is a lightweight, cost-efficient AI model by Google, optimized for low-latency, high-throughput tasks, supporting text, image, audio, and video inputs with a 1 million token context window and a knowledge cutoff in August 2024.

multimodal
fast
cost-effective
coding
reasoning
🌍 Global
1000K ctx
No API Key
Google AI

Gemini 2.0 Pro

gemini-2-0-pro-preview

Gemini 2.0 Pro is Google's advanced AI model, released on December 11, 2024, offering enhanced capabilities in coding performance and complex prompt handling. It supports multimodal inputs and outputs, including text, images, video, and audio, with a context window of 2 million input tokens and 8,192 output tokens. The model has a knowledge cutoff in August 2024 and is available through Google AI Studio and Vertex AI. Pricing is approximately $0.10 per million input tokens and $0.40 per million output tokens.

coding
reasoning
multimodal
fast
cost-effective
🌍 Global
2000K ctx
No API Key
Google AI

Gemini 2.5 Deep Think

gemini-2-5-deep-think-latest

Gemini 2.5 Deep Think is Google's advanced AI model, offering a massive 1 million-token context window and robust multimodal capabilities, excelling in complex reasoning and coding tasks.

multimodal
reasoning
coding
long-context
cost-effective
🌍 Global
1049K ctx
No API Key
Google AI

Gemini 2.5 Pro

gemini-2-5-pro-old-2

Gemini 2.5 Pro is Google's state-of-the-art AI model, excelling in complex reasoning and coding tasks with a vast context window and multimodal capabilities.

reasoning
coding
multimodal
long-context
cost-effective
🌍 Global
1000K ctx
No API Key
Google AI

Gemini 2.5 Flash

gemini-2-5-flash-preview

Gemini 2.5 Flash is Google's most efficient AI model, designed for speed and low-cost, supporting text, images, video, and audio inputs with a 1 million token context window and adaptive thinking capabilities.

multimodal
adaptive thinking
cost-effective
high-performance
🌍 Global
1000K ctx
No API Key
Google AI

Gemini 2 Flash Thinking

gemini-2-flash-thinking-old

Gemini 2.5 Flash is Google's advanced AI model designed for large-scale, low-latency tasks that require reasoning and agentic capabilities. It supports multimodal inputs, including text, images, video, and audio, with a context window of 1 million tokens and a knowledge cutoff in January 2025. The model offers a cost-effective pricing structure, with input costs at $0.30 per million tokens and output costs at $2.50 per million tokens. Its strengths include high performance, adaptive thinking capabilities, and cost efficiency, making it suitable for applications requiring complex reasoning and multimodal processing.

multimodal
adaptive thinking
high performance
cost-effective
🌍 Global
32K ctx
No API Key
Google AI

Gemini Exp 1114

gemini-exp-1114-old

Gemini 3 Pro is Google's advanced AI model, released in November 2025, known for its large context window and multimodal processing capabilities.

multimodal
long-context
cost-effective
reasoning
coding
🌍 Global
32K ctx
No API Key
Google AI

Gemini Exp 1121

gemini-exp-1121-old

Google's Gemini Exp 1121 is an AI model released on November 21, 2024, featuring a 1 million-token context window and a knowledge cutoff in August 2024. It excels in reasoning tasks with a 59% average score on LiveBench, and offers cost-effective pricing at $1.25 per million input tokens and $10 per million output tokens. Its primary strength lies in handling extensive context for research and multimodal applications.

multimodal
long-context
research
cost-effective
🌍 Global
32K ctx
No API Key
Google AI

Gemini 2 Flash Thinking

gemini-2-flash-thinking-preview

Gemini 2.5 Flash is Google's advanced AI model designed for high-volume, low-latency tasks that require reasoning and multimodal processing. It offers a 1 million-token context window, supports text, images, video, and audio inputs, and provides text outputs. The model is optimized for cost-effectiveness and speed, making it suitable for applications needing rapid and reliable AI responses.

multimodal
reasoning
coding
cost-effective
fast
🌍 Global
1000K ctx
No API Key
Google AI

Gemini 1.5 Flash

gemini-1-5-flash-8b-old

Gemini 1.5 Flash is a fast and versatile multimodal AI model by Google, optimized for high-volume, real-time tasks, supporting text, images, audio, and video inputs with a large context window of 1,048,576 tokens.

multimodal
fast
cost-effective
large context window
🌍 Global
1000K ctx
No API Key
Google AI

Gemini 1.5 Flash 001

gemini-1-5-flash-001-old

Gemini 1.5 Flash is a fast and versatile multimodal AI model by Google, optimized for diverse tasks with a 1 million token context window and support for text, images, audio, and video inputs.

multimodal
fast
cost-effective
long-context
🌍 Global
1000K ctx
No API Key
Google AI

Gemma 2

gemma-2-2b-latest

Gemma 2 9B is a 9-billion parameter language model developed by Google, released on June 28, 2024. It offers a context window of 8,192 tokens and is priced at $0.03 per million input tokens and $0.09 per million output tokens. The model's knowledge cutoff is April 1, 2024. It excels in reasoning tasks and is suitable for educational item generation, adaptation to low-resource languages, and long-context deployments. As an open-source model, it provides flexibility for various applications.

reasoning
cost-effective
open-source
educational
low-resource
🌍 Global
8K ctx
No API Key
Google AI

Gemma 2

gemma-2-9b-latest

Gemma 2 is Google's efficient second-generation open-source language model, available in 9B and 27B parameter versions, designed for tasks requiring strong reasoning capabilities.

reasoning
coding
cost-effective
educational
low-resource
specialized
🌍 Global
8K ctx
No API Key
Google AI

Gemma 2

gemma-2-27b-latest

Gemma 2 9B is a 9-billion parameter language model developed by Google, designed for efficient language understanding and generation tasks.

coding
reasoning
cost-effective
🌍 Global
8K ctx

alibaba

No API Key
No API Key
alibaba

Qwen 3 A22B Thinking

qwen-3-a22b-thinking-235b-latest

Qwen3-235B-A22B-Thinking-2507 is a large-scale AI model developed by Alibaba, excelling in complex reasoning and coding tasks with a maximum context window of 387,100 tokens.

reasoning
coding
long-context
cost-effective
🌍 Global
256K ctx
No API Key
alibaba

Qwen 3

qwen-3-32b-latest

Qwen 3 is Alibaba's advanced AI model series, excelling in complex reasoning, coding tasks, and multilingual support.

coding
reasoning
multimodal
fast
cost-effective
🌍 Global
128K ctx
No API Key
alibaba

Qwen 3

qwen-3-4b-latest

Qwen 3 is Alibaba's advanced AI model series, excelling in complex reasoning, coding, and general knowledge tasks.

coding
reasoning
multimodal
fast
cost-effective
🌍 Global
32K ctx
No API Key
alibaba

Qwen 3.5 Max

qwen-3-5-max-latest

Alibaba's top open model — exceptional multilingual and coding ability.

coding
reasoning
multimodal
agentic
cost-effective
large-context
262K ctx
low
No API Key
alibaba

Qwen 2.5 VL

qwen-2-5-vl-32b-latest

Qwen 2.5 VL is a vision-language AI model developed by Alibaba, designed for complex reasoning tasks, coding assistance, and visual processing.

multimodal
coding
reasoning
cost-effective
fast
🌍 Global
16K ctx
No API Key
alibaba

Qwen 2.5 Coder

qwen-2-5-coder-0-5b-latest

Qwen 2.5 Coder is a high-performance AI model developed by Alibaba, optimized for coding and reasoning tasks across multiple programming languages. It offers a large context window, competitive benchmark scores, and cost-effective usage, making it suitable for various software development applications.

coding
reasoning
cost-effective
high-performance
open-source
🌍 Global
32K ctx
No API Key
alibaba

Qwen 2.5

qwen-2-5-7b-latest

Qwen 2.5 is an advanced AI language model developed by Alibaba, renowned for its exceptional long-context processing capabilities and competitive performance across various benchmarks.

long-context
coding
reasoning
cost-effective
open-source
🌍 Global
128K ctx
No API Key
alibaba

Qwen 2.5 Coder

qwen-2-5-coder-32b-latest

Qwen 2.5 Coder is an advanced AI model developed by Alibaba, designed to excel in code generation, debugging, and understanding across multiple programming languages. It offers a large context window, high performance in coding benchmarks, and cost-effective usage, making it suitable for various software development tasks.

coding
reasoning
cost-effective
high-performance
open-source
🌍 Global
128K ctx
No API Key
alibaba

Qwen 2.5 Coder

qwen-2-5-coder-14b-latest

Qwen 2.5 Coder is an advanced AI model developed by Alibaba, designed to excel in code generation, debugging, and understanding across multiple programming languages. It offers a large context window, high performance in coding benchmarks, and cost-effective usage, making it suitable for various software development tasks.

coding
reasoning
cost-effective
fast
open-source
🌍 Global
128K ctx
No API Key
alibaba

Qwen 2.5 Coder

qwen-2-5-coder-7b-latest

Qwen 2.5 Coder is an advanced AI model developed by Alibaba, designed to excel in code generation, debugging, and understanding complex codebases. It offers a substantial context window of 1 million tokens, enabling it to process extensive code and documentation efficiently. With competitive benchmark scores and cost-effective usage, it serves as a valuable tool for developers and educators alike.

coding
reasoning
multimodal
fast
cost-effective
🌍 Global
128K ctx
No API Key
alibaba

QWQ Max

qwq-max-preview

Qwen3-Max is Alibaba's flagship AI model, featuring over 1 trillion parameters and a 262,144-token context window, excelling in advanced reasoning and instruction-following tasks.

reasoning
instruction-following
long-context
high-performance
cost-effective
🌍 Global
262K ctx
No API Key
alibaba

Qwen 2.5

qwen-2-5-14b-latest

Qwen 2.5 is a state-of-the-art AI language model developed by Alibaba, renowned for its exceptional performance in handling long-context tasks, coding, and reasoning. It offers cost-effective solutions for processing extensive documents and codebases.

long-context
coding
reasoning
cost-effective
🌍 Global
128K ctx
No API Key
alibaba

Qwen 2.5 Coder

qwen-2-5-coder-3b-latest

Qwen 2.5 Coder is an advanced AI model developed by Alibaba, designed to excel in coding and reasoning tasks across multiple programming languages. It offers a large context window, high benchmark scores, and cost-effective usage, making it suitable for various software development applications.

coding
reasoning
cost-effective
high-performance
open-source
🌍 Global
32K ctx
No API Key
alibaba

Qwen 3 A22B

qwen-3-a22b-235b-old

Qwen 3 A22B is a large-scale AI model developed by Alibaba, excelling in advanced reasoning and coding tasks, suitable for enterprise applications and complex problem-solving scenarios.

reasoning
coding
cost-effective
high-performance
🌍 Global
128K ctx
No API Key
alibaba

Qwen 3 Coder

qwen-3-coder-480b-latest

Qwen 3 Coder is a state-of-the-art AI model developed by Alibaba, designed for advanced coding and reasoning tasks with a vast context window and multilingual support.

coding
reasoning
multilingual
cost-effective
long-context
🌍 Global
256K ctx
No API Key
alibaba

Qwen3 A3B Thinking

qwen3-a3b-thinking-30b-latest

Qwen3-Next-80B-A3B-Thinking is a reasoning-enhanced variant of Alibaba's Qwen3-Next architecture, designed for complex problem-solving with mixture-of-experts efficiency.

reasoning
coding
cost-effective
high-performance
🌍 Global
1000K ctx
No API Key
alibaba

Qwen 2.5

qwen-2-5-32b-latest

Qwen 2.5 is an advanced AI language model developed by Alibaba, renowned for its exceptional long-context processing capabilities and competitive performance across various benchmarks.

long-context
coding
reasoning
cost-effective
🌍 Global
128K ctx
No API Key
alibaba

Qwen 2.5

qwen-2-5-0-5b-latest

Qwen 2.5 is an advanced AI language model developed by Alibaba, renowned for its exceptional long-context processing capabilities and competitive performance across various benchmarks.

long-context
coding
reasoning
cost-effective
🌍 Global
32K ctx
No API Key
alibaba

Qwen 2.5 Turbo

qwen-2-5-turbo-latest

Qwen2.5-Turbo is an AI language model developed by Alibaba, designed to handle extensive contexts efficiently and affordably, excelling in both coding and reasoning tasks.

long-context
coding
reasoning
cost-effective
🌍 Global
1000K ctx
No API Key
alibaba

Qwen 2.5

qwen-2-5-3b-latest

Qwen 2.5 is an advanced AI language model developed by Alibaba, renowned for its exceptional long-context processing capabilities and strong performance in coding and reasoning tasks.

long-context
coding
reasoning
cost-effective
🌍 Global
32K ctx
No API Key
alibaba

Qwen 2.5 Coder

qwen-2-5-coder-1-5b-latest

Qwen 2.5 Coder is a versatile AI model developed by Alibaba, designed to excel in coding and reasoning tasks across multiple programming languages, with support for long-context processing up to 1 million tokens.

coding
reasoning
multimodal
fast
cost-effective
🌍 Global
32K ctx
No API Key
alibaba

Qwen 2.5

qwen-2-5-1-5b-latest

Qwen 2.5 is an advanced AI language model developed by Alibaba, renowned for its exceptional long-context processing capabilities and competitive performance across various benchmarks.

long-context
coding
reasoning
cost-effective
🌍 Global
32K ctx
No API Key
alibaba

QwQ

qwq-32b-latest

QwQ is an open-source AI model developed by Alibaba, designed to enhance reasoning and coding capabilities. It boasts a large context window and competitive performance across various benchmarks.

reasoning
coding
open-source
high-performance
cost-effective
🌍 Global
131K ctx
No API Key
alibaba

Qwen 3 A22B

qwen-3-a22b-235b-latest

Qwen 3 A22B is a large-scale AI model developed by Alibaba, excelling in complex reasoning and coding tasks with a substantial context window and competitive pricing.

reasoning
coding
long-context
cost-effective
🌍 Global
262K ctx
No API Key
alibaba

Qwen 3 A3B

qwen-3-a3b-30b-old

Qwen 3 A3B is a reasoning-enhanced AI model by Alibaba, designed for complex problem-solving and efficient long-context processing.

reasoning
coding
long-context
cost-effective
🌍 Global
128K ctx
No API Key
alibaba

Qwen 2.5

qwen-2-5-72b-latest

Qwen 2.5 is an advanced AI language model developed by Alibaba, renowned for its exceptional long-context processing capabilities and cost efficiency.

long-context
coding
reasoning
cost-effective
🌍 Global
128K ctx
No API Key
alibaba

Qwen 3 A3B

qwen-3-a3b-30b-latest

Qwen 3 A3B is a reasoning-enhanced variant of Alibaba's Qwen 3 series, designed for complex problem-solving with mixture-of-experts efficiency.

reasoning
coding
cost-effective
high-performance
🌍 Global
1000K ctx
No API Key
alibaba

Qwen 2.5 Max

qwen-2-5-max-latest

Qwen 2.5 Max is a versatile AI model developed by Alibaba, excelling in tasks that demand extensive context understanding, coding assistance, and general knowledge retrieval. It offers a substantial context window of 131,072 tokens, enabling it to process and comprehend lengthy documents effectively. The model demonstrates strong performance in various benchmarks, including a score of 89.4 in the Arena-Hard preference benchmark, surpassing competitors like DeepSeek V3 (85.5) and Claude 3.5 Sonnet (85.2). In coding tasks, it achieves a HumanEval score of 73.2, indicating its proficiency in code generation and problem-solving. Despite its capabilities, users should be aware of its cost structure, with input tokens priced at $1.6 per million and output tokens at $6.4 per million, which may impact cost-effectiveness for extensive usage. Overall, Qwen 2.5 Max is well-suited for applications requiring deep contextual understanding and complex reasoning, such as legal document analysis, technical content generation, and advanced coding assistance.

multimodal
reasoning
coding
long-context
cost-effective
🌍 Global
32K ctx
No API Key
alibaba

QwQ Preview

qwq-preview-32b-preview

QwQ-32B-Preview is an experimental AI model developed by Alibaba's Qwen team in 2024, focusing on enhancing AI reasoning capabilities, particularly in mathematics and programming.

reasoning
coding
experimental
math
programming
🌍 Global
32K ctx

deepseek

No API Key
No API Key
deepseek

DeepSeek Coder 2

deepseek-coder-2-236b-latest

DeepSeek Coder 2 is an advanced AI model specializing in coding and reasoning tasks, offering high performance and cost efficiency.

coding
reasoning
cost-effective
high-performance
large-context
🌍 Global
128K ctx
No API Key
deepseek

DeepSeek R2

deepseek-r2-latest

Open reasoning model rivaling proprietary frontier on math and code.

reasoning
coding
cost-effective
open-weights
large-context
256K ctx
low
No API Key
deepseek

DeepSeek V4

deepseek-v4-latest

Open-weights frontier MoE — near-flagship quality at a fraction of the cost.

coding
reasoning
multimodal
fast
cost-effective
open-weights
agentic
256K ctx
low
No API Key
deepseek

DeepSeek V3.1

deepseek-v3-1-840b-latest

DeepSeek V3.1 is an advanced open-source AI model developed by DeepSeek, known for its strong performance in coding and reasoning tasks, large context window, and cost-effective operation.

coding
reasoning
cost-effective
open-source
high-performance
🌍 Global
128K ctx
No API Key
deepseek

DeepSeek 2.5

deepseek-2-5-236b-old

DeepSeek 2.5 is an AI model developed by DeepSeek, known for its strong performance in coding and reasoning tasks, cost efficiency, and extended context window capabilities.

coding
reasoning
cost-effective
long-context
open-source
🌍 Global
128K ctx
No API Key
deepseek

Deepseek 2.5

deepseek-2-5-old

DeepSeek 2.5 is an advanced AI model developed by DeepSeek, known for its cost-effectiveness and strong performance in coding and reasoning tasks. It offers a context window of 128,000 tokens, enabling efficient processing of extensive inputs. The model is open-source, allowing for customization and integration into various applications.

coding
reasoning
cost-effective
long-context
open-source
🌍 Global
128K ctx
No API Key
deepseek

R1 Distill Qwen

r1-distill-qwen-1-5b-latest

DeepSeek's R1 Distill Qwen is a distilled language model designed to enhance reasoning and coding capabilities while maintaining efficiency.

reasoning
coding
cost-effective
fast
🌍 Global
131K ctx
No API Key
deepseek

R1 Distill Qwen

r1-distill-qwen-7b-latest

DeepSeek's R1 Distill Qwen is a distilled language model designed to enhance reasoning and coding capabilities while maintaining cost efficiency.

coding
reasoning
cost-effective
🌍 Global
131K ctx
No API Key
deepseek

R1 Distill Llama

r1-distill-llama-8b-latest

R1 Distill Llama 70B is a distilled large language model by DeepSeek, offering advanced reasoning and coding capabilities at a competitive cost.

reasoning
coding
cost-effective
high-performance
🌍 Global
131K ctx
No API Key
deepseek

DeepSeek V3.1 Reasoner

deepseek-v3-1-reasoner-840b-latest

DeepSeek V3.1 Reasoner is an advanced AI model by DeepSeek, featuring a hybrid reasoning architecture that integrates chat, reasoning, and coding capabilities into a unified model. It offers a 128K token context window, a knowledge cutoff in July 2025, and competitive pricing, making it suitable for complex problem-solving and large-scale document processing tasks.

coding
reasoning
cost-effective
open-source
high-performance
🌍 Global
128K ctx
No API Key
deepseek

DeepSeek R1

deepseek-r1-671b-latest

DeepSeek R1 is an advanced AI model developed by DeepSeek, known for its strong performance in reasoning and coding tasks, cost efficiency, and long-context processing capabilities.

coding
reasoning
cost-effective
long-context
open-source
🌍 Global
128K ctx
No API Key
deepseek

DeepSeek V3

deepseek-v3-685b-latest

DeepSeek V3 is an advanced AI model developed by DeepSeek, known for its exceptional coding and reasoning capabilities, large context window, and cost-effective pricing.

coding
reasoning
cost-effective
fast
open-source
🌍 Global
128K ctx
No API Key
deepseek

R1 Distill Llama

r1-distill-llama-70b-latest

R1 Distill Llama 70B is a distilled large language model by DeepSeek, offering advanced reasoning and coding capabilities with a cost-effective pricing structure.

reasoning
coding
cost-effective
high-performance
🌍 Global
131K ctx
No API Key
deepseek

Deepseek V3

deepseek-v3-671b-old

DeepSeek V3 is an advanced AI model developed by DeepSeek, renowned for its exceptional coding and reasoning capabilities, extensive context window, and cost-effective performance.

coding
reasoning
cost-effective
high-performance
open-source
🌍 Global
128K ctx
No API Key
deepseek

R1 Distill Qwen

r1-distill-qwen-32b-latest

DeepSeek's R1 Distill Qwen is a distilled language model designed to enhance reasoning and coding capabilities while maintaining cost efficiency.

reasoning
coding
cost-effective
efficient
🌍 Global
131K ctx
No API Key
deepseek

R1 Lite Preview

r1-lite-preview-preview

DeepSeek R1 is a large language model designed for advanced reasoning and coding tasks, offering a 128,000-token context window and competitive pricing.

coding
reasoning
cost-effective
large-context
high-performance
🌍 Global
128K ctx
No API Key
deepseek

R1 Distill Qwen

r1-distill-qwen-14b-latest

DeepSeek's R1 Distill Qwen is a series of distilled language models designed to enhance reasoning and coding capabilities while maintaining cost efficiency.

reasoning
coding
cost-effective
efficient
high-performance
🌍 Global
131K ctx

Meta

No API Key
No API Key
Meta

Llama 4 Behemoth

llama-4-behemoth-2t-preview

Llama 4 Behemoth is Meta's most powerful AI model, featuring 288 billion active parameters and a context window of up to 10 million tokens, excelling in multimodal reasoning and coding tasks.

coding
reasoning
multimodal
fast
cost-effective
🌍 Global
10000K ctx
No API Key
Meta

Llama 4 Behemoth

llama-4-behemoth-latest

Meta's largest open-weights teacher model for distillation and research.

research
multimodal
mixture-of-experts
teacher-model
unreleased
1000K ctx
high
No API Key
Meta

Llama 3.3

llama-3-3-70b-latest

Meta's Llama 3.3 is a 70-billion-parameter multilingual large language model optimized for dialogue, reasoning, and coding tasks, offering high performance at a cost-effective price point.

coding
reasoning
multilingual
cost-effective
🌍 Global
128K ctx
No API Key
Meta

Llama 3.1

llama-3-1-70b-old

Llama 3.1 is Meta's latest open-weight large language model, offering enhanced performance, a vast context window, and cost-effective usage.

coding
reasoning
multimodal
fast
cost-effective
🌍 Global
128K ctx
No API Key
Meta

Llama 3.1

llama-3-1-405b-latest

Llama 3.1 is Meta's latest large language model, offering enhanced performance and efficiency across various tasks.

coding
reasoning
multimodal
fast
cost-effective
🌍 Global
128K ctx
No API Key
Meta

Llama 4 Maverick

llama-4-maverick-400b-latest

Llama 4 Maverick is Meta's advanced multimodal AI model, excelling in coding, reasoning, and vision tasks, offering high efficiency and a large context window.

multimodal
coding
reasoning
cost-effective
fast
🌍 Global
1000K ctx
No API Key
Meta

Llama 3.1

llama-3-1-8b-latest

Meta's Llama 3.1 is a state-of-the-art language model known for its extensive context window and competitive performance across various benchmarks.

coding
reasoning
multimodal
fast
cost-effective
🌍 Global
128K ctx
No API Key
Meta

Llama 3.2

llama-3-2-3b-latest

Llama 3.2 is Meta's latest AI model, offering enhanced capabilities in natural language processing tasks.

multimodal
cost-effective
open-source
high-performance
🌍 Global
128K ctx
No API Key
Meta

Llama 3.2

llama-3-2-1b-latest

Llama 3.2 is Meta's latest AI model, offering enhanced capabilities in natural language processing tasks.

multimodal
cost-effective
open-source
high-performance
🌍 Global
128K ctx
No API Key
Meta

Llama 4 Scout

llama-4-scout-109b-latest

Llama 4 Scout is Meta's efficient multimodal AI model, featuring a 10M token context window and native support for text and image inputs.

multimodal
efficient
large-context
image-understanding
reasoning
🌍 Global
10000K ctx
No API Key
Meta

Llama 3.2 (Vision)

llama-3-2-vision-11b-latest

Llama 3.2 (Vision) is a 90-billion-parameter multimodal AI model developed by Meta, designed for advanced visual reasoning and language tasks, including image captioning, visual question answering, and complex image-text comprehension.

multimodal
vision
image analysis
document processing
chatbots
autonomous systems
🌍 Global
128K ctx
No API Key
Meta

Llama 3.2 (Vision)

llama-3-2-vision-90b-latest

Llama 3.2 (Vision) is a 90-billion-parameter multimodal AI model developed by Meta, designed for advanced visual reasoning and language tasks, including image captioning, visual question answering, and complex image-text comprehension.

multimodal
vision
image analysis
document processing
chatbots
autonomous systems
🌍 Global
128K ctx

xAI

No API Key
No API Key
xAI

Grok 4.1

grok-4-1-latest

xAI's frontier model with real-time knowledge and strong agentic reasoning.

coding
reasoning
multimodal
fast
cost-effective
agentic
512K ctx
high
No API Key
xAI

Grok 3 mini

grok-3-mini-latest

Grok 3 Mini is a lightweight AI model developed by xAI, optimized for efficient performance in logic-based tasks and coding applications, offering transparent reasoning traces and cost-effective usage.

multimodal
reasoning
coding
fast
cost-effective
🌍 Global
130K ctx
No API Key
xAI

Grok 3 mini (Think)

grok-3-mini-think-preview

Grok 3 Mini is a cost-efficient AI model by xAI, optimized for advanced reasoning and problem-solving tasks, particularly in STEM fields.

coding
reasoning
multimodal
fast
cost-effective
🌍 Global
1000K ctx
No API Key
xAI

Grok 3

grok-3-latest

Grok 3 is xAI's third-generation AI model, designed to excel in reasoning, coding, and processing large-scale inputs.

reasoning
coding
multimodal
high-context
cost-effective
🌍 Global
130K ctx
No API Key
xAI

Grok 3 (Think)

grok-3-think-preview

Grok 3 is xAI's advanced AI model, known for its exceptional reasoning and coding capabilities, large context window, and cost-effective pricing.

reasoning
coding
multimodal
high-context
cost-effective
🌍 Global
1000K ctx
No API Key
xAI

Grok 2

grok-2-latest

Grok-2 is xAI's second-generation language model, offering enhanced reasoning, coding, and vision capabilities, outperforming models like GPT-4-Turbo and Claude 3.5 in various benchmarks.

coding
reasoning
multimodal
fast
cost-effective
🌍 Global
128K ctx
No API Key
xAI

Grok-2 mini

grok-2-mini-latest

Grok-2 mini is xAI's efficient language model designed for speed without sacrificing quality, offering strong performance in reasoning, coding, and conversational tasks.

coding
reasoning
multimodal
fast
cost-effective
🌍 Global
128K ctx
No API Key
xAI

Grok 4 Heavy

grok-4-heavy-latest

Grok 4 Heavy by xAI is a high-performance AI model designed for complex reasoning and coding tasks, offering a 256,000-token context window and competitive pricing.

multimodal
reasoning
coding
high-performance
cost-effective
🌍 Global
256K ctx

Mistral

No API Key
No API Key
Mistral

Mistral Large 3

mistral-large-3-latest

European frontier model — GDPR-native, strong function calling.

multimodal
open-weights
mixture-of-experts
long-context
agentic
multilingual
256K ctx
medium
No API Key
Mistral

Codestral 25.01

codestral-25-01-latest

Codestral 25.01 is Mistral AI's advanced code generation model, offering enhanced performance and efficiency for software development tasks.

coding
reasoning
fast
cost-effective
🌍 Global
256K ctx
No API Key
Mistral

Devstral Medium

devstral-medium-latest

Devstral Medium is a high-performance language model developed by Mistral AI, specializing in code generation and agentic reasoning tasks. Released on July 10, 2025, it offers a 131,072-token context window, enabling efficient processing of large codebases. The model is accessible via API and supports enterprise deployment with optional fine-tuning capabilities.

coding
reasoning
cost-effective
fast
🌍 Global
131K ctx
No API Key
Mistral

Devstral Small 1.1

devstral-small-1-1-24b-latest

Devstral Small 1.1 is a 24-billion parameter language model developed by Mistral AI, optimized for software engineering tasks and agentic coding workflows.

coding
reasoning
cost-effective
fast
🌍 Global
131K ctx
No API Key
Mistral

Mistral Large

mistral-large-123b-old

Mistral Large is a high-performance AI model developed by Mistral, designed for enterprise applications requiring extensive context processing and multilingual support.

coding
reasoning
multilingual
cost-effective
🌍 Global
128K ctx
No API Key
Mistral

Mistral Large

mistral-large-123b-latest

Mistral Large is a 123-billion-parameter language model developed by Mistral AI, designed for enterprise applications requiring extensive multilingual support and long-context processing. It features a 32,000-token context window and is priced at $2.00 per million input tokens and $6.00 per million output tokens. Released on February 26, 2024, it has achieved an MMLU score of 81.2% in a 5-shot scenario and a HellaSwag score of 89.2% in a 10-shot scenario. The model is open-source, allowing for customization and fine-tuning to specific tasks. Its strengths include strong coding capabilities, multilingual support, and efficient token usage, making it suitable for enterprise deployments requiring extensive multilingual support and long-context processing.

coding
reasoning
multilingual
enterprise
cost-effective
🌍 Global
128K ctx
No API Key
Mistral

Mistral Medium 3

mistral-medium-3-latest

Mistral Medium 3 is a high-performance, cost-effective AI model designed for enterprise applications, offering strong coding and reasoning capabilities with multimodal processing support.

coding
reasoning
multimodal
cost-effective
enterprise-ready
🌍 Global
131K ctx
No API Key
Mistral

Mistral Small 3

mistral-small-3-24b-latest

Mistral Small 3 is a 24-billion parameter language model optimized for speed and efficiency, featuring a 32,000-token context window and competitive pricing.

coding
reasoning
fast
cost-effective
🌍 Global
32K ctx
No API Key
Mistral

Pixtral Large

pixtral-large-123b-latest

Pixtral Large is a 124-billion parameter, open-weight, multimodal AI model developed by Mistral, designed to process and understand both text and images. Released on November 18, 2024, it offers a substantial context window of 131,072 tokens, enabling efficient handling of extensive inputs. The model is available under the Mistral Research License for research and educational use, and the Mistral Commercial License for commercial applications. Pixtral Large demonstrates strong performance in various benchmarks, particularly excelling in reliability and specific accuracy tasks. It achieved a 100% success rate across all benchmarks, indicating consistent and dependable outputs. In terms of accuracy, the model excels in areas requiring precise knowledge and ethical reasoning, achieving perfect 100% accuracy in both Hallucinations (Baseline) and Ethics (Baseline) benchmarks. It also performed well in General Knowledge (99.5% accuracy) and Mathematics (90.0% accuracy). While its Instruction Following (64.0% accuracy) and Reasoning (60.0% accuracy) scores are respectable, they represent areas with potential for further improvement compared to its top-tier performance in other categories. The model's performance in Email Classification (98.0% accuracy) is strong, though its duration for this benchmark was notably high. Coding accuracy (82.0%) is moderate. Overall, Pixtral Large is a highly reliable and accurate model, particularly strong in knowledge-based and ethical tasks, with competitive speed and moderate pricing. However, its higher cost per million tokens and slower processing speed compared to some alternatives may be considerations for specific use cases. The model supports image input, allowing it to process and understand documents, charts, and natural images. This capability makes it suitable for applications requiring multimodal understanding, such as document analysis, content creation, and complex query answering. Its large context window further enhances its suitability for tasks involving extensive information processing, such as long-context chat and document analysis, agent workflows with large memory windows, agentic systems with function or tool calling, workflow automation and API orchestration, multimodal applications requiring image or audio processing, and content analysis across multiple media types. In summary, Pixtral Large offers a robust set of features and capabilities, making it a strong candidate for various advanced AI applications, particularly those requiring multimodal understanding and extensive context processing.

multimodal
image processing
function calling
structured output
long-context
document analysis
workflow automation
🌍 Global
128K ctx
No API Key
Mistral

Mistral Small 3.1

mistral-small-3-1-24b-latest

Mistral Small 3.1 is a 24-billion-parameter AI model developed by Mistral, featuring advanced multimodal capabilities and an extended context window of 128,000 tokens, enabling efficient processing of long documents and complex tasks.

multimodal
cost-effective
high-context
efficient
🌍 Global
128K ctx
No API Key
Mistral

Mistral Nemo

mistral-nemo-12b-latest

Mistral NeMo is a 12-billion parameter multilingual language model developed by Mistral in collaboration with NVIDIA, designed for efficient processing of large context windows and cost-effective deployment.

multilingual
cost-effective
fast
coding
🌍 Global
128K ctx

moonshotai

No API Key
No API Key
moonshotai

Kimi K2.5

kimi-k2-5-latest

Moonshot's agentic open MoE with elite long-context recall.

coding
reasoning
multimodal
agentic
cost-effective
long-context
512K ctx
low

moonshot

No API Key
No API Key
moonshot

Kimi K2

kimi-k2-latest

Kimi K2 is an open-source AI model developed by Moonshot AI, featuring a Mixture-of-Experts (MoE) architecture with 1 trillion total parameters and 32 billion active parameters per token. It offers a 128,000-token context window, enabling it to process and understand extensive documents and conversations. Kimi K2 excels in coding tasks, achieving a 65.8% pass rate on the SWE-bench Verified benchmark, and demonstrates strong reasoning capabilities, scoring 97.4% on the MATH-500 benchmark. It is competitively priced at $0.15 per million input tokens and $2.50 per million output tokens, making it a cost-effective choice for developers and researchers. The model's knowledge is up-to-date as of July 12, 2025, and it is licensed under a modified MIT license, allowing for broad usage and modification.

coding
reasoning
agentic
open-source
cost-effective
🌍 Global
256K ctx

Cohere

No API Key
No API Key
Cohere

Command A

command-a-111b-latest

Command A is Cohere's advanced AI model optimized for enterprise applications, offering high performance and multilingual support.

agentic
multilingual
coding
enterprise
high-performance
🌍 Global
256K ctx

novasky

No API Key
No API Key
novasky

Sky T1 Preview

sky-t1-preview-32b-preview

Sky-T1-32B-Preview is an open-source AI model developed by NovaSky, offering high performance in reasoning and coding tasks at a fraction of traditional training costs.

reasoning
coding
open-source
cost-effective
high-performance
🌍 Global
32K ctx

Perplexity

No API Key
No API Key
Perplexity

R1 1776

r1-1776-latest

R1 1776 is a post-trained reasoning model by Perplexity AI, designed to provide unbiased and accurate information while maintaining high reasoning capabilities. It offers a 128K token context window and is available under the Apache 2.0 license, allowing for open-source deployment.

reasoning
uncensored
open-source
high-accuracy
cost-effective
🌍 Global
128K ctx

microsoft

No API Key
No API Key
microsoft

Phi 4

phi-4-14b-latest

Phi-4 is Microsoft's fourth-generation language model, designed for complex reasoning and coding tasks, offering high performance and efficiency.

reasoning
coding
cost-effective
efficient
high-performance
🌍 Global
32K ctx

amazon

No API Key
No API Key
amazon

Nova Pro

nova-pro-latest

Amazon Nova Pro is a highly capable multimodal AI model designed for tasks requiring advanced understanding of text, images, and videos. It offers a large context window, competitive pricing, and strong performance across various benchmarks.

multimodal
fast
cost-effective
visual reasoning
agentic workflows
🌍 Global
300K ctx
No API Key
amazon

Nova Lite

nova-lite-latest

Amazon Nova Lite is a cost-effective, fast, and efficient multimodal AI model designed for real-time interactions, document analysis, and visual question answering. It supports text, image, and video inputs with a 300K token context window, making it suitable for applications requiring quick responses and processing of diverse data types.

multimodal
fast
cost-effective
image processing
video processing
🌍 Global
300K ctx
No API Key
amazon

Nova Micro

nova-micro-latest

Amazon's Nova Micro is a text-only AI model optimized for speed and cost, featuring a 128K token context window and strong performance in core language tasks.

coding
reasoning
fast
cost-effective
🌍 Global
128K ctx

Welcome to OrchestrAIte

Your secure gateway to every frontier AI model — governed, monitored, and ready to work.

V 1.0-I-XXXI-MMXXVI