AI Models Directory
Browse and compare frontier models across OpenAI, Anthropic, Google AI Studio, Groq LPU Cloud, DeepSeek, and Mistral. Switch between any model by changing a single string in /api/v1/chat/completions.
| Provider & Model ID | Input $/1M | Output $/1M | Cache Read $/1M | Context | Avg TTFT / Speed | Capabilities | Action |
|---|---|---|---|---|---|---|---|
GPT-OSS 20B on GroqGroq LPU Cloud groq/gpt-oss-20b Very fast low-cost Groq route for high-volume subagents and iterative execution loops. | $0.07 | $0.30 | — | 131.072K | 65ms TTFT 1000 tok/s · 99.9% SLA | streamingtoolsreasoningjsoncoding | Try |
GPT-OSS 120B on GroqGroq LPU Cloud groq/gpt-oss-120b High-throughput open-weight reasoning model served on Groq for fast coding and specialist-agent work. | $0.15 | $0.60 | — | 131.072K | 90ms TTFT 500 tok/s · 99.9% SLA | streamingtoolsreasoningjsoncoding | Try |
Mistral Small 4Mistral AI mistral/mistral-small-4 Efficient hybrid Mistral model combining instruct, reasoning, coding, and multimodal support. | $0.15 | $0.60 | — | 256K | 180ms TTFT 165 tok/s · 99.9% SLA | streamingvisiontoolsreasoningjsoncoding | Try |
GPT-5.6 LunaOpenAI openai/gpt-5.6-luna Cost-sensitive GPT-5.6 tier for fast subagents, classification, extraction, and routine coding work. | $0.20 | $1.20 | — | 1,050K | 180ms TTFT 180 tok/s · 99.9% SLA | streamingvisiontoolsreasoningjsoncoding | Try |
AetherGate Smart Router (Auto)AetherGate Virtual Route aethergate/smart-router Dynamic AetherGate route that ranks current BYOK models for complexity, latency, and failover resilience. | $0.20 | $1.20 | — | 1,050K | 190ms TTFT 180 tok/s · 99.99% SLA | streamingvisiontoolsreasoningjsoncoding | Try |
DeepSeek V4.1 FlashDeepSeek deepseek/deepseek-flash Current DeepSeek fast route with thinking/non-thinking modes, agent tooling, vision, and 1M context. | $0.30 | $1.20 | $0.006 | 1,000K | 240ms TTFT 130 tok/s · 99.5% SLA | streamingvisiontoolsreasoningjsoncoding | Try |
Gemini 3.8 FlashGoogle AI Studio google/gemini-3.8-flash Production Gemini model optimized for long-horizon software engineering, autonomous agents, and complex workflows. | $0.75 | $3.75 | $0.075 | 1,048.576K | 190ms TTFT 190 tok/s · 99.9% SLA | streamingvisiontoolsreasoningjsoncoding | Try |
DeepSeek V4 ProDeepSeek deepseek/deepseek-v4-pro DeepSeek high-capability route for difficult coding, reasoning, and agentic tasks with 1M context. | $1.32 | $3.96 | $0.044 | 1,000K | 420ms TTFT 85 tok/s · 99.5% SLA | streamingtoolsreasoningjsoncoding | Try |
Mistral Medium 3.5Mistral AI mistral/mistral-medium-3-5 Frontier-class Mistral model optimized for multimodal agentic and coding workloads. | $1.50 | $7.50 | — | 256K | 290ms TTFT 105 tok/s · 99.9% SLA | streamingvisiontoolsreasoningjsoncoding | Try |
Claude Sonnet 5Anthropic anthropic/claude-sonnet-5 Current Claude Sonnet generation for agentic coding, long-horizon work, tool use, and high-throughput reasoning. | $2.00 | $10.00 | — | 1,000K | 380ms TTFT 95 tok/s · 99.9% SLA | streamingvisiontoolsreasoningjsoncoding | Try |
GPT-5.6 TerraOpenAI openai/gpt-5.6-terra Balanced GPT-5.6 tier for strong coding and reasoning at lower cost than Sol. | $2.00 | $12.00 | — | 1,050K | 300ms TTFT 120 tok/s · 99.9% SLA | streamingvisiontoolsreasoningjsoncoding | Try |
GPT-5.6 SolOpenAI openai/gpt-5.6-sol OpenAI flagship for complex professional reasoning, coding, and agentic workloads. | $4.00 | $20.00 | $0.400 | 1,050K | 430ms TTFT 90 tok/s · 99.9% SLA | streamingvisiontoolsreasoningjsoncoding | Try |
Claude Opus 5Anthropic anthropic/claude-opus-5 High-depth Claude model for difficult reasoning, architecture, verification, and long-running agentic tasks. | $5.00 | $25.00 | — | 1,000K | 620ms TTFT 58 tok/s · 99.9% SLA | streamingvisiontoolsreasoningjsoncoding | Try |