Model catalog
16 models, one key, metered per million tokens
Rates below are list prices in USD per 1M tokens. Your prepaid credits cover every model — pick per request, no per-vendor setup.
| Model | Provider | Context | Input / 1M | Output / 1M | |
|---|---|---|---|---|---|
|
GPT-4o
Popular
Flagship multimodal model for the hardest production workloads.
|
OpenAI | 128K | $2.5 | $10 | use → |
|
GPT-4o mini
Small, fast and remarkably capable — the default choice for high-volume tasks.
|
OpenAI | 128K | $0.15 | $0.6 | use → |
|
o3-mini
New
Reasoning model tuned for math, science and code with configurable effort.
|
OpenAI | 200K | $1.1 | $4.4 | use → |
|
Claude 3.5 Sonnet
Anthropic flagship — best-in-class writing, analysis and agentic coding.
|
Anthropic | 200K | $3 | $15 | use → |
|
Claude 3.5 Haiku
Fastest Claude — near-instant responses for chat and extraction.
|
Anthropic | 200K | $0.8 | $4 | use → |
|
Gemini 1.5 Pro
Long context
Two-million-token context window for whole-codebase and long-video reasoning.
|
2000K | $1.25 | $5 | use → | |
|
Gemini 1.5 Flash
Cost-efficient multimodal model with a one-million-token window.
|
1000K | $0.075 | $0.3 | use → | |
|
Llama 3.1 405B
The largest open-weights model — GPT-4-class quality you can self-host later.
|
Meta | 128K | $2.7 | $2.7 | use → |
|
Llama 3.1 70B
The open-weights workhorse for fine-tuning and private deployments.
|
Meta | 128K | $0.52 | $0.75 | use → |
|
Llama 3.1 8B
Ultra-cheap 8B model for classification, routing and simple extraction.
|
Meta | 128K | $0.05 | $0.08 | use → |
|
DeepSeek-V3
671B MoE open model with strong general and coding performance.
|
DeepSeek | 64K | $0.27 | $1.1 | use → |
|
DeepSeek-R1
New
Open reasoning model with visible chain-of-thought at a fraction of frontier pricing.
|
DeepSeek | 128K | $0.55 | $2.19 | use → |
|
Qwen2.5 72B
Strong multilingual open model with excellent Chinese and English ability.
|
Alibaba | 128K | $0.35 | $0.4 | use → |
|
Qwen2.5 Coder 32B
Specialised coding model that rivals much larger general models.
|
Alibaba | 128K | $0.18 | $0.18 | use → |
|
Mistral Large 2
European flagship with strong reasoning, code and multilingual coverage.
|
Mistral | 128K | $2 | $6 | use → |
|
Command R+
RAG-optimised enterprise model with citation support built in.
|
Cohere | 128K | $2.5 | $10 | use → |
OpenAI-compatible
Point any OpenAI SDK, LangChain or LlamaIndex client at the gateway. Keep your code, change one URL.
Failover included
A rate-limited provider is retried against spare capacity — transparently, with the same request id.
Need a model we don't list?
Tell us at [email protected] and we will wire it into the gateway.