Model catalog
3 models, one key, metered per million tokens
Rates below are list prices in USD per 1M tokens. Your prepaid credits cover every model — pick per request, no per-vendor setup.
| Model | Provider | Context | Input / 1M | Output / 1M | |
|---|---|---|---|---|---|
|
GPT-4o mini
Small, fast and remarkably capable — the default choice for high-volume tasks.
|
OpenAI | 128K | $0.15 | $0.6 | use → |
|
Claude 3.5 Haiku
Fastest Claude — near-instant responses for chat and extraction.
|
Anthropic | 200K | $0.8 | $4 | use → |
|
Gemini 1.5 Flash
Cost-efficient multimodal model with a one-million-token window.
|
1000K | $0.075 | $0.3 | use → |
OpenAI-compatible
Point any OpenAI SDK, LangChain or LlamaIndex client at the gateway. Keep your code, change one URL.
Failover included
A rate-limited provider is retried against spare capacity — transparently, with the same request id.
Need a model we don't list?
Tell us at [email protected] and we will wire it into the gateway.