Model catalog
5 models, one key, metered per million tokens
Rates below are list prices in USD per 1M tokens. Your prepaid credits cover every model — pick per request, no per-vendor setup.
| Model | Provider | Context | Input / 1M | Output / 1M | |
|---|---|---|---|---|---|
|
Llama 3.1 405B
The largest open-weights model — GPT-4-class quality you can self-host later.
|
Meta | 128K | $2.7 | $2.7 | use → |
|
Llama 3.1 70B
The open-weights workhorse for fine-tuning and private deployments.
|
Meta | 128K | $0.52 | $0.75 | use → |
|
Llama 3.1 8B
Ultra-cheap 8B model for classification, routing and simple extraction.
|
Meta | 128K | $0.05 | $0.08 | use → |
|
DeepSeek-V3
671B MoE open model with strong general and coding performance.
|
DeepSeek | 64K | $0.27 | $1.1 | use → |
|
Qwen2.5 72B
Strong multilingual open model with excellent Chinese and English ability.
|
Alibaba | 128K | $0.35 | $0.4 | use → |
OpenAI-compatible
Point any OpenAI SDK, LangChain or LlamaIndex client at the gateway. Keep your code, change one URL.
Failover included
A rate-limited provider is retried against spare capacity — transparently, with the same request id.
Need a model we don't list?
Tell us at [email protected] and we will wire it into the gateway.