Fast & cheap · OpenAI
GPT-4o mini
Small, fast and remarkably capable — the default choice for high-volume tasks.
Context window
128K
tokens
Input price
$0.15
per 1M tokens
Output price
$0.6
per 1M tokens
Routing
Included
failover + retries
Call it with any OpenAI SDK
Swap the base URL and key — request shape, streaming and tool calls stay identical.
python
from openai import OpenAI client = OpenAI( base_url="https://api.lingyuns.com/v1", api_key="ly-sk-…", ) r = client.chat.completions.create( model="gpt-4o-mini", messages=[{"role": "user", "content": "Hello"}], ) # billed at $0.15 in / $0.6 out per 1M
Specifications
Prices are list rates in USD per million tokens. Your prepaid credits cover every model in the catalog at these rates.
Alternatives
Other Fast & cheap models
Claude 3.5 Haiku
Fastest Claude — near-instant responses for chat and extraction.
Gemini 1.5 Flash
Cost-efficient multimodal model with a one-million-token window.