Gateway operational · 16 models live · GPU capacity on demand Read the docs →
LINGYUNS gateway
Model catalog

16 models, one key, metered per million tokens

Rates below are list prices in USD per 1M tokens. Your prepaid credits cover every model — pick per request, no per-vendor setup.

Model Provider Context Input / 1M Output / 1M
GPT-4o Popular
Flagship multimodal model for the hardest production workloads.
OpenAI 128K $2.5 $10 use →
GPT-4o mini
Small, fast and remarkably capable — the default choice for high-volume tasks.
OpenAI 128K $0.15 $0.6 use →
o3-mini New
Reasoning model tuned for math, science and code with configurable effort.
OpenAI 200K $1.1 $4.4 use →
Claude 3.5 Sonnet
Anthropic flagship — best-in-class writing, analysis and agentic coding.
Anthropic 200K $3 $15 use →
Claude 3.5 Haiku
Fastest Claude — near-instant responses for chat and extraction.
Anthropic 200K $0.8 $4 use →
Gemini 1.5 Pro Long context
Two-million-token context window for whole-codebase and long-video reasoning.
Google 2000K $1.25 $5 use →
Gemini 1.5 Flash
Cost-efficient multimodal model with a one-million-token window.
Google 1000K $0.075 $0.3 use →
Llama 3.1 405B
The largest open-weights model — GPT-4-class quality you can self-host later.
Meta 128K $2.7 $2.7 use →
Llama 3.1 70B
The open-weights workhorse for fine-tuning and private deployments.
Meta 128K $0.52 $0.75 use →
Llama 3.1 8B
Ultra-cheap 8B model for classification, routing and simple extraction.
Meta 128K $0.05 $0.08 use →
DeepSeek-V3
671B MoE open model with strong general and coding performance.
DeepSeek 64K $0.27 $1.1 use →
DeepSeek-R1 New
Open reasoning model with visible chain-of-thought at a fraction of frontier pricing.
DeepSeek 128K $0.55 $2.19 use →
Qwen2.5 72B
Strong multilingual open model with excellent Chinese and English ability.
Alibaba 128K $0.35 $0.4 use →
Qwen2.5 Coder 32B
Specialised coding model that rivals much larger general models.
Alibaba 128K $0.18 $0.18 use →
Mistral Large 2
European flagship with strong reasoning, code and multilingual coverage.
Mistral 128K $2 $6 use →
Command R+
RAG-optimised enterprise model with citation support built in.
Cohere 128K $2.5 $10 use →

OpenAI-compatible

Point any OpenAI SDK, LangChain or LlamaIndex client at the gateway. Keep your code, change one URL.

Failover included

A rate-limited provider is retried against spare capacity — transparently, with the same request id.

Need a model we don't list?

Tell us at [email protected] and we will wire it into the gateway.