Gateway operational · 16 models live · GPU capacity on demand Read the docs →
LINGYUNS gateway
Model catalog

5 models, one key, metered per million tokens

Rates below are list prices in USD per 1M tokens. Your prepaid credits cover every model — pick per request, no per-vendor setup.

Model Provider Context Input / 1M Output / 1M
Llama 3.1 405B
The largest open-weights model — GPT-4-class quality you can self-host later.
Meta 128K $2.7 $2.7 use →
Llama 3.1 70B
The open-weights workhorse for fine-tuning and private deployments.
Meta 128K $0.52 $0.75 use →
Llama 3.1 8B
Ultra-cheap 8B model for classification, routing and simple extraction.
Meta 128K $0.05 $0.08 use →
DeepSeek-V3
671B MoE open model with strong general and coding performance.
DeepSeek 64K $0.27 $1.1 use →
Qwen2.5 72B
Strong multilingual open model with excellent Chinese and English ability.
Alibaba 128K $0.35 $0.4 use →

OpenAI-compatible

Point any OpenAI SDK, LangChain or LlamaIndex client at the gateway. Keep your code, change one URL.

Failover included

A rate-limited provider is retried against spare capacity — transparently, with the same request id.

Need a model we don't list?

Tell us at [email protected] and we will wire it into the gateway.