One gateway between your product and every model
Lingyuns is operated by Lingyun Technology inc, a Colorado corporation based in Fruita. We run a unified inference gateway and an on-demand GPU cloud, so a small engineering team can reach frontier capacity without signing five vendor contracts.
An OpenAI-compatible gateway
60+ models behind one endpoint, with automatic failover between providers, per-key metering and a single USD balance for tokens and compute alike.
Product and platform teams
Startups shipping an AI feature this quarter, agencies routing client workloads, and platform teams that need a fallback route when a primary provider degrades.
Prepaid, transparent, no lock-in
Buy credits in USD, spend them across any model at published per-million rates. GPU capacity is billed hourly or as a 720-hour monthly reservation. Nothing auto-renews silently.
How we operate
Your data is not training data
Prompts and completions passing through the gateway are not used to train models. Retention is minimised to what billing and abuse-prevention require, and a zero-retention mode is available on request.
Published prices, no surprise invoices
Every model lists its input and output rate per million tokens on the catalogue page, and every account can see its own metering per model, per day.
Human settlement for high-value orders
Token packs and GPU reservations are confirmed by a billing specialist over email, so enterprises can pay by wire and keep an auditable paper trail.
Portable by design
Because the API is wire-compatible with the OpenAI SDK, moving in — or out — is a one-line change. We would rather earn renewals than lock accounts in.
Corporate details
Work with us
Whether you are shipping a prototype or reserving a cluster, the fastest path is a short message describing the workload.