Gateway operational · 16 models live · GPU capacity on demand Read the docs →
LINGYUNS gateway
About

One gateway between your product and every model

Lingyuns is operated by Lingyun Technology inc, a Colorado corporation based in Fruita. We run a unified inference gateway and an on-demand GPU cloud, so a small engineering team can reach frontier capacity without signing five vendor contracts.

WHAT WE RUN

An OpenAI-compatible gateway

60+ models behind one endpoint, with automatic failover between providers, per-key metering and a single USD balance for tokens and compute alike.

WHO IT IS FOR

Product and platform teams

Startups shipping an AI feature this quarter, agencies routing client workloads, and platform teams that need a fallback route when a primary provider degrades.

HOW WE CHARGE

Prepaid, transparent, no lock-in

Buy credits in USD, spend them across any model at published per-million rates. GPU capacity is billed hourly or as a 720-hour monthly reservation. Nothing auto-renews silently.

Principles

How we operate

Your data is not training data

Prompts and completions passing through the gateway are not used to train models. Retention is minimised to what billing and abuse-prevention require, and a zero-retention mode is available on request.

Published prices, no surprise invoices

Every model lists its input and output rate per million tokens on the catalogue page, and every account can see its own metering per model, per day.

Human settlement for high-value orders

Token packs and GPU reservations are confirmed by a billing specialist over email, so enterprises can pay by wire and keep an auditable paper trail.

Portable by design

Because the API is wire-compatible with the OpenAI SDK, moving in — or out — is a one-line change. We would rather earn renewals than lock accounts in.

Corporate details

Registered nameLingyun Technology inc
Entity typeCorporation — State of Colorado, USA
Entity ID20251050668
Registered office242 N Ash Street, Fruita, CO 81521, United States

Work with us

Whether you are shipping a prototype or reserving a cluster, the fastest path is a short message describing the workload.