Gateway operational · 16 models live · GPU capacity on demand Read the docs →
LINGYUNS gateway
GPU cloud

NVIDIA capacity, billed by the hour or reserved by the month

Dedicated KVM instances with root access, persistent volumes and prebuilt CUDA images. Boot in minutes; cancel an hourly instance any time.

Class GPU VRAM vCPU / RAM Storage Network Hourly Monthly
RTX 4090 Best value NVIDIA RTX 4090 24 GB 16 / 64 GB 500 GB NVMe 2 Gbps $0.55 $349.00
RTX A6000 NVIDIA RTX A6000 48 GB 24 / 96 GB 1000 GB NVMe 4 Gbps $0.89 $569.00
L40S Inference NVIDIA L40S 48 GB 32 / 128 GB 1000 GB NVMe 4 Gbps $1.29 $799.00
A100 80GB NVIDIA A100 SXM 80 GB 48 / 256 GB 2000 GB NVMe 8 Gbps $1.79 $1,149.00
H100 80GB Fastest NVIDIA H100 SXM 80 GB 64 / 320 GB 2000 GB NVMe 10 Gbps $2.99 $1,899.00
H200 141GB Frontier NVIDIA H200 SXM 141 GB 96 / 512 GB 4000 GB NVMe 10 Gbps $3.99 $2,499.00

Monthly reservations assume 720 hours (30 days) — reserve and the effective rate drops to the monthly column.

RTX 4090
NVIDIA RTX 4090
Best value
$0.55 / hour · $349.00 / month

Best price/performance for 7B–14B inference and fine-tuning.

VRAM24 GB
vCPU16
RAM64 GB
NVMe500 GB
  • 1× RTX 4090 24GB GDDR6X
  • 16 vCPU · 64GB RAM · 500GB NVMe
  • 2 Gbps network · static IP
Reserve RTX 4090
RTX A6000
NVIDIA RTX A6000
$0.89 / hour · $569.00 / month

48GB cards for larger context inference and multi-model serving.

VRAM48 GB
vCPU24
RAM96 GB
NVMe1000 GB
  • 1× RTX A6000 48GB GDDR6
  • 24 vCPU · 96GB RAM · 1TB NVMe
  • 4 Gbps network · static IP
Reserve RTX A6000
L40S
NVIDIA L40S
Inference
$1.29 / hour · $799.00 / month

Purpose-built for generative AI inference at scale.

VRAM48 GB
vCPU32
RAM128 GB
NVMe1000 GB
  • 1× L40S 48GB GDDR6 ECC
  • 32 vCPU · 128GB RAM · 1TB NVMe
  • 4 Gbps network · static IP
Reserve L40S
A100 80GB
NVIDIA A100 SXM
$1.79 / hour · $1,149.00 / month

The workhorse for training and large-model serving.

VRAM80 GB
vCPU48
RAM256 GB
NVMe2000 GB
  • 1× A100 80GB HBM2e
  • 48 vCPU · 256GB RAM · 2TB NVMe
  • 8 Gbps network · NVLink ready
Reserve A100 80GB
H100 80GB
NVIDIA H100 SXM
Fastest
$2.99 / hour · $1,899.00 / month

Fastest transformer engine for frontier training runs.

VRAM80 GB
vCPU64
RAM320 GB
NVMe2000 GB
  • 1× H100 80GB HBM3
  • 64 vCPU · 320GB RAM · 2TB NVMe
  • 10 Gbps network · NVLink ready
Reserve H100 80GB
H200 141GB
NVIDIA H200 SXM
Frontier
$3.99 / hour · $2,499.00 / month

Frontier capacity for the largest models and context windows.

VRAM141 GB
vCPU96
RAM512 GB
NVMe4000 GB
  • 1× H200 141GB HBM3e
  • 96 vCPU · 512GB RAM · 4TB NVMe
  • 10 Gbps network · NVLink ready
Reserve H200 141GB

Bring your own stack

Root SSH, Docker, CUDA 12.4 and PyTorch 2.4 images, or upload your own. Jupyter available on every instance.

Clusters on request

Multi-node NVLink and InfiniBand fabrics for training runs — talk to sales for sizing and reserved pricing.

Pay how you already pay

Same invoice flow as tokens: wire or USDT, USD-denominated, activated by a billing specialist.