Gateway operational · 16 models live · GPU capacity on demand Read the docs →
LINGYUNS gateway
Fast & cheap · Google

Gemini 1.5 Flash

Cost-efficient multimodal model with a one-million-token window.

Context window 1000K tokens
Input price $0.075 per 1M tokens
Output price $0.3 per 1M tokens
Routing Included failover + retries

Call it with any OpenAI SDK

Swap the base URL and key — request shape, streaming and tool calls stay identical.

python
from openai import OpenAI

client = OpenAI(
    base_url="https://api.lingyuns.com/v1",
    api_key="ly-sk-…",
)

r = client.chat.completions.create(
    model="gemini-1-5-flash",
    messages=[{"role": "user", "content": "Hello"}],
)
# billed at $0.075 in / $0.3 out per 1M
Full quickstart

Specifications

Model IDgemini-1-5-flash
ProviderGoogle
CategoryFast & cheap
Context window1000K tokens
Input / 1M$0.075
Output / 1M$0.3
ModalityText · streaming · tools

Prices are list rates in USD per million tokens. Your prepaid credits cover every model in the catalog at these rates.

Alternatives

Other Fast & cheap models

GPT-4o mini

Small, fast and remarkably capable — the default choice for high-volume tasks.

Context128K
Input / 1M$0.15
Output / 1M$0.6

Claude 3.5 Haiku

Fastest Claude — near-instant responses for chat and extraction.

Context200K
Input / 1M$0.8
Output / 1M$4