OpenAI-compatible inference

A gateway for large language models.

Starting with Qwen3.8-27B. Stream responses through one API, with transparent per-token pricing and a zero data retention policy. Built for OpenRouter distribution, with direct access by invitation.

OpenRouter integration is planned. Direct API access is invitation-only.

Built for developers

Inference without the overhead

Starting with Qwen3.8-27B

Use qwen/qwen3.8-27b through an OpenAI-compatible chat completions API, with streaming, non-streaming responses, and token-usage reporting.

Clear usage and pricing

Track requests, token usage, and estimated charges in your dashboard. Manage API keys per account and see exactly how your usage adds up.

Zero data retention

Our policy is simple: no persistent storage of your prompts or completions, and no use of your content for model training. Necessary account, usage, and security metadata is retained. See our data policy.

Pricing

Flat rates per million tokens

Input (uncached)

$0.24

per 1M prompt tokens

Input (cached)

$0.24

per 1M cached prompt tokens

Output

$2.15

per 1M completion tokens

These configured rates are used to estimate usage charges; payment arrangements are agreed separately. Model details and examples →

Distribution

OpenRouter first, then direct

OpenRouter is our primary planned distribution channel. Integration is not yet live.

Direct API access is available by invitation, with account-scoped keys and a usage dashboard. Commercial arrangements are agreed separately.

Service status: check current gateway availability on the status page.
Your data: zero data retention applies to inference content, not the operational metadata needed to run your account. Read our privacy policy, terms, and data policy.
Availability

Service at a glance

Gateway live & monitored Direct access by invitation OpenRouter: planned
Checking live gateway status…