A gateway for large language models.
Starting with Qwen3.8-27B. Stream responses through one API, with transparent per-token pricing and a zero data retention policy. Built for OpenRouter distribution, with direct access by invitation.
OpenRouter integration is planned. Direct API access is invitation-only.
Serving path
Requests pass through the gateway to the inference worker and stream back to you. The gateway does not persist inference content; account, usage, and security metadata is stored separately.
Inference without the overhead
Starting with Qwen3.8-27B
Use qwen/qwen3.8-27b through an OpenAI-compatible chat completions API, with streaming, non-streaming responses, and token-usage reporting.
Clear usage and pricing
Track requests, token usage, and estimated charges in your dashboard. Manage API keys per account and see exactly how your usage adds up.
Zero data retention
Our policy is simple: no persistent storage of your prompts or completions, and no use of your content for model training. Necessary account, usage, and security metadata is retained. See our data policy.
Flat rates per million tokens
Input (uncached)
$0.24
per 1M prompt tokens
Input (cached)
$0.24
per 1M cached prompt tokens
Output
$2.15
per 1M completion tokens
These configured rates are used to estimate usage charges; payment arrangements are agreed separately. Model details and examples →
OpenRouter first, then direct
OpenRouter is our primary planned distribution channel. Integration is not yet live.
Direct API access is available by invitation, with account-scoped keys and a usage dashboard. Commercial arrangements are agreed separately.