About
Large language models. One gateway.
Prefill Systems is an inference provider starting with Qwen3.8-27B: an OpenAI-compatible API, transparent pricing, and a zero data retention policy.
Why we exist
We make language-model inference straightforward to integrate and operate. OpenRouter is our primary planned distribution channel, with direct API access for invited customers.
Direct access
Invited customers can generate API keys, track token usage, and review estimated charges in their dashboard. Direct commercial arrangements are agreed separately.
Our service
| Fact | Status |
|---|---|
| Public gateway | Live, monitored, invitation-only |
| OpenRouter integration | Planned |
| Inference data policy | Zero data retention — no persistent prompt or completion storage |
| SLA / uptime commitment | None yet — see the status page |
Built around your workflow
Use familiar OpenAI-compatible clients, stream responses as they arrive, and manage access through account-scoped API keys. See models and pricing for the current offering or developer docs to integrate.
Principles
- Transparent pricing. Per-token rates and account-level usage reporting make costs easy to follow.
- Zero data retention. No persistent storage of prompts or completions. Necessary operational metadata is retained separately.
- Controlled access. Create and revoke API keys from your dashboard, with usage attributed to your account.