Pricing
STREAM BY CAMELAI
Inference,
without a meter.
Unlimited frontier intelligence for $5 a month per stream, with one visible queue per stream. No token metering, no overage charges.
Questions or buying in bulk? Contact us.
How it works
Three steps. No invoice surprises.
Subscribe
Create an account at stream.camelai.com and start the founding plan. Your API key is issued immediately — there's no approval step.
Point your client
Stream speaks the OpenAI API format. Swap the base URL and model name in the SDK you already use and keep the rest of your code.
Generate
Send as much work as you want. Each stream generates one request at a time and the rest wait in a visible queue (bursts of extra parallelism can happen, but only streams guarantee it). The price per stream stays $5.
What Stream is
One stream. One queue. No meter.
Capacity
A queue you can see
Compatibility
The format you already use
The guarantee
How we keep it smart.
You never pick a model on Stream. Keeping the fleet smart is our job, and these are the rules we run it by.
A public quality bar
Every model we serve scores 70% or higher on Terminal-Bench 2.1, or 50 or higher on the Artificial Analysis Intelligence Index. If a model can't hit those numbers, we don't serve it.
Always the newest version
Whichever model answers you, you get its newest public release. We don't serve old snapshots.
Nothing to run yourself
We run the GPUs and the routing; you point your SDK. Streaming, tool calls, and structured output work on every request, whichever model answers.
The spec sheet
The numbers.
Four numbers decide whether Stream fits your workload: how much context a request can carry, how many requests run at once, how fast tokens arrive, and how long the first one takes. Here's where the fleet stands on each.
260K
token context window
Room for a full agent session: long system prompts, tool output, history. Every request gets at least 260K tokens, and more when the model serving you supports it. Past the window we compact the middle of the conversation, never your task or your latest turns.
1
generation per stream
How unlimited works without a meter: each stream runs one generation at a time while the rest wait in a visible queue. That's the whole constraint. Bursts of extra parallelism can happen, but only adding streams guarantees it.
40+
tokens per second
Fast enough to keep an agent loop moving. We tune the fleet so the p10 holds 40 tokens a second or better, and even the p5 stays above 20.
<5s
to the first token
You're never staring at a silent cursor for long: p95 time to first token is under five seconds.
Guaranteed numbers are part of the deal. Targets are tuning goals we publish, not a service-level agreement; the fine print lives in the terms.
Founding plan
Billed monthly. Cancel anytime. Questions or need more concurrency? Contact sales.
Frequently asked questions