STREAM BY CAMELAI

Inference,
without a meter.

Unlimited frontier intelligence for $5 a month per stream, with one visible queue per stream. No token metering, no overage charges.

Questions or buying in bulk? Contact us.

How it works

Three steps. No invoice surprises.

01

Subscribe

Create an account at stream.camelai.com and start the founding plan. Your API key is issued immediately — there's no approval step.

02

Point your client

Stream speaks the OpenAI API format. Swap the base URL and model name in the SDK you already use and keep the rest of your code.

03

Generate

Send as much work as you want. Each stream generates one request at a time and the rest wait in a visible queue (bursts of extra parallelism can happen, but only streams guarantee it). The price per stream stays $5.

stream.camelai.com
model auto
context 260K
stream true
tools enabled
Tokens
38.4M
This month
$5.00
Overages
$0.00
streaming responsequeue

What Stream is

One stream. One queue. No meter.

Pricing

Flat price

$5 a month per stream, flat. There is no token allowance behind it and no overage line waiting at the end of the month.

Capacity

A queue you can see

Each stream runs one generation at a time; extra requests wait their turn in a visible queue instead of raising your bill. Add streams for more parallel capacity.

Compatibility

The format you already use

Streaming, tool calling, and structured output through the OpenAI API shape. Your SDK, your framework, your existing code.

The guarantee

How we keep it smart.

You never pick a model on Stream. Keeping the fleet smart is our job, and these are the rules we run it by.

DeepSeek
Gemini
GPT
Qwen
Muse
The fleet changes as new models launch. The floor doesn't.

A public quality bar

Every model we serve scores 70% or higher on Terminal-Bench 2.1, or 50 or higher on the Artificial Analysis Intelligence Index. If a model can't hit those numbers, we don't serve it.

Always the newest version

Whichever model answers you, you get its newest public release. We don't serve old snapshots.

Nothing to run yourself

We run the GPUs and the routing; you point your SDK. Streaming, tool calls, and structured output work on every request, whichever model answers.

The spec sheet

The numbers.

Four numbers decide whether Stream fits your workload: how much context a request can carry, how many requests run at once, how fast tokens arrive, and how long the first one takes. Here's where the fleet stands on each.

Guaranteed01

260K

token context window

Room for a full agent session: long system prompts, tool output, history. Every request gets at least 260K tokens, and more when the model serving you supports it. Past the window we compact the middle of the conversation, never your task or your latest turns.

Guaranteed02

1

generation per stream

How unlimited works without a meter: each stream runs one generation at a time while the rest wait in a visible queue. That's the whole constraint. Bursts of extra parallelism can happen, but only adding streams guarantees it.

Target03

40+

tokens per second

Fast enough to keep an agent loop moving. We tune the fleet so the p10 holds 40 tokens a second or better, and even the p5 stays above 20.

Target04

<5s

to the first token

You're never staring at a silent cursor for long: p95 time to first token is under five seconds.

Guaranteed numbers are part of the deal. Targets are tuning goals we publish, not a service-level agreement; the fine print lives in the terms.

Founding plan

$5/ month per stream
Unlimited tokens One generation per stream Frontier intelligence No overages
Get started

Billed monthly. Cancel anytime. Questions or need more concurrency? Contact sales.

Frequently asked questions

About Stream.

Get in the Stream.

Your API key is a two-minute signup away.

Get your API key