Pricing calculator
Flat-rate vs per-token AI: what would your app pay?
Per-token pricing grows with every message your users send. A flat plan costs the same however much they use it, but you pay for capacity whether you use it or not. Enter your app's numbers to see which is cheaper for you.
How the calculator works
Per-token cost is your monthly messages (daily users × messages per user × 30) times the tokens each one sends and writes, at the prices you set.
Flat-rate cost depends on how many requests your app needs to run at the same time in its busiest hour, not on tokens. We take the busiest hour's messages, multiply by how long each reply takes to write at our typical speed, and size the plan so that hour is about 60% busy. That leaves room for bursts without long queues. A plan's context size has to fit what you send plus the reply.
When per-token is cheaper
For small apps with light usage, budget per-token models are hard to beat. If the calculator says per-token wins, it probably does, and we'd rather you knew. Flat-rate makes sense once your users send enough that the token bill passes the price of the requests-at-once you need, or when you want a fixed cost you can plan around.
When flat-rate wins
- Chatty users and long conversations. Chat history is re-sent on every message, so token counts climb fast.
- Agents and background jobs. These run many calls per user action and keep a slot busy around the clock.
- Images. A single photo can count as hundreds or thousands of input tokens per request.
- Predictable budgets. A viral week costs the same as a quiet one; extra load waits in a short queue instead of raising the bill.
See the plans, or read the iOS guide if you're building an iPhone app.