Cost & Throttling
Everything your agents do with AI is paid from one place — your prepaid credit balance — and every call is metered, so you can always see what drove a spike.
This page explains what you pay for, how to see where credits go, how to spend less, and what keeps your workspace responsive when traffic surges.
What you pay for
| What | How it's paid |
|---|---|
| AI models — text and voice replies, knowledge-base indexing and search, summaries, transcription, "Ask AI" | Credits. The models are included and managed by Perfox; there's no AI key to connect. |
| Web voice | Included and managed by Perfox; voice usage is metered in credits. |
| Messaging providers you connect — Plivo (calls, SMS, WhatsApp) and your email provider | Those providers bill your own account for their own charges. Perfox's per-message rates, if any, are listed on your rate sheet. |
Every new workspace starts with 5,000 free credits, no credit card required. After that you top up with Razorpay whenever you like, or turn on auto-recharge — see Accounts & Billing. There are no plans, seats or tier caps.
Seeing where credits go
Perfox meters every AI call — chat replies, voice text and audio, knowledge-base indexing, summaries — and prices it at your rate.
- Accounts → Usage — this month's usage by channel, and which AI operations consumed the most tokens.
- Accounts → Transactions — every usage debit; expand one to see the conversations that used it.
- Estimate usage on Accounts → Billing — project a month's spend at your own rates before you launch.
How to spend less
| Lever | Effect |
|---|---|
| Model choice | Pick a lighter model on the agent's Chat Model for everyday traffic; keep heavier models for agents that need them. |
| Reply length | Shorter replies cost fewer output tokens — and are usually better on voice. |
| Persona size | The persona is sent with every turn; trim it to what matters. |
| History depth | Sending fewer past messages each turn means fewer input tokens. |
| Knowledge-base retrieval | Fewer, more relevant chunks per answer keep each turn smaller. |
Worked example — a cost spike at Acme Support
Setup: Acme's support agent has a 4,000-token persona and sends the last 60 messages with every turn. Traffic has grown steadily.
Action: Asha opens Accounts → Usage, sees that Agent (main turn) dominates, trims the persona to the 800 tokens that matter, reduces history to 20 messages, and moves the agent to a lighter model.
Result: The next week, the same number of conversations uses about a third of the credits.
What just happened: Every turn sends the persona plus recent history to the model. A shorter persona and less history mean fewer tokens on every single turn — the savings add up across every customer message.
When credits run out
A balance at or below zero pauses metered actions — agent replies, tests, calls, sends, uploads — until you top up. In the Studio, an Add credits pop-up names the action you tried. Topping up or redeeming a code resumes everything immediately. Browsing and building stay free. See Out of credits.
Load protection
Your workspace stays responsive under load:
- Bounded live turns — conversations are processed in a limited number of parallel slots, so a sudden burst queues briefly instead of overwhelming your agents. A turn waits a short time for a free slot rather than being dropped; voice calls are never held back.
- Separate lanes for heavy agents — demanding agent runs are processed separately from interactive traffic, so one heavy run never slows your live conversations.
Next steps
- Accounts & Billing — balance, top-ups, usage and invoices
- Chat Model — choose the model for each agent