Skip to content

Cost & Throttling ​

Everything your agents do with AI is paid from one place — your prepaid credit balance — and every call is metered, so you can always see what drove a spike.

This page explains what you pay for, how to see where credits go, how to spend less, and what keeps your workspace responsive when traffic surges.

What you pay for ​

WhatHow it's paid
AI models — text and voice replies, knowledge-base indexing and search, summaries, transcription, "Ask AI"Credits. The models are included and managed by Perfox; there's no AI key to connect.
Web voiceIncluded and managed by Perfox; voice usage is metered in credits.
Messaging providers you connect — Plivo (calls, SMS, WhatsApp) and your email providerThose providers bill your own account for their own charges. Perfox's per-message rates, if any, are listed on your rate sheet.

Every new workspace starts with 5,000 free credits, no credit card required. After that you top up with Razorpay whenever you like, or turn on auto-recharge — see Accounts & Billing. There are no plans, seats or tier caps.

Seeing where credits go ​

Perfox meters every AI call — chat replies, voice text and audio, knowledge-base indexing, summaries — and prices it at your rate.

  • Accounts → Usage — this month's usage by channel, and which AI operations consumed the most tokens.
  • Accounts → Transactions — every usage debit; expand one to see the conversations that used it.
  • Estimate usage on Accounts → Billing — project a month's spend at your own rates before you launch.

How to spend less ​

LeverEffect
Model choicePick a lighter model on the agent's Chat Model for everyday traffic; keep heavier models for agents that need them.
Reply lengthShorter replies cost fewer output tokens — and are usually better on voice.
Persona sizeThe persona is sent with every turn; trim it to what matters.
History depthSending fewer past messages each turn means fewer input tokens.
Knowledge-base retrievalFewer, more relevant chunks per answer keep each turn smaller.

Worked example — a cost spike at Acme Support

Setup: Acme's support agent has a 4,000-token persona and sends the last 60 messages with every turn. Traffic has grown steadily.

Action: Asha opens Accounts → Usage, sees that Agent (main turn) dominates, trims the persona to the 800 tokens that matter, reduces history to 20 messages, and moves the agent to a lighter model.

Result: The next week, the same number of conversations uses about a third of the credits.

What just happened: Every turn sends the persona plus recent history to the model. A shorter persona and less history mean fewer tokens on every single turn — the savings add up across every customer message.

When credits run out ​

A balance at or below zero pauses metered actions — agent replies, tests, calls, sends, uploads — until you top up. In the Studio, an Add credits pop-up names the action you tried. Topping up or redeeming a code resumes everything immediately. Browsing and building stay free. See Out of credits.

Load protection ​

Your workspace stays responsive under load:

  • Bounded live turns — conversations are processed in a limited number of parallel slots, so a sudden burst queues briefly instead of overwhelming your agents. A turn waits a short time for a free slot rather than being dropped; voice calls are never held back.
  • Separate lanes for heavy agents — demanding agent runs are processed separately from interactive traffic, so one heavy run never slows your live conversations.

Next steps ​