Skip to main content
Lokutor is paid from a prepaid balance in US dollars. You top up, and every session and request is charged against the balance as you use it: agent time and audio per second, the language model per token, phone lines per started minute. There is no plan, no monthly fee and no minimum. Prices are before tax. The rate card is published as JSON at GET https://api.lokutor.com/v1/pricing, with no key needed. The figures on this page are from version 2026-10-09.1; if the two ever disagree, the API is right.

The rate card

Nothing is rounded up to the minute except the phone line, which the carrier bills by the started minute. Usage is added up in fractions of a cent over a session and rounded once, when the session is charged.

What the agent minute covers

The platform price covers speech recognition, the voice, turn-taking, noise removal, orchestration and call recording. A session is billed from the moment the connection is accepted (web, app, SDK) or the call is answered (phone) until it closes or hangs up. A session left open with nobody talking is still a session: close the connection when the conversation is over.

Language model tokens

The language model is billed separately, by the token, at one price whichever model provider answers. Every request the agent makes to the model during a session counts: Each request sends the model the agent’s instructions and the conversation so far, so input tokens make up most of the count, and an agent with a long system prompt spends more of them on every turn. A typical minute of conversation uses about 20,700 input and 1,300 output tokens, which is about 1¢, so a typical minute comes to about 5.2¢ in all. Counts come from the model provider’s own report. When a provider reports none for a request that was cancelled, its input tokens are estimated and the line is marked as estimated in your usage. An agent set up with your own language-model provider and key is not billed tokens: you pay your provider directly. The platform minute still applies.

Speculative replies

With speculative replies on, the agent starts its answer while the caller pauses, before it is sure they have finished. When the caller has finished, the reply is already on its way: turns it lands in are answered in about 306 ms instead of about 660 ms. When the caller keeps talking, that early reply is thrown away, and its tokens are billed as speculative replies not used. It is on by default, and costs about a third of a cent a minute at a typical prompt (more with a long prompt, since each unused reply resends it). To turn it off for an agent, set speculative_llm to false in the agent’s config. The agent then asks the model only once the caller’s turn has ended, and its replies start about 255 ms later on average.

Phone calls

A phone call is billed as agent time and tokens, like any session, plus the phone line per started minute: a call of 61 seconds is two minutes of line. The rate depends on the direction and country: the country of your number for a call your agent receives, the destination for a call it makes. Prices in USD per minute. Mobiles in Portugal, Italy, Switzerland, Belgium, the Netherlands and Luxembourg cannot be called: their carrier rates are above the ceiling Lokutor allows, and dialling one is refused before anything is charged. In the US and Canada, mobile numbers cannot be told apart from landlines and are charged the landline rate. A call that is not answered costs nothing; a call answered by voicemail is an answered call. Phone numbers are local numbers in the US, Spain, the UK, Germany, France, Italy and Portugal, at $5 a month from the balance. The first month is charged when you buy the number, then once a month. If the balance cannot pay a month, the number is kept for 30 days; after that it is released, and a released number cannot be recovered. Calls made by your agent and phone numbers need a first top-up: the welcome credit does not pay for them.

APIs

  • Text-to-speech (/tts/synthesize, /ws/tts) is billed by the characters of text synthesized. If you stop a stream midway, the characters already synthesized are billed.
  • Speech-to-text (/stt/transcribe, /ws/stt) and noise removal (/denoise) are billed by the seconds of audio you send.
  • A request that fails on our side is not billed.

Worked examples

All at the rate card above, with the language model at a typical prompt. A 3-minute call your agent answers on a Spanish number: At 3 minutes 10 seconds the line is 4 started minutes, $0.035 more. The same call on a US number is about 7.1¢ a minute; a call your agent makes to a Spanish mobile, about 15.7¢. 1,000 minutes of conversations on the web or in your app: 42.00platform+about42.00 platform + about 10.33 of tokens (20.7 million input, 1.3 million output) = **about 52.33∗∗.Withspeculativerepliesoff,about52.33**. With speculative replies off, about 48.79. A receptionist answering 300 minutes of calls a month on a Spanish number: 300 minutes at about 8.7¢, plus the number: about $31.20 a month. Reading 1 million characters aloud with the text-to-speech API: 4.10∗∗.Transcribinganhourofaudio:∗∗4.10**. Transcribing an hour of audio: **0.084. Your own bill depends on your prompt: the dashboard’s Usage page shows what each session actually cost, line by line.

Your balance

Topping up

Top up from Billing in the dashboard, from 10to10 to 2,000 per payment. For a larger amount, write to [email protected].
  • Currency. The balance is in US dollars. At checkout you pay in your own currency, converted by Stripe, or you can choose to pay in US dollars. The dashboard shows your balance in dollars with an approximate equivalent in your currency.
  • Tax. VAT or sales tax is added at checkout, on top of the amount: a 10top−upbyanindividualinSpaincosts10 top-up by an individual in Spain costs 12.10 and adds $10 to the balance. A business can enter its VAT ID at checkout; a valid EU VAT ID outside Spain is reverse-charged.
  • Invoices. Ask for an invoice when you top up. Every payment also gets Stripe’s receipt by email.

Auto-recharge

Optional, and off until you turn it on in Billing. When a charge leaves the balance below a threshold, your saved card is charged a fixed amount: by default, below 5,add5, add 20, with a monthly cap of $500, after which auto-recharge stops until the next month and you are told. Auto-recharge is charged in US dollars; your bank converts it. If your bank asks for 3-D Secure, you get an email link to approve the payment. A declined card is retried up to three times.

Welcome credit

New accounts get $5 of credit once their email is confirmed (or on signing up with Google or GitHub), valid for 90 days: about 95 minutes of a typical web conversation. Until your first top-up, the welcome credit does not pay for calls your agent makes to phones or for phone numbers, and the account can hold 3 sessions at a time. Accounts registered before the balance existed get a credit when they move over.

Expiry and refunds

  • Purchased balance never expires. The unused part of a top-up is refunded if you ask within 14 days of it, to [email protected], in the currency you paid.
  • Credits (the welcome credit, promotional codes) expire as stated and are not refundable. Credit that expires soonest is used first; purchased balance is used last.

Itemised usage

Every session and request is itemised in the dashboard (Usage): platform time, model tokens by category, phone line, extras and API usage, with the rate card version it was charged at.

Price changes

The rate card is versioned. A session is charged entirely at the version in force when it started. Price increases are announced 30 days ahead; decreases apply at once.

Limits

Opening a session past your concurrency limit is refused with 429 auth.rate_limited (“Too many concurrent sessions”); close one and retry.

When the balance runs out

  • A call in progress is never cut for money mid-sentence. It continues while the account’s balance goes as low as -$2 (about 38 typical minutes, shared by all of the account’s calls in progress). Past that, the agent apologises, says goodbye and hangs up; it does not tell the caller why. On the web, the server closes the socket with code 1008 and the reason account balance exhausted a few seconds after the goodbye (30 seconds after asking for it, if the agent has not said it by then).
  • New sessions and API requests are refused with 402 while the balance is under 0.25.Thetext−to−speech,speech−to−textandnoise−removalAPIshavenooverdraft:theyarerefusedassoonasthebalanceisunder0.25. The text-to-speech, speech-to-text and noise-removal APIs have no overdraft: they are refused as soon as the balance is under 0.25.
  • A call to your number while the balance is under $0.25 is not answered.
  • A negative balance is paid back first by your next top-up.

Sessions nobody speaks on

A session is billed for as long as it is open, so one that nobody is using is ended for you: when neither the caller nor the agent has spoken for 5 minutes, the agent says it is ending the call and goodbye, and hangs up. On the web, the server then closes the socket with code 1000 and the reason nobody has spoken for 5 minutes. This covers a browser tab left open and a phone line nobody hung up. A conversation with pauses is never affected: any speech, from either side, starts the 5 minutes again.

Billing errors

None of these is retryable as is. On REST endpoints they arrive in the usual error format. The Call Center API answers {"code": "...", "error": "..."} with the same codes. On the WebSocket endpoints (/ws/agent, /ws/tts, /ws/stt) the refusal comes before the connection opens: the handshake answers 402 with the message as plain text, and a browser only sees the connection fail (close code 1006). If a WebSocket connection fails at the handshake, check the balance in the dashboard.