Run Agent-Fit AuditField notesAbout
00How we deliver Voice AI

A voice operator in the cloud or inside your perimeter

The demo shows how the agent sounds. This page is about what stands behind the call: what the operator is built from, the ways it can be delivered, what has to be decided before a price, and what each part costs.

01

What a voice operator is made of

There is no single "operator model" that can be handed over as a file. It is a chain of parts, and each of them works in real time.

  1. 01Telephony

    Takes the call and carries the audio both ways: a SIP trunk, Asterisk or your PBX.

    A call to your number arrives through the carrier and lands in Asterisk, which streams the audio to the agent.

  2. 02Speech recognition

    Turns the caller's voice into text while they are still talking.

    "Hi, I have bedbugs in a two-bed flat" is text before the caller has finished the sentence.

  3. 03Response model

    A language model decides what to say, from the script, your prices and your rules.

    It knows the price list, asks one question at a time, never offers a discount that does not exist, and steers back to the booking.

  4. 04Speech synthesis

    Speaks the answer in the operator's voice, with pauses, breaths and intonation.

    The first words start playing while the model is still writing the end of the sentence.

  5. 05Integrations

    Writes the outcome where your team works with it: CRM, calendar, spreadsheet, messenger.

    After the call a record appears in the CRM: name, phone, request, visit date, price quoted, outcome.

On top of the chain sits conversation control: knowing the caller has finished, not stopping at an "uh-huh", going quiet when interrupted, filling the pause when an answer is slow. The demo has all of it tuned, and that is what separates a conversation from an answering machine.

02

Where the second goes

People expect an answer within about a second of pausing. This is how that second split on a real call to our demo.

end of sentence to voice1.09 s
Detecting you have finished, and network
0.48
Model: first sentence of the answer
0.54
Synthesis: first sound
0.07
you stop talkingthe agent speaks

Measured on a demo call on 9 October 2026, cloud models. For an on-premise install this split has to be measured again on your models and your hardware: it is what decides whether the conversation feels alive.

03

Three ways to deliver it

The question "cloud or our own" settles almost everything else: timeline, budget, hardware, and how close the voice will be to the demo.

Cloud

As in the demo: models at the providers, you pay per minute of conversation.

Where the models run
At the speech and language model providers
Where data and recordings live
At the provider, with configurable retention
Voice and latency
As in the demo
Hardware
None needed
Time to first version
Fastest

When launch speed and voice quality matter, and keeping recordings at a provider is acceptable.

Your perimeter + cloud models

Telephony, conversation control, integrations and data stay with you; models are called over an API.

Where the models run
At the providers, over an API
Where data and recordings live
With you; only what is needed for the answer leaves
Voice and latency
As in the demo
Hardware
A standard server, no GPU
Time to first version
Medium

When data has to stay with you, but calling external models from inside the perimeter is allowed.

Fully on-premise

The whole chain on your servers, with nothing going out.

Where the models run
On your servers
Where data and recordings live
Only with you
Voice and latency
Re-tested: depends on the local models
Hardware
GPU servers
Time to first version
Longest: model selection and load tests

When nothing may leave, and there is budget for the hardware and time for the testing.

04

What has to be decided before a price

Price and timeline come after the specification. Without one, a thousand and one wishes surface along the way. These are the questions that need answers, with examples.

  1. 01

    Where does it run?

    Decides the delivery option, the hardware and half the budget.

    "Everything on our servers, nothing goes out" or "with us, but cloud models over an API are fine".

  2. 02

    Voice only, or a full conversation?

    Speaking a given text and an operator that listens, understands and leads a conversation are systems of very different size.

    "Read out an order status by its number" or "take a request and answer questions".

  3. 03

    What does the operator do?

    Every scenario is a prompt, knowledge, checks and tests. Without examples there is no way to size it.

    "Take a pest-control request, quote from the price list, book a time, and hand over to a person if the caller argues".

  4. 04

    How many simultaneous calls?

    Peak load decides the hardware, or the cloud bill.

    "Up to 10 on a normal day, 30 at peak, 1-2 at night".

  5. 05

    What do we connect to?

    The boundary of the work: what you connect, and what we build.

    "We run Asterisk and give you a SIP trunk; we do the CRM integration ourselves against your API".

  6. 06

    What hardware is there?

    For on-premise: we either fit models to your hardware, or hardware to the quality you need.

    "A rack in a data centre, no GPUs, a purchase budget to be agreed".

  7. 07

    How is the result accepted?

    Without acceptance criteria every change becomes an argument.

    "Ten agreed scenarios pass on our real line, median response no longer than 1.5 seconds, interruption works".

05

Different operators, different work

The word "operator" hides very different amounts of work. Two examples on the same line.

Qualifies and hands over

Finds out who is calling and why, answers routine questions, books or hands over to a person with the context already collected.

AgentBrightHome Pest Control, how can I help?

CallerWe have cockroaches, how much is it?

AgentGot it. Is it a flat? How many bedrooms?

CallerTwo.

AgentFor a two-bed it is a hundred and fifty pounds. Shall I book you in for tomorrow?

One or two scenarios, the price list, handover rules. Integration: a record in the CRM.

Sells and places the order

Handles objections, offers alternatives, checks stock and slots, places the order in your systems and confirms it.

CallerThat is expensive. Others charge less.

AgentI understand. We guarantee it for a year and the follow-up visit is free. If it is a light infestation, there is a cheaper option.

CallerFine, Saturday then.

AgentRight, let me check the technicians... Saturday, eleven in the morning is free. Booking it now.

Objection scripts, live access to schedules and stock, order creation over an API, checks before confirming. Noticeably more testing.

06

Hardware for fully on-premise

If the whole chain has to run with you, it needs compute. Options depend on the load and the model. The baseline for 20-30 simultaneous conversations is option 2, the one we would begin testing from.

Hardware options to buy

ConfigurationSimultaneous calls*Whole serverFits when
1. Pilot: 1 × L40S 48 GB or 1 × RTX PRO 6000 96 GB5-10 / 10-20$20-35kA pilot, one department, a proof of concept
2. Baseline: 2 × L40S, 32 cores, 256 GB, 2 TB NVMe20-30$35-50kProduction with a compact 7-14B parameter model
3. Baseline with headroom: 2 × RTX PRO 6000 96 GB30-50$45-60kA larger 30-70B model, or room to grow in the same box
4. Large: 4 × L40S or 4 × H200 NVL50-100+$55-75k / $160-200kA large model, 50+ calls, several sites on one cluster
5. Budget, used: A100 80 GB15-25 / 2 GPU$8-12k / GPUA tight budget; no FP8, short or no warranty

Rental: 2 × L40S by provider

ProviderPer GPU-hour2 GPUs a monthNote
Vast.ai≥ $0.54≈ $790Marketplace of community hosts, availability varies
Runpod$0.79 / $1.09≈ $1,150 / $1,590Community / Secure Cloud
Scaleway€1.47≈ €2,150A ready 2-GPU configuration, Paris only
Nebius≥ $1.55≈ $2,260Configurations sized by vCPU
OVHcloud$1.80≈ $2,630A ready 2-GPU configuration, Europe
AWS (g6e)$1.86≈ $2,720No 2-GPU size: two 1-GPU instances, or one 4-GPU (≈ $7,660)
CoreWeave$2.25≈ $13,140Only 8-GPU nodes
Immers.cloud—≈ 321 000 ₽ (2 × A100)Data in Russia. No L40S; the nearest with a public price is A100 80 GB
Selectel, Yandex Cloud, Cloud.ru, Timeweb, MTS Cloud——Data in Russia. No public L40S; RTX 6000 Ada, A6000, A100, H100 on request
Lambda, Google Cloud, Azure, Hetzner——No L40S; Google, Azure and Hetzner offer RTX PRO 6000 instead

* Call counts are assumptions, not measurements: they depend on the model size, recognition and synthesis, and are settled by a load test. Prices checked on 9 October 2026, excluding VAT; hardware and memory got dearer all through 2026. In Russia GPU prices run 1.5-2× higher and are usually made to order, 40-60 working days.

Why the number of calls alone does not size the hardware

  • What matters is how the models respond under load, when several people finish a sentence at once and wait.
  • A compact conversational model and a large one need different hardware for the same number of calls.
  • Latency is critical: a three-second answer makes a conversation feel dead, even with few calls.

The order that saves money

  1. 01Assemble the chain on rented hardware.
  2. 02Run load tests with your scenarios and your call peak.
  3. 03Fix the requirements for the box from the results.
  4. 04Only then buy the equipment.
07

Testing on your line

A voice that sounds clean in a browser crosses a mobile network with compression. If a call goes through two mobile legs, the audio degrades twice. We tested this on a model of the line: AMR-NB on each leg and G.711 between them.

The agent's voice
stays intelligible even on a bad network: address, price and phone number come through word for word
The caller's voice
short answers suffer: "a hundred by a hundred" becomes "a hundred". The agent is set to ask again when an answer sounds odd
Latency
each leg adds tens of milliseconds; it goes into the response budget

A line model is an approximation. Quality is settled by calls through your real chain, and that is part of acceptance.

Hear it in the demo on "Two legs" →
08

How the result is accepted

The criteria are fixed before work starts, so that "done" means the same thing to both sides.

Scenarios
An agreed list of conversations passes on your line end to end, including the hard ones: objections, interruptions, changing the subject.
Voice
Quality on the real line with your telephony, not in a studio and not in a browser.
Response speed
Median and 90th percentile from end of sentence to voice, under the agreed load.
Interruption
The agent goes quiet when interrupted, and does not stop for an "uh-huh".
Data
Fields after the call are right - name, phone, request, date, outcome - and land where you work with them.
09

Budget

If we take on the whole virtual call centre: Asterisk, telephony, the AI operator, integration and launch.

First working version
$25-40k
up to 20-30 simultaneous conversations, agreed scenarios and acceptance criteria
Ongoing support
$1.5-3k a month
support, scenario changes, quality monitoring
Hardware or rental
separate
for on-premise, see above
Connectivity and licences
separate
SIP trunk, numbers, cloud models where used

An order of magnitude, not a price list. The exact figure comes once the scope is fixed: an operator that qualifies and hands over and one that sells and places orders cost differently.

10

Why simpler hardware will not do

The questions that come up when people look at the table above and want to save. In short: one or two calls in a test setting run on a lot of things. Twenty or thirty simultaneous conversations answered within a second do not.

Why not a gaming GPU like an RTX 4090 or 5090?+

It has 24-32 GB of memory. One card has to hold the response model, recognition and synthesis, plus memory for every conversation in progress. For 20-30 parallel calls that is not enough, and answers start queuing. NVIDIA's driver licence also forbids GeForce cards in data centres, they have no error-correcting (ECC) memory, and their cooling is not built for 24/7 in a server rack. Fine for a one- or two-call prototype, not for production.

Why not run everything on CPUs, without a GPU?+

A language model on a CPU takes seconds per sentence. Real-time recognition and synthesis for dozens of streams hit the CPU too. The response budget is about a second, half of it for the model. CPUs miss it even with a single call.

What about one small GPU, say an L4 with 24 GB or an older T4?+

For a pilot with a few calls, possibly - it is option 1 from the table, only weaker. The trouble is peaks: when five people finish talking at once, their answers queue, and a second becomes three to five. The T4 also lacks modern compute formats and is noticeably slower. Only worth it if the peak is a couple of calls.

What about a Mac Studio or a mini PC with lots of memory?+

It runs a single model well, and is handy for development. But it is built for one stream: it serves dozens of parallel conversations poorly, because it cannot batch requests the way server GPUs do. And there is no server management, redundancy or ordinary rack operation.

Why not just pick a smaller model that fits cheaper hardware?+

A small model fits, but holds the script worse: it mixes up prices, drifts off topic, agrees to discounts that do not exist, and handles languages worse. We saw this in tests even between cloud models of one tier: one answered briefly and by the rules, another of the same speed slid into monologues. Model size is a trade-off between conversation quality and hardware, and it has to be tested on your scenarios.

Why size hardware for the peak, not the average?+

Calls arrive in bursts: in the morning, after a mailing, at lunch. Latency does not rise gently; it jumps as soon as requests outnumber what the card can process. A server sized for the average answers in three seconds at the peak, which is exactly when most people call.

Can we start cheap and add hardware later?+

Yes, and that is the right path. Rent and load-test first, then option 1 for a pilot. The architecture lets recognition, synthesis and the response model run on separate servers, so growing means adding hardware, not rebuilding the system.

Where to start

Hear the agent first. If this order of budget works for you, the next step is to describe the first delivery and fix what has to work at acceptance.

What it costs: from $99 a month · Which outbound calls does the agent make?

Comparing it with a live answering service for your trade? Answering service by industry.

Call the demoDiscuss a delivery

Assistants that answer your customers, plugged into what you already run. The automation service itself lives on inite.ai.

Email[email protected]
OfficeGlobal Remote Operations

INITE Ecosystem

  • INITE AI
  • INITE Education
  • INITE Club
  • INITE Events
  • INITE Fund
  • INITE Digital
  • INITE Rent
  • INITE Estate
  • INITE Shop
  • INITE Studio
  • INITE Health
  • INITE Travel
  • INITE Sport

Company

  • About
  • Products
  • Field notes
  • Run Agent-Fit Audit
  • How the models are used
  • How we handle AI responsibly

Projects and Solutions

  • Voice AI
  • Sales AI
  • Social AI
  • Omni AI

Automation by INITE

Sold and delivered on inite.ai

  • Services
  • Methodology
  • Service Tiers
  • Case Studies
  • Resources

© 2025 INITE SOLUTIONS. All rights reserved.

Privacy PolicyTerms and ConditionsCookie PolicyData Protection
LanguagesEnglish/Русский/Español/Português