The questions that come up when people look at the table above and want to save. In short: one or two calls in a test setting run on a lot of things. Twenty or thirty simultaneous conversations answered within a second do not.
Why not a gaming GPU like an RTX 4090 or 5090?+
It has 24-32 GB of memory. One card has to hold the response model, recognition and synthesis, plus memory for every conversation in progress. For 20-30 parallel calls that is not enough, and answers start queuing. NVIDIA's driver licence also forbids GeForce cards in data centres, they have no error-correcting (ECC) memory, and their cooling is not built for 24/7 in a server rack. Fine for a one- or two-call prototype, not for production.
Why not run everything on CPUs, without a GPU?+
A language model on a CPU takes seconds per sentence. Real-time recognition and synthesis for dozens of streams hit the CPU too. The response budget is about a second, half of it for the model. CPUs miss it even with a single call.
What about one small GPU, say an L4 with 24 GB or an older T4?+
For a pilot with a few calls, possibly - it is option 1 from the table, only weaker. The trouble is peaks: when five people finish talking at once, their answers queue, and a second becomes three to five. The T4 also lacks modern compute formats and is noticeably slower. Only worth it if the peak is a couple of calls.
What about a Mac Studio or a mini PC with lots of memory?+
It runs a single model well, and is handy for development. But it is built for one stream: it serves dozens of parallel conversations poorly, because it cannot batch requests the way server GPUs do. And there is no server management, redundancy or ordinary rack operation.
Why not just pick a smaller model that fits cheaper hardware?+
A small model fits, but holds the script worse: it mixes up prices, drifts off topic, agrees to discounts that do not exist, and handles languages worse. We saw this in tests even between cloud models of one tier: one answered briefly and by the rules, another of the same speed slid into monologues. Model size is a trade-off between conversation quality and hardware, and it has to be tested on your scenarios.
Why size hardware for the peak, not the average?+
Calls arrive in bursts: in the morning, after a mailing, at lunch. Latency does not rise gently; it jumps as soon as requests outnumber what the card can process. A server sized for the average answers in three seconds at the peak, which is exactly when most people call.
Can we start cheap and add hardware later?+
Yes, and that is the right path. Rent and load-test first, then option 1 for a pilot. The architecture lets recognition, synthesis and the response model run on separate servers, so growing means adding hardware, not rebuilding the system.