| ▲ | MitziMoto 10 hours ago | |||||||||||||
We use a Twilio integration to make it work over a real phone line, but it could just as well be an app or websocket connection in a browser. Voice transport medium aside, the actual use case is not quite what you described. Companies like Eleven Labs and Vapi provide a full end to end voice agent platform. They handle the STT -> LLM -> TTS pipeline and infrastructure for voice agents. Think of a customer talking to a virtual receptionist to schedule an appointment. On the ElevenLabs platform, you provide them a system prompt (or an entire workflow/graph of system prompts) that instruct the voice agent on how to talk to the customer, tone, guardrails, how to answer specific questions, etc. At some point we need that LLM agent, running on an infrastructure we don't control, to talk to our "CRM" (for simplicity sake). Enter MCP. The MCP server we build and host supplies the "tools", like list_appointments, schedule_appointment, cancel_appointment-- whatever they may be. When eleven Labs voice agent connects, it sees all the tools available and can use them per the instructions in the system prompt. | ||||||||||||||
| ▲ | jdironman 9 hours ago | parent [-] | |||||||||||||
How is security handled in a situation like this? (I'll end up reading up on it). But I mean things like how does the API let a specific session with a specific user only surface information for that specific user from the available tooling? | ||||||||||||||
| ||||||||||||||