Simon Moe
“For a voice agent company, they want to control their own model so that it can make sure the model actually responds by the required time. So the customer, when they're on the phone, they can ensure the agent is responding according to a SLA. And this sometimes is only you can do with your controlled intelligence because you know the whole hardware you're running and the whole system you're monitoring versus signing up for relying on your critical infrastructure with a proprietary API where they might go down anytime.”
Source→“vLLM is an inference engine. That means its job is to turn available GPUs into a running endpoint for intelligence... we support more than a thousand model architectures up to today.”
Source→AI-extracted from podcast / newsletter / paper summaries. May contain errors.