watched the live translation video very impressive
Seems like a shift from previous voice models where it sequentially processes voice to text then feeds it to LLM and then back which cant escape the clunky lag
not sure how pipecat stands now, gpt live seems like it takes audio tokens and does inference on it directly
I've wondered if this would happen, although doing inference directly on speech tokens would seem to imply an entirely different model (trained on lots and lots of actual speech).
Seems like a shift from previous voice models where it sequentially processes voice to text then feeds it to LLM and then back which cant escape the clunky lag
not sure how pipecat stands now, gpt live seems like it takes audio tokens and does inference on it directly