Voice AI industry still uses transcripts to determine the emotions of the user.
So when me and my friend were discussing about this we got the known limitations of transcriptions is that it doesn't capture the user's emotion from audio. Like, "I'm frustrated with your service" reads the same whether calm or angry.
Then I thought, why can't the decision model be tested for this to get quicker decisions on the go, knowing the limitation that it can't accept the audio file directly? So I thought of converting it to a pictorial representation and then submitting it to the clef-flash [a cheaper option]. Although the results are not too exciting, I would say it is just starting, and we can fine-tune the decision model for the same use case, and it can help a lot in several scenarios.
Worth giving it a try and checkout on your own voice.
Comments URL: [https://news.ycombinator.com/item?id=50017751](https://news.ycombinator.com/item?id=50017751)
Points: 1