Integrated Toolsets speech to text

Discussion about components and system integration.

One thing that I really enjoyed about Apple operating system is that the Speech to text input is excellent. I’m hoping to find another service that works with Wayland, I found https://voxtype.io/ which uses a model from hugging face, so in that instance is there a really good model to consider for speech to text, using a microphone input with minimal background noise.

The service works as expected and with some configuration settings it has most of the features I want. I would prefer if it inserted at cursor, however I did not see an easy way to configure that.

Good news is the focus configuration has worked inside of other applications.

I changed the model, and configured gpu support, which both have improved the experience. It’s not quite perfect, but it fills the gap.

What model are you using to make it usable?

“small.en” is about 95% reliable. However, there is some delay after the recording when it converts to processing. I’d really like it if I could get the cursor to change it’s configuration to show me when the transcription is ready.

It is too slow, I wish it were as fast as the translation on my apple devices. I like it in theory but I cannot use it to dictate.

Is it possible to make it faster?

Right now a dictated paragraph takes more than 30 seconds to convert to text. The model is exceptionally accurate in my use, just slow.