| ▲ | INTPenis 2 hours ago |
| I don't think the challenge with speech to text was size of the binary. In my experience the challenge is understanding my 84 year old Croatian father with a sagging mouth after a stroke, when he's trying to write his autobiography. I just setup Windows speech to text for him last week and it's great to see how he can write an entire page in 10 minutes, it would take him days using the keyboard. But every single sound he makes with his mouth ends up on the page too. |
|
| ▲ | ComputerGuru 2 hours ago | parent | next [-] |
| Sorry about your father. He needs a dictation model, not a general purpose speech-to-text model. They ignore umms and ahhs, change things like “an elephant, no a monkey, went up the tree” to “a monkey went up the tree,” support saying punctuation aloud sometimes, etc. Gemini team just released Gemini 3.5 Transcribe that’s supposed to be good at this; it’s available via api: https://blog.google/innovation-and-ai/models-and-research/ge... |
| |
| ▲ | cgbur an hour ago | parent | next [-] | | For essentially infinite and fast dictation I use https://github.com/cjpais/Handy on Parakeet streaming (cohere is far better, but slower and has a token output limit so you cant ramble for many minutes). And then just do a cleanup pass with a cheap LLM, it will in my experience, do far better than trying to voice control to go edit a sentence or change words. I just weave instructions into my writing. I understand this requires technical know-how, but for those with it, this is the best solution I have found to long form writing without my hands. | | |
| ▲ | Imustaskforhelp 18 minutes ago | parent [-] | | +1 for handy and then using LLM's for the cleanup pass, though what are your observations on feeling as if sharing that output though? Because I have seemingly mixed opinions on it, on one hand, I did put the effort but on the other, the output is AI generated so I am unsure about sharing it with others (because they might think its AI generated) Do you use it for very small edits (removing just the uhhm's?) or for slightly more edits. The way that I use it sometimes is that while thinking, I will write something which can sometimes make me feel as if a better re-write can better explain my thoughts or rephrasing it as such. For example. I will think about X topic, connect it to Y, then try to add some more points about X again. I found LLM's to do a really decent job at generating the final outputs as such, but as I said, I am left sometimes feeling a little confused as to sharing it or not because of it being AI generated and the end user not knowing if I put an actual effort into creation of it or not. Should I try to share the actual transcript of it as well, I really wish if some good ethics and internet ettiquette could be established about it. |
| |
| ▲ | nvtop an hour ago | parent | prev [-] | | I've been using Gemini Desktop App purely for dictation. It's a miracle! For the first time in my life I'm blown away by the quality of my (heavy accent) speech recognition. Just be sure to disable the "speak to window -> reasoning" option to make it purely dictation and stop from writing whole emails for you. |
|
|
| ▲ | yu3zhou4 2 minutes ago | parent | prev | next [-] |
| [delayed] |
|
| ▲ | boplicity an hour ago | parent | prev | next [-] |
| There are many different challenges, each requiring their own solution. I, for one, really miss the old Google Assistant on my Android phone. It would very reliably play most songs that I wanted to hear on Spotify. Gemini fails at this almost every time, and is significantly slower. It's actually a difficult problem, as the songs people want to hear are regularly being released, are often associated with uncommon names, or have words in unusual orders, so normal LLM style tools just don't cut it. |
| |
| ▲ | xp84 23 minutes ago | parent [-] | | This reminds me of how the thing iPhones had pre-Siri (so we're talking pre-2010), which was entirely offline, did a better job than even the most modern thing at "Play [one of the finite set of songs in my library]." I sometimes get absurd matches from bands I've never heard of, when the right answer is something right there in my library. |
|
|
| ▲ | yymir an hour ago | parent | prev | next [-] |
| i mean for something this small, it can be fit into a l3 cache on a cpu and be essentially always on various purposes |
|
| ▲ | testycool an hour ago | parent | prev [-] |
| Unrelated: I love your username. |