Remix.run Logo
▲ ComputerGuru 3 hours ago

Sorry about your father. He needs a dictation model, not a general purpose speech-to-text model. They ignore umms and ahhs, change things like “an elephant, no a monkey, went up the tree” to “a monkey went up the tree,” support saying punctuation aloud sometimes, etc.

Gemini team just released Gemini 3.5 Transcribe that’s supposed to be good at this; it’s available via api: https://blog.google/innovation-and-ai/models-and-research/ge...

▲cgbur 3 hours ago | parent | next [-]

For essentially infinite and fast dictation I use https://github.com/cjpais/Handy on Parakeet streaming (cohere is far better, but slower and has a token output limit so you cant ramble for many minutes). And then just do a cleanup pass with a cheap LLM, it will in my experience, do far better than trying to voice control to go edit a sentence or change words. I just weave instructions into my writing. I understand this requires technical know-how, but for those with it, this is the best solution I have found to long form writing without my hands.

▲flockonus a few seconds ago | parent | next [-]

Agreed, record and transcribe however you can and then use a LLM to clear out.

I prompt it to:

"Attached (or underneath) is the transcript of a self recording i've done with tons of rambling and some incorrect words transcriptions, please do a pass clearing out and arranging any typos or possible misunderstandings. Keep original in parenthesis when not sure if it's a misunderstanding. Do not summarize or alter the nature of the content, simply tidy the transcript."

▲dv35z 2 hours ago | parent | prev | next [-]

Another plug for Handy, and wanted to share something cool about it.

You can set it "Push to talk" mode (like a walkie-talkie radio), and when you're done talking and release the button, it can paste the text into any text field.

You can even replicate ChatGPT voice conversation mode, by having Handy as your speech input, and then (I forgot the extension) enabling a speech-to-text model for OpenCode. Surprisingly relaxing flow for certain tasks, like tweaking a website's styles.

▲Imustaskforhelp 2 hours ago | parent | prev [-]

+1 for handy and then using LLM's for the cleanup pass, though what are your observations on feeling as if sharing that output though?

Because I have seemingly mixed opinions on it, on one hand, I did put the effort but on the other, the output is AI generated so I am unsure about sharing it with others (because they might think its AI generated)

Do you use it for very small edits (removing just the uhhm's?) or for slightly more edits.

The way that I use it sometimes is that while thinking, I will write something which can sometimes make me feel as if a better re-write can better explain my thoughts or rephrasing it as such. For example. I will think about X topic, connect it to Y, then try to add some more points about X again.

I found LLM's to do a really decent job at generating the final outputs as such, but as I said, I am left sometimes feeling a little confused as to sharing it or not because of it being AI generated and the end user not knowing if I put an actual effort into creation of it or not.

Should I try to share the actual transcript of it as well, I really wish if some good ethics and internet ettiquette could be established about it.

▲asa123 an hour ago | parent [-]

what do you mean:

"what are your observations on feeling as if sharing that output though?"

and "Should I try to share the actual transcript of it as well, I really wish if some good ethics and internet ettiquette could be established about it."

could you rephrase the question

for STT, it's literally you saying it, with a model transcribing, and then another model correcting a little, and you also have control over editing it. I don't think any of the arguments on "etiquette re: sharing AI output" apply here.

▲nonmaskable 18 minutes ago | parent [-]

[flagged]

▲nvtop 3 hours ago | parent | prev [-]

I've been using Gemini Desktop App purely for dictation. It's a miracle! For the first time in my life I'm blown away by the quality of my (heavy accent) speech recognition. Just be sure to disable the "speak to window -> reasoning" option to make it purely dictation and stop from writing whole emails for you.