Remix.run Logo
▲ crazygringo 4 hours ago

I think we're about to see a dictation revolution.

I've been coding with Claude, and I pretty much do it exclusively by voice now, simply because it's so much faster. And I kind of feel like Scotty from Star Trek IV, when he holds up a Mac mouse to try to talk into it.

I've been using VoiceInk [1] which looks like it's basically the same as this, but has been around for longer.

What has really made it work for me is using a Bluetooth media control like [2], where I use Karabiner Elements to remap its play/pause button to the dictation keyboard shortcut. I map rewind to option-backspace to delete the last word, fast-forward to shift-enter to insert a newline, a lower button to enter to submit my prompt, and the volume up/down buttons to scroll up/down. I've also just ordered to Xiaomi Bluetooth remote [3] that includes a microphone itself to see if I can get it to work by speaking directly into it and using it as the microphone -- there are a couple of open source projects to turn it into a Mac microphone directly. Since I'd like to be able to talk more quietly instead of into my Mac.

But what I'm REALLY waiting for is the ability to use one button for dictation, and a second button for issuing commands for a local LLM to interpret. So I hold down the voice button which transcribes "I think we need to catch that" and then the second button and go "change catch to cache, like c-a-c-h-e". Or hold down the second button and go "switch to VS code". I don't want something as finicky as macOS Voice Control, I want a local LLM I can speak naturally to.

I can very much see a future where I spend the majority of my "work" time using a Apple TV-type remote.

[1] https://github.com/Beingpax/VoiceInk

[2] https://www.amazon.com/Satechi-Bluetooth-Multimedia-Remote-C...

[3] https://www.notebookcheck.net/Xiaomi-Bluetooth-Remote-2-Pro-...

▲polyterative 3 hours ago | parent | next [-]

I am having great results by mapping the press to dictate onto my pedal.I have been vibecoding without using my hands for the past four months with great results.It's amazing using the computer without touching it.I can go on for hours with a couple of pedals.I use them to switch windows, dictate and press enter.That's really all you need.

▲patrickk an hour ago | parent | prev | next [-]

Great ideas here thanks! I’m going to see if I can whip up a Windows equivalent, using Autohokey instead of Karabiner.

▲vardalab 2 hours ago | parent | prev | next [-]

The one thing about the Bluetooth microphones is that there is this annoying lag that it's very hard to get used to, even when using AirPods which you think would be optimized for this. It's just untenable in my opinion because I could never get used to the 3-400 millisecond delay that I had to deal with. I ended up using simple Apple wired headphones as my go-to solution eventually. They actually produce a least a modest strain on ears when being in for hours on end, at least in my situation. I do like your idea about the Bluetooth multimedia remote controller.

▲ 6 minutes ago | parent | next [-]
[deleted]
▲kartik017 2 hours ago | parent | prev [-]

i have fixed this problem in betterwispr btw ;)

▲pegasus 2 hours ago | parent [-]

You fixed bluetooth lag??

▲Grombobulous 4 hours ago | parent | prev | next [-]

I can see the appeal from a workflow perspective, though from a “social” perspective I don’t see myself enjoying talking to myself all day.

Bonus points if you work in an office. This kind of workflow would be a nightmare.

…unless it means we get our private offices back.

▲dnautics 2 hours ago | parent [-]

I have a friend who has reporposed a stenographer's mask

▲cjonas 4 hours ago | parent | prev | next [-]

i forked handy and built 2 modes. One for dictation and one for quick agent actions that uses the pi harness with local models and jev style classifiers so it can read screens and interact with elements (via accessiblity tree, screenshots, os scripts, bash, mcp, etc). Its nice because in agent mode you can just give it instructions on how to respond (like dictation that can read the context of the current page or input). I've been pretty happy with it so far and debating if i should release it. Just not sure if the world needs yet another vibe coded agent system.

▲Razengan 4 hours ago | parent | prev [-]

I'm glad we're getting to a point where voice can be the main UI, and that some people are actively using it to their benefit, but I hope it doesn't become THE primary interface, even with AIs.

I find speaking tiresome, somehow, and if I had to talk all the time to my computer with text input being a fidgety "accessibility" option, I'd retire to a monastery and just become a monk.

Agree with the need for different buttons for different voice contexts though.

Though I think it will become a dedicated AI key on some keyboards.. How about the Right Alt/Option on MacBooks? Does anybody use that? It could be the "AI Anywhere" input key, optionally defaulting to voice, and while F5 could continue to serve as a literal dictation key..

▲pmoriarty 2 hours ago | parent | next [-]

Speech also has a lack of privacy, and it can annoy people nearby.

Imagine everyone on public transport speaking in to their phones to navigate and control them.

A microphone sensitive to subvocalization could maybe get around some of these problems, but it'd still be kind of weird.

▲dnautics 2 hours ago | parent | prev [-]

I upload documents to a remarkable, mark it up, and have Claude build changes there.