| ▲ | Whistle: Speech to Text in 16.9 MB(cactuscompute.com) | ||||||||||||||||||||||||||||||||||||||||
| 133 points by gmays 2 hours ago | 35 comments | |||||||||||||||||||||||||||||||||||||||||
| ▲ | albert_e 12 minutes ago | parent | next [-] | ||||||||||||||||||||||||||||||||||||||||
What the demo does not do is show streaming output of transcribed text as we are speaking and recording (before we hit stop). That is an essential feature IMO for most general purpose live STT apps. | |||||||||||||||||||||||||||||||||||||||||
| ▲ | INTPenis an hour ago | parent | prev | next [-] | ||||||||||||||||||||||||||||||||||||||||
I don't think the challenge with speech to text was size of the binary. In my experience the challenge is understanding my 84 year old Croatian father with a sagging mouth after a stroke, when he's trying to write his autobiography. I just setup Windows speech to text for him last week and it's great to see how he can write an entire page in 10 minutes, it would take him days using the keyboard. But every single sound he makes with his mouth ends up on the page too. | |||||||||||||||||||||||||||||||||||||||||
| |||||||||||||||||||||||||||||||||||||||||
| ▲ | wkcheng 12 minutes ago | parent | prev | next [-] | ||||||||||||||||||||||||||||||||||||||||
How does this compare with Parakeet? I've been using that locally in my projects on an M-series macbook and it's been working great. It's fast and accurate enough for my use cases (meeting transcription, audio transcription for demo videos, etc.) This definitely seems lighter and faster. How does accuracy compare? | |||||||||||||||||||||||||||||||||||||||||
| ▲ | e12e 10 minutes ago | parent | prev | next [-] | ||||||||||||||||||||||||||||||||||||||||
Hm. I saw language=detect and tried some Japanese - which (given the actual list of supported languages) unsurprisingly turned into some mangled Spanish. Since it doesn't support Norwegian - I tried English - and it mis-transcribed "cleaning" for "training" - probably a failure due to context/training (Hello everyone, today we are going to do some cleaning). So, reasonable, but limited? | |||||||||||||||||||||||||||||||||||||||||
| ▲ | joewhale an hour ago | parent | prev | next [-] | ||||||||||||||||||||||||||||||||||||||||
I initially read this as whistle to text, which would be way cooler. | |||||||||||||||||||||||||||||||||||||||||
| |||||||||||||||||||||||||||||||||||||||||
| ▲ | jayshah5696 16 minutes ago | parent | prev | next [-] | ||||||||||||||||||||||||||||||||||||||||
This is actually a really great release. Congratulations team. I just tried few words. My Indian accent also was able to pick up.I'm gonna run it on my Linux Box. | |||||||||||||||||||||||||||||||||||||||||
| ▲ | andy_ppp an hour ago | parent | prev | next [-] | ||||||||||||||||||||||||||||||||||||||||
Wow certainly in English this is incredibly accurate I tried to break it and it understood me perfectly! I know it's slightly off topic but surely it must be easy by now to train a spell checker that doesn't annoy the crap out of everyone using it (looking at you here Apple)! | |||||||||||||||||||||||||||||||||||||||||
| ▲ | kamranjon 42 minutes ago | parent | prev | next [-] | ||||||||||||||||||||||||||||||||||||||||
Sooo I haven't really been super impressed with the needle models before, but this is very impressive. It transcribed multiple sentences I gave it with complex timing and words and in such a small footprint, I'm super impressed. Excited to see what types of things can be built with something like this, the performance seems very good. | |||||||||||||||||||||||||||||||||||||||||
| ▲ | mo2art 33 minutes ago | parent | prev | next [-] | ||||||||||||||||||||||||||||||||||||||||
RuntimeError: audio limit is 30 s | |||||||||||||||||||||||||||||||||||||||||
| |||||||||||||||||||||||||||||||||||||||||
| ▲ | armcat an hour ago | parent | prev | next [-] | ||||||||||||||||||||||||||||||||||||||||
Those are insane benchmarks at this size. Well done! | |||||||||||||||||||||||||||||||||||||||||
| ▲ | tecleandor an hour ago | parent | prev | next [-] | ||||||||||||||||||||||||||||||||||||||||
Spanish is not good (seems to write non existing words and/or with terrible typos...) but English seem to work good even with my (Spanish) accent... | |||||||||||||||||||||||||||||||||||||||||
| |||||||||||||||||||||||||||||||||||||||||
| ▲ | mrkn1 an hour ago | parent | prev | next [-] | ||||||||||||||||||||||||||||||||||||||||
love seeing more sub-20MB, CPU-first models. if anyone wants a CLI built on the same ethos (no GPU, no cloud), been using yapsnap streaming Zipformer ASR, plus diarization and timestamps all on CPU! It supports 10 languages. Unlimited transcription for free. | |||||||||||||||||||||||||||||||||||||||||
| |||||||||||||||||||||||||||||||||||||||||
| ▲ | aidotguru 41 minutes ago | parent | prev | next [-] | ||||||||||||||||||||||||||||||||||||||||
eager to see if working in android phones | |||||||||||||||||||||||||||||||||||||||||
| |||||||||||||||||||||||||||||||||||||||||
| ▲ | saturn8601 an hour ago | parent | prev | next [-] | ||||||||||||||||||||||||||||||||||||||||
Initial tests make this feel just like iPhone's terrible text to speech. It is the one thing I utterly hate about iPhone. Ive tried apps that try to embed themselves into the iPhone keyboard and they always don't work out well. Hopefully this gets better and we can somehow get it into the iPhone more seamlessly. | |||||||||||||||||||||||||||||||||||||||||
| |||||||||||||||||||||||||||||||||||||||||
| ▲ | agilek 27 minutes ago | parent | prev [-] | ||||||||||||||||||||||||||||||||||||||||
Can we have more languages? | |||||||||||||||||||||||||||||||||||||||||
| |||||||||||||||||||||||||||||||||||||||||