> Complete local text-to-waveform speech synthesis under 10M parameters.
In case, like me, you hoped "complete" voice might mean both stt and tts. Not to speak poorly of it, just clarifying.
> English only, with one fixed male voice. This is not zero-shot voice cloning.
(And then a bunch of statements on limitations that I read as 'quality can be spotty but if you play with it it should be fine') But like. In <10M params I'm not judging:)
I keep seeing tts stories here. Is it just an interesting subset of the llm world, or is there a huge use case I’m somehow missing?
This is amazing, the quality blow my mind for such small model! I just replaced my old onnx model with yours!
here my implementation with speech dispatcher and server: https://github.com/skorotkiewicz/inflect-speechd
thanks for shearing!
Couple highlights:
> Complete local text-to-waveform speech synthesis under 10M parameters.
In case, like me, you hoped "complete" voice might mean both stt and tts. Not to speak poorly of it, just clarifying.
> English only, with one fixed male voice. This is not zero-shot voice cloning.
(And then a bunch of statements on limitations that I read as 'quality can be spotty but if you play with it it should be fine') But like. In <10M params I'm not judging:)
The inflections are weird but this doesn't sound like a robot. Not bad!
I'd love to hear it but it seems your quota is exhausted.
This is impressive. I wish there were a voice clone option.
With so few parameters, I imagine a voice fine-tune might be readily tractable.
Amazing quality for small size, but definitely not that enjoyable to listen to.
IMHO, its at about the same quality level of historic TTS tools.
I'm not sure which historic tools you mean, but to me this sounds much better than anything older than ten years ago.
I compared the macos Samantha just now and I guess the inflect-micro is marginally better...
Ivona „Joey“, „Amy“
amazing quality for such small size!
Alternative title: Text to speech in 9.36M, English only.