Voice Generator Software

I’m curious about some sort of text to speech software or something similar to make samples with.
Unless these are samples sourced from somewhere, I assume that’s what Tzusing is using here and on some other tracks from his latest record

and here?

I’ve tried out a couple free text to speech sites I found through google, though the effect was…not cool. Any ideas?

The Dopplereffekt one sounds like a real voice run through a ring modulator, the other one might also be a real voice with some delay but the samples are cut off at the end which gives it a slightly unnatural feeling.

1 Like

These do not have a strong sonic text to speech quality, which makes me wonder if they are just actual speech processed in some way. The first example sounds the least like TOS. Dopplereffekt uses a ring modulator.

So if you want that sound, record yourself then put a ring on it :upside_down_face:

1 Like

try the ones on google that say AI text to speech, they usually have premium voices you can try free for a limited amount of text. It’s enough to make samples.

You can usually pick an accent / regional voicing and you just have to go through a bunch of them to find a useable one.

Also, you have to try abstract phonetic spellings to get the pronunciations you want, just typing the words as spelled in the dictionary won’t always yield the desired result.

Interesting advice re: abstract spelling. Gonna try that, thanks!

I still quite like Plogue’s Alter/Ego - also a bonus that its free, with three free sound-banks available.

3 Likes

sick, installing this now :slight_smile:

Spelling is part of it, but sometimes you have to break words apart if the virtual/software reader isn’t spacing the delivery of syllables the way you want, or if they aren’t enunciating the way you want them to. There’s no way to coach software on which syllable you want it to stress like there would be if you were working with a voice actor, so the only thing you can do is try various input methods and sometimes having a 3 syllable word broken down into 3 phonetic word-like sounds with spaces in-between is the only way to get the result that you want, even then sometimes it’s still impossible.

Another example is that I’m not aware of any reader software or program where you can tell the AI what mood or tone you want. So if the synthesized voice is dry and unemotive, you can’t tell it to smile, or if it’s always in a good mood you can’t tell it to sound serious. That’s the main reason you have to audition as many of the available voices as possible. Just my experience though, your mileage may vary, as they say.

{S doorslam