Calling VO-6 Masters

I’m developing a Text to speech (song?!) process. But the system needs a few examples to help guide accuracy and options.

If anyone is willing to share some pattern/kit sysex (that have quality diphthong and phoneme use) here it will expedite the process.
even just 16 steps with one or two clear words will still be really helpful.

i’ve given it a go but VO-6 is quite subjective in output, so the more approaches/variations we can incorporate the more wide ranging the possibilities.

i want to incorporate the filter for mouth shape effectively alongside the core synthesis. (then LFO for vibrato/tremolo etc)

end game is: import a MIDI file and enter lyrics to make the mono sing. (potential for 1/2/3/4/5/6 voice arrangements (plock permitting)!
(actually in poly mode 6 voice unison arrangements should be far more effective as all voices will share same trigs.

5 Likes

completely related

3 Likes

Are you going to interface with a TTS app?
I feel like what is needed is a table
Phonème => VO-6 values
No clue how phonème are encoded though, never looked into this.

everything’s in place, it’s its own system. just need a wide range of words/diphthongs/phonemes/morphemes in pattern (.syx) form to help steer.

2 Likes

example word vowel VOC1 setting VOC2 setting

lake A 93 118
leak E 40 127
like I - - not a
discreet sound (aah-y-uh)
oh O 99 48
you U - - not a
discreet sound (y-oo)
lack a 127 93
let e 124 98
lick i 91 109
lock (ah) o 127 60
luck (uh) u 112 53
luke u 71 51
look oo 93 61

2 Likes

but how these are timed and shaped are user taste.

have spent time on this but not main focus currently.

GUIDE_vo6_component_bank.txt (35.8 KB)
vo6_component_bank_bundle_all_messages.syx (6.8 KB)

if anyone fancies the challenge (and has the time) i’ve got a working system where feeding back assumptions on word components with human correction builds a working dataset.
updating the patterns take time, i find moving the pattern back or forth so your focus is on steps 1 and 2. then loop this with a pattern length of 2. tweak locks until it sounds right. move onto next 2 step segment. it’s a laborious task! i’m trying but more heads with their own subjective takes will make it better.

word components targeted
  • Monophthong or vowel hold: one steady vowel sound, like A, EH, EE, OH, UH, ER.
  • Diphthong: a vowel glide from one vowel to another, like AI, OI, OW, UE.
  • Off-glide or rhotic glide: a vowel moving into another color, often R-colored, like AIR, EAR, IR, UR.
  • Onset: the starting consonant of a syllable, like H, T, K, SJ, J.
  • Coda: the ending consonant of a syllable, like -N, -D, -K, -S.
  • Cluster: two-consonant shapes at the start or end, like ST, TR, PL, NK, RD.
  • Rime: a reusable vowel-plus-ending chunk, like -ain, -ake, -ore, -ook.
  • Morpheme: a chunk that carries meaning, usually things like prefixes or suffixes, such as re-, pre-, un-, -tion, -ness, -ly, -ive.
  • Reduced join or weak syllable: an unstressed linking chunk, often based on schwa/UH.
  • Prosody or boundary: speech contour shapes like question rise, comma hold, period close.
  • Approximation phone: a best-fit VO6 stand-in for sounds the machine does not really have directly, like W, Y, DH, ZH.