Voice Cloning
To create a high-quality custom voice, the input audio (training data) must be extremely clean.
- Recording Location:
- Reduce room echo or reverberation. Soundproof the space using a vocal booth, a closet, or by building a blanket fort.
- Equipment:
- Microphone: A professional XLR microphone is recommended (in the $150–$300 range, such as the Audio-Technica AT2020 or Rode NT1).
- Audio Interface: Use a USB interface (such as Focusrite).
- Pop filter: Essential for avoiding plosive sounds (air striking the microphone).
- Software (Software/DAW):
- Use programs like REAPER or Audacity.
- Settings: Record WAV files at a sample rate of 44.1 kHz or 48 kHz and a bit depth of 24-bit.
- Levels: Aim for audio peaks between -6 dB and -3 dB, and an average loudness of -18 dB.
- Positioning:
- Maintain a distance of approximately 20 cm (equivalent to two fists) from the microphone.
- Speak at an angle to avoid breathing directly into the microphone.
- Performance:
- Consistency: Maintain the same energy/tone throughout. Do not mix whispering and shouting in the training file.
AI replicates everything, including breaths, stutters, and hesitation sounds (like "ummm"). Make sure the performance is exactly what you want to replicate.
Go to the next article in the Makina guide.
خ