Thank you for this amazing repo and for sharing your pretrained models!
I have a few questions:
-
In your experience, how large should the dataset be to achieve reasonable quality?
-
Can the pretrained models be fine-tuned instead of training from scratch?
-
Would training a model solely on the voice audio of one specific speaker improve inference results for that speaker?
Thank you for this amazing repo and for sharing your pretrained models!
I have a few questions:
In your experience, how large should the dataset be to achieve reasonable quality?
Can the pretrained models be fine-tuned instead of training from scratch?
Would training a model solely on the voice audio of one specific speaker improve inference results for that speaker?