Skip to content

Need An Easy Inference Sample #11

Description

@fish0510

Hello,

I would like to run inference using your provided pretrained model without additional training or data binarization. However, when I tried to run it, I encountered many required files in the data/ directory. On the other hand, there are no clear examples explaining input information format. For instance, in style_transfer.py, in the example_run function, it does not specify the expected format of information in inp.

Could someone share your experience or provide guidance on this?

Additionally, if I want to use my own data (e.g., annotated data from OpenCpop) for inference with the pretrained model, what steps should I follow?

Thank you!

Activity

  1. AaronZ345 commented on Apr 2, 2025

    @AaronZ345
    Owner

    If you want to perform customizable inference, such as in style transfer, you have two options:

    1. You can replace the item_name used in inp with the information commented out in the file. For example in style_transfer.py:
    # 'text_gen': ,
    # 'note_gen': ,
    # 'note_dur_gen': ,
    # 'note_type_gen':,
    # 'text_in': ,
    # 'note_in': ,
    # 'note_dur_in': ,
    # 'note_type_in':,
    # 'ref_audio': ,
    # 'ph_durs':
    1. Alternatively, you can refer to the GTSinger metadata format to create metadata for your data. This ensures that each WAV file possesses the required attributes. Then, use item_name to facilitate the inference process.

    If you want to perform large-scale inference using OpenCPOP data, please follow GTSinger's guidelines to binarize the data. After that, run the following command:

    CUDA_VISIBLE_DEVICES=$GPU python tasks/run.py --config egs/sdlm.yaml --exp_name SDLM --infer
  2. jackien1 commented on Apr 3, 2025

    @jackien1

    I second this as well. If TCSinger could make a Google Colab tutorial or simple inference sample that would be greatly appreciated!

  3. DongJiashu commented on Apr 18, 2025

    @DongJiashu

    If you want to perform customizable inference, such as in style transfer, you have two options:

    1. You can replace the item_name used in inp with the information commented out in the file. For example in style_transfer.py:

    'text_gen': ,

    'note_gen': ,

    'note_dur_gen': ,

    'note_type_gen':,

    'text_in': ,

    'note_in': ,

    'note_dur_in': ,

    'note_type_in':,

    'ref_audio': ,

    'ph_durs':

    1. Alternatively, you can refer to the GTSinger metadata format to create metadata for your data. This ensures that each WAV file possesses the required attributes. Then, use item_name to facilitate the inference process.

    If you want to perform large-scale inference using OpenCPOP data, please follow GTSinger's guidelines to binarize the data. After that, run the following command:

    CUDA_VISIBLE_DEVICES=$GPU python tasks/run.py --config egs/sdlm.yaml --exp_name SDLM --infer

    How the 2nd way knows what to infer? like option 1 you give the audio / datas so it gets input and know what to output, but im confused with the 2nd way

  4. AaronZ345 commented on May 5, 2025

    @AaronZ345
    Owner

    If you want to perform customizable inference, such as in style transfer, you have two options:

    1. You can replace the item_name used in inp with the information commented out in the file. For example in style_transfer.py:

    'text_gen': ,

    'note_gen': ,

    'note_dur_gen': ,

    'note_type_gen':,

    'text_in': ,

    'note_in': ,

    'note_dur_in': ,

    'note_type_in':,

    'ref_audio': ,

    'ph_durs':

    1. Alternatively, you can refer to the GTSinger metadata format to create metadata for your data. This ensures that each WAV file possesses the required attributes. Then, use item_name to facilitate the inference process.

    If you want to perform large-scale inference using OpenCPOP data, please follow GTSinger's guidelines to binarize the data. After that, run the following command:

    CUDA_VISIBLE_DEVICES=$GPU python tasks/run.py --config egs/sdlm.yaml --exp_name SDLM --infer

    How the 2nd way knows what to infer? like option 1 you give the audio / datas so it gets input and know what to output, but im confused with the 2nd way

    For this way, model will find the item_name in the metadata and get all information needed.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions