MiniMax H3 Speech Verify & Align / ASR校验裁切 (EXP/T8):
The MiniMaxH3SpeechVerifyT8 node is designed to verify and align speech audio with expected text using advanced speech recognition and speaker verification techniques. This node is particularly useful for ensuring that the generated or recorded speech matches a given script or text, which is crucial in applications like automated dubbing, voice-over, and dialogue systems. By leveraging a combination of Automatic Speech Recognition (ASR) and speaker verification models, it provides a robust mechanism to check both the content and the speaker's identity. This ensures that the audio not only contains the correct words but is also spoken by the intended voice, enhancing the authenticity and accuracy of speech-based applications.
MiniMax H3 Speech Verify & Align / ASR校验裁切 (EXP/T8) Input Parameters:
audio
This parameter represents the audio input that needs to be verified. It is the primary data that the node processes to check against the expected text and speaker profile.
expected_text
The text that the audio is expected to contain. This parameter is crucial as it serves as the reference for the ASR system to verify the content of the audio.
verify_mode
Determines the mode of verification, which can affect how strictly the audio is checked against the expected text. Different modes may offer varying levels of tolerance for discrepancies.
asr_model_directory
Specifies the directory where the ASR model is located. This is essential for loading the correct model that will be used to transcribe and verify the audio content.
language
Indicates the language of the audio and expected text. This ensures that the ASR system uses the appropriate language model for transcription and verification.
min_similarity
Defines the minimum similarity threshold required for the audio to be considered a match with the expected text. A higher value means stricter verification.
beam_size
Controls the beam size used in the ASR decoding process, affecting the balance between speed and accuracy of the transcription.
cpu_threads
Specifies the number of CPU threads to be used for processing, which can impact the speed of verification, especially on multi-core systems.
unload_after_verify
A boolean parameter that determines whether the ASR model should be unloaded from memory after verification, which can help manage system resources.
strict
Indicates whether the verification should be strict, potentially affecting how minor discrepancies are handled.
pre_padding_seconds
The amount of silence to add before the audio during verification, which can help in aligning the audio with the expected text.
post_padding_seconds
The amount of silence to add after the audio during verification, aiding in proper alignment and verification.
voice_profile
An optional parameter that provides a reference audio for speaker verification, ensuring the audio is spoken by the correct voice.
speaker_check_mode
Determines the mode of speaker verification, which can range from off to various levels of strictness.
speaker_model_directory
Specifies the directory where the speaker verification model is located, necessary for loading the correct model for speaker checking.
min_speaker_similarity
Sets the minimum similarity threshold for speaker verification, ensuring the speaker's identity matches the reference profile.
unload_speaker_after_verify
A boolean parameter that decides whether the speaker model should be unloaded after verification, helping to manage memory usage.
peak_limit_dbfs
Defines the peak limit in decibels full scale for the audio, which can be used to normalize or limit the audio level during processing.
MiniMax H3 Speech Verify & Align / ASR校验裁切 (EXP/T8) Output Parameters:
verification_result
The result of the verification process, indicating whether the audio matches the expected text and speaker profile. This output is crucial for determining the success of the verification.
similarity_score
Provides a score representing the similarity between the audio and the expected text, offering insight into how closely the audio matches the script.
speaker_similarity_score
A score indicating the similarity between the speaker's voice in the audio and the reference voice profile, which is important for confirming speaker identity.
MiniMax H3 Speech Verify & Align / ASR校验裁切 (EXP/T8) Usage Tips:
- Ensure that the
expected_textclosely matches the content of the audio to improve verification accuracy. - Use an appropriate
languagesetting to match the audio content, as this significantly impacts ASR performance. - Adjust
min_similarityandmin_speaker_similaritythresholds based on the desired strictness of verification. - Consider the system's resource availability when setting
cpu_threadsandunload_after_verifyto optimize performance.
MiniMax H3 Speech Verify & Align / ASR校验裁切 (EXP/T8) Common Errors and Solutions:
"ASR model not found in specified directory"
- Explanation: The ASR model directory provided does not contain the necessary model files.
- Solution: Verify the path in
asr_model_directoryand ensure it points to the correct location with the required model files.
"Speaker model not found in specified directory"
- Explanation: The speaker model directory is incorrect or missing the necessary files for speaker verification.
- Solution: Check the
speaker_model_directorypath and ensure it contains the appropriate speaker model files.
"Audio and expected text language mismatch"
- Explanation: The language setting does not match the language of the audio or expected text.
- Solution: Ensure the
languageparameter is set to the correct language of the audio content and expected text.
