Save 4 hours! We auto-setup your workflow! Free!

Drop your workflow.json — we handle every dependency, custom node, and model. Just open the link and run.

Auto-Setup Workflow Json (Free) Now!
ComfyUI > Nodes > comfyui-minimax-h3-audio-T8 > MiniMax H3 Speech Verify & Align / ASR校验裁切 (EXP/T8)

ComfyUI Node: MiniMax H3 Speech Verify & Align / ASR校验裁切 (EXP/T8)

Class Name

MiniMaxH3SpeechVerifyT8

Category
T8/MiniMax H3/Speech/Experimental
Author
T8mars (Account age: 1708days)
Extension
comfyui-minimax-h3-audio-T8
Latest Updated
2026-08-20
Github Stars
0.75K

How to Install comfyui-minimax-h3-audio-T8

Install this extension via the ComfyUI Manager by searching for comfyui-minimax-h3-audio-T8
  • 1. Click the Manager button in the main menu
  • 2. Select Custom Nodes Manager button
  • 3. Enter comfyui-minimax-h3-audio-T8 in the search bar
After installation, click the Restart button to restart ComfyUI. Then, manually refresh your browser to clear the cache and access the updated list of nodes.

Visit ComfyUI Online for ready-to-use ComfyUI environment

  • Free trial available
  • 16GB VRAM to 80GB VRAM GPU machines
  • 400+ preloaded models/nodes
  • Freedom to upload custom models/nodes
  • 200+ ready-to-run workflows
  • 100% private workspace with up to 200GB storage
  • Dedicated Support

Run ComfyUI Online

MiniMax H3 Speech Verify & Align / ASR校验裁切 (EXP/T8) Description

Speech audio verification and alignment node using ASR and speaker verification for accurate content and speaker identity validation.

MiniMax H3 Speech Verify & Align / ASR校验裁切 (EXP/T8):

The MiniMaxH3SpeechVerifyT8 node is designed to verify and align speech audio with expected text using advanced speech recognition and speaker verification techniques. This node is particularly useful for ensuring that the generated or recorded speech matches a given script or text, which is crucial in applications like automated dubbing, voice-over, and dialogue systems. By leveraging a combination of Automatic Speech Recognition (ASR) and speaker verification models, it provides a robust mechanism to check both the content and the speaker's identity. This ensures that the audio not only contains the correct words but is also spoken by the intended voice, enhancing the authenticity and accuracy of speech-based applications.

MiniMax H3 Speech Verify & Align / ASR校验裁切 (EXP/T8) Input Parameters:

audio

This parameter represents the audio input that needs to be verified. It is the primary data that the node processes to check against the expected text and speaker profile.

expected_text

The text that the audio is expected to contain. This parameter is crucial as it serves as the reference for the ASR system to verify the content of the audio.

verify_mode

Determines the mode of verification, which can affect how strictly the audio is checked against the expected text. Different modes may offer varying levels of tolerance for discrepancies.

asr_model_directory

Specifies the directory where the ASR model is located. This is essential for loading the correct model that will be used to transcribe and verify the audio content.

language

Indicates the language of the audio and expected text. This ensures that the ASR system uses the appropriate language model for transcription and verification.

min_similarity

Defines the minimum similarity threshold required for the audio to be considered a match with the expected text. A higher value means stricter verification.

beam_size

Controls the beam size used in the ASR decoding process, affecting the balance between speed and accuracy of the transcription.

cpu_threads

Specifies the number of CPU threads to be used for processing, which can impact the speed of verification, especially on multi-core systems.

unload_after_verify

A boolean parameter that determines whether the ASR model should be unloaded from memory after verification, which can help manage system resources.

strict

Indicates whether the verification should be strict, potentially affecting how minor discrepancies are handled.

pre_padding_seconds

The amount of silence to add before the audio during verification, which can help in aligning the audio with the expected text.

post_padding_seconds

The amount of silence to add after the audio during verification, aiding in proper alignment and verification.

voice_profile

An optional parameter that provides a reference audio for speaker verification, ensuring the audio is spoken by the correct voice.

speaker_check_mode

Determines the mode of speaker verification, which can range from off to various levels of strictness.

speaker_model_directory

Specifies the directory where the speaker verification model is located, necessary for loading the correct model for speaker checking.

min_speaker_similarity

Sets the minimum similarity threshold for speaker verification, ensuring the speaker's identity matches the reference profile.

unload_speaker_after_verify

A boolean parameter that decides whether the speaker model should be unloaded after verification, helping to manage memory usage.

peak_limit_dbfs

Defines the peak limit in decibels full scale for the audio, which can be used to normalize or limit the audio level during processing.

MiniMax H3 Speech Verify & Align / ASR校验裁切 (EXP/T8) Output Parameters:

verification_result

The result of the verification process, indicating whether the audio matches the expected text and speaker profile. This output is crucial for determining the success of the verification.

similarity_score

Provides a score representing the similarity between the audio and the expected text, offering insight into how closely the audio matches the script.

speaker_similarity_score

A score indicating the similarity between the speaker's voice in the audio and the reference voice profile, which is important for confirming speaker identity.

MiniMax H3 Speech Verify & Align / ASR校验裁切 (EXP/T8) Usage Tips:

  • Ensure that the expected_text closely matches the content of the audio to improve verification accuracy.
  • Use an appropriate language setting to match the audio content, as this significantly impacts ASR performance.
  • Adjust min_similarity and min_speaker_similarity thresholds based on the desired strictness of verification.
  • Consider the system's resource availability when setting cpu_threads and unload_after_verify to optimize performance.

MiniMax H3 Speech Verify & Align / ASR校验裁切 (EXP/T8) Common Errors and Solutions:

"ASR model not found in specified directory"

  • Explanation: The ASR model directory provided does not contain the necessary model files.
  • Solution: Verify the path in asr_model_directory and ensure it points to the correct location with the required model files.

"Speaker model not found in specified directory"

  • Explanation: The speaker model directory is incorrect or missing the necessary files for speaker verification.
  • Solution: Check the speaker_model_directory path and ensure it contains the appropriate speaker model files.

"Audio and expected text language mismatch"

  • Explanation: The language setting does not match the language of the audio or expected text.
  • Solution: Ensure the language parameter is set to the correct language of the audio content and expected text.

MiniMax H3 Speech Verify & Align / ASR校验裁切 (EXP/T8) Related Nodes

Go back to the extension to check out more related nodes.
comfyui-minimax-h3-audio-T8
RunComfy
Copyright 2025 RunComfy. All Rights Reserved.

RunComfy is the premier ComfyUI platform, offering ComfyUI online environment and services, along with ComfyUI workflows featuring stunning visuals. RunComfy also provides AI Models, enabling artists to harness the latest AI tools to create incredible art.

MiniMax H3 Speech Verify & Align / ASR校验裁切 (EXP/T8)