Save 4 hours! We auto-setup your workflow! Free!

Drop your workflow.json — we handle every dependency, custom node, and model. Just open the link and run.

Auto-Setup Workflow Json (Free) Now!
ComfyUI > Nodes > comfyui-minimax-h3-audio-T8 > MiniMax H3 Speech Long Form Control / 长文本控制 (EXP/T8)

ComfyUI Node: MiniMax H3 Speech Long Form Control / 长文本控制 (EXP/T8)

Class Name

MiniMaxH3SpeechLongFormControlT8

Category
T8/MiniMax H3/Speech/Experimental
Author
T8mars (Account age: 1708days)
Extension
comfyui-minimax-h3-audio-T8
Latest Updated
2026-08-20
Github Stars
0.75K

How to Install comfyui-minimax-h3-audio-T8

Install this extension via the ComfyUI Manager by searching for comfyui-minimax-h3-audio-T8
  • 1. Click the Manager button in the main menu
  • 2. Select Custom Nodes Manager button
  • 3. Enter comfyui-minimax-h3-audio-T8 in the search bar
After installation, click the Restart button to restart ComfyUI. Then, manually refresh your browser to clear the cache and access the updated list of nodes.

Visit ComfyUI Online for ready-to-use ComfyUI environment

  • Free trial available
  • 16GB VRAM to 80GB VRAM GPU machines
  • 400+ preloaded models/nodes
  • Freedom to upload custom models/nodes
  • 200+ ready-to-run workflows
  • 100% private workspace with up to 200GB storage
  • Dedicated Support

Run ComfyUI Online

MiniMax H3 Speech Long Form Control / 长文本控制 (EXP/T8) Description

Manage long-form speech segments in speech processing workflow, evaluating and accepting based on criteria for seamless integration.

MiniMax H3 Speech Long Form Control / 长文本控制 (EXP/T8):

The MiniMaxH3SpeechLongFormControlT8 node is designed to manage and control the acceptance of long-form speech segments within a speech processing workflow. This node plays a crucial role in ensuring that the audio output aligns with the intended speech plan by evaluating and accepting segments based on various criteria such as text and speaker similarity. It facilitates the seamless integration of audio segments into the final output, providing a structured approach to handling complex speech tasks. By leveraging this node, you can efficiently manage long-form speech projects, ensuring that each segment meets the desired quality and consistency standards.

MiniMax H3 Speech Long Form Control / 长文本控制 (EXP/T8) Input Parameters:

session

The session parameter represents the current session context in which the speech processing is taking place. It is essential for maintaining state and ensuring that all operations are executed within the correct session environment.

speech_plan

The speech_plan parameter outlines the intended structure and content of the speech. It serves as a blueprint for the speech processing workflow, guiding the acceptance and integration of audio segments.

segment_index

The segment_index parameter specifies the index of the current segment being processed. It is crucial for identifying and managing individual segments within the larger speech plan.

audio

The audio parameter contains the audio data for the segment being evaluated. This data is analyzed to determine its suitability for inclusion in the final output.

transcript

The transcript parameter provides the textual representation of the audio segment. It is used to assess text similarity and ensure that the audio aligns with the intended speech content.

text_similarity

The text_similarity parameter measures the similarity between the transcript and the intended speech content. It helps in evaluating whether the audio segment accurately represents the desired text.

speaker_similarity

The speaker_similarity parameter assesses the similarity between the speaker's voice in the audio segment and the reference voice. It ensures consistency in speaker characteristics across segments.

accepted

The accepted parameter is a boolean flag indicating whether the segment has been accepted for inclusion in the final output. It is used to track the acceptance status of each segment.

replace_existing

The replace_existing parameter is a boolean flag that determines whether an existing segment should be replaced with the current one if accepted. It allows for dynamic updates to the speech plan.

MiniMax H3 Speech Long Form Control / 长文本控制 (EXP/T8) Output Parameters:

output_audio

The output_audio parameter contains the processed audio data for the accepted segment. This output is ready for integration into the final speech output, ensuring that it meets the desired quality and consistency standards.

report_json

The report_json parameter provides a detailed report of the segment evaluation process in JSON format. It includes information about the acceptance criteria, similarity scores, and any actions taken during the evaluation.

ui

The ui parameter offers a user interface representation of the output, including details such as the filename and subfolder of the audio output. This makes it easier to manage and access the processed audio files.

MiniMax H3 Speech Long Form Control / 长文本控制 (EXP/T8) Usage Tips:

  • Ensure that the speech_plan is well-defined and accurately represents the intended speech content to facilitate effective segment evaluation.
  • Regularly monitor the text_similarity and speaker_similarity scores to maintain high-quality and consistent audio output across segments.
  • Utilize the replace_existing parameter judiciously to update segments only when necessary, preserving the integrity of the original speech plan.

MiniMax H3 Speech Long Form Control / 长文本控制 (EXP/T8) Common Errors and Solutions:

"Invalid session context"

  • Explanation: This error occurs when the session parameter is not correctly initialized or is missing.
  • Solution: Ensure that the session parameter is properly set up and passed to the node before execution.

"Segment index out of range"

  • Explanation: This error indicates that the specified segment index does not exist within the speech plan.
  • Solution: Verify that the segment_index parameter is within the valid range of segments defined in the speech plan.

"Audio data missing"

  • Explanation: This error arises when the audio parameter is not provided or is empty.
  • Solution: Check that the audio data is correctly loaded and passed to the node for processing.

"Text similarity threshold not met"

  • Explanation: This error occurs when the text similarity score is below the acceptable threshold for segment acceptance.
  • Solution: Review the transcript and speech plan to ensure they align closely, and adjust the text similarity threshold if necessary.

MiniMax H3 Speech Long Form Control / 长文本控制 (EXP/T8) Related Nodes

Go back to the extension to check out more related nodes.
comfyui-minimax-h3-audio-T8
RunComfy
Copyright 2025 RunComfy. All Rights Reserved.

RunComfy is the premier ComfyUI platform, offering ComfyUI online environment and services, along with ComfyUI workflows featuring stunning visuals. RunComfy also provides AI Models, enabling artists to harness the latest AI tools to create incredible art.

MiniMax H3 Speech Long Form Control / 长文本控制 (EXP/T8)