MiniMax H3 Speech Plan / 语音规划 (EXP/T8):
The MiniMaxH3SpeechPlanT8 node is designed to facilitate the separation of spoken text from acting directions and perform language-aware chunking in audio processing. This node is particularly beneficial for projects that require precise audio planning and execution, as it ensures that the render duration is never guessed, providing a reliable and consistent output. By focusing on the separation and chunking of audio content, it allows for more accurate and efficient audio processing, making it an essential tool for AI artists working with complex audio projects. The node's ability to handle language-specific nuances in audio content further enhances its utility, ensuring that the final output is both accurate and contextually appropriate.
MiniMax H3 Speech Plan / 语音规划 (EXP/T8) Input Parameters:
speech_plan
The speech_plan parameter is a crucial input that defines the structure and content of the audio to be processed. It serves as the blueprint for separating spoken text from acting directions and guides the chunking process. This parameter directly impacts how the audio is organized and processed, ensuring that the final output aligns with the intended plan.
output_sample_rate
The output_sample_rate parameter determines the sample rate of the output audio. It offers options of 32000, 44100, and 48000 Hz, with a default value of 32000 Hz. This parameter affects the quality and size of the audio output, with higher sample rates providing better audio quality at the cost of larger file sizes.
crossfade_seconds
The crossfade_seconds parameter specifies the duration of the crossfade between audio segments. It has a default value of 0.06 seconds, with a range from 0.0 to 0.5 seconds. This parameter is essential for ensuring smooth transitions between audio segments, preventing abrupt changes that could disrupt the listening experience.
peak_limit_dbfs
The peak_limit_dbfs parameter sets the maximum allowable peak level for the audio output, with a default value of -1.0 dBFS. It ranges from -12.0 to 0.0 dBFS. This parameter is important for maintaining audio quality by preventing clipping and distortion in the final output.
audio_segments
The audio_segments parameter allows for the input of multiple audio segments, which are processed according to the speech plan. It supports a minimum of 1 and a maximum of 100 segments. This parameter is vital for projects that involve multiple audio clips, enabling the node to handle complex audio compositions efficiently.
MiniMax H3 Speech Plan / 语音规划 (EXP/T8) Output Parameters:
audio_output
The audio_output parameter provides the processed audio as per the defined speech plan. It represents the final audio product, which has been separated, chunked, and crossfaded according to the input parameters. This output is crucial for ensuring that the audio aligns with the intended plan and meets the desired quality standards.
timeline_output
The timeline_output parameter offers a timeline of the audio boundaries along with the planned-text in SRT/VTT format. This output is essential for synchronizing audio with text, making it particularly useful for projects that require precise timing and alignment between audio and visual elements.
MiniMax H3 Speech Plan / 语音规划 (EXP/T8) Usage Tips:
- Ensure that your
speech_planis well-defined and accurately reflects the intended structure of your audio project to achieve the best results. - Choose an
output_sample_ratethat balances audio quality and file size based on your project's requirements. - Adjust the
crossfade_secondsto ensure smooth transitions between audio segments, especially in projects with frequent segment changes.
MiniMax H3 Speech Plan / 语音规划 (EXP/T8) Common Errors and Solutions:
"Invalid sample rate selected"
- Explanation: This error occurs when an unsupported sample rate is chosen.
- Solution: Ensure that the
output_sample_rateis set to one of the supported values: 32000, 44100, or 48000 Hz.
"Audio segment limit exceeded"
- Explanation: This error indicates that the number of audio segments exceeds the maximum allowed.
- Solution: Reduce the number of
audio_segmentsto 100 or fewer to comply with the node's limitations.
"Peak limit out of range"
- Explanation: This error arises when the
peak_limit_dbfsis set outside the permissible range. - Solution: Adjust the
peak_limit_dbfsto be within the range of -12.0 to 0.0 dBFS.
