Save 4 hours! We auto-setup your workflow! Free!

Drop your workflow.json — we handle every dependency, custom node, and model. Just open the link and run.

Auto-Setup Workflow Json (Free) Now!
ComfyUI > Nodes > Comfyui_SynVow_Qwen3ASR > Qwen3 Forced Align

ComfyUI Node: Qwen3 Forced Align

Class Name

Qwen3ForcedAlign

Category
Qwen3-ASR
Author
shumoLR (Account age: 803days)
Extension
Comfyui_SynVow_Qwen3ASR
Latest Updated
2026-02-06
Github Stars
0.04K

How to Install Comfyui_SynVow_Qwen3ASR

Install this extension via the ComfyUI Manager by searching for Comfyui_SynVow_Qwen3ASR
  • 1. Click the Manager button in the main menu
  • 2. Select Custom Nodes Manager button
  • 3. Enter Comfyui_SynVow_Qwen3ASR in the search bar
After installation, click the Restart button to restart ComfyUI. Then, manually refresh your browser to clear the cache and access the updated list of nodes.

Visit ComfyUI Online for ready-to-use ComfyUI environment

  • Free trial available
  • 16GB VRAM to 80GB VRAM GPU machines
  • 400+ preloaded models/nodes
  • Freedom to upload custom models/nodes
  • 200+ ready-to-run workflows
  • 100% private workspace with up to 200GB storage
  • Dedicated Support

Run ComfyUI Online

Qwen3 Forced Align Description

Sophisticated node for text-to-speech alignment with precise timestamps, ideal for subtitling and audio transcription.

Qwen3 Forced Align:

Qwen3ForcedAlign is a sophisticated node designed to perform text-to-speech alignment, providing precise timestamps for each segment of text in relation to the corresponding audio. This node is particularly beneficial for applications requiring accurate synchronization between spoken words and their textual representation, such as in subtitling, language learning tools, or audio transcription services. By leveraging advanced forced alignment techniques, Qwen3ForcedAlign ensures that each word or sentence in the text is matched with its exact timing in the audio, enhancing the clarity and usability of audio-visual content. The node is capable of handling multiple languages, with a default focus on Chinese, and can segment text by sentences to improve alignment accuracy. Its integration into the Qwen3-ASR framework allows for seamless operation on CUDA-enabled devices, ensuring efficient processing even for large audio files.

Qwen3 Forced Align Input Parameters:

aligner

The aligner parameter specifies the Qwen3 ForcedAligner model to be used for the alignment process. This model is responsible for analyzing the audio and text inputs to determine the precise timing of each segment. The aligner must be pre-loaded and compatible with the Qwen3-ASR framework to ensure accurate results.

audio

The audio parameter represents the audio input that will be aligned with the provided text. It should be in a format that includes both the waveform and the sample rate. The waveform is typically a numerical representation of the audio signal, and the sample rate indicates how many samples per second are used to represent the audio. This parameter is crucial as it directly influences the alignment accuracy.

text

The text parameter is the string input that contains the textual content to be aligned with the audio. It supports multiline text, allowing for comprehensive alignment of longer passages. The text should be clear and well-structured to facilitate accurate segmentation and alignment.

language

The language parameter specifies the language of the text and audio inputs. It supports multiple languages, with a default setting of Chinese. This parameter ensures that the alignment process takes into account language-specific characteristics, which can significantly impact the accuracy of the alignment.

segment_by_sentence

The segment_by_sentence parameter is a boolean option that determines whether the text should be segmented by sentences during the alignment process. When set to true, the node will attempt to align each sentence individually, which can improve the precision of the alignment, especially in complex or lengthy texts. The default value is true.

Qwen3 Forced Align Output Parameters:

timestamps

The timestamps output provides a string containing the start and end times for each segment of text in relation to the audio. This output is essential for applications that require precise timing information, such as subtitle generation or detailed transcription.

text_list

The text_list output is a string that lists all the text segments that have been aligned. This output allows users to verify which parts of the text were successfully aligned and can be used for further processing or analysis.

start_times

The start_times output provides a string of the start times for each aligned text segment. This information is crucial for understanding when each segment begins in the audio, enabling precise synchronization.

end_times

The end_times output provides a string of the end times for each aligned text segment. This output complements the start times, offering a complete picture of the duration and timing of each text segment within the audio.

Qwen3 Forced Align Usage Tips:

  • Ensure that the audio input is clear and of high quality to improve alignment accuracy.
  • Use the segment_by_sentence option for longer texts to enhance precision by aligning each sentence individually.
  • Verify that the language parameter is set correctly to match the language of the text and audio inputs for optimal results.

Qwen3 Forced Align Common Errors and Solutions:

Model not found, downloading: <model_name>

  • Explanation: This error occurs when the specified model is not found locally and needs to be downloaded.
  • Solution: Ensure that you have a stable internet connection to allow the model to be downloaded successfully. Check the model name for any typos.

Using local model: <local_model_path>

  • Explanation: This message indicates that the node is using a locally stored model for alignment.
  • Solution: No action is needed if the local model is correct. If issues arise, verify the model's integrity and compatibility.

Alignment completed, <number> segments

  • Explanation: This message confirms that the alignment process has finished and indicates the number of segments aligned.
  • Solution: Review the output to ensure that all expected segments are included. If discrepancies are found, check the input text and audio for errors.

Qwen3 Forced Align Related Nodes

Go back to the extension to check out more related nodes.
Comfyui_SynVow_Qwen3ASR
RunComfy
Copyright 2025 RunComfy. All Rights Reserved.

RunComfy is the premier ComfyUI platform, offering ComfyUI online environment and services, along with ComfyUI workflows featuring stunning visuals. RunComfy also provides AI Models, enabling artists to harness the latest AI tools to create incredible art.

Qwen3 Forced Align