Save 4 hours! We auto-setup your workflow! Free!

Drop your workflow.json — we handle every dependency, custom node, and model. Just open the link and run.

Auto-Setup Workflow Json (Free) Now!
ComfyUI > Nodes > ComfyUI-MAINodes > H3 Audio Recover (hold-map atempo, pitch kept)

ComfyUI Node: H3 Audio Recover (hold-map atempo, pitch kept)

Class Name

H3AudioRecover

Category
audio/minimax/motion
Author
matlowai (Account age: 1004days)
Extension
ComfyUI-MAINodes
Latest Updated
2026-08-26
Github Stars
0.11K

How to Install ComfyUI-MAINodes

Install this extension via the ComfyUI Manager by searching for ComfyUI-MAINodes
  • 1. Click the Manager button in the main menu
  • 2. Select Custom Nodes Manager button
  • 3. Enter ComfyUI-MAINodes in the search bar
After installation, click the Restart button to restart ComfyUI. Then, manually refresh your browser to clear the cache and access the updated list of nodes.

Visit ComfyUI Online for ready-to-use ComfyUI environment

  • Free trial available
  • 16GB VRAM to 80GB VRAM GPU machines
  • 400+ preloaded models/nodes
  • Freedom to upload custom models/nodes
  • 200+ ready-to-run workflows
  • 100% private workspace with up to 200GB storage
  • Dedicated Support

Run ComfyUI Online

H3 Audio Recover (hold-map atempo, pitch kept) Description

Specialized node for retiming audio clips to world clock, maintaining synchronization with visual elements and incorporating reference tracks for audio fidelity.

H3 Audio Recover (hold-map atempo, pitch kept):

H3AudioRecover is a specialized node designed to retime audio clips in synchronization with a specified world clock, ensuring that the audio aligns perfectly with the intended timing of visual or other media elements. This node is particularly beneficial for applications where precise audio timing is crucial, such as in film, animation, or interactive media projects. By adjusting the audio to match a given frame-per-second (FPS) rate, H3AudioRecover ensures that the audio does not drift over time, maintaining synchronization with visual elements. The node can handle complex timing maps, known as "holds," which dictate how different segments of the audio should be stretched or compressed to fit the desired timing. Additionally, H3AudioRecover can incorporate reference audio tracks to ensure that the retimed audio maintains the desired characteristics, such as tone and quality, by blending the retimed audio with the reference track. This capability is particularly useful for maintaining audio fidelity and ensuring that the final output meets the creative and technical requirements of the project.

H3 Audio Recover (hold-map atempo, pitch kept) Input Parameters:

waveform

The waveform parameter is the audio data that you want to retime. It is typically a multi-dimensional array representing the audio signal, with dimensions corresponding to channels and samples. This parameter is crucial as it provides the raw audio that will be processed and adjusted to match the specified timing. The waveform should be provided in a format compatible with the node's processing capabilities, such as a PyTorch tensor.

sample_rate

The sample_rate parameter specifies the number of samples per second in the audio waveform. It is a critical parameter because it defines the temporal resolution of the audio data. A higher sample rate means more samples per second, resulting in higher audio quality and more precise timing adjustments. The sample rate should match the rate at which the audio was originally recorded or intended to be played back.

holds

The holds parameter is a JSON-encoded string that defines the timing map for the audio retiming process. It consists of a sequence of hold values that specify how long each segment of the audio should be held or stretched. This parameter allows for complex timing adjustments, enabling precise control over how the audio aligns with the world clock. The holds map is essential for ensuring that the audio timing matches the desired visual or media timing.

fps

The fps parameter stands for frames per second and defines the frame rate of the visual or media content with which the audio should be synchronized. This parameter is crucial for ensuring that the audio timing aligns perfectly with the visual elements, preventing drift or misalignment over time. The fps value should match the frame rate of the media content to achieve accurate synchronization.

reference

The reference parameter is an optional input that provides a reference audio track for the retiming process. It is used to ensure that the retimed audio maintains the desired characteristics, such as tone and quality, by blending the retimed audio with the reference track. This parameter is particularly useful when the retimed audio needs to closely match a specific reference in terms of sound quality or other attributes.

reference_mix

The reference_mix parameter is a float value that determines the blending ratio between the retimed audio and the reference audio track. A value of 0.0 means that the retimed audio is used exclusively, while a value of 1.0 means that the reference audio is used exclusively. Intermediate values blend the two audio tracks, allowing for a mix that retains the timing of the retimed audio while incorporating the qualities of the reference track.

H3 Audio Recover (hold-map atempo, pitch kept) Output Parameters:

waveform

The waveform output parameter is the retimed audio data that has been processed to align with the specified world clock and timing map. This output is a multi-dimensional array representing the adjusted audio signal, with dimensions corresponding to channels and samples. The retimed waveform ensures that the audio is perfectly synchronized with the visual or media content, maintaining the desired timing and quality.

H3 Audio Recover (hold-map atempo, pitch kept) Usage Tips:

  • Ensure that the sample_rate and fps parameters are correctly set to match the original recording and media content, respectively, to achieve accurate synchronization.
  • Use the reference and reference_mix parameters to maintain audio quality by blending the retimed audio with a high-quality reference track, especially when precise audio characteristics are required.
  • Carefully construct the holds map to define the desired timing adjustments for different segments of the audio, allowing for precise control over the retiming process.

H3 Audio Recover (hold-map atempo, pitch kept) Common Errors and Solutions:

retimed length <value> != world target <value>

  • Explanation: This error occurs when the length of the retimed audio does not match the expected length based on the world clock and timing map.
  • Solution: Verify that the holds map and fps parameter are correctly set to ensure that the audio is retimed accurately. Adjust the parameters as needed to achieve the desired synchronization.

mix=1.0 output is not the reference bit-for-bit

  • Explanation: This error indicates that the output audio does not match the reference audio exactly when the reference_mix is set to 1.0.
  • Solution: Ensure that the reference audio track is correctly specified and that the reference_mix parameter is set to 1.0. Check for any discrepancies in the audio data or processing that may affect the output.

non-finite samples in the retimed track

  • Explanation: This error suggests that there are invalid or non-finite values in the retimed audio waveform.
  • Solution: Check the input waveform for any anomalies or issues that may cause non-finite values. Ensure that the audio data is properly formatted and free of errors before processing.

H3 Audio Recover (hold-map atempo, pitch kept) Related Nodes

Go back to the extension to check out more related nodes.
ComfyUI-MAINodes
RunComfy
Copyright 2025 RunComfy. All Rights Reserved.

RunComfy is the premier ComfyUI platform, offering ComfyUI online environment and services, along with ComfyUI workflows featuring stunning visuals. RunComfy also provides AI Models, enabling artists to harness the latest AI tools to create incredible art.

H3 Audio Recover (hold-map atempo, pitch kept)