Save 4 hours! We auto-setup your workflow! Free!

Drop your workflow.json — we handle every dependency, custom node, and model. Just open the link and run.

Auto-Setup Workflow Json (Free) Now!
ComfyUI > Nodes > ComfyUI-MAINodes > H3 Streamed Blocks (exact low-VRAM, alpha)

ComfyUI Node: H3 Streamed Blocks (exact low-VRAM, alpha)

Class Name

H3StreamedBlocks

Category
MAINodes/VRAM Lab
Author
matlowai (Account age: 1004days)
Extension
ComfyUI-MAINodes
Latest Updated
2026-08-26
Github Stars
0.11K

How to Install ComfyUI-MAINodes

Install this extension via the ComfyUI Manager by searching for ComfyUI-MAINodes
  • 1. Click the Manager button in the main menu
  • 2. Select Custom Nodes Manager button
  • 3. Enter ComfyUI-MAINodes in the search bar
After installation, click the Restart button to restart ComfyUI. Then, manually refresh your browser to clear the cache and access the updated list of nodes.

Visit ComfyUI Online for ready-to-use ComfyUI environment

  • Free trial available
  • 16GB VRAM to 80GB VRAM GPU machines
  • 400+ preloaded models/nodes
  • Freedom to upload custom models/nodes
  • 200+ ready-to-run workflows
  • 100% private workspace with up to 200GB storage
  • Dedicated Support

Run ComfyUI Online

H3 Streamed Blocks (exact low-VRAM, alpha) Description

Optimizes processing of large data blocks in AI models for memory efficiency, chunking data for speed and accuracy.

H3 Streamed Blocks (exact low-VRAM, alpha):

H3StreamedBlocks is a specialized node designed to optimize the processing of large data blocks in AI models, particularly in scenarios where memory efficiency is crucial. This node is part of a suite of tools aimed at enhancing the performance of AI models by managing how data is streamed and processed in chunks. The primary goal of H3StreamedBlocks is to facilitate the handling of large datasets by breaking them into smaller, manageable pieces, thereby reducing the memory footprint and improving processing speed. This is particularly beneficial in environments with limited VRAM, as it allows for the efficient execution of complex models without overwhelming the system's resources. By leveraging techniques such as chunking and block replacement, H3StreamedBlocks ensures that data is processed in a streamlined manner, enhancing both the speed and accuracy of AI computations.

H3 Streamed Blocks (exact low-VRAM, alpha) Input Parameters:

block

The block parameter represents the specific data block that is being processed. It is crucial for defining the segment of data that will undergo transformation, ensuring that the node operates on the correct portion of the dataset. This parameter does not have a predefined range as it depends on the dataset being used.

x

The x parameter is the input data that the block will process. It serves as the primary data source for the node's operations, and its characteristics can significantly impact the node's performance and output. There are no specific minimum or maximum values for this parameter, as it is determined by the dataset.

t_emb

The t_emb parameter stands for time embedding, which is used to incorporate temporal information into the data processing. This parameter is essential for models that require time-based data analysis, ensuring that temporal patterns are accurately captured.

mod_segments

The mod_segments parameter defines the segments of the model that will be modified during processing. It allows for targeted adjustments within the model, enhancing flexibility and control over the data transformation process.

rope_freqs

The rope_freqs parameter refers to the rotational position encoding frequencies, which are used to encode positional information within the data. This parameter is vital for maintaining spatial awareness in the data processing, ensuring that positional relationships are preserved.

transformer_options

The transformer_options parameter provides additional configuration settings for the transformer model used in the node. It allows for customization of the model's behavior, enabling users to fine-tune the processing according to their specific needs.

q_chunk

The q_chunk parameter specifies the size of the query chunk, with a default value of 16384. This parameter determines how the data is divided into smaller pieces for processing, impacting both memory usage and processing speed.

kv_chunk

The kv_chunk parameter defines the size of the key-value chunk, also defaulting to 16384. Similar to q_chunk, it influences the division of data into manageable segments, affecting the node's efficiency and performance.

mlp_chunk

The mlp_chunk parameter sets the size of the multi-layer perceptron chunk, with a default value of 16384. This parameter plays a role in determining how the data is processed through the MLP layers, impacting the overall computational load.

kv_block

The kv_block parameter indicates the specific key-value block being processed, with a default value of 0. It is used to identify and manage the particular segment of data within the key-value structure.

probe

The probe parameter is an optional diagnostic tool that can be used to monitor and analyze the node's performance. It provides insights into the processing behavior, aiding in troubleshooting and optimization.

index

The index parameter specifies the position of the data block within the dataset. It is used to track and manage the sequence of data processing, ensuring that the correct order is maintained.

kv_int8

The kv_int8 parameter is a boolean flag that indicates whether the key-value data should be processed using 8-bit integer precision. This option can reduce memory usage but may impact precision.

kv_sage

The kv_sage parameter is a boolean flag that determines whether the SageAttention package should be used for key-value processing. It requires the sageattention package and can enhance processing capabilities.

kv_fp4

The kv_fp4 parameter is a boolean flag that specifies whether 4-bit floating-point precision should be used for key-value processing. This option can significantly reduce memory usage while maintaining reasonable precision.

kv_fp4_exact_av

The kv_fp4_exact_av parameter is a boolean flag that ensures exact average precision when using 4-bit floating-point processing. It provides a balance between memory efficiency and precision.

kv_mix

The kv_mix parameter is a boolean flag that allows for mixed precision processing of key-value data. It offers flexibility in managing memory usage and precision, adapting to the specific needs of the task.

H3 Streamed Blocks (exact low-VRAM, alpha) Output Parameters:

m

The m parameter is the primary output of the H3StreamedBlocks node. It represents the processed model or data block after undergoing the transformations specified by the input parameters. This output is crucial for subsequent stages of data processing, as it contains the refined and optimized data ready for further analysis or use in AI models.

H3 Streamed Blocks (exact low-VRAM, alpha) Usage Tips:

  • To optimize performance, adjust the q_chunk, kv_chunk, and mlp_chunk parameters based on your system's VRAM capacity. Smaller chunk sizes can reduce memory usage but may increase processing time.
  • Utilize the probe parameter to monitor the node's performance and identify potential bottlenecks or inefficiencies in the data processing pipeline.

H3 Streamed Blocks (exact low-VRAM, alpha) Common Errors and Solutions:

Model does not look like MiniMax H3 (no blocks[*].attn.qkv_proj)

  • Explanation: This error occurs when the model being used does not have the expected structure or attributes required by H3StreamedBlocks.
  • Solution: Ensure that the model is compatible with H3StreamedBlocks and contains the necessary attributes, such as attn.qkv_proj, before attempting to process it.

kv_store kvi8s needs the sageattention package; falling back to bf16 (exact)

  • Explanation: This warning indicates that the SageAttention package is required for the kv_sage option but is not available.
  • Solution: Install the SageAttention package to enable the kv_sage functionality, or proceed with the fallback option if precision is not a critical concern.

kv_store kvfp4s needs SageAttention3's fp4 extension

  • Explanation: This warning suggests that the SageAttention3's fp4 extension is necessary for the kv_fp4 option but is not installed.
  • Solution: Install the SageAttention3's fp4 extension to utilize the kv_fp4 feature, or use the fallback option if memory efficiency is not a priority.

H3 Streamed Blocks (exact low-VRAM, alpha) Related Nodes

Go back to the extension to check out more related nodes.
ComfyUI-MAINodes
RunComfy
Copyright 2025 RunComfy. All Rights Reserved.

RunComfy is the premier ComfyUI platform, offering ComfyUI online environment and services, along with ComfyUI workflows featuring stunning visuals. RunComfy also provides AI Models, enabling artists to harness the latest AI tools to create incredible art.

H3 Streamed Blocks (exact low-VRAM, alpha)