Save 4 hours! We auto-setup your workflow! Free!

Drop your workflow.json — we handle every dependency, custom node, and model. Just open the link and run.

Auto-Setup Workflow Json (Free) Now!
ComfyUI > Nodes > ComfyUI-INT4-Fast

ComfyUI Extension: ComfyUI-INT4-Fast

Repo Name

ComfyUI-INT4-Fast

Author
viralvfx (Account age: 1094 days)
Nodes
View all nodes(2)
Latest Updated
2026-07-10
Github Stars
0.03K

How to Install ComfyUI-INT4-Fast

Install this extension via the ComfyUI Manager by searching for ComfyUI-INT4-Fast
  • 1. Click the Manager button in the main menu
  • 2. Select Custom Nodes Manager button
  • 3. Enter ComfyUI-INT4-Fast in the search bar
After installation, click the Restart button to restart ComfyUI. Then, manually refresh your browser to clear the cache and access the updated list of nodes.

Visit ComfyUI Online for ready-to-use ComfyUI environment

  • Free trial available
  • 16GB VRAM to 80GB VRAM GPU machines
  • 400+ preloaded models/nodes
  • Freedom to upload custom models/nodes
  • 200+ ready-to-run workflows
  • 100% private workspace with up to 200GB storage
  • Dedicated Support

Run ComfyUI Online

ComfyUI-INT4-Fast Description

ComfyUI-INT4-Fast enhances Flux-based models in ComfyUI by employing INT4 quantization on RTX 30 and 40 series GPUs, resulting in accelerated inference and reduced VRAM consumption.

ComfyUI-INT4-Fast Introduction

Welcome to ComfyUI-INT4-Fast, a high-performance extension designed to enhance your AI art creation experience by optimizing the way diffusion models are loaded, run, and serialized. This extension is specifically tailored for models in the INT4 format, utilizing the convrot_w4a4 layout. By leveraging the power of GPU Tensor Cores, ComfyUI-INT4-Fast offers ultra-fast and memory-efficient inference, making it an ideal choice for AI artists looking to streamline their workflow and achieve faster results without compromising on quality.

The author has built this extension on the foundational work of BobJohnson24, who developed ComfyUI-INT8-Fast. By adapting key architectures such as dynamic quantization, Hadamard rotation, and model patching, ComfyUI-INT4-Fast supports native INT4 workflows, providing a seamless and efficient experience for users.

How ComfyUI-INT4-Fast Works

At its core, ComfyUI-INT4-Fast operates by quantizing model weights and activations to a 4-bit format, which significantly reduces the computational load and memory usage during model inference. This process is akin to compressing a large image file into a smaller one without losing essential details, allowing for faster processing and reduced storage requirements.

The extension employs a technique known as Hadamard rotation, which rearranges model weights to minimize outliers before quantization. This ensures that the quality of the output remains high, even with reduced precision. Additionally, ComfyUI-INT4-Fast dynamically applies LoRA (Low-Rank Adaptation) patching, which adjusts model weights on-the-fly to maintain coherence with the rotated weight basis.

ComfyUI-INT4-Fast Features

  • Fast INT4 Inference (convrot_w4a4): This feature allows you to load and run models with 4-bit weights and activations, utilizing GPU Tensor Cores for enhanced speed and efficiency.
  • Mixed-Precision Checkpoint Support: The extension intelligently routes standard INT4 layers to Tensor Cores while directing sensitive layers, such as initial and final patch projections stored in INT8 format, to optimized execution paths.
  • On-the-Fly Quantization: Instantly converts standard float checkpoints (BF16/FP16/FP32) to INT4 upon loading, saving time and resources.
  • Dynamic LoRA Patching: Automatically applies Hadamard rotation to LoRA down projection weights, ensuring they align with the rotated weight basis.
  • Checkpoint Serialization: Saves quantized model checkpoints in the standard ComfyUI quantized model format, making it easy to share and reuse models.

ComfyUI-INT4-Fast Models

ComfyUI-INT4-Fast has been tested and verified with the following pre-quantized INT4/INT8 mixed-precision model:

  • Krea2 Turbo INT4: This model is available for download here. It serves as an excellent example of the extension's capabilities, providing fast and efficient performance.

Troubleshooting ComfyUI-INT4-Fast

If you encounter any issues while using ComfyUI-INT4-Fast, here are some common problems and solutions:

  • First Generation Run Delay: The initial generation run may take longer due to model initialization and custom operator setup. This is normal, and subsequent runs will be significantly faster.
  • Model Loading Errors: Ensure that your ComfyUI is updated to the latest version and that comfy-kitchen is installed, as it provides essential execution layouts.
  • GPU Compatibility Issues: If you experience compatibility issues with your GPU, consider checking for updates or consulting community forums for specific solutions related to your hardware.

Learn More about ComfyUI-INT4-Fast

To further enhance your understanding and usage of ComfyUI-INT4-Fast, consider exploring the following resources:

  • Community Forums: Engage with other AI artists and developers to share experiences, ask questions, and find solutions to common problems.
  • Tutorials and Documentation: Look for tutorials that provide step-by-step guidance on using ComfyUI-INT4-Fast effectively in your projects.

By leveraging these resources, you can maximize the potential of ComfyUI-INT4-Fast and elevate your AI art creation process.

ComfyUI-INT4-Fast Related Nodes

RunComfy
Copyright 2025 RunComfy. All Rights Reserved.

RunComfy is the premier ComfyUI platform, offering ComfyUI online environment and services, along with ComfyUI workflows featuring stunning visuals. RunComfy also provides AI Models, enabling artists to harness the latest AI tools to create incredible art.

ComfyUI-INT4-Fast detailed guide | ComfyUI