The useful design assigns each device a job and measures every transfer between them.
A second GPU is most useful when it stops pretending to be part of the first. Two cards give a ComfyUI workflow two memory pools, two compute devices, and a transfer path between them. They do not create one larger allocation that every node can use without knowing where its tensors live. The architecture works when the graph treats one card as the primary workshop and the other as a donor with a specific job.
That job should begin at a component boundary. An image pipeline usually has a text encoder, a diffusion model, a VAE, optional conditioning models, and intermediate latents. Those components peak at different stages. Put the text encoder on the donor card, pass its conditioning output to the primary, then keep the repeated denoising loop and its growing latent workspace on the primary. A VAE can also live on the donor when final decode needs more room than the primary has left.
The placement is valuable because the transfer happens at a stage boundary instead of inside every sampling step. Moving a conditioning tensor once is different from moving model blocks across PCIe for every denoiser pass. Moving a completed latent to a donor-hosted VAE can also be reasonable. Moving the latent and VAE back and forth because two nodes disagree about their device is how a clean capacity win becomes a slow or broken graph.
Current ComfyUI can handle the simple version without this custom node. The stable v0.37.0 source includes Select Model Device, Select CLIP Device, and Select VAE Device nodes, plus a separate MultiGPU CFG Split path for distributing sampling work units (ComfyUI v0.37.0 multi-GPU source). Native component placement should now be the first test. It has the smallest compatibility surface and covers the most useful donor-card pattern: keep whole components on named devices.
ComfyUI-MultiGPU goes further. Its final 2.6.4 code wraps standard and third-party loaders with explicit device choices and adds DisTorch2 allocation for safetensors and GGUF models. DisTorch2 can assign static model blocks by exact byte amounts, ratios, or fractions across the compute GPU, another GPU, and CPU memory (project documentation). A string such as cuda:0,12gb;cuda:1,* is a placement plan for the model, not evidence that CUDA now sees one pooled device.
The project says this directly in its repository notes: the feature improves memory management, not parallel processing. Workflow stages still execute sequentially while components or model blocks live on specified devices. The capacity gain comes from removing static weights from the primary card so its fast memory remains available for activations, attention workspace, larger latents, or more frames. Speed may improve when that placement avoids repeated unload and reload cycles. It may also fall when a frequently used block has to cross a slower link.
That distinction produces a practical order of operations. Start by keeping the diffusion model and sampling workspace on the card with the best combination of compute, supported dtype, and usable VRAM. Move a whole text encoder to the donor if the prompt-encoding stage fits there. Move the VAE only after verifying that encode and decode nodes agree about the device. Place ControlNet or another one-stage component on the donor when its integration explicitly supports that route. Whole-component placement makes the transfer points visible, which makes them measurable.
Layer distribution belongs one step later. Use DisTorch2 only when the diffusion model itself will not leave enough primary VRAM for the target resolution, batch, or frame count. Prefer exact byte allocations over attractive percentages because the cards are not equal and the workflow still needs reserve on each device. Keep the hottest repeated blocks on the primary when the loader permits it. Treat CPU memory as the last donor tier, since a model that fits through host offload can still miss the working-session latency budget.
The strongest objection is that layer placement can make an otherwise impossible workflow run. That is true. A fit result still needs its own label. ComfyUI-MultiGPU’s example workflows cover several mixed-card systems, but they are project tests rather than a matched benchmark for another workstation. The project also claimed potential GGUF speed gains over its first DisTorch implementation, not a universal advantage over native loading or whole-component placement. The only result that transfers cleanly is the method: assign, measure, then keep or revert the split.
The maintenance state now changes the recommendation. The repository declares 2.6.4 as the final version, says maintenance has ended, and will archive on September 30. No successor fork is endorsed (archive notice). The project has no GitHub release or tag for 2.6.4; the version lives in package metadata, while the last functional merge landed May 8 and the September 20 change added the archive notice (version metadata). Pinning an existing working graph is defensible. Choosing this unmaintained custom node for a new workflow now requires owning the compatibility work.
That ownership is not theoretical. An open June report describes async execution changes breaking DisTorch and GGUF multi-GPU nodes (issue #203). A detailed August reproduction on two RTX 3090s reports a process-level CUDA abort when peer access is unavailable and cudaMallocAsync handles cross-device tensors; the same workload ran on one card through native Dynamic VRAM (issue #213). Neither report proves that every dual-GPU graph fails. Both prove that core version, PyTorch build, allocator, peer access, loader, and exact component path belong in the test record.
I would give every placement experiment three verdicts. Fit means the complete graph runs twice at the target resolution, batch, and precision while recording peak VRAM on both cards and peak system RAM. Speed means cold launch, prompt encoding, each sampling pass, VAE decode, and a warm rerun stay inside declared limits. Quality means the same seed, graph, model files, precision, sampler, and step count produce an accepted result after placement changes. Device assignment should not become an excuse to change five other variables at once.
The log needs enough detail to reproduce the route: ComfyUI version, custom-node commit, model and quant, primary and donor device order, per-component placement, DisTorch allocation string, P2P state, allocator, driver, and wall time by stage. Record the failure too. A clean out-of-memory error, a silent fallback to CPU, and a hard process abort describe different compatibility boundaries.
ComfyUI-MultiGPU captured the right systems idea even as its maintenance window closes. Spare VRAM becomes useful when a graph gives it a bounded responsibility. The second card is a donor, not a magic pool, and the transfer path is part of the machine.
If this was useful, forward it to one engineer who needs less noise in their feed.


