Two mismatched cards work best as a production line, not a shared bucket of VRAM.
A 5080 and a 3090 do not become one 40GB video GPU. They become a useful local video workstation when each card owns a stage. Give the 24GB card the long, memory-heavy generation job. Give the 16GB card bounded preparation and finishing work. The handoff between them is part of the architecture, not an inconvenience to hide.
The hardware points toward that split before any workflow opens. NVIDIA lists the RTX 3090 with 24GB of GDDR6X and the RTX 5080 with 16GB of GDDR7. The 5080 is the newer Blackwell card. The 3090 still has fifty percent more capacity. Video generation frequently makes capacity the first constraint, while restoration, interpolation, image preparation, and encoding can fit inside a smaller bounded worker. That makes the older card the sensible generator candidate and the newer card the sensible finishing candidate, subject to measurement on the exact nodes in use.
Video models punish vague capacity planning. Resolution, frame count, attention backend, quantization, tiling, sliding-window size, and enabled controls all change the memory bill. WanGP’s current model guide explicitly rejects fixed VRAM tiers and recommends starting with a distilled variant, testing a short clip, then adding duration, resolution, controls, and upsampling one variable at a time (WanGP model guide). A successful two-second prototype is evidence for that configuration. It is not a promise that a ten-second 720p clip will fit.
The default topology should therefore be two isolated workers. Run one ComfyUI environment with only the 3090 visible and a second environment with only the 5080 visible. ComfyUI’s current startup reference exposes --cuda-device, --port, separate input and output directories, and network binding for exactly this kind of process boundary (ComfyUI startup flags). Verify the device indexes at launch instead of assuming card zero is always the 5080. Give each worker its own Python environment, custom-node lock, cache, temp path, and output path.
Isolation buys more than cleaner memory accounting. A generator stack can pin the ComfyUI core, attention backend, model loader, and custom nodes that make one Wan, Hunyuan, or LTX workflow stable. The finishing stack can carry a different PyTorch build or restoration node without forcing the generator to move with it. Updating one worker becomes a deliberate compatibility event instead of a workstation-wide gamble.
The 3090 worker should do one job first: produce a short accepted source clip. HunyuanVideo 1.5 is a useful example of the pressure involved. Its official repository lists a 14GB minimum with model offloading enabled and notes that disabling offload can improve speed when more memory is available (HunyuanVideo 1.5). LTX’s official desktop path sets 16GB as the local-generation floor on Windows and Linux, while its ComfyUI 2.5 workflows distinguish a lighter single-stage distilled graph from a two-stage graph that adds a 2x upscale and refine pass (LTX Desktop, LTX 2.5 workflows). Those are project support boundaries, not measurements on this machine. They still show why the 3090’s extra 8GB belongs where the longest repeated sampling loop lives.
Prototype below delivery resolution. Lock the model files, seed, prompt, input image, frame count, frame rate, sampler, steps, attention backend, precision, and offload policy. Generate one short clip, inspect motion and composition, then change one dimension. More frames and more pixels multiply the spatiotemporal work before the clip has proved it deserves either. The first goal is an accepted motion plate, not a final export.
The handoff should be boring. Write the generator result to a job directory containing numbered lossless frames or a documented mezzanine file, the workflow JSON, and a small manifest. Record the model and file hashes, ComfyUI and custom-node revisions, seed, frame count, frame rate, dimensions, color assumptions, and the generator’s peak VRAM, peak system RAM, and wall time. Mark the directory complete only after the expected frame count and hashes verify. The 5080 worker should never watch a half-written output and guess that generation finished.
The 5080 worker can then prepare reference images, interpolate an accepted clip, restore selected frames, upscale, and encode the delivery file. None of those assignments is automatic. SeedVR2’s official ComfyUI project says even its 3B path can need at least 18GB in the base configuration, then offers VRAM preservation, smaller batches, FP8 models, and block swapping as ways to trade speed for fit (ComfyUI SeedVR2). A 16GB 5080 therefore gets a measured SeedVR2 experiment, not a guaranteed job. Start with a small frame batch and the target output size. Keep it only if fit, speed, and temporal quality all pass. A lighter upscaler is the correct fallback when the restoration model spends the session crossing the host bus.
Frame interpolation needs the same discipline. Run it after the generator clip is accepted, not inside every prototype. Compare the interpolated output against the source for edge doubling, occlusion errors, and motion that changes meaning. Faster hardware does not rescue a finishing stage that invents bad in-between frames. The 5080 earns the stage by reducing end-to-end delivery time while preserving the clip, not by reporting high utilization.
The two workers do not need to run simultaneously to be useful. Sequential operation already protects each environment and makes memory ownership obvious. Concurrency becomes a second experiment because both processes still share system RAM, storage bandwidth, PCIe resources, cooling, and the power envelope. Measure generator-only, finisher-only, and overlapped runs before treating pipeline overlap as free throughput.
Automation can wait until the manual handoff is dependable. ComfyUI_NetDist can drive two local instances with different ports and CUDA devices, then queue work and fetch the remote result (ComfyUI_NetDist). Its own documentation also notes rough edges around seeds and multi-machine scaling. A watched folder plus a validated manifest is less elegant and easier to debug. Add remote queueing only after the artifact contract survives retries, partial files, and a worker restart.
A single cross-device ComfyUI graph remains worth testing, especially when whole-component placement avoids reloading a large text encoder or VAE. It should not be the baseline for this workstation. Cross-device tensors, custom-node compatibility, allocator behavior, and hidden fallbacks make a monolithic graph harder to reproduce than two workers joined by files. Require a measured improvement over the staged pipeline before accepting that complexity.
The acceptance record needs three verdicts. Fit means both workers complete their assigned stage twice at the target settings without exhausting VRAM or RAM. Speed means generator time, handoff time, finishing time, and optional overlap meet a declared production budget. Quality means the source motion, restored detail, interpolated frames, and final encode pass a repeatable visual review. A load screen proves none of them.
This is an occasional production box, not a reason to manufacture a daily-video obligation. Queue a small set of deliberate clips, preserve the working environments, and spend the expensive generation pass only after a low-resolution prototype earns it. The workstation becomes realistic when every card has a bounded job and every boundary leaves evidence.
If this was useful, forward it to one engineer who needs less noise in their feed.


