Cyberpunk 2077 is already one of the best-looking games available, which makes it a difficult target for image-quality experiments. This project asks a more unusual question: can one RTX 3090 render the game, send its image to a second RTX 3090 for neural rendering, receive the enhanced result back, and then feed that result into the game’s frame generation?
The answer in this test is yes, with a real performance cost. The entire round trip happens before frame generation, so the generated presentation frames are based on the returned neural image. The result felt playable and looked visibly different. This is an intentionally impractical multi-GPU experiment, but both RTX 3090s have active roles in each enhanced frame.
What this recording does not establish is that two GPUs are faster than one when frame generation is involved. The build shown here was an early, unoptimized proof of concept. Because the enhanced image must return before frame generation, the render GPU waits for the neural GPU before the remaining game work can continue. Later tests running the same new architecture on one GPU have produced similar results, so a material dual-GPU performance benefit has not yet been demonstrated for this frame-generation path.
Watch the experiment
The main recording is just over five minutes long: 2560×1440 H.264 video captured from the full display, with AI-generated narration. It shows the live Cyberpunk capture and the neural-rendering comparison process. Three diagnostic comparisons from the same run are included with this post.
Companion measurements
The single-RTX-3090 control and matched single-versus-dual-RTX-3090 comparison were recorded with frame generation disabled. They isolate the cost of running the neural stage locally and the benefit of moving it to the second RTX 3090. Those results provide useful context, but they use a different presentation path and should not be treated as a direct benchmark of the return-before-frame-generation build shown in the main video.
The hardware
The system is built around an AMD Ryzen 9 7900, 64 GB of memory, and an ASUS TUF Gaming X870-PLUS WIFI motherboard.
The GPUs have separate jobs:
- The ASUS RTX 3090 on the chipset-fed PCIe 4.0 x4 path renders Cyberpunk 2077. In the recording it is the game adapter, running at 2560×1440 with DLSS Balanced and path tracing enabled.
- The Gigabyte RTX 3090 is connected through an M.2 OCuLink PCIe 4.0 x4 path. It is the neural-rendering adapter.
- The RTX 5070 Ti is used as the dedicated capture and H.264 encoding adapter. It does not render the game or run the neural stage in this experiment.
- The Radeon RX 7800 XT is connected over USB4 but remains idle in this recording.
This is not pooled VRAM and it is not traditional multi-GPU rendering. Each card owns a separate stage in the pipeline.
The game-rendering RTX 3090 also uses the experimental DLSSG for SM86 4× multi-frame-generation path on RTX 30-series hardware after receiving the enhanced image. The Ampere-compatible runtime reports a maximum of three generated frames per base frame, corresponding to a configured 4× path, and DLSS-G feature creation succeeds. The available telemetry does not independently count every generated frame reaching the display, so 4× describes the configured and accepted runtime path rather than a verified display-frame multiplier.
The two-card round trip before frame generation
The order of operations is the point of this build:
- The chipset-connected ASUS RTX 3090 renders Cyberpunk’s path-traced scene and produces the post-Ray-Reconstruction HDR image with matching motion and depth guides.
- Our bridge sends that image and the guides over the OCuLink PCIe 4.0 x4 route to the Gigabyte RTX 3090, which runs the neural-rendering model.
- The enhanced HDR image comes back to the ASUS RTX 3090 and replaces the relevant game image before the remaining game commands continue.
- Cyberpunk’s downstream post-processing and SM86 frame-generation path receive that enhanced base image. Frame generation can then add presentation frames between enhanced results.
- The RTX 5070 Ti captures and encodes the displayed output for the video.
That return point is essential. If neural rendering happened only in a separate preview window, Cyberpunk’s frame generator would never see its result. Here the result re-enters the game’s own pipeline before frame generation.
This return-to-game path is an experimental local development build. It is not a feature of the linked public MGPU Bridge release.
The diagnostic snapshots use a 1707×960 neural-rendering working image and return to a 2560×1440 presentation. The model was run with Style B, intensity 2, structure 2, skin 2, automatic masking enabled, and one neural pass.
The snapshots recorded approximately 47.4–47.7 ms through the measured return and transfer stage. The completed neural-output rate varies with the live scene and bridge state, and the video’s delivered capture rate is not the same thing as game FPS or unique neural outputs. The second GPU takes the model work away from the render GPU, but the system still has to move a large HDR image and its guides across a PCIe 4.0 x4 link, synchronize the two devices, and return the result before the game can continue. The experiment proves that the architecture can function; it does not make the round trip free.
Performance in practice
I found the run playable at 1440p with path tracing and Neural Rendering. The two-card round trip is expensive: with Neural Rendering armed, the second RTX 3090 completed a median of approximately 14 enhanced images per second, with 12–14 covering the middle 80% of samples. The ReShade finish-effects callback ran at a median of approximately 54 per second, with 50–56 covering the middle 80%.
The order of the stages offers a plausible explanation for the surprisingly fluid result. The SM86 path can place as many as three generated frames between enhanced base frames. Four presentation opportunities for each of roughly 14 neural results would be close to 56 per second, numerically consistent with the 54-per-second callback rate. The five-minute recording averaged 48.9 captured frames per second across both NR-on and NR-off sections.
That numerical relationship is not proof of an exact 4× displayed-frame ratio. The callback rate, capture rate, completed neural-output rate and unique monitor scanout rate are different quantities. Input latency was not measured. The supported conclusion is that the game felt playable while each enhanced base image made a round trip through the second RTX 3090 before frame generation.
What the companion tests show
With frame generation disabled and both game rendering and Neural Rendering running on one RTX 3090, the control recording averaged 17.63 game callbacks per second and 12.05 completed neural outputs per second. The active RTX 3090 averaged 99.89% load while the second RTX 3090 remained idle.
With the same saved game settings and neural configuration but Neural Rendering moved to the second RTX 3090, the observed rates rose to 26.37 game callbacks per second and 21.44 neural outputs per second. Render-GPU load averaged 96.10% and neural-GPU load averaged 98.66%. That represents 49.6% more game callbacks and 78.0% more completed neural outputs in those frame-generation-off recordings.
This is clear evidence that offloading can help when the neural result follows the bridge’s independent presentation path. It does not establish that the same gain survives when the image must return to the game before frame generation.
What the frame-generation test does not prove
In the return-before-frame-generation architecture, the game GPU cannot simply continue independently while the second GPU works. It needs the enhanced image back before downstream processing and frame generation can proceed. The remote route frees neural-compute capacity on the render GPU, but it adds transport, synchronization and a wait for the second GPU. A same-GPU route avoids that round trip but makes Neural Rendering compete with path tracing and frame generation on one device.
Later testing of the new architecture on a single RTX 3090 has produced results in a similar range to the two-RTX-3090 route. Those observations are not a formal matched benchmark, but they mean I cannot claim that the second GPU currently provides a material performance benefit when frame generation is enabled.
The original video also predates the optimization work performed since it was recorded. Those later changes are not represented in the video’s results and should not be used to reinterpret its performance.
What the comparison images show
The three captures cover Misty’s Esoterica, Tom’s Diner, and a crowded market area. The effect is deliberately restrained rather than a dramatic colour filter. The returned image can look cleaner around fine edges, faces, and lighting transitions, while broad scenes may show only a subtle difference.
The comparison PNGs are diagnostic previews generated from saved HDR data. They use a fixed preview transform, so they should not be treated as a perfect reproduction of the game’s final tone mapping.
On the top is the base image, and on the bottom the neural rendering enhanced image.
Misty’s Esoterica

Misty’s Esoterica shows the effect clearly without turning into a dramatically different image. The neural-rendered frame is slightly more contrasty and saturated, with stronger definition around Misty’s face, hair, clothing, the wallpaper, and smaller objects around the counter. At the same time, very fine high-frequency detail is reduced. The result is less like conventional sharpening and more like a redistribution of detail: small pixel-level texture is suppressed while larger, visually meaningful edges and structures become more pronounced.
A numerical comparison supports that impression. Overall brightness is almost unchanged, while luminance contrast increases by about 4% and saturation by about 3.5%. Mid-frequency detail increases substantially, while very-high-frequency energy falls. After accounting for a small alignment difference between the captured frames, structural similarity is about 0.958, so most of the original image remains intact despite the visible change.
The interesting part is around Misty herself. Her eyes, mouth, hair and clothing are more strongly defined, but some of those details are also subtly different rather than simply sharper. That is the neural-rendering trade-off visible in a single frame: the model is reconstructing what it considers plausible detail rather than merely reproducing the original pixels. Whether that looks better is subjective; what is much less ambiguous is that the second GPU is doing materially different image-processing work.
Tom’s Diner

Tom’s Diner shows a somewhat stronger transformation than Misty’s Esoterica. The neural-rendered image is darker in the shadows, more saturated and noticeably higher in local contrast. The officers in the foreground gain stronger separation in their faces, uniforms, badges and equipment, while the diner interior becomes more clearly defined around the stools, floor tiles, counter and illuminated signage. Small lettering such as the NCPD markings also appears more pronounced.
The measurements follow the visual impression. Average brightness falls by only about 1%, but luminance contrast rises by roughly 9%, saturation by 8%, and coarse edge strength by around 22%. At the same time, the finest high-frequency detail falls by about 14%, while mid-scale detail increases. Again, this is not simply a sharpening filter: the neural pass is suppressing some very fine texture while reinforcing larger structures that the eye reads as useful detail.
After correcting for the same small capture alignment offset, structural similarity is about 0.93. That is lower than the Misty comparison, indicating that the neural renderer has changed this scene more substantially. The overall geometry remains intact, but surfaces, edges and character detail have been reinterpreted enough to give the bottom image a distinctly different presentation rather than merely a cleaner copy of the top one.
Market scene

The market scene is the strongest example of the neural renderer changing the overall presentation rather than simply cleaning up isolated details. The returned image is only fractionally darker on average, but luminance contrast increases by about 8% and saturation by roughly 21%. Neon signs, shop lighting and coloured surfaces become noticeably richer, while shadows deepen and the scene gains much stronger visual separation.
Edge strength increases by around 28%, and mid-frequency image detail rises by roughly 29%. That is visible throughout the scene: people in the distance separate more clearly from the background, signage and stall structures are easier to read, and wet pavement, railings, pipes and building geometry gain stronger definition. Unlike a simple sharpening pass, the effect is spread across lighting, colour and structural detail.
After compensating for the small capture offset, structural similarity is about 0.917, making this the most substantially changed of the three examples. The geometry and composition remain the same, but the neural renderer has made a fairly assertive interpretation of the scene. Of the three captures, this one makes the subjective nature of “enhancement” clearest: the bottom image is unquestionably different and more visually pronounced, but whether the stronger contrast and colour are preferable is a matter of taste.
These still images document spatial differences in the selected frames. They do not measure temporal stability, frame pacing, generated-frame quality or input latency.
Software and attribution
The project is an experimental MGPU Bridge / Neural Coprocessor integration by Marcelo Guibout, released under the MIT license. It runs as a ReShade add-on, using ReShade by Patrick Mours/crosire under the BSD 3-Clause license.
Related work and references include OptiScaler and Dagherbou’s DLSS Neural Rendering fork, released under GPLv3, and colour-composition work derived from RenoDX by clshortfuse/Carlos Lopez Jr. under the MIT license. DLSSG for SM86 by sdli1995 provides the Ampere-compatible 4× multi-frame-generation route; its source is GPLv3, while the embedded NVIDIA components remain subject to NVIDIA’s own terms. The public DreamPunk 3.4.X archive is by NextGen Dreams.
Cyberpunk 2077 is developed and published by CD PROJEKT RED. NVIDIA supplies the DLSS and NGX technology used by the neural-rendering runtime. This experiment is an injected research integration and is not an official NVIDIA implementation or an endorsement by any upstream project.
The result
The central result is the complete two-card pipeline: one RTX 3090 renders, the other performs Neural Rendering, the enhanced image returns over PCIe x4, and the game generates presentation frames from that returned image. The cards have distinct jobs and cooperate inside the same game frame.
The image changes are visible, the run felt playable, and the transfer cost is clear in the diagnostics. The experiment proves that the return-before-frame-generation architecture can work. It does not prove that this dual-GPU route is faster than running the same architecture on one GPU.
The frame-generation-off comparison shows a clear benefit from offloading in the bridge’s independent-output architecture. With frame generation enabled, however, the render GPU must wait for the enhanced image to return, and current single- and dual-GPU observations are in a similar range. A material performance benefit from the second GPU remains unproven.
I will continue optimizing the architecture to see whether improvements can produce a material benefit from the second GPU with frame generation enabled.
