View Full Version : madVR - high quality video renderer (GPU assisted)
tp4tissue
3rd July 2019, 15:28
Now, if you take a look at the 2160p25 and 1080p25 profiles, we can see that the settings are pretty high (not maxed out, of course).
Other refresh rates and/or resolutions (e.g. 720p30) obviously have worse settings, but the question is: do we really care about those? :)
Here's what an Enthusiast should say:::
If anyone is watching anything other than 1080p remux and 2160p remux, that itself is the problem, there should never be any need to upscale 720p, because anything that's still in 720p isn't worth watching or upscaling. <turns 30 degrees, fold arms, /combative smirk>
tp4tissue
3rd July 2019, 15:33
Agreed, just buy a used (!) 1060 6GB (MSI Gaming series are really low noise/silent as well).
The heatsink looks adequate, how many vrm does this version have.
Basically there are 3+1, 4+1 and 6+1 version cards. Because most ebay cards are miners, the Safest cards are the 6+1.
el Filou
3rd July 2019, 15:50
Here's what an Enthusiast should say:::I'm a 'content enthusiast' in addition to being a video quality enthusiast, and I have content that's simply not available in HD.
(I also have DVDs I could just buy again on Blu-ray but I'd rather invest my money in hardware or new content than do that).
Vrams are not created equally.
I already told them the very same here and here, asked them for similar graphs, never got a reply
you want some here as brutal as possible thanks to mpcVR native d3d11 render path.
https://abload.de/img/d3d9copybackrnjtc.png
https://abload.de/img/d3d11nativev8j3g.png
as additional informations:
i use the slowerest available hevc ready GPU i have here to show the biggest deltas, pretty much everything should be done using d3d11 video processing which uses very little GPU resources no interop as far as i know the only difference how the data get's to the renderer.
is so little work it not even starting to boost properly this is far less then it looks like. still excellent showing from d3d11 native.
so here starts the real issue why are user using a much much faster 1080ti loosing 10-15 % with madVR?
oldpainlesskodi
3rd July 2019, 17:22
Here's what an Enthusiast should say:::
If anyone is watching anything other than 1080p remux and 2160p remux, that itself is the problem, there should never be any need to upscale 720p, because anything that's still in 720p isn't worth watching or upscaling. <turns 30 degrees, fold arms, /combative smirk>
You Sir, earn a beer for that one!
In a perfect world, would one need Madvr I guess is the point, and yet, here we all are :D
chros
3rd July 2019, 17:55
The heatsink looks adequate, how many vrm does this version have.
Basically there are 3+1, 4+1 and 6+1 version cards. Because most ebay cards are miners, the Safest cards are the 6+1.
Good question, I'm not an expert of this, but this page says (https://www.overclockers.com/msi-gtx-1060-gaming-x-6g-video-card-review/) 5+1.
Note that MSI gaming (X) series is one of the most expensive cards on the market (in a given line-up).
I'm a 'content enthusiast' in addition to being a video quality enthusiast, and I have content that's simply not available in HD.
(I also have DVDs I could just buy again on Blu-ray but I'd rather invest my money in hardware or new content than do that).
Yes, that's a good point and there can be still 720p broadcasts as well.
But that's what I meant about asking these details as well, because it can be unimportant to somebody else.
@huhn, thanks, I'll take a look at it later (I don't have access to certain sites here).
hotripper
3rd July 2019, 17:57
Yeah, I said screw it and bought an adapter from work...
Yeah I am there with you though, I would like to get a new receiver as well, it woulkd be nice to use only one cable, and not have to switch inputs all the time but I have better things to spend my money on right now, in time I will switch.
Can I recommend a nice little utility called Audio Switcher it is nice, it lets you assign an icon for audio devices, click to switch, rename, and hotkeys, etc. To be honest the best feature is having the icon show so you know which audio out is selected at any given time.
chros
4th July 2019, 10:26
you want some here as brutal as possible thanks to mpcVR native d3d11 render path.
I'm not sure that that's the best example, it's not even madVR, not to mention not NGU, but I see the same thing on the screenshot I reported here (https://www.avsforum.com/forum/26-home-theater-computers/2364113-guide-building-4k-htpc-madvr-96.html#post58177072): GPU usage 8% vs 3% -> ~5% difference.
so here starts the real issue why are user using a much much faster 1080ti loosing 10-15 % with madVR?
Not to mention that you used an SD file for the comparison (according to the screenshot), and not a ~60GB 4K remux (https://www.avsforum.com/forum/26-home-theater-computers/2364113-guide-building-4k-htpc-madvr-95.html#post58167358).
Do I miss something or misinterpreted something?
Charky
4th July 2019, 11:21
I only don't agree with this part: there's no such thing as future proofing in the world of PC, never was :) I just buy what (I think) I need at the moment.
To prove your point : the immediate future is mostly HDMI 2.1 (and 4k 60 Hz RGB full chroma) and no current GPU supports it.
I'm not sure that that's the best example, it's not even madVR, not to mention not NGU, but I see the same thing on the screenshot I reported here (https://www.avsforum.com/forum/26-home-theater-computers/2364113-guide-building-4k-htpc-madvr-96.html#post58177072): GPU usage 8% vs 3% -> ~5% difference.
this is an 960 in idle playing an UHD file... not a 1060 not a 1080 ti they are supposed to be faster you don't pay more for less performance.
it's 5 % at 1 ghz on an 960.
you want to know the difference between 2 decoding option why would you care about NGU? you want to know the processing difference between ways to get the data to the renderer nothing else.
the GPU has 0.1 watt difference in power consumption between d3d9 copyback and d3d11 native.
and now the most important thing if mpcVR can do it this fast there is no reason madVR can do it this fast or even is this fast it's just far harder to test without code debugging.
chros
4th July 2019, 12:20
this is an 960 in idle playing an UHD file... not a 1060 not a 1080 ti they are supposed to be faster you don't pay more for less performance.
it's 5 % at 1 ghz on an 960.
So there IS a diff for You as well not just for us! (the used clock speed doesn't matter until it's the same for both tests)
you want to know the difference between 2 decoding option why would you care about NGU?
Because until we don't use the *exact* same test setup we can't be sure about the outcome. (and see below)
you want to know the processing difference between ways to get the data to the renderer nothing else.
But the problem is: the bigger the data that has to go to the system ram the difference is higher!
So, in summary, if you are still interested testing this on your system, I can provide you the test file and all the madvr setting to use (!) during the weekend.
(Otherwise I don't see the point to talk about this anymore :) , no offense.)
so you can't read the memory load?
and what bigger data the decoded frame size from a youtube video with the same size as a BD have 100 % the same size. system ram what are you talking about...
yeah everyone has to use my test settings and file or there test is wrong...
chros
4th July 2019, 12:47
system ram what are you talking about...
dxva2 copyback uses sytem ram, doesn't it?
clsid
4th July 2019, 14:02
RAM is there to be used.
Native is obviously more efficient than copyback. But the performance impact, while noticeable, isn't that huge that makes it a necessity to use.
the system memory shouldn't have an impact on the GPU performance as long as it can do it in realtime. maybe if it is still not clear i tested an UHD file not an SD file here.
the down and upload operation are effecting the GPU.
clsid
4th July 2019, 14:08
the immediate future is mostly HDMI 2.1 (and 4k 60 Hz RGB full chroma) and no current GPU supports it.
DisplayPort 1.4 to HDMI 2.1 converter (https://www.anandtech.com/show/14535/realtek-demonstrates-rtd2173-displayport-14-to-hdmi-21-converter)
That would be compatible with AMD Navi, NVIDIA RTX, and Intel Gen11 (IceLake).
chros
4th July 2019, 14:35
the system memory shouldn't have an impact on the GPU performance as long as it can do it in realtime.
...
the down and upload operation are effecting the GPU.
That's correct, but the ram on the GPU is waaaay faster than the system ram, and it seems that the GPU has to wait for the data.
if it is still not clear i tested an UHD file not an SD file here.
No, it wasn't for me, maybe I missed something on the screenshot.
Native is obviously more efficient than copyback. But the performance impact, while noticeable, isn't that huge that makes it a necessity to use.
Actually, it is: that was our point with (at least) Manni.
Even the cropped picture (with black bar detection) using dxva2 copyback is slower (!) than processing the whole image with d3d11 native (using the same settings in madvr) :)
Anyway, I stop this conversion for now, this is how it works on our systems, I/we don't want to convince anybody, everyone can try it for themselves and do/think/believe whatever they like :)
el Filou
4th July 2019, 16:05
Maybe there's some kind of 'stall' somewhere that blocks the GPU doing other tasks while it's doing texture transfer for copyback, and you don't notice the difference unless you have heavy rendering settings?
On my old Core 2, copyback with 2160p24 HDR maxes out the GPU and rendering times shoot up to 50-65 ms (can't test the Haswell unfortunately as my Radeon doesn't do HEVC).
System RAM definitely has a big impact on copyback performance, so I think it doesn't depend on if you have a 960 or a 1080 Ti but on how fast your memory subsystem is. Maybe also the CPU being busy with other stuff has an impact? For example, did you test if disabling black bar detection changes anything?
there are multiply issues with that too.
you don't have PCIe 3 and so your ram should be faster then your PCIe 2.0.
the CPU speed could cripple it too but i doubt blackbar detection could effect it much because it is not the same program and you have more then 1 core there still worth investing.
the first consumer grade CPU with PCIe 3 was ivy bridge AFAIK.
and just have a look at the bus load.
you don't have to test HEVC and and HDR doesn't help here anyway h264 should be good enough.
tp4tissue
4th July 2019, 22:03
DisplayPort 1.4 to HDMI 2.1 converter (https://www.anandtech.com/show/14535/realtek-demonstrates-rtd2173-displayport-14-to-hdmi-21-converter)
That would be compatible with AMD Navi, NVIDIA RTX, and Intel Gen11 (IceLake).
Don't bet on these converters, for example, alot of dp1.4 to hdmi 2.0 converts today will output the wrong gamma curve. :scared:
tp4tissue
4th July 2019, 22:04
Has anyone tested 4K-> 8k madvr performance ? 1080ti enough ?
Asmodian
5th July 2019, 03:52
For which settings?
My 2080 Ti can do 4Kp60 to 8Kp60 with NGU medium for both chroma and luma (~12ms), NGU high takes ~16ms, while very high takes ~44ms. Using Bicubic for chroma upscaling instead drops rendering times by about 2.5ms. This is without HDR, artifact removal, post processing, or trade quality for performance options.
nsnhd
5th July 2019, 05:17
Is NGU very high for luma worth sacrificing other settings, if your GPU can manage ? Or NGU high is good enough ?
VAMET
5th July 2019, 08:11
Dear Friends
How madVR is dealing with AMD CPU and GPU? I am going to buy Ryzen 7 3700 and RX5700XT. Will it be enough for better settings in madVR?
Sincerely
nevcairiel
5th July 2019, 09:37
Noone here has touched either of those, so we don't know. That said CPU is mostly irrelevant for madVR. GPU, we'll have to see. Polaris was somehow rather bad for madVR, if that continues with NAVI we won't know until someone tests.
chros
5th July 2019, 10:05
For example, did you test if disabling black bar detection changes anything?
I don't remember :) But for me the only advantage of using copyback would be to utilise black bar detection+cropping to save performance, and that's not the case. :) Otherwise I don't mind the full image processing and it will make to write profile rules easier.
tp4tissue
5th July 2019, 13:08
For which settings?
My 2080 Ti can do 4Kp60 to 8Kp60 with NGU medium for both chroma and luma (~12ms), NGU high takes ~16ms, while very high takes ~44ms. Using Bicubic for chroma upscaling instead drops rendering times by about 2.5ms. This is without HDR, artifact removal, post processing, or trade quality for performance options.
Can it do 4k 24p + Lanczos chroma + NGU Luma + HDR-SDR + 3DLut , Not worried about 60p material, only 4K Bluray remuxes
tp4tissue
5th July 2019, 13:12
Is NGU very high for luma worth sacrificing other settings, if your GPU can manage ? Or NGU high is good enough ?
No, put everything into NGU Luma first. :cool:
nsnhd
5th July 2019, 13:55
No, put everything into NGU Luma first. :cool:
I mean Luma in my question :)
Charky
5th July 2019, 16:48
Is NGU very high for luma worth sacrificing other settings, if your GPU can manage ? Or NGU high is good enough ?Trust your own eyes. Can you see the difference ? If you can't, then it's good enough [emoji16]
tp4tissue
6th July 2019, 03:53
I mean Luma in my question :)
Between H and VH on luma, if you're more than 3 feet away, the visual difference is small, but VH is indeed sharper, and you CAN see this in fine lines like animal fur. sifu from kungfu panda , closeup on his furry ears are a good test.
For chroma, I wouldn't even recommend NGU, because it'd be a waste of electricity. Even at 3 feet, it becomes very hard to see the difference between ngu VH vs lanczos. at 3feet plus, I would say it's impossible, I tested for myself, I got it wrong 50% of the time. :D
chros
6th July 2019, 07:41
For chroma, I wouldn't even recommend NGU, because it'd be a waste of electricity. Even at 3 feet, it becomes very hard to see the difference between ngu VH vs lanczos. at 3feet plus, I would say it's impossible, I tested for myself, I got it wrong 50% of the time. :D
:) Do you use chroma 4:4:4 with your display?
nsnhd
6th July 2019, 08:16
Between H and VH on luma, if you're more than 3 feet away, the visual difference is small, but VH is indeed sharper, and you CAN see this in fine lines like animal fur. sifu from kungfu panda , closeup on his furry ears are a good test.
I'm less than 3 feet away, so I'll try to keep NGU luma on very high over other settings.
As @Asmodian stated above, NGU very high costs nearly triple on performance (44/16ms) of high, so it must use the most GPU power for a better result.
chros
6th July 2019, 08:52
Native is obviously more efficient than copyback. But the performance impact, while noticeable, isn't that huge that makes it a necessity to use.
did you test if disabling black bar detection changes anything?
I don't remember :) But for me the only advantage of using copyback would be to utilise black bar detection+cropping to save performance, and that's not the case. :) Otherwise I don't mind the full image processing and it will make to write profile rules easier.
My last report about this using GPU 1060 6GB (max, underclocked) freq is 1544Mhz:
- first 3 minutes of Shazam 23p 4k HDR BD remux (~75GB, video bitrate 76.7 Mb/s) on a 4K screen
- external srt subtitle is used (MPC-BE internal sub filter)
- LAV filters
- madvr:
-- hdr passthrough
-- only chroma upscaling is applied: NGU Sharp High
-- dithering: Error Diffusion 2
-- no trade quality option is checked
-- full screen window mode
-- 10 bit output if possible
GPU usage results (checked with nvidiainspector):
- dxva2 native: 76% - 80%
- dxva2 copy-back, - crop: 83% - 87%
- dxva2 copy-back, + crop: 83% - 91%
- d3d11 native: 73% - 77%
- d3d11 copy-back, - crop: 85% - 88%
- d3d11 copy-back, + crop: 87% - 95%
There's the ~10% difference on my system. The closest performer is dxva2 native but with its obvious flaws.
Interestingly enough, cropping (with copy-back modes) increases GPU usage and don't reduce it (it uses the same profile, so result is valid).
I'll be curious about your results/graphs with similar test case, guys, including your system (mine is in my signature).
You do realize that cropping maybe using a different profile that you have set up. Zoom Control Cropping should always have a lower usage.
QB
tp4tissue
6th July 2019, 14:10
:) Do you use chroma 4:4:4 with your display?
Of course I do, and it's not an option. 444 or GTFO. :devil:
tp4tissue
6th July 2019, 14:13
You do realize that cropping maybe using a different profile that you have set up. Zoom Control Cropping should always have a lower usage.
QB
zoom ctrl doesn't always work with NGU, so it's not a win win always.
w/ NGU it sometimes crops to 3838 then you get lanzos which obviously pushes above 39ms total w/ all the other toppings.
el Filou
6th July 2019, 17:37
you don't have to test HEVC and and HDR doesn't help here anyway h264 should be good enough.Unfortunately the Radeon 7870 doesn't even support 4K H264, so useless to test. The real limitations of copyback decoding only start to become a problem with 10-bit 4K, because it takes up 8 times the bandwidth of 8-bit FHD. With lower resolutions, using copyback or native doesn't have an impact on which madVR settings I am able to use. With 4K 10-bit it does.
So I've done some benchmarks on my HTPC...
Notes:
- 'GPU' and 'video' numbers are the frequency reported / usage % (e.g. 1290 MHz at 50% load = 'GPU 645')
- CPU and GPU usage counters of DXVA Checker are completely wrong, I don't know how it computes them. Maybe the GPU counter reports only shader usage so it could be right but not useful, but the CPU usage is always wrong. I used HWMonitor which gives the same values as other monitoring tools.
CPU @ 3500 MHz (FSB 333), RAM verified dual channel
1. 4K HEVC 10-bit. Best of 5 passes with DXVA Checker decode/playback, best of 3 runs with madVR. Playback at 1920x1080. madVR settings: scale chroma separately; no compromise on HDR quality; SSIM2D downscale; clip pre-measured for HDR; no black bars detection.
RAM @ 666:
Decode: 63,0 fps, CPU 65, GPU 1006, Bus 24
Playback: 34,9 fps, CPU 81, GPU 731, Bus 21
madVR: 439 dropped frames, avg 50,16 ms, max 78,17 ms, GPU 1772, CPU 95
RAM @ 800:
Decode: 66,8 fps, CPU 60, GPU 1017, Bus 25
Playback: 34,5 fps, CPU 84, GPU 656, Bus 21
madVR: 315 dropped frames, avg 45,68 ms, max 63,78 ms, GPU 1772, CPU 90
For reference, with Native:
Decode: 178,5 fps, CPU 46, GPU 1642, video 1467
Playback: 177,0 fps, CPU 54, GPU 1785, video 1467
madVR: 0 dropped frames, avg 34,38 ms, max 38,46 ms, GPU 1613, CPU 68
Difference of dropped frames and max render times under madVR just with 20% faster RAM is massive.
With DXVA Checker, CPU is not fully loaded with decode and only 6% faster decode with 20% faster RAM. Software/platform inefficiency?
2. Same test but with madVR 'light' settings: compromise on HDR quality checked; Bicubic downscaling instead of SSIM2D
copyback: avg 16,5 ms, max 25,04 ms, GPU 1136, CPU 78
native: avg 14,93 ms, max 17,78 ms, GPU 592, CPU 25
max render time is 40% better while GPU is two times less loaded, CPU three times less loaded. Massive performance impact.
I understand why CPU would be loaded if it has to wait for frames to be read/written from/to system RAM, but why more GPU load? Can't the GPU render a frame it has received from the renderer while the next queued frame from the decoder is transfered over the PCIe bus and back?
A single 4K P010 frame is 25 MB, at PCIe 2 x16 it should take 3,125 ms, 6,25 ms round-trip just for the time over the bus. If the rendering has stalls it could explain the difference of a few ms between copyback & native even with very high end GPUs.
3. A lighter test comparing Jellyfish clip at 1080p HEVC, same bitrate, in 8-bit and 10-bit:
decode 8-bit: 266,5 fps, CPU 54, GPU 1797, video 1430, bus 11
decode 10-bit: 210,7 fps, CPU 60, GPU 1797, video 996, bus 22
playback 8-bit: 240,2 fps, CPU 75, GPU 1797, video 1141, bus 15
playback 10-bit: 181,9 fps, CPU 75, GPU 1743, video 852, bus 20
We see 10-bit decode takes up exactly two times the bus bandwidth as 8-bit, as expected.
The 10-bit decode performance doesn't scale to 4x the speed of the 4K clip (would be 267 fps).
for reference, 10-bit native: 299,2 fps, CPU 18, GPU 1613, video 1415, bus 2
4. Just out of curiosity I underclocked the CPU to 2100 MHz (FSB 200), to be able to test more different RAM speeds:
(Jellyfish 10-bit DXVA Checker decode):
RAM @ 400: 131,6 fps (native 268,4), CPU 76, GPU 1589, video 989, bus 13
RAM @ 533: 139,6 fps (native 275,7), CPU 70, GPU 1642, video 909, bus 14
RAM @ 666: 148,8 fps (native 281,1), CPU 67, GPU 1642, video 798, bus 15
RAM @ 800: 146,0 fps (native 281,8), CPU 68, GPU 1428, video 766, bus 15
for reference, CPU @ 3500 & RAM @ 800: 210,7 fps, CPU 60, GPU 1797, video 996, bus 22
With same RAM speed but 66% faster CPU, 40-45% more fps.
With same (slow) CPU speed but 66% faster RAM, 13% more fps.
nevcairiel
6th July 2019, 19:44
70% CPU usage on Copy-Back is not a typical result, really. On NVIDIA or Intel you should see extremely low CPU usage, if you have a relatively recent CPU, since both of those will use the DMA engines to copy the image, which does not result in high CPU usage.
AMD, especially on older generations, has been notoriously bad with copy-back, and I would not recommend using it there, or using it as a testing reference for any meaning beyond those cards specifically.
Unfortunately I couldn't really determine from your post which hardware was used.
el Filou
6th July 2019, 20:46
Yes it's old it's the one from my sig, Core 2 E7400.
Is the DMA method possible starting from the CPUs with integrated memory controller?
Edit: LAV says 'cb direct', if that's useful.
tp4tissue
7th July 2019, 02:06
70% CPU usage on Copy-Back is not a typical result, really. On NVIDIA or Intel you should see extremely low CPU usage, if you have a relatively recent CPU, since both of those will use the DMA engines to copy the image, which does not result in high CPU usage.
AMD, especially on older generations, has been notoriously bad with copy-back, and I would not recommend using it there, or using it as a testing reference for any meaning beyond those cards specifically.
My HTPC build with g3258 @ 4.7ghz will go from 60-70% for 4K Dx11 Copyback.
I would expect i3 to have something like that too, but i5 and above shouldn't
i get 15 % CPU usage on an i3 4130 and it isn't even using the full clock... 2.5-3 ghz playing UHD 10 bit 59p using d3d9.
with 10 bit 23p i get about 5 % at 1.1 ghz.
is the copyback operation AVX2 optimised and that's why?
tp4tissue
7th July 2019, 14:01
is that on full bitrate 4k remux ?
it's a broadcast sample and we are talking about decoded frames here they have always the same size with a 10 mbit source or a 125 mbit.
YukonTrooper
8th July 2019, 05:58
Hi, guys. Trying to troubleshoot my framerate sync issue. For a couple months I've been getting frame drops/skips where I haven't had them before. ReClock stopped working and I've been having issues with custom resolutions but still testing.
My composition rate is 23.971. Is that normal?
svengun
8th July 2019, 09:54
My composition rate is 23.971. Is that normal?
I think ideal would be 23.976
chros
8th July 2019, 11:35
You do realize that cropping maybe using a different profile that you have set up. Zoom Control Cropping should always have a lower usage.
Thanks, but it uses the same "2160p25" profile (https://www.avsforum.com/forum/26-home-theater-computers/2364113-guide-building-4k-htpc-madvr-95.html#post58165252) (it crops the image to 3840x1608), so that's not it, but something else.
We can continue here (https://forum.doom9.org/showthread.php?t=176642).
seiyafan
8th July 2019, 17:47
Anyone got Navi yet? Curious to know how it performs. =)
just lay back and wait there are no custom designs yet and the navi blower is one of the worst stock cooler.
petran79
8th July 2019, 19:47
I encountered a crash in Potplayer but I do not know if it is due to Madvr or potplayer.
Unhandled exception occured [0xC000000D@0x000000004A535AB4] at MadVR64.ax
Problem appears if while playing the video at full screen, I press the minimize button and then the maximize button. While sound plays normally, I get a white screen after maximize and program crashes with that message.
vBulletin® v3.8.11, Copyright ©2000-2026, vBulletin Solutions Inc.