View Full Version : madVR - high quality video renderer (GPU assisted)
nevcairiel
15th April 2014, 12:32
madshi: Do you have any plans for using ivtc with 50/59/60 fps sources? It can be a large performance gain, in my case would allow nnedi3 doubling of 720p. I have some clips of 23, 25, 29 and 59 progressive in 720p59 if they are of any use.
Did you try forcing it to IVTC? It should be able to detect any cadence, even if its 6:4 instead of 3:2 due to being 60 fps.
Just toggle the content type with Ctrl-Alt-Shift-T to "Film", and it may just work.
kasper93
15th April 2014, 13:37
What we really need is auto detection when to use film mode. I'm sure madshi will do it when the time comes :)
Tapatalk 4 @ GT-I9300
DragonQ
15th April 2014, 14:14
It's probably quite high on madshi's priority list. Combined with profiles it'd mean settings don't need to manually changed per video any more. :)
madshi
15th April 2014, 14:43
madshi I have noticed something that may or not be an issue. Its not just happening with your latest windowed mode updates, happened before so let me know if you think its worth creating a bug report for. And it actually doesn't seem to negatively effect me at all but you may like to look into it.
I notice sometimes when I start a movie, sometimes the render queue starts at say 15-16/16 but after a few moment drops to 10-11/16 (always exactly this amount), and even if I leave it there for hours with <50% cpu usage, it never grows beyond 10-11/16. The only way I can get it to fill up again is to seek after which it immediately returns to 15-16/16.
I can only seem to reproduce it when actually starting a movie from scratch (I have madVR set to wait for queues to fill up) and when in non fullscreen mode (both FSE or windowed fullscreen don't seem to show it). After seeking I never see it drop below 15-16/16. All other queues arn't effected. Should I report this on the bug tracker incase it does end up causing frame drops on someones machine at some point?
Things like this are very hard to find the cause for. Maybe by looking at very long logs for a very long time I could find something, or maybe not. At the moment, I'd prefer to not touch this, as it would cost a *lot* of time to investigate, and the chance is high that there isn't anything I can do about it, anyway.
The no cadence breaks for a whole episode with new window mode was definitely an exception. Others have been 5-10 down from 15-20 but only usually happen during the black frames between segments and is not noticeable. New window has reduce frame drops from 5-15 to 0-5. Render times have dropped about 10% which gives enough to disable overlay with ivtc+smoothmotion in my case. Secondary display with a different refresh rate then primary now works fine in new window mode, no need for overlay or fse for me anymore.
Sounds good!
Hello everybody, do you think that we can join BlurayDisc "Mastered in 4k" in their full quality? what settings should I set in lav and madvr?
There's a thread on AVSForum where some users analyzed some "Mastered in 4k" discs (by using AviSynth scripts etc) and found that there isn't really any xvYCC data worth talking about. At this point in time I think it's mostly a marketing gimmick by Sony and nothing else.
HTPCs cannot output untouched YCbCr, there is generally always a RGB step in between, so its doubtful this is ever going to work.
True. However, if the display is calibrated to a bigger than BT.709 gamut (e.g. DCI) and if madVR is configured accordingly, the extended colors might still be transported to the display just fine. I'm not really 100% sure, though. The whole gamut topic will get interesting when we get the first 4K Blu-Rays with hopefully a bigger native gamut. At that point I guess I'll have to investigate how to transport that data to the display properly.
I thought xvYCC used negative values (or are values below 16 considered to be "negative"?) to expand the gamut beyond the normal range.
So:
+100% green
−100% red
−100% blue
(i.e. −100% magenta, being the opposite of green)
Would be equal to 200% green.
Which should be possible to represent in RGB if you are doing the appropriate color management and have a display capable of displaying that range of saturation.
Displays calibrated to Adobe RGB are common with photo editing for example.
The encoding is done in YCbCr, but yes, RGB values after YCbCr -> RGB conversion could become negative, in relation to the BT.709 gamut. After conversion to a bigger gamut they might get back into positive range. But I'm not sure whether all of this works correctly in madVR atm, to be honest. If the colors are just slightly outside the valid range, then there should be no problem because madVR has some headroom due to doing all math in video levels. But if the colors are waaay outside of BT.709, then I'm not sure if madVR isn't maybe clipping them, not sure...
I haven't had much time to test, but I think I found a situation where new windowed fullscreen is worse than old windowed fullscreen. It looks like the new version has slightly lower rendering times, but it starts dropping frames "earlier" (at lower frame times) than the old windowed path.
Try cranking up enough features to have average rendering time close to 1/fps. e.g. on 23.976 fps content, around 40 ms. For me, new path drops frames for the same settings that old path doesn't. (With things turned down more reasonably so we're not skirting the 90%+ render times and the GPU barely keeping up with the source, new path works perfectly.)
Yeah, that's quite possible. That will be especially true if your display refresh rate is higher than the movie framerate.
I tried a whole bunch of configs/reformat/updates/official-beta drivers.. giving up...
No go with NED 2x luma on 1920x1080 to 2560x1600 ...
Tried both my 7970cfx and 7870xt downstairs..
/Cry really loudly
I'm willing to buy a 290x.... Anyone try NED 2xluma @ 19x10 to 25x16.
You could try the interop test builds (see next post) to see if they make a difference for you. But at this point I guess your best bet might be an NVidia 750ti, because it seems that with your mainboard the AMD OpenCL interop cost is too high. Alternatively you could replace your mainboard, but I think going over to green land would be the easier solution.
In 87.7 and 87.9 MadVR does not switch into FSE correctly on the first attempt for me. Alt+Enter to go back to windowed mode, and then another time to go to full screen again works. 87.4 worked okay on the first try. I haven't tried any of the intermediate releases.
I'm using Win 7 x64 with MPC-HC 1.7.3 using the Intel HD4600 graphics in my i7-4770k.
Can you please try to find out which exact madVR build introduced this problem? Also a debug log might help figuring out why the switch fails. Please don't switch back and forth in the debug log. Just let it fail, then stop, I don't need to see it working in the log, I only need to see the fail. Please enable the OSD (Ctrl+J) while creating the debug log, because otherwise the log will not contain all important information.
Aliased?
That's irrelevant when you get 1:1 scaling.. from 1280x720 to 2560x1600(1440)
That is as much information as is in the file.. anything ONTOP of that is an approximation.:p
You're aware of that almost all 1280x720 files were downscaled from 2K or 4K masters? The downscaling is usually done with linear interpolation. Something like Lanczos. The key thing to understand here is that a video file like that does not describe rectangular pixels. You need to think of each pixel more of a circle. If you display these files in 2x resolution with Nearest Neighbor scaling you're throwing away potential image quality. Why? Because what you're doing is this:
4K master -> Blu-Ray downscale -> 720p downscale -> 1440p upscale
All the downscales were done with linear interpolation. Which means that each pixel also contains a small portion of the original neighboring pixels. The best way to watch such movies is to upscale them with a good upscaling algorithm. This will get you nearer to the way the image looked in its original resolution.
If you don't believe me, try this:
(1) Take a sharp and detailed photo.
(2) Downscale it with your favorite image editor to 50%, by using a good downscaling algorithm (e.g. Cubic or Lanczos).
(3) Upscale it again 200% to get back to the original resolution.
Now for step (3) try Nearest Neighbor scaling and compare it to e.g. Lanczos. Check which upscaled image looks nearer to the original photo. This test is very valid for video playback, too. After all you don't just want to see what is in the video file, you want to see an image which is as near to the original film scan as possible, don't you?
720p to 25x16 or 15x14 using NN is akin to LOSSLESS conversion
Lossless to the 720p file, yes. But the 720p file itself has a much lower resolution and quality compared to the original film negative. By using a good upscaling algorithm you would get nearer to the original film negative. Do you want to stay lossless to the 720p downscaled source? Or do you want to get as near to the original film negative as possible?
I'd really like to see some more development put into NNEDI to attempt to improve it, I wonder the chances of it improving in future?
NNEDI3 is what it is. There's no way for *me* to improve it. I could just post-process it (e.g. sharpen it), or alternatively I could try to create a new algorithm from the ground up, but I'm not sure if I could even reach NNEDI3 quality with such a new algorithm. If you want the NNEDI3 algorithm itself to be improved, you'd have to talk to tritical who created NNEDI3 in the first place. But I don't think you can expect big improvements.
Basically it sounds like someone needs to create an external Darby-like neural network upscaler.
Darbee is a sharpener, not an upscaler, and Darbee doesn't use a neural network. If you want sharper images, use a sharpening algorithm after NNEDI3 scaling.
Madshi, how does madVR's NNEDI settings stack up vs the default settings in Avisynth filter? Was wondering if we'd see options available for nsize and qual etc.
I've done some image quality comparisons and found an nsize setting of "8x4" to produce the least amount of artifacts for image doubling. It happens to also be the fastest nsize setting available. So the choice was simple for me. I don't plan to offer nsize or qual options in madVR, because the performance cost would not be worth the quality gain. The best way to improve quality is to increase the neuron count, so that's the only option I'm offering. No plans to change that.
1080p: 16 neurons (Not with NNEDI3 ChromaUpscaling, and it's not used often.)
720p < 24fps: 64 neurons
720p > 24fps: 32 neurons
720p > 30fps: 32 neurons
SD < 24 fps: 128 neurons
SD > 24 fps: 128 neurons
SD > 30 fps: 32 neurons
Nice. Looks faster than my AMD7770, although I haven't compared in detail.
Are you using PCI-E 2.0 or 3.0?
PCIe version is only important for AMD users.
madshi: Do you have any plans for using ivtc with 50/59/60 fps sources? It can be a large performance gain, in my case would allow nnedi3 doubling of 720p. I have some clips of 23, 25, 29 and 59 progressive in 720p59 if they are of any use.
It's on my to do list, but not for soon.
Did you try forcing it to IVTC? It should be able to detect any cadence, even if its 6:4 instead of 3:2 due to being 60 fps.
Just toggle the content type with Ctrl-Alt-Shift-T to "Film", and it may just work.
Hmmmm... You're right, it does seem to work. At least it lists 6:4 and plays just fine in 60Hz. Not sure whether the IVTC decimation timestamp manipulations will work properly, though. I guess at 24Hz it would probably play fine. But playing this at 60Hz with Smooth Motion FRC turned on might fail to achieve smooth motion.
What we really need is auto detection when to use film mode.
Yes, we do need that. Unfortunately it's not that easy to implement properly. Especially if we want to take mixed sources (e.g. film content with video overlay) into account.
@madshi:
Blaire linked to your recent workaround (your changelog for 0.87.9) for the NV driver issue and asked me via PM, if there still is a driver fix needed.
I kinda feared this would happen, since your woraround takes the pressure off of NV to fix an issue no one else (besides you and people that use NV hardware with madVR) seems to care about (that's how I interpret it).
It looks to me that they were in the process of working on the fix, but they re-checked if they have the newest madVR version to test against.
Now, what should I tell him? Some technical details would probably be helpful. Also how you (if memory serves right, a madVR user actually came up with the idea) worked around the bug, so NV knows where and what to search for.
To be honest, madVR doesn't need a fix, anymore. The workaround works fine and doesn't have any negative side effects. That said, it's a clear bug in the NVidia drivers, from what I can see, so they might still want to fix it.
Basically the old madVR builds did this:
for each video frame do
{
clTargetTex = clCreateFromD3D9TextureNV(...);
clEnqueueAcquireD3D9ObjectsNV(clTargetTex);
clSetKernelArg(clTargetTex);
clEnqueueNDRangeKernel(...);
clEnqueueReleaseD3D9ObjectsNV(clTargetTex);
clFinish(...);
clReleaseMemObject(clTargetTex);
}
With this code, older NVidia drivers worked fine, but newer NVidia drivers either do nothing, or write zeroed out data to the target texture.
The latest madVR builds now use the following approach instead, which works around the issue:
clTargetTex = clCreateFromD3D9TextureNV(...);
for each video frame do
{
clEnqueueAcquireD3D9ObjectsNV(clTargetTex);
clSetKernelArg(clTargetTex);
clEnqueueNDRangeKernel(...);
clEnqueueReleaseD3D9ObjectsNV(clTargetTex);
clFinish(...);
}
clReleaseMemObject(clTargetTex);
madshi
15th April 2014, 14:50
Here's a new test build set for AMD users wanting to do NNEDI3:
http://madshi.net/madVRinteropTest.rar
In the rar file are two madVR.ax files which use different methods to try to improve the interop problem. Unfortunately the improvement is probably not as large as I had hoped, but there should be a small improvement at least. Probably one build will work better than the other build. Please try both and let me know which build works better for you. I've intentionally removed the rendering times from the OSD (only for these test builds, of course) because due to the way these 2 test builds work, judging them by looking at the rendering times would be misleading. So please judge these builds by testing which build allows you to use higher/more quality settings.
Looking forward to your feedback!
(FWIW, I've concentrated on NNEDI3 luma doubling, with NNEDI3 chroma upscaling and NNEDI3 chroma doubling disabled. Enabling those might still work, but I've not tested that.)
James Freeman
15th April 2014, 15:11
Is there a problem with NNEDI3 and AMD?
Not long ago it was Nvidia that didn't work at all, now its AMD?
michkrol
15th April 2014, 15:21
Is there a problem with NNEDI3 and AMD?
Not long ago it was Nvidia that didn't work at all, now its AMD?
On AMD it's a performance only problem - it works correctly, just slower than it should, because of the way AMD('s driver) goes around DX->OpenCL interop.
DragonQ
15th April 2014, 15:31
Here's a new test build set for AMD users wanting to do NNEDI3:
http://madshi.net/madVRinteropTest.rar
In the rar file are two madVR.ax files which use different methods to try to improve the interop problem. Unfortunately the improvement is probably not as large as I had hoped, but there should be a small improvement at least. Probably one build will work better than the other build. Please try both and let me know which build works better for you. I've intentionally removed the rendering times from the OSD (only for these test builds, of course) because due to the way these 2 test builds work, judging them by looking at the rendering times would be misleading. So please judge these builds by testing which build allows you to use higher/more quality settings.
Looking forward to your feedback!
(FWIW, I've concentrated on NNEDI3 luma doubling, with NNEDI3 chroma upscaling and NNEDI3 chroma doubling disabled. Enabling those might still work, but I've not tested that.)
Whilst playing a 640x480p/25 file with 16 neurons and Smooth Motion enabled (60 Hz):
v0.87.9: 35-40 dropped frames per second; render queue is 1-2/8; present queue is 0-1/8; GPU load ~95%
Test 1: 1-2 dropped frames per second; render & present queues are 0-4/8 or 1-5/8 typically; GPU load ~80%
Test 2: 0 dropped frames per second; render & present queues are 4-7/8 or 5-8/8 typically; GPU load ~82%
Test 2 seems the best for me. Still can't use 32 neurons though, I get a dropped frame every few seconds and GPU usage rises to 89%.
MS-DOS
15th April 2014, 15:54
Here's a new test build set for AMD users wanting to do NNEDI3:
http://madshi.net/madVRinteropTest.rar
Let's see. On my 5870 (Win 7 x64, 13.12) the interop cost was insane, as I posted here (http://forum.doom9.org/showthread.php?p=1673786#post1673786) (the image is dead, argh).
Tested the new builds on 480 -> 1080 (+J3AR) content in FSE (new path), which gave me about ~8-10 dropped frames per second even with 16 neurons before.
TestBuild1 - Seems to work smoothly up to 64 neurons, 128 starts to give loads of presentations glitches and the playback stutters quite a lot, but it doesn't report any dropped frames, thou. GPU load is stuck at ~63%.
TestBuild2 - Seems smooth up to 128 (!) neurons with no dropped frames or presentation glitches, ~64% GPU load. Setting it to 256 neurons puts 99% load on the GPU and I'm starting to get frame drops.
The improvement overall looks very large to me, TB2 is a beast. Could you implement these two in your OpenCL benchmark? I'd really like to see the raw numbers :D
Great work!
huhn
15th April 2014, 16:24
Hmmmm... You're right, it does seem to work. At least it lists 6:4 and plays just fine in 60Hz. Not sure whether the IVTC decimation timestamp manipulations will work properly, though. I guess at 24Hz it would probably play fine. But playing this at 60Hz with Smooth Motion FRC turned on might fail to achieve smooth motion.
IVTC with something else like 3:2 normally never works fine. madvr doesn't drop the right frame with right detected 4:2:2:2 and playback is unwatchable and this on a 23 hz tv.
@tesbuilds
for me on a r9 270 the build 1 is "faster"
i tested 256 neuron 480p23 to 1080p. with the old build it is impossible with both new builds it works but with test 2 all queue drop but no frame is dropped. with test1 all queue fill up after some time so i think this is working better.
i get 82 % gpu usage test1 and 84% with test2 both drop like crazy with opend gpu-z so they should't be judge with gpu-z
TheLion
15th April 2014, 17:05
Let's see. On my 5870 (Win 7 x64, 13.12) the interop cost was insane, as I posted here (http://forum.doom9.org/showthread.php?p=1673786#post1673786) (the image is dead, argh).
Tested the new builds on 480 -> 1080 (+J3AR) content in FSE (new path), which gave me about ~8-10 dropped frames per second even with 16 neurons before.
TestBuild1 - Seems to work smoothly up to 64 neurons, 128 starts to give loads of presentations glitches and the playback stutters quite a lot, but it doesn't report any dropped frames, thou. GPU load is stuck at ~63%.
TestBuild2 - Seems smooth up to 128 (!) neurons with no dropped frames or presentation glitches, ~64% GPU load. Setting it to 256 neurons puts 99% load on the GPU and I'm starting to get frame drops.
The improvement overall looks very large to me, TB2 is a beast. Could you implement these two in your OpenCL benchmark? I'd really like to see the raw numbers :D
Great work!
This is very exciting news indeed. My 5870 prevented me from using NNEDI3 at all. I will try these test builds as soon as I can - here is hope that at least chroma upsampling for 1080p is now possible, as well as SD->1080p.
tFWo
15th April 2014, 17:14
Same as @huhn with my 270x.
Build1 is slightly better than build2. Slightly lower gpu load and (maybe) faster queue filling.
Both builds allow much higher NNEDI settings than 87.9. :)
720p24->1680x1050@60
87.9 using both chroma upscaling 32neurons and luma doubling 32neurons was just below the treshold for smooth playback (40.5ms)
new builds allow 64 neurons on both settings (around 85% load)
SD@24->1680x1050@60
87.9 128 neurons was usable on both
new builds allow 256 on both or 128 on both + 32quad for luma (also around 85%)
aminfri
15th April 2014, 17:20
About time i reported some stats too:
Using the latest test builds with Hi10 720p to 1080 and these settings:
Jinc 3 AA, Chroma upscaling,
jinc 3 AA, Image upscaling,
Catmull-Rom AA SLL, image downscaling,
Smooth Motion enabled,
Dithering, Error Diffusion 1,
I could easily get 64 Neurons with both test builds, but the usage with the first build (76%) was just a bit lower that the build 2 (78%). Previously i couldn't enable Image doubling without frame drops. So these builds are definitely huge improvements.
On 128 Neurons i was getting frame drops left and right with both builds.
System specs in sig.
seiyafan
15th April 2014, 17:32
Here's a new test build set for AMD users wanting to do NNEDI3:
http://madshi.net/madVRinteropTest.rar
How do I use it? Just paste into MadVR folder?
leeperry
15th April 2014, 17:36
How do I use it? Just paste into MadVR folder?
Backup your existing mVR folder, then copy all the files in there and alternatively rename both builds to madVR.ax
w00t, moar testing :)
I suppose that implementing those changes in the test app woulda been too much work but it didn't work on my box anyway.
Was kinda looking for a reason to avoid going green, let's see how that goes :p
Farfie
15th April 2014, 17:37
On my Win7 x64 HD5850 machine, TB1 is a very clear winner going from 720p -> 1440p. I'm able to use 64 neurons now, which is very close to my GTX680. With TB2, I get about 1 frame drop per "OSD refresh tick," and with the original I get anywhere between 3-5 per. Of course, this is at an overclock of 800mhz for the core (above 725 default), so any AMD user might want to push for this, since this was enough to get 64 neurons for luma doubling at this resolution with TB1 :)
I don't know why my results differ from DragonQ and MS-DOS with TB1 being better than TB2 very clearly. Perhaps it has to do with the resolution size. I speak for nothing though, as madshi will probably know why :)
seiyafan
15th April 2014, 17:48
Great work Madshi! 1080->1440 Before it's dropping 10-15 frames a second, now 0!
Now a question, for movies which of the following provides more visual improvement? debanding or ED?
huhn
15th April 2014, 17:51
Great work Madshi!
Now a question, for movies which of the following provides more visual improvement? debanding or ED?
if needed debanding for sure.
TheLion
15th April 2014, 17:59
Great work, madshi!
On my Win 8.1 64bit i7 system with AMD 5870 (latest beta Catalyst) both test builds show huge improvements for NNEDI3 (doubling as well as chroma upscaling).
Now I can finally use it at all - the limits to the max settings are the same for both builds. testbuild2 seems to die more gracefully when "overloaded": TB1 shows massive amounts of repeated frames in addition to the dropped.
chroma upscaling for 1080p works now up to 32 neurons - it wasn't fast enough before at all.
noee
15th April 2014, 18:11
Win7 x64, HD6570 PCI-E 2.0x16
24p 720x368 P010 (OrderedDith NNEDI luma doubling/32n/SMFRC off) => {1080 playback@59.942Hz}
879: GPU ~95+%, Render(1-4/14) - Present(5-7/10), occasional frame drop
OP1: GPU ~89+%, Render(12-14/14) - Present(7-10/10), no drops
OP2: GPU ~85+%, Render(12-14/14)/Present(8-10/10), no drops
seiyafan
15th April 2014, 18:14
if needed debanding for sure.
what if the video quality is high, like blu-ray? Would it still benefit more from debanding than dithering?
MS-DOS
15th April 2014, 18:21
I hope it's just some kind of a bug with TB1, which causes constant presentation glitches to me when GPU load is above a certain value, and can be fixed. Because, like for most people posted above, to me TB1 has slightly lower GPU cost than TB2.
I tested with SM disabled, ordered dithering, and SC80 chroma upscaling, all Q4P disabled, except subtitles optimization and "don't render frames when fade in/out detected".
James Freeman
15th April 2014, 18:24
what if the video quality is high, like blu-ray? Would it still benefit more from debanding than dithering?
When you are at the edge of the big dilemma of "Visible Quality" vs "Machine Power", I suggest to pick the one which is more visible or beneficial for the picture quality.
In that case, go for Debanding instead of a heavier and almost invisible (imo) dithering algorithm.
Same goes true for NNEDI3 vs Smooth Motion for example.
Judder free playback outweighs slight improvement in scaling aliasing a hundredfold.
baii
15th April 2014, 18:36
Also factor in fan noise when you push the gpu hard. Especially in a laptop set up.
leeperry
15th April 2014, 18:38
No night/day difference between both builds on my 7850/Haswell rig, both run quite a bit faster than 0.87.9. If anything the first picture of a movie shows up faster with the first build, GPU memory and D3D usage are also the lowest. I vote 1 :)
64x nnedi for chroma & luma 29.97 960x540@1080p:
1: http://thumbnails110.imagebam.com/32102/695090321017680.jpg (http://www.imagebam.com/image/695090321017680) 2: http://thumbnails112.imagebam.com/32102/502f8d321017681.jpg (http://www.imagebam.com/image/502f8d321017681) 0.87.9: http://thumbnails111.imagebam.com/32102/4ea789321017682.jpg (http://www.imagebam.com/image/4ea789321017682)
128x nnedi for chroma & luma 25fps 640x480@1080p:
1: http://thumbnails111.imagebam.com/32102/33658a321017715.jpg (http://www.imagebam.com/image/33658a321017715) 2: http://thumbnails109.imagebam.com/32102/883541321017713.jpg (http://www.imagebam.com/image/883541321017713) 0.87.9: http://thumbnails109.imagebam.com/32102/833901321017714.jpg (http://www.imagebam.com/image/833901321017714)
flashmozzg
15th April 2014, 18:57
No night/day difference between both builds on my 7850/Haswell rig, both run quite a bit faster than 0.87.9. If anything the first picture of a movie shows up faster with the first build, GPU memory and D3D usage are also the lowest. I vote 1 :)
Try without HW monitoring tools.
kasper93
15th April 2014, 19:31
Here's a new test build set for AMD users wanting to do NNEDI3:
Good work. Seems to be a lot faster. build2 is better for me. I can do 32 neurons on 720p->1080p while with build1 it drops frames even with 16 neurons. So my vote is for build 2 :) Comparing to stable this is BIG improvement.
Still we should somehow reach AMD and made them fix that ;/
EDIT:
Build2 is memory hungry, 3.6GB of "system commit" was freed after closing player. I needed to close some programs, because I got only 6GB RAM, 3GB pagefile, 1GB of gpu mem which is full, but I used to it already ;p Windows notified me during playback that I run out of memory. But I had already around 5GB used.
iSunrise
15th April 2014, 19:58
...To be honest, madVR doesn't need a fix, anymore. The workaround works fine and doesn't have any negative side effects. That said, it's a clear bug in the NVidia drivers, from what I can see, so they might still want to fix it.
Basically the old madVR builds did this:
...
Thanks. I just forwarded everything to Blaire. It's their decision now.
turbojet
15th April 2014, 20:39
Forcing ivtc with deint=ivtc on film in 59 fps source detects 6:4 cadence but doesn't remove frames and gpu load remains high.
29i works fine with force film mode now, it didn't last I checked months ago. Unfortunately 59i doesn't even when double framerate deinterlacing allowed it still detects 2:2 and plays at 29 fps.
leeperry
15th April 2014, 21:09
Try without HW monitoring tools.
I initially did, reason why I thought hard figures would be more meaningful.
ThurstonX
15th April 2014, 21:25
Finally found some time to do a few quick tests. AMD Radeon R7 200 Series; passively cooled (SAPPHIRE Ultimate 100368USR Radeon R7 250 1GB 128-Bit GDDR5 PCI Express 3.0); Catalyst 14.2; Core i5-3470; 8 GB RAM
Display is an old Sharp Aquos LC-32GA5U running at native 1366x768 via DVI
tl;dr
v0.87.9 couldn't run without dropping frames; Test1 ran with Luma doubling forced at 16 neurons (32 was too much) and Jinc 3 AR; Test2 could only handle Lanczos 3 AR, so I vote for Test1.
Hope this helps, and thanks for the test builds. Definitely a step in the right direction!
I started with v0.87.9 trying to force NNEDI3 to double Luma resolution using 16 neurons. Plenty of dropped frames.
Settings
Chroma upscaling: Bicubic 75 (No AR)
Image doubling: use NNEDI3 to double Luma; always - if upscaling is needed; 16 neurons
Image upscaling: Jinc 3 AR
Image downscaling: Catmull-Rom scale in linear light
No Debanding
Smooth motion: Enable, only if there would be motion judder without it...
Dithering: Ordered; use colored noise; change dither for every frame
Trade quality for performance: first five items checked
Exclusive mode settings at default
With v0.87.9 I got the following:
Queues
Decoder: 13-16/16
Upload: 6-8/8
Deinterlace: 5-8/8
Render: 2-4/8
Present: 0-2/8
Tons of dropped frames
with Test1
Decoder: 14-16/16
Upload: 7-8/8
Deinterlace: 6-8/8
Render: varied from 5-7/8; 6-7/8; 6-8/8
Present: varied from 4-5/8; 4-6/8
1 frame repeat every 3.63 secs
NO dropped frames
Source video (a VHS capture using an old Hauppauge card)
Format : MPEG-PS
File size : 8.90 GiB
Duration : 1h 40mn
Overall bit rate : 12.7 Mbps
Video
ID : 224 (0xE0)
Format : MPEG Video
Format version : Version 2
Format profile : Main@Main
Format settings, BVOP : Yes
Format settings, Matrix : Custom
Format settings, GOP : M=3, N=15
Duration : 1h 40mn
Bit rate : 12.0 Mbps
Width : 720 pixels
Height : 480 pixels
Display aspect ratio : 4:3
Frame rate : 29.970 fps
Standard : NTSC
Color space : YUV
Chroma subsampling : 4:2:0
Bit depth : 8 bits
Scan type : Interlaced
Scan order : Top Field First
Compression mode : Lossy
Bits/(Pixel*Frame) : 1.159
Time code of first frame : 00:00:00:00
Time code source : Group of pictures header
Stream size : 8.46 GiB (95%)
Audio
ID : 192 (0xC0)
Format : MPEG Audio
Format version : Version 1
Format profile : Layer 2
Duration : 1h 40mn
Bit rate mode : Constant
Bit rate : 384 Kbps
Channel(s) : 2 channels
Sampling rate : 48.0 KHz
Compression mode : Lossy
Delay relative to video : -111ms
Stream size : 275 MiB (3%)
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
Been playing an NTSC DVD (Peter Gabriel Live in Athens 1987). Same resolution, interlaced, but dropped frames using Jinc 3 AR. OK switching to Lanczos 3 AR. Didn't check any other differences between the two sources.
2nd edit:
I think the difference is that the DVD is really 16:9, while the VHS capture is 4:3. That's my guess, anyway.
Asmodian
15th April 2014, 21:56
what if the video quality is high, like blu-ray? Would it still benefit more from debanding than dithering?
ED is a large performance hit and ordered dither is quite good, if you have banding debanding + OD is better than no debanding + ED. debanding + no dither is not a reasonable option.
leeperry
15th April 2014, 22:00
I believe this was already mentioned but 48/96/192 neurons NNEDI would be pretty cool for when you can't quite run for 64/128x and still got potential cycles unused. I still kinda find NNEDI too sharp for chroma but sometimes I still don't have enough horse power to go for the next level and I can't run 256 neurons luma for SD@1080p, I would welcome the opportunity to try 192 if technically doable :)
It's even more true now that AMD boards have magically earned extra headroom. :thanks:
Fullmetal Encoder
15th April 2014, 22:13
I'd love to test the new build but I keep getting "dxva processing failed" message from madVR. I don't have any of those options selected (for dxva) in the UI and I'm using Radeon 5850 on Windows 7 Pro. I would add that this is with ED 2 selected.
kasper93
15th April 2014, 22:20
@Fullmetal Encoder: Make sure to disable all "enhancements" in gpu driver. Especially "dynamic contrast" there seem to be bug in drivers which breaks DXVA processing, for example deinterlacing will fail to initialize when opening, yet you can re-enable it later.
QBhd
15th April 2014, 22:45
Time for my report on the test builds.
System:
R9 270X factory OC to 1120/1400 (GPU/memory)
PCI-e 2.0 (GA-990FXA-UD7)
Windows 8.1
Target resolution:
1024x768 (rectangular pixels)
Source:
1280x720p24
Settings:
Chroma Upscaling - NNEDI3 x32
Image Upscaling - Jinc 3 AR
Luma Doubling - NNEDI3 x64
Chroma Doubling - NNEDI3 x32
Image Downscaling - Catmull-Rom AR LL
Debanding - med/med
With previous release I could only do Ordered Dithering (Error Diffusion just pushed over the limit of GPU)
Testbuild1 - Still dropped frames with ED
Testbuild2 - NO dropped frames with ED
So my vote goes to Testbuild2... It allows for me to go even further than any build to date
QB
Fullmetal Encoder
16th April 2014, 00:06
Scaling from 720x480 to 1920x1200 with ED2 and NNEDI doubling at 32 neurons on luma using a Radeon 5850 I am getting:
- 52 frame drops/refresh with 87.9
- 40 frame drops/refresh with test 1
- 31 frame drops/refresh with test 2
I don't know why others with the 5850 are getting so much better performance though.
Fullmetal Encoder
16th April 2014, 00:08
@Fullmetal Encoder: Make sure to disable all "enhancements" in gpu driver. Especially "dynamic contrast" there seem to be bug in drivers which breaks DXVA processing, for example deinterlacing will fail to initialize when opening, yet you can re-enable it later.
Thank you very much! I don't know how, but all of those "enhancements" were on in CCC. Although I'm not sure how they got turned on since I turned them off long ago :o
tickled_pink
16th April 2014, 00:50
Win7 x64, HD 7750 PCI-E 2.0x16
720x404@25fps with 64 neurons tested
Test build 1 uses slightly less (55% vs 57%) GPU and less graphics memory (~40 MB or 10%) than test build 2.
0.87.9 used ~60% GPU and similar amount of memory as test build 1.
Neither improved performance enough to allow more neurons but a 10% overall improvement is certainly welcome!
sajara
16th April 2014, 01:23
This test came a bit as a shock because I do remember being unable to use NNEDI3 even with 16 neurons when first release and didn't bothered to try again.
AMD 5730M 650Mhz core /800Mhz GDDR3 mem
H264 clip 720x304 -> 1366x768
87.9 - 16 Neurons ~86.7% / 32 Neurons - slideshow
Test 1 - 16 Neurons ~57.6% / 32 Neurons ~86.5%
Test 2 - 16 Neurons ~60% / 32 Neurons ~89.7%
Queues the same in test 1 and 2.
So again beyond words on the improvement and as much, amazed being able to do 32 neurons.
ryrynz
16th April 2014, 01:49
Do the test builds improve anything on Nvidia hardware at all?
Procrastinating
16th April 2014, 06:40
Considering the previous responses, and the less meaningful low render times, I changed the survey defaults to double luma, and added a default video of tears of steel. Remember that any data is good data, and this survey/spreadsheet will not only help madshi, but those interested in what the optimal media GPU for them might be.
To Madshi in particular, I think it will be interesting to see across the various hardware configurations, how the improvements appear between versions. I will probably try the AMD builds at some point.
Asmodian
16th April 2014, 07:00
Do the test builds improve anything on Nvidia hardware at all?
Yes, but only with SLI on. SLI is much better with the test builds though still not as fast as without SLI.
madVR 87.9:
1280x720p24 -> 2560x1440 @ 72Hz, Bicubic75 AR chroma, Bicubic75 AR image, NNEDI3 128 Luma doubling, No smooth motion, no debanding, Ordered Dither, 3DLUT calibration, Windowed Overlay.
SLI on 41.2ms
GPU0 81% @ 1097 MHz, 19% PCI-E, 7% memory controller
GPU1 18% @ 836 MHz, 16% PCI-E, 0% memory controller
SLI off 29.9ms
GPU0 67% @ 1097 MHz, 5% PCI-E, 7% memory controller
GPU1 00% @ 324 MHz, 0% PCI-E, 0% memory controller
GTX Titans, 3770K @ 4.6 GHz, Z77 chipset, each GPU is on PCI-E 3.0 x8, 32GB DDR3-2133CL9.
madVR interopTest1 & interopTest2 (the two are identical as far as I can tell):
SLI on
GPU0 73% @ 1097 MHz, 11% PCI-E, 7% memory controller
GPU1 09% @ 836 MHz, 7% PCI-E, 0% memory controller
SLI off
GPU0 67% @ 1097 MHz, 5% PCI-E, 7% memory controller
GPU1 00% @ 324 MHz, 0% PCI-E, 0% memory controller
I can also run my "720p24" profile with SLI on which used to drop a lot frames. Jinc3 chroma, Jinc3 Image, NNEDI3 128 luma doubling, ED2, no debanding, no smooth motion.
I did recheck 87.9 immediately after these tests and it does perform as it did before so this isn't an accidental setting or system change. :)
Nvidia Driver 337.50
I believe this was already mentioned but 48/96/192 neurons NNEDI would be pretty cool for when you can't quite run for 64/128x and still got potential cycles unused. I still kinda find NNEDI too sharp for chroma but sometimes I still don't have enough horse power to go for the next level and I can't run 256 neurons luma for SD@1080p, I would welcome the opportunity to try 192 if technically doable :)
It's even more true now that AMD boards have magically earned extra headroom. :thanks:
Sadly I don't think finer grained neuron settings are possible. From the NNEDI3 docs:
nns -
Sets the number of neurons in the predictor neural network. Possible settings are
0, 1, 2, 3, and 4. 0 is fastest. 4 is slowest, but should give the best quality. This
is a quality vs speed option; however, differences are usually small. The difference
in speed will become larger as 'qual' is increased.
0 - 16
1 - 32
2 - 64
3 - 128
4 - 256
Default: 1 (int)
Another impressive update madshi, and I don't even have an AMD GPU. Thanks again!
James Freeman
16th April 2014, 08:34
Another impressive update madshi and I don't even have an AMD GPU. Thanks again!
I'm pretty sure you're wrong, unless madshi is a real living magician.... :)
Asmodian
16th April 2014, 08:37
Huh? did you read my post?
James Freeman
16th April 2014, 08:45
Ohhhhhh.... I see.
There should be a comma there.
Like so:
Another impressive update madshi, and I don't even have an AMD GPU. Thanks again!
Not like so (what I thought):
Another impressive update, madshi and I don't even have an AMD GPU. Thanks again!
:D
Asmodian
16th April 2014, 08:51
OH! haha yes, I never saw that reading. :)
Procrastinating
16th April 2014, 11:54
Alright, after testing the new test builds, I can confirm that, on my HD6770 I go from
Old: Many drops, render times ~46ms on 720p->1080p source, using luma 32 doubling
New interop 1: ~0.2 drops per second
New interop 2: ~ 1 drop per second.
The problem now, is that I'm no longer seeing render times in the debug window (with the new builds).
I can conclude however, that I am at least getting the fastest render times from test build 1, and the improvements are at least enough to prevent noticeable framedrops on a particular source now.
romulous
16th April 2014, 11:57
The problem now, is that I'm no longer seeing render times in the debug window (with the new builds)!
Quoting from that same post in which madshi posted the download link (two lines under the link itself in fact):
I've intentionally removed the rendering times from the OSD (only for these test builds, of course) because due to the way these 2 test builds work, judging them by looking at the rendering times would be misleading. So please judge these builds by testing which build allows you to use higher/more quality settings.
Procrastinating
16th April 2014, 12:06
My bad, but the conclusion stands at least.
That said, I wonder where the difference in results for the two builds lie, with some people reporting better results in either. It doesn't appear to be related to overall architecture, so possibly clocks?
vBulletin® v3.8.11, Copyright ©2000-2026, vBulletin Solutions Inc.