Log in

View Full Version : Media Player .NET (MPDN) - D3D HQ GPU Video Renderer [v2.49.0/v1.31.0 27 Dec 2018]


Pages : 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 [60] 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96

ryrynz
17th August 2015, 11:30
Does DX12 work with the older driver?

Your D3D12 test works okay, dunno about everything else though.

Zachs
17th August 2015, 12:04
Well then there's no reason why NVIDIA needed to remove the extension!

ryrynz
17th August 2015, 12:13
Yeah it's all a bit weird.. is it broken or removed? I think Manuel was looking into it, I might ask for an update.

BTW if anyone has their Windows 10 start menu stop working like I just did (also kills right click and left functionality on the taskbar)
then go here http://www.thewindowsclub.com/start-menu-does-not-open-windows-10 and run number 4)
You'll need to run an elevated powershell
http://serverfault.com/questions/464018/run-elevated-powershell-prompt-from-command-line

Bugger if I know what caused it.. Probably should've stuck with 8.1 for a bit longer.

Belphemur
17th August 2015, 13:03
Yeah it's all a bit weird.. is it broken or removed? I think Manuel was looking into it, I might ask for an update.

BTW if anyone has their Windows 10 start menu stop working like I just did (also kills right click and left functionality on the taskbar)
then go here http://www.thewindowsclub.com/start-menu-does-not-open-windows-10 and run number 4)
You'll need to run an elevated powershell
http://serverfault.com/questions/464018/run-elevated-powershell-prompt-from-command-line

Bugger if I know what caused it.. Probably should've stuck with 8.1 for a bit longer.

I advise you to simply replace it: http://www.classicshell.net/

Since Windows 8 I'm using this Start Menu ... I tried Windows 10 Start menu for a couple of week and really ... I prefer the classic shell.

ryrynz
17th August 2015, 13:13
I advise you to simply replace it: http://www.classicshell.net/

Since Windows 8 I'm using this Start Menu ... I tried Windows 10 Start menu for a couple of week and really ... I prefer the classic shell.

Yup, well aware of that one and Vistart too but will be running StartIsBack++ I think on my main machine when I upgrade, the new menu just doesn't have the same functionality. A little disappointing but hey I guess it works for the masses well enough.

I really thought it was an update I did, or the RAM speed I changed or the driver install I did. Real pain when you have one issue that springs up after you've made quite a number of changes recently and you just have to revert them all. Sure didn't help that I didn't have a restore point >.<

Anyway.. Windows 10.. not exactly a painless experience. I feel for the guys that have had their LCD's destroyed by what appears to be Nvidia's 353.62 driver (https://forums.geforce.com/default/topic/862417/geforce-drivers/windows-10-official-353-62-drivers-are-killing-samsung-and-lg-notebook-lcd-display-panels/1/). Fun times.

aufkrawall
18th August 2015, 07:02
Yeah it's all a bit weird.. is it broken or removed?
It is removed deliberately, according to ManuelG. A beta tester for Nvidia said it won't come back. Pretty certain, but not an official statement though! However, better bury any hopes.

Since OpenCL interop is or can be bugged on AMD as well (stupid PowerPlay bugs, especially with PCIe 2.0), I think it's not worth investing any time in new features that would rely on interop.

Shiandow's NNEDI3 SM 5.0 implementation works perfectly for me. GPU usage is a bit higher, but afair power consumption isn't.
GM200 probably can do 1080p60 64 neurons doubling via SM 5.0.

Btw: Could Asynchronous Compute of GCN be useful with DX12 for NNEDI3?

Zachs
18th August 2015, 08:11
I'm working on another code path for SM5.0 NNEDI3. This may or may not be faster but I'll probably need the AMD owners to test it out since I don't have an AMD GPU.

Anyway I'll make another post when it's ready. It'll require a new version of MPDN too unfortunately.

Edit: Not sure about async compute, but from my early research into DX12, it would seem that interoperability with D3D9 had been dropped. If it's true then DX12 support won't be possible since most of the shader codes are in SM3.0.

aufkrawall
18th August 2015, 09:40
I'm working on another code path for SM5.0 NNEDI3. This may or may not be faster but I'll probably need the AMD owners to test it out since I don't have an AMD GPU.

Sounds interesting, I'm curious how fast it will be.

I don't know how much time it would cost you, but what do you think of a "MPDN Benchmark" tool?
It would run each shader as fast as possible and measure finishing times. I think it would be very beneficial for determinating the best video card for HQ HTPC purposes. :)

ryrynz
18th August 2015, 09:53
I don't know how much time it would cost you, but what do you think of a "MPDN Benchmark" tool?
It would run each shader as fast as possible and measure finishing times. I think it would be very beneficial for determinating the best video card for HQ HTPC purposes. :)

Ah, we've been down this road before. Love me some benches..

Got a good name for this too Zach..

The Determinator :cool:

trandoanhung1991
18th August 2015, 12:24
Anyone else having a problem with the reclock script? It's causing, for me, a lot of dropped and delayed frames. I'm using SVP with it, might that be the problem?

ryrynz
18th August 2015, 12:31
Anyone else having a problem with the reclock script? It's causing, for me, a lot of dropped and delayed frames. I'm using SVP with it, might that be the problem?

Yeah possibly.. why not disable it and see, it's a quick test. SVP can be pretty intensive.

Zachs
18th August 2015, 12:34
I don't think you should use Reclock with SVP. MPDN doesn't know what the actually frame rate is in that case since SVP can change it at any time.

Zachs
18th August 2015, 12:36
Ah, we've been down this road before. Love me some benches..

Got a good name for this too Zach..

The Determinator :cool:
That's definitely doable with player extensions and render scripts. I wouldn't do it yet though since the algos are still changing and being optimised.

trandoanhung1991
18th August 2015, 14:37
Yeah possibly.. why not disable it and see, it's a quick test. SVP can be pretty intensive.

Yep, disabling the script keeps the frame rate stable.

I don't think you should use Reclock with SVP. MPDN doesn't know what the actually frame rate is in that case since SVP can change it at any time.

Something happened when both of them are on which causes frame rate to drop below 59, down to as low as below 40 for brief moments, which causes drops and delays.

I have another question. Is it possible to get MPDN to play BDMVs folders?

aufkrawall
18th August 2015, 23:43
Shouldn't MPDN be able to open jpegs via LAV the same way as it can open pngs?
It can't open this one:
http://abload.de/img/testqwud0.jpg

Zachs
19th August 2015, 00:01
I just tried your JPG with GraphStudioNext (LAV Splitter Source --> LAV Video Decoder --> Video Renderer) and it couldn't display anything too.

Zachs
19th August 2015, 00:02
I have another question. Is it possible to get MPDN to play BDMVs folders?

I don't think that's supported by LAV Splitter Source?

nevcairiel
19th August 2015, 06:46
I don't think that's supported by LAV Splitter Source?

It is, tell it to open the index.bdmv (for automatic playlist selection) or one of the playlists in the PLAYLIST folder, and it'll work.

Zachs
19th August 2015, 07:10
Test Build for MPDN v2.41.0 (Build 3320) is out (use it with the latest GitHub extensions) - http://mpdn.zachsaw.com/Test%20Builds/3320/

NNEDI3 (SM5.0) now has an alternate code path under optimizations. As I only have access to an old Fermi and an Intel P4600 currently, I can't say how much better or worse it'll perform beyond these two GPUs. My old first gen Fermi (NVS4200M) doesn't like this code path at all. However, on the P4600, it has sped up NNEDI3 by up to 100% in some situations. On the Intel GPU, the SM version was already much faster than the OpenCL version without this alternate code path but it is now more than twice as fast. The Intel GPU seems to always prefer the "Scalar and Small Code" optimization with alternate path whereas the first gen Fermi absolutely hates it (twice as slow).

I'd like your help now to find out what the best option for your GPU is - and I suspect different generations of GPU even from the same vendor will prefer different optimizations.

Nvidia and Intel GPUs have starkly contrasting fortunes with the alternate code path, so I'm very curious to see what it would do with an AMD GPU as well as other generations of Nvidia GPUs.

aufkrawall
19th August 2015, 07:43
I can confirm your Fermi findings with Maxwell.
Using alternate code path doubles GPU load to 100% and I have lots of dropped frames. No matter if vector or scalar.

Anima123
19th August 2015, 08:03
720p -> 1080p, GTX 880M + HD 4600 Optimus, Alternate weight access method used

neurons 64 + 64,
Prefer Scaler: 61ms
Prefer Vector: 61ms
Prefer Scaler & Small Code: 106ms (picture errors)
Prefer Vector & Small Code: 113ms (picture errors)

The error picture looks like:
https://www.dropbox.com/s/sb621se03b00r7k/Error_Using_Small_Code.png?dl=0

Normal picture without Small Code:
https://www.dropbox.com/s/al8rhyc3cadvsn5/Normal_Image.png?dl=0

neurons 32 + 32,
Prefer Scaler: 37ms
Prefer Vector: 37ms
Prefer Scaler & Small Code: 60ms (picture error)
Prefer Vector & Small Code: 59.4ms (picture error)

neurons 32 + 16,
Prefer Scaler: 33.2 ms
Prefer Vector: 33 ms
Prefer Scaler & Small Code: 33.5ms (picture error)
Prefer Vector & Small Code: 34.2ms (picture error)

Edit: Using traditional method, avoid branch is the most efficient for my GPU,
neurons 64 + 64: 33.2 ms
neurons 32 + 32: 23.2 ms
neurons 32 + 16: 19.1 ms

Zachs
19th August 2015, 08:16
Looks like it's not just Fermi hating this new path. The whole NVIDIA product line hates it. Try it with your HD4600, you'll find the results to be the complete opposite!

huhn
19th August 2015, 08:41
my r9 270 doesn't really care maybe 0.5 ms difference. prefer scalar was the fastest mode with and without alternative weight access method.

aufkrawall
19th August 2015, 08:58
Would it make any difference to use DirectCompute instead of SM 5.0?

Zachs
19th August 2015, 10:19
It really comes down to how well the driver maps the code to the underlying resource. Theoretically they should all perform equally but as we've seen it's very erratic - the same code using different version of Microsoft shader compiler would make a lot of difference.

aufkrawall
19th August 2015, 11:32
my r9 270 doesn't really care maybe 0.5 ms difference. prefer scalar was the fastest mode with and without alternative weight access method.
Is there a huge difference for you in GPU load, clocks and render times with prefer scalar over OpenCL?

ryrynz
19th August 2015, 14:45
As you know my 750 Ti loves avoiding branches so this alternative method is of no use to me.

I will mention however that each option was slower with the alternative weights and the small code options
gave me some lovely artifacting I could probably printscreen and sell as art.

But anyway, OpenCL is still faster for me by 6ms or so so I'll just stick with that for now until I get a 960 or something.
FWIW it's quite nice getting 256/16 neurons at only 66% GPU load & 27 ms render on a low end graphics card like the 750 Ti
on 704x480 content. Some mighty bang for my buck right there, madVR has to sit on 64 neurons for the same level of performance.

madshi
19th August 2015, 15:23
FWIW it's quite nice getting 256/16 neurons at only 66% GPU load & 27 ms render on a low end graphics card like the 750 Ti
on 704x480 content. Some mighty bang for my buck right there, madVR has to sit on 64 neurons for the same level of performance.
We're talking X/Y here, right? Does 256/16 look better than 64/64? In my experience 16 neurons doesn't handle certain edge angles well. So I'm not sure if it's a good idea to use 256/16 over 64/64. Have you compared this with test images etc?

ryrynz
20th August 2015, 01:55
We're talking X/Y here, right? Does 256/16 look better than 64/64? In my experience 16 neurons doesn't handle certain edge angles well. So I'm not sure if it's a good idea to use 256/16 over 64/64. Have you compared this with test images etc?

I've done a quick comparison but not on test images, but my results echo exactly what Nevcairiel said, the second neuron count has significantly less impact than the first.
In fact I would not be able to spot the difference in a moving image it's that small.

I tested this on an anime image where lines are obviously all over the place and any drop in neuron count could be seen to negatively affect lines for fingers etc.
I know exactly what to look for having done comparisons with 64 vs 256 neurons many times and I did not see any drawbacks setting to only 16 neurons for the second value.
As is often often said, test images don't reflect real world results and although I'm sure you could find some small drawbacks, the benefits would surely outweigh them.

I think your own testing here would also conclude that allowing this second value to be set independently would be a good idea.

Zachs
20th August 2015, 04:33
Hi everyone,

Could you let me know if this MPDN test build 3338 (http://mpdn.zachsaw.com/Test%20Builds/3338/) makes OpenCL work again with the new Nvidia drivers?

Cheers.

Anima123
20th August 2015, 05:11
Hi everyone,

Could you let me know if this MPDN test build 3338 (http://mpdn.zachsaw.com/Test%20Builds/3338/) makes OpenCL work again with the new Nvidia drivers?

Cheers.

Yes, this build's OpenCL NNEDI3 works with nVidia driver 353.62. I am using this version because the next version wasn't steady enough under windows 10.

Edit: And, OpenCL NNEDI3 is a little bit slower than it's fastest shader counterparts, which is 'Avoid Branches' for nVidia 880M.

Edit2: Without knowledge of which direction is the first processed for both implementations, but the OpenCL version seems shaper than the shader version with the same neurons 128 + 16 for both.

I guess the question which direction should be put first for the best quality has never been discussed before?

ryrynz
20th August 2015, 05:15
Hi everyone,

Could you let me know if this MPDN test build 3338 (http://mpdn.zachsaw.com/Test%20Builds/3338/) makes OpenCL work again with the new Nvidia drivers?


Yes, this build's OpenCL NNEDI3 works with nVidia driver 353.62. I am using this version because the next version wasn't steady enough under windows 10.

Also confirmed fixed on 355.60, thanks!

Didn't really notice any performance hit either ^.^

Anima123
20th August 2015, 07:10
Here's screen-shot of two version of NNEDI3, both with 128 neurons for first pass and 16 for second pass, picture is from 1024x768 -> 1920x1080:

NNEDI3 shader version:
https://www.dropbox.com/s/em95mv8c0euz7ue/NNEDI3_128n16.png?dl=0

NNEDI3 OpenCL version:
https://www.dropbox.com/s/i3m2z77z7v89ggy/OpenCL128n16.png?dl=0

Zachs
20th August 2015, 07:11
Also confirmed fixed on 355.60, thanks!

Didn't really notice any performance hit either ^.^

Well that's the quick hack version - the proper one will give you performance *gains* with this new interop path.

EDIT: It may very well prevent AMD GPUs from doing any copyback too but I have no way of testing it.

Zachs
20th August 2015, 07:14
I guess the question which direction should be put first for the best quality has never been discussed before?

Which direction should be scaled first really comes down to whether you're more sensitive to vertical or horizontal lines. For example, I tend to be a lot more sensitive to horizontal lines being jagged than vertical lines.

ryrynz
20th August 2015, 07:17
Well that's the quick hack version - the proper one will give you *gains*

Oh yeah. Need those gains man, bring it.

Which direction should be scaled first really comes down to whether you're more sensitive to vertical or horizontal lines. For example, I tend to be a lot more sensitive to horizontal lines being jagged than vertical lines.

Yeah I tested this today and the differences are minimal. I prefer the OpenCL version.

nevcairiel
20th August 2015, 07:20
I've done a quick comparison but not on test images, but my results echo exactly what Nevcairiel said, the second neuron count has significantly less impact than the first.
In fact I would not be able to spot the difference in a moving image it's that small.

I probably wouldn't use such extremes though, rather try 128/32 instead of 256/16, or something like that. As madshi said, 16 can have quite odd effects on some things.

ryrynz
20th August 2015, 07:23
I probably wouldn't use such extremes though, rather try 128/32 instead of 256/16, or something like that. As madshi said, 16 can have quite odd effects on some things.

Yeah, the upgrade to 64 on the second pass doesn't cost me much so I decided to settle on that, I'm not much a fan of 16 neuron NNEDI3 anyway (prefer 64 as a minimum TBH) I was just using it for testing.

Here's screen-shot of two version of NNEDI3, both with 128 neurons for first pass and 16 for second pass, picture is from 1024x768 -> 1920x1080:


You would not notice any difference between them at least according to the image you've screenshot.
I have yet to see any content that shows any reason to prefer one over the other, so don't worry about it.

Zachs
20th August 2015, 07:37
Oh yeah. Need those gains man, bring it.

The gain isn't much but I understand why Nvidia chose to drop the old crappy interop now...

ryrynz
20th August 2015, 07:40
The gain isn't much but I understand why Nvidia chose to drop the old crappy interop now...

Do enlighten us, because besides those Nvidia driver wizards you seem to be one of the few that knows why.

Zachs
20th August 2015, 08:03
Simply because it's a more efficient API.

aufkrawall
20th August 2015, 08:13
I think it really just depends on the content if OpenCL or SM 5.0 version looks better.
Maybe vertical lines are more often a problem and thus how it currently is implemented in OpenCL could be slightly better, but maybe not.

However, I don't think it's a good idea to do one pass with many and one pass with few neurons.

64 pass 1 & 2:
http://abload.de/thumb/64jur84.png (http://abload.de/image.php?img=64jur84.png)

256 + 16:
http://abload.de/thumb/25616muqse.png (http://abload.de/image.php?img=25616muqse.png)
(both SM 5.0)

ryrynz
20th August 2015, 08:42
However, I don't think it's a good idea to do one pass with many and one pass with few neurons.


Man, I'm gonna dream about that screenshot considering how many times I've seen it... but yeah you can see some really obvious differences on Hayley's
eyelid it's not as straight as on the 256/64 screenshot.

It's already been established that 16 neurons isn't very good in comparison to the other options and this highlights it pretty well.

aufkrawall
20th August 2015, 08:53
In this picture you can easily spot both advantages and disadvantages of SM 5.0 and OpenCL implementation (both 64 neurons for both passes quadrupling).
e.g. with OpenCL, there is more aliasing at the doorframe, while with SM 5.0, vertical angles like Francine's legs show some weird doubled line contoures. They can be alo noticed with vertical real life content like antennas.
I think 64 neurons are really the minimum for each pass. However, let's not forget how bad other algorithms can look with such difficult content. We really got used to NNEDI3 quality.

madshi
20th August 2015, 08:54
I could imagine a split like 128/32 working instead of 64/64, as nevcairiel suggested. Or instead of 64/64 some users might have just enough juice in their GPUs to go to 128/64. In any case, doesn't hurt to allow different settings for X/Y, I suppose.

aufkrawall
20th August 2015, 09:44
Simply because it's a more efficient API.
Why was the old one less efficient? A GTX 780 Ti scored 1600fps with your benchmark.

Zachs
20th August 2015, 12:20
Why was the old one less efficient? A GTX 780 Ti scored 1600fps with your benchmark.

Oh the FPS isn't much of an efficiency indicator for interop. It just means with that implementation, your GPU could theoretically run 1600fps max. At which point your memory bandwidth would've been the true indicator of the frame rate. So it isn't exactly a benchmark per se.

Zachs
20th August 2015, 12:46
The new interop path is not tested on AMD so I'd appreciate it if you could try it out and let me know if it works at all. [Options -> Video Renderer -> General -> Always Use New OpenCL Interop]

If it does, it would be good to find out if it still causes a copyback - but does anyone know how to empirically determine this?

Thanks guys.

aufkrawall
20th August 2015, 13:15
Why not write a new benchmark that uses the new interop?

Zachs
20th August 2015, 13:15
Haven't got the time to do it. Don't even know if it works on AMD yet...