View Full Version : madVR - high quality video renderer (GPU assisted)
iSunrise
31st January 2014, 17:00
I've already ripped out the NVidia GPU in my development PC again. You know, it totally refuses to even acknowledge the precense of my LCD monitor, when connecting via DVI, while my Intel and AMD GPUs have no problems with that. FWIW, the NVidia does like my projector, when using DVI. I have to use VGA connection to get an image with the NVidia 650 on my LCD monitor.
Thatīs very strange, I cannot think of one setup where I ever encountered this and I have built like a couple hundred PCs with Nvidia cards and various LCD panels, especially over DVI, which has never failed me. If your LCD explicitly needs a DVI-I or analogue connection, maybe thatīs where something weird is happening.
I can put the NVidia GPU back in, but doing so costs time and effort. So I'd need a 100% sure way to reproduce any specific problem.
Itīs reproducible every time, so itīs a 100% sure way. I can reproduce it every time if I do a complete system restart with the settings NNEDI3 chroma upscaling and deinterlacing.
Just follow my strict instructions step by step in my last post.
Thanks. So it only occurs with deint + Nnedi3 chroma upscaling at the same time, correct? Is that video mode deint or forced film mode?
It seems to happen only in film mode. Like I already posted in my step by step guide, it will trigger every time if you donīt force anything, thatīs why I wrote "donīt check anything else". I was about as strict as I could be to reproduce the problem for you.
This bug was a pain to figure out.
Xaurus
31st January 2014, 17:03
madshi, just want to say thank you for this new version. I was witnessing my GTX 660 kneeling VERY quickly with these new features, wow.
I ended up on 16 neurons with nnedi3 error defusion, using Jinc 3 for luma and chroma. Great stuff although it kills the GPU. :)
madshi
31st January 2014, 17:18
So this is the absolute best for downscaling NNEDI?
"The absolute best"? You're violating the forum rules... :devil:
Getting a 3D TV with 4:4:4 at all refresh rates and no wifi will be next to impossible to find IMO.
The LG 32LA6136 seems to fit that bill. Haven't decided yet, though.
So I would say if you care about good black and color reproduction, you should choose a VA Panel (all Samsung TV and almost all Sony TV but the only model that can interest you is the Sony KDL-32HX750 if you can still find one where you live) but if you don't care about that or care more about viewing angles, choose an IPS panel (LG, Philips except high end models).
Ok, thx.
Is NNEDI too demanding for my Radeon 6650M?
It plays most material fine with Jinc3/AR on both chroma and luma but with NNEDI and all other set to bilinear it's like watching a slideshow even with 16 neurons...
I don't know. Try with 24fps SD content and 16 neurons. If that doesn't work, there's probably no hope.
It's only during video mode deinterlace. If I force film mode it goes away, if I then force video mode it's back again. So it only seems to occur when deinterlacing in video mode while using Nnedi3 chroma upscaling.
Itīs reproducible every time, so itīs a 100% sure way. I can reproduce it every time if I do a complete system restart with the settings NNEDI3 chroma upscaling and deinterlacing.
Just follow my strict instructions step by step in my last post.
It seems to happen only in film mode. Like I already posted in my step by step guide, it will trigger every time if you donīt force anything, thatīs why I wrote "donīt check anything else". I was about as strict as I could be to reproduce the problem for you.
Ok, it seems I can reproduce this problem in video mode on my 9400 mainboard. Shouldn't be too hard to fix.
huhn
31st January 2014, 17:22
Thanks! Some more options, it seems. From the reviews it seems to me that Samsung has often problems with clouding, not sure if that limits the use as a PC monitor. Found less complaints about that in the Philips reviews. What is the latest judgement about IPS vs. VA? Will look at the LG and Sony models...
clouding is luck it's like dead pixel.
and philips doesn't build any panel, so you get a sony, samsung, lg, sharp or something else and i'm pretty sure clounding comes when the panel is created and some people say is comes from bad transporting.
iSunrise
31st January 2014, 17:25
Ok, it seems I can reproduce this problem in video mode on my 9400 mainboard. Shouldn't be too hard to fix.
Finally. Iīm curious, why doesnīt this happen on an AMD card?
and philips doesn't build any panel, so you get a sony, samsung, lg, sharp or something else and i'm pretty sure clounding comes when the panel is created and some people say is comes from bad transporting.
Philips is pure crap IMHO. It says something about their quality control if they sometimes use very cheap china/korea panels in their sets and the customer never knows. Itīs like playing a card game. I would advise strongly against it.
IMHO Samsung is at the very top along with Sony (maybe even better). LG may be hit and miss, but their quality control on the Apple products is also very bad. Compared to Samsung, itīs like a night and day difference.
madshi
31st January 2014, 17:33
Finally. Iīm curious, why doesnīt this happen on an AMD card?
With NVidia's OpenCL implemention the channel order depends on the bitdepth. With 8bit the channel order is different than with 16bit. So I have to change things around depending on which bitdepth the OpenCL input has. I don't know if that is the cause of the problem, but it's likely. No such problems with Intel/AMD. There the channel order is always the same.
iSunrise
31st January 2014, 17:46
With NVidia's OpenCL implemention the channel order depends on the bitdepth. With 8bit the channel order is different than with 16bit. So I have to change things around depending on which bitdepth the OpenCL input has. I don't know if that is the cause of the problem, but it's likely. No such problems with Intel/AMD. There the channel order is always the same.
I kinda feared (like I stated in my very first post about this) that this is another OpenCL "weirdness" on Nvidia. Changing the channel order depending on the output bitdepth sounds like a really good way to make things very difficult for a programmer, because he actually has to do the thinking (or rather, implement different behaviours for different bitdepths) that normally the OpenCL implementation should do for him. Thanks for the input.
Btw, still waiting for NVīs answer to the reproduced problem by Blaire with the drivers >327.23, who has told me that he expects an answer from them tonight or Monday. One thing is for certain, though, they seem to have changed quite a lot since 327.23, because the size of both of their nvopencl.dlls (packaged for system32 and SysWOW64) went up by 50%!
PS: Iīm going to try to reproduce the scripting bug I also mentioned a couple of pages ago, but Iīm not very hopeful, since I did so many things at once when testing, that I am not sure where madVR suddenly decided not to change profiles again, whatever files I fed it after a certain breaking point. There must be a check somewhere, that ultimatey executes the profile change that didnīt seem to work anymore. It was probably not even related to my first script itself.
trip_let
31st January 2014, 17:48
I have i5-2410M + 540M. It works fine (I have to force mpc-hc to use nVidia GPU by renaming it), however 540M barely can hadle luma doubling on SD 24 fps video with 16 neurouns.
I'm on some ancient drivers (310), though.
Thanks for that.
Oddly enough, when I try some ancient drivers (306) and force Nvidia, all I get is a black screen. Even weirder, EVR's broken too and crashes out when I try to play anything using Nvidia on that old driver. Uh, this stuff used to work with old drivers so I have no idea what's going on. I think I'll blame... uh, Windows 8.1? (so I don't have to blame myself, of course :P) I'll look into it more when I have time.
Oh yeah, I did also try clean installations and some other driver versions, but no luck yet.
leeperry
31st January 2014, 18:10
The LG 32LA6136 seems to fit that bill.
That's a bunch of money for a <1K:1 CR IPS panel, OTOH viewing angles are wide and LG TV's usually come with a very impressive ISF certified colorimetry menu :)
omarank
31st January 2014, 18:24
Same here as turbojet. Playing a 60fps file with your profile setup activates the "On" profile, but Smooth Motion FRC according to the debug OSD is turned off.
Does Smooth Motion FRC disable on your PC if you remove the profiles and just switch it to "only if there would be judder..."?
Ok, I have removed the profile and switched to "only if there would be judder..." Now SM gets enabled for 29.97 fps interlaced videos (video mode) and 30 fps progressive videos (from Kodak digicam). It remains disabled however for 29.97 fps progressive videos. IIRC, SM switching was working fine in v0.86.11.
@Madshi, Turbojet: Can you check for 29.97 fps interlaced videos?
huhn
31st January 2014, 18:24
That's a bunch of money for a <1K:1 CR IPS panel, OTOH viewing angles are wide and LG TV's usually come with a very impressive ISF certified colorimetry menu :)
in germany 3d displays with >32< zoll, start at 360
you got the philips 450x aktive rest is over 400 euro and samsung no all refresh rate 4:4:4. they most likely share the same panel.
and passive only lg. cheap or no extra glasses needed.
that's a lot of money for a 1080p display and a 3d only for dev.
you are very fast over 500 euro
XMonarchY
31st January 2014, 18:34
I started a thread about broken OpenCL on GeForce forums and one person replied that its madVR that is the issue... - https://forums.geforce.com/default/topic/680591/opencl-has-been-broken-since-327-23-drivers-/?offset=3 but as stated earlier - it could be the D3D9 that causes the issue with OpenCL. Is there a way to tell though? Do other OpenCL applications have issues with new drivers?
BTW, I really like the whole "Random Question" thing of these forums, but does it do it for each and every single reply/post?
cyberbeing
31st January 2014, 18:54
Simple test: Use MS Paint, light gray background. Some diagonal black lines on top. Upscale this with 16 neurons. Ouch.
Not specific enough. Created a 1280x720 light gray image, drew some random black diagonal lines ranging from 1px-5px width on it, but was unable to reproduce.
Could you post a test image?
But aside from artificial test patterns, you said you actually saw this on real videos, right?
Either way, I'm interested in what an artifact with madVR OpenCL NNEDI3 16 neurons looks like, how noticeable it was within the original scene, and how rare the occurrence is.
iSunrise
31st January 2014, 19:21
I started a thread about broken OpenCL on GeForce forums and one person replied that its madVR that is the issue... - https://forums.geforce.com/default/topic/680591/opencl-has-been-broken-since-327-23-drivers-/?offset=3 but as stated earlier - it could be the D3D9 that causes the issue with OpenCL. Is there a way to tell though? Do other OpenCL applications have issues with new drivers?
BTW, I really like the whole "Random Question" thing of these forums, but does it do it for each and every single reply/post?
I can certainly understand that you absolutely want to fix what is broken at the moment as fast as possible. Youīre as frustrated as others by this, but we have to follow a certain way to get this fixed.
And as such, it doesnīt help if you post guesses or false information, as it doesnīt make our lives easier. How many times do we have to tell you that (a) itīs an NV opencl issue, because it doesnīt happen with older builds at all (b) it works with AMD and (c) it doesnīt help anyone if you suddenly post a workaround (that doesnīt even work) several times on different forums, so that NV becomes suspicious.
Whatīs even worse is that this one user now reports that madVR is the issue, who has no idea whatsoever, because he doesnīt know madshiīs code. Posts like that should be completely ignored or deleted.
So please, the issue itself has been reported to NV and theyīre trying to reproduce it now (which is easy). Just give it some time, will you.
Even if you would find a workaround, NV would need to fix it in their drivers, anyway. So why donīt you actually wait a bit, like we all do.
XMonarchY
31st January 2014, 20:06
I can certainly understand that you absolutely want to fix what is broken at the moment as fast as possible. Youīre as frustrated as others by this, but we have to follow a certain way to get this fixed.
And as such, it doesnīt help if you post guesses or false information, as it doesnīt make our lives easier. How many times do we have to tell you that (a) itīs an NV opencl issue, because it doesnīt happen with older builds at all (b) it works with AMD and (c) it doesnīt help anyone if you suddenly post a workaround (that doesnīt even work) several times on different forums, so that NV becomes suspicious.
Whatīs even worse is that this one user now reports that madVR is the issue, who has no idea whatsoever, because he doesnīt know madshiīs code. Posts like that should be completely ignored or deleted.
So please, the issue itself has been reported to NV and theyīre trying to reproduce it now (which is easy). Just give it some time, will you.
Even if you would find a workaround, NV would need to fix it in their drivers, anyway. So why donīt you actually wait a bit, like we all do.
I've only provided information that was provided to me. I never said it was accurate. You say they don't know madVR code and they say you don't know nV code. I'm considering all possibilities and seeking information in an attempt to create a workaround for myself and everyone else who wants to use their GPUs for madVR OpenCL along with the latest fixes and improvements in video games.
Thanks for a warm reply - I'll make sure not to post a fix if I find one. Its better to do it the right way - wait 6 more driver releases before nVidia fixes this issue and creates another one.
pirlouy
31st January 2014, 20:30
No need for quoting all post when just posting after...
The most regrettable thing is to give credits to a random guy who obviously is not a developer.
and one person (https://forums.geforce.com/default/topic/680591/geforce-drivers/opencl-has-been-broken-since-327-23-drivers-/post/4106647/#4106647) replied that its madVR that is the issue...
iSunrise
31st January 2014, 20:31
You say they don't know madVR code and they say you don't know nV code.
I meant the one guy that posted a reply to your post. NV knows about all of this already (we gave them the details that were provided by madshi himself) and they will come back to us (people that are in the Beta program, actually) if they have reproduced it themselves. They donīt need the code from madshi anyway, because they can track the error without any problem. And even if they need more info, we will know soon. That one guy that posted the replay though, can do neither of that. So itīs better to ignore him.
Thanks for a warm reply - I'll make sure not to post a fix if I find one. Its better to do it the right way - wait 6 more driver releases before nVidia fixes this issue and creates another one.
Unfortunately, thatīs exactly what needs to happen. Wouldnīt matter if I post it any other way. I just want to make sure that an official fix doesnīt get delayed any further. And that benefits us all.
5ts
31st January 2014, 20:42
Thanks Madshi for your excellent work. Is it possible to have a direct source mode feature added to madvr for those with external video processors that would disable all video processing features such as chroma upsampling, scaling and color conversion etc in madvr to effectively send the video signal untouched to the external video processor to do the heavy lifting. This would be an awesome feature if it could be incorporated into this fantastic renderer.
DarkSpace
31st January 2014, 20:57
Thanks Madshi for your excellent work. Is it possible to have a direct source mode feature added to madvr for those with external video processors that would disable all video processing features such as chroma upsampling, scaling and color conversion etc in madvr to effectively send the video signal untouched to the external video processor to do the heavy lifting. This would be an awesome feature if it could be incorporated into this fantastic renderer.
Uhm, what?
So, you want a video renderer (madVR in this case) that forwards all the video renderer work to something else (your "external video processor") so that this something else can do the work the video renderer is supposed to do?
A possibly quite stupid question, but why?
Can't you just use the "external video processor" as video renderer instead, then?
cyberbeing
31st January 2014, 21:27
Would it be possible to have madVR's traditional scaling algorithms correct any NNEDI3 pixel shift when they perform additional upscaling/downscaling afterwords?
madshi
31st January 2014, 21:34
Ok, I have removed the profile and switched to "only if there would be judder..." Now SM gets enabled for 29.97 fps interlaced videos (video mode) and 30 fps progressive videos (from Kodak digicam). It remains disabled however for 29.97 fps progressive videos. IIRC, SM switching was working fine in v0.86.11.
@Madshi, Turbojet: Can you check for 29.97 fps interlaced videos?
I think the issue you're seeing has nothing to do with the new profiling functionality, and I would guess it also occurs with v0.86.11. I think your movie framerate and your display refresh rate are simply so far away that madVR believes SM is needed. Look at the OSD (Ctrl+J) and check if refresh rate and movie framerate are really (more or less) identical. They're probably not. If you still think this is a new bug caused by profiles, then please downdate to v0.86.11 and double check that the problem *really* didn't occur with v0.86.11 because right now I think you'd get the same with v0.86.11.
I started a thread about broken OpenCL on GeForce forums and one person replied that its madVR that is the issue...
Probably a madVR hater posting that without having a clue. Just ignore him.
Could you post a test image?
But aside from artificial test patterns, you said you actually saw this on real videos, right?
When talking about image scaling I like to test with specific images which are known to showcase strengths and weaknesses of scaling algorithms. E.g. try this image:
http://madshi.net/clownOrg.png
Look at the white cars. With 16 neurons there's still some aliasing left in there. The improvement from 16 to 32 neurons is larger there than when going from 32 to 64 neurons. But as I said earlier, it depends. Sometimes the step from X to Y neurons looks bigger than the step from Y to Z neurons. And sometimes it's the other way round. It differs depending on the test image. But I've seen several instances like this "clown" image where 16 neurons still left some aliasing in the image which 32 neurons mostly took care of.
But in the end all I can give is my thoughts. And in terms of image quality analysis I don't claim to be the final judge. You're free to disagree with me and post different recommendations.
Thanks Madshi for your excellent work. Is it possible to have a direct source mode feature added to madvr for those with external video processors that would disable all video processing features such as chroma upsampling, scaling and color conversion etc in madvr to effectively send the video signal untouched to the external video processor to do the heavy lifting. This would be an awesome feature if it could be incorporated into this fantastic renderer.
That's technically not possible for 2 reasons:
(1) The GPUs in HTPCs are used to output RGB. You can switch them to YCbCr, but all that results in is that they take the Windows RGB desktop and convert it to YCbCr in the GPU drivers and then output that via HDMI. Which makes no sense since that means *more* processing than outputting RGB directly. It is very hard (with some GPUs and OSs even impossible) for madVR to output YCbCr data in such a way that the GPU and OS don't further manipulate it. So you have to live with RGB output. And once you accept that RGB output is the best option for HTPCs, you automatically have to live with all that's necessary for good quality RGB output, which is at least chroma upscaling, color conversion and then conversion to the final RGB output bitdepth.
(2) Even if it were possible for madVR to tell the GPU/OS to output some specific YCbCr data untouched, there's still the problem that HDMI up until 1.4 does not support 4:2:0 transport. Which means that every source device (including external Blu-Ray players etc) has to take the decoded video and at least partially upscale chroma to 4:2:2. So untouched output is technically possible. Ok, HDMI 2.0 finally introduced 4:2:0 support, so there's that. But no GPUs I know actually have HDMI 2.0 ports yet. And even if they had, we're back at problem (1).
But I think you really don't need to worry: I believe that madVR has better chroma upscaling, color conversion and scaling algorithms than any external video processor out there today. E.g. users with Lumagen processors have personally told me that they prefer madVR's Jinc3 AR over Lumagen's scaling. The one thing external video processors might do better today is that they might offer more reliable/easy to use deinterlacing. That is an area I find lacking myself in HTPCs, currently.
5ts
31st January 2014, 21:39
Uhm, what?
So, you want a video renderer (madVR in this case) that forwards all the video renderer work to something else (your "external video processor") so that this something else can do the work the video renderer is supposed to do?
A possibly quite stupid question, but why?
Can't you just use the "external video processor" as video renderer instead, then?
Let me clarify further. By external video processor I am referring to video processors from the likes of Lumagen and DVDO which is a hardware device. I am seeking a feature to emulate a configuration where blu ray players from the likes of Oppo for e.g. offer a direct source mode that bypass the blu ray player's internal video processing capabilities and send these to the external video processor (Lumagen, DVDO) for processing which is superior. All I am requesting is for madvr to emulate such a feature for those who have the benefit of external video processors. It is always wise to seek clarification first instead of passing judgement on the intelligence of one's question.
madshi
31st January 2014, 21:41
Would it be possible to have madVR's traditional scaling algorithms correct any NNEDI3 pixel shift when they perform additional upscaling/downscaling afterwords?
Hmmmm... That should be technically possible. I'm already doing exactly that to match the separately upscaled chroma channels to the NNEDI3 scaled luma channel. Sounds like a good idea to me, although it would further complicate my code (depending on scaling factor I'd have to sometimes correct the position of the chroma channels and sometimes the position of the luma channels, and sometimes both and sometimes none, ouch). Please add this as a "bug" to the tracker. I don't want to spend time on that atm, but I consider it a good idea, so I don't want to lose sight of that.
DarkSpace
31st January 2014, 22:24
It is always wise to seek clarification first instead of passing judgement on the intelligence of one's question.
Which is precisely why I wrote down what I understood your feature request was, and then offered an alternative to what I understood your request was, all without hostility. That was intended to give you an alternative to something I think makes little sense should my understanding be correct, and should my understanding be incorrect, it was intended to give you an opportunity to correct specifically the point where I misunderstood (here: the "external" part of the video processor).
I obviously did simply misunderstand your request, but even so, I am sorry that you perceived it as hostile. Let me assure you, there was no hostility on my end.
Finally, I did not question your intelligence. You may notice that I specifically did not include anything regarding that except for the simple question of why you would want that (saying that I don't see the reason, and not saying that there is none). In this case, the inability to see the reason behind your request originated from my misunderstanding of the request itself.
Mano
31st January 2014, 22:38
is it normal for mpc to use 1.5gb ram then another 500mb in exclusive mode playing a 322mb mkv anime?
ryrynz
31st January 2014, 23:13
Only if you have Avisynth or something similar running.
sandman7920
31st January 2014, 23:21
Which decode method produce better quality for image up scaling DXVA or Software?
Sent from my HTC One using Tapatalk
DarkSpace
31st January 2014, 23:32
By external video processor I am referring to video processors from the likes of Lumagen and DVDO which is a hardware device. I am seeking a feature to emulate a configuration where blu ray players from the likes of Oppo for e.g. offer a direct source mode that bypass the blu ray player's internal video processing capabilities and send these to the external video processor (Lumagen, DVDO) for processing which is superior.
I have done a bit of thinking, and have reached the conclusion that what you intend may indeed be possible, depending on how those hardware devices work (and how configurable they are):
If you create a file with the filename "YCbCr" (no extension, and the file itself may be empty) in the directory madVR resides in, madVR will not convert the data to RGB before presenting it. The image will be YUV, but still be treated (and flagged) as RGB by the GPU. Therefore, you will output YUV, but the GPU will signal that it's outputting RGB. If you can configure your device to treat the input as YUV regardless of what the signal thinks it is, you can let it perform the conversion to RGB instead.
In any case, however, madVR will upsample the Chroma channels. If you select Nearest Neighbor as Chroma upsampling algorithm, the pixels will be literally doubled for 4:2:0 content. If you can tell your device to only take one pixel of every 2x2 block and upsample the Chroma from that, you can also forward that task to the device.
Make sure you output the video at its original resolution, of course. Otherwise, make sure your device knows the video's original resolution and instruct it to invert Nearest Neighbor scaling from the device's input resolution to the original video resolution before doing the Chroma upsampling step. Of course, also make sure you have Nearest Neighbor selected as Image Upscaling algorithm. This may or may not work, but in my opinion (does anyone know something better?) it is your best chance of getting stuff right.
I am not sure about deinterlacing. If you can manually tell your device when to treat the input as interlaced, set your GPU to deinterlace using Weave (this will result in not deinterlacing at all, and instead treating the image as progressive), and set madVR to video mode deinterlacing. Also check the halfrate deinterlacing option in madVR's Trade quality for performance section: If you weave the image back together, doublerate deinterlacing will probably give you the same image twice in a row (I never tried this myself).
Disable all settings in the Artifacts removal tab of Processing.
You probably also want to disable SmoothMotion framerate conversion.
Since I will be assuming that your input is DVD or BluRay content, or otherwise native 8-bit content, disable dithering completely: You only use Nearest Neighbor scaling, which does not increase the image bitdepth, so you can safely round down or truncate back to 8 bit. You can find the setting in the Trade quality for performance section.
I hope I didn't forget anything!
cyberbeing
31st January 2014, 23:50
But I've seen several instances like this "clown" image where 16 neurons still left some aliasing in the image which 32 neurons mostly took care of.
So the artifacts you speak of are only 'aliasing' or can it produce other strange artifacts as well?
The clown image change from 16 neurons to 32 still seems very minor. And it actually appears that 32 has more artifacts than 16 does in that image when compared against 64. 64 seems like a sharper version of 16 with similar sub-pixel structure, while 32 seems to have a very *different* sub-pixel structure with bad guesses which are reverted by 64 & 128. 256 on the other hand seems too strong to the point it starts enhancing minor source artifacts in a couple places.
The reliable improvement sweet spot for NNEDI3 appears to be 64-128 neurons, with 16 neurons as nice low-cost speed option. I'll continue to do more testing, but 32 neurons seems rather inconsistent compared to the others. It almost makes me think there is a math error or something with 32 neurons, but maybe this is just how NNEDI3 behaves normally...
Hmmmm... That should be technically possible. I'm already doing exactly that to match the separately upscaled chroma channels to the NNEDI3 scaled luma channel. Sounds like a good idea to me, although it would further complicate my code (depending on scaling factor I'd have to sometimes correct the position of the chroma channels and sometimes the position of the luma channels, and sometimes both and sometimes none, ouch). Please add this as a "bug" to the tracker. I don't want to spend time on that atm, but I consider it a good idea, so I don't want to lose sight of that.
I think it would be worthwhile, especially when you consider the sub-pixel positioned subtitle typesetting madVR receives. Screenshot comparisons against the quality of traditional resamplers would also be more practical.
madshi
1st February 2014, 00:30
So the artifacts you speak of are only 'aliasing' or can it produce other strange artifacts as well?
The clown image change from 16 neurons to 32 still seems very minor. And it actually appears that 32 has more artifacts than 16 does in that image when compared against 64. 64 seems like a sharper version of 16 with similar sub-pixel structure, while 32 seems to have a very *different* sub-pixel structure with bad guesses which are reverted by 64 & 128. 256 on the other hand seems too strong to the point it starts enhancing minor source artifacts in a couple places.
The reliable improvement sweet spot for NNEDI3 appears to be 64-128 neurons, with 16 neurons as nice low-cost speed option. I'll continue to do more testing, but 32 neurons seems rather inconsistent compared to the others. It almost makes me think there is a math error or something with 32 neurons, but maybe this is just how NNEDI3 behaves normally...
I don't really have the time to micro-analyze this with large numbers of test images and videos atm. With the clown image I clearly prefer 32 neurons over 16 neurons. You seem to disagree, that's fine. I can only post my personal subjective impressions. And my impressions are that more neurons usually helps, and I've seen a reduction in aliasing when going from 16 to 32 neurons with one or the other test images. That's why I personally suggest to use at least 32 neurons (or more). But what I suggest is not the law, and I'm fine if people have different opinions.
I'll happily admit that different neuron settings sometimes give a little bit different results, and sometimes one neuron setting can look better than another, and it's not *always* the higher neuron setting that wins. But usually it is. And if you compare 16 neurons to 256 neurons, the difference can sometimes be quite noticeable (in favor of 256 neurons), while comparing "neighbor" neuron settings often only shows small improvements.
I think that's the last I want to say about neurons for now.
madshi
1st February 2014, 00:53
P.S: One last post:
original image (http://madshi.net/castleOrg.png) --|-- 16 neurons (http://madshi.net/castle16neurons.png) --|-- 32 neurons (http://madshi.net/castle32neurons.png)
I believe 32 neurons look *much* better and cleaner with this test image than 16 neurons. This is the kind of artifacts I had in mind, when only using 16 neurons.
cyberbeing
1st February 2014, 01:43
I don't really have the time to micro-analyze this with large numbers of test images and videos atm.
Well if you or anyone else have any further insight or opinions into the quality scaling of NNEDI3 with increasing neurons, I'd be interested to hear it. I've never really used NNEDI3 for scaling before madVR, so at this point I'm still interested in seeing any samples where NNEDI3 may prove disadvantageous.
With the clown image I clearly prefer 32 neurons over 16 neurons. You seem to disagree, that's fine.
It's not that I necessarily disagree, since 32 neurons does reduce aliasing on some small edges compared to 16 neurons on the clown image. It just doesn't seem to reduce global blurriness and refine large edges like 64 and higher do. This just makes the performance:quality ratio of going from 16->32 neurons seem a bit poor. It's one of those things where if I had make quality compromises with other madVR settings to get to 32 neurons, I'm just not yet sure that would be worth it.
How madVR settings interact with each other in terms of quality became much more complex once NNEDI3 and error-diffusion were added to the mix. The interaction of upscaling and downscaling algorithms with NNEDI3 luma-only seems to need the most re-testing.
P.S: One last post:
original image (http://madshi.net/castleOrg.png) --|-- 16 neurons (http://madshi.net/castle16neurons.png) --|-- 32 neurons (http://madshi.net/castle32neurons.png)
I believe 32 neurons look *much* better and cleaner with this test image than 16 neurons. This is the kind of artifacts I had in mind, when only using 16 neurons.
Thanks, this is a much more obvious example of how 32 neurons can do better than 16 neurons when quadrupling.
So it seems possible that 16 neurons has issues reconstructing sub-pixel edges of 1px->2px sized detail with high contrast thresholds.
Which dithering setting in madVR did you use for those images?
_____
Side question:
With NNEDI3 Luma-only, the image is YCbCr with only the Y channel being doubled.
"Image upscaling" then resizes the CbCr chroma channels only, and "Image downscaling" then resizes the Y luma channels only to destination size.
Do both the "anti-ringing" and "linear light" settings in madVR remain functional as normal?
6233638
1st February 2014, 03:54
When talking about image scaling I like to test with specific images which are known to showcase strengths and weaknesses of scaling algorithms. E.g. try this image:
http://madshi.net/clownOrg.png
Look at the white cars. With 16 neurons there's still some aliasing left in there. The improvement from 16 to 32 neurons is larger there than when going from 32 to 64 neurons. But as I said earlier, it depends. Sometimes the step from X to Y neurons looks bigger than the step from Y to Z neurons. And sometimes it's the other way round. It differs depending on the test image. But I've seen several instances like this "clown" image where 16 neurons still left some aliasing in the image which 32 neurons mostly took care of.
But in the end all I can give is my thoughts. And in terms of image quality analysis I don't claim to be the final judge. You're free to disagree with me and post different recommendations.It's true that with some images - this one included - 16 neurons can actually show more aliasing in places than if you simply used Jinc 3 AR at a lower cost. However, in most of my real-world testing, the benefit of NNEDI3 image doubling even with 16 neurons outweighs the potential downsides.
I need to do more testing before I start posting comparisons and recommendations, but from a very small amount of testing with some of the old 360p game clips I used when comparing scaling algorithms, the improvements NNEDI3 brings are absolutely ridiculous.
It also seems that, at least in the material I have been testing so far, NNEDI3 chroma doubling is almost worthless, even if you ignore the performance hit it introduces.
NNEDI3 chroma scaling is very nice when used at 32 neurons, which seems to be the sweet spot for image quality vs performance. The image quality still looks better as you increase the neuron count, but it doesn't seem to be worth the performance hit above 32. (and 16 seems fine if it's all your hardware can handle - still better than Jinc3AR)
Similar to NNEDI3 chroma doubling, luma quadrupling does not seem worth the cost at all, unless you are using 256 neurons for doubling and still have power to spare. In all the testing I've done so far, you're better off increasing the neuron count than a balance between doubling and quadrupling. (e.g. 256 neurons for doubling instead of 64 double + 32 quad)
This is when outputting at 1080p of course - perhaps the results would be different at 4K.
I will say though, while Jinc 3 AR + NNEDI3 doubling with 64 neurons looks very good, I think I still want to use SoftCubic with some lower quality DVDs - but I should investigate the profile switching capabilities further, as it seems like there is probably a way to toggle between the two with a single key.
But I think you really don't need to worry: I believe that madVR has better chroma upscaling, color conversion and scaling algorithms than any external video processor out there today.I don't think there's any question of that now with the NNEDI3 image doubling.
The one thing external video processors might do better today is that they might offer more reliable/easy to use deinterlacing. That is an area I find lacking myself in HTPCs, currently.That's possibly true with video content. I still had cadence detection issues with PAL content when I had a Lumagen Radiance (the original one) but they may have improved that since. The hardware is basically just an FPGA as I understand it, so they're able to make very big changes via updates.
Stereodude
1st February 2014, 05:23
I tried playing a few 1080p24 standard x264 blu-ray movies presented at 23.976 Hz.
Average rendering time (using Jinc 3 AR for both Luma and Chroma, Debanding low, all trade quality for performance options except OpenCL error diffusion disabled) was around 8-10 ms. The rendering time is about the same for 720p movies, so scaling for instance is rather cheap even with Jinc 3 AR.
With OpenCL error diffusion the average rendering time was raised to around 28-30 ms.
Should it really be that demanding?My Radeon HD 7790 playing back a 24000/1000fps blu-ray with the same settings you're running is around 11.5ms. Enabling smooth motion pushes that to ~13.3ms.
Turning off smooth motion and enabling switching to OpenCL error diffusion that number climbs to 26-27ms. Re-enabling smooth motion pushes that to ~44.3ms. FWIW, my "monitor" is a CRT HDTV runs that a custom resolution of 1806x1016. I'm using Windows 7 32-bit and the 13.12 Catalyst drivers.
Basically, I can only use OpenCL error diffusion for content that's 30fps or less assuming I don't want to use smooth motion.
TheProfosist
1st February 2014, 05:41
Well if you or anyone else have any further insight or opinions into the quality scaling of NNEDI3 with increasing neurons, I'd be interested to hear it. I've never really used NNEDI3 for scaling before madVR, so at this point I'm still interested in seeing any samples where NNEDI3 may prove disadvantageous.
My suggestion is install AviSynth and install NNEDI3 and the stuff it needs and start trying it as a resizer on your files.
James Freeman
1st February 2014, 08:19
@madshi
Why ""nnedi3 - OpenCL rewrite" (http://forum.doom9.org/showthread.php?t=169766) woks perfectly with Nvidia whether MadVR does not?
"nnedi3 - OpenCL rewrite" uses OpenCL and it works perfectly over here.
Isn't your code based on this version of NNEDI3?
QBhd
1st February 2014, 09:01
@madshi
Why ""nnedi3 - OpenCL rewrite" (http://forum.doom9.org/showthread.php?t=169766) woks perfectly with Nvidia whether MadVR does not?
"nnedi3 - OpenCL rewrite" uses OpenCL and it works perfectly over here.
Isn't your code based on this version of NNEDI3?
I am sure this quote from the first page has something to do with your Q:
"OpenCL device preferences.
Don't bother with running the code on CPU OpenCL devices original nnedi3 would be way faster simply due to prescreener.
For GPU AMD cards with GCN architecture are recommended. Nvidia does ok, but has disadvantage of completely using one of your CPU cores on heavy GPU computations. Intel integrated... it works there too!
Theoretical FLOPS should be good indication of performance as long as you factor in the efficiency of particular architecture. Table below provides some useful coefficients how TFLOPS scale to FPS for cards of different architectures.
In case of multiple OpenCL platforms the order of preference: AMD GPU -> any GPU -> the rest. No manual choice yet."
QB
James Freeman
1st February 2014, 09:21
I am sure this quote from the first page has something to do with your Q:
"OpenCL device preferences.
Don't bother with running the code on CPU OpenCL devices – original nnedi3 would be way faster simply due to prescreener.
The original nnedi3 is WAY slower than this version.
Nvidia is a GPU OpenCL device as shown by GPU-Z when running this NNEDI3 version.
This version runs on the GPU + one of the CPU cores.
Nvidia does OK, but has disadvantage of completely using one of your CPU cores on heavy GPU computations.
QB
This strengthen the fact that it works with Nvidia, but the version in madVR does not.
Maybe madshi did some extreme optimizations (which he obviously did) that improved some things and broke other?
omarank
1st February 2014, 09:22
I think the issue you're seeing has nothing to do with the new profiling functionality, and I would guess it also occurs with v0.86.11. I think your movie framerate and your display refresh rate are simply so far away that madVR believes SM is needed. Look at the OSD (Ctrl+J) and check if refresh rate and movie framerate are really (more or less) identical. They're probably not. If you still think this is a new bug caused by profiles, then please downdate to v0.86.11 and double check that the problem *really* didn't occur with v0.86.11 because right now I think you'd get the same with v0.86.11.
In the OSD, the refresh rate reported is ~60.0178 Hz. I just tried v0.86.11 and found that SM switching is working perfectly with that version. I have been using smooth motion since the very first release; had there been any such issue earlier, I would have reported it.
5ts
1st February 2014, 09:56
I have done a bit of thinking, and have reached the conclusion that what you intend may indeed be possible, depending on how those hardware devices work (and how configurable they are):
If you create a file with the filename "YCbCr" (no extension, and the file itself may be empty) in the directory madVR resides in, madVR will not convert the data to YUV before presenting it. The image will be YUV, but still be treated (and flagged) as RGB by the GPU. Therefore, you will output YUV, but the GPU will signal that it's outputting RGB. If you can configure your device to treat the input as YUV regardless of what the signal thinks it is, you can let it perform the conversion to RGB instead.
In any case, however, madVR will upsample the Chroma channels. If you select Nearest Neighbor as Chroma upsampling algorithm, the pixels will be literally doubled for 4:2:0 content. If you can tell your device to only take one pixel of every 2x2 block and upsample the Chroma from that, you can also forward that task to the device.
Make sure you output the video at its original resolution, of course. Otherwise, make sure your device knows the video's original resolution and instruct it to invert Nearest Neighbor scaling from the device's input resolution to the original video resolution before doing the Chroma upsampling step. Of course, also make sure you have Nearest Neighbor selected as Image Upscaling algorithm. This may or may not work, but in my opinion (does anyone know something better?) it is your best chance of getting stuff right.
I am not sure about deinterlacing. If you can manually tell your device when to treat the input as interlaced, set your GPU to deinterlace using Weave (this will result in not deinterlacing at all, and instead treating the image as progressive), and set madVR to video mode deinterlacing. Also check the halfrate deinterlacing option in madVR's Trade quality for performance section: If you weave the image back together, doublerate deinterlacing will probably give you the same image twice in a row (I never tried this myself).
Disable all settings in the Artifacts removal tab of Processing.
You probably also want to disable SmoothMotion framerate conversion.
Since I will be assuming that your input is DVD or BluRay content, or otherwise native 8-bit content, disable dithering completely: You only use Nearest Neighbor scaling, which does not increase the image bitdepth, so you can safely round down or truncate back to 8 bit. You can find the setting in the Trade quality for performance section.
I hope I didn't forget anything!
Thanks for the advice and appreciate the time taken to provide a detailed response to my request. I will implement your suggestions and confirm if it achieves the desired result.
the_weirdo
1st February 2014, 10:06
@madshi
Why ""nnedi3 - OpenCL rewrite" (http://forum.doom9.org/showthread.php?t=169766) woks perfectly with Nvidia whether MadVR does not?
"nnedi3 - OpenCL rewrite" uses OpenCL and it works perfectly over here.
Isn't your code based on this version of NNEDI3?
As madshi has answered, this issue is not caused by the NNEDI3 OpenCL implementation in madVR, but by the D3D9 <-> OpenCL interop, which seems to be broken in recent drivers of Nvidia.
madshi
1st February 2014, 10:28
Which dithering setting in madVR did you use for those images?
Should be Error Diffusion, I think.
"Image upscaling" then resizes the CbCr chroma channels only, and "Image downscaling" then resizes the Y luma channels only to destination size.
Do both the "anti-ringing" and "linear light" settings in madVR remain functional as normal?
Yes.
from a very small amount of testing with some of the old 360p game clips I used when comparing scaling algorithms, the improvements NNEDI3 brings are absolutely ridiculous.
Yes, especially with game/computer material the differences are extreme. Linear resampling algorithms simply suck with such source material. Video content usually has much less contrasty lines and much less native aliasing in the original source, though, so with real life video the advantage of NNEDI3 is quite a bit smaller.
It also seems that, at least in the material I have been testing so far, NNEDI3 chroma doubling is almost worthless, even if you ignore the performance hit it introduces.
NNEDI3 chroma scaling is very nice when used at 32 neurons, which seems to be the sweet spot for image quality vs performance.
I think we need to find clear terms when talking about these things. When I first read your post I understood it exactly the wrong way. Maybe if we want to differentiate between NNEDI3 chroma upscaling and chroma channel doubling we should refer to one of these as "4:2:0 -> 4:4:4 chroma upscaling", and to the other as "chroma channel doubling"?
I do agree with your chroma quality assessment.
The image quality still looks better as you increase the neuron count, but it doesn't seem to be worth the performance hit above 32. (and 16 seems fine if it's all your hardware can handle - still better than Jinc3AR)
I have a similar impression - although with some images going to 64 and then to 128 neurons does still result in a nice quality improvement, so I find it hard to draw any definite conclusions. It's a valid question, though, which neuron setting to recommend, since every higher setting almost doubles the performance cost.
Similar to NNEDI3 chroma doubling, luma quadrupling does not seem worth the cost at all, unless you are using 256 neurons for doubling and still have power to spare. In all the testing I've done so far, you're better off increasing the neuron count than a balance between doubling and quadrupling. (e.g. 256 neurons for doubling instead of 64 double + 32 quad)
I've found that the output of luma doubling is soft enough that using 16 neurons for quadrupling doesn't seem to produce any problems. So I think it's fine to use 16 neurons for quadrupling, which makes it more feasible from a performance point of view. Whether that brings enough quality advantage to make it worthwhile is another question, but when using really small sized sources, or when upscaling to very high resolution monitors, it might make sense.
That's possibly true with video content. I still had cadence detection issues with PAL content when I had a Lumagen Radiance (the original one) but they may have improved that since. The hardware is basically just an FPGA as I understand it, so they're able to make very big changes via updates.
They used to do their own film mode detection back in the old days, before they had a Gennum VXP chip built it. Ever since they're using the Gennum VXP chip (now Sigma owned, IIRC), I believe they're relying on the Gennum to do deinterlacing and film mode detection and all that stuff. I don't think they use their own FPGA for that, anymore.
My Radeon HD 7790 playing back a 24000/1000fps blu-ray with the same settings you're running is around 11.5ms. Enabling smooth motion pushes that to ~13.3ms.
Turning off smooth motion and enabling switching to OpenCL error diffusion that number climbs to 26-27ms. Re-enabling smooth motion pushes that to ~44.3ms. FWIW, my "monitor" is a CRT HDTV runs that a custom resolution of 1806x1016. I'm using Windows 7 32-bit and the 13.12 Catalyst drivers.
Basically, I can only use OpenCL error diffusion for content that's 30fps or less assuming I don't want to use smooth motion.
That is a little bit slower than on my PC. Maybe Windows 8.1 helps here? Or maybe the AMD Windows 8.1 drivers are slightly better optimized for OpenCL interop? I don't know...
Why ""nnedi3 - OpenCL rewrite" (http://forum.doom9.org/showthread.php?t=169766) woks perfectly with Nvidia whether MadVR does not?
That one works because it does CPU -> OpenCL -> CPU. In contrast madVR does Direct3D9 -> OpenCL -> Direct3D9. Basically the AviSynth script doesn't use D3D9 at all. Instead it manually uploads every source frame to the GPU and downloads every upscaled frame again to system RAM. madVR does everything on the GPU, using D3D9 <-> OpenCL interop.
In the OSD, the refresh rate reported is ~60.0178 Hz. I just tried v0.86.11 and found that SM switching is working perfectly with that version. I have been using smooth motion since the very first release; had there been any such issue earlier, I would have reported it.
Alright. So the next question would be, just to confirm: The problem does occur with v0.87.4 both with profiles activated and deactivated, is that correct? So it's not the profiling that is causing the problem, but the switch "only if there would be judder..." fails to work properly for you, when using v0.87.4, right? If that is correct, then please create a debug log for me with a video where SM FRC gets turned on, although it should be deactive. The log should tell me why madVR thinks it should be activated. And maybe you could also upload a small sample of that video for me to test with, so we're looking at the same thing. And if you use deint, please tell me whether you're using forced film mode or not.
madshi
1st February 2014, 10:39
Since I will be assuming that your input is DVD or BluRay content, or otherwise native 8-bit content, disable dithering completely: You only use Nearest Neighbor scaling, which does not increase the image bitdepth, so you can safely round down or truncate back to 8 bit. You can find the setting in the Trade quality for performance section.
IMHO that is a bad suggestion because already color conversion produces floating point RGB values. So dithering is needed to produce correct output.
Just thought of one additional reason why "passthrough" from a source device to the processor won't work, technically: SD content is usually encoded in 720x480 (NTSC) or 720x576 (PAL) pixels. That is an encoded aspect ratio of 1.5 respectively 1.25. There's pretty much zero content out there with such a native aspect ratio. Which means that *every* single DVD or broadcast SD source will require the source device to perform scaling to produce the correct aspect ratio. That applies to HTPCs as much as to consumer electronics DVD and Blu-Ray players. IMHO every Blu-Ray player offering to do a "pure direct" mode is simply lying because "pure direct" is technically not possible. What they mean is that they turn off any processing that is not strictly needed, but there will still be quite a bit of processing, nevertheless, especially for SD content.
And since we now established that AR correction is needed for SD sources in any case, using Nearest Neigbor for chroma upscaling (as a passthrough trick/hack) suddenly becomes a really bad idea.
The 8472
1st February 2014, 11:24
Actually my idea doesn't work. I could apply Error Diffusion in YCbCr color space, but that's not what madVR outputs. I'd have to convert back to RGB afterwards. And after the conversion I'd again have RGB floating point data. So I'd have to apply Error Diffusion yet again. I've been thinking about this some more. The point of error diffusion is to keep track of the quantization error while walking over the pixels.
This consists of 3 steps
1. find the color of the new pixel
2. calculate the error of the quantization
3. distribute the error among neighboring unprocessed pixels
Normally you want to do everything in a perceptually uniform colorspace to calculate a visually more accurate delta-E because palette-based dithering introduces a much greater quantization error than the following conversion from that colorspace back to RGB. But since madVR is dithering to 6/7/8 bits per channel we're already working with a huge color "palette", thus keeping the quantization error small so the back-conversion error would become more significant.
I think i may have a solution for that
Do 1. as float/16bit-RGB* -> integer-RGB as you do now but do 2. and 3-. in a float/16bit-LAB* to calculate the error. Then apply it to the other pixels and then transform them back to float/16bit-RGB*.
Equivalently you could also do 1. in float/16bit-LAB* -> integer-RGB, then convert the integer-RGB back to float/16bit-LAB*, calculate the error (2.) and distribute the error back to the other LAB*-pixels (3.). That would reduce the number of colorspace conversions required for each iteration.
The key point is doing 1. with intRGB as target while doing 2 and 3 in a different color space.
* I don't know where you're using float or high bit depth integer math internally, so consider the given terms as placeholders.
Maybe simply applying some dynamic weights (http://www.compuphase.com/cmetric.htm) on the errors for the individual RGB channels would also reduce some of the infamous blue specks, but I don't know how it would affect the luminance error distribution.
DarkSpace
1st February 2014, 12:16
If you create a file with the filename "YCbCr" (no extension, and the file itself may be empty) in the directory madVR resides in, madVR will not convert the data to YUV before presenting it.
Of course I meant "RGB" there, not "YUV". I edited my original post.
IMHO that is a bad suggestion because already color conversion produces floating point RGB values. So dithering is needed to produce correct output.
I specifically mentioned a way to avoid the YUV-to-RGB-conversion because of this. Did I forget any other step that increases bitdepth?
Just thought of one additional reason why "passthrough" from a source device to the processor won't work, technically: SD content is usually encoded in 720x480 (NTSC) or 720x576 (PAL) pixels. That is an encoded aspect ratio of 1.5 respectively 1.25. There's pretty much zero content out there with such a native aspect ratio. Which means that *every* single DVD or broadcast SD source will require the source device to perform scaling to produce the correct aspect ratio. That applies to HTPCs as much as to consumer electronics DVD and Blu-Ray players.
I also addressed that point, though of course, I have no idea how much sense my suggestions actually make.
IMHO every Blu-Ray player offering to do a "pure direct" mode is simply lying because "pure direct" is technically not possible. What they mean is that they turn off any processing that is not strictly needed, but there will still be quite a bit of processing, nevertheless, especially for SD content.
I have honestly no idea about that. All I was operating on was the fact that 5ts said that those players offer a pass-through mode. In my opinion, the best way to realize such a pass-through is to use NN scaling, and have the device interpret the data as 4:2:0 instead. You may very well be right, though, and in that case, of course, my suggestions become relatively worthless.
And since we now established that AR correction is needed for SD sources in any case, using Nearest Neigbor for chroma upscaling (as a passthrough trick/hack) suddenly becomes a really bad idea.
I agree. Although I still think that NN scaling (and inverted NN scaling to source resolution) might be able to circumvent that. (Not like I think that player manufacturers actually care about stuff like that, to be honest, though...)
madshi
1st February 2014, 12:29
I've been thinking about this some more. The point of error diffusion is to keep track of the quantization error while walking over the pixels.
This consists of 3 steps
1. find the color of the new pixel
2. calculate the error of the quantization
3. distribute the error among neighboring unprocessed pixels
Normally you want to do everything in a perceptually uniform colorspace to calculate a visually more accurate delta-E because palette-based dithering introduces a much greater quantization error than the following conversion from that colorspace back to RGB. But since madVR is dithering to 6/7/8 bits per channel we're already working with a huge color "palette", thus keeping the quantization error small so the back-conversion error would become more significant.
I think i may have a solution for that
Do 1. as float/16bit-RGB* -> integer-RGB as you do now but do 2. and 3-. in a float/16bit-LAB* to calculate the error. Then apply it to the other pixels and then transform them back to float/16bit-RGB*.
Equivalently you could also do 1. in float/16bit-LAB* -> integer-RGB, then convert the integer-RGB back to float/16bit-LAB*, calculate the error (2.) and distribute the error back to the other LAB*-pixels (3.). That would reduce the number of colorspace conversions required for each iteration.
The key point is doing 1. with intRGB as target while doing 2 and 3 in a different color space.
* I don't know where you're using float or high bit depth integer math internally, so consider the given terms as placeholders.
Maybe simply applying some dynamic weights (http://www.compuphase.com/cmetric.htm) on the errors for the individual RGB channels would also reduce some of the infamous blue specks, but I don't know how it would affect the luminance error distribution.
Hmmmm... It's an interesting idea. Though, I think I would stick to simple linear light instead of LAB, because I doubt that LAB would bring enough benefit to justify the extra conversion costs. Or what do you think? I could, however, calculate the error in linear light YCbCr and spread it to the neighbor pixels that way. That should (at least partially) losen the strict "each channel on its own" processing. Not sure how that would affect the "look" of error diffusion...
I specifically mentioned a way to avoid the YUV-to-RGB-conversion because of this. Did I forget any other step that increases bitdepth?
No, you're right, I forgot about your idea of dropping the YCbCr -> RGB conversion.
I also addressed that point, though of course, I have no idea how much sense my suggestions actually make.
I don't think that will work. I don't know of any automated way to tell the video processor which original the source had. I believe the source must correct the aspect ratio, otherwise the aspect ratio will stay distorted.
The 8472
1st February 2014, 13:15
Hmmmm... It's an interesting idea. Though, I think I would stick to simple linear light instead of LAB, because I doubt that LAB would bring enough benefit to justify the extra conversion costs. Or what do you think? I could, however, calculate the error in linear light YCbCr and spread it to the neighbor pixels that way. That should (at least partially) losen the strict "each channel on its own" processing. Yeah, LAB was just a suggestion based on what I normally do in imagemagick. Anything psychovisually better than RGB should already help with those blue specks.
Note that the thing should be faster if you do everything in LL YCbCr except doing LL YCbCR -> integer RGB in the quantization step. That'll only cost you 3 color space conversions per pixel instead of 7.
Step by step:
1. convert entire frame to high depth LL YCbCr if it's not already in that format [conversion 1]
Kernel:
2. convert source pixel to closest integer RGB target value [conversion 2]
3. convert target pixel back to LL YCbCr [conversion 3]
4. calculate error in source color space
5. spread to neighboring source pixels
If you do it in RGB you would have to do the conversions for the error calculation and for each neighboring pixel separately.
Not sure how that would affect the "look" of error diffusion...
For comparison purposes you could do
a) output to 3-4bit per channel to get exaggerated results
b) do a 0.0-10.0 gradient, dither it, take a screenshot and expand levels to 0-255
6233638
1st February 2014, 13:46
Yes, especially with game/computer material the differences are extreme. Linear resampling algorithms simply suck with such source material. Video content usually has much less contrasty lines and much less native aliasing in the original source, though, so with real life video the advantage of NNEDI3 is quite a bit smaller.Smaller, but still significant at times.
I think we need to find clear terms when talking about these things. When I first read your post I understood it exactly the wrong way. Maybe if we want to differentiate between NNEDI3 chroma upscaling and chroma channel doubling we should refer to one of these as "4:2:0 -> 4:4:4 chroma upscaling", and to the other as "chroma channel doubling"?Sorry, I tried to be clear with my "Chroma Doubling/Chroma Scaling" - it is easy to get the two confused though.
I have a similar impression - although with some images going to 64 and then to 128 neurons does still result in a nice quality improvement, so I find it hard to draw any definite conclusions. It's a valid question, though, which neuron setting to recommend, since every higher setting almost doubles the performance cost.I was only talking specifically about chroma with the 32 neuron recommendation. Luma definitely benefits all the way up to 256 neurons. (though some images seem to work better with lower settings, but overall higher generally seems to be better)
Some small comparisons from an old 360p video I've been testing with:
Chroma Scaling: Bicubic 75 AR | Jinc 3 AR | NNEDI3 32 Neurons
http://abload.de/img/bc75ar-chroma6cch6.pnghttp://abload.de/img/j3ar-chromafoijh.pnghttp://abload.de/img/nnedi3-32-chroma16i18.png
Luma/Chroma Scaling: Nearest Neighbor | Mitchell-Netravali | Jinc 3 AR | NNEDI3 64/32 with J3AR
http://abload.de/img/nn-dialb2i98.pnghttp://abload.de/img/mn-dial4sfzo.pnghttp://abload.de/img/j3ar-dial7ceum.pnghttp://abload.de/img/nnedi3-dialbvdjm.png
And a fullscreen example: Nearest Neighbor | Mitchell-Netravali Luma/Bicubic 75 Chroma | Jinc 3 AR | NNEDI3 64/32 with J3AR
http://abload.de/thumb/nn-fulljgfr6.jpg (http://abload.de/img/nn-fulljgfr6.jpg)http://abload.de/thumb/mn-full55c1a.jpg (http://abload.de/img/mn-full55c1a.jpg)http://abload.de/thumb/j3ar-fullmnfrq.jpg (http://abload.de/img/j3ar-fullmnfrq.jpg)http://abload.de/thumb/nnedi3-fullvdfok.jpg (http://abload.de/img/nnedi3-fullvdfok.jpg)
I'm using Mitchell-Netravali in some of the examples, as that was one of the best balances between sharpness and ringing that was available at the time when I first tested this source.
As I said, the results are ridiculously good, considering the quality of the source.
I've found that the output of luma doubling is soft enough that using 16 neurons for quadrupling doesn't seem to produce any problems. So I think it's fine to use 16 neurons for quadrupling, which makes it more feasible from a performance point of view. Whether that brings enough quality advantage to make it worthwhile is another question, but when using really small sized sources, or when upscaling to very high resolution monitors, it might make sense.It's not that performance is an issue - I just haven't seen anything where quadrupling is beneficial over simply increasing the neuron count of doubling, and in some cases it has been detrimental - but it may be because I am only scaling to 1080p
They used to do their own film mode detection back in the old days, before they had a Gennum VXP chip built it. Ever since they're using the Gennum VXP chip (now Sigma owned, IIRC), I believe they're relying on the Gennum to do deinterlacing and film mode detection and all that stuff. I don't think they use their own FPGA for that, anymore.Oh I forgot that they were using the Gennum VXP for deinterlacing, I was thinking it was only being used for noise reduction.
DragonQ
1st February 2014, 13:57
Stop posting these comparison screenshots, it makes me sad that I have no GPUs capable of NNEDI3. :(
I guess one of these days I'll want to play a game that requires a better GPU than my GTS 250 and then at least I'll be able to use NNEDI3 on my desktop...
leeperry
1st February 2014, 14:04
Yep, Joe Kane might stop complaining about 4K scaling (http://www.avsforum.com/t/1471727/) after he tries NNEDI in mVR :D
Still utterly impressed by the new build, using my 7850 I can do 30fps SD or 720p@1080p with error diffusion, low debanding(analyzing gradient angles), Bicubic 75 AR chroma, NNEDI 2X/4X@32 neurons for luma, luma upscaling@lanczos3 AR and downscaling@CR AR....I can't engage LL for the latter or J3AR for upscaling otherwise I reach >90% GPU load and it starts dropping frames.....Kinda makes me wish I bought a 7870 but my factory overclocked 7850 is almost as fast as a 7870, there's still hope that madshi could improve performance a tiny bit or I might just grab a 150€ 7950...that's ridiculous money for this kind of PQ and I could prolly swing 64 neurons all the way :scared:
For the record I use a G3420 Pentium that never reaches >30% load, using software video decoding at that.
vBulletin® v3.8.11, Copyright ©2000-2026, vBulletin Solutions Inc.