View Full Version : madVR - high quality video renderer (GPU assisted)
TinTime
4th May 2009, 03:17
With all other renderers (and software decoding) it connects using YUY2. Perhaps it's a HD YV12 connection that causes the problem? Mind you, connecting the Nvidia decoder with YUY2 instead of YV12 to ffdshow and AviSynth doesn't make any difference for me so perhaps not.
On my machine NVIDIA -> ffdshow -> any other renderer (VMR9, Overlay, EVR, & Haali) and I get a perfect picture.
I wonder why I don't then? Might be my ffdshow version (revision 2527). To be honest I don't ever use the Nvidia decoder except for odd occasions like now. Interestingly I remember that I stopped using it when I first came across an HD MPEG-2 stream because I ran into some problem. Unfortunately I can't remember what the problem was...
racerxnet
4th May 2009, 03:46
The thing is, madVR is the only renderer that has issues with the NVIDIA decoder playing high-def content
I have had the Nvidia decoder working with the Arcsoft filters in the past for DXVA acceleration on BR content, so it is possible. This is with MadVr for final output. I don't bother with HD content as I have about 600 SD DVD's which I scale to 1366 x 768. It is the panels native resolution, so I am stuck with this for now. The new 09 renderer works quit well as long as I do not push the resampling to hard. With resampling at default values I can turn off the anti tearing option. If I resample with lancoz 8 and tearing off I then get tearing on the top of the screen.
Just my .02
MAK
I
I'm sorry, I should have been more clear.
madVR 0.8 did not measure the time needed to upload the textures to the GPU. madVR 0.9 does measure this time and includes it in the GPU rendering statistics. Which means that if you compare madVR 0.8 statistics to madVR 0.9, you have to substract the madVR 0.9 "uploading textures" time from the average GPU rendering time. I'm aware that this is not really intuitive. The problem is that madVR 0.8 simply did not show the complete numbers. However, strange enough, madVR 0.9 shows lower GPU rendering times for me even without doing this math...
Does uploading textures mean updating textures?
So for MadVR 0.7 and 0.9:
avr gpu 0.9 - upd text 0.9 <> (?) avr gpu 0.7 ?
The thing is, madVR is the only renderer that has issues with the NVIDIA decoder playing high-def content (VMR9, Overlay, EVR, & Haali are all fine). Since that is the case, you would assume madVR is at fault. On my machine NVIDIA -> ffdshow -> any other renderer (VMR9, Overlay, EVR, & Haali) and I get a perfect picture. NVIDIA -> ffdshow -> madVR gets misaligned luma and chroma, as you mentioned.
whoa that brings back memories.. when i used to have an nvidia card i had that chroma offset problem pop its head up every now and again.. sorry i never figured out what the cause is.. i just got an ati..
tetsuo55
4th May 2009, 10:47
@Everyone discussing CPU for AVC decoding.
The samples mentioned are relatively easy to decode and do not require a C2D @ 3ghz.
For smooth playback you need a peak cpu usage of no higher than 75%, above this value dropped frames and/or jitter (might) become a problem.
Really difficult movies (of which very few exist) will only meet this limit on a 3ghz+ C2D or any Quadcore CPU worth its salt.
I must admit that some pretty cool updates have been made to both ffdshow and coreavc since i saw the last benchmarks and i can image the required CPU power having come down a bit since then.
I know of 2 really difficult samples that drop frames even on cheaper standalone players.
1. The "killa" or "killer" sample (have never been able to find it myself though)
2. A scene from i believe BBC's planet earth. This scene involves a pan over a heavily forested mountain range, a large group of birds fly by. The changes from one frame to the next are so large that almost none of the optimisations help out here and each and every frame needs to be completely decoded, the AVC also maxes out all the legal bluray specs.
But like i said, most AVC files don't even come close to this level of complexity/difficulty to decode and it's these scene's that require the 3ghz CPU.
tetsuo55
4th May 2009, 10:54
@Madshi:
Bugreport: Settings are not saved when madVR is run from a protected directory.
I reported earlier that the .bat install method does not get the required admin rights needed to install madVR.
I found a new problem related to installing madVR to a protected directory like "program files"
Because the filter is not run with admin rights i am unable to save the settings (because obviously "program files" is read only to any non-admin application)
madVR does not show an error that saving settings failed (i think it just assumes it will always work)
To work better in the future madVR should be limited user aware, and should ask for admin rights when saving settings.
I can also confirm the DVD decoding problem.
I will try to make a small (couple of MB) fake dvd and see if that produces the error.(which i can then upload for you to test with)
P.S.
If anyone else want's to make a small dvd sample, please go ahead, it will take a while before i will be able to do so.
any dvdmastering software could be used for this with a small free-to-distribute sample, like big buck bunny
Mark_A_W
4th May 2009, 11:38
madVR 0.9 released
Those of you for which the final refresh rate estimate is off big time, could you please reproduce the problem, and then look into the madVR folder? There should be a file named "VSync.dat" there with detailed VSync data, which will hopefully help me find out what's going wrong on your PCs. Could you please zip and upload the file? Thanks! The file is only created if madVR detects that the final estimate is probably wrong.
File attached Madshi.
It should read ~95.904hz (it's interlaced, but that's irrelevant).
Display 2 is right.
Display 3 reads 130+hz fullscreen, and 105+hz in a window.
Not sure if you are worried yet, but v0.9 is not smooth like v0.8 was.
Thanks
Mark
@Everyone discussing CPU for AVC decoding.
The samples mentioned are relatively easy to decode and do not require a C2D @ 3ghz.
For smooth playback you need a peak cpu usage of no higher than 75%, above this value dropped frames and/or jitter (might) become a problem.
Really difficult movies (of which very few exist) will only meet this limit on a 3ghz+ C2D or any Quadcore CPU worth its salt.
Not so sure about that. Depends on multi-threading optimisation as well, most codecs have difficulties exploiting more than just two cores efficiently. Besides wolfdale even though just dual is faster in some specific tests than quad, due to the wonderful cache size ;)
As for the 75% CPU consumption, that is the reason for my worries, as with mVR 1080p bluray video is just a bit below that threshold, 60% CPU on the OP scene. With Haali it is typically around 40% so if I watched blurays I'd choose Haali over current version of mVR due to less CPU power.
As for the different CPU-only codecs, you must be wrong.
I even rechecked it again with that 1080p video and as expected there's little to none difference in CPU amongst MPC internal, ffdshow or CoreAVC. They use different cpu% on different scenes so hard to say, but as a stab most of the times cpu load difference is within 5%.
As for the absolute values, 1080p AVC high-profile rendered and downscaled to 720p by mVR (shader math only) takes about 30-45% CPU. Which is substantially less than rendering same video without any rescale :rolleyes: that takes about 50-60%. Very roughly, the codecs may be ranged as follows: CoreAVC, MPC internal, ffdshow in the order of CPU load but again the difference is so small that can be discarded (for instance I don't even use CoreAVC nowadays since no CUDA gain for me anyway).
littleD
4th May 2009, 18:13
I can confirm that cpu usage when using madvr is much reduced. I tested 720p h.264 as its the only source resolution i can play on my system. And its playable now! :) I get similar result with evr actually. Tested with MPC HC 1093 on ati 3450.
Softcubic is very smooth but at least cause no ringing and jaggies. Another nice scene to test is very begining of "Flushed away" with dreamworks title appearing from balloons.:)
Edit:
Ok i was too enthusiastic :) madvr still takes about 20% more cpu than evr. But is better compared to constant 100% in previous versions.
tetsuo55
4th May 2009, 19:17
Not so sure about that. Depends on multi-threading optimisation as well, most codecs have difficulties exploiting more than just two cores efficiently. Besides wolfdale even though just dual is faster in some specific tests than quad, due to the wonderful cache size ;)
As for the 75% CPU consumption, that is the reason for my worries, as with mVR 1080p bluray video is just a bit below that threshold, 60% CPU on the OP scene. With Haali it is typically around 40% so if I watched blurays I'd choose Haali over current version of mVR due to less CPU power.
As for the different CPU-only codecs, you must be wrong.
I even rechecked it again with that 1080p video and as expected there's little to none difference in CPU amongst MPC internal, ffdshow or CoreAVC. They use different cpu% on different scenes so hard to say, but as a stab most of the times cpu load difference is within 5%.
As for the absolute values, 1080p AVC high-profile rendered and downscaled to 720p by mVR (shader math only) takes about 30-45% CPU. Which is substantially less than rendering same video without any rescale :rolleyes: that takes about 50-60%. Very roughly, the codecs may be ranged as follows: CoreAVC, MPC internal, ffdshow in the order of CPU load but again the difference is so small that can be discarded (for instance I don't even use CoreAVC nowadays since no CUDA gain for me anyway).
The problem is, that you can really only compare benchmarks if they are run with the killa sample and the difficult scene from bcc's planet earth(which would require the original bluray).
I have yet to actually encouter a problematic movie in real life, so far everything has worked just fine on a c2d mobile 2,6 ghz(although i have seen cases go over 75% for a second)
iSunrise
4th May 2009, 19:40
@madshi:
I´ve just tested v0.9 thoroughly and I must say I´m quite impressed with the lower GPU usage times, that is by comparing the numbers with _and_ without "update textures". Also, CPU usage seems to be lower, too, which is a great step forward. This is on a very fast Core i7 and a GTX260-216 with >GTX285 clocks and ZoomPlayer with Luma set to Spline36 and Chroma set to SoftCubic100.
However, not to my very liking v0.9 made one thing very obvious to me, which is good in a way, since you can hopefully reproduce and fix this.
Now, being more specific:
There is at least one very noticeable problem if you use CoreAVC with CUDA decoding, which will partly or completely go away if you either uncheck CUDA-decoding or use ffdshow as a decoder (ffmpeg-mt selected).
Here´s a small list of things that I came across. Each decoder is mentioned so you can check them step-by-step. I´ve reproduced each step a dozen times just to be sure.
Movie sample (official trailer):
wolverine-tlra_h1080p.mov
Now the decoders (madVR v0.9 as renderer):
CoreAVC with CUDA-decoding:
1) Avrg gpu rendering time is noticably higher (*1) and max gpu rendering time/update textures goes through the roof (depending on madVR settings ranging from 3-8 times higher, which is madVR settings dependant *1)
2) Display estimate 3 very often resets to 0.00000Hz and stays there for several seconds while the movie is playing or paused and Display will show [1s]
CoreAVC without CUDA-decoding:
1) Max gpu rendering time/update textures looks fine
2) Display estimate 3 often resets to 0.00000Hz and stays there for several seconds while the movie is playing or paused and Display will show [1s]
ffdshow:
1) Max gpu rendering time/update textures looks fine
2) Display estimate 3 occasionally resets to 0.00000Hz and Display will show [1s]
*1: Compared to CoreAVC without CUDA-decoding and ffdshow
It looks like 1) is a result of both (decoder and renderer) using the GPU extensively and/or it´s related to your new video->GPU uploading method. The higher avrg rendering times and the drastically higher max rendering time numbers just don´t make any sense to me. If I choose higher settings (like Lanczos8 on both Luma and Chroma) my max gpu rendering times are sometimes higher than the movie frame interval. If I´m only using software/cpu-decoding the max gpu rendering times/updating textures are 1/5 of that, so it will never reach the frame interval, regardless of the settings I choose.
I hope you can look into this. Thanks.
Finally, here´s the 3 shots (coreavc+cuda/coreavc-nocuda/ffmpeg-mt):
http://www.abload.de/thumb/wolverine_coreavc_cudamw4j.png (http://www.abload.de/image.php?img=wolverine_coreavc_cudamw4j.png) http://www.abload.de/thumb/wolverine_coreavc_nocu3o56.png (http://www.abload.de/image.php?img=wolverine_coreavc_nocu3o56.png) http://www.abload.de/thumb/wolverine_ffmpeg-mt_000o43.png (http://www.abload.de/image.php?img=wolverine_ffmpeg-mt_000o43.png)
kostik
4th May 2009, 20:27
what are the best chroma and luma resampling settings? for h264?
to get the best quality?
6233638
4th May 2009, 21:23
what are the best chroma and luma resampling settings? for h264?
to get the best quality?
I found Mitchell-Netravali to be the best compromise between sharpness and ringing for luma, and SoftCubic 50 to be best for chroma. (100 blurs it too much)
OK, now moar comparison for scaling algorithms, this time for the anime content.
I have this rather good R2 dvd so the video is unaltered in any way apart from decoding and rendering by mVR.
Scaling specified is only for luma, chroma is softcubic50 for each screenshot.
Bilinear:
http://img3.imagebanana.com/img/j35xr7x/thumb/03bl.png (http://img3.imagebanana.com/view/j35xr7x/03bl.png)
C-R:
http://img3.imagebanana.com/img/osmrlbgj/thumb/03cm.png (http://img3.imagebanana.com/view/osmrlbgj/03cm.png)
Lanc4:
http://img3.imagebanana.com/img/uwicce2d/thumb/03l4.png (http://img3.imagebanana.com/view/uwicce2d/03l4.png)
Lanc8:
http://img3.imagebanana.com/img/qqo2jy59/thumb/03l8.png (http://img3.imagebanana.com/view/qqo2jy59/03l8.png)
Splin64:
http://img3.imagebanana.com/img/k7wjx92e/thumb/03sp.png (http://img3.imagebanana.com/view/k7wjx92e/03sp.png)
I haven't taken bicubics and softcubics as they were rather soft anyway.
In order to compare you need to download these pngs and compare them in a viewer so that they would be shown in exactly same position on screen (I use acdsee and scroll with a mousewheel, so I can quickly switch the pictures).
I really find hard to describe the difference between bilinear and C-R methods so for the purposes of upscaling C-R is bad, imo. I haven't saved mitchell, unfortunately (and mVR is hard to point to the exactly same frame ;P). It wasn't bad, but somewhat in between bilinear and lanc, still too soft for such task.
My personal choice is Spline64, as I've been using splines for quite a while even with ffdshow resize. Interesting that difference between lanc8 and spl64 is very subtle, although the methods differ considerably (?). Edges are a bit softer with spl64 though but seems overall shaprness is good for any of them. Basically normal unaltered DVD anime content (read: dvd content, not the video damaged by 95% of encoders in the wild ;P) can be quite watchable when upscaled to 720p with either lanc8 or spl64. Of course it is only valid for large regions with contrast edges, smaller details like text etc looks blurry in any way (as expected, that what we need HD for:))
6233638
5th May 2009, 04:32
From that comparison, I would say:
Bilinear is too soft, and suffers from aliasing.
Lanczos8 introduces too many artefacts.
Spline64 is very similar to Lanczos4 but with marginally less ringing.
Spline64 is sharper than Catmull-Rom, at the expense of introducing more ringing into the picture.
Personally, from those examples, I would choose Catmull-Rom as I'm quite adverse to ringing. I'd like to see how Mitchell Netravali compares, as it seemed to be the best from the testing I did, but that was filmed content rather than animated. It would also be good to see an unscaled image to get an idea of how sharp those lines should be.
Thunderbolt8
5th May 2009, 10:11
madshi, where is the focus when downscaling, on keeping the colours as close to the original or (also) on sharpness (with default options)? because I'm wondering whether I now need to apply another sharpness level in ffdshow as with haali filter when I scale down from 1080p to 720p res. I have the feeling that madvr already displays the picture a little sharper then as haali, is this correct?
I found Mitchell-Netravali to be the best compromise between sharpness and ringing for luma, and SoftCubic 50 to be best for chroma. (100 blurs it too much)
And what was your input source (resolution) and display resolution? Did your resize, up or down?
6233638
5th May 2009, 11:09
And what was your input source (resolution) and display resolution? Did your resize, up or down?
I was using blu-ray at double-size just for testing. (as blu-ray tends to have better quality video than DVD with less artefacts)
madshi
5th May 2009, 11:28
Here are the same frame (from 0.8, will install 0.9 right after). [...] Add Haali renderer shots. Now it's hard to see the difference, but it's still here.
From what I can see, both EVR and Haali are slightly sharper than your madVR screenshot. That might make a difference in banding, too? Which madVR upscaling filter are you using? You may want to try a sharper one (e.g. Lanczos) to see whether the increases banding or not.
With all other renderers (and software decoding) it connects using YUY2. Perhaps it's a HD YV12 connection that causes the problem?
That's quite possible. After all, the MPC HC decoders also worked fine with YUY2 and had a bug when forced to output YV12.
Does uploading textures mean updating textures?
So for MadVR 0.7 and 0.9:
avr gpu 0.9 - upd text 0.9 <> (?) avr gpu 0.7 ?
Yes, upload=update. Not sure what you mean with your second question.
Bugreport: Settings are not saved when madVR is run from a protected directory.
[...]
To work better in the future madVR should be limited user aware, and should ask for admin rights when saving settings.
This is a problematic situation. Asking for admin rights in order to save settings would be bad user experience IMHO. Of course I could do what Microsoft suggests, namely storing the settings into the user profile directory tree. But actually I hate this logic because it scatters all the files all over the harddisk. So I'm not really sure how to handle this.
Any opinions/suggestions?
CoreAVC with CUDA-decoding:
1) Avrg gpu rendering time is noticably higher (*1) and max gpu rendering time/update textures goes through the roof (depending on madVR settings ranging from 3-8 times higher, which is madVR settings dependant *1)
Interesting numbers, thanks. It is probably to be expected that average rendering numbers are a bit higher when CUDA is used, because obviously with both CUDA + madVR active the GPU will be more busy than with just one of them active. I don't think I can do anything about that.
The max gpu rendering times look really bad, but as I've already said multiple times, in the long run the max gpu rendering times are not very important.
I found Mitchell-Netravali to be the best compromise between sharpness and ringing for luma, and SoftCubic 50 to be best for chroma. (100 blurs it too much)
I really find hard to describe the difference between bilinear and C-R methods so for the purposes of upscaling C-R is bad, imo. I haven't saved mitchell, unfortunately (and mVR is hard to point to the exactly same frame ;P). It wasn't bad, but somewhat in between bilinear and lanc, still too soft for such task.
My personal choice is Spline64
That goes to show that luma scaling algorithms are really a matter of taste. Most of the algorithms have some advantages and some disadvantages. Personally, I like Lanczos4 for its sharpness, but (obviously) I dislike the ringing. I like Mitchell, but actually I like SoftCubic50 even more, which I find very similar to Mitchell, but with less aliasing. So sometimes I'm using Lanczos4 and sometimes SoftCubic50.
For chroma I'm using SoftCubic100. Interesting that both of you guys prefer SoftCubic50 for chroma.
Interesting that difference between lanc8 and spl64 is very subtle, although the methods differ considerably (?)
Lanczos and Spline produce very similar results, I think. However, you should compare with the same number of taps. Lanczos3 and Spline36 use 3 taps. Lanczos4 and Spline64 use 4 taps. Lanczos8 uses 8 taps. Personally, I can see some differences between 3 and 4 taps. But I don't really see much of a difference between 4 and 8 taps. Except that 8 taps rings quite a bit more...
where is the focus when downscaling, on keeping the colours as close to the original or (also) on sharpness (with default options)? because I'm wondering whether I now need to apply another sharpness level in ffdshow as with haali filter when I scale down from 1080p to 720p res. I have the feeling that madvr already displays the picture a little sharper then as haali, is this correct?
madVR does not tamper with the colors, regardless of which downscaling filter you're using. madVR also does not apply any sharpness algorithm. I don't know if madVR is sharper than Haali, that probably also depends on which downscaling filter you're using. If you find madVR generally sharper than Haali, then Haali probably does something wrong.
nijiko
5th May 2009, 14:43
>>madshi
Excuse me. Madshi. Have you known the problem with NVidia Video Decoder in HD clip?
mark0077
5th May 2009, 14:58
madshi, regarding saving renderer settings, could you save to registry as another option.
That goes to show that luma scaling algorithms are really a matter of taste. Most of the algorithms have some advantages and some disadvantages. Personally, I like Lanczos4 for its sharpness, but (obviously) I dislike the ringing. I like Mitchell, but actually I like SoftCubic50 even more, which I find very similar to Mitchell, but with less aliasing. So sometimes I'm using Lanczos4 and sometimes SoftCubic50.
For chroma I'm using SoftCubic100. Interesting that both of you guys prefer SoftCubic50 for chroma.
I haven't tested chroma rescalers yet. However I don't think it is really important as chroma has originally just a quarter of information compared to luma. Besides anime specifics with gradients and solid regions also needs to be taken into account. So I basically left SoftCubic50, is it sharper than 100?
As for the splines, even though they are different they are still quite remarkably similar. Would be interesting if somebody analyzed these pictures and posted enlarged cuts displaying the difference in ringing. I kind of don't see any worth mentioning in spline64 :)
And is it possible to implement spline256 since there is lanc8 which should be equivalent?
madshi
5th May 2009, 15:15
Have you known the problem with NVidia Video Decoder in HD clip?
I guess it's caused by a bug in the NVidia decoder due to being forced to output YV12, so I don't plan to further look into this right now.
regarding saving renderer settings, could you save to registry as another option.
Well, I could try to save to the ini first, and if that fails, store the settings in HKEY_CURRENT_USER. That should work...
I haven't tested chroma rescalers yet. However I don't think it is really important as chroma has originally just a quarter of information compared to luma. Besides anime specifics with gradients and solid regions also needs to be taken into account. So I basically left SoftCubic50, is it sharper than 100?
SoftCubic100 is *very* soft, but I like that best for chroma. Most of the time there's no difference visible. But if there is, to me SoftCubic100 looks better because it gets rid of any jaggies which SoftCubic50 doesn't always do. So SoftCubic100 is madVR's default setting for chroma resampling.
Would be interesting if somebody analyzed these pictures and posted enlarged cuts displaying the difference in ringing. I kind of don't see any worth mentioning in spline64 :)
Look at the upper part of the nose (that one dark stroke between the eyes). There's a dark shadow left and right of that stroke.
And is it possible to implement spline256 since there is lanc8 which should be equivalent?
No, that's not possible because I don't know the correct formula for spline256. For Lanczos the basic formula is always the same, regardless of how many taps you use. For spline the formula's coefficients are different, depending on the number of taps.
Casshern
5th May 2009, 15:52
SoftCubic100 is *very* soft, but I like that best for chroma. Most of the time there's no difference visible. But if there is, to me SoftCubic100 looks better because it gets rid of any jaggies which SoftCubic50 doesn't always do. So SoftCubic100 is madVR's default setting for chroma resampling.
Instead of upscaling luma and chroma separatly, its much better to use luma information to upscale chroma. This would use the higher res luma information to more accurately reconstruct the lowres chroma information. The easiest would be to use a delta in luma as weights for interpolating the chroma. As it is a obvious idea, there probably is already some work on it done.
Nvidia IGP 9400 +512m, Intel 8300 2.83, Vista 32, AERO, ffdshow mt, MPC-НС 1065
madVR 0.7 reports:
luma - lancsoz3
chroma - softcubic100
1) Input - 1080x1920/23.976, video refresh rate (Nvidia control panel) - 23Hz.
Output - NO RESIZE, monitor resolution - 1920x1080
madVR 0.7 OSD says
avrg GPU rendering time - 25 (stable)
2) Input - 1080x1920/23.976, video refresh rate (Nvidia control panel) - 23Hz.
Output - RESIZE 1920x1080 --> 132,0,1766,919
madVR 0.7 OSD says:
avrg GPU rendering time - 45
CPU load
OSD OFF - 65-72%
And now for MadVR 0.9:
luma - lancsoz3, chroma - softcubic100
1)NO RESIZE
avrg GPU rendering time - 26.5-6.5(updating textures time)=20 (stable)
-20% against 0.7 :)
2)RESIZE
avrg GPU rendering time - 45.5-7.5=38
-15.5% against 0.7 :)
CPU load
OSD OFF - 60%
-14% against 0.7 :)
3) Scaling up 576 (PAL DVD) --> 1080
luma - lancsoz4, chroma - softcubic100
avrg GPU rendering time - 29.3
My IGP Nvidia 9400 can upscale PAL DVD to 1080p with sharp lancsoz4! Thank you madshi!l
madshi
5th May 2009, 16:13
Instead of upscaling luma and chroma separatly, its much better to use luma information to upscale chroma.
You say that as if it was a fact. Have you seen this done? And have you seen that it's actually "much better"?
This would use the higher res luma information to more accurately reconstruct the lowres chroma information. The easiest would be to use a delta in luma as weights for interpolating the chroma.
Doing it that way would mean making brighter pixels more saturated and darker pixels less saturated, right? I don't really see how that would be "more accurate". But I'm not an expert in this area. Is there any "scientific" reason for making brighter pixels more saturated than darker pixels?
As far as I can see, luma and chroma are independent and the luma value does not have any direct influence on chroma. So I don't really see how luma can help upsampling chroma better. But then we're talking about gamma corrected Y'CbCr and not linear light YCbCr. And IIRC I've been told that there is a bit of luminance in Cb and Cr, too. Argh, this is complicated.
@yesgrey, your opinion?
As it is a obvious idea, there probably is already some work on it done.
There is an article which handles border cases where chroma is spread to neighbor pixels, if luma is too dark or too bright to hold the upsampled chroma. I've yet to look into implementing a similar algorithm. But the article only handles such corner cases and does not *generally* reshuffle the chroma. Actually the author of that article told me that a friend of his suggested to use luma to form chroma better, but he was not convinced of his friend's efforts...
leeperry
5th May 2009, 16:21
I have the feeling that madvr already displays the picture a little sharper then as haali, is this correct?
yes I agree. mostly because the chroma is blurrier than ffdshow I think(in softcubic50 at least) so the luma looks cleaner(which is even more true considering it's processed in 16bit).
is there a way to get very blurry chroma from the ffdshow avisynth filter? I can't use either mVR/rgb3dlut() at this point as the colors are not identical to realtime dddc() on my set up :o
cyberbeing
5th May 2009, 18:34
I guess it's caused by a bug in the NVidia decoder due to being forced to output YV12, so I don't plan to further look into this right now.
I don't believe this is the issue, as it was just baseless speculation.
When outputting YUY2 from the NVIDIA decoder to FFDshow which is outputting YV12 to madVR, that misaligned luma and chroma problem still happens. This only happens with madVR. Other renderers are fine. This suggests it doesn't have anything to do with the colorspace being output by the NVIDIA decoder.
madshi
5th May 2009, 18:48
I don't believe this is the issue, as it was just baseless speculation.
When outputting YUY2 from the NVIDIA decoder to FFDshow which is outputting YV12 to madVR, that misaligned luma and chroma problem still happens. This only happens with madVR. Other renderers are fine. This suggests it doesn't have anything to do with the colorspace being output by the NVIDIA decoder.
Ok. Is the NVidia decoder freeware and does it work on ATI cards, too?
cyberbeing
5th May 2009, 19:01
Ok. Is the NVidia decoder freeware and does it work on ATI cards, too?
It's not freeware, but there is a 30 day trial here: http://www.nvidia.com/object/dvd_decoder_1.02-223-trial.html
It seems I remember hearing of people using it on ATI cards before, but not owning an ATI card myself currently, I can't confirm. Since you would be using it in software mode with madVR, I really don't see why it wouldn't work.
racerxnet
5th May 2009, 19:10
cyberbeing
Quote:
Originally Posted by madshi View Post
Ok. Is the NVidia decoder freeware and does it work on ATI cards, too?
It's not freeware, but there is a 30 day trial here: http://www.nvidia.com/object/dvd_dec...223-trial.html
It seems I remember hearing of people using it on ATI cards before, but not owning an ATI card myself currently, I can't confirm. Since you would be using it in software mode with madVR, I really don't see why it wouldn't work.
It works fine with ATI cards. I have been using this decoder since it first came out. Yes it will decode in software mode for YV12 output to MadVR. It will also provide DXVA for HD content with the Arcsoft filters. This has been shown in the AVS forum for HTPC use.
MAK
TinTime
5th May 2009, 19:14
When outputting YUY2 from the NVIDIA decoder to FFDshow which is outputting YV12 to madVR, that misaligned luma and chroma problem still happens. This only happens with madVR. Other renderers are fine. This suggests it doesn't have anything to do with the colorspace being output by the NVIDIA decoder.
Have you tried YV12 output from the Nvidia decoder into ffdshow? Does this work with other renderers for HD material?
cyberbeing
5th May 2009, 19:16
Have you tried YV12 output from the Nvidia decoder into ffdshow? Does this work with other renderers for HD material?
Yes it does work for other renderers. madVR is the only renderer with this problem so it must not be handling something correctly.
Hypernova
5th May 2009, 21:23
From what I can see, both EVR and Haali are slightly sharper than your madVR screenshot. That might make a difference in banding, too? Which madVR upscaling filter are you using? You may want to try a sharper one (e.g. Lanczos) to see whether the increases banding or not.
My answer would be no, but I don't really trust my eyes outside of choosing things for myself, so here's the shot.
http://img91.imageshack.us/img91/6246/lanczos8.th.png (http://img91.imageshack.us/img91/6246/lanczos8.png)
On other note, I use Bilinear with madVR (on both) right now because that's the only way I can get smooth playback fullscreen. If I had a choice, I probably go with lanczos4 or one of spline. Mainly because I mainly watch anime, so I prefer it sharp (but not with sharpening). madVR 0.9 certainly improve to the point that I can use Softcubic for SD contents (~35-38ms render time). So not all my hope is lost yet, maybe :)
madshi
5th May 2009, 21:40
My answer would be no, but I don't really trust my eyes outside of choosing things for myself, so here's the shot.
http://img91.imageshack.us/img91/6246/lanczos8.th.png (http://img91.imageshack.us/img91/6246/lanczos8.png)
Thanks! This madVR screenshot is now sharper than the EVR and Haali screenshots, but still madVR shows less banding compared to Haali and much less banding compared to EVR. So your screenshots are a really good example of other renderers showing more banding than madVR does.
Interestingly the colors in your madVR 0.9 screenshot are now similar to EVR. So probably your EVR settings were correct all the time...
6233638
5th May 2009, 21:53
That goes to show that luma scaling algorithms are really a matter of taste. Most of the algorithms have some advantages and some disadvantages. Personally, I like Lanczos4 for its sharpness, but (obviously) I dislike the ringing. I like Mitchell, but actually I like SoftCubic50 even more, which I find very similar to Mitchell, but with less aliasing. So sometimes I'm using Lanczos4 and sometimes SoftCubic50.
I found that SoftCubic50, while initially looking better, actually introduced more artefacts into the picture. There was "blocking" in some scenes. I'll try to get some comparison shots done later.
For chroma I'm using SoftCubic100. Interesting that both of you guys prefer SoftCubic50 for chroma.
I find that SoftCubic 100 smooths chroma out too much, which can make fine coloured details look a bit desaturated around the edges. (you'll see a similar effect if you turn up sharpness too high on your display)
Lanczos and Spline produce very similar results, I think. However, you should compare with the same number of taps. Lanczos3 and Spline36 use 3 taps. Lanczos4 and Spline64 use 4 taps. Lanczos8 uses 8 taps. Personally, I can see some differences between 3 and 4 taps. But I don't really see much of a difference between 4 and 8 taps. Except that 8 taps rings quite a bit more...
What I thought was artefacts being added by Lanczos8 last night, looks like it might actually be double-ringing. So the initial ringing is lessened but it rings twice which makes it look worse.
Here's a few examples taken from the same images. (at 10x)
4, then 8:
http://img213.imageshack.us/img213/8382/1lanc4.th.png (http://img213.imageshack.us/img213/8382/1lanc4.png) http://img93.imageshack.us/img93/9093/1lanc8.th.png (http://img93.imageshack.us/img93/9093/1lanc8.png)
http://img213.imageshack.us/img213/8493/3lanc4.png http://img359.imageshack.us/img359/8295/3lanc8.png
mark0077
5th May 2009, 21:58
Just butting in to say keep up the good work. Great stuff being done in here. madshi did you manage to get a dvd sample yet, my bb has been slowed due to going over my download limits so might be 3-4 days until I can upload a dvd (requested by testuo) for you to test / fix the macrovision errors people get playing dvd's from hdd.
EDIT: With all of the discussion about different algorithms, when madVR is being improved / finalized, if different algorithms are generally agreed to be best (produce closes to original) for upscaling and downscaling, will madVR have a smart setting or some sort of Auto, which can select the best method based on what is being done.
I am always concerned with new work like this, that an auto setting is put in there for those of us who want the benefits, without having to do lots of research on the best settings. Would a type of auto setting be a good idea out of the bag? Cheers.
nijiko
5th May 2009, 22:34
Ok. Is the NVidia decoder freeware and does it work on ATI cards, too?
It works well on ATI but probably no DXVA.
And it's a commerce software. You can learn from here:
http://www.nvidia.com/object/dvd_decoder.html
yesgrey
6th May 2009, 00:01
@yesgrey, your opinion?
Maybe it's a good idea...
From the formulas:
Y' = Kr*R' + Kg*G' + Kb*B'
Pb = B' - Y'
Pr = R' - Y'
Cb = 0.5 + Pb
Cr = 0.5 + Pr
then
Cb = 0.5 + B' - Y'
Cr = 0.5 + R' - Y'
You could try this, considering 2x2 pixels in YV12:
1)Using the Y' values from all 4 pixels, calculate the equivalent Y'eq value using the same method that is used for the chroma downsampling (this could be a problem, because there are more than one method...).
2)Calculate B' = Cb + Y'eq - 0.5 and R' = Cr + Y'eq - 0.5.
3)Upsample B' and R' values instead of Cb and Cr like you are doing now.
4) Calculate the Cb and Cr values using the upsampled B' and R' and the Y' values of each pixel.
This is just a quick thought, so maybe it's just a very dumb idea... ;)
Casshern
6th May 2009, 07:50
You say that as if it was a fact. Have you seen this done? And have you seen that it's actually "much better"?
Doing it that way would mean making brighter pixels more saturated and darker pixels less saturated, right? I don't really see how that would be "more accurate". But I'm not an expert in this area. Is there any "scientific" reason for making brighter pixels more saturated than darker pixels?
As far as I can see, luma and chroma are independent and the luma value does not have any direct influence on chroma. So I don't really see how luma can help upsampling chroma better. But then we're talking about gamma corrected Y'CbCr and not linear light YCbCr. And IIRC I've been told that there is a bit of luminance in Cb and Cr, too. Argh, this is complicated.
@yesgrey, your opinion?
There is an article which handles border cases where chroma is spread to neighbor pixels, if luma is too dark or too bright to hold the upsampled chroma. I've yet to look into implementing a similar algorithm. But the article only handles such corner cases and does not *generally* reshuffle the chroma. Actually the author of that article told me that a friend of his suggested to use luma to form chroma better, but he was not convinced of his friend's efforts...
I am not talking about border cases, nor should luma affect chroma other than in upscaling it - so the colors themselves are not altered. The absolut simplest would be to use a 4 tap horizontal window on luma deltas (L) and 4 tap horizontal actual chroma window (of the simply doubled chroma info - so its a point upsampled chroma channel to the same resolution as luma) (C). Then one could do something like this:
Destination_chroma_for_pixel_p=Sum(Ci*(Sum_up_to_i(L)/Sum_all(L)))
p is in the middle of the tap window.
This ensures that the weights are monoton increasing and normalized in total to 1. Then the chroma info is resampled using these weights.
Naturally this should be extended to the vertikal res as well, and it one could refine it by subsampling on a 0.5 pixel raster as the luma deltas are shifted by 0.5 pixels. Also the linear interpolation can easily be extended to bicubic, lanczos etc. The last thing to think about is how to actually weight the chroma channels: do in rgb light, or UV or normal RGB. But i hope you get the idea...
NOTE: It just occured to me that the designers of YUV were pretty clever in coding chroma as differentials. When upscaling these with anything better than point sampling to 4:4:4 one could get the same results as above much easier. So an alternative pipeline would be to upscale YUV 4:2:2 to YUV 4:4:4 and only then convert/upscale to the target RGB resolution. Or directly upscale the chroma differentials to the target res. This might be easier then the above. So i might just have reinvented the wheel. Any thoughts on that? Other ideas?
madshi
6th May 2009, 08:25
I found that SoftCubic50, while initially looking better, actually introduced more artefacts into the picture. There was "blocking" in some scenes. I'll try to get some comparison shots done later.
I find that SoftCubic 100 smooths chroma out too much, which can make fine coloured details look a bit desaturated around the edges. (you'll see a similar effect if you turn up sharpness too high on your display)
Comparison shots for both problems would be welcome.
madshi did you manage to get a dvd sample yet
No.
With all of the discussion about different algorithms, when madVR is being improved / finalized, if different algorithms are generally agreed to be best (produce closes to original) for upscaling and downscaling, will madVR have a smart setting or some sort of Auto, which can select the best method based on what is being done.
If you read through the few previous posts you'll notice that everybody has his own favorites. There is no "best method". It's a matter of taste.
Maybe it's a good idea...
From the formulas:
Y' = Kr*R' + Kg*G' + Kb*B'
Pb = B' - Y'
Pr = R' - Y'
Cb = 0.5 + Pb
Cr = 0.5 + Pr
then
Cb = 0.5 + B' - Y'
Cr = 0.5 + R' - Y'
You could try this, considering 2x2 pixels in YV12:
1)Using the Y' values from all 4 pixels, calculate the equivalent Y'eq value using the same method that is used for the chroma downsampling (this could be a problem, because there are more than one method...).
2)Calculate B' = Cb + Y'eq + 0.5 and R' = Cr + Y'eq + 0.5.
3)Upsample B' and R' values instead of Cb and Cr like you are doing now.
4) Calculate the Cb and Cr values using the upsampled B' and R' and the Y' values of each pixel.
This is just a quick thought, so maybe it's just a very dumb idea... ;)
That raises a number of questions for me:
(1) I understand Cb and Cr in your math above are still from the gamma corrected Y'CbCr, right? Maybe we should designate them Cb' and Cr', although that's wrong. Or maybe we should find a new name for them?
(2) What is "Pb" and "Pr" and "Kr/g/b"?
(3) You seem to be calculating B' and R' (gamma corrected blue/red), but without using any of the Rec601/709 coefficients. Wouldn't that result in incorrect results?
(4) You totally ignore G'. How can you recalculate Cb and Cr without using G'?
I just don't fully understand what your formulas are doing and why. Could you please explain? Thanks!
I am not talking about border cases, nor should luma affect chroma other than in upscaling it - so the colors themselves are not altered. The absolut simplest would be to use a 4 tap horizontal window on luma deltas (L) and 4 tap horizontal actual chroma window (of the simply doubled chroma info - so its a point upsampled chroma channel to the same resolution as luma) (C). Then one could do something like this:
Destination_chroma_for_pixel_p=Sum(Ci*(Sum_up_to_i(L)/Sum_all(L)))
p is in the middle of the tap window.
This ensures that the weights are monoton increasing and normalized in total to 1. Then the chroma info is resampled using these weights.
Naturally this should be extended to the vertikal res as well, and it one could refine it by subsampling on a 0.5 pixel raster as the luma deltas are shifted by 0.5 pixels. Also the linear interpolation can easily be extended to bicubic, lanczos etc. The last thing to think about is how to actually weight the chroma channels: do in rgb light, or UV or normal RGB. But i hope you get the idea...
I don't fully understand your formula. Maybe I'm being stupid? :o
But more importantly: I still don't see the physical/scientific reasoning for doing what you suggest. Now I don't have free time to spare. So I can't really afford to spend hours and hours on an algorithm without even understanding *why* it should be better in theory. So could you please spend a few lines on explaining why your formula should produce better results?
NOTE: It just occured to me that the designers of YUV were pretty clever in coding chroma as differentials. When upscaling these with anything better than point sampling to 4:4:4 one could get the same results as above much easier.
Well, that is exactly what all the better chroma upsampling algorithms out there (including madVR) are already doing!
Maybe it's a good idea...
From the formulas:
Y' = Kr*R' + Kg*G' + Kb*B'
Pb = B' - Y'
Pr = R' - Y'
Cb = 0.5 + Pb
Cr = 0.5 + Pr
then
Cb = 0.5 + B' - Y'
Cr = 0.5 + R' - Y'
You could try this, considering 2x2 pixels in YV12:
1)Using the Y' values from all 4 pixels, calculate the equivalent Y'eq value using the same method that is used for the chroma downsampling (this could be a problem, because there are more than one method...).
2)Calculate B' = Cb + Y'eq + 0.5 and R' = Cr + Y'eq + 0.5.
3)Upsample B' and R' values instead of Cb and Cr like you are doing now.
4) Calculate the Cb and Cr values using the upsampled B' and R' and the Y' values of each pixel.
This is just a quick thought, so maybe it's just a very dumb idea... ;)
yesgrey3
And if we did like this:
1) See above (downsample Y' --> Y'eq)
2) Transform between Y'eqCbCr --> R'G'B' (1/4 of source pixels)
3) Upsample B', R' and G' values R'G'B' --> R"G"B" (back to initial number of pixels, we could call it upscaling in gamma corrected RGB space)
4) Correct R"G"B" luminance using exact Y' values from the source Rc"=R"+d, Gc"=G"+d, Bc"=B"+d:
Kr*(R"+d) + Kg*(G"+d) + Kb*(B"+d)=Y', d=(Y'-(Kr*R"+Kg*G"+Kb*B"))/(Kr+Kg+Kb)
5) Degamma Rc"Gc"Bc" --> RGB
6) Scale
7) Gamma it RGB-->R'G'B'
mark0077
6th May 2009, 09:56
madshi, I would have to disagree (albeit with my limited knowledge) that there is no "best" method out of a group of scaling techniques for a certain operation. How can this be. Isn't there something wrong with this statement.
If we are trying to create an image thats as close to the original as possible, then (maybe with the use of a camera or some other method) to compare the original image with the various scaled ones and do some sort of comparison on those to see which is closest.
I can't think of a good way of doing this, maybe using a CRT, but to say there isn't a best method I think is wrong. One technique, at least for one specific image would be closer to the original (percieved by our eyes) than other methods.
madshi
6th May 2009, 10:10
@Mark, all the resampling filters are a compromise between aliasing, sharpness and ringing. There is no filter which is best in every category. E.g. Lanczos is the sharpest filter (good), but has the most ringng (bad). Look here:
http://forum.doom9.org/showthread.php?p=1272990#post1272990
If ringing artifacts don't bother you, Lanczos is "best". If you hate ringing artifacts, Lanczos is "worst". So obviously there is no overall "best" filter. You may also want to reread the forum rules, especially rule 12.
mark0077
6th May 2009, 10:19
Well thanks for the reply and link but mathematically one image WILL be closer to the original than another right? I am not breaking rule 12, I am just discussing something....... anyways, one image will be closer to the original than another I think, whether it would have one of the many negative effects you talk about. If part of the aim / point of madVR is to get us the original image in all of its glory, and as accurately as possible, then the whole idea of a users preference goes out the window.
OK fair enough when it comes to something that can't be done perfectly / exactly like resizing users will always have a preference, but I still think strongly that user preference should have as little impact / effect as possible if there is a way to measure how close to the original a scaled image is perceptually (and I'm sure your away of some technique for doing such a thing?). I hope you see what I am saying here.
madshi
6th May 2009, 10:41
Lanczos is nearer to the original, I believe. But it adds artifacts to the image which other filters don't (or do less). So regardless of whether you like it or not, it still is a matter of taste.
This is my last post about this topic for now.
yesgrey
6th May 2009, 13:43
(1) I understand Cb and Cr in your math above are still from the gamma corrected Y'CbCr, right? Maybe we should designate them Cb' and Cr', although that's wrong. Or maybe we should find a new name for them?
Yes, I agree with the Cb' Cr' designations. Soon we should start talking about YCbCr and we need to distinguish between them.;)
(2) What is "Pb" and "Pr" and "Kr/g/b"?
From my understanding:
Kr, Kg and Kb are the contributions or R', G' and B', when all they are equal to 1.0, that gives the white point as Y' = 1.0. (This in [0.0 1.0] range).
Pb' and Pr' (they are not primed, but I will also use the prime symbol to avoid confusion with Pb, Pr when converting from linear RGB) are analog values. Cb' and Cr' are the designations when using them in the digital format.
Pb' and Pr' are defined in the range [-0.5 0.5], so we need to add 0.5 to them to be defined in the range [0.0 1.0] and allow them to be used in the digital format Cb' Cr'.
(3) You seem to be calculating B' and R' (gamma corrected blue/red), but without using any of the Rec601/709 coefficients. Wouldn't that result in incorrect results?
No, because B' and R' are calculated using the coefficients, just not in the way you are used to see them.;)
The Rec601/709 only define the Kr and Kb coefficients. Kg is calculated from Kr+Kg+Kb = 1.
The coefficients that we are used to, are calculated based on the formulas I gave above. Those three formulas can be translated into a matrix of coefficients.
If you want more details on how to get the coefficients, see here (http://en.wikipedia.org/wiki/YCbCr):
(4) You totally ignore G'. How can you recalculate Cb and Cr without using G'?
The only part of G' that is needed is contained in Y'.;)
yesgrey3
And if we did like this:
...
4) Correct R"G"B" luminance using exact Y' values from the source Rc"=R"+d, Gc"=G"+d, Bc"=B"+d:
Kr*(R"+d) + Kg*(G"+d) + Kb*(B"+d)=Y', d=(Y'-(Kr*R"+Kg*G"+Kb*B"))/(Kr+Kg+Kb)
The only problem I see, besides the extra steps, is you considering that R", G" and B" would get the same correction from Y' values. R', G' and B' have different contributions to Y', so R", G" and B" should also be corrected by different values, which could not be computed with the suggestion you made.
I think that my method should be pretty simple to try, because does not require significative changes by madshi. Then, if it gives any improvement, maybe we can think on more complex methods.
madshi
6th May 2009, 14:06
No, because B' and R' are calculated using the coefficients, just not in the way you are used to see them.;)
But your formulas don't list any coefficients! You list Kr, Kg and Kb in the first line and then you just drop them and never use them again. Also your line "Pb = B' - Y'" doesn't match what I read in Wikipedia. There the "Pb" formular is much more complicated. You seem to have simply dropped the whole Kb coefficient! I'm confused... :confused:
Casshern
6th May 2009, 15:10
Well, that is exactly what all the better chroma upsampling algorithms out there (including madVR) are already doing!
Good! I only proposed this because you wrote that chroma and luma upsampling is completely unrelated. If you do it on the chroma differentials (heres the "hidden" connection) everything is fine - and the result should be the same. The only room to improve would be to use a psychovisually more "accurate" definition of Y and the chroma components (maybe through RGB light - then something like i proposed might be necessary again), but i doubt it would make perceptible differences. So everything is good enough for me- keep up the good work!
yesgrey
6th May 2009, 16:30
You list Kr, Kg and Kb in the first line and then you just drop them and never use them again. :confused:
Y' = Kr*R' + Kg*G' + Kb*B'
Pb' = B' - Y'
Pr' = R' - Y'
If you substitute Y' in Pb' and Pr' you'll get:
Pb' = B' - (Kr*R' + Kg*G' + Kb*B')
Pr' = R' - (Kr*R' + Kg*G' + Kb*B')
But since Y' is already known, you can use it directly. The coefficients are used to calculate Pb' and Pr' from R', G'and B'; from Y' it would be a lot simpler.
Also, Rec BT709 defines Kr = 0.2126 and Kb = 0.0722, and Rec BT601 defines Kr = 0.2990 and Kb = 0.1140.
If you want to give it a try, just trust me for now, the formulas are correct. If you prefer to understand it first, I will continue to explain it to you. If you feel this is too off-topic let me know and I will PM you instead...
I only think that the method I explained above should be pretty simple to try, because you already have all the math written; is just using it with different values, nothing more.;)
vBulletin® v3.8.11, Copyright ©2000-2026, vBulletin Solutions Inc.