Log in

View Full Version : My most exhaustive study of rescaling filters yet


Katie Boundary
5th October 2022, 07:19
The test: 201 frames of the movie "Splice", resized from 720x480 to 512x288 and then back up to 720x480. Twenty resizers were tested against each other, though not every possible combination of upscaler and downscaler was tested, and some resizers were only tested as upscalers or as downscalers, not both (for example, Arearesize was not tested as an upscaler, and Spline144 was not tested as a downscaler). PSNR averages were written down.

The results were fascinating.

https://imgur.com/vAc8957

(downscalers are listed at left; upscalers run across the top)

To put things simply, a clear dichotomy between "good" resizers and "bad" ones emerged. Whenever a "good" downscaler was used, the "good" upscalers formed a clear hierarchy from best to worst regardless of which "good" downscaler was used, and whenever a "good" upscaler was used, the "good" downscalers formed a clear hierarchy from best to worst regardless of which "good" upscaler was used. Furthermore, this hierarchy was nearly the same for downscalers as it was for upscalers. The "good" resizers began, roughly, at Catrom/Spline16/Lanczos2. This gap was the first of two major jumps in performance. The other jump was going from the 2-tap resizers to the 3-tap ones. Beyond Spline36/Lanczos3, performance differences were minimal; the PSNR difference between 2-way Catrom and 2-way Lanczos3 was almost three times as much as the difference between 2-way Lanczos3 and 2-way Lanczos6.

Arearesize, Linear, Hermite, and Mitchell-Netravali formed the "bad" downscalers, along with two experimental guest cubics designed for extreme sharpness at the expense of mathematical correctness: "b0,c1" (no-op but doesn't preserve linear gradients) and "b-1,c1" (preserves linear gradients but sharpens even when no resize is performed). There was no clear hierarchy here; for example, Hermite was a much better downscaler than Linear or Mitchell-Netravali, but a worse upscaler than either of them. What was most interesting was the love affair between the blurry filters and the c1 cubics. Linear, Hermite, and M-N all performed best as downscalers when combined with b-1,c1 as an upscaler, and vice versa, while arearesize achieved its best performance when combined with Spline144 and its second-best with b-1,c1. Many of these combinations even outperformed 2-way Catrom and/or Spline16 by narrow margins! The b0 cubic was a little more normal and slutty, achieving its best performance when combined with the "good" filters, and the blurry filters likewise did better when combined with the "best" of the "good" filters than they did with b0. However, b0 still made a better upscaler for them than Spline16 or Catrom did.

A major anomaly here was Spline144. When "bad" filters were used for downscaling, Spline144 crushed most of the competition; when paired with Arearesize or Linear, it crushed all other upscalers. However, when a "good" downscaler was used, Spline144's performance fell into the toilet next to Mitchell-Netravali. Its behavior was very similar to the c1 cubics in this way. I suspect that at high tap numbers, the Spline family increasingly favors sharpness over mathematical correctness, or perhaps something similar to Runge's Phenomenon starts to set in. I don't know. Either way, the moral of the story here is that I don't think anyone's time or brain cells should be invested into making Spline192 or Spline256 happen. On the subject of the Spline filters, Lanczos delivers better PSNR numbers than Spline at the same number of taps.

Blackman was tested, but it did very poorly from a PSNR-to-taps perspective. Blackman4 did worse than Lanczos3, and Blackman6 did worse than Lanczos5.

So, what does all this mean? First of all, whether you're upscaling or downscaling, the ideal number of taps is three. Second, the spline resizers are janky and probably shouldn't be touched. Third, AVIsynth doesn't have any 3-tap (or better) filters other than Lanczos, Blackman, and Spline36. So... remember all those people who just told us to STFU and use Lanczos3 for everything instead of giving us the "well it depends on your scaling factor and image content and mercury retrograde and blah blah" speech? They were right all along.

DTL
5th October 2022, 22:40
Testing resamplers better to do for some defined digital movie imaging system with at least fixed 2 main transforms:
1. How to encode input continuous 2D image in digital form.
2. How to restore digitally encoded image in continious 2D image form.
Current widely used industry digital movie imaging systems looks like do not have any standards on any of these 2 base conversions. So when you trying to use AVS to scale some content not known where to come from and where to send to it is typically only way to select best resampler is to control result by own eyes on some local display device (executing conversion 2 in some way).
Practically you test lots of different resamplers designed to work in different digital workflows in some mixed way. The target practical use case is not described clearly - it may be input with some HD movie, making low res rip and displaying it on some hardware+software setup with higher resolution using best selected upscaler by the result test table ?

Katie Boundary
5th October 2022, 22:54
FOLLOW-UP:

I finally figured out how to modify the Blackman window to generate windowed sinc filters in Desmos graphing calculator so I could look at the filter shapes. What I saw explained a lot. The windowing function is extremely aggressive, so for just about any tap value higher than 2, Blackman's final lobe is practically nonexistent. Blackman-4 has a nearly identical filter shape to Lanczos-3, and Blackman-3 is nearly identical to Catrom. The first lobe of Blackman-2 is nearly identical to Hermite; however, the second lobe is noticeable. It's like a cubic with a b-value of 0 and a c-value of 0.1 or 0.2

Since filters with more taps/lobes are more computationally expensive, this all means that Blackman is basically a huge waste of CPU cycles.


FOLLOW-UP 2:

It just occurred to me that I could eliminate downscaling from the equation entirely by using the resamplers as bob-deinterlacers. Running a new round of tests, and on more varied footage (every 50th frame of the first 6000 frames of the first episode of Enterprise), I got a different hierarchy of resamplers, but the overall picture was the same. Once again, there was the One-Tap Toilet, where Linear, Hermite, Lanczos-1, etc. hang out; there were the Catrom Clones (including Blackman-3); and finally there was the 3-tap plateau past which further taps just waste CPU cycles. Blackman-2 and Mitchell-Netravali once again delivered results more similar to 1-tap filters than to the rest of the 2-taps. Unsurprisingly, Spline100 once again performed horribly (between Lanczos-1 and Linear) and Spline144 couldn't do better than the Catrom clones, confirming my suspicion that something is seriously wrong with Spline100/144.

DTL
6th October 2022, 07:44
" new round of tests, and on more varied footage (every 50th frame of the first 6000 frames of the first episode of Enterprise"

It is even worse for testing of moving pictures resamplers. Additional requirements for moving pictures resamplers is to keep object's view in a sequence of frames as equal as possible if objects are only translates relative to sampling grid.
For static pictures resamplers like photo it is much less important. Typical issues of static pictures resamples on moving pictures sequencies - flickering of edges of objects or flickering of small objects if the speed of movement is not integer of sampling step. It may be result of aliasing effects that is typically not very good fixed on static pictures resamplers to make pictures more sharp in low pixels size (typical for photo or PC digital images, for sampled natural moving pictures I prefer to use word 'samples'). If you try to see the result of downscaling of some resolution test pattern you will found most of your tested resamplers left more or less non-fixed aliasing.

Better testing on a varied footage is to select small parts of different scenes of a long movie as continious sequencies of frames. At least a sets of 10..100 frames in length.

Katie Boundary
7th October 2022, 18:17
Additional requirements for moving pictures resamplers is to keep object's view in a sequence of frames as equal as possible if objects are only translates relative to sampling grid.
For static pictures resamplers like photo it is much less important. Typical issues of static pictures resamples on moving pictures sequencies - flickering of edges of objects or flickering of small objects if the speed of movement is not integer of sampling step.

You don't seem to understand. I'm NOT testing their suitability as bob-deinterlacers. I'm ONLY testing their properties as resamplers. Using them as bobbers was just a means to that end.

poisondeathray
7th October 2022, 18:51
So... remember all those people who just told us to STFU and use Lanczos3 for everything instead of giving us the "well it depends on your scaling factor and image content and mercury retrograde and blah blah" speech? They were right all along.

How did you come to that conclusion when you didn't test other scaling factors , or types of image content ? Or did you test them, but not post the results ?



DTL is right about the video vs. still image discussion, it's easy to see the failings of traditional resamplers in terms of temporal aliasing when you test larger scaling factors with video and you compare with newer generation upscalers such as machine learning

Katie Boundary
8th October 2022, 01:42
it's easy to see the failings of traditional resamplers in terms of temporal aliasing

Temporal aliasing is a product of fast motion and low framerates, not resizing, and it cannot be measured by AVIsynth's Compare() function, which is what I was using for these tests.

when you test larger scaling factors with video

While I have no doubt that some of the minor details of my findings would change with scale factor (for example, I expect Arearesize to become more accurate at extreme downscales), the overall findings fit into a pattern that shouldn't be scale-dependent.

I can do it if you want, I just don't expect the findings to be useful.

and you compare with newer generation upscalers such as machine learning

I'm not aware of any machine learning resizers for AVIsynth that work at arbitrary scale factors.

DTL
8th October 2022, 05:46
Also the PSNR metric may be not best even for static image quality compare - https://videoprocessing.ai/metrics/ways-of-cheating-on-popular-objective-metrics.html . If the researcher have more time it may be checked how different metrics work with tested resamplers.

As noted in the linked research - PSNR and SSIM are highly sensitive to rotations, spatial shifts and scalings. The resamplers may easy add some spatial (phase) shift mostly invisible at real image viewing but may significantly affect PSNR result. Typically developer of resampler may not check its software for small shifting.

Katie Boundary
8th October 2022, 23:20
Also the PSNR metric may be not best even for static image quality compare - https://videoprocessing.ai/metrics/ways-of-cheating-on-popular-objective-metrics.html . If the researcher have more time it may be checked how different metrics work with tested resamplers.

PSNR is the only one reported by compare() so that's what I had to use.

As noted in the linked research - PSNR and SSIM are highly sensitive to rotations, spatial shifts and scalings. The resamplers may easy add some spatial (phase) shift mostly invisible at real image viewing but may significantly affect PSNR result. Typically developer of resampler may not check its software for small shifting.

Simpleresize has a known scaling issue, but it was not used in any of these tests.

AVIsynth's built-in resizers are known to shift chroma planes left or right when resizing in YUV colorspaces, but I avoided that problem by converting to RGB first.

No rotations were performed.

DTL
9th October 2022, 18:57
You can check if SSIM plugin is work now http://avisynth.nl/index.php/SSIM . And compare it with PSNR metric. It looks exist as x64 build - https://forum.doom9.org/showthread.php?t=175211