Katie Boundary
5th October 2022, 07:19
The test: 201 frames of the movie "Splice", resized from 720x480 to 512x288 and then back up to 720x480. Twenty resizers were tested against each other, though not every possible combination of upscaler and downscaler was tested, and some resizers were only tested as upscalers or as downscalers, not both (for example, Arearesize was not tested as an upscaler, and Spline144 was not tested as a downscaler). PSNR averages were written down.
The results were fascinating.
https://imgur.com/vAc8957
(downscalers are listed at left; upscalers run across the top)
To put things simply, a clear dichotomy between "good" resizers and "bad" ones emerged. Whenever a "good" downscaler was used, the "good" upscalers formed a clear hierarchy from best to worst regardless of which "good" downscaler was used, and whenever a "good" upscaler was used, the "good" downscalers formed a clear hierarchy from best to worst regardless of which "good" upscaler was used. Furthermore, this hierarchy was nearly the same for downscalers as it was for upscalers. The "good" resizers began, roughly, at Catrom/Spline16/Lanczos2. This gap was the first of two major jumps in performance. The other jump was going from the 2-tap resizers to the 3-tap ones. Beyond Spline36/Lanczos3, performance differences were minimal; the PSNR difference between 2-way Catrom and 2-way Lanczos3 was almost three times as much as the difference between 2-way Lanczos3 and 2-way Lanczos6.
Arearesize, Linear, Hermite, and Mitchell-Netravali formed the "bad" downscalers, along with two experimental guest cubics designed for extreme sharpness at the expense of mathematical correctness: "b0,c1" (no-op but doesn't preserve linear gradients) and "b-1,c1" (preserves linear gradients but sharpens even when no resize is performed). There was no clear hierarchy here; for example, Hermite was a much better downscaler than Linear or Mitchell-Netravali, but a worse upscaler than either of them. What was most interesting was the love affair between the blurry filters and the c1 cubics. Linear, Hermite, and M-N all performed best as downscalers when combined with b-1,c1 as an upscaler, and vice versa, while arearesize achieved its best performance when combined with Spline144 and its second-best with b-1,c1. Many of these combinations even outperformed 2-way Catrom and/or Spline16 by narrow margins! The b0 cubic was a little more normal and slutty, achieving its best performance when combined with the "good" filters, and the blurry filters likewise did better when combined with the "best" of the "good" filters than they did with b0. However, b0 still made a better upscaler for them than Spline16 or Catrom did.
A major anomaly here was Spline144. When "bad" filters were used for downscaling, Spline144 crushed most of the competition; when paired with Arearesize or Linear, it crushed all other upscalers. However, when a "good" downscaler was used, Spline144's performance fell into the toilet next to Mitchell-Netravali. Its behavior was very similar to the c1 cubics in this way. I suspect that at high tap numbers, the Spline family increasingly favors sharpness over mathematical correctness, or perhaps something similar to Runge's Phenomenon starts to set in. I don't know. Either way, the moral of the story here is that I don't think anyone's time or brain cells should be invested into making Spline192 or Spline256 happen. On the subject of the Spline filters, Lanczos delivers better PSNR numbers than Spline at the same number of taps.
Blackman was tested, but it did very poorly from a PSNR-to-taps perspective. Blackman4 did worse than Lanczos3, and Blackman6 did worse than Lanczos5.
So, what does all this mean? First of all, whether you're upscaling or downscaling, the ideal number of taps is three. Second, the spline resizers are janky and probably shouldn't be touched. Third, AVIsynth doesn't have any 3-tap (or better) filters other than Lanczos, Blackman, and Spline36. So... remember all those people who just told us to STFU and use Lanczos3 for everything instead of giving us the "well it depends on your scaling factor and image content and mercury retrograde and blah blah" speech? They were right all along.
The results were fascinating.
https://imgur.com/vAc8957
(downscalers are listed at left; upscalers run across the top)
To put things simply, a clear dichotomy between "good" resizers and "bad" ones emerged. Whenever a "good" downscaler was used, the "good" upscalers formed a clear hierarchy from best to worst regardless of which "good" downscaler was used, and whenever a "good" upscaler was used, the "good" downscalers formed a clear hierarchy from best to worst regardless of which "good" upscaler was used. Furthermore, this hierarchy was nearly the same for downscalers as it was for upscalers. The "good" resizers began, roughly, at Catrom/Spline16/Lanczos2. This gap was the first of two major jumps in performance. The other jump was going from the 2-tap resizers to the 3-tap ones. Beyond Spline36/Lanczos3, performance differences were minimal; the PSNR difference between 2-way Catrom and 2-way Lanczos3 was almost three times as much as the difference between 2-way Lanczos3 and 2-way Lanczos6.
Arearesize, Linear, Hermite, and Mitchell-Netravali formed the "bad" downscalers, along with two experimental guest cubics designed for extreme sharpness at the expense of mathematical correctness: "b0,c1" (no-op but doesn't preserve linear gradients) and "b-1,c1" (preserves linear gradients but sharpens even when no resize is performed). There was no clear hierarchy here; for example, Hermite was a much better downscaler than Linear or Mitchell-Netravali, but a worse upscaler than either of them. What was most interesting was the love affair between the blurry filters and the c1 cubics. Linear, Hermite, and M-N all performed best as downscalers when combined with b-1,c1 as an upscaler, and vice versa, while arearesize achieved its best performance when combined with Spline144 and its second-best with b-1,c1. Many of these combinations even outperformed 2-way Catrom and/or Spline16 by narrow margins! The b0 cubic was a little more normal and slutty, achieving its best performance when combined with the "good" filters, and the blurry filters likewise did better when combined with the "best" of the "good" filters than they did with b0. However, b0 still made a better upscaler for them than Spline16 or Catrom did.
A major anomaly here was Spline144. When "bad" filters were used for downscaling, Spline144 crushed most of the competition; when paired with Arearesize or Linear, it crushed all other upscalers. However, when a "good" downscaler was used, Spline144's performance fell into the toilet next to Mitchell-Netravali. Its behavior was very similar to the c1 cubics in this way. I suspect that at high tap numbers, the Spline family increasingly favors sharpness over mathematical correctness, or perhaps something similar to Runge's Phenomenon starts to set in. I don't know. Either way, the moral of the story here is that I don't think anyone's time or brain cells should be invested into making Spline192 or Spline256 happen. On the subject of the Spline filters, Lanczos delivers better PSNR numbers than Spline at the same number of taps.
Blackman was tested, but it did very poorly from a PSNR-to-taps perspective. Blackman4 did worse than Lanczos3, and Blackman6 did worse than Lanczos5.
So, what does all this mean? First of all, whether you're upscaling or downscaling, the ideal number of taps is three. Second, the spline resizers are janky and probably shouldn't be touched. Third, AVIsynth doesn't have any 3-tap (or better) filters other than Lanczos, Blackman, and Spline36. So... remember all those people who just told us to STFU and use Lanczos3 for everything instead of giving us the "well it depends on your scaling factor and image content and mercury retrograde and blah blah" speech? They were right all along.