View Full Version : [New patch] Hadamard motion estimation
Sagekilla
4th September 2007, 18:53
What's your --deadzone settings?
Standard, I didn't modify them at all.
Edit: To be more precise, --deadzone-inter 21 --deadzone-intra 11
Razorholt
4th September 2007, 19:18
So, your settings can help me compete with H.264 and get that sort of results? -> http://images.apple.com/movies/wb/300/300-tlr1b_h480p.mov (I know the sources aren't probably the same but I'm focusing on gains retention here, especially on skins)
Objectively, and post-processing aside, you're saying that x264 and H.264 can produce the same exact results at same bitrates, correct?
I remember Sharktooth making a comment on x264 and grains retention but I can't find the post... :(
Terranigma
4th September 2007, 19:34
What aku was saying, was that x264 is h.264, and that you weren't clear on what h.264 coder or codec you were talking about. Mainconcept? Elecard? Nero Recode? Ateme? Perhaps something else?
Dark Shikari
4th September 2007, 19:39
So, your settings can help me compete with H.264 and get that sort of results? -> http://images.apple.com/movies/wb/300/300-tlr1b_h480p.mov (I know the sources aren't probably the same but I'm focusing on gains retention here, especially on skins)
Objectively, and post-processing aside, you're saying that x264 and H.264 can produce the same exact results at same bitrates, correct?
I remember Sharktooth making a comment on x264 and grains retention but I can't find the post... :(
Here's what you're saying, paraphrased.
So, your engine mods can help me compete with cars and get that sort of results?
Objectively, you're saying that your Honda Civic and cars can reach the same exact speed in the same time, correct?
x264 is an implementation of H.264. :p
Razorholt
4th September 2007, 19:41
What aku was saying, was that x264 is h.264, and that you weren't clear on what h.264 coder or codec you were talking about. Mainconcept? Elecard? Nero Recode? Ateme? Perhaps something else?
Thanks Terranigma for the clarification. I use both Mainconcept and Nero.
Dark Shikari
4th September 2007, 19:42
Thanks Terranigma for the clarification. I use both Mainconcept and Nero.Mainconcept in most tests shows as being roughly tied with x264, though I hope to change that over the next few months with my x264 modifications.
Nero is much more limited and is considerably inferior I believe due to its encoder limitations.
Sharktooth
4th September 2007, 19:43
h.264 is a standard. x264 is an implementation of the h.264 standard.
now, about encoding fine details like grain, try lowering the deadzones (between 3 and 6 for intra and between 10 and 18 for inter).
keep in mind lowering the deadzones settings will require a higher bitrate.
also avoid to overcompress (using insane settings) coz some options may sacrifice fine details for a higher compression (coz metrics and human visual system are 2 completely different things).
Terranigma
4th September 2007, 19:46
Mainconcept in most tests shows as being roughly tied with x264, though I hope to change that over the next few months with my x264 modifications.
I always wondered what settings were used with these tests, because I can't get a quality encoding with mainconcept if I tried. Even using the suggested settings from the adobe doc. :p
I prefer x264 over these other encoders mainly because it's open source and gives you the freedom to control every aspect of the encoder. Take Mainconcept and Elecard for example; it won't let me use more than 3 b-frames. :mad:
Razorholt
4th September 2007, 19:50
Here's what you're saying, paraphrased.
So, your engine mods can help me compete with cars and get that sort of results?
Objectively, you're saying that your Honda Civic and cars can reach the same exact speed in the same time, correct?
x264 is an implementation of H.264. :p
Oookay... Let me correct what I wrote. I was asking whether MeGUI - that I love and cherish - and any other x264-based encoder can match Nero, MainConcept, etc... Is that better, Mister? :p
Dark Shikari
4th September 2007, 19:53
Oookay... Let me correct what I wrote. I was asking whether MeGUI - that I love and cherish - and any other x264-based encoder can match Nero, MainConcept, etc... Is that better, Mister? :p
Yes, I would personally state that in my opinion x264, with the proper settings, is vastly superior to all other H.264 encoders due to either better quality/bitrate, better customizability, or both, except in the following cases:
1. Interlaced encoding. x264 doesn't have full MBAFF/PAFF support.
2. Film Grain Modelling. Some of the fanciest/most expensive encoders, most not available to consumers, have FGM support. This gives a considerable advantage in compressing movies at high bitrates.
Some professional compressionists are on record as stating similar; that x264 outperforms most other commercial encoders.
MeGUI can use any settings you want with x264, so if it doesn't look as good as video compressed by a different H.264 implementation, check your settings first.
Sharktooth
4th September 2007, 19:55
Ehrr... MeGUI is just a GUI... it uses x264 for encoding. So does every GUI and every software that uses x264 for encoding. There are no x264-based encoders except x264 :D
And however, yes, IMHO x264 is as good if not better than other encoders. It's just a matter of how you configure it for encoding.
akupenguin
4th September 2007, 19:59
also avoid to overcompress (using insane settings) coz some options may sacrifice fine details for a higher compression (coz metrics and human visual system are 2 completely different things).
Your intent is correct, but your statement is misleading, so I'll rephrase it:
There is no such thing as "overcompress", except maybe "pick too low of a bitrate", which is unrelated to the current discussion.
You meant "overfit". Any lossy compression has to sacrifice some types of information in favor of other types of information. Optimizing for some metric is better than not optimizing for anything, even if that metric is very approximate. But overfitting to a model of distortion that isn't identical to HVS can cause the encoder to make such tradeoffs in cases that are detrimental to perceived quality.
Terranigma
4th September 2007, 20:00
Razorholt, compare HQ-Schizo (http://forum.doom9.org/showthread.php?p=1040887#post1040887) to whatever encoder you're using and post screenshots. :p
Manao
4th September 2007, 20:06
2. Film Grain Modelling. Some of the fanciest/most expensive encoders, most not available to consumers, have FGM support. This gives a considerable advantage in compressing movies at high bitrates.No. FGM helps at all bitrates, and I'd say it helps more at low bitrates than at high bitrates.
Dark Shikari
4th September 2007, 20:07
No. FGM helps at low bitrates, not at high bitrates.Yes, you're correct. My thought process was:
a) I remove grain when I encode at low bitrates.
b) Therefore, grain is only important at high bitrates.
c) Therefore, FGM is only useful at high bitrates.
But obviously FGM is useful at low bitrates in order to avoid a).
Sharktooth
4th September 2007, 20:10
Your intent is correct, but your statement is misleading, so I'll rephrase it:
There is no such thing as "overcompress", except maybe "pick too low of a bitrate", which is unrelated to the current discussion.
You meant "overfit". Any lossy compression has to sacrifice some types of information in favor of other types of information. Optimizing for some metric is better than not optimizing for anything, even if that metric is very approximate. But overfitting to a model of distortion that isn't identical to HVS can cause the encoder to make such tradeoffs in cases that are detrimental to perceived quality.
it's exactly what i meant but i was never able to explain things in the correct way, even in my native language... so thanks.
Sagekilla
4th September 2007, 20:39
Yes, I would personally state that in my opinion x264, with the proper settings, is vastly superior to all other H.264 encoders due to either better quality/bitrate, better customizability, or both, except in the following cases:
1. Interlaced encoding. x264 doesn't have full MBAFF/PAFF support.
2. Film Grain Modelling. Some of the fanciest/most expensive encoders, most not available to consumers, have FGM support. This gives a considerable advantage in compressing movies at high bitrates.
Some professional compressionists are on record as stating similar; that x264 outperforms most other commercial encoders.
MeGUI can use any settings you want with x264, so if it doesn't look as good as video compressed by a different H.264 implementation, check your settings first.
Yup, agreed on that point that x264 is superior to other encoders. Like Akupenguin said, he's giving us enough rope to hang ourselves with x264. With that said, with proper settings x264 blows away other consumer available codecs. You don't necessarily have to use insane settings like mine, I just do that so I can get the lowest possible file size short of enabling ESA. I've personally used only one other H.264 based encoder, namely Nero Digital, and I didn't quite like it's results.. The videos looked mushy considering the bitrate I was using and the settings I used (both pretty high)
Edit: Speaking of your first point, who even uses interlacing anymore? Unless you've somehow found interlaced content that will not deinterlace properly I see no point in even using the setting. It's an old technology from the early days of analog broadcasting that shouldn't have a place in today's progressive based LCD/Plasma/DLP/whatever displays.
fields_g
5th September 2007, 12:09
This might show how little I know about FGM, but if you encode at 320x240, then play it scaled larger (for example 960x720) would the grain pixels be at the scaled playback resolution?
If it is the playback resolution, wouldn't it be an argument for FGM potentially being benificial for lower resolution encodes also?
foxyshadis
5th September 2007, 12:28
FGM is normally implemented in the decoder, so unless the decoder scales on output, it'd be same as the source. (I don't know of any that do, although it's a good reason to make one.) If FGM was implemented in mplayer/ffdshow or even a renderer, then it's possible.
akupenguin
5th September 2007, 12:32
This might show how little I know about FGM, but if you encode at 320x240, then play it scaled larger (for example 960x720) would the grain pixels be at the scaled playback resolution?
The standard only specifies a syntax and semantics for describing the grain. So the player is allowed to upscale before reconstructing it. That said, the grain syntax doesn't allow for features smaller than 1 pixel, so your upscaling player would have to either extrapolate the high frequencies or somehow know that the grain description is supposed to apply to a higher resolution than was actually encoded. Maybe SVC can signal that.
akupenguin
9th September 2007, 06:49
Optimization idea for SATD ESA (possibly also UMH, but I'm not sure):
SATD(enc,ref) = sum(abs(hadamard(diff(enc,ref)))) = sub(abs(diff(hadamard(enc),hadamard(ref))).
Despite the two hadamards, that actually decreases the amount of computation. One of the hadamards is of the input block and so can be done only once per search. The other can be reused, since each 4x4 hadamard block is shared among (16 mvs offset by multiples of 4 pixels) * (several block sizes). This sharing is possible only after the above factoring, because that's what causes the different hadamards to have the same inputs.
There's also some sharing within the computation of hadamard(ref), e.g. run 1 row transform and then 4 column transforms at 1 pixel offsets.
And after reducing SATD to prefilter+SAD, lossless SEA should work. Though I'm not sure whether SEA will be able to eliminate many mvs, since it depends on certain statistics of the image, and the hadamard filtered image will be different from a natural image.
The disadvantage is that it takes lots of memory: 32 bytes per pixel per reference frame, as compared to 4 for SAD SEA, and 0 for DIA/HEX/UMH.
Dark Shikari
9th September 2007, 07:07
Optimization idea for SATD ESA (possibly also UMH, but I'm not sure):
SATD(enc,ref) = sum(abs(hadamard(diff(enc,ref)))) = sub(abs(diff(hadamard(enc),hadamard(ref))).
Despite the two hadamards, that actually decreases the amount of computation. One of the hadamards is of the input block and so can be done only once per search. The other can be reused, since each 4x4 hadamard block is shared among (16 mvs offset by multiples of 4 pixels) * (several block sizes). This sharing is possible only after the above factoring, because that's what causes the different hadamards to have the same inputs.
There's also some sharing within the computation of hadamard(ref), e.g. run 1 row transform and then 4 column transforms at 1 pixel offsets.
And after reducing SATD to prefilter+SAD, lossless SEA should work. Though I'm not sure whether SEA will be able to eliminate many mvs, since it depends on certain statistics of the image, and the hadamard filtered image will be different from a natural image.
The disadvantage is that it takes lots of memory: 32 bytes per pixel per reference frame, as compared to 4 for SAD SEA, and 0 for DIA/HEX/UMH.32 bytes per pixel per reference frame?
That's an entire gigabyte of memory for a 1080p clip encoded with 16 reference frames... :eek:
Your method would be lossless, but are you sure it outperforms my current method, which albeit not lossless appears "mostly lossless" in most cases (its not in the patch in the original post though)?
I would think any method that requires so much memory is infeasible and impractical.
akupenguin
9th September 2007, 07:15
You don't strictly need that much memory, but without it you can only reuse results within one search (or with more complexity, across block sizes within one ref). In that restricted case, it only reduces hadamards by a factor of ~25 compared to brute force. Does your lossy method eliminate 96% of the mvs?
Plus, who'd run 1080p 16ref SATD ESA on a wimpy computer?
Dark Shikari
9th September 2007, 07:32
You don't strictly need that much memory, but without it you can only reuse results within one search (or with more complexity, across block sizes within one ref). In that restricted case, it only reduces hadamards by a factor of ~25 compared to brute force. Does your lossy method eliminate 96% of the mvs?
Plus, who'd run 1080p 16ref SATD ESA on a wimpy computer?
A factor of 25?
Isn't that a bit of an overestimation, one would think? SAD ESA doesn't do nearly that much, does it?
My method does a full SAD ESA (not SEA) and then does SATD on, I'm guessing, about 1/10 - 1/5 of those.
akupenguin
9th September 2007, 07:52
SAD SEA runs the actual SAD on between 1/4 and 1/8 of the mvs (assuming merange=16).
But SAD SEA's efficiency has a different basis: how well you can estimate SAD scores with a faster metric. My proposed SATD ESA is based on redundant computations, not estimation.
Ok, so the 25 is for 16x16 partitions (whether or not you use the lots of memory). If you do use memory then it completely eliminates hadamards from smaller partitions. If you don't use any memory then it reduces hadamards by a factor of 13 in 16x8 partitions and 6 in 8x8 partitions. Either way, there's still a SAD ESA (to add up the results of the hadamards, not like your threshold).
Dark Shikari
9th September 2007, 08:03
SAD SEA runs the actual SAD on between 1/4 and 1/8 of the mvs (assuming merange=16).
But SAD SEA's efficiency has a different basis: how well you can estimate SAD scores with a faster metric. My proposed SATD ESA is based on redundant computations, not estimation.
Ok, so the 25 is for 16x16 partitions (whether or not you use the lots of memory). If you do use memory then it completely eliminates hadamards from smaller partitions. If you don't use any memory then it reduces hadamards by a factor of 13 in 16x8 partitions and 6 in 8x8 partitions. Either way, there's still a SAD ESA (to add up the results of the hadamards, not like your threshold).It sounds like it would be a bit more efficient, and truly lossless as compared to a normal SATD ESA, but on the other hand it would require a lot more coding to implement, especially given the different behavior required for each type of partition.
akupenguin
9th September 2007, 15:13
especially given the different behavior required for each type of partition.
No difference in behavior. With memory, the hadamard computations aren't actually done during ME, they're a prefilter like SEA's integral image, and my numbers are the amortized cost. Without memory, the difference in speedup factors is because the different partitions sizes have to filter the same area. i.e. The behavior differs now, and it won't after the proposed change.
DeathTheSheep
9th September 2007, 15:28
Plus, who'd run 1080p 16ref SATD ESA on a wimpy computer?
Very true, I was thinking the same thing. :p 16 refs is as slow as molasses as is (say that 6 times fast).
Sagekilla
9th September 2007, 16:31
Very true, I was thinking the same thing. :p 16 refs is as slow as molasses as is (say that 6 times fast).
That that that that that that! (Sorry, couldn't help it :p) Anyway, who'd run a 16 ref 1080p encode to begin with? Unless you have some dual socket kentsfield with at least 4 GB of RAM I doubt you'd be doing that. Besides, 4-6 refs sounds more realistic for that kind of encode.
Dark Shikari
9th September 2007, 17:56
Apparently the UMH mode might need some slight tweaking; its still using the same early termination, which is designed around SAD. Turning it off results in a catastrophic FPS drop, which suggests that its early terminating, well, a whole lot of the time more than with SAD.
I'm going to do some testing to see if it should be modified or not.
akupenguin
10th September 2007, 03:58
x264_satd_fpel.05.diff (http://akuvian.org/src/x264/x264_satd_fpel.05.diff) Your algorithm. Mostly cosmetic changes, but I did get a 8% speedup just by changing sadarray[] from column major to row major order.
x264_satd_fpel.06.diff (http://akuvian.org/src/x264/x264_satd_fpel.05.diff) Reusing hadamards. Faster than brute-force, but slower that your threshold.
It successfully eliminates almost all of the hadamards: from 20% (yours) to 4% of the cpu time. And the 20% is with SSSE3 while the 4% is plain C. However, 16bit SAD is slower than than 8bit SAD, so just "reducing motion search to SAD ESA" isn't enough.
Dark Shikari
10th September 2007, 04:09
x264_satd_fpel.05.diff (http://akuvian.org/src/x264/x264_satd_fpel.05.diff) Your algorithm. Mostly cosmetic changes, but I did get a 8% speedup just by changing sadarray[] from column major to row major order.
x264_satd_fpel.06.diff (http://akuvian.org/src/x264/x264_satd_fpel.05.diff) Reusing hadamards. Faster than brute-force, but slower that your threshold.
It successfully eliminates almost all of the hadamards: from 20% (yours) to 4% of the cpu time. And the 20% is with SSSE3 while the 4% is plain C. However, 16bit SAD is slower than than 8bit SAD, so just "reducing motion search to SAD ESA" isn't enough.Ah, so the SAD needs to be 16-bit because the SATD scores are higher than 255.
That's quite a patch; how much slower is it than my threshold? If its not much slower it would be preferable as it is indeed lossless, while my threshold (according to some reports) can fail to provide good results in some cases.
akupenguin
10th September 2007, 04:12
brute force: 1.03 fps
reuse: 1.76 fps
threshold: 2.60 fps
Dark Shikari
10th September 2007, 04:13
brute force: 1.03 fps
reuse: 1.76 fps
threshold: 2.60 fpsWhat happens if you take the SATD scores and scale them to fit in 0-255 (and, say, cut off the top 1% to avoid dealing with extremely high maximums)?
Would the precision loss be bad enough to decrease quality? And how much of a speed boost would it give?
Additionally, is there any way to combine my thresholding with your optimization?
akupenguin
10th September 2007, 04:21
Hadamard DC coefs are in the range 0-4080. And they really use that whole range, it's not just outliers. So you could downscale everything by a factor of 16, and it would become just an 8bit SAD. But I'm sure that will lose lots of precision. Or you could keep DC coefs at full precision in a separate array, and scale/clip AC which uses a much smaller typical range. Might be ok quality, but more complicated.
Dark Shikari
10th September 2007, 04:25
Hadamard DC coefs are in the range 0-4080. And they really use that whole range, it's not just outliers. So you could downscale everything by a factor of 16, and it would become just an 8bit SAD. But I'm sure that will lose lots of precision. Or you could keep DC coefs at full precision in a separate array, and scale/clip AC which uses a much smaller typical range. Might be ok quality, but more complicated.What exactly is a DC or AC coefficient? I have never found an explanation for these terms.
akupenguin
10th September 2007, 04:31
By analogy to Direct Current / Alternating Current. In any frequency-based transform, such as FFT, DCT, or Hadamard, the DC coefficient represents the average of the input window, and all the other coefficients are AC and represent differences between samples.
akupenguin
10th September 2007, 04:54
x264_satd_fpel.07.diff (http://akuvian.org/src/x264/x264_satd_fpel.07.diff) Threshold. Another 13% speedup because you weren't using sad_x3.
Dark Shikari
10th September 2007, 05:05
x264_satd_fpel.07.diff (http://akuvian.org/src/x264/x264_satd_fpel.07.diff) Threshold. Another 13% speedup because you weren't using sad_x3.SAD_X3 has that much of a speed increase? What about SAD_X4?
akupenguin
10th September 2007, 05:32
x4 is slower that x3 here, because typical meranges result in 4n+1 columns, so 3 sads are wasted if you do multiples of 4.
There shouldn't be any significant difference in speed per sad between x3 and x4, the two versions are just to allow for whatever number is convenient.
DeathTheSheep
10th September 2007, 17:09
How would you convert from the old threshold format to the new one?
"average-3*(average-min)/4" -> "(average+3*min)>>2"
Let's say I wanted the old one to be average-1*(average-min)/3; how would I translate this for the new one?
What does ">>" do anyway?
akupenguin
10th September 2007, 17:11
average-3*(average-min)/4 = average-average*3/4+min*3/4 = average*1/4+min*3/4 = (average+3*min)/4 = (average+3*min)>>2
Dark Shikari
10th September 2007, 17:14
How would you convert from the old threshold format to the new one?
"average-3*(average-min)/4" -> "(average+3*min)>>2"
Let's say I wanted the old one to be average-1*(average-min)/3; how would I translate this for the new one?
What does ">>" do anyway?
>> is just a bitshift; personally I think its pointless to say ">>2" instead of "/4" because its less clear to the reader yet both result in the exact same code due to compiler optimization.
akupenguin
10th September 2007, 17:23
The compiler can only optimize /4 into >>2 for unsigned values, because division of negative numbers has different rounding.
DeathTheSheep
10th September 2007, 17:26
I see. What is the bitshift to result in an equivalent of /10, or /5, for instance?
Dark Shikari
10th September 2007, 18:08
I see. What is the bitshift to result in an equivalent of /10, or /5, for instance?
That is vastly more complicated; powers of 2 are easy because its binary.
If I recall correctly, to divide by 5, you do something like (x * 0x66666667) >> 1.
DeathTheSheep
10th September 2007, 18:18
Gotcha. Yeah, I'll stick with division. :)
[edit] Never mind, definitely a compiler thing...
Dark Shikari
10th September 2007, 18:32
Gotcha. Yeah, I'll stick with division. :)The problem with division being that it can take upwards of 40-80 processor cycles depending on the CPU :p
DeathTheSheep
10th September 2007, 18:40
Ouch!!
Heck, according to my compiler, "(average+3*min)/10" != "average-3*(average-min)/10"!!
Apparently, I fail at algebra. :p
akupenguin
10th September 2007, 18:43
Heck, according to my compiler, "(average+3*min)/10" != "average-3*(average-min)/10"!!
You fail at algebra.
vBulletin® v3.8.11, Copyright ©2000-2026, vBulletin Solutions Inc.