Log in

View Full Version : XviD Options and their effect on size and encoding time (XviD v.1.10b2)


Harley Quin
14th October 2005, 00:41
This test shows how different XviD settings change the size (in MB) and the encoding time (in min) of the resulting video. All tests are done as single pass constant quantizer. I used two DVD-sources with very different compressibility. Both clips are resized to 688x288 and are 21000 frames (=14min) long. Multiply the size by 10 and you get the framerate (in kBit/sec).

XviD-Settings (if not stated otherwise):
Slightly altered defaults, red marks the changes, "+" means checked, "-" means unchecked.
Main: Single Pass; Quantizer:2
Levels: unrestricted; Quantization Type:MPEG; -AdapQuant; -IntEnc; -QPel; -GMC; +B-VOPs (1/1,5/1,0); +PackedBS
Zones: +Weight:1; -K -G -O -C; BVOPsens:0;
Advanced: MSP:6; VHQ:1; -VHQforB; +ChromaM; -Turbo; FDR:0; MaxI:250

The results give a general guideline, but
- sizes and relations will vary with other sources;
- most tests are MPEG quantization at q2, the results may or may not be adaptable to higher quantizer encodings, h.263 or CQMs;
- only one setting is changed at a time, changing more settings at once may not give the equivalent result.

I got some interesting information out of it. Enjoy.
Quantization Type and Quantizers

Clip 1 Clip 2
h.263 MPEG h.263 MPEG
size % time size % time size % time size % time
q1 520,4 (284%) 29 495,4 (241%) 29 q1 235,2 (360%) 24 214,8 (282%) 23
q2 183,4 (100%) 28 205,8 (100%) 28 q2 65,4 (100%) 23 76,2 (100%) 23
q3 120,8 ( 66%) 28 124,0 ( 60%) 28 q3 44,1 ( 67%) 22 46,1 ( 61%) 22
q4 90,4 ( 49%) 27 94,4 ( 46%) 27 q4 33,6 ( 51%) 22 35,4 ( 46%) 21
q5 73,2 ( 40%) 27 76,7 ( 37%) 26 q5 27,6 ( 42%) 21 28,7 ( 38%) 21
q6 60,7 ( 33%) 26 64,4 ( 31%) 26 q6 23,0 ( 35%) 21 24,0 ( 32%) 21
q7 52,9 ( 29%) 26 55,0 ( 27%) 26 q7 20,4 ( 31%) 21 21,1 ( 28%) 21
q8 46,5 ( 25%) 26 48,7 ( 24%) 26 q8 18,1 ( 28%) 21 18,7 ( 25%) 21


Profile Options Clip 1 Clip 2

Consecutive B-frames
no B-frames (-/---/-) 252,1 29 100,3 23
1 consec. B (1/1,5/1) 205,8 28 76,2 23
2 consec. B (2/1,5/1) 197,922 28 68,917 22
3 consec. B (3/1,5/1) 197,908 28 68,921 22
4 consec. B (4/1,5/1) 197,945 28 68,900 22

B-frames Quantizer
1 consecutive B-frame
P=2, no B (-/---/-) 252,1 29 100,3 23
P=2, B=2 (1/1,0/0) 251,7 28 98,4 23
P=2, B=3 (1/1,0/1) 217,2 28 81,1 23
P=2, B=4 (1/1,5/1) 205,8 28 76,2 23
2 consecutive B-frames
P=2, no B (-/---/-) 252,1 29 100,3 23
P=2, B=2 (2/1,0/0) 255,0 28 100,3 23
P=2, B=3 (2/1,0/1) 212,2 28 76,1 22
P=2, B=4 (2/1,5/1) 197,9 28 68,9 22

Adaptive Quantization
(AQ) yes 189,7 29 73,7 23
no 205,8 28 76,2 23

Quarter Pixel
(QPel) yes 219,2 35 82,8 29
no 205,8 28 76,2 23

Global Motion Compensation
(GMC) yes 204,3 47 75,1 29
no 205,8 28 76,2 23


Advanced Options Clip 1 Clip 2

Motion Search Precision
MSP 0 588,0 17 391,6 16
MSP 1 247,035 25 97,781 21
MSP 2 247,035 25 97,781 21
MSP 3 247,035 25 97,781 21
MSP 4 212,7 25 78,9 22
MSP 5 209,1 27 77,1 22
MSP 6 205,8 28 76,2 23

VHQ modes
VHQ 0 211,7 25 78,5 20
VHQ 1 205,8 28 76,2 23
VHQ 2 201,6 35 73,5 29
VHQ 3 200,8 42 73,1 33
VHQ 4 196,9 51 71,7 40

VHQ for B-frames
yes 204,3 30 73,8 23
no 205,8 28 76,2 23

Chroma Motion
yes 205,8 28 76,2 23
no 207,2 26 77,0 22

Turbo
yes 204,8 28 74,9 22
no 205,8 28 76,2 23


Zones Options Clip 1 Clip 2

Chroma Optimizer
yes 205,852 29 76,182 23
no 205,775 28 76,178 23
Used Software:
DGindex 1.0.12, DGdecode 1.0.0, AviSynth 2.5.5, VirtualDubMod 1.5.10.1 (Fast Recompress), XviD 1.10 beta2 (Koepi)

Used Movies:
Clip 1: The Rock, uncut PAL-Version, 0:32:51 - 0:46:51
(Interrogation Room, Hotel Scene, Car Chase)
LoadPlugin("Path\dgdecode.dll")
LoadPlugin("Path\UnDot.dll")
mpeg2source("Path\rock.d2v")
trim(49281,70280)
crop(12,72,696,424)
Undot()
BicubicResize(688,288,0,0.5)

Clip 2: The Usual Suspects, PAL-Version, 0:07:32 - 0:21:32
(Line-Up, Cell-Block, 1st Clinic Scene)
LoadPlugin("Path\dgdecode.dll")
LoadPlugin("Path\UnDot.dll")
mpeg2source("Path\usual.d2v")
trim(11310,32309)
crop(8,74,704,428)
Undot()
BicubicResize(688,288,0,0.5)

TonyMi
14th October 2005, 07:00
This test shows how different XviD settings change the size (in MB) and the encoding time (in min) of the resulting video...
Can you please add PSNR to the results?

TonyMi

Harley Quin
14th October 2005, 07:35
This thread is meant as a comparison of size and speed. It's not about quality, which is much more subjective than a single Peak-Signal-to-Noise-Ratio value makes us believe...
greetings

TonyMi
14th October 2005, 07:48
This thread is meant as a comparison of size and speed. It's not about quality, which is much more subjective than a single Peak-Signal-to-Noise-Ratio value makes us believe...
greetings
I understand, but you invested time to do much of compressions and you can easily get more information from the result. It will be usefull for basic idea about time/size/quality decision for each option.

TonyMi

Teegedeck
14th October 2005, 09:44
Thanks for that comparison, Harley Quin. Noteworthy that you used clips of reasonable length. :) I agree that PSNR would only confuse things. PSNR offers results that are of very questionable value (if we went with PSNR values a clip at constant quantizer with 1 consecutive b-frames, for example, would always get worse marks than one without b-frames; a clip with 2 consecutive b-frames would get worse marks than one with 1 b-frame; etc. - that would be pretty misleading) and it is better to cast an uninhibited look on the data about efficiency you present here, first.

Your test nicely showcases some things; for example that MSP 1,2,3 are actually the same. OK, we knew that one. ;) Also it was known, but is good to repeat, that more than 2 B-frames in a row don't improve efficiency but that AQ does - but that the percentage depends very much on the source. Interesting to see that chroma-motion really offers a small improvement of efficiency, too, that was mere guesswork before.

What I miss from your setup is a test of Trellis - a tool that improves efficiency notably and does no harm - something like a 'free meal' IMHO - but does not quite get the attention and widespread use it deserves.

sysKin
14th October 2005, 10:04
As much as I agree that psnr isn't perfect, it's still better than no metric at all. Currently there's no metric at all because "constant quantizer" is a purely artificial setting without an meaning for, well, anything. Not even quantization is constant with this setting (bframes, aq, trellis), nor ANYTHING else.

Two-pass might take some time but it's the only way.

stephanV
14th October 2005, 10:07
(if we went with PSNR values a clip at constant quantizer with 1 consecutive b-frames, for example, would always get worse marks than one without b-frames; a clip with 2 consecutive b-frames would get worse marks than one with 1 b-frame; etc. - that would be pretty misleading)
I don't see how this would be misleading, the bit rates obtained are given and I don't find it questionable that a b-frame is not as good as a p-frame. Of course how much a b-frame is worse, IS questionable and PSNR measurements won't give an answer to that.

Regarding speed and "effeciency", testing VHQ modes is very interesting too IMO.

Teegedeck
14th October 2005, 10:15
As much as I agree that psnr isn't perfect, it's still better than no metric at all. Currently there's no metric at all because "constant quantizer" is a purely artificial setting without an meaning for, well, anything. Not even quantization is constant with this setting (bframes, aq, trellis), nor ANYTHING else.

Two-pass might take some time but it's the only way.
Nah, not really. You're about to turn the aim of this test around - it was about efficiency, not about quality. Testing quality is of course more meaningful than testing efficiency, it is one step beyond, because quality is what we really want in the end, but it is at the same time more complex to do and certainly not within the scope of what Harley Quin wanted with this test.

All that this test gives us is hints on what doesn't make sense to use (more than 2 b-frames), what might make sense (GMC) and what should seriously be considered (high VHQ etc.). Whether the options that are to be seriously considered yield good quality would be another question and another test.

Harley Quin
14th October 2005, 12:18
Yes, Teegedeck read my mind.
I really did not want, and don't want to feed, a quality discussion. I was curious about size and time, I tried to make it as thorough as my knowledge allows and I thought I might share it as additional information. But I will stay within the bounds of the thread title.
I know I can't make it right to everybody, as PSNR is seen very controversial. If I added PSNR the other half would tear me apart for adding useless data. ;) I decided to leave it out. I will keep it that way.

More options are about to come and I will try to maintain it, when GUI or Codec changes. For Greyscale Encoding and Interlaced Encoding I want to use other sources, Packed Bitstream and Trellis on the list... I just started with more significant options and as data grew, I wanted to get something in the air to begin with.

Syskin, I know the quantizer is not really constant in most cases. But when I only write "single-pass" it might be misunderstood as constant bitrate, which is why single-pass still has the nimbus of bad quality.
You can still read in Crusty's otherwise brilliant FAQ, that "Unless you absolutely have to go for single pass for a specific reason there really is no other way but two-pass." And I personally think, that unless you absolutely have to reach a specific file-size there is no need to do a two-pass encoding at all.
I could change it to "target quantizer", would that be more precise?

greetings

Sharktooth
14th October 2005, 13:06
Just a note on PSNR.
I dont think it is a good idea since PSNR and psychovisual techniques are not compatible. I mean psy enanchments come from altering some parts or the entire source frame or even sequences of frames quality in a way the human eye cant distinguish the differences from the source. That alteration is made to spare bits where possible to "reassign" them for more complex parts (frames, scenes etc...).
Everytime psy enanchments are used PSNR drops since the differences from the original frame are measured (even if the human eye cant distinguish the differences during playback). So PSNR is not a good test...

P.S.: Remember that B-Frames are an approach to psy too.

sysKin
14th October 2005, 16:44
You're about to turn the aim of this test around - it was about efficiency, not about quality.
Yes ok, but what *is* efficiency. The only one I know is compression VS at fixed quality, or quality at fixed filesize.

If you say that this test has any meaning at all, just exactly how do you read it? (ie is smaller filesize better or not).

If you say smaller filesize is better, I can give you an "even better" codec is a matter of minutes. Twice "better" if you want to.

I don't think you'll even convice me that a fixed quantizer in GUI has any meaning. It's just a number.


I generally would ignore such threads, if it wasn't that many people are here for some knowledge. This is too close to measuring number of UFOs at fixed number of sheep. ;) (j/k)

Teegedeck
14th October 2005, 20:20
Harley Quin laid it out as a test of efficiency - 'compression vs. speed' - of XviD's encoding tools. That means, savings from the use of single features weighted against the resulting speed penalty. I think that about summs it up. Right?

The 'quality penalty' was not tested. Doing so on the features with worthwhile results of this 'compression vs. speed' test is of course interesting.

Now let me step aside and ask: what's all the fuss about? I mean, it is not as if we all did not already know the results, didn't we? We know GMC has a big speed penalty and little capacity for saving space as well as small capacity for quality increase (yet I use it); we know that b-frames are absolutely worthwhile in 2-pass; we know that VHQ (=Very Handsome Qualigosaur) has a big speed penalty but is absolutely worth it. We're just discussing data that gives a little more foundation to things we already know from experience and we're nodding appreciatively. I can't find that so bad. It's fun, it's a social event, we learn a little, we're not trolling or twisting the truth or anything remotely as serious as that. Sheesh.

aabxx
27th July 2006, 01:23
Has anyone considered turning this thread into a sticky? It's a shame it's on page 20, when it could do much so much good if it were on page 1 all the time :)

DarkZell666
27th July 2006, 12:10
Something bugs me a little :

Turbo
yes 204,8 28 74,9 22
no 205,8 28 76,2 23


I would have thought that turbo would blow up the filesize a bit because it enables a less precise motion estimation method (which doesn't seem to be less efficient at all ... oO)

Ok it's a matter of 1MB out of ~200MB, but still ;)
Wierdly enough, turbo doesn't appear to be the right name for it either :p

One more thing :

Multiply the size by 10 and you get the framerate (in kBit/sec).
May I dare ask if you were around Venus while writting this post ? ;) I've been there a couple of times we might have met already :p

Harley Quin
27th July 2006, 13:18
May I dare ask if you were around Venus while writting this post ? ;) I've been there a couple of times we might have met already :p
Yes, you may, and, no, I wasn't. At least... I hope so...;)
Take a clip with 100 MB,
- multiply by 1024, for kilobyte
- multiply by 8, for kilobit
- multiply by 1,024, for 'kilo' above means 1024, whereas in kilobit/sec 'kilo' means 1000
- divide by 840, the number of seconds, you get 998,64 kbit/sec.
20971,52 frames would have been more precise, but I thought, what the... :)

It's really slightly outdated and not that groundbreaking an information to be a sticky... But it's nice to see it's been read once in a while since then. :)

greetings and enjoy

Manao
27th July 2006, 13:59
I would have thought that turbo would blow up the filesize a bit because it enables a less precise motion estimation method (which doesn't seem to be less efficient at all ... oO)

Ok it's a matter of 1MB out of ~200MB, but still
Wierdly enough, turbo doesn't appear to be the right name for it either No, it all comes down to what sysKin has been hamering all along the thread : file size alone means nothing. Turbo means 'speed up at the cost of efficiency', and efficiency is quality at a given filesize ( or size at a given quality ).

DarkZell666
27th July 2006, 14:16
@Harley
you get the framerate (in kBit/sec).
ok so you really were refering to bitrate then ;)

@Manao : true enough :)

Sagittaire
27th July 2006, 20:08
Nah, not really. You're about to turn the aim of this test around - it was about efficiency, not about quality. Testing quality is of course more meaningful than testing efficiency, it is one step beyond, because quality is what we really want in the end, but it is at the same time more complex to do and certainly not within the scope of what Harley Quin wanted with this test.

Your "efficiency" doesn't mean anything.

- I can choose to compare custom matrix flat16 vs flat32. Size for flat32 matrix will have the smaller size but quality for flat16 will be very higher.

- Compare size without quality for bframe setting is completely useless too : size for 1/1.5/1.0 will be better but quality for 1/1.0/0.0 will be very higher. If you want the most "efficiency" setting for bframe try 5/1.0/31 ... lol

- Same thing for Trellis. Enable trellis make higher size. Conclusion: trellis is not good. No, simply because trellis improve quality too.

Manao
27th July 2006, 20:43
I can choose to compare custom matrix flat16 vs flat32Strange, I would simply have raised the quantizer. Well, I guess that was too obvious :p
Same thing for Trellis. Enable trellis make higher sizeWrong. Trellis, depending on the quantizer, the quantization method, and the rate-distorsion tradeoff chosen, can either raise or lower the size. And it can lower it greatly, since with a very high rate-distorsion tradeoff, you would end up without any DCT coefficients to encode ( quality would be worse than Q31 though... ).

Sagittaire
27th July 2006, 21:05
Strange, I would simply have raised the quantizer. Well, I guess that was too obvious :p
Wrong. Trellis, depending on the quantizer, the quantization method, and the rate-distorsion tradeoff chosen, can either raise or lower the size. And it can lower it greatly, since with a very high rate-distorsion tradeoff, you would end up without any DCT coefficients to encode ( quality would be worse than Q31 though... ).

Well no exclusive proposition ... (ouf)

Same thing for Trellis. Enable trellis (can) make higher size then conclusion is "trellis is not good"?
No, simply because trellis improve quality too.


Strange, I would simply have raised the quantizer. Well, I guess that was too obvious

No it's constant quantizer comparison ... ;-)

Sagittaire
27th July 2006, 21:20
and efficiency is quality at a given filesize ( or size at a given quality ).

agree, agree + speed ...

efficiency is
- better speed for same quality and same size
- better quality for same speed at same size
- better size for same speed and same quality

and IMO it's impossible to make that without metric (really too difficult to make that with subjective test).

IgorC
27th July 2006, 23:14
Sagi. :) You forgot there are other interpretations of
Eficiency is also :
better quality and better speed or a a litle bit worse quality with much higher speed. (fast preset) ...etc

foxyshadis
27th July 2006, 23:57
No, slightly lower quality at much higher speed is equivalent to the same quality at somewhat higher speed (unless your quality weight is enough higher than speed to offset it), which is increased efficiency. You have to normalize at least one to make measurements against, or there's no science involved. You're free to weight the normalization though - although this is tempting it's not terribly useful:

efficiency(1) = ssim(1) * fps(1) / size(1)
efficiency(2) = ssim(2) * fps(2) / size(2)

Efficiency is simply whatever your equation gives, once you plug the numbers and weights in. Everyone has their own weights, and their own saturation points at which raising one variable no longer matters, which is why it's all so subjective. ;) (Not to mention ssim being rather imperfect.) A proper equation would be a trivariate logistic function with saturation points.

Then you can just graph it in Maple and match any codec's performance against your personal efficiency measure. :p