Log in

View Full Version : libavcodec avc decoding performance test


bond
24th August 2005, 02:05
i today wanted to find out what influences decoding of avc the most with the libavcodec avc decoder as i wanted to know what to best enable/disable in my encodes for getting realtime playback on my good old pentium3 866mhz

therefore i made a bunch of encodes with x264 r287 and benchmarked them with recent libavcodec cvs via mplayer's -vo null -benchmark option with the following "raw" results
the source is the typical matrix1 encode (smith interrogating morpheus, lobby shootout...) 7116 frames:

640x256 B0 Ref3 p8x8-i4x4 loop-5 cabac 103.999s 68.44fps
720x288 B2 Ref3 p8x8-i4x4 cabac 105.682s 67.33fps
720x288 B2 Ref3 p4x4-i4x4 cabac 107.274s 66.33fps
640x256 B3-Ref Ref5 p4x4-i4x4 loop-5 WBP cabac 111.700s 63.71fps
640x256 B3-Ref Ref5 p4x4-i8x8 loop-5 WBP cabac 113.954s 62.45fps
720x288 B3-Ref Ref5 p4x4-i8x8 WBP cabac 114.845s 61.96fps
640x256 B2 Ref3 p8x8-i4x4 loop-5 cabac 117.338s 60.65fps
720x288 B3-Ref Ref5 p4x4-i8x8 loop-5 WBP 119.702s 59.45fps
640x256 B2 Ref3 p8x8-i4x4 loop-5 w(b)p cabac 122.646s 58.02fps
720x288 B0 Ref5 p4x4-i8x8 loop-5 WBP cabac 123.567s 57.59fps
640x256 B2 Ref5 p4x4-i4x4 loop-5 cabac 127.723s 55.70fps
720x288 B3-Ref Ref1 p4x4-i4x4 loop-5 WBP cabac 128.334s 55.45fps
720x288 B2 Ref3 p8x8-i4x4 loop-5 cabac 133.842s 53.17fps
720x288 B1 Ref5 p4x4-i8x8 loop-5 WBP cabac 134.824s 52.78fps
720x288 B3-Ref Ref3 p4x4-i4x4 loop-5 cabac 136.626s 52.08fps
720x288 B3-Ref Ref3 p4x4-i4x4 loop-5 WBP cabac 139.901s 50.86fps
720x288 B3-Ref Ref5 p4x4-i4x4 loop-5 WBP cabac 142.535s 49.92fps
720x288 B3-Ref Ref5 p4x4-i8x8 loop-5 WBP cabac 142.565s 49.91fps
720x288 B3 Ref5 p4x4-i8x8 loop-5 WBP cabac 143.767s 49.50fps
720x288 B3-Ref Ref5 p4x4-i8x8 loop+6 WBP cabac 150.636s 47.24fpsgrouping the comparable things together we get the following:

------------------------------------------------------------------------
------------------------------------------------------------------------
blocksizes

p8x8 vs. p4x4

720x288 B2 Ref3 p8x8-i4x4 cabac 105.682s 67.33fps
720x288 B2 Ref3 p4x4-i4x4 cabac 107.274s 66.33fps
1.00

------------------------------------------------------------------------
i4x4 vs. i8x8

720x288 B3-Ref Ref5 p4x4-i4x4 loop-5 WBP cabac 142.535s 49.92fps
720x288 B3-Ref Ref5 p4x4-i8x8 loop-5 WBP cabac 142.565s 49.91fps
0.01

640x256 B3-Ref Ref5 p4x4-i4x4 loop-5 WBP cabac 111.700s 63.71fps
640x256 B3-Ref Ref5 p4x4-i8x8 loop-5 WBP cabac 113.954s 62.45fps
1.26

------------------------------------------------------------------------
------------------------------------------------------------------------
multiple reference frames

ref1 vs. ref5

720x288 B3-Ref Ref1 p4x4-i4x4 loop-5 WBP cabac 128.334s 55.45fps
720x288 B3-Ref Ref5 p4x4-i4x4 loop-5 WBP cabac 142.535s 49.92fps
5.53

------------------------------------------------------------------------
ref1 vs. ref3

720x288 B3-Ref Ref1 p4x4-i4x4 loop-5 WBP cabac 128.334s 55.45fps
720x288 B3-Ref Ref3 p4x4-i4x4 loop-5 WBP cabac 139.901s 50.86fps
4.59

------------------------------------------------------------------------
ref3 vs. ref5

720x288 B3-Ref Ref3 p4x4-i4x4 loop-5 WBP cabac 139.901s 50.86fps
720x288 B3-Ref Ref5 p4x4-i4x4 loop-5 WBP cabac 142.535s 49.92fps
0.94

------------------------------------------------------------------------
------------------------------------------------------------------------
weighted (bi)prediction

wbp vs. nowbp

720x288 B3-Ref Ref3 p4x4-i4x4 loop-5 cabac 136.626s 52.08fps
720x288 B3-Ref Ref3 p4x4-i4x4 loop-5 WBP cabac 139.901s 50.86fps
1.22

------------------------------------------------------------------------

wp+wbp vs. nowp+nowbp

640x256 B2 Ref3 p8x8-i4x4 loop-5 cabac 117.338s 60.65fps
640x256 B2 Ref3 p8x8-i4x4 loop-5 w(b)p cabac 122.646s 58.02fps
2.63

------------------------------------------------------------------------
------------------------------------------------------------------------
b-frames

0 B-frames vs. 3 B-frames

720x288 B0 Ref5 p4x4-i8x8 loop-5 WBP cabac 123.567s 57.59fps
720x288 B3 Ref5 p4x4-i8x8 loop-5 WBP cabac 143.767s 49.50fps
8.09

------------------------------------------------------------------------
0 B-frames vs. 2 B-frames

640x256 B0 Ref3 p8x8-i4x4 loop-5 cabac 103.999s 68.44fps
640x256 B2 Ref3 p8x8-i4x4 loop-5 cabac 117.338s 60.65fps
7.79

------------------------------------------------------------------------
0 B-frames vs. 1 B-frames

720x288 B0 Ref5 p4x4-i8x8 loop-5 WBP cabac 123.567s 57.59fps
720x288 B1 Ref5 p4x4-i8x8 loop-5 WBP cabac 134.824s 52.78fps
4.81

------------------------------------------------------------------------
1 B-frames vs. 3 B-frames

720x288 B1 Ref5 p4x4-i8x8 loop-5 WBP cabac 134.824s 52.78fps
720x288 B3 Ref5 p4x4-i8x8 loop-5 WBP cabac 143.767s 49.50fps
3.28

------------------------------------------------------------------------
B-ref vs. no B-ref

720x288 B3-Ref Ref5 p4x4-i8x8 loop-5 WBP cabac 142.565s 49.91fps
720x288 B3 Ref5 p4x4-i8x8 loop-5 WBP cabac 143.767s 49.50fps
0.41

------------------------------------------------------------------------
------------------------------------------------------------------------
loop

loop-5 vs. no loop

720x288 B2 Ref3 p8x8-i4x4 cabac 105.682s 67.33fps
720x288 B2 Ref3 p8x8-i4x4 loop-5 cabac 133.842s 53.17fps
14.16

720x288 B3-Ref Ref5 p4x4-i8x8 WBP cabac 114.845s 61.96fps
720x288 B3-Ref Ref5 p4x4-i8x8 loop-5 WBP cabac 142.565s 49.91fps
12.05

------------------------------------------------------------------------
loop-5 vs. loop+6

720x288 B3-Ref Ref5 p4x4-i8x8 loop-5 WBP cabac 142.565s 49.91fps
720x288 B3-Ref Ref5 p4x4-i8x8 loop+6 WBP cabac 150.636s 47.24fps
2.67

------------------------------------------------------------------------
------------------------------------------------------------------------
cabac

cabac vs. no cabac

720x288 B3-Ref Ref5 p4x4-i8x8 loop-5 WBP 119.702s 59.45fps
720x288 B3-Ref Ref5 p4x4-i8x8 loop-5 WBP cabac 142.565s 49.91fps
9.54

------------------------------------------------------------------------
------------------------------------------------------------------------
cabac + noloop vs. loop-5 + nocabac


720x288 B3-Ref Ref5 p4x4-i8x8 WBP cabac 114.845s 61.96fps
720x288 B3-Ref Ref5 p4x4-i8x8 loop-5 WBP 119.702s 59.45fps
2.51

------------------------------------------------------------------------
------------------------------------------------------------------------
resolution

720x288 vs. 640x256

640x256 B3-Ref Ref5 p4x4-i8x8 loop-5 WBP cabac 113.954s 62.45fps
720x288 B3-Ref Ref5 p4x4-i8x8 loop-5 WBP cabac 142.565s 49.91fps
12.54

------------------------------------------------------------------------
------------------------------------------------------------------------summing the findings up imho i think it can be said that there are three groups of features/things which impact speed:

1) the ones that impact speed a lot:
loop, resolution, cabac and high numbers of b-frames (eg 3), of which cabac and b-frames take less decoding time

2) the ones that hardly impact speed (but in sum might do too):
blocksizes, weighted prediction, weighted biprediction and b-references (which is actually faster than b-frames without b-ref)

3) the ones in the middle:
small numbers of b-frames (eg 1) and multiple reference frames

hope someone might find this useful (eg maybe for tuning the decoder?) :)


edit: loop was set with a strenght of -5,-5

Caroliano
24th August 2005, 02:19
Very intersting. I have an 1.7 celeron and besides that I want that my friends that have even slower computers play my encodes well too.

And why you dont use the defaut value for in-loop filter? It would make any diference?

CiNcH
24th August 2005, 04:34
I haven't done speed tests yet but my Intel Pentium M 1.6 GHz (Dothan) computer stays at 600 MHz (lowest SpeedStep) when playing back an AVC / aacPlus v2 encode. CPU usage ranges between 50 and 70%.

Video: 720 x 304 / x264 High Profile (CABAC, 5 ref. frames, 3 cons. b-frames, b-pyramid, weighted b-pred., in-loop -2/-2, all mb partitions, 8x8DCT)
Audio: aacPlus v2 (AAC + SBR + PS)
Container: MP4

Demuxer: Haali Media Splitter
Decoder: ffdshow, CoreAAC

ggab
24th August 2005, 09:13
CiNcH, which aacPlus v2 encoder did u use? thks

hworldjj
24th August 2005, 10:17
Hi, bond,

Did you compare the performance between Moonlight, Nero and ffmpeg H.264 decoder? It seems ffmpeg a little worse than the other two.

CiNcH
24th August 2005, 14:38
CiNcH, which aacPlus v2 encoder did u use? thks

I used the Coding Technologies reference encoder. Check! (http://forum.doom9.org/showthread.php?t=99071)

bond
25th August 2005, 18:35
i now added a few more test:
- i moved multiple reference frames to a "middle" group as their impact on speed is indeed noticeable
- using a higher strength of loop indeed influences speed more than a lower strength (as loop is less used with a lower strength)
- the higher the number of b-frames, the slower the decoding

about the question to also benchmark other decoders, i would of course be very interested in this too, but i dunno how to accurately measure the time with these filters?

is it possible to decode via directshow filters in mplayer?

dimzon
26th August 2005, 11:19
about the question to also benchmark other decoders, i would of course be very interested in this too, but i dunno how to accurately measure the time with these filters?

You can use my method (check my signature)

bond
26th August 2005, 12:07
You can use my method (check my signature)hm thx thats indeed a possibility, i will have a look :)

Sergey A. Sablin
26th August 2005, 12:25
hm thx thats indeed a possibility, i will have a look :)
You can try to use this direct show filter to measure DShow decoders preformance.
Before starting the graph you need to uncheck Graph->Use clock menu. After graph stopped open property page of this filter and you will see the performance.

bond
26th August 2005, 13:37
You can try to use this direct show filter to measure DShow decoders preformance.
Before starting the graph you need to uncheck Graph->Use clock menu. After graph stopped open property page of this filter and you will see the performance.great stuff, thx a lot!!!

btw if you want me to test your decoder too, it might be a nice if you could send me a unlocked copy, cause i cant get things to work with your player as described here (http://www.elecard.ru/forum/viewtopic.php?t=550&highlight=) :(

Sergey A. Sablin
26th August 2005, 13:48
great stuff, thx a lot!!!

btw if you want me to test your decoder too, it might be a nice if you could send me a unlocked copy, cause i cant get things to work with your player as described here (http://www.elecard.ru/forum/viewtopic.php?t=550&highlight=) :(
Our current decoder is a little bit slower than we have in mooolight time, so we don't want to test it now. But anyway I'll try to help you with decoder evaluation.

BTW when you measure DShow decoders performance you should be aware of output mediatype. Elecard Chegepuga (filter attached above) connects on any mediatype, but different decoders use different media types as first output mediatype - some of them use YV12 (as simplest), some of them use YUY2 or UYVY as fastest to render via graphic cards. Of course second variant is a little bit slower in pure performance (I mean just decoding - not displaying)

EDIT: you can specify media type on ech's property page or use null-in-place filter for that purpose

bond
26th August 2005, 14:49
Our current decoder is a little bit slower than we have in mooolight time, so we don't want to test it now. But anyway I'll try to help you with decoder evaluation.hm what do you mean with "moonlight time"? does moonlight use another decoder than elecard?

BTW when you measure DShow decoders performance you should be aware of output mediatype. Elecard Chegepuga (filter attached above) connects on any mediatype, but different decoders use different media types as first output mediatype - some of them use YV12 (as simplest), some of them use YUY2 or UYVY as fastest to render via graphic cards. Of course second variant is a little bit slower in pure performance (I mean just decoding - not displaying)

EDIT: you can specify media type on ech's property page or use null-in-place filter for that purpose thx for the info! :)

Sergey A. Sablin
26th August 2005, 14:59
hm what do you mean with "moonlight time"? does moonlight use another decoder than elecard?
Elecard developed and supported all the products for moonlight, but in december of 2004 our contract was finished. After that moonlight stopped it's activities and now it is under liquidation process.
So now we are working only on our own products.

bond
26th August 2005, 15:02
Elecard developed and supported all the products for moonlight, but in december of 2004 our contract was finished. After that moonlight stopped it's activities and now it is under liquidation process.
So now we are working only on our own products.ah ic, so as moonlight was swallowed by mainconcept i assume this means that moonlight/mainconcept will not use elecards avc products in the future, but the ones produced by mainconcept?

Sergey A. Sablin
27th August 2005, 10:24
ah ic, so as moonlight was swallowed by mainconcept i assume this means that moonlight/mainconcept will not use elecards avc products in the future, but the ones produced by mainconcept?
Moonlight is saling out its IP - http://www.moonlight.co.il/rfp.php
Elecard is in the process of merging with Mainconcept. So there will be joint products of elecard-mainconcept.

bond
27th August 2005, 19:21
ah ic, mixed things up

bond
27th August 2005, 19:56
will/did elecard also write an own avc decoder seperately from moonlight?

Sergey A. Sablin
29th August 2005, 04:35
will/did elecard also write an own avc decoder seperately from moonlight?
As I said previously
Our current decoder is a little bit slower than we have in mooolight time, so we don't want to test it now.
so it is now under last stage of development - performance tuning ;)

bond
29th August 2005, 11:43
ah i understand, well keep things coming ;)

Sirber
29th August 2005, 12:22
Great comparison bond! :D

bond
29th August 2005, 13:27
thx sirber :)

sergey, am i right when i assume your filter tests the decoding performance of both decoder and system and not only the decoder (eg mplayer can measure what the decoder alone needs and what the system needs)

Sergey A. Sablin
29th August 2005, 13:39
thx sirber :)

sergey, am i right when i assume your filter tests the decoding performance of both decoder and system and not only the decoder (eg mplayer can measure what the decoder alone needs and what the system needs)
filter just calculates the number of frames per second which decoder decodes at the time you ran it. It is just renderer filter that receives decoded frames, summarize they count and devides by spending time. So if system was busy at that time then results will be worse.

bond
29th August 2005, 13:48
filter just calculates the number of frames per second which decoder decodes at the time you ran it. It is just renderer filter that receives decoded frames, summarize they count and devides by spending time. So if system was busy at that time then results will be worse.hm ic, i asked because the results of your filter dont match exactly with the cpu time used by graphedit shown by the task manager (there is a difference of about 1 sec)

Sergey A. Sablin
29th August 2005, 13:55
hm ic, i asked because the results of your filter dont match exactly with the cpu time used by graphedit shown by the task manager (there is a difference of about 1 sec)
it will be more understandable if you will tell the percentage number of difference ;) cause it may be just 1% of error or 10% that make sense.
Also I recommend you to close all another applications when you test performance. And do not forget that there are another filters in the graph that can use CPU time.

bond
29th August 2005, 14:16
it will be more understandable if you will tell the percentage number of difference ;) cause it may be just 1% of error or 10% that make sense.the difference is like 3%

Also I recommend you to close all another applications when you test performance. And do not forget that there are another filters in the graph that can use CPU time. i noticed that i am able to connect the filter directly to the parsers too (mpeg2video), is it possible that way to measure the performance of a parser?

Sergey A. Sablin
29th August 2005, 14:52
the difference is like 3%

i noticed that i am able to connect the filter directly to the parsers too (mpeg2video), is it possible that way to measure the performance of a parser?
yes sure. but the filter will measure the number of packets received from parser per second. Packets are commonly much lesser than encoded frames (but with different parsers size of packets may vary), so you can't compare decoding time with parsing time.
You need to calculate whole time spended to parsing and whole time spended to decoding then you can compare them.

EDIT: I think 3% isn't to much - it maybe thread switching / io operations or like that

bond
29th August 2005, 14:59
yes sure. but the filter will measure the number of packets received from parser per second. Packets are commonly much lesser than encoded frames (but with different parsers size of packets may vary), so you can't compare decoding time with parsing time.
You need to calculate whole time spended to parsing and whole time spended to decoding then you can compare them.hm would it be possible to take the "frames per second" value shown by the filter to compare the speed of different parsers of any kind?
will this be comparable with potentially different packets?
will this be comparable with different containers, eg placing the same stream in different containers?

Sergey A. Sablin
29th August 2005, 15:07
hm would it be possible to take the "frames per second" value shown by the filter to compare the speed of different parsers of any kind?
will this be comparable with potentially different packets?
will this be comparable with different containers, eg placing the same stream in different containers?
no
no
no
Commonly packets have fixed size, so for all these cases filter have to know what does mean frame/picture for this format and parse receiving packets to recognize when frame began and when it finished - too much logic for simple measurement tool initially developed for encoders/decoders which use frames/pictures as output.

bond
29th August 2005, 15:12
ic thx

bond
19th December 2005, 22:11
You can try to use this direct show filter to measure DShow decoders preformance.
Before starting the graph you need to uncheck Graph->Use clock menu. After graph stopped open property page of this filter and you will see the performance.sergey, i noticed a problem with your filter:
i tried to measure the decoding speed of an .avi with uncompressed YUY2:

when rendering the avi normally the graph looks as follows:
source -> avi splitter -> avi decompressor -> renderer
there is no colorspace conversion

when i want to replace the renderer with your Chegepuga filter avi decompressor makes a yuy2 -> rgb32 conversion
when i enforce yuy2 in chegepuga some other filters are placed in between doing strange colorspace conversions to rgb4 and rgb8 or so before connecting to chegepuga

any idea how i can get chegepuga connect correctly to the yuy2 avi?

Sergey A. Sablin
20th December 2005, 04:45
you can connect directly chegepuga to avi splitter if you have any problems with avi decompressor:

source -> avi splitter -> chegepuga

if it is uncompressed avi file than avi splitter will deliver uncompressed frames in original color space.

bond
20th December 2005, 12:13
indeed, but i still wondered why its not possible to get yuy2 from the avi decompressor (i use real's yuv codec in vfw)