View Full Version : MPEG-4 AVC/H.264 decoder comparison
bond
30th August 2005, 19:15
one of the main advantages of open standards, like mpeg-4 avc, is that it leads to competition between various codec producers as the consumer will always tend to use the product with the best price/performance relationship.
while prices are pretty well known for the consumer there is one thing often lying in the dark: the actual performance
while for the encoding side quality comparisons exist, like the one from doom9, decoders are often overlooked. of course, as every decoder should output exactly the same picture and speed also didnt really play a role as existing formats are pretty easy to decode.
but with a format, like avc, which needs a lot of resources while decoding, this might be different for a lot of people who dont own top of the notch pcs.
so as avc is very speed consuming and actually only one free avc decoder is existing (libavcodec) it indeed makes sense to compare the performance of existing decoders, before the user decides what to spend his money on
in this comparison i will compare the following decoders:
- ateme
- elecard
- libavcodec
- mainconcept
- moonlight
- nero
- videosoft
imho the quality of a decoder is defined by four things:
A) price
B) supported features of the format
C) decoding speed
D) postprocessing
A) PRICE
its pretty easy to check out the prices of the available decoders:
- libavcodec (http://www1.mplayerhq.hu/cgi-bin/cvsweb.cgi/ffmpeg/libavcodec/h264.c?cvsroot=FFMpeg): 0 USD as being opensource, but you have to compile it yourself, supported in various players
- videosoft (http://secure.softwarekey.com/solo/products/trigdetail.asp?T=2345): 19.95 USD, available in their "VSS H.264/AVC Codec Baseline" tool
- moonlight (http://www.elecard.net.ru/products/): 20 USD, available in the "Moonlight-Elecard MPEG Player 3.0"
- nero (http://www.nero.com/enu/Nero_Digital_Pro_InfoPage.html): 29.90 USD, available in their "Nero Digital Pro" tool
- elecard (http://www.mainconcept.com/h264_encoder.shtml): 499 USD, available in the mainconcept "H.264 Encoder v2" tool
- mainconcept (http://www.mainconcept.com/h264_encoder.shtml): 499 USD, not available anymore, has been available in their "H.264 Encoder v1" tool (v2 uses the elecard decoder)
- ateme: not publically available
B) SUPPORTED FEATURES
according to my tests the following features are NOT supported by the different decoders (version number described below):
| ateme | elecard | libavcodec | mainconcept | moonlight | nero | videosoft |
------------------------------------------------------------------------------------------------------
BASELINE PROFILE| | | | | | | |
------------------------------------------------------------------------------------------------------
p4x4,b8x8,i4x4 | | | | | | | |
------------------------------------------------------------------------------------------------------
loop | | | | | | | |
------------------------------------------------------------------------------------------------------
multireferences | | | | | | | |
------------------------------------------------------------------------------------------------------
adaptive quant | | | | | | | |
------------------------------------------------------------------------------------------------------
------------------------------------------------------------------------------------------------------
MAIN PROFILE | | | | | | | |
------------------------------------------------------------------------------------------------------
b-frames | | | | | | | |
------------------------------------------------------------------------------------------------------
b-references | | | | | | | x |
------------------------------------------------------------------------------------------------------
wp,wbp | | | | | | | |
------------------------------------------------------------------------------------------------------
cabac | | | | | | | |
------------------------------------------------------------------------------------------------------
fields-only | | | | * | | | |
------------------------------------------------------------------------------------------------------
paff | | | | * | * | | x |
------------------------------------------------------------------------------------------------------
mbaff | | | | * | x | | x |
------------------------------------------------------------------------------------------------------
------------------------------------------------------------------------------------------------------
HIGH PROFILE | | | | x | | | x |
------------------------------------------------------------------------------------------------------
8x8dct | | | | x | | | x |
------------------------------------------------------------------------------------------------------
i8x8 | | | | x | | | x |
------------------------------------------------------------------------------------------------------
cqm | | | | x | | | x |
------------------------------------------------------------------------------------------------------
lossless | | x | | x | x | x | x |
------------------------------------------------------------------------------------------------------so the only decoder which supported everything was the not publically available decoder from ateme
second place goes to libavcodec (not supporting interlacing), elecard and nero (not supporting lossless avc)
the not anymore developed decoders from mainconcept and moonlight had problems with their interlacing support. moonlight shows artefacts with paff. mainconcept shows artefacts with all interlacing modes.
videosoft didnt support b-references/arbitrary frame orders
it also gets interesting when looking at the high profile support, whose tested features were supported more or less by libavcodec, moonlight, elecard and nero but fully only by ateme
C) DECODING SPEED
as all decoders are available as directshow decoders i used this interface for testing the decoders
additionally i also tested libavcodec in mplayer to honor its advantage of being useable also in potentially better and faster interfaces than directshow
my cpu is a pentium3 866mhz
the speed has been measured with elecard's great Chegepuga filter, available here (http://forum.doom9.org/showthread.php?p=703999#post703999). yv12 output has been enforced, which all decoders supported
the graph has been setup as
file parser -> decoder -> chegepuga
as some decoders are limited to or work best with specific file formats, eg mainconcept only works with .mpg, i decided to use the file parsers of the decoder manufacturers too. this also helped avoiding interoperability problems. of course faster parsers can influence the output speed positively, which has to be taken into account when looking at the values described below
the following filter/parser combinations have been used:
ateme: ateme mp4 parser 1.2.5.3 / ateme decoder 2.2.1.0
ateme_old: ateme mp4 parser 1.2.2.0 / ateme decoder 1.1.2.0
elecard: elecard mp4 parser 1.4.0 b50929 / mainconcept decoder 1.01.00.07
libav-ffdshow: haali mp4 parser Sep 03 2005 / ffdshow-libavcodec decoder Oct 13 2005
libav-ffdshow_old: haali mp4 parser Aug 18 2005 / ffdshow-libavcodec decoder Aug 22 2005
libav-mplayer: mplayer mp4 parser cvs aug 12 2005 / mplayer-libavcodec decoder cvs aug 12 2005
mainconcept: mainconcept mpg parser 1.00.01.02 / mainconcept decoder 1.01.00.04
moonlight: elecard mp4 parser 1.3.5 b50823 / moonlight decoder 0.9.0 b50208
nero: nero mp4 parser 2.0.2.49 / nero decoder 2.0.2.46
videosoft: m$ avi parser 6.5.1.902 / videosoft decoder 2.3.1.5
videosoft_old: m$ avi parser 6.5.1.902 / videosoft decoder 2.0.2.3the method for measuring the libavcodec performance with mplayer is described here (http://forum.doom9.org/showthread.php?t=99131)
the source was the typical matrix 1 clip often used for comparisons (smith interrogating morpheus, lobby shootout), ~7000 frames
various resolutions and codec settings have been used (the details can be seen in the raw results attached below), the target bitrate was ~700kbps, which can be seen as the typical 1 CDR DVD backup bitrate
as encoder x264 has been used, because its one of the most complete encoders and easy to configure
also its able to output directly to .mp4, which most decoders supported
.mpg has been muxed from .mp4 with ffmpeg cvs june 24 2005
the following amount of samples have been encoded:
high profile: 9
main profile: 13
baseline profile: 1
Results: (measured in frames per second)
ALL SAMPLES
average of all samples except 1 with cqm
ateme: 58.78
libav-mplayer: 58.22
moonlight: 55.48
libav-ffdshow: 52.15
libav-ffdshow_old: 52.11
nero: 50.74
elecard: 44.04
HIGH PROFILE
average of all samples except 1 with cqm
libav-mplayer: 55.11
ateme: 53.77
moonlight: 49.57
libav-ffdshow: 49.44
libav-ffdshow_old: 48.84
nero: 46.37
elecard: 40.48
ateme_old: high profile not supported
mainconcept: high profile not supported
videosoft: high profile not supportedMAIN PROFILE
ateme: 60.83
libav-mplayer: 59.00
moonlight: 57.59
libav-ffdshow_old: 53.24
libav-ffdshow: 53.08
nero: 52.58
elecard: 45.61
mainconcept: 43.40
videosoft: not all samples tested (b-ref not supported)
ateme_old: not all samples tested (b-ref not supported)BASELINE PROFILE
moonlight: 75.34
libav-mplayer: 72.88
ateme: 72.28
ateme_old: 70.63
videosoft: 64.01
libav-ffdshow_old: 63.47
libav-ffdshow: 61.83
nero: 61.80
elecard: 52.02
mainconcept: 51.16
videosoft_old: 33.46- libavcodec in mplayer and ateme are nearly always the fastest
when looking only at the directshow decoders you see that
- moonlight is faster than ffdshow and nero (but not than mplayer)
- ffdshow is as good as always faster than nero
- videosoft is performing in the middle
- elecard and mainconcept are slow
D) POST PROCESSING
imho post processing doesnt really play a role with avc, as the format itself is already good enough to provide a pretty good picture without the need for enhancing it during playback
as post processing would also require more processor power and the output quality would have to be judged subjectively i left it away in this comparison totally
CONCLUSION
the results show once again that opensource development is very powerful when it comes to providing excellent quality as libavcodec was able to provide great speeds when being used in the opensource mplayer and also supports as good as all tested avc features except interlacing
the not publically available decoder from ateme also provided excellent results outperforming the other decoders on many samples and being the only one which supported all tested coding features
moonlight was a surprise for me as it provided excellent speed results, supported nearly all avc features i tested and is also relatively cheap available. also moonlight offers the possibility to play avc in .mp4 and .mpg
moonlight is bankrupt and selling its tools, including the decoder, atm. i hope someone will buy it and continue developing it, its really worth it
nero was a disappointment imho as it was till now always seen as the benchmark, a role it clearly wasnt able to play, being slower than moonlight and also slightly slower than ffdshow, not to speak of mplayer
it also has to be mentioned that nero uses a very fast mp4 parser, whereas ffdshow and moonlight used slower ones
still nero also supports a lot of avc features
judging from the price mainconcept charges for their avc implementation (499 USD) i had high expectations for both mainconcepts original decoder (available in v1) and the new elecard decoder its uses since v2, which it surely didnt meet, as it performed pretty poorly
i know elecard's decoder is very new and i hope they will continue improving it so it can keep up speedwise with the other decoders and also with the price you have to pay for it. it already supports a lot of avc features
videosoft's decoder was a midperformer, with the downside of not supporting b-references, paff/mbaff and high profile. i know vss is working on high profile, but i wasnt able to test it. i hope the bref thing will get fixed soon so i can rank the decoder speedwise in main profile too
bond
30th August 2005, 19:22
i finally found the time to calculate how the different decoders perform on different coding tools, eg which decoder is the fastest on decoding cabac, etc...
the results shown below can give you an idea on how the decoders perform, of course the results are only valid on the specific two clips i compared for deriving the shown value (in fps) telling the decoding speed difference between two clips (one with the specific feature enabled and the other one with the feature disabled)
i ranked the decoders by the % by which the decoding speed decreases when an additional features is enabled
the higher the shown value, the worse the performance of the decoder for the specific feature:
------------------------------------------------------------------------
------------------------------------------------------------------------
blocksizes
p8x8 vs. p4x4
720x288 B2 Ref3 i4x4 cabac
nero: 18,4% 12,00
ateme_old: 9,7% 5,13
ateme: 9,2% 5,62
videosoft: 5,7% 3,24
libav-ffdshow: 1,7% 1,03
libav-mplayer: 1,5% 1,00
elecard: 1,5% 0,74
libav-ffdshow_old: 1,4% 0,84
moonlight: 1,1% 0,68
mainconcept: 0,8% 0,37
------------------------------------------------------------------------
i4x4 vs. i8x8
720x288 B3-Ref Ref5 p4x4 loop-5 WBP cabac
ateme: 4,8% 2,57
elecard: 3,1% 1,20
libav-ffdshow: 1,3% 0,59
libav-ffdshow_old: 0,9% 0,40
libav-mplayer: 0,0% 0,01
moonlight: -0,1% -0,06
nero: -0,7% -0,28
640x256 B3-Ref Ref5 p4x4 loop-5 WBP cabac
libav-ffdshow_old: 4,8% 2,87
elecard: 4,3% 2,12
ateme: 3,0% 1,94
libav-mplayer: 2,0% 1,26
moonlight: 1,3% 0,70
nero: 0,4% 0,21
libav-ffdshow: -5,8% -3,09
------------------------------------------------------------------------
------------------------------------------------------------------------
multiple reference frames
ref1 vs. ref5
720x288 B3-Ref p4x4-i4x4 loop-5 WBP cabac
nero: 10,8% 5,06
libav-mplayer: 10,0% 5,53
elecard: 7,1% 2,95
libav-ffdshow: 7,0% 3,42
ateme: 6,7% 3,85
libav-ffdshow_old: 6,7% 3,22
mainconcept: 5,5% 2,12
moonlight: 5,1% 2,46
------------------------------------------------------------------------
ref1 vs. ref3
720x288 B3-Ref p4x4-i4x4 loop-5 WBP cabac
libav-mplayer: 7,6% 4,59
nero: 7,3% 4,33
ateme: 5,9% 3,97
libav-ffdshow: 5,7% 3,12
elecard: 5,7% 2,69
libav-ffdshow_old: 5,3% 3,02
mainconcept: 4,4% 1,99
moonlight: 2,7% 1,86
------------------------------------------------------------------------
ref3 vs. ref5
720x288 B3-Ref p4x4-i4x4 loop-5 WBP cabac
libav-mplayer: 1,8% 0,94
nero: 1,7% 0,73
moonlight: 1,3% 0,60
elecard: 0,7% 0,26
libav-ffdshow: 0,7% 0,30
libav-ffdshow_old: 0,4% 0,20
mainconcept: 0,4% 0,13
ateme: -0,2% -0,12
------------------------------------------------------------------------
------------------------------------------------------------------------
weighted (bi)prediction
wbp vs. nowbp
720x288 B3-Ref Ref3 p4x4-i4x4 loop-5 cabac
moonlight: 20,0% 11,54
mainconcept: 10,4% 4,27
elecard: 8,4% 3,55
nero: 6,2% 2,80
libav-mplayer: 2,3% 1,22
libav-ffdshow_old: 2,1% 0,96
libav-ffdshow: 1,1% 0,49
ateme: 0,9% 0,46
------------------------------------------------------------------------
wp+wbp vs. nowp+nowbp
640x256 B2 Ref3 p8x8-i4x4 loop-5 cabac
videosoft: 35,6% 19,74
moonlight: 23,6% 16,18
nero: 13,7% 8,16
mainconcept: 12,3% 5,58
elecard: 11,3% 5,37
ateme_old: 9,1% 5,25
libav-mplayer: 4,3% 2,63
libav-ffdshow_old: 2,9% 1,61
ateme: 2,4% 1,65
libav-ffdshow: 0,4% 0,20
------------------------------------------------------------------------
------------------------------------------------------------------------
b-frames
0 B-frames vs. 3 B-frames
720x288 Ref5 p4x4-i8x8 loop-5 WBP cabac
moonlight: 27,7% 17,49
nero: 27,6% 15,52
elecard: 18,3% 8,30
ateme: 16,5% 9,95
libav-mplayer: 14,0% 8,09
libav-ffdshow: 12,5% 6,37
libav-ffdshow_old: 11,8% 5,90
------------------------------------------------------------------------
0 B-frames vs. 2 B-frames
640x256 Ref3 p8x8-i4x4 loop-5 cabac
ateme_old: 20,0% 14,35
mainconcept: 13,6% 7,17
libav-ffdshow: 12,9% 8,08
ateme: 12,7% 9,80
elecard: 12,6% 6,81
nero: 12,4% 8,40
libav-ffdshow_old: 12,3% 7,89
libav-mplayer: 11,4% 7,79
moonlight: 9,5% 7,21
videosoft: 9,0% 5,50
------------------------------------------------------------------------
0 B-frames vs. 1 B-frames
720x288 Ref5 p4x4-i8x8 loop-5 WBP cabac
moonlight: 23,1% 14,55
nero: 21,8% 12,23
elecard: 15,4% 6,97
ateme: 13,6% 8,20
libav-mplayer: 8,4% 4,81
libav-ffdshow: 8,3% 4,25
libav-ffdshow_old: 5,8% 2,92
------------------------------------------------------------------------
1 B-frames vs. 3 B-frames
720x288 Ref5 p4x4-i8x8 loop-5 WBP cabac
nero: 7,5% 3,29
libav-ffdshow_old: 6,3% 2,98
libav-mplayer: 6,2% 3,28
moonlight: 6,1% 2,94
libav-ffdshow: 4,5% 2,12
elecard: 3,5% 1,33
ateme: 3,4% 1,75
------------------------------------------------------------------------
B-ref vs. no B-ref
720x288 B3 Ref5 p4x4-i8x8 loop-5 WBP cabac
nero: 3,4% 1,39
libav-ffdshow: 0,9% 0,38
libav-mplayer: 0,8% 0,41
ateme: 0,7% 0,34
libav-ffdshow_old: 0,4% 0,16
elecard: 0,2% 0,08
moonlight: -0,1% -0,04
------------------------------------------------------------------------
------------------------------------------------------------------------
loop
loop-5 vs. no loop
720x288 B2 Ref3 p8x8-i4x4 cabac
libav-mplayer: 21,0% 14,16
libav-ffdshow: 20,5% 12,28
nero: 19,8% 12,89
libav-ffdshow_old: 18,7% 10,85
elecard: 18,4% 9,34
mainconcept: 15,8% 7,48
videosoft: 14,3% 8,18
moonlight: 6,9% 4,34
ateme: 6,2% 3,80
ateme_old: 4,9% 2,62
720x288 B3-Ref Ref5 p4x4-i8x8 WBP cabac
libav-ffdshow: 20,6% 11,67
libav-mplayer: 19,4% 12,05
libav-ffdshow_old: 17,3% 9,30
elecard: 17,0% 7,62
nero: 16,5% 8,33
ateme: 6,2% 3,37
moonlight: 5,8% 2,80
------------------------------------------------------------------------
loop-5 vs. loop+6
720x288 B3-Ref Ref5 p4x4-i8x8 WBP cabac
ateme: 14,7% 7,45
moonlight: 11,9% 5,44
nero: 7,4% 3,11
libav-ffdshow_old: 5,5% 2,43
libav-mplayer: 5,3% 2,67
libav-ffdshow: 4,4% 2,00
elecard: 4,2% 1,55
------------------------------------------------------------------------
------------------------------------------------------------------------
cabac
cabac vs. no cabac
720x288 B3-Ref Ref5 p4x4-i8x8 loop-5 WBP
libav-mplayer: 16,0% 9,54
libav-ffdshow_old: 15,1% 7,89
libav-ffdshow: 13,8% 7,20
ateme: 10,3% 5,80
moonlight: 9,3% 4,68
nero: 5,2% 2,31
elecard: 4,9% 1,90
------------------------------------------------------------------------
------------------------------------------------------------------------
cabac + noloop vs. loop-5 + nocabac
720x288 B3-Ref Ref5 p4x4-i8x8 WBP
elecard: 12,8% 5,72
nero: 12,0% 6,02
libav-ffdshow: 7,9% 4,47
libav-mplayer: 4,1% 2,51
libav-ffdshow_old: 2,6% 1,41
moonlight: -3,9% -1,88
ateme: -4,5% -2,43
------------------------------------------------------------------------
------------------------------------------------------------------------
resolution
720x288 vs. 640x256
B3-Ref Ref5 p4x4-i8x8 loop-5 WBP cabac
nero: 22,8% 12,43
libav-ffdshow_old: 21,1% 11,87
elecard: 20,4% 9,49
libav-ffdshow: 20,3% 11,48
ateme: 20,2% 12,83
libav-mplayer: 20,1% 12,54
moonlight: 17,2% 9,46
------------------------------------------------------------------------
------------------------------------------------------------------------
custom quant matrix
cqm qmatrix vs. no cqm
720x288 B3-Ref Ref5 p4x4-i8x8 loop-5 WBP cabac
nero: 7,5% 2,68
ateme: 3,5% 1,71
elecard: 3,3% 1,50
moonlight: 1,7% 0,67
------------------------------------------------------------------------
------------------------------------------------------------------------
bond
30th August 2005, 21:39
raw results from which all the above values are derived:
high profile:
x264_hp_2pass_640x256_B3-Ref_Ref5_p4x4-i8x8_loop-5_WBP_cabac
ateme: 63.45
libav-mplayer: 62.45
libav-ffdshow: 56.46
libav-ffdshow_old: 56.33
moonlight: 55.02
nero: 54.47
elecard: 46.62
ateme_old: high profile not supported
mainconcept: high profile not supported
videosoft: high profile not supported
x264_hp_2pass_720x288_B0_Ref5_p4x4-i8x8_loop-5_WBP_cabac
moonlight: 63.09
ateme: 60.23
libav-mplayer: 57.59
nero: 56.17
libav-ffdshow: 50.97
libav-ffdshow_old: 50.20
elecard: 45.35
ateme_old: high profile not supported
mainconcept: high profile not supported
videosoft: high profile not supported
x264_hp_2pass_720x288_B1_Ref5_p4x4-i8x8_loop-5_WBP_cabac
libav-mplayer: 52.78
ateme: 52.03
moonlight: 48.54
libav-ffdshow_old: 47.28
libav-ffdshow: 46.72
nero: 43.94
elecard: 38.38
ateme_old: high profile not supported
mainconcept: high profile not supported
videosoft: high profile not supported
x264_hp_2pass_720x288_B3_Ref5_p4x4-i8x8_loop-5_WBP_cabac
ateme: 50.28
libav-mplayer: 49.50
moonlight: 45.60
libav-ffdshow: 44.60
libav-ffdshow_old: 44.30
nero: 40.65
elecard: 37.05
ateme_old: high profile not supported
mainconcept: high profile not supported
videosoft: high profile not supported
x264_hp_2pass_720x288_B3-Ref_Ref5_p4x4-i8x8_loop-5_WBP
libav-mplayer: 59.45
ateme: 56.42
libav-ffdshow_old: 52.35
libav-ffdshow: 52.18
moonlight: 50.24
nero: 44.35
elecard: 39.03
ateme_old: high profile not supported
mainconcept: high profile not supported
videosoft: high profile not supported
x264_hp_2pass_720x288_B3-Ref_Ref5_p4x4-i8x8_loop-5_WBP_cabac
ateme: 50.62
libav-mplayer: 49.91
moonlight: 45.56
libav-ffdshow: 44.98
libav-ffdshow_old: 44.46
nero: 42.04
elecard: 37.13
ateme_old: high profile not supported
mainconcept: high profile not supported
videosoft: high profile not supported
x264_hp_2pass_720x288_B3-Ref_Ref5_p4x4-i8x8_loop-5_WBP_cabac_cqm-qmatrix
ateme: 48.91
moonlight: 44.89
nero: 39.36
elecard: 35.63
libav-ffdshow: cqm not supported
libav-mplayer: cqm not supported
ateme_old: high profile not supported
mainconcept: high profile not supported
videosoft: high profile not supported
x264_hp_2pass_720x288_B3-Ref_Ref5_p4x4-i8x8_loop+6_WBP_cabac
libav-mplayer: 47.24
ateme: 43.17
libav-ffdshow: 42.98
libav-ffdshow_old: 42.03
moonlight: 40.12
nero: 38.93
elecard: 35.58
ateme_old: high profile not supported
mainconcept: high profile not supported
videosoft: high profile not supported
x264_hp_2pass_720x288_B3-Ref_Ref5_p4x4-i8x8_WBP_cabac
libav-mplayer: 61.96
libav-ffdshow: 56.65
ateme: 53.99
libav-ffdshow_old: 53.76
nero: 50.37
moonlight: 48.36
elecard: 44.75
ateme_old: high profile not supported
mainconcept: high profile not supported
videosoft: high profile not supported
main profile:
x264_mp_2pass_640x256_B3-Ref_Ref5_p4x4-i4x4_loop-5_WBP_cabac
ateme: 65.39
libav-mplayer: 63.71
libav-ffdshow_old: 59.20
moonlight: 55.72
nero: 54.68
libav-ffdshow: 53.37
elecard: 48.74
mainconcept: 46.47
ateme_old: b-ref not supported
videosoft: b-ref not supported
x264_mp_2pass_720x288_B2_Ref3_p4x4-i4x4_cabac
libav-mplayer: 66.33
moonlight: 61.99
libav-ffdshow: 58.73
libav-ffdshow_old: 57.31
ateme: 55.33
videosoft: 54.05
nero: 53.08
elecard: 51.58
ateme_old: 47.86
mainconcept: 47.86
x264_mp_2pass_720x288_B2_Ref3_p8x8-i4x4_cabac
libav-mplayer: 67.33
nero: 65.08
moonlight: 62.67
ateme: 60.95
libav-ffdshow: 59.76
libav-ffdshow_old: 58.15
videosoft: 57.29
ateme_old: 52.99
elecard: 50.84
mainconcept: 47.49
x264_mp_2pass_720x288_B2_Ref3_p8x8-i4x4_loop-5_cabac
moonlight: 58.33
ateme: 57.15
libav-mplayer: 53.17
nero: 52.19
ateme_old: 50.37
videosoft: 49.11
libav-ffdshow: 47.48
libav-ffdshow_log: 47.30
elecard: 41.50
mainconcept: 40.01
x264_mp_2pass_720x288_B2-Ref_Ref3_p4x4-i4x4_cabac
libav-mplayer: 65.36
moonlight: 62.09
libav-ffdshow: 59.51
libav-ffdshow_old: 57.33
ateme: 57.29
nero: 55.63
elecard: 52.50
mainconcept: 49.27
ateme_log: b-ref not supported
videosoft: b-ref not supported
x264_mp_2pass_720x288_B3-Ref_Ref1_p4x4-i4x4_loop-5_WBP_cabac
ateme: 57.04
libav-mplayer: 55.45
libav-ffdshow: 48.99
libav-ffdshow_old: 48.08
moonlight: 47.96
nero: 46.82
elecard: 41.28
mainconcept: 38.78
ateme_old: b-ref not supported
videosoft: b-ref not supported
x264_mp_2pass_720x288_B3-Ref_Ref3_p4x4-i4x4_loop-5_WBP_cabac
ateme: 53.07
libav-mplayer: 50.86
moonlight: 46.10
libav-ffdshow: 45.87
libav-ffdshow_old: 45.06
nero: 42.49
elecard: 38.59
mainconcept: 36.79
ateme_old: b-ref not supported
videosoft: b-ref not supported
x264_mp_2pass_720x288_B3-Ref_Ref3_p4x4-i4x4_loop-5_cabac
moonlight: 57.64
ateme: 53.53
libav-mplayer: 52.08
libav-ffdshow: 46.36
libav-ffdshow_old: 46.02
nero: 45.29
elecard: 42.14
mainconcept: 41.06
ateme_old: b-ref not supported
videosoft: b-ref not supported
x264_mp_2pass_720x288_B3-Ref_Ref5_p4x4-i4x4_loop-5_WBP_cabac
ateme: 53.19
libav-mplayer: 49.92
libav-ffdshow: 45.57
moonlight: 45.50
libav-ffdshow_old: 44.86
nero: 41.76
elecard: 38.33
mainconcept: 36.66
ateme_log: b-ref not supported
videosoft: b-ref not supported
x264_mp_2pass_640x256_B0_Ref3_p8x8-i4x4_loop_cabac
ateme: 77.31
moonlight: 75.64
ateme_old: 71.90
libav-mplayer: 68.44
nero: 67.89
libav-ffdshow_old: 64.37
libav-ffdshow: 62.63
videosoft: 60.95
elecard: 54.19
mainconcept: 52.72
videosoft_old: 31.87
x264_mp_2pass_640x256_B2_Ref5_p4x4-i4x4_loop_cabac
ateme: 67.19
libav-mplayer: 55.70
moonlight: 54.29
libav-ffdshow_old: 53.07
libav-ffdshow: 52.87
ateme_log: 49.60
nero: 47.79
elecard: 43.91
mainconcept: 41.60
videosoft: 38.97
videosoft_old: 24.25
nero_2pass-777kbps_Cabac_Deblock-5-adapt_B2_Ref3_noWPred_Qpel_p8x8_cartoon_psy2_extra
moonlight: 68.43
ateme: 67.51
libav-mplayer: 60.65
nero: 59.49
ateme_old: 57.55
libav-ffdshow_old: 56.48
videosoft: 55.45
libav-ffdshow: 54.55
elecard: 47.38
mainconcept: 45.55
nero_2pass-777kbps_Cabac_Deblock-5-adapt_B2_Ref3_WPred+wbp_Qpel_p8x8_cartoon_psy2_extra
ateme: 65.86
libav-mplayer: 58.02
libav-ffdshow_old: 54.87
libav-ffdshow: 54.35
ateme_old: 52.30
moonlight: 52.25
nero: 51.33
elecard: 42.01
mainconcept: 39.97
videosoft: 35.71
baseline profile:
x264_bp_720x288_B0_Ref5_p8x8-i4x4_loop-5_wbp
moonlight: 75.34
libav-mplayer: 72.88
ateme: 72.28
ateme_log: 70.63
videosoft: 64.01
libav-ffdshow_old: 63.47
libav-ffdshow: 61.83
nero: 61.80
elecard: 52.02
mainconcept: 51.16
videosoft_old: 33.46
Sharktooth
30th August 2005, 21:43
:eek:
sorry, i need to buy a pair of googles...
celtic_druid
31st August 2005, 03:20
@bond, did you use my ffdshow build as is? Because the libavcodec is built with gcc and generic flags. It might be interesting to see the same test with a version compiled with flags specific for your CPU, since as you say ffdshow is not that much slower than moonlight.
bill_baroud
31st August 2005, 10:53
Thanks for this usefull test !
I would just say that, for people like me who didn't go read the thread about libavcodec speed measurement, your results numbers mean absolutly nothing... You could write somewhere that you are talking about Frame Per Second (and not cpu occupation or whatelse)... That puzzled me until i decided to read the thread linked in your post.
alexcyn
31st August 2005, 14:03
Hi Bond.
It seems you are using very-very old videosoft decoder version 2.0.2.3. New much faster version 2.2 is available for already long time. Free evaluation of VSS h264 DirectShow decoder filter 2.2 is available here:
http://www.vsofts.com/h264/decoders.html
See item "Installation". Or direct link to installer:
http://www.vsofts.com/h264/pub/vssh3dec.exe
The DirectShow decoder should support any format, not only AVI.
Commercial version of the decoder is included into both Base ($20) and Main ($99) codec 2.3 consumer packages, so the price of the decoder is only $20.
The newest VSS H264 decoder 3.0 (currently available only in Professional products) supports all features from Main & Baseline profiles as well as High Profile. Also it has performance ~15% better than version 2.2.
By the way, what hardware (CPU, memory) are you using?
=Alexey Doilnitsyn, VSS Inc.
Manao
31st August 2005, 15:22
i didnt know what to expect from the ateme decoder and i also have to say i wasnt really positively surprised as it performed in the middle or slow
i have to note that ateme is working on a new decoder also supporting high profile, which i wasnt able to test, i hope they will also enhance decoding speedfor the main profile that are decoded by ffdshow / moonlight / nero / ateme, the average speeds are respectively 55.93, 61.94, 56.69 and 54.65. So indeed, moonlight is above the others, but the others rank the same.
Also, I wouldn't have used average speeds but, inverse of average of inverse speeds ( hence, the sum of decoding time ).
Finally, including mplayer is great for promotting mplayer ( or vlc, btw ), but serves no purposes in your comparison : you're comparing decoders, not players ( and -vo null is almost cheating if it does what I think :p )
IgorC
31st August 2005, 16:41
Here Nero decoder is still faster than ffdshow+haali and mplayer. Maybe it depends of settings. I also use hard settings (high profile, ref 8-16, weight, bframes 2-3, high values for mvrange for x264/H.264 etc.) And probably it depens of CPU. In my case SSE2.
hpn
31st August 2005, 16:43
and -vo null is almost cheating if it does what I think :p
"-vo null" means mplayer will output no frame to a video device. I've never tried the Chegepuga filter that Bond has used with the paid decoders, but I guess It does the same thing (plus the fps measuring itself), so unless I'm missing something I don't see any cheating here.
akupenguin
31st August 2005, 16:53
baseline profile:
x264_720x288_B0_Ref5_p8x8-i4x4_loop-5_wbp_cabac
cabac isn't baseline. (wpred isn't either, but wbp without B-frames doesn't matter)
Also, how about lossless and interlacing in the feature comparison?
Manao
31st August 2005, 17:05
Well, I've got a decoder available both in a ds filter and in a standalone application that can behave like mplayer -vo null. The ds overhead for the filter is roughly a copy of a picture, which hardly matters at the framerate i'm testing. In one case ( standalone ), i get 48 fps, in the second ( ds filter + chegepuga ), i get 42 fps, so it's a 12.5 % speed gain ( or 11% speed loss :p ).
DirectShow itself is responsible for that loss ( which is huge, because I've got a 2800+ with a fairly fast memory ).
So perhaps it's not -vo null that creates the difference, yet comparing ds filters to mplayer isn't that useful when you want to compare decoders.
SeeMoreDigital
31st August 2005, 17:12
Great work bond....
It will interesting to see if things change as time rolls by and refinements are made....
Cheers
hworldjj
1st September 2005, 06:45
Thanks! Good reading
bond
1st September 2005, 20:34
first of all thx for all the responses and the interest! :)
@bond, did you use my ffdshow build as is? Because the libavcodec is built with gcc and generic flags. It might be interesting to see the same test with a version compiled with flags specific for your CPU, since as you say ffdshow is not that much slower than moonlight.yep i used your build as is. if you could make a build compiled for my pentium3 866mhz i would be happy to test it too :)
your results numbers mean absolutly nothing... You could write somewhere that you are talking about Frame Per Second (and not cpu occupation or whatelse)... That puzzled me until i decided to read the thread linked in your post.indeed, i will add this
It seems you are using very-very old videosoft decoder version 2.0.2.3. New much faster version 2.2 is available for already long time. Free evaluation of VSS h264 DirectShow decoder filter 2.2 is available here:i used indeed an old version, and that was mainly done because i did the test to find out what decoder i could use for my encodes, so i used the last version you released which was unlimited
i simply wasnt really interested in testing a decoder which becomes useless for me after 30 days :(
The DirectShow decoder should support any format, not only AVI.hm i think your decoder works with the VSSH and H264 fourcc. actually i dont know a splitter which outputs this (eg from .mp4 or .mpg) your decoder could connect to, so theoretically your decoder can of course work with any format, but practically its not so easy till now
Commercial version of the decoder is included into both Base ($20) and Main ($99) codec 2.3 consumer packages, so the price of the decoder is only $20.so the 20$ baseline version includes a main profile decoder?
The newest VSS H264 decoder 3.0 (currently available only in Professional products) supports all features from Main & Baseline profiles as well as High Profile. Also it has performance ~15% better than version 2.2.sounds indeed very powerful and i would love to test it, its just that the 30days limit makes it pretty useless for me, but if you send me an unlimited copy i will of course test it ;) :)
By the way, what hardware (CPU, memory) are you using?a good old pentium3 866mhz
Alexey Doilnitsyn, VSS Inc. great to have you around on doom9!
Also, I wouldn't have used average speeds but, inverse of average of inverse speeds ( hence, the sum of decodinghm the clips all have the same lenght, in what way would the sum tell us more?
Finally, including mplayer is great for promotting mplayer ( or vlc, btw ), but serves no purposes in your comparison : you're comparing decoders, not players ( and -vo null is almost cheating if it does what I think :p )well the point is you have to see this from a users point of view and the user asks "what player should i play my avc clips with"
for the user it doesnt matter what codec interface stands behind it, be it directshow or whatever
i am actually also planning to add decoding of qt7 to the comparison, but i am simply too lazy to set it up
the thing is if the commercial decoders are limited to directshow, its the commercials decoders problem and not libavcodec's, which is simply available in a form which allows it to be used in superior platforms than directshow, that should be honored
and as you saw i also included ffdshow and it still performed great
"-vo null" means mplayer will output no frame to a video device. I've never tried the Chegepuga filter that Bond has used with the paid decoders, but I guess It does the same thing (plus the fps measuring itself), so unless I'm missing something I don't see any cheating here. chegepuga is also a null renderer with fps measurement, exactly the same as what mplayer does imho
cabac isn't baseline. (wpred isn't either, but wbp without B-frames doesn't matter)my fault, the baseline clip i tested of course didnt use cabac, altough i noted it (must be a copy paste error or so)
actually i noticed that x264 enables the wbp flag in the pps altough it sets the baseline profile in the sps correctly
Also, how about lossless and interlacing in the feature comparison?yep i thought about that too, i was too lazy to do it actually, also i lacked the time to really make some good comparable interlaced clips with the reference
maybe i will simply test the decoders capabilities on the clips i have lying around without speed measurement
Nil Einne
1st September 2005, 21:05
I could be wrong but I was under the impression not all decoders are equal quality-wise. Frequently they use estimations etc to improve speed which is fine but when quality can very IMHO a straight out speed test is not so meaningful for many people. I could make a decoder that is very very fast but the output it so crap it isn't worth it. Of course I'm aware quality issues are a problem since it's so subjective so instead, I would recommend you get the reference implementation of h.264 which I assume is completely accurate and then do a mathematical comparison. Of course, this can be a bit misleading since a smart decoder may have higher mathmatical difference but lower noticable difference but it's better then nothing.
Manao
1st September 2005, 21:11
No, none of these codecs do such a thing. Not decoding picture as they should leads immediately to errors that are easily spottable.
Only libavcodec got a patch, very recently, that allowed to disable deblocking at very low qps, but it's not enabled by default, and it's highly recommended not to use it.
Nil Einne
1st September 2005, 22:41
Okay thanks for clarifying. I was always under the impression that there was some (minor) differences in the output quality of non post-processing decoders due to factors such as estimation, dropping least significant bits in some cases, different interpretations of the algorythm etc. Now I know. So basically you can take any decoder and use it to save an uncompressed raw video files and the file would be exactly the same! That is good :-)
Sergey A. Sablin
2nd September 2005, 06:26
chegepuga is also a null renderer with fps measurement, exactly the same as what mplayer does imho
Bond, I dunno what exactly does mplayer with "-vo null" option, but I exactly know what does chegepuga - DShow decoder should copy output frame into this renderer. So if mplayer with "-vo null" doesn't require this from decoder, then it is the difference for the test, cause for real playback decoders in mplayer also should do the copy.
In that way real performance can be different from what you have measured.
Does somebody know how mplayer process with this option? with or without the copy?
bond
2nd September 2005, 10:42
well the mplayer devs recommend when wanting to do benchmarks to use the -vo null and -benchmark options, so i assume the output values are useable!?
Sergey A. Sablin
2nd September 2005, 11:00
well the mplayer devs recommend when wanting to do benchmarks to use the -vo null and -benchmark options, so i assume the output values are useable!?
Sure it usable. But if mplayer doesn't require output frame copy, then this values are only usable to compare different decoders inside mplayer - not with decoders inside DShow environment, do you agree?
But if mplayer also require to do this copy, then results are also comparable to DS filters.
(I just mean that if devs recommend this way to measure performance than we can't say that it measure like another tools)
akupenguin
2nd September 2005, 11:01
Does somebody know how mplayer process with this option? with or without the copy?
libavcodec returns a pointer into the same decoded picture buffer used for inter prediction. Then real vos copy it to video memory, and -vo null doesn't. (In some formats (including ASP but not yet implemented for H.264), frames that don't need to be kept (B-frames or Intra-only) can be decoded directly into video memory, or incrementally into a video filter if some filtering is performed.)
You're saying that in DShow, instead the vo passes a pointer to the decoder, and real vos pass a pointer to video memory while chegepuga passes a pointer to some dummy buffer that's never read?
Sergey A. Sablin
2nd September 2005, 11:16
libavcodec returns a pointer into the same decoded picture buffer used for inter prediction. Then real vos copy it to video memory, and -vo null doesn't. (In some formats (including ASP but not yet implemented for H.264), frames that don't need to be kept (B-frames or Intra-only) can be decoded directly into video memory, or incrementally into a video filter if some filtering is performed.)
You're saying that in DShow, instead the vo passes a pointer to the decoder, and real vos pass a pointer to video memory while chegepuga passes a pointer to some dummy buffer that's never read?
I mean that in DShow decoder after frame decoding should copy the output frame into another location which is indicated by pointer passed to decoder by downstream filter (which is in this case chegepuga and in real playback is video renderer), but I don't know whether mplayer requires this too.
akupenguin
2nd September 2005, 11:20
MPlayer does not require that. Neither -vo null nor the real playback do that copy.
alexcyn
2nd September 2005, 14:01
To: Bond
> i used indeed an old version, and that was mainly done because i
> did the test to find out what decoder i could use for my encodes,
> so i used the last version you released which was unlimited
unlimited free version may violate h264 patents, thats why we had to limit it by 30 days :-(.
> i simply wasnt really interested in testing a decoder which becomes
> useless for me after 30 days :(
we will think, may be it makes sense to give you unlimited version for testing.
> hm i think your decoder works with the VSSH and H264 fourcc.
> actually i dont know a splitter which outputs this (eg from .mp4 or .mpg)
> your decoder could connect to, so theoretically your decoder can of
> course work with any format, but practically its not so easy till now
Yes, version 2.0 was limited by 2 fourccs, but version 2.2 accepts any fourcc (type=video, subtype=[any]).
> so the 20$ baseline version includes a main profile decoder?
YES
=Alexei, VSS
bond
2nd September 2005, 14:06
thx for the info! :)
Haali
2nd September 2005, 14:28
I mean that in DShow decoder after frame decoding should copy the output frame into another location which is indicated by pointer passed to decoder by downstream filter (which is in this case chegepuga and in real playback is video renderer), but I don't know whether mplayer requires this too.
DShow can also act in the same way as mplayer. In dshow buffer management is done in some nontrivial way. First you negotiate an allocator with the downstream filter, obviously in case of video renderer you want to use the renderer's allocator since it knows how to work with video memory. Second you set the allocator parameters like number of buffers, alignment and buffer size. Then during playback you call allocator's GetBuffer() and filter's Receive() after you are done processing. If you specify AM_GBF_NOTASYNCPOINT in GetBuffer() call, then it will return the buffer with unchanged contents from the previous frame. So to avoid extra copying overhead you can request one buffer from the allocator and use AM_GBF_NOTASYNCPOINT. In this case renderers like overlay mixer will return a pointer to the overlay's video memory. Most decoders do it that way, buf ffdshow still performs an extra copy internally afaik.
TheBashar
3rd September 2005, 03:35
I could be wrong but I was under the impression not all decoders are equal quality-wise.
I concur on this point. After reading your comparison, I tried out the Moonlight-Elecard MPEG Player that Bond linked to. In testing with some of my AVC high-profile encodes, I found it periodically produced dark macroblocks which did not appear when using nero's decoder.
Being fast is great, but not at the cost of decoding artifacts.
Here's a sample of what I'm talking about:
http://img352.imageshack.us.nyud.net:8090/img352/9422/mempeg43dn.png
stephanV
3rd September 2005, 06:53
This is a bug, not a difference in quality. :)
bobololo
4th September 2005, 01:19
For those who are interested in decoding filter benchmarking, Haali was kind enough to write a little dshow application extremely convenient for this purpose.
It's available at the url: http://haali.cs.msu.ru/mkv/timeCodec.exe
It requires the latest version of Haali Media Splitter available here (http://haali.cs.msu.ru/mkv/MatroskaSplitter.exe).
With some decoding filters (like ateme's one), it may requires an additionnal filter you can find here (you have to register it manually): http://haali.cs.msu.ru/mkv/m2r.dll
I did a quick test using a clip posted during ateme hp beta (batman-ateme-3500k-hp.mp4 (ftp://mood.ateme.com/beta/batman-ateme-3500k-hp.mp4)) and I got those figures:
nero (nve 3.1.0.16): 46.4 fps
moonlight (0.9.0 build 50208 beta): 47.6 fps
ffdshow (20050822 - cd build): 48.1 fps
ateme (2.2.1.0): 56.5 fps
The tests were done on a Pentium 4 @ 3.0 GHz (with HT).
For all decoders, timeCodec uses Haali Splitter to parse the file and to feed the decoder filter.
This test used an updated version of ateme decoder filter that be provided in the next beta release.
Many thanks to Haali again for his great work !
ps: I've tried moonlight decoder on other clips, and I had some decoding issues like with this one (ftp://mood.ateme.com/beta/cinderella-ateme-3000k-hp.mp4). Also it doesn't seem to decode mbaff correctly ?
EDIT: I found my problem with ffdshow, the postprocessing was enabled. It's much better now :) and I udpdated all the results with a 3.0 GHz cpu measurements.
Sergey A. Sablin
4th September 2005, 09:16
MPlayer does not require that. Neither -vo null nor the real playback do that copy.
Well, just a two questions:
1. Did you mean that decoders use video memory for storing reference pictures? How about reading from video memory?
2. Did you mean that all decoders use YV12 colorspace for rendering? YV12 is slower for rendering than YUY2 and UYVY on all video cards I know.
DShow can also act in the same way as mplayer. In dshow buffer management is done in some nontrivial way. First you negotiate an allocator with the downstream filter, obviously in case of video renderer you want to use the renderer's allocator since it knows how to work with video memory. Second you set the allocator parameters like number of buffers, alignment and buffer size. Then during playback you call allocator's GetBuffer() and filter's Receive() after you are done processing. If you specify AM_GBF_NOTASYNCPOINT in GetBuffer() call, then it will return the buffer with unchanged contents from the previous frame. So to avoid extra copying overhead you can request one buffer from the allocator and use AM_GBF_NOTASYNCPOINT. In this case renderers like overlay mixer will return a pointer to the overlay's video memory. Most decoders do it that way, buf ffdshow still performs an extra copy internally afaik.
Did you mean decoding directly to video memory? If yes - than try to use at least two decoders with such technique. It'll be very interesting.
Reference pictures are also can't be decoded into video memory, so they need a copy anyway.
Haali
4th September 2005, 10:10
What I meant is decoders usually request only one buffer from the allocator, I don't know if they do an extra copy or not.
akupenguin
5th September 2005, 12:36
Well, just a two questions:
1. Did you mean that decoders use video memory for storing reference pictures? How about reading from video memory?
libavcodec supports:
Store picture in application specified pointer, usually video memory. Used for non-referenced pictures when no filtering is needed. If you use this mode for a referenced picture, it will try to read it back from the buffer.
Callback after each row of macroblocks with a pointer to the decoded slice. Used for referenced pictures or simple filters.
Return a pointer to the decoded frame. Used for MEncoder, -vo null, or complex filters.
The decoder does no copies in any of those. The -vo (other than null) does a copy in (2) and (3), or filters/encoders may use the (read-only) buffer as is.
2. Did you mean that all decoders use YV12 colorspace for rendering? YV12 is slower for rendering than YUY2 and UYVY on all video cards I know.
Yes, all decoders output the same pixel format as the video actually stores (so usually YV12). All video cards I know of are fast enough to do scaling + YV12->RGB for any video resolution they support at all, so you can free a little CPU time by not doing software YV12->YUY2 conversion. If for some reason you or the -vo need another pixel format, then MPlayer will insert a conversion filter (or maybe it's not always automatic; anyway, it's equivalent to "-vf scale").
Sergey A. Sablin
5th September 2005, 14:02
libavcodec supports:
[list=1] Store picture in application specified pointer, usually video memory. Used for non-referenced pictures when no filtering is needed. If you use this mode for a referenced picture, it will try to read it back from the buffer.
reading from video memory is very low -> it is very unefficient to use it when postprocessing is used. That means it is very unefficient for H.264 (at least), cause deblocking are used for most cases.
Yes, all decoders output the same pixel format as the video actually stores (so usually YV12). All video cards I know of are fast enough to do scaling + YV12->RGB for any video resolution they support at all, so you can free a little CPU time by not doing software YV12->YUY2 conversion. If for some reason you or the -vo need another pixel format, then MPlayer will insert a conversion filter (or maybe it's not always automatic; anyway, it's equivalent to "-vf scale").
For modern video cards yes. But for old YV12 is much slower.
Also decoding directly to video memory (i.e. in YV12) for some video cards produce many errors - try to use two instances of ffdshow with YV12 output enabled to decode some stream and you will see these artifacts. (I've tried mpeg-2 sequence with matrox parhelia via libavcodec and libmpg2. You can either use two instances in one process or in two different processes - it doesn't matter)
I don't want to say that decoding directly to video memory or rendering in YV12 format is slower everywhere, but in most cases it is still either slower or buggy and comparing decoding methods that doesn't work correctly in all cases is not so correct.
bond
5th September 2005, 23:16
updated some values with a new version of ateme:
ateme-new: ateme mp4 parser 1.2.5.3 / ateme decoder 2.2.1.0
HIGH PROFILE
x264_hp_2pass_720x288_B0_Ref5_p4x4-i8x8_loop-5_WBP_cabac.mp4
moonlight: 63.09
ateme-new: 60.23
libav-mplayer: 57.59
nero: 56.17
libav-ffdshow: 50.20
x264_hp_2pass_720x288_B3-Ref_Ref5_p4x4-i8x8_loop-5_WBP.mp4
libav-mplayer: 59.45
ateme-new: 56.42
libav-ffdshow: 52.35
moonlight: 50.24
nero: 44.35
x264_hp_2pass_720x288_B3-Ref_Ref5_p4x4-i8x8_loop-5_WBP_cabac_cqm-qmatrix.mp4
ateme-new: 48.91
moonlight: 44.89
nero: 39.36
libav-ffdshow: cqm not supported
libav-mplayer: cqm not supported
x264_hp_2pass_720x288_B3-Ref_Ref5_p4x4-i8x8_loop+6_WBP_cabac.mp4
libav-mplayer: 47.24
ateme-new: 43.17
libav-ffdshow: 42.03
moonlight: 40.12
nero: 38.93
x264_hp_2pass_720x288_B3-Ref_Ref5_p4x4-i8x8_WBP_cabac.mp4
libav-mplayer: 61.96
ateme-new: 53.99
libav-ffdshow: 53.76
nero: 50.37
moonlight: 48.36MAIN PROFILE
x264_mp_2pass_720x288_B2_Ref3_p8x8-i4x4_cabac.mp4
libav-mplayer: 67.33
nero: 65.08
moonlight: 62.67
ateme-new: 60.95
libav-ffdshow: 58.15
ateme: 52.99
mainconcept: 47.49
x264_mp_2pass_720x288_B2-Ref_Ref3_p4x4-i4x4_cabac.mp4
libav-mplayer: 65.36
moonlight: 62.09
libav-ffdshow: 57.33
ateme-new: 57.29
nero: 55.63
mainconcept: 49.27
ateme: b-ref not supported
nero_2pass-777kbps_Cabac_Deblock-5-adapt_B2_Ref3_WPred+wbp_Qpel_p8x8_cartoon_psy2_extra.mp4
ateme-new: 65.86
libav-mplayer: 58.02
libav-ffdshow: 54.87
ateme: 52.30
moonlight: 52.25
nero: 51.33
mainconcept: 39.97BASELINE PROFILE
x264-r285_bp_720x288_B0_Ref5_p8x8-i4x4_loop-5_wbp_cabac_mp4box.mp4
moonlight: 75.34
libav-mplayer: 72.88
ateme-new: 72.28
ateme: 70.63
libav-ffdshow: 63.47
nero: 61.80
mainconcept: 51.16
videosoft: 33.46very good, but also varying results, often very fast, but also often on par with other good decoders
Haali
6th September 2005, 07:53
Also decoding directly to video memory (i.e. in YV12) for some video cards produce many errors
That's because they support only one YV12 overlay, so two instances conflict when using the same hardware resource. AFAIK overlay is the only place where planar YV12 is supported by video hardware, even modern cards don't support YV12 textures.
bond
6th September 2005, 13:40
and some more findings with a new videosoft decoder:
videosoft-new: m$ avi parser 6.5.1.902 / videosoft decoder 2.3.1.5
MAIN PROFILE
x264_mp_2pass_720x288_B2_Ref3_p4x4-i4x4_cabac.mp4
libav-mplayer: 66.33
moonlight: 61.99
libav-ffdshow: 57.31
videosoft-new: 54.05
nero: 53.08
ateme: 47.86
mainconcept: 47.86
x264_mp_2pass_720x288_B2_Ref3_p8x8-i4x4_cabac.mp4
libav-mplayer: 67.33
nero: 65.08
moonlight: 62.67
ateme-new: 60.95
libav-ffdshow: 58.15
videosoft-new: 57.29
ateme: 52.99
mainconcept: 47.49
x264_mp_2pass_720x288_B2_Ref3_p8x8-i4x4_loop-5_cabac.mp4
moonlight: 58.33
libav-mplayer: 53.17
nero: 52.19
ateme: 50.37
videosoft-new: 49.11
libav-ffdshow: 47.30
mainconcept: 40.01
x264_mp_2pass_640x256_B0_Ref3_p8x8-i4x4_loop_cabac.mp4
moonlight: 75.64
ateme: 71.90
libav-mplayer: 68.44
nero: 67.89
libav-ffdshow: 64.37
videosoft-new: 60.95
mainconcept: 52.72
videosoft: 31.87
x264_mp_2pass_640x256_B2_Ref5_p4x4-i4x4_loop_cabac.mp4
libav-mplayer: 55.70
moonlight: 54.29
libav-ffdshow: 53.07
ateme: 49.60
nero: 47.79
mainconcept: 41.60
videosoft-new: 38.97
videosoft: 24.25
BASELINE PROFILE
x264-r285_bp_720x288_B0_Ref5_p8x8-i4x4_loop-5_wbp_cabac_mp4box.mp4
moonlight: 75.34
libav-mplayer: 72.88
ateme-new: 72.28
ateme: 70.63
videosoft-new: 64.01
libav-ffdshow: 63.47
nero: 61.80
mainconcept: 51.16
videosoft: 33.46the new vss decoder can definitely keep up in some cases with the others, but is also worse than the others in other cases
its definitely much better than the first version is tested (the last one freely available)
i also found some things:
- it can connect to the nero and haali parser, but doesnt show anything when playing (outputs MPEG2Video)
- it works fine with the avi parser (outputs H264 and VSSH)
- it also works with the ateme mp4 parser (outputs H264)
- b-ref crash the decoder
- high profile is not supported, altough videosoft already has a hp decoder, which i wasnt able to test
saratoga
21st September 2005, 20:07
a good old pentium3 866mhz
That CPU lacks SSE2 which is the new standard for floating point calculations on present x86 hardware. Its reasonable to think the SSE2 code is probably better developed and supported as x87 is about to be depreciated. Its possible x87 fp is provided purely as legacy support in some codecs, particularly ones like Nero which are aimed at commerical use.
It would be interesting to see if the relative results change any when you run the same test on an SSE2 capable processor.
akupenguin
21st September 2005, 21:16
Codecs have nothing to do with floating-point. There aren't any x87 or SSE2 instructions at all in libav's H.264 decoder.
Manao
21st September 2005, 21:55
Indeed.
And, furthermore, only amd64 and P4 have SSE2, and, for most amd64 ( if not all ), SSE2 ops are as slow as their MMX counterparts.
So, basically, SSE2 helps only for P4 and very recent amd64.
saratoga
22nd September 2005, 01:44
Codecs have nothing to do with floating-point. There aren't any x87 or SSE2 instructions at all in libav's H.264 decoder.
You're correct. Change x87 to MMX and my post will make a little more sense.
And, furthermore, only amd64 and P4 have SSE2, and, for most amd64 ( if not all ), SSE2 ops are as slow as their MMX counterparts.
So you mean they're only fast for the overwhelming majority of machines ;)
I'm still intersted on figures from a newer P4 or A64. While the P3 numbers are relevent, a control would be nice.
bond
16th October 2005, 17:34
ok i finally found the time to finalize the comparison
the following additions have been made:
- ateme
- videosoft
- ffdshow with pentium3 specific compile flags (thx celtic_druid!)
- elecard
- values showing how the different decoders perform on specific coding tools (maybe gives devs a hint on what needs tuning)
i hope you guys find it interesting
Manao
16th October 2005, 18:05
The way you're computing decoder's efficiency for each tools is flawed : you have to measure the loss of time, not the loss of fps. Because losing 1sec on a 5 sec decoding time at 100 fps means a drop of 16.66 fps, while losing 1 sec on a 5 sec decoding time at 50 fps means a drop of 8.33 fps.
Enabling a feature almost always add a constant time, not a proportionnal fps loss.
Also, do consider that incertitude on fps figures you're giving are at least 2 fps.
Finally, what was the average quantizer of the clip you used ( out of curiosity, i think it might explain why ateme's decoder goes faster when deblocking -5 is enabled... )
bond
16th October 2005, 19:20
The way you're computing decoder's efficiency for each tools is flawed : you have to measure the loss of time, not the loss of fps. Because losing 1sec on a 5 sec decoding time at 100 fps means a drop of 16.66 fps, while losing 1 sec on a 5 sec decoding time at 50 fps means a drop of 8.33 fps.
Enabling a feature almost always add a constant time, not a proportionnal fps loss.
Also, do consider that incertitude on fps figures you're giving are at least 2 fps.right, i now ranked the decoders by the % by which the decoding speed decreases when an additional features is enabled
the ranking didnt change much, as most decoders perform in the same range (additionally there is the incertitude you mentioned)
Finally, what was the average quantizer of the clip you used ( out of curiosity, i think it might explain why ateme's decoder goes faster when deblocking -5 is enabled... ) the quants of the files are around 20
bond
16th October 2005, 22:54
almost forgot to mention: another interesting thing i found was that ffdshow and/or libavcodec seems to have gotten slower since the last test
and that altough celtic_druid made a pentium3 specific build, which should give faster results than the old ffdshow build i used the first time (without p3 specific stuff)
still the new ffdshow showed on not so few samples worse results than with the old version...
bond
23rd October 2005, 21:38
ok i now added info about interlacing support of the decoders:
ateme, nero and elecard support fields-only, paff and mbaff (also with i8x8 of the high profile)
mainconcept theoretically does this too, but seems to be very buggy as it shows lots of artefacts with all modes
moonlight supports fields-only and paff (also with i8x8 of the high profile), but shows artefacts with paff too
videosoft supports fields-only
libavcodec doesnt support interlacing at all
Beave
29th October 2005, 01:52
Could you be interested in benchmarking some HD content? 720p in High Profile is not playable on my AMD64 3000+. I wonder if there is some filter fast enough for this.
Manao
29th October 2005, 06:11
There is. But teasing bond like that is a shame, since he only has a p3 866. But, of course, nothing prevents you from doing the test yourself.
bond
29th October 2005, 12:13
yep testing more and higher resolutions (i mainly tested D1) would be very interesting, but i am not the right one to talk to cause a pentium3 866mhz might not be representative at all for this
Inventive Software
31st October 2005, 14:58
There is. But teasing bond like that is a shame, since he only has a p3 866. But, of course, nothing prevents you from doing the test yourself.
Hey! I got a Celeron 800. I struggle to play H.264 DVD resolution content, let alone HD resolutions.
But yeah, like Manao said, there's nothing stopping you testing it yourself! :D
vBulletin® v3.8.11, Copyright ©2000-2026, vBulletin Solutions Inc.