NikosD
16th November 2011, 21:59
Latest update with LAV Video x64 0.64 in DXVA native and pure decode mode, using latest ASICs like VP7 from Nvidia GTX 960 and QuickSync 3 from Haswell.
Added also AMD Polaris RX 470 results and just one result of Pascal GTX 1060 VP8 decoder.
All Intel CPUs from Haswell to Kabylake have exactly the same 4K HW H.264 decoder, they differ only in clock speed.
Take a look here:
http://forum.doom9.org/showthread.php?p=1712350#post1712350
I've recently flashed my Radeon 5750 BIOS with 6750 BIOS.
The two cards use the same UVD2.2
But it seems that 6750 BIOS on a 5750, can lead to a UVD2.2 overclocking.
The default 5750 BIOS put UVD2.2 in standard UVD mode at Core/GPU = 400MHz / 900MHz
The 6750 BIOS on a 5750 card, put UVD2.2 in 3D mode at Core/GPU = 710MHz / 1160MHz
So I have two UVD2.2 systems benchmarked.
One plain UVD2.2 (400/900) and one UVD2.2 OC (710/1160)
I did my tests with the new DXVA checker x86 v2.7.0 http://bluesky23.yukishigure.com/en/index.html
Two systems tested:
1) My signature system:
Win 7 x64 SP1 - C2D@2.83 GHz - Radeon (6)750 - Catalyst 12.1 preview,
RAM configuration for AMD system: (mostly for DXVA-CB comparisons)
4GB (2 x 2GB) of DDR2 at FSB: DRAM = 1:1
Speed = 4-4-4-12@566 MHz (283x2)
2) Intel/ Nvidia system:
Win 7 SP1 x86 - Core i5-2400 (3.1GHz) - Geforce GT 440 (DDR5) - Nvidia beta 290.53 - Intel HD 2000 - Intel drivers v.2622
RAM configuration for Intel/ Nvidia system: (mostly for DXVA-CB comparisons, QuickSync decoder)
4GB (2 x 2GB) of DDR3 at FSB: DRAM = 1:5
Speed = 9-9-9-24@1338 MHz (669x2)
The decoders used are:
CoreAVC 3.0.1 (both modes - DXVA native, NVCUVID)
LAV Video 0.47 (in all modes - DXVA2 native, DXVA2 copy-back, NVCUVID, QS)
MS DS/MFT
FFDShow v4322 (QS)
For VC-1/ WMV3 I used the AMD Playback Decoder MFT and for CPU results I used the built-in WMVideo Decoder DMO (because is faster than LAV slow VC-1/WMV decoder)
I used five Reference H.264 files from here:
http://forum.doom9.org/showthread.php?t=159486
and I added five new reference files.
You can find every sample posted (from 1 to 10) here:
ftp://helpedia.com/pub/multimedia/x264/testvideos/
6.Avatar-1080p60fpsRef4-44.9Mbps
7.Vortexx_1088p24fpsRef3-109Mpbs
8.Birds_1080p24fpsRef4-112Mbps
9.Ducks.Take.Off.1080p30fpsRef5-108Mbps
10.Crowd.Run.1080p25Ref4-116Mbps
Also I used VC-1 and WMV3 files from here:
http://forum.doom9.org/showthread.php?t=156660
For CPU results (Core 2 Duo - Core i5) I used LAV Video 0.47.
Every benchmark mode used EVR renderer.
The results:
First is the Video Processor - QuickSync (QS), UVD2.2, VP4, CPU etc
Second is the decoder - MS DS (Microsoft's DirectShow), MS MFT (Microsoft's Media Foundation), LAV Video etc
Third is the decoder's mode - Native DXVA, Copy-Back (CB) DXVA, Quicksync (QS) etc
H.264
1. Twinpeaks-30fps
1. QS MS DS 401/401/401
QS MS MFT 390/395/400
QS CoreAVC 368/375/383
QS LAV NATIVE 366/374/375
CPU Core i5@3.1 253/264/274
QS LAV QS 200/201/202
QS FFDShow 158/161/163
VP5 LAV CUDA 130/139/143
VP5 MS DS 133/138/141
QS LAV CB 131/137/140
VP5 MS MFT 128/137/141
VP5 CoreCUDA 84/89/93
CPU C2D@2.83 73/85/96
VP4 LAV CUDA 80/84/88
VP4 MS MFT 80/84/88
VP4 LAV CB 81/84/87
VP4 MS DS 80/84/87
VP4 LAV NATIVE 77/79/82
UVD2.2 OC MS MFT 73/77/88
UVD2.2 OC LAV NATIVE 75/77/83
UVD2.2 OC MS DS 76/77/80
VP4 CoreCUDA 62/65/66
UVD2.2 LAV NATIVE 57/58/62
UVD2.2 OC LAV CB 56/57/59
UVD2.2 MS DS 52/57/66
UVD2.2 MS MFT 51/57/65
VP4 CoreAVC 55/56/57
UVD2.2 LAV CB 48/53/55
UVD2.2 OC CoreAVC 46/53/56
UVD2.2 CoreAVC 44/51/55
2. Samsung-30fps
1. QS MS MFT 234/271/341
QS MS DS 224/266/333
QS CoreAVC 229/263/321
QS LAV NATIVE 219/259/325
QS LAV QS 134/162/190
QS FFDShow 136/150/160
CPU Core i5@3.1 99/132/197
VP5 CUDA 82/115/129
VP5 CoreCUDA 89/107/123
QS LAV CB 86/105/128
VP5 MS DS 93/105/121
UVD2.2 OC MS MFT 52/62/75
UVD2.2 OC LAV NATIVE 55/62/71
UVD2.2 OC MS DS 50/62/72
UVD2.2 OC LAV CB 53/55/58
VP4 MS MFT 34/55/91
VP4 LAV CB 34/55/84
VP4 LAV CUDA 31/55/90
VP4 LAV NATIVE 35/54/83
VP4 MS DS 33/54/82
VP4 CoreCUDA 32/51/77
UVD2.2 OC CoreAVC 39/49/57
CPU C2D@2.83 34/49/81
UVD2.2 MS MFT 35/46/62
UVD2.2 LAV CB 35/46/56
UVD2.2 LAV NATIVE 35/45/54
UVD2.2 MS DS 32/45/56
VP4 CoreAVC 27/41/62
UVD2.2 CoreAVC 30/38/44
3. Basket-60fps
DXVA checker 2.8.0b3 used for DXVA-CB, QS, NVCUVID and CPU.
1. QS CoreAVC 461/504/550
QS MS MFT 455/502/567
QS MS DS 455/502/553
QS LAV NATIVE 439/483/536
CPU Core i5@3.1 265/286/317
QS LAV QS 200/203/209
QS FFDShow 160/151/172
VP5 CoreCUDA 137/149/164
VP5 MS DS 118/133/149
QS LAV CB 109/115/120
VP4 CoreCUDA 77/84/106
VP4 LAV CB 75/82/99
VP4 MS MFT 74/82/107
CPU C2D@2.83 72/82/103
VP4 MS DS 75/81/89
VP4 LAV CUDA 73/81/99
VP4 LAV NATIVE 71/81/103
UVD2.2 OC LAV NATIVE 74/76/77
UVD2.2 OC MS MFT 70/76/82
UVD2.2 OC MS DS 62/76/78 *
UVD2.2 OC CoreAVC 56/57/58
UVD2.2 LAV NATIVE 55/57/59
UVD2.2 LAV CB 54/57/58
UVD2.2 MS MFT 53/57/65
VP4 CoreAVC 54/57/66
UVD2.2 MS DS 41/57/59 *
UVD2.2 CoreAVC 42/44/50
4. Girls-60fps
1. QS MS DS 410/423/436
QS MS MFT 410/420/456
QS CoreAVC 401/414/430
QS LAV NATIVE 397/412/430
CPU Core i5@3.1 198/209/234
QS LAV QS 193/200/203
QS FFDShow 160/166/171
VP5 LAV CUDA 141/143/146
VP5 CoreCUDA 110/128/144
VP5 MS DS 114/122/133
QS LAV CB 93/97/109
VP4 LAV CB 74/76/80
VP4 LAV NATIVE 74/76/79
VP4 LAV CUDA 74/76/78
VP4 MS MFT 73/76/79
UVD2.2 OC MS MFT 72/76/81
UVD2.2 OC MS DS 72/76/77 *
VP4 MS DS 73/75/77
UVD2.2 OC LAV NATIVE 71/75/77
VP4 CoreCUDA 67/73/81
CPU C2D@2.83 63/70/85
UVD2.2 OC LAV CB 51/59/60
UVD2.2 MS MFT 52/57/62
UVD2.2 LAV NATIVE 54/56/59
UVD2.2 LAV CB 54/56/58
UVD2.2 OC CoreAVC 55/56/57
VP4 CoreAVC 54/56/58
UVD2.2 MS DS 50/56/59 *
UVD2.2 CoreAVC 42/45/50
5. Cat-60fps
No MFT splitter for M2TS files
1. QS CoreAVC 394/402/411
QS MS DS 383/400/412
QS LAV NATIVE 372/381/387
QS LAV QS 190/194/197
CPU Core i5@3.1 157/187/206
QS FFDShow 156/161/166
VP5 LAV CUDA 138/141/147
VP5 CoreCUDA 134/138/145
QS LAV CB 117/118/125
VP5 MS DS 78/98/106
UVD2.2 OC MS DS 69/74/76
VP4 CoreCUDA 68/74/81
UVD2.2 OC LAV NATIVE 70/73/74
VP4 LAV CUDA 68/72/78
VP4 LAV CB 67/72/78
VP4 MS DS 68/71/79
VP4 LAV NATIVE 65/71/77
CPU C2D@2.83 54/62/71
UVD2.2 OC LAV CB 57/58/59
UVD2.2 OC CoreAVC 54/55/56
UVD2.2 LAV NATIVE 54/55/56
UVD2.2 LAV CB 53/55/57
UVD2.2 MS DS 50/55/58
VP4 CoreAVC 48/50/56
UVD2.2 CoreAVC 42/44/46
6. Avatar-60fps
MS MFT crashes DXVA Checker all versions, all platforms
1. QS MS DS 328/345/367
QS CoreAVC 322/330/336
QS LAV NATIVE 322/329/337
QS LAV QS 189/193/195
QS FFDShow 155/162/166
CPU Core i5@3.1 143/159/175
QS LAV CB 108/114/122
VP4 MS DS 65/76/84
VP4 LAV CUDA 64/76/82
VP4 LAV CB 64/76/82
VP4 LAV NATIVE 67/73/80
UVD2.2 OC LAV NATIVE 68/70/73
UVD2.2 OC MS DS 68/70/72 *
UVD2.2 OC LAV CB 55/57/58
UVD2.2 OC CoreAVC 53/55/57
UVD2.2 LAV CB 48/53/57
UVD2.2 LAV NATIVE 49/52/54
CPU C2D@2.83 44/52/63
UVD2.2 MS DS 39/52/55 *
VP4 CoreAVC 46/51/57
UVD2.2 CoreAVC 40/42/44
7. Vortex-24fps
1. QS FFDShow 122/125/129
QS LAV QS 119/121/124
QS MS DS 119/120/120
QS MS MFT 119/120/120
QS LAV NATIVE 118/120/122
QS CoreAVC 118/119/122
QS LAV CB 112/117/121
VP5 LAV CUDA 71/73/76
VP5 CoreCUDA 72/73/76
VP5 MS DS 71/73/75
VP5 MS MFT 72/73/76
CPU Core i5@3.1 56/59/61
UVD2.2 OC LAV CB 35/37/44
UVD2.2 OC MS MFT 34/37/42
UVD2.2 OC LAV NATIVE 36/36/39
UVD2.2 OC MS DS 36/36/37
UVD2.2 OC CoreAVC 28/29/31
UVD2.2 LAV NATIVE 26/28/28
UVD2.2 LAV CB 25/27/32
UVD2.2 MS DS 25/26/29
UVD2.2 MS MFT 24/26/34
CPU C2D@2.83 19/24/27
VP4 LAV CUDA 21/22/26
VP4 LAV CB 21/22/25
UVD2.2 CoreAVC 21/22/25
VP4 MS MFT 21/22/24
VP4 CoreCUDA 21/22/24
VP4 LAV NATIVE 20/22/24
VP4 MS DS 19/22/24
VP4 CoreAVC 17/18/21
8. Birds-24fps
[U]1. QS MS MFT 110/118/133
QS CoreAVC 110/117/129
QS FFDShow 112/115/120
QS LAV NATIVE 109/114/122
QS LAV QS 108/114/122
QS MS DS 108/113/122
QS LAV CB 78/83/88
VP5 CoreCUDA 71/77/87
VP5 LAV CUDA 59/69/77
VP5 MS DS 59/63/71
CPU Core i5@3.1 50/54/59
UVD2.2 OC MS MFT 32/38/47
UVD2.2 OC LAV NATIVE 36/37/43
UVD2.2 OC LAV CB 34/37/41
UVD2.2 MS DS 35/37/39
UVD2.2 OC CoreAVC 28/30/34
UVD2.2 LAV NATIVE 26/27/34
UVD2.2 LAV CB 25/27/35
UVD2.2 MS DS 25/27/35
UVD2.2 MS MFT 23/27/33
UVD2.2 CoreAVC 21/23/27
VP4 CoreCUDA 20/22/29
VP4 MS MFT 19/22/30
VP4 LAV CUDA 10/22/35 (A lot of breaks)
VP4 LAV CB 19/21/29
VP4 LAV NATIVE 19/21/28
VP4 MS DS 18/21/24
CPU C2D@2.83 18/21/23
VP4 CoreAVC 15/18/25
9. Ducks -30fps
1. QS CoreAVC 119/136/147
QS MS MFT 117/135/152
QS LAV NATIVE 118/134/144
QS LAV QS 116/130/139
QS FFDShow 129/129/132
QS MS DS 123/125/126 *
VP5 CoreCUDA 74/83/95
VP5 LAV CUDA 71/82/92
QS LAV CB 67/75/91
CPU Core i5@3.1 56/63/71
VP5 MS DS 48/57/84
UVD2.2 OC LAV NATIVE 36/41/48
UVD2.2 OC MS MFT 36/41/48
UVD2.2 OC LAV CB 35/41/48
UVD2.2 OC MS DS 30/39/43 *
UVD2.2 OC CoreAVC 29/33/37
UVD2.2 LAV CB 25/30/39
UVD2.2 MS MFT 24/30/38
UVD2.2 LAV NATIVE 26/29/35
UVD2.2 MS DS 21/27/32 *
VP4 LAV CUDA 15/26/39
UVD2.2 CoreAVC 22/25/29
VP4 MS MFT 21/25/34
VP4 CoreCUDA 21/25/31
VP4 LAV CB 20/25/34
VP4 LAV NATIVE 20/25/30
CPU C2D@2.83 20/24/29
VP4 MS DS 19/24/31 *
VP4 CoreAVC 18/21/27
10. Crowd Run-25fps
1. QS FFDShow 118/118/118
QS MS DS 112 (Only Average result)
QS MS MFT 109/109/110
QS CoreAVC 109/109/109
QS LAV QS 109/109/109
QS LAV NATIVE 107/107/107
QS LAV CB 71/72/73
VP5 LAV CUDA 68/70/72
VP5 CoreCUDA 68/69/70
VP5 MS DS 59/59/59
CPU Core i5@3.1 52/53/54
UVD2.2 OC LAV CB 32/33/37
UVD2.2 OC LAV NATIVE 32/33/33
UVD2.2 OC MS MFT 31/33/35
UVD2.2 OC MS DS 25/31/35
UVD2.2 OC CoreAVC 26/27/28
UVD2.2 LAV CB 23/24/27
UVD2.2 MS MFT 23/24/25
UVD2.2 LAV NATIVE 23/24/24
UVD2.2 MS DS 16/22/24
VP4 LAV CUDA 20/21/23
VP4 MS MFT 20/21/23
VP4 LAV CB 20/21/23
VP4 CoreCUDA 20/21/22
CPU C2D@2.83 19/21/22
VP4 LAV NATIVE 20/20/22
UVD2.2 CoreAVC 20/20/21
VP4 MS DS 18/20/22 *
VP4 CoreAVC 17/18/19
* MS DS decoder has a lot of artifacts at the beginning of the decoding, resulting low min value and probably lower average value
VC-1/WMV3
A) VC-1 - Devil May Cry 1080/60p-40Mbps
1. QS LAV QS 214/218/222
QS FFDShow 158/167/170
UVD2.2 OC AMD MFT 87/88/89
UVD2.2 OC LAV NATIVE 87/88/89
VP4 LAV CUVID 80/80/84
VP4 LAV CB 79/80/84
VP4 LAV NATIVE 76/80/84
CPU Core i5@3.1 68/76/96
UVD2.2 AMD MFT 65/66/67
UVD2.2 LAV NATIVE 64/66/69
UVD2.2 OC LAV CB 51/57/59
UVD2.2 LAV CB 48/55/56
CPU C2D@2.83 41/44/47
B) WMV3 - MP10 Digital Life 1080/24p-10Mbps
1. QS LAV QS 242/247/252
QS FFDShow QS 170/172/175
CPU Core i5@3.1 89/102/110
UVD2.2 OC AMD MFT 87/90/91
UVD2.2 OC LAV NATIVE 87/90/91
VP4 LAV CUVID 84/89/101
VP4 LAV CB 82/83/86
VP4 LAV NATIVE 81/83/84
UVD2.2 AMD MFT 68/68/69
UVD2.2 LAV NATIVE 68/68/69
CPU C2D@2.83 56/63/68
UVD2.2 OC LAV CB 57/58/59
UVD2.2 LAV CB 55/57/59
Comments:
1) The performance of QuickSync HW is beyond any competition, using native DXVA mode with every decoder used (MS DS/MFT, CoreAVC, LAV NATIVE).
The performance of copy-back mode using Intel's MSDK QuickSync decoder v0.28 software (FFDshow, LAV QS) is heavily multi-threaded and optimized but for some reason is a lot slower than QS decoder v0.20 (more than 20%) and a lot slower from DXVA2 native with the above system configuration in high frame rate clips - 60fps and/or low bitrate clips (Clips 1 to 6)
But it's very good and sometimes faster than native DXVA when used for high birate - low frame rate clips (Clips 7 to 10)
The performance of LAV DXVA2 copy-back is simply awful. It's slower than VP5 most of the times!
For laptop users, or for people who want their CPU and GPU load as low as possible during playback mode, native DXVA2 decoders (MS DS/MFT, CoreAVC, LAV NATIVE) are BY FAR the most efficient decoders.
2) VP5 is about 2 times faster than VP4 in "easy" low bitrate clips from 1 to 6. But it's more than 3 times faster in "difficult" high bitrate clips from 7 to 10, as it is built for 4K x 2K decoding. For those huge bandwidth clips, it closes the gap with QuickSync, but the distance is still obvious.
3) VP4 has a lot of problems at huge bandwidths starting from clip 7 up to 10 like 5750 UVD2.2, although the latter is a little faster. It seems absolutely reasonable for Nvidia to go for VP5 in order to support 4K x 2K and large bandwidths.
UVD2.2 OC can play easily every clip from 1 to 10!
4) UVD2.2 OC is about 35% - 42% faster than 5750 UVD2.2 in H.264 and 32% faster in VC-1/WMV3 and it's faster than VP4 and Core2Duo too in bandwidth heavy files starting from clip 7.
Added also AMD Polaris RX 470 results and just one result of Pascal GTX 1060 VP8 decoder.
All Intel CPUs from Haswell to Kabylake have exactly the same 4K HW H.264 decoder, they differ only in clock speed.
Take a look here:
http://forum.doom9.org/showthread.php?p=1712350#post1712350
I've recently flashed my Radeon 5750 BIOS with 6750 BIOS.
The two cards use the same UVD2.2
But it seems that 6750 BIOS on a 5750, can lead to a UVD2.2 overclocking.
The default 5750 BIOS put UVD2.2 in standard UVD mode at Core/GPU = 400MHz / 900MHz
The 6750 BIOS on a 5750 card, put UVD2.2 in 3D mode at Core/GPU = 710MHz / 1160MHz
So I have two UVD2.2 systems benchmarked.
One plain UVD2.2 (400/900) and one UVD2.2 OC (710/1160)
I did my tests with the new DXVA checker x86 v2.7.0 http://bluesky23.yukishigure.com/en/index.html
Two systems tested:
1) My signature system:
Win 7 x64 SP1 - C2D@2.83 GHz - Radeon (6)750 - Catalyst 12.1 preview,
RAM configuration for AMD system: (mostly for DXVA-CB comparisons)
4GB (2 x 2GB) of DDR2 at FSB: DRAM = 1:1
Speed = 4-4-4-12@566 MHz (283x2)
2) Intel/ Nvidia system:
Win 7 SP1 x86 - Core i5-2400 (3.1GHz) - Geforce GT 440 (DDR5) - Nvidia beta 290.53 - Intel HD 2000 - Intel drivers v.2622
RAM configuration for Intel/ Nvidia system: (mostly for DXVA-CB comparisons, QuickSync decoder)
4GB (2 x 2GB) of DDR3 at FSB: DRAM = 1:5
Speed = 9-9-9-24@1338 MHz (669x2)
The decoders used are:
CoreAVC 3.0.1 (both modes - DXVA native, NVCUVID)
LAV Video 0.47 (in all modes - DXVA2 native, DXVA2 copy-back, NVCUVID, QS)
MS DS/MFT
FFDShow v4322 (QS)
For VC-1/ WMV3 I used the AMD Playback Decoder MFT and for CPU results I used the built-in WMVideo Decoder DMO (because is faster than LAV slow VC-1/WMV decoder)
I used five Reference H.264 files from here:
http://forum.doom9.org/showthread.php?t=159486
and I added five new reference files.
You can find every sample posted (from 1 to 10) here:
ftp://helpedia.com/pub/multimedia/x264/testvideos/
6.Avatar-1080p60fpsRef4-44.9Mbps
7.Vortexx_1088p24fpsRef3-109Mpbs
8.Birds_1080p24fpsRef4-112Mbps
9.Ducks.Take.Off.1080p30fpsRef5-108Mbps
10.Crowd.Run.1080p25Ref4-116Mbps
Also I used VC-1 and WMV3 files from here:
http://forum.doom9.org/showthread.php?t=156660
For CPU results (Core 2 Duo - Core i5) I used LAV Video 0.47.
Every benchmark mode used EVR renderer.
The results:
First is the Video Processor - QuickSync (QS), UVD2.2, VP4, CPU etc
Second is the decoder - MS DS (Microsoft's DirectShow), MS MFT (Microsoft's Media Foundation), LAV Video etc
Third is the decoder's mode - Native DXVA, Copy-Back (CB) DXVA, Quicksync (QS) etc
H.264
1. Twinpeaks-30fps
1. QS MS DS 401/401/401
QS MS MFT 390/395/400
QS CoreAVC 368/375/383
QS LAV NATIVE 366/374/375
CPU Core i5@3.1 253/264/274
QS LAV QS 200/201/202
QS FFDShow 158/161/163
VP5 LAV CUDA 130/139/143
VP5 MS DS 133/138/141
QS LAV CB 131/137/140
VP5 MS MFT 128/137/141
VP5 CoreCUDA 84/89/93
CPU C2D@2.83 73/85/96
VP4 LAV CUDA 80/84/88
VP4 MS MFT 80/84/88
VP4 LAV CB 81/84/87
VP4 MS DS 80/84/87
VP4 LAV NATIVE 77/79/82
UVD2.2 OC MS MFT 73/77/88
UVD2.2 OC LAV NATIVE 75/77/83
UVD2.2 OC MS DS 76/77/80
VP4 CoreCUDA 62/65/66
UVD2.2 LAV NATIVE 57/58/62
UVD2.2 OC LAV CB 56/57/59
UVD2.2 MS DS 52/57/66
UVD2.2 MS MFT 51/57/65
VP4 CoreAVC 55/56/57
UVD2.2 LAV CB 48/53/55
UVD2.2 OC CoreAVC 46/53/56
UVD2.2 CoreAVC 44/51/55
2. Samsung-30fps
1. QS MS MFT 234/271/341
QS MS DS 224/266/333
QS CoreAVC 229/263/321
QS LAV NATIVE 219/259/325
QS LAV QS 134/162/190
QS FFDShow 136/150/160
CPU Core i5@3.1 99/132/197
VP5 CUDA 82/115/129
VP5 CoreCUDA 89/107/123
QS LAV CB 86/105/128
VP5 MS DS 93/105/121
UVD2.2 OC MS MFT 52/62/75
UVD2.2 OC LAV NATIVE 55/62/71
UVD2.2 OC MS DS 50/62/72
UVD2.2 OC LAV CB 53/55/58
VP4 MS MFT 34/55/91
VP4 LAV CB 34/55/84
VP4 LAV CUDA 31/55/90
VP4 LAV NATIVE 35/54/83
VP4 MS DS 33/54/82
VP4 CoreCUDA 32/51/77
UVD2.2 OC CoreAVC 39/49/57
CPU C2D@2.83 34/49/81
UVD2.2 MS MFT 35/46/62
UVD2.2 LAV CB 35/46/56
UVD2.2 LAV NATIVE 35/45/54
UVD2.2 MS DS 32/45/56
VP4 CoreAVC 27/41/62
UVD2.2 CoreAVC 30/38/44
3. Basket-60fps
DXVA checker 2.8.0b3 used for DXVA-CB, QS, NVCUVID and CPU.
1. QS CoreAVC 461/504/550
QS MS MFT 455/502/567
QS MS DS 455/502/553
QS LAV NATIVE 439/483/536
CPU Core i5@3.1 265/286/317
QS LAV QS 200/203/209
QS FFDShow 160/151/172
VP5 CoreCUDA 137/149/164
VP5 MS DS 118/133/149
QS LAV CB 109/115/120
VP4 CoreCUDA 77/84/106
VP4 LAV CB 75/82/99
VP4 MS MFT 74/82/107
CPU C2D@2.83 72/82/103
VP4 MS DS 75/81/89
VP4 LAV CUDA 73/81/99
VP4 LAV NATIVE 71/81/103
UVD2.2 OC LAV NATIVE 74/76/77
UVD2.2 OC MS MFT 70/76/82
UVD2.2 OC MS DS 62/76/78 *
UVD2.2 OC CoreAVC 56/57/58
UVD2.2 LAV NATIVE 55/57/59
UVD2.2 LAV CB 54/57/58
UVD2.2 MS MFT 53/57/65
VP4 CoreAVC 54/57/66
UVD2.2 MS DS 41/57/59 *
UVD2.2 CoreAVC 42/44/50
4. Girls-60fps
1. QS MS DS 410/423/436
QS MS MFT 410/420/456
QS CoreAVC 401/414/430
QS LAV NATIVE 397/412/430
CPU Core i5@3.1 198/209/234
QS LAV QS 193/200/203
QS FFDShow 160/166/171
VP5 LAV CUDA 141/143/146
VP5 CoreCUDA 110/128/144
VP5 MS DS 114/122/133
QS LAV CB 93/97/109
VP4 LAV CB 74/76/80
VP4 LAV NATIVE 74/76/79
VP4 LAV CUDA 74/76/78
VP4 MS MFT 73/76/79
UVD2.2 OC MS MFT 72/76/81
UVD2.2 OC MS DS 72/76/77 *
VP4 MS DS 73/75/77
UVD2.2 OC LAV NATIVE 71/75/77
VP4 CoreCUDA 67/73/81
CPU C2D@2.83 63/70/85
UVD2.2 OC LAV CB 51/59/60
UVD2.2 MS MFT 52/57/62
UVD2.2 LAV NATIVE 54/56/59
UVD2.2 LAV CB 54/56/58
UVD2.2 OC CoreAVC 55/56/57
VP4 CoreAVC 54/56/58
UVD2.2 MS DS 50/56/59 *
UVD2.2 CoreAVC 42/45/50
5. Cat-60fps
No MFT splitter for M2TS files
1. QS CoreAVC 394/402/411
QS MS DS 383/400/412
QS LAV NATIVE 372/381/387
QS LAV QS 190/194/197
CPU Core i5@3.1 157/187/206
QS FFDShow 156/161/166
VP5 LAV CUDA 138/141/147
VP5 CoreCUDA 134/138/145
QS LAV CB 117/118/125
VP5 MS DS 78/98/106
UVD2.2 OC MS DS 69/74/76
VP4 CoreCUDA 68/74/81
UVD2.2 OC LAV NATIVE 70/73/74
VP4 LAV CUDA 68/72/78
VP4 LAV CB 67/72/78
VP4 MS DS 68/71/79
VP4 LAV NATIVE 65/71/77
CPU C2D@2.83 54/62/71
UVD2.2 OC LAV CB 57/58/59
UVD2.2 OC CoreAVC 54/55/56
UVD2.2 LAV NATIVE 54/55/56
UVD2.2 LAV CB 53/55/57
UVD2.2 MS DS 50/55/58
VP4 CoreAVC 48/50/56
UVD2.2 CoreAVC 42/44/46
6. Avatar-60fps
MS MFT crashes DXVA Checker all versions, all platforms
1. QS MS DS 328/345/367
QS CoreAVC 322/330/336
QS LAV NATIVE 322/329/337
QS LAV QS 189/193/195
QS FFDShow 155/162/166
CPU Core i5@3.1 143/159/175
QS LAV CB 108/114/122
VP4 MS DS 65/76/84
VP4 LAV CUDA 64/76/82
VP4 LAV CB 64/76/82
VP4 LAV NATIVE 67/73/80
UVD2.2 OC LAV NATIVE 68/70/73
UVD2.2 OC MS DS 68/70/72 *
UVD2.2 OC LAV CB 55/57/58
UVD2.2 OC CoreAVC 53/55/57
UVD2.2 LAV CB 48/53/57
UVD2.2 LAV NATIVE 49/52/54
CPU C2D@2.83 44/52/63
UVD2.2 MS DS 39/52/55 *
VP4 CoreAVC 46/51/57
UVD2.2 CoreAVC 40/42/44
7. Vortex-24fps
1. QS FFDShow 122/125/129
QS LAV QS 119/121/124
QS MS DS 119/120/120
QS MS MFT 119/120/120
QS LAV NATIVE 118/120/122
QS CoreAVC 118/119/122
QS LAV CB 112/117/121
VP5 LAV CUDA 71/73/76
VP5 CoreCUDA 72/73/76
VP5 MS DS 71/73/75
VP5 MS MFT 72/73/76
CPU Core i5@3.1 56/59/61
UVD2.2 OC LAV CB 35/37/44
UVD2.2 OC MS MFT 34/37/42
UVD2.2 OC LAV NATIVE 36/36/39
UVD2.2 OC MS DS 36/36/37
UVD2.2 OC CoreAVC 28/29/31
UVD2.2 LAV NATIVE 26/28/28
UVD2.2 LAV CB 25/27/32
UVD2.2 MS DS 25/26/29
UVD2.2 MS MFT 24/26/34
CPU C2D@2.83 19/24/27
VP4 LAV CUDA 21/22/26
VP4 LAV CB 21/22/25
UVD2.2 CoreAVC 21/22/25
VP4 MS MFT 21/22/24
VP4 CoreCUDA 21/22/24
VP4 LAV NATIVE 20/22/24
VP4 MS DS 19/22/24
VP4 CoreAVC 17/18/21
8. Birds-24fps
[U]1. QS MS MFT 110/118/133
QS CoreAVC 110/117/129
QS FFDShow 112/115/120
QS LAV NATIVE 109/114/122
QS LAV QS 108/114/122
QS MS DS 108/113/122
QS LAV CB 78/83/88
VP5 CoreCUDA 71/77/87
VP5 LAV CUDA 59/69/77
VP5 MS DS 59/63/71
CPU Core i5@3.1 50/54/59
UVD2.2 OC MS MFT 32/38/47
UVD2.2 OC LAV NATIVE 36/37/43
UVD2.2 OC LAV CB 34/37/41
UVD2.2 MS DS 35/37/39
UVD2.2 OC CoreAVC 28/30/34
UVD2.2 LAV NATIVE 26/27/34
UVD2.2 LAV CB 25/27/35
UVD2.2 MS DS 25/27/35
UVD2.2 MS MFT 23/27/33
UVD2.2 CoreAVC 21/23/27
VP4 CoreCUDA 20/22/29
VP4 MS MFT 19/22/30
VP4 LAV CUDA 10/22/35 (A lot of breaks)
VP4 LAV CB 19/21/29
VP4 LAV NATIVE 19/21/28
VP4 MS DS 18/21/24
CPU C2D@2.83 18/21/23
VP4 CoreAVC 15/18/25
9. Ducks -30fps
1. QS CoreAVC 119/136/147
QS MS MFT 117/135/152
QS LAV NATIVE 118/134/144
QS LAV QS 116/130/139
QS FFDShow 129/129/132
QS MS DS 123/125/126 *
VP5 CoreCUDA 74/83/95
VP5 LAV CUDA 71/82/92
QS LAV CB 67/75/91
CPU Core i5@3.1 56/63/71
VP5 MS DS 48/57/84
UVD2.2 OC LAV NATIVE 36/41/48
UVD2.2 OC MS MFT 36/41/48
UVD2.2 OC LAV CB 35/41/48
UVD2.2 OC MS DS 30/39/43 *
UVD2.2 OC CoreAVC 29/33/37
UVD2.2 LAV CB 25/30/39
UVD2.2 MS MFT 24/30/38
UVD2.2 LAV NATIVE 26/29/35
UVD2.2 MS DS 21/27/32 *
VP4 LAV CUDA 15/26/39
UVD2.2 CoreAVC 22/25/29
VP4 MS MFT 21/25/34
VP4 CoreCUDA 21/25/31
VP4 LAV CB 20/25/34
VP4 LAV NATIVE 20/25/30
CPU C2D@2.83 20/24/29
VP4 MS DS 19/24/31 *
VP4 CoreAVC 18/21/27
10. Crowd Run-25fps
1. QS FFDShow 118/118/118
QS MS DS 112 (Only Average result)
QS MS MFT 109/109/110
QS CoreAVC 109/109/109
QS LAV QS 109/109/109
QS LAV NATIVE 107/107/107
QS LAV CB 71/72/73
VP5 LAV CUDA 68/70/72
VP5 CoreCUDA 68/69/70
VP5 MS DS 59/59/59
CPU Core i5@3.1 52/53/54
UVD2.2 OC LAV CB 32/33/37
UVD2.2 OC LAV NATIVE 32/33/33
UVD2.2 OC MS MFT 31/33/35
UVD2.2 OC MS DS 25/31/35
UVD2.2 OC CoreAVC 26/27/28
UVD2.2 LAV CB 23/24/27
UVD2.2 MS MFT 23/24/25
UVD2.2 LAV NATIVE 23/24/24
UVD2.2 MS DS 16/22/24
VP4 LAV CUDA 20/21/23
VP4 MS MFT 20/21/23
VP4 LAV CB 20/21/23
VP4 CoreCUDA 20/21/22
CPU C2D@2.83 19/21/22
VP4 LAV NATIVE 20/20/22
UVD2.2 CoreAVC 20/20/21
VP4 MS DS 18/20/22 *
VP4 CoreAVC 17/18/19
* MS DS decoder has a lot of artifacts at the beginning of the decoding, resulting low min value and probably lower average value
VC-1/WMV3
A) VC-1 - Devil May Cry 1080/60p-40Mbps
1. QS LAV QS 214/218/222
QS FFDShow 158/167/170
UVD2.2 OC AMD MFT 87/88/89
UVD2.2 OC LAV NATIVE 87/88/89
VP4 LAV CUVID 80/80/84
VP4 LAV CB 79/80/84
VP4 LAV NATIVE 76/80/84
CPU Core i5@3.1 68/76/96
UVD2.2 AMD MFT 65/66/67
UVD2.2 LAV NATIVE 64/66/69
UVD2.2 OC LAV CB 51/57/59
UVD2.2 LAV CB 48/55/56
CPU C2D@2.83 41/44/47
B) WMV3 - MP10 Digital Life 1080/24p-10Mbps
1. QS LAV QS 242/247/252
QS FFDShow QS 170/172/175
CPU Core i5@3.1 89/102/110
UVD2.2 OC AMD MFT 87/90/91
UVD2.2 OC LAV NATIVE 87/90/91
VP4 LAV CUVID 84/89/101
VP4 LAV CB 82/83/86
VP4 LAV NATIVE 81/83/84
UVD2.2 AMD MFT 68/68/69
UVD2.2 LAV NATIVE 68/68/69
CPU C2D@2.83 56/63/68
UVD2.2 OC LAV CB 57/58/59
UVD2.2 LAV CB 55/57/59
Comments:
1) The performance of QuickSync HW is beyond any competition, using native DXVA mode with every decoder used (MS DS/MFT, CoreAVC, LAV NATIVE).
The performance of copy-back mode using Intel's MSDK QuickSync decoder v0.28 software (FFDshow, LAV QS) is heavily multi-threaded and optimized but for some reason is a lot slower than QS decoder v0.20 (more than 20%) and a lot slower from DXVA2 native with the above system configuration in high frame rate clips - 60fps and/or low bitrate clips (Clips 1 to 6)
But it's very good and sometimes faster than native DXVA when used for high birate - low frame rate clips (Clips 7 to 10)
The performance of LAV DXVA2 copy-back is simply awful. It's slower than VP5 most of the times!
For laptop users, or for people who want their CPU and GPU load as low as possible during playback mode, native DXVA2 decoders (MS DS/MFT, CoreAVC, LAV NATIVE) are BY FAR the most efficient decoders.
2) VP5 is about 2 times faster than VP4 in "easy" low bitrate clips from 1 to 6. But it's more than 3 times faster in "difficult" high bitrate clips from 7 to 10, as it is built for 4K x 2K decoding. For those huge bandwidth clips, it closes the gap with QuickSync, but the distance is still obvious.
3) VP4 has a lot of problems at huge bandwidths starting from clip 7 up to 10 like 5750 UVD2.2, although the latter is a little faster. It seems absolutely reasonable for Nvidia to go for VP5 in order to support 4K x 2K and large bandwidths.
UVD2.2 OC can play easily every clip from 1 to 10!
4) UVD2.2 OC is about 35% - 42% faster than 5750 UVD2.2 in H.264 and 32% faster in VC-1/WMV3 and it's faster than VP4 and Core2Duo too in bandwidth heavy files starting from clip 7.