View Full Version : H.264 DXVA Benchmarks: QuickSync vs UVD 2.2 vs VP4 vs VP5
NikosD
16th November 2011, 21:59
Latest update with LAV Video x64 0.64 in DXVA native and pure decode mode, using latest ASICs like VP7 from Nvidia GTX 960 and QuickSync 3 from Haswell.
Added also AMD Polaris RX 470 results and just one result of Pascal GTX 1060 VP8 decoder.
All Intel CPUs from Haswell to Kabylake have exactly the same 4K HW H.264 decoder, they differ only in clock speed.
Take a look here:
http://forum.doom9.org/showthread.php?p=1712350#post1712350
I've recently flashed my Radeon 5750 BIOS with 6750 BIOS.
The two cards use the same UVD2.2
But it seems that 6750 BIOS on a 5750, can lead to a UVD2.2 overclocking.
The default 5750 BIOS put UVD2.2 in standard UVD mode at Core/GPU = 400MHz / 900MHz
The 6750 BIOS on a 5750 card, put UVD2.2 in 3D mode at Core/GPU = 710MHz / 1160MHz
So I have two UVD2.2 systems benchmarked.
One plain UVD2.2 (400/900) and one UVD2.2 OC (710/1160)
I did my tests with the new DXVA checker x86 v2.7.0 http://bluesky23.yukishigure.com/en/index.html
Two systems tested:
1) My signature system:
Win 7 x64 SP1 - C2D@2.83 GHz - Radeon (6)750 - Catalyst 12.1 preview,
RAM configuration for AMD system: (mostly for DXVA-CB comparisons)
4GB (2 x 2GB) of DDR2 at FSB: DRAM = 1:1
Speed = 4-4-4-12@566 MHz (283x2)
2) Intel/ Nvidia system:
Win 7 SP1 x86 - Core i5-2400 (3.1GHz) - Geforce GT 440 (DDR5) - Nvidia beta 290.53 - Intel HD 2000 - Intel drivers v.2622
RAM configuration for Intel/ Nvidia system: (mostly for DXVA-CB comparisons, QuickSync decoder)
4GB (2 x 2GB) of DDR3 at FSB: DRAM = 1:5
Speed = 9-9-9-24@1338 MHz (669x2)
The decoders used are:
CoreAVC 3.0.1 (both modes - DXVA native, NVCUVID)
LAV Video 0.47 (in all modes - DXVA2 native, DXVA2 copy-back, NVCUVID, QS)
MS DS/MFT
FFDShow v4322 (QS)
For VC-1/ WMV3 I used the AMD Playback Decoder MFT and for CPU results I used the built-in WMVideo Decoder DMO (because is faster than LAV slow VC-1/WMV decoder)
I used five Reference H.264 files from here:
http://forum.doom9.org/showthread.php?t=159486
and I added five new reference files.
You can find every sample posted (from 1 to 10) here:
ftp://helpedia.com/pub/multimedia/x264/testvideos/
6.Avatar-1080p60fpsRef4-44.9Mbps
7.Vortexx_1088p24fpsRef3-109Mpbs
8.Birds_1080p24fpsRef4-112Mbps
9.Ducks.Take.Off.1080p30fpsRef5-108Mbps
10.Crowd.Run.1080p25Ref4-116Mbps
Also I used VC-1 and WMV3 files from here:
http://forum.doom9.org/showthread.php?t=156660
For CPU results (Core 2 Duo - Core i5) I used LAV Video 0.47.
Every benchmark mode used EVR renderer.
The results:
First is the Video Processor - QuickSync (QS), UVD2.2, VP4, CPU etc
Second is the decoder - MS DS (Microsoft's DirectShow), MS MFT (Microsoft's Media Foundation), LAV Video etc
Third is the decoder's mode - Native DXVA, Copy-Back (CB) DXVA, Quicksync (QS) etc
H.264
1. Twinpeaks-30fps
1. QS MS DS 401/401/401
QS MS MFT 390/395/400
QS CoreAVC 368/375/383
QS LAV NATIVE 366/374/375
CPU Core i5@3.1 253/264/274
QS LAV QS 200/201/202
QS FFDShow 158/161/163
VP5 LAV CUDA 130/139/143
VP5 MS DS 133/138/141
QS LAV CB 131/137/140
VP5 MS MFT 128/137/141
VP5 CoreCUDA 84/89/93
CPU C2D@2.83 73/85/96
VP4 LAV CUDA 80/84/88
VP4 MS MFT 80/84/88
VP4 LAV CB 81/84/87
VP4 MS DS 80/84/87
VP4 LAV NATIVE 77/79/82
UVD2.2 OC MS MFT 73/77/88
UVD2.2 OC LAV NATIVE 75/77/83
UVD2.2 OC MS DS 76/77/80
VP4 CoreCUDA 62/65/66
UVD2.2 LAV NATIVE 57/58/62
UVD2.2 OC LAV CB 56/57/59
UVD2.2 MS DS 52/57/66
UVD2.2 MS MFT 51/57/65
VP4 CoreAVC 55/56/57
UVD2.2 LAV CB 48/53/55
UVD2.2 OC CoreAVC 46/53/56
UVD2.2 CoreAVC 44/51/55
2. Samsung-30fps
1. QS MS MFT 234/271/341
QS MS DS 224/266/333
QS CoreAVC 229/263/321
QS LAV NATIVE 219/259/325
QS LAV QS 134/162/190
QS FFDShow 136/150/160
CPU Core i5@3.1 99/132/197
VP5 CUDA 82/115/129
VP5 CoreCUDA 89/107/123
QS LAV CB 86/105/128
VP5 MS DS 93/105/121
UVD2.2 OC MS MFT 52/62/75
UVD2.2 OC LAV NATIVE 55/62/71
UVD2.2 OC MS DS 50/62/72
UVD2.2 OC LAV CB 53/55/58
VP4 MS MFT 34/55/91
VP4 LAV CB 34/55/84
VP4 LAV CUDA 31/55/90
VP4 LAV NATIVE 35/54/83
VP4 MS DS 33/54/82
VP4 CoreCUDA 32/51/77
UVD2.2 OC CoreAVC 39/49/57
CPU C2D@2.83 34/49/81
UVD2.2 MS MFT 35/46/62
UVD2.2 LAV CB 35/46/56
UVD2.2 LAV NATIVE 35/45/54
UVD2.2 MS DS 32/45/56
VP4 CoreAVC 27/41/62
UVD2.2 CoreAVC 30/38/44
3. Basket-60fps
DXVA checker 2.8.0b3 used for DXVA-CB, QS, NVCUVID and CPU.
1. QS CoreAVC 461/504/550
QS MS MFT 455/502/567
QS MS DS 455/502/553
QS LAV NATIVE 439/483/536
CPU Core i5@3.1 265/286/317
QS LAV QS 200/203/209
QS FFDShow 160/151/172
VP5 CoreCUDA 137/149/164
VP5 MS DS 118/133/149
QS LAV CB 109/115/120
VP4 CoreCUDA 77/84/106
VP4 LAV CB 75/82/99
VP4 MS MFT 74/82/107
CPU C2D@2.83 72/82/103
VP4 MS DS 75/81/89
VP4 LAV CUDA 73/81/99
VP4 LAV NATIVE 71/81/103
UVD2.2 OC LAV NATIVE 74/76/77
UVD2.2 OC MS MFT 70/76/82
UVD2.2 OC MS DS 62/76/78 *
UVD2.2 OC CoreAVC 56/57/58
UVD2.2 LAV NATIVE 55/57/59
UVD2.2 LAV CB 54/57/58
UVD2.2 MS MFT 53/57/65
VP4 CoreAVC 54/57/66
UVD2.2 MS DS 41/57/59 *
UVD2.2 CoreAVC 42/44/50
4. Girls-60fps
1. QS MS DS 410/423/436
QS MS MFT 410/420/456
QS CoreAVC 401/414/430
QS LAV NATIVE 397/412/430
CPU Core i5@3.1 198/209/234
QS LAV QS 193/200/203
QS FFDShow 160/166/171
VP5 LAV CUDA 141/143/146
VP5 CoreCUDA 110/128/144
VP5 MS DS 114/122/133
QS LAV CB 93/97/109
VP4 LAV CB 74/76/80
VP4 LAV NATIVE 74/76/79
VP4 LAV CUDA 74/76/78
VP4 MS MFT 73/76/79
UVD2.2 OC MS MFT 72/76/81
UVD2.2 OC MS DS 72/76/77 *
VP4 MS DS 73/75/77
UVD2.2 OC LAV NATIVE 71/75/77
VP4 CoreCUDA 67/73/81
CPU C2D@2.83 63/70/85
UVD2.2 OC LAV CB 51/59/60
UVD2.2 MS MFT 52/57/62
UVD2.2 LAV NATIVE 54/56/59
UVD2.2 LAV CB 54/56/58
UVD2.2 OC CoreAVC 55/56/57
VP4 CoreAVC 54/56/58
UVD2.2 MS DS 50/56/59 *
UVD2.2 CoreAVC 42/45/50
5. Cat-60fps
No MFT splitter for M2TS files
1. QS CoreAVC 394/402/411
QS MS DS 383/400/412
QS LAV NATIVE 372/381/387
QS LAV QS 190/194/197
CPU Core i5@3.1 157/187/206
QS FFDShow 156/161/166
VP5 LAV CUDA 138/141/147
VP5 CoreCUDA 134/138/145
QS LAV CB 117/118/125
VP5 MS DS 78/98/106
UVD2.2 OC MS DS 69/74/76
VP4 CoreCUDA 68/74/81
UVD2.2 OC LAV NATIVE 70/73/74
VP4 LAV CUDA 68/72/78
VP4 LAV CB 67/72/78
VP4 MS DS 68/71/79
VP4 LAV NATIVE 65/71/77
CPU C2D@2.83 54/62/71
UVD2.2 OC LAV CB 57/58/59
UVD2.2 OC CoreAVC 54/55/56
UVD2.2 LAV NATIVE 54/55/56
UVD2.2 LAV CB 53/55/57
UVD2.2 MS DS 50/55/58
VP4 CoreAVC 48/50/56
UVD2.2 CoreAVC 42/44/46
6. Avatar-60fps
MS MFT crashes DXVA Checker all versions, all platforms
1. QS MS DS 328/345/367
QS CoreAVC 322/330/336
QS LAV NATIVE 322/329/337
QS LAV QS 189/193/195
QS FFDShow 155/162/166
CPU Core i5@3.1 143/159/175
QS LAV CB 108/114/122
VP4 MS DS 65/76/84
VP4 LAV CUDA 64/76/82
VP4 LAV CB 64/76/82
VP4 LAV NATIVE 67/73/80
UVD2.2 OC LAV NATIVE 68/70/73
UVD2.2 OC MS DS 68/70/72 *
UVD2.2 OC LAV CB 55/57/58
UVD2.2 OC CoreAVC 53/55/57
UVD2.2 LAV CB 48/53/57
UVD2.2 LAV NATIVE 49/52/54
CPU C2D@2.83 44/52/63
UVD2.2 MS DS 39/52/55 *
VP4 CoreAVC 46/51/57
UVD2.2 CoreAVC 40/42/44
7. Vortex-24fps
1. QS FFDShow 122/125/129
QS LAV QS 119/121/124
QS MS DS 119/120/120
QS MS MFT 119/120/120
QS LAV NATIVE 118/120/122
QS CoreAVC 118/119/122
QS LAV CB 112/117/121
VP5 LAV CUDA 71/73/76
VP5 CoreCUDA 72/73/76
VP5 MS DS 71/73/75
VP5 MS MFT 72/73/76
CPU Core i5@3.1 56/59/61
UVD2.2 OC LAV CB 35/37/44
UVD2.2 OC MS MFT 34/37/42
UVD2.2 OC LAV NATIVE 36/36/39
UVD2.2 OC MS DS 36/36/37
UVD2.2 OC CoreAVC 28/29/31
UVD2.2 LAV NATIVE 26/28/28
UVD2.2 LAV CB 25/27/32
UVD2.2 MS DS 25/26/29
UVD2.2 MS MFT 24/26/34
CPU C2D@2.83 19/24/27
VP4 LAV CUDA 21/22/26
VP4 LAV CB 21/22/25
UVD2.2 CoreAVC 21/22/25
VP4 MS MFT 21/22/24
VP4 CoreCUDA 21/22/24
VP4 LAV NATIVE 20/22/24
VP4 MS DS 19/22/24
VP4 CoreAVC 17/18/21
8. Birds-24fps
[U]1. QS MS MFT 110/118/133
QS CoreAVC 110/117/129
QS FFDShow 112/115/120
QS LAV NATIVE 109/114/122
QS LAV QS 108/114/122
QS MS DS 108/113/122
QS LAV CB 78/83/88
VP5 CoreCUDA 71/77/87
VP5 LAV CUDA 59/69/77
VP5 MS DS 59/63/71
CPU Core i5@3.1 50/54/59
UVD2.2 OC MS MFT 32/38/47
UVD2.2 OC LAV NATIVE 36/37/43
UVD2.2 OC LAV CB 34/37/41
UVD2.2 MS DS 35/37/39
UVD2.2 OC CoreAVC 28/30/34
UVD2.2 LAV NATIVE 26/27/34
UVD2.2 LAV CB 25/27/35
UVD2.2 MS DS 25/27/35
UVD2.2 MS MFT 23/27/33
UVD2.2 CoreAVC 21/23/27
VP4 CoreCUDA 20/22/29
VP4 MS MFT 19/22/30
VP4 LAV CUDA 10/22/35 (A lot of breaks)
VP4 LAV CB 19/21/29
VP4 LAV NATIVE 19/21/28
VP4 MS DS 18/21/24
CPU C2D@2.83 18/21/23
VP4 CoreAVC 15/18/25
9. Ducks -30fps
1. QS CoreAVC 119/136/147
QS MS MFT 117/135/152
QS LAV NATIVE 118/134/144
QS LAV QS 116/130/139
QS FFDShow 129/129/132
QS MS DS 123/125/126 *
VP5 CoreCUDA 74/83/95
VP5 LAV CUDA 71/82/92
QS LAV CB 67/75/91
CPU Core i5@3.1 56/63/71
VP5 MS DS 48/57/84
UVD2.2 OC LAV NATIVE 36/41/48
UVD2.2 OC MS MFT 36/41/48
UVD2.2 OC LAV CB 35/41/48
UVD2.2 OC MS DS 30/39/43 *
UVD2.2 OC CoreAVC 29/33/37
UVD2.2 LAV CB 25/30/39
UVD2.2 MS MFT 24/30/38
UVD2.2 LAV NATIVE 26/29/35
UVD2.2 MS DS 21/27/32 *
VP4 LAV CUDA 15/26/39
UVD2.2 CoreAVC 22/25/29
VP4 MS MFT 21/25/34
VP4 CoreCUDA 21/25/31
VP4 LAV CB 20/25/34
VP4 LAV NATIVE 20/25/30
CPU C2D@2.83 20/24/29
VP4 MS DS 19/24/31 *
VP4 CoreAVC 18/21/27
10. Crowd Run-25fps
1. QS FFDShow 118/118/118
QS MS DS 112 (Only Average result)
QS MS MFT 109/109/110
QS CoreAVC 109/109/109
QS LAV QS 109/109/109
QS LAV NATIVE 107/107/107
QS LAV CB 71/72/73
VP5 LAV CUDA 68/70/72
VP5 CoreCUDA 68/69/70
VP5 MS DS 59/59/59
CPU Core i5@3.1 52/53/54
UVD2.2 OC LAV CB 32/33/37
UVD2.2 OC LAV NATIVE 32/33/33
UVD2.2 OC MS MFT 31/33/35
UVD2.2 OC MS DS 25/31/35
UVD2.2 OC CoreAVC 26/27/28
UVD2.2 LAV CB 23/24/27
UVD2.2 MS MFT 23/24/25
UVD2.2 LAV NATIVE 23/24/24
UVD2.2 MS DS 16/22/24
VP4 LAV CUDA 20/21/23
VP4 MS MFT 20/21/23
VP4 LAV CB 20/21/23
VP4 CoreCUDA 20/21/22
CPU C2D@2.83 19/21/22
VP4 LAV NATIVE 20/20/22
UVD2.2 CoreAVC 20/20/21
VP4 MS DS 18/20/22 *
VP4 CoreAVC 17/18/19
* MS DS decoder has a lot of artifacts at the beginning of the decoding, resulting low min value and probably lower average value
VC-1/WMV3
A) VC-1 - Devil May Cry 1080/60p-40Mbps
1. QS LAV QS 214/218/222
QS FFDShow 158/167/170
UVD2.2 OC AMD MFT 87/88/89
UVD2.2 OC LAV NATIVE 87/88/89
VP4 LAV CUVID 80/80/84
VP4 LAV CB 79/80/84
VP4 LAV NATIVE 76/80/84
CPU Core i5@3.1 68/76/96
UVD2.2 AMD MFT 65/66/67
UVD2.2 LAV NATIVE 64/66/69
UVD2.2 OC LAV CB 51/57/59
UVD2.2 LAV CB 48/55/56
CPU C2D@2.83 41/44/47
B) WMV3 - MP10 Digital Life 1080/24p-10Mbps
1. QS LAV QS 242/247/252
QS FFDShow QS 170/172/175
CPU Core i5@3.1 89/102/110
UVD2.2 OC AMD MFT 87/90/91
UVD2.2 OC LAV NATIVE 87/90/91
VP4 LAV CUVID 84/89/101
VP4 LAV CB 82/83/86
VP4 LAV NATIVE 81/83/84
UVD2.2 AMD MFT 68/68/69
UVD2.2 LAV NATIVE 68/68/69
CPU C2D@2.83 56/63/68
UVD2.2 OC LAV CB 57/58/59
UVD2.2 LAV CB 55/57/59
Comments:
1) The performance of QuickSync HW is beyond any competition, using native DXVA mode with every decoder used (MS DS/MFT, CoreAVC, LAV NATIVE).
The performance of copy-back mode using Intel's MSDK QuickSync decoder v0.28 software (FFDshow, LAV QS) is heavily multi-threaded and optimized but for some reason is a lot slower than QS decoder v0.20 (more than 20%) and a lot slower from DXVA2 native with the above system configuration in high frame rate clips - 60fps and/or low bitrate clips (Clips 1 to 6)
But it's very good and sometimes faster than native DXVA when used for high birate - low frame rate clips (Clips 7 to 10)
The performance of LAV DXVA2 copy-back is simply awful. It's slower than VP5 most of the times!
For laptop users, or for people who want their CPU and GPU load as low as possible during playback mode, native DXVA2 decoders (MS DS/MFT, CoreAVC, LAV NATIVE) are BY FAR the most efficient decoders.
2) VP5 is about 2 times faster than VP4 in "easy" low bitrate clips from 1 to 6. But it's more than 3 times faster in "difficult" high bitrate clips from 7 to 10, as it is built for 4K x 2K decoding. For those huge bandwidth clips, it closes the gap with QuickSync, but the distance is still obvious.
3) VP4 has a lot of problems at huge bandwidths starting from clip 7 up to 10 like 5750 UVD2.2, although the latter is a little faster. It seems absolutely reasonable for Nvidia to go for VP5 in order to support 4K x 2K and large bandwidths.
UVD2.2 OC can play easily every clip from 1 to 10!
4) UVD2.2 OC is about 35% - 42% faster than 5750 UVD2.2 in H.264 and 32% faster in VC-1/WMV3 and it's faster than VP4 and Core2Duo too in bandwidth heavy files starting from clip 7.
pirlouy
16th November 2011, 22:58
Interesting results.
If I'm not wrong, after flashing your BIOS, all sample can be read without dropped frames (fps always > sample fps).
Not the kind of stuff I'd try with my 5770 though. :-)
NikosD
31st December 2011, 12:32
Only VP5 results are missing...
I would like to see some results especially for clips from 7 to 10.
wanezhiling
2nd January 2012, 11:53
clip 3 4 5 6... I doubt the result.
Use PotPlayer DXVA, AMD never reach 60fps, even on HD6990... nVidia and Intel are 60 easily~
NikosD
2nd January 2012, 12:03
clip 3 4 5 6... I doubt the result.
Use PotPlayer DXVA, AMD never reach 60fps, even on HD6990... nVidia and Intel are 60 easily~
Don't!
Flashing 5750 BIOS with 6750 BIOS put UVD 2.2 in a "special" PowerPlay mode with 3D clocks (core:710 MHz/ Memory:1160 MHz - Sapphire Vapor-X edition).
So there is an overclock to UVD2.2, only allowed in this strange situation (5750 flashed by 6750 BIOS)
When I tried to edit my 5750 BIOS in order to work in UVD mode at 3D clocks like the flashed 5750, it worked but with no performance advantage.
The only solution seems to be a flashed 5750 card.
I don't have a real 6750 card to check.
And VP4 going to 93% to 97% utilization during playback in PotPlayer DXVA for 5th and 6th clip is not easily!
wanezhiling
2nd January 2012, 14:56
http://www.gokuai.com/f/2u7UmrSwOj6ehL26
http://www.gokuai.com/f/S9yFt7d55975qK38
NikosD, here's two 1080p60fps samples, please just use "your (6)750 + PotPlayer DXVA + Fraps" to test it.:rolleyes:
mariush
2nd January 2012, 15:40
For those that have problems using Rapidshare or Hotfile or can't download large files without disconnections, I'm still hosting these files and the previous ones here:
ftp://helpedia.com/pub/multimedia/x264/testvideos/
Resume supported, can use several download threads to speed things up etc...
NikosD
2nd January 2012, 17:12
http://www.gokuai.com/f/2u7UmrSwOj6ehL26
http://www.gokuai.com/f/S9yFt7d55975qK38
NikosD, here's two 1080p60fps samples, please just use "your (6)750 + PotPlayer DXVA + Fraps" to test it.:rolleyes:
1) Using PotPlayer you don't need Fraps.
PotPlayer has internal statistics and file information activated by pressing Tab key.
2) Your second file called "SNSD_-_Tell_Me_Your_Wish_(Genie)" is EXACTLY the same file called "Girls" the 4th clip of my collection. It took me 30 minutes to download the same clip I had.
PLEASE PAY ATTENTION to what I write at my first post, read the whole text, download the files and CHECK THINGS by YOURSELF.
3) Your first file is a little strange.
Although it reaches bitrates up to 188Mbps, it is not that difficult after all.
But it crashes MS MFT decoder in DXVA checker.
5750 MS DS 54/56/60
6750 MS DS 72/75/79
As you can see it's completely playable under (6)750
4) You should notice the performance of (6)750 not only in clips from 3 to 6, but from 7 to 10 which are unplayable by 5750 UVD and VP4 in realtime.
They are more difficult clips that show the performance of (6)750 and make the difference from 5750 UVD and VP4.
Maybe AMD graphics card holders should make some "noise" with my accidental finding, about how AMD "handles" UVD2.2/UVD3 in their cards.
Take a look here:
http://www.agile-news.com/news-323468-Photo:-no-fear-of-4K-x-2K!-Fire-whirlwind-2-HD6570-2G-Daniel-Edition-hardware-solution-2160p!.html
How is it possible for UVD3 to accelerate 4K x 2K video files with 238Mbps bitrate and 14% CPU usage and the same card cannot play 1080p60 ???
wanezhiling
3rd January 2012, 01:33
Fraps is useful, I don't like PotPlayer's OSD info, it has many mistakes..
"It took me 30 minutes to download the same clip I had." --> My fault:p
Cuz it really shocked me when I glanced at your (6)750's performance...
I never reach 60fps with "HD6850 + PotPlayer DXVA" just like 5750, so maybe I should flash my BIOS to test this again..
6750 MS DS 72/75/79 --> What I can say is congratulation~
In fact, I had made this comparison(PureVideo VS QuickSync VS UVD) for several times on my forum, I got same result with you except (6)750 .:)
http://forum.doom9.org/showthread.php?p=1546068#post1546068 This is VP5 DXVA 2160P from my forum,you must remember it.:)
btw, I hate english!:devil:as a remote eastern person, I can understand every word you say, but..:devil:
NikosD
3rd January 2012, 09:09
For me - 5750 UVD2.2 - it's not working even though DXVA checker says about the Device decoders:
"ModeH264_VLD_NoFGT: DXVA2, 720x480 / 1280x720 / 1920x1080 / 3840x2160"
"ModeH264_VLD_NoFGT_Flash: DXVA2, 720x480 / 1280x720 / 1920x1080 / 3840x2160"
So from the side of hardware and driver, UVD 2.2 is capable of 4K x 2K and if CoreAVC DXVA is capable of 4K x 2K, as it is clearly seen by your screenshots, then the problem must be an artificial restriction in everything else but VP5 inside the code of CoreAVC DXVA.
DXVA checker reports for clips beyond 1080p and CoreAVC DXVA the following:
"ModeUnknown (NV12): DXVA1 (VMR)"
and of course it's not working.
In fact, I had made this comparison(PureVideo VS QuickSync VS UVD) for several times on my forum, I got same result with you except (6)750 .:)
Which is your forum ?
This is VP5 DXVA 2160P from my forum,you must remember it.:)
Do you have access on VP5 hardware or could someone else from your forum post some benchmark results for clips 1 to 10 or at least 7 to 10 ?
btw, I hate english!:devil:as a remote eastern person, I can understand every word you say, but..:devil:
No problem at all!
nm
3rd January 2012, 09:40
Take a look here:
http://www.agile-news.com/news-323468-Photo:-no-fear-of-4K-x-2K!-Fire-whirlwind-2-HD6570-2G-Daniel-Edition-hardware-solution-2160p!.html
How is it possible for UVD3 to accelerate 4K x 2K video files with 238Mbps bitrate and 14% CPU usage and the same card cannot play 1080p60 ???
Weren't you guys talking about 6750, which isn't the same card at all. 6750 has UVD 2.2, but the agile-news article is about 6570, which has UVD 3.
NikosD
3rd January 2012, 11:36
Weren't you guys talking about 6750, which isn't the same card at all. 6750 has UVD 2.2, but the agile-news article is about 6570, which has UVD 3.
I was referring to a previous comment of Wanezhiling
Use PotPlayer DXVA, AMD never reach 60fps, even on HD6990... nVidia and Intel are 60 easily~
It is true BTW.
Take whatever 5xxx or 6xxx card and try to run a demanding 1080p60fps clip in any player you want.
I don't have a real 6xxx card but take a look here when renq tried the first 3 clips:
http://forum.doom9.org/showthread.php?p=1489378#post1489378
I rewrite the results of UVD3 here by renq:
DIVX ver is 9.01.21
Arcsoft ver 2.27.319.108
FFDShow DXVA ver is 3800
All post-processing turned OFF.
1. Clip
MS MFT - 48/57/85
DIVX DXVA - 48/58/85
Arcsoft - 45/57/69
Ffdshow - 44/57/70
2. Clip
DIVX - 38/48/75
MS DTV/DVD - 31/48/86
Arcsoft - Wouldn't play
FFDSHow - Same
3. Clip
DIVX - 51/57/84
MS DTV/DVD - 48/57/79
Arcsoft - 53/57/64
ffdshow - 54/57/71
So when you see the original UVD2.2 performance by 5750 and the performance of UVD3 posted above which is the same as of 5750, how is it possible the same UVD3 card to perform like here:
http://www.agile-news.com/news-323468-Photo:-no-fear-of-4K-x-2K!-Fire-whirlwind-2-HD6570-2G-Daniel-Edition-hardware-solution-2160p!.html
In fact I can't actually even play in DXVA mode a simple clip beyond 1080p like 2048 x 1200 in any combination of codec/player I have tried.
What I'm trying to say is that I'm looking for a suitable H.264 DXVA codec and player (DXVA checker is fine for me) to check the performance in 4K x 2K of "original" UVD2.2, my "Frankenstein" UVD2.2+ and UVD3, VP5 by other users.
Also I'm trying to say that the perfomance of UVD2.2 and UVD3 is burried by AMD by not suitable BIOS/ drivers.
I would really like to know the codec and player used by http://www.agile-news.com to play in DXVA mode 4K x 2K video clip file with UVD3
wanezhiling
3rd January 2012, 18:38
Which is your forum ?
A Chinese PotPlayer fan's forum, I'm not administrator,just a moderator.
I totally agree with your interesting things: QS's amazing speed, VP4's poor performance at huge bandwidths(I only have Duck takes off 1080p30p), AMD never reach 60fps expect your (6)750 and so on.
Because I tested hd3850 - m hd4650 - hd5770 - hd6850, 8600gt - 9300m gs - gt240 - gts450, I5 2300(hd2000)
Do you have access on VP5 hardware or could someone else from your forum post some benchmark results for clips 1 to 10 or at least 7 to 10 ?
I had contacted the geforce 410m's owner (http://forum.doom9.org/showthread.php?p=1546068#post1546068), but he say he had deleted those samples(screenshots) which I given to him..
And I had to tell you one thing: we can't download clips 1 to 10, we even cannot browse those download link page because of national policy...sigh...
Fortunately, 410m's owner promised me to re-download 2 sample(Duck takes off 1080p30p and clip4"girls"). Tomorrow I will post his benchmark results here.:)
I found you always look for a way to make AMD's 4k x 2k, my notebook, m hd4650, DXVA checker also says it support 3840x2160...but even on HD6850, I never seen it decoded 2160p, I only seen BSOD....:devil:
So from the side of hardware and driver, UVD 2.2 is capable of 4K x 2K and if CoreAVC DXVA is capable of 4K x 2K, as it is clearly seen by your screenshots, then the problem must be an artificial restriction in everything else but VP5 inside the code of CoreAVC DXVA.
No, VP5 can't decode 4K x 2K by CoreAVC DXVA. CoreAVC CUDA is ok. you can see that screenshots. In fact besides CoreAVC CUDA, VP5 can docode 4K x 2K by PotPlayer self DXVA(note: old version,at least before 2011.11), TMT5, MainConcept(Broadcast) AVC/H.264 deocoder.
Failure lists: PowerDVD11 DXVA, ffdshow DXVA, MPC-HC DXVA, CoreAVC DXVA, MS DTV-DVD, LAV CUVID
nm
3rd January 2012, 19:29
I was referring to a previous comment of Wanezhiling
Ah, I see now. Thanks for the full explanation. :)
NikosD
3rd January 2012, 19:57
And I had to tell you one thing: we can't download clips 1 to 10, we even cannot browse those download link page because of national policy...sigh...
Sorry to hear that...I' ve heard in the news about the Chinese government's restrictions on Internet use and you confirm the bad news.
Maybe you could try the link below:
ftp://helpedia.com/pub/multimedia/x264/testvideos/
It has all 10 clips.
No, VP5 can't decode 4K x 2K by CoreAVC DXVA. CoreAVC CUDA is ok. you can see that screenshots. In fact besides CoreAVC CUDA, VP5 can docode 4K x 2K by PotPlayer self DXVA(note: old version,at least before 2011.11), TMT5, MainConcept(Broadcast) AVC/H.264 deocoder.
Failure lists: PowerDVD11 DXVA, ffdshow DXVA, MPC-HC DXVA, CoreAVC DXVA, MS DTV-DVD, LAV CUVID
Thanks for the info.
wanezhiling
4th January 2012, 06:07
Maybe you could try the link below:
ftp://helpedia.com/pub/multimedia/x264/testvideos/
OK, I'll try it @30k/s..:( So the complete result should be here tomorrow.. I hate the ISP!!:mad:
Thanks for the info.
It doesn't matter.
OK, this's VP5's Benchmark on Clip4 and Clip9 (http://www.gokuai.com/f/9hs278K80ZB7X6cJ).
I5 2430M
nVidia 410M
Driver: ForceWare 267.54/Win7 64
4G DDR3
4. Girls-60fps
VP5 CoreCUDA 110/128/144
VP5 CUVID 141/143/146
VP5 MS DS 114/122/133
9. Ducks -30fps
VP5 CoreCUDA 74/83/95
VP5 CUVID 71/82/92
VP5 MS DS 48/57/84
;)Absolutely, VP5 is much faster than VP4. Think about one thing: 410M is a so "weak" graphics card, imagine if GTX560ti is VP5:D
nevcairiel
4th January 2012, 07:47
Think about one thing: 410M is a so "weak" graphics card, imagine if GTX560ti is VP5:D
The speed of the card is of little importance for progressive video.
The Video decoder is separate from the 3D engine, and always runs at the same speed. Its not faster in faster cards.
Assuming there is enough memory bandwidth, every card using VP5 should run at the same speed.
For example (from VP4):
A GT430 and a GTX570 have the same video decoding performance, despite their massive differences in 3D power.
Anyhow, NVIDIA said that VP5 would be double the speed of VP4 approximately.
Sadly, thats still way below the Intel decoder. And IVB which is coming in 3-4 month will once again increase that performance.
Now all Intel has to do is improve their drivers and iron out the last remaining issues. :)
NikosD
4th January 2012, 09:40
The speed of the card is of little importance for progressive video.
The Video decoder is separate from the 3D engine, and always runs at the same speed. Its not faster in faster cards.
Assuming there is enough memory bandwidth, every card using VP5 should run at the same speed.
This is maybe true for Video decoder hardware inside Nvidia cards and ATI cards (with the exception of my Frankenstein (6)750 card), but it's not true for Intel Video Hardware.
For ATI and 5xxx series I know that UVD2.2 clocks are 400MHz for Video processor and 900 MHz for memory.
For AMD and 6xxx series (including 67xx cards) I assume that UVD2.2/UVD3 is clocked at the same frequency of 400/900 MHz, from the performance I see by UVD3.
From my GT440 I see that VP4 is clocked at max 3D clocks of the card.
I know nothing about other VP4 cards, but I believe Nevcariel saying that they all clock at the same speed (frequency).
So for progressive video - not interlaced - the pure video decoding performance will be the same for every AMD/ Nvidia graphics card using the same video processor, assuming there are no tricks from AMD/ Nvidia regarding BIOS/drivers for specific models and cards (favoring specific models over some others by clocking the video processor to different speeds for example)
For the interlaced video, the deinterlacing process is executed by GPU shaders or computational units or whatever name you call them.
So the GPU performance in general and not video processor alone, does matter for interlaced video.
Also GPU performance does matter regarding video processing like De-noise, De-blocking, Edge-Enhancement etc.
These post-processing filters and many more are executed by AMD GPU shaders at hardware/driver level for AMD cards.
Of course those kind of video filters and many others can be executed by software like FFDshow in the CPU if you have a weak graphics card and a powerful processor.
For Intel the QuickSync video engine has the same speed of the GPU inside the processor, which varies from 650 MHz for the low end processors, up to 1350 MHz for the upper class processors working in GPU turbo mode.
For example my Core i5-2400 uses QS from 850 MHz to 1100MHz when Core i7-2600K uses the same QS from 850 MHz to 1350 MHz.
So for Intel hardware the choice of the CPU/GPU has some difference in the performance of video decoders.
Anyhow, NVIDIA said that VP5 would be double the speed of VP4 approximately.
Sadly, thats still way below the Intel decoder. And IVB which is coming in 3-4 month will once again increase that performance.
From the figures of Wanezhiling it seems that for "easy" clips - meaning clips with low bitrate like 4. Girls - VP5 has less than double the performance of VP4.
But for "heavy" clips with huge bitrates like 9. Ducks, VP5 has more than 3 times the performance of VP4 which is great.
Because VP5 seems to be a very well balanced video processor, pushing the performance figures where it is needed - to large bitrates required by 4K x 2K and special encoded clips like from 7 to 10 in my collection.
And for those clips - 7 to 10 - the performance of VP5 is close to QS 1st generation.
Of course Ivy will extend the gap even more in order to support multiple 4K x 2K streams simultaneously and 4K x 4K (square resolution)
wanezhiling
4th January 2012, 10:07
nev, NikosD, I really appreciate it that you correct my wrong opinion.:thanks:
And I always wonder know this reason (http://forum.doom9.org/showthread.php?p=1529788#post1529788)... Since the ForceWare 285.62, this became worse and worse...:(
btw, can you tell me the principle of Intel's amazing speed(more than several times the performance of nVidia/AMD)? Thx.
NikosD
4th January 2012, 11:32
btw, can you tell me the principle of Intel's amazing speed(more than several times the performance of nVidia/AMD)? Thx.
This is a hard question to be answered in detail and very technical.
Maybe Egur (Eric Gur) at http://forum.doom9.org/showthread.php?p=1523738#post1523738 can answer this with the help of "inside" Intel information.
First of all Intel's MFX engine (QS) has an obvious advantage over UVDx and VPx.
Greater speed (frequency).
MFX works in the range of 850 MHz - 1100 MHz for most of the desktop processors.
AMD UVD2.2 works at 400MHz ! with no dynamic change of frequency, which means that no matter if you play or benchmark a clip (full load), UVD2.2 always works at 400 MHz
UVD3, VP4 and I'm sure VP5 have dynamic change of working frequency depending on the load.
VP4 reaches 820MHz during benchmarking (full load).
Of course it's not only about frequency, for example because of the integration inside a very fast processor MFX can take advantage of very fast access to memory/ caches.
In general I could say that your question seems of the same principle as of saying why SandyBridge is faster than Bulldozer, or why AMD 79xx series graphics cards (Tahiti architecture) are faster than Nvidia's 5xx series (GF110 architecture)
Of course Sandy, Bulldozer, Southern Island, GF110 architecture chips are extremely complicated chips inside, they are monsters with billions transistors, but executing in the end the same kind of x86 instructions (more or less) for CPU and same kind of D3D, OpenGL, OpenCL instructions for GPUs but with a much different way between different architectures.
Video processors like VP4/5, UVD2.2/3, QuickSync - to be more exact the video decoding/encoding engine of QuickSync is called MFX engine (Multi-Format Codec) - are a lot lot simpler processors than a modern CPU or a GPU processor.
They use fixed function logic, not general purpose logic like CPUs and GPUs (in our days), and they have a form of an ASIC.
They do actually just one thing but they do it extremely fast with very low power consumption, comparing to CPU when you see the low frequency and low number of transistors used by Video processors.
That thing is decoding Video compression algorithms like MPEG-2, MPEG-4 ASP, MPEG-4 AVC, VC-1.
If you study those algorithms you will see that most of the resources needed to decode them are used for Inverse Discrete Cosine Transformations (iDCT) which is a mathematical equation/ transformation.
So if we go deeper, the performance of MFX engine as of every fast video decoder has to do about how quickly performs iDCT, but of course this is a very simple approach of video processor performance.
For example VP5 increased the performance of huge bitrate video clips much more than low bitrate comparing to VP4.
That point has to do with internal changes to access in memory/ caches and wider buses etc.
We need hardware experts and specialized knowledge to go deeper from here, I think!
wanezhiling
4th January 2012, 15:08
Thanks NikosD, I have got some useful info by your words.
http://www.dvbsupport.net/
PotPlayer 1.5.31298 is coming!
Added Nvidia CUDA Decoder (MPEG1/2, H264/AVC, VC1 are supported)
nevcairiel
4th January 2012, 15:12
I wonder how much of that is stolen from my decoder, you know, they aren't known for inventing their own stuff. :p
wanezhiling
4th January 2012, 17:40
410M owner failed to download Clip2..
VP5's Benchmark on Clip1-10(expect Clip2) (http://www.gokuai.com/f/LsNI0R2Wwq4QOV2h):
1. Twinpeaks-30fps
VP5 CUVID 129/139/143
VP5 MS DS 133/138/141
VP5 MS MFT 128/137/141
VP5 CoreCUDA 83/89/92
3. Basket-60fps
VP5 CoreCUDA 108/145/163
VP5 CUVID failed by unknown error
VP5 MS DS 95/131/148
4. Girls-60fps
VP5 CUVID 141/143/146
VP5 CoreCUDA 110/128/144
VP5 MS DS 114/122/133
5. Birds-60fps
VP5 CUVID 102/138/148
VP5 MS DS 81/110/119
VP5 MS MFT 86/108/119
VP5 CoreCUDA 87/98/104
6. Cat-60fps
VP5 CUVID 105/138/151
VP5 CoreCUDA 133/137/142
VP5 MS DS 78/98/106
7. Vortex-24fps
VP5 MS DS 71/73/75
VP5 CUVID 71/73/75
VP5 CoreCUDA 72/73/75
8. Birds-24fps
VP5 CoreCUDA 71/77/85
VP5 CUVID 59/69/77
VP5 MS DS 59/63/71
9. Ducks -30fps
VP5 CoreCUDA 74/83/95
VP5 CUVID 71/82/92
VP5 MS DS 48/57/84
10. Crowd Run-25fps
VP5 CUVID 68/69/70
VP5 CoreCUDA 68/69/70
VP5 MS DS 59/59/59
So we needn't Clip2's benchmark result, this's enough.:)
I have a question: Why MFT only shows on Clip1 and 5?
NikosD
4th January 2012, 18:08
So we needn't Clip2's benchmark result, this's enough.:)
I have a question: Why MFT only shows on Clip1 and 5?
Nice work !
If you are sure that the owner of VP5 has done the tests 3 times in a row for each clip and he didn't take the figures by the first run, I can put the results to my first post to be easily compared with the rest of the hardware.
In order to bench MS MFT you have to install an MFT splitter for MKV files.
The only one I know and use is the DivX MFT splitter for MKV video container.
For MOV and MP4 containers, Win 7 has built-in MFT Splitter support.
CruNcher
4th January 2012, 18:22
I wonder how much of that is stolen from my decoder, you know, they aren't known for inventing their own stuff. :p
Hehe though there where also already nvcuvid based Decoder from Asia before yours and even CoreCodecs but alot of those operate in the dark or have no real feedback chain or doesn't want foreign feedback @ all and so they are staying crap except implementations of the bigger isvs ;) Most improvements due to feedback came from Donald (Doom9) anyways and found their way into the Driver and so got pushed to everyone in the ecosystem early on :)
I don't think Potplayer needs to steal that (their Direct3D Render and XML GUI implementations seem also very unique, not sure where they should have stolen that) whoever works on the DXVA part seems to know what hes doing quiet well so surely he will have had no problems implementing nvcuvid and for sure MSDK is next ;D
It's true they steal or better borrow a lot of the subtitle stuff from MPC-HCs development though ;)
And even if it's stolen they put the pieces together (ffmpeg,dxva,splitter,subtitle,gui,Osd, now they start with 3rd party apis) very efficiently and advance very fast as a whole in one Media Player :)
compared to their Asian competition like Baofeng or KMP (older potplayer integration) they advance faster especially in stability :)
wanezhiling
4th January 2012, 18:35
If you are sure that the owner of VP5 has done the tests 3 times in a row for each clip and he didn't take the figures by the first run, I can put the results to my first post to be easily compared with the rest of the hardware.
3 times..maybe tomorrow.. Now it's 1:33 in China, too late..I should go to bed..
In order to bench MS MFT you have to install an MFT splitter for MKV files.
The only one I know and use is the DivX MFT splitter for MKV video container.
For MOV and MP4 containers, Win 7 has built-in MFT Splitter support.
:thanks:Just as you say "It seems that MS DS and MS MFT have almost the same speed"
I add a DucksTakeOff 2160P figure, you can download it again.
VP5 CoreCUDA 2160P 26/28/31
It's hard to keep Avg 30fps.. btw,PotPlayer DXVA could keep.
NikosD
4th January 2012, 18:46
http://www.dvbsupport.net/
PotPlayer 1.5.31298 is coming!
VP5 CoreCUDA 2160P 26/28/31
It's hard to keep Avg 30fps.. btw,PotPlayer DXVA could keep.
Can you ask them why did they stopped supporting PotPlayer's internal DXVA for 4K x 2K resolution in latest builds - you said that we need previous versions before 11/2011 - and if they are going to support it in next versions including ATI's hardware UVD2.2/3 and not only VP5 ?
wanezhiling
4th January 2012, 19:12
Can you ask them why did they stopped supporting PotPlayer's internal DXVA for 4K x 2K resolution in latest builds - you said that we need previous versions before 11/2011
The developer adds a DXVA compatibility check like MPC-HC because of the feedback about BSOD.
But the truth is that PotPlayer cancelled the support of beyond 1080,not a DXVA compatibility check!..sigh..
and if they are going to support it in next versions including ATI's hardware UVD2.2/3 and not only VP5 ?
I also wanna PotPlayer return 2160P's support, but it seems that korean has no this plan from the reply I sent email.
Maybe you can try this: Send it to "ahahlive@hanmail.net"(PotPlayer's official address)
oh no, I must go to bed..It's 2:12..
wanezhiling
5th January 2012, 11:44
NikosD, a good news: PotPlayer 1.5.31323 released (http://cafe.daum.net/pot-tool/AZMV/8566).
+ DXVA 사용 조건에 해상도 추가
It means that 2160P DXVA return!:)
English version will come soon.
NikosD
5th January 2012, 12:20
That's good.
I have exchanged 9 emails with "ahahlive@hanmail.net" from yesterday, trying to persuade them to support DXVA 4K x 2K.
I'm waiting for the English version to check it out.
wanezhiling
5th January 2012, 14:17
If you are sure that the owner of VP5 has done the tests 3 times in a row for each clip and he didn't take the figures by the first run, I can put the results to my first post to be easily compared with the rest of the hardware.
Now it's ok. 3 times in a row for each clip, and choose the best one
VP5's Benchmark on Clip1-10 (http://www.gokuai.com/f/O0Y7g91x6w0n53co):
1. Twinpeaks-30fps
VP5 CUVID 130/139/143
VP5 MS DS 133/138/141
VP5 MS MFT 128/137/141
VP5 CoreCUDA 84/89/93
2. Samsung-30fps
VP5 CUVID 82/115/129
VP5 CoreCUDA 89/107/123
VP5 MS DS 93/105/121
3. Basket-60fps
VP5 CoreCUDA 137/149/164
VP5 CUVID failed by unknown error
VP5 MS DS 118/133/149
4. Girls-60fps
VP5 CUVID 141/143/146
VP5 CoreCUDA 110/128/144
VP5 MS DS 114/122/133
5. Birds-60fps
VP5 CUVID 132/142/148
VP5 MS DS 107/111/118
VP5 MS MFT 104/109/117
VP5 CoreCUDA 89/99/106
6. Cat-60fps
VP5 CUVID 138/141/147
VP5 CoreCUDA 134/138/145
VP5 MS DS 78/98/106
7. Vortex-24fps
VP5 CUVID 71/73/76
VP5 CoreCUDA 72/73/76
VP5 MS DS 71/73/75
VP5 MS MFT 72/73/76
8. Birds-24fps
VP5 CoreCUDA 71/77/87
VP5 CUVID 59/69/77
VP5 MS DS 59/63/71
9. Ducks -30fps
VP5 CoreCUDA 74/83/95
VP5 CUVID 71/82/92
VP5 MS DS 48/57/84
10. Crowd Run-25fps
VP5 CUVID 68/70/72
VP5 CoreCUDA 68/69/70
VP5 MS DS 59/59/59
extra. Ducks - 30fps - 2160P
VP5 CoreCUDA 26/28/31
PotPlayer 1.5.31323 en ver is available. (http://www.dvbsupport.net/download/index.php?act=view&id=239)
NikosD
5th January 2012, 14:51
Thanks!
I' ll put the results in the first post.
I tried the player the moment it got out, but with no luck.
I forced it to always use for every resolution DXVA and I got a fullscreen GREEN frame during DXVA H.264 playback, but otherwise the playback is normal.
I can hear the sound and the slider is moving with the right speed and I get realtime fps/playback for most of my 4K x 2K clips.
Also PotPlayer says that is using DXVA during H.264 4K playback.
The same Green frame appears when I enable DXVA MPEG2_VLD and DXVA MPEG4ASP_VLD with the help of DXVA checker from within hidden features of the driver.
It seems that all these features (MPEG2 VLD, MPEG4 ASP VLD and H.264 over 1080) are not "officially" supported by UVD2.2, although you can enable them in the drivers.
Try it the new PotPlayer with VP5 and I'm curious if someone could try it with UVD3.
wanezhiling
5th January 2012, 16:09
I tried the player the moment it got out, but with no luck.
I forced it to always use for every resolution DXVA and I got a fullscreen GREEN frame during DXVA H.264 playback, but otherwise the playback is normal.
I can hear the sound and the slider is moving with the right speed and I get realtime fps/playback for most of my 4K x 2K clips.
Also PotPlayer says that is using DXVA during H.264 4K playback.
So is my m hd4650(UVD2.0).. And I saw BSOD again when playing a 4K x 2K sample.:devil:
The same Green frame appears when I enable DXVA MPEG2_VLD and DXVA MPEG4ASP_VLD with the help of DXVA checker from within hidden features of the driver.
It seems that all these features (MPEG2 VLD, MPEG4 ASP VLD and H.264 over 1080) are not "officially" supported by UVD2.2, although you can enable them in the drivers.
I also can enable DXVA MPEG2_VLD and DXVA MPEG4ASP_VLD with the help of DXVA checker,but there's nothing but greenscreen..
HD6850 really could DXVA MPEG4ASP by PotPlayer DXVA, Click this (http://we.pcinlife.com/data/attachment/forum/201201/05/2241151stzspeskzwsrjsw.jpg).;) In fact UVD3.0 could do this by TMT5, AMD Hardware MFT Playback Decoder(AMDhwDecoder.dll) and so on.
nVidia VP4/VP5 also can DXVA MPEG4ASP only by LAV CUVID Decoder (http://forum.doom9.org/showthread.php?t=160290)
Try it the new PotPlayer with VP5 and I'm curious if someone could try it with UVD3.
We have tested VP5, nothing changed~~ It's still easy to DXVA 2160P (http://forum.doom9.org/showthread.php?p=1546068#post1546068):cool:
Offer a new figure "Girls 2160p (http://we.pcinlife.com/data/attachment/forum/201201/05/2231341lwsgsswqeezqoww.jpg)", this sample is horrible 120fps.. VP5 can only decode up to 30fps..sigh..
Once I had a chance to test a HD6850, I tried every player/decoder I know to play 4K x 2K, but with no luck.. The only succeed one is a 1440p sample, click this (http://we.pcinlife.com/data/attachment/forum/201109/24/2217435znkoorncncarcov.jpg). but only 32fps..
PS: VP5 could reach 60fps,:cool: click this (http://we.pcinlife.com/data/attachment/forum/201201/05/2231358ga877p8v5v5ba2j.jpg)
btw, I'm a nVidia fan.
NikosD
5th January 2012, 16:59
So is my m hd4650(UVD2.0).. And I saw BSOD again when playing a 4K x 2K sample.:devil:
Although HD4650 uses UVD2.2, it's not the same as 5750 UVD 2.2 which is not the same as 6750 UVD2.2!
It's been since last summer to see a BSOD for that reason.
I also can enable DXVA MPEG2_VLD and DXVA MPEG4ASP_VLD with the help of DXVA checker,but there's nothing but greenscreen..
Me too
nVidia VP4/VP5 also can DXVA MPEG4ASP only by LAV CUVID Decoder (http://forum.doom9.org/showthread.php?t=160290)
LAV CUDIV as a separate decoder doesn't exist anymore.
Nevcariel put LAV CUVID inside LAV Video, but without porting MPEG4 ASP capability
Offer a new figure "Girls 2160p (http://we.pcinlife.com/data/attachment/forum/201201/05/2231341lwsgsswqeezqoww.jpg)", this sample is horrible 120fps.. VP5 can only decode up to 30fps..sigh..
I want to download this monster sample of 120fps.
Where can I find it ?
btw, I'm a nVidia fan.
BTW, I'm an AMD, Intel, Nvidia fan.
wanezhiling
5th January 2012, 17:17
Although HD4650 uses UVD2.2
No, mine is not HD4650(UVD2.2), it's Mobility HD 4650(UVD2.0). http://en.wikipedia.org/wiki/Unified_Video_Decoder
LAV CUDIV as a separate decoder doesn't exist anymore.
Nevcariel put LAV CUVID inside LAV Video, but without porting MPEG4 ASP capability
I know that, http://forum.doom9.org/showthread.php?p=1546801#post1546801
I want to download this monster sample of 120fps.
Where can I find it ?
ed2k://|file|Girls.Generation.Oh.4in1.201002.HDTV.x264.2160p.120fps.DTSES.6.1ch.mkv|2706942815|70cf7fd7e1dd0b61cae96dff22d99bc5|h=bwg7va6gmcdomin6orcreagnxvr5adlb|/
Do you know this (http://we.pcinlife.com/data/attachment/forum/201201/06/004128cmmgggyg7y7mu1ma.jpg)? It seems that WMVideo Decoder only support IDCT, not VLD...
NikosD
5th January 2012, 17:58
I don't use e2dk, so I have managed to download only the YouTube version which is a "light" one.
Only 30fps at 3840 x 2160 and 42.5 Mbps file.
It's 1 GB.
I can play it in DXVA mode with pure sound and Green frame at full speed though.
The original is 120 fps 3840 x 2160 and 100Mbps file.
The size is 2.52GB for 3 minutes :)
I'll keep trying to find it.
Yes Microsoft's decoders do not provide VC-1/ WMV3 VLD DXVA acceleration that's why AMD offers its own AMD Playback decoder MFT for VC1/ WMV3 and AMD Hadware MFT Playback Decoder for MPEG4 ASP
wanezhiling
5th January 2012, 18:30
I'll keep trying to find it.
http://115.com/file/clozg3pv#
Girls.Generation.Oh.4in1.201002.HDTV.x264.2160p.120fps.DTSES.6.1ch.part1.rar
http://115.com/file/aqajzuuw#
Girls.Generation.Oh.4in1.201002.HDTV.x264.2160p.120fps.DTSES.6.1ch.part2.rar
http://115.com/file/aqaj7hys#
Girls.Generation.Oh.4in1.201002.HDTV.x264.2160p.120fps.DTSES.6.1ch.part3.rar
The way to download them (http://we.pcinlife.com/data/attachment/forum/201201/06/012533pk55ajryualyb2z5.png) You also can use download manager software like IDM:)
PS:
http://www.gokuai.com/f/B05Yj637KS6yyu75
4096 x 3072, maybe you could collect it.
Yes Microsoft's decoders do not provide VC-1/ WMV3 VLD DXVA acceleration that's why AMD offers its own AMD Playback decoder MFT for VC1/ WMV3 and AMD Hadware MFT Playback Decoder for MPEG4 ASP
PotPlayer can't DXVA WMV3 for AMD :devil:
NikosD
5th January 2012, 18:59
Thanks!
I have already found it in Mediafire in 15 parts.
I'm in the end now and yes I use IDM too.
PS:
http://www.gokuai.com/f/B05Yj637KS6yyu75
4096 x 3072, maybe you could collect it.
Yes that too.
PotPlayer can't DXVA WMV3 for AMD :devil:
The reason PotPlayer doesn't offer WMV3 acceleration is because it supports only iDCT and MoComp, not VLD.
On the other hand, AMD's hardware supports only VLD for VC-1/WMV3.
So...
As I have said to tetsuo55, WMV3 and VC-1 are fully compatible.
So if one has DXVA VLD support for VC-1 it could use the same decoder for WMV3.
I don't know if MPC-HC has corrected that - because they have VC1 VLD support but they didn't seem to use it for WMV3, but PotPlayer for an unknown reason to me, doesn't want to offer VLD acceleration for WMV3 as it does for VC-1
wanezhiling
5th January 2012, 19:21
but PotPlayer for an unknown reason to me, doesn't want to offer VLD acceleration for WMV3 as it does for VC-1
AMD's hardware supports only VLD for WMV3 by AMD Playback decoder(MFT)
PotPlayer doesn't support MFT, just DS.
DXVA Checker never say AMD has ModeWMV9_IDCT, this is why PotPlayer doesn't offer WMV3 DXVA support for AMD.
===================================
All my 2160p Clips is here (http://xhmikosr.1f0.de/index.php?folder=c2FtcGxlcy8yMTYwcA==).
Another 2:
http://115.com/file/be83t4l4#
HD.Club-4K-Chimei-inn-50mbps
http://115.com/file/e7zwgsgg#
Crowd_Run_2160p_50fps_275Mbps
Too late, I should sleep, bye~
NikosD
5th January 2012, 19:42
AMD's hardware supports only VLD for WMV3 by AMD Playback decoder(MFT)
PotPlayer doesn't support MFT, just DS.
DXVA Checker never say AMD has ModeWMV9_IDCT, this is why PotPlayer doesn't offer WMV3 DXVA support for AMD.
It seems that you have misunderstood something...
If you run DXVA checker, it shows the Decoder devices installed by the driver for the underline hardware UVD2.2.
So for UVD2.2 it doesn't have WMV9 AT ALL.
It has only VC-1.
So, in order to accelerate WMV3 files in either DS or MFT you have to use VC-1 decoder device.
But the VC-1 decoder device, the ModeVC1_VLD you see in DXVA Checker is VLD ONLY.
It doesn't matter if you use DS or MFT VC-1 decoder to accelerate WMV3.
The only things that matter for WMV3 acceleration are:
1) Use VC-1 decoder device, because there is no WMV9
2) Use it in VLD mode ONLY.
AMD installs an MFT VC-1/WMV3 decoder which uses VC1_VLD
PotPlayer, if you check the "Built-in video decoder settings" has three options for WMV3:
1) Disable
2) IDCT
3) MoComp
It should have a 4th option:
4) VLD (bit stream decoder) like VC-1 which is below WMV3 settings
Check the DXVA preferences of the built-in video decoder settings and you will understand better what I'm saying to you.
=================================
All my 2160p Clips is here (http://xhmikosr.1f0.de/index.php?folder=c2FtcGxlcy8yMTYwcA==).
Another 2:
http://115.com/file/be83t4l4#
HD.Club-4K-Chimei-inn-50mbps
http://115.com/file/e7zwgsgg#
Crowd_Run_2160p_50fps_275Mbps
Too late, I should sleep, bye~
I have downloaded them all !
wanezhiling
5th January 2012, 19:58
Check the DXVA preferences of the built-in video decoder settings and you will understand better what I'm saying to you
I understand what you say, thanks. I'll keep learning.:)
but PotPlayer for an unknown reason to me, doesn't want to offer VLD acceleration for WMV3 as it does for VC-1
ahahlive@hanmail.net ;) This's always the best way.
egur
5th January 2012, 22:26
Can you benchmark ffdshow-QS on the VC1 and wmv9 clips?
renq
6th January 2012, 12:03
As soon as I get my ASUS mobo back from RMA, I'll (re)test all the 10 clips and perhaps some more on the 6950's UVD3 ;)
wanezhiling
6th January 2012, 12:47
NikosD, http://cafe.daum.net/pot-tool/AZMV/8571
+ 내장 DXVA 디코더에 WMV9 VLD모드 추가
Add support for WMV9 VLD...
You won again:D
wanezhiling
6th January 2012, 16:04
VP5 DXVA 4096 x 3072 (http://we.pcinlife.com/data/attachment/forum/201201/06/230440e20a0a3zeo0zaecz.jpg)
NikosD
6th January 2012, 16:14
Can you benchmark ffdshow-QS on the VC1 and wmv9 clips?
I updated the VC-1 results but I didn't have the choice inside FFDShow video decoder settings of Intel QuickSync for WMV3 files.
NikosD
6th January 2012, 16:15
As soon as I get my ASUS mobo back from RMA, I'll (re)test all the 10 clips and perhaps some more on the 6950's UVD3 ;)
Very good, I'll put the results at my first post.
nevcairiel
6th January 2012, 16:15
I updated the VC-1 results but I didn't have the choice inside FFDShow video decoder settings of Intel QuickSync for WMV3 files.
The QuickSync WMV3 decoder is a software decoder, and its rather slow, so don't bother with that. ;)
NikosD
6th January 2012, 16:18
NikosD, http://cafe.daum.net/pot-tool/AZMV/8571
Add support for WMV9 VLD...
You won again:D
This time it took me 4 emails only :)
Did you try it for ATi cards ?
Because initially they told me that only Nvidia cards could benefit from WMV9 VLD.
I'll try it myself later when I will have access to my ATi card again.
wanezhiling
6th January 2012, 16:23
I had tried on my m hd4650 and gt240, yep, VLD, Cheers~~;)
PS: Intel is still IDCT...(why?)
wanezhiling
9th January 2012, 14:52
http://cafe.daum.net/pot-tool/AZMV/8571
wanezhiling 12.01.07. 20:53
It seems that Nvidia and AMD benefit from WMV9 VLD.
Intel is still WMV9 IDCT, why? And VC1 is IDCT too... God...
팟플.개발자 19:27
Intel's VC1 VLD mode(ClearVideo) is diffrent.
So it cannot support.
different... Does anyone know the difference?:confused:
NikosD
9th January 2012, 15:38
Intel has two different Decoder devices for VC-1 installed regarding Intel HD graphics (sandy):
ModeVC1_VLD_ClearVideo
ModeVC1_VLD_2_ClearVideo
I don't know the difference.
I know that egur (Eric Gur) posted that VC-1 hardware inside Sandy can use HW acceleration for Advanced profile of VC-1 only.
The other profiles (Simple and Main) are handled by software decoding.
WMV3 is compliant with Simple and Main VC-1 profiles, so it can't get any HW acceleration (just like Simple and Main VC-1 profiles)
This is the same reason that I couldn't benchmark WMV9 as Egur asked me to do so in a previous post.
There is no WMV9 HW acceleration.
And I'm a little curious why Egur asked me to benchmark WMV9, BTW :rolleyes:
I don't know if this is a hardware/drivers/Media SDK or other kind limitation.
CruNcher
9th January 2012, 15:55
Arcsofts Decoder supports also WMV3 streams it shows the Decoder being the VC-1 Bitstream (ClearVideo), though i didn't found a bitstream yet that shows a big difference decoding wise vs either Libavs Software Decoder or WMV Video Decoder. Though what really surprised me VC-1 (WVC1) decodes more efficient with Intels MSDK Decoder then a DXVA Decoder it seems it shows much lower utilization even with the Memory copy going on but im still looking into that.
So for WMV3 MP@HL it looks even DXVAed it's pretty much the same power consumption as Libavs Decoder (so on the first sight pretty useless)
http://img824.imageshack.us/img824/9205/arcsoftdxvawmv3mphl.png
Though the WVC1 decoding result was surprising
NikosD
9th January 2012, 16:09
We are talking about Full bitstream acceleration (VLD), not HW acceleration in general.
Intel's drivers support DXVA decoder devices of ModeVC1_IDCT and ModeWMV9_IDCT.
We are not talking about them, they are hardware assisted (partial) acceleration.
PotPlayer added WMV9 VLD support for suitable hardware/drivers.
CruNcher
9th January 2012, 16:33
I see the on paar results of Libav and Arcsofts results Decoding WMV3 Main show also that they pretty similar most probably indeed Software Decoding so comparable to Libav
Also the difference between Arcsofts WVC1 decoding and MSDK (LAV Video,ffdshow-quicksync) result could indicate that Arcsoft only uses the old VC1 mode that is not such performant else i wouldn't understand why MSDK has better Performance with the Memory Copy then Arcsofts DXVA Decoder
wanezhiling
9th January 2012, 19:28
I5 2300(HD2000), DXVA VC1(WVC1) failed by MPC HC..
We know MPC HC's DXVA is based on VLD, so this proved that Korean was right, Intel's VC1_VLD(ClearVideo)must be unusual.
Hope developer will explain this I ask to him tomorrow..
-------------------------------------------
Because Intel's VC1 VLD is not standard struct DXVA.
But intel's IDCT mode is standard DXVA.
I request document VC1 VLD mode to intel, but intel cannot do anything...
:devil:Korean developer said nothing...
Maybe Egur can answer this..
NikosD
10th January 2012, 10:09
A lot of people report problems with Intel Video Hardware/drivers working together with pure, vanilla DXVA.
That's why Egur is trying to build a new approach of using Intel's video hardware to accelerate decoding through Intel's Media SDK - the QS decoder project.
It's like a requirement to use Media SDK in order to HW accelerate video decoding of MPEG-2, VC-1, H.264 through DXVA for Intel.
On the other hand, Nvidia and AMD can work without problems nowdays with direct, pure DXVA.
From my limited tests I had a lot of problems with MS DS H.264 decoder and QuickSync hardware (corrupted images and crashing during benchmarking in DXVA checker), but no problem at all with MS MFT H.264 decoder and QuickSync.
Unfortunately Intel's MFT decoders for MPEG-2, VC-1, H.64 are broken, they don't even enumerate in DXVA checker.
A lot of effort is needed by the teams of Intel Graphics Driver development and Intel Media SDK development to reach the maturity of video drivers by AMD and Nvidia.
mariner
10th January 2012, 17:12
...
Another 2:
http://115.com/file/be83t4l4#
HD.Club-4K-Chimei-inn-50mbps
http://115.com/file/e7zwgsgg#
Crowd_Run_2160p_50fps_275Mbps
...
Greetings wanezhiling and NikosD.
Thanks for posting these interesting clips. Was able to get hold of a Sapphire 6670 DDR5 and did some quick tests:
1. Could not get DXVA to work in Potplayer playing these 2160p clips, using both internal and external filters. What settings are required?
2. Was able to get AS video decoder to run in DXVA using mpc and dxva_checker, but with corrupted output and low frame rate.
3. Benchmark for the other clips are in line with 6750's results. A few fps higher, but lower than VP5.
4. The Duck clip (9) failed to enumerate.
5. Finally, how does one identify if UVD3 is in use?
Many thanks and best regards.
NikosD
10th January 2012, 17:34
1. Could not get DXVA to work in Potplayer playing these 2160p clips, using both internal and external filters. What settings are required?
If you force enable of DXVA use in PotPlayer latest version for every resolution and still can't play 2160p files, then there is nothing you can do about.
But you have to force it by checking "Always use" in "Resolution limit".
Maybe you could try CoreAVC DXVA as external filter to PotPlayer.
2. Was able to get AS video decoder to run in DXVA using mpc and dxva_checker, but with corrupted output and low frame rate.
ATI is still limiting UVD3 I suppose for 4K playback.
3. Benchmark for the other clips are in line with 6750's results. A few fps higher, but lower than VP5.
So UVD3 is faster than normal UVD2.2 and UVD2.2+
Can you post your results here ?
I could update them to first post too in order to have an all around comparison.
5. Finally, how does one identify if UVD3 is in use?
The easiest way for me is by observing the gpu/memory frequency.
From idle low clocks, goes to UVD mode clocks during video playback/benchmarking and then back to idle clocks.
BTW, can you see if the clocks of GPU change during benchmarking ?
I mean from lower to higher as the clip moves on.
Another way could be by using AMD GPU clock Tool which has UVD status report and it's for older cards and , but I think it works for yours too.
wanezhiling
10th January 2012, 18:00
HD7970 is still UVD3.0, just adds a VCE, sigh..
I'm waiting for Kepler,hope it's VP5+:D
"The Way it's meant to be played" :p
I never expect AMD.
NikosD
10th January 2012, 18:08
UVD2.x/3.x has a lot more power than it is exposed by current BIOS/drivers.
Don't understimate AMD's hardware, only get angry with their tactics and policies.
mariner
10th January 2012, 18:19
Thanks for the reply, NikosD.
If you force enable of DXVA use in PotPlayer latest version for every resolution and still can't play 2160p files, then there is nothing you can do about.
But you have to force it by checking "Always use" in "Resolution limit".
Maybe you could try CoreAVC DXVA as external filter to PotPlayer.
I meant Potplayer reverted to YUY2 when forced to use DXVA, for both internal and external decoders, Arcsoft's included.
Using version 31393.
ATI is still limiting UVD3 I suppose for 4K playback.
How did this chap do it?
http://www.agile-news.com/news-323468-Photo:-no-fear-of-4K-x-2K!-Fire-whirlwind-2-HD6570-2G-Daniel-Edition-hardware-solution-2160p!.html
So UVD3 is faster than normal UVD2.2 and UVD2.2+
Can you post your results here ?
I could update them to first post too in order to have an all around comparison.
The 6670 (don't know if it's UVD3 or not) is faster by about 5fps. Will try to download all the clips.
Facing a couple of issues here:
1. AMD MFT crashed with the h264 mp4 clips,
2. Duck clip would not enumerate.
Catalyst 11.12
The easiest way for me is by observing the gpu/memory frequency.
From idle low clocks, goes to UVD mode clocks during video playback/benchmarking and then back to idle clocks.
BTW, can you see if the clocks of GPU change during benchmarking ?
I mean from lower to higher as the clip moves on.
Another way could be by using AMD GPU clock Tool which has UVD status report and it's for older cards and , but I think it works for yours too.
Will take a look. Do you mean this one?
http://www.techpowerup.com/downloads/1128/AMD_GPU_Clock_Tool_v0.9.8.html
Thanks.
wanezhiling
10th January 2012, 18:29
HD6000 is UVD3.0 expect 6770/6750.. So HD6670 is 3.0.
NikosD, I got a HD7970's DXVA Checker (http://we.pcinlife.com/data/attachment/forum/201201/11/0119101z55hxsx2sfihhf2.jpg):D, compared with HD6850 (http://we.pcinlife.com/data/attachment/forum/201201/11/011910fvb3zb1ljbfyql3f.jpg), no difference at all..
NikosD
10th January 2012, 18:41
I meant Potplayer reverted to YUY2 when forced to use DXVA, for both internal and external decoders, Arcsoft's included.
Using version 31393.
Then it seems to fall back to software mode.
CoreAVC, Cyberlink DXVA maybe
How did this chap do it?
http://www.agile-news.com/news-323468-Photo:-no-fear-of-4K-x-2K!-Fire-whirlwind-2-HD6570-2G-Daniel-Edition-hardware-solution-2160p!.html
I would really like to know.
I'm guessing with special driver/BIOS and/or special Video player.
Facing a couple of issues here:
1. AMD MFT crashed with the h264 mp4 clips
You have to disable AVIVO transcoding from Catalyst Control Center.
Do you mean this one?
http://www.techpowerup.com/downloads/1128/AMD_GPU_Clock_Tool_v0.9.8.html
Yes.
Better use a desktop gadget like GPU Observer for GPU/VPU info.
I use it.
6670 has UVD3 inside.
NikosD
10th January 2012, 18:50
HD6000 is UVD3.0 expect 6770/6750.. So HD6670 is 3.0.
NikosD, I got a HD7970's DXVA Checker (http://we.pcinlife.com/data/attachment/forum/201201/11/0119101z55hxsx2sfihhf2.jpg):D, compared with HD6850 (http://we.pcinlife.com/data/attachment/forum/201201/11/011910fvb3zb1ljbfyql3f.jpg), no difference at all..
It has one more line at the end, I don't know what it is stands for.
Don't expect much from the DXVA Checker screenshot, because the main info from that is the codecs and resolutions accelerated.
There are no more codecs to be accelerated.
MPEG-2, MPEG-4 ASP, MPEG-4 AVC, MVC, VC-1, WMV3 are already in hardware.
I can't think of any other codec right now in real need of HW acceleration.
The thing that matters right now is speed for 4K playback with H.264.
wanezhiling
10th January 2012, 19:04
The thing that matters right now is speed for 4K playback with H.264.
VP5 has showed that, I believe Kepler will be faster in terms of speed.
As mariner saied "6670 is faster by about 5fps than (6)750", so assuming that 6670 could dxva 4K, how do you think it's capable of this (http://xhmikosr.1f0.de/index.php?folder=c2FtcGxlcy8yMTYwcC9EdWNrc1Rha2VPZmY=)(2160p,370M bitrate,50fps)?:rolleyes:
NikosD
10th January 2012, 19:12
VP5 has showed that, I believe Kepler will be faster in terms of speed.
I think Kepler, whenever comes, will have exactly same speed because it will have exactly same VPU - VP5
As mariner saied "6670 is faster by about 5fps than (6)750", so assuming that 6670 could dxva 4K, how do you think it's capable of this (http://xhmikosr.1f0.de/index.php?folder=c2FtcGxlcy8yMTYwcC9EdWNrc1Rha2VPZmY=)(2160p,370M bitrate,50fps)?:rolleyes:
That's what I'm saying all the time.
DON'T BELIEVE ATI!
The hardware is more capable than ATI wants you to think!
wanezhiling
10th January 2012, 19:33
I think Kepler, whenever comes, will have exactly same speed because it will have exactly same VPU - VP5
Compared with GT240, GTX560TI does have same speed, so maybe you're right, but maybe a little surprise, just let it happen.
The hardware is more capable than ATI wants you to think!
:p
btw,if possible, I'll try to contact with hd7970's owner to test the power of "UVD3.0+VCE".
mariner
12th January 2012, 07:36
144mbps 4k/60p benchmark
Now that the JVC GY-HMQ10 has been released, would like to see 4k/60p benchmark posted as well.
It seems UVD3 would most likely fail. Lets see if VP5 or QS (or is it actually vlc?) would make the cut.
Thanks.
http://pro.jvc.com/prof/attributes/features.jsp?model_id=MDL102132
CruNcher
12th January 2012, 16:00
hehe QS in Sandy Bridge @ least is hardly capable of QFHD @ 50Mbps 30 fps so this 60 fps will stutter even more, though their is no possibility to use that resolution for DXVA yet which could ultimately confirm the Decoder isn't capable of it, but for now it looks like it's to weak and Ivy Bridge will support it.
NikosD
12th January 2012, 18:36
QS not capable of 4K ?
Says who ?
Have you done any tests by yourself and proved that is not capable ?
QS is clearly faster than VP5.
So if VP5 is capable of 4K, how is it possible a faster HW like QS not being able to accelerate 4K ?
nevcairiel
12th January 2012, 18:39
Speed is not the only concern, especially with fixed-function hardware, it may just not have been designed for 4K frame sizes.
At least the Driver thinks its not capable, because it doesn't expose a 4K DXVA2 mode.
NikosD
12th January 2012, 19:14
So they have to "open" driver to support up to 4K.
And of course MSDK too, because it seems that everything has to pass through MSDK in order to work DXVA properly.
NikosD
12th January 2012, 20:57
I did some tests today with QuickSync HW:
1) 4K playback
4K is not possible in HW due to driver's and MSDK restrictions.
The software fallback works OK with PotPlayer using both pure DXVA internal codec and QuickSync "internal" codec.
Also the fallback works OK with CoreAVC DXVA, LAV video and QS decoder, but MSDK uses slowest software decoding than ALL the others (Pot internal, LAV, CoreAVC)
So, if you play out of Intel's HW DXVA video files, AVOID the QS decoder software MSDK routines, they are not optimized as FFMpeg, CoreAVC, LAV.
MS DS/MFT doesn't provide software fallback and crashes DXVA checker.
2) CoreAVC although it says it uses MSDK, it doesn't. It's pure DXVA and very fast implementation.
3) QS decoder is faster than LAV video QS and MS DS/ CoreAVC are both a lot faster than both MSDK solutions in 60fps clips.
But MS MFT is the fastest of all, only usable in WMP12 though.
4) For normal playback PotPlayer's internal DXVA codecs work like a charm for both MPEG-2/ H.264 progressive or interlaced.
Minimum power consumtion - at least 50% down from MSDK solutions (LAV QS, QS decoder) and easy playback of all up to 1080p video files.
You only need QS decoder for VC-1, because PotPlayer is not able to utilise VC1_VLD mode, only VC1_IDCT.
UPDATE:
I sent an email to ahahlive@hanmail.net, suggesting a method to utilise VC1_VLD with DXVA and QuickSync.
Let's hope it's feasible and see it in next PotPlayer.
nevcairiel
12th January 2012, 21:08
Obviously "pure" DXVA codecs are faster and more efficient o.O
NikosD
15th January 2012, 21:57
The last few days I was thinking and reading about OVD API - which means OpenVideo Decode API.
It's a Video Acceleration API similar to NVCUVID from Nvidia, but of course it's not using CUDA - which is exclusive for Nvidia's hardware.
OVD uses DXVA through OpenCL which like OpenGL is not exclusive to any HW or Platform and it was created by ATI/ AMD for ATI GPU.
But there is OpenCL support for GPU HW by both Nvidia and ATI, and after IvyBridge launch, Intel is going to support OpenCL for GPU HW too.
So although ATI started the project, it could probably integrated to Nvidia and Intel in the near future.
Anyway, for ATI i think that OVD API could resolve two major issues:
1) Fast Frame Copy version - useful for anyone who would like to use a different renderer than EVR, like madVR.
Using Frame Copy version through OVD API, which means through OpenCL, is the only way to have zero penalty during the frame copy process (From GPU to CPU).
Unfortunately there is no other fast way for that kind of process, regarding ATI's HW.
2) OVD API could lead to major performance increase, even greater than direct, pure DXVA.
The reason for this, is the clock of the GPU which means the clock of UVD.
Because using OpenCL pushes the clocks to maximum 3D clocks, which means a twofold increase for many cards out there.
Direct, pure DXVA put GPU in UVD mode where the clock is only 400MHz.
But the GPU is capable of more than 800MHz in 3D mode for many ATI cards.
The only drawback of that approach (OVD) is the increase in power consumption and maybe the unknown - at least to me - difficulty of OVD API.
CruNcher
16th January 2012, 01:48
I did some tests today with QuickSync HW:
1) 4K playback
4K is not possible in HW due to driver's and MSDK restrictions.
The software fallback works OK with PotPlayer using both pure DXVA internal codec and QuickSync "internal" codec.
Also the fallback works OK with CoreAVC DXVA, LAV video and QS decoder, but MSDK uses slowest software decoding than ALL the others (Pot internal, LAV, CoreAVC)
So, if you play out of Intel's HW DXVA video files, AVOID the QS decoder software MSDK routines, they are not optimized as FFMpeg, CoreAVC, LAV.
MS DS/MFT doesn't provide software fallback and crashes DXVA checker.
2) CoreAVC although it says it uses MSDK, it doesn't. It's pure DXVA and very fast implementation.
3) QS decoder is faster than LAV video QS and MS DS/ CoreAVC are both a lot faster than both MSDK solutions in 60fps clips.
But MS MFT is the fastest of all, only usable in WMP12 though.
4) For normal playback PotPlayer's internal DXVA codecs work like a charm for both MPEG-2/ H.264 progressive or interlaced.
Minimum power consumtion - at least 50% down from MSDK solutions (LAV QS, QS decoder) and easy playback of all up to 1080p video files.
You only need QS decoder for VC-1, because PotPlayer is not able to utilise VC1_VLD mode, only VC1_IDCT.
UPDATE:
I sent an email to ahahlive@hanmail.net, suggesting a method to utilise VC1_VLD with DXVA and QuickSync.
Let's hope it's feasible and see it in next PotPlayer.
http://forum.doom9.org/showpost.php?p=1551696&postcount=2121
Works also with Cyberlink though as i said the information how to utilize it is most probably under NDA and only so far with this nice fake Adapter which looks very similar to VC1Tweak from Madshi just that it supports WM ASF Reader you can utilize it for .wmv (though some bitstreams still fail (only a hand few), not sure why yet they are progressive and dont seem so much complex or different @ all, could be mux dependent) :)
PS: Got it working in MPC-HC also now via AV Splitter (it has a special interoperability table setting dynamic vc-1 tweak though not only for vc-1) though parkyoy.wmv seems the only one working many others fail (black screen audio plays) much more compared to Potplayers Adapter ;)
Here you can find some measurement done with MPC-HC trunk (EVR Custom) comparable to Potplayer (EVR Custom) MPC-HC experimental gave marginally better results (lower render overhead) :)
http://forum.doom9.org/showpost.php?p=1551882&postcount=553
pirlouy
16th January 2012, 13:15
Why would OVD be quicker than DXVA if is is based on DXVA ? :/
From your second point, it seems, it will just change AMD profile in order to go to a mode which uses more potential from the card. But that could be fixed for DXVA too if they fixed Drivers/firmware/hardware.
If it is based on DXVA, it can't be better. The only way to be better than DXVA is if you use your own way of communication with GPU (i.e. DXVA independent). I guess it's what CUDA allows...
nevcairiel
16th January 2012, 13:48
You can easily force your GPU into 3D mode, but that would not make decoding faster, just consumes more power.
CUDA does the same, it forces full performance mode (for actual CUDA functions), however it doesn't make decoding faster - the decoders run on their own clock domain, separate from the main GPU or the CUDA cores. Thats the same for NVIDIA and ATI - may be different for Intel. The only thing that benefits from faster clocks is deinterlacing, but the driver should automatically bring the card to full performance mode if the load is high enough.
The frame copying is never "zero penality", copying the frame always takes time (and CPU cycles). I cannot judge if OVD would make that faster then it is with DXVA2, noone really can without trying. One can at least hope!
Note that on the NGC-based graphics card (7xxx series), the DXVA2 frame copy is also quite a lot faster then on previous models.
OVD has one main disadvantage - its not based on D3D, which means you cannot use it for a "native" DXVA mode, and you cannot access the cards post-processing/deinterlacing functions.
Also, the API is indeed difficult to use. I don't know if anyone is actually using it.
Anyway, since i cannot prove this to you without writing or finding a decoder that uses OVD, i'll just leave it at this. :)
PS:
CUDA also isn't any faster then DXVA, its just a hell of a lot easier to use and has a much better error resilience.
NikosD
16th January 2012, 15:53
Why would OVD be quicker than DXVA if is is based on DXVA ? :/
From your second point, it seems, it will just change AMD profile in order to go to a mode which uses more potential from the card. But that could be fixed for DXVA too if they fixed Drivers/firmware/hardware.
True, but they DON'T FIX it.
So until then, forcing 3D clocks right now using OpenCL is not going to make OVD faster than DXVA, only the card itself - in particular the UVD ASIC will be faster (from 400MHz to 3D clocks).
But a faster UVD means faster decoding.
You can easily force your GPU into 3D mode, but that would not make decoding faster, just consumes more power.
CUDA does the same, it forces full performance mode (for actual CUDA functions), however it doesn't make decoding faster - the decoders run on their own clock domain, separate from the main GPU or the CUDA cores. Thats the same for NVIDIA and ATI - may be different for Intel.
Well, from my experiments with 5750 card as i have written before, when I push GPU clocks to 3D mode with original BIOS nothing happens in the decoding mode, just consumes more power, as you wrote.
BUT after flashing my 5750 with a BIOS of 6750, without changing clocks in BIOS, allowed the same card go in ACTUAL 3D mode clocks FOR UVD ENGINE.
So in my (6)750 I have decoding performance of a UVD2.2 clocked in 710 MHz and memory clocked in 1160MHz (GDDR5) instead of a 400/900 UVD2.2 which is the default frequency mode.
I don't know the reason, but I suspect that Firmware/Driver restrictions are keeping down the internal clocks of UVDx engine and after the BIOS flashing the restrictions stopped working.
Because my card is not recognized as a 5750 nor 6750.
It's both at the same time depending on the software.
That's why I call it (6)750 or Frankenstein card.
PS:
CUDA also isn't any faster then DXVA, its just a hell of a lot easier to use and has a much better error resilience.
Early benchmark results based on VP5 HW at my first post, not done by me, seem to indicate that CUDA sometimes is faster than DXVA - at least from the results/ codecs tested.
Of course "faster than DXVA" can't really be done when CUDA and OVD are based on DXVA.
But maybe poor/ slow implementations of direct DXVA decoders can be slower for specific clips than optimized CUVID decoders and in the future optimized OVD decoders.
So, from your words it seems that even if you own an AMD card, you wouldn't even try to write a decoder based on OVD?
nevcairiel
16th January 2012, 16:18
So, from your words it seems that even if you own an AMD card, you wouldn't even try to write a decoder based on OVD?
I would not, the API is not worth it.
I actually own a AMD card now, i just don't have it in my main dev PC. :)
NikosD
17th January 2012, 10:36
The OVD profiles of my card as reported by DXVA Checker are not even capable yet of H.264 L5.1 profile.
Profile
OVD_H264_Baseline_41: Yes
OVD_H264_Main_41: Yes
OVD_H264_High_41: Yes
OVD_H264_Baseline_51: No
OVD_H264_Main_51: No
OVD_H264_High_51: No
OVD_H264_Stereo_High: No
OVD_VC1_Simple: Yes
OVD_VC1_Main: Yes
OVD_VC1_Advanced: Yes
OVD_MPEG2_VLD: Yes
Last line - which exposes MPEG2_VLD for HD5000 series - forced ATI to declare that it's a bug of AMD APP :)
NikosD
24th January 2012, 17:29
NikosD, I got a HD7970's DXVA Checker (http://we.pcinlife.com/data/attachment/forum/201201/11/0119101z55hxsx2sfihhf2.jpg):D, compared with HD6850 (http://we.pcinlife.com/data/attachment/forum/201201/11/011910fvb3zb1ljbfyql3f.jpg), no difference at all..
Can you upload somewhere a DXVA Checker screenshot of your UVD 2.0 ?
I suppose you have Win 7 and recent Catalyst drivers, right ?:p
wanezhiling
24th January 2012, 17:57
http://i.imgur.com/VeTDw.png
http://i.imgur.com/IgXlZ.png
PS: http://www.gokuai.com/f/1Iy283I8786B8FG2;)
renq
28th January 2012, 20:16
Some Beta test results to follow;)
Rig:
I5-2500K
2x4GB 1600MHz
HD6950 w/ shaders unlocked @ 800/1250 (def 6950 clocks:o)
Windows 8 Dev Preview
Default Win8 AMD driver- 8.88.5.3 dated 20.09.2011
DXVAChecker properties:
Renderer: Enhanced Video Renderer (DirectShow)
Decoder: Microsoft DTV-DVD Video Decoder
Decoder Device: ModeH264_VLD_NoFGT
Processor Device: -
Results:
1.
Time: 00:14.293
Average FPS: 56,041
Min/Max FPS: 49 / 64
2.
Time: 01:03.296
Average FPS: 48,139
Min/Max FPS: 39 / 63
3.
Time: 01:24.346
Average FPS: 57,442
Min/Max FPS: 45 / 63
4.
Time: 03:25.727
Average FPS: 58,082
Min/Max FPS: 40 / 66
5.
Time: 00:42.109
Average FPS: 56,686
Min/Max FPS: 40 / 60
6.
Time: 00:52.456
Average FPS: 57,686
Min/Max FPS: 52 / 60
7.
Time: 00:11.802
Average FPS: 28,724
Min/Max FPS: 27 / 32
8.
Time: 00:18.077
Average FPS: 29,319
Min/Max FPS: 26 / 38
9.
Time: 00:13.095
Average FPS: 30,622
Min/Max FPS: 21 / 36
10.
Time: 00:04.600
Average FPS: 23,913
Min/Max FPS: 16 / 27
wanezhiling
28th January 2012, 21:02
3.
Average FPS: 57,442
4.
Average FPS: 58,082
5.
Average FPS: 56,686
6.
Average FPS: 57,686
http://forum.doom9.org/showpost.php?p=1548287&postcount=4 ;):D
NikosD's (6)750, what a Frankenstein card...
So nev, I think you needn't spend more time on DXVA2(copy-back) for 60fps movies, because it's AMD's issue.:p
NikosD
29th January 2012, 16:23
The results of Renq above, indicate that UVD3 "seems" to have same the performance as 5750 and not (6)750.
ATI makes UVD3 to look like UVD2.2 in terms of raw performance.
At the same time UVD3 according to ATI is capable of 4K HW acceleration :sly: and according to this (http://www.agile-news.com/news-323468-Photo:-no-fear-of-4K-x-2K!-Fire-whirlwind-2-HD6570-2G-Daniel-Edition-hardware-solution-2160p!.html)
Of course the above results are coming from a Beta OS like Windows 8.
But I think that Win 7 SP1 would give similar results.
I have no explanation for the contradicting results, statements and facts
renq
29th January 2012, 19:06
Of course the above results are coming from a Beta OS like Windows 8.
But I think that Win 7 SP1 would give similar results.
I have no explanation for the contradicting results, statements and facts
Well, let us see;)
Same rig.
Windows 7 X64 SP1
DXVAChecker 2.6.2 64bit (previous Windows 8 DP results were with 32bit IIRC, 95% certain of it)
LAVFilter 0.45
AMD Catalyst 12.1 WHQL (8.93-111205a-132104C-ATI)
Default Vid Quality settings in CCC:
* Edge Enh- 10
* De-noise 64
* Mosquito NR 50
* De-block 50
* Dyn contrast enabled, ESVP (Enforce smooth video playback) enabled
PS! LAV results are ALL Software mode results (i5-2500K @ 4,2GHz)
AVG Results:
1.
LAV video - 486,152
DS- 73,500
MFT- 68,980
2.
LAV video - 177,882
DS - 63,797
3.
LAV- FAIL
DSD - 76,834
4.
LAV - 334,202
DS - 76,559
5.
LAV - 257,164
DS - 75,807
MFT - 75,276
6.
LAV - 267,098
DS- 77,147
7.
LAV - 69,567
DS- 45,479
MFT - 37,715
8.
LAV - 64,448
DS- 40,021
9.
LAV - 76,746
DS - 43,753
10.
LAV - 62,916
DS - 36,268
More thorough results:
1.
Renderer: Enhanced Video Renderer (DirectShow)
Decoder: LAV Video Decoder
Decoder Device: -
Processor Device: -
Time: 00:01.697
Average FPS: 486,152
Min/Max FPS: 481 / 481
Renderer: Enhanced Video Renderer (DirectShow)
Decoder: Microsoft DTV-DVD Video Decoder
Decoder Device: ModeH264_VLD_NoFGT
Processor Device: -
Time: 00:10.898
Average FPS: 73,500
Min/Max FPS: 46 / 103
NOTE- No Video in Preview
Renderer: Enhanced Video Renderer (Media Foundation)
Decoder: Microsoft H264 Video Decoder MFT
Decoder Device: ModeH264_VLD_NoFGT
Processor Device: -
Time: 00:11.960
Average FPS: 68,980
Min/Max FPS: 49 / 88
2.
Renderer: Enhanced Video Renderer (DirectShow)
Decoder: LAV Video Decoder
Decoder Device: -
Processor Device: -
Time: 00:17.208
Average FPS: 177,882
Min/Max FPS: 134 / 243
CPU Usage (%): Avg: 00 Min: 00 Max: 00
GPU Usage (%): Avg: 03 Min: 00 Max: 20
NOTE- random(ish) peaks of gpu usage
Renderer: Enhanced Video Renderer (DirectShow)
Decoder: Microsoft DTV-DVD Video Decoder
Decoder Device: ModeH264_VLD_NoFGT
Processor Device: -
Time: 00:47.761
Average FPS: 63,797
Min/Max FPS: 46 / 93
NOTE- No VIdeo in preview window
3.
LAV- Doesn't do anything
Renderer: Enhanced Video Renderer (DirectShow)
Decoder: Microsoft DTV-DVD Video Decoder
Decoder Device: ModeH264_VLD_NoFGT
Processor Device: -
Time: 01:03.058
Average FPS: 76,834
Min/Max FPS: 71 / 80
NOTE- Green blocks corruption for first few secs of preview
4.
Renderer: Enhanced Video Renderer (DirectShow)
Decoder: LAV Video Decoder
Decoder Device: -
Processor Device: -
Time: 00:36.050
Average FPS: 334,202
Min/Max FPS: 316 / 376
CPU Usage (%): Avg: 00 Min: 00 Max: 00
GPU Usage (%): Avg: 11 Min: 00 Max: 35
NOTE- The GPU usage goes up to 35% occasionally
Renderer: Enhanced Video Renderer (DirectShow)
Decoder: Microsoft DTV-DVD Video Decoder
Decoder Device: ModeH264_VLD_NoFGT
Processor Device: -
Time: 02:36.075
Average FPS: 76,559
Min/Max FPS: 72 / 91
NOTE- No Video in Preview window
5.
Renderer: Enhanced Video Renderer (DirectShow)
Decoder: LAV Video Decoder
Decoder Device: -
Processor Device: -
Time: 00:09.667
Average FPS: 257,164
Min/Max FPS: 219 / 296
CPU Usage (%): Avg: 00 Min: 00 Max: 00
GPU Usage (%): Avg: 06 Min: 00 Max: 30
NOTE- The GPU usage goes up to 30% once per run (usually in the middle of the video/run/test)
Renderer: Enhanced Video Renderer (DirectShow)
Decoder: Microsoft DTV-DVD Video Decoder
Decoder Device: ModeH264_VLD_NoFGT
Processor Device: -
Time: 00:31.488
Average FPS: 75,807
Min/Max FPS: 73 / 85
NOTE- No Video in Preview window
Renderer: Enhanced Video Renderer (Media Foundation)
Decoder: Microsoft H264 Video Decoder MFT
Decoder Device: ModeH264_VLD_NoFGT
Processor Device: -
Time: 00:33.025
Average FPS: 75,276
Min/Max FPS: 70 / 78
6.
Renderer: Enhanced Video Renderer (DirectShow)
Decoder: LAV Video Decoder
Decoder Device: -
Processor Device: -
Time: 00:11.449
Average FPS: 267,098
Min/Max FPS: 244 / 294
CPU Usage (%): Avg: 00 Min: 00 Max: 00
GPU Usage (%): Avg: 12 Min: 00 Max: 30
NOTE- Up to 30% GPU usage 2-3 times per run!
Renderer: Enhanced Video Renderer (DirectShow)
Decoder: Microsoft DTV-DVD Video Decoder
Decoder Device: ModeH264_VLD_NoFGT
Processor Device: -
Time: 00:39.224
Average FPS: 77,147
Min/Max FPS: 74 / 99
NOTE- No Video in Preview window
7.
Renderer: Enhanced Video Renderer (DirectShow)
Decoder: LAV Video Decoder
Decoder Device: -
Processor Device: -
Time: 00:05.333
Average FPS: 69,567
Min/Max FPS: 68 / 71
Renderer: Enhanced Video Renderer (DirectShow)
Decoder: Microsoft DTV-DVD Video Decoder
Decoder Device: ModeH264_VLD_NoFGT
Processor Device: -
Time: 00:07.454
Average FPS: 45,479
Min/Max FPS: 17 / 79
NOTE- No Video in Preview window
Renderer: Enhanced Video Renderer (Media Foundation)
Decoder: Microsoft H264 Video Decoder MFT
Decoder Device: ModeH264_VLD_NoFGT
Processor Device: -
Time: 00:09.837
Average FPS: 37,715
Min/Max FPS: 36 / 41
8.
Renderer: Enhanced Video Renderer (DirectShow)
Decoder: LAV Video Decoder
Decoder Device: -
Processor Device: -
Time: 00:08.534
Average FPS: 64,448
Min/Max FPS: 58 / 72
Renderer: Enhanced Video Renderer (DirectShow)
Decoder: Microsoft DTV-DVD Video Decoder
Decoder Device: ModeH264_VLD_NoFGT
Processor Device: -
Time: 00:13.243
Average FPS: 40,021
Min/Max FPS: 20 / 94
NOTE- No Video in Preview window
9.
Renderer: Enhanced Video Renderer (DirectShow)
Decoder: LAV Video Decoder
Decoder Device: -
Processor Device: -
Time: 00:06.502
Average FPS: 76,746
Min/Max FPS: 65 / 93
CPU Usage (%): Avg: 00 Min: 00 Max: 00
GPU Usage (%): Avg: 03 Min: 00 Max: 10
Renderer: Enhanced Video Renderer (DirectShow)
Decoder: Microsoft DTV-DVD Video Decoder
Decoder Device: ModeH264_VLD_NoFGT
Processor Device: -
Time: 00:09.165
Average FPS: 43,753
Min/Max FPS: 23 / 83
NOTE- No Video in Preview window
10.
Renderer: Enhanced Video Renderer (DirectShow)
Decoder: LAV Video Decoder
Decoder Device: -
Processor Device: -
Time: 00:03.306
Average FPS: 62,916
Min/Max FPS: 59 / 64
CPU Usage (%): Avg: 00 Min: 00 Max: 00
GPU Usage (%): Avg: 04 Min: 00 Max: 09
Renderer: Enhanced Video Renderer (DirectShow)
Decoder: Microsoft DTV-DVD Video Decoder
Decoder Device: ModeH264_VLD_NoFGT
Processor Device: -
Time: 00:03.033
Average FPS: 36,268
Min/Max FPS: 33 / 41
NOTE- No Video in Preview window
nevcairiel
29th January 2012, 19:17
These DXVA numbers are impossible for an ATI card.
Are you sure it didn't use your Intel GPU or software or something? :) (It'll always use the primary GPU)
renq
29th January 2012, 19:43
These DXVA numbers are impossible for an ATI card.
Are you sure it didn't use your Intel GPU or software or something? :) (It'll always use the primary GPU)
Ur, right, it's SW mode, dunno why DXVAChecker shows 0% CPU util.
:o
NikosD
29th January 2012, 20:30
You have to check your ATI's clocks.
From idle mode clocks they have to go to UVD mode clocks (usually 400/900) or more.
Use GPU Observer gadget.
It works for Nvidia and AMD cards.
wanezhiling
29th January 2012, 20:58
http://forum.doom9.org/showpost.php?p=1554385&postcount=85
I never doubt it, it's true AMD.;)
http://forum.doom9.org/showpost.php?p=1554567&postcount=88
Obviously, made by I5 2500K(cpu mode)
0% CPU util, just a bug~~
renq
30th January 2012, 07:03
You have to check your ATI's clocks.
From idle mode clocks they have to go to UVD mode clocks (usually 400/900) or more.
Use GPU Observer gadget.
It works for Nvidia and AMD cards.
500/1250 according to GPU-Z:)
NikosD
30th January 2012, 13:09
Use DXVA Checker 32bit, not x64
NikosD
30th January 2012, 21:24
PS:
http://www.gokuai.com/f/B05Yj637KS6yyu75
4096 x 3072, maybe you could collect it.
VP5 DXVA 4096 x 3072 (http://we.pcinlife.com/data/attachment/forum/201201/06/230440e20a0a3zeo0zaecz.jpg)
These clips (first post - second post) are the same ?
Because I downloaded the first one and it says that is a Sorenson Spark clip at 117.586 fps.
Sorenson Spark is very close to H.263, not H.264.
So how can VP5 accelerate a clip in DXVA like the screenshot you posted at the second reference ?
Maybe PotPlayer uses MPEG1 or MPEG2 HW acceleration of VP5, because H.263 and Sorenson Spark is close to MPEG1/MPEG2, not H.264.
Can you confirm that the clip of second post - with VP5 screenshot - uses Sorenson Pack codec ?
nevcairiel
30th January 2012, 21:27
Sorenson Spark and H.263 are close to MPEG-4, not MPEG1/2
NikosD
30th January 2012, 21:35
True.
Sorenson Video (first versions) was close to MPEG1/MPEG2.
But H.263 and Sorenson Spark is close to MPEG-4.
And this makes things more complicated, because DXVA MPEG-4 is not even supported by PotPlayer for Nvidia HW :confused: (except H.264 of course)
EDIT:
-------------
PotPlayer says Input: AVC1, although in File info (from MediaInfo tool) says Sorenson Spark.
Somehow the Sorenson Spark format is HW accelerated via H.264.
CruNcher
30th January 2012, 23:54
and Sorenson SVQ is near to H.264 ;) you see the pattern, its a very common one see RealVideo and other codec implementer whenever they send some white paper to MPEG they work on something similar and sometimes a own codec comes up, and sometimes they don't agree and create something complete different see Microsoft (though they are back stabers its not really that unique,except psy ideas), Leadtools, Iterate and others ;)
Its a evolutionary process and only very few times something unique comes up somewhere else apart from the Mpeg Chain (of monster patents) (Dirac and Snow being one of the last unique ones in our current time, that made some news their where some other Wavelett based ones XWD,loco but Dirac and Snow are the most known Dirac even made it to a SMPTE standard, Iterate being the craziest different one (fractal) that still no one yet reverse engineered) ;)
wanezhiling
31st January 2012, 02:00
You could ask developer and he will reply what nev has said.:)
PS: VP5 could DXVA 4096 x 4096 smoothly if clips are light.
No, VP5 can't decode 4K x 2K by CoreAVC DXVA. CoreAVC CUDA is ok. you can see that screenshots. In fact besides CoreAVC CUDA, VP5 can docode 4K x 2K by PotPlayer self DXVA(note: old version,at least before 2011.11), TMT5, MainConcept(Broadcast) AVC/H.264 deocoder.
Failure lists: PowerDVD11 DXVA, ffdshow DXVA, MPC-HC DXVA, CoreAVC DXVA, MS DTV-DVD, LAV CUVID
Add a new member: LAV DXVA2(copy-back), it showed active when playing duckstakeoff 2160p though very very slow.:(
nevcairiel
31st January 2012, 08:15
LAV DXVA2(copy-back), it showed active when playing duckstakeoff 2160p though very very slow.:(
The problem here is that the only VP5 card is the 520, which is just a very slow card. It would probably do OK with pure DXVA, but the copy-back is probably too much for it. ;)
If NVIDIA releases their new generation of GPUs, i'll get one and see what can be done. ;)
CruNcher
31st January 2012, 08:46
http://www.gokuai.com/f/B05Yj637KS6yyu75
that's definitely not spark,whyever MediaInfo detects it as such seems to be a bug
x264 - core 79 r1360 528c100 - H.264/MPEG-4 AVC codec - Copyleft 2003-2009 - http://www.videolan.org/x264.html - options: cabac=1 ref=1 deblock=1:0:0 analyse=0x3:0x133 me=hex subme=6 psy=1 psy_rd=1.0:0.0 mixed_ref=0 me_range=16 chroma_me=1 trellis=1 8x8dct=1 cqm=0 deadzone=21,11 chroma_qp_offset=-2 threads=12 nr=0 decimate=1 mbaff=0 constrained_intra=0 bframes=3 b_pyramid=0 b_adapt=1 b_bias=0 direct=3 wpredb=1 wpredp=2 keyint=250 keyint_min=25 scenecut=40 rc_lookahead=40 rc=2pass mbtree=1 bitrate=6900 ratetol=1.0 qcomp=0.60 qpmin=10 qpmax=51 qpstep=4 cplxblur=20.0 qblur=0.5 ip_ratio=1.40 aq=1:1.00
Strange even with the Metadata Stripped it's still detected as Spark
NikosD
1st February 2012, 09:35
Metadata and the 32 first video frames of the file are saying this is Sorenson (CodecID = 2).
H.264 is after the 32 first video frames.
The file is buggy.
MediaInfo currently displays information about the first video frames.
Case solved.
egur
16th February 2012, 14:56
@NikosD
ffdshow 4322 and LAV 0.47 offer significant performance improvements with regard to QS.
The 2622 driver also raises performance.
Maybe you should update the scores?
NikosD
16th February 2012, 20:25
I was going to update the information of benchmarks, but I'll wait for the release of new LAV 0.47 (with faster QS than 0.46 and maybe the new DXVA2 "native" mode)
I want to test every mode on every system possible.
(DXVA2 copy-back, CUDA, QS, DXVA2 "native" etc)
Also I would like to test CoreAVC 3.1, which should be already available.(Both CUDA and DXVA "native")
egur
20th February 2012, 09:14
Very well. Nev and I got some progress fine tuning our code so LAV 0.47 (QS 0.28) shows a great improvement (http://docs.google.com/spreadsheet/ccc?key=0Ajo8vvjNtaZ5dC1abjBSeVlmcnZXSjYwampfamk3ZWc#gid=0).
You should update to driver v2622 for better results. Also post your RAM speed. This has a significant impact on results.
egur
21st February 2012, 08:29
@NikosD
The ffdshow QS scores are very old and don't represent current performance...
You can use LAV 0.47 for QS benchmarks as well, it's slightly more optimized than ffdshow.
NikosD
23rd February 2012, 07:23
Updated everything.
Some odd results appeared. :rolleyes:
CruNcher
23rd February 2012, 10:23
Sorry but i cant agree rating the Quality especially of a DXVA Decoder only based on it's speed is wrong DXVA implementations are different and some ISVs fix Bitstream issues that other doesn't which depending on how they do it adds overhead. For example Microsofts DXVA or Cyberlinks wont ever playback some Bitstreams that are easily done by CoreAVC or Arcsoft (without needing to fallback to Software), especially if you have made streams with older x264 versions which had bugs you are screwed with most DXVA decoder ;)
So this "this is the fastest" is only 1 Side of the Medal but shouldn't be the only factor taken into account, playback stability imho weights more then speed, though if you really want almost 100% stability you have to use a Software Decoder they cope with mostly any bitstream that a DXVA decoder would fail on also direct API implementations by the Vendor do like nvcuvid,intels quicksync or amds ovd (though even they have issues in their underlaying implementations with several bitstreams, but they are most of the times reacting fast to fix them).
BetaBoy
23rd February 2012, 13:51
CruNcher... Agreed. Also note it looks like we identified a slow down for DXVA which would explain the results here. We are trying to get it added into CoreAVC 3.1.
NikosD
23rd February 2012, 16:19
Quality of H.264 decoding should be the same for every bugless decoder.
As a matter of fact there is no such thing as "quality" for H.264.
The bugless DXVA implementation of H.264 decoding, should be handled with the same perception as H.264 decoding.
So, decoding speed for DXVA H.264 acceleration is the number one factor.
Of course it's not the only one, because bugs always exist.
nevcairiel
23rd February 2012, 18:10
Bugs can also exist in the H.264 stream itself, which technically makes those streams invalid (not standards conform), but trying to gracefully skip/recover those is whats important sometimes. Of course, sometimes you're just at the mercy of the GPU driver or the hardware.
mark0077
23rd February 2012, 19:52
Hi guys, I'm putting together some benchmarks on my own machine today using LAV / ffdshow decoders, and am finding very interesting results.
Although QS is extremely fast, and I'm getting very high decoding rates with it in both LAV and ffdshow, theres still 2 things that I'm confused about.
1) Some particular files show a large difference in performance between QS in LAV and ffdshow, sometimes with ffdshow being much faster, sometimes LAV being slightly faster. Is it worth me posting the details of these and my system to the likes of yourself nev?
2) Although performance of QS is tremendous and I'm amazed by it, especially compared to my gpu decoding, it uses alot more cpu than other software decoders (about 1.5x - 2x) on my core i7 920. I thought decoding speed and cpu usage might have been more closely tied, but instead I see the opposite of what I expected, higher performance from QS using more cpu.
I was hoping to use FPS output from graphedit to decide what the best decoder was for my system for various types of content, but because I'm using avisynth scripts that I have tuned over time to reach about 90% cpu usage max, the increased cpu usage of QS over other decoders pushes the cpu that extra bit higher causing dropped frames... So its improved performance seems to actually have a negative effect in this case. Is increase cpu usage by using QS normal or should it be as low or lower than other sofware decoders?
nevcairiel
23rd February 2012, 20:02
QS is a GPU decoder, its not a software decoder. I noticed you say i7 920 - that CPU does not contain a GPU thats QuickSync compatible. You must be confusing some things.
For the record, on my i7 2600k, i get around 500-550 FPS at 100% CPU usage in software decoding, and around 400 fps with QuickSync at ~30% CPU usage (testing done with the Twin Peaks sample from this thread).
During normal playback, QuickSync usually hovers at 1-2%, possibly up to 3-4% for 60fps material.
mark0077
23rd February 2012, 20:14
doh! I feel stupid. I didn't know about it being it being a GPU decoder :S aaah,
I just tried it on my older core i7 920 and get massive performance figures with QS, much higher than other decoders, even the other dxva ones in LAV on some particular files... I don't know how its supposed to behave on my core i7 but its really fast which is attractive. I guess its still confusing as to why it uses more cpu than other decoding options CUVID / DXVA / other pure software decoders. Maybe I'll ask the developer as to how it should behave on older cpus like mine.
nevcairiel
23rd February 2012, 21:22
The QS decoder shouldn't even be used when there is no GPU support, it would fallback to LAVs software decoder.
mark0077
23rd February 2012, 22:57
Ah right. When I use the QS decoder in ffdshow at least, I can see IntelQuicksyncDecoder.dll instances in Process Explorer mpc-hc Threads list. I'll try to clarify with the author what the decoder does on cpus like mine where theres no internal gpu. Its definitely being used by ffdshow and giving massive FPS gains. If I get any interesting feedback I'll post it here.
EDIT: Using LAV Video set to QuickSync, shows the quicksync decoder dll in process explorer also FWIW
fairchild
3rd March 2012, 21:35
Is there anyway someone can upload twinpeaks.1080.wmv to a working filesharing site like mediafire, the rapidfire link is slow as molasses. I tried googling and came up short. I'd like to run some vc-1 1080p tests and can't seem to find any good 10s clips to use.
nevcairiel
3rd March 2012, 21:53
Is there anyway someone can upload twinpeaks.1080.wmv to a working filesharing site like mediafire, the rapidfire link is slow as molasses. I tried googling and came up short. I'd like to run some vc-1 1080p tests and can't seem to find any good 10s clips to use.
http://files.1f0.de/samples/twinpeaks.1080.wmv
fairchild
3rd March 2012, 22:17
Sweet thx Nev.
wanezhiling
9th March 2012, 09:24
I got such a broken image (http://i.imgur.com/ODlZ3.jpg) with CyberLink DXVA or HAM on my GTS450(295.73whql) while no problem with my 4650m(12.3).
Did anyone reproduce the issue with Samsung clip?:)
NikosD
6th May 2012, 17:29
Added WMV3 results for DXVA QuickSync
vanden
14th May 2012, 16:08
Hello I have a HD 6850 and after some test my results are close to your original 5750 !
Avatar 60fps (MS DS) :
Renderer: Enhanced Video Renderer (DirectShow)
Decoder: Microsoft DTV-DVD Video Decoder
Decoder Device: ModeH264_VLD_NoFGT
Processor Device: -
Time: 00:38.471
Average FPS: 51,363
Min/Max FPS: 48 / 54
CPU Usage (%): Avg: 02 Min: 00 Max: 03
GPU Usage (%): Avg: 00 Min: 00 Max: 02
Generally UVD launch with gpu 300mhz / memory 1100mhz ...
In 3d I use AMD overdirve (900/1150).
If I modify my bios (with RBE) is there a chance that it works as well as your (6)750 ?
NikosD
14th May 2012, 17:12
It's been months since my last attempt to benchmark and test my UVD.
I remember that I had tried the procedure you propose, to manually change the UVD frequency with RBE, with no success.
Unfortunately my latest attempts to put UVD in 3D mode with my (6)750 have all failed, too.
I think the problem is latest Catalysts (12.4) which have probably closed that "door" to UVD overclocking.
I'll do further tests and report back.
pokazene_maslo
14th May 2012, 19:08
Hi.
How do I enable 4K h264 DXVA on my ATI HD6970? Whatever I set in MPC-HC or Pot player the DXVA simply doesn't kick in.
Catalyst 12.4, win7 x64
vanden
14th May 2012, 21:16
I'll do further tests and report back.
Thanks
Hi.
How do I enable 4K h264 DXVA on my ATI HD6970? Whatever I set in MPC-HC or Pot player the DXVA simply doesn't kick in.
Catalyst 12.4, win7 x64
No 2160p for AMD card ...
It's exaggerate, AMD announce support of 4k (April 2010/Catalyst 10.4) !!
We are in May 2012 !!
vanden
14th May 2012, 22:42
I have found a software (iTurbo On HIS site). It's possible to disable 2D clock and always force 3D clock (900/1150 or more).
I have little improvement :
Average FPS: 55,736
Min/Max FPS: 51 / 58
But it's not as significant in that your tests
pokazene_maslo
15th May 2012, 11:05
No 2160p for AMD card ...
It's exaggerate, AMD announce support of 4k (April 2010/Catalyst 10.4) !!
We are in May 2012 !!
why is then dxva checker displaying such support (3840x2160)?
nevcairiel
15th May 2012, 11:09
why is then dxva checker displaying such support (3840x2160)?
Because no tool is perfect? :p
In theory at least the 7xxx series cards support it, but when you try to decode a 4K image, only half is decoded and the other half is green. :D
wanezhiling
15th May 2012, 11:47
4650M UVD2.0 (http://i1131.photobucket.com/albums/m545/wanezhiling/2012-05-15_183345.png)
Dangerous, when it went to 00:00:04, system hung..:p
vanden
15th May 2012, 14:58
I have tested catalyst 12.1 with 3D clock forced : same result as catalyst 12.4 ...
NikosD
15th May 2012, 20:17
I have tested catalyst 12.1 with 3D clock forced : same result as catalyst 12.4 ...
You have a 6850 card which uses UVD3.
UVD3 is different than UVD2.2 used by my 5750/6750 card, because UVD3 uses dynamic change of frequency.
Which means that even if you change the UVD3 frequency manually, ATI can put UVD3 back to its default frequency, even if you see the GPU going to 3D clocks.
Internally UVD3 will work to a different (default) frequency.
They can control frequency completely.
Moreover, even in UVD2.2 which doesn't have dynamic frequency, in the original BIOS of 5750 when I put manually UVD2.2 in 3D clocks, I couldn't really put UVD2.2 in 3D mode.
I only saw GPU going to 3D clocks, but UVD2.2 worked internally in default clocks.
It was only after changing my BIOS from 5750 to 6750, that I saw - without altering any setting of BIOS - that UVD2.2 was going to 3D clocks by itself, automatically.
Drivers must had gone crazy with that hybrid card and allowed UVD2.2 to work in 3D clocks, statically.
If UVD2.2 was going to 3D mode, the whole card was remaining to that frequency, until restart.
Unfortunately that is no longer available.
I talked too much here :) and to ATI engineers about the subject and they managed to close that "hole".
With Catalysts 12.4 - I don't know with earlier versions between 12.1 and 12.3, I haven't tried them - it's impossible to put UVD2.2 to 3D clocks.
They have simply closed that "hole" :mad:
The only thing remaining is to try to change 6750 BIOS manually and put UVD2.2 in 3D clocks, but latest version of RBE can't recognise 6750 BIOS completely and I don't have the time or mood to do that.
If I do it, I will report here.
UPDATE 14/06
It seems that DXVA Checker 2.8.2 can put UVD2.2 of my fake (6)750 to 3D clocks and the benchmarks show that it can put it in real!
So we are back again with UVD2.2 working in real 710/1160 MHz!
pokazene_maslo
3rd June 2012, 13:41
Because no tool is perfect? :p
In theory at least the 7xxx series cards support it, but when you try to decode a 4K image, only half is decoded and the other half is green. :D
Maybe it requires stream with 2 slices? Have you tested 4K with multiple slices? On my HD6970 8 slices doesn't help.
wanezhiling
6th June 2012, 05:37
http://i.imgur.com/vvvT0.jpg
http://xhmikosr.1f0.de/samples/2160p/DucksTakeOff/DucksTakeOff_2160p50.x264.CRF24.mkv
I5 3450
HD2500(650-1100)
PotPlayer, evr cp
VP5 has only 30fps.
NikosD
6th June 2012, 12:14
Very impressive QuickSync performance for Ivy.
Even HD2500 can put Core i5 3450 to a very low 8% at 2GHz - lowest frequency I think for Core i5 3450 which has a base frequency of 3.1GHz and 3.5GHz max.
But GPU clock is at max 1100MHz and GPU Load is 88%.
I don't know if GPU load is actually VPU load for this clip but it seems that HD2500 goes to the limit to play this clip.
You could check for sure the VPU load with GPA tool from Intel.
I think that a good benchmark comparison with DXVA Checker between HD2500 and HD4000 would be useful to see the maximum performance and the potential of the cards regarding 4K H.264 clips and the differences between them.
NikosD
6th August 2012, 16:33
Just tried Win 8 Pro x64 with built-in drivers for ATI 5750.
Unfortunately latest Catalysts for Win 8 (release preview) didn't install - so I don't have CCC, OpenGL , OpenCL etc.
But DXVA acceleration works good with built-in Win 8 drivers and DXVA Checker v2.9.1 has exactly the same decoders.
The performance is more or less the same as with Win 7 and WMP12 and PotPlayer work the same.
I'm writing this post with IE10 (not metro version)
UPDATE:
I managed to install Win 8 release preview drivers, which are older than Microsoft drivers.
I have now OpenGL, OpenCL, CCC etc.
It's interesting that WMP12 and Video application of Metro UI with both drivers (Catalyst, Microsoft) try to play 4K H.264 clips in DXVA mode, not in software (there is no software fallback).
I get a black screen, but the sound is OK.
I'm still waiting ATI to allow 4K playback with future drivers.
NikosD
7th August 2012, 13:56
More info about Win 8
Win 7 vs Win 8
Win 8 boots faster, has faster shutdown, less footprint (both in memory and disk) and it's faster overall.
IE10 is very fast
It's a modern, more secure OS but has Metro too :confused:
Start button and gadgets ( I use Network/ CPU/ GPU gadgets all the time) are the big losses - for me.
All applications work good so far - and Win 7 drivers too.
About Win 8 DXVA
First of all, with WDDM v1.2 all hardware acceleration - 2D, 3D, video (DXVA) - goes below the same API D3D11.1
So DXVA no longer uses D3D9 for Video Acceleration, it uses D3D11 like 2D and 3D acceleration.
AMD Catalysts didn't provide AMD MFT codecs (none), because Microsoft did it.
It seems that Microsoft - at last - has put VLD mode for VC-1/WMV3 in MFT codec and added a strange Media subtype to its MS MFT H.264 codec, named:
MediaType_Video 3f40f4f0-5622-4ff8-b6d8-a17a584bee5e
which probably is used for Mpeg4 ASP video format (DivX/ Xvid)
So, with an MFT player - like WMP12 or Video (new player of Metro) you can have HW acceleration for everything (MPEG-2, MPEG-4 ASP, H.264, VC-1/WMV3) with Microsoft's codecs -even with built-in drivers.
I did some preliminary benchmarking and it seems that MS DS/MFT codecs have the same performance as LAV 0.51.3 regarding DXVA.
In pure CPU mode, both codecs DS/ MFT H.264 have the same performance as Win 7 (no optimizations) which is far below LAV Video 0.51.3.
I tested again 4K H.264 clips and they are decoded in HW with both WMP12 and Video applications (no software fallback) and both drivers (Microsoft, Catalyst).
The result is a black screen with audio.
Catalyst drivers gave a small boost in Windows Index for 2D and 3D - from 6.9 to 7.2 - compared to newer built-in drivers.
NikosD
17th September 2012, 13:22
AFAIK the 2 most "difficult clips" of H.264 4K resolution are:
1) The 3840 x 2160 version of "Ducks Take Off" at 50 fps and 370Mbps bitrate posted above and
2) The 3840 x 2160 version of "Girls Generation Oh 4in1" at 120fps and 100Mbps at L5.1 ReF 16
I haven't found L5.2 yet.
nevcairiel
17th September 2012, 13:27
For the record, the fps of the file doesn't really matter for benchmarking.
100Mbps at 120fps is 0.83Mb per frame, while 370Mbps at 50fps is 7.4Mbps per frame. Clearly, the second is much more complicated to decode.
If you want the most complicated clip, get one with a high bitrate per frame. The FPS does not matter for benchmarking (it only matters for playback).
NikosD
17th September 2012, 13:49
I disagree with you because my experience with both benchmarking and playback on various VPU hardware says that high FPS does matter.
There are HW units and SW implementations that have problems in both benchmarking and playback with high FPS but without problems at all with high bandwidths.
For example my normal UVD2.2 can play in realtime and benchmark faster than VP4 various H.264 1080p clips with huge bandwidths >100Mbps that VP4 can't, but at the same time VP4 is faster playing lower bandwidth H.264 1080p clips with high FPS like 60 fps.
Normal - default clocked - UVD2.2 can not touch 60fps with H.264 1080p60 fps no matter what the bandwidth is - low or high.
The example of low performance of SW implementation with high FPS is QuickSync decoder - I think not because of Eric's excellent job but because of Intel's MSDK and driver implementation and HW restrictions of memory operations.
As you know if you try to benchmark or play a H.264 1080p60FPS clip - low or high bandwidth, it doesn't matter - on QuickSync hardware using vanilla DXVA, the results will be excellent.
Very low CPU utilization, very low Power and more than enough FPS.
The very same clip at 60fps with QuickSync decoder (FFDShow or LAV implementation) will give much more CPU utilization, high Power and much lower fps.
So, as a conclusion, "difficult" clips for playback and benchmarking with current HW and SW implementations are those that have high bandwidth, high FPS or both.
nevcairiel
17th September 2012, 15:08
The speed of the decoder itself is only influenced by the normalized bandwidth. But higher fps usually has lower (normalized) bandwidth, so in that sense it has an indirect influence - but it all comes back to bandwidth. Sadly bandwidth is usually expressed in a "per second" value, which includes the fps, instead of a normalized/averaged "per frame" value.
When it come to playback, you will see differences though.
Like you said, a high-fps and low-bandwidth clip has higher CPU usage in some implementations. Yes, thats to be expected, because more frames are copied per second from the GPU (like with QuickSync or DXVA2-CB)
But during benchmarking, you just get as many frames as the decoder can output, independent of the fps.
NikosD
17th September 2012, 15:30
Interesting reply.
But according to your scheme how could you explain the UVD2.2 vs VP4 situation ?
If you see the benchmarks or if you try to play at normal speed (x1) a H.264 1080p30 fps clip with huge bandwidth ~100Mbps, you will see that VP4 is slower than UVD2.2.
That means - according to your scheme - that UVD2.2 decoder is generally faster than VP4, because it can handle larger normalized bandwidth than VP4.
But that is not true.
Because if you benchmark or play H.264 1080p60 fps clip, independantly of bandwidth, you will see that VP4 is faster than UVD2.2.
Could you explain it according to your scheme ?
vivan
17th September 2012, 17:40
That means - according to your scheme - that UVD2.2 decoder is generally faster than VP4, because it can handle larger normalized bandwidth than VP4.No. I think he wasn't comparing different decoders.
Simple example: you can change video fps while remuxing - e.g. increase it by 2 times. Would it change benchmark result? No. However such video will have 2 times higher bitrate.
So when you're using for benchmarking 2 videos and getting some numbers, bitrate doesn't mean anything (since you can get any bitrate by just changing speed of the video).
However when playing video in real time - both bitrate and fps matters.
NikosD
17th September 2012, 20:49
I get your point, as I got it from nevcariel.
But still, how can you both explain benchmark results of different decoders like VP4 and UVD2.2 like the ones I mentioned before ?
Why UVD2.2 can output more decoded frames from VP4 at H.264 1080p huge bitrate-low fps clips, but lags at high fps-low bitrate H.264 1080p clips ?
Moreover, the real purpose of video benchmarking is to expose and see the potential of an existing decoder implementation of HW and SW during playback.
So for example if the speed of video decoding during benchmarking for a 1080p30 fps file is 65 fps, then it means that the implementation can almost sure play both 2 such files at the same time (with proper driver and player support) or it can play that file at 2x speed.
On the other hand, if during benchmarking a file that is encoded at 60 fps, gets 40fps then you can be sure that normal playback will be slower than realtime playback.
In my opinion, if you want to check the potential of an existing video decoder (SW & HW) and get the whole picture of its own and compared to the others , you have to check it with difficult clips, so you have to check both parameters - bandwidth and fps.
That's why I proposed two clips.
One with huge bandwidth and high fps and the other one with extreme high fps.
NikosD
17th January 2013, 23:11
A lot of new DS filters for Catalyst 13.1 - AMD MJPEG decoder, ATI Ticker, ATI Video Rotation filter, ATI Video Scaler filter.
I haven't seen them before.
Also, officially AMD pulls back 4K DXVA acceleration - only 1920x1080 for Catalyst 13.1
huhn
3rd April 2013, 03:10
intel 3770k @4 ghz turbo hd4000 stock clock
lavfilter 55.3:
1
dxvan 733 qs 266
2
dxvan 508 qs 163
3
dxvan 797 qs 284
4
dxvan 729 qs 276
5
dxvan 687 qs 269
6
dxvan 649 qs 265
7
dxvan 288 qs 234
8
dxvan 280 qs 213
9
dxvan 311 qs 226
10
dxvan 253 qs 207
4k
crowedrun2160p50 crf24
dxvan 165 qs-
duckstakeoff2160p50 crf24
dxvan 118 qs-
intotree2160p50 crf24
dxvan 146 qs-
oldtowncross2160p50 crf24
dxvan 137 qs-
parkjoy2160p50 crf24
dxvan 159 qs-
and for the lols
amiga long play dune 241 kbit crf 30 640x400p25
dxvan 3088 qs 2081 cpu 1902 30 % cpu
like i said in the quicksync thread i can't use quicksync for 4k sry about that.
in my eyes intel dxva is so damn fast that it doesn't matter anymore...
NikosD
3rd April 2013, 11:05
Excellent work!
Although we are comparing different versions of LAV (native & QS) between your results and mine, the difference between the two QuickSync generations is huge.
Of course you have overclocked your 3770K, but if you haven't change Base Clock I think the iGPU doesn't change its clock by simply changing the multiplier.
The comparison of QS v1.0 of Core i5 - 2400 (SNB) vs QS v2.0 of Core i7-3770K (IVB) is interesting because they work at almost same clock speed.
Max QS v1.0 clock is 1100MHz and Max QS v2.0 clock is 1150MHz.
Did you see which is the actual clock of your GPU during benchmarking ?
You can check it out with GPU-Z.
In huge bitrate clips the DXVA Native performance is more than doubled, the QS is slightly less.
In smaller bitrate clips, QS performance is close.
Native performance is close to double again.
4K performance is excellent too.
I have some 4K clips with large bitrates and high FPS (60fps & 120fps) but I don't think that QS 2.0 will have a problem in DXVA native mode.
huhn
3rd April 2013, 17:22
the new ivy bridge driver fixed my 4k problem.
duckstakeoff 2160
dxvan 100 qs 60
with the new driver my duckstakeoff benchmark was a lot slower don't know why.
with quicksync:
the mhz in gpu-z goes up to 850 mhz and then stays at 650 mhz so there is a lot of room left. and the gpu power is at max 1.8 watt
the gpu load from gpu-z shows at max 30%
with dxva:
the igpu goes up to 1150 and uses alot more power then quicksync 6.6 watt
if i can trust gpu-z then quicksync is not that fast but it needs "lot" less power.
NikosD
3rd April 2013, 21:03
Well, it is known that QS software has its limitations regarding performance, especially for low bitrate clips or high fps clips.
The most obvious bottleneck is copy- back procedure of rendered frames between CPU and GPU.
That's why you see low clocks, because of memory performance bottleneck. It doesn't use QuickSync hardware a lot and it's slower.
Also in real playback mode and benchmark mode it uses CPU a lot more than pure DXVA and it consumes more power than native DXVA.
Check the power consumption during playback of the whole chip (CPU, GPU, QS) and you will find out that QS consumes more power.
QS is useful mainly for movie compatibility reasons and the support of VC-1.
New driver seems that doesn't provide good performance results.
So it's not a good fix, if you fix something and break something else :)
Your numbers provided as benchmark results are average. Right ?
The benchmark figures of 4K sample of ducks are different between your posts.
The image of DXVA checker at the other thread of QS software says 83-90 fps on average but your first post here say a lot more with the old drivers and the overclocked CPU.
Which is right and under what benchmark conditions ?
huhn
3rd April 2013, 22:17
at the moment 100 is normal the frist screen was made on my hd6770 quicksync ist faster on his own intel gpu.
and the monitor was running in ycbcr so no need for the ycbcr - rgb coversation
NikosD
4th April 2013, 23:54
I've just realized that the 4K ducks 50fps clip has so large video bitrate that you need a really fast Hard Disk to keep up, especially in benchmark mode - where you can "play" it in double speed or more with DXVA native.
So, if your hard disk was busy during benchmarking or fragmented, maybe it was a bottleneck for maximum performance.
It would be really interesting if you could move the sample to an SSD - if you have - and run it again to see the results.
huhn
5th April 2013, 19:31
i do this later i got a ssd with over 500 mb read.
i may even downgrading to the old driver to double check but not sure this is a bit of work.
the file are at the moment on a nas which is pretty fast...
huhn
7th April 2013, 23:03
nothing changed the frist run is faster but i ignore the first run
http://s3.imgimg.de/uploads/nothingchangedae1782f0png.png
a slight offtopic question, but what windows player support 4k resolution source files together with QS/Dshow filter?
NikosD
9th April 2013, 21:06
Well I did manage to overclock Core i5-2400 GPU from 1100 MHz to 1450 MHz - 32% raise.
Also I managed to overclock memory from 1333MHz to 1373MHz with an increase to CPU from 3.1GHz to 3.2GHz.
After all these I ran dxva benchmarks of clips 2, 3 and 8 of my collection.
I used latest drivers and DXVA Checker 2.9.1 along with LAV 0.55.3 Native
Clips 2 and 3 didn't go up to expected performance, although of course they had an increase.
I analyzed the benchmarking procedure with Intel GPA 2013 R1 to find out that those 2 clips didn't fully utilize the MFX engine.
Also the combination of latest drivers, dxva checker and LAV video was slower than the one of the original benchmarks of my first page of benchmarks.
But clip 8 was a different story.
The performance went from 118fps to 159fps. Big improvement.
Intel GPA said that utilization of MFX engine was 99% for clip 8.
Now Eric with my overclocked GPU, I'm ready for 4K :)
@TEB
PotPlayer can play 4K in all modes (native, cb, CUDA, QS)
huhn
10th April 2013, 20:43
i'm back to the old driver with an clean windows at frist i was shock at the dxva performance 91 was the best i can get with the old driver even lower then the new one...
then i changed all media settings to application settings and it is now at 130 fps! even higher then before: http://s3.imgimg.de/uploads/duckstakeoff0f7bc83epng.png quicksync still crashes.
the first clip is even faster: about 875 http://s3.imgimg.de/uploads/dxva20c02e518png.png
it is not easy at all to test dxva performance...
sry about that but the nr. before are all to low...
NikosD
10th April 2013, 21:25
Interesting...
What are your settings to media settings in order to try them with my Core i5 ?
And if you install again the new driver with the fast media settings, maybe you will get faster results than the old one ?
huhn
11th April 2013, 16:45
the parts under color/image enhancement i set them all to application settings.
and i found that
Several optimizations to reduce the system power consumption have been added in this driver. These optimizations might result in minor performance drop in some games while yielding power savings. The user can choose to override these power optimizations by selecting the “Maximum Performance” power mode in Intel control panel or “High Performance” mode in the Windows power setting.
source: http://www.legitreviews.com/news/15354/
NikosD
12th April 2013, 19:43
I tried your settings and got a serious performance boost.
I have to ask Eric why Intel has default values that lower performance so much.
Of course, if you benchmark with DXVAChecker using EVR renderless, it doesn't matter what are your media settings.
It only matters when you are using EVR (as in real world playback)
Also, I managed to overclock GPU without doing anything special, just a BIOS setting and I pushed the clock to 2100MHz (from 1100 MHz max)
The result using LAV 0.56 and DXVAChecker 2.9.1+ fast media settings with clip 8 was a huge 220 fps.
Close to IvyBridge.
wanezhiling
25th May 2013, 16:28
By the way, WMP12 seems to be able to play back 4K files as well now. Not sure if it was able to do that before.
Its sw decoding, not hw decoding.
NikosD
25th May 2013, 16:36
Using MPC-HC 1.6.7.7114 x86 (@ EVR) + LAV Filters 0.57.0 x86 (@ DXVA2 native) + Windows 7 SP1 x64 + GIGABYTE GV-N520SL-1GI (http://www.gigabyte.com/products/product-page.aspx?pid=4006#ov) (@ 320.18 WHQL) unfortunately does not play back the 2160p30 "Roast Duck (http://forum.doom9.org/showthread.php?p=1591912#post1591912)" clip for example smoothly.
Playback is slow and stuttery.
Downloading the file right now and I'm gonna test it with my a little slower GT 610 card (535MHz memory vs 600MHz memory)
AFAIK, 370 Mbps is out of spec for H.264 High Profile.
H.264 High Profile can only go up to 300 Mpbs @ Level 5.1/5.2 according to official H.264 specification AFAIK.
True, is out of spec.
It's exactly the same with MPC-HC.
EVR CP performance is worse than EVR performance.
Maybe the Koreans of PotPlayer can do a miracle :eek:
PS:
By the way, WMP12 seems to be able to play back 4K files as well now. Not sure if it was able to do that before.
True
NikosD
25th May 2013, 16:37
Its sw decoding, not hw decoding.
No, it's hardware decoding.
Microsoft's codecs both DS and MFT can decode 4K with VP5.
Probably true for Ivy, too. (I don't have one)
NikosD
25th May 2013, 17:07
[QUOTE=jq963152;1630141
Using MPC-HC 1.6.7.7114 x86 (@ EVR) + LAV Filters 0.57.0 x86 (@ DXVA2 native) + Windows 7 SP1 x64 + GIGABYTE GV-N520SL-1GI (http://www.gigabyte.com/products/product-page.aspx?pid=4006#ov) (@ 320.18 WHQL) unfortunately does not play back the 2160p30 "Roast Duck (http://forum.doom9.org/showthread.php?p=1591912#post1591912)" clip for example smoothly.
Playback is slow and stuttery.
[/QUOTE]
I've already had that file called "Chimei-inn"
For reasons unknown to me, it is true that this 50Mbps file can put VP5 to 99% utilization, but I can play it almost flawlessly@30 fps with PotPlayer :p
If you benchmark it with DXVA Checher you will get an average of little more than 30fps like
26/30/33
If you play it with DXVA Checker you will also get a 26/28/30 and the same feeling I get with PotPlayer.
You can't possibly say that an average 28fps of a 30fps clip is slow and stuttered :confused:
NikosD
25th May 2013, 20:30
Well, I have about 40 4K clips - from 2048x1280 up to 4096x2304 - and from 15 fps up to 120 fps.
The great majority of them ~90% can be decoded from around 30fps to max ~42-45fps.
To check the power of VP5 try to google a version of "duck" file which is 3840x2160@30fps and of average bitrate of 243Mbps close to theoretical maximum of 300Mbps.
It's a 493MB file on HD.
If you benchmark it or play it with DXVA Checker you will get these figures - which for me are amazing:
25/28/30
The performance is not as fluid as it could, but it's more than good enough to watch the clip.
Most of the people here believe or have said that VP5 is almost useless for 4K decoding which is by FAR WRONG.
It is not created for low bitrate 4K clips - as almost everyone believes - but on the contrary it manages to play 243Mbps 3840x2160 clips up to 28 fps on average using EVR.
If you use it with EVR-CP, it is slow.
This is what I'm trying to say.
The way to benchmark DXVA performance of various codecs is described here at an old thread by me:
http://forum.doom9.org/showthread.php?t=156660
NikosD
25th May 2013, 21:50
As I 've written before GT520 = GT610 (GT610 is a rebranded GT520)
My card has 810MHz clock for GPU cores and 535 MHz = 1070 MHz DDR3 memory clock - little slower than yours.
For DXVA native playback/ benchmarking, we have actually the same card.
NikosD
25th May 2013, 22:40
I really can't follow you with your last post.
GT520 cards do exist with GF108 chip but are very rare variants.
GF108 is basically a GT440 card which has a VP4 video processor.
But, what all these have to do with our cards ?
It's a completely irrelevant statement.
We can both play in HW 4K H.264 clips, so we have both a GF119 GPU with a VP5 processor with the same performance.
What are you trying to prove and what are you looking for ?
Download a lot of different 4K clips, do some benchmarking with DXVA checker and check out your results.
Then, compare your own results with mine and come back posting your findings.
NikosD
26th May 2013, 08:09
Is it normal that the "GPU Usage:" indicator in DXVA Checker 2.9.1 is greyed out and doesn't show anything while benchmarking and playing back in DXVA Checker 2.9.1?
Same goes here.
ATI/AMD cards usually don't have such issues.
Probably for DXVA Checker developer is easier to correct bugs with ATI/AMD cards than Nvidia.
Or maybe it's an Nvidia driver's bug.
Also, is it normal that the "GPU Acceleration" menu entry is greyed out and not selectable for "[DS] LAV Video Decoder [DXVA2] [AVC1 3840x2160]" in DXVA Checker 2.9.1?
I think that is something that has to do with internal structure of LAV Video decoder.
Because MS DS/MFT codecs and others have no such issues.
But it doesn't stop you or bother you for anything regarding playing/ benchmarking with DXVA Checker.
Both of the above are not a problem for playing/ benchmarking clips with DXVA Checker.
My interest now is if you or someone else could check out VP5 performance using EVR-CP and VP5 performance using EVR renderer with a Kepler card (from GT640 and above - any GTX 600 series)
It would be interesting to see if there are any benefits for VP5 performance by using faster VRAM, wider memory bus and more GPU cores embedded of course in a different GPU architecture (Kepler)
wanezhiling
26th May 2013, 08:24
EVR-CP costs more performance than vanilla EVR on all graphics cards, its just normal.
The more modern and the more powerful your GPU is, the less increase (vanilla EVR --> EVR-CP).
NikosD
26th May 2013, 08:38
In general this is true.
I'm interested in performance numbers - a quantitative approach.
NikosD
26th May 2013, 13:06
Playback of 4K files with VP5 seems like playback of 1080p60 fps files with ATI cards.
It's doable but not perfect.
Use PotPlayer with internal DXVA filters and use EVR renderer - not the default EVR CP.
I have a 90% success of all my 4K files and 100% of all youtube downloaded 4K files with that little card.
I can even play 4K clips with <65% VPU utilization.
It's the best HTPC card I've ever had, fanless with a TDP of 29W.
The conversation ends here by me because we have same numbers but different impressions of those same numbers :D
NikosD
11th June 2013, 15:01
EVR-CP costs more performance than vanilla EVR on all graphics cards, its just normal.
The more modern and the more powerful your GPU is, the less increase (vanilla EVR --> EVR-CP).
The above statement is false.
At last I got a little time to check it in more detail and found out these:
1) VP4, UVD, UVD2.2 have no performance hit at all using EVR-CP instead of EVR with PotPlayer.
As a matter of fact using EVR-CP with UVD2.2 (Radeon 5750) had a better sense of fluid playback than EVR vanilla with 1080p60 clips
2) Only VP5 has a TREMENDOUS performance hit using the default EVR-CP renderer instead of EVR.
I only made the observation, I don't know the reason.
So, I have to repeat myself once more.
Use only EVR renderer with PotPlayer and VP5 for demanding playback situations (like 4K playback).
DO NOT USE the default EVR-CP renderer of PotPlayer with VP5.
I didn't check performance of MPC-HC with VP5 in any situation (EVR/ EVR-CP)
Maybe it's slower than PotPlayer with EVR and VP5, causing the stuttering mentioned above.
NikosD
19th June 2013, 10:06
A performance preview of Snapdragon 800 at Anandtech:
Along with the x2 performance of Adreno 330 vs Adreno 320 we have the first mobile SoC capable of 4K H.264 Hardware Encoding / Decoding (playback).
It is capable of Hardware encoding/decoding of H.264 3840x2160@25fps with 120Mbps bitrate available to smartphones and tablets :eek:
4K samples here 3840x2160@25fps-120Mbpshttp://images.anandtech.com/reviews/gadgets/qualcomm/MDP8974/VID_20130618_161251.renametomp4.zip (]http://anandtech.com/show/7082/snapdragon-800-msm8974-performance-preview-qualcomm-mobile-development-tablet[/URL)
and YouTube here 3840x2160@25fps-56Mbps[URL=]http://youtu.be/H2eoSEPIQPQ
I'm sure that Ivy and Haswell can play both of them easily, but I have to say that my tiny VP5 can play both of them too - with an almost zero CPU utilization of my poor Core2Duo ;)
wanezhiling
8th July 2013, 03:57
http://bluesky23.yu-nagi.com/en/DXVAChecker.html
DXVA Checker 3.0.0
NikosD
8th July 2013, 18:24
A lot of changes...Really nice app.
NikosD
10th August 2013, 08:34
For those interested, I did some tests regarding multiple streams decoding on various video hardware and various video formats.
The test platform was:
Win 8 x64
MPC-HC v1.6.8 (with internal codecs using DXVA native)
Video Processors (with latest drivers and default clocks)
UVD2.2 (Radeon 5750)
VP4 (Geforce 440GT)
VP5 (Geforce 610GT) and
QuickSync 1 (HD 2000 - SandyBridge)
The video formats were all 1080p (full HD 1920x1080 progressive):
WMV HD (a 1080p24fps@10Mbps clip - the toughest WMV3 clip I have)
VC-1 (a 1080p24fps@18Mbps clip - a medium to hard VC-1 clip)
MPEG-2 (a 1080p24fps@45Mbps clip - a hard MPEG-2 (.ts) clip)
H.264 (a 1080p24fps@11Mbps clip - a Main@L4.1 "normal" clip)
The results are below - very interesting in my opinion.
WMV3
UVD2.2 - 2 streams
VP4 - 3 streams ~89% VPU utilization
VP5 - 4 streams ~89~ VPU utilization
QuickSync - Not supported
VC-1
UVD2.2 - 2 streams
VP4 - 3 streams ~95% VPU utilization
VP5 - 4 streams ~95~ VPU utilization
QuickSync - Not supported
MPEG-2
UVD2.2 - Not supported
VP4 - 3 streams ~80% VPU utilization
VP5 - 5 streams ~92~ VPU utilization
QuickSync - >9 streams
H.264
UVD2.2 - 2 streams
VP4 - 3 streams ~88% VPU utilization
VP5 - 4 streams ~99~ VPU utilization
QuickSync - >9 streams
Comments:
1)
UVD2.2 supports only 2 streams for all formats - odd behavior (driver?).
VP4, VP5 are the only video processors supporting all of 4 formats
I stopped QuickSync multiple streams at 9, not because of frame dropping or performance reasons, but because of limited space in my one and only screen (21.5" - 1920x1080). I think that maybe 10 or 11 streams were doable.
2) It's a pity that so fast hardware like QuickSync supports only 2 formats (MPEG-2, H.264) in DXVA native mode even AFTER 3 GENERATIONS/ VERSIONS of QuickSync.
UVD3.0 supports 5 formats (MPEG-2, MPEG-4 ASP, H.264, VC-1, WMV3) and VP4, VP5 support 4 and a half - because MPEG-4 ASP is supported only in CUVID mode.
3)When trying to measure QuickSync VPU utilisation with Intel GPA monitor, I had major frame dropping with the window open in front. Only in minimized mode it worked well - strange.
From what I've seen, the major bottleneck was EU utilisation and not MFX utilisation.
Faster QuickSync versions with more EUs and faster clocks can go up to a lot more streams I believe.
nevcairiel
10th August 2013, 08:39
VC1 works on Intel if you use the QuickSync decoder from egur in LAV Video or ffdshow. You could also implement it with dxva native, just no one has done it, its not a hw limitation.
MPEG4 ASP works also in DXVA on Nvidia but not sure which decoder implements dxva mpeg4 anyway.
NikosD
10th August 2013, 08:46
I know that VC-1 works with QuickSync decoder of Eric, but QuickSync decoder is not DXVA native.
It's more like DXVA copy-back.
I think I made it clear in my description that I used DXVA native only.
If you know any DXVA decoder that works with Nvidia hardware, tell me and I'll try it.
I have found none.
Until then, MPEG-4 ASP is not supported by Nvidia hardware in DXVA mode.
wanezhiling
10th August 2013, 08:52
You could also implement it with dxva native, just no one has done it, its not a hw limitation.
Is it a technical problem? I really hope LAV could be the first one to implement Intel VC-1 DXVA.:)
NikosD
10th August 2013, 09:08
You could also implement it with dxva native, just no one has done it, its not a hw limitation.
Regarding to your answer in Intel Media SDK - Quick Sync Video thread on February 2013 which is here
@rtabrah
Eric Sardella once wrote a whitepaper how to properly use the (proprietary) ClearVideo H.264 interface back on the older Intel GPUs, which elaborated on the specific differences required to use it. Luckily SNB and IVB now implement the standard H.264 DXVA decoder, and its no longer required.
http://software.intel.com/en-us/articles/using-h264avc-directx-video-acceleration-with-the-intel-g45gm45-express-chipsets/
However, if the same kind of information present in this document could also be provided for VC-1, i would be happy to implement it for full Intel VC-1 DXVA2 support, which would be contributed back to ffmpeg for everyone to use.
I don't need a fancy document, i just need to know what to do. :)
there is no VC-1 DXVA2 documentation from Intel, but if there was, you would be very happy to implement it - according to your words ;)
Is it true ?
Is there documentation from Intel side to implement DXVA VC-1 and are you going to do it ?
nevcairiel
10th August 2013, 17:31
There is only documentation missing how to do it, but that doesn't mean its not possible. Could also try to reverse engineer from the Media SDK, but thats annoying work.
The Hardware supports it, its only software thats missing.
Even then, you don't need to manually implement it, you can use the Media SDK and use it as a native DXVA decoder. Someone just needs to do it.
I believe some of the commercial decoders implement it (most likely using the MSDK in exactly this way), Cyberlink or ArcSoft, i forgot which one, maybe both? Its a Blu-ray format, so they usually try to support it on new hardware.
wanezhiling
25th November 2013, 09:58
http://wccftech.com/amd-kaveri-apu-steamrollerb-core-features-20-cpu-30-gpu-performance-uplift-richland-platform-details-unveiled/
http://cdn2.wccftech.com/wp-content/uploads/2013/11/AMD-Kaveri-APU-Platform-Details-635x476.png
UVD4.2 :eek:
So current GCN GPU(HD7xxx and R7/R9)'s UVD is 4.0 not 3.0?
NikosD
25th November 2013, 11:18
It looks like a typo.
I have seen it in all other places as UVD3.2.
wanezhiling
25th November 2013, 14:51
http://www.amd.com/US/PRODUCTS/DESKTOP/PROCESSORS/A-SERIES/Pages/a-series-apu.aspx#3
UVD3.2 was already introduced in Trinity APU.:p
nevcairiel
25th November 2013, 15:46
Since there isn't even a "NEW" in front of the UVD thing, it may still be a typo.
littleD
25th November 2013, 19:48
Why not 4.2. It's annouced that AMD R9 2x0 cards have improved UVD. I cant find that anybody would number it though. If we take in consideriation tighter componet placement on HSA chip, new VCE, then UVD on kaveri (2.0?) might get even higher number too.
Funny fact. I always considered AMD HW decode chips inferior to nvidia, especially in the times of lack of 5.1 level support. But now it's only AMD whos chips can hardware decode 10-bit h264 videos. Check this out all anime maniacs http://www.phoronix.com/scan.php?page=news_item&px=MTQ5NjU
nevcairiel
25th November 2013, 20:39
But now it's only AMD whos chips can hardware decode 10-bit h264 videos. Check this out all anime maniacs http://www.phoronix.com/scan.php?page=news_item&px=MTQ5NjU
There is no consumer hardware today that can properly decode 10-bit H.264.
All they did was hack the existing 8-bit decoder to simply decode 10-bit material. You can get a picture out of it that looks mostly correct, but it will exhibit occasional visible corruption, and it will never be 100% correct. There are other precedences for this happening on other hardware, so its nothing new.
Wrong decoding is worse then no decoding, IMHO.
NikosD
1st December 2013, 00:06
http://www.amd.com/US/PRODUCTS/DESKTOP/PROCESSORS/A-SERIES/Pages/a-series-apu.aspx#3
UVD3.2 was already introduced in Trinity APU.:p
I was definitely wrong, I thought we were talking about Richland.
Kaveri will have indeed a new decoding engine -UVD 4.2
The interesting thing is how much really "new" will be.
It won't have HEVC decoding in hardware, but let's hope it will have 4K H.264 decoding, at last.
I'm not sure about UVD 4.0, if has ever used in a GPU.
Wikipedia has stopped to UVD3 and I haven't found any other info.
wanezhiling
1st December 2013, 05:39
VLIW5/VLIW4 architecture
HD6000/1st Llano UVD3.0
2nd Trinity/Richland UVD3.2
GCN architecture
HD7000 UVD4.0?
3rd Kaveri UVD4.2
:p
NikosD
1st December 2013, 21:02
Another industry first for Qualcomm.
After the 4K H.264 hardware accelerated video in Snapdragon 800 - for smartphones and tablets - Qualcomm is the first company to bring HW accelerated H.265 content in new Snapdragon 805, available in 2014.
Details here:
http://anandtech.com/show/7537/qualcomms-snapdragon-805-25ghz-128bit-memory-interface-d3d11class-graphics-more
NikosD
18th February 2014, 15:45
Maxwell is out with a promise of Nvidia of HW assisted H.265 decoding (not full HW acceleration) in the near future.
It has a faster decoder VP6 (?) probably capable of 4K@60 fps.
Radeon 7750 once again fails to HW accelerate 4K clips and uses software decoding.
Read here:
http://www.anandtech.com/show/7764/the-nvidia-geforce-gtx-750-ti-and-gtx-750-review-maxwell/9
NikosD
25th February 2014, 22:05
I've installed latest iGPU drivers for my Pentium G3420 yesterday v.3412 and find out today two extraordinary things that I would never expect.
1) QuickSync transcoding - yes decoding and encoding - is available for the first time in a Pentium CPU (Haswell-DT core) !!
I tried it and it really works!
QuickSync encoding with Pentium G3420!
2) HEVC_VLD_Main appeared while enumerating DXVA device decoders !!
qtwebkit
1st March 2014, 21:35
Hi, I've just made a simple benchmark for recent popular decoders - LAV Video Decoder, MPC Video Decoder, Microsoft DTV-DVD Video Decoder and ArcSoft Video Decoder. I failed to test CyberLink Video Decoder in DXVA Checker, as it was high-lighted in DSF/MFT Viewer list and would not be checked when choosing a sample. Any one knows the reason?
Here's the result and screenshots (https://onedrive.live.com/redir?resid=C134CFE2F5253FB2%21214).
It seems MPC Video Decoder lead in most 720P Samples, while ArcSoft Video Decoder performance better in 1080P Samples.Sadlly I found LAV take more CPU usage than BE's standalone filter and ArcSoft Video Decoder in DXVA mode. As I was told the Microsoft Decoder is much more effecient in Windows 8, It's werid that it performanced the worst of all in my test.
I wonder if you have made such tests recently, since my result is far beyond my expectation, I wish you and Nev would take a look at the performance and give me some help:).
NikosD
1st March 2014, 22:08
Welcome to the club :)
It's been a really long long time since my last overall codecs test -with all available codecs tested.
I've settled down with LAV filters and feel lazy to test anything else!
I failed to test CyberLink Video Decoder in DXVA Checker, as it was high-lighted in DSF/MFT Viewer list and would not be checked when choosing a sample. Any one knows the reason?
IIRC there were specific CyberLink video decoder versions that could be used outside the app.
Maybe you could try to check the properties of that decoder and find out anything there that could possibly help.
Also, just in case, try the today's release of DXVA Checker v3.0.1
Sadlly I found LAV take more CPU usage than BE's standalone filter and ArcSoft Video Decoder in DXVA mode.
Only Nev could answer this.
As I was told the Microsoft Decoder is much more effecient in Windows 8, It's werid that it performanced the worst of all in my test.
I think Win 8.1 has even faster decoders, but I'm not sure if they are faster than LAV filters.
This is an easy test for me, since I have Win 8.1 Pro x64, so I'll test it tomorrow and see the performance difference.
For me the good thing is that free decoders are very close and sometimes faster than the best commercial ones.
This is a huge progress.
nevcairiel
1st March 2014, 23:19
LAV is optimized for smooth and fluid playback, not benchmarking. I could probably make it slightly faster in benchmarking, but past experience has shown that aggressive optimization to get 1% more performance in a benchmark can result in playback being no longer smooth on some systems - especially low powered ones.
qtwebkit
2nd March 2014, 08:22
Due to hardware and other limitations, I could not make an overall comparison as yours. Actually, I've been making MPC-HC as my default player for months, and also, thanks to the superior LAV Filters.
IIRC there were specific CyberLink video decoder versions that could be used outside the app.
Maybe you could try to check the properties of that decoder and find out anything there that could possibly help.
Also, just in case, try the today's release of DXVA Checker v3.0.1
Yes, I got the CyberLink standalone filters from wanezhiling's sharing. But CyberLink appears the same as 3.0.0 in 3.0.1. Maybe I could ask him for help. Thanks you all the same!;)
I'll test it tomorrow and see the performance difference.
Waiting for your test.:)
NikosD
2nd March 2014, 08:31
As I was told the Microsoft Decoder is much more effecient in Windows 8, It's werid that it performanced the worst of all in my test.
I wonder if you have made such tests recently, since my result is far beyond my expectation, I wish you and Nev would take a look at the performance and give me some help:).
I think Win 8.1 has even faster decoders, but I'm not sure if they are faster than LAV filters.
This is an easy test for me, since I have Win 8.1 Pro x64, so I'll test it tomorrow and see the performance difference.
I could probably make it slightly faster in benchmarking, but past experience has shown that aggressive optimization to get 1% more performance in a benchmark can result in playback being no longer smooth on some systems - especially low powered ones.
Well, I was wrong, has been a really long time since my last non LAV video tests.
MS DS and MS MFT decoders of Windows 8.1 are a lot faster than LAV video in both 1080p and 4K H.264 clips
To be more exact MS DS is the faster one, followed by MS MFT.
The performance advantage of MS DS is about 15% (10% to 20%) over LAV Video latest nightly !
P.S
@qtwebkit
I've just seen your screenshots and I have to say that you must never stay to one run, one test.
You must run at least three to five times each test and try to find out the average, skipping the first test, which almost always is not accurate.
qtwebkit
2nd March 2014, 08:47
LAV is optimized for smooth and fluid playback, not benchmarking.
Clear now:). I was missleaded by the "performance" and "speed". Thanks for all you have done to provide us the great freeware!
BTW, are there any specific optimization for Windows 8 in recent nightly codes?
qtwebkit
2nd March 2014, 09:08
MS DS and MS MFT decoders of Windows 8.1 are a lot faster than LAV video in both 1080p and 4K H.264 clips
To be more exact MS DS is the faster one, followed by MS MFT.
I've once got the similiar result, but not this time. In fact, I have tried MS DS and MS MFT for more than 5 times but merely with no difference.
It seems there's something wrong with my MS Decoder, just as CyberLink, high-lighted red in DSF/MFT Viewer. Maybe it's the problem of my Windows 8 LITE. So I would change to another environment and try again.
BTW, have you test the latest MPC Video Decoder(MPC-BE), which I would like to know your result either.:D
nevcairiel
2nd March 2014, 13:45
MS DS and MS MFT decoders of Windows 8.1 are a lot faster than LAV video in both 1080p and 4K H.264 clips
Are you talking about DXVA or Software?
I just tested LAV Video vs MS DirectShow in DXVA, and the performance is about the same on my Haswell, and on high-bitrate clips LAV was even slightly faster (ie. on high-bitrate Birds in 4K, 180 fps LAV at 3-4% CPU vs. 160 fps MS at 7-8% CPU)
All tests on Windows 8.1 on a HD4600 with DXVAChecker 3.0.1, LAV in DXVA2-Native mode of course.
qtwebkit
2nd March 2014, 14:18
@nevcairiel
I was puzzled by my test result, which 8bit+i420+720P consume more cpu than 8bit+i420+1080P. I know litte about it, could you explain it briefly? THX!
NikosD
2nd March 2014, 15:26
Are you talking about DXVA or Software?
@nev
@qtwebkit
Well,
in order to have comparable results I will tell you what I did:
I tested on my signature system various clips in DXVA native.
As a matter of fact, I don't know any way to put MS DS or MS MFT in DXVA copy-back mode.
I just tested LAV Video vs MS DirectShow in DXVA, and the performance is about the same on my Haswell, and on high-bitrate clips LAV was even slightly faster (ie. on high-bitrate Birds in 4K, 180 fps LAV at 3-4% CPU vs. 160 fps MS at 7-8% CPU)
All tests on Windows 8.1 on a HD4600 with DXVAChecker 3.0.1, LAV in DXVA2-Native mode of course.
Using DXVAChecker v3.0.1, I used as benchmark clips two of my 1080p video posted in the first page.
I chose No5 (5.Cat-1080p60fpsRef4-25Mbps) and No7 (7.Vortexx_1088p24fpsRef3-109Mpbs) because both clips can be tested with all three codecs (LAV, MS DS and MS MFT).
The .mkv clips can't be tested without an MKV MFT splitter.
So, for clip No5 - high frame-rate, medium bandwidth (Lav latest nightly build)
LAV Avg: 410fps
MS DS: Avg: 440fps
For clip No7 - low frame-rate, huge bandwidth
LAV Avg: 210fps
MS DS Avg: 250fps
I hope it's clear now.
Waiting for your HD4000 and HD4600 results.
One last important thing.
You have to disable every VPP function in control panel, even the hidden ones like VIDEO - > color enhancement, VIDEO -> Image enhancement, VIDEO -> Image scaling
You have to toggle everything to "Application settings" and "OFF"
nevcairiel
2nd March 2014, 15:36
For the two clips you mentioned. Benchmarked with DXVAChecker 3.0.1 in EVR renderless mode, on my HD4600
Cat:
LAV DXVA2-Native: 981 fps
MS DS: 960 fps
Vortex:
LAV DXVA2-Native: 410 fps
MS DS: 372 fps
Important to note is that the MS DS decoder seems to generally have at least twice as much CPU usage according to DXVA Checker.
Of course all VPP options are disabled, I double checked just in case.
NikosD
2nd March 2014, 15:38
No.
Not EVR renderless, in EVR option.
nevcairiel
2nd March 2014, 15:45
That only brings the results closer together
Cat: LAV 783 fps, MS DS 778 fps
Vortex: LAV 291 fps, MS DS 290 fps
Note that I run every file 6 times, then take the average from the last 5 runs (ignoring the first, since it always varies a lot)
qtwebkit
2nd March 2014, 19:05
I've also tested the two clips, in nevcairiel's method - 6 times for each and abandon the first time.
OS: Windows 8
Graphic: HD4000
Render: EVR
Mode: DXVA2 Native
LAVFilters: latest nightly
DXVA Checker : 3.0.1
Cat:
LAV : 256.4fps 9%
MS DS : 262fps 12.4%
Vortexx:
LAV : 180.6fps 11.2%
MS DS : 157.4fps 17.6%
Hmm...I got a result different from both of you
If you would like to see results of MPC-BE and ArcSoft, see the detailed material here (https://onedrive.live.com/redir?resid=C134CFE2F5253FB2%21214).
NikosD
2nd March 2014, 19:11
Your results are too low for both clips, for both decoders and for HD 4000.
I'm pretty sure you haven't disabled all VPP functions in Intel's control panel.
qtwebkit
2nd March 2014, 19:21
I checked it before testing and made sure the options you mentioned is set all right.
I think maybe it's the driver problem. I've tried many times to update the Intel graphic driver but all failed with BSOD or black screen... hence I have to use the old oem driver plubished in June 2012.:(
I'm going to update to Windows 8.1 and try again if the new driver works properly.
NikosD
2nd March 2014, 19:26
OEM driver from 2012?
You have to update it ASAP.
Maybe it's time for a format or even better a Win 8.1 upgrade.
The progress is great between those 2012 drivers and the nowadays drivers.
qtwebkit
2nd March 2014, 19:32
Yes, I'll handle it immediatly. As soon as I finnish, I will test again and post my result, maybe tomorrow.:)
Things become more complicated as I upgrade to Windows 8.1. The performance seems benifit little from the new driver(Jan 28,2014), about 277fps for in Cat and 180fps in Vortexx for LAV, merely no improvement in Vortexx.
I've just found something else maybe ralated to this weird problem. I try to run the decoders with EVR Renderless, the result is much closer to nevcairiel's.
Cat:
LAV : 677fps
MS DS : 682fps
Vortexx
LAV : 297fps
MS DS :253fps
But EVR Renderless shows nothing but a black screen during benchmark, while EVR will show video but end up with a very low score:confused:
qtwebkit
3rd March 2014, 20:30
A friend tested the two clips with his HD4000, Windows 8.1 x64, both x64 and x86 filters, but result very close to mine.
performance of x86 filters(x64 scores quite close to x86)
Cat:
LAV : 290fps
MS DS : 293fps
Vortexx:
LAV : 198fps
MS DS : 187fps
Despite the small gap between our results, we are both too slow comparing with yours.
And the same problem, or bug? Black screen in EVR renderless.:(
nevcairiel
3rd March 2014, 20:47
EVR renderless is not supposed to show an image.
qtwebkit
3rd March 2014, 21:10
EVR renderless is not supposed to show an image.
Hmm...but our score in EVR renderless are extraordinary high (for me Votexx could reach 1500fps with no cpu usage), for some specific clips.
NikosD
3rd March 2014, 21:11
Things become more complicated as I upgrade to Windows 8.1. The performance seems benifit little from the new driver(Jan 28,2014), about 277fps for in Cat and 180fps in Vortexx for LAV, merely no improvement in Vortexx.
Can you post again which is your CPU ?
Can you see the clock of iGPU with GPU-Z for example, during benchmarking ?
HD4000 should be faster than my GT1.
Are you sure you have an HD4000 and not an HD2500 ?
I've just found something else maybe ralated to this weird problem. I try to run the decoders with EVR Renderless, the result is much closer to nevcairiel's.
Cat:
LAV : 677fps
MS DS : 682fps
Vortexx
LAV : 297fps
MS DS :253fps
Not even these numbers are fast.
For me LAV is 340 fps for EVR renderless and Vortexx
EVR renderless is not supposed to show an image.
Right :)
Hmm...but our score in EVR renderless are extraordinary high (for me Votexx could reach 1500fps with no cpu usage), for some specific clips.
EVR renderless shows maximum potential of decoder, because it doesn't have to display the decoded frames, that's why the screen is black and you could go to 3000 fps or more for low resolution/ low bandwidth clips.
qtwebkit
3rd March 2014, 21:26
My CPU is i3-3110m
http://gpuz.techpowerup.com/14/03/03/ff4.png
Cat
http://gpuz.techpowerup.com/14/03/03/9ff.png
Vortexx
http://gpuz.techpowerup.com/14/03/03/h3d.png
Hope these information would help.
NikosD
3rd March 2014, 21:37
Well you do have HD4000 with 16 EU's while I have 10.
Your iGPU goes to maximum at 1GHz which is 10% slower than mine which goes up to 1.1GHz.
All of the above don't explain your inferior performance compared with my GT1.
I have to repeat myself just one more time because I can't think of anything else.
Are you sure that ALL VPP functions are disabled in control panel ?
They can push performance down A LOT.
qtwebkit
3rd March 2014, 22:28
Yes,I'm quite sure everyting in the 3 options
VIDEO - > color enhancement, VIDEO -> Image enhancement, VIDEO -> Image scaling
are toggled to "Application settings" and "OFF"
I've been trying to exclude my omissions and turn to my friends for help, who also did all the settings as you said, but end up with the same problem .
Thank you for your concern on my problem, as well as nevcairiel. :)
qtwebkit
4th March 2014, 16:41
I wonder if it‘s the CPU's problem? Since we both use mobile processors, while yours is for destop.
NikosD
4th March 2014, 18:21
Can you benchmark this ?
ftp://helpedia.com/pub/multimedia/x264/testvideos/2160p%20samples/DucksTakeOff_2160p50.x264.CRF24.mkv
It's an UHD file so you have to check 4K box in LAV properties.
qtwebkit
5th March 2014, 02:53
It seem DucksTakeOff could not be tested in MS DS.
EVR
LAV : 67fps
MS DS : failed
EVR Renderless
LAV : 126fps
MS DS : failed
NikosD
5th March 2014, 06:30
Mine are: 99/146
There is a bottleneck in your system, but I can't tell where.
Yups
5th March 2014, 13:14
I've installed latest iGPU drivers for my Pentium G3420 yesterday v.3412 and find out today two extraordinary things that I would never expect.
1) QuickSync transcoding - yes decoding and encoding - is available for the first time in a Pentium CPU (Haswell-DT core) !!
I tried it and it really works!
QuickSync encoding with Pentium G3420!
2) HEVC_VLD_Main appeared while enumerating DXVA device decoders !!
Can you check this driver: https://downloadcenter.intel.com/Detail_Desc.aspx?agr=Y&DwnldID=23644
It is the first official driver with API 1.8 support, maybe there is something different for decoding as well.
NikosD
5th March 2014, 13:18
I have seen that driver at the Intel support forum, but I think is only for gaming.
No difference in other things.
What is the API 1.8 and where did you read that ?
Never heard before.
Yups
5th March 2014, 13:27
This is not just for gaming, it's a new driver with other new features as well. API 1.x is the Media SDK API for Quicksync features. In order to support all new (Hardware) Media SDK 2014 features API 1.8 is required. Media SDK 2013 R2 had API version 1.7. For example you can check the API version in the Handbrake log.
NikosD
5th March 2014, 13:44
Thanks for the info!
I didn't know.
After googling a little I found out this:
http://software.intel.com/en-us/forums/topic/499189
It seems that my assumptions were right. There is a plan from Intel to support a Hybrid CPU+GPU HEVC decoder for this generation processors.
Also API 1.8 is mainly for Broadwell but as Peter Larsson from Intel says about API 1.8:
For instance HEVC codec id, deinterlace control, audio and lookahead control improvements are not dependent on next gen. Core.
He also says:
On that note, there are plans to release a "hybrid" (utilizing EUs of the Core Processor graphics unit) HEVC decoder later this year.
The last post of that page refers to my post at AVS Forum!
Anyway, this driver and the interesting comments on that link above is mainly for developers like Nevcairiel.
I'm not ;)
qtwebkit
5th March 2014, 14:50
There is a bottleneck in your system, but I can't tell where.
:o It's really hard to tell. I'll find more people to help with the test. Once I had new findings, I'll feed back to you.
Yups
5th March 2014, 14:57
I've checked it.
15.33.8.64.3345
http://s14.directupload.net/images/140305/rwzp4esa.png
15.36.0.64.3380
http://s14.directupload.net/images/140305/ncdcaz3o.png
15.33.14.64.3412
http://s1.directupload.net/images/140305/qkmbzsma.png
15.33.0.3464
http://s1.directupload.net/images/140305/k23wj6yj.png
It's gone in 3464 for some reason. Maybe they intend to bring it back in 15.36 driver series, the 3380 alpha already has more DXVA entries than any 15.33.
nevcairiel
5th March 2014, 15:16
It was more then likely a mistake to have it there in the first place. Something that slipped through the cracks.
I wouldn't be surprised if their "hybrid" approach only works with the Media SDK, and not through official DXVA.
NikosD
5th March 2014, 17:16
So, only Eric's QS decoder will be able to utilize it ?
NikosD
7th March 2014, 08:34
The question of an Intel forum member referring to my AVSForum post, was moved to a separate post inside Intel forums, regarding HEVC_VLC_Main support.
It's here:
http://software.intel.com/en-us/forums/topic/506797
Interesting reply of the Intel rep:
Hi,
The support in the driver is not fully validated and Intel will be careful to not expose the interface in the future, unless it is a validated feature.
Yups
9th March 2014, 21:04
From the answer it sounds like the HEVC feature will come officially with 15.36 drivers series which is currently in alpha or beta state.
NikosD
9th March 2014, 21:14
From the screenshots you posted there is only one 15.36 driver which is the first one that has HEVC support.
In the next driver they got busted by me and they removed it in the latest beta.
BTW, where did you find 15.36.3380 ? :)
I think the one appeared in Windows Update after 3345 was 3379 but I don't remember if it was 15.36 or 15.33.
You are probably right about 15.36 and HEVC, we'll see...
Yups
10th March 2014, 14:26
There were no 15.36 drivers through Windows update. For example you could download from here: http://www.station-drivers.com/index.php/articles/763-intel-hd-iris-graphics-version-15-36-0-3380-alpha
HEVC decode now supported by the driver and video players can now take advantage of the GPU accelerated decode support offered by Intel
https://downloadcenter.intel.com/Detail_Desc.aspx?agr=Y&ProdId=3720&DwnldID=23885&ProductFamily=Grafik&ProductLine=Desktop-Grafikcontroller&ProductProduct=Intel%C2%AE+Core%E2%84%A2+Prozessoren+der+vierten+Generation+mit+Intel%C2%AE+HD-Grafik+4600&lang=eng
Anyone tried this out? Does anyone has a link to a HEVC video?
NikosD
4th June 2014, 14:53
Thanks!
The driver is latest official and not beta.
Supports Intel® Iris™ graphics, Intel® Iris™ Pro graphics, and Intel® HD graphics on:
4th Generation Intel® Core™ Processor Platform
4th Generation Intel® Core™ Processor U Series-based Platform
4th Generation Intel® Core™ Processor Y Series-based Platform
3rd Generation Intel® Core™ Processor Platform
3rd Generation Intel® Core™ Processor U Series-based Platform
3rd Generation Intel® Core™ Processor Y Series-based Platform
Supports Intel® HD graphics on:
Intel® Pentium® Processor 1403 v2/1405 v2/2020M/2030M/2117U/
2129Y/2127U/A1018/G2010/G2020/G2020T/G2030/G2030T/G2100T/
G2120/G2120T/G2130/G2140
Intel® Celeron® Processor 927UE/1000M/1005M/1007U/1017U/1019Y/
1020E/1020M/1037U/1047UE/G1610/G1610T/G1620/G1620T/G1630
Intel® Pentium® Processor 3550M/3556U/3558U/3560Y/3561Y/G3220/
G3220T/G3320TE/G3420/G3420T/G3430/3560M/G3240/G3420T/
G3440/3440T/G3450/G3258
Intel® Celeron® Processor 2000E/2002E/2950M/2955U/2957U/2961Y/
2980U/2981U/G1820/G1820T/G1820TE/G1830/2970M/G1840/G1840T/G1850
Unfortunately I didn't find a HEVC device decoder using DXVA Checker, but it's the first official driver that supports HEVC decoding (but not transcoding)
HEVC decode now supported by the driver and video players can now take advantage of the GPU accelerated decode support offered by Intel
Transcoding to HEVC is currently not supported.
I have some HEVC files to check, but first developers like Nevcairiel or video players like PotPlayer should build the HEVC decoder utilizing GPU support.
I have this in DXVAChecker:
http://s14.directupload.net/images/140604/lu4yuscf.png
NikosD
4th June 2014, 17:03
Unfortunately I have this:
http://s1.directupload.net/images/140604/o53zp4jt.png
Win 8 I assume. My test is from Win 7 on a HD4600. Shouldn't make a difference though. Maybe Pentium or GT1 in general isn't supported because not enough EU power? Or a bug?
How can I check on a HEVC video if it's working?
NikosD
4th June 2014, 17:51
It seems that the HW decoder uses EUs and that could be a reason of not supporting it on my Intel HD graphics (only 10 EUs)
I don't think it's my Win 8.1 OS.
About HEVC video, you can't test it because as I wrote you before I don't know any video player/ decoder utilizing the new Intel's HW decoder.
Check LAV Video decoder and PotPlayer for updates.
wanezhiling
4th June 2014, 17:52
Yups, You can't check now because currently there is no available hevc dxva decoders.
btw I've reported to PotPlayer devs.
clsid
4th June 2014, 17:56
Current hardware does not support HEVC, so keep on dreaming. It will simply do software decoding.
NikosD
4th June 2014, 19:06
Is there something not clear in the release notes of Intel's driver ?
It clearly says GPU accelerated support.
Obviously partial acceleration using mainly EUs and CPU and probably not QuickSync.
Or OpenCL/CUDA.
pretty much sure we will get this feature for free.
CUDA is Nvidia only and for HEVC decoding Intel surely won't use OpenCL. They made it clear in the past that execution units are used for this as Quicksync doesn't support it in the current hardware. EUs are part of the (GPU) hardware so the comment from clsid is clearly nonsense.
NikosD
4th June 2014, 19:57
Exactly (I would give a point to OpenCL, although only AMD has clearly said that will follow this path)
nevcairiel
4th June 2014, 23:47
Intels hardware decoder is quite flexible, I bet they can at least do parts of the decoding process in hardware, and the remaining parts in the EUs or on the CPU. We'll see how it performs once someone implements it, but for me it'll be a while since I'll be out of the country for a couple weeks soon, and somehow I doubt someone else will beat me to it, its quite a complex endavour and entirely new code, can't steal it from some place. ;)
GTPVHD
5th June 2014, 14:51
https://software.intel.com/en-us/forums/topic/499189
Intel Core Processor platforms currently do not have fixed function HW HEVC codec capability. This is a feature that may be supported in future Intel Processors. On that note, there are plans to release a "hybrid" (utilizing EUs of the Core Processor graphics unit) HEVC decoder later this year.
You won't see power efficient full HEVC hardware decoding until Skylake, even Broadwell does not support full HEVC hardware decoding.
nevcairiel
5th June 2014, 14:57
We're all well aware that there is no full fixed function decoding of HEVC, as discussed just a couple posts above.
This is the hybrid variant, which off-loads certain tasks to the hardware.
We're all well aware that there is no full fixed function decoding of HEVC, as discussed just a couple posts above.
This is the hybrid variant, which off-loads certain tasks to the hardware.
Exactly this. We knew since months that QS of Haswell doesn't support it. The Intel answer is old, already linked here and known. We are happy enough that it works via the shader somehow. Not as power efficient and fast as a pure fixed function solution but it should help to lower the CPU dependence, especially for slower mobile Dualcore systems this is important.
NikosD
5th June 2014, 21:41
Intels hardware decoder is quite flexible, I bet they can at least do parts of the decoding process in hardware, and the remaining parts in the EUs or on the CPU.
It seems that no part of the decoding process is going to be executed in fixed-function HW decoder (QuickSync), although as nevcairiel has said, it is quite flexible (but probably not enough for HEVC decoding).
VLD will be done by CPU and the other parts by EUs.
I think that is what GTPVHD wanted to say.
Not as power efficient and fast as a pure fixed function solution but it should help to lower the CPU dependence, especially for slower mobile Dualcore systems this is important.
Unfortunately, my dual core Pentium processor doesn't have HEVC DXVA enabled.
There are lots of mobile Core i3, i5, i7 2/4 low TDP/clocked models in the market. Not sure if it's intended that Pentium models don't have HEVC enabled, if this is the case Intel should make it clear in their changelog. You could try to ask them in the Media SDK forum. In their graphics section I doubt you will get a useful answer.
NikosD
6th June 2014, 10:22
Their excuse could be "Intel HD Graphics" with only 10 EUs, that all Pentium iGPUs have.
On the contrary, most - if not all - Core iX processors have 20 EUs, like the ones you mentioned above.
But for Pentiums, is more useful and necessary to have GPU support for HEVC decoding, due to slower CPU performance compared to other Core iX dual cores.
vortex_hl
17th June 2014, 19:16
fyi geforce 340.43 drivers support HEVC_VLC_Main profile on GTX 750 Ti
http://i.imgur.com/iloiSBq.png
NikosD
18th June 2014, 06:46
Very interesting.
It seems that Nvidia with VP6 is closer to Intel than ever.
Sulik
18th June 2014, 08:41
NikosD: any plans to include VP6 numbers in your benchmark results ?
NikosD
18th June 2014, 11:11
If I had one, I would do it immediately :)
Anteys
21st June 2014, 07:52
Unless I'm misinterpreting something, my Titan with VP5 on 340.43 shows the same support for HEVC_VLD_Main:
http://i.imgur.com/07ccL3w.jpg
NikosD
21st June 2014, 09:46
Then I'll have to check my poor GT610 with VP5, to see if HEVC_VLD_Main depends on fixed-function HW (VP5&VP6) or as I suspect, on the number of shaders.
Which means that my card - GT610 - won't have it.
Wait just for 10 min.
Update:
Unfortunately I was right
http://s10.postimg.org/effteamc9/GT610_340_43.png
NikosD
21st June 2014, 10:15
Now, only one more test is missing.
A VP4 card in the middle or upper class, like the popular GT440 or better, in order to test if VP5 or VP6 has anything to do with HEVC_VLD_main, or it is something about shaders and CPU only (as I suspect)
Of course, marketing reasons could stop Nvidia of installing HEVC_VLD_Main on older Fermi cards with VP4, even if Fermi cards with a lot of shaders, could be capable of running this mode.
Intel for example, decided to remove HEVC_VLD_Main support from Haswell Pentium/Celeron for marketing reasons.
Wow, so my 750 Ti and Intel HD 4400 supports HEVC too :goodpost:
Can't wait to get the latest driver :D :D
But how to test it? any HEVC DxVA decoder??
NikosD
30th June 2014, 15:41
This is not a good sign.
Intel cut every codec to 1080p with latest beta driver 3652!
http://s27.postimg.org/scbz7bfer/DXVA_Checker_3652_Pentium.png
Shiandow
30th June 2014, 19:05
Now, only one more test is missing.
A VP4 card in the middle or upper class, like the popular GT440 or better, in order to test if VP5 or VP6 has anything to do with HEVC_VLD_main, or it is something about shaders and CPU only (as I suspect)
I haven't really been following the discussion, but I happen to have a GTX560Ti, which is a VP4 card. So does this screenshot help?
http://i.imgur.com/1FntopP.png
Nvidia GTX560Ti, driver version 335.23.
Windows 7 x64 SP1.
Same result after updating the driver to 340.43
NikosD
30th June 2014, 19:40
Yes, this is exactly the test I wanted to see.
So only Kepler and Maxwell cards are included for HEVC.
Fermi cards are excluded regardless the number of shaders.
foxyshadis
1st July 2014, 00:53
This is not a good sign.
Intel cut every codec to 1080p with latest beta driver 3652!
I saw your posts on the Intel board. I'm not convinced this is going to be final behavior, but I appreciate you pressing them to not give up just because it's technically difficult (or worse, marketing, as you suspect).
I have to say I'm becoming quite annoyed at needing separate drivers for every Intel system at this point. I have multiple generations of PCs I keep reasonably up-to-date, especially when reinstalling, and now I'll have to keep separate installers for Core2, Sandy, Ivy, and Haswell.
FWIW, HEVC_VLD_Main isn't just missing from Pentium brands, this is my Haswell i7 with the beta:
NikosD
1st July 2014, 12:25
Thanks.
Yes, it's true, maintenance of different systems even from the same brand is getting more and more difficult.
We have already different drivers for Sandy.
If they separate Ivy from Haswell it will be even harder.
NikosD
3rd August 2014, 14:11
Interesting HEVC OpenCL solution from Strongene here:
http://xhevc.com/en/downloads/downloadCenter.jsp
It's a DS filter working with discrete AMD cards (HD 5000 or better) or APU.
I haven't tried it with Nvidia/Intel, it could work.
Using my old Core2Duo and HD 5750 card, it works for most of my H.265 samples, but with no clear performance benefits over the CPU only decoding.
My platform is very old supporting PCI-E v1.1x4, so maybe that's the problem.(it needs PCI-E v2.0)
It's free, you can try it with better AMD cards/platform or with Nvidia/Intel cards.
huhn
3rd August 2014, 14:49
tried it with a r 9 270 i can't see a a real GPU usage but cpu usage is about 20-50% lower but memory usage was a lot higher.
i try it with EVR next.
i have a haswell hd 4400, a ivy hd 4000, a 760 gtx and the r9 270 in working pc right now to test.
Edit: gpu uasge is 0-31 % in highest powerstate with EVR
NikosD
3rd August 2014, 14:55
Try to use DXVA Checker in benchmark mode in order to see GPU usage going high and compare with the CPU only decoding.
What renderer did you use if not EVR ?
huhn
3rd August 2014, 15:14
i normally use madVR.
here a small test:
openCL32 187 130-255
lav32 87 50-127
lav64 142 82-198
openCL decoder is totally limited by my CPU a i3 4130 it just uses the GPU to accelerate the decoding nothing more.
NikosD
3rd August 2014, 15:39
I can't understand the figures.
What is the first and what is the second.
Did you try a 1080p HEVC clip ?
huhn
3rd August 2014, 15:47
decoder avgfps minfps-maxfps
so openCL 32 bit has an avg of 187 fps with min of 130 and a maximum of 255 fps.
the openCL decoder isn't working with h264 so yeah i used a HEVC clip.
the forum removed the formation of the "chart" so no wonder is hard to understand the format was fine while i was tipping.
P.J
3rd August 2014, 16:08
It crashes while playing a 4K HEVC video with 750Ti
huhn
3rd August 2014, 16:37
i tested it with a intel hd 4400.
it crashes when the AMD gpu is disabled!
it has about 23% gpu usage when run with lavfilter
and it has ~88% gpu uses when used with the openCL decoder.
Renderer: Enhanced Video Renderer (DirectShow)
Decoder: Lentoid HEVC Decoder (OpenCL)
[1]
Decoder Device: -
Processor Device: BF752EF6-8CC4-457A-BE1B-08BD1CAEEE9F
Time: 22.427
Frames: 2183
Avg FPS: 97,336fps (Min-Max: 54-101fps)
Avg CPU Usage: 31% (Min-Max: 16-56%)
Avg GPU Usage: -
[2]
Decoder Device: -
Processor Device: BF752EF6-8CC4-457A-BE1B-08BD1CAEEE9F
Time: 22.268
Frames: 2183
Avg FPS: 98,032fps (Min-Max: 73-100fps)
Avg CPU Usage: 33% (Min-Max: 17-58%)
Avg GPU Usage: -
as you can see the cpu was limited by the gpu this time.
my guess the decoder crashes when no amd GPU is found but works in theory with intel too.
edit: the gpu uasge is this time from gpu-z because it is not shown with in dxva tester.
clsid
3rd August 2014, 16:53
Performance test in GraphStudioNext (with null renderer). TearsOfSteal 1080p sample.
OpenCL: 278 fps
LAV x86: 133 fps
LAV x64: 280 fps
huhn
3rd August 2014, 16:58
how high where your gpu usage and what gpu?
NikosD
3rd August 2014, 17:30
@huhn My figures are like yours, see below (the analogy, not the absolute values)
Also Nvidia works fine without AMD card in the system :)
@P.J Try a 1080p clip
@clsid GPU ? CPU ?
I tried my Core i5-2400 with an Nvidia GT440, Win 8.1 x64
Using DXVA Checker x86 (32bit) in benchmark mode and EVR on TearsOfSteel 1080p HEVC sample
LAV x86: 59/90/229
OpenCL x86: 152/224/264
Using DXVA Checker x64 (64bit) in benchmark mode and EVR on TearsOfSteel 1080p HEVC sample
LAV x64: 105/180/307
We have a clear winner which is OpenCL of course.
The GPU load during OpenCL benchmarking was about 45%-50%, during LAV (x86/x64) was about 10%-15% due to EVR renderer.
LAV x64 is not optimized at all for multithreading.
The CPU usage was lower (a lot) from both LAV x86 and OpenCL decoder.
I'll try my brand new Core i7-4790 later (without discrete card)
JohnLai
3rd August 2014, 18:11
O.o? I didn't notice you also active in this forum.
So, since you are able to actually run the opencl decoder, how is the picture/image quality?
huhn
3rd August 2014, 18:20
it's a decoder it's suppose to be bit identical to other decoder...
NikosD
3rd August 2014, 20:16
@JohnLai I'm active here since 2010.
I was about to write to your post, when I saw you here.
You found me before I find you :)
Yes the quality is the same just like any other HEVC decoder, I suppose.
Only with my old platform of PCI-E v1.0 I had some problems with not so smooth decoding.
But with Nvidia I had no problem at all.
I'll try to put the 5750 on the Core i5 platform to see the difference.
5750 should be faster, maybe a lot faster than GT440 in OpenCL.
Now, regarding to my Core i7-4790 CPU with HD 4600@1.2GHz GPU, the results were not good.
It's a 4C/8C CPU, but only one core was active (!)@ 4.0GHz - no multithreading at all.
I don't know why, probably due to low OpenCL performance of HD 4600 or maybe it wasn't recognized properly from the decoder.
The threads were to "Auto" from codec properties, I tried to put them manually to 4 and then 8 threads, but nothing changed.
Only one core was active and couldn't even load the GPU properly.
I got a clock of GPU @750MHz with 77% utilization when max clock is 1.200MHz
The result was disappointing:
OpenCL x86 69/112/128
LAV x86 88/130/208 (all 8 threads active@3.8GHz)
huhn
3rd August 2014, 20:29
looks totally normal when i compare it to my hd 4400 at 98 fps the cpu was at about 1 thread too.
but i don't know if we should call this a openCL decoder i mean most work is done by the CPU...
NikosD
4th August 2014, 04:31
It is OpenCL decoder, because decoding on CPU only (Core i5-2400), LAV x86 goes up to 90 fps (avg), where OpenCL Decoder using GT440 on the same CPU (Core i5-2400) goes 224 fps (avg).
The difference is exactly 2.5 times faster.
So it seems that OpenCL decoder has been optimized to offload special difficult parts of HEVC decoding to GPU or LAV x86 is extremely slow or OpenCL CPU part is extremely optimized.
Any of three or all of them could be true.
But in realtime playback the GPU usage was low and looking at LAV x64 with 180 fps (avg), then LAV x86 is definitely slow with half speed and OpenCL decoder looks like an optimized CPU decoder more than OpenCL decoder.
When I go back to Core i5 system, I'll test more difficult clips in real time decoding, like 4K HEVC to check GPU usage.
I still don't understand why is using 1 thread only with Intel iGPU.
NikosD
4th August 2014, 06:29
I have finally found out a real use and real strength of Lentoid HEVC Decoder (OpenCL) for Intel iGPU systems.
I tried a 4K HEVC sample (It's Ducks Take off - 3840x2160@25fps - 32.3Mbps - HEVC Main@L5.1) with my Core i7-4790 (iGPU HD 4600) and all 8 threads were used.
Benchmark x86
LAV 7/18/20 CPU Usage 88% (8 threads)
OpenCL 11/35/47 CPU Usage 55% (8 threads) - GPU clock/usage 900MHz@88%
Benchmark x64
LAV 14/25/27 CPU Usage 85% (8 threads)
Real-time playback
LAV x86 19fps (avg) - CPU 92% (non real-time decoding - struggling playback)
OpenCL x86 25fps (avg) - CPU 32% (very smooth playback) GPU idle clock/usage 600MHz@66%
LAV x64 25fps (avg) - CPU 88% (smooth playback)
I think it's clear which is the most efficient HEVC decoder right now.
JohnLai
4th August 2014, 06:47
Hmm........I wonder what portion of decoding that being offloaded to GPU.....resizing? deblocking?
Ah, back to further elaboration of my original question on Lentoid opencl HEVC Decoder. Reason I asked about the decoder output quality is due to past h264 decoders.
When h264 decoder was in its infancy, developers used a lot of quality-damaging 'shortcut' in order to ensure smooth playback, example, the famous 'skip deblocking' resulted in smooth playback but blocky image output.
NikosD
4th August 2014, 06:54
For the first part of your question about what is offloaded to GPU, I think only Strongene could answer for sure and accurately, but I doubt if they would give such details about their decoder.
For the second part of your question, I totally agree with you about H. 264 decoders.
I don't know if they can "cheat" with H.265 decoders, too.
P.J
4th August 2014, 21:48
Also Nvidia works fine without AMD card in the system :)
@P.J Try a 1080p clip
1080p HEVC is useless since my i3 4130 can handle it ~%40 easily. HEVC is mostly for 4K.
I'll try HD4400 instead of 750 Ti...
P.J
4th August 2014, 22:41
Ok, seems it crashes while playing 60fps 4k h.265 but no problem with Ducks Take off 4K HEVC
Then I tried with both HD4400 and 750Ti and the CPU usage is still 99% but smoother than LAV x64 (RAM usage: 1GB)
HD4400: ~60% @ 600MHz
750Ti: ~60% @ 200MHz (1176MHz Max)
Then the decoder only uses a small and specific part of any GPU...
Anyway to benchmark it in DXVA Checker? It doesn't recognize the h.265 video at all
NikosD
5th August 2014, 04:42
What do you mean it doesn't recognise the H. 265 video at all ?
If you mean the OpenCL decoder, you have to run the x86 version of DXVA Checker, not the x64.
If you mean the Ducks Take Off 4K HEVC clip, you have to install LAV splitter.
P.J
5th August 2014, 12:00
Worked now:
Renderer: Enhanced Video Renderer (DirectShow)
Decoder: Lentoid HEVC Decoder (OpenCL)
Decoder Device: -
Processor Device: 6CB69578-7617-4637-91E5-1C02DB810285
Time: 22.882
Frames: 500
Avg FPS: 21.851fps (Min-Max: 0-26fps)
Avg CPU Usage: 94% (Min-Max: 46-100%)
Avg GPU Usage: -
Renderer: Enhanced Video Renderer (DirectShow)
Decoder: LAV Video Decoder x64
Decoder Device: -
Processor Device: 6CB69578-7617-4637-91E5-1C02DB810285
Time: 53.390
Frames: 500
Avg FPS: 9.365fps (Min-Max: 3-11fps)
Avg CPU Usage: 97% (Min-Max: 48-100%)
Avg GPU Usage: -
Waiting for better OpenCL optimization...
nevcairiel
5th August 2014, 13:28
OpenCL is in general not all that super useful for video decoding. People tried before with H.264 to leverage it, and it just didn't help all that much.
The decoder is probably just better optimized for CPU decoding, so its faster. Optimizations for the HEVC decoder in ffmpeg (which LAV uses) are coming basically every day, as clsid mentioned in another thread it got 20% faster just recently, and there is still a lot of potential for improvements.
NikosD
5th August 2014, 15:31
Decoder: Lentoid HEVC Decoder (OpenCL)
Avg FPS: 21.851fps (Min-Max: 0-26fps)
Avg CPU Usage: 94% (Min-Max: 46-100%)
Decoder: LAV Video Decoder x64
Avg FPS: 9.365fps (Min-Max: 3-11fps)
Avg CPU Usage: 97% (Min-Max: 48-100%)
Waiting for better OpenCL optimization...
It's obvious comparing your results that CPU usage hasn't changed using OpenCL decoder.
BUT it's more than obvious that with the same CPU usage you have 2.5 times faster result, exactly as I described it in a previous post!
With OpenCL decoder you can just play the clip in almost real-time, while using LAV x64 the playback should be awful.
OpenCL decoder is not perfect, but it's 2.5 times faster than LAV.
I think, that's a huge difference between those two free decoders.
I suggest to download and try PowerDVD 14 which supports OpenCL HEVC for AMD, but probably will work for other GPUs too and tell us your results.
Optimizations for the HEVC decoder in ffmpeg (which LAV uses) are coming basically every day, as clsid mentioned in another thread it got 20% faster just recently, and there is still a lot of potential for improvements.
Downloaded and tried LAV x64 0.62.0.3287 4/8/2014 compilation including latest FFMpeg update and the results were disappointing.
Exactly the same performance for 4K HEVC:
LAV x64 0.62 (Jun-2014) : 14/25/27
LAV x64 0.62 (Aug-2014) : 10/25/30
I don't see where you saw the 20% improvement.
P.J
5th August 2014, 16:39
PowerDVD works quite different, faster than LAV but not perfect... and I think it's not OpenCL based, almost no GPU usage
Waiting for a proper DXVA2 decoder :)
OpenCL is in general not all that super useful for video decoding. People tried before with H.264 to leverage it, and it just didn't help all that much.
The decoder is probably just better optimized for CPU decoding, so its faster. Optimizations for the HEVC decoder in ffmpeg (which LAV uses) are coming basically every day, as clsid mentioned in another thread it got 20% faster just recently, and there is still a lot of potential for improvements.
What about QuickSync and CUDA?
huhn
5th August 2014, 17:01
you can't program in quicksync.
cuda should have the same problems as openCL i guess it has something to do with the huge used numbers but what did i know ^^
the "dxva" HEVC decoder where planned as openCL. for true DXVA support we need new ASIC with HEVC support so new cards.
the planned "DXVA" decoder aren't comparable with a true hardware DXVA decoder.
so this comment was very true http://forum.doom9.org/showpost.php?p=1682920&postcount=247
you can write a decoder for everything in openCL does that mean we "have" hardware decoder for everything?
NikosD
5th August 2014, 19:25
... and I think it's not OpenCL based, almost no GPU usage
That "almost no GPU usage" is the same scenario for Strongene's OpenCL decoder too.
Probably both decoders leverage GPU decoding capabilities using the enormous parallel processing of GPU processors in small parts of the HEVC algorithm, but a lot times faster than CPU decoding.
Waiting for a proper DXVA2 decoder :)
Proper DXVA2 decoder needs proper HW.
Broadwell will not have QuickSync engine capable of H.265 decoding and VP6 from Nvidia doesn't support H.265 decoding too.
So, for the next couple of years, you have to play and fight with implementations like OpenCL decoder or Intel's and Nvidia's HEVC_MAIN_VLD decoder which use CPU+GPU in a similar way like OpenCL decoder (meaning that they leverage the EUs or SMXs or whatever you want to call the computing processors of GPUs)
so this comment was very true http://forum.doom9.org/showpost.php?p=1682920&postcount=247
That comment was and is completely inaccurate at the second part.
Like OpenCL decoder, the Nvidia's and Intel's HEVC_MAIN_VLD implementations are leveraging both CPU+GPU compute engine, in order to decode HEVC.
If OpenCL decoder is SW solution, then how come during playback of 4K HEVC, I need 92% of CPU power of Core i7-4790 for real-time playback using LAV x64, when OpenCL decoder achieves same result with 32% CPU usage ?
Is the OpenCL decoder almost 3 times faster due to CPU optimizations only ?
I doubt it.
UPDATE:
After the above post I did some more "weird" tests and maybe the OpenCL decoder is indeed a lot faster than LAV in CPU decoding.
Look here:
http://forum.doom9.org/showthread.php?p=1689049#post1689049
huhn
5th August 2014, 19:34
it is GPU assisted software yes.
is GPU assisted software now a hardware decoder?
is everything running on a GPU hardware code not software?
does this help at all?
P.J
5th August 2014, 19:48
I believe PowerDVD HEVC decoder only uses CPU while Strongene OpenCL HEVC decoder uses both CPU and GPU.
HEVC_MAIN_VLD should be better solution if it can use the full potential of GPU and not only a small part of it.
NikosD
5th August 2014, 19:48
@huhn
Yes it helps.
My CPU using LAV x64 needs 92% usage for real-time playback of 4K HEVC and only 32% CPU using OpenCL decoder.
Do you want any further proof ?
The only thing is how much the GPU really assists for dropping CPU usage and how much of this drop happens due to a lot better CPU decoding of OpenCL decoder.
NikosD
5th August 2014, 19:51
@P.J
I haven't tried myself PowerDVD, I only said what I read.
Maybe it's for AMD GPUs only.
P.J
5th August 2014, 20:10
@P.J
I haven't tried myself PowerDVD, I only said what I read.
Maybe it's for AMD GPUs only.
Where did you read that?
http://www.cyberlink.com/products/powerdvd-ultra/spec_en_GB.html
NikosD
5th August 2014, 20:12
From Anandtech's presentation and review of A10-7800:
The other feature using the GCN cores is HEVC Compute support with PowerDVD 14, using OpenCL to speed up decoding for high definition content. With a soon-to-be released update, AMD Fluid Motion Video should also be supported
AMD also points out in its release that PowerDVD 14 is fully supporting HEVC compute via OpenCL on AMD APUs, with also AMD Fluid Motion Video in a later update.
P.J
5th August 2014, 20:21
I don't believe them :D
huhn
5th August 2014, 20:39
@huhn
Yes it helps.
does this help at all?
this was pointed at discussing about hardware and software decoder not if the openCL version is or can be faster.
this openCL decoder is currently fast if not blocked by a terrible bad GPU i don't argue that!
foxyshadis
6th August 2014, 01:38
The Strongene decoder may very well be the same decoder as what's in PowerDVD, although PowerDVD might have licensed an older version. From their release notes:
The OpenCL version is developed jointly by Strongene and AMD; [...] This version supports the OpenCL devices like AMD HD 5000 and above discrete GPUs, and AMD APUs (like Richland and Kaveri).
The fact that it works at all on nVidia or Intel is fantastic, since it obviously wasn't optimized for them in any way. At least it wasn't specifically disabled on them.
Edit: Downloaded the PowerDVD trial, it's definitely using different code for its OpenCL.
clsid
6th August 2014, 14:00
Has anyone tested if the decoder also works without OpenCL? For example by temporarily switching to the Generic VGA Driver.
It would be interesting to see its pure CPU performance.
P.J
6th August 2014, 18:06
Edit: Downloaded the PowerDVD trial, it's definitely using different code for its OpenCL.
Would you share some results? Does it use any GPU at all compared to LAV?
NikosD
6th August 2014, 19:23
@ clsid
Good idea.
I disabled from Device Manager the HD 4600 driver so I used "Microsoft Basic Render Driver" - you don't have to uninstall your driver for anyone would like to try.
Initially I used DXVA Checker in benchmark mode as usual using EVR renderer.
The results with OpenCL decoder look like it didn't work (very low CPU usage ~7 fps performance on Core i7-4790)
Then I used EVR Renderless and I did two tests:
One with 4C/4T configuration (HT OFF) and the other at 4C/8T (HT ON)
Also, in order to use OpenCL Decoder in OpenCL mode, I did the exact same tests using EVR Renderless mode with OpenCL (HD 4600 driver)
LAV x86/x64 had exactly the same results, with or without OpenCL.
But OpenCL decoder, although used in EVR Renderless mode, had a small difference using OpenCL
The results reveal a lot of things.
Core i7-4790 - EVR Renderless
4C/4T
LAV x86 10/17/20 CPU usage: 98%
LAV x64 17/23/28 CPU usage: 97%
OpenCL x86 29/40/44 CPU usage: 96% OpenCL OFF
OpenCL x86 23/41/49 CPU usage: 86% GPU usage 600MHz@60% OpenCL ON
4C/8T
LAV x86 12/19/22 CPU usage: 85%
LAV x64 19/26/30 CPU usage: 78%
OpenCL x86 30/44/51 CPU usage: 82% OpenCL OFF
OpenCL x86 15/44/55 CPU usage: 76% GPU usage: 600@50% OpenCL ON
My comments:
1) CPU usage of LAV x86/x64 and OpenCL x86 on a 4C/4T CPU is excellent - Almost 100% !
2) CPU usage of LAV x64 on a 4C/8T could be optimized better.
LAV x86 has 10% more CPU usage than x64 and more than OpenCL x86 too.
3) For LAV x86/x64 there was no difference between EVR and EVR renderless performance with or without OpenCL enabled.
But for OpenCL x86 using EVR renderless gives it a boost with OpenCL enabled or disabled due to larger CPU utilization and lower GPU usage.
4) It is clear that OpenCL decoder is mainly a CPU decoder and it's faster as CPU decoder on a 4C/8T Core i7-4790, than a CPU/GPU decoder on the same CPU.
5) The use of GPU (HD 4600) as HEVC OpenCL decoder, doesn't make faster the HEVC decoding on a Core i7-4790 4C/8T , but drops the CPU usage a lot by offloading a part of the HEVC algorithm to the GPU during real-time playback and benchmarking with EVR (not EVR renderless)
6) The huge difference between OpenCL decoder and LAV x86/x64 is not the CPU usage (actually OpenCL decoder has less CPU usage compared to LAV x86) but the more optimized use of the CPU - maybe use of different vector instruction set (?)
Asmodian
7th August 2014, 04:28
Thanks for the benchmarks, very interesting results. :)
I don't think you have any data to back up point two, both show a similar drop in CPU utilization with hyper-threading on. It may simply be a limitation of hyper-threading, the extra threads have to share most of the hardware with the original four. Of course there probably is more optimization possible for both decoders.
Edit: Sorry, LAV x64 does have enough lower utilization to be interesting but it might be more due to the nature of hyper-threading instead of less optimization.
I strongly agree with point four, epically looking at minimum frame rates which are the most important for real time display. When using OpenCL the OpenCL decoder min frame rate was half of the same decoder in pure CPU mode! It even dropped below the min frame rate of LAV x64. Interop penalty?
Edit: I hope this interop penalty would not show itself on an AMD APU though I suspect it would still run faster in pure CPU on your i7-4790.
NikosD
7th August 2014, 05:26
I agree with everything you wrote at your post.
LAV x64 probably is pushing HT to its limits on a Core i7-4790, because is a lot faster than LAV x86 per thread.
But looking at the test of clsid measuring fps per thread, maybe you can get a few fps more by using a multiplier of x2 than x1.5 that LAV is using now.
Main issue for LAV is the per thread performance of CPU decoding compared to OpenCL decoder.
The difference is so huge that I think - as I wrote above - that OpenCL decoder is using vectorised code a lot better than LAV.
About the min value of OpenCL, most of the times was 0!
It maybe is an interoperability penalty or EVR renderless is causing the whole thing, because with EVR using OpenCL the min value is normal.
Regarding OpenCL GPU performance, it would be interesting to add a fast GPU to a fast CPU and repeat the tests, especially an AMD GPU.
Because all of my tests were done on the iGPU of Core i7-4790 which is a HD4600.
But because HD4600 is clearly underutilised even by a strong CPU, I doubt that even a R9 290 would make any significant difference.
But we have to test it as always.
huhn
7th August 2014, 06:59
on my i3 4120 the hd 4400 is working like a handbrake for the openCL assisted decoder max CPU usage was about 40%.
but i use a difference file like the rest of you
NikosD
7th August 2014, 07:26
@huhn
I had the same problem with some files, probably due to incompatibility of OpenCL decoder.
Even LAV has problems with some HEVC files displaying black screen or the problems are within the files (first early samples of HEVC encoding)
Google 4K HEVC Ducks Take off sample and try again.
clsid
7th August 2014, 14:08
The raw CPU performance of the OpenCL decoder is a good indication of what we can expect of LAV in the future. The HEVC decoder in FFmpeg is still under heavy development, and there is still lots of room for improvement.
huhn
7th August 2014, 17:04
my data is worthless. it always uses the AMD GPU not my HD 4000 and when i disable the AMD gpu is crashes same goes for EVR renderless it simply stops at one point.
my 1080p file still works like a handbrake even through the AMD GPU is used when the intel GPU is active. EVR renderless gets stuck at one point and but shows the same CPU usage of under 50 %. intel GPU is still at ~88% doesn't make a lot of sense to me...
when the AMD gpu is active everything works fine.
Asmodian
7th August 2014, 22:33
when the AMD gpu is active everything works fine.
How is the min frame rate with the AMD GPU active?
NikosD
8th August 2014, 11:21
Moving my Radeon 5750 from the old Core2Duo PCI-E v1.0 x4 platform to the Core i5-2400 PCI-E v2.0 x16 platform with clean installed Win OS, the picture of Strongene's OpenCL decoder, changed once again (and of LAV too)
1080p HEVC TearsofSteel movie - EVR - OpenCL ON - Benchmarking mode of DXVA Checker
LAV x86 (June 2014): 59/90/229
LAV x86 (Aug 2014): 64/102/239
GT440 OpenCL x86: 152/224/264 CPU usage: 86%
5750 OpenCL x86: 71/109/187 CPU usage: 28% GPU usage max clock@90%
LAV x64 (June 2014): 105/180/307
LAV x64 (Aug 2014): 131/224/360
4K HEVC Ducks Take off - Same configuration as above
LAV x86 (Aug 2014): 6/13/15 CPU 97%
5750 OpenCL x86 4/24/29 CPU 68% GPU 62%
LAV x64 (Aug 2014): 11/18/21 CPU 97%
My comments:
1) We should pay more attention when the developer says that OpenCL decoder is optimized for AMD cards and for each card different resolution could be supported.
2) The OpenCL decoder clearly drops a lot the CPU usage of Core i5 when decoding 1080p clips with 5750, but is not optimized for 5750 & 4K clips.
The CPU usage goes high as the GPU usage drops on 4K.
3) LAV Aug seems more optimized for 1080p HEVC clips than 4K and it's definitely faster than June.
4) Min values still go even to 0 fps (just for an instant) when using AMD card (5750)
huhn
8th August 2014, 16:10
How is the min frame rate with the AMD GPU active?
from which file did you like to know that?
P.J
8th August 2014, 21:09
Renderer: Enhanced Video Renderer (DirectShow)
Decoder: LAV Video Decoder x64 0.62
Decoder Device: -
Processor Device: 6CB69578-7617-4637-91E5-1C02DB810285
Time: 40.965
Frames: 500
Avg FPS: 12.205fps (Min-Max: 5-15fps)
Avg CPU Usage: 97% (Min-Max: 72-100%)
Avg GPU Usage: -
Only 3fps more... waiting for proper solution from Intel/Nvidia ;)
NikosD
9th August 2014, 07:45
Performance test in GraphStudioNext (with null renderer). TearsOfSteal 1080p sample.
OpenCL: 278 fps
LAV x86: 133 fps
LAV x64: 280 fps
I tried GraphStudioNext(with null renderer) version 0.61.265 with OpenCL x86 decoder, but I didn't find the OpenCL decoder filter.
I mean, I can build the graph with the sample, decoder and null renderer, but when I select View -> Performance Test... only a few of the DS filters appear (only the "pure" CPU decoders I think)
There is no OpenCL decoder in that list.
How can I add the OpenCL decoder to that list or how can I test performance of the OpenCL decoder from the graph that I built ?
clsid
9th August 2014, 14:57
Build the performance test graph with LAV Video, and then manually edit the graph to swap the decoder.
wxhyn
9th August 2014, 18:42
Moving my Radeon 5750 from the old Core2Duo PCI-E v1.0 x4 platform to the Core i5-2400 PCI-E v2.0 x16 platform with clean installed Win OS, the picture of Strongene's OpenCL decoder, changed once again (and of LAV too)
1080p HEVC TearsofSteel movie - EVR - OpenCL ON - Benchmarking mode of DXVA Checker
LAV x86 (June 2014): 59/90/229
LAV x86 (Aug 2014): 64/102/239
GT440 OpenCL x86: 152/224/264 CPU usage: 86%
5750 OpenCL x86: 71/109/187 CPU usage: 28% GPU usage max clock@90%
LAV x64 (June 2014): 105/180/307
LAV x64 (Aug 2014): 131/224/360
4K HEVC Ducks Take off - Same configuration as above
LAV x86 (Aug 2014): 6/13/15 CPU 97%
5750 OpenCL x86 4/24/29 CPU 68% GPU 62%
LAV x64 (Aug 2014): 11/18/21 CPU 97%
My comments:
1) We should pay more attention when the developer says that OpenCL decoder is optimized for AMD cards and for each card different resolution could be supported.
2) The OpenCL decoder clearly drops a lot the CPU usage of Core i5 when decoding 1080p clips with 5750, but is not optimized for 5750 & 4K clips.
The CPU usage goes high as the GPU usage drops on 4K.
3) LAV Aug seems more optimized for 1080p HEVC clips than 4K and it's definitely faster than June.
4) Min values still go even to 0 fps (just for an instant) when using AMD card (5750)
I have tested the OpenCL decoder on my i7 3960X + R9 290X. The results are as follows:
4K Elysium with VMR-9 renderless
LAV x64
Avg FPS: 92.334fps (Min-Max: 56-175fps)
Avg CPU Usage: 70% (Min-Max: 58-87%)
Avg GPU Usage: 0% (Min-Max: 0-0%)
OpenCL disable GPU
Avg FPS: 153.480fps (Min-Max: 103-215fps)
Avg CPU Usage: 77% (Min-Max: 67-88%)
Avg GPU Usage: 0% (Min-Max: 0-0%)
OpenCL enable GPU
Avg FPS: 164.095fps (Min-Max: 105-250fps)
Avg CPU Usage: 57% (Min-Max: 16-72%)
Avg GPU Usage: 44% (Min-Max: 0-89%)
Looks like the GPU does take effect and the speed is higher.
NikosD
9th August 2014, 19:27
You must test EVR in order to fully accelerate OpenCL decoding.
Renderless modes (VMR/EVR) are used mainly to see the pure decoding performance, without serious affect of GPU.
But regarding OpenCL GPU performance we want the opposite!
Full involvement of the GPU.
wxhyn
10th August 2014, 06:12
I tried OpenCL decoder with EVR renderless. Very weird, it frequently jammed while testing and the speed is not as fast as VMR for both GPU disable and enable, sometimes only 1 or 2 threads are used, clearly not all the speed potential is unleashed. Seems there's confliction between the decoder and EVR. For realtime playback, both EVR and VMR works fine.
I have collected the CPU and GPU usage while playing back the 4K Elysium.
OpenCL enable GPU
Frame Rate: Avg: 24 Min: 24 Max 24
CPU Usage: Avg: 04 Min: 00 Max 07
GPU Usage: Avg: 40 Min: 00 Max 100
OpenCL disable GPU
Frame Rate: Avg: 24 Min: 24 Max 24
CPU Usage: Avg: 07 Min: 00 Max 13
GPU Usage: Avg: 06 Min: 00 Max 100
LAV x64
Frame Rate: Avg: 24 Min: 24 Max 24
CPU Usage: Avg: 11 Min: 02 Max 26
GPU Usage: Avg: 04 Min: 00 Max 100
The OpenCL has the lowest CPU usage figure.
Asmodian
10th August 2014, 06:42
from which file did you like to know that?
Any 4K HEVC file really, I am simply curious if the OpenCL decoder is able to avoid the odd stalls/interop on AMD GPUs.
Looking at wxhyn's benchmarks it might, at least on new AMD cards.
@wxhyn, are you using an HEVC file, I think 4K Elysium is actually AVC isn't it?
NikosD
10th August 2014, 07:58
I tried OpenCL decoder with EVR renderless
You have tested everything besides what's most common/useful :)
Benchmark with EVR and give us your results with OpenCL ON. (OpenCL OFF is not working OK with EVR)
NikosD
10th August 2014, 11:02
Build the performance test graph with LAV Video, and then manually edit the graph to swap the decoder.
It worked.
But it makes me wonder, if you knew the method to completely isolate the GPU decoding (OpenCL decoding) by using the Null renderer, why did you ask for someone to test the OpenCL's decoder pure CPU performance by disabling GPU (OpenCL) when you have already done that ? ( by testing it with Null Renderer - which is the same)
What did you want to see ?
My results with GSN using Null renderer
(pure CPU performance without involving GPU at all)
1080p Tears of Steel (Avg fps)
LAV x86 (Aug): 166,3 fps
OpenCL x86: 341,8 fps
LAV x64 (Aug): 370,3 fps
LAX x64 (Aug): 237,5 fps using EVR with DXVA Checker
4K Duck Take Off (Avg fps)
LAV x86 (Aug): 19,7 fps
OpenCL x86: 45,4 fps
LAV x64 (Aug): 27,3 fps
JFYI, the next DXVA Checker v3.1.0 that I have tried in beta, has the exact same results with GSN using "DXVA decoding" which is the new method for testing pure CPU decoding performance
(It replaces the VMR/EVR renderless mode)
@clsid
As I've written before, I've seen lots of times a huge drop in performance of LAV decoder when using EVR renderer (real world use/test) vs null renderer.
I mean all the decoders have a performance hit using EVR renderer compared to a null renderer, but for LAV is there something to be optimized better for real world use of actual renderers like EVR ?
wxhyn
10th August 2014, 11:48
Any 4K HEVC file really, I am simply curious if the OpenCL decoder is able to avoid the odd stalls/interop on AMD GPUs.
Looking at wxhyn's benchmarks it might, at least on new AMD cards.
@wxhyn, are you using an HEVC file, I think 4K Elysium is actually AVC isn't it?
The 4K Elysium I used is HEVC encoded by NGCodec. From my observation, there are problems when do performance test with EVR renderless. Playback looks normal no matter with VMR or EVR.
wxhyn
10th August 2014, 12:35
You have tested everything besides what's most common/useful :)
Benchmark with EVR and give us your results with OpenCL ON. (OpenCL OFF is not working OK with EVR)
OpenCL enable GPU
Avg FPS: 127.515fps (Min-Max: 70-154fps)
Avg CPU Usage: 43% (Min-Max: 19-68%)
Avg GPU Usage: 59% (Min-Max: 0-95%)
LAV x64
Avg FPS: 107.338fps (Min-Max: 64-189fps)
Avg CPU Usage: 71% (Min-Max: 55-86%)
Avg GPU Usage: 0% (Min-Max: 0-0%)
The OpenCL performance is lower with EVR than VMR.
NikosD
10th August 2014, 12:47
R9 290X seems underutilised even in 4K and even fed up by i7-3960X.
And the fps are lower than using CPU alone.
The behaviour is like 5750...
I think it shouldn't, it should be more optimized for 4K.
Could you try the 1080p "Tears of steel" HEVC movie in EVR benchmarking mode ?
wxhyn
10th August 2014, 14:23
R9 290X seems underutilised even in 4K and even fed up by i7-3960X.
And the fps are lower than using CPU alone.
The behaviour is like 5750...
I think it shouldn't, it should be more optimized for 4K.
Could you try the 1080p "Tears of steel" HEVC movie in EVR benchmarking mode ?
Sorry, I don't have the 1080p tears of steel video clip. I think it's too early to draw a conclusion that R9 290X is underutilized in 4K. I observed that the OpenCL decoder with GPU enable/disable has almost the same performance with each other, i.e. the slow decoding compare to using VMR is not only with GPU only, but also with pure CPU decoding. Seems there's a wall that the speed can't get over. Another very interesting result is I tested the speed in EVR with several other 4K clips. All speeds are almost the same. I guess maybe Strongene may not deal with the EVR correctly. It looks like waiting for something finished before it can go any further.
NikosD
10th August 2014, 15:49
Sorry, I don't have the 1080p tears of steel video clip.
When I said movie, I didn't mean commercial movie.
Tears of Steel is a free to download movie.
You can grab it from here in HEVC 1080p format.
http://trailers.divx.com/hevc/TearsOfSteelFull12min_1080p_24fps_27qp_1474kbps_GPSNR_42.29_HM11.mkv
I think it's too early to draw a conclusion that R9 290X is underutilized in 4K. I observed that the OpenCL decoder with GPU enable/disable has almost the same performance with each other, i.e. the slow decoding compare to using VMR is not only with GPU only, but also with pure CPU decoding.
I guess maybe Strongene may not deal with the EVR correctly. It looks like waiting for something finished before it can go any further.
You have to forget VMR, it's meaningless to use it.
It's a very old renderer from Windows XP time.
Since Vista, only EVR is used.
You can't DXVA HW accelerate anything in Vista and above using VMR, you have to use EVR which is a very light but full of capabilities renderer, especially the EVR-CP (custom presenter)
DXVA Checker v3.1.0 is out officially, with a lot of useful changes.
There is no VMR/EVR and VMR/EVR Renderless anymore.
Check the http://bluesky23.yu-nagi.com/en/DXVAChecker.html page for changes.
wxhyn
11th August 2014, 12:20
I have collected the results with 1080p tears of steel with DXVA decoding and DXVA processing modes
Strongene OpenCL DXVA decoding
FPS: 559.518 [350-704] fps
CPU Usage: 45 [23-57] %
GPU Usage: 68 [0-95] %
Strongene OpenCL DXVA processing
FPS: 189.263 [125-191] fps
CPU Usage: 14 [8-25] %
GPU Usage: 73 [0-100] %
Strongene CPU version DXVA decoding
FPS: 305.755 [172-442] fps
CPU Usage: 25 [17-29] %
GPU Usage: 0 [0-30] %
Strongene CPU version DXVA processing
FPS: 187.171 [126-200] fps
CPU Usage: 15 [6-20] %
GPU Usage: 55 [0-100] %
LAV x64 DXVA decoding DXVA decoding
FPS: 349.292 [226-507] fps
CPU Usage: 42 [30-51] %
GPU Usage: 0 [0-0] %
LAV x64 DXVA decoding DXVA processing
FPS: 303.971 [215-349] fps
CPU Usage: 37 [9-54] %
GPU Usage: 87 [0-100] %
The results looks normal when using DXVA decoding mode. But when using DXVA processing mode, the Strongene decoders show weird results which the highest FPS can't get over 200 FPS no matter how fast the decoding is. I have tested several other 1080p clips, all of the average FPS are close to 190 FPS. Seems it's a bug of the decoders (both the CPU and OpenCL version). The LAV decoder looks normal.
NikosD
11th August 2014, 12:51
The DXVA processing method has changed and uses native scaling by default pushing the GPU usage a lot, even a pure CPU decoder like LAV has huge GPU usage.
Change the scaling manually by putting 640x480 resolution and try the processing method again.
wxhyn
11th August 2014, 13:05
I tested another 1080p video clip. I don't know the name, it's a Korea singing group, singing and dancing.
Strongene OpenCL DXVA decoding
FPS: 204.664 [151-315] fps
CPU Usage: 33 [20-43] %
GPU Usage: 28 [0-83] %
Strongene OpenCL DXVA processing
FPS: 182.384 [122-200] fps
CPU Usage: 30 [10-38] %
GPU Usage: 70 [0-100] %
Strongene CPU version DXVA decoding
FPS: 143.681 [110-245] fps
CPU Usage: 21 [13-25] %
GPU Usage: 0 [0-0] %
Strongene CPU version DXVA processing
FPS: 124.531 [91-187] fps
CPU Usage: 17 [11-21] %
GPU Usage: 40 [0-100] %
LAV x64 DXVA decoding DXVA decoding
FPS: 161.381 [130-288] fps
CPU Usage: 43 [21-59] %
GPU Usage: 0 [0-6] %
LAV x64 DXVA decoding DXVA processing
FPS: 147.937 [129-296] fps
CPU Usage: 44 [30-59] %
GPU Usage: 51 [0-100] %
This time the results looks normal. When the pure decoding speed not exceeds 200FPS too much. The processing results are very close to decoding results. Hope Strongene can fix this bug very soon.
NikosD
11th August 2014, 13:16
I don't think there is a bug in Strongene's decoder or probably an inefficency of the benchmarking tool.
In your OpenCL version the maximum fps are limited by the GPU.
The CPU version is not existant because you disable the driver, so the results can't be accurate and predictable.
Try with GraphStudioNext and null renderer for pure CPU results of Strongene's decoder.
wxhyn
11th August 2014, 13:52
I don't think there is a bug in Strongene's decoder or probably an inefficency of the benchmarking tool.
In your OpenCL version the maximum fps are limited by the GPU.
The CPU version is not existant because you disable the driver, so the results can't be accurate and predictable.
Try with GraphStudioNext and null renderer for pure CPU results of Strongene's decoder.
No, there is a CPU version decoder on Strongene's website. The one I used is the CPU version. So GPU is enabled when I tested.
NikosD
11th August 2014, 13:57
That decoder is slower than using Strongene's OpenCL as a CPU decoder and probably has more bugs.
Try GSN to see the difference.
wxhyn
11th August 2014, 15:18
That decoder is slower than using Strongene's OpenCL as a CPU decoder and probably has more bugs.
Try GSN to see the difference.
What does GSN refer to?
wxhyn
11th August 2014, 15:19
What does GSN refer to?
I guess it stands for GraphStudioNext.
NikosD
13th August 2014, 12:36
I built a testing collection of artificially high bandwidth H.264 and H.265 files using x264 and x265 encoders respectively.
I call them monster files.
It's a collection of 1080p and 3840x2160 (4K) files with bandwidths of 50Mbps, 100Mbps, 200Mbps. ..up to 600Mbps for both resolutions and both codecs.
They are 28 files in total with a footprint on disk of about 15GB.
It's not possible to upload all of them, but if someone has a specific need for a file, we'll see what we'll do.
nevcairiel
18th August 2014, 09:19
Since you guys like benchmarking software decoder as well now, here is a nightly build of LAV with HEVC improvements merged from the OpenHEVC project. On a 4K sample I tried, it increased performance over 100% (37 -> 85) in the 64-bit build.
The performance increase will vary greatly on the features used in the encode, ie. if SAO is used or not, etc. (for reference, the 1080p Tears Of Steel encode from DivX does NOT use SAO, so the improvement will be smaller).
32-bit: http://files.1f0.de/lavf/LAVFilters-0.62-14-g1c3f78b.zip
64-bit: http://files.1f0.de/lavf/LAVFilters-0.62-14-g1c3f78b-x64.zip
Also note that extremely high bitrates distort the result, as most of the time is then spent in CABAC decoding (ie. bitstream parsing), and not image reconstruction.
Real-world bitrates are the most sensible to test, as thats what will matter the most in the end as well.
NikosD
18th August 2014, 17:19
DXVA Checker v3.1.0 - DXVA decoding - Signature system
1080p - (1920_ProRes_2mbps.mkv)
LAV x64 (.14) 314/406/455 CPU: 80%
LAV x64 (.13) 113/258/343 CPU: 67%
LAV x86 (.14) 134/176/256 CPU: 92%
LAV x86 (.13) 90/148/192 CPU: 87%
3840 x 2160 - Ducks Take Off
LAV x64 (.14) 53/57/62 CPU: 78%
LAV x64 (.13) 21/27/33 CPU: 77%
LAV x86 (.14) 27/34/38 CPU: 87%
LAV x86 (.13) 14/20/22 CPU: 82%
Really huge performance advantage of the new version (.14) over the old one (.13)
LAX x86 .14 is even faster than LAV x64 .13 in 4K decoding (!!)
I think I was right talking about vectorized code, or not ? ;)
foxyshadis
18th August 2014, 23:52
Wow! The x64 version is now faster than OpenCL! That's incredible.
nevcairiel
19th August 2014, 08:00
Its fascinating how much SSE2 IDCT and SSSE3 SAO can do for performance, huh.
NikosD
19th August 2014, 12:43
Even with my Core2Duo (Wolfdale with SSE4.1 support) I see huge gains on 4K clips between the two latest versions:
60% for x86
70% for x64.
If you see the above tables, Haswell has even better results:
70% for x86
111% for x64.
Now the question is:
Still no AVX2 optimizations ? Why ?
huhn
19th August 2014, 14:28
Still no AVX2 optimizations ? Why ?
this is most likely one of the last steps to add this, most CPU doesn't support this and using AVX1/2 doesn't mean you get an huge performance improvement. so they use the common ones first or those that give an good improvement. as you can see they get huge improvements with this so they are totally right.
kasper93
19th August 2014, 15:35
I did quick test on 4096x1720 24fps with 2157Kbps sample. And compared to Lentoid decoder.
LAV x86: 29.3461 FPS
LAV x64: 89.4249 FPS
Lentoid: 81.1206 FPS
Not bad, but x86 is really lagging behind which can be a problem for madVR users.
huhn
19th August 2014, 16:12
I did quick test on 4096x1720 24fps with 2157Kbps sample. And compared to Lentoid decoder.
LAV x86: 29.3461 FPS
LAV x64: 89.4249 FPS
Lentoid: 81.1206 FPS
Not bad, but x86 is really lagging behind which can be a problem for madVR users.
I don't see a huge problem with madVR and 32 bit h265 is young and a 64 bit version of madVR will come sooner or later so it's fine for the time been.
NikosD
19th August 2014, 17:07
this is most likely one of the last steps to add this, most CPU doesn't support this and using AVX1/2 doesn't mean you get an huge performance improvement. so they use the common ones first or those that give an good improvement. as you can see they get huge improvements with this so they are totally right.
I think AVX can't help here because we are talking about integers mostly.
But AVX2, although only for Haswell and better, could be a lot faster than SSEx with 256bit registers vs 128bit registers.
And Haswell got an implementation of AVX2 right from the beginning, not like SSE2 and Pentium 4.
A lot of programs optimised for SIMD get AVX2 optimizations now, they don't have to wait for other CPUs to come with AVX2 support.
clsid was saying that in the future (!) ffmpeg HEVC decoder will eventually reach Strongene's speed.
Well, after my pressing about vectorised code missing from LAV's ffmpeg HEVC decoder, someone looked at OpenHEVC that already had vectorised code.
And the funny thing is that OpenHEVC is a fork of FFMpeg.
So eventually, the future is now.
nevcairiel
19th August 2014, 17:16
There is a whole bunch of reasons why there is no avx2 yet, and many rather technical that going into them is not worth it. However, there will also not be such a huge boost as you might think, the biggest boost was from the SSE stuff. Its not going to be twice as fast just because its in theory double the register size.
The next step will have to be to rewrite these optimizations into proper ASM instead of compiler intrinsics, and contribute them to FFmpeg proper, only after that avx2 is likely to appear. And that's neither an easy nor a fast task. Optimizing algorithms like this is very specialized knowledge.
NikosD
19th August 2014, 17:24
ASM itself is a very hard thing on its own.
No doubt about that.
But I think the developers involved are special too.
I would definitely like to read the technical reasons of not having AVX2 yet (besides what is already mentioned) and the technical reasons why the gain wouldn't be so much compared to SSEx's boost.
NikosD
23rd August 2014, 08:36
I'll answer myself, my previous question based on the main developer of x264 project - dark_shikari - words:
Here is an introduction to AVX2 optimizations in x264 project (about 1 month before the actual release of Haswell)
http://www.scribd.com/doc/137419114/Introduction-to-AVX2-optimizations-in-x264
Here is a comment of Dark_Shikari about autovectorization of modern compilers regarding SIMD instructions
https://news.ycombinator.com/item?id=5603406
...and finally here is the actual gain of Haswell vs Ivy measured by him (15%-20% per clock and from that figure, only 5% due to AVX2 optimizations only) on June 2013.
I don't know if further optimizations regarding AVX2 code have been made in x264 project, since June 2013
http://forum.doom9.org/showthread.php?p=1632275#post1632275
I don't know if x265 can benefit more of AVX2 and if H.265 decoding is something a lot different and can benefit even more.
NikosD
31st August 2014, 09:49
It seems that in 2 days from now (2nd of September), a new card from AMD will be reviewed by technical sites.
I read some interesting info regarding the new multimedia engine of Tonga GPU (R9 285) which is considered as a new iteration of GCN architecture v1.2
"The GCN 1.2-based GPUs will also feature a new multimedia engine – which comprises of universal video decoder 6.0 (UVD 6) and video encoder engine 3.1 (VCE 3.1) technologies – as well as a new high-quality scaler for video. There is no word about support for ultra-high-definition (UVD) video codecs, such as H.265/HEVC or VP9, so it looks like the new GPUs will not support them."
Let's hope that finally, AMD will support in HW 4K H.264 and possibly a GPU assisted H.265 format
Yups
31st August 2014, 10:19
I'm looking forward to Skylake-S which can do H265 encoding in hardware.
NikosD
31st August 2014, 10:55
2016 ?
We are on August of 2014 !
Yups
31st August 2014, 10:59
Mid-2015 for Skylake-S. Q2 2015 in the last Roadmap.
nevcairiel
31st August 2014, 11:03
Broadwell isn't coming out before Q1 2015, if you still believe Skylake is up for anywhere in 2015..... ;)
NikosD
31st August 2014, 11:16
Broadwell for desktop is for Q2 2015.
No way Intel will release two different architectures at the same time !
GTPVHD
31st August 2014, 13:57
Nvidia Maxwell second generation GM204 will launch on September 19, time to see if it supports HEVC decoding.
NikosD
31st August 2014, 14:01
I guess not.
VP6 of 750 is very new and there was no time to add such a feature since then.
But is Nvidia is a big brand.
We'll see.
Yups
31st August 2014, 16:49
Broadwell isn't coming out before Q1 2015, if you still believe Skylake is up for anywhere in 2015..... ;)
Your are not well informed. Broadwell for desktop is practically nonexistent next year, it's only Broadwell-K with GT3e graphics. The mainstream successor of Haswell non-K will be Skylake-S next year. These CPUs are non-K SKUs with GT2 graphics. And yes Broadwell-K and Skylake-S are both supposed to come mid-2015. Many people don't understand that both are aimed for different segments which allows Intel to bring Broadwell and Skylake at the same time on desktop.
http://www.hardwareluxx.com/index.php?option=com_content&view=article&id=31758%20%20&catid=34&Itemid=99
http://www.kitguru.net/components/cpu/anton-shilov/intel-to-release-first-skylake-microprocessors-in-q2-2015-says-report/
http://www.cpu-world.com/news_2014/2014052601_Intel_Skylake_desktop_CPUs_to_launch_in_Q2_2015.html
NikosD
31st August 2014, 18:00
Nice!
We'll see.
GTPVHD
1st September 2014, 22:58
http://game24.nvidia.com/
GM204 confirmed launching on the 19th.
NikosD
2nd September 2014, 19:23
At last!
The first AMD card with 4K H.264 fixed-function HW decoding support.
"New video decode block. The R9 285 also includes a new video decoder block for full hardware decode of H.264 4K streams. H.264 base, main, and high profile up to 5.2 are all supported. That means the new block can handle decoding 4096×2304 streams at 60 fps. There’s also a new fixed function video transcoder unit (VCE) that supports full hardware encoding to H.264, but AMD didn’t provide additional details on features or capabilities that distinguish this new unit from previous generation hardware."
nevcairiel
2nd September 2014, 19:32
Didn't they claim support before? :d
NikosD
2nd September 2014, 19:39
My Radeon 5750 had 4K H.264 support since 2010! :D
But this time they admit that it's the first time they support 4K H.264 in fixed-function decoding and all of the review sites are saying the same thing.
Yups
9th September 2014, 19:14
Intel is talking about GPU accelerated HEVC decoding and encoding in Broadwell. I thought Skylake was supposed to bring HEVC GPU encoding support for the first time in an Intel product.
https://intel.activeevents.com/sf14/connect/fileDownload/session/F7EBB2F85164E31DF41CDD863B86286F/SF14_SPCS002_101f.pdf
Page 77 and 88.
NikosD
9th September 2014, 22:14
Thanks for the document.
Nothing new actually.
GPU accelerated decoding & encoding means hybrid approach, not HW accelerated in ASIC.
It means it uses the EUs and CPU - probably not at all the ASIC or just a little.
The approach could be similar to the existent HEVC_MAIN_VLD for Haswell if we had an appropriate decoder to check it out now before Broadwell (in case of course that HEVC_MAIN_VLD mode actually works)
You can clearly say that I'm right, looking at the page 88, where it says "supported in higher performance SKU", meaning is not ASIC accelerated.
We have to wait for Skylake or Nvidia/AMD, whoever gets it first.
NikosD
9th September 2014, 22:22
...and also it's only 4K30 for HEVC in higher performance SKU.
But my Core i7-4790 already does 4K60 in SW thanks to the assembly OpenHEVC project optimizations that Nevcairiel has embedded in LAV filters.
GTPVHD
10th September 2014, 07:56
http://www.anandtech.com/show/8510/idf-2014-intel-demonstrates-skylake-due-h22015
Intel confirmed Skylake will be out second half of 2015, it will probably have the power efficient, fixed function HEVC decoder unlike Broadwell's hybrid EU/CPU decoder.
NikosD
11th September 2014, 17:59
Latest UVD from AMD inside R9 285 is 3 times faster than VP5 and even faster than VP6 according to Anandtech in both 1080p & 4K H.264 files using Microsoft's decoder (LAV filters don't provide support yet)
Here is the article:
http://www.anandtech.com/show/8460/amd-radeon-r9-285-review/4
It looks like a 4K80-85fps decoder to me.
wanezhiling
12th September 2014, 00:22
impossible...impossible...impossible...
http://i.imgur.com/H0ctjK9.gif
GTPVHD
12th September 2014, 06:37
http://tpucdn.com/reviews/Sapphire/R9_285_Dual-X_OC/images/power_bluray.gif
AMD cards are a joke when it comes to power consumption. 51W for Blu-ray playback vs 6W for Maxwell.
huhn
12th September 2014, 09:04
if you have such a powerful GPU you normally have a good CPU too. and in this case it isn't rare you can spare power by using the software decoding. the higher powerstate of such a GPU can easily take more power as software decoding.
wanezhiling
12th September 2014, 09:25
http://tpucdn.com/reviews/Sapphire/R9_285_Dual-X_OC/images/power_bluray.gif
AMD cards are a joke when it comes to power consumption. 51W for Blu-ray playback vs 6W for Maxwell.
:D lol...No red in upper half part...
STaRGaZeR
12th September 2014, 10:07
Yep, they ramp up clocks and voltage when using DXVA (and with multi monitor too). Good thing that I use manual profiles in Afterburner :D
RainyDog
12th September 2014, 12:41
Power usage for video playback is easily the biggest issue I have with my R9 290. PC pulls nearly 200w using madVR and my preferred settings and it really bothers me.
If I switch from DXVA decoding to software with LAV and MPC-BE, total power consumption drops considerably. But is there any settings in madVR that require DXVA decoding? I' sure nnedi3 didn't seem to be working when I tried software last time.
nevcairiel
12th September 2014, 12:48
There is even features in madVR that do NOT work when using DXVA decoding, certainly none that require it to work - especially since not every video can be decoded through DXVA, that would be quite the silly limitation.
NikosD
15th September 2014, 11:36
I'm trying to collect HEVC decoders in SW, in order to start a new thread here and evaluate their performance.
Free/Open source/commercial, everything is accepted.
I was thinking to test basically two resolutions 1080p & 4K (UHD) in various clips for both of them, maybe a few more clips in 4K resolution which I think will be the target resolution of HEVC, especially after the official 4K BluRay specs.
I was thinking also to test them using various CPU architectures from Core2Duo up to Haswell Core i7.
The problem is that I didn't find a lot of decoders in the form that I could test them, e.g in binary form (not source) and in DS filter (not embedded in players or plug-ins etc.)
I have already found two.
LAV filters and Strongene's SW decoder.
Looking forward for your suggestions, if I can test at least 5 of them I could start a new thread.
huhn
15th September 2014, 12:00
I suggest to wait for about 6 month so the decoder can mature in the meanwhile.
HEVC decoding is simple not important right now.
NikosD
15th September 2014, 13:41
It's been already more than a year since the final official specification and about two years or more since the first H.265 encoders.
Yes, I agree that probably we have to wait even more than 6 months, about a year or more for the next year's 4K Bluray arrival which will be a boost for HEVC, but at that time we should have HW decoders too.
Anyway, if I manage to get a few SW HEVC decoders, I'll do it.
GTPVHD
16th September 2014, 13:33
http://videocardz.com/52362/only-at-vc-nvidia-geforce-gtx-980-final-specifications
GTX 980 and GTX 970, world's first HDMI 2.0 graphics cards!
NikosD
16th September 2014, 20:05
Impressive specs, but it's going to be slower than 780 Ti in several cases, which is something not easily understood and justified.
nevcairiel
17th September 2014, 16:23
I managed to benchmark the hybrid HEVC hardware decoder on a NVIDIA GTX 780, and its not very fast - but it does have low CPU usage.
1080p is at 110-120 fps (15% cpu at benchmark, 1-2% during playback at 24p), and 4K only at 30-35 fps (20% benchmark cpu usage).
For comparison, the 32-bit LAV software decoder achieved 140 and 30 fps on the two clips i tested (of course 100% CPU), 64-bit is of course much faster.
Of course the design is more likely targeted at playback and not at fast transcoding, but on any decent CPU the pure software decoders are just faster (but do use a bit more CPU during playback), especially if you can use 64-bit.
NikosD
17th September 2014, 16:36
If you are talking about Strongene's OpenCL decoder, I had exactly the same experience with very low CPU usage during playback, but slower benchmark performance than fast CPUs.
But IIRC x86 LAV decoder was slower than hybrid decoder.
By the way, are you interested in implementing hybrid DXVA HW H.265 decoder for Nvidia and Intel ?
I remember you saying that it was scheduled (by you) since the first beta drivers of Intel offering DXVA H.265 support.
nevcairiel
17th September 2014, 16:47
I'm not talking about OpenCL, but NVIDIAs own hybrid decoder included in the driver, the same thing thats exposed through DXVA2.
Its not available through any software that I know of so far though, so maybe mine will be the first.
NikosD
17th September 2014, 17:05
Nice!
The reason I asked you is because I thought that hybrid DXVA decoders would have the same experience/behavior like OpenCL decoders and you just confirm that!
Is there any VP5 assistance for H.265 or it's just SMXs and CPU ?
Because 780 has a lot of shaders and other Kepler GPUs will be a lot slower than that.
Have you tried Intel's DXVA H.265 decoder ?
nevcairiel
18th September 2014, 17:05
Its unclear which features it really uses of the GPU, and I've been busy implementing it instead of worrying about driver details.
In any case, I have both LAV's CUVID and DXVA2 working on NVIDIA, for some reason DXVA2 still is broken on Intel, more to do in the coming days to get to the bottom of this. QuickSync (or more specifically, Intel's MediaSDK) doesn't seem to expose HEVC support yet, as far as i can see.
Drivers only expose 8-bit modes so far.
JohnLai
18th September 2014, 17:24
It would be great if Intel GPU assisted HEVC DXVA can work considering that most ordinary users tend to use Intel Haswell CPU/IGPU these days......Too bad I heard Intel won't implement such DXVA workaround with past Ivy Bridge IGPU (particularly HD 2000 and 4000)
NikosD
18th September 2014, 18:02
Its unclear which features it really uses of the GPU, and I've been busy implementing it instead of worrying about driver details.
In any case, I have both LAV's CUVID and DXVA2 working on NVIDIA, for some reason DXVA2 still is broken on Intel, more to do in the coming days to get to the bottom of this. QuickSync (or more specifically, Intel's MediaSDK) doesn't seem to expose HEVC support yet, as far as i can see.
Drivers only expose 8-bit modes so far.
Thanks for the info.
It's bad for Intel to implement first officially a hybrid decoder, but give it broken.
Did you try latest official drivers v. 3907 ?
About Nvidia I guess they will not have implemented a hybrid decoder with fixed-function assistance.
I think for low power Kepler/Maxwell cards it will make a difference.
It's easy to check it out by using a monitor tool like GPU-Z or a gadget while playing/ benchmarking the hybrid decoder.
Nvidia offers a VPU load metric different than GPU load which shows exactly that - fixed function load.
You definitely don't have to look inside the driver, instead you look at the decoding itself.
nevcairiel
18th September 2014, 18:05
I didn't say theirs is broken, it may as well be my code not being finished yet, its hard to judge if you have no alternate reference implementation to test if the hardware/drivers work.
nevcairiel
19th September 2014, 08:49
Maxwell v2 has a hardware HEVC encoder, but still only the hybrid decoder. I'll be getting one of those cards, maybe the hybrid decoder got faster on that architecture at least.
NikosD
19th September 2014, 10:39
Anandtech says it uses only shaders and fixed-function HW for H.265 decoding.
Is it accurate ?
Is there an ETA for the release of your new decoder including HEVC hybrid ?
I would like to test it myself on a Kepler 740M card inside a girlfriend's laptop.
It would be better if I could test Intel's H.265 too :D
Thanks.
Yups
9th January 2015, 13:34
HEVC Decoding works on my Haswell iGPU with MPC 1.7.7 via DXVA2 by the way. CPU utilization much lower than before.
NikosD
9th January 2015, 14:36
There is a new thread opened by me for HEVC decoding.
There is a lot of discussion and figures for HEVC decoding.
Look at my signature.
NikosD
28th February 2015, 21:06
Now that all three companies (Intel, Nvidia, AMD) have updated their fixed-function video decoders, I wonder what is their performance in H.264 1080p clips.
If someone has an Nvidia 960 GTX or AMD R9 285 or IvyBridge or Broadwell iGPU, I would like to see their results using DXVA Checker x64 v3.3.2 in pure decode mode and LAV x64 v0.64 in DXVA native for the clips from the first post.
1 to 6 is here: ftp://helpedia.com/pub/multimedia/x264/testvideos/2011%20-%2002%20-%20H.264%20CPU%20DXVA%20codec%20comparison%20-%20Core2Duo%20vs%20UVD%202.2/
7 to 10 is here:ftp://helpedia.com/pub/multimedia/x264/testvideos/2012%20-%2001%20-%20QuickSync%20vs%20UVD%202.2%20vs%20VP4/
Now that all three companies (Intel, Nvidia, AMD) have updated their fixed-function video decoders, I wonder what is their performance in H.264 1080p clips.
If someone has an Nvidia 960 GTX or AMD R9 285 or Broadwell iGPU, I would like to see their results using DXVA Checker x64 v3.3.2 in pure decode mode and LAV x64 v0.64 in DXVA native for the clips with huge bitrate from 7 to 10 from the first post.
ftp://helpedia.com/pub/multimedia/x264/testvideos/2012%20-%2001%20-%20QuickSync%20vs%20UVD%202.2%20vs%20VP4/
Any of those clips would definitely give us an idea of H.264 decoding performance.
1. Twinpeaks-30fps
VP7 LAV NATIVE 475/475/475
2. Samsung-30fps
VP7 LAV NATIVE 293/347/404
3. Basket-60fps
VP7 LAV NATIVE 558/562/565
4. Girls-60fps
VP7 LAV NATIVE 534/540/548
5. Birds-60fps
VP7 LAV NATIVE 474/503/527
6. Cat-60fps
VP7 LAV NATIVE 510/518/524
7. Vortex-24fps
VP7 LAV NATIVE 169/171/170
8. Birds-24fps
VP7 LAV NATIVE 171/177/176
9. Ducks -30fps
VP7 LAV NATIVE 181/191/192
10. Crowd Run-25fps
VP7 LAV NATIVE 152/154/152
Just be aware of the Core/Memory clocks ;)
I set 'Repeat Count' to '8' because it's very slow in thr first benchmark
NikosD
6th March 2015, 20:28
Using DXVA Checker x64 v3.3.2 in pure decode mode and LAV Video x64 v0.64 in DXVA native mode, I did the tests again with modern and older ASICs for the whole collection of clips from 1 to 10.
The results of VP7 (Nvidia GTX 960) are from P.J which has a GPU with high clocks (GPU clock = 1.47GHz).
The results of Neet009 using a different GTX 960 are a little lower.
Results:
1. Twinpeaks-30fps
QS3 1070/1070/1070
CPU 4790 515/515/515
QS1 484/484/484
VP7 475/475/475
RX 470 414/419/414 [Playback] - New driver
RX 470 312/338/380 [Decode] - New driver
RX 470 288/308/326 [Playback]
RX 470 279/289/310 [Decode]
VP5 125/139/143
Intel GMA HD 85/128/142
VP4 80/84/86
UVD2.2 46/56/62
UVD+ 2/51/60
2. Samsung-30fps
QS3 767/808/872
VP7 293/347/404
QS1 293/340/376
CPU 4790 218/268/349
RX 470 218/265/328 [Playback] - New driver
RX 470 199/240/306 [Decode] - New driver
RX 470 168/209/259 [Playback]
RX 470 158/200/263 [Decode]
VP5 109/119/125
Intel GMA HD 51/74/103
VP4 34/55/82
UVD2.2 35/46/55
3. Basket-60fps
QS3 1105/1155/1184
RX 470 530/588/617 [Playback] - New driver
QS1 540/581/633
VP7 558/562/565
CPU 4790 471/539/610
RX 470 422/450/471 [Playback]
RX 470 345/362/432 [Decode] - New driver
RX 470 317/337/398 [Decode]
VP5 136/144/157
Intel GMA HD 75/103/134
VP4 71/82/104
UVD2.2 54/57/59
4. Girls-60fps
QS3 1023/1060/1091
RX 470 532/548/561 [Playback] - New driver
VP7 534/540/548
QS1 487/500/514
CPU 4790 411/427/441
RX 470 412/425/444 [Playback]
RX 470 340/346/440 [Decode] - New driver
RX 470 310/333/415 [Decode]
VP5 135/137/139
Intel GMA HD 90/105/112
VP4 74/76/79
UVD2.2 55/56/58
6. Cat-60fps
QS3 959/967/970
VP7 510/518/524
RX 470 438/471/489 [Playback] - New driver
QS1 404/424/435
CPU 4790 347/379/400
RX 470 344/368/383 [Playback]
RX 470 313/326/383 [Decode] - New driver
RX 470 285/299/328 [Decode]
VP5 131/137/143
Intel GMA HD 68/92/96
VP4 67/76/83
UVD2.2 48/52/55
UVD+ 0/53/56
7. Vortex-24fps
QS3 358/359/358
VP7 169/171/170
QS1 156/159/158
CPU 4790 97/113/119
RX 470 107/110/110 [Playback & Decode] - new driver
RX 470 83/85/85 [Playback & Decode]
VP5 72/73/74
Intel GMA HD 35/46/49
UVD2.2 0/26/58
UVD+ 0/25/29
VP4 19/22/24
8. Birds-24fps
QS3 351/360/358
VP7 171/177/176
QS1 151/160/161
RX 470 105/111/111 [Playback & Decode] - new driver
CPU 4790 100/110/113
RX 470 75/86/92 [Playback & Decode]
VP5 71/77/79
Intel GMA HD 35/42/47
UVD2.2 13/27/47
UVD+ 0/26/29
VP4 19/22/28
9. Ducks -30fps
QS3 413/413/413
VP8 249/258/249
VP7 181/191/192
QS1 168/183/183
RX 470 105/126/134 [Playback & Decode] - new driver
CPU 4790 115/125/134
RX 470 80/98/110 [Playback & Decode]
VP5 74/84/91
Intel GMA HD 25/48/58
UVD2.2 0/30/58
VP4 21/25/30
10. Crowd Run-25fps
QS3 328/330/328
VP7 152/154/152
QS1 143/145/143
RX 470 98/101/105 [Playback & Decode] - new driver
CPU 4790 87/98/102
RX 470 74/78/78 [Playback & Decode]
VP5 66/68/69
Intel GMA HD 17/35/43
UVD2.2 20/23/24
UVD+ 0/22/28
VP4 18/21/21
Comments:
CPU 4790 = Core i7 4790@3.8GHz using 16 threads (LAV properties) - Win 8.1 Pro x64
QS3 = Haswell HD 4600@1.5GHz drivers: 4080
QS1 = SandyBridge HD 2000@1.5GHz drivers: 4101
VP8 = Nvidia GTX 1060 (Zotac Mini) drivers: 368.95
VP7 = Nvidia GTX 960@1.47GHz drivers: 347.52
RX 470 = AMD Radeon RX 470 drivers: 17.2.1 / New driver >17.4
VP5 = Nvidia GT610@0.81GHz drivers: 347.52
VP4 = Nvidia GT440@0.82GHz drivers: 347.52
Intel GMA HD = Core i5 520M (Arrandale)@0.77GHz drivers: 3268
UVD2.2 = Radeon 5750@0.4GHz drivers: 14.12
UVD+ = Radeon 3650@0.72GHz drivers: 13.9
P.S
1) For Arrandale I used MS DS x64 decoder because LAV video is not compatible with Arrandale.
Arrandale uses GPU shaders a lot, but CPU ~2%
For clips 3,4,6 the GPU load was 100% with a GPU clock of max 777MHz
For clips 1,2,7,8,9,10 the GPU load was ~70% with a GPU clock of min 372MHz
The GMA HD looks like a rather hybrid decoder using both ASIC+GPU, than a pure ASIC decoder.
2) Looking forward for Radeon R9 285, IvyBridge, Broadwell results
3) For RX 470 it was used DXVA Checker v3.15, LAV filters 0.69, Win 10 x64 and Playback mode used a 1280x720 scaling.
leonccyiu
18th March 2015, 03:48
I hope I am not posting this in the wrong place
I recently purchased an intel pentium g3258 thinking the hd graphics should be able to decode 4k h264 video, however my cpu is being used.
Upon running DXVA Checker my suspicions were confirmed
H264_VLD_NoFGT: DXVA2/D3D11, SD / HD / FHD
I looked through earlier posts in the thread and noticed that QFHD used to be included but after a driver update isn't anymore.
Does anyone know more about this? Or if it's still possible to revert to an older driver?
Thanks
NikosD
18th March 2015, 07:50
H264_VLD_NoFGT: DXVA2/D3D11, SD / HD / FHD
Really ?
That's a very bad thing for Intel.
I remember when I had a Pentium G3420 installed that a beta driver disabled 4K H.264 for Pentium but that was a beta driver and then I switched to Core i7.
I don't believe that they did that on official drivers!
If I were you I would download and install latest driver v.4156 and complain to Intel forums if 4K is still disabled.
Here is the download link and the forum:
https://communities.intel.com/thread/61436
P.J
18th March 2015, 17:54
Strange, my N2830 can play 4K H.264
leonccyiu
19th March 2015, 02:56
Really ?
That's a very bad thing for Intel.
I remember when I had a Pentium G3420 installed that a beta driver disabled 4K H.264 for Pentium but that was a beta driver and then I switched to Core i7.
I don't believe that they did that on official drivers!
If I were you I would download and install latest driver v.4156 and complain to Intel forums if 4K is still disabled.
Here is the download link and the forum:
https://communities.intel.com/thread/61436
My Driver version is 15.33.32.64.4061
According to the Intel Driver update utility that is the latest version. I downloaded the driver in your link, but it said it wasn't compatible with my hardware.
I believe the beta drivers you installed for your Pentium G3420 is the time in which Intel decided to disable 4k h264 decoding on hd graphics.
Do you mind if I use your screenshot to complain?
and Do you know where I can complain to Intel?
Thanks
NikosD
19th March 2015, 06:50
My Driver version is 15.33.32.64.4061
According to the Intel Driver update utility that is the latest version. I downloaded the driver in your link, but it said it wasn't compatible with my hardware.
You are doing something wrong.
Your Pentium G3258 belongs to Haswell family.
First of all you should install new family drivers 15.36 and not 15.33.
15.33 is the old family which is compatible with Ivy and Haswell and 4061 are the latest, when 15.36 is the new family that supports Haswell and Broadwell and 4156 are the latest drivers.
Maybe you downloaded the wrong link 32bit vs 64bit.
If you can't install 4156 drivers, ask for help from the link above that I gave you.
Do you mind if I use your screenshot to complain?
and Do you know where I can complain to Intel?
Don't use something from me, take a screenshot of your system using DXVA checker and paste it in the link above that I gave you.
It's from Intel forums, I don't know any other way.
HarryM
26th June 2016, 21:34
You are doing something wrong.
Your Pentium G3258 belongs to Haswell family.
First of all you should install new family drivers 15.36 and not 15.33.
15.33 is the old family which is compatible with Ivy and Haswell and 4061 are the latest, when 15.36 is the new family that supports Haswell and Broadwell and 4156 are the latest drivers.
15.40 family is the newest.
NikosD
28th February 2017, 14:10
Added today at this post https://forum.doom9.org/showthread.php?p=1712350#post1712350 my new results for AMD RX 470 card.
The behavior was totally strange, as playback performance was better or even a lot better using playback mode (scale to 1280x720) than decode mode, but after clip 7 where the bitrate is huge, the playback performance was the same like pure decode performance.
For "normal" bitrate clips up to 30Mbps, the decoder is very fast but above 100Mbps the performance drops a lot, more than any other decoder.
Another strange thing is that 4K H.264 decoding performance looks exactly like 1080p.
Low bitrate clips at 4K resolution have almost same decoding speed like 1080p, but when the bitrate rises ~100Mbps, the decoding performance drops a lot.
The HW H.264 decoder of Polaris cards looks like it's not affected by higher resolution compared to high (>100Mbps) bitrate.
The last strange result is that clips from 7 to 9 show a 100% utilisation of 1 core (25% CPU usage of my quad core Core i5-2400) but only using "Decode" mode.
Playback mode is using CPU <5% like the decode and playback mode of all the other clips.
Even clip 10 has no problem at all in CPU usage in "Decode" and "Playback" mode.
What a strange HW decoder...
NikosD
16th July 2017, 10:45
New results for the new drivers (>17.4) of AMD RX 400/500 series that improved a lot the HW decoding performance of both H.264/H.265 clips
Results here:
https://forum.doom9.org/showthread.php?p=1712350#post1712350
vBulletin® v3.8.11, Copyright ©2000-2026, vBulletin Solutions Inc.