View Full Version : H.264 CPU/DXVA codec comparison - Core2Duo vs UVD 2.2
NikosD
12th February 2011, 11:22
UPDATE 22/4/2011: UVD+ (Radeon 3650 results added)
UPDATE 31/3/2011: CoreAVC 2.5.1 (CPU & DXVA results added)
This is my second post regarding codec performance/ benchmarking.
The first one is here:
http://forum.doom9.org/showthread.php?t=156660
This time DXVA and CPU codecs are included too.
All tests have been done on
Win 7 SP1 x64 - Core 2 Duo @ 2.83GHz - Radeon 5750 (UVD 2.2) - Catalyst 11.1a
Second system is (for DXVA only):
Win XP SP3 32bit - Core 2 Duo @ 2.83GHz - Radeon 3650 AGP (UVD+) - Catalyst 11.2
The benchmark tool is DXVAChecker v2.4.0 (32bit)
Home page: http://bluesky23.yu-nagi.com/en/
You can find all reference video files here:
ftp://helpedia.com/pub/multimedia/x264/testvideos/
1.Twinpeaks1080p30fps-27Mbps
2.Samsung.Demo.Oceanic.Life-1080p30fpsRef16-40Mbps
3.Basketball - 1088p60fpsRef8-10Mbps
4.Girls.YoonYoon-1080p60fpsRef5-21Mbps
5.Birds_1080p60fpsReF2-30Mbps
6.Cat-1080p60fpsRef4-25Mbps
Benchmark instructions are here:
http://forum.doom9.org/showthread.php?t=156660
Codecs included in comparison:
CoreAVC v2.5.1 - CPU & DXVA
CoreAVC v2.0 - CPU only
DiAVC v1.2.2 - CPU only
FFMpeg-mt v52.110.0 (rev3757) - CPU only
DivX H.264 v1.2.1 Build 9.0.1.21 - CPU & DXVA
FFDshow DXVA rev3757 - DXVA only
MPC-HC v1.5.1.2910 (32bit - standalone filter) - CPU & DXVA
Cyberlink PowerDVD 10 v1.0.2229 (latest as of 12th Feb) - CPU & DXVA
Microsoft DirectShow H.264 (built-in Win 7) - CPU & DXVA
Microsoft MediaFoundation H.264 (built-in Win 7) - CPU & DXVA
Four comments:
1) CoreAVC v2.5.1 CPU is the fastest codec. Second best is CoreAVC again -previous version v2.0
2) CoreAVC v2.5.1 DXVA has almost identical results with MPC-HC DXVA & FFDShow DXVA
3) Core2Duo@2.83 GHz is faster (with optimized codecs) than UVD 2.2 in H.264 decoding
4) UVD 2.2 is very close in MIN and AVG frame rate to UVD+ (in Radeon 3650) in supported video clips by UVD+ (BluRay spec only)
Results:
A. Twinpeaks-30fps
Codec Codec type Min/Avg/Max fps
1) CoreAVC v2.5.1 CPU 80/97/107
2) CoreAVC v2.0 CPU 75/93/101
3) FFMpeg-mt CPU 68/85/94
4) DiAVC CPU 64/74/79
5) Microsoft DS CPU 55/69/95
6) PowerDVD CPU 56/68/76
7) Microsoft MFT CPU 54/66/86
8) MPC-HC CPU 44/62/69
9) Microsoft DS DXVA 50/60/100
10) DivX DXVA 49/60/86
11) FFDShow DXVA 49/60/86
12) CoreAVC v2.5.1 DXVA 50/59/83
13) MPC-HC DXVA 50/59/83
14 ) PowerDVD DXVA 50/59/78
15) Microsoft MFT DXVA 50/59/76
DivX CPU ---
B. Samsung-30fps
1) CoreAVC v2.5.1 CPU 32/49/90
2) FFMpeg-mt CPU 32/49/89
3) DiAVC CPU 34/49/81
4) CoreAVC v2.0 CPU 32/48/96
5) DivX DXVA 37/46/79
6) Microsoft DS DXVA 32/46/80
7) MPC-HC DXVA 37/46/75
8) CoreAVC v2.5.1 DXVA 37/46/74
9) DivX CPU 31/46/86
10) PowerDVD DXVA 32/45/65
11) PowerDVD CPU 22/43/79
12) Microsoft DS CPU 23/40/78
13) MPC-HC CPU 19/28/60
Microsoft MFT CPU ---
Microsoft MFT DXVA ---
FFDShow DXVA ---
C. Basket-60fps
1) CoreAVC v2.5.1 CPU 71/89/110
2) CoreAVC v2.0 CPU 72/88/111
3) DiAVC CPU 75/84/103
4) FFMpeg-mt CPU 73/83/104
5) DivX CPU CPU 70/83/98
6) PowerDVD CPU 67/78/97
7) Microsoft DS CPU 43/60/85
8) PowerDVD DXVA 55/58/70
9) Microsoft DS DXVA 50/57/107
10) DivX DXVA 55/57/81
11) MPC-HC DXVA 55/57/79
12) CoreAVC v2.5.1 DXVA 54/57/77
13) FFDShow DXVA 52/57/76
14) MPC-HC CPU 40/48/68
Microsoft MFT CPU ---
Microsoft MFT DXVA ---
D. Girls-60fps
1) CoreAVC v2.0 CPU 58/73/94
2) CoreAVC v2.5.1 CPU 59/72/90
3) DivX CPU 62/69/79
4) DiAVC CPU 58/68/83
5) FFMpeg-mt CPU 58/65/82
6) PowerDVD CPU 52/62/82
7) MPC-HC DXVA 56/57/80
8) DivX DXVA 55/57/82
9) CoreAVC v2.5.1 DXVA 55/57/81
10) FFDShow DXVA 55/57/80
11) PowerDVD DXVA 55/57/79
12) Microsoft DS DXVA 43/57/82
13) Microsoft DS CPU 43/52/72
14) MPC-HC CPU 37/40/55
Microsoft MFT CPU ---
Microsoft MFT DXVA ---
E. Birds-60fps
1) Microsoft MFT DXVA 51/163/404
2) PowerDVD CPU 52/68/71
3) DiAVC CPU 54/63/69
4) Microsoft MFT CPU 39/61/93
5) DivX CPU 53/61/71
6) CoreAVC v2.5.1 CPU 53/59/68
7) CoreAVC v2.0 CPU 41/57/71
8) PowerDVD DXVA 52/56/68
9) DivX DXVA 52/55/77
10) CoreAVC v2.5.1 DXVA 52/55/75
11) FFDShow DXVA 52/55/74
12) Microsoft DS DXVA 51/55/89
13) MPC-HC DXVA 50/55/80
14) FFMpeg-mt CPU 48/54/67
15) Microsoft DS CPU 37/48/68
16) MPC-HC CPU 30/35/41
F. Cat-60fps
1) CoreAVC v2.5.1 CPU 66/70/76
2) CoreAVC v2.0 CPU 66/70/75
3) DiAVC CPU 64/70/74
4) FFMpeg-mt CPU 63/67/72
5) PowerDVD DXVA 52/57/69
6) FFDShow DXVA 54/56/76
7) CoreAVC v2.5.1 DXVA 48/56/85
DivX CPU ---
PowerDVD CPU ---
DivX DXVA ---
Microsoft DS DXVA ---
Microsoft MFT DXVA ---
MPC-HC DXVA ---
Microsoft MFT CPU ---
Microsoft DS CPU ---
MPC-HC CPU ---
Second system:
FFDShow DXVA rev3828
A. FFDShow DXVA 43/53/61
B. Corrupted image due to L5.1 (not supported by UVD/UVD+)
C. Corrupted image due to L5.1 (not supported by UVD/UVD+)
D. FFDShow DXVA 50/55/60
E. FFDShow DXVA 49/53/59
F. FFDShow DXVA 50/55/61
Feel free to add your comments/ results.
altruist
31st March 2011, 03:51
This is an excellent study. I can't believe no one else has thanked you for this.
I recently noticed CoreAVC now supports ATI hardware decoding through DXVA. Thought you might want to know if you don't already :)
NikosD
31st March 2011, 08:43
Thanks for your comments.
I know that CoreAVC 2.5.1 has DXVA support, but there is no demo version AFAIK.
When I get the new version, I will definitely try it.
UPDATE: Results for CoreAVC v2.5.1 added (CPU & DXVA)
bobdynlan
31st March 2011, 15:21
UPDATE 31/3/2011:
E. Birds-60fps
1) Microsoft MFT DXVA 51/163/404
Did you retest this? Does not seem like a valid result, maybe you should remove it from the top of the list.
CruNcher
31st March 2011, 16:18
According to that test the Samsung clip seems the most complex one even more complex then the 60 fps ones due to the high bitrate + cabac most likely and ref frames
are any of them sliced ?
Cyberlink seems also to have to fight with it quiet Hard and FFMPEG MT can even survive against CoreAVC and DiAVC on that one interesting.
You should add Arcsoft and Mainconcept to that list, many falsely belive DivX and Mainconcepts implementation are identical that is false though their are differences, especialy as Mainconcept is @ SDK 8.8 and DivX still somewhere @ the 8.5-8.7 codebase :)
Also Elecard released their H.264 DXVA implementation that seems also fast :)
Also please add on which Power Profile you tested on Win 7 :)
the latest PowerDVD decoder is also 1.0.0.2610
BetaBoy
31st March 2011, 18:40
Thanx for the comparison. I'd hold of on 'true' DXVA comp stats with CoreAVC 2.5.x as we are about to release more DXVA features in upcoming releases. We have not even begun optimizations for DXVA... and we already know of bottlenecks that should make it even better/faster.
On the DivX / Main concept diffs.... they still use the same 'cores' from what others have posted here on D9.
pirlouy
31st March 2011, 19:57
Not sure to understand.
From these stats, can we say that CPU is better than GPU (DXVA) for 24fps movie ??
neoufo51
31st March 2011, 21:42
Edit: Never mind
mark0077
31st March 2011, 21:55
Hi,
Is it worth adding scores for the new lav cuid decoder http://forum.doom9.org/showthread.php?t=160290
Excellent thread btw, very useful.
neoufo51
31st March 2011, 22:07
Hi,
Is it worth adding scores for the new lav cuid decoder http://forum.doom9.org/showthread.php?t=160290
Excellent thread btw, very useful.
He can't because that's only for Nvidia cards and he is doing this on ATI cards.
NikosD
1st April 2011, 07:43
Did you retest this? Does not seem like a valid result, maybe you should remove it from the top of the list.
All tests were done 3 times and the numbers are all true.
You can see an analysis and possible explanation of those strange figures here:
http://forum.doom9.org/showthread.php?t=156660&page=7
Check out the posts of a member named hwti and my comments
NikosD
1st April 2011, 07:58
According to that test the Samsung clip seems the most complex one even more complex then the 60 fps ones due to the high bitrate + cabac most likely and ref frames
The Samsung clip is the famous difficult Samsung clip with 16 ReFrames, huge bitrate etc, etc but both CPU and DXVA codecs manage to stay above the 30fps by 50% - 49fps on average
On the other hand, because of the double frame rate (60fps) needed by the other clips, none of the DXVA & CPU codecs manage to stay above the 60fps by 50% - which means 90fps on average
are any of them sliced ?
Sorry I don't get it.
Cyberlink seems also to have to fight with it quiet Hard and FFMPEG MT can even survive against CoreAVC and DiAVC on that one interesting.
True
You should add Arcsoft and Mainconcept to that list, many falsely belive DivX and Mainconcepts implementation are identical that is false though their are differences, especialy as Mainconcept is @ SDK 8.8 and DivX still somewhere @ the 8.5-8.7 codebase :)
Also Elecard released their H.264 DXVA implementation that seems also fast :)
I have some results of ArcSoft here:
http://forum.doom9.org/showthread.php?t=156660
Mainconcept and Elecard are not so popular. If you give me some links I'll try it :)
Also please add on which Power Profile you tested on Win 7 :)
It's High Performance, but what's the difference ? They are all tested under the same conditions.
the latest PowerDVD decoder is also 1.0.0.2610
Mine was the latest as of 12th February - date of my first post
kypec
1st April 2011, 08:02
@NikosD: thanks for your efforts put into this extensive test rounds. Could you please post the results in more flexible format, like Google spreadsheet perhaps? I think that would provide better edit options for you and better sorting options for viewers as well.
NikosD
1st April 2011, 08:25
Not sure to understand.
From these stats, can we say that CPU is better than GPU (DXVA) for 24fps movie ??
Not better, faster.
A Core2Duo@2.83GHz supported by optimized and multithreaded codecs (like CoreAVC, FFmpeg-mt, DiAVC etc) is faster than UVD 2.2.
So, a modern Core i7 with 4 or 6 cores would be a lot, lot faster than UVD or VPx (Nvidia)
But as long as the UVD or VPx (Nvidia) or Intel's hardware solution, manage to play the clips with a minimum frame rate above the frame rate of x1 - say 24fps or 30fps or 60fps - then it's working.
The main reason of using DXVA and dedicated hardware for decoding video formats is power (laptops) - because Core2Duo consumes 10 times more power than UVD.
Of course, if you have a slow CPU then speed does matter, too.
And of course, during playback on the dedicated fixed function hardware inside the GPU (UVD, VPx etc), the CPU is free of doing other things, because the CPU utilization is <5%.
So, with a dedicated hardware in GPU and DXVA, you have an extra dedicated extremely low power consuming processor in your system capable of decoding several video formats like H.264, VC-1, MPEG2, WMV, MPEG4 ASP(DivX, Xvid) besides your CPU.
pirlouy
1st April 2011, 12:10
Ok. For me DXVA is dangerous, because too much dependant of Nvidia or ATI or Intel development. And I rather trust ffmpeg coding than those 3.
I've tried DXVA (through MPC-HC, ffdshow and another decoder), but was not really convinced, because it was sometimes jerky.
Like I don't use laptop, power is not something to consider (even for Earth, since I don't watch HD videos all day). I prefer using GPU to upscale only (madVR renderer).
Like I don't do any post-processing stuff, and from your results, it confirms that it's better (in my case) to use CPU to decode.
Thanks for your tests.
NikosD
1st April 2011, 12:23
I forgot to say that on ATI recent hardware with UVD 2.x and above, you can do postprocessing on driver's level within the Catalyst suite during playback by DXVA decoding, using GPU shaders (not used by DXVA) with 0% CPU utilization for all video formats supported by UVD.
Even if you don't use post-proc or a laptop, you pay the bill for the power you consume :)
pirlouy
1st April 2011, 13:26
I've done a new test with DXVA. If I have a bit rate too high (mt2s with 20 Mbps for example), I have a lot of dropped frames (jerky videos). Yet I have quite the same GPU than you. Strange... I have catalyst from February (preview 2). But I won't search more. Everything is ok with CPU.
ps: Even if I'd watch 100 HD videos in a year, it would be ridiculous in terme of difference of power. 2€ for me and that won't do a difference for Global warming. :-)
Nikos,all the samples are offline.Can you re-up them please ?
NikosD
1st April 2011, 16:22
None of them is offline.
I think you do something wrong.
NikosD
1st April 2011, 16:28
ps: Even if I'd watch 100 HD videos in a year, it would be ridiculous in terme of difference of power. 2€ for me and that won't do a difference for Global warming. :-)
I'm not trying to convince you on anything, but you have a dedicated co-processor in your system designed to do one thing better than anyone else and you prefer to give that thing to your main processor who has a lot more to do and it is definitely not designed to do that job.
Your choice.
P.S And of course video acceleration is all over Internet because of Flash videos and the new version 10.2 which uses hardware acceleration for H.264 HD videos with minimum CPU utilization.
The same goes for HTML5, too.
CruNcher
1st April 2011, 19:15
That 4 Girls clip is also nice some peaks their really kill VP2 @ least with 95% DSP utilization :D its nicely @ the Edge (slightly over) of what VP2 is capable with VMR7 you don't see the frame drops as it trys to hold it stable and jitters with 10ms and plays @ 50fps @ VMR9 you see every drop nice to test the max edge of VP2, here for that clip on VP2 i would definitely switch automatically to Software Playback as with CUDA La CUVID/CoreAVC it jitters even more 24ms (no surprise) :)
Also a nice clip to test Intels HD2000/3000 Decoder with and compare vs VPx and UVD :)
Though its interesting comparing that to the Sony Playstation Net Wipeout 60 fps clip that plays super fluid and seems to perfectly fit into the VP2 specs, so i wonder what the Encoder of that won leaving specs in terms of Visual Quality i guess not really much, but that's sadly with x264 encoders become standard thinking (encode as complex as possible don't care about specs and force Hardware Manufactures to become better, even if they dont want to and are happy with being able to play Blu-Rays ;) )
NikosD
1st April 2011, 20:01
I'm looking forward for VP4 and Intel benchmark results, too.
Where are you guys ? :)
CruNcher
1st April 2011, 20:20
There is no real Benchmark for this Full Hardware Playback/ Live PostPro Framework yet :)
Cyberlink DXVA (VP2) + VMR9 Renderless + Sharpen Complex 2 (simple) + High Resolution Timing
http://img863.imageshack.us/img863/5233/gpulivesharpencyberlink.png
None of them is offline.
I think you do something wrong.
Correct ! :D
nevcairiel
1st April 2011, 20:48
I'm looking forward for VP4 and Intel benchmark results, too.
Where are you guys ? :)
Due to stupid HotFiles limitations, i only have some clips so far.
This is VP4 on the 270.51 driver (GTX 570 - but all VP4 chips should run the same speed, its not dependent on the clock domain of the main GPU)
CoreAVC 2.5.1 crashed the driver in DXVA mode, so only CUDA tested.
1. Twinpeaks
MS MFT DXVA - 79/88/104
MS DS DXVA - 81/84/86
MPC-HC DXVA - 81/84/86
CoreAVC 2.5.1 CUDA - 67/70/72
LAV CUVID - 78/81/84
3. Basketball
MS MFT DXVA - ---
MS DS DXVA - 74/82/103
MPC-HC DXVA - 69/82/99
CoreAVC 2.5.1 CUDA - 75/84/106
LAV CUVID - 71/84/106
5. Birds
MS MFT DXVA - 62/75/87
MS DS DXVA - 70/77/84
MPC-HC DXVA - 70/77/85
CoreAVC 2.5.1 CUDA - 55/61/67
LAV CUVID - 68/77/85
Let me take the opportunity to question the results of your MS MFT DXVA test on the Birds sample, its just not in line with any other results.
I think its also interesting that my totally unoptimized decoder (in a unoptimized debug build, too!) is faster then CoreAVC CUDA in some samples, although they should be using the same decode engine.
Oh, and just for fun, some CPU decoding values using CoreAVC 2.5.1 and ffdshow (using ffmpeg-mt)
This is on a Core i7 2600K in stock (turbo) speeds.
Of course not comparable to your CPU measurements.
Twin Peaks:
CoreAVC - 520 fps
ffdshow - 476 fps
I think the twin peaks sample gives false values because its completly decoded so fast that the fps calculation is probably wrong.
Basketball:
CoreAVC - 459 fps
ffdshow - 473 fps
Birds
CoreAVC - 268 fps
ffdshow - 265 fps
pirlouy
2nd April 2011, 01:04
I'm not trying to convince you on anything, but you have a dedicated co-processor in your system designed to do one thing better than anyone else and you prefer to give that thing to your main processor who has a lot more to do and it is definitely not designed to do that job.
I disagree. :-)
CPU can do everything and all statistics in this thread shows the CPU suffers less than GPU.
It is confirmed in my case. My E8400 decodes 20Mbs without problem (30% CPU) whereas my HD5770 renders jerky videos. Maybe it's a driver problem, but it's a problem I have not using CPU. And I can't use "hardware" decoders, i don't think there are for ATI...
For HTML5 and Flash, it's different, CPU is also used by javascript, css and all other stuff, so it can be helpful to have GPU.
I hope I won't piss you off, but I was not sure I read results correctly, but I think I was correct. Especially with new CPU, if you don't do a lot of post-processing, that's better to use CPU than DXVA.
NikosD
2nd April 2011, 08:50
Let me take the opportunity to question the results of your MS MFT DXVA test on the Birds sample, its just not in line with any other results.
You can see an analysis and possible explanation of those strange figures here:
http://forum.doom9.org/showthread.php?t=156660&page=7
Check out the posts of a member named hwti and my comments
NikosD
2nd April 2011, 10:17
Oh, and just for fun, some CPU decoding values using CoreAVC 2.5.1 and ffdshow (using ffmpeg-mt)
This is on a Core i7 2600K in stock (turbo) speeds.
It would be extremely interesting to post some benchmark results of Intel Clear Video HD hardware.
I've seen really big numbers using Quick Sync in decoding - encoding (transcoding) applications.
nevcairiel
2nd April 2011, 10:34
I have a P67 board, i cannot use the integrated GPU.
In any case, Quick Sync is only for encoding, the decode engine is probably not significantly faster then previous generations of ClearVideo HD.
As the encoding is done in the GPU, decoding with the CPU for re-encoding tasks is probably better. (260 fps decoding, hooray)
NikosD
2nd April 2011, 11:30
Well according to Intel and several hardware review sites Quick Sync is not only for encoding but for decoding, too.
Because, as I wrote to my previous post, transcoding is a two phase task. Decode and encode.
So if you want the fastest transcoding engine possible (like Quick Sync), you have to optimize both parts.
That's why Intel boosted decode performance a lot in Sandy Bridge processors and built for the first time a fixed-function hardware encoder.
The previous version of Clear Video decoding hardware is a lot weaker than Sandy Bridge, I don't know the difference.
More details here:
http://www.anandtech.com/show/4083/the-sandy-bridge-review-intel-core-i7-2600k-i5-2500k-core-i3-2100-tested/8
As for your second thought that maybe CPU decoding is faster than hardware decoding, it seems that it's definitely not true for Quick Sync according to here:
http://www.tomshardware.com/reviews/video-transcoding-amd-app-nvidia-cuda-intel-quicksync,2839-7.html
We have to wait for actual results posted by someone who wants to help the debate :)
Anyone ?
nevcairiel
2nd April 2011, 11:35
Maybe Quick Sync can directly access the decoded image in the GPUs memory, saving the memory transfer there would make it alot faster. I still think that the raw decode performance isn't that much greater. In any case, i will upgrade to a Z68 board once its out, so i can use the integrated GPU. If until then no-one else does some tests, i can do it then.
NikosD
2nd April 2011, 11:58
OK!
Looking forward to UVD 3 results,too. Because I'm not going to upgrade to a Radeon 6xxx card soon :)
renq
2nd April 2011, 12:55
OK!
Looking forward to UVD 3 results,too. Because I'm not going to upgrade to a Radeon 6xxx card soon :)
I'll download the clips and get to it then:)
Since I'm not a premium user at hotfile. it might take a while, but since I already have downloaded the first clip:
MS MFT DXVA: 51/57/61
Configuration:
Windows 7 x64 SP1
Phenom II X2 560 @ X4 B60 3885MHz
4GB DDR3
2GB HD6950 with unlocked shaders (stock clocks)
32bit dxva checker 2.4.0.0.0.0
Catalyst 11.4 beta
CruNcher
2nd April 2011, 13:19
I disagree. :-)
CPU can do everything and all statistics in this thread shows the CPU suffers less than GPU.
It is confirmed in my case. My E8400 decodes 20Mbs without problem (30% CPU) whereas my HD5770 renders jerky videos. Maybe it's a driver problem, but it's a problem I have not using CPU. And I can't use "hardware" decoders, i don't think there are for ATI...
For HTML5 and Flash, it's different, CPU is also used by javascript, css and all other stuff, so it can be helpful to have GPU.
I hope I won't piss you off, but I was not sure I read results correctly, but I think I was correct. Especially with new CPU, if you don't do a lot of post-processing, that's better to use CPU than DXVA.
It heavily depends it's not said per see that it's always more efficient Power Consumption wise to use a Full GPU Framework especially as you still bound to CPU Overhead even on Win 7 ;)
You rather have to carefully balance out pro/cons and combine them smartly on the possibilities of the OS to get best results (Power Consumption saving).
But per se saying its overall less efficient is wrong also we are just @ the beginning of how to manage the GPU efficiently inside Windows and WDDM 1.1 is not really much different in those regards then Windows XP in the future we hopefully gonna see better GPU/CPU Management and less overhead ;)
I have a P67 board, i cannot use the integrated GPU.
In any case, Quick Sync is only for encoding, the decode engine is probably not significantly faster then previous generations of ClearVideo HD.
As the encoding is done in the GPU, decoding with the CPU for re-encoding tasks is probably better. (260 fps decoding, hooray)
Yeah some reviews show 25W with full 1080p Blu-Ray playback (i think it was Xbitlabs review) :)
not that bad though nothing which Nvidia or ATI couldn't achieve (if such a low power discrete DSP only card would exist from them it consumes roughly 3W (not the DSP only measurement of that is quiete impossible todo it will be somewhere @ 500mw or roughly 1W) for 1080p Blu-Ray where Intels old Decoder Core consumed somewhere 8W (though also not the Decoder IP alone but all the sourunding CPU logic that needed to be active @ the same time in that case Atom, which came from Imagination http://www.imgtec.com/powervr/powervr-technology.asp ;) ), also i don't believe most reviews in Nvidia Encoder results because they don't set it up for maximum Performance (they all rely on data they gather from ISV implementations) and Nvidia is continuously improving their Encoder they just in the 270. driver fixed a very bad Motion Estimation bug (one of the Main Developers of Nvidias GPU Encoder Core is a ex MSU Researcher http://www.linkedin.com/pub/anton-obukhov/6/527/78b he also was one of the main Brains behind the FRExt implementation in the Nvidia Encoder that even still Mainconcept has to fight with) :)
And as you said the interesting thing is how does Intels DSP and EUs work together and are they more efficient then what Nvidia/AMD/ATI have currently discrete are the shorter paths really help allot (or is Nvidias/ATIs Optimization so good that it can cope even with that, Winning @ the Driver level) ?
No one really benched that yet except for Encoding but their it was even from THGs test (which is currently no doubt the most reliable test existing) only tested out of a consumer level pov :P
NikosD
2nd April 2011, 15:38
the first clip:
MS MFT DXVA: 51/57/61
Configuration:
Windows 7 x64 SP1
Phenom II X2 560 @ X4 B60 3885MHz
4GB DDR3
2GB HD6950 with unlocked shaders (stock clocks)
32bit dxva checker 2.4.0.0.0.0
Catalyst 11.4 beta
This is a complete disaster!
Your score is lower than mine!
You could check out two things:
1) Do not have any default options enabled in Video settings of Catalyst Control Center. You have to disable every default post-processing filter that Catalyst apply after driver's installation like Color Vibrance, Flesh tone correction, Dynamic range, Edge-enhancement, De-noise, Mosquito noise reduction etc etc.
You have to disable them all.
2) Select the null renderer (black screen) during benchmarking mode in DXVAchecker.
Could you retest it?
NikosD
2nd April 2011, 16:09
Maybe Quick Sync can directly access the decoded image in the GPUs memory, saving the memory transfer there would make it alot faster. I still think that the raw decode performance isn't that much greater
It heavily depends it's not said per see that it's always more efficient Power Consumption wise to use a Full GPU Framework especially as you still bound to CPU Overhead even on Win 7 ;)
You rather have to carefully balance out pro/cons and combine them smartly on the possibilities of the OS to get best results (Power Consumption saving).
But per se saying its overall less efficient is wrong also we are just @ the beginning of how to manage the GPU efficiently inside Windows and WDDM 1.1 is not really much different in those regards then Windows XP in the future we hopefully gonna see better GPU/CPU Management and less overhead ;)
I reply to both as an opportunity to clarify one thing.
GPU video acceleration is a misleading and confusing term, describing something that does not exist the last 4 years.
When ATI and Nvidia tried to accelerate video decoding in their cards, they did it using the same piece of hardware they used for 2D and 3D acceleration.
The results were limited and very hard to implement in video player applications.
In 2007 ATI used for the first time a dedicated piece of logic - UVD - in ATI HD 2000 series which completely offloaded both CPU & GPU shaders and it was completely independent from the rest of the card. It uses of course the video card memory, but it doesn't use GPU shaders and CPU at all (<5% CPU utilization & 0% GPU utilization)
GPU in general has nothing to do with video decoding nowadays.
When we say "hardware acceleration" or DXVA or GPU acceleration referring to video files, we always mean the independent, dedicated, added fixed function logic circuit which exists on the same die of what we call "GPU"
It's a little piece of IC, extremely fast and extremely low power consuming processor that we call UVD in ATI GPU, VP2, VP3, VP4 in Nvidia GPU. Intel has not given a particular name AFAIK.
Quick Sync is something a little different because it's a dedicated fixed function decoder and for the first time encoder, too. So Quick Sync it's the first dedicated fixed function transcoder, a bold move of Intel.
AMD and Nvidia rely on GPU shaders to encode video (or CPU) and fixed function processor (UVD and VPx) to decode video.
So, after all these I can say that there is no CPU overhead, GPU framework etc.
My Power Scheme is High Performance but I have enabled in BIOS the SpeedStep feature of Core2Duo and during DXVA playback my 2.83GHz processor goes in deep sleep mode running at the minimum clock speed of 1.7GHz.
Because it is not used at all!
CruNcher
2nd April 2011, 16:49
If you would have read the 2nd part you would know that im pretty aware of what you just explained you misunderstood me i talked about Framework not Decoding alone and surely there is a CPU overhead and depending on your view of it it is important to you or not, surely for most consumer it isn't ;)
Post Process on every GPU Nvidia/Intel/AMD/ATI/Mobile ones is dependent on the either so called EUs or Shaders if you can keep paths short as possible theoretically you should have a performance advantage that's what nev asked himself and if you believe that encoding with Quicksync shows you as low CPU utilization as if you where Decoding that is wrong there is overhead ;)
I could also explain you
http://www.www.xbitlabs.com/images/video/intel-hd-graphics-2000-3000/transcode-1.png
http://www.www.xbitlabs.com/images/video/intel-hd-graphics-2000-3000/transcode-2.png
where this massive difference comes from between Cyberlinks and Arcsofts Encoding Framework especially the Nvidia part ;)
Playback power consumption on SB without Post Processing
http://www.www.xbitlabs.com/images/video/intel-hd-graphics-2000-3000/power-5.png
The Picture btw i showed above with Complex Sharpen 2 consumes about 45W (only the Card no CPU) on a 9800 GT non fermi cores GPU Load is @ avg 30% still frames get lost also because the Nvidia VP2 as i said before cant cope with this stream and that's why jitter is also so high with 34ms its a perfect benchmark for Vista/7 Aero (DWM) and also Sandy Bridge :)
NikosD
2nd April 2011, 17:47
These graphs you posted are puzzles to be solved by them (Cyberlink, Arcsoft etc)
The reviewer as you have read in the article had no answer and Cyberlink and Arcsoft had no answer, too !
Who am I to have an answer:)
I think nobody has a solid answer.
But I was referring to DXVA only - which means decoding only - because this is the subject of the whole thread and the question of Pirlouy I was answering.
We compare CPU vs DXVA codecs, so by definition we are referring to decoding only.
I think it's clear now.
renq
2nd April 2011, 18:11
This is a complete disaster!
Your score is lower than mine!
You could check out two things:
1) Do not have any default options enabled in Video settings of Catalyst Control Center. You have to disable every default post-processing filter that Catalyst apply after driver's installation like Color Vibrance, Flesh tone correction, Dynamic range, Edge-enhancement, De-noise, Mosquito noise reduction etc etc.
You have to disable them all.
2) Select the null renderer (black screen) during benchmarking mode in DXVAchecker.
Could you retest it?
DIVX ver is 9.01.21
Arcsoft ver 2.27.319.108
FFDShow DXVA ver is 3800
All post-processing turned OFF.
1. Clip
MS MFT - 48/57/85
DIVX DXVA - 48/58/85
Arcsoft - 45/57/69
Ffdshow - 44/57/70
2. Clip
DIVX - 38/48/75
MS DTV/DVD - 31/48/86
Arcsoft - Wouldn't play
FFDSHow - Same
3. Clip
DIVX - 51/57/84
MS DTV/DVD - 48/57/79
Arcsoft - 53/57/64
ffdshow - 54/57/71
CruNcher
2nd April 2011, 18:14
And sure playback wise SB could be more efficient then VP2 or Vp4 especially power consumption wise as you dont have the discrete GFX card consumption to take into account (default idle consumption) anymore with a high Performance Decoder such as CoreAVC and with certain streams you wont have any chance and can't use only the DSP as they either wont be compatible or the DSP to weak for Realtime Playback but i tried to explain to pirolouy that this isn't always a good thing if you work in Frameworks where Post Processing is added which tasks the CPU another time and if you mix the most efficient stuff of a system in a balanced way you can get better results especially if you want to keep also 3D Performance @ the same time which Sandy Bridge cant deliver ;)
Because he thinks CPU is always more efficient and that is per see wrong even on a very Powerful low power CPU like Sandy Bridge, and now you can just imagine how Powerful it can be to mix Intels Media SDK 2.0 Ecosystem and Nvidias Nvcuvid,Nvcuvenc together into 1 Framework and utilizing both fixed functions and Shader units for different tasks @ the same time ;)
And yes i have that answer that xbitlabs and lol Arcsoft doesn't have (you can also find it on doom9) ;)
NikosD
2nd April 2011, 18:28
It seems that UVD 3 has the same performance as UVD 2.2, or worse in minimum framerate.
The results - in terms of raw performane of UVD 3 - I have to say are a little disappointing.
Thanks renq.
CruNcher
2nd April 2011, 18:54
ATI/AMD never claimed that Performance is better just that the featureset was enhanced, though the same on Nvidias side
seems the 4 Girls stream is practicaly the same Performance as on Nvidia VP2 with UVD 2.2 11) PowerDVD DXVA 55/57/79 it doesn't reach the full 60 fps so no big difference though i reach only 50 fps without benching with DXVAChecker but in Real but also on XP with VMR7 windowed lets see what DXVAChecker is saying.
pirlouy
2nd April 2011, 19:18
Thanks NikosD and CruNcher for enlightenments. I admit I don't know everythink about DXVA, that's why my first post was a question. :-)
I know it's a bit a pity not to use this module (UVD in my case), but I don't know if it's my module which has a problem, or if it's a driver problem, but in my case, it's safer to use CPU (no dropped frames).
Are there any OpenCL video decoders in the market ? Maybe it's a young technology/language, but it could be better than CUDA (for general users, not inevitably Nvidia users), since it's not reserved for a company...
nevcairiel
2nd April 2011, 19:27
Are there any OpenCL video decoders in the market ? Maybe it's a young technology/language, but it could be better than CUDA (for general users, not inevitably Nvidia users), since it's not reserved for a company...
Thats a common misconception. A CUDA video decoder does not use CUDA to decode. CUDA offers a special API to access the video decoder - you just "use" CUDA to access that API, the decoder is not written in CUDA. Thats why i named my decoder CUVID (the name of the API), and try to avoid using the terms CUDA in general when talking about it.
Because of this, OpenCL does not qualify for this. The "common" API on Windows is DXVA, sadly it does come with limitations, but because its mostly meant for playback, i don't see the vendors investing in a common API anytime soon.
NikosD
2nd April 2011, 19:29
The situation seems like MLAA.
ATI introduced MLAA as Radeon 6xxx series feature only - but later added MLAA to 5xxx series, too.
When ATI introduced UVD 2.2 they said it had all the features that UVD 3 has. Full MPEG2 acceleration, MPEG4 ASP support (DivX, Xvid)
I'm pretty sure that UVD 3 & UVD 2.2 are the same hardware.
They decided to expose the "extra" features of both versions to UVD 3 only.
Please ATI fanboys, don't answer to that guess. It's just an evil thought of me :)
But I don't think that they will act the same way as MLAA
NikosD
2nd April 2011, 19:39
Thats a common misconception. A CUDA video decoder does not use CUDA to decode. CUDA offers a special API to access the video decoder - you just "use" CUDA to access that API, the decoder is not written in CUDA. Thats why i named my decoder CUVID (the name of the API), and try to avoid using the terms CUDA in general when talking about it.
Because of this, OpenCL does not qualify for this. The "common" API on Windows is DXVA, sadly it does come with limitations, but because its mostly meant for playback, i don't see the vendors investing in a common API anytime soon.
This is not exactly right.
The latest OpenCL 1.1 specification has added direct calls to DXVA, like CUDA.
I'm not a developer to know exactly the differences between CUDA and OpenCL access to DXVA, but I'm pretty sure that OpenCL will finally catch up CUDA in DXVA access, if it hasn't done that already.
But I'm not aware of any "OpenCL" decoder.
nevcairiel
2nd April 2011, 20:11
This is not exactly right.
Oh, but it is.
The latest OpenCL 1.1 specification has added direct calls to DXVA, like CUDA.
It does not.
OpenCL 1.1 spec: http://www.khronos.org/registry/cl/specs/opencl-1.1.pdf
Go find me any reference to video decoding.
ATI is trying to be smart, and invented an API called "OpenVideo Decode", but its only ATI, and not "Open", and its still in its very early days. (and not related to OpenCL, although it has interop with OpenCL like CUVID has with CUDA)
CruNcher
2nd April 2011, 20:20
VP2 WIN XP SP3 Forceware 270.51 9800 GT
4 Girls
DXVA:
Renderer: Video Mixing Renderer 7
Decoder: CyberLink Video Decoder
Decoder Device: ModeH264_VLD_NoFGT
Processor Device: -
Time: 03:47.269
Average FPS: 53,012
Min/Max FPS: Min: 50 Max: 58
Renderer: Video Mixing Renderer 9
Decoder: CyberLink Video Decoder
Decoder Device: ModeH264_VLD_NoFGT
Processor Device: -
Time: 03:47.192
Average FPS: 53,030
Min/Max FPS: Min: 50 Max: 58
Renderer: Video Mixing Renderer 7
Decoder: CoreAVC Video Decoder
Decoder Device: ModeH264_VLD_NoFGT
Processor Device: -
Time: 03:53.328
Average FPS: 51,635
Min/Max FPS: Min: 47 Max: 56
Renderer: Video Mixing Renderer 9
Decoder: CoreAVC Video Decoder
Decoder Device: ModeH264_VLD_NoFGT
Processor Device: -
Time: 03:52.397
Average FPS: 51,842
Min/Max FPS: Min: 48 Max: 56
Nvcuvid via Dshow Surprise Surprise ;) (not really)
Renderer: Video Mixing Renderer 7
Decoder: CoreAVC Video Decoder
Decoder Device: -
Processor Device: -
Time: 04:02.820
Average FPS: 49,617
Min/Max FPS: Min: 44 Max: 55
Renderer: Video Mixing Renderer 9
Decoder: CoreAVC Video Decoder
Decoder Device: -
Processor Device: -
Time: 04:02.846
Average FPS: 49,612
Min/Max FPS: Min: 44 Max: 54
Renderer: Video Mixing Renderer 7
Decoder: LAV CUVID Decoder
Decoder Device: -
Processor Device: -
Time: 04:24.156
Average FPS: 45,609
Min/Max FPS: Min: 43 Max: 52
Renderer: Video Mixing Renderer 9
Decoder: LAV CUVID Decoder
Decoder Device: -
Processor Device: -
Time: 04:23.916
Average FPS: 45,651
Min/Max FPS: Min: 43 Max: 53
Renderer: Video Mixing Renderer 7
Decoder: CUDA Video Decoder
Decoder Device: -
Processor Device: -
Time: 04:41.867
Average FPS: 42,733
Min/Max FPS: Min: 40 Max: 46
Renderer: Video Mixing Renderer 9
Decoder: CUDA Video Decoder
Decoder Device: -
Processor Device: -
Time: 04:41.806
Average FPS: 42,742
Min/Max FPS: Min: 40 Max: 46
CPU:
Renderer: Video Mixing Renderer 7
Decoder: CoreAVC Video Decoder
Decoder Device: -
Processor Device: -
Time: 00:56.136
Average FPS: 214,622
Min/Max FPS: Min: 200 Max: 234
Renderer: Video Mixing Renderer 9
Decoder: CoreAVC Video Decoder
Decoder Device: -
Processor Device: -
Time: 00:55.837
Average FPS: 215,771
Min/Max FPS: Min: 201 Max: 248
NikosD
2nd April 2011, 20:23
I meant this:
http://developer.amd.com/gpu/amdappsdk/pages/default.aspx
"Support for UVD video hardware component through OpenCL"
I thought it was OpenCL 1.1 feature, not ATI-only.
I'm pretty sure that if ATI did it, Khronos will do it also in a future spec.
What do you think?
nevcairiel
2nd April 2011, 20:52
I don't think it'll be standardized in OpenCL.
pirlouy
2nd April 2011, 21:05
Of course I don't have nevcairiel's knowledge, but for me, OpenCL is better than CUDA because even if you seem to disagree, it seems more open than CUDA.
And if people starts to develop applications for CUDA, it's the beginning of a company dependence.
Imagine if there were a C for Intel and a C for AMD, it would be boring for everyone.
From what I've read, OpenCL can use DXVA (with its limitations), but it seems it's up to GPU companies to allow direct connections between OpenCL and their H.264/VC1 etc. modules. And from NikosD's link, it seems it's the case now with AMD.
But indeed, all this is young and should have a lot of bugs. But in some years, I hope there won't be a lot of applications "optimized for Nvidia"....
NikosD
3rd April 2011, 10:34
The only reason for Nvidia not to implement a solution for "OpenCL" video decoder ever, is CUDA.
If Nvidia had the choice, they would prefer to not implement anything in OpenCL. CUDA is their exclusive API, they try to spread everywhere and to everyone in order to sell more cards.
Eventually we will see an "OpenCL" video decoder, but for ATI only.
pirlouy
3rd April 2011, 11:33
I think they support OpenCL, but through CUDA (OpenCL -> CUDA -> hardware). So it should work, but not as powerful as if there were a direct "connection". But at least they try...
thuan
4th April 2011, 04:50
I run everything on my home system through an UPS and with its software, my power usage between accelerated playback with DXVA and ffdshow with EVR-CP are not that much different, hovering around 123w and 125w respectively. This and problems I have with accelerated video playback are more than enough for me to abandon it.
CruNcher
4th April 2011, 11:01
hmm using FFMpeg-MT, CoreAVC or DiAVC ? and yes its a valid question if you only look @ Playback power consumption especially as CPUs become more and more energy efficient, nobody yet measured how much energy for example does CoreAVC need for Playback on SB with the integrated GPU not doing anything else and with the GPU doing the decode though in this non discrette case and with the Logic combined on 1 Die its almost clear that it would save a lot of energy even compared to a High efficient Decoder like CoreAVC (it could look completely different for a discrette system where you have a base idle consumption + decoder dsp consumption + other cards components consumption while decoding) ;)
Though for other System configurations it would loss efficiency depending on several factors, though a Desktop system with discrete GPU is mostly also used for Gaming so the GPU is active so or so including the Logic you may not need, and being able to utilize that to save some power in certain tasks or just offload the CPU is nothing wrong and yeah Decoding alone can be a very small power save depending on the System but if you combine more and more tasks smart and efficiently you can save a lot.
nevcairiel
4th April 2011, 11:08
Of course once CPUs become more powerful and efficient, the difference will always be smaller. However, alot of people have fairly weak CPUs in their dedicated HTPCs, which might just manage to decode it in software, but do that at nearly full load, which adds noise and heat.
Or people want to do post-processing (ok, which does not work with DXVA, but does work with for example CUVID based decoders), and need the CPU for that.
There will always be valid use-cases for accelerated decoding, just not for everyone.
thuan
4th April 2011, 11:44
Agree and I was using ffdshow so ffmpeg-mt. Tested video was h264 encode of the last part of Planet Earth ep 5 where there were a lot of locusts.
CruNcher
4th April 2011, 12:11
Of course once CPUs become more powerful and efficient, the difference will always be smaller. However, alot of people have fairly weak CPUs in their dedicated HTPCs, which might just manage to decode it in software, but do that at nearly full load, which adds noise and heat.
Or people want to do post-processing (ok, which does not work with DXVA, but does work with for example CUVID based decoders), and need the CPU for that.
There will always be valid use-cases for accelerated decoding, just not for everyone.
Not entirely true Roozhou works on DXVA frame grabing based stuff which could be also utilized for Post Processing :)
True the first tests don't look as efficient as Nvcuvid on XP at least see Roozhou test http://forum.doom9.org/showthread.php?t=160371 but testing and real life results are sometimes not near @ each other see my Performance issues with LAV CUVID and Haali/Madvr which i yet cant explain either why i lose so heavily efficiency compared to VMR on my System configuration where theoreticaly LAV CUVID + Madvr should be somewhere performance wise equal in theory as when rendering onto VMR (where LAV CUVID on VMR9 renderless also shows problems here) but somewhere is a bottleneck might be inside my system configuration or how LAV CUVID,CoreAVC CUDA, CUDA Video Decoder provide the decoded frames i dunno yet another explanation could be high resolution timing.
I mean if you see that
LAV CUVID + MadVR
http://img189.imageshack.us/img189/5263/disappearingframes.png
LAV CUVID + VMR9 renderless
http://img190.imageshack.us/img190/4484/lavcuvidvmr9.png
and this you would also ask you these same questions ;)
Cyberlink DXVA + VMR9 renderless
http://img23.imageshack.us/img23/9450/justwow.png
nevcairiel
4th April 2011, 12:26
Especially ATI cards are *really slow* when downloading the contents of a D3D Texture (where the DXVA decoded image ends up), so it will cut into the performance. (Especially on XP this shows, the new driver model used by Vista/7 seems to improve this somewhat - but then XP is really getting old, personally i don't care for any statistics in it.)
CruNcher
4th April 2011, 16:44
Could somebody Benchmark this in Vista/7 with Forceware 270.51 ?
DXVA + VMR9 Renderless + Complex Sharpening 2
The 4 Girls stream can be found @ the first post
http://img860.imageshack.us/img860/1104/shaderperformancebench.png
Don't expect to get that Stream Playing Back Realtime neither Nvidia nor ATI Hardware can do that (or some overhead doesn't allow them too, someone with VDPAU Linux could test ?) according (also for Vista/7) to the Benchmarks here and my own, though Intel HD2000/3000 results on Vista/7 would still be needed, someone with a Broadcom Crystal or any Hardware based Windows Decoder setup and results from their would also be great :)
nevcairiel
4th April 2011, 18:25
Don't expect to get that Stream Playing Back Realtime neither Nvidia nor ATI Hardware can do that (or some overhead doesn't allow them too, someone with VDPAU Linux could test ?)
Obviously you're wrong about that
http://images.gammatester.com/pics/7ab488d8c82576b826ccdd54a1d8e0f8.png
Since i'm on Win7, i had to use EVR Custom for DXVA (VMR9 DXVA is only XP)
So, this is MPC-HC DXVA decoder, EVR Custom + sharpen complex 2, on a VP4 card (GTX 570)
Btw, the shaders shouldn't influence the performance, as long as you don't overload the shader unit. The decoder is completly separate from the shaders.
CruNcher
4th April 2011, 18:34
Nice :) so the major question now is VP4 responsible or Win 7 eg WDDM 1.1 and DWM
do you get the same results with LA CUVID (nvcuvid) instead of DXVA ?
VP2 is known to be clocked @ 400 MHz :D
And that you don't show a picture of the decoded content i guess means you can't capture the surface on EVR Custom ?
thuan
4th April 2011, 18:48
Remember I have a VP2 9800GT and with new driver, DXVA playback (and through CUDA for that matter) on Windows 7 is fubar. I say whether they works or not will always be a combination of factors. As I don't want to play roulette anymore, I'm giving up on accelerated playback for now :sign:.
nevcairiel
4th April 2011, 18:58
do you get the same results with LA CUVID (nvcuvid) instead of DXVA ?
I get about 50 dropped frames with my CUVID decoder over the whole file, but i attribute that to its not being optimized. Didn't test CoreAVC in CUDA mode.
CruNcher
4th April 2011, 19:12
I get about 50 dropped frames with my CUVID decoder over the whole file, but i attribute that to its not being optimized. Didn't test CoreAVC in CUDA mode.
Not bad :) did you compared CPU utilization in both cases already ?
do you use also the 64 bit version like thuan ?
Remember I have a VP2 9800GT and with new driver, DXVA playback (and through CUDA for that matter) on Windows 7 is fubar. I say whether they works or not will always be a combination of factors. As I don't want to play roulette anymore, I'm giving up on accelerated playback for now :sign:.
Do you get the same results as nevcairiel with the same playback setup from the 4 girls clip ?
MPC-HC DXVA decoder + EVR Custom + sharpen complex 2
you could also try
Cyberlink DXVA + EVR Custom + sharpen complex 2
or
LAV CUVID/CoreAVC CUDA/CUDA Video Decoder + EVR Custom + sharpen complex 2
Not sure why you give up it shouldn't be a problem to get it working and if 270.51 doesn't work on your 9800GT and Win 7 then try the latest Win 7 WHQL driver for the 9800GT that should work.
Would be this one currently for the 9800GT http://us.download.nvidia.com/Windows/266.58/266.58_desktop_win7_winvista_64bit_international_whql.exe
If you mean by Fubar the Performance is bad then please tell us by how much, and what do you mean by a combination of factors (MB Bios,Cpu Frequency switching, Ram timings/frequency, Gfx Card Bios, GFX Card Frequency switching ?) ? :)
thuan
5th April 2011, 03:51
I don't want to do another test because I have done it enough but my observation is like this:
As I said on another thread, only on driver version 258.96 and older, I can get accelerated playback plays at acceptable/normal performance. With newer driver and my VP2 card, performance is worse and on certain high bitrate 1080p stream, it simply goes like a slideshow. This consistently happens with any method I use to play video, regardless it's CUDA based or DXVA based decoder that is used with madVR or EVR-CP, with 258.96 and older it works, with newer driver it does not.
So I think it's a combination of my VP2 card and driver, but I'm not so sure whether that is correct in my case, still I believe it is so.
CruNcher
5th April 2011, 17:00
Im slowly also going into that direction of the Driver for my issues with LAV CUVID + MadVR + XP i tried Kernel timing now and other things still i lose a heavy amount of Frames cant be normal though Benchmarks show no issues (though DXVAchecker only seems to bench VMR7/9 windowed) @ all but as soon video output seems to come into play it crawls awaythough i have 0 problems with DXVA except the very heavy streams most others also have issues with here.
Though we still have to less results here nevcairiels results beat everything else in terms of Performance currently on his 570 with VP4 and Win 7 :) for both NVCUVID and DXVA on EVR-CP very impressive results (number wise) we have to trust him on the Real Performance ;)
PS: Im currently in a state where i don't understand the results anymore (LAV CUVID,MADVR CoreAVC CUDA + MADVR + Haali)
Here is the IMHO most direct way to test the DSP with as low overhead as possible directly via nvcuvid on XP
Display overhead
http://img846.imageshack.us/img846/8108/directdspuse.png
Without display overhead
http://img52.imageshack.us/img52/8686/disabledisplayout.png
Though this pure performance results absolutely don't come through for me with LAV CUVID or any other NVCUVID directshow implementation on any renderer in XP only with DXVA i get to the same Decoding Performance Level (obviously only on VMR7/9 windowed).
It matches with the VMR7 windowed results i get with Cyberlinks DXVA in MPC-HC
Huh after closing DGDecNV and retrying it in MPC-HC with VMR7 Renderless (3D surface) ahh ok that is explainable its not using DXVA @ all it shows it though in the Status Display ;) but it doesn't use it :P
http://img696.imageshack.us/img696/9230/vmr7renderless.png
VMR7 windowed right after it
http://img268.imageshack.us/img268/3197/vmr7windowed.png
VMR9 Renderless (no shaders) (Alternate Sync, 3D surfaces,Bicubic,VMR9-mixer mode)
http://img854.imageshack.us/img854/1049/vmr9renderless.png
VMR9 windowed
http://img851.imageshack.us/img851/3816/vmr9windowed.png
LAV CUVID + VMR7 windowed
http://img703.imageshack.us/img703/9936/lavcuvidvmr7windowed.png
LAV CUVID + VMR9 windowed
http://img651.imageshack.us/img651/6042/lavcuvidvmr9windowed.png
LAV CUVID + VMR7 Renderless (any mode) doesn't work falls back to Video Renderer eg VMR7 windowed see above result
LAV CUVID + VMR9 Renderless (no shaders) (Alternate Sync, 3D surfaces,Bicubic,VMR9-mixer mode)
http://img849.imageshack.us/img849/4875/lavcuvidvmr9renderless.png
LAV CUVID + MADVR (NV12) (heavy frame dropping fps comparable to Haali + CoreAVC CUDA (YUY2))
http://img813.imageshack.us/img813/6388/lavcuvidmadvr.png
Haali + CoreAVC CUDA (YUY2)
http://img851.imageshack.us/img851/5713/haaliyuy2coreavccuda.png
CoreAVC CUDA (NV12) VMR7 windowed (comparable LA CUVID result)
http://img20.imageshack.us/img20/579/coreavcnv12vmr7windowed.png
Currently best NVCUVID dshow result -22 fps compared to the DXVA result on VMR9 Renderless
CoreAVC CUDA (NV12) + VMR9 Renderless (no shaders) (Alternate Sync, 3D surfaces,Bicubic,VMR9-mixer mode)
http://img69.imageshack.us/img69/8035/coreavccudavmr9renderle.png
So quiet normal expected results the windowed modes are faster especially VMR7 windowed (not without a reason the default XP renderer)
VMR9 renderless obviously has the highest overhead with -8 fps compared to VMR7 windowed though not that much to worry about therefore it has all the benefits of being dxva capturable and support subtitles and the alternate sync + shader customization. Still this makes it unexplainable why i get so bad results with LAV CUVID + MadVR/VMR9 renderless or CoreAVC CUDA + MadVR especially looking @ the bandwith NV12 results.
D3D9 Surface speed test:
NV12: upload 403 fps, download 532 fps, trick download failed
YV12: upload 75 fps, download 15 fps, trick download failed
A8R8G8B8: upload 387 fps, download 248 fps, trick download failed
DXVA Surface speed test:
NV12: upload 425 fps, download 526 fps, trick download failed
YV12: upload 75 fps, download 14 fps, trick download failed
A8R8G8B8: upload 383 fps, download 242 fps, trick download failed
So maybe someone else has a idea why the performance of LAV CUVID + MADVR on XP 9800 GT Forceware 270.51 could be so extreme slow even with NV12
The major difference between LA CUVID VMR7/9 windowed (especially 9 windowed that's not overlay @ all) and MadVR is what i try to understand, or is it really the missing hardware memcopy improvements that came after G92 that play @ major role here and explain nevcariels result or is my bandwith not enough @ all ?
http://img541.imageshack.us/img541/6451/cudabandwith.png
IgorC
6th April 2011, 17:10
NikosD,
I was looking for such information. Thank you for it.
Also it will be great to see laptop's battery life scenario for different decoders.
One particular decoder can be fastest but can consume more power. It depends on how smart resources are used.
NikosD
7th April 2011, 10:03
Thanx, good to hear.
Unfortunately, I don't have a laptop to test it.
renq
8th April 2011, 10:52
Tested the clips on my old laptop (Core 2 Duo T5600 1,83GHz, W7 Pro X32 sp1 beta, 7600Go)
A. Twinpeaks-30fps
CoreAVC v2.5.1[/B] CPU 59/74/82
Microsoft MFT CPU 42/48/54
B. Samsung-30fps
CoreAVC v2.5.1 CPU 21/35/77
C. Basket-60fps
CoreAVC 2.5.1 CPU 55/66/92
D. Girls-60fps
CoreAVC 2.5.1 CPU 46/51/72
E. Birds-60fps
CoreAVC 2.5.1 CPU 31/38/49
F. Cat-60fps
CoreAVC 2.5.1 CPU 45/49/53
thuan
10th April 2011, 18:03
Seems like I have found my issue, I don't know what nvidia did but when I checked GPU-Z with a H264 video played back using MPC DXVA on 258.96 I get typically around 20% lower video engine load compared to newer driver. This gave me a clue that it is actually something nvidia did in their driver that I hope in later driver version will be fixed, if ever. Does any of you guys know the place I can report this issue to nvidia?
Rain1
10th April 2011, 18:15
Does any of you guys know the place I can report this issue to nvidia?
Try this (http://nvidia-submit.custhelp.com/cgi-bin/nvidia_submit.cfg/php/enduser/std_alp.php)
NikosD
2nd June 2011, 16:10
Maybe Quick Sync can directly access the decoded image in the GPUs memory, saving the memory transfer there would make it alot faster. I still think that the raw decode performance isn't that much greater. In any case, i will upgrade to a Z68 board once its out, so i can use the integrated GPU. If until then no-one else does some tests, i can do it then.
Z68 boards are out!
Tell me that you did the upgrade in order to see some results of Quicksync...
nevcairiel
2nd June 2011, 16:24
I did upgrade, here some quick tests, with only the MS DS H264 decoder - Min/Avg/Max values.
1. Twin Peaks: 193/195/198
3. Basketball: 193/202/207
4. Girls: 199/200/213
5. Birds: 198/200/200
Now only if the actual GPU would be fast enough to do some madVR processing with that..
NikosD
2nd June 2011, 16:41
Extremely fast as I told you!
Although the results seem a little odd because they are all around 200 fps for every clip and every value (min/avg/max)
Did you try the newest DXVAchecker v2.5 ?
Maybe it can handle Quicksync better.
Are you sure that DXVA only mode is used by Quicksync and no DXVA by GPU or software CPU ?
I'm sure the answer is DXVA Quicksync only, but I'm just wondering...
nevcairiel
2nd June 2011, 16:46
Cpu load was consistently low. Constant fps results are to be expected for a good hardware decoder, they usually don't care how complex a movie is.
Looking forward to the next nvidia decoder generation, some catching up todo performance wise.
And yes, dxvachecker2.5
NikosD
2nd June 2011, 16:57
So it seems that Quicksync is about 3 times faster than VP4 and 2-2.5 times slower than Core i7 2600 in H.264 decoding.
Looking forward to ATI next generation video decoder, too :)
Thanks for your results.
CruNcher
2nd June 2011, 19:14
@ nev
could you bench this also with Intels supplied Quicksync decoder in the Media SDK ?
dukey
2nd June 2011, 19:38
How does cuda compare using CoreAVC ?
CruNcher
2nd June 2011, 20:58
as NikosD said around 3 times slower :) Though 200 fps is a lot that mostly never someone gonna use it makes sense for the Quicksync Encoder though as it allows a lot of parallel stream encoding like Intel showed it off, and that at very acceptable quality see Tomshardware or last MSU test :)
Nev could you measure the Power consumption @ those 200 fps ?
Also 1 great thing with Quicksync is Desktop Live Recording you could easily do 60 fps recordings of the Full Windows Desktop (compressed) running even high precision timer applications without any dropped frames (in theory) :)
nevcairiel
2nd June 2011, 21:09
How does cuda compare using CoreAVC ?
CUDA gets about the same performance as the DXVA decoders, as its limited by the hardware, not any software interface.
dukey
2nd June 2011, 21:30
Strange,
when i tried CUDA vs DXVA, Cuda came out about 2x as fast. Or used 2x less CPU.
CruNcher
2nd June 2011, 23:53
That would be indeed very strange as CPU usage should be almost the same only performance should be a tad better for DXVA
dukey
3rd June 2011, 00:52
I average something like 20~ % CPU usage on this video with Cuda, and around 45% with DXVA, so CUDA totally blows DXVA away. DXVA on my system seems to offer almost no improvement at all. Perhaps it's just this video I am testing.
XinHong
3rd June 2011, 07:05
CUDA's performance is more dependent than DXVA. On my "old" system with a 8600GT DXVA is faster than CUDA.
NikosD
21st June 2011, 14:08
It seems that with new Fusion APU, there is a chance for fast DXVA + madVR decoding/ rendering due to UVD3 + No data copy between GPU memory and main memory.
We need tests, as always.
roozhou
22nd June 2011, 06:37
DXVA usually uses less memory than CUVID.
NikosD
22nd June 2011, 06:51
It doesn't matter.
Up to now, PotPlayer + DXVA renderless mode + madVR, doesn't work due to slow memory copies in ATI cards.
There is a hope that Fusion APU can do better.
CruNcher
22nd June 2011, 07:07
It doesn't matter.
Up to now, PotPlayer + DXVA renderless mode + madVR, doesn't work due to slow memory copies in ATI cards.
There is a hope that Fusion APU can do better.
Yes it should be both Brazos as well as Liano should be capable of it the lack though in Brazos is Encoding performance as they cutted of needed SIMD functions so vs Quicksync at least Brazos has no chance looks different for Liano. Though AMDs GPU encoder research is anyways to say it with nice words "minmalistic" and also in Decoder bug issues they didn't earned good community response over the last years for UVD ;)
NikosD
22nd June 2011, 07:44
UVD3 is fully featured and quite fast.
The bugs were always in software.
No serious bugs after Catalyst 10.12 + AMD MFT codecs for DivX, VC-1, WMV3
MatLz
19th November 2011, 20:34
Hi,
I am in trouble with the "Twinpeaks" sample in the first post.
FFdshow or ffms2 produce artifacts at 4s to 6s.
VLCrap is less crappy than I thought because it plays the file without problem.
Can someone confirm this ?
the_weirdo
20th November 2011, 03:24
VLC 1.2.0 nightly build (http://nightlies.videolan.org/build/win32/last/) as well as latest LAVFilters and mplayer2 produce artifacts with that sample too. So maybe it's a regression of libavcodec.
MatLz
20th November 2011, 03:36
My VLC is 1.1.9, april 2011.
NikosD
9th December 2011, 20:04
Cpu load was consistently low. Constant fps results are to be expected for a good hardware decoder, they usually don't care how complex a movie is.
Looking forward to the next nvidia decoder generation, some catching up todo performance wise.
And yes, dxvachecker2.5
Hello. I have some questions regarding QuickSync and DXVA.
I've recently read that the bitstream formatting stage is being done on the CPU even in SandyBridge with Quicksync.
Can you confirm it?
What is the situation regarding DXVA support of real DXVA capable video players like WMP12, MPC-HC(internal codecs), PotPlayer(internal codecs) and QuickSync ?
Are the players compatible with QuickSync DXVA decoding ?
What about DXVA codecs like DivX, CoreAVC, MS DS, MS MFT ?
Are there any DS or MFT codecs made by Intel ?
Is there a full DXVA VLD VC-1 support exposed by drivers for QuickSync?
And my last question:
Is it possible to post a DXVA checker screenshot (latest version/ latest drivers ?) of QuickSync ?
Thank you in advance!
vivan
10th December 2011, 10:40
What is the situation regarding DXVA support of real DXVA capable video players like WMP12, MPC-HC(internal codecs), PotPlayer(internal codecs) and QuickSync ?I'm using only MPC-HC and it support QS with H264/AVC (DXVA) internal decoder (however it produce artifacts on 1080p with 16 ReFrames video).
What about DXVA codecs like DivX, CoreAVC, MS DS, MS MFT ?At least CoreAVC.
Btw, 2.Samsung.Demo.Oceanic.Life-1080p30fpsRef16-40Mbps using CoreAVC DXVA:
Average FPS: 269,454
Min/Max FPS: 237 / 330
Are there any DS or MFT codecs made by Intel ?Maybe http://forum.doom9.org/showthread.php?t=162442
Is it possible to post a DXVA checker screenshot (latest version/ latest drivers ?) of QuickSync ?http://2.firepic.org/2/images/2011-12/10/63telyjuv6ve.png
NikosD
10th December 2011, 13:31
I'm using only MPC-HC and it support QS with H264/AVC (DXVA) internal decoder (however it produce artifacts on 1080p with 16 ReFrames video).
At least CoreAVC.
Btw, 2.Samsung.Demo.Oceanic.Life-1080p30fpsRef16-40Mbps using CoreAVC DXVA:
Average FPS: 269,454
Min/Max FPS: 237 / 330
Maybe http://forum.doom9.org/showthread.php?t=162442
http://2.firepic.org/2/images/2011-12/10/63telyjuv6ve.png
Thank you for your replies.
It seems that more work has to be done in drivers, at least.
It is good that VC-1 VLD is present at DXVA checker.
Could you do some more benchmarks with samples from these links ?
The second VC-1 clip from here: (The tough one with 1080p60 fps at 40Mbps)
http://forum.doom9.org/showthread.php?t=156660
And two very demanding H.264 clips in terms of bandwidth from here: (7th and 8th clip)
http://forum.doom9.org/showthread.php?t=163110
The link that you provided is referring to an impressive work of Eric Gur, which involves Intel Media SDK and DXVA.
It's not a pure DXVA solution.
During benchmarking, please take a look at the frequency of your processor- it should be go down to ~1.6GHz.
What is the CPU usage in DXVA checker during benchmarking ? Min/Avg/Max
Thanks
vivan
11th December 2011, 11:12
About H/W - I have i5-2410M CPU, it has 2,3 Ghz base clock speed, 2,9 Ghz Max Turbo and 800 Mhz lowest clock speed.
During playback it runs at minimum clock speed (800 Mhz). During benchmarks - at base clock speed (2,3 Ghz). If I choose "Power saving mode" it runs at 800 Mhz but benchmark results are lower - 65-70 fps for 7-8 files.
VC-1
Renderer: Video Mixing Renderer
Decoder: ffdshow Video Decoder
ffdshow VMR
Time: 00:16.983
Average FPS: 201,201
Min/Max FPS: 194 / 204
CPU Usage (%): Avg: 27 Min: 25 Max: 30
Renderer: Video Mixing Renderer
Decoder: WMVideo Decoder DMO
Time: 00:50.363
Average FPS: 67,847
Min/Max FPS: 60 / 88
CPU Usage (%): Avg: 43 Min: 39 Max: 47
Renderer: Enhanced Video Renderer (Media Foundation)
Decoder: WMVideo Decoder MFT
Time: 01:12.893
Average FPS: 46,863
Min/Max FPS: 43 / 54
CPU Usage (%): Avg: 25 Min: 24 Max: 27
8. AVC
Renderer: Enhanced Video Renderer (DirectShow)
Decoder: CoreAVC Video Decoder
Decoder Device: ModeH264_VLD_NoFGT_ClearVideo
Time: 00:04.268
Average FPS: 128,866
Min/Max FPS: 121 / 129
CPU Usage (%): Avg: 06 Min: 02 Max: 15
Renderer: Enhanced Video Renderer (DirectShow)
Decoder: Microsoft DTV-DVD Video Decoder
Decoder Device: ModeH264_VLD_NoFGT
Time: 00:04.275
Average FPS: 123,977
Min/Max FPS: 116 / 127
CPU Usage (%): Avg: 26 Min: 21 Max: 37
Renderer: Video Mixing Renderer
Decoder: ffdshow Video Decoder
Time: 00:04.671
Average FPS: 117,748
Min/Max FPS: 99 / 120
CPU Usage (%): Avg: 30 Min: 26 Max: 40
I have some problems with file number 7, DXVA Checher showed only one decoder:
7. AVC
Renderer: Enhanced Video Renderer (Media Foundation)
Decoder: Microsoft H264 Video Decoder MFT
Decoder Device: ModeH264_VLD_NoFGT_ClearVideo
Time: 00:02.846
Average FPS: 130,358
Min/Max FPS: 129 / 130
CPU Usage (%): Avg: 24 Min: 22 Max: 27
But when I remuxed it to mkv:
7. AVC (remuxed to mkv)
Renderer: Enhanced Video Renderer (DirectShow)
Decoder: Microsoft DTV-DVD Video Decoder
Decoder Device: ModeH264_VLD_NoFGT
Time: 00:02.587
Average FPS: 131,040
Min/Max FPS: 127 / 131
CPU Usage (%): Avg: 27 Min: 20 Max: 33
Renderer: Enhanced Video Renderer (DirectShow)
Decoder: CoreAVC Video Decoder
Decoder Device: ModeH264_VLD_NoFGT_ClearVideo
Time: 00:02.894
Average FPS: 128,196
Min/Max FPS: 124 / 129
CPU Usage (%): Avg: 11 Min: 06 Max: 16
Renderer: Video Mixing Renderer
Decoder: ffdshow Video Decoder
Time: 00:03.271
Average FPS: 113,421
Min/Max FPS: 98 / 120
CPU Usage (%): Avg: 28 Min: 26 Max: 31
NikosD
12th December 2011, 10:04
Good job!
These are VERY INTERESTING results.
It seems obvious to me now that indeed, even in latest HW of Intel (QuickSync), there is no pure DXVA solution in video decoding.
It reminds me the debate to my other post regarding the same thing of a missing stage in H.264 decoding pipeline that has to be done in CPU when processed by G45 chipset.
That debate had a lot of intense comments about Intel's HW/drivers efficiency in DXVA decoding.
Not a lot of things have changed since, regarding CPU dependency.
For a pure HW decoding implementation like UVDx or VPx, the CPU always run in lowest frequency - even in benchmark mode - and the results are always the same, no matter what the CPU frequency is.
This clearly doesn't happen in QS implementation, because higher CPU frequency leads to higher framerates in benchmark mode.
That bitstream formating stage is being processed by CPU.
It's also clear that FFDShow VC-1 codec (Intel Media SDK) is using QS but with a large CPU usage ~25%
CoreAVC is definitely the way to go for DXVA decoding with H.264 and QS - fast performance with minimum CPU usage, although I think I read that is not pure DXVA solution for Intel.
It's using Intel Media SDK too like FFDShow QS.
Interesting comparison between FFDshow QS vs CoreAVC in DXVA (Intel Media SDK) mode
After your results I have to ask Intel one question:
Where are your DXVA DS/ MFT decoders for H.264/ VC-1 ?
Thank you Vivan for your time and the results.
NikosD
12th December 2011, 17:21
@vivan
Any particular reason that you tested both VC-1 and H.264 clips with FFDShow QS version in VMR renderer ?
Why don't you use EVR with FFDShow QS and add the results to the above post?
It should be faster...
CruNcher
14th December 2011, 19:48
Unfortunately its not that easy to bench the performance as DXVA Checker also has problems with several DXVA Decoder + Splitter combinations you sometimes get very weird benchmark or playback results some Decoder splitter combinations even fail completely in a 0 bench result (they just brake up returning 1000 fps) (happens also on the plain Graph level with Software like Graphstudio).
Also CoreAVC is some kind of a special DXVA implementation it can avoid some x264 bitstream issues other DXVA Decoder like Cyberlinks or Arcsoft would fail on, i would prefer it anytime in terms of overall bitstream stability though even better for DXVA is Mirillis Splash Players DXVA + Direct3d Render combination its pure awesome (it even avoids extreme bitstream issues not 1 glitch where CoreAVC would show @ least 1 on intels Hardware) and on the other side it's also the most resource efficient Decoder + Render combination i yet saw on Windows :)
Mirillis is really a company to look out for very talented Windows Multimedia Coder (with very deep Codec knowledge) they even beat the Asians on some stuff,they surely gonna give CoreCodec a run for it's money in the future, just recently they started to attack FRAPS Lossy Intra Codec Part for now and they don't do bad (also they have very talented GUI designers and Artists as well, remembers a little bit of DivX awesome Web Designer (who created the Cow and Alien Stuff) back then just for the Desktop) ;)
Though his style was is awesome http://web.archive.org/web/20041216090025im_/http://images.divx.com/home/banner_lotr.jpg :) not for nothing he won several awards :D
Normally im not so into Design of applications but usability and efficiency though if you do all of those right or very balanced it's awesome ;) http://mirillis.com/gfx/action/mirillis_action_window_desktop_recording.jpg
NikosD
14th December 2011, 22:36
I think that latest versions of DXVA checker 2.6.1/2.6.2 are rock solid in terms of stability of results and bug free more or less.
I'm talking about ATI hardware UVD and Microsoft+AMD decoders (DS&MFT).
Do you have any examples of bad behavior of DXVA checker and funny results ?
For Mirillis I could say that when I first wrote here about Splash Pro in the DiAVC post, nobody knew the program and they all tried to lower it by talking about only for bugs that time.
It's good to see that other people than me see talent in that company.
vivan
16th December 2011, 12:38
Any particular reason that you tested both VC-1 and H.264 clips with FFDShow QS version in VMR renderer ?
Why don't you use EVR with FFDShow QS and add the results to the above post?
It should be faster...Actuallty it was slightly slower and less stable.
7.
Renderer: Enhanced Video Renderer (DirectShow)
Decoder: ffdshow Video Decoder
Decoder Device: ModeH264_VLD_NoFGT
Time: 00:03.303
Average FPS: 112,322
Min/Max FPS: 96 / 117
CPU Usage (%): Avg: 27 Min: 25 Max: 30
8.
Renderer: Enhanced Video Renderer (DirectShow)
Decoder: ffdshow Video Decoder
Decoder Device: ModeH264_VLD_NoFGT
Time: 00:04.932
Average FPS: 111,517
Min/Max FPS: 79 / 115
CPU Usage (%): Avg: 27 Min: 25 Max: 31
And with vc-1 ffdshow qs failed.
If I choose play -> EVR - it plays it smoothly (http://2.firepic.org/2/images/2011-12/16/zoslxl93k364.png).
But if I choose benchmark -> EVR - it just show blank window (http://2.firepic.org/2/images/2011-12/16/ztjkn2iam60i.png). And forces turning Aero off in 90% cases.
Looks like DXVAChecker bug...
And about CPU usage egur recently wrote in his thread:
http://forum.doom9.org/showthread.php?p=1544519#post1544519
Anyway, CPU usage is low enough and decoder is quite fast. Plus it's possible to use madVR :)
NikosD
16th December 2011, 13:00
Strange behavior justified only by the non pure DXVA implementation of FFDShow QS.
Pure DXVA uses EVR only.
It could be a bug of DXVA Checker or FFDShow QS or both.
The work that has already be done by Egur is impressive but Intel has to do two things:
1) Make pure working DXVA decoders (DS/MFT).
2) Make HW acceleration (DXVA) COMPLETELY CPU INDEPENDANT.
It's a combination of drivers/ hardware that I hope to see in Ivy Bridge which will have the capability of multiple streams of H.264 hw decoding up to 4K resolution.
It will support 4K x 4K too.
nevcairiel
16th December 2011, 13:07
2) Make HW acceleration (DXVA) COMPLETELY CPU INDEPENDANT.
Their GPU is in their CPU. :P
For the record, using a good DXVA decoder (or a QuickSync decoder) will already result in very low CPU usage, allowing it to remain in the lowest performance state possible.
NikosD
16th December 2011, 13:55
Their GPU is in their CPU. :P
For the record, using a good DXVA decoder (or a QuickSync decoder) will already result in very low CPU usage, allowing it to remain in the lowest performance state possible.
The CPU resources used by a very powerful CPU like SB or Ivy etc when decoding in benchmark mode even the toughest H.264 clips should be near 0%.
With my poor CPU Core2Duo at 1.6GHz when I benchmark even the thoughest H.264 clips with UVD2.2, I never go beyond 8%-9% at the max. The average is 3%-5% !!! With a dual core@1.6GHz!
They have to implement the whole pipeline of H.264 (including bitstream format stage) in drivers/ hardware in Ivy Bridge.
BTW, DXVA decoder equals or should equal QS decoder, meaning that QuickSync decoder should use pure DXVA decoder only (not using Intel Media SDK)
The latter (Intel Media SDK) is far more flexible and useful for some people than pure DXVA , but it should exist an option for pure speed (DXVA only) by Intel.
vivan
16th December 2011, 14:13
should exist an option for pure speed (DXVA only) by Intel.lowest fps is ~120 fps on 8th video while both nvidia ant ati has much slower perfomance ;)
And for "pure speed" there is an option called "CoreAVC", at least.
Also madVR is much better option than higher benchmark results :)
nevcairiel
16th December 2011, 15:03
The CPU resources used by a very powerful CPU like SB or Ivy etc when decoding in benchmark mode even the toughest H.264 clips should be near 0%.
Since i don't believe any benchmark i didn't do myself, here it is.
I used the "Girls" clip, because i still had it on my disc.
NVIDIA GTX 570
Renderer: Enhanced Video Renderer (DirectShow)
Decoder: Microsoft DTV-DVD Video Decoder
Decoder Device: ModeH264_VLD_NoFGT
Processor Device: -
Time: 02:36.734
Average FPS: 76,237
Min/Max FPS: 74 / 79
CPU Usage (%): Avg: 01 Min: 00 Max: 02
Intel i7 2600k:
Renderer: Enhanced Video Renderer (DirectShow)
Decoder: Microsoft DTV-DVD Video Decoder
Decoder Device: ModeH264_VLD_NoFGT
Processor Device: -
Time: 00:28.663
Average FPS: 416,879
Min/Max FPS: 398 / 427
CPU Usage (%): Avg: 03 Min: 02 Max: 05
I'll leave everyone to judge these results all they want, but a 5-fold increase of decoding speed will just need more CPU, no questions asked. 1% at 80fps is just 5% at 400 fps.
Not sure why this clip is so extremely fast, though, others play at around ~200-300, but still at a maximum of 4-5% CPU.
So, no, they don't need to change anything, and they also don't need their own decoder, the MS decoder is doing just fine. :)
PS:
The Intel Media SDK is just a wrapper around DXVA2, there is no special API like there is with CUDA for NVIDIA.
NikosD
16th December 2011, 17:40
If we compare different things we are not going to extract useful and valid results.
First of all you didn't write your CPU frequency during benchmarking.
5% of what CPU ? Frequency!
The toughest clips I know and Vivan tried are clips 7. and 8. and the CPU usage was ~25% according to him during benchmarking.
This is a DAMN HIGH CPU USAGE FOR A QUAD CORE CPU AS POWERFUL AS SANDY.
Try to benchmark clips 7. and 8. WRITING THE PROCESSOR FREQUENCY.
But above of all these...
When Vivan forced the CPU frequency down to 800MHz, he got half fps in clips 7. and 8. !!
This is UNACCEPTABLE.
CPU frequency and cpu decoding ability should have nothing to do with decoding performance when using DXVA.
It's that simple!
And for "pure speed" there is an option called "CoreAVC", at least.
It uses Intel Media SDK too, I think.
Pure DXVA means no CUDA or OpenVideo or Intel Media SDK involved.
nevcairiel
16th December 2011, 18:30
You sure are an unfriendly fellow
Anyhow, here are results for sample 8.
Intel i7 2600k
Renderer: Enhanced Video Renderer (DirectShow)
Decoder: Microsoft DTV-DVD Video Decoder
Decoder Device: ModeH264_VLD_NoFGT
Processor Device: -
Time: 00:03.700
Average FPS: 143,243
Min/Max FPS: 133 / 148
CPU Usage (%): Avg: 08 Min: 07 Max: 09
CPU goes to full clock during benchmark (3.8Ghz)
In "Playback" mode, CPU stays at lowest setting (1.6Ghz), at around 3-4% usage
The NVIDIA GTX570 doesn't even manage to reach 24 fps on it:
Renderer: Enhanced Video Renderer (DirectShow)
Decoder: Microsoft DTV-DVD Video Decoder
Decoder Device: ModeH264_VLD_NoFGT
Processor Device: -
Time: 00:23.928
Average FPS: 22,150
Min/Max FPS: 20 / 24
CPU Usage (%): Avg: 03 Min: 02 Max: 05
Again, a 5 to 6 fold performance advantage, with only tripple the CPU usage (which is partly from the demuxer and other DirectShow components).
I don't know what more proof you need, the decoding is completely done in hardware, and its the fastest out there. The relatively high cpu load in benchmark mode is only natural from the very high fps achieved.
Don't forget that you still need a CPU to pump 120 fps of a very high bitrate clip from your HDD through DirectShow to the decoder.
DXVA only does the decoding, but the splitter still needs to process the video, and if its very high bitrate, there also is more data to process --> higher CPU load.
120mbit at 24 fps equals 600mbit at 120 fps, quite alot of data to shuffle around!
Oh well, i'm done. Believe what you want, but the hardware decoder in the Intel CPUs is some fine work.
Now if they only fixed some of the driver bugs, it would be even better. :)
NikosD
16th December 2011, 18:41
Keep up the good work with LAV Filters.
Hope you add MFT MKV splitter and DXVA for all hardware (ΑΤΙ, Nvidia, Intel) soon :)
NikosD
28th December 2011, 20:58
I have recently installed latest beta drivers of Nvidia 290.53 on a Geforce GT440 under Windows 7 Home x86 and I didn't get any codecs by Nvidia.
I thought that Nvidia provided "NVIDIA Video Decoder MFT" in their drivers.
Does any know what happened and Nvidia stopped providing MFT decoders ?
Which is the last driver with Nvidia MFT decoders ?
Where can I find NVIDIA MFT decoders ?
Thanks!
nevcairiel
28th December 2011, 21:57
NVIDIA didn't provide their own decoder for quite a long time because the MS decoder works just fine with their cards.
I don't know when the last driver was that shipped the decoder, but it wasn't in the 2xx series of drivers, so no driver that supports the 440 will have it. :p
NikosD
29th December 2011, 08:53
Thanks for the info.
But then, if Nvidia decided so, how could someone use VC-1/WMV3 DXVA VLD decoder ?
Because MS decoders(DS/MFT) for DXVA VC-1/WMV3 don't support VLD decoding.
MS decoders provide full acceleration (VLD) only for MPEG2 and H.264.
That's why AMD decided - that was a very nice move by AMD - to implement its own MFT decoders for VLD VC-1/WMV3/ MPEG4ASP (DivX, Xvid).
BTW, how could someone use DXVA MPEG4ASP with VP4 ?
I think that official DivX codec only works with UVD3 and your CUVID decoder is discontinued.
Is there any other way ?
JohnnyFu
3rd August 2012, 08:50
Hi guys,
I have a question regarding DXVA and I believe this is a good place to ask.
During evaluation of h264 software decoders we realized we cannot decode four h264 streams at the same without violating our software specification requirements for CPU load.
I tried CoreAVC, Mainconcept, Elecard and Microsoft. Our i7-620M plattform is above 70% load when decoding more then two streams (1080p 25fps, avg 20Mbit/s) at the same time. And under 100% load when decoding four streams at the same time.
I noticed when decoding a single stream using EVR renderer, at least Elecard, Mainconcept and Microsoft make use of DXVA reducing CPU load to almost 0%.
However, I also noticed when decoding more then one stream at the same time, only one GraphStudio seems to make use of DXVA.
So I wonder, if it is possible to make use of DXVA with multiple decoding processes at the same time? Or does simple not work by design?
Any tip on how to decode up to four HD streams at the same time on our i7-620M plattform without causing CPU load higher then 60% would be greatly appreciated! Our goal are streams at 1080p, 20mbps, 23-30fps.
EDIT: just noticed, using MPC-HC both process seems to make use of DXVA as they are both causing just 2% load instead of 30%. However, the video sutters horrible...
NikosD
3rd August 2012, 09:43
Your CPU/Platform (i7-620M) uses Nehalem architecture for mobile, called Arrandale.
Arrandale doesn't have a powerful integrated GPU, to be more exact VPU (Video Processing Unit), to process at the same time four 1080p30fps H.264 streams.
It is possible that you can't even process two streams of the above type with your hardware.
So whatever software/ codec you try (CoreAVC, Mainconcept etc) it doesn't change anything, it's the hardware the limiting factor of the decoding performance of your platform.
Using my discrete card - ATI Radeon 5750 - I can only decode in HW (DXVA) two streams of 1080p30fps H.264 at 20Mbps.
Not even three.
But I can do four 720p30fps H.264 streams at 10Mbps.
QuickSync hardware available inside most Sandy Bridge and Ivy Bridge Intel processors and maybe VP5 discrete Nvidia cards can decode four 1080p30 fps - 20Mbps H.264 streams.
For QuickSync I'm sure.
JohnnyFu
3rd August 2012, 10:01
We are able to decode two to three 1080p streams at 20-30mpbs using CoreAVC in software mode. However this violates our CPU load SRS.
When using DXVA with two streams, both Microsoft and Elecard seems to make use of DXVA for both streams. However it looks like they are competing for DXVA API in some way as the video stutters horrible as soon as I start the second stream.
I just realized our encoding plattform uses Sandy Bridge, I'll reconfigure the lab and see if there is any stuttering when decoding two streams using DXVA.
vivan
3rd August 2012, 10:14
Run benchmark using DXVA checker. If you'll get less than 60 fps on your video - than your hardware is too slow for decoding even 2 such streams.
JohnnyFu
3rd August 2012, 10:17
Wow... I just decoded four streams using Microsoft's Decoder without any problems, 10% CPU load with Sandy Bridge! Thank you very much, huge step forward for me!
fashionman
28th November 2012, 10:10
hi,
I develop video decoder with media foundation and DXVA2, after decode frame, the GPU usage about increase 20%, but no video display on screen, could you give me advise,
thanks
NikosD
28th November 2012, 12:51
Sorry I'm not a developer.
Maybe Egur (Eric) or Nevcariel (Hendrik) could help you.
rubait
23rd July 2020, 00:28
Can you give me access to ftp://helpedia.com/pub/multimedia/x264/testvideos/? Wanted to test these out on a Icelake system for comparison.
NikosD
23rd July 2020, 11:11
Can you give me access to ftp://helpedia.com/pub/multimedia/x264/testvideos/? Wanted to test these out on a Icelake system for comparison. The server seems to be down, but it's not mine.
I think a user called @mariush had uploaded all those files on that server.
Maybe you could reach him.
vBulletin® v3.8.11, Copyright ©2000-2026, vBulletin Solutions Inc.