Log in

View Full Version : H.264 DXVA Benchmarks: QuickSync vs UVD 2.2 vs VP4 vs VP5


Pages : 1 2 3 4 5 6 [7] 8 9

clsid
6th August 2014, 14:00
Has anyone tested if the decoder also works without OpenCL? For example by temporarily switching to the Generic VGA Driver.

It would be interesting to see its pure CPU performance.

P.J
6th August 2014, 18:06
Edit: Downloaded the PowerDVD trial, it's definitely using different code for its OpenCL.

Would you share some results? Does it use any GPU at all compared to LAV?

NikosD
6th August 2014, 19:23
@ clsid

Good idea.

I disabled from Device Manager the HD 4600 driver so I used "Microsoft Basic Render Driver" - you don't have to uninstall your driver for anyone would like to try.

Initially I used DXVA Checker in benchmark mode as usual using EVR renderer.

The results with OpenCL decoder look like it didn't work (very low CPU usage ~7 fps performance on Core i7-4790)

Then I used EVR Renderless and I did two tests:
One with 4C/4T configuration (HT OFF) and the other at 4C/8T (HT ON)

Also, in order to use OpenCL Decoder in OpenCL mode, I did the exact same tests using EVR Renderless mode with OpenCL (HD 4600 driver)

LAV x86/x64 had exactly the same results, with or without OpenCL.

But OpenCL decoder, although used in EVR Renderless mode, had a small difference using OpenCL

The results reveal a lot of things.

Core i7-4790 - EVR Renderless

4C/4T

LAV x86 10/17/20 CPU usage: 98%

LAV x64 17/23/28 CPU usage: 97%

OpenCL x86 29/40/44 CPU usage: 96% OpenCL OFF
OpenCL x86 23/41/49 CPU usage: 86% GPU usage 600MHz@60% OpenCL ON

4C/8T

LAV x86 12/19/22 CPU usage: 85%

LAV x64 19/26/30 CPU usage: 78%

OpenCL x86 30/44/51 CPU usage: 82% OpenCL OFF
OpenCL x86 15/44/55 CPU usage: 76% GPU usage: 600@50% OpenCL ON


My comments:

1) CPU usage of LAV x86/x64 and OpenCL x86 on a 4C/4T CPU is excellent - Almost 100% !

2) CPU usage of LAV x64 on a 4C/8T could be optimized better.
LAV x86 has 10% more CPU usage than x64 and more than OpenCL x86 too.

3) For LAV x86/x64 there was no difference between EVR and EVR renderless performance with or without OpenCL enabled.
But for OpenCL x86 using EVR renderless gives it a boost with OpenCL enabled or disabled due to larger CPU utilization and lower GPU usage.

4) It is clear that OpenCL decoder is mainly a CPU decoder and it's faster as CPU decoder on a 4C/8T Core i7-4790, than a CPU/GPU decoder on the same CPU.

5) The use of GPU (HD 4600) as HEVC OpenCL decoder, doesn't make faster the HEVC decoding on a Core i7-4790 4C/8T , but drops the CPU usage a lot by offloading a part of the HEVC algorithm to the GPU during real-time playback and benchmarking with EVR (not EVR renderless)

6) The huge difference between OpenCL decoder and LAV x86/x64 is not the CPU usage (actually OpenCL decoder has less CPU usage compared to LAV x86) but the more optimized use of the CPU - maybe use of different vector instruction set (?)

Asmodian
7th August 2014, 04:28
Thanks for the benchmarks, very interesting results. :)

I don't think you have any data to back up point two, both show a similar drop in CPU utilization with hyper-threading on. It may simply be a limitation of hyper-threading, the extra threads have to share most of the hardware with the original four. Of course there probably is more optimization possible for both decoders.
Edit: Sorry, LAV x64 does have enough lower utilization to be interesting but it might be more due to the nature of hyper-threading instead of less optimization.

I strongly agree with point four, epically looking at minimum frame rates which are the most important for real time display. When using OpenCL the OpenCL decoder min frame rate was half of the same decoder in pure CPU mode! It even dropped below the min frame rate of LAV x64. Interop penalty?

Edit: I hope this interop penalty would not show itself on an AMD APU though I suspect it would still run faster in pure CPU on your i7-4790.

NikosD
7th August 2014, 05:26
I agree with everything you wrote at your post.

LAV x64 probably is pushing HT to its limits on a Core i7-4790, because is a lot faster than LAV x86 per thread.

But looking at the test of clsid measuring fps per thread, maybe you can get a few fps more by using a multiplier of x2 than x1.5 that LAV is using now.

Main issue for LAV is the per thread performance of CPU decoding compared to OpenCL decoder.

The difference is so huge that I think - as I wrote above - that OpenCL decoder is using vectorised code a lot better than LAV.

About the min value of OpenCL, most of the times was 0!

It maybe is an interoperability penalty or EVR renderless is causing the whole thing, because with EVR using OpenCL the min value is normal.

Regarding OpenCL GPU performance, it would be interesting to add a fast GPU to a fast CPU and repeat the tests, especially an AMD GPU.

Because all of my tests were done on the iGPU of Core i7-4790 which is a HD4600.

But because HD4600 is clearly underutilised even by a strong CPU, I doubt that even a R9 290 would make any significant difference.

But we have to test it as always.

huhn
7th August 2014, 06:59
on my i3 4120 the hd 4400 is working like a handbrake for the openCL assisted decoder max CPU usage was about 40%.

but i use a difference file like the rest of you

NikosD
7th August 2014, 07:26
@huhn

I had the same problem with some files, probably due to incompatibility of OpenCL decoder.

Even LAV has problems with some HEVC files displaying black screen or the problems are within the files (first early samples of HEVC encoding)

Google 4K HEVC Ducks Take off sample and try again.

clsid
7th August 2014, 14:08
The raw CPU performance of the OpenCL decoder is a good indication of what we can expect of LAV in the future. The HEVC decoder in FFmpeg is still under heavy development, and there is still lots of room for improvement.

huhn
7th August 2014, 17:04
my data is worthless. it always uses the AMD GPU not my HD 4000 and when i disable the AMD gpu is crashes same goes for EVR renderless it simply stops at one point.

my 1080p file still works like a handbrake even through the AMD GPU is used when the intel GPU is active. EVR renderless gets stuck at one point and but shows the same CPU usage of under 50 %. intel GPU is still at ~88% doesn't make a lot of sense to me...

when the AMD gpu is active everything works fine.

Asmodian
7th August 2014, 22:33
when the AMD gpu is active everything works fine.

How is the min frame rate with the AMD GPU active?

NikosD
8th August 2014, 11:21
Moving my Radeon 5750 from the old Core2Duo PCI-E v1.0 x4 platform to the Core i5-2400 PCI-E v2.0 x16 platform with clean installed Win OS, the picture of Strongene's OpenCL decoder, changed once again (and of LAV too)

1080p HEVC TearsofSteel movie - EVR - OpenCL ON - Benchmarking mode of DXVA Checker

LAV x86 (June 2014): 59/90/229

LAV x86 (Aug 2014): 64/102/239

GT440 OpenCL x86: 152/224/264 CPU usage: 86%

5750 OpenCL x86: 71/109/187 CPU usage: 28% GPU usage max clock@90%

LAV x64 (June 2014): 105/180/307

LAV x64 (Aug 2014): 131/224/360


4K HEVC Ducks Take off - Same configuration as above


LAV x86 (Aug 2014): 6/13/15 CPU 97%

5750 OpenCL x86 4/24/29 CPU 68% GPU 62%

LAV x64 (Aug 2014): 11/18/21 CPU 97%


My comments:


1) We should pay more attention when the developer says that OpenCL decoder is optimized for AMD cards and for each card different resolution could be supported.


2) The OpenCL decoder clearly drops a lot the CPU usage of Core i5 when decoding 1080p clips with 5750, but is not optimized for 5750 & 4K clips.
The CPU usage goes high as the GPU usage drops on 4K.


3) LAV Aug seems more optimized for 1080p HEVC clips than 4K and it's definitely faster than June.


4) Min values still go even to 0 fps (just for an instant) when using AMD card (5750)

huhn
8th August 2014, 16:10
How is the min frame rate with the AMD GPU active?

from which file did you like to know that?

P.J
8th August 2014, 21:09
Renderer: Enhanced Video Renderer (DirectShow)
Decoder: LAV Video Decoder x64 0.62
Decoder Device: -
Processor Device: 6CB69578-7617-4637-91E5-1C02DB810285
Time: 40.965
Frames: 500
Avg FPS: 12.205fps (Min-Max: 5-15fps)
Avg CPU Usage: 97% (Min-Max: 72-100%)
Avg GPU Usage: -

Only 3fps more... waiting for proper solution from Intel/Nvidia ;)

NikosD
9th August 2014, 07:45
Performance test in GraphStudioNext (with null renderer). TearsOfSteal 1080p sample.

OpenCL: 278 fps
LAV x86: 133 fps
LAV x64: 280 fps

I tried GraphStudioNext(with null renderer) version 0.61.265 with OpenCL x86 decoder, but I didn't find the OpenCL decoder filter.

I mean, I can build the graph with the sample, decoder and null renderer, but when I select View -> Performance Test... only a few of the DS filters appear (only the "pure" CPU decoders I think)
There is no OpenCL decoder in that list.

How can I add the OpenCL decoder to that list or how can I test performance of the OpenCL decoder from the graph that I built ?

clsid
9th August 2014, 14:57
Build the performance test graph with LAV Video, and then manually edit the graph to swap the decoder.

wxhyn
9th August 2014, 18:42
Moving my Radeon 5750 from the old Core2Duo PCI-E v1.0 x4 platform to the Core i5-2400 PCI-E v2.0 x16 platform with clean installed Win OS, the picture of Strongene's OpenCL decoder, changed once again (and of LAV too)

1080p HEVC TearsofSteel movie - EVR - OpenCL ON - Benchmarking mode of DXVA Checker

LAV x86 (June 2014): 59/90/229

LAV x86 (Aug 2014): 64/102/239

GT440 OpenCL x86: 152/224/264 CPU usage: 86%

5750 OpenCL x86: 71/109/187 CPU usage: 28% GPU usage max clock@90%

LAV x64 (June 2014): 105/180/307

LAV x64 (Aug 2014): 131/224/360


4K HEVC Ducks Take off - Same configuration as above


LAV x86 (Aug 2014): 6/13/15 CPU 97%

5750 OpenCL x86 4/24/29 CPU 68% GPU 62%

LAV x64 (Aug 2014): 11/18/21 CPU 97%


My comments:


1) We should pay more attention when the developer says that OpenCL decoder is optimized for AMD cards and for each card different resolution could be supported.


2) The OpenCL decoder clearly drops a lot the CPU usage of Core i5 when decoding 1080p clips with 5750, but is not optimized for 5750 & 4K clips.
The CPU usage goes high as the GPU usage drops on 4K.


3) LAV Aug seems more optimized for 1080p HEVC clips than 4K and it's definitely faster than June.


4) Min values still go even to 0 fps (just for an instant) when using AMD card (5750)


I have tested the OpenCL decoder on my i7 3960X + R9 290X. The results are as follows:

4K Elysium with VMR-9 renderless
LAV x64
Avg FPS: 92.334fps (Min-Max: 56-175fps)
Avg CPU Usage: 70% (Min-Max: 58-87%)
Avg GPU Usage: 0% (Min-Max: 0-0%)


OpenCL disable GPU
Avg FPS: 153.480fps (Min-Max: 103-215fps)
Avg CPU Usage: 77% (Min-Max: 67-88%)
Avg GPU Usage: 0% (Min-Max: 0-0%)


OpenCL enable GPU
Avg FPS: 164.095fps (Min-Max: 105-250fps)
Avg CPU Usage: 57% (Min-Max: 16-72%)
Avg GPU Usage: 44% (Min-Max: 0-89%)

Looks like the GPU does take effect and the speed is higher.

NikosD
9th August 2014, 19:27
You must test EVR in order to fully accelerate OpenCL decoding.

Renderless modes (VMR/EVR) are used mainly to see the pure decoding performance, without serious affect of GPU.

But regarding OpenCL GPU performance we want the opposite!

Full involvement of the GPU.

wxhyn
10th August 2014, 06:12
I tried OpenCL decoder with EVR renderless. Very weird, it frequently jammed while testing and the speed is not as fast as VMR for both GPU disable and enable, sometimes only 1 or 2 threads are used, clearly not all the speed potential is unleashed. Seems there's confliction between the decoder and EVR. For realtime playback, both EVR and VMR works fine.
I have collected the CPU and GPU usage while playing back the 4K Elysium.

OpenCL enable GPU
Frame Rate: Avg: 24 Min: 24 Max 24
CPU Usage: Avg: 04 Min: 00 Max 07
GPU Usage: Avg: 40 Min: 00 Max 100

OpenCL disable GPU
Frame Rate: Avg: 24 Min: 24 Max 24
CPU Usage: Avg: 07 Min: 00 Max 13
GPU Usage: Avg: 06 Min: 00 Max 100

LAV x64
Frame Rate: Avg: 24 Min: 24 Max 24
CPU Usage: Avg: 11 Min: 02 Max 26
GPU Usage: Avg: 04 Min: 00 Max 100

The OpenCL has the lowest CPU usage figure.

Asmodian
10th August 2014, 06:42
from which file did you like to know that?

Any 4K HEVC file really, I am simply curious if the OpenCL decoder is able to avoid the odd stalls/interop on AMD GPUs.

Looking at wxhyn's benchmarks it might, at least on new AMD cards.

@wxhyn, are you using an HEVC file, I think 4K Elysium is actually AVC isn't it?

NikosD
10th August 2014, 07:58
I tried OpenCL decoder with EVR renderless

You have tested everything besides what's most common/useful :)

Benchmark with EVR and give us your results with OpenCL ON. (OpenCL OFF is not working OK with EVR)

NikosD
10th August 2014, 11:02
Build the performance test graph with LAV Video, and then manually edit the graph to swap the decoder.

It worked.

But it makes me wonder, if you knew the method to completely isolate the GPU decoding (OpenCL decoding) by using the Null renderer, why did you ask for someone to test the OpenCL's decoder pure CPU performance by disabling GPU (OpenCL) when you have already done that ? ( by testing it with Null Renderer - which is the same)

What did you want to see ?

My results with GSN using Null renderer
(pure CPU performance without involving GPU at all)


1080p Tears of Steel (Avg fps)


LAV x86 (Aug): 166,3 fps

OpenCL x86: 341,8 fps

LAV x64 (Aug): 370,3 fps
LAX x64 (Aug): 237,5 fps using EVR with DXVA Checker



4K Duck Take Off (Avg fps)


LAV x86 (Aug): 19,7 fps

OpenCL x86: 45,4 fps

LAV x64 (Aug): 27,3 fps


JFYI, the next DXVA Checker v3.1.0 that I have tried in beta, has the exact same results with GSN using "DXVA decoding" which is the new method for testing pure CPU decoding performance
(It replaces the VMR/EVR renderless mode)

@clsid

As I've written before, I've seen lots of times a huge drop in performance of LAV decoder when using EVR renderer (real world use/test) vs null renderer.

I mean all the decoders have a performance hit using EVR renderer compared to a null renderer, but for LAV is there something to be optimized better for real world use of actual renderers like EVR ?

wxhyn
10th August 2014, 11:48
Any 4K HEVC file really, I am simply curious if the OpenCL decoder is able to avoid the odd stalls/interop on AMD GPUs.

Looking at wxhyn's benchmarks it might, at least on new AMD cards.

@wxhyn, are you using an HEVC file, I think 4K Elysium is actually AVC isn't it?

The 4K Elysium I used is HEVC encoded by NGCodec. From my observation, there are problems when do performance test with EVR renderless. Playback looks normal no matter with VMR or EVR.

wxhyn
10th August 2014, 12:35
You have tested everything besides what's most common/useful :)

Benchmark with EVR and give us your results with OpenCL ON. (OpenCL OFF is not working OK with EVR)

OpenCL enable GPU
Avg FPS: 127.515fps (Min-Max: 70-154fps)
Avg CPU Usage: 43% (Min-Max: 19-68%)
Avg GPU Usage: 59% (Min-Max: 0-95%)

LAV x64
Avg FPS: 107.338fps (Min-Max: 64-189fps)
Avg CPU Usage: 71% (Min-Max: 55-86%)
Avg GPU Usage: 0% (Min-Max: 0-0%)

The OpenCL performance is lower with EVR than VMR.

NikosD
10th August 2014, 12:47
R9 290X seems underutilised even in 4K and even fed up by i7-3960X.

And the fps are lower than using CPU alone.

The behaviour is like 5750...
I think it shouldn't, it should be more optimized for 4K.

Could you try the 1080p "Tears of steel" HEVC movie in EVR benchmarking mode ?

wxhyn
10th August 2014, 14:23
R9 290X seems underutilised even in 4K and even fed up by i7-3960X.

And the fps are lower than using CPU alone.

The behaviour is like 5750...
I think it shouldn't, it should be more optimized for 4K.

Could you try the 1080p "Tears of steel" HEVC movie in EVR benchmarking mode ?

Sorry, I don't have the 1080p tears of steel video clip. I think it's too early to draw a conclusion that R9 290X is underutilized in 4K. I observed that the OpenCL decoder with GPU enable/disable has almost the same performance with each other, i.e. the slow decoding compare to using VMR is not only with GPU only, but also with pure CPU decoding. Seems there's a wall that the speed can't get over. Another very interesting result is I tested the speed in EVR with several other 4K clips. All speeds are almost the same. I guess maybe Strongene may not deal with the EVR correctly. It looks like waiting for something finished before it can go any further.

NikosD
10th August 2014, 15:49
Sorry, I don't have the 1080p tears of steel video clip.


When I said movie, I didn't mean commercial movie.
Tears of Steel is a free to download movie.

You can grab it from here in HEVC 1080p format.

http://trailers.divx.com/hevc/TearsOfSteelFull12min_1080p_24fps_27qp_1474kbps_GPSNR_42.29_HM11.mkv


I think it's too early to draw a conclusion that R9 290X is underutilized in 4K. I observed that the OpenCL decoder with GPU enable/disable has almost the same performance with each other, i.e. the slow decoding compare to using VMR is not only with GPU only, but also with pure CPU decoding.

I guess maybe Strongene may not deal with the EVR correctly. It looks like waiting for something finished before it can go any further.

You have to forget VMR, it's meaningless to use it.
It's a very old renderer from Windows XP time.

Since Vista, only EVR is used.
You can't DXVA HW accelerate anything in Vista and above using VMR, you have to use EVR which is a very light but full of capabilities renderer, especially the EVR-CP (custom presenter)

DXVA Checker v3.1.0 is out officially, with a lot of useful changes.
There is no VMR/EVR and VMR/EVR Renderless anymore.

Check the http://bluesky23.yu-nagi.com/en/DXVAChecker.html page for changes.

wxhyn
11th August 2014, 12:20
I have collected the results with 1080p tears of steel with DXVA decoding and DXVA processing modes

Strongene OpenCL DXVA decoding
FPS: 559.518 [350-704] fps
CPU Usage: 45 [23-57] %
GPU Usage: 68 [0-95] %

Strongene OpenCL DXVA processing
FPS: 189.263 [125-191] fps
CPU Usage: 14 [8-25] %
GPU Usage: 73 [0-100] %

Strongene CPU version DXVA decoding
FPS: 305.755 [172-442] fps
CPU Usage: 25 [17-29] %
GPU Usage: 0 [0-30] %

Strongene CPU version DXVA processing
FPS: 187.171 [126-200] fps
CPU Usage: 15 [6-20] %
GPU Usage: 55 [0-100] %

LAV x64 DXVA decoding DXVA decoding
FPS: 349.292 [226-507] fps
CPU Usage: 42 [30-51] %
GPU Usage: 0 [0-0] %

LAV x64 DXVA decoding DXVA processing
FPS: 303.971 [215-349] fps
CPU Usage: 37 [9-54] %
GPU Usage: 87 [0-100] %

The results looks normal when using DXVA decoding mode. But when using DXVA processing mode, the Strongene decoders show weird results which the highest FPS can't get over 200 FPS no matter how fast the decoding is. I have tested several other 1080p clips, all of the average FPS are close to 190 FPS. Seems it's a bug of the decoders (both the CPU and OpenCL version). The LAV decoder looks normal.

NikosD
11th August 2014, 12:51
The DXVA processing method has changed and uses native scaling by default pushing the GPU usage a lot, even a pure CPU decoder like LAV has huge GPU usage.

Change the scaling manually by putting 640x480 resolution and try the processing method again.

wxhyn
11th August 2014, 13:05
I tested another 1080p video clip. I don't know the name, it's a Korea singing group, singing and dancing.

Strongene OpenCL DXVA decoding
FPS: 204.664 [151-315] fps
CPU Usage: 33 [20-43] %
GPU Usage: 28 [0-83] %

Strongene OpenCL DXVA processing
FPS: 182.384 [122-200] fps
CPU Usage: 30 [10-38] %
GPU Usage: 70 [0-100] %

Strongene CPU version DXVA decoding
FPS: 143.681 [110-245] fps
CPU Usage: 21 [13-25] %
GPU Usage: 0 [0-0] %

Strongene CPU version DXVA processing
FPS: 124.531 [91-187] fps
CPU Usage: 17 [11-21] %
GPU Usage: 40 [0-100] %

LAV x64 DXVA decoding DXVA decoding
FPS: 161.381 [130-288] fps
CPU Usage: 43 [21-59] %
GPU Usage: 0 [0-6] %

LAV x64 DXVA decoding DXVA processing
FPS: 147.937 [129-296] fps
CPU Usage: 44 [30-59] %
GPU Usage: 51 [0-100] %

This time the results looks normal. When the pure decoding speed not exceeds 200FPS too much. The processing results are very close to decoding results. Hope Strongene can fix this bug very soon.

NikosD
11th August 2014, 13:16
I don't think there is a bug in Strongene's decoder or probably an inefficency of the benchmarking tool.

In your OpenCL version the maximum fps are limited by the GPU.

The CPU version is not existant because you disable the driver, so the results can't be accurate and predictable.

Try with GraphStudioNext and null renderer for pure CPU results of Strongene's decoder.

wxhyn
11th August 2014, 13:52
I don't think there is a bug in Strongene's decoder or probably an inefficency of the benchmarking tool.

In your OpenCL version the maximum fps are limited by the GPU.

The CPU version is not existant because you disable the driver, so the results can't be accurate and predictable.

Try with GraphStudioNext and null renderer for pure CPU results of Strongene's decoder.

No, there is a CPU version decoder on Strongene's website. The one I used is the CPU version. So GPU is enabled when I tested.

NikosD
11th August 2014, 13:57
That decoder is slower than using Strongene's OpenCL as a CPU decoder and probably has more bugs.

Try GSN to see the difference.

wxhyn
11th August 2014, 15:18
That decoder is slower than using Strongene's OpenCL as a CPU decoder and probably has more bugs.

Try GSN to see the difference.

What does GSN refer to?

wxhyn
11th August 2014, 15:19
What does GSN refer to?

I guess it stands for GraphStudioNext.

NikosD
13th August 2014, 12:36
I built a testing collection of artificially high bandwidth H.264 and H.265 files using x264 and x265 encoders respectively.

I call them monster files.

It's a collection of 1080p and 3840x2160 (4K) files with bandwidths of 50Mbps, 100Mbps, 200Mbps. ..up to 600Mbps for both resolutions and both codecs.

They are 28 files in total with a footprint on disk of about 15GB.

It's not possible to upload all of them, but if someone has a specific need for a file, we'll see what we'll do.

nevcairiel
18th August 2014, 09:19
Since you guys like benchmarking software decoder as well now, here is a nightly build of LAV with HEVC improvements merged from the OpenHEVC project. On a 4K sample I tried, it increased performance over 100% (37 -> 85) in the 64-bit build.
The performance increase will vary greatly on the features used in the encode, ie. if SAO is used or not, etc. (for reference, the 1080p Tears Of Steel encode from DivX does NOT use SAO, so the improvement will be smaller).

32-bit: http://files.1f0.de/lavf/LAVFilters-0.62-14-g1c3f78b.zip
64-bit: http://files.1f0.de/lavf/LAVFilters-0.62-14-g1c3f78b-x64.zip

Also note that extremely high bitrates distort the result, as most of the time is then spent in CABAC decoding (ie. bitstream parsing), and not image reconstruction.
Real-world bitrates are the most sensible to test, as thats what will matter the most in the end as well.

NikosD
18th August 2014, 17:19
DXVA Checker v3.1.0 - DXVA decoding - Signature system

1080p - (1920_ProRes_2mbps.mkv)


LAV x64 (.14) 314/406/455 CPU: 80%

LAV x64 (.13) 113/258/343 CPU: 67%

LAV x86 (.14) 134/176/256 CPU: 92%

LAV x86 (.13) 90/148/192 CPU: 87%


3840 x 2160 - Ducks Take Off


LAV x64 (.14) 53/57/62 CPU: 78%

LAV x64 (.13) 21/27/33 CPU: 77%

LAV x86 (.14) 27/34/38 CPU: 87%

LAV x86 (.13) 14/20/22 CPU: 82%


Really huge performance advantage of the new version (.14) over the old one (.13)

LAX x86 .14 is even faster than LAV x64 .13 in 4K decoding (!!)


I think I was right talking about vectorized code, or not ? ;)

foxyshadis
18th August 2014, 23:52
Wow! The x64 version is now faster than OpenCL! That's incredible.

nevcairiel
19th August 2014, 08:00
Its fascinating how much SSE2 IDCT and SSSE3 SAO can do for performance, huh.

NikosD
19th August 2014, 12:43
Even with my Core2Duo (Wolfdale with SSE4.1 support) I see huge gains on 4K clips between the two latest versions:

60% for x86
70% for x64.

If you see the above tables, Haswell has even better results:

70% for x86
111% for x64.

Now the question is:

Still no AVX2 optimizations ? Why ?

huhn
19th August 2014, 14:28
Still no AVX2 optimizations ? Why ?

this is most likely one of the last steps to add this, most CPU doesn't support this and using AVX1/2 doesn't mean you get an huge performance improvement. so they use the common ones first or those that give an good improvement. as you can see they get huge improvements with this so they are totally right.

kasper93
19th August 2014, 15:35
I did quick test on 4096x1720 24fps with 2157Kbps sample. And compared to Lentoid decoder.

LAV x86: 29.3461 FPS
LAV x64: 89.4249 FPS
Lentoid: 81.1206 FPS

Not bad, but x86 is really lagging behind which can be a problem for madVR users.

huhn
19th August 2014, 16:12
I did quick test on 4096x1720 24fps with 2157Kbps sample. And compared to Lentoid decoder.

LAV x86: 29.3461 FPS
LAV x64: 89.4249 FPS
Lentoid: 81.1206 FPS

Not bad, but x86 is really lagging behind which can be a problem for madVR users.

I don't see a huge problem with madVR and 32 bit h265 is young and a 64 bit version of madVR will come sooner or later so it's fine for the time been.

NikosD
19th August 2014, 17:07
this is most likely one of the last steps to add this, most CPU doesn't support this and using AVX1/2 doesn't mean you get an huge performance improvement. so they use the common ones first or those that give an good improvement. as you can see they get huge improvements with this so they are totally right.

I think AVX can't help here because we are talking about integers mostly.

But AVX2, although only for Haswell and better, could be a lot faster than SSEx with 256bit registers vs 128bit registers.

And Haswell got an implementation of AVX2 right from the beginning, not like SSE2 and Pentium 4.

A lot of programs optimised for SIMD get AVX2 optimizations now, they don't have to wait for other CPUs to come with AVX2 support.

clsid was saying that in the future (!) ffmpeg HEVC decoder will eventually reach Strongene's speed.

Well, after my pressing about vectorised code missing from LAV's ffmpeg HEVC decoder, someone looked at OpenHEVC that already had vectorised code.

And the funny thing is that OpenHEVC is a fork of FFMpeg.

So eventually, the future is now.

nevcairiel
19th August 2014, 17:16
There is a whole bunch of reasons why there is no avx2 yet, and many rather technical that going into them is not worth it. However, there will also not be such a huge boost as you might think, the biggest boost was from the SSE stuff. Its not going to be twice as fast just because its in theory double the register size.

The next step will have to be to rewrite these optimizations into proper ASM instead of compiler intrinsics, and contribute them to FFmpeg proper, only after that avx2 is likely to appear. And that's neither an easy nor a fast task. Optimizing algorithms like this is very specialized knowledge.

NikosD
19th August 2014, 17:24
ASM itself is a very hard thing on its own.

No doubt about that.

But I think the developers involved are special too.

I would definitely like to read the technical reasons of not having AVX2 yet (besides what is already mentioned) and the technical reasons why the gain wouldn't be so much compared to SSEx's boost.

NikosD
23rd August 2014, 08:36
I'll answer myself, my previous question based on the main developer of x264 project - dark_shikari - words:

Here is an introduction to AVX2 optimizations in x264 project (about 1 month before the actual release of Haswell)

http://www.scribd.com/doc/137419114/Introduction-to-AVX2-optimizations-in-x264

Here is a comment of Dark_Shikari about autovectorization of modern compilers regarding SIMD instructions

https://news.ycombinator.com/item?id=5603406

...and finally here is the actual gain of Haswell vs Ivy measured by him (15%-20% per clock and from that figure, only 5% due to AVX2 optimizations only) on June 2013.
I don't know if further optimizations regarding AVX2 code have been made in x264 project, since June 2013

http://forum.doom9.org/showthread.php?p=1632275#post1632275

I don't know if x265 can benefit more of AVX2 and if H.265 decoding is something a lot different and can benefit even more.

NikosD
31st August 2014, 09:49
It seems that in 2 days from now (2nd of September), a new card from AMD will be reviewed by technical sites.


I read some interesting info regarding the new multimedia engine of Tonga GPU (R9 285) which is considered as a new iteration of GCN architecture v1.2


"The GCN 1.2-based GPUs will also feature a new multimedia engine – which comprises of universal video decoder 6.0 (UVD 6) and video encoder engine 3.1 (VCE 3.1) technologies – as well as a new high-quality scaler for video. There is no word about support for ultra-high-definition (UVD) video codecs, such as H.265/HEVC or VP9, so it looks like the new GPUs will not support them."

Let's hope that finally, AMD will support in HW 4K H.264 and possibly a GPU assisted H.265 format

Yups
31st August 2014, 10:19
I'm looking forward to Skylake-S which can do H265 encoding in hardware.

NikosD
31st August 2014, 10:55
2016 ?

We are on August of 2014 !