Log in

View Full Version : Evaluation of HEVC decoders (SW, Hybrid and HW)


Pages : 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 [26] 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52

NikosD
31st August 2016, 13:12
The fastest Kabylake processor Core i7 7700K will try to stand against the 8 core/ 16 thread ZEN in the early 2017.

It has no chance at all, it will be eaten for breakfast by ZEN in all apps besides the heavily optimized AVX/ AVX2 ones.

Of course Kabylake has an iGPU inside where ZEN 8 core/ 16 thread will be a CPU only SoC.

But that's a problem of Intel to release 8 core CPUs without iGPUs for mainstream Core i7 and not for the High-End Desktop (HEDT) terribly expensive systems only.

huhn
31st August 2016, 14:26
the marketing from intel is terrible in general.

they have mainstream processors without an iGPU but they usually cost extra making them pointless.

so there is absolutely no reason to buy a mainstream CPU without iGPU. you can use the hardware decoder or the iGPU just in case you need it.

AMD ZEN can spare a lot of silicon making them a lot cheaper in production and the TDP will look better on paper.

NikosD
31st August 2016, 14:39
Didn't understand a word of what you've written, but OK...Maybe others do.

Stereodude
31st August 2016, 16:12
The fastest Kabylake processor Core i7 7700K will try to stand against the 8 core/ 16 thread ZEN in the early 2017.

It has no chance at all, it will be eaten for breakfast by ZEN in all apps besides the heavily optimized AVX/ AVX2 ones.

Of course Kabylake has an iGPU inside where ZEN 8 core/ 16 thread will be a CPU only SoC.

But that's a problem of Intel to release 8 core CPUs without iGPUs for mainstream Core i7 and not for the High-End Desktop (HEDT) terribly expensive systems only.
I know I always enjoy reading your off topic delusional AMD fanboy posts. Please keep it up.

NikosD
31st August 2016, 16:29
Just facts, as always.

Yups
31st August 2016, 17:50
Zen will lose in 99% of games because it cannot compete with the IPC+clockspeed of Kabylake. Just one example where you are wrong.

NikosD
31st August 2016, 17:59
I'm not so sure about games like you.

Because new games using Vulkan and DX12 tend to use more and more cores and threads, it will be an interesting fight with games.

In office applications, production applications, rendering, general purpose benchmarks and Cinebench it will vanish Kabylake for sure.

Yups
31st August 2016, 18:57
Doesn't matter really. Over 99% of the games will still use DX11 or even DX9 by next year. And for non gaming apps there is lots of stuff where you don't see improvements over 4 cores. It would be very very sad if a 8 core CPU can't demolish a 4 core CPU in Cinebench. This isn't a fanboy thread for AMD CPUs, you may search for another place to discuss this further.

P.J
31st August 2016, 21:43
No more off topic, please @@

http://www.notebookcheck.com/Kaby-Lake-Core-i7-7500U-im-Test-Skylake-auf-Steroiden.172422.0.html

Notebookcheck has a DXVA checker screenshot from the Kaby Lake HD650 Graphics

HEVC Main 10 as well as VP9 Profile 0 and 10 bit are now decoded up to 8k.

http://www.notebookcheck.com/fileadmin/Notebooks/MSI/CX72-7QL/dxva_kaby.png

Saved for Kaby Lake :thanks:

NikosD
1st September 2016, 06:12
Polaris DXVA Checker
https://s9.postimg.org/p38ccn03j/dxvachecker_msi_rx470.jpg

CruNcher
3rd September 2016, 01:08
Most Probably behind one of the uknowns is the VP9 Decoder some time now past since their announcement that it's going to be enabled and yet nothing happened officially ;)

While Intel as well as Nvidia are running, the disadvantage of low resources to work with you could guess.

Overall Polaris DSP really looks like just being in the R&D state of the Previous VPX (Maxwell) and Intel far far ahead of both again (Tegra,Maxwell,Pascal,Polaris).

I really wonder how much sense it really makes to keep that fixed function unit of Nvidias and AMds in a Intel System it makes no real sense efficiency wise especially in the upcoming Kaby Lake Systems it's just pure unneeded overhead on the discrete side and the valuable Wafer space could be used for more 3d and compute logic or to make the grid even denser.


I guess AMD has planned to release/enable VP9 Decoding support officially with the upcoming Polaris 10 Notebook Reviews and supplied Drivers.

NikosD
3rd September 2016, 04:17
Yes, but is the late Polaris VP9 decoder a fixed-fuction decoder like Maxwell's with 960 card which Nvidia enabled some time later than the initial release of 960 or a hybrid decoder ?

Of course as long as Chrome official releases don't support VP9 HW decoder, all the other uses of VP9 are insignificant.

Only Edge supports VP9 HW decoder.

But the difference between a hybrid and a HW decoder is not insignificant.

Trevonn
3rd September 2016, 06:04
I really wonder how much sense it really makes to keep that fixed function unit of Nvidias and AMds in a Intel System it makes no real sense efficiency wise especially in the upcoming Kaby Lake Systems it's just pure unneeded overhead on the discrete side and the valuable Wafer space could be used for more 3d and compute logic or to make the grid even denser.

It makes complete sense considering Intel has only just delivered a Hardware decoder for HEVC 10-bit, VP9 when you could have had one from NVIDIA all the way back in January 2015 (VP9 Enabled in December 2015). Plus nobody upgrades their CPU that regularly unlike GPUs because it's a waste of money.

Let's not leave yet another industry to stagnating Intel

rbej
4th September 2016, 07:47
I really like all of you guys that after you tell me your off-topic story you call me for writing off-topic.

But you really overcome everyone here by telling me what to say in my own thread and to search somewhere else to write.

It couldn't be more funny.

The truth is that AMD is going to kick for good some @sses and all the fan boys of Nvidia and Intel are desperate and nervous.

Vega will just destroy Pascal for sure, it's another class it will show no mercy.

The same thing is gonna happen to Kabylake.

All of the AAA titles are already DX12 and Vulkan or they are going to get a patch soon.

The rest of the games are simply not interesting.

Most of the apps nowadays are getting a serious boost going hyperthteading, meaning getting 8 threads for 4 real cores.

Just imagine how fast will be 8 real ZEN cores with 16 threads.

Poor Intel...

Poor Nvidia...

2017 will be the year of AMD.

Case closed.

Very funny joke . Today is 1 April??.

You must very "love" AMD, if you realy believe in this.

ashlar42
6th September 2016, 10:57
What would currently be the most cost effective solution to build a Windows 10 based HTPC, with 4K and HEVC decoding capabilities? I ask here since you seem to be trying all combinations. Thank you.

huhn
6th September 2016, 14:39
the cheapest card is the rx 460. for UHD presentation you should get a GPU with at least 3GB of vram.

JohnLai
7th September 2016, 04:02
the cheapest card is the rx 460. for UHD presentation you should get a GPU with at least 3GB of vram.

I disagree with rx 460....the better card would be nvidia gtx 1050 (4gb variant, not yet released) or 1060 3gb.

Don't forget Pascal series support HEVC 8,10 and 12 bits decoding in hardware. Beside, it has been confirmed through nvidia Video Codec SDK that pascal also support VP9 Profile 0 decoding in hardware too.

AMD Polaris only supports 8 and 10 bit HEVC....and there is conflicting info about vp9 support whether it is hybrid or full fixed hardware mode.

EDIT: Oh great, assuming if VCE used in Bristol Ridge is the same as Polaris, then according to http://techreport.com/review/30619/amd-unwraps-its-seventh-generation-desktop-apus-and-am4-platform , VP9 support is limited to 1080p content. Meanwhile Pascal could support up to 4k and 8 k VP9 decoding....., what a bummer....

huhn
7th September 2016, 15:47
you should read the question.

and the RC 460 is not the cheapest card that can decode HEVC?

JohnLai
7th September 2016, 16:01
you should read the question.

and the RC 460 is not the cheapest card that can decode HEVC?

:p Cheapest but not the best value for video decoding functionality.

huhn
7th September 2016, 16:25
so card for double the money that can theoretically decode 12 bit which isn't even part of the DXVA spec is better?

P.J
7th September 2016, 20:12
:p Cheapest but not the best value for video decoding functionality.
Maybe good for 4k h.265 @60fps while 960/950 are better.

I would get 1050/1060 for video or waiting for Kaby Lake.

JohnLai
8th September 2016, 03:40
so card for double the money that can theoretically decode 12 bit which isn't even part of the DXVA spec is better?

What make you think Microsoft won't update the spec?

Even if Microsoft doesn't update the spec, Nvidia might do it via their CUVID/NVDECODE API. (Seem nvidia does update its cuda decoder)

huhn
8th September 2016, 08:19
there is no commercial 12 bit content. even if they are adding it you are unlikely to really need it any time soon.

the last CUVID "update" i know of was just a repack. CUVID has some serious issues on windows 10 on top.

they have a lot todo maybe adding 10 bit is a start...
the problem that only nvidia can do it in theory will slow this down too. commercial product will stay away from it for better compatibility.

Paul Tronc
8th September 2016, 08:41
native
madVR GPU queue at 4 present queue 3 1600 mb BLACK SCREEN in full screen.
EVR CP queue 4 1780 mb

a difference of 100 of mb each time are pretty normal.
get more than 2 gb that's all i have to say to this.

the default GPU queue for madVR is 8 which is unusable for UHD with 2 GB Vram. the default GPU queue for EVR is 5 or 6.

GPU usages was ~70% with both EVR and madVR. madVR was at default except queue and FSE mode.

the cheapest card is the rx 460. for UHD presentation you should get a GPU with at least 3GB of vram.

Hello, I just have purchased a 960GTX 2Gb in order to get full hardware 4K HEVC decoding in madvr. Are you saying that it won't work because of lack of memory?

huhn
8th September 2016, 12:04
if you are going to use it to display UHD AT UHD resolution.

you could run out of Vram.

in term of madVR you are relatively lucky. you can lower the GPU queue. but deinterlanced content can't be deinterlanced by madVR like this and other issues can show up.

with default madVR settings it will not work.
i'm really regretting buying the 2G version of the 960.

you are totally fine on a 1080p screen (i planned this card for a 1080p screen). just to make that clear.

Paul Tronc
8th September 2016, 13:26
you are totally fine on a 1080p screen (i planned this card for a 1080p screen). just to make that clear.

I'm relieved to read this. Actually my brand new projector is 1080p, so I want to downscale UHD content. But it still have to figure out how the video decoding chain works, because I'm using SVP (frc tool triggered via avisynth) together with madvr. I can't say whether SVP is working on 4K resolution or 1080p on my current setup, when I play 4k on my 1080p diffuser. I'd like to get SVP working smoothly on UHD content, together with madvr. I just received the 960/2Gb, paid ~100 bucks (used). Still have to clean all radeon software and switch to this Nvidia card before I can tell how it behaves.

Paul Tronc
9th September 2016, 08:53
All right,

My first tests using GTX 960 + HEVC + Madvr 32 + Lav Filter 32 + 1080p diffuser are very positive. That's simple, I could'nt find the decoding limit of my setup. I launched the http://jell.yfish.us/ 400mbps 4k uhd hevc 10bit video, I get a perfect rendering using my standard Madvr settings. The only limitation is the downscaling algorithm, it seems like I have to stay on bicubic. Any other algorithm will cause frame dropping.

I suspect the jellyfish samples to be not fully representative of a high bitrate video, it looks too easy to be true. I tried another 2016p 60fps 50Mbps video, no problem : http://demo-uhd3d.com/fiche.php?cat=uhd&id=90

Next steps, I want to try :

- DXVA checker benchmark
- same tests on my 1440p display
- adding frc (SVP) to get 60fps (I don't understand why SVP/avisynth don't see HEVC files, there must be a wrong setting somewhere)

If you have another suggestion for a stronger stress test...

v0lt
9th September 2016, 19:50
If you have another suggestion for a stronger stress test...
Netflix_TunnelFlag_4096x2160_60fps_x265_8bit_700Mbit.mkv (https://yadi.sk/i/swMdlJHMuB6yR)

huhn
9th September 2016, 20:27
close i get ~59 FPS in decode mode.

Paul Tronc
10th September 2016, 11:00
Netflix_TunnelFlag is not smooth on my 1440p display, I'll check later on my 1080p one. I tried with MPC HC 32, MPC-BE 64, PotPLayer 64, each time with Madvr.

I'm currently encoding a 1080p BluRay to HEVC at 190fps, CPU usage 7% , GPU usage 6%, GPU temp 50°C with the fans off. Pretty impressive.

CruNcher
10th September 2016, 21:02
It seems that Nvidias Hybrid H.265 Decoder is more efficient with CUVID Decoding especially on intra frames DXVA Native has the tendency to cause latency spikes (no matter which decoder used) on edge decoding cases i guess it has todo with DXVA Natives Frequency Scaling issues in conjunction with Nvidias Boost Power Management System.

I frames seem to heavy fluctuate in my tests with Nvidias Hybrid Decoder GM204 and a DXVA frequency of around 759 MHz @ 0,862v Real 4K 23.976 Fps

CUVID doesn't show these Latency I frame fluctuations @ 1126 MHz @ 1.025v


NVAPI Hook readouts (MSI Afterburner hooked NVAPI results gained after test)

Though Unwinders code is super nice it itself creates almost 0 fluctuations when polling of Nvapi is kept in sane ranges ;)

Im mainly using as seen MPC-BE Sync Graph which is awesome fine grained for a almost Realtime Graph especially in Analyze Mode 2 when CPU overhead is way lower then with the Full OSD and additional timer polling (NVAPI ?) eliminated :)

@V0lt

Please add a Sync Graph only Analyze Mode without any additional OSD timer running maybe it can further lower the Overhead and it's caused EVR Latency Fluctuations that become visible in the Graph itself increasing the Audio and Video distance from the middle with DWMs tripple buffering :)
same happens with Energy Saving Modes the lines move away when timer precision is lowered and CPU overhead gets higher with lower Frequency :)


PS: About the Intel vs AMD thing Intel had a major advantage with shrinking up to 22nm before AMD could do it thus is the only reason they won the Mobile Space with Baytrail and Cherrytrail adding more CUs each time this advantage is slowly over, but it gave Intel the time to improve the GPU each time and close the gap to AMD and Intels GPU Core being now more feature advanced then all of them including the Video Asic part, where we will see surely another interesting development but i guess not with Kabylake in terms of Intels capability with their Video Decode/Encode time to market updates in the future ;)

At least with 14nm we reach pretty much parity for the first time between virtualy all of them for a longer time period.

Intel = 22nm/14nm/ in the future targeting 11nm but the shrinking advantage will be lower as ever before
Nvidia =28nm/16nm/14nm
AMD = 28nm/14nm

so 14nm is the common denominator on that a lot architecturally will happen finally and no shrinking race advantage anymore for a longer period :)

aufkrawall
10th September 2016, 21:33
Who cares about minor fluctuations when you typically have more than a dozen frames queued ahead by decoder and even can have madVR render them ahead?

Paul Tronc
11th September 2016, 00:22
It seems that Nvidias Hybrid H.265 Decoder is more efficient with CUVID Decoding especially on intra frames DXVA Native has the tendency to cause latency spikes ...

Indeed : I switched to CUVID, the Netflix_TunnelFlag is almost playable. Globally smooth, only 170 frames dropped for the entire sequence.

Paul Tronc
11th September 2016, 00:28
Im mainly using as seen MPC-BE Sync Graph which is awesome fine grained for a almost Realtime Graph especially in Analyze Mode 2 when CPU overhead is way lower then with the Full OSD and additional timer polling (NVAPI ?) eliminated :)

@V0lt

Please add a Sync Graph only Analyze Mode without any additional OSD timer running maybe it can further lower the Overhead and it's caused EVR Latency Fluctuations that become visible in the Graph itself increasing the Audio and Video distance from the middle with DWMs tripple buffering ...

I'll have to dig into this 'Sync Graph', for now I don't understand what it is.

CruNcher
11th September 2016, 17:54
It's pretty simple a Graph that shows you the current Render Jitter and this is dependent on many System factors the lower the jitter the smoother the Rendering playback ;)

And the Graph is very sensitive to System Problems and fluctuations in the whole WDDM Playback Chain because it's own Render Overhead isn't really high though it works only in the Render context of D3D9 in MPC-BE for now it's though muich nicer then some static numbers you can't bring into context over time of what you seeing even if it's slightly delayed :)

Push for example the Print Screen button and see if you get a spike ;)

Indeed : I switched to CUVID, the Netflix_TunnelFlag is almost playable. Globally smooth, only 170 frames dropped for the entire sequence.

Not bad if you take into account that you have only 2 GB

Though you wouldn't really expect anything else from VP7/8

Nvidia is very picky about memory see the last AV1 Optimizations they proposed for the De-ringing Filter or their latest AGGA R&D ;)

Nvidias Hybrid Decoder failing not even Broadcast Complexity but high enough to get it performing very unstable ;)

http://i1.sendpic.org/t/cQ/cQIaWZlkF5cxXQEjibEXl65h6Wo.jpg (http://sendpic.org/view/1/i/eb4nMTYvvZQ2XsfTV3F4BbVbFrX.png)

13 sm vs 4 Intel Cores to the rescue ;)

http://i1.sendpic.org/t/pb/pbWkzApGJyjH9CmLG9jP8D3YjcK.jpg (http://sendpic.org/view/1/i/kT8krcaaWAS7gj9Qhau7iM1Mq4T.png)


PS: Did some RX 460,470 480 user tried if Strongene Lentoids Hybrid OpenCL Decoder still works, as each IDCT,MC and PP part is generation individually optimized and chosen at runtime i wonder if it still works @ all on Polaris GCN 4.0 supporting the Decoding or fails or if it's going to fallback to the Tonga\Hawai or some Generic GPU/CPU Path ?

Lentoid can even partly outrun lav video in complex bitstream parts when the CPU starts to starving under heavy load and push out more stable results then but therefore it also uses lot more System Ram and it performs equal in 32 bit as in 64 bit no slowdown like lav video ;)



Fast Scenechange sudden Bitrate Peak test

Lav Video CPU

http://i1.sendpic.org/t/f4/f4Oppp66Wexv9MiIAESTetRV4w3.jpg (http://sendpic.org/view/1/i/gw5BsSf8m04CHQeCGnKR1PK2Yx2.png)


Lentoid CPU

http://i1.sendpic.org/t/t6/t6eICbkKQw85HNV8BkZiuh6MpCq.jpg (http://sendpic.org/view/1/i/x6OZDjn8t6JPWlHQfmRblcLFUch.png)

Paul Tronc
19th September 2016, 09:16
I'm back, to share my findings. I'm trying to optimize my CPU and GPU power in order to get the maximum HEVC decoding result.
For now, the best thing I can achieve on a 4K 40Mbps HEVC sample, using a 2600K@4.4 and GTX960, on a 1080p display :

- GPU HEVC decoding : OK.
- Madvr : Jinc and other "rather good" algoritms
- FRC (Frame Rate Conversion) using SVP, from 24p to 60p

The GPU assumes most Madvr + LAV work, while the CPU is used for frame interpolation.

I'm currently tuning the settings to get a rock stable smooth experience.
The 1080p output helps a lot, on my 1440p the rendering time grows quickly.

wanezhiling
7th October 2016, 14:15
http://i.imgur.com/CYpKdmp.jpg
http://i.imgur.com/opaTGVd.jpg
Pascal 8K VP9

NikosD
7th October 2016, 14:23
And 10bit VP9 ?

wanezhiling
7th October 2016, 14:28
Do you have 10-bit VP9 clips?

NikosD
7th October 2016, 14:44
No and I haven't found an encoder

aufkrawall
7th October 2016, 15:23
Any VP9 decoding still not enabled with Polaris?

huhn
7th October 2016, 18:28
here a 3 sec VP9 10 bit file: http://filehorst.de/download.php?file=bDyyChti
i encoded it with VPXENC

edit: i will check out 16.10.1 for polaris VP9 support

NikosD
7th October 2016, 18:32
You own a polaris card ?

huhn
7th October 2016, 18:45
yes and no VP9 support in 16.10.1.

NikosD
7th October 2016, 18:47
It would be useful if you could post H.264 & H.265 (8bit & 10bit) benchmarks of the usual clips of this thread.

Which one by the way ?

huhn
7th October 2016, 18:56
RX 480 4Gb.

it's broken or at least i have major problem with it to even get an Constant image have to replace before i can do any proper test.

i will do some test with the card in some of my other PC this weekend to make sure it is really the card. even though it is pretty clearly the card.

in short no performance tests for now. and most likely not any time soon.

NikosD
7th October 2016, 19:07
Do you have problems in 2D and 3D besides video ?

huhn
7th October 2016, 19:17
i get the problem in madVR and in browser. so it is generally.

nothing special just bad luck.

i could work around it and do some tests but testing this card it self is just more important right now.

trust me i'm very very upset about this.

NikosD
7th October 2016, 19:22
Brand ?

huhn
7th October 2016, 19:40
XFX.

i don't know why this should matter.