View Full Version : What is current status for hardware H.265 encoding.
pandy
16th September 2015, 12:43
I'm currently aware only of 2 relatively cheap technologies with H.265 HW encoding capabilities.
1. Nvidia NVENC http://developer.download.nvidia.com/compute/nvenc/v5.0_beta/NVENC_DA-06209-001_v06.pdf
2. Intel QuickSync https://software.intel.com/sites/default/files/managed/3b/f2/Intel_HEVCwhitepaper_v1.45_06Apr2015.pdf
My goal is to have real time encoding of H.265 with 4k, 50fps, 10 bit per component.
(but 8 bit depth for component is also acceptable - depends on many factors)
nevcairiel
16th September 2015, 13:27
AFAIK both NVENC and QuickSync are 8-bit only for HEVC encoding, at least on current hardware.
pandy
17th September 2015, 09:38
AFAIK both NVENC and QuickSync are 8-bit only for HEVC encoding, at least on current hardware.
Yep, NVENC seem to be limited to 8 bit, QSV is not clear to me as i have a bit confusing information's from Intel.
----
I will change conditions in my question to 8 and 10 per component.
kolak
17th September 2015, 13:00
There is this:
https://communities.intel.com/community/itpeernetwork/datastack/blog/2015/09/10/new-intel-visual-compute-accelerator-makes-its-debut
which is close to 2x realtime for UHD at decent quality apparently. Ittiam has ready solution which works with it.
stax76
17th September 2015, 16:14
I'm currently aware only of 2 relatively cheap technologies with H.265 HW encoding capabilities.
1. Nvidia NVENC http://developer.download.nvidia.com/compute/nvenc/v5.0_beta/NVENC_DA-06209-001_v06.pdf
2. Intel QuickSync https://software.intel.com/sites/default/files/managed/3b/f2/Intel_HEVCwhitepaper_v1.45_06Apr2015.pdf
My goal is to have real time encoding of H.265 with 4k, 50fps, 10 bit per component.
(but 8 bit depth for component is also acceptable - depends on many factors)
If you are looking for a GUI you can take a look at StaxRip, it supports H.265 hardware encoding using the command line tools QSVEncC and NVEncC.
H.265 encoding with QSVEncC requires Intel Skylake.
H.265 encoding with NVEncC requires NVIDIA Maxwell but I'm not sure if all Maxwell cards support it, it definitely works with GTX 960.
pandy
18th September 2015, 15:21
There is this:
https://communities.intel.com/community/itpeernetwork/datastack/blog/2015/09/10/new-intel-visual-compute-accelerator-makes-its-debut
which is close to 2x realtime for UHD at decent quality apparently. Ittiam has ready solution which works with it.
Thx for this - i saw this at IBC - Intel engineers says "is completely new solution only 2 day old"...
That's why i'm trying to collect all those puzzles - in case of Intel it is extremely difficult as they have different divisions talking about same product group.
If you are looking for a GUI you can take a look at StaxRip, it supports H.265 hardware encoding using the command line tools QSVEncC and NVEncC.
H.265 encoding with QSVEncC requires Intel Skylake.
H.265 encoding with NVEncC requires NVIDIA Maxwell but I'm not sure if all Maxwell cards support it, it definitely works with GTX 960.
GUI not needed if commands are sufficiently documented, from my perspective quality is also not so important - i would have available up to 40 - 50Mbps per stream and it will be used internally (so over 1GbE link).
More important is to accept some sources and provide basic ip streaming capabilities (ffmpeg/gstreamer looks like perfect tools from my perspective).
kolak
21st September 2015, 13:25
Thx for this - i saw this at IBC - Intel engineers says "is completely new solution only 2 day old"...
That's why i'm trying to collect all those puzzles - in case of Intel it is extremely difficult as they have different divisions talking about same product group.
Development was lead by Polish developer, Tomasz Madajczak (I think). It's quite easy to find direct email :)
pandy
22nd September 2015, 11:57
Development was lead by Polish developer, Tomasz Madajczak (I think). It's quite easy to find direct email :)
I've contacted with Fred Fan from Intel - He said that Intel VCA should be available from now on Intel distributors list...
Need to find one...
btw slightly ot but it will be appreciated to hear some opinions about Intel analyzer and conformance/stress bitstream library (and encoder).
pieter3d
22nd September 2015, 18:21
I actually wrote the original Intel Pro Analyzer from scratch when I worked there, so my opinion is of course biased. However I think it's quite good, I designed it to address the issues I've always had with other tools I've had to use before. The competing product Parabola Explorer is quite nice too.
Parabola
22nd September 2015, 19:02
... Parabola Explorer is quite nice too.
Thanks Pieter! We saw the Intel analyzer at IBC and it looked nice on a large 4k screen. In many ways it now seems remarkably similar to our Parabola Explorer but Intel is asking an order of magnitude more $ for a license.
Regarding the conformance streams, consider also those from Argon Design (no connection). Also there is a freely available set of conformance streams on the ITU website. Parabola and others have found gaps in the coverage of the ITU set and contributed our own streams to cover specific corner cases.
I understand Intel's HEVC encoding is reasonably good too although I have not had a chance to try it.
pandy
23rd September 2015, 13:49
Thx guys - my question related to analyzer will be easier to understand when You think MTS4xxx from Tek as a reference point (both capabilities and price) so in other words - searching something similar but more affordable from $$$ perspective.
Streams - this is slightly more complex than just $$$ (but reference is of course Allegro where company i work is close to buy - i'm not convinced if with this kind of money this most optimal solution - don't get me wrong - i know that this is very niche product and due of this high price) as i need fit inside particular distribution chain where short streams looping very bad (have no clue why - i have no access to system, can't verify or change anything) - that's why Intel configurable encoder where i can generate long time sequence with particular stress will be very nice to have - stress not only to decoder but also to remain part of system, also to measure power consumption and thermal signature...
iwod
23rd December 2015, 10:07
I'm currently aware only of 2 relatively cheap technologies with H.265 HW encoding capabilities.
1. Nvidia NVENC http://developer.download.nvidia.com/compute/nvenc/v5.0_beta/NVENC_DA-06209-001_v06.pdf
2. Intel QuickSync https://software.intel.com/sites/default/files/managed/3b/f2/Intel_HEVCwhitepaper_v1.45_06Apr2015.pdf
My goal is to have real time encoding of H.265 with 4k, 50fps, 10 bit per component.
(but 8 bit depth for component is also acceptable - depends on many factors)
From the Intel PDF the SDK is not available on Apple OSX AGAIN!
Any one could take a guess why Intel / Apple are disallowing the use of 3rd party QuickSync on OSX.
XMonarchY
30th December 2015, 03:02
So even if my HDTV is using 12bit output and the game I am recording uses deep color, like Alien: Isolation (10bit+), then will I still only get 8bit with NVENC?
pandy
4th January 2016, 11:46
So even if my HDTV is using 12bit output and the game I am recording uses deep color, like Alien: Isolation (10bit+), then will I still only get 8bit with NVENC?
Seem that yes - also QSV is 8 bit limited, for 10 bit seem CPU work in kind of hybrid mode...
From the Intel PDF the SDK is not available on Apple OSX AGAIN!
Any one could take a guess why Intel / Apple are disallowing the use of 3rd party QuickSync on OSX.
Isn't this tightly coupled with Apple business model?
Blue_MiSfit
13th April 2016, 01:46
I'm surprised with how good NVENC HEVC is! I encoded my usual 1080p24 test clip in 4 Mbps VBR (15 Mbps peak) and it looked surprisingly good.
The encode ran at 150 fps on my GTX 960 / 4 Ghz quad core Skylake.
JohnLai
13th April 2016, 04:48
I'm surprised with how good NVENC HEVC is! I encoded my usual 1080p24 test clip in 4 Mbps VBR (15 Mbps peak) and it looked surprisingly good.
The encode ran at 150 fps on my GTX 960 / 4 Ghz quad core Skylake.
......Be ready to cry with Nvenc HEVC lack of b-frame, SAO, 64 CTU size.
4Mbps FOR 1080P24? Use adaptive quantization and use twopass mode (VBR2Pass is really slow)
Note: I am not sure about skylake hevc encoding standard support because I don't have one.
benwaggoner
14th April 2016, 18:46
Note: I am not sure about skylake hevc encoding standard support because I don't have one.
I believe that Skylake is the first chipset to support 10-bit HEVC decode, so perhaps it's the first that might support 10-bit encode. But I don't actually know if it does.
I don't see any near-term change in the fundamental rule that HW and GPU based encoding can be great to get decent quality quickly or with low power, but that high-quality/efficiency encoding is primarily going to be done on CPU. There are some things about HEVC that may make GPU accelerated encoding more feasible than with H.264, but there are also lots more logical choices between the many more different ways to do things that need to be made, which happen on the CPU.
I'm more excited about AVX-512 to improve encoding speed (and thus allow higher quality with the same performance). It won't do much for H.264, but should be quite helpful with HEVC.
sneaker_ger
14th April 2016, 18:55
Skylake only does hybrid 10 bit HEVC decoding. Kaby Lake will bring fixed function 10 bit decoding. Don't know about encoding.
AVX-512 is really complicated. Some chips will have it but only partially. It seems we won't get widespread AVX-512 for end users until Cannonlake.
nevcairiel
14th April 2016, 23:23
Only high-end Skylake Xeons will have it (2017), Cannonlake will be the first consumer generation with AVX-512, at least thats rumored so far.
JohnLai
15th April 2016, 04:07
I believe that Skylake is the first chipset to support 10-bit HEVC decode, so perhaps it's the first that might support 10-bit encode. But I don't actually know if it does.
I don't see any near-term change in the fundamental rule that HW and GPU based encoding can be great to get decent quality quickly or with low power, but that high-quality/efficiency encoding is primarily going to be done on CPU. There are some things about HEVC that may make GPU accelerated encoding more feasible than with H.264, but there are also lots more logical choices between the many more different ways to do things that need to be made, which happen on the CPU.
I'm more excited about AVX-512 to improve encoding speed (and thus allow higher quality with the same performance). It won't do much for H.264, but should be quite helpful with HEVC.
Well, for some, it is all speed and quality regardless of file size (how much does 4TB HDD cost anyway?)
The only problem I have with hardware based encoder is the standard support. I was planning to buy new budget Skylake CPU to replace my old Pentium Dual Core E5700. But after reading Intel response about B-frame.......I kinda hesitate...
https://software.intel.com/en-us/forums/intel-media-sdk/topic/623600
So, what is Low Delay B-frames (LDB) or Generalized P/B (GPB) ???
hyongmin
1st June 2016, 04:51
Sorry, I didn't see the 2x number in this blog.
Would you please point it out for me? Thanks
There is this:
https://communities.intel.com/community/itpeernetwork/datastack/blog/2015/09/10/new-intel-visual-compute-accelerator-makes-its-debut
which is close to 2x realtime for UHD at decent quality apparently. Ittiam has ready solution which works with it.
kolak
1st June 2016, 20:50
??
I know people who have access to it.
FancyMouse
5th June 2016, 00:48
So, what is Low Delay B-frames (LDB) or Generalized P/B (GPB) ???
I have seen somewhere that the term low-delay B-frames means B frame referencing only previous frames (i.e. IPPB where B refers to P1P2, but not traditional IPBPI hierarchy). This is low-latency because such B frames only rely on already-output frames which don't increase latency.
Not sure it's the same thing in your context, though.
bajarwas
14th July 2016, 04:57
has anyone tested the encoding capabilities of the new graphics cars from Nvidia and AMD ?
GTX 10s series (1070/1080)
RX480
pandy
17th July 2016, 10:55
has anyone tested the encoding capabilities of the new graphics cars from Nvidia and AMD ?
GTX 10s series (1070/1080)
RX480
Nope... but AMD seem to be out of area of my interests (checked VCE specification and all information's available and it looks like no match to even current available consumer solutions i.e. Intel and NVidia - this is sad as nowadays even FPGA receive H.265 encoder as a part of SoC e.g. Xilinx Zynq) - i will try to buy 1060 (all i need is HW encoding) - currently use 980 and IMHO NVenc functionality is better than QSV - Intel is somehow less stable and more vague also less fps produced - my comment is not about quality but easy to use and how stable it is - big plus for NVidia on this (my goal is to create real time video services for testing - this don't need to be hq video and as i have almost 40Mbps available then it will be hq video).
IMHO 1060 (or 1070) should be first consumer HW encoder capable to provide 60 fps on H.265 (HEVC) - i observe in my setup that 980 is not capable to provide more than 30 fps on H.265 - QSV is slower than NVidia.
NikosD
19th July 2016, 21:21
Nope... but AMD seem to be out of area of my interests (checked VCE specification and all information's available and it looks like no match to even current available consumer solutions i.e. Intel and NVidia - this is sad as nowadays even FPGA receive H.265 encoder as a part of SoC e.g. Xilinx Zynq)
Polaris 10 (RX 480 & 470) and Polaris 11 (RX 460) offer 4K60fps HEVC HW encoding and 4K120fps HEVC HW decodong.
i will try to buy 1060 (all i need is HW encoding) - currently use 980 and IMHO NVenc functionality is better than QSV - Intel is somehow less stable and more vague also less fps produced - my comment is not about quality but easy to use and how stable it is - big plus for NVidia on this (my goal is to create real time video services for testing - this don't need to be hq video and as i have almost 40Mbps available then it will be hq video).
Skylake's HW HEVC encoder and Haswell and onwards HW H.264 encoder are far more flexible and with more encoding options than Nvidia or AMD.
Better quality and better speed using smaller files.
IMHO 1060 (or 1070) should be first consumer HW encoder capable to provide 60 fps on H.265 (HEVC) - i observe in my setup that 980 is not capable to provide more than 30 fps on H.265 - QSV is slower than NVidia.
Do you have a skylake to compare HW H.265 encoding with Nvidia HW H.265 encoding ?
Which app do you use ?
You should try QSVEncC in order to find out that Skylake has faster and better quality HW H.265 encoding than Nvidia.
Also as I wrote on my first post Polaris 10 and 11 have already brought us 4K60 fps HW H.265 encoding.
easyfab
19th July 2016, 22:14
Polaris 10 (RX 480 & 470) and Polaris 11 (RX 460)
Skylake's HW HEVC encoder and Haswell and onwards HW H.264 encoder are far more flexible and with more encoding options than Nvidia or AMD.
Better quality and better speed using smaller files.
And Skylake HEVC vs nvidia pascal with new NVENC SDK 7.0 ?
Which has the best quality ?
I see in Encoder Application Note -> https://developer.nvidia.com/nvidia-video-codec-sdk
new features for pascal cards like SAO ( with this comment : Significantly improves encoded video quality for HEVC. )
But no B-frames with HEVC :(
If someone can do tests It would be very interseting to see if it really give some nice quality boost. rigaya has new NVEnc 2.08 for that.
JohnLai
20th July 2016, 03:59
And Skylake HEVC vs nvidia pascal with new NVENC SDK 7.0 ?
Which has the best quality ?
I see in Encoder Application Note -> https://developer.nvidia.com/nvidia-video-codec-sdk
new features for pascal cards like SAO ( with this comment : Significantly improves encoded video quality for HEVC. )
But no B-frames with HEVC :(
If someone can do tests It would be very interseting to see if it really give some nice quality boost. rigaya has new NVEnc 2.08 for that.
~.~ and no 64x64 CU size too. Larger CU size is critical for 4K video.
Implemented look-ahead function is just being used to determine where to place I (HEVC and H264) and B (H264) frames optimally. Something like 'scenecut', it does its job well.....I-Frames are inserted during every scenechange from what I can check using HEVC bitstream analyser.
At least there is quality improvement from using look-ahead and Long-Term Reference pictures.
The only weird option from the SDK is ;
uint16_t targetQuality; /**< [in]: Target CQ (Constant Quality) level for VBR mode (range 0-51 with 0-automatic) */
It said, CQ for VBR, does this mean there is no need to specify vbr bitrate value anymore? How about the Initial QP value? Or QP min and Max?
Constant Quality VBR? --> Does this vary bitrate or quantizer for nvenc case? Documentation mentions Enables Constant Quality Mode where in video quality can be chosen by a quality factor.
pandy
23rd July 2016, 17:07
Polaris 10 (RX 480 & 470) and Polaris 11 (RX 460) offer 4K60fps HEVC HW encoding and 4K120fps HEVC HW decodong.
Source please as AMD is quite silent and only some bits where decoding is mentioned but encoding for 3840x2160@Main10 60fps is not mentioned at all.
Skylake's HW HEVC encoder and Haswell and onwards HW H.264 encoder are far more flexible and with more encoding options than Nvidia or AMD.
Better quality and better speed using smaller files.
Maybe yes but it is less stable than NVenc side to this seem performance is lower.
Do you have a skylake to compare HW H.265 encoding with Nvidia HW H.265 encoding ?
Yes, Intel Core i7-6700K
Which app do you use ?
ffmpeg
You should try QSVEncC in order to find out that Skylake has faster and better quality HW H.265 encoding than Nvidia.
I need to generate in real time transport stream and ffmpeg seem to be best for my needs - side to this i assume ffmpeg i just wrap-app around Intel plugin that need to be anyway explicitly loaded.
Also as I wrote on my first post Polaris 10 and 11 have already brought us 4K60 fps HW H.265 encoding.
And once again please provide me some reliable source for this information - for today NVidia provide best developer support, later Intel and AMD... well - silence.
Roph
27th August 2016, 08:44
Just thought I'd share this stuff here too since I don't see anyone else encoding H.265 on their AMD Polaris GPU yet.
Source Sample (http://dump.roph.eu/vce/hevc/Source%20(1080p).mkv) - ~15mbps H.264 1080p, bluray rip.
AMD HEVC VCE Encodes: (You will have to save these files and view them locally most likely, only Microsoft Edge seems to be able to play HEVC in-browser)
1080p: 3mbps (http://dump.roph.eu/vce/hevc/3mbit%20(1080p).mp4) / 2.5mbps (http://dump.roph.eu/vce/hevc/2.5mbit%20(1080p).mp4) / 2mbps (http://dump.roph.eu/vce/hevc/2mbit%20(1080p).mp4) [edit] 3mbps with 150 GOP (http://dump.roph.eu/vce/hevc/3mbit%20(1080p)%20(150%20GOP).mp4) / 3mbps with 300 GOP (http://dump.roph.eu/vce/hevc/3mbit%20(1080p)%20(300%20GOP).mp4)
720p: 1.5mbps (http://dump.roph.eu/vce/hevc/1.5mbit%20(720p).mp4) / 1mbps (http://dump.roph.eu/vce/hevc/1mbit%20(720p).mp4)
Example of settings:
http://dump.roph.eu/vce/hevc/Settings.png
What's also incredible is the encoding speed. Sure, GPU H265 won't match software x265, but at ~5-10fps on my CPU there is no comparison to getting 360fps on the GPU:
https://i.imgur.com/l2CJTAh.png
Encoding speed varies with resolution, over 500fps for low resolution stuff like DVD resolution. The 360fps I achieved in the screenshot was encoding 960x540. 720p yields about 220fps, 1080p about 120, 4K about ~40. However I think I'm software/CPU limited on some of these.
The Tool I'm using is called A's video converter, officially it only supports H.264 via AMD's VCE though the developer kindly provided a test build with H.265 support for polaris - they don't own a polaris GPU yet.
I assumed they used AMD's recently released Media SDK though apparently it doesn't support HEVC yet. Another developer did some poking and assumes they're querying the GPU through media foundation - which may be another encoding speed bottleneck factor.
Overall I'm very impressed with the quality. This also isn't the best quality that an AMD GPU can encode HEVC with, there is no use of the 2-Pass functionality which would greatly improve quality further.
This will probably be my only post here as the doom9 forum post requirements with these questions are ridiculous. Instead I'm happy to post more at videohelp: http://forum.videohelp.com/threads/380081-AMD-Polaris-(Radeon-RX-4xx)-H265-Encoding-Samples
NikosD
27th August 2016, 09:02
Indeed AMD has recently moved from its Media SDK to AMF (Advanced Media Framework) SDK which is part of GPUOPEN and it is open source.
It's OS and Framework agnostic and provides access to decoding UVD and encoding VCE fixed-function units along with pre/post processing.
From MediaSDK v1.1 we now have AMF v1.3 with HEVC support.
The runtime is inside the drivers for easy update and maintenance.
BTW the developer of A's converter is the same of the DXVA Checker.
Doom9 I think asks questions only the first time, then you could post freely.
Thanks for your post anyway.
JohnLai
27th August 2016, 11:58
Just thought I'd share this stuff here too since I don't see anyone else encoding H.265 on their AMD Polaris GPU yet.
So...if you ever return here........The sample:
No B-Frame.
No SAO.
Okay....., full PU 4x4 until 32x32 for Intra PU. As for Inter PU, 64x64 LCU available. But it is not complete.
Inter Frame PU sizes:
8x8
8x16
16x8
16x16
16x32
32x16
32x32
32x64
64x32
64x64
P-Frame still refers to 1 previous frame instead of multiple preceding frames. (Nvidia Nvenc also have similar problem/bug/feature)
Yups
29th August 2016, 19:37
IMHO 1060 (or 1070) should be first consumer HW encoder capable to provide 60 fps on H.265 (HEVC) - i observe in my setup that 980 is not capable to provide more than 30 fps on H.265 - QSV is slower than NVidia.
QSV HEVC with TU7 is easily faster than NVENC comparing 6700k and GTX 970. With Pascal may be Nvidia is the faster one.
And Skylake HEVC vs nvidia pascal with new NVENC SDK 7.0 ?
Which has the best quality ?
Pascal has better quality because even Maxwell has with the newest SDK. Intel needs Kabylake and a new Media SDK update with Lookahead support to regain the lead.
Yups
29th August 2016, 19:48
Encoding speed varies with resolution, over 500fps for low resolution stuff like DVD resolution. The 360fps I achieved in the screenshot was encoding 960x540. 720p yields about 220fps, 1080p about 120, 4K about ~40. However I think I'm software/CPU limited on some of these.
Pretty slow compared to Intel, i7-6700k @HD530 gets over 250 fps with TU7 there. GTX 970 over 160 fps. I didn't check quality.
pandy
29th August 2016, 20:30
QSV HEVC with TU7 is easily faster than NVENC comparing 6700k and GTX 970. With Pascal may be Nvidia is the faster one.
As i struggle still on HW encoding (still some options are not work as described or work completely opposite) perhaps you can share command line (ffmpeg) to compare NVenc vs QSV - for today my observations are rather solid that QSV is somehow slightl slower than NVEnc and also QSV seem to be less robust (NVEnc work like a charm - just start and almost imeediately encoding session started, Intel is somehow slow in starting and need some time between starting session - i assume to close previous session - not sure how this is related to ffmpeg itself).
NikosD
29th August 2016, 20:34
We never use ffmpeg for HW encoding.
Only rigaya's CLI tools usually via a nice GUI like StaxRip.
Rigaya and StaxRip provide the interface for the best HW encoding in terms of quality, speed and robustness for all platforms - AMD, Intel, Nvidia
ilovejedd
30th August 2016, 00:57
Well, for some, it is all speed and quality regardless of file size (how much does 4TB HDD cost anyway?)
In which case, probably best to stick to original source. I've yet to make a transparent encode at fast speed that wasn't the same size or even larger than the original file (1080p Blu-ray source).
I'm currently testing StaxRip/QSVEncC (ICQ at default settings) on my new Skylake laptop (i7-6500U). I'm using ICQ 20 and comparing frame by frame, there's noticeable smoothing on the encode. Of course, it's not really something I actually notice when watching the film.
It's not what I'd consider fast either. It encodes at ~1.5-2x real-time, sure, but I've been spoiled by the ~200FPS I get with Handbrake/Intel QSV H.264 (High@L4.0, Balanced CQ 20). On the upside, ICQ 20 on H.265 QSV seems closer in quality to ICQ 18 on H.264 QSV at just ~70% file size so that's definitely good.
Mind, my only reason for encoding is playback on mobile devices so speed and small file size at OK quality is my primary objective.
Roph
30th August 2016, 09:20
So...if you ever return here........The sample:
No B-Frame.
No SAO.
Okay....., full PU 4x4 until 32x32 for Intra PU. As for Inter PU, 64x64 LCU available. But it is not complete.
Inter Frame PU sizes:
8x8
8x16
16x8
16x16
16x32
32x16
32x32
32x64
64x32
64x64
P-Frame still refers to 1 previous frame instead of multiple preceding frames. (Nvidia Nvenc also have similar problem/bug/feature)
How were you able to anaylze the file? Mediainfo / streameye don't show such info, I'd like to poke around some more.
Pretty slow compared to Intel, i7-6700k @HD530 gets over 250 fps with TU7 there. GTX 970 over 160 fps. I didn't check quality.
The author is making the GPU encode HEVC somewhat of a hack job via the outdated Media Foundation interface which predates polaris/H.265. True HEVC support will come soon in AMD's recently released media SDK, which I assume would run at the full performance. Media foundation has performance issues even with H264, which is why the built-in VCE support of OBS Studio is derided due to its poor performance.
In some cases, I'm also CPUlimited rather than GPU limited. FX-8320 :(
(Still the ridiculous post question limit, come on - this isn't 1998 :/ )
pandy
30th August 2016, 10:21
We never use ffmpeg for HW encoding.
Only rigaya's CLI tools usually via a nice GUI like StaxRip.
Rigaya and StaxRip provide the interface for the best HW encoding in terms of quality, speed and robustness for all platforms - AMD, Intel, Nvidia
Ok, i use synthetic video - internal ffmpeg source as such files can't provide required functionality, also i don't need gui.
hajj_3
30th August 2016, 10:51
Have any of you guys tried comparing the best quality h265 hardware encode settings with the best quality software encoded x264 settings to see which offers the best quality and filesize? If hardware h265 looks better then i think a lot of people would want to switch as they would save a lot of time.
JohnLai
30th August 2016, 15:10
How were you able to anaylze the file? Mediainfo / streameye don't show such info, I'd like to poke around some more.
https://software.intel.com/en-us/intel-video-pro-analyzer
Just use the trial version to find out everything.....
hajj_3
30th August 2016, 15:25
Intel just announced the Kaby Lake media capabilities: http://www.anandtech.com/show/10610/intel-announces-7th-gen-kaby-lake-14nm-plus-six-notebook-skus-desktop-coming-in-january/3
It supports full fixed function 8-bit encode and 8/10-bit decode for VP9 codec and 8/10bit encode and decode for h265.
Yups
30th August 2016, 16:42
As i struggle still on HW encoding (still some options are not work as described or work completely opposite) perhaps you can share command line (ffmpeg) to compare NVenc vs QSV - for today my observations are rather solid that QSV is somehow slightl slower than NVEnc and also QSV seem to be less robust (NVEnc work like a charm - just start and almost imeediately encoding session started, Intel is somehow slow in starting and need some time between starting session - i assume to close previous session - not sure how this is related to ffmpeg itself).
ffmpeg is slow, no wonder. Make sure you are using the fixed function decoder for encoding and QSVEnc is the best.
Intel just announced the Kaby Lake media capabilities: http://www.anandtech.com/show/10610/intel-announces-7th-gen-kaby-lake-14nm-plus-six-notebook-skus-desktop-coming-in-january/3
It supports full fixed function 8-bit encode and 8/10-bit decode for VP9 codec and 8/10bit encode and decode for h265.
Intel also said that the HEVC Encoding quality is more than 10% improved over Skylake. The hardware is fine, they only have to add Lookahead.
JohnLai
30th August 2016, 16:51
Intel also said that the HEVC Encoding quality is more than 10% improved over Skylake. The hardware is fine, they only have to add Lookahead.
Hardware is not fine.
Aside from Lookahead, Intel should add P-Frame support, SAO and all PU sizes. ~.~
For some unknown reason, all big 3 (Intel, Nvidia and AMD) hardware encoders have problem encoding dark / black scene properly. I can see those big dark blocky artifacts on dark scene even with 15Mbps on 1080p. Bright scene is perfect though.....
Yups
30th August 2016, 17:04
Hardware is fine imho. It will never support the same features as a software encoder. I'm only interested in the end result and Intels biggest problem with HEVC is scene change. H264 is better in this regard because there is Lookahead for VBR or ICQ. Especially with low bitrate videos this is a big problem. Lookahead is just a software thing. Aside from this HEVC is very much better than H264 from Intel.
ilovejedd
31st August 2016, 00:17
For some unknown reason, all big 3 (Intel, Nvidia and AMD) hardware encoders have problem encoding dark / black scene properly. I can see those big dark blocky artifacts on dark scene even with 15Mbps on 1080p.
That's not limited to HEVC though. My QSVEncC and Handbrake H.264 QuickSync encodes exhibit the same issue. At similar bitrates, blocking on QSV HEVC actually isn't as bad as on QSV H.264.
Yups
10th September 2016, 11:49
I changed my GTX 970 for a GTX 1080. One thing I can say is that the performance of 8 bit HEVC video encoding is much improved.
GTX 970
Video 1: 163 fps
Video 2: 175 fps
GTX 1080
Video 1: 288 fps +76,7%
Video 2: 308 fps +76%
The super high clock rate of Pascal is the main reason in order to speed up the video engine. 1962 Mhz for my GTX 1080 and 1291 Mhz for GTX 970.
NikosD
10th September 2016, 11:58
GTX 970 has a hybrid HEVC decoder and it's a lot slower than 960 GTX fixed-fuction decoder.
You should compare Pascal against 960 GTX Maxwell.
JohnLai
10th September 2016, 13:21
I changed my GTX 970 for a GTX 1080. One thing I can say is that the performance of 8 bit HEVC video encoding is much improved.
GTX 970
Video 1: 163 fps
Video 2: 175 fps
GTX 1080
Video 1: 288 fps +76,7%
Video 2: 308 fps +76%
The super high clock rate of Pascal is the main reason in order to speed up the video engine. 1962 Mhz for my GTX 1080 and 1291 Mhz for GTX 970.
Hmm? Something isn't right with your gtx 970 figures.
Did you use CUVID to decode the source video? Or CPU to decode?
CUVID ---> Zero-copy to NVENC for encoding assuming if the video is DXVA-decodable.
You could easily get 200 - 220 fps on 1920x1080p source video to 1920x1080p encoded video.
Simplified:
CPU decoding = HDD--->cpu--->RAM--->GPU--->HDD
Pure hardware GPU decode + encode = HDD ---> GPU--->HDD
Yups
10th September 2016, 15:00
I have been using NVEncC (avcuvid native) in Staxrip which is the fastest alongside QSVEnC (Intel). I don't think there is something wrong.
JohnLai
10th September 2016, 15:08
I have been using NVEncC (avcuvid native) in Staxrip which is the fastest alongside QSVEnC (Intel). I don't think there is something wrong.
Hmm.........in that case, I have no further comment. :cool:
aegisofrime
11th September 2016, 04:09
I have a GTX 1070 as well and am wondering how best to tune it for maximum quality while maintaining decent speed.
JohnLai
11th September 2016, 04:18
I have a GTX 1070 as well and am wondering how best to tune it for maximum quality while maintaining decent speed.
Use staxrip default CQP rate control plus --lookahead 32 and 10bit HEVC encoding...
Or.... a variation of VBR Constant Quality ---> select normal VBR (not vbr2 as nvidia said it is designed for low latency 2 pass), set maximum bitrate 17500kbps (don't worry, it won't actually use 17500kbps) --qp-init 1 --lookahead 32 + 10bit encoding --aq --vbr-quality 26 (26 or 25) should be sufficient
Note: In order to use VBRCQ (known as --vbr-quality) , must use VBR, max bitrate 17500 and --qp-init 1.
-lookahead 32 and --aq are optional, but lookahead and adaptive quantization improve quality,so why not?
**VBR2 doesn't work in conjunction with --vbr-quality, it will maxed out the 17500 bitrate instead of readjusting the bitrate/quality based on --vbr-quality)
RainyDog
12th September 2016, 12:00
Use staxrip default CQP rate control plus --lookahead 32 and 10bit HEVC encoding...
Or.... a variation of VBR Constant Quality ---> select normal VBR (not vbr2 as nvidia said it is designed for low latency 2 pass), set maximum bitrate 17500kbps (don't worry, it won't actually use 17500kbps) --qp-init 1 --lookahead 32 + 10bit encoding --aq --vbr-quality 26 (26 or 25) should be sufficient
Note: In order to use VBRCQ (known as --vbr-quality) , must use VBR, max bitrate 17500 and --qp-init 1.
-lookahead 32 and --aq are optional, but lookahead and adaptive quantization improve quality,so why not?
**VBR2 doesn't work in conjunction with --vbr-quality, it will maxed out the 17500 bitrate instead of readjusting the bitrate/quality based on --vbr-quality)
Hi John, do you know what I need to do in order to get the new Pascal encoding options to work through StaxRip on my GTX 1060 please?
If I try adding --main 10 or --lookahead 32 in the custom command lines box then it just comes up with an error and won't start encoding.
I expect it's because the latest StaxRip test build doesn't contain the latest rigaya CLI?
Thanks.
JohnLai
12th September 2016, 16:37
Hi John, do you know what I need to do in order to get the new Pascal encoding options to work through StaxRip on my GTX 1060 please?
If I try adding --main 10 or --lookahead 32 in the custom command lines box then it just comes up with an error and won't start encoding.
I expect it's because the latest StaxRip test build doesn't contain the latest rigaya CLI?
Thanks.
Yes. Just replace the Nvencc with the latest build from rigaya http://rigaya34589.blog135.fc2.com/blog-entry-814.html
Don't forget to update your driver too. Minimum driver version is 368.69
Edit: By the way....why you add "--main 10"? I thought rigaya nvencc command should be "--profile main10"?
aegisofrime
12th September 2016, 17:05
Yes. Just replace the Nvencc with the latest build from rigaya http://rigaya34589.blog135.fc2.com/blog-entry-814.html
Don't forget to update your driver too. Minimum driver version is 368.69
Edit: By the way....why you add "--main 10"? I thought rigaya nvencc command should be "--profile main10"?
Thanks for your reply. I actually tried it, speeds are great but unsurprisingly quality still falls far below x265. Well, can't have the best of both worlds I guess.
Anyway, the switch for 10-bit should be --output-depth 10.
JohnLai
12th September 2016, 17:33
Thanks for your reply. I actually tried it, speeds are great but unsurprisingly quality still falls far below x265. Well, can't have the best of both worlds I guess.
Anyway, the switch for 10-bit should be --output-depth 10.
Can't blame me for not knowing :D After all, I don't own pascal gpu. (Only maxwell gpu for me)
Can't expect much from hardware encoders.
For example, x265 crf 20 at very fast preset for live action film normally encoded with average QPI 18, QPP 20, QPB 23 or so depending on the content.
Now, since nvenc doesn't support B-frame, you might wanna adjust Nvencc CQP to QPI 18 and QPP 20 respectively. Unfortunately, the file size is going to be quite big.
General compression ratio for each frame is:
I : 100% normally non-compressed or slightly compressed.
P : 50% of I frame size
B : 25% of I frame size
Based on above.....let say we have GOP (group of picture) of 240 with IPBBB type (3 B-frames) where only I frame being inserted one time for each 240)
We got:
1 I -frame = 1Mb
59.75 P-frames = 0.5Mb X 59.75 = 29.875Mb
179.25 B-frames = 0.25Mb X 179.25 = 44.8125Mb
Total sizes = 1Mb + 29.875Mb + 44.8125Mb = 75.6875Mb
Without B-frame:
1 I-frame = 1Mb
239 P-frames = 0.5Mb x 239 = 119.5Mb
*Note: there is no fractional kind of frame, above is just a rough calculation
[(119.5 / 75.6875 ) -1] X 100% = 57.886% larger size......
Edit: Only Intel QSV HEVC kinda follow the rule of 100%,50% and 25% by using B-frame as P-frame, my finding here http://forum.doom9.org/showpost.php?p=1775316&postcount=1474
trip_let
12th September 2016, 20:24
Quick confirmation: only Pascal (Nvidia 10 series) and Kaby Lake (Intel 7 series Core) so far has any kind of total or hybrid hardware 10 bit HEVC encode? For consumer stuff, that is.
NikosD
12th September 2016, 20:26
Polaris has also HEVC 10 bit pure HW encoding
Roph
12th September 2016, 20:35
Polaris has also HEVC 10 bit pure HW encoding
You sure about that? I've only seen 10-bit HEVC decoding through UVD. No mention of encoding with VCE.
Also, this AMD employee seems certain that Polaris actually has a regression in that it does not support B-frames when encoding: https://github.com/GPUOpen-LibrariesAndSDKs/AMF/issues/8
I'm most interested in the 2-pass encoding on Polaris, though as of yet it's still AWOL in their SDK.
JohnLai
13th September 2016, 03:46
Quick confirmation: only Pascal (Nvidia 10 series) and Kaby Lake (Intel 7 series Core) so far has any kind of total or hybrid hardware 10 bit HEVC encode? For consumer stuff, that is.
Pascal supports 10bit hevc encoding in fully hardware mode (The only other functionality being offloaded to CUDA cores are lookahead, two pass rate control and adaptive quantization stuff --> doesn't use much of CUDA cores, around 4-10% for 4k 8bit hevc encode with CQP mode for lookahead using GTX970 )
Upcoming Kaby Lake claims to support 10bit HEVC encode, since product isn't out yet.....nobody can be sure....
You sure about that? I've only seen 10-bit HEVC decoding through UVD. No mention of encoding with VCE.
Also, this AMD employee seems certain that Polaris actually has a regression in that it does not support B-frames when encoding: https://github.com/GPUOpen-LibrariesAndSDKs/AMF/issues/8
I'm most interested in the 2-pass encoding on Polaris, though as of yet it's still AWOL in their SDK.
https://www.youtube.com/watch?v=hvD37UUcdIo&feature=youtu.be&t=154
Polaris should support 10bit HEVC encoding...... Unless if my hearing got problem....
About AMD B-frame....it is kinda weird for Polaris series not to support B-Frame for H264...so...since you have polaris....why not try an h264 encoding with B-frame command specified by using Rigaya VCEencc, then use the bitstream analyzer to find out?
Roph
13th September 2016, 06:44
https://www.youtube.com/watch?v=hvD37UUcdIo&feature=youtu.be&t=154
Polaris should support 10bit HEVC encoding...... Unless if my hearing got problem....
About AMD B-frame....it is kinda weird for Polaris series not to support B-Frame for H264...so...since you have polaris....why not try an h264 encoding with B-frame command specified by using Rigaya VCEencc, then use the bitstream analyzer to find out?
Yeah I've interacted with Robert a few times on reddit, he's an AMD Marketing guy. He has mentioned VCE and its capabilities a couple times in his comments, even mentioned he made a few of AMD's presentation slides.
Unfortunatly a few developers of third party software - and an AMD developer who actually works on their VCE Media SDK (See that github issue link) all say that Polaris lacks B-frames support.
Querying the GPU's capabilities via the SDK shows B-frames are not supported. Asking it to encode them anyway results in errors thrown.
Funnily enough I read up on VCEenc recently, it doesn't seem to have been updated yet to support AMD's new Media SDK; the author notes that AMD's SDK hasn't been updated since Jan 2015 in their blog post. Nevertheless I tried forcing b-frames with VCEenc via staxrip, and nope: (Staxrip's VCEEnc is outdated, I dropped in the latest 2.0 binary)
-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-
Encoding using VCEEncC 1.03v2 x64
-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-
F:\tmp\Staxrip\Apps\VCEEncC\VCEEncC64.exe --quality slow --gop-len 120 --bframes 3 --max-bitrate 9000 --vbv-bufsize 9000 --vbr 7090 -i "D:\Work\Encode\VCE\HEVC Test\Media\Source (1080p)_temp\Source (1080p)_new.avs" -o "D:\Work\Encode\VCE\HEVC Test\Media\Source (1080p)_temp\Source (1080p)_new_out.h264"
VCEEnc 2.00 (x64) / Windows 10 (x64)
CPU: AMD FX(tm)-8320 Eight-Core Processor (4C/8T)
GPU: Radeon RX 470 Graphics [Ellesmere 2000MHz (2117.13 (VM))]
Input Info: Avisynth 2.60 yv12->nv12[AVX], 1920x1024p, 24000/1001 fps
Output: H.264/AVC High @ Level 4.1
1920x1024p 23.976fps (24000/1001fps)
Quality: slow
VBR: 7090 kbps, Max 9000 kbps
QP: Min: 0, Max: 51
VBV Bufsize: 9000 kbps
Bframes: 0 frames, b-pyramid: off
Motion Est: Q-pel
Slices: 1
GOP Len: 120 frames
Others: deblock hrd
encoded 1280 frames, 59.59 fps, 7222.98 kbps, 45.97 MB
encode time 0:00:21, CPULoad: 20.44
m_encoder->SetProperty(BPicturesDeltaQP) failed Error:AMF_ALREADY_INITIALIZED
m_encoder->SetProperty(BPicturesPattern) failed Error:AMF_OUT_OF_RANGE
m_encoder->SetProperty(ReferenceBPicturesDeltaQP) failed Error:AMF_ALREADY_INITIALIZED
Start: 06:38:43 AM
End: 06:39:06 AM
Duration: 00:00:22
General
Complete name : D:\Work\Encode\VCE\HEVC Test\Media\Source (1080p)_temp\Source (1080p)_new_out.h264
Format : AVC
Format/Info : Advanced Video Codec
File size : 46.0 MiB
Video
Format : AVC
Format/Info : Advanced Video Codec
Format profile : High@L4.1
Format settings, CABAC : Yes
Format settings, ReFrames : 4 frames
Format settings, GOP : M=1, N=60
Width : 1 920 pixels
Height : 1 024 pixels
Display aspect ratio : 1.85:1
Frame rate : 23.976 (24000/1001) fps
Standard : Component
Color space : YUV
Chroma subsampling : 4:2:0
Bit depth : 8 bits
Scan type : Progressive
Color range : Limited
JohnLai
13th September 2016, 07:29
Yeah I've interacted with Robert a few times on reddit, he's an AMD Marketing guy. He has mentioned VCE and its capabilities a couple times in his comments, even mentioned he made a few of AMD's presentation slides.
Unfortunatly a few developers of third party software - and an AMD developer who actually works on their VCE Media SDK (See that github issue link) all say that Polaris lacks B-frames support.
Querying the GPU's capabilities via the SDK shows B-frames are not supported. Asking it to encode them anyway results in errors thrown.
Funnily enough I read up on VCEenc recently, it doesn't seem to have been updated yet to support AMD's new Media SDK; the author notes that AMD's SDK hasn't been updated since Jan 2015 in their blog post. Nevertheless I tried forcing b-frames with VCEenc via staxrip, and nope: (Staxrip's VCEEnc is outdated, I dropped in the latest 2.0 binary)
Oh darn, what a setback, not having B-Frame for HEVC is one thing, dropping B-Frame support for H264 is another when previous iteration of VCE supported B-frame for H264 just fine.
:(
trip_let
13th September 2016, 20:38
Okay, thanks for all the info.
I just tried HEVC encode on NVEnc via Staxrip on my GTX 960 using the VBR setting and the quality for a given size was still noticeably worse than x264 @ 8 bits, medium preset on the one sample I looked at. I used CRF on x264 with the parameter changed to match the bitrate of the HEVC encode (came out to CRF 22.4). Maybe I used the wrong settings, but the difference wasn't that small. Kind of disappointing other than the blazing speed, seeing as the medium and fast-ish presets in x264 are already plenty fast enough for my needs.
Well, at least there's lower power consumption, and not every system has a 4C8T kind of processor at high clocks.
bladerunner1982
12th January 2017, 13:54
Hi,
can anyone post some hardware encoded Kaby Lake HEVC samples?
What fps do you get at what quality settings?
What platform has the biggest potential to match x265 quality in the near future? Intel, Nvidia or AMD?
Thanks...
CruNcher
14th January 2017, 14:55
Nvidia leads Discrete Balance wise and Intel leads on the Internal CPU side and Quality :)
AMD is a good question they just trying to fix a lot so far it doesn't make a bad impression comparable to Nvidias state but far away from Intel yet for the Power Consumption this wont change neither with Ryzen if AMD doesn't improved it's APU Encoder significantly.
But the Main Goals are still low latency for Nvidia and AMD so i wouldn't expect big steps anyways and AMD is already marketing CPU Encoding with their super efficient Internal Async comunication (so you buy their 8 core Ryzen) so you see that they themselves prepare no big jumps at all the biggest jump was the 2 pass and lookahead with Polaris but if you remember Nvidia had this already with Maxwell so not really anything interesting ;)
AMD just made the move to improve it's Streaming Quality Efficiency with Polaris and is now pretty much on par with Nvidia.
I don't think we will see that much more improvements for VEGA and RyZen alone though i hope i be wrong ;)
But RyZen in combination with VEGA that becomes something to look out for indeed ;)
Heaviest overhead for Nvidia is indeed the two pass rate control and it's visual efficiency overall is questionable (vs the performance impact) depending on the bitrate target and your actual goal.
ShogoXT
17th January 2017, 23:37
https://github.com/GPUOpen-LibrariesAndSDKs/AMF/releases
AMD just released AMF 1.4. I am eager to see how it's hevc encoding stacks up finally.
Edit: No b frames and now no 10 bit encoding. I'm starting to feel burned having bought a Rx 480. It was one of my reasons for buying this product. I did not find out until months later about these feature regressions.
JohnLai
18th January 2017, 04:03
https://github.com/GPUOpen-LibrariesAndSDKs/AMF/releases
AMD just released AMF 1.4. I am eager to see how it's hevc encoding stacks up finally.
Edit: No b frames and now no 10 bit encoding. I'm starting to feel burned having bought a Rx 480. It was one of my reasons for buying this product. I did not find out until months later about these feature regressions.
Hope someone provides a sample with 7 reference frames (Polaris) for me to analyse after this....to verify if the P frame can use multiple preceding frames.
Before AMF1.4 is out, every VCE sample I verified only make use of single reference frame just like nvidia nvenc.
NikosD
18th January 2017, 08:54
I have already contacted rigaya to take a look on the new AMF v1.4, although I'm pretty sure he had already seen that.
The moment he releases his updated VCEENC with HEVC support, I'll post here a few samples with different encoding options and speeds.
JohnLai
18th January 2017, 10:11
I have already contacted rigaya to take a look on the new AMF v1.4, although I'm pretty sure he had already seen that.
The moment he releases his updated VCEENC with HEVC support, I'll post here a few samples with different encoding options and speeds.
I am interested with HEVC_MAX_NUM_REFRAMES, HEVC_RATE_CONTROL_PREANALYSIS_ENABLE and HEVC_ENABLE_VBAQ as well as type of supported motion partitions.
Hmmm.....documentation doesn't mention anything about SAO and CU size. Only bilinear and bicubic hardware resizer.
But based on A's Video Converter sample without proper AMF last time.....I don't put too much hope on it.
But UVD can support AMFVideoDecoderHW_H265_MAIN10. Kinda weird for encoder section not to mention anything about HEVC 10bit encoding.
:confused:
*by the way, does Kaby lake hevc encoder support SAO and 64x64 CU yet?*
NikosD
18th January 2017, 11:26
But UVD can support AMFVideoDecoderHW_H265_MAIN10. Kinda weird for encoder section not to mention anything about HEVC 10bit encoding.
:confused:
It seems that HEVC encoder is only 8 bit for Polaris.
Maybe the upcoming VEGA can do something better on this.
https://github.com/GPUOpen-LibrariesAndSDKs/AMF/issues/51#issuecomment-269660790
NikosD
18th January 2017, 11:43
For some unknown reason, all big 3 (Intel, Nvidia and AMD) hardware encoders have problem encoding dark / black scene properly. I can see those big dark blocky artifacts on dark scene even with 15Mbps on 1080p. Bright scene is perfect though.....
That was an issue for AMD that managed to fix in 16.9.2 drivers according to this:
https://github.com/GPUOpen-LibrariesAndSDKs/AMF/issues/28#issue-179300712
NikosD
19th January 2017, 22:15
But the Main Goals are still low latency for Nvidia and AMD so i wouldn't expect big steps anyways and AMD is already marketing CPU Encoding with their super efficient Internal Async comunication (so you buy their 8 core Ryzen) so you see that they themselves prepare no big jumps at all the biggest jump was the 2 pass and lookahead with Polaris but if you remember Nvidia had this already with Maxwell so not really anything interesting ;)
Indeed.
Infinity Fabric (what a terrible marketing name) will make a difference according to AMD in the interconnections in RyZen and VEGA (and all the next products actually)
x264 and possibly x265 (although that has a lot of AVX2 optimizations that are slower to RyZen compared with Haswell and onward) have been demonstrated by AMD as applications that favor RyZen over Intel HEDT processors.
But probably VEGA has or should have some new features/ better performance in HW encoding that Polaris hasn't.
sephirotic
20th January 2017, 02:48
Wrong post.
Pitou
20th January 2017, 21:39
Hello all,
It's been a long time since I came here and I have a little question.
I'm playing with hevc and tried x265 and NVEnc (using ffmpeg) in Archlinux, trying to re-encode a 1080p VC1 movie.
I was reading some of the @JohnLai posts and got some great hints.
However, I would like to use NVEnc has it is faster. I'm using it with a Nvidia 1060 (Pascal) video card.
What would be the absolute best settings to have the best quality and reasonable filesize?
I can try using StaxRip instead of ffmpeg if needed.
One thing is that with ffmpeg I can specify "-hwaccel cuvid -c:v vc1_cuvid", so that he GPU is used to decode the source. I get an incredible speed gain using this.
Can StaxRip do this as well?
Does StaxRip (NVEncC) have more options that will result in better quality than ffmpeg?
Here are 2 sample cmds I tried to encode the video:
ffmpeg -hwaccel cuvid -c:v vc1_cuvid -i movie.mkv -qmin 0 -qmax 20 -preset slow -rc vbr_2pass -rc-lookahead 32 -c:v hevc_nvenc -c:a copy movieout.mkv
ffmpeg -hwaccel cuvid -c:v vc1_cuvid -i movie.mkv -preset medium -profile:v main10 -spatial_aq 1 -rc-lookahead 32 -rc constqp -global_quality 22 -c:v hevc_nvenc -c:a copy movieout.mkv
Thank you.
Pitou!
JohnLai
21st January 2017, 03:49
Hello all,
It's been a long time since I came here and I have a little question.
I'm playing with hevc and tried x265 and NVEnc (using ffmpeg) in Archlinux, trying to re-encode a 1080p VC1 movie.
I was reading some of the @JohnLai posts and got some great hints.
However, I would like to use NVEnc has it is faster. I'm using it with a Nvidia 1060 (Pascal) video card.
What would be the absolute best settings to have the best quality and reasonable filesize?
I can try using StaxRip instead of ffmpeg if needed.
One thing is that with ffmpeg I can specify "-hwaccel cuvid -c:v vc1_cuvid", so that he GPU is used to decode the source. I get an incredible speed gain using this.
Can StaxRip do this as well?
Does StaxRip (NVEncC) have more options that will result in better quality than ffmpeg?
Here are 2 sample cmds I tried to encode the video:
ffmpeg -hwaccel cuvid -c:v vc1_cuvid -i movie.mkv -qmin 0 -qmax 20 -preset slow -rc vbr_2pass -rc-lookahead 32 -c:v hevc_nvenc -c:a copy movieout.mkv
ffmpeg -hwaccel cuvid -c:v vc1_cuvid -i movie.mkv -preset medium -profile:v main10 -spatial_aq 1 -rc-lookahead 32 -rc constqp -global_quality 22 -c:v hevc_nvenc -c:a copy movieout.mkv
Thank you.
Pitou!
Hmm, staxrip + nvencc. Yes, nvencc has CUVID + NPP resizers too.
http://forums.guru3d.com/showthread.php?t=411509
Kinda lazy to re-type everything....read from beginning until the end. Problem with ffmpeg default high quantizers value for I & P (global_quality flag + lookahead issue).
Pitou
21st January 2017, 04:32
Thanks very much for the reply, I'll read the entire thread for sure.
In the meantime, I tried StaxRip + nvencc for my VC1 movie. So far I'm decoding it in software because when trying to use cuvid, I'm getting this error:
avcuvid: codec
avcuvid: unable to decode by cuvid.
Failed to open input file.
Is it because ncencc doesn't support hardware decoding using cuvid for VC1?
Would it be ok for a h264 movie?
Thanks again!
Pitou!
JohnLai
21st January 2017, 04:59
Thanks very much for the reply, I'll read the entire thread for sure.
In the meantime, I tried StaxRip + nvencc for my VC1 movie. So far I'm decoding it in software because when trying to use cuvid, I'm getting this error:
avcuvid: codec
avcuvid: unable to decode by cuvid.
Failed to open input file.
Is it because ncencc doesn't support hardware decoding using cuvid for VC1?
Would it be ok for a h264 movie?
Thanks again!
Pitou!
Eh? By right, Pascal should be able to decode VC1 codec.
https://developer.nvidia.com/nvidia-video-codec-sdk under NVDEC - Hardware-Accelerated Video Decoding section
Strange indeed.
Well, there are two CUVID modes for nvencc decoder, one is "Native" and the other is "CUDA". Have you try both?
Pardon me, I just read rigaya nvencc documentation and it appears the developer doesn't implement VC-1 decoding support
http://rigaya34589.blog135.fc2.com/blog-entry-739.html
The developer only enabled CUVID decoding for MPEG1, MPEG2 and H.264/AVC
Pitou
21st January 2017, 05:08
Yes tried both, but native or avisynth is rather slow. Getting 40fps instead of around 250fps with cuda.
Just tried with a h264 source and it works fine. I'm getting around 250fps.
With ffmpeg, vc1_cuvid works fine to decode VC1 with cuda. I'm getting around 250fps there also.
(ffmpeg -hwaccel cuvid -c:v vc1_cuvid)
Pitou!
JohnLai
21st January 2017, 05:10
Yes tried both, but native or avisynth is rather slow. Getting 40fps instead of around 250fps with cuda.
Just tried with a h264 source and it works fine. I'm getting around 250fps.
With ffmpeg, vc1_cuvid works fine to decode VC1 with cuda. I'm getting around 250fps there also.
(ffmpeg -hwaccel cuvid -c:v vc1_cuvid)
Pitou!
How about :
http://i.imgur.com/hwfccMc.jpg
Pitou
21st January 2017, 12:05
I get much better speed now, about 80fps. I'll use that for VC1 sources and cuvid for h264 sources
Thanks for the hint!
Pitou!
JohnLai
21st January 2017, 13:51
I get much better speed now, about 80fps. I'll use that for VC1 sources and cuvid for h264 sources
Thanks for the hint!
Pitou!
Only 80fps with ffmpeg(dxva) for VC1? :confused: Got a feeling ffmpeg auto-select wrong GPU for decoding. (Maybe ffmpeg dxva auto select intel igpu, there is a bug with ffmpeg dxva decoding using Intel IGPU)
The copy-back operation shouldn't be that taxing.
EDIT:
If you have intel integrated GPU enabled (plus using windows 8/10)....you can select "QSVEncC (Intel)" as decoder.
It turned out QSVENCC decoder supports MPEG2, H264, VC1 too.
Note: Intel hardware decoder is the fastest compared to Nvidia and AMD.
~Using Intel IGPU decoder to pipe the decoded video to Nvidia Nvenc for encoding.~
CruNcher
21st January 2017, 17:22
The TS Parser that Rigaya uses by default in Nvenc 3.02 makes me crazy.
it fails with
Complete name : F:\hevc\10bit\fail wip\Samsung_UHD_Ride_on_Board.ts
Format : MPEG-TS
File size : 992 MiB
Duration : 2 min 41 s
Overall bit rate mode : Constant
Overall bit rate : 51.6 Mb/s
Video
ID : 257 (0x101)
Menu ID : 1 (0x1)
Format : HEVC
Format/Info : High Efficiency Video Coding
Format profile : Main 10@L5.1@High
Codec ID : 36
Duration : 2 min 40 s
Width : 3 840 pixels
Height : 2 160 pixels
Display aspect ratio : 16:9
Frame rate : 59.940 (60000/1001) FPS
Color space : YUV
Chroma subsampling : 4:2:0
Bit depth : 10 bits
Writing library : ATEME Titan KFE 3.6.2 (4.6.1.9)
Audio
ID : 258 (0x102)
Menu ID : 1 (0x1)
Format : AAC
Format/Info : Advanced Audio Codec
Format version : Version 4
Format profile : LC
Muxing mode : ADTS
Codec ID : 15
Duration : 2 min 40 s
Channel(s) : 2 channels
Channel positions : Front: L R
Sampling rate : 48.0 kHz
Frame rate : 46.875 FPS (1024 spf)
Compression mode : Lossy
Result on the mux side:
avout: failed to write header for output file: Invalid argument
[mp4 @ 000000000902d680] sample rate not set
also ProRes seems not supported @ all
Lets see if it works on 3.05 now by default :)
uhhh Rigaya updated to Video Codec SDK 7.1.9 since 7.01 it seems
seems 375.63 stopped working now and now 375.95 is minimum
Glimpse of 8.0 coming ?
Failed to create instance of nvEncodeAPI, please consider updating your GPU driver.
1) Enhancements to H.264 motion estimation only mode:
a) Ability to select specific motion vector partitions and intra mode enable/disable for motion estimation only mode.
b) Performance enhancement for stereo mode motion-estimation.
2) Streamlined the nomenclature of rate control modes.
3) Quality improvement for H.264 Temporal Adaptive Quantization(TAQ).
Nothing really interesting for HEVC though only H.264 and MVC improvements overall (officially), nothing about bug fixes in terms of HEVC Encoding or general Improvements.
Most important improvements where introduced already with 7.01
JohnLai
21st January 2017, 18:20
[@CruNcher]
Muxer error = Seem like there is a problem with the audio. If ffmpeg is being used...then -ar 48000 is used. In case of NVENCC, maybe --audio-samplerate 48000 ?
NVENC SDK 7.1 requires NVIDIA Windows display driver 375.95 or newer.
CruNcher
21st January 2017, 18:33
Every muxer for Rigayas Nvenc TS input shows the same problem, very unreliable standalone and ProRes fails Decoding completely.
JohnLai
21st January 2017, 18:46
Well....you could contact rigaya or create a ticket https://github.com/rigaya/NVEnc/issues about the parser issue.
CruNcher
21st January 2017, 19:37
I'll do it when im sure 3.05 has still the same issue it looks like it copying the new ffmpeg library components to the 3.02 binary but i dunno how both are internally working together and before i can test 3.05 i need to change the driver and im very picky with driver changing overall (especially Nvidia Drivers since some strategy updates), even if it looks basically like a very small change going from 375.63 up to 375.95 being forced currently through this API change that seems to have happened with the 7.1 SDK.
I want to produce some test files first and compare if they really have been no noticeable bitstream differences on the HEVC side of things and some test files for H.264 as well looking at the overall improvements there.
So far i found 1 visual issue on Nvidias Encoder side in form of Nvenc im overall not happy with it's visual retranscoding outcome for the Performance, especially not vs H.264 CPU but overall it doesn't do bad for UHD almost 50 fps realtime at reasonable bitrates, i mostly hit 30 fps though on GM204 :)
If that result should have improved with the Driver update it would be really interesting :D
Hmm it seems to work without --audio-copy so the direct copy transfer seems to fail interesting this is being shown when aborting the execution.
[mp4 @ 00000000091fd680] sample rate not set
avout: failed to write header for output file: Invalid argument
Ignoring attempt to set invalid timebase 1/0 for st:1
encoded 0 frames, 0.00 fps, -nan(ind) kbps, 0.00 MB
encode time 0:00:03 / CPU Usage: 75.59%
@JohnLai
you could be right it seems to be rather a muxer issue with that ADTS track, but then practicaly every muxer fails geez.
with the default audio transcoding overhead ~0.5x Realtime and the internal 10->8bit conversion
[26.2%] 2572 frames: 30.78 fps, 24654.49 kb/s, remain 0:03:55
https://devtalk.nvidia.com/default/topic/987496/video-technologies/nvenc-diagram-correction/
hehe yeah funny though that the Encoder is as powerful as GM206 but the Decoder is Hybrid and pretty damn weak on the Performance GTX 970.
So some files the Encoder creates most GM204 can't even Playback flawless but GM206 can :D
And the part about the Decoding support in the Matrix doesn't fit either for it based on the Assumption that NVCUVID is now transformed to NVDEC ;)
Normally it should be as well shown as supported in the NVDEC Matrix even if heavier CPU Depending as on GM206 ;)
Note: For Video Codec SDK 7.0, NVCUVID has been renamed to NVDECODE API.
JohnLai
22nd January 2017, 04:15
[@CruNcher]
30fps hevc encoding seems to be about right due to cpu decoding. Can't expect high quality visual transcoding from GM204 fixed function encoder.
Encoder itself could handle 8bit HEVC 4k at 60fps just fine. Then again, not even core i7-3770k at 4.5ghz (nor my i5-3570k 4.2ghz) can decode the source 10bit HEVC 4K60 50Mbps without dropping frame.
Very ironic indeed.
However, there is a bug with color range for nvidia hevc encoding https://devtalk.nvidia.com/default/topic/958132/video-technologies/nvenc-hevc-with-full-range-colors-/
This particular bug doesn't happen with H264.
CruNcher
22nd January 2017, 11:59
I find it already visual very pleasing for the Performance it achieves and you shouldn't forget this is not yet a full GPU developed Video Codec, that thing is currently working only in Nvidias Lab ;)
That Intel would beat it at even lower Power in every of their High End Notebooks @ around 35W is a really masterful architectural achievement that even AMD has a hard time to reach yet with their APU integration and Fusion idea (HSA) ;)
Im throwing currently around 150W out of the Gen2 System for that 30 fps result as much as for the Decoding actually @ flawless non dropping 60 fps @ around 70W :)
It will be interesting to see what i can achieve with VP9 and AV1 at those 150W against Kyrion and Titan ;)
https://youtu.be/BaiPRAPOnjA?t=184
JohnLai
22nd January 2017, 13:03
VCEEnc 3.00 by rigaya is out!
Polaris HEVC encoding is supported. Only 8bit encoding.
EDIT:
Eh? It appears rigaya-san bought second hand rx460 just to ensure everything works.
Talk about amount of dedication....:eek:
NikosD
22nd January 2017, 13:09
I've just read his email he sent me.
You are fast!
He also wrote this:
Details:
https://github.com/rigaya/VCEEnc/releases/tag/3.00
"Please note that VCEEnc still lacks some features compared to QSVEnc/NVEnc, like HW HEVC decode or sw decoding by libavcodec."
I'll try to post some encodings later today.
CruNcher
22nd January 2017, 13:44
Please try todo the best retranscode you can of Samsungs Journey of Colors in 8 bit and a target of 25 Mbps :)
gona release my test result shortly in the Decoder Evaluation thread of it
encoded 6646 frames, 32.25 fps, 22602.53 kbps, 298.75 MB
encode time 0:03:26 / CPU Usage: 72.36%
frame type IDR 31
frame type I 31, avgQP 23.48, total size 15.90 MB
frame type P 6615, avgQP 28.70, total size 282.85 MB
Format : MPEG-4
Format profile : Base Media / Version 2
Codec ID : mp42 (isom/iso2/mp41)
File size : 299 MiB
Duration : 1 min 50 s
Overall bit rate : 22.6 Mb/s
Writing application : NVEncC (x64) 3.02
Video
ID : 1
Format : HEVC
Format/Info : High Efficiency Video Coding
Format profile : Main@L5@High
Codec ID : hev1
Codec ID/Info : High Efficiency Video Coding
Duration : 1 min 50 s
Bit rate : 22.6 Mb/s
Width : 3 840 pixels
Height : 2 160 pixels
Display aspect ratio : 16:9
Frame rate mode : Constant
Frame rate : 59.940 (60000/1001) FPS
Color space : YUV
Chroma subsampling : 4:2:0
Bit depth : 8 bits
Scan type : Progressive
Bits/(Pixel*Frame) : 0.045
Stream size : 299 MiB (100%)
after testing the new Driver result and comparing it
JohnLai
22nd January 2017, 13:46
I've just read his email he sent me.
You are fast!
He also wrote this:
Details:
https://github.com/rigaya/VCEEnc/releases/tag/3.00
"Please note that VCEEnc still lacks some features compared to QSVEnc/NVEnc, like HW HEVC decode or sw decoding by libavcodec."
I'll try to post some encodings later today.
Prepared to be saddened for MikhailAMD said Polaris VCE encoder only support 8bit HEVC encoding.......T_T.....
EDIT:
One can use directshow filter in conjunction with LAVfilter DXVA Copy back to get hardware acceleration decode and piping it to vceencc.
EDIT2:
Don't forget to use "--quality slow" and "--pre-analysis auto". You know, cause last time intel TU has different min/max rectangle partitions for each target usage. Worry if AMD also has same issue.
Pitou
22nd January 2017, 14:31
Only 80fps with ffmpeg(dxva) for VC1? Got a feeling ffmpeg auto-select wrong GPU for decoding. (Maybe ffmpeg dxva auto select intel igpu, there is a bug with ffmpeg dxva decoding using Intel IGPU)
The copy-back operation shouldn't be that taxing.
EDIT:
If you have intel integrated GPU enabled (plus using windows 8/10)....you can select "QSVEncC (Intel)" as decoder.
It turned out QSVENCC decoder supports MPEG2, H264, VC1 too.
Note: Intel hardware decoder is the fastest compared to Nvidia and AMD.
~Using Intel IGPU decoder to pipe the decoded video to Nvidia Nvenc for encoding.~
I only have a Q9450 CPU, which is very old and doesn't have accelleration. That's why I need my Nvidia to do all the work. But apprently, NVencC doesn't decode VC1 on GPU. Only FFMpeg does.
Pitou!
CruNcher
22nd January 2017, 14:51
Prepared to be saddened for MikhailAMD said Polaris VCE encoder only support 8bit HEVC encoding.......T_T.....
Wonder why everyone is so surprised about it it was clear from day 1 that RX 470/480 (Polaris) is just Designed to beat GTX 970 (GM204) and be able to compete with GTX 1060 (GP104) on the Mass Mainstream market ;)
And Amd is lacking behind on UVD/VCE they just started to keep up :)
Mikhail has a interesting Background as a PMTS :)
https://www.youtube.com/watch?v=6GncOvnc_S0
MTUCI not MSU ;)
JohnLai
22nd January 2017, 15:24
I only have a Q9450 CPU, which is very old and doesn't have accelleration. That's why I need my Nvidia to do all the work. But apprently, NVencC doesn't decode VC1 on GPU. Only FFMpeg does.
Pitou!
I just dropped an email to rigaya asking about VC1, HEVC and VP9 NVDEC(cuvid) missing support. Don't put high hope on it.
Wonder why everyone is so surprised about it it was clear from day 1 that RX 470/480 (Polaris) is just Designed to beat GTX 970 (GM204) and be able to compete with GTX 1060 (GP104) on the Mass Mainstream market ;)
Because almost all reddit users at /AMD actually believe Polaris has 10bit HEVC encoding support. ~.~......
NikosD
22nd January 2017, 15:27
Please try todo the best retranscode you can of Samsungs Journey of Colors in 8 bit and a target of 25 Mbps :)
I'm ready to start testing new VCEEnc.
Give me a link for the above clip.
I just dropped an email to rigaya asking about VC1, HEVC and VP9 NVDEC(cuvid) missing support. Don't put high hope on it.
@JohnLai
I urgently suggest you to open a new thread here in Doom9 regarding GPU/HW encoding of H.264/ H.265 for all (AMD, Nvidia, Intel)
You are the most suitable guy to do that, since you dedicate a lot of hours testing, reading, evaluating mostly Nvidia HW encoders but also you are clearly interested in all HW encoders.
You are very helpful and detailed suggesting solutions to users and you will get a lot of feedback because I think HW encoding is a trend right now.
That thread is clearly missing from doom9.
Go on and we will all contribute to that, starting from my results of Polaris HEVC encoding :)
gona release my test result shortly in the Decoder Evaluation thread of it
No, don't post such info in the Decoder Evaluation thread.
It would be much better to do it in the new thread regarding HW encoding that JohnLai will open soon :)
JohnLai
22nd January 2017, 15:37
@JohnLai
I urgently suggest you to open a new thread here in Doom9 regarding GPU/HW encoding of H.264/ H.265 for all (AMD, Nvidia, Intel)
No way....ain't my specialty.:scared:
Speaking of H264....seem like only Intel IGPU (haswell onwards) and AMD VCE (except Polaris) support b-pyramid. Nvidia never support b-pyramid. I wonder why.
NikosD
22nd January 2017, 15:41
No way....ain't my specialty.:scared:
It is exactly your specialty, at least from all the users writing here.
And what do you mean specialty ?
We aren't exactly professionals, we only want to help users.
Come on!
You are the best on HW encoding and you don't have to know everything.
Just make the beginning and all the rest will follow.
mariush
22nd January 2017, 15:44
This VCEEnc 3.0 actually works on my Windows 7 machine, with my XFX RX 470 card. A's video converter didn't show the HEVC option.
So if you want me to do some test, paste the command line and a link to the sample input file (if any particular one is desired) and I can do tests for you guys.
I did a test encode of a 1280x536 5mbps vbr ~1min video and it encoded it at around 135 fps 1600 kbps ... it was the absolute minimum command line parameters to get it to work, like -c hevc --avcee (or something like that) -i file.mp4 -o file.hevc
NikosD
22nd January 2017, 15:45
Don't post anything here until JohnLai opens a thread regarding HW encoding [emoji14]
CruNcher
22nd January 2017, 15:48
http://demo-uhd3d.com/fiche.php?cat=uhd&id=91
It's a relative old Ateme Encoder result ;)
JohnLai
22nd January 2017, 16:15
Don't post anything here until JohnLai opens a thread regarding HW encoding [emoji14]
Sorry, I can't.
By the way, why does rigaya explicitly mention --vbaq (H.264 only)? AMF 1.4 simply states By default, disable VBAQ. It should be possible to enable it for HEVC or if this is just an oversight by rigaya?
It appears Long Term Reference picture support is not implemented too.
EDIT: Assuming if AMD VCE HEVC P frames can refers to multiple preceding frames, LTR support is even more crucial.
NikosD
22nd January 2017, 17:12
http://demo-uhd3d.com/fiche.php?cat=uhd&id=91
It's a relative old Ateme Encoder result ;)
Downloading...
Sorry, I can't.
I really, really, really can't understand you...
By the way, why does rigaya explicitly mention --vbaq (H.264 only)? AMF 1.4 simply states By default, disable VBAQ. It should be possible to enable it for HEVC or if this is just an oversight by rigaya?
It appears Long Term Reference picture support is not implemented too.
EDIT: Assuming if AMD VCE HEVC P frames can refers to multiple preceding frames, LTR support is even more crucial.
I tried VBAQ with HEVC and get a reply:
"VBAQ is not supported with HEVC encoding, disabled."
Yes, I think the first version of VCEEnc supporting AMF has some small bugs and is missing a few things.
While I'm downloading the 10bit HEVC file, I did some tests with this source:
ftp://helpedia.com/pub/multimedia/testvideos/x264/2012%20-%2001%20-%20QuickSync%20vs%20UVD%202.2%20vs%20VP4/9.Ducks.Take.Off.1080p30fpsRef5-108Mbps.mkv
For all my tests I used --cqp 25 along with --preset-analysis auto
I only changed -u (quality) using three options fast, balanced, slow.
The results:
(All HEVC encoded samples have half size of the original H.264 file with quality you can see by yourselves)
Ducks_HEVC_CQP25_fast.mkv
https://www.sendspace.com/file/9djrh3
DX9: List of adapters:
0: Device ID: 67DF [Radeon (TM) RX 470 Graphics]
DX9 : Chosen Device 0: Device ID: 67DF [Radeon (TM) RX 470 Graphics]
VCEEnc 3.00 (x64) / Windows 10 (x64)
CPU: Intel Core i5-2400 @ 3.10GHz [TB: 3.20GHz] (4C/4T)
GPU: \\.\DISPLAY1 [Ellesmere 1300MHz (2236.10)]
Input Info: avcodec video: H.264/AVC, 1920x1080, 30000/1001 fps
Output: H.265/HEVC main @ Level 4.1
1920x1080p 1:1 29.970fps (30000/1001fps)
avwriter: hevc => matroska
Quality: balanced
CQP: I:25, P:25
VBV Bufsize: 20000 kbps
Bframes: 0 frames
Motion Est: Q-pel
Slices: 1
GOP Len: 300 frames
Others: deblock hrd pre-analysis:auto
encoded 500 frames, 62.77 fps, 57915.24 kbps, 115.18 MB
encode time 0:00:08, CPULoad: 26.13%
frame type IDR 2
frame type I 2, total size 0.60 MB
frame type P 498, total size 114.58 MB
Ducks_HEVC_CQP25_balanced.mkv
https://www.sendspace.com/file/yfii43
DX9: List of adapters:
0: Device ID: 67DF [Radeon (TM) RX 470 Graphics]
DX9 : Chosen Device 0: Device ID: 67DF [Radeon (TM) RX 470 Graphics]
VCEEnc 3.00 (x64) / Windows 10 (x64)
CPU: Intel Core i5-2400 @ 3.10GHz [TB: 3.30GHz] (4C/4T)
GPU: \\.\DISPLAY1 [Ellesmere 1300MHz (2236.10)]
Input Info: avcodec video: H.264/AVC, 1920x1080, 30000/1001 fps
Output: H.265/HEVC main @ Level 4.1
1920x1080p 1:1 29.970fps (30000/1001fps)
avwriter: hevc => matroska
Quality: balanced
CQP: I:25, P:25
VBV Bufsize: 20000 kbps
Bframes: 0 frames
Motion Est: Q-pel
Slices: 1
GOP Len: 300 frames
Others: deblock hrd pre-analysis:auto
encoded 500 frames, 55.80 fps, 58366.81 kbps, 116.08 MB
encode time 0:00:09, CPULoad: 25.84%
frame type IDR 2
frame type I 2, total size 0.60 MB
frame type P 498, total size 115.48 MB
Ducks_HEVC_CQP25_quality.mkv
https://www.sendspace.com/file/cteioc
DX9: List of adapters:
0: Device ID: 67DF [Radeon (TM) RX 470 Graphics]
DX9 : Chosen Device 0: Device ID: 67DF [Radeon (TM) RX 470 Graphics]
VCEEnc 3.00 (x64) / Windows 10 (x64)
CPU: Intel Core i5-2400 @ 3.10GHz [TB: 3.30GHz] (4C/4T)
GPU: \\.\DISPLAY1 [Ellesmere 1300MHz (2236.10)]
Input Info: avcodec video: H.264/AVC, 1920x1080, 30000/1001 fps
Output: H.265/HEVC main @ Level 4.1
1920x1080p 1:1 29.970fps (30000/1001fps)
avwriter: hevc => matroska
Quality: balanced
CQP: I:25, P:25
VBV Bufsize: 20000 kbps
Bframes: 0 frames
Motion Est: Q-pel
Slices: 1
GOP Len: 300 frames
Others: deblock hrd pre-analysis:auto
During the encoding of all three, GPU utilisation and GPU Memory controller was 0%, GPU clock was at low 751MHz but GPU Memory speed was at max 2000MHz
GPU core only power consumption was about 7.7W to 12.3W
As you can see Quality says always balanced which is probably a bug.
Also, the size of balanced and slow samples is exactly the same.
Using other samples of 1080p H.264 files as a source (~45Mbps), I managed a speed ~100fps for HEVC encoding.
Bframes reported are always 0, no matter what.
JohnLai
22nd January 2017, 18:35
Darn, I forgot to send a request to rigaya about adding AMF_VIDEO_ENCODER_HEVC_MAX_NUM_REFRAMES support after enumerating AMFCaps interface of AMF_VIDEO_ENCODER_HEVC_CAP_MAX_REFERENCE_FRAMES.
Anyway, currently checking Ducks_HEVC_CQP25_quality.mkv.
Intra PU sizes
4x4 8x8 16x16 32x32
Inter PU sizes
8x8
8x16
16x8
16x16
16x32
32x16
32x32
32x64
64x32
64x64
Hmm......
VCE "quality" sample from nikosd
max_transform_hierarchy_depth_inter 4
max_transform_hierarchy_depth_intra 4
transform_skip_enabled_flag 0
cu_qp_delta_enabled_flag 0
pps_loop_filter_across_slices_enabled_flag 0
NvenC using GTX970
max_transform_hierarchy_depth_inter 3
max_transform_hierarchy_depth_intra 0
transform_skip_enabled_flag 1
cu_qp_delta_enabled_flag 1
pps_loop_filter_across_slices_enabled_flag 1
Well, as usual, no SAO for Polaris.
QSV TU1 Skylake
max_transform_hierarchy_depth_inter 2
max_transform_hierarchy_depth_intra 2
transform_skip_enabled_flag 0
cu_qp_delta_enabled_flag 1
pps_loop_filter_across_slices_enabled_flag 0
NikosD
22nd January 2017, 18:37
I will send him a thorough email with various small bugs I have found out.
Tell me to add features from the AMF.
JohnLai
22nd January 2017, 18:56
I will send him a thorough email with various small bugs I have found out.
Tell me to add features from the AMF.
LTR, REF, HRD conformance, option to enable VBAQ (sdk said disable by default, it doesn't mean it can't be enabled or is there serious bug with it?)
HEVC_DE_BLOCKING_FILTER_DISABLE, this one should be set to true or false if I wanna keep deblocking active?
Judging from the AMF sdk, that about it.
CruNcher
22nd January 2017, 19:36
NVEnc 3.02 (x64), using NVENC API v7.0
OS Version Windows 7 (x64)
CPU Intel Core i5-2400 @ 3.10GHz [TB: 3.30GHz] (4C/4T)
GPU #0: GeForce GTX 970 (13 EU) @ 1266 MHz (375.63)
Input Buffers CUDA, 32 frames
Input Info avsw: hevc(yv12(10bit))->nv12 [SSE2], 3840x2160, 60000/1001 fps
Vpp Filters copyHtoD
Output Info H.265/HEVC main @ Level auto
3840x2160p 1:1 59.940fps (60000/1001fps)
avwriter: hevc => mp4
Rate Control VBR2
Bitrate 25000 kbps (Max: 30000 kbps)
Initial QP I:20 P:23 B:25
VBV buf size auto
Lookahead on, 16 frames, Adaptive I, B Insert
GOP length 600 frames
B frames 0 frames
Ref frames 3 frames, LTR: on
AQ off
MV Quality Q-pel
CU max / min 32 / 8
encoded 6646 frames, 30.14 fps, 22650.44 kbps, 299.38 MB
encode time 0:03:40 / CPU Usage: 72.37%
frame type IDR 31
frame type I 31, avgQP 23.52, total size 15.46 MB
frame type P 6615, avgQP 28.61, total size 283.93 MB
For highest Quality Playback Efficiency use a Renderer with Realtime Debanding option (MadVR/MPDotNet/MPv) or push it through a TV PP Decoder Pipeline Sony/Samsung e.c.t :)
https://www.sendspace.com/file/declpg
NikosD
22nd January 2017, 19:38
LTR, REF, HRD conformance, option to enable VBAQ (sdk said disable by default, it doesn't mean it can't be enabled or is there serious bug with it?)
HEVC_DE_BLOCKING_FILTER_DISABLE, this one should be set to true or false if I wanna keep deblocking active?
Judging from the AMF sdk, that about it.
Eventually, I read both PDFs for AVC and HEVC from AMF docs and I added all the small bugs I have found out and I sent a huge email to rigaya!
Let's see what he could manage to add and fix.
thanks!
NikosD
22nd January 2017, 21:36
Anyway, currently checking Ducks_HEVC_CQP25_quality.mkv.
Intra PU sizes
4x4 8x8 16x16 32x32
Inter PU sizes
8x8
8x16
16x8
16x16
16x32
32x16
32x32
32x64
64x32
64x64
Hmm......
VCE "quality" sample from nikosd
max_transform_hierarchy_depth_inter 4
max_transform_hierarchy_depth_intra 4
transform_skip_enabled_flag 0
cu_qp_delta_enabled_flag 0
pps_loop_filter_across_slices_enabled_flag 0
Two of the bugs reported to rigaya were:
1) HW decoding of VCE v3.0 is not working like v2.0, because it uses ~25% of a 4C/4T CPU which means that one core is used at 100%, while v2.0 uses HW decoding with ~2% CPU
But v3.0 is slightly faster than v2.0
2) The Quality reported by the runtime info is always at balanced no matter what.
I mean even if I choose fast or slow (quality), the Quality info line says "Balanced"
I thought it was cosmetic but when I tried H.264 encoding it says Quality fast, balanced or slow and the variation in speed is a lot more than H.265 encoding between different presets.
So, hold your horses about the "quality" sample until rigaya replies what's really going on.
CruNcher
22nd January 2017, 23:37
375.95 Installed checking for result differences
So indeed only unification of things no change like the release notes also stated on the HEVC side.
VBR 2Pass now called like Nvidias Internal naming convention VBR High Quality and so on ;)
NVEnc 3.02 (x64), using NVENC API v7.0
OS Version Windows 7 (x64)
CPU Intel Core i5-2400 @ 3.10GHz [TB: 3.30GHz] (4C/4T)
GPU #0: GeForce GTX 970 (13 EU) @ 1266 MHz (375.63)
Input Buffers CUDA, 32 frames
Input Info avsw: hevc(yv12(10bit))->nv12 [SSE2], 3840x2160, 60000/1001 fps
Vpp Filters copyHtoD
Output Info H.265/HEVC main @ Level auto
3840x2160p 1:1 59.940fps (60000/1001fps)
avwriter: hevc => mp4
Rate Control VBR2
Bitrate 25000 kbps (Max: 30000 kbps)
Initial QP I:20 P:23 B:25
VBV buf size auto
Lookahead on, 16 frames, Adaptive I, B Insert
GOP length 600 frames
B frames 0 frames
Ref frames 3 frames, LTR: on
AQ off
MV Quality Q-pel
CU max / min 32 / 8
encoded 6646 frames, 30.14 fps, 22650.44 kbps, 299.38 MB
encode time 0:03:40 / CPU Usage: 72.37%
frame type IDR 31
frame type I 31, avgQP 23.52, total size 15.46 MB
frame type P 6615, avgQP 28.61, total size 283.93 MB
NVEnc 3.02 (x64), using NVENC API v7.0
OS Version Windows 7 (x64)
CPU Intel Core i5-2400 @ 3.10GHz [TB: 3.30GHz] (4C/4T)
GPU #0: GeForce GTX 970 (13 EU) @ 1266 MHz (375.95)
Input Buffers CUDA, 32 frames
Input Info avsw: hevc(yv12(10bit))->nv12 [SSE2], 3840x2160, 60000/1001 fps
Vpp Filters copyHtoD
Output Info H.265/HEVC main @ Level auto
3840x2160p 1:1 59.940fps (60000/1001fps)
avwriter: hevc => mp4
Rate Control VBR2
Bitrate 25000 kbps (Max: 30000 kbps)
Initial QP I:20 P:23 B:25
VBV buf size auto
Lookahead on, 16 frames, Adaptive I, B Insert
GOP length 600 frames
B frames 0 frames
Ref frames 3 frames, LTR: on
AQ off
MV Quality Q-pel
CU max / min 32 / 8
encoded 6646 frames, 29.87 fps, 22650.44 kbps, 299.38 MB
encode time 0:03:42 / CPU Usage: 70.54%
frame type IDR 31
frame type I 31, avgQP 23.52, total size 15.46 MB
frame type P 6615, avgQP 28.61, total size 283.93 MB
NVEnc 3.05 (x64), using NVENC API v7.1
OS Version Windows 7 (x64)
CPU Intel Core i5-2400 @ 3.10GHz [TB: 3.30GHz] (4C/4T)
GPU #0: GeForce GTX 970 (13 EU) @ 1266 MHz (375.95)
Input Buffers CUDA, 32 frames
Input Info avsw: hevc(yv12(10bit))->nv12 [SSE2], 3840x2160, 60000/1001 fps
Vpp Filters copyHtoD
Output Info H.265/HEVC main @ Level auto
3840x2160p 1:1 59.940fps (60000/1001fps)
avwriter: hevc => mp4
Rate Control VBRHQ
Bitrate 25000 kbps (Max: 30000 kbps)
Initial QP I:20 P:23 B:25
VBV buf size auto
Lookahead on, 16 frames, Adaptive I, B Insert
GOP length 600 frames
B frames 0 frames
Ref frames 3 frames, LTR: on
AQ off
MV Quality Q-pel
CU max / min 32 / 8
encoded 6646 frames, 30.27 fps, 22650.44 kbps, 299.38 MB
encode time 0:03:39 / CPU Usage: 72.38%
frame type IDR 31
frame type I 31, avgQP 23.52, total size 15.46 MB
frame type P 6615, avgQP 28.61, total size 283.93 MB
btw
[mpegts @ 00000000003779a0] start time for stream 1 is not set in estimate_timin
gs_from_pts
[mpegts @ 00000000003779a0] Could not find codec parameters for stream 1 (Audio:
aac ([15][0][0][0] / 0x000F), 0 channels, fltp): unspecified sample rate
Consider increasing the value for the 'analyzeduration' and 'probesize' options
so it's not really Rigayas fault
NikosD
23rd January 2017, 15:14
OK, I got some interesting replies from rigaya.
He will add REF in next version and probably add after a time LTR.
Regarding HRD, it is always enabled - no need to disable it
Deblocking filter always enabled, otherwise bad video quality.
VBAQ for HEVC always disabled, otherwise bad video quality.
Regarding HEVC -u options for quality always showing "balanced" he had no clue, but will investigate.
It could be a bug or limitation.
NikosD
23rd January 2017, 15:20
I also told him to add:
HEVC tier (main, high)
Full range color (H.264 only)
HW HEVC decoding
He could probably add them layer.
JohnLai
23rd January 2017, 16:05
Gotcha, NikosD.
Once rigaya-san adds the ref + ltr support, we can finally know if multiple reference frames are used.
I was thinking VBAQ to be acronym of Variable Bitrate Adaptive Quantization, clearly it is not....since developer said it produces bad quality.
Hmm, reading through AMF sdk source code.....where is AMD promised Two Pass encoding? There is nothing about two-pass in AMF SDK.
CruNcher
23rd January 2017, 20:39
NikosD how does the Samsung 10->8bit retranscode comes forward for you in direct compare vs the the Nvidia one i posted at the 2x reduction target for it any watchable results yet ? :)
im trying to get some result out of ffmpeg currently but that is somehow tricky with that parsing issue together not as easy as i thought and i wonder if that .ts is corrupt overall but Lav Splitter has 0 issues with it and also MPV and others show no issues and adjusting the probesize doesn't fix this very weired
i didn't got any encoding result for nvenc out of ffmpeg yet it behaves crazy with that input overall.
and it tells me it can't find any nvenc device at all tried different pixel formats and things but it acts totally weired -gpu list detects it correctly overall, pretty frustrating.
NikosD
23rd January 2017, 20:44
I'll wait a little for some basic fixes of rigaya regarding VCEENC before I try that sample.
You can see the results of HEVC encoding of the source and the samples I posted.
CruNcher
23rd January 2017, 21:17
@NikosD
Ok so only ducks, jellyfish and some desktop results to directly compare for now, not ideal but better then nothing ;)
lets hope some Intel user joins in Skylake/Kabylake then we can throw each other results around and compare though obviously neither of us would have any chance vs Quicksync overall ;)
What the ????
[graph 0 input from stream 0:0 @ 0000000000357140] Setting 'video_size' to value
'3840x2160'
[graph 0 input from stream 0:0 @ 0000000000357140] Setting 'pix_fmt' to value '7
2'
[graph 0 input from stream 0:0 @ 0000000000357140] Setting 'time_base' to value
'1/90000'
[graph 0 input from stream 0:0 @ 0000000000357140] Setting 'pixel_aspect' to val
ue '1/1'
[graph 0 input from stream 0:0 @ 0000000000357140] Setting 'sws_param' to value
'flags=2'
[graph 0 input from stream 0:0 @ 0000000000357140] Setting 'frame_rate' to value
'60000/1001'
[graph 0 input from stream 0:0 @ 0000000000357140] w:3840 h:2160 pixfmt:yuv420p1
0le tb:1/90000 fr:60000/1001 sar:1/1 sws_param:flags=2
[format @ 0000000000358620] compat: called with args=[yuv420p|nv12|p010le|yuv444
p|yuv444p16le|bgr0|rgb0|cuda]
[format @ 0000000000358620] Setting 'pix_fmts' to value 'yuv420p|nv12|p010le|yuv
444p|yuv444p16le|bgr0|rgb0|cuda'
[auto_scaler_0 @ 0000000000358b80] Setting 'flags' to value 'bicubic'
[auto_scaler_0 @ 0000000000358b80] w:iw h:ih flags:'bicubic' interl:0
[format @ 0000000000358620] auto-inserting filter 'auto_scaler_0' between the fi
lter 'Parsed_null_0' and the filter 'format'
[AVFilterGraph @ 0000000002f661e0] query_formats: 4 queried, 2 merged, 1 already
done, 0 delayed
[auto_scaler_0 @ 0000000000358b80] picking p010le out of 7 ref:yuv420p10le alpha
:0
[auto_scaler_0 @ 0000000000358b80] w:3840 h:2160 fmt:yuv420p10le sar:1/1 -> w:38
40 h:2160 fmt:p010le sar:1/1 flags:0x4
[hevc_nvenc @ 0000000003269020] Loaded Nvenc version 7.1
[hevc_nvenc @ 0000000003269020] Nvenc initialized successfully
[hevc_nvenc @ 0000000003269020] 1 CUDA capable devices found
[hevc_nvenc @ 0000000003269020] [ GPU #0 - < GeForce GTX 970 > has Compute SM 5.
2 ]
[hevc_nvenc @ 0000000003269020] 10 bit encode not supported
[hevc_nvenc @ 0000000003269020] No NVENC capable devices found
[hevc_nvenc @ 0000000003269020] Nvenc unloaded
Stream mapping:
Stream #0:0 -> #0:0 (hevc (native) -> hevc (hevc_nvenc))
Error while opening encoder for output stream #0:0 - maybe incorrect parameters
such as bit_rate, rate, width or height
[AVIOContext @ 00000000003bf520] Statistics: 0 seeks, 0 writeouts
[AVIOContext @ 00000000006b90a0] Statistics: 29381488 bytes read, 8 seeks
So it thinks i want todo a actuall 10bit encode and cancels ?
ok got it ffmpegs picky parser ;)
First result doesn't make the same quality level impression as rigayas nvencc output i posted above currently need to tweak it to the same level first options wise -preset slow itself doesn't seem on that level alone.
easyfab
24th January 2017, 20:53
for information, new Media Server Studio 2017 R2 for intel QSV
https://software.intel.com/en-us/forums/intel-media-sdk/topic/708917
JohnLai
25th January 2017, 16:25
http://rigaya34589.blog135.fc2.com/blog-entry-891.html
VCEEnc 3.01 is out.
Google translate version of changelog:
Added functions and fixed bugs.
[Common]
· Check the function of VCE at the time of execution and check the parameters.
- Added option to specify reference distance. (- ref <int>)
- Added option to specify the number of LTR frames. (- ltr <int>)
· Added H.264 Level 5.2.
- Version of AMF added to version information.
[VCEEncC]
· Fixed spelling error etc in help.
· Added option to check the function of VCE. (- check-features)
The function of HEVC can not be displayed normally.
Is this ...?
· Added HW decoding of HEVC (8 bit).
· Since wmv3's HW decoding does not work properly, it is deleted.
NikosD
25th January 2017, 16:30
I know.
I've already exchanged a few emails with rigaya :)
JohnLai
25th January 2017, 16:32
I know.
I've already exchanged a few emails with rigaya :)
:devil:
Awaiting hevc samples.
One with 6 refs
One with 6 refs + 6 LTR
NikosD
25th January 2017, 16:33
Still a few critical bugs.
I'll try though.
JohnLai
25th January 2017, 16:37
Still a few critical bugs.
I'll try though.
Don't use --pre-analysis
For some reason, this option causes my system with R7 260X H264 encoding to blue screen.
NikosD
25th January 2017, 16:39
--pre-analysis works for me, there are other bugs.
BTW, pre-analysis is the two-pass encoding you mentioned in a previous post.
JohnLai
25th January 2017, 16:42
--pre-analysis works for me, there are other bugs.
BTW, pre-analysis is the two-pass encoding you mentioned in a previous post.
That is two-pass? .....kinda disappointing ....judging from h264 vbr video quality.....
By the way, it turned out there is a bug about the blue screen with pre-analysis activated at amf github too.
https://github.com/GPUOpen-LibrariesAndSDKs/AMF/issues/62
NikosD
25th January 2017, 16:47
That is two-pass? .....kinda disappointing ....judging from h264 vbr video quality.....
I think that's what says here:
https://github.com/GPUOpen-LibrariesAndSDKs/AMF/issues/2
Have you tried CQP with pre-analysis?
JohnLai
25th January 2017, 16:54
I think that's what says here:
https://github.com/GPUOpen-LibrariesAndSDKs/AMF/issues/2
Have you tried CQP with pre-analysis?
Nope. I only tested 4 times with VBR mode. (5000kbps and 2500kbps, pre-analysis 'full' and 'none')
But it randomly blue screen out of nowhere.
Kinda not worth turning pre-analysis on with random system blue screen before it even start encoding.
By the way, are you sure pre-analysis option is really two-pass?
Cause from the video quality output....it doesn't seem so.
NikosD
25th January 2017, 17:11
Did you read the link above ?
My impression by reading that ticket closed is that is two pass encoding.
I don't have any other source for that info.
JohnLai
25th January 2017, 17:17
Did you read the link above ?
My impression by reading that ticket closed is that is two pass encoding.
I don't have any other source for that info.
I did read it.
But it is a conjecture by Xaymar (OBS studio plugin developer)
That option doesn't seem to be 'lookahead' either. Hmm....oh well, no point thinking too much about it.
NikosD
25th January 2017, 17:23
But AMD didn't reply him differently.
I'll ask rigaya if he knows something more.
CruNcher
25th January 2017, 20:27
Hmm i wonder why FFMPEG NVENC behaves so different overall then Rigayas NVEnCC Encoder i have to say i like Rigayas decission overall more even though it doesn't hit the bitrate as exact as FFMPEG NVENC does in every configured way.
NVENCC
Format : MPEG-4
Format profile : Base Media / Version 2
Codec ID : mp42 (isom/iso2/mp41)
File size : 299 MiB
Duration : 1 min 50 s
Overall bit rate : 22.7 Mb/s
Writing application : NVEncC (x64) 3.02
Video
ID : 1
Format : HEVC
Format/Info : High Efficiency Video Coding
Format profile : Main@L5@High
Codec ID : hev1
Codec ID/Info : High Efficiency Video Coding
Duration : 1 min 50 s
Bit rate : 22.7 Mb/s
Width : 3 840 pixels
Height : 2 160 pixels
Display aspect ratio : 16:9
Frame rate mode : Constant
Frame rate : 59.940 (60000/1001) FPS
Color space : YUV
Chroma subsampling : 4:2:0
Bit depth : 8 bits
Scan type : Progressive
Bits/(Pixel*Frame) : 0.046
Stream size : 299 MiB (100%)
FFMPEG NVENC
Format : MPEG-4
Format profile : Base Media
Codec ID : isom (isom/iso2/mp41)
File size : 302 MiB
Duration : 1 min 50 s
Overall bit rate : 22.8 Mb/s
Writing application : Lavf57.62.100
Video
ID : 1
Format : HEVC
Format/Info : High Efficiency Video Coding
Format profile : Main@L5@High
Codec ID : hev1
Codec ID/Info : High Efficiency Video Coding
Duration : 1 min 50 s
Bit rate : 22.8 Mb/s
Width : 3 840 pixels
Height : 2 160 pixels
Display aspect ratio : 16:9
Frame rate mode : Constant
Frame rate : 59.940 (60000/1001) FPS
Color space : YUV
Chroma subsampling : 4:2:0
Bit depth : 8 bits
Scan type : Progressive
Bits/(Pixel*Frame) : 0.046
Stream size : 302 MiB (100%)
Target for Rigaya NVENCC was 25 Mbps target for FFMPEG NVENC was 23 Mbps FFMPEG hit it pretty perfectly Rigaya missed it by a alot ;)
Rigaya is almost 2 Mbps off
So for Rigayas current setup you better calculate with that offset by default
NikosD
25th January 2017, 21:28
NikosD how does the Samsung 10->8bit retranscode comes forward for you in direct compare vs the the Nvidia one i posted at the 2x reduction target for it any watchable results yet ? :)
I tried all the sources of StaxRip but they can't demux that 10bit HEVC file and VCEEncC can't demux it too.
Find me a way to demux it in order to test it.
:devil:
Awaiting hevc samples.
One with 6 refs
One with 6 refs + 6 LTR
For you and everyone else who wants to see Polaris encoder in action.
With this source:
https://www.sendspace.com/file/zg01zb
I got these encoding results:
In order to achieve more than half of the original bit rate I used CQP 27 in all tests
1)H.264 (VCE) Polaris encoder (VCEEncC v3.01)
1st Test H.264
Encoding options:
Quality -> Balanced
Ref -> 6
LTR -> 2 (That's the MAX for both H.264 & H.265)
VBAQ -> ON
Pre-analysis -> Full
Profile -> High
Everything else set to default.
Speed -> ~100fps
Result:
https://www.sendspace.com/file/0fbm6y
2nd Test H.264
Encoding options:
Quality -> Slow
Ref -> 16
Everything else, same as 1st test.
Speed -> 55 fps
Result:
https://www.sendspace.com/file/breejv
2) H.265 (HEVC) Polaris encoder
1st Test H.265
Encoding options:
Quality -> Balanced
Ref -> 6
LTR -> 2 (That's the MAX for both H.264 & H.265)
Pre-analysis -> Auto
Everything else set to default.
Speed -> ~91fps
Result:
https://www.sendspace.com/file/dejr6x
2nd Test H.265
Encoding options:
Quality -> Slow
Ref -> 16
Everything else, same as 1st test.
Speed -> 91 fps
Result:
https://www.sendspace.com/file/kh28qr
For HEVC the problem of not being able to test other than balanced quality (like fast or slow) still exists, so although I choose Slow in the 2nd test, the output says Balanced
All of the encoded samples look great to me.
Waiting for your feedback of my samples and your quality tests comparison on the above source using Nvidia, Intel or AMD older cards.
JohnLai
26th January 2017, 12:35
Source frame number 18 (P-frame) is selected for PSNR and MSSIM analysis (0th frame is I-Frame, so I skipped it, and b-frame isn't suitable for comparison because Polaris HEVC encoding only has I and P frame) (P-frame comparison only).
1st Test H.264
PSNR
Y 15.7386
Cb 10.7270
Cr 10.7073
MSSIM
Y 0.0826
Cb 0.0422
Cr 0.0423
2nd Test H.264
PSNR
Y 15.7369
Cb 10.7270
Cr 10.7071
MSSIM
Y 0.0826
Cb 0.0421
Cr 0.0424
1st Test H.265
PSNR
Y 15.7008
Cb 10.7266
Cr 10.7066
MSSIM
Y 0.0822
Cb 0.0422
Cr 0.0424
2nd Test H.265
PSNR
Y 15.7008
Cb 10.7266
Cr 10.7066
MSSIM
Y 0.0822
Cb 0.0422
Cr 0.0424
1st Test H.264
MAX 0.0961 MSSIM / 28.5 PSNR
MIN 0.0418 MSSIM / 10.7 PSNR
2nd Test H.264
MAX 0.0961 MSSIM / 28.4 PSNR
MIN 0.0419 MSSIM / 10.7 PSNR
1st Test H.265
MAX 0.0958 MSSIM / 25.8 PSNR
MIN 0.0416 MSSIM / 10.7 PSNR
2nd Test H.265
MAX 0.0958 MSSIM / 25.8 PSNR
MIN 0.0416 MSSIM / 10.7 PSNR
Hmm...so I transcode the source video using* my current pc GTX 970 into H264 and HEVC.
GTX 970 LOOKAHEAD 32 H264 CQP 20 23 25 AQ (not using temporal aq, it just normal spatial aq) File size : 147668 KB (Probably due to B-frame usage)
PSNR
Y 15.6819
Cb 10.7267
Cr 10.7057
MSSIM
Y 0.0817
Cb 0.0424
Cr 0.0425
MAX 0.0953 MSSIM / 25.8 PSNR
MIN 0.0416 MSSIM / 10.7 PSNR
GTX 970 LOOKAHEAD 32 HEVC CQP 20 23 25 AQ File size : 193856KB
PSNR
Y 15.6819
Cb 10.7267
Cr 10.7057
MSSIM
Y 0.0817
Cb 0.0424
Cr 0.0425
MAX 0.0953 MSSIM / 25.8 PSNR
MIN 0.0416 MSSIM / 10.7 PSNR
GTX 970 LOOKAHEAD 32 HEVC CQP 29 29 29 AQ File size 79748KB
PSNR
Y 15.6818
Cb 10.7255
Cr 10.7057
MSSIM
Y 0.0819
Cb 0.0424
Cr 0.0424
MAX 0.0955 MSSIM / 25.8 PSNR
MIN 0.0417 MSSIM / 10.7 PSNR
NikosD
26th January 2017, 13:01
What are the sizes of the output files using CQP 20 23 25 ?
JohnLai
26th January 2017, 13:31
What are the sizes of the output files using CQP 20 23 25 ?
Up there....in the red color.....
NikosD
26th January 2017, 13:45
Ok...Your low numbered CQP H.264 file is double than mine and your HEVC file is almost triple, the comparison is by far not equal.
Your HEVC file is larger than the original AVC source (!)
When viewing the files of Polaris and GTX 970 which do you think look better ?
But of course at the same file size.
JohnLai
26th January 2017, 14:15
Ok...Your low numbered CQP H.264 file is double than mine and your HEVC file is almost triple, the comparison is by far not equal.
Your HEVC file is larger than the original AVC source (!)
When viewing the files of Polaris and GTX 970 which do you think look better ?
But of course at the same file size.
Comparing Nvidia HEVC CQP I 29 : P 29 with VCE? In motion, definitely Nvidia HEVC output.
In general nvidia HEVC is gonna be larger than AMD HEVC for sure.
This is due to more I-frame insertion + adaptive GOP (from Lookahead) + AQ + limited maximum CU 32x32.
By the way, VCE HEVC reference frame usage --> Also use 1 preceding frame as reference frame similar to NVENC HEVC. Your H264 samples also the same where the P-frame only makes use of 1 reference frame.
NumNegativePics 1
NumPositivePics 0
NumDeltaPocs 1
UsedByCurrPicS0 1
UsedByCurrPicS1
DeltaPocS0 -1
DeltaPocS1
EDIT: All of the VCE samples GOP = 600 , only 1 I-frame being inserted every 600 frames........
NikosD
26th January 2017, 14:31
From all the options rigaya has implemented I haven't used only the ones regarding b-frames, that I think are useless for Polaris encoder.
But there are at least two parameters of AMF that I don't know if they could be useful and rigaya hasn't implemented them so far: (for AVC and HEVC)
1) headers insertion spacing
2) IDR period
Cab you upload the small HEVC sample to sendspace ?
JohnLai
26th January 2017, 14:42
From all the options rigaya has implemented I haven't used only the ones regarding b-frames, that I think are useless for Polaris encoder.
But there are at least two parameters of AMF that I don't know if they could be useful and rigaya hasn't implemented them so far: (for AVC and HEVC)
1) headers insertion spacing
2) IDR period
Cab you upload the small HEVC sample to sendspace ?
Header = no idea
IDR period = "--gop-len" available in VCEenc, but what we actually need is adaptive gop length. Doubt AMD gonna implement it anytime soon.
Sample? This gonna take a long time using 50kb/s ADSL upload....
EDIT: wow....30 minutes to upload.....
EDIT2: Hmm? I just noticed nvencc lookahead also decides fixed 600 gop is optimal for the source material?
EDIT3: Interesting.....nvidia adaptive quantization is something to be feared, even if I set CQP of I29 and P29.....it actually varies the QP for each CU, 25 being the lowest and 32 being the highest. You can check CU QP value in the bitstream of the video I uploaded later. (52% uploading...unless if it disconnected again....)
EDIT4 : -.-....For all the VCE samples......every CU is using CU QP 27........
NikosD
26th January 2017, 14:46
I think you mean 50KByte/s which ~512Kbit/s.
It's not that bad!
JohnLai
26th January 2017, 15:25
I think you mean 50KByte/s which ~512Kbit/s.
It's not that bad!
Finally.... https://www.sendspace.com/file/mghgaq
NikosD
26th January 2017, 15:40
Thank you.
My eyes can't tell a difference between your sample and mine.
I mean no difference at all.
So, my next question is:
What is your encoding speed (FPS) for that HEVC small file ?
JohnLai
26th January 2017, 15:52
Thank you.
My eyes can't tell a difference between your sample and mine.
I mean no difference at all.
So, my next question is:
What is your encoding speed (FPS) for that HEVC small file ?
You can't tell the difference when it is in motion. Pausing it at certain frame...and you will notice it.
This is the speed.
encoded 1998 frames, 144.57 fps, 19596.57 kbps, 77.87 MB
NikosD
26th January 2017, 15:53
No, I can't tell any difference even in still images.
JohnLai
26th January 2017, 16:38
No, I can't tell any difference even in still images.
Really?
The tree? The leaf?
Facial texture?
Is there any website that can mouse over to compare two different image?
Other than the dreaded http://screenshotcomparison.com/ , can't even upload to this site.
EDIT:
Since almost all image hosters tend to autoconvert png file....
AvatarHEVCslow_track1_und.mkv_snapshot_00.03_[2017.01.26_23.24.51]
https://www.sendspace.com/file/0rffgz
6.Avatar-1080p60fps NVENC CQP I29P29 LOOKAHEAD 32 AQ.mp4_snapshot_00.03_[2017.01.26_23.25.04]
https://www.sendspace.com/file/6vg0ah
Yups
26th January 2017, 18:46
With this source:
https://www.sendspace.com/file/zg01zb
HD 630 1150 Mhz
b-frames 4, reference frames 2, Target usage 7
HEVC
CQP= 370 fps
VBR= 290 fps
JohnLai
26th January 2017, 18:53
HD 630 1150 Mhz
b-frames 4, reference frames 2, Target usage 7
HEVC
CQP= 370 fps
VBR= 290 fps
Dear Yups,
Can provide Intel QSV 10bit HEVC, b-frame 16, ref 5, TU1, scenechange sample with that source video?
I wanna check the bitstream......:)
Edit: Just use either VBR or ICQ to ensure the size is around 75 - 90 Mb.
Yups
28th January 2017, 10:52
B-frames 4 is the limit for HEVC, otherwise you run into issues.
CruNcher
28th January 2017, 11:08
Scenecuts can cause interesting fluctuations between NVEnCC and FFMPEGs NVENC on the Decoder side at the same avg bitrate target but i have no idea where this difference originates from in it's bitrate decisions yet.
Yups could you please benchmark the bitstream i posted here on your Kaby-Lake Decoder (in a rather clean system state with minimum of 2 runs)
https://www.sendspace.com/file/declpg
Nvidias GM204 Cuda Decoder Core
http://i1.sendpic.org/t/aC/aCT2XKCGnjRfhCQiPdTBwxU2U7p.jpg (http://sendpic.org/view/1/i/sq7WBQbNha5WxmPuzdYcoPvPd1M.png)
NVENCC (stream above)
Format : MPEG-4
Format profile : Base Media / Version 2
Codec ID : mp42 (isom/iso2/mp41)
File size : 299 MiB
Duration : 1 min 50 s
Overall bit rate : 22.7 Mb/s
Writing application : NVEncC (x64) 3.02
Video
ID : 1
Format : HEVC
Format/Info : High Efficiency Video Coding
Format profile : Main@L5@High
Codec ID : hev1
Codec ID/Info : High Efficiency Video Coding
Duration : 1 min 50 s
Bit rate : 22.7 Mb/s
Width : 3 840 pixels
Height : 2 160 pixels
Display aspect ratio : 16:9
Frame rate mode : Constant
Frame rate : 59.940 (60000/1001) FPS
Color space : YUV
Chroma subsampling : 4:2:0
Bit depth : 8 bits
Scan type : Progressive
Bits/(Pixel*Frame) : 0.046
Stream size : 299 MiB (100%)
SSIM
Y:0.942599 (12.410791)
U:0.967868 (14.930601)
V:0.964379 (14.482894)
All:0.950440 (13.048714)
Encode Speed = ~30 FPS (CPU/GPU)
FFMPEG NVENC (WIP)
Format : MPEG-4
Format profile : Base Media
Codec ID : isom (isom/iso2/mp41)
File size : 298 MiB
Duration : 1 min 50 s
Overall bit rate : 22.6 Mb/s
Writing application : Lavf57.62.100
Video
ID : 1
Format : HEVC
Format/Info : High Efficiency Video Coding
Format profile : Main@L5@High
Codec ID : hev1
Codec ID/Info : High Efficiency Video Coding
Duration : 1 min 50 s
Bit rate : 22.6 Mb/s
Width : 3 840 pixels
Height : 2 160 pixels
Display aspect ratio : 16:9
Frame rate mode : Constant
Frame rate : 59.940 (60000/1001) FPS
Color space : YUV
Chroma subsampling : 4:2:0
Bit depth : 8 bits
Scan type : Progressive
Bits/(Pixel*Frame) : 0.045
Stream size : 298 MiB (100%)
SSIM
Y:0.940291 (12.239596)
U:0.966039 (14.690139)
V:0.961858 (14.185987)
All:0.948177 (12.854755)
Encode Speed = ~33 FPS (CPU/GPU)
JohnLai
28th January 2017, 18:36
Scenecuts can cause interesting fluctuations between NVEnCC and FFMPEGs NVENC on the Decoder side at the same avg bitrate target but i have no idea where this difference originates from in it's bitrate decisions yet.
Nvidias GM204 Cuda Decoder Core
http://i1.sendpic.org/t/aC/aCT2XKCGnjRfhCQiPdTBwxU2U7p.jpg (http://sendpic.org/view/1/i/sq7WBQbNha5WxmPuzdYcoPvPd1M.png)
Hmm....seem like GTX970 hybrid decoding nature is the bottleneck.
http://i.imgur.com/FHozcid.png
Frame 2012 is the I-frame in the screenshot you produced
Checking the Coded Picture Buffer graph....guess the limit of CPB for our gtx970 hybrid decoding is 30 000 000, anymore than that, stutter~~~~
CruNcher
29th January 2017, 11:18
Yeah i wonder how this actually behaves between 970,980,980 TI :)
Also Cyberlinks HAM Decoder seems to be a little more performant in this case
Also currently i become really skeptical about Nvidias multipass (2Pass VBR aka VBRHQ) now testing it further on both encoder at least for non Realtime 2nd Generation source.
JohnLai
29th January 2017, 12:08
Yeah i wonder how this actually behaves between 970,980,980 TI :)
Also Cyberlinks HAM Decoder seems to be a little more performant in this case
Also currently i become really skeptical about Nvidias multipass (2Pass VBR aka VBRHQ) now testing it further on both encoder at least for non Realtime 2nd Generation source.
VBR vs VBRHQ?
--vbr 4500 --codec h265 --ref 6 --level 5.1 --lookahead 32
No adaptive quantization being used.
The I-frame:
VBR I-frame = http://i.imgur.com/7pLjtma.png
VBR2 I-frame = http://i.imgur.com/f8LuWXC.png
The 4th frame (P-frame) being chosen:
VBR P-frame = http://i.imgur.com/gTigDZP.png
VBR2 P-frame = http://i.imgur.com/wCjDOiM.png
Edit: crap, uploaded the wrong I-frame, done replacing*
NikosD
29th January 2017, 15:02
I did read it.
But it is a conjecture by Xaymar (OBS studio plugin developer)
That option doesn't seem to be 'lookahead' either. Hmm....oh well, no point thinking too much about it.
I just saw a comment from that developer in his OBS Studio forum:
It still is only a fake Two-Pass encoding that only adjusts the qp values for a given macroblock to store information better.
CruNcher
29th January 2017, 15:03
im also skeptical about the overall efficiency of such a high lookahead as 32 seing that the lookahead is cuda as well ;)
the balance with 16 frames is still aceeptable but with 32 its drifting extremely
and here i would really like to see and compare AMDs OpenCL implementation on Polaris and especially Vega :)
JohnLai
29th January 2017, 15:34
I just saw a comment from that developer in his OBS Studio forum:
It still is only a fake Two-Pass encoding that only adjusts the qp values for a given macroblock to store information better.
'Storing information better' might be the wrong term, 'storing information efficiently' would be better.
So, in a manner, VCE two pass may operate more or less similar to nvenc two pass.
In the nvenc screenshot I provided above, there are more intra block with VBR and less intra block with VBR2 in P-frame (probably because nvenc decided flat anime surface doesn't require the usage of better quality intra block, just my conjecture)
im also skeptical about the overall efficiency of such a high lookahead as 32 seing that the lookahead is cuda as well ;)
the balance with 16 frames is still aceeptable but with 32 its drifting extremely
From nvidia samples with SAO support, higher lookahead value such as 32 seem to make the output even worse (to be precise, it get blurrier and QP value is higher). 16 seems to be optimal lookahead for pascal gpu. (I compared pascal hevc samples with 8, 16 and 32 lookahead)
I am more skeptical about nvenc low-complexity SAO algorithm and its relation to higher lookahead value.
Meanwhile, for maxwell case, higher lookahead value doesn't seem to produce that much blurrier output compared to pascal.
EDIT: grammar correction....
NikosD
29th January 2017, 15:43
'Storing information better' might be the wrong term, 'storing information efficiently' would be better.
So, in a manner, VCE two pass may operate more or less similar to nvenc two pass.
But how did AMD manage to implement a two-pass encoding (pre-analysis) so flexible that can be used using CQP rate mode besides VBR?
I was under the impression that two pass encoding is used for VBR in order to hit closer the target bitrate (more efficiently)
Using CQP there is no target bitrate.
JohnLai
29th January 2017, 16:23
But how did AMD manage to implement a two-pass encoding (pre-analysis) so flexible that can be used using CQP rate mode besides VBR?
I was under the impression that two pass encoding is used for VBR in order to hit closer the target bitrate (more efficiently)
Using CQP there is no target bitrate.
Well....if you ask me....I would say.....isn't the way Xaymar describing AMD pre-analysis method similar to Nvidia adaptive quantization? (Fun fact, Nvidia AQ works in CQP mode too)
But, based on your VCE H264 1st Test H.264 with Pre-analysis set to Full, all macroblock is using QP of 27..........., no variation in QP value for each macroblock at all. Thus, I doubt AMD pre-analysis actually works..........
Is this sample encoded with VCE CQP method?
In NVENC case, there is variation in QP value for each macroblocks for all different rate control mode (CQP, VBR,VBR2,CBR) edit: If AQ is activated.
NikosD
29th January 2017, 16:29
Yes, if you read my settings the only difference between 1st and 2nd test is the quality parameter and the number of ReF
1st tests are using "balanced" quality with low REF and 2nd tests are using "slow" quality with max REF 16.
JohnLai
29th January 2017, 16:51
Yes, if you read my settings the only difference between 1st and 2nd test is the quality parameter and the number of ReF
1st tests are using "balanced" quality with low REF and 2nd tests are using "slow" quality with max REF 16.
KK, so based on your original question on "But how did AMD manage to implement a two-pass encoding (pre-analysis) so flexible that can be used using CQP rate mode besides VBR? " and xaymar description on pre-analysis.....the answer is....pre-analysis option doesn't do anything in all of four samples provided by you. All macroblock (H264) and CU (HEVC) is using same QP value of 27. Perhaps the reason is CQP rate control. (as its name implied, constant QP)
So...can I have four new samples using VBR rate control, please? (using the avatar video source)
1)H264 without pre-analysis
2)H264 with pre-analysis full
3)HEVC without pre-analysis
4)HEVC with pre-analysis auto
Let see if these four samples have differences.
*ignore the reference frame, it is proven Polaris VCE only make use of single reference frame no matter how many ref being specified.
NikosD
29th January 2017, 16:57
So...can I have four new samples using VBR rate control, please? (using the avatar video source)
1)H264 without pre-analysis
2)H264 with pre-analysis full
3)HEVC without pre-analysis
4)HEVC with pre-analysis auto
Let see if these four samples have differences.
OK, I'll do it later today.
*ignore the reference frame, it is proven Polaris VCE only make use of single reference frame no matter how many ref being specified.
OK, but I think Media Info says ReF 4 and ReF 16 in file properties.
JohnLai
29th January 2017, 17:28
OK, I'll do it later today.
:thanks:
OK, but I think Media Info says ReF 4 and ReF 16 in file properties.
Nope, bitstream of all four samples indicated only one preceding frame is used as reference frame.
For two of your H264 samples:
num_ref_idx_l0_default_active_minus1 0
Value should be 4 or 16 instead of 0.
num_ref_idx_override_flag 0
ref_pic_list_reordering_flag_l0 0
Override flag should be 1 and reordering flag l0 should be 4 or 16
I already posted the HEVC bitstream ref at https://forum.doom9.org/showpost.php?p=1794695&postcount=137
EDIT: media info probably check for max_num_ref_frames.
CruNcher
29th January 2017, 18:48
~-5 FPS i measured now on several UHD tests for VBR2
VBR2(HQ)
encoded 10741 frames, 35.87 fps, 23719.33 kbps, 506.69 MB
encode time 0:04:59 / CPU Usage: 75.71%
frame type IDR 41
frame type I 41, avgQP 25.41, total size 8.26 MB
frame type P 10700, avgQP 25.22, total size 498.43 MB
VBR
encoded 10741 frames, 41.08 fps, 23739.44 kbps, 507.12 MB
encode time 0:04:21 / CPU Usage: 81.61%
frame type IDR 41
frame type I 41, avgQP 25.98, total size 7.97 MB
frame type P 10700, avgQP 25.25, total size 499.15 MB
NikosD
29th January 2017, 21:30
KK, so based on your original question on "But how did AMD manage to implement a two-pass encoding (pre-analysis) so flexible that can be used using CQP rate mode besides VBR? " and xaymar description on pre-analysis.....the answer is....pre-analysis option doesn't do anything in all of four samples provided by you. All macroblock (H264) and CU (HEVC) is using same QP value of 27. Perhaps the reason is CQP rate control. (as its name implied, constant QP)
So...can I have four new samples using VBR rate control, please? (using the avatar video source)
1)H264 without pre-analysis
2)H264 with pre-analysis full
OK, so first my H.264 samples.
All of the below samples for H.264 HW encoding, use the Avatar source and they have these parameters in common:
Rate control: VBR 21000 Kbps
Max bitrate: 40000 Kbps
vbv buffer size: 20000 Kbps
Motion estimation: Full-pel
VBAQ: Enabled
1) Quality: Balanced
a) Pre-analysis: NONE
https://www.sendspace.com/file/8gbjyp
b) Pre-analysis: FULL
https://www.sendspace.com/file/g9ejib
2) Quality: Slow
a) Pre-analysis: NONE
https://www.sendspace.com/file/a97obr
b) Pre-analysis: FULL
https://www.sendspace.com/file/h565ft
Waiting for your feedback regarding pre-analysis, motion estimation and if you see differences between balanced and slow quality for H.264
JohnLai
30th January 2017, 11:52
A picture is worth a thousand words. Multiple screenshot capture is nightmare.
There are 7 frames.
I-frame, 1st P-frame, 2nd P-frame, 3rd P-frame, 4th P-frame, 5th P-frame and suddenly jump to 20th P-frame
Avatar_bal_VBR_Full_mot_est_full_preanalysis_vbaq
http://i.imgur.com/S45MBKW.jpg
http://i.imgur.com/bEvJDxg.jpg
http://i.imgur.com/uq39H9U.jpg
http://i.imgur.com/jmVleor.jpg
http://i.imgur.com/kdxRiyz.jpg
http://i.imgur.com/seOhw4W.jpg
http://i.imgur.com/RpOX7br.jpg
Avatar_bal_VBR_Full_mot_est_NO_preanalysis_vbaq
http://i.imgur.com/rEwSdJ7.jpg
http://i.imgur.com/Wo1mkXF.jpg
http://i.imgur.com/Y2Fc0Nr.jpg
http://i.imgur.com/nw48mhB.jpg
http://i.imgur.com/MQIYeyq.jpg
http://i.imgur.com/x2zSiJI.jpg
http://i.imgur.com/zKl2sJ0.jpg
Avatar_slow_VBR_NO_preanalysis
http://i.imgur.com/vItqBNF.jpg
http://i.imgur.com/2fRFis2.jpg
http://i.imgur.com/TfzPtZ0.jpg
http://i.imgur.com/r0dNdto.jpg
http://i.imgur.com/3HxYXiL.jpg
http://i.imgur.com/CExWjSh.jpg
http://i.imgur.com/UaQvaqw.jpg
Avatar_slow_VBR_preanalysis_full
http://i.imgur.com/Dewk8O5.jpg
http://i.imgur.com/aAL4zRX.jpg
http://i.imgur.com/d3FZLzc.jpg
http://i.imgur.com/Wxnlh9r.jpg
http://i.imgur.com/FMaL6gy.jpg
http://i.imgur.com/h2dTCqO.jpg
http://i.imgur.com/bPhVVMB.jpg
*Phew...done uploading*
Next....from motion vectors and QP comparison......it seem like Pre-analysis doesn't do anything at all. Changing from Balanced to Slow does make a difference....
EDIT:
20th P-frame from your previous VCE H264 CQP
AvatarAVC.mkv
http://i.imgur.com/IGEZ0vJ.jpg
AvatarAVCslow
http://i.imgur.com/AgiLUvD.jpg
AvatarHEVC
http://i.imgur.com/211j6BQ.jpg
AvatarHEVCslow
http://i.imgur.com/CdStTH7.jpg
So...yeah...as you said:
For HEVC the problem of not being able to test other than balanced quality (like fast or slow) still exists, so although I choose Slow in the 2nd test, the output says Balanced
Both HEVC samples have exactly the same QP, CU and motion vectors.
NikosD
30th January 2017, 12:41
Next....from motion vectors and QP comparison......it seem like Pre-analysis doesn't do anything at all. Changing from Balanced to Slow does make a difference....
I don't understand a lot from your screenshots, but do you think that motion estimation or VBAQ could be incompatible with pre-analysis and somehow deactivate it?
Although in the runtime info, it says pre-analysis full.
Or maybe it isn't implemented properly by AMF or rigaya.
JohnLai
30th January 2017, 12:55
I don't understand a lot from your screenshots, but do you think that motion estimation or VBAQ could be incompatible with pre-analysis and somehow deactivate it?
Although in the runtime info, it says pre-analysis full.
Or maybe it isn't implemented properly by AMF or rigaya.
No idea.
Need confirmation, do two of Avatar_slow H264 VBR samples have VBAQ enabled as well?
NikosD
30th January 2017, 13:18
Yes, look at my common settings.
All these parameters are enabled for all 4 samples.
JohnLai
30th January 2017, 14:24
Yes, look at my common settings.
All these parameters are enabled for all 4 samples.
If that the case, I am disappointed by Amd VCE AMF software support. The VCE hardware block is fine (first to support HEVC 64x64 CTU) even without bframe, but AMD seriously need to do something about its rate control.
Or pre-analysis isn't enabled by AMD driver yet?
Or there is issue with rigaya VCEenc implementation?
NikosD
30th January 2017, 14:53
VCEENC v3.02 is out.
I'll upload some HEVC files this time.
JohnLai
30th January 2017, 15:03
VCEENC v3.02 is out.
I'll upload some HEVC files this time.
So.....fix for "quality"/Slow setting for HEVC and H264 'quarter' pre-analysis? (But, you mentioned your h264 samples are using "full", not "quarter"? Plus pre-analysis 'none' and 'full' doesn't change anything to the bitstream)
CruNcher
1st February 2017, 16:14
VBR
- Better bitrate target hit
- Higher Performance
- Higher Metrics
VBR2(HQ)
- Worse bitrate target hit
- Lower Performance
- Lower Metrics
Really strange, i still have to find a input where it is showing some visible improvement that would somehow justify it's Performance cost.
And the difference on the input itself im testing with is so visual minimal that no one would see that difference ever per block, the only thing you will hardly recognize is the Performance cost, i guess this was really optimized entirely for Game Streaming scenarios.
JohnLai
1st February 2017, 16:56
VBR
- Better bitrate target hit
- Higher Performance
- Higher Metrics
VBR2(HQ)
- Worse bitrate target hit
- Lower Performance
- Lower Metrics
Really strange, i still have to find a input where it is showing some visible improvement that would somehow justify it's Performance cost.
And the difference on the input itself im testing with is so visual minimal that no one would see that difference ever per block, the only thing you will hardly recognize is the Performance cost, i guess this was really optimized entirely for Game Streaming scenarios.
Which is why I often recommend the usage of CQP and standard VBR in conjunction with lookahead and adaptive quantization. :D
Back in the day where lookahead, adaptive GOP and AQ didn't exist yet, VBR2Pass is useful for optimizing (allocating proper QP value / distributing bit) every frame due to fixed GOP nature. Imagine this, there is only one I-frame for every 300 frames (for 30fps video). If there is screen transition within these fixed GOP, then video quality will suffer because there is no 'proper' high quality I-frame for subsequent P-frame to refer with.
With existence of Lookahead and AQ, there is no reason to use two pass for transcoding unless one requires extremely low latency encoding such as video conferencing.
I prefer to use Nvenc Unrestrainted VBR Constant Quality mode + AQ + Lookahead than CQP.
*Note for newcomers who read this: Hardware based encoder "two pass" works differently than software based encoder.
edit:
VCEEnc 3.03 is out....
Google Translation of changelog
[Common]
· Allow VBAQ in vbr / cbr mode on HEVC.
[VCEEnc.auo]
· Also be able to use HW resizer from VCEEnc.auo.
· Fixed that Level is not saved correctly when HEVC encoder is on.
- Fixed the problem that - vbr does not work properly with HEVC encod of VCEEncC.
[VCEEncC]
· Avsw reader supports YUV 420 10 bit reading (encoding is 8 bit).
· When using avsw reader, colors are shifted depending on resolution.
NikosD
1st February 2017, 19:25
After the last emails exchanged with rigaya, I have come to some conclusions regarding VCEEncC and Polaris HW encoder.
* The quality option -u (slow, balanced, fast) is not working for HEVC. It was fixed only as runtime/ log info, but the output is the same for all quality options.
* HW 10bit HEVC decoding is already supported by LAV Video, PotPlayer and other DXVA decoders and AMF has the relevant parameter included in the API, but when rigaya actually tries to use it, he gets a message "not supported".
That's why he added SW decoding of 10bit HEVC (--avsw)
* REF and pre-analysis don't look like working as expected.
For all of the above, we really don't know if it's a driver/ API limitation (bug?) or hardware limitation.
Documentation and runtime info, like Nvidia and Intel don't say everything.
He will concentrate on bug fixes from now on for VCEEnc, because he thinks he has implemented most of the really useful/necessary parameters of AMF API.
The last somewhat major bug (?) of VCEEnc is the high cpu usage (like working in copy-back mode) which VCEEncC v2.00 didn't have, that he is trying to fix.
I will probably try some VBAQ VBR HEVC encodings and upload them to check.
CruNcher
2nd February 2017, 04:36
Which is why I often recommend the usage of CQP and standard VBR in conjunction with lookahead and adaptive quantization. :D
Back in the day where lookahead, adaptive GOP and AQ didn't exist yet, VBR2Pass is useful for optimizing (allocating proper QP value / distributing bit) every frame due to fixed GOP nature. Imagine this, there is only one I-frame for every 300 frames (for 30fps video). If there is screen transition within these fixed GOP, then video quality will suffer because there is no 'proper' high quality I-frame for subsequent P-frame to refer with.
With existence of Lookahead and AQ, there is no reason to use two pass for transcoding unless one requires extremely low latency encoding such as video conferencing.
I prefer to use Nvenc Unrestrainted VBR Constant Quality mode + AQ + Lookahead than CQP.
*Note for newcomers who read this: Hardware based encoder "two pass" works differently than software based encoder.
edit:
VCEEnc 3.03 is out....
Google Translation of changelog
[Common]
· Allow VBAQ in vbr / cbr mode on HEVC.
[VCEEnc.auo]
· Also be able to use HW resizer from VCEEnc.auo.
· Fixed that Level is not saved correctly when HEVC encoder is on.
- Fixed the problem that - vbr does not work properly with HEVC encod of VCEEncC.
[VCEEncC]
· Avsw reader supports YUV 420 10 bit reading (encoding is 8 bit).
· When using avsw reader, colors are shifted depending on resolution.
Hmm interesting is how AQ lowers the overall Decoding complexity but unrestrained seems not a good idea i had to much frame drops on the Hardware side that way and a again the visual win wasn't worth it.
JohnLai
2nd February 2017, 06:17
Hmm interesting is how AQ lowers the overall Decoding complexity but unrestrained seems not a good idea i had to much frame drops on the Hardware side that way and a again the visual win wasn't worth it.
LOL, I just realize I used the wrong name. It should be "Unconstrained".
UVBR-CQ encoding speed is around 80-90% of CQP.
targetQuality only works if initialqp is set to 1, maximum maxBitRate = averageBitRate for chosen profile level.
Most of the time, I stick with CQ value of 26.
Another quirk of UVBR-CQ mode is dark scene kinda look better compared to CQP. Hmmm.....
*Note: lookahead + AQ are enabled.
NikosD
2nd February 2017, 12:07
NikosD how does the Samsung 10->8bit retranscode comes forward for you in direct compare vs the the Nvidia one i posted at the 2x reduction target for it any watchable results yet ? :)
OK, finally after VCEEnc v3.03 I managed to demux/ decode 10bit HEVC streams using --avsw (SW decoding) but unfortunately even the latest version doesn't allow me to encode HEVC in UHD/4K resolution.
Rigaya hasn't managed to reproduce that bug yet.
So, I tried 4K H.264 encoding by enabling almost everything.
VCEEncC v3.03
VBR 22500, max bitrate/ vbv buffer 50000
VBAQ
Motion estimation/ pre-analysis full
Full range
Level 5.2
Quality best/slow
The final encoding is a little less than 300MB and it's here:
https://www.sendspace.com/file/ytj5jt
And this one is simpler, with most parameters at default:
VCEEncC v3.03
VBR 22500, max bitrate/ vbv buffer 50000
Quality best/slow
So, vbaq, pre-analysis, full range, are all set to NONE by default and motion estimation is quarter pel by default.
The final encoding with default settings is here:
https://www.sendspace.com/file/0l6pyj
JohnLai
3rd February 2017, 06:39
300Mb!? X__X
This gonna take a while......
Frame 44
Samsung_Journey_H264.mkv
Motion vector : https://k60.imgup.net/MVstandard8c4b.PNG
QP value on top left : https://i76.imgup.net/QP17032.PNG
QP value on bottom left : https://h86.imgup.net/QPA8f9a.PNG
Samsung_Journey_H264_Default.mkv
Motion vector : https://r85.imgup.net/MVdefaultfa42.PNG
QP value on top left : https://x82.imgup.net/QPDefault1dbb.PNG
QP value on bottom left : https://h11.imgup.net/QPD8aa0.PNG
No much variation in QP.
Strange thing is Samsung_Journey_H264_Default.mkv seem to make use of more motion vectors than Samsung_Journey_H264.mkv
EDIT: Compared to cruncher gtx 970 nvenc HEVC:
Motion vector : https://w53.imgup.net/NVENCCRUNCd4c4.PNG --> Using different analyser because usual analyser ended with with error after trying to rescale fit to screen, see for yourself https://b16.imgup.net/NVENCCRUNC3936.PNG
QP value on everything : https://y85.imgup.net/NVENCCRUNC3d52.PNG Seem like cruncher is using CQP value of I:21 and P:24 without AQ XD
NikosD
3rd February 2017, 17:34
300Mb!? X__X
This gonna take a while......
VCEEnc v3.04 is out fixing most, if not all, bugs mentioned by me to rigaya.
Later today or early tomorrow, I'm going to upload 4K HEVC encodings of that Samsung Journey 10bit HEVC 4K sample.
It seems that the most important parameter is -u quality (slow, fast, balanced) which unfortunately still doesn't work for HEVC (fast is a little different, but slow and balanced are exactly the same)
CruNcher
3rd February 2017, 23:33
Yeah that version was still fixed Q :)
Performance retranscode 33 fps input 60 playback 60
http://i1.sendpic.org/t/1j/1jwQfj1zCtUxlpmyPdtuTT4dTvz.jpg (http://sendpic.org/view/1/i/v5qTm7CNkmiLr1H284kAUBLzJdy.png)
VBR2
encoded 10783 frames, 33.50 fps, 24717.28 kbps, 530.07 MB
encode time 0:05:21 / CPU Usage: 52.76%
frame type IDR 20
frame type I 20, avgQP 20.65, total size 6.27 MB
frame type P 10763, avgQP 23.74, total size 523.80 MB
VBR
encoded 10783 frames, 38.95 fps, 24799.58 kbps, 531.83 MB
encode time 0:04:36 / CPU Usage: 56.79%
frame type IDR 20
frame type I 20, avgQP 20.55, total size 6.34 MB
frame type P 10763, avgQP 23.76, total size 525.50 MB
CQP
encoded 10783 frames, 38.91 fps, 24807.54 kbps, 532.01 MB
encode time 0:04:37 / CPU Usage: 56.86%
frame type IDR 20
frame type I 20, avgQP 21.00, total size 6.08 MB
frame type P 10763, avgQP 24.00, total size 525.93 MB
CBR
encoded 10783 frames, 38.85 fps, 23587.17 kbps, 505.83 MB
encode time 0:04:37 / CPU Usage: 57.07%
frame type IDR 20
frame type I 20, avgQP 20.45, total size 6.27 MB
frame type P 10763, avgQP 23.82, total size 499.57 MB
CBRHQ
encoded 10783 frames, 33.42 fps, 24915.55 kbps, 534.32 MB
encode time 0:05:22 / CPU Usage: 52.67%
frame type IDR 20
frame type I 20, avgQP 20.80, total size 6.18 MB
frame type P 10763, avgQP 23.53, total size 528.14 MB
JohnLai
4th February 2017, 13:00
https://www.techpowerup.com/230360/first-intel-processor-with-amd-radeon-graphics-within-2017
Hopefully Intel won't replace its QSV with AMD VCE........
NikosD
4th February 2017, 15:07
A bug in VCEEnc with high bitrate CBR, VBR doesn't allow me to encode the sample in HEVC.
Waiting for VCEEnc v3.05 (?)...
CruNcher
4th February 2017, 23:57
CBRHQ
encoded 3780 frames, 26.69 fps, 25070.41 kbps, 376.94 MB
encode time 0:02:21 / CPU Usage: 58.62%
frame type IDR 20
frame type I 20, avgQP 20.45, total size 23.50 MB
frame type P 3760, avgQP 22.21, total size 353.45 MB
SSIM
Y:0.985357 (18.343621)
U:0.988216 (19.287139)
V:0.988070 (19.233419)
All:0.986285 (18.628190)
VBRHQ
encoded 3780 frames, 26.54 fps, 25181.62 kbps, 378.62 MB
encode time 0:02:22 / CPU Usage: 57.78%
frame type IDR 20
frame type I 20, avgQP 19.35, total size 26.09 MB
frame type P 3760, avgQP 21.71, total size 352.53 MB
SSIM
Y:0.988181 (19.274309)
U:0.988890 (19.543008)
V:0.988778 (19.499126)
All:0.988399 (19.355007)
Highest VBRHQ result so far with a slight overflow
VBRHQ
Using Hardware Decoding over Cuvid
encoded 3780 frames, 38.28 fps, 25181.62 kbps, 378.62 MB
encode time 0:01:38 / CPU Usage: 20.83%
frame type IDR 20
frame type I 20, avgQP 19.35, total size 26.09 MB
frame type P 3760, avgQP 21.71, total size 352.53 MB
VBR
Using Hardware Decoding over Cuvid
encoded 3780 frames, 47.82 fps, 25383.44 kbps, 381.65 MB
encode time 0:01:19 / CPU Usage: 24.99%
frame type IDR 20
frame type I 20, avgQP 19.95, total size 25.23 MB
frame type P 3760, avgQP 21.64, total size 356.42 MB
SSIM
Y:0.988352 (19.337383)
U:0.989007 (19.588642)
V:0.988774 (19.497570)
All:0.988531 (19.404814)
1 click H.264->Nvidia H.265 1080 p 2nd Generation transcoding test hitting the exact same bitrate 8000 kbps
SSIM
Y:0.895877 (9.824539)
U:0.985295 (18.325234)
V:0.979849 (16.956954)
All:0.924775 (11.236397)
and here manually overriding Rigayas RC Decisions getting a slightly lower bitrate result -500 ~7500 kbps
SSIM
Y:0.895970 (9.828434)
U:0.985184 (18.292787)
V:0.979744 (16.934516)
All:0.924802 (11.237925)
Reducing in bigger steps now towards the 50% Goal (not sure which Encoder was used for this H.264 input, random sample decision)
encoded 51445 frames, 162.21 fps, 5617.28 kbps, 1377.97 MB
encode time 0:05:17 / CPU Usage: 24.62%
frame type IDR 206
frame type I 206, avgQP 23.39, total size 14.26 MB
frame type P 51239, avgQP 23.70, total size 1363.70 MB
SSIM
Y:0.895992 (9.829321)
U:0.985270 (18.317935)
V:0.979906 (16.969430)
All:0.924857 (11.241128)
You can see that the SSIM Metric is rising now ;)
NikosD
5th February 2017, 21:24
New NVenc v3.06 adds HW decoding of HEVC/VP8/VP9
CruNcher
5th February 2017, 22:13
The H.264 High Predictive one seems damaged get strange artifacts on several blocks in the lower water parts with it compared to --avsw on my parkjoy test sample i created back then with x264 lossless but i guess it's a Generic Nvidia Decoder problem.
nevcairiel
5th February 2017, 22:21
The H.264 High Predictive one seems damaged get strange artifacts on several blocks in the lower water parts with it compared to --avsw on my parkjoy test sample i created back then with x264 lossless but i guess it's a Generic Nvidia Decoder problem.
x264 used to produce "wrong" lossless files not matching the specification properly, ffmpeg/avcodec compensates for that if it detects that a file was encoded with such a x264 version, but any other decoders will likely just show corrupted output.
It was fixed a long time ago, but who knows when your file was encoded.
CruNcher
5th February 2017, 22:37
Ahh that would explain it btw would you support the predictive files on lav cuvid please it falls back to avcodec not sure if you made that maybe on purpose for those wrong files though now not pushing them through to the cuvid decoder because of that :)
General
Complete name : E:\parkjoy.mp4
Format : MPEG-4
Format profile : JVT
Codec ID : avc1 (isom/avc1)
File size : 785 MiB
Duration : 10 s 0 ms
Overall bit rate : 658 Mb/s
Encoded date : UTC 2010-05-09 14:55:11
Tagged date : UTC 2010-05-09 14:55:11
Video
ID : 1
Format : AVC
Format/Info : Advanced Video Codec
Format profile : High 4:4:4 Predictive@L4.2
Format settings, CABAC : Yes
Format settings, ReFrames : 1 frame
Format settings, GOP : M=1, N=48
Codec ID : avc1
Codec ID/Info : Advanced Video Coding
Duration : 10 s 0 ms
Bit rate : 658 Mb/s
Maximum bit rate : 699 Mb/s
Width : 1 920 pixels
Height : 1 080 pixels
Display aspect ratio : 16:9
Frame rate mode : Constant
Frame rate : 50.000 FPS
Color space : YUV
Chroma subsampling : 4:2:0
Bit depth : 8 bits
Scan type : Progressive
Bits/(Pixel*Frame) : 6.348
Stream size : 785 MiB (100%)
Writing library : x264 core 94 r1583 7608d73
Encoding settings : cabac=1 / ref=1 / deblock=1:-1:-1 / analyse=0x3:0x13 / me=dia / subme=2 / psy=0 / mixed_ref=0 / me_range=16 / chroma_me=1 / trellis=0 / 8x8dct=1 / cqm=0 / deadzone=21,11 / fast_pskip=0 / chroma_qp_offset=0 / threads=3 / sliced_threads=0 / nr=0 / decimate=1 / interlaced=0 / constrained_intra=0 / bframes=0 / weightp=2 / keyint=48 / keyint_min=25 / scenecut=0 / intra_refresh=0 / rc=cqp / mbtree=0 / qp=0
Encoded date : UTC 2010-05-09 14:55:11
Tagged date : UTC 2010-05-09 14:57:12
nevcairiel
6th February 2017, 07:11
Ahh that would explain it btw would you support the predictive files on lav cuvid please it falls back to avcodec not sure if you made that maybe on purpose for those wrong files though now not pushing them through to the cuvid decoder because of that :)
CUVID only supports 4:2:0 decoding, not 4:4:4 (or more precisely, it cannot output anything but 4:2:0, so using it for 4:4:4 decoding would obliterate the chroma quality)
GrandPa
6th February 2017, 09:49
HD 630 1150 Mhz
b-frames 4, reference frames 2, Target usage 7
HEVC
CQP= 370 fps
VBR= 290 fps
@YUPS: Would it be possible to post the result of QSVEncC64 --check -features with your Kaby Lake CPU to see the difference with my 'old' Skylake 6700K?
(I'm new to this forum, but already did tons of H.264 and H.265 encoding with StaxRip and Hybrid. Now I'm interested to see whether H.265 HW encoding quality and efficiency with HD 630 could catch up a bit. While I found H.264 HW encoding with HD 530 quite acceptable in quality, there is currently no real alternative to X.265 SW encoding for archiving purposes)
Many thanks for any help
NikosD
7th February 2017, 13:37
Really?
The tree? The leaf?
Facial texture?
Is there any website that can mouse over to compare two different image?
Other than the dreaded http://screenshotcomparison.com/ , can't even upload to this site.
I found out a working site (I think) with frame comparison:
https://juxtapose.knightlab.com/
CruNcher
8th February 2017, 02:03
CUVID only supports 4:2:0 decoding, not 4:4:4 (or more precisely, it cannot output anything but 4:2:0, so using it for 4:4:4 decoding would obliterate the chroma quality)
Yeah kinda strange you can encode it but not decode it natively sounds after bandwidth saving though i wonder what is the reason that 4:2:2 isn't supported at all neither Encoding nor Decoding.
http://i1.sendpic.org/i/hN/hNHDvUAldhhJZF66LR6YWJ625Qu.png
And i wonder what they mean with the 3 asterisk at all, they appear nowhere in the Diagram makes no sense
you have 1,2 but 3 is like it's not specific but meant for everything based on the codecs or do they meant resolution and forgot the asterisks there ?
also color lossless seems confusing since when lossless is a color format ?
Im not quiet there, more exact data what they benched and how would be also nice
encoded 51445 frames, 162.21 fps, 5617.28 kbps, 1377.97 MB
encode time 0:05:17 / CPU Usage: 24.62%
frame type IDR 206
frame type I 206, avgQP 23.39, total size 14.26 MB
frame type P 51239, avgQP 23.70, total size 1363.70 MB
http://i1.sendpic.org/i/gb/gbLfj8UzAymwO4owHqXHWrggGMw.png
Dual Pass = VBR2(HQ)
i wonder how much SAO improves the result over Maxwell on the SSIM side for Pascal :)
easyfab
11th February 2017, 13:06
@YUPS: Would it be possible to post the result of QSVEncC64 --check -features with your Kaby Lake CPU to see the difference with my 'old' Skylake 6700K?
here Check-features : qsvencc 2.62 show api 1.19 but it's 1.22 so I d'ont know if all features are correct ?
10 bit Depth has an X but I can encode in 10bit with --profile main10.
Does qsvencc need an update to show new api features?
GPU Info Intel HD Graphics 615 (24EU) 300-900MHz [4W] (21.20.16.4589)
Media SDK QuickSyncVideo (hardware encoder) PG, 1st GPU, API v1.22
QSVEncC (x64) 2.62 (r1192) by rigaya, Jan 8 2017 23:11:24 (VC 1900/Win/avx2)
reader: raw, avi, avs, vpy, avqsv [H.264/AVC, HEVC, MPEG2, VC-1, VP8, VP9]
Environment Info
OS : Windows 10 (x64)
CPU: Intel Core m3-7Y30 @ 1.00GHz [TB: 1.61GHz] (2C/4T) <Skylake>
RAM: Used 2127 MB, Total 3976 MB
GPU: Intel HD Graphics 615 (24EU) 300-900MHz [4W] (21.20.16.4589)
Media SDK Version: Hardware API v1.19
Supported Enc features:
Codec: H.264/AVC
CBR VBR AVBR QVBR CQP VQP LA LAHRD ICQ LAICQ VCM
RC mode o o o o o o o o o o o
10bit depth x x x x x x x x x x x
Fixed Func o o o o o o x x x x o
Interlace o o o o o o o o o o x
SceneChange o o o o o o x x o x o
VUI info o o o o o o o o o o o
Trellis o o o o o o o o o o o
Adaptive_I x x x x x x x x x x x
Adaptive_B x x x x x x x x x x x
WeightP x x x x x x x x x x x
WeightB x x x x x x x x x x x
FadeDetect x x x x x x x x x x x
B_Pyramid o o o o o x o x o o o
+Scenechange x x x x x x x x x x x
+ManyBframes o o o o o x x x o x o
PyramQPOffset x x x x x x x x x x x
Ext_BRC o o o o x x x x o x o
MBBRC o o o o x x x x o x o
LA Quality x x x x x x o o x o x
QP Min/Max o o o o o o o o o o o
IntraRefresh x x x x x x x x x x x
No Debloc x x x x x x x x x x x
No GPB x x x x x x x x x x x
Windowed BRC x x x x x x o o x x x
PerMBQP(CQP) x x x x o o x x x x x
DirectBiasAdj x x x x x x x x x x x
MVCostScaling x x x x x x x x x x x
Codec: HEVC
CBR VBR AVBR QVBR CQP VQP LA LAHRD ICQ LAICQ VCM
RC mode o o x x o o x x o x o
10bit depth x x x x x x x x x x x
Fixed Func x x x x x x x x x x x
Interlace x x x x x x x x x x x
SceneChange o o x x o o x x o x o
VUI info o o x x o o x x o x o
Trellis x x x x x x x x x x x
Adaptive_I x x x x x x x x x x x
Adaptive_B x x x x x x x x x x x
WeightP x x x x x x x x x x x
WeightB x x x x x x x x x x x
FadeDetect x x x x x x x x x x x
B_Pyramid x x x x x x x x x x x
+Scenechange x x x x x x x x x x x
+ManyBframes x x x x x x x x x x x
PyramQPOffset x x x x o o x x x x x
Ext_BRC o o x x x x x x o x o
MBBRC o o x x x x x x o x o
LA Quality x x x x x x x x x x x
QP Min/Max x x x x x x x x x x x
IntraRefresh o o x x o o x x o x o
No Debloc o o x x o o x x o x o
No GPB x x x x x x x x x x x
Windowed BRC x x x x x x x x x x x
PerMBQP(CQP) o o x x x x x x o x o
DirectBiasAdj x x x x x x x x x x x
MVCostScaling x x x x x x x x x x x
Codec: MPEG2
CBR VBR AVBR QVBR CQP VQP LA LAHRD ICQ LAICQ VCM
RC mode o o o x o o x x x x x
10bit depth x x x x x x x x x x x
Fixed Func o o o x o o x x x x x
Interlace o o o x o o x x x x x
SceneChange o o o x o o x x x x x
VUI info o o o x o o x x x x x
Trellis o o o x o o x x x x x
Adaptive_I o o o x o o x x x x x
Adaptive_B o o o x o o x x x x x
WeightP o o o x o o x x x x x
WeightB o o o x o o x x x x x
FadeDetect x x x x x x x x x x x
B_Pyramid o o o x o x x x x x x
+Scenechange x x x x x x x x x x x
+ManyBframes o o o x o x x x x x x
PyramQPOffset x x x x x x x x x x x
Ext_BRC o o o x x x x x x x x
MBBRC o o o x x x x x x x x
LA Quality x x x x x x x x x x x
QP Min/Max o o o x o o x x x x x
IntraRefresh o o o x o o x x x x x
No Debloc o o o x o o x x x x x
No GPB x x x x x x x x x x x
Windowed BRC o o o x o o x x x x x
PerMBQP(CQP) x x x x x x x x x x x
DirectBiasAdj o o o x o o x x x x x
MVCostScaling o o o x o o x x x x x
Supported Vpp features:
Resize o
Deinterlace o
Scaling Quality o
Denoise o
Rotate o
Mirror x
Detail Enhancement o
Proc Amp. o
Image Stabilization x
Video Signal Info o
FPS Conversion o
FPS Conversion (Adv.) o
CruNcher
11th February 2017, 19:03
Could you do the same 10->8 bit Samsung Journey retranscode please and post it's result trying to hit the same bitrate target :)
http://demo-uhd3d.com/fiche.php?cat=uhd&id=91
Using the VBR RC
https://forum.doom9.org/showpost.php?p=1794206&postcount=108
https://forum.doom9.org/showpost.php?p=1795597&postcount=176
Though im not quiet sure if the target was a little to heavy with 300 mb for Nvidia and RC overall so restricted it seems not really good balanced i would have gone with slightly higher end result for visual perception/performance now without SAO but to late i guess to change to a closer to 25 mbps hit ;)
as we have those results now for Nvidia (HEVC, no SAO) and AMDs (H.264) current SDK NVENCC status ready only Intel results missing either H.264/HEVC or both :)
also going slightly higher would need another hosting so in some sense it's even a realistic target not going over that 300 mb ;)
easyfab
11th February 2017, 19:45
Will do when I downloaded the file, http://demo-uhd3d.com is slow.
What do you want exactly :
HEVC 8 bit @ VBR 22500K with slower/best setting ( or fast/balanced ) ?
Also AVC and HEVC 10bit ?
CruNcher
11th February 2017, 19:55
Of course a 10 bit HEVC one on one retranscode version would be nice without the 8 bit conversion step also :)
as fast as possible i dunno the current speed data on your system it would be able to achieve at those settings but the target would be 1x so 60 FPS :)
and if it should be going slower it would be around 40 and 30 fps
benchmark data of --avsw instead of the Quicksync Hardware Decoding would be also interesting :)
Reaching 25/30/50/60 is usable in some way ;)
easyfab
11th February 2017, 20:08
I don't believe that I will be that fast. I only have a small 4W Intel CPU (m3-7Y30) not a 7700K. @2K It's only 50fps for HEVC vs 250fps for AVC so for 4K ....
Download ETA 13 min
CruNcher
11th February 2017, 20:22
Even more exciting you have much lower overall latency in that notebook/mobile design and your very power limited on the ~10W :)
only 50 fps what should i say pushing much more W out to reach that on the top ~40 FPS result above (granted lots of that coming from the Decoding + conversion) ;)
though of course a compare of this would be not very nice and in Nvidias case a compare with a 1050 Notebook more fair then but overall the result of the encoder output counts first ;)
Also you dont have to forget that your 615 is also slighty restricted on it's ddr3 interface
Having a m3-7Y30 + Nvidia 1050 would be a really nice combination :D
though personally my decision would have gone towards
http://ark.intel.com/products/95442/Intel-Core-i3-7100U-Processor-3M-Cache-2_40-GHz-
http://ark.intel.com/products/95443/Intel-Core-i5-7200U-Processor-3M-Cache-up-to-3_10-GHz
http://www.cpubenchmark.net/compare.php?cmp%5B%5D=2865&cmp%5B%5D=2879&cmp%5B%5D=2864
What is your notebooks overall power out decoding/encoding that target and completing the task (without display) ?
easyfab
11th February 2017, 21:19
QSVENCC 2.62 HEVC 10 bit for Samsung_Journey
QSVEncC (x64) 2.62 (r1192) by rigaya, Jan 8 2017 23:11:24 (VC 1900/Win/avx2)
OS Windows 10 (x64)
CPU Info Intel Core m3-7Y30 @ 1.00GHz [TB: 1.61GHz] (2C/4T) <Skylake>
GPU Info Intel HD Graphics 615 (24EU) 300-900MHz [4W] (21.20.16.4589)
Media SDK QuickSyncVideo (hardware encoder) PG, 1st GPU, API v1.22
Async Depth 5 frames
Buffer Memory d3d9, 1 input buffer, 22 work buffer
Input Info avqsv video: HEVC, 3840x2160, 60000/1001 fps
Output HEVC main10 @ Level auto
3840x2160p 1:1 59.940fps (60000/1001fps)
avwriter: hevc => matroska
Target usage 4 - balanced
Encode Mode Bitrate Mode - VBR
Bitrate 24000 kbps
Max Bitrate 30000 kbps
QP Limit min: none, max: none
Trellis Auto
Ref frames 4 frames
Bframes 3 frames, B-pyramid: off
Max GOP Length 600 frames
Scene Change off
Ext. Features PerMBRC
encoded 6642 frames, 12.47 fps, 22039.38 kbps, 291.13 MB
encode time 0:08:52, CPULoad: 2.12%
frame type IDR 1
frame type I 12, total size 3.91 MB
frame type P 1661, total size 211.23 MB
frame type B 4969, total size 76.00 MB
As slow as expected. and a little smaller ( 291 MB)
The quality doesn't sound to bad, but I don't know if it's as good as AMD/Nvidia quality.
I can try with best preset if needed.
https://www.sendspace.com/file/qbod0d
NikosD
11th February 2017, 21:33
QSVEnc needs an update because it reports your CPU as Skylake, but the API has the correct version v1.22
Your CPU usage is extremely low, ~2% for 1GHz 2C/4T is very impressive.
The bottleneck could be the GPU though, as it has 24EUs but only at 0.9GHz
Someone with a Pascal GPU could try to transcode that clip to 10bit HEVC in balanced speed/quality using VBR and latest NVEnc v3.06 with HW HEVC decoding to see the difference.
easyfab
11th February 2017, 21:50
QSVENCC 2.62 AVC (Kaby lake api 1.22) for Samsung_Journey
QSVEncC (x64) 2.62 (r1192) by rigaya, Jan 8 2017 23:11:24 (VC 1900/Win/avx2)
OS Windows 10 (x64)
CPU Info Intel Core m3-7Y30 @ 1.00GHz [TB: 1.61GHz] (2C/4T) <Skylake>
GPU Info Intel HD Graphics 615 (24EU) 300-900MHz [4W] (21.20.16.4589)
Media SDK QuickSyncVideo (hardware encoder) PG, 1st GPU, API v1.22
Async Depth 6 frames
Buffer Memory d3d9, 1 input buffer, 32 work buffer
Input Info avqsv video: HEVC, 3840x2160, 60000/1001 fps
VPP Enabled ColorFmtConvertion: nv12(10bit) -> nv12
Output H.264/AVC High @ Level 5.2
3840x2160p 1:1 59.940fps (60000/1001fps)
avwriter: h264 => matroska
Target usage 1 - best
Encode Mode Bitrate Mode - VBR
Bitrate 25000 kbps
Max Bitrate 37500 kbps
QP Limit min: none, max: none
Trellis Auto
Ref frames 3 frames
Bframes 3 frames, B-pyramid: on
Max GOP Length 600 frames
Scene Change off
encoded 6642 frames, 33.74 fps, 21907.80 kbps, 289.39 MB
encode time 0:03:17, CPULoad: 27.08%
frame type IDR 12
frame type I 12, total size 3.14 MB
frame type P 1661, total size 228.03 MB
frame type B 4969, total size 58.23 MB
With best preset.
~3x faster than HEVC balanced preset.
CPULoad: 27.08% Must be the 10bit -> 8bit conversion.
https://www.sendspace.com/file/87hvix
easyfab
11th February 2017, 22:23
though personally my decision would have gone towards
http://ark.intel.com/products/95442/Intel-Core-i3-7100U-Processor-3M-Cache-2_40-GHz-
http://ark.intel.com/products/95443/Intel-Core-i5-7200U-Processor-3M-Cache-up-to-3_10-GHz
[url]http://www.cpubenchmark.net/compare.php?cmp%5B%5D=2865&cmp%5B%5D=2879&cmp%5B?
Yep but I choose a fanless laptop :) and I dont find 7100u or 7200u without fan
CruNcher
11th February 2017, 22:29
Yeah cooling the higher 14nm Chips passively a no go with its base 15 TDP rating which can surely top out at 25W
First time we see HEVC B-frames coming out of a GPU implemented Encoder :D
could you try without b-frames and also a run with --avsw and could you reach with the fastest preset the AVC FPS best preset result ?
also don't forget to disable any power saving functions, though you would be fast throttled no matter what with that Fanless Design.
PS: Your balanced encodes banding result is very nice that's visible immediately, much lower banding perception in the sky scene :)
Overall the Decoding complexity is higher compared to my encode Nvidias Cuda Decoder spikes out a lot more under the 60 FPS.
I wonder if that banding result difference comes from a more efficient 10->8 bit Hardware Dithering conversion compared to the default --avsw in NVENCC
Input Info avsw: hevc(yv12(10bit))->nv12 [SSE2], 3840x2160, 60000/1001 fps
also indeed it's rather strange that NVENCC shows here yv12(10bit) instead of p010(10bit)
Does rigaya wrongly do p010->yv12->nv12 ?
easyfab
11th February 2017, 23:57
HEVC 8bit fastest preset without Bframes :
QSVEncC (x64) 2.62 (r1192) by rigaya, Jan 8 2017 23:11:24 (VC 1900/Win/avx2)
OS Windows 10 (x64)
CPU Info Intel Core m3-7Y30 @ 1.00GHz [TB: 1.61GHz] (2C/4T) <Skylake>
GPU Info Intel HD Graphics 615 (24EU) 300-900MHz [4W] (21.20.16.4589)
Media SDK QuickSyncVideo (hardware encoder) PG, 1st GPU, API v1.22
Async Depth 6 frames
Buffer Memory d3d9, 1 input buffer, 31 work buffer
Input Info avqsv video: HEVC, 3840x2160, 60000/1001 fps
VPP Enabled ColorFmtConvertion: nv12(10bit) -> nv12
Output HEVC main @ Level auto
3840x2160p 1:1 59.940fps (60000/1001fps)
avwriter: hevc => matroska
Target usage 7 - fastest
Encode Mode Bitrate Mode - VBR
Bitrate 23500 kbps
Max Bitrate 30000 kbps
QP Limit min: none, max: none
Trellis Auto
Ref frames 4 frames
Bframes none
Max GOP Length 600 frames
Scene Change off
Ext. Features PerMBRC
encoded 6642 frames, 38.97 fps, 22514.68 kbps, 297.41 MB
encode time 0:02:50, CPULoad: 26.93%
frame type IDR 1
frame type I 12, total size 3.07 MB
frame type P 6630, total size 294.34 MB
Quality is not that good and encoding time is slower without Bframes 38.97 vs 46.71 with Bframes
https://www.sendspace.com/file/mm5lj0
for --avsw : I stopped after 2 min but ~16fps
encoded 2187 frames, 15.90 fps, 55032.13 kbps, 239.36 MB
encode time 0:02:18, CPULoad: 80.07%
for fastest HEVC encoding on my little cpu
QSVEncC (x64) 2.62 (r1192) by rigaya, Jan 8 2017 23:11:24 (VC 1900/Win/avx2)
OS Windows 10 (x64)
CPU Info Intel Core m3-7Y30 @ 1.00GHz [TB: 1.49GHz] (2C/4T) <Skylake>
GPU Info Intel HD Graphics 615 (24EU) 300-900MHz [4W] (21.20.16.4589)
Media SDK QuickSyncVideo (hardware encoder) PG, 1st GPU, API v1.22
Async Depth 6 frames
Buffer Memory d3d9, 1 input buffer, 34 work buffer
Input Info avqsv video: HEVC, 3840x2160, 60000/1001 fps
VPP Enabled ColorFmtConvertion: nv12(10bit) -> nv12
Output HEVC main @ Level auto
3840x2160p 1:1 59.940fps (60000/1001fps)
avwriter: hevc => matroska
Target usage 7 - fastest
Encode Mode Bitrate Mode - VBR
Bitrate 24000 kbps
Max Bitrate 30000 kbps
QP Limit min: none, max: none
Trellis Auto
Ref frames 2 frames
Bframes 3 frames, B-pyramid: off
Max GOP Length 600 frames
Scene Change off
Ext. Features PerMBRC
encoded 6642 frames, 46.71 fps, 22262.31 kbps, 294.08 MB
encode time 0:02:22, CPULoad: 27.58%
with HEVC 10bit : it's slower ~35fps
CruNcher
12th February 2017, 00:08
Ohh your encode is 10 bits im surprised now it plays on the Cuda Decoder
ahh ok it wasn't it switched to avcodec that explains the complexity difference ;)
and so no 8 bit conversion was done at all that explains also the banding result difference.
FPS wise a nice result for such a Power Target :)
Yeah the CPU Decoding of that Bitstream lowers the overall Encoding Performance significantly
Decoding of that fastest result bitstream is extremely hard for Nvidias Cuda Decoder throughout the whole bitstream i reach only 30 fps at tops where i reach a lot of times 60 with my encode
could you try to reach 30 fps encoding with maybe 1 preset up TU 6 (fast) and no b-frames but with the 8 bit version
and make a balanced TU 4 version of that 8 bit to of course one with best preset would be a nice compare point :)
So 8 bit VBR version of
TU 1 = ??
TU 4 = ??
TU 6 = ??
TU 7 = 38.97 fps (available) https://forum.doom9.org/showpost.php?p=1796814&postcount=202
without B-frames and the FPS you achieve on your m3
JohnLai
12th February 2017, 03:41
Wait....KabyLake finally adds P-frame support for its HEVC?
EDIT:
easyfab,
Could you provide a small 1080p sample with TU1, ICQ=23 (Maybe test with ICQ-LA, who knows if Kaby supported it?), 10bit, HEVC, B-frames=6 or 8, Ref=6, B-pyramid=on (just input the command, if it doesnt work, it will fallback to off), Trellis=all.
CruNcher
12th February 2017, 05:36
That TU 7 encode has some heavier visual issues (failings) compared to my Nvidia encode and is overall more Decoding complex, TU 7 seems no good compare base at all especially fades fail.
Lets see what the SSIM result difference is and how that fade failing impacts it
Easyfabs TU 7 8 Bit Encode
SSIM
Y:0.939540 (12.185293)
U:0.959835 (13.961549)
V:0.955864 (13.552082)
All:0.945643 (12.647447)
Easyfabs TU 4 10 Bit Encode with B-frames
SSIM
Y:0.949048 (12.928351)
U:0.970515 (15.304042)
V:0.967422 (14.870695)
All:0.955688 (13.534776)
it reflects well in the visual stability and better banding result and all that at lower overall bitrate
easyfab
12th February 2017, 10:52
@JohnLai
la-icq -> not supported
2k sample (crowd_run) with your settings 55MB https://www.sendspace.com/file/dsswn6
B pyramid is not supported on current platform, disabled.
trellis is not supported on current platform, disabled.
cop.PicTimingSEI value changed off -> auto by driver
cop3.DirectBiasAdjustment value changed off -> auto by driver
cop3.GlobalMotionBiasAdjustment value changed off -> auto by driver
QSVEncC (x64) 2.62 (r1192) by rigaya, Jan 8 2017 23:11:24 (VC 1900/Win/avx2)
OS Windows 10 (x64)
CPU Info Intel Core m3-7Y30 @ 1.00GHz [TB: 1.61GHz] (2C/4T) <Skylake>
GPU Info Intel HD Graphics 615 (24EU) 300-900MHz [4W] (21.20.16.4589)
Media SDK QuickSyncVideo (hardware encoder) PG, 1st GPU, API v1.22
Async Depth 5 frames
Buffer Memory d3d9, 3 input buffer, 31 work buffer
Input Info y4m: yv12->nv12[AVX2], 1920x1080, 50/1 fps
VPP Enabled ColorFmtConvertion: nv12 -> nv12(10bit)
Output HEVC main10 @ Level auto
1920x1080p 1:1 50.000fps (50/1fps)
avwriter: hevc => matroska
Target usage 1 - best
Encode Mode ICQ (Intelligent Const. Quality)
ICQ Quality 23
QP Limit min: none, max: none
Trellis Auto
Ref frames 6 frames
Bframes 8 frames, B-pyramid: off
Max GOP Length 500 frames
Scene Change off
encoded 500 frames, 3.73 fps, 46430.36 kbps, 55.35 MB
encode time 0:02:14, CPULoad: 4.02%
frame type IDR 1
frame type I 1, total size 0.53 MB
frame type P 56, total size 12.43 MB
frame type B 443, total size 42.39 MB
Don't look at the speed I used wifi to access the source.
@CruNcher
There is only 3 speed presets because faster=fastest ....
I encoded the 500 first frames of Samsung_journey with :
- HEVC 8bit no Bframes preset fastest
- HEVC 8bit no Bframes preset balanced
- HEVC 8bit no Bframes preset best
- HEVC 10bit no Bframes preset best
I hope it's enough for you to test
https://www.sendspace.com/file/nex2x8
CruNcher
12th February 2017, 11:57
@JohnLai
la-icq -> not supported
2k sample (crowd_run) with your settings 55MB https://www.sendspace.com/file/dsswn6
B pyramid is not supported on current platform, disabled.
trellis is not supported on current platform, disabled.
cop.PicTimingSEI value changed off -> auto by driver
cop3.DirectBiasAdjustment value changed off -> auto by driver
cop3.GlobalMotionBiasAdjustment value changed off -> auto by driver
QSVEncC (x64) 2.62 (r1192) by rigaya, Jan 8 2017 23:11:24 (VC 1900/Win/avx2)
OS Windows 10 (x64)
CPU Info Intel Core m3-7Y30 @ 1.00GHz [TB: 1.61GHz] (2C/4T) <Skylake>
GPU Info Intel HD Graphics 615 (24EU) 300-900MHz [4W] (21.20.16.4589)
Media SDK QuickSyncVideo (hardware encoder) PG, 1st GPU, API v1.22
Async Depth 5 frames
Buffer Memory d3d9, 3 input buffer, 31 work buffer
Input Info y4m: yv12->nv12[AVX2], 1920x1080, 50/1 fps
VPP Enabled ColorFmtConvertion: nv12 -> nv12(10bit)
Output HEVC main10 @ Level auto
1920x1080p 1:1 50.000fps (50/1fps)
avwriter: hevc => matroska
Target usage 1 - best
Encode Mode ICQ (Intelligent Const. Quality)
ICQ Quality 23
QP Limit min: none, max: none
Trellis Auto
Ref frames 6 frames
Bframes 8 frames, B-pyramid: off
Max GOP Length 500 frames
Scene Change off
encoded 500 frames, 3.73 fps, 46430.36 kbps, 55.35 MB
encode time 0:02:14, CPULoad: 4.02%
frame type IDR 1
frame type I 1, total size 0.53 MB
frame type P 56, total size 12.43 MB
frame type B 443, total size 42.39 MB
Don't look at the speed I used wifi to access the source.
@CruNcher
There is only 3 speed presets because faster=fastest ....
I encoded the 500 first frames of Samsung_journey with :
- HEVC 8bit no Bframes preset fastest
- HEVC 8bit no Bframes preset balanced
- HEVC 8bit no Bframes preset best
- HEVC 10bit no Bframes preset best
I hope it's enough for you to test
https://www.sendspace.com/file/nex2x8
500 frames is a bit to low in the scene count and reactions it captures, especially of that very visible fade problems with TU7.
The first 500 frames are mostly black to color and straight cuts only but better then nothing thx anyways :)
TU7 though shows heavy problems with transitional blend fades.
And in all so far the Seeking behavior is a catastrophe, you should really enable scene change
What you can see in the first 500 frames immediately is that every preset works with a lower partition blocksize decision then Nvidias Encoder does.
Overall though it's not quiet fair to compare without working scenechange but yeah more finer details luminance as well as chrominance seem to be better preserved then my current encode uploaded even vs the fastest TU7 with it's overall more visual stability problems that pushes down the SSIM :)
CruNcher NVENCC (Maxwell)
http://i1.sendpic.org/t/kU/kU7pkD21hUTBaYLLdDJHEGixINX.jpg (http://sendpic.org/view/1/i/eTqrGhgqONmfMVY9C57mzv3Bd69.png)
Easyfab QSVENCC (Kabylake) TU7
http://i1.sendpic.org/t/3n/3nmp7l4SL8h8jRe6n7PliZtVrC2.jpg (http://sendpic.org/view/1/i/i7DpsljkPThhDasEaCcaJE2SzbN.png)
CruNcher NVENCC (Maxwell)
http://i1.sendpic.org/t/mD/mDVVPTgDDp2TF3mZsIpfcJcpk58.jpg (http://sendpic.org/view/1/i/avhEQtZcrUhVSxQPavR5jsqCBx5.png)
Easyfab QSVENCC (Kabylake) TU7
http://i1.sendpic.org/t/97/97q7dLFwjIu7UBycTdqk1nnNH2A.jpg (http://sendpic.org/view/1/i/czbAGVYrQW6PmCkrWljQgG0LvbL.png)
CruNcher NVENCC (Maxwell)
http://i1.sendpic.org/t/jK/jKjWVTHAnOjz1ePtCpXQsW5EAA9.jpg (http://sendpic.org/view/1/i/7d2hW3SiPP2IjCgoJqqi8YDpFdM.png)
Easyfab QSVENCC (Kabylake) TU7
http://i1.sendpic.org/t/hq/hqMCakYj8kQCdh5SW6F20ICe2Xm.jpg (http://sendpic.org/view/1/i/7ka8kz9EJDaoT2pcCwIARAsxTHv.png)
Though somehow it looks that overall the 10->8 bit conversion was also more efficiently done inside of QSVENCC compared to --avsw.
There is a really big visual perceptive difference in the luminance and resulting chrominance also, which results in a complete different lighting result of shadow scenes.
Shadow areas become brighter on Easyfabs Intel Output results like they are indirect lighten.
JohnLai
12th February 2017, 14:18
@JohnLai
la-icq -> not supported
2k sample (crowd_run) with your settings 55MB https://www.sendspace.com/file/dsswn6
Analysis....
Contrary to the log detail on P-frame.......There is no P-FRAME in the bitstream, still GPB style as usual....unless if one views those P-frame detail in the log as Generalised B-frame.
https://k55.imgup.net/ActualB-Frf7cb.PNG
There are 8 actual B-frames in the bitstream. However, only 4 active frames being referred by currently displayed B-frame. Guess the limit for refs is 4. (Please have a look on DPB "Used by current"
https://b06.imgup.net/PGeneralisab20.PNG
Next, for P-frame (a B-frame that is treated as P-frame in Generalized B-Frame mode), this P-frame can only refer up to 3 reference frames.
EDIT:
No SAO support.
Intra PU sizes
4x4 8x8 16x16 32x32
Inter PU sizes
4x8 8x4 8x8 8x16 16x8 16x16 32x32
GrandPa
14th February 2017, 16:27
here Check-features : qsvencc 2.62 show api 1.19 but it's 1.22 so I d'ont know if all features are correct ?
10 bit Depth has an X but I can encode in 10bit with --profile main10.
Does qsvencc need an update to show new api features?
Quite strange results. On my machine it reports <Kabylake> for my Skylake CPU:
QSVEncC (x64) 2.62 (r1192) by rigaya, Jan 8 2017 23:11:24 (VC 1900/Win/avx2)
reader: raw, avi, avs, vpy, avqsv [H.264/AVC, HEVC, MPEG2, VC-1, VP8, VP9]
Environment Info
OS : Windows 10 (x64)
CPU: Intel Core i7-6700K @ 4.00GHz [TB: 4.19GHz] (4C/8T) <Kabylake>
RAM: Used 2400 MB, Total 16261 MB
GPU: Intel HD Graphics 530 (24EU) 1150MHz (21.20.16.4590)
Media SDK Version: Hardware API v1.19
But besides that, there is just one single difference in features compared to your post. It reports 10bit capability for my HD530:
Codec: HEVC
CBR VBR AVBR QVBR CQP VQP LA LAHRD ICQ LAICQ VCM
10bit depth o o x x o o x x o x o
The rest is identical. Hmmm, maybe wait for an QSVEncC update to see the real feature list of KabyLake?
(I just updated the display driver to v.4590, with older versions, feature list differs a little more)
NikosD
15th February 2017, 19:21
We got an explanation from AMD of what VBAQ means.
"VBAQ stands for “Variance Based Adaptive Quantization”.
The basic idea of VBAQ:
Human visual system is typically less sensitive to artifacts in highly textured area.
In VBAQ mode, we use pixel variance to indicate the complexity of spatial texture.
This allows us to allocate more bits to smoother areas.
Enabling such feature leads to improvements in subjective visual quality with some content."
From tests, I have seen that there is no speed impact using this feature.
Also, new VCEENC v3.05 is out fixing HEVC CBR/VBR mode, so I will upload Samsung Journey HEVC encoding by Polaris 8bit HEVC encoder.
NikosD
15th February 2017, 20:22
The last VCEEnc v3.05 has partially fixed the HEVC VBR bug, you can't still go more than 20000 Kbps.
So, my Samsung Journey HEVC encoding using VBAQ and VBR 20000, Balanced quality, pre-analysis is here:
https://www.sendspace.com/file/clckxm
The file is ~258MB < 300MB due to max VBR 20000
CruNcher
16th February 2017, 11:54
Target is far off from the 299 we basically agreed upon, so i would see your encode as handicaped not reaching that goal ;)
Hmm not sure if it's the VBAQ but perceptively scene changes are very problematic to high quantized before they become stable in your encode even fast cuts.
So overall perceptive stability is fluctuating the heaviest on the first sighting in motion currently but it needs to be seen as handicapped as well compared to easyfabs or my result which hit the target bitrate and have a clear distribution advantage.
Y:0.933980 (11.803249)
U:0.961267 (14.119180)
V:0.957400 (13.705949)
All:0.942431 (12.398135)
NikosD
16th February 2017, 12:49
So overall perceptive stability is fluctuating the heaviest on the first sighting in motion currently but it needs to be seen as handicaped as well compared to easyfabs or my result which hit the target bitrate and have a clear distribution advantage.
Picture/ video quality is sometimes very objective and because I can't see anything of what you say, can you point min:sec of the video to me for all the low quality parts of the video ?
CruNcher
16th February 2017, 12:52
I think we wont need that because the bitstream here can be very nicely separated into scenes with naming conventions :)
I would say the first scenes are pretty uniform on all encoders if we dont count percepted sharpness at all.
The first scene where visual impactfull differences become visible is the "Walking on the Edge" in the Sky you might see Banding or not, though all of our encodes are banding more or less in 8 bit :)
Then you have the transition to the "CGI Chrome Window with the Sun" where the transition is very problematic visually in your encode followed by the transition to the "Standing forest(jungle) looking into the Sky" where the transition is also problematic as it shows a to high latency from blurry to sharp which makes the impression of a DOF change but there is no DOF to beginn with.
Then comes the "Walking in the Rosefield" transition also shows problems from the "Walking on the Green Grass" scene and followed by the "Big Rose to Face Closeup" transition shows also problems really heavy prediction problems in the blend both are heavily percepted.
Then the "Rose bloom flying out of the CGI Chrome Window" shows the "DOF" effect in it's transition again.
KabyLakes TU7 has also problems with those blend transitions but not really as heavy as in your Encode result, but i count your encode result as handicapped for now as i said.
What is also really interesting your Encode shows the same Chroma/Luminance Level result as mine only Easyfabs KabyLake QSVENC result differs heavily from it's overall brightness perception.
And in those regards also our Banding result is very identical Easyfabs differs significantly as well.
"Walking on the Edge" = 00:20.000
"CGI Chrome Window Sun" = 00:27.000
"Standing forest(jungle) looking into the Sky" = 00:29.000
"Walking on the Green Grass" = 00:57.000
"Walking in the Rosefield" = 01:00.000
"Big Rose to Face Closeup" = 01:08.000
"Rose bloom flying out of the CGI Chrome Window" = 01:34.000
I didn't analyze it frame by frame but it looks like they're some heavy prediction errors maybe chroma related, that i absolutely don't percept as bad the same way in Easyfabs TU7 Kabylake or my Nvidia Maxwell Encode result, including that blurry->sharp "DOF" effect on some cuts and transitions.
easyfab
16th February 2017, 18:39
To complete the collection
Encode with :
HEVC 8bit
BF 0
Preset T4
Scene change ( I needed to use --avsw for that ) Is it better ?
https://www.sendspace.com/file/8ttalh
CruNcher
16th February 2017, 20:01
It becomes slowly indistinguishable between yours and mine on the first sight there is still the banding and brightness difference and i still find my transitional blend results at a closer look better (more stable) but overall it becomes really close without looking into the details on your TU4 result in terms of perceptable stability.
Seeking wise yes much better, the seeking results also practically match now, though your seeking time is still a tad higher it seems to the final result but you are using mkv and me mp4 so this could be a reason together with overall tad higher decoding complexity.
A typical Blend Transition difference between our encode results
Easyfab QSVENCC TU4
http://i1.sendpic.org/t/df/dfKSEG8CwXTncpwuerZmaPIYgXP.jpg (http://sendpic.org/view/1/i/a8oVdqXQqvQHRwdo6D4Xy8vTQGU.png)
CruNcher NVENCC
http://i1.sendpic.org/t/uV/uVPDfUMLvFgBVz7MphNRjFHPLyo.jpg (http://sendpic.org/view/1/i/irRdKXG8GVp8ErcyoX3XsFjcaKt.png)
After the transition
Easyfab QSVENCC TU4
http://i1.sendpic.org/t/if/ifqIZMAG4DC6K28qxyCYahtLmBD.jpg (http://sendpic.org/view/1/i/8sajVi2fA1qgx4zqB3Tc7fnBI9u.png)
CruNcher NVENCC
http://i1.sendpic.org/t/sn/snhwI1GaypuGl0FcHSkVYYMKa18.jpg (http://sendpic.org/view/1/i/83vTNgJdDXRNTgp6nSmfnP3tBSh.png)
After TU4 i guess your result gonna beat Nvidia Maxwell overall and become the reference like your Best result did pretty much before in 10 bit
And as soon as this happens wee neeed a Pascal user, though i still try to squish a little bit more out in terms of overall encoding speed lose, so i waiting for the result that destroys mine now completely from you, though overall you beating my system to death already with that power output efficiency doing this on your fanless laptop ;)
For completeness
Easyfab QSVENCC TU4
SSIM
Y:0.941868 (12.355859)
U:0.963671 (14.397485)
V:0.960298 (14.011911)
All:0.948574 (12.888148)
lets recompare that with NikoSD AMD VCEEncC Encode which i personally also find perceptively overall inferior in it's current state going into the direction of even unacceptable in terms of stability
Y:0.933980 (11.803249)
U:0.961267 (14.119180)
V:0.957400 (13.705949)
All:0.942431 (12.398135)
CruNcher
17th February 2017, 18:29
@NikosD
Did you tried a retranscode of your current VCENCC result without that VBAQ ?
easyfab
17th February 2017, 18:49
@CruNcher
Thnaks for the test
QSVENCC TU4 Bframes 0
All:0.948574 (12.888148)
QSVENCC TU 4 10 Bit Encode with B-frames
All:0.955688 (13.534776)
As 10bit shouldn't give so must improvement, Bframes is a real advantages for intel VS Nvidia and AMD.
Now I hope than 1 day some mores features like look-ahead will be add.
NikosD
17th February 2017, 19:34
@NikosD
Did you tried a retranscode of your current VCENCC result without that VBAQ ?
Here you are:
https://www.sendspace.com/file/g52pd0
CruNcher
17th February 2017, 22:44
I guess Nvidias Spatial AQ and AMDs VBAQ will be pretty much the same as Fionas VAQ spatial
A small Metric lose overall most issues stay the same like in the VBAQ enabled result so VBAQ isn't at least responsible for those stability problems.
NOVBAQ
SSIM
Y:0.932798 (11.726161)
U:0.961508 (14.146268)
V:0.957671 (13.733606)
All:0.941728 (12.345423)
JohnLai
18th February 2017, 05:39
@CruNcher
AMD VBAQ still doesn't work.......nor is preanalysis......
NikosD
18th February 2017, 09:05
The first scene where visual impactfull differences become visible is the "Walking on the Edge" in the Sky you might see Banding or not, though all of our encodes are banding more or less in 8 bit :)
Then you have the transition to the "CGI Chrome Window with the Sun" where the transition is very problematic visually in your encode followed by the transition to the "Standing forest(jungle) looking into the Sky" where the transition is also problematic as it shows a to high latency from blurry to sharp which makes the impression of a DOF change but there is no DOF to beginn with.
Then comes the "Walking in the Rosefield" transition also shows problems from the "Walking on the Green Grass" scene and followed by the "Big Rose to Face Closeup" transition shows also problems really heavy prediction problems in the blend both are heavily percepted.
Then the "Rose bloom flying out of the CGI Chrome Window" shows the "DOF" effect in it's transition again.
KabyLakes TU7 has also problems with those blend transitions but not really as heavy as in your Encode result, but i count your encode result as handicapped for now as i said.
What is also really interesting your Encode shows the same Chroma/Luminance Level result as mine only Easyfabs KabyLake QSVENC result differs heavily from it's overall brightness perception.
And in those regards also our Banding result is very identical Easyfabs differs significantly as well.
"Walking on the Edge" = 00:20.000
"CGI Chrome Window Sun" = 00:27.000
"Standing forest(jungle) looking into the Sky" = 00:29.000
"Walking on the Green Grass" = 00:57.000
"Walking in the Rosefield" = 01:00.000
"Big Rose to Face Closeup" = 01:08.000
"Rose bloom flying out of the CGI Chrome Window" = 01:34.000
I didn't analyze it frame by frame but it looks like they're some heavy prediction errors maybe chroma related, that i absolutely don't percept as bad the same way in Easyfabs TU7 Kabylake or my Nvidia Maxwell Encode result, including that blurry->sharp "DOF" effect on some cuts and transitions.
Sorry, but I can't see differences between NVIDIA, AMD and the original source.
I have already posted a free link before for JohnLai, where you can upload and compare still images by hovering your mouse on them.
Maybe it would be better if we could compare still images that way, in order to find out differences that I can't see.
(the sun is very bright here, probably I have to test again in complete dark room with screen light only)
CruNcher
18th February 2017, 11:09
Interesting they're actually even more issues but i don't count them because they are visually the same in all 3 encodes "Walking on the Edge" = 00:20.000 is one of those even if it comes very differently out in easyfabs result then our encodes.
Ok ill show you the heaviest stability issues that get immediately perceived by me on your encode result as being disturbing braking the visual perception flow for myself in motion compared to easyfabs my encode result and the source, here we go.
"CGI Chrome Window Sun" = 00:27.000
http://i1.sendpic.org/t/il/ilvPQNzOEpg5uAg3Wkc0XccOaVB.jpg (http://sendpic.org/view/1/i/nqUtq00xH1bozeHpYJP3K23ve5A.png)
Looks pretty much like the Keyframe itself is total mess and the whole thing recovers slowly
"Standing forest(jungle) looking into the Sky" = 00:29.000
http://i1.sendpic.org/t/42/42ADsUCfvvXJAbml2Z9jwiCvKD4.jpg (http://sendpic.org/view/1/i/kXCSb3zGRKw8ikwnYzbgdgJ0wif.png)
This here is actual the hardest where i feel i would have big problems explaining someone why this plays out bad in motion when he just sees that on the first look super nice spatial result, and it is the way the disolve itself plays out at it's peak creating strange darker looking blocks inside of the transition.
Though im rather surprised you didn't perceive it the same way as not comfortable compared to the other results and way different even.
"Walking on the Green Grass" = 00:57.000
"Walking in the Rosefield" = 01:00.000
http://i1.sendpic.org/t/mz/mzcbTD5cWgJlzPeFoFGC2sUxChs.jpg (http://sendpic.org/view/1/i/gVcIA69guQ2ohCHFW8n5zNLPuzA.png)
This is hopefully pretty obvious and im really heavily surprised you didn't perceived it as uncomfortable compared again to the other results.
"Big Rose to Face Closeup" = 01:08.000
http://i1.sendpic.org/t/hE/hEjLV2UPkbv2jFfHyhmU1Lh1DZm.jpg (http://sendpic.org/view/1/i/4hNNgYyxxoknBFcWilmd6ChodEF.png)
Also again Keyframe messy and recovering slowly with prediction errors visible
"Rose bloom flying out of the CGI Chrome Window" = 01:34.000
http://i1.sendpic.org/t/yy/yy4Atj0hkop0lgEfm4EZ6e7PF33.jpg (http://sendpic.org/view/1/i/1VnEMB2SeMspUPyDITOy3Leccfl.png)
These are the things that make me feel there is a lot going wrong in your encode in stability on the encoder side and im not sure if even more bits would help really.
NikosD
19th February 2017, 16:16
@JohnLai and CruNcher
What is the difference in terms of quality only, between Maxwell 2nd generation HEVC encoder vs Maxwell GM206 vs Pascal ?
HEVC quality only, not speed at all.
CruNcher
19th February 2017, 22:09
Between 2nd Maxwell and GM206 virtually 0 between both and Pascal it's SAO and with it i guess the Encoder could beat Intels without it ;)
i try now to get closer to easyfabs TU1 result and see if i can somehow avoid to leave the 300 mb limit and how that will come out perceptually and hoping to stay in the 0.5x range ;)
But so far Nvidia doesn't do overly bad especially staying close on the Stability side of that balanced TU4 without B-frames is a nice achievement so far :)
Though if Intel further optimizes that will be gone fast as well ;)
JohnLai
20th February 2017, 04:04
Exactly the same as CruNcher said ~ Smooth All Object ~
Between maxwell and pascal, pascal encode quality is highly dependent on scene type due to SAO. (SAO can't be turned off)
Example, with SAO, blue sky is smooth and look 'natural'. But SAO causes somehow blurry seawater wave (or any fluid based) movement.......
NikosD
21st February 2017, 21:30
"CGI Chrome Window Sun" = 00:27.000
http://i1.sendpic.org/t/il/ilvPQNzOEpg5uAg3Wkc0XccOaVB.jpg (http://sendpic.org/view/1/i/nqUtq00xH1bozeHpYJP3K23ve5A.png)
Looks pretty much like the Keyframe itself is total mess and the whole thing recovers slowly
No problem here.
"Standing forest(jungle) looking into the Sky" = 00:29.000
http://i1.sendpic.org/t/42/42ADsUCfvvXJAbml2Z9jwiCvKD4.jpg (http://sendpic.org/view/1/i/kXCSb3zGRKw8ikwnYzbgdgJ0wif.png)
This here is actual the hardest where i feel i would have big problems explaining someone why this plays out bad in motion when he just sees that on the first look super nice spatial result, and it is the way the disolve itself plays out at it's peak creating strange darker looking blocks inside of the transition.
Though im rather surprised you didn't perceive it the same way as not comfortable compared to the other results and way different even.
No problem here.
"Walking on the Green Grass" = 00:57.000
No problem here.
"Walking in the Rosefield" = 01:00.000
http://i1.sendpic.org/t/mz/mzcbTD5cWgJlzPeFoFGC2sUxChs.jpg (http://sendpic.org/view/1/i/gVcIA69guQ2ohCHFW8n5zNLPuzA.png)
This is hopefully pretty obvious and im really heavily surprised you didn't perceived it as uncomfortable compared again to the other results.
Yes, this is the only one I can see difference, but if you hadn't mentioned that, I wouldn't see it at all.
"Big Rose to Face Closeup" = 01:08.000
http://i1.sendpic.org/t/hE/hEjLV2UPkbv2jFfHyhmU1Lh1DZm.jpg (http://sendpic.org/view/1/i/4hNNgYyxxoknBFcWilmd6ChodEF.png)
Also again Keyframe messy and recovering slowly with prediction errors visible
No problem here.
"Rose bloom flying out of the CGI Chrome Window" = 01:34.000
http://i1.sendpic.org/t/yy/yy4Atj0hkop0lgEfm4EZ6e7PF33.jpg (http://sendpic.org/view/1/i/1VnEMB2SeMspUPyDITOy3Leccfl.png)
No problem here.
CruNcher
22nd February 2017, 22:07
Text is always for the lower spatial output result
But interesting if the Decoding is different and you don't see that prediction errors especially the colored blocks even going through frame by frame in the "Rose bloom flying out of the CGI Chrome Window" Scene something goes wrong.
Im also pretty sure by now Rigaya took some heavy bad decisions on his encoder side default tuning for NVENC on perceptual ground i find his decisions questionable, though this encode sample shows that not so heavily as other will do.
Even if Metrics want me to tell otherwise.
But im still behind that and trying to understand why Rigaya does what he does on NVENCC ;)
CruNcher
24th February 2017, 18:18
My highest 23 Mbits Transcoding speed result so far of 60 FPS Hevc 4K Transcoding 8 bit Decoding(CPU)/Encoding(GPU/SIP)
0.64x
NikosD
25th February 2017, 16:04
This is the same "Samsung Journey" file transcoded to HEVC 8 bit by VCEEnc v3.05v2, but this time using Avisynth LWLibav for dithering/demuxing/decoding inside latest StaxRip.
I don't know if it has any difference from CLI VCEEnc v3.05v2 using ffmpeg SW demuxer/decoder.
I'm a little blind regarding to those small differences :)
https://www.sendspace.com/file/bubcxt
CruNcher
26th February 2017, 08:49
subjective indeed but a colored block that doesn't belong initially in the scene at all is no small difference for me personally and a quantization that fluctuates so wide and recovers slowly also isn't to be rated "small difference" imho.
Small artifacts here and there are overall acceptable but not such complete fails.
NikosD
26th February 2017, 10:00
My highest 23 Mbits Transcoding speed result so far of 60 FPS Hevc 4K Transcoding 8 bit Decoding(CPU)/Encoding(GPU/SIP)
0.64x
Using latest VCEEnc v3.05v2 as a benchmark tool, which leverages AMF API for HW decoding/ encoding, I did some tests today on the HW H.264 decoder/encoder of Polaris RX 470 card.
It seems that the HW H.264 encoder doesn't change its encoding speed by using different rate modes (CBR/VBR, CQP) or different bitrate for CBR/VBR (maximum bitrate for VBR/CBR H.264 encoding is 100Mbps and above 700Mbps for CQP)
The one and only parameter that affects HW encoding speed is quality.
So, for 1080p H.264 clip the encoding speed is:
fast -> ~140 fps
balanced -> ~ 100 fps
slow -> ~ 55 fps
The HW H.264 decoder can keep up with the fast quality setting(~140 fps at 1080p) with clips at 60Mbps.
After that limit, the HW H.264 decoder drops its speed reaching ~95fps at 1080p100Mbps H.264 clip.
My Core i5 2400 can reach 1080p at 140fps when bitrate is 20Mbps.
So, although the Polaris HW H.264 decoder seems a little slow, it's still 3 times faster than Core i5 2400.
Can you test your Maxwell 2nd gen HW H.264 decoder/encoder that way ?
Similar HEVC tests are not feasible at the moment, as the quality parameter for HEVC is still not working (it has the same speed for all settings slow/balanced/fast)
CruNcher
26th February 2017, 19:10
i'll look into it after the HEVC part
Btw could you test something ?
Could you show me your Systems/Decoder/Render Jitter response when the wave is crashing down once AMD DXVA and Once CPU (Decoder of Choice) ?
http://i1.sendpic.org/t/yh/yhO3mo0GcHIhbqaaOXmaeo7QV5p.jpg (http://sendpic.org/view/1/i/roFjjM2d36Uedz2nyZFXWrlw9Jk.png)
https://www.sendspace.com/file/cr4n8n
PS: That StaxRip Transcode didn't come out any better then the direct VCEEnc it suffers from the same Encoder Issues.
Want to slowly leave The Jorurney ;)
Sony A7
Sony XAVC->Nvidia HEVC
https://www.sendspace.com/file/5l4m3v
This should also Play very well on GM204 Hybrid Accelerated via CUVID and DXVA CopyBakck not so good via DXVA (native), maybe though even there if you use the new "Optimize for Compute" option in the newer Drivers same for Cyberlink HAM though better then DXVA native ;)
So if you have to much dropped frames issues switch to DXVA Copy Back,CUVID,HAM and/or try that new Optimize for Compute Option.
CruNcher
28th February 2017, 10:16
One of the the very firs Public Encoder Streams Retranscoded at Half the Bitrate
DivX/Mainconcept->Nvidia Hevc
Speed = 1.7x
Perceptually partly really interesting should be 2nd or 3rd Generation now
https://www.sendspace.com/file/050xe5
PSNR
y:42.833979
u:46.333830
v:47.580143
average:43.819431
min:18.154593
max:73.970364
SSIM
Y:0.987446 (19.012031)
U:0.985141 (18.280104)
V:0.990760 (20.343423)
All:0.987614 (19.070676)
NikosD
28th February 2017, 14:26
i'll look into it after the HEVC part
Reading Anandtech's R9 285 H.264 HW decoder results here http://www.anandtech.com/show/8460/amd-radeon-r9-285-review/4, I can see that it's a little faster than 750 Ti.
I think that the HW H.264 decoder of R9 285 is the same like mine, RX 470 and that 750 Ti has the same Purevideo decoder (VP6) like yours Maxwell 2nd generation.
So, I updated my table of H.264 benchmarks here https://forum.doom9.org/showthread.php?p=1712350#post1712350 with some useful/ strange comments here https://forum.doom9.org/showthread.php?p=1798958#post1798958 regarding Polaris HW H.264 decoder.
Is it possible to test your VP6 HW H.264 decoder on those clips ?
You can find all samples here ftp://helpedia.com/pub/multimedia/x264/testvideos/2011%20-%2002%20-%20H.264%20CPU%20DXVA%20codec%20comparison%20-%20Core2Duo%20vs%20UVD%202.2/ and here ftp://helpedia.com/pub/multimedia/x264/testvideos/2012%20-%2001%20-%20QuickSync%20vs%20UVD%202.2%20vs%20VP4/
JohnLai
28th February 2017, 14:32
This should also Play very well on GM204 Hybrid Accelerated via CUVID and DXVA CopyBakck not so good via DXVA (native), maybe though even there if you use the new "Optimize for Compute" option in the newer Drivers same for Cyberlink HAM though better then DXVA native ;)
So if you have to much dropped frames issues switch to DXVA Copy Back,CUVID,HAM and/or try that new Optimize for Compute Option.
Strange.:confused:
There is no hybrid HEVC CUVID decoding as far as I know.
CUVID is designed for pure fixed function hardware utilization.
NikosD
28th February 2017, 14:39
CUVID uses DXVA2, so if DXVA2 implementation is hybrid then CUVID is hybrid too.
If it's pure fixed-fuction, then it's fixed-fuction.
It depends on DXVA2
JohnLai
28th February 2017, 15:35
CUVID uses DXVA2, so if DXVA2 implementation is hybrid then CUVID is hybrid too.
If it's pure fixed-fuction, then it's fixed-fuction.
It depends on DXVA2
:confused:
Nope, I don't think CUVID uses DXVA2.
CUVID is supposed to be different in implementation than DXVA2.
Beside, for all HEVC 8bit samples with GM204, LAVFilter displays "active decoder : avcodec"
If it works, LAVFilter should displays "active decoder : cuvid"
NikosD
28th February 2017, 15:42
I'm under the impression that nevcairiel, the developer of LAV filters, had once told me that NVCUVID is like a wrapper of DXVA and has nothing to do with CUDA.
It's been years since.
I could remember falsely.
nevcairiel
28th February 2017, 16:00
CUVID can work in two modes, either direct hardware access or through DXVA. In DXVA mode, CUVID can access the hybrid DXVA decoders, but that mode doesn't work on Windows 10, only the pure CUVID mode works on 10, which only has fixed-function decoders.
JohnLai
28th February 2017, 16:05
CUVID can work in two modes, either direct hardware access or through DXVA. In DXVA mode, CUVID can access the hybrid DXVA decoders, but that mode doesn't work on Windows 10, only the pure CUVID mode works on 10, which only has fixed-function decoders.
Thanks for the clarification.
I am using win10, so that is the reason on why I only see "avcodec" when decoding 8bit hevc using GM204.
:goodpost:
CruNcher
28th February 2017, 22:46
Overall it also makes no sense the Hybrid Decoder will die and CPU Multithreading be more efficient alone (especialy on a relative low piower one) at a certain complexity range to be system efficient with it you would need to dynamically switch decoders based on the overall bitstream complexity,or get a nice async flow with lowest CPU overhead. ;)
In that regards HEVC is a tad better Parallelizeable but not really by that much that it could threaten the CPU or ASIC efficiently yet.
NikosD
1st March 2017, 04:20
Vega has Virtualized Encode feature, which according to Anandtech means:
AMD has implemented an optimized video encoding path for virtualized environments on their GPUs.
A game streaming service requires that the contents of upwards of several virtual machines be encoded quickly, so this would be the logical next step for AMD’s on-board video encoder (VCE) by making it efficiently work with virtualization.
There is also a session with Mikhail Mironov in GDC
http://schedule.gdconf.com/session/true-audio-next-and-multimedia-amd-apis-in-games-and-vr-applications-development-presented-by-amd
CruNcher
1st March 2017, 10:33
We will also discuss research projects based on Advanced Media Framework (AMF). The key focus of this presentation is integration with game engines for better VR immersion.
Makes sense to virtualize it also for Content Protection Security Reasons, most can run pretty nicely Async on GCN then virtualized it should practically never interfere with the Game Threads and causing lock situations like it happens still to often currently.
Also it fits to the virtualization of the new Shader Model 6 running via LLVM
And overall it fits to the Security Model that was being on the Roadmap since 2006 now ;)
JohnLai
1st March 2017, 12:27
Virtualized Encode eh....
Sounds good. But can AMD actually make sure its VCE software stack actually works?
Both rigaya and xaymar have headache with AMD AMF implementation.
NikosD
1st March 2017, 12:31
That Mikhail Mironov guy who is a lead developer of AMF API and replies on github, looks like he is doing a good job.
AMF is far better than previous AMD MediaSDK.
Next version will polish bugs and make features work, I think.
CruNcher
1st March 2017, 23:40
Though surprisingly he's not from the MSU ;)
CruNcher
4th March 2017, 14:54
using the Hardware Decoder/Encoder same time both NVENC Engines (Cores) at 50% (half their utilization maximum)
Pre CUDA Maxwell Driver Sheduler optimization (3D tasks AERO/DWM and Firefox)
0.931x (57 FPS)
Almost reached stable 60 FPS for UHD now @ 20 Mbits X264->Nvidia HEVC (~-3 FPS)
http://i1.sendpic.org/i/bP/bPW89uDcWkoLwoDorkuSi10Zg0V.png
~50W idle inc peripherals USB3/2 ...
~120W encoding
Overall ~70W for NVDEC/NVENC + CUDA +CPU +Memory (System/Storage) together
Maximum possible System output is somewhere ~350-380W
Node 0 = 3D/Shader Compiler
Node 1 = Cuda/Compute
Node 2 = ?
Node 3 = NVDEC
Node 4 = ?
Node 5 = ?
Node 6 = Copy
Node 7 = NVENC
Node 8 = NVENC
Node 9 - 14 = ?
Almost 60 FPS Decoding Stable (only Intra Frame Latency issues resist) on the Hybrid (GM204) CUDA Decoder (i guess most problems will be solved with the new Driver Compute Sheduler mode)
8 Bit Decoding Complexity is around 3 Sandy Bridge Cores and around 1.5 with the GM204 CUDA Decoder (so halved).
http://i1.sendpic.org/i/2T/2TBEHSsEIOg6tAvjgdZDGbncDtR.png
Zetto
4th March 2017, 18:33
anyone expects 1080 ti to have an updated nvenc engine?
sneaker_ger
4th March 2017, 18:45
Weren't video capabilities in the past related to the code(name) of the chip? GTX 1080 Ti is supposed to be GP102 like Titan X so I would not expect new capabilities. Nor VP9 10 bit (like GTX 1050 (Ti) aka GP107) for that matter. GTX 980 Ti also had less capabilities than GTX 960 despite the later release date.
CruNcher
4th March 2017, 19:00
The More CUDA Power the more tasks you can route their like 10 bit Decoding overhead ASIC is only of concern mostly on Power efficiency it makes the most sense for underpowered Shader cards or Mobile depending on the Target audience ;)
Nvidia did this in the past on many levels i remember Motion Adaptive Deinterlacing only enabled on 1 Specific Hardware and by acccident inside a Beta Driver for every Shader count, it was daunting slow and unusable on the lower Shader card though so Economicaly nonsense to keep it and possible support overhead following it and then later it became the default, as the shader power rose for the more consumer targeted cards ;)
Im pretty sure Nvidia also tests inside CUDA their future ASIC optimizations and have a pretty nice efficient conversion workflow from CUDA(GPU)->ASIC :)
The Encoder itself is still updated via CUDA additions like OpenCL can be used for x264 Lookahead Nvidia uses Cuda for their Lookahead additionally for their AQ and 2pass and most probably coming weight-b as well :)
It's fully up to Nvidias Buisness decisions what they enable where and what makes sense for which Platform and Predicted target audience System Platform resources (nowadays gathered by Nvidia via direct telemetry system data of their target customers) ;)
So if Nvidia want's to they could enable 10 Bit VP9 Decoding\Encoding Cuda accelerated on the GP102 without needing the ASIC at all, it's a pure business and resource control decision and how much sense it makes with the Shader count available for certain Targets and Usage Scenarios of Customers and Competition advances and their decisions and progress.
CruNcher
5th March 2017, 13:16
@NikosD
Nvidias HEVC Encoder already performs better then Nvidias AVC result wise HEVC is overall often sharper especially vs Nvidia AVC with B-Frames
But Metrics neither PSNR nor SSIM agree with me they would rate Nvidia AVC + B-frames results as overall better.
I cant agree.
There are 2 Main factors immediately visible Higher Sharpness (Detail retention) and better Banding avoidance at the same Target Hit without even touching AQ.
The Better Banding avoidance surprised me a little as i expected to see that mainly from 10 Bit but nope in Nvidias case their overall HEVC tuning does it better by default vs their AVC + B-frames even already heavily visible at 8 bit.
And you endup at pretty much the same Encoding speed in this case slow (AVC) vs fast (HEVC) gained almost the same overall Performance results with Nvidia HEVC leading visually overall even if PSNR/SSIM disagrees by a small margin towards Nvidia AVC.
Though surprisingly in Rigayas Encoder tuning case the Better banding avoidance by default gets pretty much lost completely and he achieves a higher PSNR/SSIM.
the highest speed i reached so far with the GTX 970 (GM204)
on the i5-2400
1080p 8mbits = 235 FPS ~9.4x @ 25 FPS
4K DCI 20 mbits = 57 FPS ~0.93x @ 59,940 FPS (almost identical to AVC slow with b-frames)
I never saw such failures in either Rigayas, FFMpegs or Nvidias own sample Encoder as your AMD result produced on the Journey sample with the Polaris Encoder no matter the bitrate.
Im pretty unsure about Rigayas work though he seems to follow a specific targeted goal that might be problematic, he seems to be to much fixed on something i really don't hope he just blindly tunes based on Metrics though.
Though i couldn't test Temporal AQ yet with AVC + B-frames it could have overall some rather big impact.
NikosD
8th March 2017, 13:20
I have exchanged a few emails today with the developer of OBS Studio (xaymar) regarding AMF (AMD's API) and Polaris HW.
He confirmed these issues:
1) Only HEVC content is affected by Pre-Analysis for now and only with CBR/VBR.
You can only really measure the impact using PSNR and SSIM, since it's visual impact is pretty much non-existent except for a few test scenes.
2) VBAQ also only works on CBR/VBR and the effects are pretty much non-existent, at least perceived effects.
3) HEVC does not really have Quality Presets. You can see a slight difference between Speed and Balanced/Quality, that is all.
But he disagrees that you can't set REF frames on H.264 encoding that he has already tested.
He told me that you just need to make sure that you are within either H264 or H265 limits.
Well, I don't know about the last one but for the first three issues we are all agree.
NikosD
8th March 2017, 15:37
Using latest VCEEnc v3.05v2 as a benchmark tool, which leverages AMF API for HW decoding/ encoding, I did some tests today on the HW H.264 decoder/encoder of Polaris RX 470 card.
It seems that the HW H.264 encoder doesn't change its encoding speed by using different rate modes (CBR/VBR, CQP) or different bitrate for CBR/VBR (maximum bitrate for VBR/CBR H.264 encoding is 100Mbps and above 700Mbps for CQP)
The one and only parameter that affects HW encoding speed is quality.
So, for 1080p H.264 clip the encoding speed is:
fast -> ~140 fps
balanced -> ~ 100 fps
slow -> ~ 55 fps
The HW H.264 decoder can keep up with the fast quality setting(~140 fps at 1080p) with clips at 60Mbps.
After that limit, the HW H.264 decoder drops its speed reaching ~95fps at 1080p100Mbps H.264 clip.
My Core i5 2400 can reach 1080p at 140fps when bitrate is 20Mbps.
So, although the Polaris HW H.264 decoder seems a little slow, it's still 3 times faster than Core i5 2400.
Can you test your Maxwell 2nd gen HW H.264 decoder/encoder that way ?
Similar HEVC tests are not feasible at the moment, as the quality parameter for HEVC is still not working (it has the same speed for all settings slow/balanced/fast)
Doing the same test with HW H.265 decoder, it seems that it has a very similar performance to HW H.264 decoder and it's a tad faster.
The HW H.265 decoder can keep up with the fast quality setting(~140 fps at 1080p) of H.264 encoder with clips at 60Mbps.
After that limit, the HW H.265 decoder drops its speed reaching ~108fps at 1080p110Mbps H.265 clip.
So it's faster in higher bitrates than HW H.264 decoder.
Benchmark tests were done based on Jellyfish (H.264 & H.265) and Birds (H.264) bitrate samples.
JohnLai
8th March 2017, 18:25
But he disagrees that you can't set REF frames on H.264 encoding that he has already tested.
He told me that you just need to make sure that you are within either H264 or H265 limits.
So, one Polaris VCE hevc or h264 1080p sample with 3 reference frame? Smaller size and shorter duration please....:D
Surely this is within the 'limits' of level 4.1?
Note: for AMD r7 260x, VCE H264 B-frame can make use of 2 reference frames just fine. However, setting B-frame to 0 = each P-frame only makes use of one preceding frame as reference.
Since Polaris VCE doesn't have B-frame support for H264....well....:(
NikosD
9th March 2017, 17:20
Note: for AMD r7 260x, VCE H264 B-frame can make use of 2 reference frames just fine. However, setting B-frame to 0 = each P-frame only makes use of one preceding frame as reference.
Since Polaris VCE doesn't have B-frame support for H264....well....:(
I have sent you a PM, but probably you didn't see it.
The reply regarding this issue was:
I-Frames may only reference the last I-Frame if in Intra-Refresh/Slice mode
P-Frames may reference the previous I-Frame or P-Frame
P-Frames may not reference the next I-Frame or P-Frame
P-Frames can reference I-Frames and P-Frames further in the past (Reference Frame range)
B-Frames must reference the next P- or I-Frame
B-Frames must reference the previous P- or I-Frame
So you can end up with the following: IPBBBPBBBPBBB.
Your Reference Frame group restarts with the I-Frame and slowly builds up to 16 frames, until it exceeds 16 frames and one of the P or B frames is dropped out.
This continues on until the next I-Frame, where it restarts again.
I have verified this behavior myself on R9 285, R9 390 and RX 480 using ffprobe and Elecard StreamEye.
JohnLai
9th March 2017, 18:58
I have sent you a PM, but probably you didn't see it.
The reply regarding this issue was:
I saw it. No time to reply yet.
P-Frames can reference I-Frames and P-Frames further in the past (Reference Frame range)
The problem = decoded picture buffer in elecard shown otherwise for Polaris HEVC sample. One can set 5 or 16 reference frames for VCE, but decoded picture buffer clearly has only ONE preceding frame as reference in use to display current frame.
Your Reference Frame group restarts with the I-Frame and slowly builds up to 16 frames, until it exceeds 16 frames and one of the P or B frames is dropped out.
For this one.....something like nvidia hevc sample. It has a lot of previously decoded frames stored in DPB, but only make use of one preceding frame as reference to display current frame.
NikosD
9th March 2017, 19:12
The problem = decoded picture buffer in elecard shown otherwise for Polaris HEVC sample. One can set 5 or 16 reference frames for VCE, but decoded picture buffer clearly has only ONE preceding frame as reference in use to display current frame.
He was clearly insisted on H.264 encoding, not HEVC.
Have you analyzed VCE H.264 samples for REF frames ?
NikosD
10th March 2017, 05:50
One can set 5 or 16 reference frames for VCE, but decoded picture buffer clearly has only ONE preceding frame as reference in use to display current frame.
I got his final reply, you were right about that from the beginning.
The AVC encoded content by default only references Index 0 (the last possible reference frame).
This is likely by design and hasn't been improved on, the VCE firmware has not seen many modifications since ATI days.
Same thing for HEVC encoded content, though the HEVC core is rather buggy at the moment.
JohnLai
10th March 2017, 09:19
He was clearly insisted on H.264 encoding, not HEVC.
Have you analyzed VCE H.264 samples for REF frames ?
Based on your Polaris VCE H264 samples last, same issue for P-frame in those samples. One ref only.
Bonaire and Hawaii VCE H264 supports B-frame where each B-frames can properly make use of 2 preceding frames as references.
Nvidia Nvenc H264 B-frame also can refer up to 4 preceding IPB frames. However, if NVENC is set not to use B-frame, then its P-frame only makes use of 1 preceding P-frame as reference.
Only Intel QSV reference frame actually somehow works for both H264 (B-frame can refer to I, P and B, meanwhile P-frame can only refer to another preceding P-frame) and HEVC (But it is GBP, not exactly IPB).
There is no perfect fixed function encoder :( . Intel, AMD and Nvidia fixed function encoders omit a lot of 'features'. Speed/quality tradeoff.
Stick with software encoders for the best quality per bitrate plus flexibility.
NikosD
10th March 2017, 09:43
There is no perfect fixed function encoder :( . Intel, AMD and Nvidia fixed function encoders omit a lot of 'features'. Speed/quality tradeoff.
Stick with software encoders for the best quality per bitrate plus flexibility.
This is not an option for H.265 encoding.
The x265 SW encoder, although it has extreme optimizations it's still extremely slow.
x264 nowadays with multicore and high frequency processors is OK.
But still, HW transcoders provide full HW hardware transcoding with 0% CPU utilization with good enough results regarding visual quality and bitrate.
CruNcher
11th March 2017, 08:09
0% is a little extreme even if you wouldn't count copy related overhead (transcoding with very low sub 1% CPU utilization) ;)
Soon new results for both Nvidia and Intel on both of their newest Cores available.
I decided to get a combined Mobile x86 Platform todo further testing not as low power limited overall as Easyfabs though (35W/45W target with roughly i expect max of 120W).
Before Raven Ridge arrives.
Should be enough Power to reach the 4K 60 fps target compared to my 170W Desktop output currently ;)
so as i said Kaby Lake + 1050 TI (GP107) both newest time to market VPU and pretty much GPU Cores as well ;)
And then later a Mobile Raven Ridge System in compare where im much much more excited about overall to see and hope AMD will have most of the Problems in the Encoder on the Software side solved by then ;)
Though the loss of the N17P-G1 is heavy compared to a 1060 you can calculate ~60% shader efficiency loss that balance feels wrong for just 10/12 bit VP9/HEVC Decoding/Encoding.
only with really good code optimization you end at the exact 50% loss
this would also impact and hit the encoder overall in terms of its cuda performance parts by a not so negligible performance amount.
Which also pretty much is the Performance of my GTX 970 Desktop GM204 that 50% loss should be currently what i drive @ 170W out of my higher Maxwell Shader count on the Encoder side.
Which shows once again Pascal is nothing more then a Paxwell ;)
CruNcher
19th March 2017, 12:02
@easyfab
could you please post your current dxvachecker output
it is identical to this right ?
http://www.notebookcheck.com/fileadmin/Notebooks/MSI/CX72-7QL/dxva_kaby.png
easyfab
19th March 2017, 13:28
@CruNcher
here the mine : http://pastebin.com/sfx4z608
It's the same with some name modification ( newer version ? )
easyfab
19th March 2017, 13:52
I aslo tried some AVC encode with vaapi.
And with this (https://github.com/01org/intel-vaapi-driver/pull/74) It become really good. I gain more than 1db for SSIM 16->17 . I need to test a litlle more but it could be better than QSV. Vaapi HEVC is not so good for the moment.
I also tried VP9 HW encode with libyami but I Couldn't set bitrate correcltly. I will wait that libav got it in (https://lists.libav.org/pipermail/libav-devel/2017-March/082884.html) to make some more tests.
CruNcher
19th March 2017, 13:52
yes its from the Ultrabook CPU line introduction end of 2016 by reviewers little older also before
Added "VP9_VLD_Intel" as an alternative name of Decoder Device
though ultrbook cpu seems to be rather rare for flexible Mobiles with at least thunderbolt 3 connection and Nvidia GPU combination
or i dont see it in the list im studying because the products are ridiculously higher priced ;)
when we go lower then 120W Max Mobile Platform with still pretty nice and sufficient overall CPU omph ;)
Though 120W Max ist still more efficient then the PS4 Pro overall, though also far away in price *grml* ;)
above that you go into Game Marketing Branding area and then it starts to become really expensive for shit you don't really want ;)
my current favors are
i5-7300HQ
i5-7200U
i5-7300U
:)
The Lenovo Yoga 720-13IKB was my favorite though with 1050 ti practically for a very small uptake now in addition for the HQ series it seems a very bad overall decision now even with that higher tdp on the CPU side and max of 120W TSP is still very nice.
a 7200U/7300U featuring the GTX 1050/TI would be nice though :)
Though investments for all those current OEM seems to high i wonder if Microsoft could pull it of based on their SurfaceBook Designs they should have the know how needed to bring it in, though it would be more expensive again, way more expensive ;)
Though maybe i go total insane and build the entire thing based on the new i7-7500U 2in1s and the whole GTX 1050 TI GPU chain sideways with Thunderbolt 3 and transfer it into a 3in1 that way with ~20W higher efficiency then a i7-7700HQ series CPU for the Core Unit ;)
http://www.notebookcheck.com/fileadmin/Notebooks/MSI/CX72-7QL/hevc10.png
The cost for that efficiency would still somewhat acceptable and i would have some future headroom for some years (if everything goes well) with that system though very high overall investment to amortize i would surely get somewhere in the range of a current Surfacebook for the complete setup (never paid or even thought to pay so much for a x86 build) :D
Though not really fair for a future Raven Ridge AM4 compare far away from AMDs Target overall, though would be interesting to see how their 4 Cores will holdup vs the overall efficiency.
Definitely the small lead they had buildup with their Carizzo UVD decision or should we rather say wrong time to market planing faded away by now entirely.
i5/i7-7x00U will become very popular hard to bite through it for AMD without a overall good balance and pricing on the Mobile side again.
easyfab
19th March 2017, 19:09
Just for fun a preview of HW VP9 Encode ( with experimental VP9_vaapi ). speed 40fps on my little hd graphics 615. I thought it would be worse
https://www.sendspace.com/file/qees6f
CruNcher
19th March 2017, 20:54
Must be a premiere first VP9 bitstream out of a GPU Accelerated Encoder :)
i guess quality and stability wont be up to what we showcased here so far on the HEVC side and you on TU7 ?
PS: Ho looks partly even more stable then NikosD current VCE result, but obviously being worked on ;)
easyfab
19th March 2017, 21:45
And another expimental HW VP9_vaapi encode but with -bf 3 option ( 3 bframes or whatever it's called for VP9 )
IMO it's a little better with it than without. I will check with SSIM/PSNR later be confirm.
https://www.sendspace.com/file/ei5z1y
CruNcher
20th March 2017, 00:23
This is a also a very nice efficient combination overall
Lenovo created a very sane configuration here but it didn't showed up in my list after the Yoga so surely not that cheap and most probably they will have eradicated the thunderbolt port :D
https://www.youtube.com/watch?v=kK84987AOwg
It will be super interesting to see if AMD really can get up to that efficiency with Zen now :)
Pretty awesome results on that channel if you think about the max power constraints.
Nvidias DCE is so freakting efficient
Some pretty nice low latency recordings FBC most probably going directly into NVENC :D
https://www.youtube.com/watch?v=mwiY2OOz1sY
Predator,Aorus,Rog, MSI Gaming and now we have the Legion Brand.
luigizaninoni
22nd March 2017, 18:59
excuse me, where can I find an explanation of the various options of intel h265 hardware encoder (kaby lake) ? I am using qsvencc via staxrip, but there are over 70 parameters and I really can't understand what most of them do
JohnLai
30th April 2017, 18:14
Hmm....seem like AMF 1.4.2 was released few days ago.....
https://github.com/GPUOpen-LibrariesAndSDKs/AMF/commit/c7f29fec2326a253d319bb569cbc183b24504df9
NikosD
30th April 2017, 18:25
I'm following this link since the beginning:
https://github.com/GPUOpen-LibrariesAndSDKs/AMF/issues
In almost every new driver, AMD updates with fixes the AMF runtime API, but unfortunately this is not the case lately.
They have allocated a lot of resources (developers) to VEGA drivers and RyZen and so AMF is low priority nowadays for AMD.
JohnLai
30th April 2017, 18:39
I'm following this link since the beginning:
https://github.com/GPUOpen-LibrariesAndSDKs/AMF/issues
In almost every new driver, AMD updates with fixes the AMF runtime API, but unfortunately this is not the case lately.
They have allocated a lot of resources (developers) to VEGA drivers and RyZen and so AMF is low priority nowadays for AMD.
Same case with Nvidia.
Been waiting for SDK8.0 with these two interesting features:
•High-bit-depth (10/12-bit) decoding (VP9/HEVC)
•Weighted Prediction
Said to be enabled for 378.66 (driver released on 2017.2.14).
No news till now. The only official statement is " SDK 8.0 will be released shortly. "
Also reported the weird chrominance and luminance full range/limited range bug caused by nvdec + nvenc transcoding since last year, but still not yet fixed.
Meanwhile, Intel.....well.....totally no news on new features. Where is lookahead algorithm for QSV hevc? It would be great if Intel adds adaptive quantization.:(
nevcairiel
30th April 2017, 21:32
NVIDIA GTC is next week, maybe they'll use the chance to release the new SDK.
JohnLai
8th May 2017, 18:31
NVIDIA GTC is next week, maybe they'll use the chance to release the new SDK.
Well, you are right.
Video Codec SDK 8.0 was released just now......
Meanwhile....@NikosD,I leave it to you for informing rigaya about the update. :p
NikosD
8th May 2017, 18:33
Meanwhile....@NikosD,I leave it to you for informing rigaya about the update. :p
Don't worry!
He knows everything before us ;)
JohnLai
9th May 2017, 15:27
Huh...seem like spatial AQ for NVENC HEVC had been enabled at driver (378.66) level.
Playing around with rigaya nvencc --aq-strength 1 to 15 and driver aq default (using rate control = CQP I20:P23)
Driver AQ =minqp 12 maxqp 27 size 29247KB
AQ1 = 19 23 28315
AQ4 = 15 25 28193
AQ8 = 12 27 29247
AQ12 = 9 29 34994
AQ14 = 7 30 41234
AQ15 = 7 31 45051
EDIT 30-05-2017: Apparently, weighted prediction feature is using CUDA cores. I wonder why Nvidia makes this feature exclusive to Pascal only. If it is cuda cored based, it can be backported to maxwell series or even kepler.
hajj_3
25th September 2017, 09:47
https://i.imgur.com/RPDt9iA.png
check out the small print for the new 8th gen intel coffee lake cpu's announced today. Those big figures on the slide are when comparing to 4000 series intel chips not 7000 series chips. Ridiculously misleading. Would be nice to know if there are any encoding/decoding performance/quality improvements in these 8th gen chips but no info has been released so far to my knowledge.
nevcairiel
25th September 2017, 09:55
Please don't post enormous images like that, it blows up the entire forum layout.
hajj_3
25th September 2017, 10:02
Please don't post enormous images like that, it blows up the entire forum layout.
doom9 should upgrade the forum software so that it automatically scales images to the size of the post with the ability to enlarge.
WhatZit
26th September 2017, 00:15
Those big figures on the slide are when comparing to 4000 series intel chips not 7000 series chips. Ridiculously misleading.
4th Generation Intel is the median that most people own, and most of those "most people" are due to think about an upgrade right about now. For example, I average a 4-5 year upgrade cycle, so "now" is perfect timing for my 3rd Gen gear.
What Intel are desperate to do is market an Intel solution for the 4th Gen owners rather than have them explore an AMD solution.
Would be nice to know if there are any encoding/decoding performance/quality improvements in these 8th gen chips but no info has been released so far to my knowledge.
I'd LOVE to see a QSV MAIN10 vs MAIN HEVC quality shootout, but I can't find anything even remotely like it online. Guess I'll be seeing for myself once I upgrade in a couple of months.
NikosD
27th September 2017, 20:14
Would be nice to know if there are any encoding/decoding performance/quality improvements in these 8th gen chips but no info has been released so far to my knowledge.
According to anandtech.com and various sources, the iGPU of "8th" generation Core is the same as the iGPU of the 7th generation Core.
The only difference is clock speed.
ShogoXT
18th February 2018, 02:10
Hi everyone sorry for the bit of a old bump, but I feel this thread is still relevant today.
Live x264 encoding works right now, but I feel like because of how complex x265, vp9, and av1 will become, hardware encoders will surely take over vs expensive stream pcs in terms of "adequate" quality vs speed. I was pleasantly surprised to hear HEVC Nvenc was actually nearly equal to x264 on medium.
Now for my question. I been playing with stream settings on the program OBS, messing with custom x264 commands and such. Now I have been testing out nvenc more.
Does anyone know what 2 pass does exactly? Is it valuable for live encoding purposes? For Twitch it doesn't need to be super low latency (usually hits view at 15-20 seconds) setting which I assume those presets are meant for Nvidia Shield and such.
I always thought the asic encoders weren't that great at self analysis as they have more strict limitations, so most of the time you wanted to run it without that option, and the second encoder was meant for dual streaming.
The only relevant post I could find was this:
https://obsproject.com/forum/threads/nvenc-in-0-14-1.46986/page-3#post-211863
Thanks
Selur
18th February 2018, 21:01
Does anyone know what 2 pass does exactly?
It encodes each frame twice, was what it said in one of the NVEnc SDK pdfs iirc. :)
I doubt that it's useful for lice encoding, but 'usually hits view at 15-20 seconds' might be enough to use this.
IgorC
18th February 2018, 21:08
I was pleasantly surprised to hear HEVC Nvenc was actually nearly equal to x264 on medium.
Thanks
Intel H.265 encoder with speed@60fps is on par with x264 Placebo.
And that's actually old their H.265 encoder. Their new one should be even better.
http://www.compression.ru/video/codec_comparison/hevc_2016/
http://www.compression.ru/video/codec_comparison/hevc_2016/figures/graph3.png
nevcairiel
18th February 2018, 22:34
Note that the Intel Media Server Studio (MSS) HEVC Encoder is not what you get when you do "QuickSync" Encoding, its a separate and commercial product you have to buy (its one of the features missing from the free Community-edition of Intel Media Server Studio) - and its expensive.
benwaggoner
19th February 2018, 01:46
As codecs have gotten more complex, they've become less suitable for ASIC and GPU acceleration. So many coding options to choose between means latency between main and HW memory becomes a huge slowdown. And CPU's look a lot more like DSP these days with many cores and SIMD functionality like AVX2 and beyond.
Fixed-function encoders need to tape out well before they hit market, and psychovisual tuning of more complex standards takes a long time and makes a huge difference. So no fixed-function solution will be able to approach quality of the software encoders available by the time they are available in products.
Broadcast-grade encoders have increasingly moved to CPU and software defined solutions over the last decade.
I can see FPGA accelerated potentially being competitive, but we haven't seen any competitive real-world implementations yet.
IgorC
20th February 2018, 23:18
Note that the Intel Media Server Studio (MSS) HEVC Encoder is not what you get when you do "QuickSync" Encoding, its a separate and commercial product you have to buy (its one of the features missing from the free Community-edition of Intel Media Server Studio) - and its expensive.
Yes, it's most likely cheaper to buy 16-32 core CPU and use x265 at that point.
hajj_3
21st February 2018, 00:55
Note that the Intel Media Server Studio (MSS) HEVC Encoder is not what you get when you do "QuickSync" Encoding, its a separate and commercial product you have to buy (its one of the features missing from the free Community-edition of Intel Media Server Studio) - and its expensive.
Does their commercial product have better compression than their quicksync then? Are there are reviews comparing them or public specifications of the differences, would be nice to know.
easyfab
21st February 2018, 10:50
Last time I try, h264_qsv was better than hevc_qsv and 2x faster.
h264_qsv has more options available ( look-ahead, B-pyramid .... )
I see that with latest drivers b-pyramid and weighted-b frame is in. I will try to do some tests to see the improvements.
nevcairiel
21st February 2018, 10:57
Does their commercial product have better compression than their quicksync then? Are there are reviews comparing them or public specifications of the differences, would be nice to know.
Its entirely unrelated products, so yes, much better. Why would anyone buy it otherwise? :)
The Intel MSS Encoder is basically a software encoder with hardware acceleration, while QuickSync is a full hardware encoder. Full hardware is faster, but as outlined in various posts above, also quite limited in quality.
hajj_3
5th April 2018, 13:53
Nvidia 8.1 SDK is out, supports b-frames in h264 and other improvements: https://developer.nvidia.com/nvidia-video-codec-sdk
shades
17th April 2018, 00:35
As codecs have gotten more complex, they've become less suitable for ASIC and GPU acceleration. So many coding options to choose between means latency between main and HW memory becomes a huge slowdown. And CPU's look a lot more like DSP these days with many cores and SIMD functionality like AVX2 and beyond.
Fixed-function encoders need to tape out well before they hit market, and psychovisual tuning of more complex standards takes a long time and makes a huge difference. So no fixed-function solution will be able to approach quality of the software encoders available by the time they are available in products.
Broadcast-grade encoders have increasingly moved to CPU and software defined solutions over the last decade.
I can see FPGA accelerated potentially being competitive, but we haven't seen any competitive real-world implementations yet.
What about stuff like this?
http://www.advantech.com/products/pci-express-cards/sub_half-length_pcie_card
And there are cheap x264 cards on eBay
Has anyone had any luck (reasonable qualtiy output) with this sort of stuff?
foxyshadis
17th April 2018, 04:10
What about stuff like this?
http://www.advantech.com/products/pci-express-cards/sub_half-length_pcie_card
And there are cheap x264 cards on eBay
Has anyone had any luck (reasonable qualtiy output) with this sort of stuff?
Broadcast encoding is never particularly efficient, but they get around that by throwing tons of bitrate at the problem. When they don't, quality suffers enormously. Those cards are basically NvEnc on steroids, using simple searches and no RDO with extremely fast memory -- and a few bottlenecks turned into inflexible hardware paths -- to get better-than-CPU speed, rather than being hyper-efficiently designed around the nuances of the spec. RDO in particular is still a parallel-killer, since the best decision depends on the exact encoding of the last best decision.
When predictable speed is all that matters, hardware will always be key, but there's a low ceiling for the maximum quality you can get out of an FPGA/ASIC without reducing it to near-CPU speed.
Intel MSS is a different beast; it's essentially a good software encoder with lots of knobs to twiddle with a few hardware-accelerated paths, but it gets most of its speed gain by being paired with insanely fast eDRAM on Iris Pro models. Otherwise, it's just SW encoder speed to go with SW encoder quality. Nvidia or AMD could dedicate their GDDR5 to it to crunch a lot harder for the same fps, but they haven't seemed willing to so far; cheap and good enough is good enough.
foxyshadis
17th April 2018, 04:12
Note that the Intel Media Server Studio (MSS) HEVC Encoder is not what you get when you do "QuickSync" Encoding, its a separate and commercial product you have to buy (its one of the features missing from the free Community-edition of Intel Media Server Studio) - and its expensive.
That changed last week, Intel released their full HEVC GPU module in the Windows 2018 R1 community release. They haven't yet done it in the Linux version for some reason.
shades
18th April 2018, 02:27
That changed last week, Intel released their full HEVC GPU module in the Windows 2018 R1 community release. They haven't yet done it in the Linux version for some reason.
I just downloaded the Linux linked version from the Intel site if you're interested.
MediaServerStudioEssentials2018R1.tar.gz
shades
18th April 2018, 02:35
Broadcast encoding is never particularly efficient, but they get around that by throwing tons of bitrate at the problem. When they don't, quality suffers enormously. Those cards are basically NvEnc on steroids, using simple searches and no RDO with extremely fast memory -- and a few bottlenecks turned into inflexible hardware paths -- to get better-than-CPU speed, rather than being hyper-efficiently designed around the nuances of the spec. RDO in particular is still a parallel-killer, since the best decision depends on the exact encoding of the last best decision.
When predictable speed is all that matters, hardware will always be key, but there's a low ceiling for the maximum quality you can get out of an FPGA/ASIC without reducing it to near-CPU speed.
Intel MSS is a different beast; it's essentially a good software encoder with lots of knobs to twiddle with a few hardware-accelerated paths, but it gets most of its speed gain by being paired with insanely fast eDRAM on Iris Pro models. Otherwise, it's just SW encoder speed to go with SW encoder quality. Nvidia or AMD could dedicate their GDDR5 to it to crunch a lot harder for the same fps, but they haven't seemed willing to so far; cheap and good enough is good enough.
VEGA-3310
4K HEVC Broadcast Video Encoding/ Decoding / Transcoding Card
35W seems a lot less than nVidia card, still, I haven't tried one. Here's a blurb from the Datasheet.
"The technology behind VEGA-3310 can do the same task in under 35W, and VEGA-3310 can also support
up to 4Kp120 high frame rate for next generation sports broadcasts and 360 degree VR applications..
This card feature a simple-to-use API and example code for FFmpeg and GStreamer multimedia frameworks to streamline product development and integration into existing applications."
xabregas
19th April 2018, 13:45
The biggest improvements might be the new amazon hevc CRAP when comparing to their old H264.
Congrats to all the people that are working in HEVC. So many years to finnaly sell something thats worst in every way against previous stuff. Instead of improving...
Yeah Yeah You save bandwith at the cost of our eyes. Same quality my ass. Its a macroblock and banding feast in every official BS stream i put my eyes on.
Even some UHD blurays suffer....
shades
22nd April 2018, 23:47
The biggest improvements might be the new amazon hevc CRAP when comparing to their old H264.
Congrats to all the people that are working in HEVC. So many years to finnaly sell something thats worst in every way against previous stuff. Instead of improving...
Yeah Yeah You save bandwith at the cost of our eyes. Same quality my ass. Its a macroblock and banding feast in every official BS stream i put my eyes on.
Even some UHD blurays suffer....
You must be doing something different to me. I can see a clear video improvemnt at the same bitrate using handbrake and x265 over x264. Mind you, it's using software encoding, not hardware.
This is using a BluRay rip of media, which is already encoded. So, If I rip say 50 gig, then recode it down to 3 gig, x265 wins every single time hands down. Recoding to x264 at 3 gig size looks terrible in comparison.
TEB
23rd April 2018, 08:46
Anyone know if the Intel HW encoder or the Nvidia HW encoders support HEVC with Constant Quality for realtime encoding?
cakuhnen
24th April 2018, 23:33
Anyone know if the Intel HW encoder or the Nvidia HW encoders support HEVC with Constant Quality for realtime encoding?
Encoding with Intel Media SDK 2018 R1, encoding with Intel Media SDK GPU Accelerated plugin i get 19 fps for 1080p vídeo and the quality are great
RanmaCanada
27th April 2018, 16:25
Anyone know if the Intel HW encoder or the Nvidia HW encoders support HEVC with Constant Quality for realtime encoding?
That is all they support, as CRF is not in their capabilities, AFAIK.
JohnLai
30th April 2018, 05:58
K, everyone. Slightly unrelated....but here nvidia stuff for reading.
http://on-demand.gputechconf.com/gtc/2018/presentation/s8601-nvidia-gpu-video-technologies.pdf
ShogoXT
2nd May 2018, 05:49
I have been treating this thread as an all in one encoding hardware implementations thread anyway.
I read that PDF,a couple of things. It's nice that they still are expanding features, but most of it seems to be about cuda accellerated encoding in the later half. Wasn't opencl and cuda acceleration shown to be worse for quality? Also those YouTube pictures are funny is it trying to demonstrate blocking artifacts.
https://blog.parsecgaming.com/nvidia-nvenc-outperforms-amd-vce-on-h-264-encoding-latency-in-parsec-co-op-sessions-713b9e1e048a
More evidence that vce is very far behind. On Reddit I try to help people with encoding settings on OBS and most people don't believe me that nvenc is decent and vce is not so good compared to say x264 very fast.
https://en.wikipedia.org/wiki/Video_Core_Next
Has anyone seen this? Apparently it was put into the Raven Ridge APU.
JohnLai
2nd May 2018, 16:14
I have been treating this thread as an all in one encoding hardware implementations thread anyway.
I read that PDF,a couple of things. It's nice that they still are expanding features, but most of it seems to be about cuda accellerated encoding in the later half. Wasn't opencl and cuda acceleration shown to be worse for quality? Also those YouTube pictures are funny is it trying to demonstrate blocking artifacts.
https://blog.parsecgaming.com/nvidia-nvenc-outperforms-amd-vce-on-h-264-encoding-latency-in-parsec-co-op-sessions-713b9e1e048a
More evidence that vce is very far behind. On Reddit I try to help people with encoding settings on OBS and most people don't believe me that nvenc is decent and vce is not so good compared to say x264 very fast.
https://en.wikipedia.org/wiki/Video_Core_Next
Has anyone seen this? Apparently it was put into the Raven Ridge APU.
So far, nvidia adaptive GOP, adaptive IPB frame placement, lookahead and AQ (spatial and temporal) which technically are based on CUDA seem to work well. It depends on the implementation I guess.
Intel also has some sort of adaptive GOP, lookahead and IPB placements.
The greatest feature added by Nvidia and Intel is "constant quality" rate control mode (vbr-quality = nvidia , ICQ = intel)
No idea what is AMD doing with its VCE.
The youtube example is funny indeed, but the second picture sets with a picture of a man = Elecard HEVCAnalyzer? Before it was discontinued and acquired by adobe? Hmm....
ShogoXT
30th May 2018, 20:08
Looks like the trend of hardware accelerated encoding is not stopping. Adobe Premiere has added Quicksync encoding acceleration, not the asic chip, but using the IGPU for processing for speeding up encoding times, since it is otherwise completely unused on higher end rigs. If I still had a Intel system I would myself keep the IGPU enabled on Windows 10 since having 2 gpus enabled is no longer an issue, for more options.
This had fueled fanboy wars of course in the AMD vs Intel vs Nvidia world:
https://www.gamersnexus.net/guides/3310-adobe-premiere-benchmarks-rendering-8700k-gpu-vs-ryzen
I was going to start commenting about it in the other forums, but I thought x264 and cpu encoding in general got rid of OpenCL and CUDA acceleration because it is only usable for lookahead and usually makes the quality worse anyway.
Both Intel and Nvidia just in the last few months have put forth these features again. Did they fix the downsides? Does Adobe just say "they wont notice the quality difference anyway"?
Id really like to know if this is the future again or a temporary fad. Or is it mainly for newer encoders like VP9?
Also AMD needs to catch up, this was literally the whole reason they came up with the HSA platform idea, so what are they doing?
EDIT: Id like to bring up these discussions on reddit and am hoping to learn a few things here first.
EDIT2: https://forums.adobe.com/thread/2473774
Looks like it is for h264.
benwaggoner
5th June 2018, 17:41
It's nice that they still are expanding features, but most of it seems to be about cuda accellerated encoding in the later half. Wasn't opencl and cuda acceleration shown to be worse for quality?
There's nothing intrinsically worse about CUDA and OpenCL encoding. The problem has been the latency between GPU and CPU got in the way of lots of fast-feedback loops, which are increasingly important as the number of tools a codec can use expands, so more things need to be tried quickly in parallel with most options quickly terminated. CABAC is historically hard to multithread (although HEVC WPP makes it much more feasible; one CABAC thread per 64 pixel vertical row), so peak single-thread performance was/is a key limiting factor (CABAC can take up a good fraction of encode and decoding total MIPS).
GPU gives great SIMD support with hundreds of parallel cores. But a modern Intel CPU has much better per-core SIMD performance with the AVX family, and are getting ever more cores per CPU, all tied together with fast caches in unified memory
CUDA and OpenCL work great for more waterfall-like processes, where the GPU doesn't need to constantly report back to the CPU and get modified instructions.
FPGA looks like it has some good potential in video encoding.
The only grail for CPU/GPU is to have a fast CPU AND a fast GPU on the same die, sharing the same shared memory and shared caches. Intel has fast CPUs but not powerful OpenCL integrated GPUs; just using AVX-2/512 is generally faster for the compute part (although using GPU for decode, preprocessing, and lookahead can help).
It seems like AMD could do something amazing here, although they do multiple dies in the same package; the Zen microarchitecture doesn't have shared memory, or even have the HBM from prior AMD processors. Not a big issue for gaming, but it is for encoding and some other kinds of GPU compute tasks.
foxyshadis
5th June 2018, 21:38
It's kind of funny that the ideal video encoder would look something like the PS3's Cell; a beefy CPU, a bunch of lightweight SPEs directly connected for specialized workloads, and a GPU to farm out the most repetitive tasks to. Internally, a lot of CPUs do seem kind of like that these days, it's just not exposed to the programmer as explicitly and thus harder to take advantage of. The PS3 was just a little too ahead of its time.
NikosD
6th June 2018, 10:55
@benwaggoner
Great post, your last one.
ShogoXT
16th September 2018, 20:16
https://www.nvidia.com/content/dam/en-zz/Solutions/design-visualization/technologies/turing-architecture/NVIDIA-Turing-Architecture-Whitepaper.pdf
Bitrate efficiency and quality has gone up! This is a big deal see page 29.
ReinerSchweinlin
13th October 2018, 23:30
That changed last week, Intel released their full HEVC GPU module in the Windows 2018 R1 community release. They haven't yet done it in the Linux version for some reason.
So the community edition now includes the advanced HEVC-Encoder, not only the QSVENC SDK which is available in handbrake and others?
zub35
7th November 2018, 16:09
Test new NVENC-HEVC encoder (+ B frames) RTX2070
x264, x265, nvenc, qsv
test 1 https://rigaya34589.blog.fc2.com/blog-entry-1069.html
test 2 https://rigaya34589.blog.fc2.com/blog-entry-1070.html
New HEVC encoder - very good for game streams (twitch etc.)
Of course, this will not replace x264 (medium+) on a second computer, but will improve the quality of broadcasts, especially for low bitrates.
p.s. I not the author of the test's
ReinerSchweinlin
8th November 2018, 11:11
Thank you, very interesting.
hajj_3
5th December 2018, 14:13
Nvidia codec SDK 9.0 has just been released, it has some nice improvements for their new turing architecture: https://developer.nvidia.com/nvidia-video-codec-sdk
videoh
5th December 2018, 16:54
Your link says "coming soon" and only 8.2 is linked. Did I miss something?
Yups
13th December 2018, 21:04
Next year could be interesting.
Other improvements include a new HEVC Quick Sync Video engine that provides up to a 30% bitrate reduction over Gen9 (at the same or better visual quality)
For the media block, Intel says that the Gen11 design includes a ground up HEVC encoder design, with high quality encode and decode support.
https://www.tomshardware.com/reviews/intel-sunny-cove-gen11-xe-gpu-foveros,5932-3.html
https://www.anandtech.com/show/13699/intel-architecture-day-2018-core-future-hybrid-x86/3
ReinerSchweinlin
14th December 2018, 09:35
thank you for the interesting news :)
I assume this will require new hardware then? If Intel puts these engines in the lower end CPUs (Like G4560 Pentiums Kaby lake), building a cheap and fast encoding machine will be possible :)
Yups
14th December 2018, 16:27
Yes it requires a new hardware which includes the new encoder, this is a Gen11 presentation, means Icelake or Lakefield.
NikosD
12th May 2019, 12:10
@JohnLai
@CruNcher
Anyone else here with a Turing card ?
I used NVEncC v4.38 today just for a few H.265 encodings and I'm really impressed by the speed and quality.
Any particular settings for H.265 on Turing cards ?
NikosD
12th May 2019, 16:31
CUVID can work in two modes, either direct hardware access or through DXVA. In DXVA mode, CUVID can access the hybrid DXVA decoders, but that mode doesn't work on Windows 10, only the pure CUVID mode works on 10, which only has fixed-function decoders.
It's been more than 2 years (!) since this post.
Does CUVID have support in both modes on Win 10 nowadays ?
nevcairiel
12th May 2019, 17:09
No. But they also haven't made a hybrid decoder in a long time.
NikosD
12th May 2019, 17:39
Actually they don't need to.
nVidia accelerates everything in HW (MPEG1, MPEG2, MPEG4 ASP, H.264, H.265 (8/10/12 bit), VC-1, VP8, VP9 (8/10/12 bit)
We could expect a hybrid AV1 decoder sometime, I guess.
NikosD
14th May 2019, 19:38
So, I did a small benchmark regarding speed, not quality, using this source:
ftp://helpedia.com/pub/multimedia/x264/testvideos/2011%20-%2002%20-%20H.264%20CPU%20DXVA%20codec%20comparison%20-%20Core2Duo%20vs%20UVD%202.2/6.Cat-1080p60fpsRef4-25Mbps.m2ts
It's an H.264 1080p60fps file with an average of 25Mbps bitrate using two platforms:
1) Skylake Core i5 6500 under Win 10 x64 v17763 using 6576 drivers (API v1.28) and the HD 530 iGPU
Using QSVEncC v3.20 - GPU Clock/ Video Clock -> 1050 MHz
Encoding speeds:
H.265
ICQ/ Best -> 33,25 fps
ICQ/ Balanced -> 66,76 fps
ICQ/ Fastest -> 211,59 fps
H.264
ICQ/ Best -> 79,35 fps
ICQ/ Balanced -> 165,60 fps
ICQ/ Fastest -> 190,45 fps
2) nVidia GTX 1660 - Win 10 x64 v17763 - Drivers 430.39 (NVENC API v9.0 - CUDA 10.1)
Using NVEncC v4.38 - GPU Clock -> 1965 MHz - Video Clock -> 1815 MHz
Encoding speeds:
H.265
VBRHQ/ Quality -> 165,14 fps
VBRHQ/ Default -> 298,40 fps
VBRHQ/ Performance -> 441,89 fps
H.264
VBRHQ/ Quality -> 241,52 fps
VBRHQ/ Default -> 413,76 fps
VBRHQ/ Performance -> 667,61 fps
Using other encoding modes (VBR/ CBR/ CBRHQ/ CQP) the encoding speed didn't change.
As you can see the performance difference is huge, because Turing encoder is really fast.
I'm not sure if HD 630 is a lot faster than HD 530 and if Pascal encoder is even faster than Turing due to more encoding engines (2 vs 1)
Forteen88
14th May 2019, 21:17
@NikosD. Could you please also do a small test on metric quality (SSIM, and maybe VMAF on a proper video-source), if possible, on the Turing-GPU encoder vs x265@slower (with the same file-size)?
I saw that there are just a few Nvidia-cards that supports Turing HEVC-encoding with B-frames,
https://developer.nvidia.com/video-encode-decode-gpu-support-matrix
NikosD
15th May 2019, 07:41
@Forteen88
Sorry, but I don't have/use tools and apps regarding quality metrics and benchmarks.
I'm here mainly for speed benchmarks.
ReinerSchweinlin
23rd May 2019, 22:41
Encoding with Intel Media SDK 2018 R1, encoding with Intel Media SDK GPU Accelerated plugin i get 19 fps for 1080p vídeo and the quality are great
I just ran across this ... Do you have any links with a "how to" ? I´d like to try out the GPU Accelerated HEVC Encoder in the free edition.
RanmaCanada
27th May 2019, 04:37
So, I did a small benchmark regarding speed, not quality, using this source:
ftp://helpedia.com/pub/multimedia/x264/testvideos/2011%20-%2002%20-%20H.264%20CPU%20DXVA%20codec%20comparison%20-%20Core2Duo%20vs%20UVD%202.2/6.Cat-1080p60fpsRef4-25Mbps.m2ts
It's an H.264 1080p60fps file with an average of 25Mbps bitrate using two platforms:
1) Skylake Core i5 6500 under Win 10 x64 v17763 using 6576 drivers (API v1.28) and the HD 530 iGPU
Using QSVEncC v3.20 - GPU Clock/ Video Clock -> 1050 MHz
Encoding speeds:
H.265
ICQ/ Best -> 33,25 fps
ICQ/ Balanced -> 66,76 fps
ICQ/ Fastest -> 211,59 fps
H.264
ICQ/ Best -> 79,35 fps
ICQ/ Balanced -> 165,60 fps
ICQ/ Fastest -> 190,45 fps
2) nVidia GTX 1660 - Win 10 x64 v17763 - Drivers 430.39 (NVENC API v9.0 - CUDA 10.1)
Using NVEncC v4.38 - GPU Clock -> 1965 MHz - Video Clock -> 1815 MHz
Encoding speeds:
H.265
VBRHQ/ Quality -> 165,14 fps
VBRHQ/ Default -> 298,40 fps
VBRHQ/ Performance -> 441,89 fps
H.264
VBRHQ/ Quality -> 241,52 fps
VBRHQ/ Default -> 413,76 fps
VBRHQ/ Performance -> 667,61 fps
Using other encoding modes (VBR/ CBR/ CBRHQ/ CQP) the encoding speed didn't change.
As you can see the performance difference is huge, because Turing encoder is really fast.
I'm not sure if HD 630 is a lot faster than HD 530 and if Pascal encoder is even faster than Turing due to more encoding engines (2 vs 1)
Turing should actually be faster than Pascal and better quality as Pascal can only use 1 encoding engine at a time. What it allows for is you to have 2 encodes going, using 1 engine each (consumer and low end Quadro cards are limited to 2 simultaneous encodes at a time though there are fixes for that) vs 1 engine doing 2 encodes, like on the cards with only 1 engine (anything below the 1070Ti.)
tyee
4th August 2019, 01:54
Just got a 1660ti and have started encoding 4k today. I use avisynth+ (latest version) for decoding and NVencC64.exe for encoding. I'm getting about 150fps for 1080p and only about 25fps for 4k. For 4k, my cpu is only 50% occupied. Any way to speed up everything??
RanmaCanada
4th August 2019, 05:13
Just got a 1660ti and have started encoding 4k today. I use avisynth+ (latest version) for decoding and NVencC64.exe for encoding. I'm getting about 150fps for 1080p and only about 25fps for 4k. For 4k, my cpu is only 50% occupied. Any way to speed up everything??
Stop piping through avisynth as it's actually slowing down your encode. It can't feed the asic fast enough. Also, you're encodes are going to be garbage compared to just using your processor. Hardware encoding still hasn't caught up to software, at least with HEVC.
sneaker_ger
4th August 2019, 11:07
I got myself a GTX 1660 Ti. Just a short test in case anyone is interested. NVEncC64 -c hevc --output-depth 10 -b 5 --bref-mode middle --aq --aq-temporal --vbrhq 0 --vbr-quality 30 -u quality vs. x264 --preset veryslow --tune film 2pass at ~2200 kbps.
Input and output files:
https://mega.nz/#F!BlNDGIgJ!7lBUBs61l1oiIIBScrMyYg
(Source is some part of Netflix' "Meridian" encoded to 1080p lossless H.264.)
Screens (s=source, nv=1660 Ti, x=x264):
https://abload.de/img/s014jj13.png
https://abload.de/img/nv01w7j73.png
https://abload.de/img/x01s0kjv.png
https://abload.de/img/s02cxjdu.png
https://abload.de/img/nv02ctkup.png
https://abload.de/img/x025nj5e.png
https://abload.de/img/s03eqj3a.png
https://abload.de/img/nv0383jk4.png
https://abload.de/img/x03hhj8k.png
https://abload.de/img/s04dcj3m.png
https://abload.de/img/nv04pfjaq.png
https://abload.de/img/x046njag.png
https://abload.de/img/s05itkot.png
https://abload.de/img/nv05b9j42.png
https://abload.de/img/x05lnkfh.png
Sharc
4th August 2019, 14:38
I got myself a GTX 1660 Ti. Just a short test in case anyone is interested. NVEncC64 -c hevc --output-depth 10 -b 5 --bref-mode middle --aq --aq-temporal --vbrhq 0 --vbr-quality 30 -u quality vs. x264 --preset veryslow --tune film 2pass at ~2200 kbps.
x264 is still superior w.r.t. reproduction of details, but NV (NVEncC) has improved significantly. My choice except for "archiving" purpose.
NikosD
4th August 2019, 15:52
I got myself a GTX 1660 Ti.
Just a short test in case anyone is interested.
NVEncC64 -c hevc --output-depth 10 -b 5 --bref-mode middle --aq --aq-temporal --vbrhq 0 --vbr-quality 30 -u quality
vs.
x264 --preset veryslow --tune film 2pass at ~2200 kbps.
Nice...Did you keep notes regarding encoding speed (fps) and what is the CPU ?
sneaker_ger
4th August 2019, 16:15
CPU is an old i5-2500K. It cannot even decode the source sample in real-time. I didn't keep notes but just did a short test just for encoding speed (with sample recoded to H.264 and decoded by HW to focus on pure encoding speed). x264 CRF is in the ball-park of 7 fps. NVENC HEVC 1080p maxed out everything about 160 fps. With HEVC default settings about 330 fps. H.264 default setting between 500 and 600 fps, may need a longer sample to measure reliably. Not sure if decoding or encoding is the bottleneck.
Sharc
4th August 2019, 16:47
Not sure if decoding or encoding is the bottleneck.
.. or data transfer rate from/to storage device (HDD, SSD, USB …)?
sneaker_ger
4th August 2019, 17:14
For the speed test I encoded a lossy sample, H.264, 10 MB, SSD.
Sharc
25th September 2019, 18:21
Source is some part of Netflix' Meridian encoded to 1080p lossless H.264. I can upload later if anyone is actually interested but be warned cause it's 2.7 GB.
May I ask you to upload the source? I am interested in doing some AVC 8-bit tests with my Pascal GTX1050Ti. I think that Pascal/NVEncC have also improved a lot recently.
Thanks.
sneaker_ger
25th September 2019, 18:27
https://mega.nz/#F!BlNDGIgJ!7lBUBs61l1oiIIBScrMyYg (full source is available at xiph.org (https://media.xiph.org/video/derf/meridian/MERIDIAN_SHR_C_EN-XX_US-NR_51_LTRT_UHD_20160909_OV/))
Sharc
25th September 2019, 18:30
Thank you.
pacuro
31st October 2019, 09:44
Where is JohnLai? Did I miss something? Hope for his update with best hevc nvenc settings for Turing. Is it enough just to enable B-frames with his last best settings?
ReinerSchweinlin
17th December 2019, 11:09
I gave myself a christmas present and got a RTX 2060. First tests with NVENC in different GUIs give me good enough quality for my 1080p recodes when speed is important (around 140fps).
So far, I only used default settings. Like pacuro, I was wondering what settings I could use to max out quality (enabling b-frames gave me a bigger filesize than defaults in hybrid). Is there some detailied documentation of the NVENC in RTX Cards I could look into?
tyee
29th December 2019, 07:41
I'm using this command line app and it has very nice quality and I also get around 150fps for 1080p, about 50fps for 4k
https://github.com/rigaya/NVEnc
Is John still around the forums? I saw his settings somewhere, he was using Staxrip.
pacuro
14th March 2020, 10:11
@ReinerSchweinlin you probably doing something wrong if enabling B-frames makes your encodes bigger. In my case just enabling b-frames + b as ref each makes files lighter - as it should be in theory.
https://bpccdn.fra1.digitaloceanspaces.com/original/3X/9/7/972e44a699e4bd0cec8c883a2a98eea8f4b7c608.png
craigpro
18th April 2020, 14:30
These are the settings I'm using and my 720p encodes are always a few hundred MB larger than similar shows from the scene HEVC encodes. Does anyone have any suggestions for better settings please? Thank you.
--cqp 18:20:22 --codec h265 --preset quality --level 5.1 --output-depth 10 --qp-init 20 --qp-max 22 --qp-min 18 --max-bitrate 5000 --aq --aq-temporal --gop-len 240 --lookahead 16 --slices 2 --multiref-l0 4 --multiref-l1 4 --strict-gop --nonrefp --weightp
benwaggoner
20th April 2020, 02:06
These are the settings I'm using and my 720p encodes are always a few hundred MB larger than similar shows from the scene HEVC encodes. Does anyone have any suggestions for better settings please? Thank you.
--cqp 18:20:22 --codec h265 --preset quality --level 5.1 --output-depth 10 --qp-init 20 --qp-max 22 --qp-min 18 --max-bitrate 5000 --aq --aq-temporal --gop-len 240 --lookahead 16 --slices 2 --multiref-l0 4 --multiref-l1 4 --strict-gop --nonrefp --weightp
Having a QP range of just 18-22 seems like a really narrow band if your content is ever VBV constrained. 4 QP isn't enough for even a 2x difference in bitrate.
RanmaCanada
20th April 2020, 05:38
These are the settings I'm using and my 720p encodes are always a few hundred MB larger than similar shows from the scene HEVC encodes. Does anyone have any suggestions for better settings please? Thank you.
--cqp 18:20:22 --codec h265 --preset quality --level 5.1 --output-depth 10 --qp-init 20 --qp-max 22 --qp-min 18 --max-bitrate 5000 --aq --aq-temporal --gop-len 240 --lookahead 16 --slices 2 --multiref-l0 4 --multiref-l1 4 --strict-gop --nonrefp --weightp
With current tech you will not get close to what software encoding can do. If you want better results, use a CPU. I know it's not what you want to hear, but it's the truth. GPU is quick and dirty.
benwaggoner
22nd April 2020, 19:48
With current tech you will not get close to what software encoding can do. If you want better results, use a CPU. I know it's not what you want to hear, but it's the truth. GPU is quick and dirty.
Another way to frame that is that the quality @ perf for GPU is higher at very high speeds, but when encoding time/resources can be more, SW pulls ahead in quality as GPU encoding's peak quality is low.
GPU encoding is great for things like Twitch where it is recording screen activity while gaming. It takes minimal CPU away from the game, and is able to use the frames in the GPU as source without having to copy the pixels to main memory. Just writing 4Kp60 RGB frames to the CPU on top of everything else further stresse the memory bandwidth.
It's also good for very low power embedded solutions where a Tegra or Atom processor can be used with little CPU but with a HW encoder available.
There's been work done on hybrid encoders, which use the GPU for the first pass to be refined in software, which can speed up encoding 20-25%.
Yups
29th November 2020, 00:41
I'm currently testing Quicksync on Iris Xe LP. Compared to Gen9.5 they have greatly improved the fixed function encoder, both quality and speed. On Gen9.5 graphics the fully fixed function encoder was more or less useless (no b-pyramid, no bframes), means the hybrid was the better choice and for H265 Gen9.5 didn't even support FF. On Iris Xe LP I don't really see a quality difference between hybrid and the fully fixed function (H265) encode, the difference is minor I would say but the speed difference is really huge. So at this point it makes the h265 hybrid somehow obsolete.
RanmaCanada
29th November 2020, 05:32
I'm currently testing Quicksync on Iris Xe LP. Compared to Gen9.5 they have greatly improved the fixed function encoder, both quality and speed. On Gen9.5 graphics the fully fixed function encoder was more or less useless (no b-pyramid, no bframes), means the hybrid was the better choice and for H265 Gen9.5 didn't even support FF. On Iris Xe LP I don't really see a quality difference between hybrid and the fully fixed function (H265) encode, the difference is minor I would say but the speed difference is really huge. So at this point it makes the h265 hybrid somehow obsolete.
Are you using the Intel Media SDK? That is what is recommended to get the best encodes out of Quicksync. At least last time I checked.
Yups
29th November 2020, 11:51
Are you using the Intel Media SDK? That is what is recommended to get the best encodes out of Quicksync. At least last time I checked.
As for the hardware encoder it doesn't matter, it just needs the graphics driver, it has all the required files. The Media SDK is required for the software encoder.
Yups
29th November 2020, 22:46
I finished my Tigerlake speed test and also some other results just for reference. QSVEnc/NVEnc (h265) used for the GPUs and Handbrake nightly for the x265 run. Test sample is a 1080p 24fps 25 Mbit video which is converted down to ~2500 Kbit.
Iris Xe LP 1100 Mhz GPU+fixed function
CBR slowest= 68 fps
CBR balanced= 131 fps
CBR fastest= 226 fps
Iris Xe LP 1100 Mhz fixed function
CBR slowest= 260 FPS
CBR balanced= 325 FPS
CBR fastest= 507 FPS
Iris Xe LP 1100 Mhz GPU+ fixed function
CQP slowest= 75 fps
CQP balanced= 144 fps
CQP fastest= 263 fps
Iris Xe LP 1100 Mhz fixed function
CQP slowest= 271 fps
CQP balanced= 344 fps
CQP fastest= 518 fps
For reference some other GPUs and x265 ultrafast:
HD 630 1150 Mhz GPU + fixed function
CBR slowest= 43 fps
CBR balanced= 100 fps
CBR fastest= 302 fps
i7-1165G7 2800 Mhz (Turbo disabled)
x265 ultrafast= 58 fps
GTX 1080 2000 Mhz
CBR quality= 317 fps
CBR fastest= 502 fps
RanmaCanada
3rd December 2020, 02:49
But what was the quality comparison to NVENC? Speeds mean nothing if the output is garbage (Looking at you AMD VCE) :) If you have not done a comparison, would it be possible to do one?
Yups
3rd December 2020, 13:53
In this speed test NVENC quality is worse than Iris Xe but I will use different settings when I go for quality on both, for example I didn't use lookahead on NVENC and keep in mind this isn't Turing. I will do a quality comparison later and also I have to test Turing which I don't have for now. I'm testing on Iris Xe mainly at the moment. As for Iris Xe CQP is the best bitrate control method when it comes to quality, better than ICQ and much better than CBR/VBR. CQP works very good with many bframes+bpyramid on Iris Xe unlike ICQ.
RanmaCanada
3rd December 2020, 21:49
In this speed test NVENC quality is worse than Iris Xe but I will use different settings when I go for quality on both, for example I didn't use lookahead on NVENC and keep in mind this isn't Turing. I will do a quality comparison later and also I have to test Turing which I don't have for now. I'm testing on Iris Xe mainly at the moment. As for Iris Xe CQP is the best bitrate control method when it comes to quality, better than ICQ and much better than CBR/VBR. CQP works very good with many bframes+bpyramid on Iris Xe unlike ICQ.
Thanks!
benwaggoner
5th December 2020, 01:06
In this speed test NVENC quality is worse than Iris Xe but I will use different settings when I go for quality on both, for example I didn't use lookahead on NVENC and keep in mind this isn't Turing. I will do a quality comparison later and also I have to test Turing which I don't have for now. I'm testing on Iris Xe mainly at the moment. As for Iris Xe CQP is the best bitrate control method when it comes to quality, better than ICQ and much better than CBR/VBR. CQP works very good with many bframes+bpyramid on Iris Xe unlike ICQ.
The Turing NVENC added B-frames, IIRC, so is quite a lot better.
utack
5th December 2020, 13:57
I am honestly surprised how much praise nvenc gets "around the internet"
The H.264 encoder certainly isn't bad for a hardware encoder, but people comparing it to x264 medium or even slow need to tone it down a notch.
In a "non realtime" scenario with aq, slowest preset and full bframes it did not seem competitive to x264 medium preset in quality in some gaming footage I've tested
Asmodian
6th December 2020, 02:22
It is from comparing it to the previous hardware encoders, Turing was a big upgrade.
It is actually somewhat watchable now! Credit where it is due. ;)
aegisofrime
6th December 2020, 02:35
Does Ampere have any upgrades over Turing? I have a 3080 and 3060Ti coming, so I'm curious about this.
foxyshadis
6th December 2020, 16:39
Does Ampere have any upgrades over Turing? I have a 3080 and 3060Ti coming, so I'm curious about this.
Would you be willing to give the HEVC challenge a shot once they arrive? https://forum.doom9.org/showthread.php?t=175776 At least then we have a thing we can compare to. Staxrip can handle the encoding to nvenc quite well.
nevcairiel
6th December 2020, 16:46
Does Ampere have any upgrades over Turing? I have a 3080 and 3060Ti coming, so I'm curious about this.
No, the encoder is the same. Only the decoder was enhanced to support AV1.
benwaggoner
7th December 2020, 17:43
I am honestly surprised how much praise nvenc gets "around the internet"
The H.264 encoder certainly isn't bad for a hardware encoder, but people comparing it to x264 medium or even slow need to tone it down a notch.
In a "non realtime" scenario with aq, slowest preset and full bframes it did not seem competitive to x264 medium preset in quality in some gaming footage I've tested
Yeah, HW encoders haven't ever really approached the "medium" speed of the best SW encoder, and the gap gets bigger as codecs get more complex. CPUs with lots of cores, lots of cache, and powerful SIMD are really good at trying lots of different ways to encode each block, with early exits.
We've seen high-end broadcast encoders move from ASIC-based encoding to SW and SW-defined encoding over the last decade or so for this reason.
The sweet spot for HW encoders these days seems to be parallel encoding of CPU-intensive stuff with minimal perf hit. Streaming games is the key scenario; a SW encoder taking "only" 30% of CPU when gaming could have a big negative impact on fps. Plus encoding frames already in GPU memory doesn't add to memory throughput pressure much.
aegisofrime
10th December 2020, 09:02
Would you be willing to give the HEVC challenge a shot once they arrive? https://forum.doom9.org/showthread.php?t=175776 At least then we have a thing we can compare to. Staxrip can handle the encoding to nvenc quite well.
Sure, can do, I have the 3060Ti already just waiting to install it for Cyberpunk :D Not to doubt you nevcairiel, but it's worth a shot eh?
varnav
11th December 2020, 17:27
Some benchmarks done by me recently. I'm getting surprisingly good results from nvenc and even Intel compared to xlib265. Any comments?
System configuration
CPU: Intel Core i7 10750H
GPU 1: Intel UHD Comet Lake GT2
GPU 2: Mobile nVidia GeForce RTX 2060 (Turing microarchitecture)
Transcoding:
Windows 10 20H2
ffmpeg 2020-12-01-git-ba6e2a2d05-essentials_build-www.gyan.dev
HEVC encoder version 3.4+27-g5163c32d7
VMAF:
Ubuntu 20.04 WSL2
ffmpeg N-100180-gdbf8a16
VMAF 1.5.3
Source: http://ftp.nluug.nl/pub/graphics/blender/demo/movies/ToS/ToS-4k-1920.mov
File size: 704 MB
h264 (High) (avc1 / 0x31637661), yuv420p, 1920x800 [SAR 1:1 DAR 12:5], 7862 kb/s, 24 fps, 24 tbr, 24 tbn, 48 tbc (default)
VMAF command line:
./build.ffmpeg.sh
ffmpeg -i result.mkv -i ToS-4k-1920.mov.mkv -lavfi libvmaf="model_path=/usr/local/share/model/vmaf_v0.6.1.pkl" -report -f null -
Software
ffmpeg -xerror -i .\ToS-4k-1920.mov -map_metadata 0 -movflags use_metadata_tags -vcodec libx265 -crf 20 -preset slow -acodec copy ToS_cpu.mov
encoded 17620 frames in 1695.38s (10.39 fps), 4568.08 kb/s, Avg QP:23.91
File size: 416 MB
VMAF score: 98.601471
nVidia
ffmpeg -xerror -vsync 0 -i .\ToS-4k-1920.mov -rc-lookahead 25 -map_metadata 0 -movflags use_metadata_tags -preset p6 -spatial-aq 1 -temporal_aq 1 -cq 26 -vcodec hevc_nvenc -acodec copy ToS_nvenc.mov
fps=220 q=19.0 Lsize=416202kB time=00:12:14.12 bitrate=4644.3kbits/s speed=9.15x
File size: 406 MB
VMAF score: 97.881867
Intel
ffmpeg -xerror -hwaccel auto -i .\ToS-4k-1920.mov -vcodec hevc_qsv -b:v 6M -acodec copy ToS_intel.mov
frame=17620 fps=82 q=-0.0 Lsize= 440303kB time=00:12:14.09 bitrate=4913.5kbits/s speed=3.43x
File size: 429 MB
VMAF score: 97.592410
benwaggoner
12th December 2020, 01:03
VMAF >95 isn't that sensitive to subjective quality differences. I wouldn't predict actual real-world quality from scores between 97.6 and 98.6. Especially for clips more than 20 seconds or so, a 98.6 can look worse than a 97.6 (although all should look quite good). Reencoding from a source only 50% higher bitrate is also an odd case since the source artifacts could easily be more prominent than artifacts added in reencoding.
Testing with all around 3 Mbps would probably give some more interesting quality comparisons. Although it all depends on the scenario in the end. There are absolutely use cases where a 50% higher bitrate for 10x faster encoding is a great tradeoff.
varnav
12th December 2020, 01:51
I agree transcoding from already compressed video is not a clean thing, I will grab raw source someday and try this with it.
For subjective part, here is set of screenshots. I mixed them for the blind test. Who can tell what's what? Several guys could not.
https://drive.google.com/open?id=1ZCV2Xb9p15uqP4OTYMBnUWRGf1Ah_XEM
My scenario is recompression of home collection of short (but still 1080p 60 fps) videos, I am making a tool for this:
https://github.com/varnav/filmcompress
foxyshadis
12th December 2020, 02:36
I agree transcoding from already compressed video is not a clean thing, I will grab raw source someday and try this with it.
For subjective part, here is set of screenshots. I mixed them for the blind test. Who can tell what's what? Several guys could not.
https://drive.google.com/open?id=1ZCV2Xb9p15uqP4OTYMBnUWRGf1Ah_XEM
My scenario is recompression of home collection of short (but still 1080p 60 fps) videos, I am making a tool for this:
https://github.com/varnav/filmcompress
We've pretty much settled on the idea that comparing screenshots is pointless, and only clips is valid, despite how much more difficult that is.
That said, in each triplet of yours, one is noticeably softer, but whether that's innate to the coder or an artifact of the options chosen is impossible to tell. The level of softness is well below what I'd consider a threshold for noticing in motion, though. Overall, 4.5Mb/s is probably just not nearly enough to stress any of these chips anymore.
hajj_3
12th December 2020, 09:24
VMAF 2.0 was released a few days ago btw: https://github.com/Netflix/vmaf/releases/tag/v2.0.0
varnav
12th December 2020, 16:49
I can cut clips if needed. I can run more tests too if someone wants.
But IMO NVENC in later hardware revisions of NV chips produces very good results. Software is still better, but speed difference is 10-20 times.
Yups
13th December 2020, 00:36
On Intel both HD630 and Iris Xe the constant bitrate modes seem to work a lot better with a low bitrate budget. For now only a PSN-Y test using this (https://drive.google.com/file/d/1YX1V0SeSkYaq6Ui41vv1wcOatbLnuLSL/view?usp=sharing) sample, all of them HEVC main 8 bit 1920x1080 at ~2.5 Mbit 24 fps/240 gop. varna, you might use this sample.
Iris Xe:
CBR 2500 Hybrid quality= 37.97 dB (60 FPS, 3 bframes+pyramid)
CBR 2500 FF/low power quality= 37.60 dB (265 fps, 5 bframes+pyramid)
ICQ 25 Hybrid quality= 39.52 dB (66 fps, 5 bframes +pyramid)
ICQ 25 Hybrid speed= 39.34 dB (213 fps, 5 bframes +pyramid)
CQP 22_23_26 Hybrid quality= 40.32 dB (67 fps, 16 bframes+pyramid, offset 1_6_8)
CQP 23_23_26 FF/low power quality= 40.06 dB (272 fps, 16 brames+pyramid, offset 2_6_8)
Best on HD630 (Gen9.5):
CQP 23_24_28 Hybrid quality= 38.98 dB (42 fps, 16 bframes+pyramid, offset 2_6_8)
Iris Xe CQP vs HD630 CQP
https://s10.directupload.net/images/201212/cs73djrc.png
The difference is really big, Gen9 looks much worse in the video.
Yups
13th December 2020, 17:44
Same video on my GTX 1080 (8 bit HEVC, 2.5 Mbit, 240 gop....)
CBR quality= 36.93 dB
CBR HQ quality= 37.37 dB (Lookahead 32)
VBR HQ quality= 38.05 dB (Lookahead 32)
CQP quality= 38.84 dB (CQP 22_33_36, Lookahead 32)
I've also tried other settings but the PSNR-Y went down or didn't improve.
easyfab
20th February 2021, 13:31
I got my new nuc 11 with i7-1165G7 CPU.
I did some tests only with 1 file for the moment ( Kimono1_1920x1080_24.y4m ) but it give me a 1st idea.
SW encoding :
enc.time s. size ko vmaf psnr ssim
x264 fast 4.25 3954 93.76 39.94 12.70
x264 veryslow 31.7 3618 90.47 39.65 12.54
x265 superfast 10b 5.58 3698 95.49 40.63 13.01
x265 superfast 10b 5.74 4042 96.20 40.86 13.16
x265 medium 10b 11.24 3202 93.91 40.29 12.85
x265 medium 10b 12.26 4232 96.16 41.03 13.35
Svt-hevc preset 8 10b 5.97 3970 95.81 40.61 13.17
rav1e preset 6 q100 8b 84.24 4020 97.01 41.34 13.68
rav1e preset 6 q110 10b 238.35 3493 96.25 41.05 13.45
aom cpu 6 crf 32 10b 69.61 3733 97.70 41.61 13.76
svt-av1 speed 7 q34 10b 29.08 4012 97.30 41.31 13.55
svt-av1 speed 8 q35 10b 27.03 3782 96.94 41.14 13.43
other HW encoding :
enc.time s. size ko vmaf psnr ssim
Nvencc(1050 gtx) quality 10b 2.65 3721 92.72 39.85 12.49
qsvencc (m3-7y30) best 10b 9.00 3826 93.28 40.07 12.69
veencc (4800u) slow 10b 4.00 3794 93.22 39.71 12.38
and new i7-1165G7 tiger lake :
enc.time s. size ko vmaf psnr ssim
Qsvencc h264 best 8bit 1.00 3658 93.59 39.45 12.32
Qsvencc hevc best 8bit 4.00 3746 96.97 40.93 13.21
Qsvencc hevc best 8bit ff 1.00 3894 96.69 40.76 13.11
Qsvencc hevc best 10bit 4.00 3719 96.85 40.91 13.17
Qsvencc hevc best 10bit ff 1.00 3845 96.50 40.75 13.07
QSV on Tiger is a great jump in quality vs my older HW and can compare to SW encoding (but 10x faster/ less than 40w).
I don't have newer nvidia to compare.
Yups
20th February 2021, 22:08
Did you use VBR/CBR for i7-1165G7? FF/Hybrid quality difference is rather low with a big speed advantage in favour of FF.
easyfab
20th February 2021, 23:10
only cqp e.g. :
-c hevc --cqp 36:38:39 --profile main10 --fixed-func -b 6 --ref 6
Yups
20th February 2021, 23:25
Ok thanks, can you test it with --qp-offset 2:6:8 --b 16
You have to lower your cpq values as well with this offset to get a similar output size.
easyfab
21st February 2021, 09:18
without :
size 3845 psnr 40.75 ssim 13.07
with --qp-offset 2:6:8 -b 16:
size 3806 psnr 40.74 ssim 13.19
And qp-offset seem to be in on by default but don't know with what values.
for info :
cop3.DirectBiasAdjustment value changed off -> auto by driver
cop3.GlobalMotionBiasAdjustment value changed off -> auto by driver
QSVEncC (x64) 4.13 (r2009) by rigaya, Feb 17 2021 13:49:19 (VC 1928/Win/avx2)
OS Windows 10 x64 (19042) [UTF-8]
CPU Info 11th Gen Intel Core i7-1165G7 @ 2.80GHz [TB: 4.69GHz] (4C/8T) <Tigerlake>
GPU Info Intel Iris(R) Xe Graphics (96EU) 100-1300MHz [28W] (27.20.100.9313)
Media SDK QuickSyncVideo (hardware encoder) FF, 1st GPU, API v1.34
Async Depth 6 frames
Buffer Memory d3d9, 1 input buffer, 46 work buffer
Input Info avqsv: H.264/AVC, 1920x1080, 24/1 fps
VPP Enabled ColorFmtConvertion: nv12 -> p010
AVSync cfr
Output HEVC main10 @ Level 5 (high tier)
1920x1080p 1:1 24.000fps (24/1fps)
avwriter: hevc => matroska
Target usage 4 - balanced
Encode Mode Constant QP (CQP)
CQP Value I:32 P:35 B:36
QP Limit min: 22, max: 63
Trellis Auto
Ref frames 6 frames
Bframes 6 frames, B-pyramid: on
Max GOP Length 240 frames
Ext. Features WeightP WeightB QPOffset tskip ctu:64 sao:all
Yups
21st February 2021, 11:17
It's on by default but they are worse. SSIM jumped from 13.07 to 13.19 with a smaller file size, clearly better I would say. You did use b frames 6 by the way unless this log is old. From my test in #369 Tigerlake CQP with this settings reached PSNR/SSIM results which are comparable to x265 medium/slow, CQP on Tigerlake is superb. And because Iris Xe is so new we may see driver improvements in the future.
benwaggoner
22nd February 2021, 21:31
It's on by default but they are worse. SSIM jumped from 13.07 to 13.19 with a smaller file size, clearly better I would say. You did use b frames 6 by the way unless this log is old. From my test in #369 Tigerlake CQP with this settings reached PSNR/SSIM results which are comparable to x265 medium/slow, CQP on Tigerlake is superb. And because Iris Xe is so new we may see driver improvements in the future.
The subjective correlation of PSNR and SSIM aren't really good enough to confidently state that different encode parameters yielding a 0.12 dB difference is a material quality improvement. There are perceptual optimizations that can reduce PSNR and SSIM by >1 dB while improving subjective quality.
Yups
22nd February 2021, 23:20
I have tested these settings a lot. They are better subjective and objective. These settings allow lower CQP values over the default ones. Of course there is a limit, a too big offset may cause problems like more noise depending on the content. And in this case the SSIM/PSNR may increase while subjective there are some downsides.
Yups
27th February 2021, 18:42
Another Tigerlake CQP test with yvid_SAM_0235-UHD_Sample1 (HEVC main 8 bit 3840x2160 80 Mbit 30 fps)
PSNR-Y SSIM-Y speed bitrate
Tigerlake CQP FF quality 29.79 0.875 67 fps 3317 kbit
Tigerlake CQP quality 29.94 0.878 35 fps 3344 kbit
x265 medium 29.40 0.871 5.4 fps 3335 kbit
x265 slow 29.92 0.876 2.4 fps 3330 kbit
x264 slow 26.80 0.811 9.5 fps 3480 kbit
I'm using these Tigerlake CQP settings (Staxrip+QSVEnc):
--codec hevc --quality best --profile main --bframes 16 --gop-len 240 --b-pyramid --open-gop --qp-offset 2:6:8 --d3d9 --fixed-func --cqp 35:38:40
ukmark
27th February 2021, 23:18
I don't have Tiger Lake, but it's predecessor Ice Lake i7-1065G7 and am very impressed with CQP fixed-function encoding. I have posted my findings on StaxRip forum here:-
http://forum.doom9.net/showthread.php?p=1936660#post1936660
It appears the main improvement with Tiger Lake is the ability to use b-frames and b-pyramid with fixed-function encoding. With Ice Lake these are only available with GPU EU encoding (which is MUCH slower).
Suffice it to say that I no longer use x265. With blu ray (or reasonably high bit-rate) sources, I see no advantages for x265. I really find it very difficult to see any quality differences (that's after re-encoding about 50 of my blu ray collection using HW QuickSync encoding.)
I also posted a few screen grabs comparing x265 with QSV CQP (at 720p resolution) here (scroll down the page to see):-
https://github.com/rigaya/QSVEnc/issues/46
Yups
28th February 2021, 01:00
With Ice Lake these are only available with GPU EU encoding (which is MUCH slower).
It is but you can try non FF TU7/fastest preset where you can use 16 bframes+bpyramid, it's fast and the quality is better than TU1/best FF no bframes+bpyramid according to my tests. Same video from #379:
TU7 16 bframes+pyramid PSNR-Y: 29.45---3337 kbit---116 fps
TU1 FF no bframes/pyramid PSNR-Y: 28.95---3432 kbit---81 fps
RanmaCanada
28th February 2021, 02:01
Thanks for these tests, as it is extremely difficult to find quicksync encode information on forums. Everything is usually NVENC or x265.
ukmark
28th February 2021, 15:34
It is but you can try non FF TU7/fastest preset where you can use 16 bframes+bpyramid, it's fast and the quality is better than TU1/best FF no bframes+bpyramid according to my tests. Same video from #379:
TU7 16 bframes+pyramid PSNR-Y: 29.45---3337 kbit---116 fps
TU1 FF no bframes/pyramid PSNR-Y: 28.95---3432 kbit---81 fps
Thanks - I'll give that a go!
Nico8583
28th February 2021, 21:26
Very interesting, I'm a SW user but I would like to try Intel HW. Is it possible with an Intel 8500 (UHD 630) ? Is the quality will be the same than newer CPU, just the speed slower ? Thank you.
Tenkei
28th February 2021, 22:37
I'd love to see some tests at higher bitrates, like 10 Mbps+.
Yups
1st March 2021, 00:05
Very interesting, I'm a SW user but I would like to try Intel HW. Is it possible with an Intel 8500 (UHD 630) ? Is the quality will be the same than newer CPU, just the speed slower ? Thank you.
No it won't be the same unfortunately, HD630 is Gen9.5 based which is old gen. Intel upgraded their HEVC encoder with Icelake (Gen11): https://forum.doom9.org/showpost.php?p=1860006&postcount=317
With Tigerlake there is another upgrade to the HEVC FF encoder (bframes+bpyramid support).
Nico8583
1st March 2021, 00:34
Thank you, I'll try it in the future if I buy a Gen11 Intel CPU
Yups
1st March 2021, 00:37
Thank you, I'll try it in the future if I buy a Gen11 Intel CPU
If you buy new, you should better buy something with Gen12/Iris Xe graphics.
Nico8583
1st March 2021, 00:47
If you buy new, you should better buy something with Gen12/Iris Xe graphics.
Thank you, I don't plan to buy a new setup now but thanks for the information ;)
benwaggoner
1st March 2021, 19:59
If you buy new, you should better buy something with Gen12/Iris Xe graphics.
Are those available for anything other than laptops yet?
Yups
2nd March 2021, 00:35
Are those available for anything other than laptops yet?
Desktop CPU generation called Rocketlake-S is coming this month with Xe based iGPU, the desktop models have only 32 EUs instead of 96, even though the media capabilities should be the same. Also there are some NUC devices in the market based on Tigerlake-U.
hajj_3
2nd March 2021, 13:31
Are those available for anything other than laptops yet?
30th march: https://wccftech.com/intel-rocket-lake-11th-gen-desktop-cpus-officially-launching-on-30th-march/
Yups
4th March 2021, 10:16
It's a pity GPUs are so expensive these days, even a small TU116 Turing with newest NVENC costs 350-500€ which is ridiculous, it makes modern iGPUs much more valuable.
Nico8583
4th March 2021, 12:37
Desktop CPU generation called Rocketlake-S is coming this month with Xe based iGPU, the desktop models have only 32 EUs instead of 96, even though the media capabilities should be the same. Also there are some NUC devices in the market based on Tigerlake-U.
It will be Gen 11, not 12 :confused:
Intel H410 and B460 will not be compatible with Intel Gen 11 CPU, it's a shame.
Yups
4th March 2021, 13:14
It will be Gen 11, not 12 :confused:
Intel H410 and B460 will not be compatible with Intel Gen 11 CPU, it's a shame.
Rocketlake-S iGPU is based on Xe architecture:
Enhanced Intel UHD graphics featuring Intel Xe Graphics architecture.
https://newsroom.intel.com/wp-content/uploads/sites/11/2020/10/Intel-Rocket-Lake-S-Architecture.pdf
Yups
6th March 2021, 01:00
Another test. I'm using ToS_1920x800_xdither from this thread (https://forum.doom9.org/showthread.php?t=175776) and a low bitrate of ~1 Mbit with max 120 gop (Quicksync 121, Intel recommended 5*fps+1 for Quicksync some years ago)
VMAF speed bitrate
Quicksync H265 (27.20.100.9316)
HD630 CQP 86.78 52 fps 1023 kbit
Iris Xe CQP FF best 90.04 344 fps 1028 kbit
Iris Xe CQP best 90.55 97 fps 1014 kbit
NVENC H265 (461.81)
GTX1080 CQP best 85.86 215 fps 1023 kbit
x265 (Staxrip 2.1.8.4)
i7-1165G7 ultrafast 87.23 81 fps 1017 kbit
i7-1165G7 ultrafast CRF 84.80 83 fps 1023 Kbit
i7-1165G7 ultrafast QP 83.90 83 fps 1001 Kbit
i7-1165G7 medium 90.13 28.5 fps 1013 Kbit
i7-1165G7 slow 92.47 10.7 fps 1014 Kbit
I did us CQP 27_27_29 offset 2_3_7 bframes 16 for Iris Xe if someone is curious, same on HD 630 except for the CQP parameter.
Tenkei
6th March 2021, 14:26
Another test. I'm using ToS_1920x800_xdither from this thread (https://forum.doom9.org/showthread.php?t=175776) and a low bitrate of ~1 Mbit with max 120 gop (Quicksync 121, Intel recommended 5*fps+1 for Quicksync some years ago)
Can you test higher bitrates (around 10 or more)?
Yups
6th March 2021, 18:12
Yes I can but 10Mbit on this video is pointless, it's too much for making a difference. In general the differences will narrow down with more bitrate. I may try a 4K 50 Mbit sample converted to 20 Mbit later one.
Yups
6th March 2021, 23:41
I'm using this this sample (https://4kmedia.org/samsung-phantom-flex-uhd-4k-demo/) with a bitrate between 18.0-18.8 Mbit and 10 bit. All of them have a good quality with this high bitrate, a faster Quicksync preset makes sense in this case. I think I could even try fastest FF.
VMAF speed bitrate
Quicksync H265 (27.20.100.9316)
HD630 CQP 97.45 15 fps 18.5 Mbit
Iris Xe CQP FF best 98.33 67 fps 18.8 Mbit
Iris Xe CQP fastest 98.09 107 fps 18.2 Mbit
Iris Xe CQP best 98.46 29.5 fps 18.7 Mbit
NVENC H265 (465.51)
GTX1080 CQP best 96.43 87 fps 18.0 Mbit
x265 (Staxrip 2.1.8.4)
i7-1165G7 medium 97.70 4.8 fps 18.4 Mbit
i7-1165G7 slow 98.44 2.1 fps 18.5 Mbit
Tenkei
7th March 2021, 19:02
Iris Xe really holds up against software encoding, at least in metrics. Do you see any visual differences between Iris Xe best and x265 slow? NVENC tends to remove details to get better compression for the image as a whole. Could you upload the encoded clips?
Yups
7th March 2021, 22:13
I don't see any visual differences, even compared to the original the difference looks relatively small which I would expect from such high VMAF scores, the original looks a bit more detailed. I have to go much lower with the bitrate.
ReinerSchweinlin
7th March 2021, 22:16
Rocketlake-S iGPU is based on Xe architecture:
https://newsroom.intel.com/wp-content/uploads/sites/11/2020/10/Intel-Rocket-Lake-S-Architecture.pdf
Thanx for posting infos about the performance of the new iGPUs from Intel and the Encoder. Appreciate it, usually there is not much substantial on the net about these..
ReinerSchweinlin
18th March 2021, 16:01
I am trying to find solid information on whether the Encoding Engine of the new UHD Quicksync Encoders in the Tiger Lake i3 are the same as the XE-Variants of the ULV i5 and I7. My thoughts are these:
Given the power draw of CPUs and the possibility of "quite good XE hardware-encodings" it might be worth a shot to switch to GEN12 Intel iGPU for transcoding a lot of files... Speed is not that important, as long as it´s effiecency is good. I don´t need 300fps for 1080p, but if I get more than realtime for a few watts - then a sloar-powered transcoding rig might be possible.
I´d simply buy a NUC at the moment, but they seem hard to get...
@YUPS
Could you upload a small comparison encode of your XE Encoding with a very low bitrate? Let´s say a few seconds of a 720p cartoon at 200kbit/s, compared to x265 slow/anime/bframes8.. or similar? I think an extreme setting like that could reveal the differences the best.
RanmaCanada
19th March 2021, 03:57
I am trying to find solid information on whether the Encoding Engine of the new UHD Quicksync Encoders in the Tiger Lake i3 are the same as the XE-Variants of the ULV i5 and I7. My thoughts are these:
Given the power draw of CPUs and the possibility of "quite good XE hardware-encodings" it might be worth a shot to switch to GEN7 Intel iGPU for transcoding a lot of files... Speed is not that important, as long as it´s effiecency is good. I don´t need 300fps for 1080p, but if I get more than realtime for a few watts - then a sloar-powered transcoding rig might be possible.
I´d simply buy a NUC at the moment, but they seem hard to get...
@YUPS
Could you upload a small comparison encode of your XE Encoding with a very low bitrate? Let´s say a few seconds of a 720p cartoon at 200mbit/s, compared to x265 slow/anime/bframes8.. or similar? I think an extreme setting like that could reveal the differences the best.
I would say that the encode engine is not the same, as the engine is more than likely baked into the Xe Iris graphics.
Second, 200mbit/s is NOT a low bitrate. Maybe you mean 200kbit/s? And as for solar a solar powered unit, I would suggest just to get an older i3-8130u laptop. I am currently using one as my Plex/Emby server and it rarely goes above 15 watts in cpu usage. I do not know the full power usage, but it only has a 45 watt adapter.
If an older laptop is not your thang, then you could easily get a 10th gen i3 NUC for well under $400 USD (https://www.amazon.ca/dp/B083GH7KTX/?coliid=I1ZBAYVXNNNLMC&colid=18US8TN9FA5SO&psc=1&ref_=lv_ov_lig_dp_it). Or a 10th gen laptop for a few dollars more. (https://www.amazon.ca/dp/B08H4YTTLP/?coliid=I1X9GXI33BT7E2&colid=18US8TN9FA5SO&psc=1&ref_=lv_ov_lig_dp_it)
Yups
19th March 2021, 13:09
I am trying to find solid information on whether the Encoding Engine of the new UHD Quicksync Encoders in the Tiger Lake i3 are the same as the XE-Variants of the ULV i5 and I7. My thoughts are these:
Given the power draw of CPUs and the possibility of "quite good XE hardware-encodings" it might be worth a shot to switch to GEN7 Intel iGPU for transcoding a lot of files... Speed is not that important, as long as it´s effiecency is good. I don´t need 300fps for 1080p, but if I get more than realtime for a few watts - then a sloar-powered transcoding rig might be possible.
I´d simply buy a NUC at the moment, but they seem hard to get...
@YUPS
Could you upload a small comparison encode of your XE Encoding with a very low bitrate? Let´s say a few seconds of a 720p cartoon at 200mbit/s, compared to x265 slow/anime/bframes8.. or similar? I think an extreme setting like that could reveal the differences the best.
Yes I can on the weekend. Do you have a small cartoon sample? I would expect all Xe based iGPUs have the same quicksync encoder, the datasheet (https://d2pgu9s4sfmw1s.cloudfront.net/UAM/Prod/Done/a062E00001Zc09kQAB/bed7d2ba-23ee-44e3-a1bb-b874e7c5f746?Expires=1616155978&Key-Pair-Id=APKAJKRNIMMSNYXST6UA&Signature=cV5g091pMQei4oe2Ne4DzYlB6q80AB47n6Zr3l1JkheoTc2y0D0mM90uxhAm8eJtridkXsSnPmPp8s2IF1Cc4y7hv58JV~qL9lcTls9~EroVt7nf-VuSSG4dbK0UdfJLAHhuiRP9Y59Ma-z5Vbm4sNWyBI-t2Y-q9PlPN3FLn4kmfoUa65NPb4taK-Mt1q2jznr5zMi6wwZQRKZwAOqERaC5BAFqlGSkUXuACToOwhbZ3UDUWCO6jEkVIOdVbqAtzVQvB7zc6UeJXrJo0VDfolWPwi0KFH9zH-G4sBD1D3fY0gkO8IF6xN7caIYDI7WzvciLPbtC8UEVhco5Jj2wFg__) is relatively clear on this.
If it's branded UHD or Iris Xe graphics is not relevant when they both are Xe based and RKL-S won't get Xe brand either and you are talking about Gen7 which is Ivy Bridge based.
RanmaCanada
19th March 2021, 19:39
@YUPS the only creative commons or open source "modern cartoon" I know of is Sol Levante (https://opencontent.netflix.com/) from Netflix. There are of course the standards like Buck Bunny, Tears of Steel, but they are over a decade old at this point and do not represent the current quality of animation, be it CGI or hand drawn. The data sheet you've linked also appears to be broken!
benwaggoner
20th March 2021, 01:15
@YUPS the only creative commons or open source "modern cartoon" I know of is Sol Levante (https://opencontent.netflix.com/) from Netflix. There are of course the standards like Buck Bunny, Tears of Steel, but they are over a decade old at this point and do not represent the current quality of animation, be it CGI or hand drawn. The data sheet you've linked also appears to be broken!
I don't know that Sol Levante represents "modern animation" so much as "future animation." It's UHD HDR and a serious compression stress test with all the sharp circular lines moving around.
It's the hardest-to-encode publicly available source without grain I can think of. x265 gets serious artifacts with it at some sections, even in 1080p24 at 9 Mbps.
benwaggoner
20th March 2021, 01:20
Just looking at a test encode, with --crf 20.5 and --vbv-maxrate of 900, I get >700 frames with a QP >40. I set up a 25 Mbps peak placebo encode to run overnight to see what's possible here.
Yups
20th March 2021, 03:48
@YUPS the only creative commons or open source "modern cartoon" I know of is Sol Levante (https://opencontent.netflix.com/) from Netflix. There are of course the standards like Buck Bunny, Tears of Steel, but they are over a decade old at this point and do not represent the current quality of animation, be it CGI or hand drawn. The data sheet you've linked also appears to be broken!
Try this link: https://cdrdv2.intel.com/v1/dl/getContent/631121
If it doesn't work you have to go to Intel ark. I was trying the Sol Levante....on both Quicksync and NVENC there is no hardware decoding and therefore high CPU utilization when I encode it. It's a 12bit ProRes 4444 video, maybe that's why. Should I use CRF or bitrate mode for x265?
RanmaCanada
20th March 2021, 06:13
Try this link: https://cdrdv2.intel.com/v1/dl/getContent/631121
If it doesn't work you have to go to Intel ark. I was trying the Sol Levante....on both Quicksync and NVENC there is no hardware decoding and therefore high CPU utilization when I encode it. It's a 12bit ProRes 4444 video, maybe that's why. Should I use CRF or bitrate mode for x265?
From what I take away from the Intel Documentation is that only the Xe graphics have the new advanced features. As for what to try, the poster wanted 200kbit encode, so I would just do that, if possible.
Even looking at the processors, the core i3 (https://ark.intel.com/content/www/us/en/ark/products/208920/intel-core-i3-1115g4-processor-6m-cache-up-to-4-10-ghz-with-ipu.html) 11th gen only have UHD graphics, while the core i5 (https://ark.intel.com/content/www/us/en/ark/products/208659/intel-core-i5-1140g7-processor-8m-cache-up-to-4-20-ghz-with-ipu.html)and above have Iris Xe.
Yups
20th March 2021, 11:02
From what I take away from the Intel Documentation is that only the Xe graphics have the new advanced features. As for what to try, the poster wanted 200kbit encode, so I would just do that, if possible.
Even looking at the processors, the core i3 (https://ark.intel.com/content/www/us/en/ark/products/208920/intel-core-i3-1115g4-processor-6m-cache-up-to-4-10-ghz-with-ipu.html) 11th gen only have UHD graphics, while the core i5 (https://ark.intel.com/content/www/us/en/ark/products/208659/intel-core-i5-1140g7-processor-8m-cache-up-to-4-20-ghz-with-ipu.html)and above have Iris Xe.
All of the Tigerlake models are Xe architecture based, depending on the EU count they do get the Iris Xe branding or just UHD like for Rocket Lake-S. The datasheet says up to 96 EUs, there is no restriction for the lower EU count Xe, the FF/media unit should be the same.
i7 Tigerlake= 96 EUs Xe
i5 Tigerlake= 80 EUs Xe
i3 Tigerlake= 48 EUs Xe
Rocketlake-S= 32/24 EUs Xe
200 kbit on this UHD HDR 10bit video is not a feasible bitrate, the content isn't easy. I think 1500 kbit is a more realistic very low starting point. What VMAF scores are acceptable for you?
RanmaCanada
20th March 2021, 23:30
All of the Tigerlake models are Xe architecture based, depending on the EU count they do get the Iris Xe branding or just UHD like for Rocket Lake-S. The datasheet says up to 96 EUs, there is no restriction for the lower EU count Xe, the FF/media unit should be the same.
i7 Tigerlake= 96 EUs Xe
i5 Tigerlake= 80 EUs Xe
i3 Tigerlake= 48 EUs Xe
Rocketlake-S= 32/24 EUs Xe
200 kbit on this UHD HDR 10bit video is not a feasible bitrate, the content isn't easy. I think 1500 kbit is a more realistic very low starting point. What VMAF scores are acceptable for you?
interesting. And I dunno, we need to ask ReinerSchweinlin as they were the ones that wanted the tests in the first place!
Yups
20th March 2021, 23:52
@YUPS
Could you upload a small comparison encode of your XE Encoding with a very low bitrate? Let´s say a few seconds of a 720p cartoon at 200mbit/s, compared to x265 slow/anime/bframes8.. or similar? I think an extreme setting like that could reveal the differences the best.
I have tested Sol Levante on my GPUs and also did one x265 run. I've choosen x265 slow main10 10 bit and max 150 gop (for all). And software decoding on all GPUs, the decoding bottleneck on Iris Xe FF is quite big, it's usually a lot faster than the GPU+FF version.
Sol Levante VMAF speed bitrate
Quicksync H265 (27.20.100.9316)
HD630 CQP best 63.83 15 fps 1353 kbit
Iris Xe CQP FF best 71.09 22 fps 1338 kbit
Iris Xe CQP best 71.70 20 fps 1320 kbit
NVENC H265 (470.05)
GTX1080 CQP best 60.57 24 fps 1328 kbit
x265 (Staxrip 2.1.8.5)
i7-1165G7 slow 68.55 1.4 fps 1326 kbit
Link: https://drive.google.com/file/d/12oTKgEr2SaMt9Tnnq7IRl3jcVol1Pmzg/view?usp=sharing
VMAF below 80 is not that great, I would choose a higher bitrate.
ReinerSchweinlin
21st March 2021, 12:47
I would say that the encode engine is not the same, as the engine is more than likely baked into the Xe Iris graphics.
Second, 200mbit/s is NOT a low bitrate. Maybe you mean 200kbit/s? [/URL]
Thanx for catching the type, of course kbit/s :)
@YUPS Thanks for taking the effort to compare and make a testsample. The results are very interesting, indeed. Too bad I sold all my RTX Cards at the moment - a GEN20 NVENC Encode to compare would be interesting.
The reason why I mentioned a low bitrate cartoon was that this is my usual workflow for determining the limits and capabilities of an encoder setting... Of course higher bitrates are always better for the result, but to see what an encoder does if he is really limited gives me a better point for judgement. I am fully aware that this is biased somehow and that someone not caring about efficiency or drive space or bandwith might have other priorities.
The used Example really is a tough one :) With cartoon, I meant something like Family Guy (line art style) - which is a little easier to judge (for me at least). Modern Anime rarely has anything from a classic cartoon. But the example still shows that Iris XE seems quite capable..
I will look for a suitable, free sample :)
hm, the question remains if the smaller models of the tiger lake series with the UHD called iGPUs have the exact same encoders. The argument of course is valid, that its derived from XE Grafix, but I can´t find 100% evidence if that´s the case. IMHO the datasheet leaves room for interpreation (unless I miss something, maybe you can the point me to the page). I remember cases like the 1650, which officially is TURING Generation, but the one thing that wasn`t was the NVENC Engine, which is VOLTA, so no B-Frame support..
If the smaller and cheaper Tiger Lake had the same Encoding Engine, that could really make up for a very nice low power encoding/Streaming/transcoding rig... If it´s a little less powerful than the Variants with the higher EU Count - wouldn´t harm my usecase :)
@RanmaCanada
Thanx for the suggestions about older laptops.. I have a bunch of older machines with low power chips, NUCs, SOCs, Laptops, etc... and these are the current candidates I use for solar powered usecases... So far, the Quality of the hardware-encoding engines wasn´t en par with x265 for a given bitrate, but with the XE Engines it seems to become competetive (while still being not power hungry)... The GEN20 NVENC isn´t too bad either... but I know of now system with lets say a 1660 which consumes as less as an XE bases system promisses to be capable off..
ReinerSchweinlin
21st March 2021, 13:08
This free video file of a blender Project would be interesting to test with a very low bitrate. I am fully aware that the results won`t be what one wants as end result - it´s more of a test how the encoder is dealing with very limited conditions.
https://upload.wikimedia.org/wikipedia/commons/a/a9/HERO_-_Blender_Open_Movie-full_movie.webm
Yes, i´ts already compressed, but thats fine for me and represents a usecase for many of us anyway :) This video has a nice mix of "cartoon style" still scenes with a little gradient in the back, defined, sharp lines in the front, scenes with a little more action, some scrolling/paning, some textures, etc... The resolution is fine, if it´s no trouble for you, I´d love to see it in 720p. (There is a reason for that: Modern encoders seem to be more optimized for higher resolutions and also the usecase calls for smaller bandwiths, which results in smaller resolutions... I found some interesting effects in these scenarios: While the same footage in 4K looks very good at edges in all encoders used - scaling it down to 720p and then viewing it on the same screen again (upscaled while playing) reveals artefacts at sharp edges - which now are much more visible than before - even if the footage wasn´t "visually super sharp" and from the viewing distance looks almost identical in terms of "optical resolution"... To deal with this effect on smaller resolutions, we all "know" that lowering the Q-Factor in constant-Q Encodings is a good idea..
So I came up with this tesprocedure of mine - encoding in a low resolution with a low bitrate reveals "a lot I need to know about an encoder to judge it".. Then i trust my eyes :)
Thanx for helping me out!
And thanx @all for commenting, of course.
ReinerSchweinlin
21st March 2021, 13:20
While looking at you encodes, I noticed something:
The hardware-encoders all were able to detect the end-credits as simple paning from bottom->up and therefore use very little bitrate (as it should be), while the x265 encode uses a LOT more bitrate for this section.. Ill try to replicate that and dig a little into the search patterns. I am sure that this makes an impact of the overall score, because if almost 1/4 of the clip gets a magnitude more bitrate than needed, the demanding rest of the clip is starving...
BTW: Back in "the days", there was "bitrate viewer 1.4" for MPEG2 files... Any successor around?
just downlaoding "SOL LEVANTE" - the prores version is so big and the download speed so slow (not maxing out my Internet at all) - which version did you use? (Or did you wait 16 hours?).
butterw2
21st March 2021, 13:48
Intel 11th Gen Desktop (Rocket Lake S, 14nm), only the i5 and up have Xe graphics (UHD-750 and UHD-730 for the i5-11400).
They have a fraction of the igpu Execution Units of the 10nm mobile Tiger Lake parts, but it is assumed that the media encoder/decoder block is the same (AV1 hw decoder, hevc encoder).
! The i3 and pentium are just rebranded current gen.
The 6core/12threads i5-11500 and i5-11400 are stated to have 65W TDP, and the T parts 35W TDP. They are sub-200$ msrp parts which can be used with the new B560 motherboards (PCIe-4 and DDR4 3200MHz).
ReinerSchweinlin
21st March 2021, 13:59
thanks for adding.
Intel 11th Gen Desktop (Rocket Lake S, 14nm), ...... but it is assumed that the media encoder/decoder block is the same
And thats exactly where clarification would be welcome :)
butterw2
21st March 2021, 14:19
The feature set is confirmed as being the same, whether the result is 100% the same (after the 14nm backport) we will likely only learn when the chips become widely available.
The main takeaway for me was that the new decoder/encoder will not be available for the cheapest chips/motherboards right now.
UHD-750/730 also has HDMI 2.0b support.
https://en.wikipedia.org/wiki/Rocket_Lake#GPU
https://en.wikipedia.org/wiki/Intel_Graphics_Technology#Integrated
Yups
21st March 2021, 17:24
hm, the question remains if the smaller models of the tiger lake series with the UHD called iGPUs have the exact same encoders. The argument of course is valid, that its derived from XE Grafix, but I can´t find 100% evidence if that´s the case. IMHO the datasheet leaves room for interpreation (unless I miss something, maybe you can the point me to the page). I remember cases like the 1650, which officially is TURING Generation, but the one thing that wasn`t was the NVENC Engine, which is VOLTA, so no B-Frame support..
RKL-S datasheet volume 1 isn't available yet but not sure if you will find any more concrete in it. The thing is Intels media engine version is tied to the graphics architecture up to now, there is no differentiation like on Nvidia. That's why you most likely won't find something more concrete in the RKL-S datasheet because there is no need for it. The feature list (https://github.com/intel/media-driver/blob/master/docs/media_features.md#media-features-summary) from Intels open source driver is identical on Tigerlake and Rocketlake.
And by the way this might change in the future starting with Alder Lake, the various IP blocks can be upgraded at a different cadence regardless of the graphics architecture, the Xe LP graphics in future generations may get a media or display upgrade.
The arrival of this new display architecture coincides with a general disaggregation of Intel GPUs' architecture version numbering for the different component IP blocks. Going forward it isn't accurate to talk about a platform using INTEL_GEN() anymore since the various IP blocks (graphics, media, display) are moving to independent internal numbering schemes that may have different granularity and move at different cadences; the hardware teams have asked us to start tracking these values separately for "graphics," "media," and "display" such that anywhere that we need to do a numerical comparison on the architecture version, we should need to use an IP-specific version number instead of INTEL_GEN().
https://patchwork.freedesktop.org/series/87886/
Intel 11th Gen Desktop (Rocket Lake S, 14nm), only the i5 and up have Xe graphics (UHD-750 and UHD-730 for the i5-11400).
They have a fraction of the igpu Execution Units of the 10nm mobile Tiger Lake parts, but it is assumed that the media encoder/decoder block is the same (AV1 hw decoder, hevc encoder).
! The i3 and pentium are just rebranded current gen.
The 6core/12threads i5-11500 and i5-11400 are stated to have 65W TDP, and the T parts 35W TDP. They are sub-200$ msrp parts which can be used with the new B560 motherboards (PCIe-4 and DDR4 3200MHz).
The rebranded i3 are Cometlake and 10th Gen, they are not sold as RKL-S 11th Gen.
Yups
21st March 2021, 17:36
While looking at you encodes, I noticed something:
The hardware-encoders all were able to detect the end-credits as simple paning from bottom->up and therefore use very little bitrate (as it should be), while the x265 encode uses a LOT more bitrate for this section.. Ill try to replicate that and dig a little into the search patterns. I am sure that this makes an impact of the overall score, because if almost 1/4 of the clip gets a magnitude more bitrate than needed, the demanding rest of the clip is starving...
BTW: Back in "the days", there was "bitrate viewer 1.4" for MPEG2 files... Any successor around?
just downlaoding "SOL LEVANTE" - the prores version is so big and the download speed so slow (not maxing out my Internet at all) - which version did you use? (Or did you wait 16 hours?).
I haven't tried CRF on x265 for this video because it's so slow, the bitrate allocation might differ there. CQP is usually the best choice on my GPUs. Sol Levante is 37.8GB big, it was fast when I downloaded.
Yups
21st March 2021, 21:33
This free video file of a blender Project would be interesting to test with a very low bitrate. I am fully aware that the results won`t be what one wants as end result - it´s more of a test how the encoder is dealing with very limited conditions.
https://upload.wikimedia.org/wikipedia/commons/a/a9/HERO_-_Blender_Open_Movie-full_movie.webm
Yes, i´ts already compressed, but thats fine for me and represents a usecase for many of us anyway :) This video has a nice mix of "cartoon style" still scenes with a little gradient in the back, defined, sharp lines in the front, scenes with a little more action, some scrolling/paning, some textures, etc... The resolution is fine, if it´s no trouble for you, I´d love to see it in 720p.
There is a also a 720p (https://upload.wikimedia.org/wikipedia/commons/transcoded/a/a9/HERO_-_Blender_Open_Movie-full_movie.webm/HERO_-_Blender_Open_Movie-full_movie.webm.720p.vp9.webm) version on this site. The CRF mode is a lot better than the bitrate mode for this type of content.
HERO - Blender Open Movie VMAF speed bitrate
Quicksync H265 (27.20.100.9316)
Iris Xe CQP FF best 86.56 550 fps 199 kbit
Iris Xe CQP best 88.59 167 fps 200 kbit
x265 (Staxrip 2.1.9.0)
i7-1165G7 slow 84.61 26 fps 200 kbit
i7-1165G7 CRF slow 88.18 34 fps 200 kbit
Link: https://drive.google.com/file/d/1zNPEmhvDEvOZBBX9hFnDTSROZgHiLMG4/view?usp=sharing
Yups
23rd March 2021, 16:54
There is a also a 720p (https://upload.wikimedia.org/wikipedia/commons/transcoded/a/a9/HERO_-_Blender_Open_Movie-full_movie.webm/HERO_-_Blender_Open_Movie-full_movie.webm.720p.vp9.webm) version on this site.
Some more reference points for this. CRF mode only for x264/265, it's much better than bitrate mode for this test case.
HERO - Blender Open Movie VMAF speed bitrate
Quicksync H265 (27.20.100.9316)
Iris Xe CQP FF best 86.56 550 fps 199 kbit
Iris Xe CQP FF balanced 85.66 850 fps 201 Kbit
Iris Xe CQP FF speed 76.25 1550 fps 200 Kbit
Iris Xe CQP best 88.59 167 fps 200 kbit
Iris Xe CQP balanced 87.18 280 fps 201 Kbit
Iris Xe CQP speed 86.22 520 fps 200 Kbit
x265 (Staxrip 2.1.9.0)
i7-1165G7 CRF slower 90.25 8 fps 200 Kbit
i7-1165G7 CRF slow 88.18 34 fps 200 kbit
i7-1165G7 CRF medium 85.72 56 fps 200 Kbit
i7-1165G7 CRF very fast 83.36 73 fps 200 Kbit
x264 (Staxrip 2.1.9.0)
i7-1165G7 CRF slower 75.37 81 fps 200 Kbit
Fixed function speed preset does not support bframes on Iris Xe with current driver, 16 bframes for the other presets. Offset mostly 2:6:8, on some a bit lower depending on the output bitrate.
benwaggoner
23rd March 2021, 19:39
I haven't tried CRF on x265 for this video because it's so slow, the bitrate allocation might differ there. CQP is usually the best choice on my GPUs. Sol Levante is 37.8GB big, it was fast when I downloaded.
Sol Levante is my new favorite encoder stress test, as it has a lot of complex content of a type that encoders aren't typically optimized for. It's definitely the most challenging to encode source without film grain I've played with.
I'd expect HW encoders to lie down and cry (aka have a lot of visible artifacts) without using High Tier or raising level above the minimum required for frame size/fps. x265 placebo falls totally apart doing a 1080p24 Level 4.0 encode.
ReinerSchweinlin
24th March 2021, 14:15
The main takeaway for me was that the new decoder/encoder will not be available for the cheapest chips/motherboards right now.
Thanx, thats too bad. Oh well, an i5 then probably would be the way to go.
Fixed function speed preset does not support bframes on Iris Xe with current driver, 16 bframes for the other presets. Offset mostly 2:6:8, on some a bit lower depending on the output bitrate.
Thank you for testing "hero". Very interesting. Also thanx for the info about b-frames. Given the speed gain, this really looks like a usable alternative to software encoding for my planed "low power encoding rig". I expected Software Encoding to be more efficient in terms of quality/bitrate, but given the speeds and power consumption - the new XE Encoders seem "good enough". Expecting your testencodes visually, they really look good!
Yups
24th March 2021, 20:41
Iris Xe CQP is really good (for a hardware encoder), CBR/VBR are not that good due to some missing features (Lookahead). I wonder if Turing/Ampere can match or beat Iris Xe with its constant rate mode, is there a x265 comparison somewhere, how is it in comparison? Bframes support for fixed function speed preset might come in a later driver because their Linux Media SDK (https://github.com/Intel-Media-SDK/MediaSDK/blob/a636bfae6373b4e40cd800da4b5f3e8d8c97dba4/doc/mediasdk_release_notes.md) enabled it:
HEVC encode
Extended B frames support across all target usage with LowPower on
(LowPower= fixed function)
Tenkei
24th March 2021, 21:28
Iris Xe CQP is really good (for a hardware encoder), CBR/VBR are not that good due to some missing features (Lookahead). I wonder if Turing/Ampere can match or beat Iris Xe with its constant rate mode, is there a x265 comparison somewhere, how is it in comparison? Bframes support for fixed function speed preset might come in a later driver because their Linux Media SDK (https://github.com/Intel-Media-SDK/MediaSDK/blob/a636bfae6373b4e40cd800da4b5f3e8d8c97dba4/doc/mediasdk_release_notes.md) enabled it:
HEVC encode
Extended B frames support across all target usage with LowPower on
(LowPower= fixed function)
Wouldn't 2 pass negate the lack of lookahead?
benwaggoner
25th March 2021, 00:15
While looking at you encodes, I noticed something:
The hardware-encoders all were able to detect the end-credits as simple paning from bottom->up and therefore use very little bitrate (as it should be), while the x265 encode uses a LOT more bitrate for this section.. Ill try to replicate that and dig a little into the search patterns. I am sure that this makes an impact of the overall score, because if almost 1/4 of the clip gets a magnitude more bitrate than needed, the demanding rest of the clip is starving....
Yes, x265 with default settings spends an inexplicable amount of bits on title cards and scrolling credits. My guess is a defect in the Rate Factor implementation which is way overestimating how low QPs need to be for that kind of content. It's nigh-impossible to tell the difference between credits at CRF 20 and CRF 40.
It's possible it is a better match for x264. With 4x4 CUs, --amp, --rect, and --tskip, x265 can encode details down to a single pixel width, and SAO is quite good at suppressing ringing artifacts for text and line art. Intra-frame prediction with 1/8th pel is also excellent for big pages of text in the same font; basically each repeated letter gets a near-perfect prediction without residual.
benwaggoner
25th March 2021, 00:16
Wouldn't 2 pass negate the lack of lookahead?
Exactly.
Historically, it's more that lookhead partially negated the need for 2 passes, of course ;)
ReinerSchweinlin
25th March 2021, 04:27
Iris Xe CQP is really good (for a hardware encoder), CBR/VBR are not that good due to some missing features (Lookahead). I wonder if Turing/Ampere can match or beat Iris Xe with its constant rate mode, is there a x265 comparison somewhere, how is it in comparison? Bframes support for fixed function speed preset might come in a later driver because their Linux Media SDK (https://github.com/Intel-Media-SDK/MediaSDK/blob/a636bfae6373b4e40cd800da4b5f3e8d8c97dba4/doc/mediasdk_release_notes.md) enabled it:
HEVC encode
Extended B frames support across all target usage with LowPower on
(LowPower= fixed function)
It´s really too bad I don´t have my RTX Cards any more (Sold them to get bigger ones - and then... well not buying one right now ....the remaining AMD Cards will do for the moment). The Turing Encoder really isn´t that bad either, but as far as I remember it wasn´t as good as what XE can do. In lack of a direct comparison, this is only "my feeling"...
2-pass with nvencc: As far as I rmemeber its not really 2-pass in the traiditonal sense of doing the whole file twice / the encoder takes a GOP or other small number of frames and runs them twice internally. The hybrid encoder from mainconcept offloads rate control to the CPU while using nvenc of Turing - but that resultet in much lower speeds (pure nvencc in my tests above 150fps - same input files with the hybrid encoder: around 30fps).
Maybe someone with a Turing Encoder wants to jump in and encode the examples from above?
I could test my AMD Encoder - but without b-frames we can predict the outcome :)
Thanx for confirming. I couldn´t find a "sane" setting for normal content which also works "as intented" on credit-szenes with x265.
rwill
25th March 2021, 22:13
Wouldn't 2 pass negate the lack of lookahead?
It depends on how its implemented.
Lookahead shows the encoder how things probably will develop
short term wise while 2-pass will show an encoder the average
rate @ quantizer ( rate@CRF ), among other things, over the
whole sequence plus the dirty per frame details.
Lookahead is needed for short term decisions like estimating
the video buffer development over the next N frames. This can
be done up to the end of sequence in the second pass but also
over the next N frames with buffering delay in the first pass.
If it is implemented good 2-pass can 'almost' negate the lack
of lookahead in a first pass. It depends on how well the
"rate@quantizer" predictors for short term decisions are done.
If one predicts short term frame sizes from rate@quantizer
from the previous pass or rate@distortion_cost from the current
pass or a mixture of both and how good the overall system is...
rate@quantizer correlation might go bad if quantizer differs too
much between passes and rate@distortion_cost might be bad due
to bad rate correlation with distortion_cost.
It depends.
2 pass will always beat 1 pass though, just imagine a sequence
with 5000 frames of "bitrate breaking action" followed by 5000
frames of "solid black". Then imagine what happens if you swap
the display order of the two 5k parts around.
Yups
25th March 2021, 23:36
Yes, x265 with default settings spends an inexplicable amount of bits on title cards and scrolling credits. My guess is a defect in the Rate Factor implementation which is way overestimating how low QPs need to be for that kind of content. It's nigh-impossible to tell the difference between credits at CRF 20 and CRF 40.
x265 Bitrate mode is not good for this sample, possibly because of the scrolling credits which is quite long. I will add CRF scores later this week, CRF scores are a lot higher.
excellentswordfight
26th March 2021, 09:01
Playing a bit bit with quicksync on Xe, when i try to use ICQ mode with lookahead on an i7-1165G7 with QSVEncC i get the following message:
"LA-ICQ (Intelligent Const. Quality with Lookahead) mode is not supported on current platform"
Anyone know why? I thought that Xe should have support for most (all?) features?
I also noticed that VBV isnt supported in ICQ mode, is that a feature missing on the intel side or on the encoder side? Is it on any roadmap to support it? I'm not a fan of doing encodes that are not vbv compliant with selected level, cause with lookahead I would assume that it should be possible.
Yups
26th March 2021, 16:02
Are you trying Lookahead with H265? Intel does not support Lookahead on H265, only with H264 and it works there. About VBV, this is what I found in the Media SDK:
For variable bitrate control, the MaxKbps parameter specifies the maximum bitrate at which the encoded data enters the Video Buffering Verifier (VBV) buffer. If MaxKbps is equal to zero, the value is calculated from bitrate, frame rate, profile, level, and so on.
There is a max-bitrate option in QSVEnc, I think it works for VBR and LA_VBR (h264).
Yups
28th March 2021, 13:13
As promised, here my x265 CRF results from Sol Levante at extremely low 4k bitrates. The best Iris Xe CQP preset is basically comparable to x265 medium CRF which I believe is the worst I have tested so far (based on the VMAF scores).
Sol Levante VMAF speed bitrate
Quicksync H265 (27.20.100.9316)
Iris Xe CQP best 72.45 20 fps 1378 kbit
x265 (Staxrip 2.1.9.0)
i7-1165G7 CRF slow 75.44 1.7 fps 1365 kbit
i7-1165G7 CRF medium 72.63 3.8 fps 1378 Kbit
i7-1165G7 CRF very fast 69.30 5.9 fps 1371 Kbit
https://drive.google.com/file/d/1GZ6aIfh1jubKkhKuuIIJVV6k51Wr5GS1/view?usp=sharing
main10 10 bit
gop 120
CQP 47_47_50
offset 2_5_10
ReinerSchweinlin
29th March 2021, 10:56
Thanx for all the work!!
Yups
5th April 2021, 01:47
Another x265/x265 CRF vs Iris Xe CQP comparison using this (https://drive.google.com/file/d/1YX1V0SeSkYaq6Ui41vv1wcOatbLnuLSL/view?usp=sharing) sample.
Intel Demo Clip VMAF PSNR SSIM VQM speed bitrate
Quicksync H265
Iris Xe CQP best 91.76 41.85 0.9748 0.790 67 fps 2438 kbit
x265 (Staxrip 2.1.9.0)
i7-1165G7 CRF slow 93.20 41.90 0.9754 0.800 8 fps 2430 kbit
i7-1165G7 CRF medium 91.05 41.35 0.9744 0.861 19 fps 2440 Kbit
i7-1165G7 CRF very fast 89.99 40.97 0.9726 0.894 32 fps 2440 Kbit
x264 (Staxrip 2.1.9.0)
i7-1165G7 CRF slow 89.38 40.21 0.9693 0.974 30 fps 2425 Kbit
https://drive.google.com/file/d/1onLozSdINzy2xIsHhqq09fgwENHMFYEx/view?usp=sharing
Open gop 120 for all, as for Iris Xe I was using these settings:
CQP 23_24_26
offset 2_5_8
bframes 16
Interestingly Iris Xe wins VQM metric against x265 slow, whereas VMAF/PSNR/SSIM prefer x265 slow over Iris Xe. Personally I prefer VMAF.
benwaggoner
5th April 2021, 21:24
Interestingly Iris Xe wins VQM metric against x265 slow, whereas VMAF/PSNR/SSIM prefer x265 slow over Iris Xe. Personally I prefer VMAF.
Can you expand on that? I've not really compared VMAF and VQM in depth myself. Why do you prefer VMAF?
Of course, all frame-scored metrics suffer from the problem of how to extrapolate from a score for a <10 seconds to a clip of meaningful duration that captures the impact of quality variation throughout a video.
Yups
5th April 2021, 22:00
I never use 10 seconds clips for this reason. As for your question I prefer their scoring system over the others, a higher score means better quality unlike with VQM. 100 means highest, 0 means lowest quality, it's simple.
With SSIM tiny differences in the scores could have big visual subjective differences, there is a tiny 0.001 difference between slow and medium in my last example. Apart from the scoring system I have more trust in VMAF, I believe it's closer to subjective quality than the others. However I wouldn't say it's always closest to subjective quality, there are surely cases where VQM or SSIM will do better.
benwaggoner
6th April 2021, 19:05
I never use 10 seconds clips for this reason. As for your question I prefer their scoring system over the others, a higher score means better quality unlike with VQM. 100 means highest, 0 means lowest quality, it's simple.
With SSIM tiny differences in the scores could have big visual subjective differences, there is a tiny 0.001 difference between slow and medium in my last example. Apart from the scoring system I have more trust in VMAF, I believe it's closer to subjective quality than the others. However I wouldn't say it's always closest to subjective quality, there are surely cases where VQM or SSIM will do better.
And VMAF isn't really a "metric" in the traditional sense. It's a ML model to predict subjective scores based on several relatively simple objective metrics. The ML is trained on a variety of test encodes, and Netflix has periodically improved the ML training, and has added a new objective metric at least once. So the same clip measured with 2018 VMAF would have a different score with the 2021 VMAF.
VMAF is also limited by the sorts of content that was subjectively rated in their training set. Early VMAF didn't seem to have tested different adaptive quantization modes, and so VMAF wasn't accurate in predicting subjective quality between AQ approaches. And it rated dark scenes overly high for some reason, perhaps due to training content limitations. VMAF's underlying objective metrics are all luma-only, so everything gets compared as a black-and-white film. Automated tuning based on VMAF would naturally shift bits from chroma to luma more than is psychovisually appropriate.
VMAF scores are also not based on just the encode itself, but are relative to the resolution compared to. For example a 720p encode compared to a 720p source would have a significantly higher VMAF than the exact same stream compared to the source at 1080p.
I've not played with the latest VMAF much yet, and I presume it is improved in some ways. Which is good! But anyone giving a VMAF score needs to state what the comparison resolution was and what VMAF version was used.
That said, VMAF is absolutely the least-bad metric we've ever had, and is getting better.
Yups
7th April 2021, 20:55
But anyone giving a VMAF score needs to state what the comparison resolution was and what VMAF version was used.
All my VMAF results are based on VMAF 2.0.0 (model 0.6.1) and resolution is unchanged, original resolution for all.
I found an Iris Xe HEVC/AVC comparison from Intel btw: https://dgpu-docs.intel.com/devices/iris-xe-max-graphics/guides/media.html
https://abload.de/img/dg1_hevc_quality_s_cubvk1s.png
benwaggoner
7th April 2021, 21:34
All my VMAF results are based on VMAF 2.0.0 (model 0.6.1) and resolution is unchanged, original resolution for all.
I found an Iris Xe HEVC/AVC comparison from Intel btw: https://dgpu-docs.intel.com/devices/iris-xe-max-graphics/guides/media.html
The metrics from that article are based on Luma PSNR, which isn't something x265 is optimized for (unless you use --tune psnr). The BDRATE differences in luma PSNR is within the range where subjective quality could be quite different; it just isn't that great a metric. And x265 defaults to a lot of psychovisual optimizations that reduce PSNR in favor of improving subjective qualtiy.
That said, these suggest a generally competent encoder for high speed use (presumably why --preset medium was the top option). If you used x265 with --preset slower --tune psnr, x265 likely would win by a fair margin.
Yups
7th April 2021, 21:51
That being said, Intel didn't use the highest quality preset in this (there is another image with quality preset). But of course x265 slower would win in almost every case unless their Ubuntu FFMPEG environment is better than my Windows QSVEnc environment which I don't think it is.
Yups
8th April 2021, 17:12
I could test a GTX 1660 Super next week, it's already running on the current 7th gen (https://developer.nvidia.com/video-encode-and-decode-gpu-support-matrix-new) Nvenc generation, I'm curious how it compares to Iris Xe. Is there a settings tutorial somewhere? Is it correct that 5 bframes is the maximum number of bframes on Nvenc?
benwaggoner
8th April 2021, 20:14
That being said, Intel didn't use the highest quality preset in this (there is another image with quality preset). But of course x265 slower would win in almost every case unless their Ubuntu FFMPEG environment is better than my Windows QSVEnc environment which I don't think it is.
I'm surprised they didn't use their best setting.
An always-interesting question is where the crossover point in speed/quality is between HW and SW encoders.
A key use of GPU encoders is for game streaming, where even 25% CPU utilization would hurt FPS in a lot of games.
Yups
8th April 2021, 21:38
They did use the best setting in the other chart: average BDRATE computed across 27 standard short sequences generated in both CBR and VBR
https://abload.de/img/dg1_hevc_quality_highb4jpe.png
benwaggoner
8th April 2021, 22:51
They did use the best setting in the other chart: average BDRATE computed across 27 standard short sequences generated in both CBR and VBR
Are the axes mislabled or am I misreading? I really doubt that x265 efficiency gets worse with slower presets!
Although faster presets do use less psychovisual optimization, and mainly make choices based on SAD, which maps to PSNR better...
Yups
8th April 2021, 23:04
Are the axes mislabled or am I misreading? I really doubt that x265 efficiency gets worse with slower presets!
Left side of the chart: Bit-rate savings (higher is better).
13.7% bitrate saving for x265 slow over medium and 11.0% higher bitrate required for very fast preset over medium. VME quality is the best Quicksync preset.
Yups
9th April 2021, 17:03
Earlier than expected I got this:
https://abload.de/img/nvencltk3j.png
I will try CQP+Lookahead 32+bframes 5+quality preset later, if there is any other important setting I should use let me know. Bframes 5 is indeed the maximum on Turing.
Yups
10th April 2021, 00:29
I have finished my first GTX 1660 test from my last (https://forum.doom9.org/showpost.php?p=1939967&postcount=437) video sample. I have tried lots of different settings and this is the best I could find (b-frame ref middle gave me a nice score boost).
Intel Demo Clip 1080p VMAF PSNR SSIM VQM speed bitrate
Quicksync H265
Iris Xe CQP best 91.76 41.85 0.9748 0.790 67 fps 2438 kbit
NVENC H265
GTX 1660S CQP best 90.62 41.28 0.9699 0.885 150 fps 2439 Kbit
x265 (Staxrip 2.1.9.0)
i7-1165G7 CRF slow 93.20 41.90 0.9754 0.800 8 fps 2430 kbit
i7-1165G7 CRF medium 91.05 41.35 0.9744 0.861 19 fps 2440 Kbit
i7-1165G7 CRF very fast 89.99 40.97 0.9726 0.894 32 fps 2440 Kbit
x264 (Staxrip 2.1.9.0)
i7-1165G7 CRF slow 89.38 40.21 0.9693 0.974 30 fps 2425 Kbit
https://drive.google.com/file/d/1NzlheKfYJNei6wGJRm3tp3HrnTkn_kaH/view?usp=sharing
https://drive.google.com/file/d/1onLozSdINzy2xIsHhqq09fgwENHMFYEx/view?usp=sharing
Metric scores are a mixed bag, respectable VMAF and PSNR scores but not that good at VQM and especially SSIM metrics. Subjective frame to frame comparison it's obvious detail preservation is a lot worse compared to Iris Xe (VME/GPU) and x265.
Yups
10th April 2021, 15:35
Blender Open Movie from here: https://forum.doom9.org/showpost.php?p=1938861&postcount=423
HERO - Blender Open Movie VMAF speed bitrate
Quicksync H265 (27.20.100.9316)
Iris Xe CQP FF best 86.56 550 fps 199 kbit
Iris Xe CQP FF balanced 85.66 850 fps 201 Kbit
Iris Xe CQP FF speed 76.25 1550 fps 200 Kbit
Iris Xe CQP best 88.59 167 fps 200 kbit
Iris Xe CQP balanced 87.18 280 fps 201 Kbit
Iris Xe CQP speed 86.22 520 fps 200 Kbit
NVENC H265 (470.14)
GTX 1660S CQP best 84.52 430 fps 200 Kbit
GTX 1660S CQP default 83.55 990 fps 201 Kbit
GTX 1660S CQP performance 79.22 1130 fps 200 Kbit
x265 (Staxrip 2.1.9.0)
i7-1165G7 CRF slower 90.25 8 fps 200 Kbit
i7-1165G7 CRF slow 88.18 34 fps 200 kbit
i7-1165G7 CRF medium 85.72 56 fps 200 Kbit
i7-1165G7 CRF very fast 83.36 73 fps 200 Kbit
x264 (Staxrip 2.1.9.0)
i7-1165G7 CRF slower 75.37 81 fps 200 Kbit
Disabled b-adapt is better for this video. Turing CQP cannot reach Iris Xe CQP quality, subjective and objective the difference is large.
Turing has two downsides, only 5 bframes versus 16 bframes on Iris Xe and there is no GPU equivalent mode which is more flexible than a fully fixed function solution, however even the FF mode from Iris Xe looks better. It might look different with CBR vs CBR which I haven't tried. That said, the H265 CQP results from Turing are really good for a hardware encoder, something like x265 fast-faster with extremely fast encoding times, the CQP quality from Iris Xe is just insane.
Tenkei
10th April 2021, 17:41
Is there any reason to use CQP instead of ICQ with QuickSync. Never used it but it seems that ICQ is CRF equivalent. Did you try --ctu 64 and --ref X?
Yups
10th April 2021, 20:03
CQP with custom offset offers higher quality than ICQ, this old Quicksync bitrate method overview is still valid:
Constant QP (CQP) provides the most control and best performance. Without question, the best coding efficiency with Intel codecs can be obtained via CQP plus custom content analysis. CQP often has significant performance advantages as well. CQP operates most closely to reference implementations. It is the most direct way to access codec capabilities and measure the effects of encoder parameter/algorithm trade-offs and also is the clearest way to evaluate against other codec algorithm implementations.
https://software.intel.com/content/www/us/en/develop/articles/common-bitrate-control-methods-in-intel-media-sdk.html
For a basic user ICQ is easier to handle, there is just one global setting and that's it. Furthermore ICQ does not really scale over 5 bframes (16 bframes can be worse than 5 at low bitrate) whereas CQP scales really good beyond 5 bframes even at low bitrate. Here I did include both ICQ and CQP: https://forum.doom9.org/showpost.php?p=1930663&postcount=369
On Iris Xe it automatically uses ctu 64 (Gen 9 ctu 32), this can't be changed at the moment. Tskip and SAO are also enabled on Tigerlake which I can't disable. Reference frames best leave it auto, with 16 bframes+bpyramid Intel sets it to 6 reference frames, I've tried 8 reference frames but there is no improvement.
Yups
17th April 2021, 23:05
CQP best runs ~5% faster with Intels new driver build 9466 (https://downloadcenter.intel.com/download/30381/Intel-Graphics-Windows-10-DCH-Drivers) on Iris Xe.
I was searching for CQP improvements and noticed that 14 and 15 bframes offers slightly better scores at a slightly lower bitrate on Iris Xe over 16 bframes which I was using. This is tested on Intel Demo Clip and Blender Open Movie low bitrate.
13 bframes and lower gradually decreases bitrate efficiency, it's a big degradation from 14 to 13 bframes. 14 bframes is a tiny bit better than 15. For better context:
Intel Demo Clip 1080p VMAF PSNR SSIM VQM speed bitrate
Quicksync H265 27.20.100.9466
Iris Xe CQP best 16 bframes 91.76 41.85 0.9748 0.790 73 fps 2438 kbit
Iris Xe CQP best 14 bframes 91.98 41.96 0.9754 0.780 74 fps 2426 Kbit
NVENC H265
GTX 1660S CQP best 90.62 41.28 0.9699 0.885 150 fps 2439 Kbit
x265 (Staxrip 2.1.9.0)
i7-1165G7 CRF slow 93.20 41.90 0.9754 0.800 8 fps 2430 kbit
i7-1165G7 CRF medium 91.05 41.35 0.9744 0.861 19 fps 2440 Kbit
i7-1165G7 CRF very fast 89.99 40.97 0.9726 0.894 32 fps 2440 Kbit
x264 (Staxrip 2.1.9.0)
i7-1165G7 CRF slow 89.38 40.21 0.9693 0.974 30 fps 2425 Kbit
Bitrate goes down from 2438 to 2426 and scores go up, it's a clear win.
MGarret
18th April 2021, 21:51
CQP with custom offset offers higher quality than ICQ, this old Quicksync bitrate method overview is still valid:
https://software.intel.com/content/www/us/en/develop/articles/common-bitrate-control-methods-in-intel-media-sdk.html
For a basic user ICQ is easier to handle, there is just one global setting and that's it. Furthermore ICQ does not really scale over 5 bframes (16 bframes can be worse than 5 at low bitrate) whereas CQP scales really good beyond 5 bframes even at low bitrate. Here I did include both ICQ and CQP: https://forum.doom9.org/showpost.php?p=1930663&postcount=369
On Iris Xe it automatically uses ctu 64 (Gen 9 ctu 32), this can't be changed at the moment. Tskip and SAO are also enabled on Tigerlake which I can't disable. Reference frames best leave it auto, with 16 bframes+bpyramid Intel sets it to 6 reference frames, I've tried 8 reference frames but there is no improvement.
What a load of misleading information in that Intel article. CQP is not even a "proper" rate control. Of course it operates closely to reference implementations because, guess what: reference software don't even have real rate control algorithms. So, yeah, CQP with some custom software that analyzes characteristics and complexity of every scene is what huge content providers like Netflix and others use and play around to squeeze more optimized encodes to save bandwidth. I see it as an entire pipeline where several tools are used and not just one rate control algorithm. When this kind of algorithm is implemented inside the encoder, then it won't be called CQP but something else.
Yups
18th April 2021, 22:01
Who cares? The point is CQP offers higher quality than ICQ, the Intel article is correct on this.
MGarret
18th April 2021, 23:57
The point is you obviously don't have any proof for your claim. I just pointed out what I think about some dubious claims from that article.
If you believe that just using fixed qp's is going to create efficient encoding in terms of file size/picture quality... well, I'm not going to persuade you otherwise. We all have our specific needs in regards to using encoders. Keep believing in your vmaf/psnr/ssim numbers.
Yups
19th April 2021, 15:49
The point is you obviously don't have any proof for your claim. I just pointed out what I think about some dubious claims from that article.
What claim? The claim that ICQ is worse than CQP on Iris Xe? Dude if ICQ would offer better results over CQP I wouldn't use CQP. It would be insane if ICQ would offer even better results over CQP. The claim that CQP offers best coding efficiency on Intel isn't dubious, dubious are your postings.
If you believe that just using fixed qp's is going to create efficient encoding in terms of file size/picture quality... well, I'm not going to persuade you otherwise. We all have our specific needs in regards to using encoders. Keep believing in your vmaf/psnr/ssim numbers.
Nonsense. Subjective and objective CQP is clearly better than ICQ, I told this more than once and I have uploaded several samples. If you don't agree prove it!
If you don't believe in vmaf/psnr/ssim numbers download the sample or ask for the upload if I didn't upload any wanted sample, I'm always checking for subjective quality my results. I don't force you to believe any numbers.
Bframes scaling with CQP is way better than with ICQ or CBR/VBR on Iris Xe. CBR/VBR 7 bframes are best and on ICQ there is no real improvement over 5 bframes. With CQP it scales up to 14-16 bframes, this is one reason why ICQ can't reach CQP. This is something you can't know because you are obviously clueless.
About fixed CQP, I mean depending on the bitrate/quality target different quantization parameter are required, this is no different to any other constant rate factor. I don't understand your problem to be honest, I mean it's super easy on Intel (easier than on Nvidia).
MGarret
20th April 2021, 12:51
What claim? The claim that ICQ is worse than CQP on Iris Xe? Dude if ICQ would offer better results over CQP I wouldn't use CQP. It would be insane if ICQ would offer even better results over CQP. The claim that CQP offers best coding efficiency on Intel isn't dubious, dubious are your postings.
Nonsense. Subjective and objective CQP is clearly better than ICQ, I told this more than once and I have uploaded several samples. If you don't agree prove it!
If you don't believe in vmaf/psnr/ssim numbers download the sample or ask for the upload if I didn't upload any wanted sample, I'm always checking for subjective quality my results. I don't force you to believe any numbers.
Bframes scaling with CQP is way better than with ICQ or CBR/VBR on Iris Xe. CBR/VBR 7 bframes are best and on ICQ there is no real improvement over 5 bframes. With CQP it scales up to 14-16 bframes, this is one reason why ICQ can't reach CQP. This is something you can't know because you are obviously clueless.
About fixed CQP, I mean depending on the bitrate/quality target different quantization parameter are required, this is no different to any other constant rate factor. I don't understand your problem to be honest, I mean it's super easy on Intel (easier than on Nvidia).
Hey chum, did I struck the nerve somehow?
Here, read about about etiquette on this forum.
4) Be nice to each other and respect the moderator. Profanity and insults will not be tolerated.
I just stated my opinion and you found yourself offended because I somehow negated your findings? Are you that sensitive? I don't even know about your samples because this thread is couple of years long and I read it occasionally. I still stand by common knowledge that CQP is inefficient and "dumb" type of rate control. I don't know what Intel is doing differently and I don't care. Enjoy your Iris XE or whatever.
Tenkei
20th April 2021, 18:44
Both of you can be correct. Usually RCF mode means better quality for the same bitrate, but if Intel's ICQ is broken it might reduce too much, like hevc-aq does in x265. Can't really conclude anything, without comparing clips of several scenes, one scene is not enough. Hand-picker quantitizer might be better for one scene, but it won't be an efficient way to encode hour long video.
Also, one frame might look better on CQ and other would be better on ICQ, so it depends on what are we comparing. Do you see quality difference in running clip or just a single random frame you chose.
ReinerSchweinlin
24th April 2021, 15:24
Blender Open Movie from here: https://forum.doom9.org/showpost.php?p=1938861&postcount=423
.
.
.
.
Disabled b-adapt is better for this video. Turing CQP cannot reach Iris Xe CQP quality, subjective and objective the difference is large.
Thanx for investigating this :) I am just waiting for the i5 NUCs to be available :)
Yups
24th April 2021, 21:21
Both of you can be correct. Usually RCF mode means better quality for the same bitrate, but if Intel's ICQ is broken it might reduce too much, like hevc-aq does in x265. Can't really conclude anything, without comparing clips of several scenes, one scene is not enough. Hand-picker quantitizer might be better for one scene, but it won't be an efficient way to encode hour long video.
ICQ isn't broken, it's how it is. That being said, H265 ICQ isn't even fully supported in fixed function mode from the hardware (rate factor can't be changed). In the end you can use whatever you want, personally I won't bother too much with ICQ because CQP offers me a better quality and slightly higher performance.
HERO - Blender Open Movie VMAF PSNR SSIM VQM speed bitrate
Quicksync H265 27.20.100.9466
Iris Xe CQP best 14 bframes 88.94 41.13 0.9864 0.544 191 fps 198 kbit
Iris Xe ICQ best 14 bframes 86.72 40.02 0.9841 0.609 174 fps 198 Kbit
@ReinerSchweinlin, hopefully we can get DG2 this year. They could aim for something higher in GPU hybrid mode, I mean this GPU comes with much more EUs and higher GPU clock speeds. If not it should run a lot faster in hybrid mode than Iris Xe iGPU.
ukmark
9th May 2021, 11:59
Does the Intel graphics driver version have any effect on the quality/speed of QuickSync fixed function (HEVC) encoding?
I know that the API version is linked to the graphics driver and affects quality/speed of non-FF encoding, but I am guessing that fixed function encoding can only be improved by upgrading the GPU?? (i.e. upgrade your laptop etc).
The reason I ask is that I'm on the latest driver as of 9/5/2021 (27.20.100.9466) and I did a quick test using FF encoding and it appeared to me that the quality (in terms of facial detail retention) has improved from a couple of months ago. Not a scientific test, but just wondering if anyone knew of a link between graphics driver version and FF encoding quality (if any).
TIA
I don't know about FF mode (haven't checked), on CQP Hybrid with Tigerlake there is a difference between some of the drivers. 9466 runs ~5% faster which I have mentioned in #454 and 30.0.100.9563 (inside preview driver) adds another 5% speedup on top of this. My VMAF scores did improve slightly over time when I compare it with a 6 months old driver, but these changes are not regularly from what I have seen. On my old Gen9 I have't seen differences for a long time, because this GPU architecture is very old.
ReinerSchweinlin
16th May 2021, 15:04
@ReinerSchweinlin, hopefully we can get DG2 this year. They could aim for something higher in GPU hybrid mode, I mean this GPU comes with much more EUs and higher GPU clock speeds. If not it should run a lot faster in hybrid mode than Iris Xe iGPU.
just checked techpowerup release info page, no updates so far. Let´s hope for the best :)
ukmark
27th May 2021, 13:18
and 30.0.100.9563 (inside preview driver) adds another 5% speedup on top of this.
How do you get your hands on the preview drivers? Thx
EDIT: NVM I managed to get it from "station-drivers" (https://www.station-drivers.com/index.php?option=com_kunena&view=topic&defaultmenu=860&Itemid=858&catid=17&id=468&lang=en&limitstart=12#5005)
benwaggoner
28th May 2021, 00:59
I'm finally getting my RTX a6000 tomorrow, if there are any tests to run on that. I generally expect HW encoding quality to be identical to other RTX 3000 series GPUs.
StormMeows
1st June 2021, 01:39
Deleted
benwaggoner
3rd June 2021, 19:03
I've been wondering how many 8-bit only HEVC decoders are out in the wild still. In the living room, there were a few early SDR UHD Smart TVs that possibly had only 8-bit decoders. There were some phone SoCs that were 8-bit HEVC only around five years ago.
Anyone have a guess as to when and what the last devices to ship with Main but not Main10 HEVC decode were? Are there any significant clusters of devices still in use with those limitations?
Blue_MiSfit
4th June 2021, 05:35
Android phones in the 5-7 year age range are probably the biggest cohort I can think of. The software fallback decoder in Android is very limited as well ;)
benwaggoner
4th June 2021, 19:23
Android phones in the 5-7 year age range are probably the biggest cohort I can think of. The software fallback decoder in Android is very limited as well ;)
Right, that was the biggest group of 8-bit only decoders I knew of. Of course with mobile replacement rates being what they are, most of those devices are already out of use and the remaining number is presumably shrinking rapidly.
Living room devices, PCs, and tablets get replaced more slowly, so 8-bit only devices (if any) in those categories may be the more enduring issue. I can't think of any living-room devices that were 8-bit only, but I'm sure some tablets were based on the same 8-bit SoCs used in some phones.
pandy
9th July 2021, 16:01
I've been wondering how many 8-bit only HEVC decoders are out in the wild still. In the living room, there were a few early SDR UHD Smart TVs that possibly had only 8-bit decoders. There were some phone SoCs that were 8-bit HEVC only around five years ago.
Anyone have a guess as to when and what the last devices to ship with Main but not Main10 HEVC decode were? Are there any significant clusters of devices still in use with those limitations?
Not direct answer but terrestrial TV in Germany use H.265 1080p50 at 8 bit not 10 bit - weird but this absolute minimum for receivers in DE so it may be situation where at least some SoC's will not support 10 bit. Unclear to me why Germany stuck with 8 bit.
benwaggoner
9th July 2021, 20:16
Not direct answer but terrestrial TV in Germany use H.265 1080p50 at 8 bit not 10 bit - weird but this absolute minimum for receivers in DE so it may be situation where at least some SoC's will not support 10 bit. Unclear to me why Germany stuck with 8 bit.
Any idea what SoC is used in those?
pandy
9th July 2021, 23:22
Any idea what SoC is used in those?
Nope - there is quite actual list of accepted devices (but there is everything, even antennas there) - https://tv-plattform.de/en/themen/dvb-t2-hd/geraeteliste/ .
SoC's itself may be even 10 bit compliant but firmware specifically targeting German requirements may support only 8 bit...
RedDwarf1
26th July 2021, 01:52
Hi everyone,
I asked a question in the Hybrid encoder thread but so far no one has answered my question(s). This thread might also fit my question because it is nVidia hardware GPU (GTX1660 NOT Super) encoding related.
With regard to HEVC/H.265:
Hybrid keeps outputting CAVLC and not CABAC entropy encoding video. I would much prefer CABAC. The question I have is: is this because nVidia GPU's can only encode with CAVLC? Has anyone managed to encode HEVC/H.265 to CABAC using a nVidia GPU? My GPU is fairly recent and there has only been one update to the NVENC encoder since then to my knowledge.
Has anyone tested nVidia HEVC encoding on Turing GPU's or even noticed it since MediaInfo and similar programs do not seem to show the entropy encoding of HEVC video. I use H264/H265 BS Analyser v3.0 which does the very basics but is a bit limited especially with large files.
https://github.com/latelee/H264BSAnalyzer/releases
Can anyone shed any light on this?
benwaggoner
26th July 2021, 19:48
Hi everyone,
I asked a question in the Hybrid encoder thread but so far no one has answered my question(s). This thread might also fit my question because it is nVidia hardware GPU (GTX1660 NOT Super) encoding related.
With regard to HEVC/H.265:
Hybrid keeps outputting CAVLC and not CABAC entropy encoding video. I would much prefer CABAC. The question I have is: is this because nVidia GPU's can only encode with CAVLC? Has anyone managed to encode HEVC/H.265 to CABAC using a nVidia GPU? My GPU is fairly recent and there has only been one update to the NVENC encoder since then to my knowledge.
I don't believe HEVC and x265 even support CAVLC; it's just H.264. Are you sure you're testing with HEVC?
RedDwarf1
26th July 2021, 22:47
Yes it is HEVC. I have just software encoded using x265 and the video shows as having CABAC entropy encoding whereas the nVidia GPU encoding always shows as CAVLC according to H264/H265 BSAnalyser.
It might be worth you testing for yourself and checking with the stream analyser. MediaInfo etc does not show this information for HEVC
Here is the info copied from the analyser>
H.265/HEVC File Information
Picture Size : 640x480
- Cropping Left : 0
- Cropping Right : 0
- Cropping Top : 0
- Cropping Bottom : 0
Video Format : YUV420 Luma bit: 8 Chroma bit: 8
Stream Type : Main Profile @ Level 3(90) Tier Main
Encoding Type : CABAC
Max fps : 30.000
Frame Count : 4936
It does seem to be getting the number of frames incorrect because there should be 5001.
benwaggoner
26th July 2021, 23:47
Yes it is HEVC. I have just software encoded using x265 and the video shows as having CABAC entropy encoding whereas the nVidia GPU encoding always shows as CAVLC according to H264/H265 BSAnalyser.
It might be worth you testing for yourself and checking with the stream analyser. MediaInfo etc does not show this information for HEVC
Here is the info copied from the analyser>
H.265/HEVC File Information
Picture Size : 640x480
- Cropping Left : 0
- Cropping Right : 0
- Cropping Top : 0
- Cropping Bottom : 0
Video Format : YUV420 Luma bit: 8 Chroma bit: 8
Stream Type : Main Profile @ Level 3(90) Tier Main
Encoding Type : CABAC
Max fps : 30.000
Frame Count : 4936
It does seem to be getting the number of frames incorrect because there should be 5001.
Can you give an example of the nvenc output with CAVLC?
RedDwarf1
28th July 2021, 05:44
Can you give an example of the nvenc output with CAVLC?
Thanks for your contribution.
I ran into some problems which has prevented me from adding to this thread. Firstly Hybrid encoder would not run. It crashed on startup. An install with the latest version which I had not got around to install before then fix that, at least it now runs & starts up.
Next I ran into a problem encoding the small video that I used with the x265 encoder with MeGUI. I used the same avisynth script which worked fine with MeGUI but Hybrid kept failing to encoder it and I have not yet managed to sort that out.
----------------------
I have managed to get it encoded, I do not know why it was not working when I tried previously because I did not alter anything when I re-tried and it did successfully encode the video.
H.265/HEVC File Information
Picture Size : 640x480
- Cropping Left : 0
- Cropping Right : 0
- Cropping Top : 0
- Cropping Bottom : 0
Video Format : YUV420 Luma bit: 8 Chroma bit: 8
Stream Type : Main Profile @ Level 5.2(156) Tier High
Encoding Type : CAVLC
Max fps : 30.000
Frame Count : 5001
So you can see it is saying this video encoded with the nVidia NVEnc is CAVLC whereas the same video encoded with x265 comes out with CABAC.
tonemapped
28th July 2021, 06:10
Thanks for your contribution.
I ran into some problems which has prevented me from adding to this thread. Firstly Hybrid encoder would not run. It crashed on startup. An install with the latest version which I had not got around to install before then fix that, at least it now runs & starts up.
Next I ran into a problem encoding the small video that I used with the x265 encoder with MeGUI. I used the same avisynth script which worked fine with MeGUI but Hybrid kept failing to encoder it and I have not yet managed to sort that out.
----------------------
I have managed to get it encoded, I do not know why it was not working when I tried previously because I did not alter anything when I re-tried and it did successfully encode the video.
H.265/HEVC File Information
Picture Size : 640x480
- Cropping Left : 0
- Cropping Right : 0
- Cropping Top : 0
- Cropping Bottom : 0
Video Format : YUV420 Luma bit: 8 Chroma bit: 8
Stream Type : Main Profile @ Level 5.2(156) Tier High
Encoding Type : CAVLC
Max fps : 30.000
Frame Count : 5001
So you can see it is saying this video encoded with the nVidia NVEnc is CAVLC whereas the same video encoded with x265 comes out with CABAC.
Why L5.2?
RedDwarf1
28th July 2021, 07:31
Why L5.2?
I think that the encoder chose that because it was set to auto.
benwaggoner
28th July 2021, 17:35
I have managed to get it encoded, I do not know why it was not working when I tried previously because I did not alter anything when I re-tried and it did successfully encode the video.
H.265/HEVC File Information
Picture Size : 640x480
- Cropping Left : 0
- Cropping Right : 0
- Cropping Top : 0
- Cropping Bottom : 0
Video Format : YUV420 Luma bit: 8 Chroma bit: 8
Stream Type : Main Profile @ Level 5.2(156) Tier High
Encoding Type : CAVLC
Max fps : 30.000
Frame Count : 5001
So you can see it is saying this video encoded with the nVidia NVEnc is CAVLC whereas the same video encoded with x265 comes out with CABAC.
Color me baffled! I thought CAVLC straight-up wasn't available with HEVC.
I just did a test encode with Adobe Media Encoder in HW encoder mode, which also uses nvenc under the hood, and MediaInfo in "Details 5" tells me that it is CABAC:
Line 523: 00002D8 cabac_init_present_flag: Yes
Line 1158: 001AB78 cabac_init_present_flag: Yes
Line 3966: 00002D9 cabac_init_present_flag: Yes
Line 4601: 001AB7A cabac_init_present_flag: Yes
Without any reference to CAVLC at all.
What tool are you using to get your report data above? Can you double-check with MediaInfo set to Details 5?
RedDwarf1
29th July 2021, 06:17
Color me baffled! I thought CAVLC straight-up wasn't available with HEVC.
I just did a test encode with Adobe Media Encoder in HW encoder mode, which also uses nvenc under the hood, and MediaInfo in "Details 5" tells me that it is CABAC:
Line 523: 00002D8 cabac_init_present_flag: Yes
Line 1158: 001AB78 cabac_init_present_flag: Yes
Line 3966: 00002D9 cabac_init_present_flag: Yes
Line 4601: 001AB7A cabac_init_present_flag: Yes
Without any reference to CAVLC at all.
What tool are you using to get your report data above? Can you double-check with MediaInfo set to Details 5?
Ah that sounds promising but I cannot replicate it. I used the H264/H265 BSAnalyser which I linked to on the previous page. It is a free bitstream analyser
I have tried Details 5 and I get a totally blank page with any file I open be it x264 or x265/HEVC :( That is on the Debug menu isn't it? I have [View] set to Text view BTW. Changing the view to another and back to Text does seem to refresh it and display output. However after doing that it does not show the entropy encoding. I have copied all the long output and searched for CABAC & CAVLC without finding either.
For the above I did upgrade to the latest portable version. I have checked a few files with the same result. I do not have too many HEVC files having recently started using it.
I do not have that Adobe application, I will have to try Handbrake or see if I can find some other program which supports hardware encoding.
Handbrake shows CAVLC when the raw video is opened with H264/H265 BSAnalyser. I can only see one way around this ATM. Can you test your video with the following program to see what it shows? BTW It does not like large files so you might need to chop the file up with something. It requires raw video and will not open video in a container (Mp4/Mkv).
https://github.com/latelee/H264BSAnalyzer/releases
benwaggoner
29th July 2021, 17:24
I have tried Details 5 and I get a totally blank page with any file I open be it x264 or x265/HEVC :( That is on the Debug menu isn't it? I have [View] set to Text view BTW. Changing the view to another and back to Text does seem to refresh it and display output. However after doing that it does not show the entropy encoding. I have copied all the long output and searched for CABAC & CAVLC without finding either.
Yeah, you found a fun MediaInfo defect. You need to load the file after setting the correct settings. Here's the full process
Open MediaInfo
Set Debug > Details 5
Set View > Text
Drag your file into the MediaInfo window
Mister XY
1st August 2021, 14:54
Are there settings for qsvenc, which are similar to crf 20 and the preset medium?
tonemapped
1st August 2021, 15:25
Are there settings for qsvenc, which are similar to crf 20 and the preset medium?
Use ICQ and setting the quality to 17 gives results somewhat comparable to x265 at crf 20. If you have a noisy source, don't use QSV.
Mister XY
1st August 2021, 19:22
I have tested with icq 17 but i have also blur/artefact in dark moments. Here are my settings for staxrip.
--codec hevc --quality best --profile main10 --slices 1 --bframes 14 --ref 5 --b-pyramid --open-gop --weightb --weightp --aud --icq 18
tonemapped
1st August 2021, 19:42
I have tested with icq 17 but i have also blur/artefact in dark moments. Here are my settings for staxrip.
There are other settings in Staxrip you can change and I'll paste the options I use, but I don't have access to my system with QSV set up. You should increase the lookahead frames to the maximum (think it's either 30 or 50 - can't remember).
Mister XY
1st August 2021, 20:19
lookahead frames means LA ICQ or? That are for x264. I have not see this in hevc.
Yups
1st August 2021, 22:41
I have tested with icq 17 but i have also blur/artefact in dark moments. Here are my settings for staxrip.
What Intel GPU are you using? This makes a big difference. On Gen9 you cannot reach x265 medium quality. With ICQ there isn't much you can do to improve. CQP is higher quality but you have to use a custom offset, 2:4:8 or 2:5:8 works quite good.
tonemapped
2nd August 2021, 00:01
What Intel GPU are you using? This makes a big difference. On Gen9 you cannot reach x265 medium quality. With ICQ there isn't much you can do to improve. CQP is higher quality but you have to use a custom offset, 2:4:8 or 2:5:8 works quite good.
I'm using Gen '9.5' (UHD 605) and it's possible to reach medium x265, but it requires more bitrate (to be expected).
Mister XY
2nd August 2021, 03:41
I have a Intel Gen 11 with UHD 750
tonemapped
2nd August 2021, 04:43
I have a Intel Gen 11 with UHD 750
Sorry for the delayed reply. I'm using my 15W HTPC (see signature) for Intel testing as I have the iGPU disabled in all other systems. Overall, I find QSV to offer better results for any grainy source, or even most sources if you're prepared for encodes to be slower.
Here's something to start you off in Staxrip 2.4.x (should apply to 2.6.x and beyond)
1. Select Intel - H265 from the menu
2. On the 'Basic' screen, select:
- mode: ICQ
- Preset: Best
- Profile: Main 10
- Level: Automatic
- Quality: 20 is a reasonable start. As with everything, it depends on the source.
Then make sure these options are enabled (you can use the search box at the bottom left of the screen): --codec hevc --quality best --profile main10 --trellis all --ctu 32 --la-quality slow --la-window-size 50 --bframes 8 --ref 4 --b-pyramid --b-adapt --i-adapt --open-gop --weightb --weightp --input-buf 16 --vpp-denoise 10 --vpp-detail-enhance 20 --sao none --fallback-rc --tskip --icq 22
That has produced perfectly reasonable results, although you will generally end up using more bitrate than x265. Some example comparisons (using older, less refined options) can be found in this thread (http://forum.doom9.net/showthread.php?t=183035).
Since you're not a purist - like me - otherwise you'd use software encoding, you can increase the perceived quality by using '--vpp-denoise 10 --vpp-detail-enhance 20' as noted above.
Again, this is just a starting point and you may find that some features are not support, but Staxrip lets the driver decide and it will automatically disabled unsupported features (unlikely given you have Gen 11).
I'd be very interested to see some screenshots of your results. Hope that's helped a bit.
Yups
2nd August 2021, 10:27
I have a Intel Gen 11 with UHD 750
That's fine, it's Gen12/Xe based. On Xe many bframes is no benefit with ICQ, actually bframes 5 was the optimum with ICQ in my tests. This is the only real change you could try.
However for highest quality you have to switch to CQP.
- CQP bitrate mode
- preset best
- bframes 14-16 (bframes 14 is slighty better in my tests)
- open gop
- b-pyramid (automatically enabled anyway)
- no fixed function
- QP offset I/P/B 2:4:8 or 2:5:8 works good
- QP I -1/-2
- QP P xx (depends on your bitrate/compression needs)
- QP B: +1/+2
QP P heavily depends on resolution and your bitrate/size budget, something between 20-25 (8 bit) is a good try. QP I needs to be 1-2 lower than P and QP B 1-2 higher than P. This works quite good with the offset from above.
If you are using 10 bit you have to add +12 on the CP quantizer for roughly the same bitrate. For exampe 23_24_26 8 bit is (roughly) the equivalent of 35_36_38 10 bit.
@tonemapped, almost all of your settings are useless because they have no effect, you can't enable/disable tskip, weight, trellis (h264 only), ctu, b-adapt, la (no LA for h265!).
Mister XY
2nd August 2021, 20:14
Thats are my settings now
--codec hevc --quality best --profile main10 --bframes 14 --open-gop --cqp 28:29:30 --qp-offset 2:5:8
Its look like better than icq.
Mister XY
5th August 2021, 18:56
Here is a screenshot with the macro blocks what i mean.
The block formation can be seen very clearly in dark areas.
Here are the settings.
--codec hevc --quality best --profile main10 --bframes 14 --open-gop --cqp 28:29:30 --qp-offset 2:4:8
Yups
5th August 2021, 20:54
We would need the original source or a small sample of this to check this out. The screenshot looks low quality to me but there is not much to say without a sample.
Mister XY
6th August 2021, 04:18
OK, here is a screenshot from the original source.
Mister XY
9th August 2021, 06:27
So, that are my settings for 4K and FHD encoding.
--codec hevc --quality best --profile main10 --bframes 14 --ref 6 --b-pyramid --open-gop --weightb --weightp --cqp 27:28:29 --qp-offset 2:4:8.
I have tested some 4k Movies and FHD Movies.
These settings are roughly the same size as with crf 20 slow.
In some movies, you see macroblocks.
If you concentrate closely, you will see macroblocks. This is noticeable from time to time, especially when there are many consistent colors.
But I also think that if SAO could be deactivated, the blocks would also disappear. Unfortunately, SAO is always on.
But for now I'm satisfied when you consider how fast I encode with the GPU instead of the CPU.
Yups
9th August 2021, 11:35
It's a very grainy original source and SAO might have an effect there. I don't know if some faster presets (including low power mode) automatically disable SAO, they are usually lower quality but in such specific case if some presets don't use SAO it might be worth a try. I wonder if it's a qsvenc issue or something else why we can't disable sao. According to qsvenc feature check it's not supported but the encoding log says SAO all.
Mister XY
9th August 2021, 18:31
what do you mean with low power mode?
Detected avaliable features for hw API v1.35, HEVC, Constant QP (CQP)
RC mode o
10bit depth o
Fixed Func o
Interlace o
VUI info o
Trellis x
Adaptive_I x
Adaptive_B x
WeightP o
WeightB o
FadeDetect o
B_Pyramid o
+ManyBframes o
PyramQPOffset o
MBBRC x
ExtBRC x
Adaptive_LTR x
LA Quality x
QP Min/Max x
IntraRefresh x
No Deblock o
No GPB o
Windowed BRC x
PerMBQP(CQP) o
DirectBiasAdj x
MVCostScaling x
SAO x
Max CTU Size x
TSkip x
HEVC SAO is not supported on current platform, disabled.
Yups
9th August 2021, 22:45
It's called fixed function in QSVEnc. Intel calls it low power.
--fixed-func use fixed func instead of GPU EU (default: off)
Mister XY
10th August 2021, 04:21
with --fixed-func qsvenc has a crash, but i see in log
HEVC SAO is not supported on current platform, disabled.
Yups
10th August 2021, 10:04
SAO isn't supported by QSVEnc and that's why you get this message. This is unrelated to the FF/low power mode.
Mister XY
11th August 2021, 19:27
How good is the ryzen 5600g in hardware encode?
RanmaCanada
12th August 2021, 03:43
How good is the ryzen 5600g in hardware encode?
AMD has always sucked (https://obsproject.com/forum/resources/ultimate-encoder-quality-analysis-2020-nvenc-vs-amf-vs-quicksync-vs-x264.998/). The 5600G runs Vega 7, and uses VCN 2.0. If you want to do hardware encoding, Intel Quicksync, especially the newest hardware, runs complete circles around it. Heck even Nvenc on Maxwell is better than AMD in some instances.
Mister XY
13th August 2021, 20:27
Ok then I know that the Ryzen is not good. Then I ask the question, what is the best APU / GPU for hardware encoding in x265 format?
Yups
13th August 2021, 22:54
CBR/VBR= Turing/Ampere
CQP= Iris Xe
Nivida has a big advantage with Lookahead in CBR/VBR mode but overall CQP offers best quality on Iris Xe.
Mister XY
14th August 2021, 07:06
That means that I currently have the best hardware for encoding because I don't use cbr or vbr.
Yups
14th August 2021, 10:22
It could be yes, although it's probably much slower than Iris Xe unless it's running in low power mode. I mean you can try and upload a sample, if you tell me the settings I can check out if it's the same. This (https://forum.doom9.org/showpost.php?p=1939967&postcount=437) or this (https://forum.doom9.org/showpost.php?p=1938769&postcount=422) you can use.
Mister XY
14th August 2021, 12:54
How do you check the vmaf?
Yups
14th August 2021, 13:08
I don't know. I'm using Video Quality Measurement Tool which isn't free. If the file size is identical there is no need for a quality metrics check.
Mister XY
14th August 2021, 13:15
OK, what for a sample do you need? A sample from surce ir from encode?
Yups
14th August 2021, 13:18
I need the encoded sample and settings you were using and driver version. If it's a known source for me I don't need that, so it depends.
Yups
14th August 2021, 22:53
Posting it here because some people might be interested. We made a comparison with the exact same settings between TGL-U Iris Xe and RKL-S UHD 750. It's a CQP best quality setting and software decoding.
File size/bitrate:
Iris Xe= 494.834 KB/8307.92 kbps
UHD750= 494.837 KB/8307.90 kbps
VMAF:
Iris Xe= 97.358
UHD 750= 97.358
Very minor file size/bitrate difference and same quality. From this I can say it's the same encoder which was expected but good to know it's the same. I was thinking the minor bitrate difference is decoding related. Another run with hardware decoding:
Iris Xe= 493.277 Kb
UHD 750= 493.286 Kb
Minor size difference as well. Iris Xe has two decoder unlike UHD750, this minor size difference could still related to decoding somehow. Doesn't matter in the end, same quality. Speed difference wasn't that big by the way.
Yups
11th September 2021, 09:34
From the release notes of Intels latest driver:
Support for H264 and HEVC DX12 video encode on Microsoft Windows® 11 for 11th Generation Intel® Core™ Processors.
https://downloadmirror.intel.com/648245/ReleaseNotes_100.9864.pdf
What is DX12 video encode and why Windows 11 only?
nevcairiel
11th September 2021, 09:53
Windows 11 adds a encoding API to DX12, which would be interesting as a cross-vendor way to access video encoding hardware, but the documentation on it seems to still be very much work in progress, so we'll see how usable that is once someone figures out how to use it.
Balling
14th September 2021, 18:17
From the release notes of Intels latest driver:
https://downloadmirror.intel.com/648245/ReleaseNotes_100.9864.pdf
What is DX12 video encode and why Windows 11 only?
dx11va is old. directx11over12 decoder will allow to do dx12 only and already works in Chrome under a flag.
Yups
9th December 2021, 15:20
Windows 11 adds a encoding API to DX12, which would be interesting as a cross-vendor way to access video encoding hardware, but the documentation on it seems to still be very much work in progress, so we'll see how usable that is once someone figures out how to use it.
Microsoft released it, there is a blog post: https://devblogs.microsoft.com/directx/announcing-new-directx-12-feature-video-encoding/
And on github: https://github.com/microsoft/DirectX-Specs/blob/master/d3d/D3D12VideoEncoding.md
tonemapped
16th December 2021, 22:36
dx11va is old. directx11over12 decoder will allow to do dx12 only and already works in Chrome under a flag.Chrome barely operates correctly with AV1 :D One of the recent updates changed my flag and decided my Atom-powered HTPC would love 4K AV1 :(
pacuro
19th February 2022, 18:55
It's called fixed function in QSVEnc. Intel calls it low power.
--fixed-func use fixed func instead of GPU EU (default: off)
I would like to test FF mode but notice crush each time I start encoding with --fixed-func.
--------------------------- Video encoding ---------------------------
QSVEnc 6.03
R:\StaxRip-v2.10.0-x64\Apps\Encoders\QSVEnc\QSVEncC64.exe --avsdll R:\StaxRip-v2.10.0-x64\Apps\FrameServer\AviSynth\AviSynth.dll --codec hevc --quality best --bframes 14 --b-pyramid --open-gop --sar 64:45 --sao none --fixed-func --cqp 20:22:24 --qp-offset 2:4:8 -i "R:\ninjago s03e30 dvd_temp\ninjago s03e30 dvd.avs" -o "R:\ninjago s03e30 dvd_temp\ninjago s03e30 dvd_out.hevc"
--------------------------------------------------------------------------------
R:\ninjago s03e30 dvd_temp\ninjago s03e30 dvd_out.hevc
--------------------------------------------------------------------------------
HEVC SAO is not supported on current platform, disabled.
cop3.DirectBiasAdjustment value changed off -> auto by driver
cop3.GlobalMotionBiasAdjustment value changed off -> auto by driver
QSVEncC (x64) 6.03 (r2496) by rigaya, Sep 25 2021 05:36:09 (VC 1929/Win)
OS Windows 10 x64 (19043) [UTF-8]
CPU Info 11th Gen Intel Core i7-11700K @ 3.60GHz [TB: 4.00GHz] (8C/16T) <Tigerlake>
GPU Info Intel UHD Graphics 750 (32EU) 350-1300MHz [125W] (30.0.101.1191)
Media SDK QuickSyncVideo (hardware encoder) FF, 1st GPU, API v2.06
Async Depth 3 frames
Buffer Memory d3d11, 24 work buffer
Input Info AviSynth+ 3.7.0 r3382(yv12)->nv12 [AVX2], 720x576, 25/1 fps
AVSync cfr
Output HEVC(yuv420) main @ Level 3
720x576p 64:45 25.000fps (25/1fps)
Target usage 1 - best
Encode Mode Constant QP (CQP)
CQP Value I:20 P:22 B:24
QP Limit min: 10, max: 51
Trellis Auto
Ref frames 5 frames
Bframes 14 frames, B-pyramid: on
Max GOP Length 250 frames
Ext. Features WeightP WeightB QPOffset
MFXENCODE: EncodeFrameAsync error: device operation failure..
Break in task MFXENCODE: device operation failure..
encoded 2 frames, 1.53 fps, 2608.80 kbps, 0.02 MB
encode time 0:00:01, CPULoad: 0.1
frame type IDR 1
frame type I 2, total size 0.05 MB
QSVEncC.exe finished with error!
I have noticed --fixed-func sets QP Limit min: 10 when EU encoding mode sets QP Limit min: 1.
Is it important? How to enable FF encoding properly?
Yups
20th February 2022, 22:32
Try out the newest driver, if it doesn't help you need a bios update if there is a new one available. There was a bug on RKL-S which affected FF encoding: https://github.com/Intel-Media-SDK/MediaSDK/issues/2784
QP Limit is not important, it can't be changed manually anyways.
pacuro
24th February 2022, 22:44
Try out the newest driver, if it doesn't help you need a bios update
Both updated right now. Still getting error. Thanks for hint anyway. Now I know there is global problem with hevc ff encoding on rkl. Hope folks are working on that in intel.
UPDATE: Asrock prepared alpha bios with some fix which did not help. I have replaced microcode to last version "50" for A0671 and FF encoding started working. 320 fps processing 1080p video.
lt8nk
28th April 2022, 22:29
Hello everyone,
I am reading this post since few months because I would like to upgrade my actual configuration and this is the only forum where people are talking about QuickSync and its quality in hardware encoding. Here most people are talking about UHD750. But do you know if UHD710 (Pentium G7400T) is also able to obtain the same result ?
By the way, I saw a post asking how to get the VMAF value. Here is a software that calculates it : https://github.com/fifonik/FFMetrics
Thank you again for all these informations
Yups
29th April 2022, 01:14
In the worst case GT1 graphics only supports the fully fixed function mode, otherwise it should have the same quality. It's the same architecture in the end.
lt8nk
29th April 2022, 01:21
GT1 is Haswell. I don't understand. But what would be the consequences of only supporting the full fixed function ?
Sorry, even if I read the thread, I may have silly questions.
Yups
29th April 2022, 12:29
For example the slower Atom from Intel supported Quicksync but only the fixed function encoding and not the shader supported Hybrid encoding from the faster Core lineup. I don't know if this is the case for UHD710 with only 16EUs, just saying this is the worst case. You can check my posting history, I posted several VMAF tests with FF and hybrid. Sure Hybrid has a better quality, although the difference isn't really big, however the speed difference is huge. Hybrid encoding might not even make sense therefore.
lt8nk
30th April 2022, 14:29
Thank you. I found your post (https://forum.doom9.org/showpost.php?p=1940526&postcount=451). I remember it now. I read it before. Hybrid is better but FF should be enough. Right now I am encoding in x265 medium with a Ryzen 1500x and it takes A LOT of time.
I used to hesitate with the 12100T but I'll try the G7400T or the G7400 as soon as the price of the 1700 motherboards goes down a bit. But I think it will not be for a few months. And I will come back here to share my tests with you.
RanmaCanada
2nd May 2022, 01:10
Thank you. I found your post (https://forum.doom9.org/showpost.php?p=1940526&postcount=451). I remember it now. I read it before. Hybrid is better but FF should be enough. Right now I am encoding in x265 medium with a Ryzen 1500x and it takes A LOT of time.
I used to hesitate with the 12100T but I'll try the G7400T or the G7400 as soon as the price of the 1700 motherboards goes down a bit. But I think it will not be for a few months. And I will come back here to share my tests with you.
Why not just upgrade your processor to a zen3 like a 5600? You will get a significant increase in encode speed without having to spend much money on a new platform.
perrinpages
21st May 2022, 18:40
What is the quality of the Apple M1 based HW encoders? I haven't seen any real tests or comparisons done.
RanmaCanada
23rd May 2022, 02:27
What is the quality of the Apple M1 based HW encoders? I haven't seen any real tests or comparisons done.
From reading on the macforums (https://forums.macrumors.com/threads/mac-mini-m1-h-265-encoding.2269815/), it's garbage tier as it's even worse than Pascal NVENC.
perrinpages
23rd May 2022, 08:07
From reading on the macforums (https://forums.macrumors.com/threads/mac-mini-m1-h-265-encoding.2269815/), it's garbage tier as it's even worse than Pascal NVENC.
Meh. I don't trust them, no real evidence. They all seem to be using handbrake...I would rather see results using compressor.
RanmaCanada
24th May 2022, 05:15
Meh. I don't trust them, no real evidence. They all seem to be using handbrake...I would rather see results using compressor.
I doubt changing the software would make that big of a difference when we're talking about hardware encoding as the API should be pretty much the same across all platforms. But I don't use Mac's so I could quite possibly be very wrong. Take it or leave it.
Ritsuka
24th May 2022, 06:03
I remember that it was mostly tuned for psnr, and right, unless Apple hided a secret "enable better quality" switch in Compressor, HandBrake will have the same exact output.
perrinpages
24th May 2022, 06:07
I doubt changing the software would make that big of a difference when we're talking about hardware encoding as the API should be pretty much the same across all platforms. But I don't use Mac's so I could quite possibly be very wrong. Take it or leave it.
Handbrake's settings are fussy and the defaults are not sane at all. Yes, Handbrake and Compressor both use the videotoolbox framework, but we haven't seen any jsons or logs so we will never know if their conclusions on the quality are purely encoder based or if it's handbrake's filtering.
Ritsuka
24th May 2022, 06:36
VideoToolbox has got exactly zero useful settings, I highly doubt I screwed something up when I implemented VideoToolbox support in HandBrake.
But if you know how to improve it, feel free to point it out.
RanmaCanada
24th May 2022, 22:23
Handbrake's settings are fussy and the defaults are not sane at all. Yes, Handbrake and Compressor both use the videotoolbox framework, but we haven't seen any jsons or logs so we will never know if their conclusions on the quality are purely encoder based or if it's handbrake's filtering.
Well at this point, since you're being so difficult, maybe you should buy an M1 Mac and report back as it appears nothing will satisfy you unless you do it yourself. Logs are very rarely even posted here, and most people requesting said comparisons don't go to the level you are in regards to demanding proof.
Sorry but no one cares about Macs here it appears, so you'll need to do the legwork that others did in this thread.
perrinpages
25th May 2022, 21:01
VideoToolbox has got exactly zero useful settings, I highly doubt I screwed something up when I implemented VideoToolbox support in HandBrake.
But if you know how to improve it, feel free to point it out.
Your implementation is fully optimized for HDR?
perrinpages
25th May 2022, 21:23
Well at this point, since you're being so difficult, maybe you should buy an M1 Mac and report back as it appears nothing will satisfy you unless you do it yourself. Logs are very rarely even posted here, and most people requesting said comparisons don't go to the level you are in regards to demanding proof.
Sorry but no one cares about Macs here it appears, so you'll need to do the legwork that others did in this thread.
Why are you so pressed? I asked a question and you replied with a useless macrumors forum that I've already read. In my first post I asked if someone had done real comparisons. Anyone who claims 2.5 Mb/s using x265 is an acceptable bitrate for HD content cannot be trusted to make quality assessments.
BuccoBruce
26th May 2022, 04:40
On Intel both HD630 and Iris Xe the constant bitrate modes seem to work a lot better with a low bitrate budget. For now only a PSN-Y test using this (https://drive.google.com/file/d/1YX1V0SeSkYaq6Ui41vv1wcOatbLnuLSL/view?usp=sharing) sample, all of them HEVC main 8 bit 1920x1080 at ~2.5 Mbit 24 fps/240 gop. varna, you might use this sample.
/snip/
https://s10.directupload.net/images/201212/cs73djrc.png
The difference is really big, Gen9 looks much worse in the video.
Another x265/x265 CRF vs Iris Xe CQP comparison using this (https://drive.google.com/file/d/1YX1V0SeSkYaq6Ui41vv1wcOatbLnuLSL/view?usp=sharing) sample.
/snip/
Would I be mistaken in assuming this would be the easiest sample to provide AMD VCE results for, for comparison's sake? Once I figure out a way to streamline running multiple encodes and then testing those metrics, and have enough time to play with settings.
I flipped through a few pages and didn't see any AMD results - or any single test clip with as broad a range of encodes/settings (thank you Yups). I have older Polaris Radeon/Radeon Pro WX cards with VCE 3.4 (technically Lexa, Ellesmere, and "Polaris 20" - they should differ only in encode speed, I'll double check VCEEnc's output but will most likely use the Ellesmere card) and a Radeon RX 5700 with VCN 2.0. I was always under the impression that while AMD's HW H.264 encodes were kind of rubbish, H.265 was "OK".
RanmaCanada
26th May 2022, 07:38
Why are you so pressed? I asked a question and you replied with a useless macrumors forum that I've already read. In my first post I asked if someone had done real comparisons. Anyone who claims 2.5 Mb/s using x265 is an acceptable bitrate for HD content cannot be trusted to make quality assessments.
Because over the almost decade that this thread has been alive, no one has mentioned Mac's at all. Common sense would dictate that obviously no one here cares for them (they aren't mentioned in any threads that I remember), or sees them as viable. You also did not mention that you had read the forums that I posted. I guess that means you've also watched this video? (https://www.youtube.com/watch?v=RzhvdRisNZM) and his subsequent video on it.
Sorry for trying to give you some information about a subject no one here cares about.
Ritsuka
26th May 2022, 08:15
That video has got nothing to do with the hardware encoder, and like 99% of YouTube videos is probably a bunch of half-truth and comparing apples to oranges.
Anyway, like I said there should not be any difference between HandBrake a Compressor. There is some data on https://forum.handbrake.fr/viewtopic.php?p=193484#p193484 , I don't know any more recent and decent data on this topic.
Mister XY
26th May 2022, 08:46
Here are my settings for HB. I use ist for UHD sources. But it's look not bad with FHD sources. Not Perfekt, but not bad for this time.
Would I be mistaken in assuming this would be the easiest sample to provide AMD VCE results for, for comparison's sake? Once I figure out a way to streamline running multiple encodes and then testing those metrics, and have enough time to play with settings.
I flipped through a few pages and didn't see any AMD results - or any single test clip with as broad a range of encodes/settings (thank you Yups). I have older Polaris Radeon/Radeon Pro WX cards with VCE 3.4 (technically Lexa, Ellesmere, and "Polaris 20" - they should differ only in encode speed, I'll double check VCEEnc's output but will most likely use the Ellesmere card) and a Radeon RX 5700 with VCN 2.0. I was always under the impression that while AMD's HW H.264 encodes were kind of rubbish, H.265 was "OK".
I don't have AMD. Yes this sample is easy to compare.
Here are my settings for HB. I use ist for UHD sources. But it's look not bad with FHD sources. Not Perfekt, but not bad for this time.
Unfortunately Handbrake doesn't support open gop yet which improves the bitrate efficiency a bit. It is planned though.
BuccoBruce
27th May 2022, 02:30
I don't have AMD. Yes this sample is easy to compare.
Thank you! I was offering to add to your already extensive results. Looks like ffmpeg might be able to output simple PSNR/SSIM/VMAF scores, all of the pretty GUI tools I can find seem to be expensive, professional, or both, and that Intel one looks like it's been discontinued. The MSU one would be ideal if it wasn't limited to 720p...
I'll try and get results up in a few days or so, with added 10-bit output from the Navi card since it supports it.
If you upload the encoded sample I can calculate it with MSU, otherwise it's not comparable to my scores.
BuccoBruce
27th May 2022, 08:46
Here's a preliminary look at how AMD fares with the Intel test clip, and it doesn't look too good. CQP 22:24 is VCEEncC's default, and you need RDNA2 for B-frames, but I'm not sure they'd help. Pre-analysis isn't supported for HEVC. There's a "Pre-Encode" rate control feature, but I'm not sure it'll help either. So here's "Default" for now, with just changing the preset from balanced to either slow or fast.
The "higher end" GPU produced larger files across the board, with the same exact settings. "Balanced" and "Slow" seem to produce identical output on both cards!
Quality/GPU VMAF PSNR SSIM Speed Bitrate Size
Bal/WX2100 95.3114 42.5086 0.9782 147.43 fps 7096.70 kbps 412.70 MB
Bal/WX5100 96.1465 42.9197 0.9785 121.80 fps 7963.64 kbps 463.12 MB
Fast/WX2100 95.2958 42.5546 0.9783 152.93 fps 7158.12 kbps 416.28 MB
Fast/WX5100 96.2284 43.1607 0.9798 122.83 fps 7890.48 kbps 458.86 MB
Slow/WX2100 95.3114 42.5086 0.9782 148.30 fps 7096.70 kbps 412.70 MB
Slow/WX5100 96.1465 42.9197 0.9785 121.26 fps 7963.64 kbps 463.12 MB
https://i.postimg.cc/SNKSL2X3/allVMAF.png (https://postimages.org/)
https://i.postimg.cc/gkBYgz8J/PSNR.png (https://postimages.org/)
https://i.postimg.cc/bwDqM79R/VMAF.png (https://postimages.org/)
"Lexa PRO GL" in the Radeon Pro WX 2100 is a Polaris 12 (https://www.techpowerup.com/gpu-specs/amd-lexa.g806) or "Lexa" GPU, and isn't very different from the RX 550. "Polaris 10 PRO GL (https://www.techpowerup.com/gpu-specs/amd-ellesmere.g795)" in the WX 5100 is Polaris 10, a single slot, toned down Radeon RX 480. They're both GCN 4.0 GPUs with VCE 3.4. I didn't expect them to perform differently, unless someone at AMD thinks ~15% is within the margin of error for a non-deterministic encode. I wouldn't be surprised if Polaris 20 XL (https://www.techpowerup.com/gpu-specs/radeon-rx-570.c2939) gave entirely different results too, even though it too reports as codename "Ellesmere" and also has VCE 3.4. Might try it later.
-------------------------------------------------------------------------------------
VCEEnc (x64) 7.00 (r1066) by rigaya, Apr 30 2022 18:34:01 (VC 1931/Win)
GPU: Radeon Pro WX 5100, AMF Runtime 1.4.23 / SDK 1.4.24
Input Info: AviSynth+ 3.7.2 r3661(yv12)->nv12 [AVX], 1920x1080, 24/1 fps
Vpp Filters copyHtoD
Output: H.265/HEVC main @ Level 4 (main tier)
1920x1080p 0:0 24.000fps (24/1fps)
Quality: slow
CQP: I:22, P:24
VBV Bufsize: 12000 kbps
Bframes: 0 frames
Pre Analysis: off
Motion Est: Q-pel
Slices: 1
GOP Len: 240 frames
VUI: matrix:bt709,colorprim:bt709,transfer:bt709
Others: deblock
The only change between the three runs was "Quality" which was either slow, balanced, or fast. Everything else was left at its default settings (for now). I might try another run with PE enabled, 2.5 Mbps VBR. And yes, I quadruple checked that the "Slow" and "Balanced" runs weren't mislabeled, I had log files from running the encodes. FFMetrics keeps crashing, although I did ask it to save CSV frame data if anyone cares for it. :logfile: I just can't play with the plots directly within it anymore without re-running an hour's worth of tests and I doubt it's worth opening Excel for.
Doesn't seem like RDNA2 fares much better either. (https://codecalamity.com/hardware-encoding-4k-hdr10-videos/)
BuccoBruce
27th May 2022, 08:55
If you upload the encoded sample I can calculate it with MSU, otherwise it's not comparable to my scores.
Whoops, I just saw that while I was writing that post. Here you go! The clips are still uploading https://mega.nz/folder/SzwDlR5D#SbfBKtLuD3rOQPjTUZ0P8g
The quality is quite poor considering it's converted into 8 Mbit. Bframes should make a noticeably difference, although the difference appears so big they need much more improvements.
BuccoBruce
27th May 2022, 15:54
The quality is quite poor considering it's converted into 8 Mbit. Bframes should make a noticeably difference, although the difference appears so big they need much more improvements.
Yeah, I'm fairly disappointed in how it turned out. I thought AMD VCE/VCN being rubbish was just a meme but I guess it's true. Almost tempted to dig a Sandy/Ivy Bridge CPU with QSV out and compare it to 8 Mbps AVC on that thing.
Yeah, I'm fairly disappointed in how it turned out. I thought AMD VCE/VCN being rubbish was just a meme but I guess it's true. Almost tempted to dig a Sandy/Ivy Bridge CPU with QSV out and compare it to 8 Mbps AVC on that thing.
I have an older Kabylake CPU, I'm sure it's much better at 8 Mbps using Quicksync. Actually 8 Mbps is too high on a half decent solution for this sample, that's why I went down to 2.5 Mbps in my testing.
PatchWorKs
1st June 2022, 06:51
Here are my settings for HB. I use ist for UHD sources. But it's look not bad with FHD sources. Not Perfekt, but not bad for this time.
I'm testing very low QSV encoding parameters (35-40) for FHD source and I strongly suggest to use rigaya's QSVEnc instead: mutch better results.
You can easily test it through FastFlix (https://fastflix.org/):
https://raw.githubusercontent.com/cdgriffith/FastFlix/master/docs/gui_preview.png
lt8nk
25th June 2022, 17:22
Why not just upgrade your processor to a zen3 like a 5600? You will get a significant increase in encode speed without having to spend much money on a new platform.
Sorry for the late answer. I thought I will receive notifications for new posts...
In fact, I hadn't given all the info so as not to overload my post. The 1500x is on the classic PC. But I also have an old Synology NAS of almost 10 years, and a mini windows PC based on Atom x5-z8350 (which doesn't allow to encode in h265, only in reading), connected to the TV. This mini PC also hosts a Jellyfin server, a TV card (so NextPVR) and a bunch of other servers. Honestly, even if I tried to optimize everything, this little Atom surprises me. As I'm running out of space with 4TB in Raid 1, I was thinking of making a big change to replace both the mini PC and the NAS in a single device that would allow me for example to encode TV recordings in H265 automatically, and to increase space with a kind of Raid 5 with SnapRaid.
I think I will wait beginning of september before buying it. Actually, it would cost me about 240 € for the whole without the new HDD.
Now you know everything.
RanmaCanada
1st July 2022, 01:40
Sorry for the late answer. I thought I will receive notifications for new posts...
In fact, I hadn't given all the info so as not to overload my post. The 1500x is on the classic PC. But I also have an old Synology NAS of almost 10 years, and a mini windows PC based on Atom x5-z8350 (which doesn't allow to encode in h265, only in reading), connected to the TV. This mini PC also hosts a Jellyfin server, a TV card (so NextPVR) and a bunch of other servers. Honestly, even if I tried to optimize everything, this little Atom surprises me. As I'm running out of space with 4TB in Raid 1, I was thinking of making a big change to replace both the mini PC and the NAS in a single device that would allow me for example to encode TV recordings in H265 automatically, and to increase space with a kind of Raid 5 with SnapRaid.
I think I will wait beginning of september before buying it. Actually, it would cost me about 240 € for the whole without the new HDD.
Now you know everything.
Ah makes sense. What I would do in your situation then is make an UNRAID server for storage and use the docker functions to host your media server. As long as you have an Intel CPU that supports quicksync, you're pretty much good to go. An i3 or even a Pentium Gold should more than do the job. I would personally pick at least 10th gen, but if you can swing it, 12th gen would be superior.
Good luck!
lt8nk
6th July 2022, 12:14
At the beginning I also thought about using UNRAID but I am a bit afraid of compatibility of my USB TV card and external sound card with passthrough. I am even not sure of compatibility with Windows 11... Concerning CPU, I will definetely go for 12th generation. But I still hesistate between Pentium G7400 and i3-12100T.
butterw2
6th July 2022, 18:05
Dual-core G7400 with UHD-710 (av1 hw-dec and h265 hw-enc) seems interesting for a low power machine running on stock cooler (46W TDP, no boost). I would like to see some quick sync testing vs 12100.
For software encoding (ex: x264) with only 4 threads, it will not be able to compete with a 12100. But from what I figure it does beat old quadcore i5s (ex: i5-6500) in terms of performance.
Yups
28th July 2022, 19:23
Here's a preliminary look at how AMD fares with the Intel test clip, and it doesn't look too good. CQP 22:24 is VCEEncC's default, and you need RDNA2 for B-frames, but I'm not sure they'd help. Pre-analysis isn't supported for HEVC.
B-frames would help, how much no idea. There are some b-frames tests for H264 but not for h265. On Intel Iris Xe CQP is very effective with many bframes. I tried what bitrate I need to match or beat your WX5100 scores with a bitrate of almost 8 Mbit. I need around 4.5 Mbit.
https://abload.de/img/irisxeo5jak.png
Settings
--avhw --codec hevc --quality best --profile main --ctu 64 --bframes 14 --gop-len 240 --b-pyramid --open-gop --no-tskip --sao none --cqp 19:22:24
https://drive.google.com/file/d/1Jp3Ns8-XsQB47NDsxKPcqCyGz_KdAIUs/view?usp=sharing
ReinerSchweinlin
10th August 2022, 11:20
Dual-core G7400 with UHD-710 (av1 hw-dec and h265 hw-enc) seems interesting for a low power machine running on stock cooler (46W TDP, no boost). I would like to see some quick sync testing vs 12100.
For software encoding (ex: x264) with only 4 threads, it will not be able to compete with a 12100. But from what I figure it does beat old quadcore i5s (ex: i5-6500) in terms of performance.
The G7400 looks interesting as a cheap HW-Encoding machine. Do we have any info whether the Encoder in the smaller Intel CPUs are the same as the above mentioned XE ones that did well?
I am still looking for some power efficient System that could do two things:
x265 encode with a very good power efficiency (speed is not important, efficiency is king - it can run for days)
HW-Encode with the best quality efficiency..
Notebooks or NUCs with I5-1135.. CPUs came to mind. Do you guys know of any other cheap candidates?
Intel Core i3-1210U seems very promising in terms of power efficiency, but I guess its too new, can´t find any systems with it....
perrinpages
10th August 2022, 22:55
The G7400 looks interesting as a cheap HW-Encoding machine. Do we have any info whether the Encoder in the smaller Intel CPUs are the same as the above mentioned XE ones that did well?
I am still looking for some power efficient System that could do two things:
x265 encode with a very good power efficiency (speed is not important, efficiency is king - it can run for days)
HW-Encode with the best quality efficiency..
Notebooks or NUCs with I5-1135.. CPUs came to mind. Do you guys know of any other cheap candidates?
Intel Core i3-1210U seems very promising in terms of power efficiency, but I guess its too new, can´t find any systems with it....
Framework is selling their mainboards with an i5-1240P for $449, but you need a case to house it in and the connectivity situation might not be ideal. It is much cheaper than any nuc I've seen that's actually in stock...
https://frame.work/products/mainboard-12th-gen-intel-core?v=FRANGACP04
Mister XY
11th August 2022, 04:30
B-frames would help, how much no idea. There are some b-frames tests for H264 but not for h265. On Intel Iris Xe CQP is very effective with many bframes. I tried what bitrate I need to match or beat your WX5100 scores with a bitrate of almost 8 Mbit. I need around 4.5 Mbit.
https://abload.de/img/irisxeo5jak.png
Settings
--avhw --codec hevc --quality best --profile main --ctu 64 --bframes 14 --gop-len 240 --b-pyramid --open-gop --no-tskip --sao none --cqp 19:22:24
https://drive.google.com/file/d/1Jp3Ns8-XsQB47NDsxKPcqCyGz_KdAIUs/view?usp=sharing
You can e.g. from minute 6:10 wonderfully recognize the artifacts or the block formation.
Otherwise the picture looks very good for the size.
Yups
12th August 2022, 17:19
You can e.g. from minute 6:10 wonderfully recognize the artifacts or the block formation.
Otherwise the picture looks very good for the size.
Keep in mind it's a low bitrate. Did you try --no-tskip --sao none? I think it's better, the size and VMAF score are definitely.
benwaggoner
13th August 2022, 01:06
Keep in mind it's a low bitrate. Did you try --no-tskip --sao none? I think it's better, the size and VMAF score are definitely.
Have you found a source/example where --no-tskip actually improves quality? Not that --tskip is enabled in any of the presets.
ReinerSchweinlin
16th August 2022, 13:23
Framework is selling their mainboards with an i5-1240P for $449, but you need a case to house it in and the connectivity situation might not be ideal. It is much cheaper than any nuc I've seen that's actually in stock...
https://frame.work/products/mainboard-12th-gen-intel-core?v=FRANGACP04
Thanx, thats a good price for the given performance, indeed.
IF the smaller CPUs with built in XE or UHD7xx provide the same quality, even a smaller system could be an option (for the HW encoding part...)
RanmaCanada
19th August 2022, 21:55
Thanx, thats a good price for the given performance, indeed.
IF the smaller CPUs with built in XE or UHD7xx provide the same quality, even a smaller system could be an option (for the HW encoding part...)
AFAIK the asics are the same for given generations as per the wikipedia (https://en.wikipedia.org/wiki/Intel_Quick_Sync_Video) entry for Intel Quicksync. Meaning Alder Lake Mobile should have the same ASICS as the desktop variant as they are not separated between mobile and desktop.
lt8nk
25th August 2022, 15:26
Maybe the Intel ARC A310 card could be a good investment if associated with a low power cpu. As it can also encode in AV1, I hope that the HW encoding part for H265 would at minimum be the same (if not better) than 12th gen . But the power consumption of the card has to be low.
perrinpages
27th August 2022, 01:51
Thanx, thats a good price for the given performance, indeed.
IF the smaller CPUs with built in XE or UHD7xx provide the same quality, even a smaller system could be an option (for the HW encoding part...)
Intel's processors generally share the same graphics architecture for a given CPU generation...the only difference would be performance for non-fixed-function encoding (dependent on the number of EUs in the SKU).
https://www.anandtech.com/show/17278/intel-launches-alder-lake-u-and-p-series-processors-ultraportable-laptops-coming-in-march
Yups
27th August 2022, 17:00
Have you found a source/example where --no-tskip actually improves quality? Not that --tskip is enabled in any of the presets.
I've tried two samples and in both of them it was slightly better according to the VMAF score. It seems to be enabled in all Quicksync presets by the way.
I hope I can try out Alchemist soon. HEVC encoder seems to be improved slightly. There is only one downside, they removed the Hybrid mode. Only the full fixed function mode is supported.
ReinerSchweinlin
29th August 2022, 15:26
AFAIK the asics are the same for given generations as per the wikipedia (https://en.wikipedia.org/wiki/Intel_Quick_Sync_Video) entry for Intel Quicksync. Meaning Alder Lake Mobile should have the same ASICS as the desktop variant as they are not separated between mobile and desktop.
Yes, I assume the same, but I am not sure...
I assume that the encoding performance of h265 has not changed after GEN11, so assuming the generation of the first XE`s offsprings like Pentium Gold 7505 SHOULD be cheap HW-Encoding horses...
RanmaCanada
29th August 2022, 17:30
Yes, I assume the same, but I am not sure...
I assume that the encoding performance of h265 has not changed after GEN11, so assuming the generation of the first XE`s offsprings like Pentium Gold 7505 SHOULD be cheap HW-Encoding horses...
If you want to go to the extreme, Framework is selling their laptop motherboards with Adler Lake processors, and they have created a bios that allows it to work without a battery. Meaning that with a 3d printed case, you can have an extremely low power hardware encoding machine.
perrinpages
30th August 2022, 16:32
Thanx, thats a good price for the given performance, indeed.
IF the smaller CPUs with built in XE or UHD7xx provide the same quality, even a smaller system could be an option (for the HW encoding part...)
You might be interested in the LattePanda 3 Delta. It has a low power 11th gen cpu (N5105).
https://www.lattepanda.com/lattepanda-3-delta
You probably won't find readily available low power 12th gen for a while. The low end is on a slower refresh cycle.
benwaggoner
30th August 2022, 18:24
I've tried two samples and in both of them it was slightly better according to the VMAF score. It seems to be enabled in all Quicksync presets by the way.
I wouldn't expect VMAF to be a particularly acute metrics for comparing those differences. It can really bias towards smoothing out flatter areas.
I hope I can try out Alchemist soon. HEVC encoder seems to be improved slightly. There is only one downside, they removed the Hybrid mode. Only the full fixed function mode is supported.
Last I looked Hybrid didn't really have a compelling quality @ speed advantage over the faster x265 presets.
Yups
31st August 2022, 11:24
Last I looked Hybrid didn't really have a compelling quality @ speed advantage over the faster x265 presets.
I'm talking about Quicksync. Hybrid encoding mode has a higher quality than pure FF, although the pure FF from Arc seems to be better than Hybrid on Iris Xe. I will test it out soon.
CQP with many bframes should have a superior quality to the faster x265 presets by the way. And sure hybrid runs relatively slow on a iGPU due to low clock speed and shader count.
ReinerSchweinlin
14th September 2022, 18:18
You might be interested in the LattePanda 3 Delta. It has a low power 11th gen cpu (N5105).
https://www.lattepanda.com/lattepanda-3-delta
You probably won't find readily available low power 12th gen for a while. The low end is on a slower refresh cycle.
Thanx. I hust looked up prices. The LattePanda is around the same ballpark as a 11GEN full laptop on sale with i5-1135G7, 8GBRAM, SSD, etc... Around 350 euros seems to be the entrypoint at the moment for at least GEN11 Intel GPU full working systems.
Addition: The N5105 seems more like a GEN10 CPU... The internal GPU is one generation older than the XE ones, we talked about intensively. Its hard to get a grasp of which quality can be achieved with which version, I was under the impression that at least "XE Generation" was needed...
If the N5105 is a candidate, there is a very low cots Intel NUC :)
Yups
17th September 2022, 16:25
N5105 comes with a Gen11 iGPU. The hybrid encoder should be the same, low power is one generation behind. Also this Atom SKUs might only support low power mode. In this case it's much worse.
Intel upgraded their HEVC encoder with Icelake (Gen11): https://forum.doom9.org/showpost.php?p=1860006&postcount=317
With Tigerlake there is another upgrade to the HEVC FF encoder (bframes+bpyramid support).
I have something new to test:
https://abload.de/img/a380m2dl1.png
Quick comparison versus Iris Xe by using this (https://drive.google.com/file/d/1YX1V0SeSkYaq6Ui41vv1wcOatbLnuLSL/view?usp=sharing) sample and CQP highest quality. As already mentioned Intel Arc only supports low power mode, there is no hybrid mode.
Intel Demo Clip VMAF speed bitrate
Quicksync H265
A380 CQP FF best 92.26 309 fps 2351 kbit
Iris Xe CQP best 91.96 67 fps 2312 kbit
Iris Xe CQP FF best 91.60 249 fps 2382 kbit
Exact same settings for both:
--avhw --codec hevc --quality best --profile main --bframes 14 --gop-len 120 --b-pyramid --open-gop --weightb --weightp --no-tskip --d3d11 --sao none --fallback-rc --fixed-func --cqp 23:24:30
Arc doesn't need Hybrid mode, it's as good or slightly better than Iris Xe hybrid with the advantage of the much higher performance.
Kurtnoise
17th September 2022, 17:12
Could you try the AV1 hardware encoder please ?
Yups
17th September 2022, 23:09
Could you try the AV1 hardware encoder please ?
I haven't tested much but in this video and this bitrate AV1 is not as good as HEVC. AV1 reviews from rigaya or THG should be accurate. The software seems immature at the moment, they may improve it over time. Some things doesn't work like mbbrc or extbrc and CQP even on lowest quantizer parameter can't go to low bitrates, so I couldn't even properly compare CQP H265 to CQP AV1. However a tweaked H265 Quicksync with 14 bframes CQP is hard to beat, even with software x265 it would require the slower presets to really make a difference.
Intel added bframes+bpyramid support for the low power mode on h264, previously it was only supported in hybrid mode. This is a very important improvement for game streaming on Arc dGPUs since hybrid mode is no real option if it's the primary GPU because it lowers the gaming performance quite a bit due to the shader usage. The fully fixed H264 encoder is on the same level as hybrid on Iris Xe now. There is no Lookahead mode though, it required GPU support.
Of course AV1 is much better than h264, twitch should support AV1 as soon as possible.
easyfab
18th September 2022, 19:18
Intel Demo Clip VMAF speed bitrate
Quicksync H265
A380 CQP FF best 92.26 309 fps 2351 kbit
Iris Xe CQP best 91.96 249 fps 2312 kbit
Iris Xe CQP FF best 91.60 67 fps 2382 kbit
Are the speed in good order for Iris Xe ?
A380 speed isn't so imprerssive.
Yups
18th September 2022, 22:27
Intel Demo Clip VMAF speed bitrate
Quicksync H265
A380 CQP FF best 92.26 309 fps 2351 kbit
Iris Xe CQP best 91.96 249 fps 2312 kbit
Iris Xe CQP FF best 91.60 67 fps 2382 kbit
Are the speed in good order for Iris Xe ?
A380 speed isn't so imprerssive.
I have a PCIe 3 system without rBAR, I cannot rule out that it effects the encoding performance. The speed is impressive imho. H264, H265, AV1 all very similar in performance. In the sample above it goes over 500 fps in balanced preset with a small quality loss.
Iris Xe was very fast already in low power mode: https://forum.doom9.org/showpost.php?p=1940526&postcount=451
It's faster than a RTX 3090 Ti: https://www.tomshardware.com/reviews/intel-arc-a380-review/5
ReinerSchweinlin
24th September 2022, 10:11
Intel Demo Clip VMAF speed bitrate
Quicksync H265
A380 CQP FF best 92.26 309 fps 2351 kbit
Iris Xe CQP best 91.96 249 fps 2312 kbit
Iris Xe CQP FF best 91.60 67 fps 2382 kbit
Are the speed in good order for Iris Xe ?
A380 speed isn't so imprerssive.
309 fps is quite impressive, most other cards only get this fast when run in "Fast" settings with much lower quality...
Yups
3rd October 2022, 00:41
I finished another test including VP9, H264, H265, AV1 CBR and also ICQ+CQP from my A380. For this test I'm using QSVEnc 7.21 and I decided to use out of the box settings, I only changed the gop length to 120 and bframes/gop-ref-dist depending on the codec and bitrate mode. gop-ref-dist 8 for AV1, bframes 5 for H265 CBR etc. H265 CQP is the only exception, I tweaked it (bframes 14, open gop, SAO+tskip off), basically this is the max possible quality. However, this optimized CQP quality isn't reachable for most tools, Handbrake doesn't even support open gop.
ToS_1920x800_xdither.y4m (https://forum.doom9.org/showthread.php?p=1853595#post1853595) PSNR SSIM VMAF speed bitrate
Quicksync H265
Intel Arc A380 CBR H264 38.965 0.9636 91.98 210 fps 1898 Kbit
Intel Arc A380 CBR VP9 38.430 0.9611 92.67 148 fps 1901 Kbit
Intel Arc A380 CBR H265 40.461 0.9710 94.92 208 fps 1898 Kbit
Intel Arc A380 CBR AV1 40.358 0.9713 94.24 205 fps 1899 Kbit
Intel Arc A380 ICQ AV1 41.570 0.9736 94.29 215 fps 1900 Kbit
Intel Arc A380 CQP AV1 41.584 0.9740 94.18 219 fps 1895 Kbit
Intel Arc A380 CQP H265 (optimized) 42.163 0.9756 95.59 218 fps 1909 Kbit
Iris Xe Hybrid CBR H264 37.625 0.9622 91.42 206 fps 1897 Kbit
Iris Xe FF CBR H264 37.441 0.9549 90.24 284 fps 1897 Kbit
Iris Xe FF CBR H265 40.166 0.9700 94.00 285 fps 1898 Kbit
VMAF scores are odd, CQP/ICQ difference to CBR is usually a lot higher, PSNR+SSIM differ a lot more.
Hardware decoding doesn't work on this RAW sample and that's why the A380 is relatively slow and missing rBAR doesn't help.
Intel improved H265 slightly over Iris Xe and apparently they improved H264 quite a bit as well which surprised me. I mean even Iris Xe was really good on h264 compared to Nvidia.
benwaggoner
3rd October 2022, 00:56
VMAF scores are odd, CQP/ICQ difference to CBR is usually a lot higher, PSNR+SSIM differ a lot more.
In terms of perceptual correlation, VMAF > SSIM > PSNR (although even VMAF still has mediocre correlation).
But for more than a few seconds of video, taking the mean of the scores of individual frames is only broadly indicative of overall quality. A video that's a stable VMAF of 80 is a lot more perceptually pleasant than one that oscillates between 60 and 100. Visual stability and interframe coherency are a huge deal, although hard to capture in simple metrics.
This is a particularly big deal for CBR versus VBR comparisons. In essence, CBR maintains a stable bitrate by varying quality, and VBR maintains a stable quality by varying bitrate. Even if CBR yields the same bitrate and average VMAF as a VBR, the VBR will often look markedly better overall.
The closing gap between CBR and VBR could come from better CBR rate control with a longer lookahead so the encoder has more flexibility to optimally allocate bits over a several second period.
This can have a big impact. Try comparing two identical CBR x265 encodes, one where --rc-lookahead=keyint, and another where --rc-lookahead=1. The former will generally look quite a bit better due to better rate control. Or compare 1-pass CBR to 2-pass CBR.
Yups
4th October 2022, 20:03
Intel Flex 140 with 4 decoder+encoder, twice as much as A380 :D
https://abload.de/img/intel-innovation-flexrhciv.jpg
https://www.nextplatform.com/2022/10/04/different-gpu-horses-for-different-datacenter-courses/
rwill
4th October 2022, 21:10
8/16.8 FLOPs is a bit low. Still faster than Zuse Z3 but the Z3 is somewhat old.
Blue_MiSfit
5th October 2022, 01:04
Quite promising for dense and power efficient live encoding!
TEB
7th October 2022, 11:29
Is the Videopipeline on the A770 the same as the FLEX 170 ?
Yups
7th October 2022, 22:41
Is the Videopipeline on the A770 the same as the FLEX 170 ?
By the looks of it they have the same media unit. DG2/ATSM have the same encoding/decoding/processing features, they are in the same table: https://github.com/intel/media-driver/blob/9de86486b373575e1ab1270443a52104edeab61a/docs/media_features.md
Yups
8th October 2022, 21:13
Actually there is a github page with lots of informations and tests from Flex series: https://github.com/intel/media-delivery/blob/ddfbad8bc5d3134b7d2ee2ae17d8aacb226a717b/doc/benchmarks/intel-data-center-gpu-flex-series/intel-data-center-gpu-flex-series.rst
Yups
10th October 2022, 15:45
Currently testing Sol Levante on Arc A380 in CBR mode on different bitrates. AV1 is a lot better than HEVC in this case, at least in these very low bitrate settings I have tried.
https://abload.de/img/levantetrflg.png
benwaggoner
10th October 2022, 17:11
Currently testing Sol Levante on Arc A380 in CBR mode on different bitrates. AV1 is a lot better than HEVC in this case, at least in these very low bitrate settings I have tried.
Have you done a visual comparison? Libaom received specific VMAF tuning, so encoders based on it tend to yield somewhat higher VMAF scores than a double-blind comparison would yield.
That said, AV1's strength has primarily been in <<1 Mbps bitrates. I've not found example much >500 Kbps where AV1 offers a consistent perceptual improvement over well-tuned HEVC.
Yups
10th October 2022, 17:27
Have you done a visual comparison? Libaom received specific VMAF tuning, so encoders based on it tend to yield somewhat higher VMAF scores than a double-blind comparison would yield.
That said, AV1's strength has primarily been in <<1 Mbps bitrates. I've not found example much >500 Kbps where AV1 offers a consistent perceptual improvement over well-tuned HEVC.
Yes I made some visual comparisons and it looks noticeably better on AV1, less blocky. This is basically default CBR, nothing special.
AV1= https://abload.de/img/av1n3de3.png
https://abload.de/img/av1_2qpcqf.png
HEVC= https://abload.de/img/hevcwlcyc.png
https://abload.de/img/hevc_20qcm0.png
Yups
11th October 2022, 18:37
I could improve HEVC with open gop and further improve with VBR, it comes closer to AV1 but the gap is still relatively big. I also improved AV1 slightly with gop-ref-dist 4 instead of 8. AV1 VBR isn't better and CQP/ICQ is worse for this test case. AV1 QSV is better than x265 slow ABR in this test.
SolLevante_SDRv2_1080p24_8bit (https://forum.doom9.org/showthread.php?p=1853595#post1853595)
Quicksync QSVEnc PSNR SSIM VMAF VQM Bitrate
Arc A380 CBR H264 32.763 0.9018 63.52 1.956 1380 Kbit
Arc A380 CBR VP9 33.052 0.9032 68.43 1.930 1375 Kbit
Arc A380 CBR H265 33.996 0.9173 71.66 1.676 1384 Kbit
Arc A380 CBR H265 open gop 33.997 0.9174 71.86 1.672 1384 Kbit
Arc A380 VBR H265 open gop 34.169 0.9202 73.17 1.630 1381 Kbit
Arc A380 CBR AV1 gop-ref-dist 8 34.328 0.9221 74.21 1.578 1379 Kbit
Arc A380 CBR AV1 gop-ref-dist 4 34.416 0.9223 74.56 1.572 1383 Kbit
x264 slow ABR Handbrake nightly 33.170 0.9086 64.62 2.083 1381 Kbit
x265 slow ABR Handbrake nightly 34.361 0.9187 72.48 1.729 1379 Kbit
tormento
12th October 2022, 15:43
I could improve HEVC with open gop
Can we say that ARC is encoding HEVC better than x265 (at least slow)?
excellentswordfight
12th October 2022, 15:46
Can we say that ARC is encoding HEVC better than x265 (at least slow)?
I doubt that its true as a general statement. I did a test recently with hevc on a HD770 (xe igpu) and i preferred x264 over it in that case.
I've not use vmaf that much, but isnt low 70 rather low quality either way?
rwill
12th October 2022, 16:45
I doubt that its true as a general statement. I did a test recently with hevc on a HD770 (xe igpu) and i preferred x264 over it in that case.
I've not use vmaf that much, but isnt low 70 rather low quality either way?
70 is eye cancer territory. And the sequence has around 28% black/white text only credits so VMAF for normal scenes is probably even lower. The numbers Yups published here cannot and should not be used to compare the different encoders realistically. Besides the unrealistic crap quality you also cannot compare CBR and ABR encodes, thats not how the objective metrics he used work. The results are useless.
Using some clean Tears of Steel 1080p source and target ABR rates from 1500k to 5000k, in maybe 4 or 5 steps, with a tune for PSNR and using global PSNR for comparison might give a better indication of the encoder implementations compression efficiency.
Yups
12th October 2022, 18:54
Can we say that ARC is encoding HEVC better than x265 (at least slow)?
No I wouldn't say. To make a statement like this it requires much more samples and much more variation in bitrate, video samples etc. I would expect that x265 slow is better in most cases because it's more stable. I may post other videos where x265 slow will be better.
I doubt that its true as a general statement. I did a test recently with hevc on a HD770 (xe igpu) and i preferred x264 over it in that case.
I've not use vmaf that much, but isnt low 70 rather low quality either way?
What software did you use? The software is a big problem for hardware encoder. Most of the software is very basic without any options and some software like Handbrake are using poor settings. For example Handbrake is using a very low gop size of 1s, on a 24fps video like this the gop size of 24 is low when x265 default max gop is 250 afaik. It can be changed in Handbrake with custom settings but nobody does it. Open gop isn't even supported in Handbrake, for x265 it's enabled on default afaik.
70 is eye cancer territory. And the sequence has around 28% black/white text only credits so VMAF for normal scenes is probably even lower. The numbers Yups published here cannot and should not be used to compare the different encoders realistically. Besides the unrealistic crap quality you also cannot compare CBR and ABR encodes, thats not how the objective metrics he used work. The results are useless.
Using some clean Tears of Steel 1080p source and target ABR rates from 1500k to 5000k, in maybe 4 or 5 steps, with a tune for PSNR and using global PSNR for comparison might give a better indication of the encoder implementations compression efficiency.
I can post a 5000 kbit bitrate result with tuned PSNR. From what I can see it doesn't change much though.
Yups
12th October 2022, 20:02
Here the result. It's not fundamentally different to the low bitrate results.
SolLevante_SDRv2_1080p24_8bit (https://forum.doom9.org/showthread.php?p=1853595#post1853595)
Quicksync QSVEnc PSNR SSIM VMAF VQM Bitrate
Arc A380 VBR H265 open gop 37.546 0.9570 90.98 0.754 5256 Kbit
Arc A380 VBR AV1 gop-ref-dist 4 38.259 0.9602 91.67 0.724 5255 Kbit
x265 slow ABR Handbrake PSNR tune 38.270 0.9568 90.68 0.774 5262 Kbit
tormento
12th October 2022, 22:43
Here the result.
Can you link the resulting videos?
Are you using the best parameters on ARC or "medium/default" ones?
Yups
13th October 2022, 12:38
Can you link the resulting videos?
Are you using the best parameters on ARC or "medium/default" ones?
I use best preset/TU1. Command line looks like this: --avsw --codec hevc --output-depth 8 --quality best --gop-len 120 --vbr 6500 --bframes 5 --open-gop
Here the samples: https://drive.google.com/file/d/1BMAEsC10KPdD6eTuElq5upWGi45iVifS/view?usp=sharing
edit: Davinci Resolve Quicksync AV1 is broken, it stutters. Maybe same issues as in Handbrake: https://github.com/HandBrake/HandBrake/issues/4570
I guess they are not even aware of this lol. Also it's slow because hardware decoding doesn't work. keyframe 1 on default poor quality, it can be changed though. Gop ref dist can't be changed, this is a problem assuming they are using dist 1. As I said the software is a big problem. edit2: frame ordering is the culprit for Davinci Resolve. It's enabled by default, the stutter is gone without frame ordering.
outhud
14th October 2022, 08:38
I'm close to buying a cheap Intel mini PC for realtime 1080p H265 encoding.
I'm trying to decide between Gemini Lake N4020 and newer Jasper Lake N5105. Both have Intel UHD Graphics.
Considering both can do realtime 1080p H265 HW encodes, is there any reason to think that there would be a quality improvement using the newer processor?
tormento
14th October 2022, 13:37
Here the samples
Can you please tell us the encoding speed?
Can you increase/decrease quality factor to see how the samples scale?
Yups
14th October 2022, 17:34
I'm close to buying a cheap Intel mini PC for realtime 1080p H265 encoding.
I'm trying to decide between Gemini Lake N4020 and newer Jasper Lake N5105. Both have Intel UHD Graphics.
Considering both can do realtime 1080p H265 HW encodes, is there any reason to think that there would be a quality improvement using the newer processor?
Gemini Lake N4020 features a Gen9 GPU and Jasper Lake Gen11. Gen11 is better of course, it comes with an upgraded HEVC encoder.
The problem might be Gen9/Gen11 doesn't support bframes in fully fixed function mode. Icelake-U didn't have bframes support in FF which is Gen11 based. Unless Jasper Lake differs this isn't optimal because without bframes the quality will drop quite a bit, same for Gemini Lake. And maybe the hybrid mode won't work on Atom based SKUs anyways. If you can wait the upcoming ADL-N is better.
Can you please tell us the encoding speed?
Can you increase/decrease quality factor to see how the samples scale?
On this sample about 160 fps for AV1/HEVC QSV. This is a RAW sample, hardware decoding doesn't work there. Balanced preset isn't faster therefore, the encoding engine is not the limiting factor. I haven't checked x265 speed and I only have a quadcore CPU, this isn't too meaningful.
tormento
15th October 2022, 10:19
On this sample about 160 fps for AV1/HEVC QSV.
:eek:
Will the different Arc cards feature the same encoding speed (such as nvidia) or it will be dependent to the number of cores?
Yups
15th October 2022, 13:16
:eek:
Will the different Arc cards feature the same encoding speed (such as nvidia) or it will be dependent to the number of cores?
They have the same encoder/decoder. It can differ depending on the clock speed and memory speed differences if it isn't a fixed function encoder+decode. I would assume the speed is similar if it's a fully fixed function encode. There is a good test from techgage (https://techgage.com/article/intel-arc-a750-a770-workstation-review/). Similar speed in Handbrake (https://techgage.com/viewimg/?img=https://techgage.com/wp-content/uploads/2022/10/Intel-Arc-A770-and-A750-Performance-HandBrake-Transcode.jpg&desc=Intel%20Arc%20A770%20and%20A750%20Performance%20(HandBrake%20Transcode)) which I think is representative for a fully fixed function transcode. There are bigger differences in Vegas Pro and Adobe, hard to say why. ProRes encode in Adobe is certainly without hardware decoding, so maybe the faster memory from the bigger Arc cards helps. The other Adobe encode is with heavy effects, I guess it uses the GPU for it, in such a case the difference will be big.
outhud
15th October 2022, 19:38
Gemini Lake N4020 features a Gen9 GPU and Jasper Lake Gen11. Gen11 is better of course, it comes with an upgraded HEVC encoder.
Thanks for this info! I see the encoder has been upgraded to support higher resolutions, but I take from your post that it also has been upgraded to give better quality at 1080p?
If the 1080p quality is the same, then the N4020 will be my best choice because it beats N5105 on price.
1080p H265 is my only use case for the CPU, I'll be running it as a headless Linux server running only tvheadend to transcode digital TV to H265 in realtime.
Yups
15th October 2022, 22:29
Thanks for this info! I see the encoder has been upgraded to support higher resolutions, but I take from your post that it also has been upgraded to give better quality at 1080p?
If the 1080p quality is the same, then the N4020 will be my best choice because it beats N5105 on price.
1080p H265 is my only use case for the CPU, I'll be running it as a headless Linux server running only tvheadend to transcode digital TV to H265 in realtime.
There is big quality difference, I mean it's Gen11 vs Gen9.
Other improvements include a new HEVC Quick Sync Video engine that provides up to a 30% bitrate reduction over Gen9 (at the same or better visual quality)
For the media block, Intel says that the Gen11 design includes a ground up HEVC encoder design, with high quality encode and decode support.
https://www.tomshardware.com/reviews/intel-sunny-cove-gen11-xe-gpu-foveros,5932-3.html
https://www.anandtech.com/show/13699/intel-architecture-day-2018-core-future-hybrid-x86/3
Atom based Gen11 SKUs only support the FF mode, I don't think there is bframes support. Not ideal but of course better than Gen9 iGPU.
RanmaCanada
16th October 2022, 04:34
I'm close to buying a cheap Intel mini PC for realtime 1080p H265 encoding.
I'm trying to decide between Gemini Lake N4020 and newer Jasper Lake N5105. Both have Intel UHD Graphics.
Considering both can do realtime 1080p H265 HW encodes, is there any reason to think that there would be a quality improvement using the newer processor?
If you're in the states, why not look on ebay for an 11th gen laptop with a smashed screen (https://www.ebay.com/itm/394288111429?_trkparms=amclksrc%3DITM%26aid%3D1110006%26algo%3DHOMESPLICE.SIM%26ao%3D1%26asc%3D20201210111314%26meid%3D1aa20a900add40d1a93aa564e745e6d3%26pid%3D101195%26rk%3D1%26rkt%3D12%26sd%3D354323643789%26itm%3D394288111429%26pmt%3D1%26noa%3D0%26pg%3D2047675%26algv%3DSimplAMLv11WebTrimmedV3MskuAspectsV202110NoVariantSeedKnnRecallV1%26brand%3DLenovo&_trksid=p2047675.c101195.m1851&amdata=cksum%3A3942881114291aa20a900add40d1a93aa564e745e6d3%7Cenc%3AAQAHAAABIMFr2e4EmAnM%252ByHZkULYKDIJ4L66fOjNL0iupgt%252BzO1%252F3AE1t3mNirUYB96NktMCicMagiS6mbeTl0xquGODv9l8iKOqtJXlKg2qgQBwOmNjHsQkhJ1ju%252FQXecFbzwuG6WTigWFgGu25g4p5f5YcE0FWEhtU6tWalHRoLnVZuZXSqqrl%252BENED49wlbSc0AWxDfeNhhoDZGgYg4RAzRCQLbSZV%252BxBuNORsVLMUpOWy3eneCCFtkRnOzEnhObbacQ7fmTkWO2p3CLjd8t%252FZGKXFTOp9XGs1dWuIgY17TBfDE9nrUZFufYAKUgcHnEtZJ8kowM7tg%252BDHtyRuOq7WYdwoFofvXVEMvb0dyBBwBiIZKQC6HIZBJvKX076HtWSHz9h6w%253D%253D%7Campid%3APL_CLK%7Cclp%3A2047675) or refurbished (https://www.ebay.com/itm/275225296169?hash=item4014b4bd29:g:FXUAAOSwuvdjIiwU&amdata=enc%3AAQAHAAAAsB5ABgwTztmoBFS3MsPcV0tP%2FI%2B4vCJohV5fp2vqvDLCHTh1tb2ATcBRjZY1ij2ZJqTmdzkQyDHNqkV%2B2pkj3qJynAdNYNYinzBTnMTF0mb8NDkcNAtSWRQvadDA8r5w06Y1qodr%2F%2Fnpbq04fg4hmhXyS0zD1Kl%2BxIcREpXStvP9CXddklD%2FthAc5%2B%2FqOMvzTmV9ZkKrsy2Y4fNsyxb4omLl3aUd14UDC0VWJ0mcKJzS%7Ctkp%3ABk9SR8qTvO37YA)?
I picked up an i3-1115G4 laptop off eBay for $120 as it had a smashed screen, and I'm using it as my Emby/Plex server. External out works, and you can just install remote software like Anydesk or whatever you like, and Bob's your uncle.
outhud
17th October 2022, 10:30
There is big quality difference, I mean it's Gen11 vs Gen9.
https://www.tomshardware.com/reviews/intel-sunny-cove-gen11-xe-gpu-foveros,5932-3.html
https://www.anandtech.com/show/13699/intel-architecture-day-2018-core-future-hybrid-x86/3
Atom based Gen11 SKUs only support the FF mode, I don't think there is bframes support. Not ideal but of course better than Gen9 iGPU.
Great! Thanks a lot. I think the Gen11 wins out so after reading those links.
tormento
20th October 2022, 18:39
40xx series capabilities (https://github.com/rigaya/NVEnc/blob/master/GPUFeatures/rtx4090.txt)
GrandPa
21st October 2022, 16:34
Did anybody recently test the new Intel ARC A380 on Intel's 8th generation CPU platform?
Several reviews praised the video encoding capabilities of the A380, but on the other hand explicitely mentioned severe problems to get the card up and running on "old" platforms without rBAR feature enabled. Intel doesn't mention platforms below 10th generation CPUs in the compatibility list.
@YUPS: Which platform did you use?
Since nearly 3 years, I was using HEVC HW encoding on a GTX-1660 card (Turing chip) with remarkably high encoding speed -compared to SW encoding- and good quality and compression rate, mainly on 720p HDTV material. I'm using a i7-8700K CPU on a GA Z370-HD3P motherboard.
Now, I would like to give the ARC A380 with its HEVC and AV1 capabilities a try, but I'm wondering whether the card would run on my system.
Any remarks are welcome.
RanmaCanada
21st October 2022, 20:26
Did anybody recently test the new Intel ARC A380 on Intel's 8th generation CPU platform?
Several reviews praised the video encoding capabilities of the A380, but on the other hand explicitely mentioned severe problems to get the card up and running on "old" platforms without rBAR feature enabled. Intel doesn't mention platforms below 10th generation CPUs in the compatibility list.
@YUPS: Which platform did you use?
Since nearly 3 years, I was using HEVC HW encoding on a GTX-1660 card (Turing chip) with remarkably high encoding speed -compared to SW encoding- and good quality and compression rate, mainly on 720p HDTV material. I'm using a i7-8700K CPU on a GA Z370-HD3P motherboard.
Now, I would like to give the ARC A380 with its HEVC and AV1 capabilities a try, but I'm wondering whether the card would run on my system.
Any remarks are welcome.
Intel specifically states you need an rBAR system for the card to work. The reason they don't mention them, is because they "aren't compatible", and you're on your own if you attempt to use them.
https://www.intel.com/content/www/us/en/support/articles/000091128/graphics.html
It is quite possible the card will work on your older system, but Intel does not guarantee any type of performance.
https://www.techpowerup.com/review/intel-arc-a770-pcie-3-resizable-bar/ Sorry no encoding tests.
All you can do is buy it and try it, and report back.
GrandPa
21st October 2022, 22:27
Intel specifically states you need an rBAR system for the card to work. The reason they don't mention them, is because they "aren't compatible", and you're on your own if you attempt to use them.
https://www.intel.com/content/www/us/en/support/articles/000091128/graphics.html
It is quite possible the card will work on your older system, but Intel does not guarantee any type of performance.
https://www.techpowerup.com/review/intel-arc-a770-pcie-3-resizable-bar/ Sorry no encoding tests.
All you can do is buy it and try it, and report back.
Many thanks for the link with GPU-Z. It helped me to find out all the rBAR requirements and to discover that the BIOS of my MoBo actually does support rBAR (it was hidden behind the "Above 4G Decode" setting, which I hadn't enabled yet).
Thus, a new A380 card is ordered, let's wait and see. I will report about it later.
Yups
21st October 2022, 23:05
Did anybody recently test the new Intel ARC A380 on Intel's 8th generation CPU platform?
Several reviews praised the video encoding capabilities of the A380, but on the other hand explicitely mentioned severe problems to get the card up and running on "old" platforms without rBAR feature enabled. Intel doesn't mention platforms below 10th generation CPUs in the compatibility list.
@YUPS: Which platform did you use?
Since nearly 3 years, I was using HEVC HW encoding on a GTX-1660 card (Turing chip) with remarkably high encoding speed -compared to SW encoding- and good quality and compression rate, mainly on 720p HDTV material. I'm using a i7-8700K CPU on a GA Z370-HD3P motherboard.
Now, I would like to give the ARC A380 with its HEVC and AV1 capabilities a try, but I'm wondering whether the card would run on my system.
Any remarks are welcome.
I have an older 7th Gen CPU i7-7700k+Z170. I don't use it for gaming, although I think missing rBAR might effect encoding performance in some cases with bigger PCIe copy work, for example when it's not using hardware decoding.
ukmark
22nd October 2022, 11:26
@Yups The MOBO on my Ice Lake laptop gave up the ghost, but I managed to get a great deal on an Acer i5-1135 G7 laptop.
It seems to me that FF encoding with this gives very good quality. As you say, Hybrid does give better quality, but how noticeable is the difference? I did a quick test here and I struggled to actually see any difference between Hybrid and FF for the same CQP values. The VMAF, PSNR and SSIM scores for Tiger Lake HEVC seem very close between FF and Hybrid from your earlier tests. Where do you see the differences in quality between FF and Hybrid, (is it facial details etc)?? Or, is it that the differences are so marginal, that FF is the way to go with Tiger Lake??
Here are my encoding options from StaxRip v2.13:-
--avhw --codec hevc --quality best --profile main10 --bframes 14 --b-pyramid --adapt-ltr --open-gop --async-depth 4 --no-repeat-pps --no-tskip --sao none --fixed-func --cqp 30:32:33 --qp-offset 2:4:8
TIA and BTW thanks for all your testing and advice!!
Yups
22nd October 2022, 12:34
Hybrid has a little more details in direct screenshot comparisons and the file size is a little smaller but it's not a big difference. Because of the big speed difference FF is the preferred choice in most cases. You can try without the custom offset by the way, maybe it's slightly better. CQP with 14 bframes is really good for sure. I tested with FFMPEG the last days, it's really good. But there is no SAO option to turn it off.
ukmark
22nd October 2022, 12:58
Hybrid has a little more details in direct screenshot comparisons and the file size is a little smaller but it's not a big difference. Because of the big speed difference FF is the preferred choice in most cases. You can try without the custom offset by the way, maybe it's slightly better. CQP with 14 bframes is really good for sure. I tested with FFMPEG the last days, it's really good. But there is no SAO option to turn it off.
Thanks for the reply. I really like the speed of FF with TGL. I'll try having a go with no qp offsets as you suggest and may also try using my original CQP settings but lowering the main CQP I/P/B settings to 29/31/32 from 30/32/33 and see if the size difference is significant. Thanks again. I just want a "set it and forget" preset that just works. If it means a slightly bigger file size, but same quality as Hybrid, then that's no problem.
ukmark
23rd October 2022, 16:56
I did some more testing of fixed-function vs hybrid and for the life of me I can't see any differences. I used the same CQP settings for both encodes (the only difference being fixed-function was turned on in one encode and turned off in the other). I used the StaxRip tool "Video Comparison" which allows you to compare the exact same frames from multiple files, in this case just the two files. Although you can't see a side-by-side view, the tool has a tab at the top of the gui (one tab for each file). If you click on each tab back and forth, you can usually see any differences, as the screen just overlays the current frame from that video.
In my testing, it was like a mirror image. I tested a video which had a lot of dark and bright scenes, fast and slow movement and I just could not see any differences. I just dragged the slider in the gui to random parts of the video and must have looked at over 200 frames. Also when playing back both files, there was just no observable differences. So, to my eyes, Tiger Lake has made hybrid redundant, I'm just sticking with fixed-function encoding.
I know that AV1 will eventually take over from HEVC, but right now I am very happy with 10-bit HEVC fixed-function encodes on Tiger Lake.
RanmaCanada
23rd October 2022, 17:13
I did some more testing of fixed-function vs hybrid and for the life of me I can't see any differences. I used the same CQP settings for both encodes (the only difference being fixed-function was turned on in one encode and turned off in the other). I used the StaxRip tool "Video Comparison" which allows you to compare the exact same frames from multiple files, in this case just the two files. Although you can't see a side-by-side view, the tool has a tab at the top of the gui (one tab for each file). If you click on each tab back and forth, you can usually see any differences, as the screen just overlays the current frame from that video.
In my testing, it was like a mirror image. I tested a video which had a lot of dark and bright scenes, fast and slow movement and I just could not see any differences. I just dragged the slider in the gui to random parts of the video and must have looked at over 200 frames. Also when playing back both files, there was just no observable differences. So, to my eyes, Tiger Lake has made hybrid redundant, I'm just sticking with fixed-function encoding.
I know that AV1 will eventually take over from HEVC, but right now I am very happy with 10-bit HEVC fixed-function encodes on Tiger Lake.
I am going to have to agree with you. In my limited testing, I encoded the same 4k movie, and I could not tell the difference between x265 and my Tiger Lake encode, other than the speeds. Tiger Lake managed 15+fps, while my x265 encode was going at around 1-1.5. I think it's time I get a larger ssd for my laptop so I can start encoding more stuff directly instead of dumping it on the network.
ukmark
23rd October 2022, 18:03
I am going to have to agree with you. In my limited testing, I encoded the same 4k movie, and I could not tell the difference between x265 and my Tiger Lake encode, other than the speeds. Tiger Lake managed 15+fps, while my x265 encode was going at around 1-1.5. I think it's time I get a larger ssd for my laptop so I can start encoding more stuff directly instead of dumping it on the network.
Was that Tiger Lake encode fixed function or hybrid?? I am getting around 235 fps when encoding a blu ray to 1080p using fixed function. Using latest Intel graphics driver (31.0.101.3729). My encode settings for StaxRip are:-
--avhw --codec hevc --quality best --profile main10 --bframes 16 --b-pyramid --adapt-ltr --open-gop --async-depth 4 --no-repeat-pps --no-tskip --sao none --fixed-func --cqp 32:33:34 --qp-offset 2:4:8
benwaggoner
24th October 2022, 02:18
I'm pretty sure you can't have --bframes 16 for a Blu-ray compatible encode!
RanmaCanada
24th October 2022, 02:18
Was that Tiger Lake encode fixed function or hybrid?? I am getting around 235 fps when encoding a blu ray to 1080p using fixed function. Using latest Intel graphics driver (31.0.101.3729). My encode settings for StaxRip are:-
--avhw --codec hevc --quality best --profile main10 --bframes 16 --b-pyramid --adapt-ltr --open-gop --async-depth 4 --no-repeat-pps --no-tskip --sao none --fixed-func --cqp 32:33:34 --qp-offset 2:4:8
I used fixed function, and it was a 4k HDR movie, not 1080p.
ukmark
24th October 2022, 09:25
I'm pretty sure you can't have --bframes 16 for a Blu-ray compatible encode!
Not sure what this means. I ripped a blu ray disc to an mkv file to play back on my S905x3 box running CoreELEC. This plays back perfectly on my 4k TV. All my encodes are done this way. The source can be a blu ray disc or a higher bit rate x264 file, or my recorded TV shows etc, pretty much anything. I use StaxRip to encode. If I have a blu ray disc to encode, I decrypt first using MakeMKV and then encode the decrypted file.
All my encodes are kept on a 4TB HDD and played back on a laptop (via USB cable) or the S905x3 to a 4k TV (via HDMI cable). I don't play anything using a blu ray player. I have an external USB Blu ray drive connected to laptop, that I use (with MakeMKV) to decrypt and rip any blu ray discs I own.
Can u pls explain more? TIA
ukmark
24th October 2022, 09:28
I used fixed function, and it was a 4k HDR movie, not 1080p.
I would still have thought the speed would have been quicker. 4k is 4x size of blu ray, so a rough guess would have given between 50-60 fps (if I'm getting over 200 fps with 1080p). Maybe the HDR slows things down a lot. I have never ripped a 4k HDR file, so would not know. Still it's 10-15x faster than x265.
ukmark
24th October 2022, 16:28
Hybrid has a little more details in direct screenshot comparisons and the file size is a little smaller but it's not a big difference. Because of the big speed difference FF is the preferred choice in most cases. You can try without the custom offset by the way, maybe it's slightly better. CQP with 14 bframes is really good for sure. I tested with FFMPEG the last days, it's really good. But there is no SAO option to turn it off.
I have also been playing around with ICQ and 16 bframes using the latest Intel driver (31.0.101.3729) - (not sure if the graphics driver version makes a difference if you are using fixed-function encoding??) Admittedly a fairly quick test, but it does appear that ICQ and 16 bframes looks good (at least to my eyes). I also remember you saying that 14 bframes gave better results than 16 bframes and also with a little better compression. Again, using latest driver, it does appear that 16 bframes gives better compression than 14 bframes (CQP or ICQ). Just thought I'd pass it on - maybe worth a revisit? CQP still encodes quite a bit faster than ICQ using StaxRip and its output is so reliable.
Yups
24th October 2022, 23:29
I have also been playing around with ICQ and 16 bframes using the latest Intel driver (31.0.101.3729) - (not sure if the graphics driver version makes a difference if you are using fixed-function encoding??) Admittedly a fairly quick test, but it does appear that ICQ and 16 bframes looks good (at least to my eyes). I also remember you saying that 14 bframes gave better results than 16 bframes and also with a little better compression. Again, using latest driver, it does appear that 16 bframes gives better compression than 14 bframes (CQP or ICQ). Just thought I'd pass it on - maybe worth a revisit? CQP still encodes quite a bit faster than ICQ using StaxRip and its output is so reliable.
Do you make a PSNR/SSIM/VMAF comparison or how do you know 16 is better? Did you check the file size? bframes 14 file size is slightly smaller than bframes 16 in my testing, this is also an indication that bframes 14 is slightly more efficient. It could depend on the sample if this isn't the same for you. The difference is small anyways.
ICQ isn't bad but not as good as CQP. CQP with bframes 14-16 is extremely effective on Xe LP.
benwaggoner
24th October 2022, 23:35
I would still have thought the speed would have been quicker. 4k is 4x size of blu ray, so a rough guess would have given between 50-60 fps (if I'm getting over 200 fps with 1080p). Maybe the HDR slows things down a lot. I have never ripped a 4k HDR file, so would not know. Still it's 10-15x faster than x265.
HDR itself doesn't slow anything down. HDR requires 10-bit, and 10-bit is slower than 8-bit encoding, but that's Main versus Main10, not anything to do with color volume.
benwaggoner
24th October 2022, 23:38
Do you make a PSNR/SSIM/VMAF comparison or how do you know 16 is better? Did you check the file size? bframes 14 file size is slightly smaller than bframes 16 in my testing, this is also an indication that bframes 14 is slightly more efficient. It could depend on the sample if this isn't the same for you. The difference is small anyways.
ICQ isn't bad but not as good as CQP. CQP with bframes 14-16 is extremely effective on Xe LP.
Those are all single-frame metrics, so it's hard to interpret them with changes in B-frames. Harmonic mean of the prior two seconds isn't bad, but it can get complicated due to the different nature of reference and non-reference frames.
Yups
25th October 2022, 02:36
Those are all single-frame metrics, so it's hard to interpret them with changes in B-frames. Harmonic mean of the prior two seconds isn't bad, but it can get complicated due to the different nature of reference and non-reference frames.
The quality changes massively and you are saying they cannot interpret them, what? Subjective quality changes massively and of course the metric scores go up.
ukmark
25th October 2022, 07:57
Do you make a PSNR/SSIM/VMAF comparison or how do you know 16 is better? Did you check the file size? bframes 14 file size is slightly smaller than bframes 16 in my testing, this is also an indication that bframes 14 is slightly more efficient. It could depend on the sample if this isn't the same for you. The difference is small anyways.
ICQ isn't bad but not as good as CQP. CQP with bframes 14-16 is extremely effective on Xe LP.
Purely based on file size. So far 14 bframes gives a slightly bigger file size than 16 bframes (all other settings are the same). I don't compare PSNR/VMAF/SSIM, just an observation re file size and bframes. I've only tried a few files but each time, 14 bframes file size is bigger than 16 bframes with FF encoding. As you say, the difference is very small. I can't say that 14 bframes is better/worse quality than 16 bframes. To me the quality looks the same, but I have no metrics to qualify that, just that it appears to compress slightly better with 16 bframes and latest driver.
ukmark
25th October 2022, 09:13
Do you make a PSNR/SSIM/VMAF comparison or how do you know 16 is better? Did you check the file size? bframes 14 file size is slightly smaller than bframes 16 in my testing, this is also an indication that bframes 14 is slightly more efficient. It could depend on the sample if this isn't the same for you. The difference is small anyways.
ICQ isn't bad but not as good as CQP. CQP with bframes 14-16 is extremely effective on Xe LP.
Metrics for the first 10 minutes from my Star Trek into Darkness blu ray encode.
PSNR SSIM VMAF SPEED SIZE
14 bframes 36.22 0.9713 94.18 238 412 Mb
16 bframes 36.21 0.9711 94.13 239 403 Mb
Encode settings using StaxRip below:-
--avhw --codec hevc --quality best --profile main10 --bframes 14(or 16) --b-pyramid --adapt-ltr --open-gop --async-depth 4 --no-repeat-pps --no-tskip --sao none --fixed-func --cqp 30:32:33 --qp-offset 2:4:8
Differences are very small as can be seen. I have chosen to stick with 14 bframes, as even though the differences are small, I'd rather have that tiny bit extra quality and accept that file sizes could be around 2% bigger.
benwaggoner
25th October 2022, 21:42
The quality changes massively and you are saying they cannot interpret them, what? Subjective quality changes massively and of course the metric scores go up.
If you see massive subjective improvements, you don't need metrics! But sometimes subjective quality goes up and metrics go down, and vise versa.
For example, adaptive quantization in general tends to reduce metrics but (properly applied) improves subjective quality.
Mister XY
26th October 2022, 06:30
With all due respect to this effort and this work, using a score to orientate oneself brings 0 points. So you might have reduced a file to 200 MB, but especially in dark scenes you can see a lot of artefacts and blocks. So the only thing left to do is to check for yourself whether the settings have any effect or not. Example: Encoded the Intel film over and over again with different settings. In the x264 format, in the x265 (via software and hardware) format and also in the AV1 format. The funny thing was that my AV1 film and my x265 film were identical in size, but the AV1 film had a better score. Now comes the funny thing, the AV1 film had a very high level of blocking in dark scenes, even though it was better than the other formats according to the metrics. Therefore, these score values have no meaning for me. The only benchmark that really exists is to take a film and encode it again and again, with different settings etc. and then you pay attention to fast transitions, to fine details in the film and to dark scenes and only when these are then in the encode still all are available, then you have reasonable settings. But using a score as a guide goes in the wrong direction. I can have a 100% score but not a viewable image because maybe the single image is 100% but not in conjunction with 23 other images that follow.
rwill
26th October 2022, 07:38
Yeah, single frame metrics combined with mean... questionable when there are multiple frames. Useless when the sequence contains more than one scene.
GrandPa
2nd November 2022, 10:42
All you can do is buy it and try it, and report back.
The A380 card has arrived a few days ago, installation went without any problem, and I spent many hours with coding tests.
Summary
Configuration: i7-8700K on Z370 platform, GTX-1660 and A380 cards installed, BAR enabled, QSV HW decoding enabled
Basic preconditions for the coding tests: No visible blocks or artefacts in fast motion scenes or large dark, still areas
Procedure: Basis for comparison is a SW encode with x.265 at a reasonably high quality, VMAF score is taken as reference
(Absolute VMAF figures will not help here, just referring to the x.265 coding results)
Observations:
QSV is in most cases around 15-20% slower than NV (except with 4K material)
QSV is considerably faster than NV during downscaling (4K-to-720p, 1080p-to-720p)
QSV delivers better quality in most cases or better compression rate at same quality
StaxRip parameters:
--avhw --codec hevc --quality best --profile main10 --bframes 16 --b-pyramid --adapt-ltr --open-gop --async-depth 4
--colormatrix bt709 --colorprim bt709 --transfer bt709 --fixed-func --cqp ??:??:?? --qp-offset 2:4:8
A few sample results:
4K (Jellyfish) Speed [fps] File size [MB] VMAF
x.265 Reference 9 57,6 87,81
Nvenc 56 57,0 88,05
QSVenc 74 46,1 88,40
1080p (SolLevante) Speed [fps] File size [MB] VMAF
x.265 Reference 40 194,4 91,66
Nvenc 356 171,9 91,80
QSVenc 300 139,4 91,45
720p (HDTV Clip) Speed [fps] File size [MB] VMAF
x.265 Reference 136 23,9 89,33
Nvenc 855 23,0 89,13
QSVenc 629 17,9 89,89
(No audio track contained, CQP parameters tuned accordingly aiming at the reference quality)
Overall, the A380 card is a good deal, although there seems to be still room left for optimization, maybe with future
driver releases.
ukmark
4th November 2022, 22:36
I love StaxRip but it appears that development has stopped. It is easily the best gui I have used and does everything very well. I'm mainly interested in the gpu hardware quick sync side of things, rather than the software/CPU encoding.
Does anybody know of any other current encoder gui that will incorporate up-to-date support of quick sync and x265 etc along with audio support etc??
I know Handbrake is still going, but you are limited in the advanced options that can be set (unless things have changed). For example the "qp-offsets" can be set in StaxRip and these really help in maintaining quality whilst keeping file size as low as possible. This option is not available in Handbrake. This is not the basic QP I,P,B offsets but the
offsets under "rate control" section of the StaxRip gui. It may well be that even in its current state StaxRip will suffice for years to come as regards HEVC quick sync support, but wanted to know if there was anything else out there with the same advanced features etc.
I have looked on the web but could not find anything. Free or paid does not matter, but would prefer free.
TIA
Yups
5th November 2022, 03:58
Metrics for the first 10 minutes from my Star Trek into Darkness blu ray encode.
PSNR SSIM VMAF SPEED SIZE
14 bframes 36.22 0.9713 94.18 238 412 Mb
16 bframes 36.21 0.9711 94.13 239 403 Mb
Encode settings using StaxRip below:-
--avhw --codec hevc --quality best --profile main10 --bframes 14(or 16) --b-pyramid --adapt-ltr --open-gop --async-depth 4 --no-repeat-pps --no-tskip --sao none --fixed-func --cqp 30:32:33 --qp-offset 2:4:8
Differences are very small as can be seen. I have chosen to stick with 14 bframes, as even though the differences are small, I'd rather have that tiny bit extra quality and accept that file sizes could be around 2% bigger.
In all my tests bframes 14 file size was slightly smaller and scores a bit higher, for example in the Intel video. In your test there is not much difference anyways. 14-16 are equally good overall.
Did you try without custom offset with like 30:32:36? I think disabling tskip might not be better in every case.
You don't need LTR and repeat pps with CQP, it makes no difference.
I love StaxRip but it appears that development has stopped. It is easily the best gui I have used and does everything very well. I'm mainly interested in the gpu hardware quick sync side of things, rather than the software/CPU encoding.
Does anybody know of any other current encoder gui that will incorporate up-to-date support of quick sync and x265 etc along with audio support etc??
Maybe have a look to Hybrid (https://forum.doom9.org/showthread.php?t=153035)
It's not up to date (no AV1 support yet) but it's not stopped like Staxrip.
ukmark
6th November 2022, 19:24
Maybe have a look to Hybrid (https://forum.doom9.org/showthread.php?t=153035)
It's not up to date (no AV1 support yet) but it's not stopped like Staxrip.
Cheers.
RanmaCanada
6th November 2022, 22:07
I love StaxRip but it appears that development has stopped. It is easily the best gui I have used and does everything very well. I'm mainly interested in the gpu hardware quick sync side of things, rather than the software/CPU encoding.
Does anybody know of any other current encoder gui that will incorporate up-to-date support of quick sync and x265 etc along with audio support etc??
I know Handbrake is still going, but you are limited in the advanced options that can be set (unless things have changed). For example the "qp-offsets" can be set in StaxRip and these really help in maintaining quality whilst keeping file size as low as possible. This option is not available in Handbrake. This is not the basic QP I,P,B offsets but the
offsets under "rate control" section of the StaxRip gui. It may well be that even in its current state StaxRip will suffice for years to come as regards HEVC quick sync support, but wanted to know if there was anything else out there with the same advanced features etc.
I have looked on the web but could not find anything. Free or paid does not matter, but would prefer free.
TIA
You can always update the encoder files (https://github.com/rigaya/QSVEnc) manually in staxrip and then it will ask you what the new version fo the software is. Best way to stick with a gui you like and know how to use.
ukmark
7th November 2022, 12:59
Thanks for that - yes I do update QSVEnc regularly and am using the latest version (v7.23). Maybe stax76 will return and continue with StaxRip in the future. It's definitely the best GUI I have used so far.
ReinerSchweinlin
19th November 2022, 11:43
So my ARC 770 arrived :)
Where do I start? From reading this thread, I remember that with Intel XE, the QSVENCC from rigaya, called upon by staxrip, was the best solution for quality. The ffmpeg implementation did have some problems? A´S Video Encoder seems to beusing the intel provided encoder, which offers a lot less knobs to turn.... Handbrake also offers only the bare minimum to tune.
SO using hybrid/staxrip or any other QSVENCC GUI is still the best option? (for the goal of getting max quality out of encodes..)
Yups
20th November 2022, 02:00
QSVEnc, FFmpeg and Handbrake. FFmpeg HEVC works fine, the only missing option for me is SAO. I think QSVEnc GUI is the best to begin with.
ReinerSchweinlin
21st November 2022, 18:02
Thanx, so I´ll probably use hybrid as GUI for QSVENC :)
GrandPa
17th December 2022, 22:47
Overall, the A380 card is a good deal, although there seems to be still room left for optimization, maybe with future
driver releases.
Indeed, with the latest driver release 3959, coding speed has been considerably improved in several test cases. I'll continue monitoring the promising evolution.
Lan4
3rd March 2023, 01:41
Hello! Help me please. What is VQEnhancer? I can't find a description for this.
DuyNghia.Net
22nd July 2023, 23:57
Hello everyone, I am about to buy Intel UHD 770 12th/13th Gen for encoding HEVC QSV Encoding 1080p/4K.
How fast is it compared to A380? Anyone, Thanks!
My 9th gen Intel 9400 UHD630 gets Constant 60fps 1080p transcoding H264->HEVC using QSVEnc which is quite slow!
RanmaCanada
26th July 2023, 01:48
Hello everyone, I am about to buy Intel UHD 770 12th/13th Gen for encoding HEVC QSV Encoding 1080p/4K.
How fast is it compared to A380? Anyone, Thanks!
My 9th gen Intel 9400 UHD630 gets Constant 60fps 1080p transcoding H264->HEVC using QSVEnc which is quite slow!
They should be about the same as they have the same ASICS as far as I understand. The dedicated graphics might be a bit faster due to cpu overhead. If I am mistaken, can someome please correct me as I do not want to be spreading lies.
Yups
19th August 2023, 23:38
Hello everyone, I am about to buy Intel UHD 770 12th/13th Gen for encoding HEVC QSV Encoding 1080p/4K.
How fast is it compared to A380? Anyone, Thanks!
My 9th gen Intel 9400 UHD630 gets Constant 60fps 1080p transcoding H264->HEVC using QSVEnc which is quite slow!
HD630 features a hybrid only HEVC encoder which is slow. UHD 770 can do both hybrid and fully fixed encoding. UHD 770 and A380 are much faster and better quality. Note that A380 als supports AV1 encoding.
You can see here:
https://forum.doom9.org/showpost.php?p=1937452&postcount=396
Iris Xe has the same encoder as UHD 770.
easyfab
24th September 2024, 19:55
New intel lunar lake with Xe2 is out now.
Is better encoding with HEVC or AV1 expected ?
RanmaCanada
25th September 2024, 23:08
New intel lunar lake with Xe2 is out now.
Is better encoding with HEVC or AV1 expected ?
The current wikipedia (https://en.wikipedia.org/wiki/Intel_Quick_Sync_Video) page only goes up to Arrow Lake (who's page doesn't exist), so uncertain. It would be great, but with everything that is exploding at Intel lately, more than likely not.
Edit: since I can't strike this through, I'll add this pdf (https://cdrdv2-public.intel.com/824434/2024_Intel_Tech%20Tour%20TW_Xe2%20and%20Lunar%20Lakes%20GPU.pdf) from Intel. Apparently the new Lunar Lake has hardware VVC encoding, which would potentially mean new ASICS and possibly better encoding for everything else. Until someone buys one, and tests it, we won't know for sure.
ReinerSchweinlin
26th September 2024, 10:04
Seems VCC Decode is on board, I see no clear mentioning of encoding via hardware.
nevcairiel
26th September 2024, 10:26
Lunar Lake has VVC Decode, not Encode. Arrow Lake is expected to get 1st Gen Xe graphics only, so no VVC there.
RanmaCanada
26th September 2024, 17:12
I should have read the entire pdf till the end. Stupid of Intel of hide that tidbit on the very last page. My apologies. I presumed that since it was mentioned on slide 47 and beyond in the codec section that it included both encode and decode. Putting the facts on the last page, after the F1 gaming "demo", is dirty.
Again, I'm sorry.
burnix
3rd December 2024, 20:14
hi.
Can someone can share ffmpeg script to make low bitrate good quality video with hevc_qsv ???
is it possible to avoid big artifact on hi-moving scene even in veryslow preset ????
Thanks for your advice.
benwaggoner
11th December 2024, 18:37
hi.
Can someone can share ffmpeg script to make low bitrate good quality video with hevc_qsv ???
is it possible to avoid big artifact on hi-moving scene even in veryslow preset ????
Thanks for your advice.
You need to give a lot more specifics to get any useful advice.
What is "low bitrate?"
What is the content?
What is the frame size and fps?
What do you mean by "big artifact?" Highly visible? Spatially large?
What have you tried so far, with what results?
burnix
11th December 2024, 19:04
You need to give a lot more specifics to get any useful advice.
What is "low bitrate?" 800k (video) / 128k audio
Film
[QUOTE=What is the frame size and fps? 1920*XXXX / 23 or 25 fps
big square visible on high moving scene
[QUOTE=What have you tried so far, with what results?
i have tested hevc enc with ffmpeg and qsvenc (qsvenc 50% faster than ffmpeg on basic hevc encoding)
But as i say in another thread, i'm far to be an expert like you
ffmpeg.exe -hwaccel qsv -init_hw_device qsv=qsv -filter_hw_device qsv -hwaccel_output_format qsv -i %1 -pix_fmt nv12 -c:v hevc_qsv -preset veryslow -b:v 800k -ac 2 -c:a aac -b:a 128k -vsync vfr -f mp4 %~n1-h265%~x1
QSVEncC64.exe --avhw -i %1 -o %~n1-qsv-i5%~x1 --codec hevc --cbr 800 --avsync vfr --chapter-copy --audio-stream :stereo --audio-bitrate 128 --audio-codec aac:aac_coder=twoloop --quality best --extbrc --mbbrc --i-adapt --b-adapt --direct-bias-adjust --adapt-ref --adapt-ltr --adapt-cqm --fade-detect --tskip --sao all --hevc-gpb
benwaggoner
12th December 2024, 01:22
800k (video) / 128k audio
Film
1920*XXXX / 23 or 25 fps
big square visible on high moving scene
i have tested hevc enc with ffmpeg and qsvenc (qsvenc 50% faster than ffmpeg on basic hevc encoding)
But as i say in another thread, i'm far to be an expert like you
ffmpeg.exe -hwaccel qsv -init_hw_device qsv=qsv -filter_hw_device qsv -hwaccel_output_format qsv -i %1 -pix_fmt nv12 -c:v hevc_qsv -preset veryslow -b:v 800k -ac 2 -c:a aac -b:a 128k -vsync vfr -f mp4 %~n1-h265%~x1
QSVEncC64.exe --avhw -i %1 -o %~n1-qsv-i5%~x1 --codec hevc --cbr 800 --avsync vfr --chapter-copy --audio-stream :stereo --audio-bitrate 128 --audio-codec aac:aac_coder=twoloop --quality best --extbrc --mbbrc --i-adapt --b-adapt --direct-bias-adjust --adapt-ref --adapt-ltr --adapt-cqm --fade-detect --tskip --sao all --hevc-gpb
800 Kbps is really low for typical TV/movie 1080p content. There won't be any way to do it without artifacts. You might be able to get it with fewer artifacts, particularly if you're willing to do 2-pass x265 using slow settings, and wait >10x as long to encode. Even with all that, getting decent quality below a 4 Mbps average bitrate can be impossible, especially if there's much film grain in the source. Clean animation can go quite a bit lower.
burnix
12th December 2024, 08:03
800 Kbps is really low for typical TV/movie 1080p content. There won't be any way to do it without artifacts. You might be able to get it with fewer artifacts, particularly if you're willing to do 2-pass x265 using slow settings, and wait >10x as long to encode. Even with all that, getting decent quality below a 4 Mbps average bitrate can be impossible, especially if there's much film grain in the source. Clean animation can go quite a bit lower.
After reading much thread and most of yours i can acheive good visual quality with nearly zero artifact at 800k in hevc.
I have found that the "veryfast" profile of ffmpeg integrate a lots of settings helping me achieve this best result (but i add some x265 settings more).
So its not impossible and pretty decent to view it in a 42' led tv
ReinerSchweinlin
8th March 2025, 09:50
AMD claims that the new Hardware Encoder in the RX 9070 is very good. Does anyone happen to have one and is using it for video-transcoding / Encoding and took a closer look at the quality ?
Iij33920
6th April 2025, 19:37
Has anyone run any tests with recent (3-5000 models) Nvidia nvencc? Looking for a starting place for encoding to keep alongside the source (kind of like flac + m4a). Mostly 1080p film sources, older stuff (late 80s, 90s and early 2000s) so some noise/grain possible. Not looking for transparent encodes otherwise I'd be using software.
Z2697
6th April 2025, 21:36
Has anyone run any tests with recent (3-5000 models) Nvidia nvencc? Looking for a starting place for encoding to keep alongside the source (kind of like flac + m4a). Mostly 1080p film sources, older stuff (late 80s, 90s and early 2000s) so some noise/grain possible. Not looking for transparent encodes otherwise I'd be using software.
https://rigaya.github.io/vq_results/
The author of the popular command line tools for NVENC, QSV and AMF/VCE.
ReinerSchweinlin
13th November 2025, 16:39
We had some talks here abput the Intel iGPUs and "lowe power mode". While settiing up i3-n305 and fiddling around with handrake, I stumbled across this setting again. I tried to find a in depth ressource or the differences of normal and low-power mode, but could no find any. Can anyone help ?
DuyNghia.Net
15th June 2026, 19:00
They should be about the same as they have the same ASICS as far as I understand. The dedicated graphics might be a bit faster due to cpu overhead. If I am mistaken, can someome please correct me as I do not want to be spreading lies.
HD630 features a hybrid only HEVC encoder which is slow. UHD 770 can do both hybrid and fully fixed encoding. UHD 770 and A380 are much faster and better quality. Note that A380 als supports AV1 encoding.
You can see here:
https://forum.doom9.org/showpost.php?p=1937452&postcount=396
Iris Xe has the same encoder as UHD 770.
Thank You very much for your replies & sorry for this late reply as in 2023-2025 I lost my job & rent a new apartment.
I have UHD770 from (Intel 13500 & 12700K)
Also I just bought 2 old laptops for $US151 $US209 that have 1135G7 and 1145G7 both have TWO Multi-Format Codec Engines according to Intel specs
The QSVEncC64.exe encoding speed are very fast when converting from 1080p MP4 H264 to 1080p H265 all 4 of them are very similar which are ~290-320 fps (sometimes a bit lower or higher)
The gen11 13.3" laptops eat about 20w constantly when encoding
13500 eats 40-65w, 12700K eats 20w more than 13500 I do work on PC while waiting for them to encode my programming tutorials
DuyNghia.Net
23rd June 2026, 20:51
After using iGPU Intel Iris Xe Graphics G7 on 1145G7 (HP G8 laptop) + 1135G7 (Lenovo Gen 2) for quite sometimes I noticed some weird issues comparing to UHD770 on 13500 + 12700K
- I can run 2-3 instances of QSVEncC64.exe simultaneously on UHD770 without any warnings or errors
- Xe G7 can likely give me warnings/errors as image attached shows https://freeimage.host/i/CTtOASp
- Some command arguments cause Xe G7 performance reduce to like ~240-250fps instead of ~290fps
- In some case both UHD770 and Iris XeG7 can drop encoding fps to like 130-140fps, normally 290-320fps
RAM and NVMe isn't bottleneck at all because I have 64GB desktop & 16GB laptop respectively, TaskManager shows those resources are still free by a lot
The good thing is that all 4 iGPU produce Exact Same bit-to-bit (binary comparision) results, so their encoders on HEVC are the same!
rwill
23rd June 2026, 21:03
- In some case both UHD770 and Iris XeG7 can drop encoding fps to like 130-140fps, normally 290-320fps
Maybe thermal throttling? You might want to check CPU load and GPU temperature in the Task Manager..
DuyNghia.Net
23rd June 2026, 21:49
Maybe thermal throttling? You might want to check CPU load and GPU temperature in the Task Manager..
Idle is quite high, my room temp is 31 Celsius (°C)
Encoding ~above 60
https://freeimage.host/i/CTD1Bvs
Idle
https://freeimage.host/i/CTD1V6P
It sometimes jumps to 80-90 when runs, loads But keeps stable 60-64 when encodes
vBulletin® v3.8.11, Copyright ©2000-2026, vBulletin Solutions Inc.