View Full Version : Which processor to encode x265 4K ?
Nico8583
24th June 2019, 21:15
Hi :)
I would like to change my old X6 1090T to encode x265 1080p/4K.
I know Ryzen 3 will be out soon but I don't know if it would be a good value for money choice.
Currently, which processors are good value for money ? Ryzen 7 2700x ? i7 9700k ?
Thank you !
Groucho2004
24th June 2019, 21:43
Have a look here (https://www.techpowerup.com/review/intel-core-i7-9700k/6.html) (scroll down to "H.265 Media Encoding").
mariush
24th June 2019, 22:08
For comparison, your x6 1090t is somewhere around i3-7100 .. Ryzen 3 1200 in those charts on Techpowerup.
Nico8583
24th June 2019, 22:13
Thank you, i7 9700k and R7 2700X seem to be very close in performances.
i7 is much expensive and need a CPU cooler but 2700x need a graphic card.
Nico8583
24th June 2019, 22:20
Thank you for comparison, so I can expect FPS x 4 with 9700k or 2700X
RanmaCanada
24th June 2019, 23:00
I would say wait and see what happens with Ryzen 3000 series after they are released. You've waited this long to upgrade, so seriously, what's another 2-3 weeks (for PROPER reviews). It is highly possible that with their increased IPC that they might catch Intel for x265. And if not, well the Ryzen 2000 series will be a LOT cheaper by then, as AMD does whatever they can to get rid of old stock, where as Intel likes to keep prices on old stock stupidly high.
Asmodian
25th June 2019, 01:17
Actually Intel is expected to reduce the prices of their CPUs by up to 15%. Zen2 is too good for them to ignore this time.
It really is not the right time to get a CPU, 2-3 weeks should offer better CPUs per dollar from both AMD and Intel.
Nico8583
25th June 2019, 07:26
Yes, you're right both, I'll not buy anything before 2 weeks, but I start to search an alternative if Ryzen 3 are not as good as expected ;) (because if there are discount prices after Ryzen 3 release, I wish to buy immediatly and not thinking if it is a good choice). Now I know 9700K and 2700X are good values.
Thank you !
hajj_3
25th June 2019, 08:26
zen 2 supports proper 256bit avx 2 like intel, zen 1 used 2 x 128bit. 256bit AVX2 support increases performance of x265 by quite a bit, maybe 10% or something so i wouldn't buy zen+ if you are going to be doing quite a bit of x265 encoding.
Atak_Snajpera
25th June 2019, 11:45
Yes, you're right both, I'll not buy anything before 2 weeks, but I start to search an alternative if Ryzen 3 are not as good as expected ;) (because if there are discount prices after Ryzen 3 release, I wish to buy immediatly and not thinking if it is a good choice). Now I know 9700K and 2700X are good values.
Thank you !
Zen2 will be much faster than Zen1 due to AVX256! x265 really loves AVX256. That's the main reason why 10C 4GHz Intel is only slightly slower than ThreadRipper 1950x@4GHz.
2990WX has very low fps because of broken windows scheduler.
Just compare those results.
65.0 fps - 2 x Intel Xeon E5-4660 v3 @ 2.9GHz^ ( 14C / 28T ) MMX2 SSE2Fast SSSE3 SSE4.2 AVX AVX2 FMA3 LZCNT BMI2
63.6 fps - Intel Core i9-7920X @ 4.6GHz^ ( 12C / 24T ) MMX2 SSE2Fast SSSE3 SSE4.2 AVX AVX2 FMA3 LZCNT BMI2
54.3 fps - Intel Core i9-7900X @ 4.7GHz^ ( 10C / 20T ) MMX2 SSE2Fast SSSE3 SSE4.2 AVX AVX2 FMA3 LZCNT BMI2
52.5 fps - AMD Threadripper 2990WX @ 3.4GHz ( 32C / 64T ) MMX2 SSE2Fast SSSE3 SSE4.2 AVX AVX2 FMA3 LZCNT BMI2
50.3 fps - AMD Threadripper 1950X @ 4.0GHz^ ( 16C / 32T ) MMX2 SSE2Fast SSSE3 SSE4.2 AVX AVX2 FMA3 LZCNT BMI2
44.6 fps - Intel Core i9-7900X @ 4.0GHz ( 10C / 20T ) MMX2 SSE2Fast SSSE3 SSE4.2 AVX AVX2 FMA3 LZCNT BMI2
43.6 fps - AMD Threadripper 1950X @ 3.4GHz ( 16C / 32T ) MMX2 SSE2Fast SSSE3 SSE4.2 AVX AVX2 FMA3 LZCNT BMI2
36.8 fps - Intel Core i7-8700K @ 5.0GHz^ ( 6C / 12T ) MMX2 SSE2Fast SSSE3 SSE4.2 AVX AVX2 FMA3 LZCNT BMI2
35.3 fps - Intel Core i7-8700K @ 4.9GHz^ ( 6C / 12T ) MMX2 SSE2Fast SSSE3 SSE4.2 AVX AVX2 FMA3 LZCNT BMI2
33.9 fps - Intel Core i7-5960X @ 4.4GHz^ ( 8C / 16T ) MMX2 SSE2Fast SSSE3 SSE4.2 AVX AVX2 FMA3 LZCNT BMI2
31.8 fps - Intel Core i7-8700K @ 4.3GHz ( 6C / 12T ) MMX2 SSE2Fast SSSE3 SSE4.2 AVX AVX2 FMA3 LZCNT BMI2
29.8 fps - Intel Xeon E5-2675 v3 @ 1.8GHz ( 16C / 32T ) MMX2 SSE2Fast SSSE3 SSE4.2 AVX AVX2 FMA3 LZCNT BMI2
25.5 fps - AMD Ryzen 7 1700 @ 3.7GHz^ ( 8C / 16T ) MMX2 SSE2Fast SSSE3 SSE4.2 AVX AVX2 FMA3 LZCNT BMI2
24.8 fps - Intel Core i7-6700K @ 4.8GHz^ ( 4C / 8T ) MMX2 SSE2Fast SSSE3 SSE4.2 AVX AVX2 FMA3 LZCNT BMI2
24.3 fps - AMD Ryzen 7 1700 @ 3.5GHz^ ( 8C / 16T ) MMX2 SSE2Fast SSSE3 SSE4.2 AVX AVX2 FMA3 LZCNT BMI2
23.8 fps - Intel Core i5-8400 @ 3.8GHz ( 6C / 6T ) MMX2 SSE2Fast SSSE3 SSE4.2 AVX AVX2 FMA3 LZCNT BMI2
23.7 fps - Intel Core i7-7700K @ 4.8GHz^ ( 4C / 8T ) MMX2 SSE2Fast SSSE3 SSE4.2 AVX AVX2 FMA3 LZCNT BMI2
23.3 fps - Intel Core i7-6700K @ 4.7GHz^ ( 4C / 8T ) MMX2 SSE2Fast SSSE3 SSE4.2 AVX AVX2 FMA3 LZCNT BMI2
22.6 fps - Intel Core i7-6700K @ 4.5GHz^ ( 4C / 8T ) MMX2 SSE2Fast SSSE3 SSE4.2 AVX AVX2 FMA3 LZCNT BMI2
21.6 fps - Intel Core i7-7700K @ 4.4GHz^ ( 4C / 8T ) MMX2 SSE2Fast SSSE3 SSE4.2 AVX AVX2 FMA3 LZCNT BMI2
20.6 fps - AMD Ryzen 5 1600 @ 3.8GHz^ ( 6C / 12T ) MMX2 SSE2Fast SSSE3 SSE4.2 AVX AVX2 FMA3 LZCNT BMI2
18.6 fps - Intel Core i7-6600K @ 4.5GHz^ ( 4C / 4T ) MMX2 SSE2Fast SSSE3 SSE4.2 AVX AVX2 FMA3 LZCNT BMI2
16.4 fps - Intel Xeon E5-2690 @ 2.9GHz ( 8C / 16T ) MMX2 SSE2Fast SSSE3 SSE4.2 AVX
15.8 fps - Intel i7-6770HQ @ 2.6GHz ( 4C / 8T ) MMX2 SSE2Fast SSSE3 SSE4.2 AVX AVX2 FMA3 LZCNT BMI2
15.8 fps - Intel Core i5-4690K @ 4.2GHz^ ( 4C / 4T ) MMX2 SSE2Fast SSSE3 SSE4.2 AVX AVX2 FMA3 LZCNT BMI2
15.6 fps - Intel Xeon E3 1231 v3 @ 3.4GHz ( 4C / 8T ) MMX2 SSE2Fast SSSE3 SSE4.2 AVX AVX2 FMA3 LZCNT BMI2
14.7 fps - Intel Xeon E5-2670 @ 2.6GHz ( 8C / 16T ) MMX2 SSE2Fast SSSE3 SSE4.2 AVX
14.7 fps - Intel Core i7-3930K @ 3.2GHz ( 6C / 12T ) MMX2 SSE2Fast SSSE3 SSE4.2 AVX
14.6 fps - AMD Ryzen 5 2400G @ 4.0GHz^ ( 4C / 8T ) MMX2 SSE2Fast SSSE3 SSE4.2 AVX AVX2 FMA3 LZCNT BMI2
13.9 fps - Intel Core i5-7400 @ 3.5GHz ( 4C / 4T ) MMX2 SSE2Fast SSSE3 SSE4.2 AVX AVX2 FMA3 LZCNT BMI2
13.8 fps - Intel Core i5-6500 @ 3.2GHz ( 4C / 4T ) MMX2 SSE2Fast SSSE3 SSE4.2 AVX AVX2 FMA3 LZCNT BMI2
13.5 fps - AMD Ryzen 5 1500X @ 3.5GHz ( 4C / 8T ) MMX2 SSE2Fast SSSE3 SSE4.2 AVX AVX2 FMA3 LZCNT BMI2
13.3 fps - Intel Core i5-4570S @ 3.6GHz^ ( 4C / 4T ) MMX2 SSE2Fast SSSE3 SSE4.2 AVX AVX2 FMA3 LZCNT BMI2
12.7 fps - 2 x Intel Xeon X5550 @ 3.0GHz ( 4C / 8T ) MMX2 SSE2Fast SSSE3 SSE4.1 Cache64
12.4 fps - AMD Ryzen 5 2200G @ 4.0GHz^ ( 4C / 4T ) MMX2 SSE2Fast SSSE3 SSE4.2 AVX AVX2 FMA3 LZCNT BMI2
12.1 fps - AMD FX-8320 Eight-Core @ 4.32GHz^ ( 4C / 8T ) MMX2 SSE2Fast SSSE3 SSE4.2 AVX XOP FMA4 FMA3 LZCNT BMI1
12.0 fps - Intel Core i5-4460 @ 3.2GHz ( 4C / 4T ) MMX2 SSE2Fast SSSE3 SSE4.2 AVX AVX2 FMA3 LZCNT BMI2
11.2 fps - Intel Core i7-3770K @ 3.5GHz ( 4C / 8T ) MMX2 SSE2Fast SSSE3 SSE4.2 AVX
9.7 fps - Intel Core i3-7100 @ 3.9GHz ( 2C / 4T ) MMX2 SSE2Fast SSSE3 SSE4.2 AVX AVX2 FMA3 LZCNT BMI2
8.8 fps - AMD Ryzen 3 1300X @ 3.5GHz ( 4C / 4T ) MMX2 SSE2Fast SSSE3 SSE4.2 AVX AVX2 FMA3 LZCNT BMI2
8.3 fps - Intel Core i5-2400 @ 3.7GHz^ ( 4C / 4T ) MMX2 SSE2Fast SSSE3 SSE4.2 AVX
7.8 fps - Intel i7-3612QM @ 2.1GHz ( 4C / 8T ) MMX2 SSE2Fast SSSE3 SSE4.2 AVX
7.6 fps - Intel Core i7-7500U @ 3.5GHz^ ( 2C / 4T ) MMX2 SSE2Fast SSSE3 SSE4.2 AVX AVX2 FMA3 LZCNT BMI2
6.1 fps - Intel Celeron G3900 @ 4.0GHz^ ( 2C / 2T ) MMX2 SSE2Fast SSSE3 SSE4.2 LZCNT
5.9 fps - AMD Athlon X4 760K Quad Core @ 4.5GHz^ ( 2C / 4T ) MMX2 SSE2Fast SSSE3 SSE4.2 AVX XOP FMA4 FMA3 LZCNT BMI1
5.6 fps - Intel Pentium G3258 @ 4.2GHz^ ( 2C / 2T ) MMX2 SSE2Fast SSSE3 SSE4.2 LZCNT
5.5 fps - Intel Xeon X5470 @ 3.33GHz ( 4C / 4T ) MMX2 SSE2Fast SSSE3 SSE4.1 Cache64
4.8 fps - Intel Core i3-3220 @ 3.3GHz ( 2C / 4T ) MMX2 SSE2Fast SSSE3 SSE4.2 AVX
4.4 fps - Intel Core2 Quad Q8200 @ 2.8GHz^ ( 4C / 4T ) MMX2 SSE2Fast SSSE3 SSE4.1 Cache64
4.2 fps - Intel Core i3-2100 @ 3.1GHz ( 2C / 4T ) MMX2 SSE2Fast SSSE3 SSE4.2 AVX
3.7 fps - Intel Core2 Quad Q8200 @ 2.33GHz ( 4C / 4T ) MMX2 SSE2Fast SSSE3 SSE4.1 Cache64
3.6 fps - Intel Core i3-4005U @ 1.7GHz ( 2C / 4T ) MMX2 SSE2Fast SSSE3 SSE4.2 AVX AVX2 FMA3 LZCNT BMI2
1.9 fps - AMD Athlon II X4 620 @ 2.6GHz ( 4C / 4T ) MMX2 SSE2Fast LZCNT
Nico8583
25th June 2019, 12:52
Thank you !
So the best way to encode x265 is Ryzen 3000 or i7 8700K/9700K(/i9 9900K but too much expensive) ?
Atak_Snajpera
25th June 2019, 13:30
Thank you !
So the best way to encode x265 is Ryzen 3000 or i7 8700K/9700K(/i9 9900K but too much expensive) ?
It is just a matter of your budget. Ryzen 3600 looks good if you OC to 4.5GHz.
benwaggoner
25th June 2019, 14:47
It is just a matter of your budget. Ryzen 3600 looks good if you OC to 4.5GHz.
Also, doesn't AVX512 help performance with slower presets at 4K resolutions? Optimum processor might differ some based on preset.
Operating costs also factor in, ala pixels per picojoule.
Atak_Snajpera
25th June 2019, 15:26
Also, doesn't AVX512 help performance with slower presets at 4K resolutions? Optimum processor might differ some based on preset.
Operating costs also factor in, ala pixels per picojoule.
Have you seen 9900k with AVX-512 support recently?
Asmodian
25th June 2019, 21:14
The i9-9900X supports AVX-512 (https://ark.intel.com/content/www/us/en/ark/products/189124/intel-core-i9-9900x-x-series-processor-19-25m-cache-up-to-4-50-ghz.html), not the k. :p
It is odd that list of benchmarks isn't using AVX-512 on the i9-7900X or anything else, it would be nice to know how much impact it has.
benwaggoner
25th June 2019, 22:19
The i9-9900X supports AVX-512 (https://ark.intel.com/content/www/us/en/ark/products/189124/intel-core-i9-9900x-x-series-processor-19-25m-cache-up-to-4-50-ghz.html), not the k. :p
It is odd that list of benchmarks isn't using AVX-512 on the i9-7900X or anything else, it would be nice to know how much impact it has.
I'm not sure. Perhaps because avx512 needs to be explicitly turned on?
I've got my dual Xeon 6140 workstation coming next week (basically a physical EC2 c5.16xlarge), and I can do some benchmarks when it arrives.
And I've got some 8K test sources too, so I can benchmark those with/without. I would expect that the value may go up with resolution, as avx512 is only supposed to become useful at 2160p.
Blue_MiSfit
26th June 2019, 03:40
I've got my dual Xeon 6140 workstation coming next week (basically a physical EC2 c5.16xlarge), and I can do some benchmarks when it arrives.
drool...
Asilurr
26th June 2019, 04:12
Just compare those results.
[...]I would advise a healthy dose of skepticism whenever you see results without extensive coverage of the testing methodology. Did they use the same hardware reference frame (same memory modules, same SSD)? Did they use the same software version (and contemporary versions of x265 too, not [e.g.] two years old ones)? Did they use the same parameter configuration? Did they use the same content source? Did they test multiple scenarios (i.e. various bit depths, various chroma subsampling types, various resolutions)? Did they use an OS which allows meaningful comparison of highly-threaded CPUs (not repeating the unfortunate case of Threadrippers on Windows)?
Atak_Snajpera
26th June 2019, 09:58
I would advise a healthy dose of skepticism whenever you see results without extensive coverage of the testing methodology. Did they use the same hardware reference frame (same memory modules, same SSD)? Did they use the same software version (and contemporary versions of x265 too, not [e.g.] two years old ones)? Did they use the same parameter configuration? Did they use the same content source? Did they test multiple scenarios (i.e. various bit depths, various chroma subsampling types, various resolutions)? Did they use an OS which allows meaningful comparison of highly-threaded CPUs (not repeating the unfortunate case of Threadrippers on Windows)?
Oh boy! You and your questions...
Dude! Those results come from my benchmark
http://forum.pclab.pl/topic/1184884-x265-FHD-Benchmark/
NikosD
26th June 2019, 12:14
Ryzen 3000 will have no clock penalty (aka no performance penalty) leveraging AVX2 instructions according to AMD.
So, it's possible Ryzen 3000 to be a lot faster with heavy use of AVX2 compared to all Intel processors so far, due to no clock restrictions.
excellentswordfight
26th June 2019, 18:03
I'm not sure. Perhaps because avx512 needs to be explicitly turned on?
I've got my dual Xeon 6140 workstation coming next week (basically a physical EC2 c5.16xlarge), and I can do some benchmarks when it arrives.
And I've got some 8K test sources too, so I can benchmark those with/without. I would expect that the value may go up with resolution, as avx512 is only supposed to become useful at 2160p.
I’ve done some tests at 2160p with a few skylake-sp platforms. Even at 2160p i had a hard time getting the same utilization as without avx512, and even without that it’s hard to see any big gains cause of the ~500Mhz clockspeed penalty.
RanmaCanada
27th June 2019, 02:31
Oh boy! You and your questions...
Dude! Those results come from my benchmark
http://forum.pclab.pl/topic/1184884-x265-FHD-Benchmark/
It's obvious some people don't do their "research" before they post :D
I did something similar to The Stilt, when he was wrong and just called him a random haha.
Thanks for posting everything in one place! It will make things far more easier to compare once we get some runs from Ryzen 2!
nevcairiel
27th June 2019, 10:52
some runs from Ryzen 2!
Since we're liking to be accurate:
Its either Zen 2, or Ryzen 3000. Ryzen 2 is misleading, and not a term used by AMD, since Ryzen 3/5/7/9 are model-classes in their lineup, so it could easily be mistaken for a lower-end model - and next generation there would be even more real confusion. :)
Asilurr
27th June 2019, 14:52
Dude! Those results come from my benchmark
http://forum.pclab.pl/topic/1184884-x265-FHD-Benchmark/The link you've provided redirects to Mediafire (http://www.mediafire.com/file/nm5n7xg2nb22zl7/x265_FHD_Benchmark.7z), where one can download the presumably latest version (?) of the benchmark. Id est the package uploaded on July 26th, 2018. Inside the package, one can find:
1. An old version of x265 [2.2+15-a18ab7656c30].
2. An old version of FFmpeg [N-82889-g54931fd].
3. The five test sources, which are all 8-bit, all 4:2:0, all 16:9 1080p.
As for the benchmark itself:
1. As one can't test a given CPU in vacuo, one actually tests an ensemble of CPU+MB+RAM+IO. Rephrasing: one adds degrees of freedom to the system, together those serve the purpose of propagating the uncertainty.
2. As one doesn't test x265 on its own, due to the way the benchmark is designed, one actually tests the pair of FFmpeg+x265. Another degree of freedom, another potential path for the propagation of uncertainty.
3. As the benchmark uses an old version of x265, all the results which are obtained are actually "watermarked" by the common reference frame of a 2.2 x265 encoder. One can't extrapolate the results to a 3.1 x265 encoder, and assume them to hold true by default.
4. As the benchmark tests the default configuration of parameters for x265 (i.e. --preset medium), all the results which are obtained are actually "watermarked" by the common reference frame of that particular preset. One can't extrapolate the results to another preset (or another parameter configuration), and assume them to hold true by default.
5. The benchmark doesn't test a single low complexity source, or a single 10-bit source, or a single non-4:2:0 source, or a single non-FHD source. All the results are again "watermarked" by a highly specific encoding scenario. One can't extrapolate them to another encoding scenario, and assume them to hold true by default.
Do you understand what is the actual result of the benchmark you are providing? It enables one to say: I am encoding this particular source of this particular complexity (at this particular resolution, this particular bit depth, this particular chroma subsampling), within this particular software environment (OS, x265, FFmpeg), on this particular machine (CPU, MB, RAM, IO). It certainly does not enable one to say: CPU MMM achieves an average of abc% better FPS than CPU NNN, across all possible/conceivable encoding scenarios.
At this point, I'd simply reiterate my initial statement: I would advise a healthy dose of skepticism whenever you see results without extensive coverage of the testing methodology.
Atak_Snajpera
27th June 2019, 16:03
The link you've provided redirects to Mediafire (http://www.mediafire.com/file/nm5n7xg2nb22zl7/x265_FHD_Benchmark.7z), where one can download the presumably latest version (?) of the benchmark. Id est the package uploaded on July 26th, 2018. Inside the package, one can find:
1. An old version of x265 [2.2+15-a18ab7656c30].
2. An old version of FFmpeg [N-82889-g54931fd].
3. The five test sources, which are all 8-bit, all 4:2:0, all 16:9 1080p.
As for the benchmark itself:
1. As one can't test a given CPU in vacuo, one actually tests an ensemble of CPU+MB+RAM+IO. Rephrasing: one adds degrees of freedom to the system, together those serve the purpose of propagating the uncertainty.
2. As one doesn't test x265 on its own, due to the way the benchmark is designed, one actually tests the pair of FFmpeg+x265. Another degree of freedom, another potential path for the propagation of uncertainty.
3. As the benchmark uses an old version of x265, all the results which are obtained are actually "watermarked" by the common reference frame of a 2.2 x265 encoder. One can't extrapolate the results to a 3.1 x265 encoder, and assume them to hold true by default.
4. As the benchmark tests the default configuration of parameters for x265 (i.e. --preset medium), all the results which are obtained are actually "watermarked" by the common reference frame of that particular preset. One can't extrapolate the results to another preset (or another parameter configuration), and assume them to hold true by default.
5. The benchmark doesn't test a single low complexity source, or a single 10-bit source, or a single non-4:2:0 source, or a single non-FHD source. All the results are again "watermarked" by a highly specific encoding scenario. One can't extrapolate them to another encoding scenario, and assume them to hold true by default.
Do you understand what is the actual result of the benchmark you are providing? It enables one to say: I am encoding this particular source of this particular complexity (at this particular resolution, this particular bit depth, this particular chroma subsampling), within this particular software environment (OS, x265, FFmpeg), on this particular machine (CPU, MB, RAM, IO). It certainly does not enable one to say: CPU MMM achieves an average of abc% better FPS than CPU NNN, across all possible/conceivable encoding scenarios.
At this point, I'd simply reiterate my initial statement: I would advise a healthy dose of skepticism whenever you see results without extensive coverage of the testing methodology.
Are you trying to convince me that AMD's AVX128 in Zen1 is able to compete with full fat Intel's AVX256. x256 is highly optimized for AVX2 and that's why ZEN1 sucks in this encoder. Period! Luckily Zen2 finally has AVX256 so I'm expecting 3600 to have similar performance as stock 8700k.
BTW. Your requirements for "proper" benchmark are just insane and totally unrealistic! I advise you to stop using any software because there are countless variables affecting performance of your machine (including phases of the moon).
Asilurr
28th June 2019, 06:01
You're being deliberately obtuse, I suppose I can attempt a simplistic analogy by referring to a system with only two degrees of freedom.
I want to investigate what happens to water when I heat it up. To be rigorous, that's the so-called "pure water" which is known today as UPW (https://en.wikipedia.org/wiki/Ultrapure_water). At a pressure of 1 atm, I observe that water is gaseous at 400 K. Conducting a second experiment, I observe that water is gaseous at 500 K too. At this point I ask myself: can I consider the previous results a good predictor of what would happen to water when it's heated to 600 K instead? Can I simply assume that it would be gaseous, thus not having to actually perform a test? The answer to that is no, a proper phase diagram (https://upload.wikimedia.org/wikipedia/commons/0/08/Phase_diagram_of_water.svg) shows what happens to water heated at 600 K and highlights the critical importance of the other parameter of this simplistic system.
To return to the discussion of benchmarking encoding times, we're now looking at a system with dozens of degrees of freedom: the parametric space of the hardware, the parametric space of the software, the parametric space of the source to be encoded, the parametric space of the encoding process itself. This entire thread was initiated by explicitly mentioning UHD/4K encoding scenarios. It's right in front of your eyes, mentioned both in the title and the leading post. To which you've replied (here (http://forum.doom9.org/showthread.php?p=1877898#post1877898)) by implicitly assuming that the results of FHD benchmarking scenarios (highly specific FHD encoding scenarios, as already pointed out) will simply hold true even if one particular parameter varies drastically, claiming that increasing the resolution four times doesn't alter the expected result. To which I replied that an implicit assumption is not sound, as there is an explicit need of additional testing to account for the variable parameter(s).
NikosD
28th June 2019, 11:24
To return to the discussion of benchmarking encoding times, we're now looking at a system with dozens of degrees of freedom: the parametric space of the hardware, the parametric space of the software, the parametric space of the source to be encoded, the parametric space of the encoding process itself. This entire thread was initiated by explicitly mentioning UHD/4K encoding scenarios. It's right in front of your eyes, mentioned both in the title and the leading post. To which you've replied (here (http://forum.doom9.org/showthread.php?p=1877898#post1877898)) by implicitly assuming that the results of FHD benchmarking scenarios (highly specific FHD encoding scenarios, as already pointed out) will simply hold true even if one particular parameter varies drastically, claiming that increasing the resolution four times doesn't alter the expected result. To which I replied that an implicit assumption is not sound, as there is an explicit need of additional testing to account for the variable parameter(s). Hello.
Could you propose an existent benchmark, covering your needs of proper x265 encoding benchmarking procedure for multithreaded CPUs ?
Or do you have a suggestion how to build one ?
Atak_Snajpera
28th June 2019, 11:57
You're being deliberately obtuse, I suppose I can attempt a simplistic analogy by referring to a system with only two degrees of freedom.
I want to investigate what happens to water when I heat it up. To be rigorous, that's the so-called "pure water" which is known today as UPW (https://en.wikipedia.org/wiki/Ultrapure_water). At a pressure of 1 atm, I observe that water is gaseous at 400 K. Conducting a second experiment, I observe that water is gaseous at 500 K too. At this point I ask myself: can I consider the previous results a good predictor of what would happen to water when it's heated to 600 K instead? Can I simply assume that it would be gaseous, thus not having to actually perform a test? The answer to that is no, a proper phase diagram (https://upload.wikimedia.org/wikipedia/commons/0/08/Phase_diagram_of_water.svg) shows what happens to water heated at 600 K and highlights the critical importance of the other parameter of this simplistic system.
To return to the discussion of benchmarking encoding times, we're now looking at a system with dozens of degrees of freedom: the parametric space of the hardware, the parametric space of the software, the parametric space of the source to be encoded, the parametric space of the encoding process itself. This entire thread was initiated by explicitly mentioning UHD/4K encoding scenarios. It's right in front of your eyes, mentioned both in the title and the leading post. To which you've replied (here (http://forum.doom9.org/showthread.php?p=1877898#post1877898)) by implicitly assuming that the results of FHD benchmarking scenarios (highly specific FHD encoding scenarios, as already pointed out) will simply hold true even if one particular parameter varies drastically, claiming that increasing the resolution four times doesn't alter the expected result. To which I replied that an implicit assumption is not sound, as there is an explicit need of additional testing to account for the variable parameter(s).
Yes! You can easily extrapolate results from 1920x1080 to 3840x2160. Encoding speed will drop 4 times. 4x more pixels to examine means 4 times more cpu cycles has to be used. End of story. I'm done with you!
RanmaCanada
29th June 2019, 04:31
Since we're liking to be accurate:
Its either Zen 2, or Ryzen 3000. Ryzen 2 is misleading, and not a term used by AMD, since Ryzen 3/5/7/9 are model-classes in their lineup, so it could easily be mistaken for a lower-end model - and next generation there would be even more real confusion. :)
Ya got me there! and thank you.
Now hopefully someone will leak out something other than stupid geekbench results. I seriously want to see how well this generation will do in encoding. I don't care about games..I want encoding benchmarks.
Sadly, I think we are going to have to wait till one of the forum members gets one as we know review sites don't exactly know how to benchmark when it's not games.
benwaggoner
1st July 2019, 22:00
Yes! You can easily extrapolate results from 1920x1080 to 3840x2160. Encoding speed will drop 4 times. 4x more pixels to examine means 4 times more cpu cycles has to be used. End of story.
Not entirely. CABAC can take up a decent chunk of CPU, and since bitrate isn't linear to pixel count, in the real world you get less than a 4x increase for 4x the pixels. Plus the more pixels there are, the less each one matters so some different encoding technique.
Changing resolution can also really impact threading; more pixels mean fewer frame threads are needed on smaller core counts, reducing overhead.
And of course, how well does encoding scale across multiple cores? Across the same or different NUMA nodes?
mandarinka
4th July 2019, 04:06
https://www.ptt.cc/bbs/PC_Shopping/M.1561539742.A.61C.html
This has a (unconfirmed) results for x265 benchmark used on HWBot. I think it has an older binary so gains from AVX256 might be a bit subdued, but the same would be true for the Intel chip.
The benchmark tests and reports max single thread turbo clock at initialization so I think the screenshot means that these are scores for stock Ryzen 5 3600 and Core i7-8700K.
Basically looks nice for Ryzen 3000 and I would wait for it and not get anything else, if this is confirmed.
Asmodian
4th July 2019, 04:28
Very nice results, I hope they are true! It is looking like they probably are, with more leaks from other areas (https://www.gamersnexus.net/news-pc/3485-hw-news-intel-asks-if-its-screwed-displayport-20-and-more).
Forteen88
4th July 2019, 08:22
https://www.ptt.cc/bbs/PC_Shopping/M.1561539742.A.61C.htmlUnfortunately, they don't give information about the speed of the RAM used.
mandarinka
5th July 2019, 22:29
https://imgur.com/a/YkoOCgM
Check that Handbrake result :) Pretty nice if correct.
Nico8583
6th July 2019, 11:19
Interesting :)
What is the difference between i9 9900K 95W and i9 9900K ? 95W is the stock TDP and both are 3,6Ghz. Perhaps the second is O/C ?
sneaker_ger
6th July 2019, 11:25
3.6 GHz is only the advertised base clock. The i9-9900K goes up to 5 GHz dynamically. If it's not limited to the default "95W" setting it can keep turbo clocks longer.
https://www.anandtech.com/show/13591/the-intel-core-i9-9900k-at-95w-fixing-the-power-for-sff
benwaggoner
6th July 2019, 20:59
Very nice results, I hope they are true! It is looking like they probably are, with more leaks from other areas (https://www.gamersnexus.net/news-pc/3485-hw-news-intel-asks-if-its-screwed-displayport-20-and-more).
It's be helpful if they documented what parameters were actually being tested. "1080p" is obviously not a x265 --preset.
hajj_3
6th July 2019, 21:37
It's be helpful if they documented what parameters were actually being tested. "1080p" is obviously not a x265 --preset.
Zen2 chips and the reviews are out tomorrow so not long until we have some good x265 and x264 benchmarks for zen2 to compare with latest intel chips.
hajj_3
7th July 2019, 14:19
x265 and x264 benchmarks of the new Ryzen 3000 series chips:
https://images.anandtech.com/graphs/graph14605/111203.png
https://images.anandtech.com/graphs/graph14605/111201.png
https://img.purch.com/r/711x457/aHR0cDovL21lZGlhLmJlc3RvZm1pY3JvLmNvbS9VL04vODQ0Nzk5L29yaWdpbmFsL2ltYWdlMDEzLnBuZw==
https://img.purch.com/r/711x457/aHR0cDovL21lZGlhLmJlc3RvZm1pY3JvLmNvbS9VL1AvODQ0ODAxL29yaWdpbmFsL2ltYWdlMDEyLnBuZw==
handbrake v1.2.2:
https://techreport.com/r.x/2019_07_06_AMD_s_Ryzen_7_3700X_and_Ryzen_9_3900X_CPUs_reviewed/Graphs_handbrake.png
https://hothardware.com/ContentImages/Article/2873/content/x265.png
staxrip x264 and x265 benchmarks: https://nl.hardware.info/reviews/9397/10/amd-ryzen-7-3700x-a-ryzen-9-3900x-review-intel-voorbij-benchmarks-video--en-audio-encoding-x264-x265-en-flac
birdie
7th July 2019, 15:11
The most relevant graph for encoding/rendering:
https://techreport.com/wp-content/uploads/2019/07/power-taskenergy-FIXED.png
Atak_Snajpera
7th July 2019, 16:24
https://p.xfastest.com/~sinchen/AMD-Ryzen-7-3700X-9-3900X/AMD-Ryzen-7-3700X-9-3900X-34.jpg
Boulder
7th July 2019, 18:08
Looks like a 3700X is the new 1700X as the best bang for buck. Hopefully we'll see a benchmark for the most used presets like slow, slower and veryslow.
We should really thank AMD for keeping the socket the same for all these years, upgrading gets very cheap.
hajj_3
7th July 2019, 18:12
Looks like a 3700X is the new 1700X as the best bang for buck. Hopefully we'll see a benchmark for the most used presets like slow, slower and veryslow.
We should really thank AMD for keeping the socket the same for all these years, upgrading gets very cheap.
3600 is the best bang for buck, 6 cores, 12 threads - $199. Those weren't given to reviewers though as AMD obviously wants people to buy their more expensive chips.
Stereodude
7th July 2019, 18:40
Do any of these sites run any sort of standardized x265 test that I can download and run on my system to see how my current system compares to the new Ryzen 3xxx CPUs?
Atak_Snajpera
7th July 2019, 18:56
Do any of these sites run any sort of standardized x265 test that I can download and run on my system to see how my current system compares to the new Ryzen 3xxx CPUs?
Guys from https://news.xfastest.com/review/review-02/66816/amd-ryzen-7-3700x-ryzen-9-3900x-review/
used my x265 fhd benchmark
http://forum.pclab.pl/topic/1184884-x265-FHD-Benchmark/
That guy also used the same benchmark but instead of fps showed elapsed time in seconds. To get fps you have to do this 2500/seconds.
https://youtu.be/rcFUGSElZEs?t=8m18s
Stereodude
7th July 2019, 19:34
Guys from https://news.xfastest.com/review/review-02/66816/amd-ryzen-7-3700x-ryzen-9-3900x-review/
used my x265 fhd benchmark
http://forum.pclab.pl/topic/1184884-x265-FHD-Benchmark/
That guy also used the same benchmark but instead of fps showed elapsed time in seconds. To get fps you have to do this 2500/seconds.
https://youtu.be/rcFUGSElZEs?t=8m18s
Thanks! Is the build of x265 in your package new enough to take advantage of all the new (to AMD) features in the Zen 2 architecture?
Intel Xeon E5-2687W v2 @ 3.4GHz ( 8C / 16T )
y4m [info]: 1920x1080 fps 50/1 i420p8 sar 1:1 unknown frame count
raw [info]: output file: NUL
x265 [info]: HEVC encoder version 2.2+15-a18ab7656c30
x265 [info]: build info [Windows][GCC 6.2.0][64 bit] 8bit
x265 [info]: using cpu capabilities: MMX2 SSE2Fast SSSE3 SSE4.2 AVX
x265 [info]: Main Still Picture profile, Level-4.1 (Main tier)
x265 [info]: Thread pool created using 16 threads
x265 [info]: Slices : 1
x265 [info]: frame threads / pool features : 5 / wpp(17 rows)
x265 [info]: Coding QT: max CU size, min CU size : 64 / 8
x265 [info]: Residual QT: max TU size, max depth : 32 / 1 inter / 1 intra
x265 [info]: ME / range / subpel / merge : hex / 57 / 2 / 2
x265 [info]: Keyframe min / max / scenecut / bias: 50 / 500 / 40 / 5.00
x265 [info]: Lookahead / bframes / badapt : 20 / 4 / 2
x265 [info]: b-pyramid / weightp / weightb : 1 / 1 / 0
x265 [info]: References / ref-limit cu / depth : 3 / on / on
x265 [info]: AQ: mode / str / qg-size / cu-tree : 1 / 1.0 / 32 / 1
x265 [info]: Rate Control / qCompress : CRF-28.0 / 0.60
x265 [info]: tools: rd=3 psy-rd=2.00 rskip signhide tmvp strong-intra-smoothing
x265 [info]: tools: lslices=6 deblock sao
encoded 2500 frames in 131.23s (19.05 fps), 7025.74 kbps, Avg QP:37.21
So the Ryzen 3900X is ~3x as fast with x265 encoding as my Xeon E5-2687W v2? :confused:
I realize this is only one particular x265 test scenario, but if I can get a 3x speed improvement with the much slower preset I normally use for 1080p x265 I should be in the car heading to Microcenter to buy a Ryzen 3900X system right now.
Atak_Snajpera
7th July 2019, 19:44
Thanks! Is the build of x265 in your package new enough to take advantage of all the new (to AMD) features in the Zen 2 architecture?
Intel Xeon E5-2687W v2 @ 3.4GHz ( 8C / 16T )
y4m [info]: 1920x1080 fps 50/1 i420p8 sar 1:1 unknown frame count
raw [info]: output file: NUL
x265 [info]: HEVC encoder version 2.2+15-a18ab7656c30
x265 [info]: build info [Windows][GCC 6.2.0][64 bit] 8bit
x265 [info]: using cpu capabilities: MMX2 SSE2Fast SSSE3 SSE4.2 AVX
x265 [info]: Main Still Picture profile, Level-4.1 (Main tier)
x265 [info]: Thread pool created using 16 threads
x265 [info]: Slices : 1
x265 [info]: frame threads / pool features : 5 / wpp(17 rows)
x265 [info]: Coding QT: max CU size, min CU size : 64 / 8
x265 [info]: Residual QT: max TU size, max depth : 32 / 1 inter / 1 intra
x265 [info]: ME / range / subpel / merge : hex / 57 / 2 / 2
x265 [info]: Keyframe min / max / scenecut / bias: 50 / 500 / 40 / 5.00
x265 [info]: Lookahead / bframes / badapt : 20 / 4 / 2
x265 [info]: b-pyramid / weightp / weightb : 1 / 1 / 0
x265 [info]: References / ref-limit cu / depth : 3 / on / on
x265 [info]: AQ: mode / str / qg-size / cu-tree : 1 / 1.0 / 32 / 1
x265 [info]: Rate Control / qCompress : CRF-28.0 / 0.60
x265 [info]: tools: rd=3 psy-rd=2.00 rskip signhide tmvp strong-intra-smoothing
x265 [info]: tools: lslices=6 deblock sao
encoded 2500 frames in 131.23s (19.05 fps), 7025.74 kbps, Avg QP:37.21
So the Ryzen 3900X is ~3x as fast with x265 encoding as my Xeon E5-2687W v2? :confused:
I realize this is only one particular x265 test scenario, but if I can get a 3x speed improvement with the much slower preset I normally use for 1080p x265 I should be in the car heading to Microcenter to buy a Ryzen 3900X system right now.
Ivybridge does not have AVX2 hence poor performance in x265.
Stereodude
7th July 2019, 20:08
Ivybridge does not have AVX2 hence poor performance in x265.
Well, the x264 results for the 3900X are almost double this system too and AFAIK AVX2 makes less of a difference with x264.
encoded 2500 frames, 38.21 fps, 22398.07 kb/s
benwaggoner
7th July 2019, 20:16
The most relevant graph for encoding/rendering:
https://techreport.com/r.x/2019_07_06_AMD_s_Ryzen_7_3700X_and_Ryzen_9_3900X_CPUs_reviewed/power-taskenergy-FIXED.png
The picture's URL is 404.
Stereodude
7th July 2019, 20:18
The picture's URL is 404.
Works here.
RanmaCanada
7th July 2019, 21:40
The 3900x is impressive, but I would also like to see some real world results. Benchmarks are great, but no one uses crf 28 to encode, or at least they shouldn't haha. I as well would like to see some crf 18-20 at slow, slower, very slow, slowest and placebo. Right now I encode my TV/Movie blurays on slow (1080P gets 5-7 fps), and anime on slower (1080P gets 0.8-1 fps if I am lucky!), with a Ryzen 2700. I haven't tried 4k yet as I don't own a compatible drive to rip the few 4k movies I own.
I score 21.97 fps on Atak's benchmark, and 3.03 on Sagitare's.
mandarinka
7th July 2019, 21:46
Some other x265 tests I found:
https://www.techpowerup.com/review/amd-ryzen-9-3900x/13.html
TPU seems to meassure plain x265 and x264 at crf mode and slow preset, instead of frontends like handbrake.
https://nl.hardware.info/reviews/9397/10/amd-ryzen-7-3700x-a-ryzen-9-3900x-review-intel-voorbij-benchmarks-video--en-audio-encoding-x264-x265-en-flac
Staxrip x264 + x265. They have interactive graphs and comparison with a lot of CPUs.
birdie
7th July 2019, 23:06
The picture's URL is 404.
They changed the URL after I posted it. I've edited my post and fixed the URL.
Stereodude
8th July 2019, 00:13
The 3900x is impressive, but I would also like to see some real world results. Benchmarks are great, but no one uses crf 28 to encode, or at least they shouldn't haha. I as well would like to see some crf 18-20 at slow, slower, very slow, slowest and placebo. Right now I encode my TV/Movie blurays on slow (1080P gets 5-7 fps), and anime on slower (1080P gets 0.8-1 fps if I am lucky!), with a Ryzen 2700. I haven't tried 4k yet as I don't own a compatible drive to rip the few 4k movies I own.
I score 21.97 fps on Atak's benchmark, and 3.03 on Sagitare's.
TPU used CRF 20 and slow and the 3900X looks to be almost twice as fast as your 2700.
I've been encoding 1080p at CRF 16 placebo 10-bit (8-bit source). It's ~0.25fps on my 8 or 10 core Ivy Bridge E5's. I have a lot of cores, but an old'ish CPU architecture.
nevcairiel
8th July 2019, 00:20
3600 is the best bang for buck, 6 cores, 12 threads - $199. Those weren't given to reviewers though as AMD obviously wants people to buy their more expensive chips.
GamersNexus reviewed a 3600
https://www.gamersnexus.net/hwreviews/3489-amd-ryzen-5-3600-cpu-review-benchmarks-vs-intel
RanmaCanada
8th July 2019, 04:40
TPU used CRF 20 and slow and the 3900X looks to be almost twice as fast as your 2700.
I've been encoding 1080p at CRF 16 placebo 10-bit (8-bit source). It's ~0.25fps on my 8 or 10 core Ivy Bridge E5's. I have a lot of cores, but an old'ish CPU architecture.
No where do they list how long the video is they use to encode, so it's pretty much useless. Yes it does show it is almost twice as fast, but fps is a far better indicator of speed. For all we know they could have gotten 8fps on the 3900x and 4.5fps on the 2700x. Time to complete something means nothing without a reference point, especially when we have no idea what the reference material was. Was it an extremely complex scene, or was it just a test pattern (which is stupidly easy).
I know not everyone will agree with me, but as someone who encodes and encodes, fps is far more important to me as every fps gained is an extra 2 or so minutes of footage encoded per hour (if I did the math right). Hence why I would like the actual results in that instead of time.
The most relevant graph for encoding/rendering:
https://techreport.com/wp-content/uploads/2019/07/power-taskenergy-FIXED.png
Thank you for that! :)
You are completely right, everything else is irrelevant.
Those Ryzen 3000 efficiency results are extremely impressive.
mariush
8th July 2019, 10:05
Slightly off topic but apparently the hardware HEVC encoder in the Navi RX 5700 is way powerful, much faster than h264 implementatin
See https://www.youtube.com/watch?v=Yi-_T3vsv-Q and https://www.youtube.com/watch?v=fY2fbAzFiUE
Atak_Snajpera
8th July 2019, 18:00
Ryzen 1600 vs 2600 vs 3600 in x265 FHD Benchmark
https://www.youtube.com/watch?v=-ZrurVMEAVs&feature=youtu.be&t=3m28s
Stereodude
8th July 2019, 18:37
Ryzen 1600 vs 2600 vs 3600 in x265 FHD Benchmark
https://www.youtube.com/watch?v=-ZrurVMEAVs&feature=youtu.be&t=3m28s
The video doesn't even have the right screenshots with the right label and it's rather hard to see the results. To save someone else the difficulty...
Ryzen 5 1600 - 18.73 fps
Ryzen 5 2600 - 21.10 fps
Ryzen 5 3600 - 32.71 fps
The Ryzen 9 3900X has twice as many cores as the Ryzen 5 3600 and is almost twice as fast. It looks to be pretty close to linear. So the Ryzen 9 3950X should be about 33% faster than the 3900X in the benchmark (presuming the TDP doesn't clobber it).
Atak_Snajpera
8th July 2019, 18:49
It is easy to predict that 3950x will be on top of the chart with score around 70fps beating 2 x Intel Xeon E5-4660 v3 @ 2.9GHz ( 14C / 28T ). I just can't wait to see TR 64C/128T result ;) However I personally doubt that five x265 encoders running simultaneously will be enough to feed that beast.
Nico8583
8th July 2019, 19:19
So after Ryzen 3000 release, do you think 3700x is a good CPU to encode to x265 ? Or do you think it's better to wait 9900K at a lower price ? Thank you !
NikosD
8th July 2019, 19:47
Ryzen 5 1600 - 18.73 fps
Ryzen 5 2600 - 21.10 fps
Ryzen 5 3600 - 32.71 fps
So, in other words Ryzen 5 2600 is ~13% faster in x265 encoding than Ryzen 5 1600, but Ryzen 5 3600 is ~55% faster than Ryzen 5 2600 and ~75% faster than Ryzen 5 1600.
Faster clock and mainly AVX2 implementation makes big difference for x265 obviously.
Atak_Snajpera
8th July 2019, 19:53
So after Ryzen 3000 release, do you think 3700x is a good CPU to encode to x265 ? Or do you think it's better to wait 9900K at a lower price ? Thank you !
Again
https://p.xfastest.com/~sinchen/AMD-Ryzen-7-3700X-9-3900X/AMD-Ryzen-7-3700X-9-3900X-34.jpg
RanmaCanada
8th July 2019, 20:36
So after Ryzen 3000 release, do you think 3700x is a good CPU to encode to x265 ? Or do you think it's better to wait 9900K at a lower price ? Thank you !
If you're going to debate about that, might as well get the 3900x as it's the same price as a 9900k, with far better performance. I mean if the 9900k was in your budget, and you're not just looking for a reason to cheap out :P haha.
Nico8583
8th July 2019, 20:39
Thank you, I just ordered a Ryzen 7 3700X :D now I'm searching a motherboard and DDR4. I think I'll go with a B450M. Do you know the motherboard and RAM used on this test ?
hajj_3
8th July 2019, 22:52
Thank you, I just ordered a Ryzen 7 3700X :D now I'm searching a motherboard and DDR4. I think I'll go with a B450M. Do you know the motherboard and RAM used on this test ?
ryzen 3000 chips won't work on some non-x570 motherboards at all and some won't work without a bios update which may require an older am4 cpu to flash the bios. Research boards before ordering.
Nico8583
8th July 2019, 23:48
I found the MSI B450 Mortar and it can be flashed without CPU ;)
Boulder
9th July 2019, 09:13
Regarding memory, take a look at this:
https://i.ibb.co/7YHtnML/upload-2019-7-5-8-9-28.png (https://ibb.co/YR6Bcnw)
No need to go overboard with memory, the difference is quite small but you could end up paying quite a lot extra. Personally I run my 24GB at 2866C14, which means the calculated latency is ~9.76ns. I'm not going to buy faster memory in case I upgrade to Zen 2, unless I get a real bargain.
NikosD
9th July 2019, 12:29
From rigaya, the developer of all modern, free and best hardware video encoders
https://blog-imgs-130.fc2.com/r/i/g/rigaya34589/benchmark_3700x_default_x264.png
https://blog-imgs-130.fc2.com/r/i/g/rigaya34589/benchmark_3700x_default_x265.png
mariush
9th July 2019, 13:54
I found the MSI B450 Mortar and it can be flashed without CPU ;)
Look at this list of motherboards and pick one that has at least 2 green squares : https://docs.google.com/spreadsheets/d/1d9_E3h8bLp-TXr-0zTJFqqVxdCR9daIVNyMatydkpFA/htmlview?sle=true#gid=639584818
Nico8583
9th July 2019, 16:15
Look at this list of motherboards and pick one that has at least 2 green squares : https://docs.google.com/spreadsheets/d/1d9_E3h8bLp-TXr-0zTJFqqVxdCR9daIVNyMatydkpFA/htmlview?sle=true#gid=639584818
Thank you, I already looked at it and I'll go with the MSI B450M Mortar :)
Nico8583
9th July 2019, 16:18
From rigaya, the developer of all modern, free and best hardware video encoders
https://blog-imgs-130.fc2.com/r/i/g/rigaya34589/benchmark_3700x_default_x264.png
https://blog-imgs-130.fc2.com/r/i/g/rigaya34589/benchmark_3700x_default_x265.png
Thank you, very interesting.
IMHO the 2700X is missing in this benchmark, it would be interesting to compare 2700X and 3700X.
Atak_Snajpera
9th July 2019, 16:43
Thank you, very interesting.
IMHO the 2700X is missing in this benchmark, it would be interesting to compare 2700X and 3700X.
2700x will be only slightly faster than 1700. 2700x still has avx128.
RanmaCanada
10th July 2019, 01:25
2700x will be only slightly faster than 1700. 2700x still has avx128.
This. I have a 2700 and it's not much faster than the 1700 that is in these benchmarks. In fact I am currently doing a 1080p slow encode and only getting 6.07 fps, which matches up with their results.
Time to upgrade! So much great information in this thread.
Nintendo Maniac 64
10th July 2019, 07:19
Thank you, I already looked at it and I'll go with the MSI B450M Mortar :)
Just keep in mind that MSI decided to be weird and requires you to first install BIOS v17 before you're able to successfully install BIOS v18 (which is the newest).
And yes this applies to when using their CPU-less BIOS flash function as well.
EDIT: It's since come to my attention that the v17 and v18 version numbering only applies to the B450 Tomahawk.
Therefore, a more accurate way of phrasing my statement would be that MSI requires the user to first install the March 2019 BIOS before one can then successfully install the latest June/July 2019 BIOS.
Nico8583
10th July 2019, 07:45
Just keep in mind that MSI decided to be weird and requires you to first install BIOS v17 before you're able to successfully install BIOS v18 (which is the newest).
And yes this applies to when using their CPU-less BIOS flash function as well.
Thank you so I must install v17 BIOS with CPU-less function ? And I don't see v18 on MSI Website, I can see only v17 and it supports Ryzen 3000 :
- Support Ryzen 3000 series CPU.
Ryzen 9 3900X/Ryzen 7 3800X/Ryzen 7 3700X/Ryzen 5 3600X/Ryzen 5 3600/Ryzen 5 3400G/Ryzen 3 3200G
Nintendo Maniac 64
10th July 2019, 08:50
Thank you so I must install v17 BIOS with CPU-less function ? And I don't see v18 on MSI Website, I can see only v17 and it supports Ryzen 3000 :
- Support Ryzen 3000 series CPU.
Ryzen 9 3900X/Ryzen 7 3800X/Ryzen 7 3700X/Ryzen 5 3600X/Ryzen 5 3600/Ryzen 5 3400G/Ryzen 3 3200G
Oh, hmm. Comparing it to MSI's more popular motherboards like the Tomahawk, it looks like the version numbering on the Mortar is one value behind for some reason...
Therefore you actually might have to install BIOS v16 first before being able to successfully install BIOS v17.
And for reference, you would have to do this even if you were updating the BIOS via the traditional method.
EDIT: OK yeah, it looks like the v17 and v18 version numbering only applies to the B450 Tomahawk.
Therefore, a more accurate way of phrasing my previous statement would be that MSI requires the user to first install the March 2019 BIOS before one can then successfully install the latest June/July 2019 BIOS.
benwaggoner
10th July 2019, 18:01
And bear in mind that fastest is going to depend more on per-core performance for lower resolutions, but will be near linear to number of cores for higher resolution. My new Dual Intel Xeon Gold 6140 workstation should be arriving tomorrow. I look forward to testing how well it scales at 4K and 8K.
Nintendo Maniac 64
10th July 2019, 23:43
Just came across some more x265 benchmark results via Handbrake 1.2.2:
https://techgage.com/article/amd-ryzen-7-3700x-ryzen-9-3900x-workstation-performance/3/[/url]"]
Protip: click the image to enlarge
https://archive.is/eZQ3R/0ddea02ae31ec821f760201184bc1ca2f239d7f8.png (https://archive.is/Q87Bx/83e2fee17c7b4daade7218f7cef6b667264de4db.png)
Stereodude
11th July 2019, 13:42
And bear in mind that fastest is going to depend more on per-core performance for lower resolutions, but will be near linear to number of cores for higher resolution. My new Dual Intel Xeon Gold 6140 workstation should be arriving tomorrow. I look forward to testing how well it scales at 4K and 8K.
What's lower and higher resolution in the context of your comment? Lower is 1080p and less higher is 2160p and up?
Atak_Snajpera
11th July 2019, 18:45
Here is good explanation why Zen2 is soooo good
https://i.imgsafe.org/63/63a2e61a55.png
FlopsCPU
http://forum.pclab.pl/topic/1105978-FlopsCPU-klasyczny-benchmark-z-1992-roku-w-nowej-oprawie/
Stereodude
11th July 2019, 19:43
From rigaya, the developer of all modern, free and best hardware video encoders
https://blog-imgs-130.fc2.com/r/i/g/rigaya34589/benchmark_3700x_default_x264.png
https://blog-imgs-130.fc2.com/r/i/g/rigaya34589/benchmark_3700x_default_x265.png
Based on that the performance differences with x265 and x264 become larger with the slower presets, not smaller.
Atak_Snajpera
11th July 2019, 20:05
Based on the that the performance differences with x265 and x264 become larger with the slower presets, not smaller.
Most likely due to better cpu utilization on slower presets.
jd17
12th July 2019, 07:14
Does anyone know more than just rumors regarding Zen2 APUs?
Will we get some before the end of the year?
I would really like to upgrade to Zen2, but I also like iGPUs, especially considering global system and idle efficiency.
It's hard to beat my i5-7500 in that regard.
The 3400G is a bit of a slap in the face to be honest...
nevcairiel
12th July 2019, 10:36
Does anyone know more than just rumors regarding Zen2 APUs?
Will we get some before the end of the year?
The "3000" Zen+ APUs only released quite recently, so it will probably be quite a while.
Atak_Snajpera
12th July 2019, 10:36
Apus are always 1gen behind. You will have to wait for zen3 to gdy zen2 apu.
jd17
12th July 2019, 12:03
My hope was based on this:
https://wccftech.com/exclusive-amds-plans-for-7nm-ryzen-apus/
If true, 7nm APUs could come earlier than their predecessors.
I just hope it's actually an AMD source and not just rumors...
nevcairiel
12th July 2019, 12:07
Don't get your hopes up on so-called leaks from wccftech. :)
Look at their previous "leaks" from a supposed source inside AMD, majority has been proven quite wrong once the actual releases rolled around.
jd17
12th July 2019, 12:25
OK then... I'll try to live with the 7500 a year longer...
Lyris
15th July 2019, 06:40
Wow - so the new Ryzen chips are faster than last year’s best Threadripper? Will have to keep an eye on this.
NikosD
15th July 2019, 08:21
In general, the IPC of AMD's Ryzen 3000 series is higher even than latest Coffee Lake 9000 series from Intel after many, many years.
So, in order to find the fastest CPU you have to look at the number of threads utilized by the app (single threaded vs multi threaded) and the clock.
Generally speaking, Intel has still the clock advantage but AMD has both IPC and more threads than Intel.
x265 encoding runs a lot faster nowadays for Ryzen 3000 than Core 9000 at the same price:
Ryzen 3900X vs Core i9 9900K
benwaggoner
17th July 2019, 00:17
What's lower and higher resolution in the context of your comment? Lower is 1080p and less higher is 2160p and up?
It's relative to the number of cores in question, and the quality/speed tradeoffs as well. But yeah, HD versus UHD is a pretty good demarcation. The jump from 1080p to 2160p is a lot bigger than 720p to 1080p. And with HEVC WPP gives us parallelism proportional to frame height.
blublub
21st July 2019, 20:05
Ohh I am so excited about the new TR generation. If I get a 16c or 24c clocking with 4,3Ghz PBO that'll kick ass....
NikosD
22nd July 2019, 18:35
Ohh I am so excited about the new TR generation. If I get a 16c or 24c clocking with 4,3Ghz PBO that'll kick ass.... You have to wait for September for 16C and October for 24C/32C/48C/64C.
Stereodude
22nd July 2019, 19:59
Isn't the 16C in September the 3950x, not a Threadripper II?
Asmodian
22nd July 2019, 20:03
Yes, that will be a normal consumer chip with few PCIe and memory lanes.
Stereodude
22nd July 2019, 20:07
Yes, that will be a normal consumer chip with few PCIe lanes and memory lanes.
But still has $300-400 motherboards. :(
Asmodian
22nd July 2019, 20:21
But they do have PCIe 4.0. :)
Stereodude
22nd July 2019, 20:58
But they do have PCIe 4.0. :)
Which I don't think is of much use for video encoding. However, there seems to be about a 10% performance hit using a X470 chipset mobo instead of a X570 for a new Ryzen 3xxx for video encoding. :confused:
Nintendo Maniac 64
22nd July 2019, 21:48
But still has $300-400 motherboards. :(
You shouldn't need a $300+ motherboard unless you're going to be overclocking that 3950X (source (https://www.youtube.com/watch?v=zuyuS04lD4o)).
mandarinka
22nd July 2019, 22:35
You have to wait for September for 16C and October for 24C/32C/48C/64C.
Yeah, 3950X comes out in september (official info), Threadripper possibly in october (unofficial rumor from Taiwan).
But still has $300-400 motherboards. :(
MSI B450 Tomahawk (or B450 Tomahawk Max which has larger flash for BIOS, so that fufure BIOS updates shoudl be more painless, also this board is made for Ryzen 3000 out of the box) is rather affordable.
Or Asus Crosshair VI Hero (X370) - that one has very beefy VRM and is deeply discounted if you can find it.
Stereodude
23rd July 2019, 03:10
MSI B450 Tomahawk (or B450 Tomahawk Max which has larger flash for BIOS, so that fufure BIOS updates shoudl be more painless, also this board is made for Ryzen 3000 out of the box) is rather affordable.
Or Asus Crosshair VI Hero (X370) - that one has very beefy VRM and is deeply discounted if you can find it.
If you want to lose ~12% of your Ryzen 3xxx's performance right off the top I guess you could use a B450 motherboard.
excellentswordfight
23rd July 2019, 08:03
If you want to lose ~12% of your Ryzen 3xxx's performance right off the top I guess you could use a B450 motherboard.
Please provide a source to that claim, cause I have seen lots of people running older mb with the new 3000-series and non have reported lower clockspeeds/performance. As long as your mb can provide enough power, the only loss is pcie 4, which I guess most of us can live without.
Stereodude
23rd July 2019, 11:31
Please provide a source to that claim, cause I have seen lots of people running older mb with the new 3000-series and non have reported lower clockspeeds/performance. As long as your mb can provide enough power, the only loss is pcie 4, which I guess most of us can live without.
Yup, I just totally made it up with absolutely no backing. You got me. :rolleyes:
Or not: https://www.anandtech.com/show/14603/the-msi-meg-x570-ace-motherboard-review/6
excellentswordfight
23rd July 2019, 13:54
Yup, I just totally made it up with absolutely no backing. You got me. :rolleyes:
Or not: https://www.anandtech.com/show/14603/the-msi-meg-x570-ace-motherboard-review/6
Well... From your article:
"For our motherboard reviews, we use our short form testing method. These tests usually focus on if a motherboard is using MultiCore Turbo (the feature used to have maximum turbo on at all times, giving a frequency advantage), or if there are slight gains to be had from tweaking the firmware. We put the memory settings at the CPU manufacturers suggested frequency, making it very easy to see which motherboards have MCT enabled by default."
I dont see how that test confirms higher performance in general for the x570 platform, just that the bios settings on different boards are tweaked differently. So the performance difference will vary depending on what boards you test, and might not matter at all if you are willing to tweak some settings for yourself.
https://www.techpowerup.com/review/amd-ryzen-3900x-3700x-tested-on-x470/6.html
"The question on your mind will now be whether you lose anything by the way of performance or overclocking headroom. We pulled out an MSI X470 Gaming M7 to find out just that. We are happy to report that you don't lose any performance."
Stereodude
23rd July 2019, 15:17
Well... From your article:
"For our motherboard reviews, we use our short form testing method. These tests usually focus on if a motherboard is using MultiCore Turbo (the feature used to have maximum turbo on at all times, giving a frequency advantage), or if there are slight gains to be had from tweaking the firmware. We put the memory settings at the CPU manufacturers suggested frequency, making it very easy to see which motherboards have MCT enabled by default."
I dont see how that test confirms higher performance in general for the x570 platform, just that the bios settings on different boards are tweaked differently. So the performance difference will vary depending on what boards you test, and might not matter at all if you are willing to tweak some settings for yourself.
https://www.techpowerup.com/review/amd-ryzen-3900x-3700x-tested-on-x470/6.html
"The question on your mind will now be whether you lose anything by the way of performance or overclocking headroom. We pulled out an MSI X470 Gaming M7 to find out just that. We are happy to report that you don't lose any performance."
Try to spin it all you want, but the X470 and the B450 delivered 9-12 percent less performance in Anandtech's Handbrake x264 and x265 tests. Plop in a Ryzen 3xxx CPU in them and that's what you get. Maybe you can narrow the gap if you have a lot of free time to tweak every BIOS setting.
The TechPowerUp test isn't clear what their H.264 and H.265 benchmark is, so its hard to give it any consideration.
excellentswordfight
23rd July 2019, 16:24
Try to spin it all you want, but the X470 and the B450 delivered 9-12 percent less performance in Anandtech's Handbrake x264 and x265 tests. Plop in a Ryzen 3xxx CPU in them and that's what you get. Maybe you can narrow the gap if you have a lot of free time to tweak every BIOS setting.
The TechPowerUp test isn't clear what their H.264 and H.265 benchmark is, so its hard to give it any consideration.
No, thats not what the test says, the test says that the MSI MEG X570 is faster then the ITX ASRock B450 and the Gigabyte x470 one. But a x470 board can still be faster then an x570 one cause it just a matter of how the bios is set up, just look at the powerdraw on that board, its just an factory OC. And yes, by default most b450 boards will be a bit slower, but your statment is way to generell cause it will differ from board to board, we dont even know if PBO is working correctly on that b450 board! And eitherway, I would say that the performance gap would need to be even bigger then that to compensate for the price and the fact that most board has an fan to be able to cool the chipset
https://img.purch.com/r/711x457/aHR0cDovL21lZGlhLmJlc3RvZm1pY3JvLmNvbS8yL04vODQ1MDg3L29yaWdpbmFsL2ltYWdlMDExLnBuZw==
Nintendo Maniac 64
23rd July 2019, 19:55
Perhaps I was too subtle with my post, but there are even X570 motherboards for $200 or less (and no I don't just mean "$199" either) so you can avoid the whole 400-series vs 500-series motherboard debate altogether, and these are being recommended by a renowned extreme overclocker (Buildzoid (https://www.youtube.com/channel/UCrwObTfqv8u1KO7Fgk-FXHQ/videos)) to boot so it's not like he doesn't know what he talking about:
https://www.youtube.com/watch?v=zuyuS04lD4o#t=21m01s
Of the $200-and-less X570 boards Buildzoid recommends, I'd go for the Gigabyte models due to their chipset fan being positioned lower and therefore won't be suffocated by a discrete GPU like has been reported on some Asus and ASRock motherboards (MSI's boards also have a lower-positioned chipset fan which would work just as well, but Buildzoid didn't recommend any specific MSI X570 boards, so...).
tl;dw: (these are PCPartPicker links for your convenience):
full ATX @ $200: Gigabyte X570 Aorus Elite (https://pcpartpicker.com/product/nHxbt6/gigabyte-x570-aorus-elite-atx-am4-motherboard-x570-aorus-elite) (best all-arounder at this price-point)
full ATX @ $170: Gigabyte X570 Gaming X (https://pcpartpicker.com/product/LJxbt6/gigabyte-x570-gaming-x-atx-am4-motherboard-x570-gaming-x) (cheaper variant with features stripped out)
mini ITX @ $220: Gigabyte X570 i Aorus Pro WiFi (https://pcpartpicker.com/product/NQ7p99/gigabyte-x570-i-aorus-pro-wifi-mini-itx-am4-motherboard-x570-i-aorus-pro-wifi) (...it's mini ITX, 'nuff said; also has wifi)
For reference, this video is the same one I included in my previous post but I had put it as a hyperlink on the text "source". As I mentioned, perhaps that was a bit too subtle on my part and it flew under the radar...
Alternatively maybe it's because the cheaper X570 boards are only covered near the end of the video, and it being a 30 minute video may mean that people didn't realize you don't have to watch the entire video (there are even timestamps provided right in the video itself in the bottom-right corner). Therefore I've modified the link so that it'll start directly at the point where Buildzoid begins talking about $200 X570 motherboards and then the segment right after that is sub-$200 X570 motherboards.
DJATOM
23rd July 2019, 20:49
You shouldn't need a $300+ motherboard unless you're going to be overclocking that 3950X (source (https://www.youtube.com/watch?v=zuyuS04lD4o)).
Yeah, I've bought ROG Crosshair VII Hero for that purpose (~ $330 in my country). Still waiting for 3900X to arrive, no plans for 3950X this year. Probably will pick it with next gen CPUs out, hope there will be reasonable price drop.
Stereodude
24th July 2019, 00:31
No, thats not what the test says, the test says that the MSI MEG X570 is faster then the ITX ASRock B450 and the Gigabyte x470 one. But a x470 board can still be faster then an x570 one cause it just a matter of how the bios is set up, just look at the powerdraw on that board, its just an factory OC. And yes, by default most b450 boards will be a bit slower, but your statment is way to generell cause it will differ from board to board, we dont even know if PBO is working correctly on that b450 board! And eitherway, I would say that the performance gap would need to be even bigger then that to compensate for the price and the fact that most board has an fan to be able to cool the chipset
Still, you're using a motherboard designed to handle processors with up to 8 cores for processors that are now up to 16 cores. I think I'd tend to favor a low end X570 over a B450 or X470 even with the fan on the X570.
Regardless, I'm in no hurry so I can wait and see how things settle out. I'm waiting at least for the 3950X, maybe even for a TR2.
DJATOM
26th July 2019, 14:48
So I've got 3900X onto my hands and we made an encoding speed comparison between it and 2x set of Xeon E5-2687W 0 (Jensen's setup)
SAO3 NCOP1 (full filtering script (AA, denoise, deband) + encode)
3900X: encoded 2208 frames in 0:18:21.28 (2.00 fps), 25531.44 kb/s, Avg QP:15.20
2687W x2: encoded 2208 frames in 0:32:32.68 (1.13 fps), 25506.87 kb/s, Avg QP:15.21
SAO3 NCOP1 (filtering script w/o AA part + encode)
3900X: encoded 2208 frames in 0:07:30.00 (4.91 fps), 26553.97 kb/s, Avg QP:14.98
2687W x2: encoded 2208 frames in 0:15:14.86 (2.41 fps), 26540.64 kb/s, Avg QP:14.99
I'm pretty satisfied with results. Now imagine that there are 16 core chips will come this autumn, it will be even faster.
Stereodude
26th July 2019, 15:37
So I've got 3900X onto my hands and we made an encoding speed comparison between it and 2x set of Xeon E5-2687W 0 (Jensen's setup)
SAO3 NCOP1 (full filtering script (AA, denoise, deband) + encode)
3900X: encoded 2208 frames in 0:18:21.28 (2.00 fps), 25531.44 kb/s, Avg QP:15.20
2687W x2: encoded 2208 frames in 0:32:32.68 (1.13 fps), 25506.87 kb/s, Avg QP:15.21
SAO3 NCOP1 (filtering script w/o AA part + encode)
3900X: encoded 2208 frames in 0:07:30.00 (4.91 fps), 26553.97 kb/s, Avg QP:14.98
2687W x2: encoded 2208 frames in 0:15:14.86 (2.41 fps), 26540.64 kb/s, Avg QP:14.99
I'm pretty satisfied with results. Now imagine that there are 16 core chips will come this autumn, it will be even faster.
Is that a first gen E5-2687W (not E5-2687Wv2)? What resolution and x265 preset is the data for?
So it's almost twice as fast as a dual processor E5-2687W system?
DJATOM
26th July 2019, 18:06
Is that a first gen E5-2687W (not E5-2687Wv2)? What resolution and x265 preset is the data for?
So it's almost twice as fast as a dual processor E5-2687W system?
Yeah, it's not v2, just E5-2687W 0 (that's how it's identified in apps).
Script - https://pastebin.com/j9GQ7Ykq
Resolution: 1080p Blu-ray (anime).
cmd: .\vaporPortable-r44\vspipe.exe -y NCOP1.vpy - | x265-x64-v3.1+8-aMod-gcc911-opt-zen2 -F 12 --pools "" --hme --hme-search umh,star,star --limit-modes --open-gop --cbqpoffs -2 --crqpoffs -2 --no-rskip --no-tskip --keyint 240 --no-cutree --ref 4 --bframes 9 --bframe-bias 0 --b-pyramid --b-adapt 2 --no-sao --no-sao-non-deblock --aq-mode 4 --aq-strength 0.86 --deblock 1:-1 --tu-intra-depth 2 --tu-inter-depth 2 --me 2 --wpp --subme 5 --crf 15 --qcomp 0.72 --b-pyramid --merange 48 --weightp --weightb --rd 4 --psy-rd 2 --rdoq-level 2 --psy-rdoq 4 --sar 1:1 --info --colorprim bt709 --transfer bt709 --colormatrix bt709 --output "17.hevc" --csv-log-level 2 --csv "17.txt" --y4m -
I also made x265-x64-v3.1+8-aMod-gcc911-opt-sandybridge for encoding on Jensen's server.
RanmaCanada
26th July 2019, 18:53
Still, you're using a motherboard designed to handle processors with up to 8 cores for processors that are now up to 16 cores. I think I'd tend to favor a low end X570 over a B450 or X470 even with the fan on the X570.
Regardless, I'm in no hurry so I can wait and see how things settle out. I'm waiting at least for the 3950X, maybe even for a TR2.
Then by your logic anyone using an x99 retail motherboard with Xeons are losing performance. Which they are NOT. I mean even my old x79 based Sabertooth was "only designed for 4 cores", yet it worked perfectly fine with up to 10 cores (mine has an 8 core e5-2670).
The motherboard mfg's knew that with AM4 they would have to be prepared for multiple cores, and thus they designed them with the VRMs to handle the power draw.
Stereodude
27th July 2019, 00:20
Then by your logic anyone using an x99 retail motherboard with Xeons are losing performance. Which they are NOT. I mean even my old x79 based Sabertooth was "only designed for 4 cores", yet it worked perfectly fine with up to 10 cores (mine has an 8 core e5-2670).
The motherboard mfg's knew that with AM4 they would have to be prepared for multiple cores, and thus they designed them with the VRMs to handle the power draw.
Or not...
Note - The ASRock B450 Gaming ITX-ac model crashed instantly every time the small FFT torture test within Prime95 was initiated. At anything on the CPU VCore above 1.35 V would result in instant instability. The Ryzen Master auto-overclocking function failed every time it tried to dial in settings, but it does however operate absolutely fine at stock, and with Precision Boost Overdrive enabled. Either the firmware is the issue, or the board just isn't capable of overclocking the Ryzen 3700X with extreme workloads with what is considered a stable overclock on the X570 chipset. We will re-test this in the future.
RanmaCanada
27th July 2019, 05:49
Or not...
I'd say that is more a problem of the mITX platform just not having any where near enough air moving to keep things cool or the BIOS being horrible. It's funny you'd cherry pick the absolutely worse case scenario to try to prove your point. If the board does work fine with a 2700x, then at that point it would be more a problem with the BIOS lacking, which we know is a serious issue on x4xx platforms as the mfg don't want to "waste" time fixing hardware and software issues when they have fancy new stuff for people to buy. MSI is already telling people to just buy the 570 platform because they don't want to bother fixing the issues that they created in the x4xx series boards.
ASROCK is also the only mfg with 3 phase VRM's on their x4xx series boards. Everyone else has a minimum of 4.
mparade
2nd August 2019, 23:20
2990WX has very low fps because of broken windows scheduler.
Does anyone know if there is any solution for this yet? The same is happening to my 2990WX at the moment. Maybe if using Coreprio?
NikosD
3rd August 2019, 04:23
Does anyone know if there is any solution for this yet? The same is happening to my 2990WX at the moment. Maybe if using Coreprio? Did you try latest Windows 10 official version 1903 ?
It has an updated scheduler, but I think not so enhanced for Threadripper to wait for miracles.
If I had a 2990WX I would definitely try Coreprio.
RanmaCanada
3rd August 2019, 04:30
Does anyone know if there is any solution for this yet? The same is happening to my 2990WX at the moment. Maybe if using Coreprio?
Switch to linux? The scheduler problem doesn't exist there. You can always run windows in an emulator, or use UNRAID as your base and have both Linux and Windows available on your system in native mode, with hardware passthrough.
Atak_Snajpera
3rd August 2019, 10:03
Does anyone know if there is any solution for this yet? The same is happening to my 2990WX at the moment. Maybe if using Coreprio?
According to guy from level1techs ( https://youtu.be/cTmnyOZ0pE4?t=34m ) you should
1) install Linux
2) setup virtual machine on linux
3) install windows 10 on virtual machine
Done.
TEB
8th August 2019, 11:05
https://www.phoronix.com/scan.php?page=article&item=amd-epyc-7502-7742&num=4
Looks like PhoronixŽs benchmark isnt able to saturate all these cores... but nevertheless.. impressive numbers for the new EPYC 7xx2 series ;)
jfcarbel
14th September 2019, 19:10
Again
https://p.xfastest.com/~sinchen/AMD-Ryzen-7-3700X-9-3900X/AMD-Ryzen-7-3700X-9-3900X-34.jpg
Was this done on a Windows 7 or Windows 10 machine?
I think I read that Windows 10 performs much faster for the Ryzen 3000
Atak_Snajpera
14th September 2019, 21:09
Was this done on a Windows 7 or Windows 10 machine?
I think I read that Windows 10 performs much faster for the Ryzen 3000
Within margin of error. Do not expect miracles.
jfcarbel
14th September 2019, 21:41
Within margin of error. Do not expect miracles.
Thanks, so not significant.
So is your windows 7 updater able to get Windows 7 running well with the Ryzen 3000 CPUs?
-QfG-
14th September 2019, 21:52
Example for a bad Encoding CPU (i7700k) ~70hours for an 2h video
https://s17.directupload.net/images/190914/v78ur2h9.png
I would say, minimum 8 cores (I9900k or Ryzen 7 3xxx for example)
Atak_Snajpera
14th September 2019, 22:40
Thanks, so not significant.
So is your windows 7 updater able to get Windows 7 running well with the Ryzen 3000 CPUs?
Yep. However If you have x570 then you will have to add this step to get USB 3 ports to work
https://www.sevenforums.com/drivers/419688-will-new-amd-x570-chipset-mobos-have-win-7-chipset-drivers-post3441921.html?#post3441921
NikosD
16th September 2019, 18:57
For those who are still wondering which is the fastest platform for HEVC encoding:
8K real-time HEVC encoding at 79fps with just one CPU (!)
https://www.tomshardware.com/news/amd-epyc-rome-8k-real-time-encoding,40400.html
archer75
16th October 2019, 16:55
I have a question that hopefully someone can help me with.
I have two computers, a PC I built and a 2018 mac mini. Using handbrake on both, using the same settings with the same file and encoding with HEVC 10bit I get very different results.
The PC uses an Intel 5820 and the Mac uses an Intel 8700.
The encode that comes out of the macs 8700 is quite a bit smaller file size than the one produced by the 5820 and it also has noticeably better image quality. Why is that? I'm considering an upgrade and knowing why the file is both smaller with better quality would help me in choosing which processor for a new build.
Nico8583
16th October 2019, 16:59
Do you use the same handbrake version with the same third-party tools version ?
The OS are different so executable must be different (I don't know MAC OS very well).
Nico8583
16th October 2019, 18:29
I think the OS but I'm not sure.
RanmaCanada
17th October 2019, 04:52
Is one using quicksync and the other using pure software? Realistically there should be no difference between the 2 if you are using the exact same settings, profiles, etc.
benwaggoner
21st October 2019, 00:02
How big is the file size and quality gap? Does one of the systems have way more cores than the other? There used to be quality regressions using frame threading, although this has gotten quite a big better in recent years.
DotJun
1st November 2019, 11:47
Does anyone know why avx2 is being used when set to auto detect on my intel 7820x? I typically set it to avx512, but forgot to do so on my last batch.
Boulder
1st November 2019, 11:56
Because AVX512 is not generally recommended, so you need to enable it manually if you want to use it. Basically you should test to see if you benefit from it or not, seems to depend on the case quite a lot.
DotJun
1st November 2019, 12:13
Because AVX512 is not generally recommended, so you need to enable it manually if you want to use it. Basically you should test to see if you benefit from it or not, seems to depend on the case quite a lot.
Ok, I didnt know it was excluded from auto detection, thanks for the info and I tested it a while back. My results showed a slight improvement in FPS, even with a negative offset, so I just stuck with it.
hajj_3
1st November 2019, 13:24
Ok, I didnt know it was excluded from auto detection, thanks for the info and I tested it a while back. My results showed a slight improvement in FPS, even with a negative offset, so I just stuck with it.
AVX512 makes the processor too hot so the clock speed of the AVX portion is reduced and as a result it isn't very fast. I think Intel might have fixed that in their latest processors.
DotJun
1st November 2019, 13:45
AVX512 makes the processor too hot so the clock speed of the AVX portion is reduced and as a result it isn't very fast. I think Intel might have fixed that in their latest processors.
Yep, thats the reason for the negative offset. From the small amount of testing I did, avx512 still came out with slightly higher FPS even with a couple hundred mghz lower clock due to offset.
Though I will admit that I wouldnt be able to run 512 if my computer wasnt located in a 65-70f temp room.
RanmaCanada
2nd November 2019, 00:43
Yep, thats the reason for the negative offset. From the small amount of testing I did, avx512 still came out with slightly higher FPS even with a couple hundred mghz lower clock due to offset.
Though I will admit that I wouldnt be able to run 512 if my computer wasnt located in a 65-70f temp room.
Ha with winter coming, you can just place it outside and run without an offset :P haha.
DotJun
9th November 2019, 08:16
I had a chance to do a small sample test and here are my results:
system: 7820x oc'd to 4.4ghz with 32g of memory. A -3 offset was used for both avx2 and avx512 as explained below.
10 minute 1080p sample on auto detect: 5.38fps
10 minute 1080p sample on avx512: 5.72fps
5 minute 4k sample on auto detect: 0.84fps
5 minute 4k sample on avx512: 0.92fps
crf and slower settings were used on all samples.
The odd part is that both auto detect (avx2) and avx512 needed the same -3 offset or windows becomes unstable and crash. Also, avx2 ran 5-10c hotter than avx512, which is the opposite of what I thought would happen.
Results might be different on a longer sample clip, but I really didn't want to invest too much time into it since for whatever reason avx2 was pushing out more heat than what I am comfortable with.
aymanalz
9th November 2019, 11:05
The odd part is that both auto detect (avx2) and avx512 needed the same -3 offset or windows becomes unstable and crash. Also, avx2 ran 5-10c hotter than avx512, which is the opposite of what I thought would happen.
Results might be different on a longer sample clip, but I really didn't want to invest too much time into it since for whatever reason avx2 was pushing out more heat than what I am comfortable with.
This may have more to do with your processor being overclocked, than any quirks of AVX2.
Have you tried the same tests on stock clocks?
nevcairiel
9th November 2019, 11:39
Have you tried the same tests on stock clocks?
Noone runs a Skylake X series on stock clocks, that would be a waste of those CPUs, since they OC extremely well. As such any tests on stock really don't tell you anything useful.
For the question at hand, yes, AVX512 can run overall cooler then AVX2, because it does twice the work in the same time, or needs less time to do the same work, which means the AVX units are overall less busy, and don't heat up as much, since software like x265 isn't pushing pure AVX512 load, it still has loads of other things to compute.
DotJun
9th November 2019, 13:16
Noone runs a Skylake X series on stock clocks, that would be a waste of those CPUs, since they OC extremely well. As such any tests on stock really don't tell you anything useful.
For the question at hand, yes, AVX512 can run overall cooler then AVX2, because it does twice the work in the same time, or needs less time to do the same work, which means the AVX units are overall less busy, and don't heat up as much, since software like x265 isn't pushing pure AVX512 load, it still has loads of other things to compute.
Ok that makes sense. Is that also why I'm seeing increased performance with avx512, because it has time to do other things since it's overall less busy due to finishing avx stuff sooner?
nevcairiel
9th November 2019, 14:24
Ok that makes sense. Is that also why I'm seeing increased performance with avx512, because it has time to do other things since it's overall less busy due to finishing avx stuff sooner?
Thats generally how all SIMD speedup works, you make it spend less time in typical DSP functions, which are easy to optimize, so the overall process runs faster.
aymanalz
10th November 2019, 05:39
Noone runs a Skylake X series on stock clocks, that would be a waste of those CPUs, since they OC extremely well. As such any tests on stock really don't tell you anything useful.
Yes, running an X series on stock clocks would be a waste of money. But I was suggesting that as a test to eliminate overclocking issues as a likely culprit in the unstability and crashing that he mentioned. I don't see how AVX-2 can cause crashes at this point, as x265 has been pretty well optimized for AVX2 by now. A bad overclock on the other hand...
If those crashes and unstability occur at stock settings, then he can be sure that it's not related to hardware issues.
DotJun
10th November 2019, 13:19
Yes, running an X series on stock clocks would be a waste of money. But I was suggesting that as a test to eliminate overclocking issues as a likely culprit in the unstability and crashing that he mentioned. I don't see how AVX-2 can cause crashes at this point, as x265 has been pretty well optimized for AVX2 by now. A bad overclock on the other hand...
If those crashes and unstability occur at stock settings, then he can be sure that it's not related to hardware issues.
It wouldn't crash at stock speeds if it isn't crashing at 4.4 with a -3 offset.
NikosD
15th November 2019, 12:58
The Ryzen 9 3900X 12C/24T was already too fast, faster than any Intel desktop processor.
But the Ryzen 9 3950X 16C/32T is even faster, a lot faster.
But you can always wait for the Threadrippers that will be released this month. (if you can afford the price premium)
https://i.postimg.cc/63T8y7Gq/handbrake-3950x.png
Stereodude
15th November 2019, 13:26
It's looking good. Can you give us more details on that graph? Like what resolution / settings were used in handbrake?
Personally, I want to see how the 3960X and 3970X compare before buying something. I presume reviews will drop by the 25th when they all go on sale though I don't plan to be a day 1 buyer anyhow.
NikosD
15th November 2019, 16:07
It's looking good.
Can you give us more details on that graph?
Like what resolution / settings were used in handbrake? It's a graph from legitreviews.com review.
We used Big Buck Bunny as our input file, which has become one of the world standards for video benchmarks.
For our benchmark scenario we used a standard 2D 4K (3840Ś2160) 60 FPS clip in the MP4 format and used Handbrake version 1.2.2
Stereodude
15th November 2019, 17:17
It's a graph from legitreviews.com review.
So based on their picture in the review they're taking a 2160p60 input clip and converting it to 1080p30 (x264) using the "Fast 1080p" preset in Handbrake 1.2.2. I'm guessing the x265 test is also a 1080p encode.
Edit: This is probably not correct.
NikosD
15th November 2019, 17:38
By reading almost 10 reviews of Ryzen 9 3950X today, I can clearly say that this particular mainstream desktop CPU is the best processor of AMD and generally of x86/x64 platform of all time.
It's the fastest single-thread CPU of all Intel and AMD processors and the fastest gaming CPU of AMD ever (still Intel has a slight advantage for games at 1080p)
The multi-thread performance is most of the times better than the 18 core Intel HEDT and a lot faster than second generation 16 core Threadripper.
All of that using the same power consumption of 9900K (~140W in all core turbo real performance) which has half cores (!) - only 8 - making 3950X an extremely efficient processor.
You can even drop the TDP (base frequency) to 65W instead of 105W with a performance loss of 10% - 15%
All of those with 750$.
AMD at its best after many, many years.
Blue_MiSfit
15th November 2019, 19:23
^^ I'm SUPER impressed with it.
I have a 9900k at work and the improvement relative to this may get me to upgrade from my 7700k at home, especially with me getting more interested in SVT AV1 testing :)
16 cores / 32 threads is perfect for x265 as well, especially all in 1 NUMA node.
Stereodude
15th November 2019, 20:04
Do we know yet if the new Zen 2 based Threadrippers (3960X / 3970X) will have more than one NUMA node splitting the cores up?
Atak_Snajpera
15th November 2019, 20:44
Do we know yet if the new Zen 2 based Threadrippers (3960X / 3970X) will have more than one NUMA node splitting the cores up?
3980x (48C/96T) and 3990x (64C/128T) will be seen by windows as 2 numa nodes unless you disable SMT.
Stereodude
15th November 2019, 20:58
3980x (48C/96T) and 3990x (64C/128T) will be seen by windows as 2 numa nodes unless you disable SMT.
Are you sure? The 64 core Epyc Rome is only 2 NUMA nodes with 8 chiplets so it seems like 4 chiplets in single NUMA node should be possible.
Atak_Snajpera
15th November 2019, 21:31
Are you sure? The 64 core Epyc Rome is only 2 NUMA nodes with 8 chiplets so it seems like 4 chiplets in single NUMA node should be possible.
Yes 48T and 64T will be seen as single NUMA by windows. Higher number of threads will give you extra node in task manager.
RanmaCanada
16th November 2019, 02:52
Well if anyone here has a Zen 2, aka Ryzen 3000 series chip, they should run the benchmark Sagitare created and post their results there. That way we can get some real world numbers, and not the canned crap that benchmarking sites use. Though the exe files need to be updated as they're over a year old now. I honestly wish all sites would do crf 20 slow, then slower, slowest and placebo. No one uses the settings they benchmark with so the results they produce are useless.
https://forum.doom9.org/showthread.php?t=174393
Stereodude
16th November 2019, 17:53
So based on their picture in the review they're taking a 2160p60 input clip and converting it to 1080p30 (x264) using the "Fast 1080p" preset in Handbrake 1.2.2. I'm guessing the x265 test is also a 1080p encode.
So I decided to download HandBrake 1.2.2 and the 2160p60 version of Big Buck Bunny and try this test myself to see how my E5-2687Wv2 computer compares to the systems Legit Reviews used to see what sort of improvement I might get. I got an average FPS of 47.3 using x264 and the settings shown in their screenshot. They're clearly not using the HandBrake settings shown in the screenshot. There's no way a 8C/16T Xeon without AVX2 is going to best a 8C/16T 9th Gen i9 by over 50% when the i9 has the same number of cores, a sizeable clock speed advantage, and ~6 generations of IPC improvements and support for additional instructions.
So, no idea what settings they're using, but it's definitely not what's shown in the screenshot.
RanmaCanada
16th November 2019, 18:25
So I decided to download HandBrake 1.2.2 and the 2160p60 version of Big Buck Bunny and try this test myself to see how my E5-2687Wv2 computer compares to the systems Legit Reviews used to see what sort of improvement I might get. I got an average FPS of 47.3 using x264 and the settings shown in their screenshot. They're clearly not using the HandBrake settings shown in the screenshot. There's no way a 8C/16T Xeon without AVX2 is going to best a 8C/16T 9th Gen i9 by over 50% when the i9 has the same number of cores, a sizeable clock speed advantage, and ~6 generations of IPC improvements and support for additional instructions.
So, no idea what settings they're using, but it's definitely not what's shown in the screenshot.
And that is why we can't trust reviews for encoding benchmarks if they don't give us the command lines used, and I suggested that any one of the numerous people here who have a zen 2, should run Sagitare's benchmark so we can at least get an apples to apples comparison. Though as the exe's are actually 2 years old, Sagitare needs to update it (I tried and failed miserably haha).
Warrex
25th November 2019, 17:52
Threadripper 39xxX:
https://p.xfastest.com/~sinchen/AMD-Ryzen-Threadripper-3960X-3970X/AMD-Ryzen-Threadripper-3960X-3970X-33.jpg
For comparison:
https://p.xfastest.com/~sinchen/AMD-Ryzen-7-3700X-9-3900X/AMD-Ryzen-7-3700X-9-3900X-34.jpg
Intel SVT (AV1, etc.) performance:
https://techgage.com/article/amd-threadripper-3960x-3970x-intel-i9-10980xe-linux/2/
Stereodude
25th November 2019, 21:45
And what does the XFASTEST x264/x265 test consist of?
Their data for x265 looks promising, but shows it's not faster in x264.
Others show improvements in both (from Legit Reviews):
https://i.imgur.com/oDRhNTr.png
Warrex
25th November 2019, 22:54
And what does the XFASTEST x264/x265 test consist of?
As already discussed in this thread XFASTEST use Atak's benchmarks:
https://forum.pclab.pl/topic/1184884-x265-FHD-Benchmark/
https://www.guru3d.com/files-details/x264-fhd-benchmark-v1-1-64bit.html
RanmaCanada
26th November 2019, 05:35
I personally feel no benchmark is valid unless they are using at least the slow preset. I will wait until I see someone running with those, like techpowerup did (https://www.techpowerup.com/review/amd-ryzen-7-3700x/13.html) for the Ryzen 3700x and 3900x. Though they used time to encode, and did not give us the fps, nor the time of the actual clip they used. I really hate how no one properly benches encoding. Tell us how long the clip is, how many frames it has, and what type of media it is. I would honestly say using one of the test pattern clips, or even any of the open source/Creative Commons movies, like Tears of Steel. I mean how else are we the public supposed to replicate the results if we don't know what reviewers are using to test with.
nevcairiel
26th November 2019, 08:50
I mean how else are we the public supposed to replicate the results if we don't know what reviewers are using to test with.
You are not, because otherwise there would be fanbois from both sides showing up playing point and counter-point showing how the results the website posted is wrong in either direction.
Thats why you keep benchmarks at least slightly opaque. If you want to do your own benchmarks, you need to make your own references to benchmark against.
Atak_Snajpera
26th November 2019, 15:54
Even 3950X is extremely good in x265
https://i.postimg.cc/KYgqT9Cf/Untitled-1.png
NikosD
27th November 2019, 02:23
And SVT-AV1 too
https://i.postimg.cc/Bn7dyfhg/y-V94uwn-ZCpkfk4-B3z-UR4-Td-650-80.png
Nintendo Maniac 64
27th November 2019, 19:50
And SVT-AV1 too
I know you can't compare setups, but how the heck did Tom's get such ridiculously faster numbers than Phoronix? Windows vs Linux shouldn't cause that much of a difference!
https://www.phoronix.com/scan.php?page=article&item=amd-linux-3960x-3970x&num=8
Atak_Snajpera
27th November 2019, 20:17
I know you can't compare setups, but how the heck did Tom's get such ridiculously faster numbers than Phoronix? Windows vs Linux shouldn't cause that much of a difference!
https://www.phoronix.com/scan.php?page=article&item=amd-linux-3960x-3970x&num=8
Easier to compress sample? Different ENC mode? Hard to tell...
NikosD
27th November 2019, 20:35
I know you can't compare setups, but how the heck did Tom's get such ridiculously faster numbers than Phoronix? I know nothing of SVT-AV1 but I must say that according to Phoronix Enc4 vs Enc8 has one order of magnitude (x10) difference regarding absolute numbers (!)
I think we must focus on relative performance for each test of Tom's and Phoronix regarding to Enc4.
For example, in both tests of Enc4 the new Threadripper 32C is about two times faster than Core i9 18C and 2.5 to 3 times faster than previous gen Threadripper 32C.
The absolute performance could probably be more important for those who actually use the specific AV1 encoder and not using it just for benchmarks.
TEB
12th December 2019, 14:51
So... i got access to a Lenovo SR630 with Epyc 7742 on at work now running RHEL 8.1. Any benchmarks u guys want me to run on it?
processor : 127
vendor_id : AuthenticAMD
cpu family : 23
model : 49
model name : AMD EPYC 7742 64-Core Processor
stepping : 0
microcode : 0x830101c
cpu MHz : 2332.947
cache size : 512 KB
physical id : 0
siblings : 128
RanmaCanada
12th December 2019, 17:49
Try running Sagitare's benchmark. https://forum.doom9.org/showthread.php?t=174393 It's a little old.
You could also run a 1080p or 4k benchmark of say Tears of Steel at CRF 18-20 and go from Placebo all the way to ultra fast (if you have time) Because this would be a real world workload and let people see the difference in each preset.
Atak_Snajpera
18th December 2019, 23:14
This https://forum.pclab.pl/topic/1184884-x265-FHD-Benchmark/
TEB
18th December 2019, 23:30
This https://forum.pclab.pl/topic/1184884-x265-FHD-Benchmark/
Running RHEL 8.x
Atak_Snajpera
18th December 2019, 23:42
Running RHEL 8.x
Windows in VM maybe?
blublub
20th January 2020, 20:43
If anyone wants me to run a test on a Threadripper 3960X let me know.
Conditions:
Windows based
easy to setup - best specify the source with link, programm to encode with wanted preset
RanmaCanada
21st January 2020, 01:33
If anyone wants me to run a test on a Threadripper 3960X let me know.
Conditions:
Windows based
easy to setup - best specify the source with link, programm to encode with wanted preset
For simplicity I would say to do Tears of Steel with each preset at CRF 16-18 (what most people here use). If you have time do 1080p to 1080p and 4k to 4k. (https://mango.blender.org/download/)
If that's too long, do the good old Park Joy (https://media.xiph.org/video/derf/).
You could use handbrake, or staxrip, your choice. No filtering, no extras, just let the presets do what they are supposed to.
The reason I ask this is because I feel Atak's benchmark uses a preset that no one will use in real life, and every site that does benchmark it does not give us the length, frames, etc of the source they are using, which makes their benchmarks useless.
By using these well known, open source/royalty free videos, the results can be properly compared with other processors, and results can be replicated.
Atak_Snajpera
21st January 2020, 11:46
For simplicity I would say to do Tears of Steel with each preset at CRF 16-18 (what most people here use). If you have time do 1080p to 1080p and 4k to 4k.
You also make blind assumptions that most people use your settings. You have zero data to prove your claim so stop saying that my benchmark using default settings is unrealistic. Furthermore yours CRF16-18 is total overkill for most users (bitrate will go through the roof). Another issue. handbreak won't saturate all 48 threads on 1080p and probably also in 4k. Chunked encoding is the only effecting way of achieving constant 100% CPU usage. With 3990x chunked encoding is basicaly required.
RanmaCanada
21st January 2020, 17:16
You also make blind assumptions that most people use your settings. You have zero data to prove your claim so stop saying that my benchmark using default settings is unrealistic. Furthermore yours CRF16-18 is total overkill for most users (bitrate will go through the roof). Another issue. handbreak won't saturate all 48 threads on 1080p and probably also in 4k. Chunked encoding is the only effecting way of achieving constant 100% CPU usage. With 3990x chunked encoding is basicaly required.
Just go through all the treads in the various subs here and almost EVERYONE uses slow or slower for their encodes, with an average CRF of 18. You will have the odd person who is "blind" and thinks 22 or higher are fine, but the majority of people use CRF's in the teens.
Remember, this is a forum for the 1% of encoders and enthusiasts. Most people want archival quality of their encodes.
I am sorry you feel offended by what I said about your benchmark, but in real world scenarios your benchmark doesn't hold up. For example on your benchmark it states my 2700 gets 21+ fps, which I've never EVER seen in encoding at slow or slower. Not even at medium have I seen it at that speed, and I run with a pretty bare command line. Benchmarks are supposed to be about real world performance, and no one who cares about quality would be using fast and higher presets.
microchip8
21st January 2020, 17:25
Just go through all the treads in the various subs here and almost EVERYONE uses slow or slower for their encodes, with an average CRF of 18. You will have the odd person who is "blind" and thinks 22 or higher are fine, but the majority of people use CRF's in the teens.
Remember, this is a forum for the 1% of encoders and enthusiasts. Most people want archival quality of their encodes.
I am sorry you feel offended by what I said about your benchmark, but in real world scenarios your benchmark doesn't hold up. For example on your benchmark it states my 2700 gets 21+ fps, which I've never EVER seen in encoding at slow or slower. Not even at medium have I seen it at that speed, and I run with a pretty bare command line. Benchmarks are supposed to be about real world performance, and no one who cares about quality would be using fast and higher presets.
I'm not "blind" and use CRF 21 in 10-bits and it looks totally fine. In fact, I can't see a (major) difference between 21 and 18 or 19, except for the bitrate. I must be blind, then?
Atak_Snajpera
21st January 2020, 17:59
Just go through all the treads in the various subs here and almost EVERYONE uses slow or slower for their encodes, with an average CRF of 18. You will have the odd person who is "blind" and thinks 22 or higher are fine, but the majority of people use CRF's in the teens.
Remember, this is a forum for the 1% of encoders and enthusiasts. Most people want archival quality of their encodes.
I am sorry you feel offended by what I said about your benchmark, but in real world scenarios your benchmark doesn't hold up. For example on your benchmark it states my 2700 gets 21+ fps, which I've never EVER seen in encoding at slow or slower. Not even at medium have I seen it at that speed, and I run with a pretty bare command line. Benchmarks are supposed to be about real world performance, and no one who cares about quality would be using fast and higher presets.
Because this benchmark uses samples containing a lot of motion and details!
crowd_run_1080p50.yuv
ducks_take_off_1080p50.yuv
in_to_tree_1080p50.yuv
old_town_cross_1080p50.yuv
park_joy_1080p50.yuv
Looks at this as worse case scenario.
Besides, You are probably encoding at cropped 1920x800 resolution instead of full 1920x1080.
This benchmark shows you what you can expect from CPU A vs CPU B. For example. Should I buy Ryzen 3950x or Intel core i9 10980xe. Do not look at raw numbers because it does not make sense.
PS. 2700x gets 28 fps in my benchmark so be more precise next time ,ok?
https://i.postimg.cc/G2K5grgN/Capture.png
Asmodian
21st January 2020, 18:02
Fast/slow/slower?
CRF 21 is not bad with x265. I use lower for x264 but with x265 below that usually results in bitrates near where I would also get transparent encodes using x264. However, I think a CPU benchmark should use slower not fast. Hardware encoding is good enough today that unless you are using slower+ a Turing GPU encode is probably a better option.
microchip8
21st January 2020, 18:06
Fast/slow/slower?
CRF 21 is not bad with x265. I use lower for x264 but with x265 below that usually results in bitrates near where I would also get transparent encodes using x264. However, I think a CPU benchmark should use slower not fast. Hardware encoding is good enough today that unless you are using slower+ a Turing GPU encode is probably a better option.
Custom settings that have features from both slow and slower preset
blublub
21st January 2020, 20:50
Fast/slow/slower?
CRF 21 is not bad with x265. I use lower for x264 but with x265 below that usually results in bitrates near where I would also get transparent encodes using x264. However, I think a CPU benchmark should use slower not fast. Hardware encoding is good enough today that unless you are using slower+ a Turing GPU encode is probably a better option.I haven't tried GPU encoding for ages. Is it really good looking now?
I will see what I can bench test over the weekend.
From my perspective it doesn't matter what presets are used for a benchmark.
Consistency is important to compare CPUs
RanmaCanada
21st January 2020, 20:54
Because this benchmark uses samples containing a lot of motion and details!
crowd_run_1080p50.yuv
ducks_take_off_1080p50.yuv
in_to_tree_1080p50.yuv
old_town_cross_1080p50.yuv
park_joy_1080p50.yuv
Looks at this as worse case scenario.
Besides, You are probably encoding at cropped 1920x800 resolution instead of full 1920x1080.
This benchmark shows you what you can expect from CPU A vs CPU B. For example. Should I buy Ryzen 3950x or Intel core i9 10980xe. Do not look at raw numbers because it does not make sense.
PS. 2700x gets 28 fps in my benchmark so be more precise next time ,ok?
https://i.postimg.cc/G2K5grgN/Capture.png
I said 2700, not 2700x
Atak_Snajpera
22nd January 2020, 00:02
I said 2700, not 2700x
According to this graph you should get 24+ FPS then
https://cdn.mos.cms.futurecdn.net/KfSHM4wcaYKM9VWgPk9FdS-650-80.png
Indeed huge error from my side ...
RanmaCanada
22nd January 2020, 00:29
Which I do see with an updated version of x265. Still no where near what I get in real world situations, hence the request for slow and lower presets as that is what the majority of us who encode would be using. Depending on material at slow I see 4-7 fps, and on slower I'm lucky if I get 2, but usually hang around 1.3-1.5 fps. This is for 1080p content. For 4k, on slow I'm lucky if I hit 1fps.
blublub
22nd January 2020, 06:05
Hi
Ok. Here is one result:
Tears of Steel 4K, x265 preset slow: 6,32fps
Software: RipBot with HEVC encoder version 3.2+34-8e6db24c1517
settings:
x265_x64.exe" --colorprim bt709 --transfer bt709 --colormatrix bt709 --crf 18 --fps 24 --min-keyint 24 --keyint 240 --frames 17616 --sar 1:1 --profile main10 --output-depth 10 --preset slow --ctu 64 --merange 57
@Atak_Snapjera:
I tried to download your benchmark but the download site kept generating one "key" after another but never presented the download link
RanmaCanada
22nd January 2020, 06:57
Thank you!
Atak_Snajpera
22nd January 2020, 12:57
Hi
Ok. Here is one result:
Tears of Steel 4K, x265 preset slow: 6,32fps
Software: RipBot with HEVC encoder version 3.2+34-8e6db24c1517
settings:
x265_x64.exe" --colorprim bt709 --transfer bt709 --colormatrix bt709 --crf 18 --fps 24 --min-keyint 24 --keyint 240 --frames 17616 --sar 1:1 --profile main10 --output-depth 10 --preset slow --ctu 64 --merange 57
@Atak_Snapjera:
I tried to download your benchmark but the download site kept generating one "key" after another but never presented the download link
Important question. Was CPU usage at 100%
blublub
22nd January 2020, 13:11
No, but close. Since I used the preset slower it was around 80 to 85%. On medium it is around 65
Atak_Snajpera
22nd January 2020, 15:49
No, but close. Since I used the preset slower it was around 80 to 85%. On medium it is around 65
Try again with active Distributed Encoding mode
Use these settings
https://i.postimg.cc/Gt0TtbFX/Capture.png
https://i.postimg.cc/FzjdJwGN/Capture2.png
Flat 100% CPU usage guaranteed!
blublub
22nd January 2020, 18:10
Hi, I have trouble getting it to encode. I disabled all NICs despite one, disabled FW and Uninstaller AV but the encode still won't start.
blublub
24th January 2020, 09:25
OK, I did get distributed encoding to start correctly.
CPU load was mostly between 80 and 90%. That is still an increase of 10 to 15% in cpu load but no matter what I can't get it constantly close to 100%.
I have no clue what the issue here is.
I run the complete encode on a NVME Samsung Pro SSD, so that can't be it.
From what I know from encoding it shouldn't be limited by memory speed and I have quad channel 2933 with more than 50 gb/s, so I guess that's not it either.
May be its a bottleneck within the CPU, i.e AVX capacity.
I don't know, but from what I am seeing here it kinda does not make sense to buy a CPU with more than 24, for encoding at the moment.
Atak_Snajpera
24th January 2020, 11:57
Have you tried with 3rd server?
blublub
24th January 2020, 12:56
Not yet.
1 encode with the settings has a 60 to 80% usage. Since a 2nd instances only marginally improves the usage I doubt a 3rd encode will, there seems to be some sort of bottleneck.
Atak_Snajpera
24th January 2020, 13:03
run my benchmark and see if you get at least similar fps like in my data base.
https://i.postimg.cc/KYgqT9Cf/Untitled-1.png
blublub
24th January 2020, 13:45
Can u post an alternate download link?
I already tried to download it 2 days ago but the website did not generate a link.
Atak_Snajpera
24th January 2020, 13:55
Maybe you should disable adblock? Change/update browser? Works fine here. Just checked.
blublub
24th January 2020, 13:59
K, I'll try tonight
blublub
24th January 2020, 14:30
run my benchmark and see if you get at least similar fps like in my data base.
https://i.postimg.cc/KYgqT9Cf/Untitled-1.png
Score is 108
Atak_Snajpera
24th January 2020, 14:36
So everything is ok then. Activate 3rd server and see if cpu usage goes up.
blublub
25th January 2020, 07:54
So everything is ok then. Activate 3rd server and see if cpu usage goes up.
Hi
I did have to fiddle aroud to get distributed encoding to work and while doing that I did not set the process priority correctly. After setting it to normal the usage is mostly 100%.
thx
Atak_Snajpera
8th February 2020, 14:33
ULTIMATE ENCODING CPU
https://i.postimg.cc/zGGxXvZk/Untitled-1.png
Source -> https://hk.xfastest.com/45759/amd-ryzen-threadripper-3990x-performance-test-kill-intel-i9-10980xe/
Emulgator
8th February 2020, 16:36
Muuuhahahaa !
Still so much improvement above the 3970, unbelievable.
excellentswordfight
8th February 2020, 18:23
Muuuhahahaa !
Still so much improvement above the 3970, unbelievable.
What? Twice the price, twice the number of cores, 18% perf increase.
If its not a saturation/thread utilization issue, i wonder were the bottlneck is? Cache/ram bandwith?
Atak_Snajpera
8th February 2020, 18:37
What? Twice the price, twice the number of cores, 18% perf increase.
I suspect that 5 instances of x265 encoding 1080p at the same time might not be enough for constant 100% cpu usage. If you overclock 3990X to 3.7 GHz then you should get extra ~50% vs 3970x.
Stereodude
8th February 2020, 19:54
What? Twice the price, twice the number of cores, 18% perf increase.
If its not a saturation/thread utilization issue, i wonder were the bottlneck is? Cache/ram bandwith?
Could be the OS. Most versions of Windows 10 only allow for 64 threads in a single processor group.
https://www.anandtech.com/show/15483/amd-threadripper-3990x-review/3
Sagittaire
8th February 2020, 23:35
I suspect that 5 instances of x265 encoding 1080p at the same time might not be enough for constant 100% cpu usage. If you overclock 3990X to 3.7 GHz then you should get extra ~50% vs 3970x.
I think that 5 instance for x265 1080p encoding can't saturate 128 thread. Perhaps time for update at 8 instance or even more with this really particular CPU. Even 3970x with 64 thread seem have problem. In good multithreaded application like cinebench, you have 100% improvement between 3950x and 3970x.
Atak_Snajpera
8th February 2020, 23:47
I think that 5 instance for x265 1080p encoding can't saturate 128 thread. Perhaps time for update at 8 instance or even more with this really particular CPU. Even 3970x with 64 thread seem have problem. In good multithreaded application like cinebench, you have 100% improvement between 3950x and 3970x.
The simplest solution would be reduction of ctu size from 64 to 32.
RanmaCanada
9th February 2020, 02:16
Or someone could test it on Linux, as it doesn't have the issues that windows does when it comes to high core count CPU's, even with "enterprise" edition it still has problems accessing all that power properly.
Nintendo Maniac 64
9th February 2020, 07:01
It could additionally be interesting to test it on normal non-enterprise Windows 10 but with SMT disabled.
NikosD
9th February 2020, 09:26
Summary:
Someone with access to a 3990X reviewer/ tester could ask for benchmarks using:
1) Windows 10 Pro
2) Windows 10 Enterprise
3) Linux
4) SMT disabled
excellentswordfight
9th February 2020, 10:35
I think that 5 instance for x265 1080p encoding can't saturate 128 thread. Perhaps time for update at 8 instance or even more with this really particular CPU. Even 3970x with 64 thread seem have problem. In good multithreaded application like cinebench, you have 100% improvement between 3950x and 3970x.
And "good multithreaded application" or maybe its more suted to call it calculations, are better of running on an GPU.
I suspect that 5 instances of x265 encoding 1080p at the same time might not be enough for constant 100% cpu usage. If you overclock 3990X to 3.7 GHz then you should get extra ~50% vs 3970x.
Yes that could be the case. But not sure about the practicality of that scenario... In one of the reviews I read is that they reached a maximum of 3.6Ghz for all core under load, that with a 280mm AIO and pullng 650W from the wall.
Could be the OS. Most versions of Windows 10 only allow for 64 threads in a single processor group.
https://www.anandtech.com/show/15483/amd-threadripper-3990x-review/3
Yes I read that article as well, and alot of others. It's still a guess game of why the resaults in that benchmark screenshot posted are the way they are.
3990X is a very cool CPU, but tbh its not a good buy for a workstation CPU *. Look at the benchmarks, you get good performance scaling in very few applications, 3970X is faster in quite a few cases cause of higher clock speeds and the applications were the scaling is good like in cinebench, yeah I rather just get octane and 3x 2080Ti if I were an c4d user...
* in general that is, iaf you have workloads were this makes sense for you, then yeah good for you.
Emulgator
9th February 2020, 23:42
What? Twice the price, twice the number of cores, 18% perf increase.
If its not a saturation/thread utilization issue, i wonder were the bottlneck is? Cache/ram bandwith?
I would have thought that thermal restrictions of these 9-die monsters would have capped any progress above the TR3960X around 5..10% per generation, so I am positively surprised by these +18% at the same 280W TDP.
(Price blanked out by me of course.)
I see the TR3990X base clock significantly limited to 2,9GHz for that reason...
(TR3960X was 3,8GHz, TR3970X was still 3,7GHz)
Well, maybe a bit unfair to compare multi-die CPUs against single-die CPUs anyway, but the race is on.
For bigger dies the manufacturer loses the bet against failures more often.
Yields of smaller dies seem more rewarding.
The best bet seems to be after the race: to compile a CPU out of the smaller die survivers;-)
Atak_Snajpera
10th February 2020, 15:27
Well, maybe a bit unfair to compare multi-die CPUs against single-die CPUs anyway, but the race is on.
Q6600 was also a multi-die CPU.
https://www.behardware.com/medias/photos_news/00/18/IMG0018284.jpg
Stereodude
24th February 2020, 16:09
I was able to benchmark my new R9 3950X system against my E5-2687W v2 in x265 using the settings I've been playing around with for high image quality encodes. Both are running stock clocks.
x265 version:
x265 [info]: HEVC encoder version 3.1.1+1-04b37fdfd2dc
x265 [info]: build info [Windows][GCC 9.1.1][64 bit] 8bit+10bit+12bit
x265 [info]: (libavcodec 58.53.101)
x265 [info]: (libavformat 58.28.101)
x265 [info]: (libavutil 56.30.100)
x265 [info]: (lsmash 2.16.1)
x265 options:
--pools 4 -F 1 --crf 16.0 -p veryslow --no-sao --aq-mode 1 --aq-strength 1.15 --vbv-maxrate 25000 --vbv-bufsize 25000 --level 5.0 --keyint 125 --open-gop -D 10 --colorprim "bt709" --transfer "bt709" --colormatrix "bt709" --sar 1:1
The source is a piece of a 1080i60 concert blu-ray that was restored to 1080p25 (1920x1080). I encoded the same 2500 frame segment multiple times and measured the time to encode. Four simultaneous encodes will effectively saturate the E5-2687W v2. It took an average time of 24001 seconds to encode the 2500 frames (four simultaneous encodes) on the E5-2687W v2. That gives a combined fps of the 4 simultaneous encodes of 0.417 fps. Four simultaneous encodes will hit 55-60% CPU on the R9 3950X. It took an average time of 7191 seconds to encode the 2500 frames (four simultaneous encodes) on the R9 3950X. That gives a combined fps of the 4 simultaneous encodes of 1.391 fps. Eight simultaneous encodes will effectively saturate the R9 3950X. It took an average time of 11302 seconds to encode the 2500 frames (eight simultaneous encodes) on the R9 3950X. That gives a combined fps of the 8 simultaneous encodes of 1.770 fps.
As a point of reference x265_FHD_Benchmark said my R9 3950X is ~4x faster than the E5-2687W v2 (77.49 fps vs. 19.05 fps).
TLDR / summary:
E5-2687W v2 - 0.417 fps (4 simultaneous encodes combined)
R9 3950X - 1.391 fps (4 simultaneous encodes combined)
R9 3950X - 1.770 fps (8 simultaneous encodes combined)
jfcarbel
15th March 2020, 00:25
Yep. However If you have x570 then you will have to add this step to get USB 3 ports to work
https://www.sevenforums.com/drivers/419688-will-new-amd-x570-chipset-mobos-have-win-7-chipset-drivers-post3441921.html?#post3441921
Hi Atak, finally getting to this build this weekend for 3900X
Is this still neccessary as I think the Jan 15 update you added x570 drivers. Doing this on a Asrock Taichi x570. Do the USB ports work after installing from an ISO updated using you latest now?
Seems there are modded drivers that will even allow for USB to work. Are these the ones you used?
https://www.win-raid.com/t4960f52-Solution-Win-Drivers-for-USB-Controllers-of-new-AMD-Chipset-Systems-5.html#msg99977
Also will this work with IDE SATA set to AHCI mode?
This is going to be installed on a Samsung NVMe drive.
Also any tips on a dual boot Win7 and Win10
As I notice that you mention to enable CSM on Motherboards to be able to install Win7, but this noticed this at AMD site. Is this true or is much of what they are saying negligible as well?
Also not sure if Win7/10 can peacefully co-exist on an NVMe. As I keep reading that in order for Win10 install to see NVMe drive you need UEFI.
For users running the AMD Ryzen Processor with Radeon Vega Graphics, AMD strongly recommends that your motherboard firmware (UEFI) be configured full UEFI Mode to ensure optimal performance, compatibility, and stability with the Windows 10 operating system.
Best pick for x265 4k encoding, Ryzen 5 3600 or Ryzen 7 2700x? How much fps I can expect from those? How much performance I would get if I go for Ryzen 7 3700x?
excellentswordfight
4th May 2020, 09:13
Best pick for x265 4k encoding, Ryzen 5 3600 or Ryzen 7 2700x? How much fps I can expect from those? How much performance I would get if I go for Ryzen 7 3700x?
3700x i quite a bit faster then 2700x, maybe about 30-40% cause of much better avx performance that was pretty limited in Ryzen 1000 and 2000-series. So i would say its worth the money.
How fast will it be? How long is a piece of string?
Between Ryzen 7 2700x and Ryzen 5 3600, which will give better performance? I have no clue about x265. I have to rip my 1080p and 4K BD collection.
Groucho2004
4th May 2020, 14:33
Between Ryzen 7 2700x and Ryzen 5 3600, which will give better performance? I have no clue about x265. I have to rip my 1080p and 4K BD collection.This (https://www.techpowerup.com/review/amd-ryzen-5-3600/13.html) might help.
Forteen88
5th May 2020, 00:06
This (https://www.techpowerup.com/review/amd-ryzen-5-3600/13.html) might help.And the comparisons here gives similar results,
https://forum.doom9.org/showthread.php?p=1908148#post1908148
MeteorRain
5th May 2020, 13:56
I have to say, zen2 is a great die.
crystalfunky
9th May 2020, 10:18
I have a 3700x and a B450 chipset with Crucial 3600 DDR4 16GB Ram.
I only get 2 fps when encoding 4K HDR10 videos. Is that normal? I'm encoding with Staxrip on preset slow.
microchip8
9th May 2020, 10:24
I have a 3700x and a B450 chipset with Crucial 3600 DDR4 16GB Ram.
I only get 2 fps when encoding 4K HDR10 videos. Is that normal? I'm encoding with Staxrip on preset slow.
Yes, it's normal
Sagittaire
9th May 2020, 11:22
I have a 3700x and a B450 chipset with Crucial 3600 DDR4 16GB Ram.
I only get 2 fps when encoding 4K HDR10 videos. Is that normal? I'm encoding with Staxrip on preset slow.
and it's really good speed for 4K HDR10 encoding.
You have 100% CPU usage?
RanmaCanada
9th May 2020, 21:48
I have a 3700x and a B450 chipset with Crucial 3600 DDR4 16GB Ram.
I only get 2 fps when encoding 4K HDR10 videos. Is that normal? I'm encoding with Staxrip on preset slow.
Not only is it normal, it's actually quite fast. What preset are you using?
crystalfunky
10th May 2020, 10:33
and it's really good speed for 4K HDR10 encoding.
You have 100% CPU usage?
Most of the time it's around 95 and 100.
Not only is it normal, it's actually quite fast. What preset are you using?
slow
blublub
15th May 2020, 06:55
Most of the time it's around 95 and 100.
slowFor preset slow that sounds about right.
crystalfunky
15th May 2020, 13:29
I'm encoding Black Hawk Down 4K HDR right now.
Preset slow
CRF 24, I get around 17 Mbit/s
FPS: around 2,6
59 % and it still says 9 hrs remaining.
Nico8583
26th June 2020, 10:10
2 FPS with a 3700X but anyone knows with a 3900X ?
excellentswordfight
26th June 2020, 13:43
2 FPS with a 3700X but anyone knows with a 3900X ?
No, but I'm using a 7402P, I get around 5fps for single file encodes using preset slow for 2160p content. Not 100% utilization though, around 80% (24C/48T).
7402P is actually very impressive, keeps clocks at 3.3Ghz in an 1u server under this rather heavy avx2 load. No intel high core count server/workstation system I've encoded on has been able to do so.
Boulder
26th June 2020, 19:04
With my 3900X and a script which has some motion-compensated denoising + downscaling to 1440p in linear light at 32bit precision and debanding, all in 16bit depth, I get over 2fps in x265 using --preset slower + hme-search umh,umh,star and slight tweaks. Using --rskip 2 makes things go faster and seems to be higher quality as the onion artifacts are less apparent. Then again, I use --rd-refine which should make things much slower.
benwaggoner
26th June 2020, 19:08
With my 3900X and a script which has some motion-compensated denoising + downscaling to 1440p in linear light at 32bit precision and debanding, all in 16bit depth, I get over 2fps in x265 using --preset slower + hme-search umh,umh,star and slight tweaks. Using --rskip 2 makes things go faster and seems to be higher quality as the onion artifacts are less apparent. Then again, I use --rd-refine which should make things much slower.
It'd be interesting to compare --preset ultrafast and see how much of the fps is x265 versus the preprocessing. It may well be the long pole isn't encoding.
Boulder
26th June 2020, 20:40
It'd be interesting to compare --preset ultrafast and see how much of the fps is x265 versus the preprocessing. It may well be the long pole isn't encoding.
Sure, I'll try to remember to do some tests after the chain of encodes finishes. Right now I'm processing a video like I described, and the frameserver part uses maybe an average of 20% of total CPU. Sometimes more, sometimes much less depending on how much the denoising part actually has to work on the frame. x265 uses 40-75% of CPU.
charliebaby
27th June 2020, 10:22
Me Benchmark Ryzen Threadrriper 2990wx 32 core 3700mhz x265 3.4.7 :-)
https://i.goopics.net/yWoZD.png
https://i.goopics.net/Z0mqG.png
Boulder
28th June 2020, 09:03
I did a little test of 500 frames, my standard Avisynth script template and encoded at 1440p (2560 x 1064).
Without any tweaks other than the HDR metadata, --preset ultrafast was 5.90 fps, --preset slower 2.81 fps.
With my x265 settings template, --preset ultrafast was 5.71 fps, --preset slower 2.63 fps.
Without any processing, i.e. just loading and decoding the source with DGSource and encoding at 4K (3840 x 1600), --preset ultrafast was 13.42 fps and --preset slower 1.54 fps using my x265 template.
Forteen88
29th June 2020, 10:28
Me Benchmark Ryzen Threadrriper 2990wx 32 core 3700mhz x265 3.4.7 :-)
https://i.goopics.net/yWoZD.png
https://i.goopics.net/Z0mqG.pngDo you normally encode with "Placebo" preset? When doing benchmarks, it's better to encode with the preset that you (or people in general) use mostly. I use almost only "Slower" preset.
charliebaby
29th June 2020, 17:23
Do you normally encode with "Placebo" preset? When doing benchmarks, it's better to encode with the preset that you (or people in generally) use mostly. I use almost only "Slower" preset.
YES juste PLACEBO Very Best :-)
DMD
18th January 2022, 22:18
Good evening.
I need to upgrade PC with CPU Ryzen 9 5900X, is it a good choice to encode H265 with reasonable fps?
Thank you
excellentswordfight
18th January 2022, 22:26
Good evening.
I need to upgrade PC with CPU Ryzen 9 5900X, is it a good choice to encode H265 with reasonable fps?
Thank you
Yes, either that or 12700k (with DDR4 if you want better bang for the buck).
https://tpucdn.com/review/intel-core-i7-12700k-alder-lake-12th-gen/images/encode-h265.png
I get about 20fps for 1080p Bluray re-encodes with x265 preset slow on a 12700k for reference, 5900X should be rather similar in terms of performance.
DMD
18th January 2022, 22:48
Yes, either that or 12700k (with DDR4 if you want better bang for the buck).
https://tpucdn.com/review/intel-core-i7-12700k-alder-lake-12th-gen/images/encode-h265.png
Personally I am interested in H265 4K encoding, I have no idea how much lower fps can be compared to 1080p (medium preset or slower?)
thank you
excellentswordfight
18th January 2022, 22:56
Personally I'm interested in H265 4K encoding, I have no idea how much lower fps can be than 1080p.
thank you
It will be much lower, I saw that you were talking about preset slower in the other thread as well, it will take a looong time to encode full length movies of thats your plan. I would guess that you only would get a few fps on an 5900x.
What speed is usable is individual I guess, but for me there isnt a use case for doing uhd bluray re-encodes with x265 at presets lower then slow for private use.
DMD
19th January 2022, 09:18
It will be much lower, I saw that you were talking about preset slower in the other thread as well, it will take a looong time to encode full length movies of thats your plan. I would guess that you only would get a few fps on an 5900x.
What speed is usable is individual I guess, but for me there isnt a use case for doing uhd bluray re-encodes with x265 at presets lower then slow for private use.
Good morning.
To understand and have a reference with the hardware, I did a test with sample video @ 23.976fps with the current PC (Intel Core i5 9600K \ GeForce RTX 2070).
I set the StaxRip preset to Slower, the process is 0.1 FPS.
MISSION IMPOSSIBLE!!! :(
https://i.postimg.cc/tRVz0PsG/Screenshot-19-01-2022-08-34-42.png
Making a prediction of around 2 FPS (I hope) with a Ryzen 9 5900X, I made this reasoning, if that's correct.
In the duration of a second of movie there are about 24fps and the processor processes 2 FPS, it means that the ratio is 1:12, so a movie with a duration of 2 hours needs 24 hours of processing, is that correct?
If necessary, it is possible to divide the source file into several parts, in order to do the processing at different times and then carry out the final union?
Thank you
Boulder
19th January 2022, 09:30
Some speedups: remove --hme (EDIT: note not --umh but --hme) and --rd-refine, set --no-amp --limit-modes 3 --ref 4. I'm not sure --pmode and --pme help at all. Set --ctu 32 for better threading (no, it won't reduce quality).
And set --no-sao unless you like to have all the details removed.
Blue_MiSfit
19th January 2022, 09:42
SAO is generally quite useful at lower bitrates. x265's implementation is not perfect, but it's net positive a lot of the time :)
When you're gunning for total transparency it may indeed be good to disable it.
excellentswordfight
19th January 2022, 09:54
Good morning.
Making a prediction of around 2 FPS (I hope) with a Ryzen 9 5900X, I made this reasoning, if that's correct.
In the duration of a second of movie there are about 24fps and the processor processes 2 FPS, it means that the ratio is 1:12, so a movie with a duration of 2 hours needs 24 hours of processing, is that correct?
If necessary, it is possible to divide the source file into several parts, in order to do the processing at different times and then carry out the final union?
Thank you
Yes, I would guess that 2fps isnt that far off (although I would think that it would be closer to 1). And yes that calculation would be correct. Slow would be about 2x faster then Slower, if you actually do some tests I would think you will see that it will not be worth the speed penalty.
And yes there are ways were you can do it in parts. Some easy (like pausing the encode), and some more complex (chunk encoding).
edit. Did a quick test and 1,2fps for 2160p with preset slower, and 4fps for preset slow. 12700k@125w. I noticed though with x265 that speed is a bit dependent on source complexity so your mileage may vary.
Set --ctu 32 for better threading (no, it won't reduce quality).
As we talking about UHD at low presets CPU saturation shouldnt be an issue even with --ctu 64 for 5900X/12700k. I've found scaling to be decent up to about 16c/32t at default settings (1080p up to about 8c/16t), then the performance boost starts to take a drastic turn.
RanmaCanada
20th January 2022, 02:00
I know on my 5800x it will take about 24 hours to do a 4k encode with preset slow. If OP wants to make an encoding only machine, I would honestly say at this point get a Threadripper or EPYC and use Ripbot to encode. But that's insane talk as no one wants to spend that much on a machine that will be beaten in say 5 years by consumer grade hardware, unless of course this is a tax write off haha.
benwaggoner
22nd January 2022, 02:32
I know on my 5800x it will take about 24 hours to do a 4k encode with preset slow. If OP wants to make an encoding only machine, I would honestly say at this point get a Threadripper or EPYC and use Ripbot to encode. But that's insane talk as no one wants to spend that much on a machine that will be beaten in say 5 years by consumer grade hardware, unless of course this is a tax write off haha.
Systems used for professional encoding get replaced a lot more often than every five years due to the big improvements in encoding time and watts/pixel over just a couple CPU generations. Heck, power savings alone would likely justify more frequent updates for a hobbyist that has their computers doing video encoding even 50% of the time.
Nico8583
14th February 2022, 13:27
Hi :)
Now there is a new Intel gen, what is the best choice to encode 4K ? I'm still on a B450M + 3700X combo but I would like to know which one between 5900X or i9 12900 would be the best choice if I upgrade ? The 5900X could be used on the B450M platform. Or perhaps a 3900X used if price is good.
Thanks !
D3X
14th February 2022, 14:09
Is there a possibility to render x265 on the GPU?
Cuda would be very nice. I don't want a lecture on why it would be bad,
just wondering if anyone has that bitbucket.org CUDA on GPU x265 repo or so.
Thanks!
D3X
14th February 2022, 14:10
https://bitbucket.org/vovagubin/x265-hevc-opencl-or-cuda-encoder
RanmaCanada
14th February 2022, 17:45
Is there a possibility to render x265 on the GPU?
Cuda would be very nice. I don't want a lecture on why it would be bad,
just wondering if anyone has that bitbucket.org CUDA on GPU x265 repo or so.
Thanks!
If you're going to use a GPU to encode, just use quicksync. It will be far faster and the quality will be pretty darn close, especially with the new ASICS in Alder Lake.
RanmaCanada
14th February 2022, 17:51
Hi :)
Now there is a new Intel gen, what is the best choice to encode 4K ? I'm still on a B450M + 3700X combo but I would like to know which one between 5900X or i9 12900 would be the best choice if I upgrade ? The 5900X could be used on the B450M platform. Or perhaps a 3900X used if price is good.
Thanks !
The 12900 beats a 5900x (https://www.techpowerup.com/review/intel-core-i9-12900k-alder-lake-12th-gen/14.html), albeit with power usage out the wazoo (https://www.techpowerup.com/review/intel-core-i9-12900k-alder-lake-12th-gen/20.html). But then you also get access to quicksync. If the cost of your electricity is extremely high, ie you live in Europe, the 12900 is a really bad choice for software encoding as it uses almost twice the power of a 5900x when encoding.
Nico8583
14th February 2022, 19:56
Thanks, I live in Europe... :rolleyes:
D3X
15th February 2022, 03:28
If you're going to use a GPU to encode, just use quicksync. It will be far faster and the quality will be pretty darn close, especially with the new ASICS in Alder Lake.
Okay, let's pretend I'm a total newb with the QS commands.
What would you recommend?
I have a RTX 3070 Ti
tormento
15th February 2022, 12:51
Or wait for a discrete Intel GPU with QuickSync.
asarian
15th February 2022, 13:40
Defintely the i9 12900K, hands down. Is almost twice as fast as the i7 10700K I (briefly) had before, when it comes to x265 encoding.
SquallMX
16th February 2022, 17:32
Making a prediction of around 2 FPS (I hope) with a Ryzen 9 5900X, I made this reasoning, if that's correct.
In the duration of a second of movie there are about 24fps and the processor processes 2 FPS, it means that the ratio is 1:12, so a movie with a duration of 2 hours needs 24 hours of processing, is that correct?
If necessary, it is possible to divide the source file into several parts, in order to do the processing at different times and then carry out the final union?
Thank you
I can confirm that, with a 5900x without OC you get 2.xx-3.xx fps depending on the movie content, usually grainy ones are the slowest, --medium preset with a few tweaks gets you 5.xx-6.xx fps without noticeable differences.
And no, Hardware Encoding (QS / nVidia NVENC) is not in the same ballpark quality-wise.
RanmaCanada
16th February 2022, 23:38
Okay, let's pretend I'm a total newb with the QS commands.
What would you recommend?
I have a RTX 3070 Ti
I can't give you any commands that will work with QS as your 3070Ti doesn't have it. It has NVENC, totally different ASICS.
D3X
17th February 2022, 10:47
I can't give you any commands that will work with QS as your 3070Ti doesn't have it. It has NVENC, totally different ASICS.
That's why I desperately need that BitBucket repo I linked to earlier. :)
excellentswordfight
17th February 2022, 13:27
That's why I desperately need that BitBucket repo I linked to earlier. :)
Well that didnt use nvenc/ASICS, if I remembered correctly it was just some experimental fork of x265 that offloaded some work utilizing opencl/cuda that never got any real development afaik. Even if you find it, it will be based on an very old x265 version that could give you some unexpected results.
This is an nvenc based encoder https://github.com/rigaya/NVEnc/releases
Defintely the i9 12900K, hands down. Is almost twice as fast as the i7 10700K I (briefly) had before, when it comes to x265 encoding.
Although 12900k is indeed very fast, its worth having in mind what RanmaCanada wrote. When the alder lake models goes above about 125W it starts to throw the efficiency out the window, and if you take the efficiency in account 12700k is much more compelling option for encoding as you probably want to keep pl1 rather conservative anyway. When my 12700k drops from pl2 (~200W) to pl1 (125W) after about 60s the decrease in performance is no were near the drop in power consumption.
5900X is probably still the king for consumer CPUs when it comes to x264/x265 encoding giving the efficiency factor. But I would say that 12700k is still an very valid option. I would only go for the 12900k if pure speed is the pretty much the only criteria.
Nico8583
17th February 2022, 15:34
Thanks all for feedback about 12700K / 12900K / 5900X
Selur
18th February 2022, 10:14
if I remembered correctly it was just some experimental fork of x265 that offloaded some work utilizing opencl/cuda that never got any real development afaik.
I agree, as far as I know there never was any usable code in it.
benwaggoner
18th February 2022, 18:17
I agree, as far as I know there never was any usable code in it.
Yeah, I think it had a maximum theoretical speed boost of only 20%, and there was lower, more portable low-hanging fruit for optimization. All the analysis save/load, reusing of data from an existing h.264 encode, etc all can deliver bigger performance boosts. And using a HW encoder as input to refine is pretty similar conceptually, and can work on most GPUs that would support sufficiently performant CUDA or OpenCL.
speedy
23rd February 2022, 23:04
Is there somewhere I can go to see how my i9-9900K compares with newer processors for 4K x265 encoding or does anyone here have data on this? Thanks!
Blue_MiSfit
24th February 2022, 01:20
I have a 9900k at work and a Ryzen 9 5900x at home (12 core). I'll do a quick test with Tears of Steel 4k :)
Both have pretty powerful AIO water coolers so they can maintain decent boost clocks.
RanmaCanada
24th February 2022, 02:16
Is there somewhere I can go to see how my i9-9900K compares with newer processors for 4K x265 encoding or does anyone here have data on this? Thanks!
https://www.techpowerup.com/review/intel-core-i9-9900k/6.html it's a little faster than a 2700x, and slower than a 3900x https://www.techpowerup.com/review/amd-ryzen-9-3900x/13.html and roughly half the speed of a 5900x https://www.techpowerup.com/review/amd-ryzen-9-5900x/14.html at 1080p so 4k would be of course "about" the same.
You can also always run the old benchmark to see how you stack up in "real world" terms..
https://forum.doom9.org/showthread.php?t=174393
Blue_MiSfit
24th February 2022, 03:26
Ok, so I did preset slower and crf 20 (everything else at defaults). This was using a recent ffmpeg build with x265 3.5. The source is the 4k SDR 4:2:0 8 bit ~80 Mbps H.264 version available from the Tears of Steel site.
- The 9900k gets about 0.8 fps
- The 5900x gets about 1.4 fps.
Looks like the better IPC and 12 cores of the Ryzen 9 vs 8 cores compensates well for the lower clock speed.
The 9900k can sustain 4.68 GHz all-core turbo whereas the 5900x can hover around 4.2. Both systems have beefy water coolers, the Corsair HX115i if memory serves...
CPU usage was 100% in both cases.
excellentswordfight
24th February 2022, 15:21
Ok, so I did preset slower and crf 20 (everything else at defaults). This was using a recent ffmpeg build with x265 3.5. The source is the ~80 Mbps H.264 version available from the Tears of Steel site.
- The 9900k gets about 0.8 fps
- The 5900x gets about 1.4 fps.
Looks like its better IPC and 12 cores vs 8 cores offsets the lower clock speed. The 9900k can sustain 4.68 GHz all-core turbo whereas the 5900x can hover around 4.2.
CPU usage was 100% in both cases.
Although one should be careful to do comparisons with so many different variables. I get about 1.4 when not power limited and 1.2 @ 125W on a 12700k with preset slower for ToS, and that seems to line up with the x265 results techpowerup publishes when compared to your numbers.
Although I still dont think consumer CPU:s is fast enough for preset 'slower' to make sense for UHD-encoding for private use, we are definitely starting to getting usable performance for 'slow' now (~4fps, so about 12h encoding time for 2h title).
speedy
24th February 2022, 15:53
That's impressive. Looks like I'm due for an upgrade soon. I'm sure I can find a home for my 9900K in another computer.
speedy
9th March 2022, 22:56
I just finished upgrading my 9900K to a 12900K and am now wondering how to get the most out of it while encoding with x265.
Any advice?
Thanks!
RanmaCanada
10th March 2022, 16:24
I just finished upgrading my 9900K to a 12900K and am now wondering how to get the most out of it while encoding with x265.
Any advice?
Thanks!
If you're not seeing 100% utilization then use Ripbot to divy up the job and do it in chunks?
speedy
18th March 2022, 20:26
Does anyone else here have experience with x265 encoding on Intel 12th Gen Alder Lake CPUs?
I'm running into issues with my encodes being moved to the E-cores (Efficiency cores) when disconnect from my Windows user session? I typically RDP (Remote Desktop Protocol) into the encoding server and check on and setup my encodes, but after disconnecting it seems like the encode workloads are moved to the E-cores. If I stay logged in then the encodes run on the P-cores (Performance Cores). I've been able to workaround this for now by using Chrome Remote Desktop which makes Windows think I'm logging into the machine locally because when I disconnect it doesn't realize I'm no longer "remoted in".
Encodes seem to take 4 times longer running on the E-cores so I'd really like to figure this out.
P.S. I'm using Handbrake 1.5.1
mastrboy
20th March 2022, 01:06
Does anyone else here have experience with x265 encoding on Intel 12th Gen Alder Lake CPUs?
I'm running into issues with my encodes being moved to the E-cores (Efficiency cores) when disconnect from my Windows user session? I typically RDP (Remote Desktop Protocol) into the encoding server and check on and setup my encodes, but after disconnecting it seems like the encode workloads are moved to the E-cores. If I stay logged in then the encodes run on the P-cores (Performance Cores). I've been able to workaround this for now by using Chrome Remote Desktop which makes Windows think I'm logging into the machine locally because when I disconnect it doesn't realize I'm no longer "remoted in".
Encodes seem to take 4 times longer running on the E-cores so I'd really like to figure this out.
P.S. I'm using Handbrake 1.5.1
Not really a solution, but maybe a decent workaround is using Process Lasso to pin the cores: https://bitsum.com
RanmaCanada
20th March 2022, 20:12
Does anyone else here have experience with x265 encoding on Intel 12th Gen Alder Lake CPUs?
I'm running into issues with my encodes being moved to the E-cores (Efficiency cores) when disconnect from my Windows user session? I typically RDP (Remote Desktop Protocol) into the encoding server and check on and setup my encodes, but after disconnecting it seems like the encode workloads are moved to the E-cores. If I stay logged in then the encodes run on the P-cores (Performance Cores). I've been able to workaround this for now by using Chrome Remote Desktop which makes Windows think I'm logging into the machine locally because when I disconnect it doesn't realize I'm no longer "remoted in".
Encodes seem to take 4 times longer running on the E-cores so I'd really like to figure this out.
P.S. I'm using Handbrake 1.5.1
I would try other programs like staxrip, hybrid, fastflix, xmedia recode, etc, to see if it's program related or windows related.
benwaggoner
23rd March 2022, 03:49
If you're not seeing 100% utilization then use Ripbot to divy up the job and do it in chunks?
Does Ripbot properly use --chunk-start and --chunk-end to avoid quality issues at stitch points?
Ala https://patentimages.storage.googleapis.com/29/05/b6/ed3668747b35b0/US10863179.pdf
Blue_MiSfit
24th March 2022, 02:19
Nice patent, Ben ;)
RanmaCanada
25th March 2022, 05:13
Does Ripbot properly use --chunk-start and --chunk-end to avoid quality issues at stitch points?
Ala https://patentimages.storage.googleapis.com/29/05/b6/ed3668747b35b0/US10863179.pdf
I believe it does but Atak_Snajpera would be the best person to answer this question.
speedy
13th April 2022, 03:23
Not really a solution, but maybe a decent workaround is using Process Lasso to pin the cores: https://bitsum.com
I would try other programs like staxrip, hybrid, fastflix, xmedia recode, etc, to see if it's program related or windows related.
I ended up figuring out that you need to run these two commands in Windows for Intel 12th Gen (Alder Lake) CPUs:
POWERCFG /POWERTHROTTLING DISABLE /PATH "c:\Program Files\HandBrake\HandBrake.Worker.exe"
POWERCFG /POWERTHROTTLING DISABLE /PATH "c:\Program Files\HandBrake\HandBrake.exe"
DMD
11th October 2022, 13:55
Good morning.
Should debut i9-13900K in days, I ask where I could find a comparison with x265 encoding vs Ryzen 9 7950X
Thanks
excellentswordfight
11th October 2022, 14:10
Good morning.
Should debut i9-13900K in days, I ask where I could find a comparison with x265 encoding vs Ryzen 9 7950X
Thanks
Of the big review sites I think that techpowerup does the best x265 test (using preset slow and crf 20), so I would check there once its reviewed.
https://www.techpowerup.com/review/amd-ryzen-9-7950x/17.html
tormento
12th October 2022, 15:32
I'd really like to see some Alder Lake or other Intel AVX 512 vs Zen 4 with AVX512 results.
Anyone? :)
rwill
12th October 2022, 16:58
I'd really like to see some Alder Lake or other Intel AVX 512 vs Zen 4 with AVX512 results.
Anyone? :)
Alder Lake does not support AVX 512 ?
tormento
12th October 2022, 22:43
Alder Lake does not support AVX 512 ?
Early version does, disabling E-cores.
RanmaCanada
13th October 2022, 05:03
I'd really like to see some Alder Lake or other Intel AVX 512 vs Zen 4 with AVX512 results.
Anyone? :)
You won't be able to get it as Intel in their wisdom has decided end users do not deserve to have access to it. Very few chips support it, and those that did will also have it disabled if users update their BIOS, with no means to flash back, as per Intel's instructions to manufacturers.
There is no point in the benchmark as the chips are unobtainable at this time.
excellentswordfight
13th October 2022, 11:14
Early version does, disabling E-cores.
In that case AVX512 will be slower, you loose about 30% performance by disabling the e-cores for x265-encoding on i7/i9. I very much doubt that AVX512 gives that kinds of performance gains, it never even accommodated for the frequency loss for all the AVX512 test ive done on xeons. On something like an i9 that already is pushed power-wize, if it can keep the frequency under avx512 power consumption must go through the roof.
For zen4, the test iŽve seen the avx512 implementation seems rather good from a power-perspektive, but there doesnt seems to result in that much gain for x265.
RanmaCanada
15th October 2022, 05:37
In that case AVX512 will be slower, you loose about 30% performance by disabling the e-cores for x265-encoding on i7/i9. I very much doubt that AVX512 gives that kinds of performance gains, it never even accommodated for the frequency loss for all the AVX512 test ive done on xeons. On something like an i9 that already is pushed power-wize, if it can keep the frequency under avx512 power consumption must go through the roof.
For zen4, the test iŽve seen the avx512 implementation seems rather good from a power-perspektive, but there doesnt seems to result in that much gain for x265.
In 2018 Intel released a white paper (https://www.intel.com/content/www/us/en/developer/articles/technical/accelerating-x265-with-intel-advanced-vector-extensions-512-intel-avx-512.html) that touted the speed increase AVX512 would give processors in encoding HEVC. It's funny how they not only took this away from consumers, but it didn't generate the speed increases they tried to market.
tormento
15th October 2022, 10:16
In that case AVX512 will be slower, you loose about 30% performance by disabling the e-cores for x265-encoding on i7/i9.
Watch derbau8r video on youtube about alder lake speed in avx512 :)
benwaggoner
16th October 2022, 22:23
In 2018 Intel released a white paper (https://www.intel.com/content/www/us/en/developer/articles/technical/accelerating-x265-with-intel-advanced-vector-extensions-512-intel-avx-512.html) that touted the speed increase AVX512 would give processors in encoding HEVC. It's funny how they not only took this away from consumers, but it didn't generate the speed increases they tried to market.
The AVX512 instruction set is quite powerful, and really does improve performance per clock with x265, particularly higher resolutions. Intel's problem is that using AVX512 quickly lowers the clock speed a lot.
The instruction set will probably become as essential as AVX2 eventually, as thermal issues get addressed. AVX2 had much the same problem when it first launched, which was much improved in the next major revision.
But it's not like the typical consumer would miss AVX2 instructions either for the most part. It's going to be much more valuable for workstations and compute-optimized servers.
frencher
17th October 2022, 01:28
Hello all,
AMD Threadripper 3990X or Intel i5 13600k for 4K encoding ?
What would be the frame rate for the i5 13600k in 4k encoding ?
A result under x265 FHD Benchmark ?
Thanks
excellentswordfight
17th October 2022, 08:48
Watch derbau8r video on youtube about alder lake speed in avx512 :)
I have, and I dont remember him testing x265. Just cause one load gets a huge speedup by avx512 doesnt means another does. As some that has benchmarked avx512 on xeon I can tell you that the speedup of avx512 in x265 is not big.
Here is a x265 benchmark for avx512 on Zen4
https://i.ibb.co/qCJy7hL/x265.jpg
And in case you would just assume that zen4 has bad avx512 performance in general here is one with 3DPM:
https://images.anandtech.com/graphs/graph17585/130235.png
benwaggoner
18th October 2022, 19:31
MCW's guidence some years back was that AVX512 was generally a speed regression on Intel CPUs of that era for anything below 4K --preset veryslow.
I've got some 8K --preset placebo test scripts that I can try with AVX512 on/off when I get back from my work trip this Friday.
benwaggoner
18th October 2022, 19:50
I have, and I dont remember him testing x265. Just cause one load gets a huge speedup by avx512 doesnt means another does. As some that has benchmarked avx512 on xeon I can tell you that the speedup of avx512 in x265 is not big.
Here is a x265 benchmark for avx512 on Zen4
Do you know what resolution was being encoded with what preset? AVX512 should help more with more pixels and with more complex presets.
excellentswordfight
19th October 2022, 08:44
Do you know what resolution was being encoded with what preset? AVX512 should help more with more pixels and with more complex presets.
No, but judging from the fps its either 1080p at a slow preset, or 2160p at maybe medium? But as the speed is really depandant on source complexity its hard to say,
Whats interesting is that zen4 doesnt suffer from downclocking. It seems like Zen4 has a very different implementation of AVX512 which makes it much more power efficient.
"But rather going for a 512-bit FPU data path and the possibility of reduced clock frequencies and power/thermal concerns, they employed a 256-bit "double pumping" strategy."
"Here is a look at the CPU peak frequency across the entire span of benchmarks tested... No real change compared to without AVX-512"
"Likewise, the CPU power consumption was similar when running the AVX-512 enabled software. For nearly all the tests the AVX2 vs. AVX-512 results were almost identical"
https://www.phoronix.com/review/amd-zen4-avx512/6
This is the speed I got on Xeon the last time i did a test (ToS 2160p @ preset slow):
Intel Xeon "Cascade Lake Refresh" 2x6226R 16c/32T 150W MSRP 1300USD (each).
avx256: 4,85fps
avx512: 4,49fps
AMD EPYC "ROME" 7502P 32c/64t 180W MSRP 2300USD
avx256: 6,26fps
Unfortunately we dont have any Ice lake-SP models I can test on as we have started to switch to Epyc. I would love to see a comparison between Xeon Gold 6314U & EPYC 7543P.
excellentswordfight
20th October 2022, 15:16
Good morning.
Should debut i9-13900K in days, I ask where I could find a comparison with x265 encoding vs Ryzen 9 7950X
Thanks
https://tpucdn.com/review/intel-core-i9-13900k/images/encode-h265.png
Thats a win for 7950X, 13900k needs close to 400W be able to push infront of it, in stock its still a bit more power hungry than 7950X.
benwaggoner
20th October 2022, 19:39
https://tpucdn.com/review/intel-core-i9-13900k/images/encode-h265.png
Thats a win for 7950X, 13900k needs close to 400W be able to push infront of it, in stock its still a bit more power hungry than 7950X.
Did you try --avx512 on the 13900? I'm curious how it might have improved the thermal throttling behaviors to increase throughput.
Zebulon84
20th October 2022, 20:29
Did you try --avx512 on the 13900?
If I understand correctly intel i9-13900K specifications (https://www.intel.com/content/www/us/en/products/sku/230496/intel-core-i913900k-processor-36m-cache-up-to-5-80-ghz/specifications.html#specs-1-0-7), the Instruction Set Extensions include AVX2 but not AVX 512.
benwaggoner
20th October 2022, 22:29
If I understand correctly intel i9-13900K specifications (https://www.intel.com/content/www/us/en/products/sku/230496/intel-core-i913900k-processor-36m-cache-up-to-5-80-ghz/specifications.html#specs-1-0-7), the Instruction Set Extensions include AVX2 but not AVX 512.
Oh, right.
When are the 13th gen equivalent Xeons coming out? Those should still have AVX512.
excellentswordfight
21st October 2022, 06:46
Oh, right.
When are the 13th gen equivalent Xeons coming out? Those should still have AVX512.
Sapphire Rapids is the code name for next gen xeon-sp, and the follow up for Ice lake-sp. It will be built on the same node and have the golden cove cores found in gen12 ”core” processors (which are pretty much the same as 13th gen) I think its been delayed again to q1/q2 next year. It will also use MCM (multi-chip module), so core count in top models should increase by a alot.
4th-gen Epyc, Genoa, is due for release soon as weel and will also feature avx512. And given that zen4 on ryzen doesnt downclock at all that will be rather intresting as well.
tormento
22nd October 2022, 09:39
4th-gen Epyc, Genoa, is due for release soon as weel and will also feature avx512.
Notice that AMD implementation of AVX512 is a sort of AVX256*2.
I have yet to see the tests.
tormento
22nd October 2022, 09:40
Thats a win for 7950X, 13900k needs close to 400W be able to push infront of it, in stock its still a bit more power hungry than 7950X.
I am so curious to see a x265 encode with the so called "eco mode" enabled for both.
Stereodude
22nd October 2022, 14:42
https://tpucdn.com/review/intel-core-i9-13900k/images/encode-h265.png
Thats a win for 7950X, 13900k needs close to 400W be able to push infront of it, in stock its still a bit more power hungry than 7950X.
If you're concerned about power efficiency with x265 encoding the 5950X is likely the best choice if your running stock default CPU settings. It's not 2x slower, but uses like half the power of these new CPUs.
tormento
22nd October 2022, 16:09
If you're concerned about power efficiency with x265 encoding the 5950X is likely the best choice if your running stock default CPU settings. It's not 2x slower, but uses like half the power of these new CPUs.
Raptor Lake has a 90W mode, Zen4 a 80ish one.
I mean those eco modes.
hajj_3
22nd October 2022, 19:57
https://images.anandtech.com/graphs/graph17601/130512.png
https://images.anandtech.com/graphs/graph17601/130513.png
https://cdn.mos.cms.futurecdn.net/kYvxyXsFiuRG3jTwNRwjQa-1200-80.png.webp
https://cdn.mos.cms.futurecdn.net/8Ke2LVJ2vA2CGpq4dor5Va-1200-80.png.webp
New intel raptor lake looks quite good but they use way more power than the 125w that they claim:
https://images.anandtech.com/graphs/graph17601/130462.png
nevcairiel
23rd October 2022, 06:55
Raptor Lake has a 90W mode, Zen4 a 80ish one.
I mean those eco modes.
You can also just input whatever power limit you want to run at, like an actual number.
Boulder
23rd October 2022, 12:45
New intel raptor lake looks quite good but they use way more power than the 125w that they claim
That's TDP which is different from the actual power usage.
DMD
3rd November 2022, 13:57
Of the big review sites I think that techpowerup does the best x265 test (using preset slow and crf 20), so I would check there once its reviewed.
https://www.techpowerup.com/review/amd-ryzen-9-7950x/17.html
Thank you.
I need to evaluate the trade-off between performance and power dissipation.
For Ryzen 95°C is a very high temperature, especially for a processor built on the 5 nm node,
that's what I read.
https://www.techpowerup.com/review/amd-ryzen-9-7950x-cooling-requirements-thermal-throttling/
Boulder
3rd November 2022, 14:48
From what I know, AMD went all-in with performance and the power usage (= heat) part was neglected. I think quite a few tests have been made and if you lower the PPT limit in PBO, you get almost the same performance for much less consumed watts.
Running a proper Curve Optimizer set will lower the power usage even further, but it will take some days to fine tune it. Then you can put some negative core voltage offset as well, it seems. My 5950X runs happily with -0.0825V without any performance degradation, then I did the CO tuning on top of that with many cores going all the way to -30 there.
DMD
3rd November 2022, 15:18
So you can still act on the core rension parameters to get a good compromise between heat and performance?
Because it was the excessive power dissipation that was worrying me a bit.
excellentswordfight
3rd November 2022, 17:39
Thank you.
I need to evaluate the trade-off between performance and power dissipation.
For Ryzen 95°C is a very high temperature, especially for a processor built on the 5 nm node,
that's what I read.
https://www.techpowerup.com/review/amd-ryzen-9-7950x-cooling-requirements-thermal-throttling/
So you can still act on the core rension parameters to get a good compromise between heat and performance?
Because it was the excessive power dissipation that was worrying me a bit.
"The biggest problem is probably psychological. For years we have been trained that "95°C is bad". This is no longer true. 95°C is the new 65°C. "
These new processors are designed to boost until they hit these temperatures as long as the power target allows for it if. So wouldnt worry about it.
If you are worried about power dissipation look at the power draw, not the temperature, if a CPU consumes 200W it will dissipates as much heat/energy if it running at 65C as 95C.
And as Boulder mentioned you can limiit the powerdraw on most new processors without much loss i performance. I have PL1 set at 200W (60s duration max) and PL2 at 125W (long time load) on my 12700k, when encoding the performance difference is not big when it dropps down to 125W, and fans goes pretty much silient with air cooler.
DMD
4th November 2022, 08:17
"The biggest problem is probably psychological. For years we have been trained that "95°C is bad". This is no longer true. 95°C is the new 65°C. "
These new processors are designed to boost until they hit these temperatures as long as the power target allows for it if. So wouldnt worry about it.
If you are worried about power dissipation look at the power draw, not the temperature, if a CPU consumes 200W it will dissipates as much heat/energy if it running at 65C as 95C.
And as Boulder mentioned you can limiit the powerdraw on most new processors without much loss i performance. I have PL1 set at 200W (60s duration max) and PL2 at 125W (long time load) on my 12700k, when encoding the performance difference is not big when it dropps down to 125W, and fans goes pretty much silient with air cooler.
I thank you for reassuring me.
Then I will consider what kind of dissipation to use , whether fan or liquid (AIO), but that is another topic.
HD MOVIE SOURCE
10th November 2022, 07:22
Just wondering, can any cpu encode 2 hours worth of footage in 24 hours or less and use at least very slow with x265?
hajj_3
10th November 2022, 08:07
Just wondering, can any cpu encode 2 hours worth of footage in 24 hours or less and use at least very slow with x265?
the resolution and bitrate of the source video would be needed to answer that and also the output resolution.
RanmaCanada
10th November 2022, 16:25
Just wondering, can any cpu encode 2 hours worth of footage in 24 hours or less and use at least very slow with x265?
Quite possibly an 5950x or 7950x with chunking being used. But again it also depends on the resolution and the bitrate, along with any switches. I know my 5800x will take 2-3 days to encode 4k anime at very slow, averaging about 1fps.
excellentswordfight
10th November 2022, 17:26
Quite possibly an 5950x or 7950x with chunking being used. But again it also depends on the resolution and the bitrate, along with any switches. I know my 5800x will take 2-3 days to encode 4k anime at very slow, averaging about 1fps.
At 4k I usually get up to 80-100% usage on 24C/48T when leaving thread-affecting switches as default, so I dont think its a necessity to use chunk encoding for 16C models (although it will probably give you a slight speed boost).
With that said, might be possible on a high core count threadripper using chunk encoding. I get about 1fps on a 24C epyc for standard complexity content at veryslow (complexity has a huge impact on speed in these cases). So I dont 7950X is enough to break 2fps.
But if HD MOVIE SOURCE still does encodes at 98Mbps CBR, I dont see any reason of using veryslow in the first place. If the output has visual issues, its not cause slower is used over veryslow.
benwaggoner
10th November 2022, 22:13
Yeah, these things can be hard to predict. When using --crf, a clean source like anime can encode a lot faster than a grainy content with the same parameters. Entropy coding is significant, and that's proportional to net bitrate. And getting a good skip match early-exits a whole lot of compute that random noise precludes. Although in the grainy 4K case, --rd 4 can sometimes deliver better quality than the --rd 6 used in --preset slower and above, and it goes quite a bit faster. Reducing --frame-threads often doesn't have that much of a speed impact at high resolutions as the overhead of frame threading makes each thread slower than without frame threading.
Using one of my "Xeon Gold 6240 CPU @ 2.60GHz, 2594 Mhz, 18 Cores, 36 Logical Processors" with my default 4K settings for undefined content tuned for "take more time wherever it makes a potentially visible improvement" does a little more than an hour a day. I'm sure a more modern CPU could do better. 16 cores at a higher clock speed would be a bit faster yet; it seems to use about 12 cores on average, with spikes up and down.
HD MOVIE SOURCE
14th November 2022, 17:42
Okay, my assumption is that with x265 encoding we're talking about content that is 4K, but I should have made that clear. So, what kind of CPU would be needed for 24 to 48 hour encode, with a 2 hour movie, and if you need bit-rate, my target bit-rate would be 98 Mbps, just like 4K UHD-BD.
benwaggoner
14th November 2022, 18:32
Okay, my assumption is that with x265 encoding we're talking about content that is 4K, but I should have made that clear. So, what kind of CPU would be needed for 24 to 48 hour encode, with a 2 hour movie, and if you need bit-rate, my target bit-rate would be 98 Mbps, just like 4K UHD-BD.
With that bitrate, you can target a lot higher performance without visual loss. With --vbv-maxrate 98000, a simple --crf 14 --preset medium should be pretty well transparent.
Making 4K look good with a 10 Mbps peak is a lot harder, which is where extra tools with extra performance implications kick in. Throwing 9.8x more bits at the problem means the encoder starts out with much, much lower QPs, which is the classic brute force way of improving quality.
excellentswordfight
15th November 2022, 13:54
Okay, my assumption is that with x265 encoding we're talking about content that is 4K, but I should have made that clear. So, what kind of CPU would be needed for 24 to 48 hour encode, with a 2 hour movie, and if you need bit-rate, my target bit-rate would be 98 Mbps, just like 4K UHD-BD.
As its seems like you are ignoring a lot o the answers provided to you, but im gonna give it a lost shot.
First of all I answered your CPU question above, with high complexity 4k material no CPU on the market is likely to break 2fps at veryslow, it can be done with chunk encoding if you go for something like a Threadripper 5975WX or 5995WX, if go above the 24h requirement something like a 7950X will probably get you there without chunkencoding, I will guestimate that you might get close to 2fps for preset veryslow.
Secondly, the gains of presets like veryslow are very much in the diminishing returns territory, and when we are talking about such high bitrates its even beyond overkill imo. As I said, if you have issues with the output at 98Mbps at something like slow or slower, I would look at other parameters then going to preset veryslow. All uhd-bluray encodes ive done has been visually lossless at slow and slower at much lower bitrates.
Its still very unclear what the goal is here, complying to the uhd-bd specifications are only limitations, there is no real reason to use those unless you really are authoring for actual physical discs. If the reason is of more academic purpose, I still dont really understand it given the abnormal avrage bitrate that doesnt really have any real world application. Cause again "98Mbps" is not "UHD-BD", the specifications mandates a max bitrate of 100Mbps, and the average bitrate will be based on the size constrains of the physical medium. Close to zero titles will actually have a avrage bitrate that high, so its nothing that represent the uhd-bd format, and a part form that setting 98Mbps togheter with something like 98 for vbv-limitis that you need to have a uhd-bd compliant encode you are pretty much doing CBR-encoding that can have some bitrate distribution sideeffects that could actually hurt quality.
If the interest is more a question of what getting the absolute max out of x265 (i.e. using all the tools in the kitchen sink without caring to much of the practicality of the encoding speed) it makes so much more sense to try to find the limit were it becomes visually transperent instead of just setting a huge abr were most tuning and cpu-consuming tools just becomes irrelevant.
HD MOVIE SOURCE
21st November 2022, 18:17
As its seems like you are ignoring a lot o the answers provided to you, but im gonna give it a lost shot.
First of all I answered your CPU question above, with high complexity 4k material no CPU on the market is likely to break 2fps at veryslow, it can be done with chunk encoding if you go for something like a Threadripper 5975WX or 5995WX, if go above the 24h requirement something like a 7950X will probably get you there without chunkencoding, I will guestimate that you might get close to 2fps for preset veryslow.
Secondly, the gains of presets like veryslow are very much in the diminishing returns territory, and when we are talking about such high bitrates its even beyond overkill imo. As I said, if you have issues with the output at 98Mbps at something like slow or slower, I would look at other parameters then going to preset veryslow. All uhd-bluray encodes ive done has been visually lossless at slow and slower at much lower bitrates.
Its still very unclear what the goal is here, complying to the uhd-bd specifications are only limitations, there is no real reason to use those unless you really are authoring for actual physical discs. If the reason is of more academic purpose, I still dont really understand it given the abnormal avrage bitrate that doesnt really have any real world application. Cause again "98Mbps" is not "UHD-BD", the specifications mandates a max bitrate of 100Mbps, and the average bitrate will be based on the size constrains of the physical medium. Close to zero titles will actually have a avrage bitrate that high, so its nothing that represent the uhd-bd format, and a part form that setting 98Mbps togheter with something like 98 for vbv-limitis that you need to have a uhd-bd compliant encode you are pretty much doing CBR-encoding that can have some bitrate distribution sideeffects that could actually hurt quality.
If the interest is more a question of what getting the absolute max out of x265 (i.e. using all the tools in the kitchen sink without caring to much of the practicality of the encoding speed) it makes so much more sense to try to find the limit were it becomes visually transperent instead of just setting a huge abr were most tuning and cpu-consuming tools just becomes irrelevant.
Thank you for the response, I am reading it, I just wanted to make it clear what I'm doing. I'll have to take a look at thread ripper, I've heard a few people talking about it before. Thank you.
The maxed-out spec for 4K-UHD-BD is 100 Mbps, but after authoring I don't believe anyone can make it work. It doesn't mux properly from what I've heard. However, I believe the absolute bleeding edge maxed is 98.5 Mbps, but I just round down to 98 Mbps.
I've only seen a handful of encodes on 4K-UHD-BD that I believe are absolutely perfect, Django being one of them. That was encoded at 96 Mbps average bit-rate, and the grain never broke up or was ever blurred. It was released in 2021 and the transfer still cannot be beaten today.
With that bitrate, you can target a lot higher performance without visual loss. With --vbv-maxrate 98000, a simple --crf 14 --preset medium should be pretty well transparent.
Making 4K look good with a 10 Mbps peak is a lot harder, which is where extra tools with extra performance implications kick in. Throwing 9.8x more bits at the problem means the encoder starts out with much, much lower QPs, which is the classic brute force way of improving quality.
Yeah, definitely. Making 10 Mbps content look good, simply requires one serious setup to make it work.
I've encoded Sol Levante from Netflix at 98 Mbs, and optimized the encoding settings to finish the encode with an average QP of 7.60. But, as you say, there's transparent, and then there's different levels of transparency.
I have found that the biggest differences that have a meaningful impact of picture quality and detail is...Obviously having enough bits to make the image look good, but no-sao, no-deblock, no-strong-intra-smoothing, subme=7 (more sharpness), aq-mode=3 with aq-strength=1.7 is the best combination for Sol Levante that I've found, and I've tested a lot of combinations.
I've also found that ipratio=1.00 & pbratio=1.00 allow for a much more consistent image from a QP point of view. They are all just about equal.
I've also found that using b-adapt=0 actually leads to better utilization of bits, because my bframes are higher quality that my I and P frames, and having b frames at 75% cheaper bits is simply better than having the system sometimes use bframes, and then at other times not. I personally think maximizing the use of bframes is extremely important for formats like 4K-UHD-BD because they can only use 3 bframes. I have never found that using b-adapt= (any setting above) leads to better results ever.
Obviously digitally made content is easier to get QP levels down lower. Film grain material will probably be around the 20s for QP, but the biggest issue getting film grain consistency to look perfect.
Just some thoughts I have.
benwaggoner
22nd November 2022, 01:12
Are you encoding HDR Sol Levante with --aq-mode 3?
excellentswordfight
26th November 2022, 21:51
The maxed-out spec for 4K-UHD-BD is 100 Mbps, but after authoring I don't believe anyone can make it work. It doesn't mux properly from what I've heard. However, I believe the absolute bleeding edge maxed is 98.5 Mbps, but I just round down to 98 Mbps.
Thats an vbv issue for x265, not a general truth for hevc/uhd-bd, but yes x265 can overshot so its usually best to leave some overhead so setting 98Mbps for vbv if authoring for uhd-bd is fine. I've used other hevc-encoders without this issue.
I've only seen a handful of encodes on 4K-UHD-BD that I believe are absolutely perfect, Django being one of them. That was encoded at 96 Mbps average bit-rate, and the grain never broke up or was ever blurred. It was released in 2021 and the transfer still cannot be beaten today.
Yes, but most movies that get uhd-bd releases are not 90min with mono audio that gets put on 100GB BD-R, so what I dont really understand/the point im trying to make is; if your UHD-BD compatibility requirement is to try to tune x265 for UHD-BD authoring, why not use a more realistic bitrate that is more typical for the format? But if you just wanna burn short test sequences to bluray for playback at the max quality, then I guess its fine to max out bandwidth....
If thats not the case and the goal is quality, why shot yourself in the foot and comply to uhd-bd? Just set level 5.1 high teir and a low crf and tune for that.
You are ofc free to do whatever you want, I'm just trying to understand your use case/rational here, cause I cant really wrap my head around what your trying to achieve.
Obviously having enough bits to make the image look good, but no-sao, no-deblock, no-strong-intra-smoothing, subme=7 (more sharpness), aq-mode=3 with aq-strength=1.7 is the best combination for Sol Levante that I've found, and I've tested a lot of combinations.
Would you mind uploading a sample of your encode?
benwaggoner
29th November 2022, 00:52
In particular, it seems really unlikely that --subme 7 would do anything to improve detail at those bitrates. I've not been able to demonstrate any visible improvement using 7 even at <500 Kbps bitrates. Not even --preset placebo goes that high, because MCW didn't find it offered quality/speed improvements valuable even at placebo. --preset veryslow only uses 4.
And all >5 does is add an extra half and quarter pel iteration, which only would matter if the second iteration found something better than the first did. Higher subme values can shave a little of QPs, but at a bitrate where you're already probably getting mean QP around 10, QP could go up or down by 2-3 without any visible change. Getting an extra 0.1 QP isn't visible even at very low bitrates. You'd need individual frames getting at least a 1 QP difference before seeing much.
HD MOVIE SOURCE
30th November 2022, 07:11
In particular, it seems really unlikely that --subme 7 would do anything to improve detail at those bitrates. I've not been able to demonstrate any visible improvement using 7 even at <500 Kbps bitrates. Not even --preset placebo goes that high, because MCW didn't find it offered quality/speed improvements valuable even at placebo. --preset veryslow only uses 4.
And all >5 does is add an extra half and quarter pel iteration, which only would matter if the second iteration found something better than the first did. Higher subme values can shave a little of QPs, but at a bitrate where you're already probably getting mean QP around 10, QP could go up or down by 2-3 without any visible change. Getting an extra 0.1 QP isn't visible even at very low bitrates. You'd need individual frames getting at least a 1 QP difference before seeing much.
For low bit-rate content, I wouldn't use a high subme setting. More detail requires more bits and using more detail with ultra-low bit-rate isn't a good idea in my opinion.
From my understanding of bit-rate distribution, is, you want low detail so the encoders don't have to try and waste bits on the ultra-high detail that it will be picking out. So, using low subme settings, intra smoothing, sao and deblock should give better results for low bit-rate content.
Picking out blades of grass and fine facial detail wouldn't be distributing bits in the best way when you have a very restricted bit-rate. That's how I see it, sounds right to me, unless I'm missing something.
I use aqmode 3 for Sol yes, I found that it improved microbanding. I didn't see improvements in picture quality with aqmode 1 or 2 personally even though they typically should be used for this type of content. As you said, even though minor QP changes will not be seen with the eye, aqmode 3 was the lowest I could get vs mode 1 and 2.
Asmodian
30th November 2022, 16:32
For low bit-rate content, I wouldn't use a high subme setting. More detail requires more bits and using more detail with ultra-low bit-rate isn't a good idea in my opinion.
From my understanding of bit-rate distribution, is, you want low detail so the encoders don't have to try and waste bits on the ultra-high detail that it will be picking out. So, using low subme settings, intra smoothing, sao and deblock should give better results for low bit-rate content.
This does not make sense. High subme gives the codec better predictions so it takes less data to save that extra detail. You are correct that smoothing the content before sending it to the encoder will be easier to encode well (at least compared to the smoothed source), but once you are in the encoder you always want the best quality options you can afford speed wise.
What happens that is bad when using high subme at low bitrate?
Subme is a pure quality v.s. speed option, using low subme as a way to smooth the video is a terrible idea! All it is doing is using a worse prediction, so the codec has to quantize the differences more to maintain the same bitrate. This raises the QP of all predicted blocks, lowering quality across the board, not only dropping details.
benwaggoner
1st December 2022, 01:52
This does not make sense. High subme gives the codec better predictions so it takes less data to save that extra detail. You are correct that smoothing the content before sending it to the encoder will be easier to encode well (at least compared to the smoothed source), but once you are in the encoder you always want the best quality options you can afford speed wise.
What happens that is bad when using high subme at low bitrate?
Subme is a pure quality v.s. speed option, using low subme as a way to smooth the video is a terrible idea! All it is doing is using a worse prediction, so the codec has to quantize the differences more to maintain the same bitrate. This raises the QP of all predicted blocks, lowering quality across the board, not only dropping details.
++ on this. I can't see any theoretical upside to using less --subme at lower bitrates. I use higher subme at lower resolutions and bitrates, since even slight improvements can make a difference, and the extra compute overhead is proportional to frame area.
There are some exceptions to what seem to be pure speed/quality tradeoffs. For example, grainy content at 4K can look at lot better with --rd 4 than --rd 6 for reasons I've not yet root caused to my satisfaction. It seems to be related to increased QP fluctuation between adjacent CTUs. And I speculate that it's a bigger issue in 4K + grain as the single-to-noise ratio can be a lot worse than lower resolutions, as the grain may be the only highly detailed element in a GOP. 80's/90's Super 35 titles like Ghostbusters or Men in Black have maybe 720p of actual detail with a 4K overly of temporally and spatially random noise.
Downscaling is a powerful low-pass filter, so the same source encoded at 1080p has the same amount of signal but a quarter the noise.
HD MOVIE SOURCE
4th December 2022, 06:42
++ on this. I can't see any theoretical upside to using less --subme at lower bitrates. I use higher subme at lower resolutions and bitrates, since even slight improvements can make a difference, and the extra compute overhead is proportional to frame area.
There are some exceptions to what seem to be pure speed/quality tradeoffs. For example, grainy content at 4K can look at lot better with --rd 4 than --rd 6 for reasons I've not yet root caused to my satisfaction. It seems to be related to increased QP fluctuation between adjacent CTUs. And I speculate that it's a bigger issue in 4K + grain as the single-to-noise ratio can be a lot worse than lower resolutions, as the grain may be the only highly detailed element in a GOP. 80's/90's Super 35 titles like Ghostbusters or Men in Black have maybe 720p of actual detail with a 4K overly of temporally and spatially random noise.
Downscaling is a powerful low-pass filter, so the same source encoded at 1080p has the same amount of signal but a quarter the noise.
Okay, that makes sense, if it's reducing bit-rates also then that's great. I didn't realize it could actually lower data rates also and increase detail.
HD MOVIE SOURCE
4th December 2022, 06:46
Thats an vbv issue for x265, not a general truth for hevc/uhd-bd, but yes x265 can overshot so its usually best to leave some overhead so setting 98Mbps for vbv if authoring for uhd-bd is fine. I've used other hevc-encoders without this issue.
Yes, but most movies that get uhd-bd releases are not 90min with mono audio that gets put on 100GB BD-R, so what I dont really understand/the point im trying to make is; if your UHD-BD compatibility requirement is to try to tune x265 for UHD-BD authoring, why not use a more realistic bitrate that is more typical for the format? But if you just wanna burn short test sequences to bluray for playback at the max quality, then I guess its fine to max out bandwidth....
If thats not the case and the goal is quality, why shot yourself in the foot and comply to uhd-bd? Just set level 5.1 high teir and a low crf and tune for that.
You are ofc free to do whatever you want, I'm just trying to understand your use case/rational here, cause I cant really wrap my head around what your trying to achieve.
Would you mind uploading a sample of your encode?
This is the link to my encode of Sol Levante from Netflix it's open source, (just letting others know). https://drive.google.com/file/d/1ySVk9evxhLM1q6PXMCvfpEPgNteteAXa/view?usp=sharing
asarian
4th December 2022, 08:26
To answer the question for me, definitely the i9 12900K. It does the encoding near twices as fast as my old i9 10700K. But, it's imperative you use P-Cores only! If not, the E-Cores will become all saturated, leaving all P-Cores near idle (as they're all waiting for the E-Cores to finish their chinks). See:
CPU loads (https://1drv.ms/u/s!AhSxhQ9g_mrM2XGDgkaCApXbEi3l?e=gDXOp5)
excellentswordfight
5th December 2022, 21:04
This is the link to my encode of Sol Levante from Netflix it's open source, (just letting others know). https://drive.google.com/file/d/1ySVk9evxhLM1q6PXMCvfpEPgNteteAXa/view?usp=sharing
Thank you.
I had a look at that encode, and it displayed what I suspected, a lot of your unorthodox setting really doesnt do you any good.
I compared it to a uhd-bluray encode I did a year ago at about 50Mbps avg bitrate with pretty much just preset slower and some minor teaks.
https://screenshotcomparison.com/comparison/30065
(mouse on for my enocode, no tone mapping applied and cropped to 1080p).
This is ofc just a screenshot, but I couldnt find any improvments with your encode at higher a bitrate. Most of the more easily compressed sections looks fine (as expected at this bitrate), but its quite obvious that your are doing stuff that hurt compression in the though scenes. How exactly did you come to the conclusion to disable strong-intra-smoothing, sao & deblock for this source? Even keeping SAO fully enabled is great for this source (i've played with selective-sao for it, but its a bit of a tradeoff). I have all those turned on, and still my version is sharper...
HD MOVIE SOURCE
6th December 2022, 06:08
Thank you.
I had a look at that encode, and it displayed what I suspected, a lot of your unorthodox setting really doesnt do you any good.
I compared it to a uhd-bluray encode I did a year ago at about 50Mbps avg bitrate with pretty much just preset slower and some minor teaks.
https://screenshotcomparison.com/comparison/30065
(mouse on for my enocode, no tone mapping applied and cropped to 1080p).
This is ofc just a screenshot, but I couldnt find any improvments with your encode at higher a bitrate. Most of the more easily compressed sections looks fine (as expected at this bitrate), but its quite obvious that your are doing stuff that hurt compression in the though scenes. How exactly did you come to the conclusion to disable strong-intra-smoothing, sao & deblock for this source? Even keeping SAO fully enabled is great for this source (i've played with selective-sao for it, but its a bit of a tradeoff). I have all those turned on, and still my version is sharper...
Do you have a link to your encode? soa makes the image look too soft for my taste. There's never a point where Id use it for any encode. This may just be personal taste but I really dislike soa and what it does to the image. I don't think at any point in the 4 mins think that adding deblock would be a good idea. If you can show me a point where you see breakup Id like the timestamps. Strong intra-smoothing is something I'm still playing with on digital sources, and I agree it probably does need to be on for this type of content.
I'm happy with the way it looks, but if there are any errors with it I'd like to take a look and see what settings would improve it.
HD MOVIE SOURCE
6th December 2022, 17:00
++ on this. I can't see any theoretical upside to using less --subme at lower bitrates. I use higher subme at lower resolutions and bitrates, since even slight improvements can make a difference, and the extra compute overhead is proportional to frame area.
There are some exceptions to what seem to be pure speed/quality tradeoffs. For example, grainy content at 4K can look at lot better with --rd 4 than --rd 6 for reasons I've not yet root caused to my satisfaction. It seems to be related to increased QP fluctuation between adjacent CTUs. And I speculate that it's a bigger issue in 4K + grain as the single-to-noise ratio can be a lot worse than lower resolutions, as the grain may be the only highly detailed element in a GOP. 80's/90's Super 35 titles like Ghostbusters or Men in Black have maybe 720p of actual detail with a 4K overly of temporally and spatially random noise.
Downscaling is a powerful low-pass filter, so the same source encoded at 1080p has the same amount of signal but a quarter the noise.
You mentioned --rd 4 / --rd 6 and how it affects film grain. From my understanding the rd setting is trying to lower the bit-rate and to get the perceived quality to be exactly the same, is that correct?
If this is the case, would it be true that rd=6 is trying to lower the bit-rate so much it's actually causing it to miss some film grain and cause inconsistencies in the film grain? So for example rd=1 isn't really trying to lower the bit-rates so it's actually retaining much of the image. But, I'm quite new to encoding, and I'm just trying to think why this is happening.
Another similar thought is that the algorithm used by x265 is simply not good enough to attempt to lower the bit-rates on film grain content without loss.
Would --ipratio=1.00 --pbratio=1.00 -- psy-rd=0.00 -- psy-rdoq=0.00 -- rdoq-level=0 help?
The thought process here to use equal QP, and to have zero influence from psy and rdoq, which is typically used on film grain content to increase detail, (from my understanding). I wonder if any of the psy settings are having an effect with the rd setting?
If people are worried about film grain detail, there are other settings that can increase film grain detail without going into psy and rdoq.
benwaggoner
7th December 2022, 20:07
You mentioned --rd 4 / --rd 6 and how it affects film grain. From my understanding the rd setting is trying to lower the bit-rate and to get the perceived quality to be exactly the same, is that correct?
If this is the case, would it be true that rd=6 is trying to lower the bit-rate so much it's actually causing it to miss some film grain and cause inconsistencies in the film grain? So for example rd=1 isn't really trying to lower the bit-rates so it's actually retaining much of the image. But, I'm quite new to encoding, and I'm just trying to think why this is happening.
Another similar thought is that the algorithm used by x265 is simply not good enough to attempt to lower the bit-rates on film grain content without loss.
Good questions. After some rumination, I have a new working theory. The big problem with noisy --rd 6 is increasingly visible differences between CUs in QP and shape. It may not be --rd 6 itself that is doing it, but features of other modes that kick in at --rd >4.
--amp, --no-amp
Enable analysis of asymmetric motion partitions (75/25 splits, four directions). At RD levels 0 through 4, AMP partitions are only considered at TU sizes 32x32 and below. At RD levels 5 and 6, it will only consider AMP partitions as merge candidates (no motion search) at 64x64, and as merge or inter candidates below 64x64.
[QUOTE]Would --ipratio=1.00 --pbratio=1.00 -- psy-rd=0.00 -- psy-rdoq=0.00 -- rdoq-level=0 help?
The thought process here to use equal QP, and to have zero influence from psy and rdoq, which is typically used on film grain content to increase detail, (from my understanding). I wonder if any of the psy settings are having an effect with the rd setting?
If people are worried about film grain detail, there are other settings that can increase film grain detail without going into psy and rdoq.
Those parameters could well help when making a mezzanine encode, where inter- and intra-frame quality variations should be eliminated as much as possible. They would be counterproductive for any kind of rate-limited encode distribution encode, as the required bitrate for a given level of psychovisual quality goes up a lot without psychovisual optimizations. Having lower QPs for frames higher in the reference hierarchy also pays off in bang-for-the-bit as the better quality in referenced frames gets cascaded down to frames that reference those frames. Making an IDR better typically makes the whole GOP look better. With flat per-frame QP, at a given bitrate IDR, I, and P QPs would go up. The lower QPs for B and B frames generally can't make back up the quality lost in the reference frames. After all, if there's an average B frame sequence length of say six frames, that means 1 of 7 frames is IDR, I, or P, and 6/7th of frames are B or b. Lowering QP on 1/7th frames is a lot more bit-efficiency than lowering them for 6/7th frames.
(Yes, lowering QP of an IDR or I takes several times more bits than lowering QP of a P frame, but there's generally only one of those frames per GOP, with all frames of a GOP benefiting from a better IDR)
(And yes, the biggest IDR benefits apply to variable GOP where the whole GOP is a single continuous shot.)
HD MOVIE SOURCE
9th December 2022, 18:23
Good questions. After some rumination, I have a new working theory. The big problem with noisy --rd 6 is increasingly visible differences between CUs in QP and shape. It may not be --rd 6 itself that is doing it, but features of other modes that kick in at --rd >4.
[QUOTE]--amp, --no-amp
Enable analysis of asymmetric motion partitions (75/25 splits, four directions). At RD levels 0 through 4, AMP partitions are only considered at TU sizes 32x32 and below. At RD levels 5 and 6, it will only consider AMP partitions as merge candidates (no motion search) at 64x64, and as merge or inter candidates below 64x64.
Those parameters could well help when making a mezzanine encode, where inter- and intra-frame quality variations should be eliminated as much as possible. They would be counterproductive for any kind of rate-limited encode distribution encode, as the required bitrate for a given level of psychovisual quality goes up a lot without psychovisual optimizations. Having lower QPs for frames higher in the reference hierarchy also pays off in bang-for-the-bit as the better quality in referenced frames gets cascaded down to frames that reference those frames. Making an IDR better typically makes the whole GOP look better. With flat per-frame QP, at a given bitrate IDR, I, and P QPs would go up. The lower QPs for B and B frames generally can't make back up the quality lost in the reference frames. After all, if there's an average B frame sequence length of say six frames, that means 1 of 7 frames is IDR, I, or P, and 6/7th of frames are B or b. Lowering QP on 1/7th frames is a lot more bit-efficiency than lowering them for 6/7th frames.
(Yes, lowering QP of an IDR or I takes several times more bits than lowering QP of a P frame, but there's generally only one of those frames per GOP, with all frames of a GOP benefiting from a better IDR)
(And yes, the biggest IDR benefits apply to variable GOP where the whole GOP is a single continuous shot.)
Thanks for the response, some really interesting ideas. Referring to the consistent I, P and B frames...Would inconsistencies in those frames from a QP standpoint be one of the reasons why the film grain can look inconsistent? My real question is, if one frame is say QP 25, and a bframe is QP 13 is that alone why odd inconsistencies can be seen?
I think the biggest thing with film grain content is that it's absolutely consistent from a pixel point of view. What I mean is, take an old movie like Django for instance. It's a constant barrage of film grain all the time. There's things moving all of the time. It never ends. When the encoder sees grain it's almost like movement right? Or it's actually worse than movement? At least with movement, if something pans from left to right the encoder can just move the background from one side to the other side and save a good amount of bits on prediction. The issue with film grain is it's completely unpredictable from teh encoder. If we look back from afar at a screen with film grain on it, it looks generally consistent. But to the encoder, every single pixel is always changing.
Having inconsistent QP values is like dialing into a radio station but getting static noise. The I frame is super loud, the P frame is less loud and the B frame is quiet. This is just how I'm picturing this in my head as a way to understand what's happening to the frames. quiet frames would be the b frames and low QPs, and the noisier scenes are louder at higher QPs.
Just another thought, I wonder if certain predictions like amp or rect, rd, and there could be others...that actually hinder film grain or have limited effect. I'm not saying that they are specifically, but as you said, maybe some type of combination could be affecting it.
Another thought and I say this because I always use CRF. My thought is that CRF could actually help with film grain because it always aims for a certain level of quality rather than bit-rate. Now, I know you're going to say, well, you can't control bit-rate with CRF, but you actually can. You can still use bv-maxrate=**** and vbv-bufsize=**** as I do with my encodes. Even if you set the CRF=0 you can still use bv-maxrate=12000 vbv-bufsize=15000 or so to target 12 Mbps.
The biggest issue with this and even 2 pass encoding is available bit-rate. Once you start encoding at really high bit-rates like 90 Mbps film grain really balances out, especially with equal QP. Unfortunately for the streaming world and even people trying to lower the bit-rates of their own movies, it's just not viable to run bit-rates anywhere near that high.
I'd like to try and encode some film grain content. I'd like some that isn't Tears of Steel. TOS has a digital-looking grain that's been applied and is very inconsistent. Even the DCP is blurry at times. You can watch the credits at the end, and the grain looks like blurry noise.
Do you know of any other open-source content that was either shot on film or has film grain that's pretty consistent looking? I'd like to play around with some settings without going heavy into the psy settings, but taking a look at some other settings and just seeing if they have a visible effect on film grain.
Thanks, I really enjoy this conversation, because making film grain look good and consistent with breaking up is a fun challenge, especially with x265, which many say is far harder to make film grain look good vs x264. I wonder, have you run any tests with x265 and found settings that exactly emulate how an image looks vs x264? I wonder if anything could be learned from the way that encoder does things, what settings are good, and applying those ideas to x265?
Thanks.
benwaggoner
9th December 2022, 22:42
Thanks for the response, some really interesting ideas. Referring to the consistent I, P and B frames...Would inconsistencies in those frames from a QP standpoint be one of the reasons why the film grain can look inconsistent? My real question is, if one frame is say QP 25, and a bframe is QP 13 is that alone why odd inconsistencies can be seen?
Yes, the noisier the content, the lower --ipratio and --pbratio need to be.
I think the biggest thing with film grain content is that it's absolutely consistent from a pixel point of view. What I mean is, take an old movie like Django for instance. It's a constant barrage of film grain all the time. There's things moving all of the time. It never ends. When the encoder sees grain it's almost like movement right? Or it's actually worse than movement? At least with movement, if something pans from left to right the encoder can just move the background from one side to the other side and save a good amount of bits on prediction. The issue with film grain is it's completely unpredictable from teh encoder. If we look back from afar at a screen with film grain on it, it looks generally consistent. But to the encoder, every single pixel is always changing.
Yeah, grain is nearly the same as random noise. Grain is sharp-edged and both spatially and temporally random. So, high spatial and temporal frequencies, both of which are hard to encode, without any ability to leverage interframe prediction. And there can be lots of false motion matches, where an encoder finds a potential match between two patches of random grain which gets a motion vector. That can cause the "grain swirlies" where there is apparent motion when there shouldn't be. Modern encoders do a lot of stuff to prevent that as much as feasible.
Having inconsistent QP values is like dialing into a radio station but getting static noise. The I frame is super loud, the P frame is less loud and the B frame is quiet. This is just how I'm picturing this in my head as a way to understand what's happening to the frames. quiet frames would be the b frames and low QPs, and the noisier scenes are louder at higher QPs.
Yep, that's exactly what the default --ipratio and --pbratio can yield with very noisy content. Higher --psy-rd and --psy-rdoq values can also help, as does lowering --aq-strength.
Just another thought, I wonder if certain predictions like amp or rect, rd, and there could be others...that actually hinder film grain or have limited effect. I'm not saying that they are specifically, but as you said, maybe some type of combination could be affecting it.
I doubt amp or rect will do much in grainy regions, as the energy is mostly noise versus the underlying structure that amp and rect can help with. As the grain complexity is consistent across all of the area of all frames, anything that causes adjoining TUs to have visibly different QP can cause some very annoying visual distortions. I suspect --rd 6 is doing adaptive CU/TU sizing and adaptive quantization more aggressively which is in general superior but not when the grain single/noise ratio is too high.
Another thought and I say this because I always use CRF. My thought is that CRF could actually help with film grain because it always aims for a certain level of quality rather than bit-rate. Now, I know you're going to say, well, you can't control bit-rate with CRF, but you actually can. You can still use bv-maxrate=**** and vbv-bufsize=**** as I do with my encodes. Even if you set the CRF=0 you can still use bv-maxrate=12000 vbv-bufsize=15000 or so to target 12 Mbps.
If you look at the --csv-log-level 2 of good looking grainy content, you'll see that there's a lot less bitrate variation than with cleaner content. CRF is fine to use, but doesn't offer as much benefit as with other content, and one can wind up spending a lot of time at the VBV cap where CRF doesn't really do anything.
The biggest issue with this and even 2 pass encoding is available bit-rate. Once you start encoding at really high bit-rates like 90 Mbps film grain really balances out, especially with equal QP. Unfortunately for the streaming world and even people trying to lower the bit-rates of their own movies, it's just not viable to run bit-rates anywhere near that high.
Yep. If you take a clean source and then add heavy grain to it, the required bitrate for transparent quality can more than double. This is why the Film Grain Synthesis approach introduced by AV1 is so promising. Instead of trying to replicate the random noise, it gets parameterized and removed, clean frames encoded, and then new random grain of a similar texture gets added after decode. This requires solving several hard problems simultaneous, and I'e not seen any implementations reliable enough to leave on by default. And some real-world AV1 players have defects in their implementation which can yield inconsistent final experiences. But it is fundamentally a sound approach I anticipate will be an assumed part of encoders and decoders 10 years from now.
The FGS decoder module is out of loop of the AV1 decoder, and it appears lots of hardware will make the AV1 FGS it available for other codecs, like VVC. And I think there are some that can do it for HEVC using SEI metadata.
I'd like to try and encode some film grain content. I'd like some that isn't Tears of Steel. TOS has a digital-looking grain that's been applied and is very inconsistent. Even the DCP is blurry at times. You can watch the credits at the end, and the grain looks like blurry noise.
Do you know of any other open-source content that was either shot on film or has film grain that's pretty consistent looking? I'd like to play around with some settings without going heavy into the psy settings, but taking a look at some other settings and just seeing if they have a visible effect on film grain.
Yeah, ToS has good digital grain for its age, but it's not the same as real film grain. Nor is its grain particularly heavy.
I can't think of any broadly available very grainy sources, alas. The lack of good grainy sources in standard test libraries is probably one reason issues in grain handling aren't more directly addressed in reference encoders.
Thanks, I really enjoy this conversation, because making film grain look good and consistent with breaking up is a fun challenge, especially with x265, which many say is far harder to make film grain look good vs x264. I wonder, have you run any tests with x265 and found settings that exactly emulate how an image looks vs x264? I wonder if anything could be learned from the way that encoder does things, what settings are good, and applying those ideas to x265?
Yeah, it's some good nerdy fun. Less fun when trying to deliver the highest quality at the lowest bitrate to millions of customers, but it's certainly a bracing challenge, and a fabulous feeling when executed well.
excellentswordfight
10th December 2022, 14:53
Thanks, I really enjoy this conversation, because making film grain look good and consistent with breaking up is a fun challenge, especially with x265, which many say is far harder to make film grain look good vs x264. I wonder, have you run any tests with x265 and found settings that exactly emulate how an image looks vs x264? I wonder if anything could be learned from the way that encoder does things, what settings are good, and applying those ideas to x265?
These are my standard settings for 1080p24 SDR standard film encodes, and imo grain characteristic are absolutely similar to normal x264 settings (--preset slower --tune film)
--preset slow --profile main10 --level-idc 41 --keyint 240 --min-keyint 24 --rc-lookahead 48 --aq-mode 1 --no-sao --deblock -1:-1 --psy-rd 2.5 --psy-rdoq 4 --colorprim bt709 --transfer bt709 --colormatrix bt709
https://screenshotcomparison.com/comparison/30156 (mouse on for x265)
Encoded with crf17 with x264, and crf for x265 is set to match average bitrate.
so --aq-mode 1 --no-sao --deblock -1:-1 and a slight psy tweak and you have something similar to x264 with tune film.
HD MOVIE SOURCE
10th December 2022, 23:08
Yes, the noisier the content, the lower --ipratio and --pbratio need to be.
Yeah, grain is nearly the same as random noise. Grain is sharp-edged and both spatially and temporally random. So, high spatial and temporal frequencies, both of which are hard to encode, without any ability to leverage interframe prediction. And there can be lots of false motion matches, where an encoder finds a potential match between two patches of random grain which gets a motion vector. That can cause the "grain swirlies" where there is apparent motion when there shouldn't be. Modern encoders do a lot of stuff to prevent that as much as feasible.
Yep, that's exactly what the default --ipratio and --pbratio can yield with very noisy content. Higher --psy-rd and --psy-rdoq values can also help, as does lowering --aq-strength.
I doubt amp or rect will do much in grainy regions, as the energy is mostly noise versus the underlying structure that amp and rect can help with. As the grain complexity is consistent across all of the area of all frames, anything that causes adjoining TUs to have visibly different QP can cause some very annoying visual distortions. I suspect --rd 6 is doing adaptive CU/TU sizing and adaptive quantization more aggressively which is in general superior but not when the grain single/noise ratio is too high.
If you look at the --csv-log-level 2 of good looking grainy content, you'll see that there's a lot less bitrate variation than with cleaner content. CRF is fine to use, but doesn't offer as much benefit as with other content, and one can wind up spending a lot of time at the VBV cap where CRF doesn't really do anything.
Yep. If you take a clean source and then add heavy grain to it, the required bitrate for transparent quality can more than double. This is why the Film Grain Synthesis approach introduced by AV1 is so promising. Instead of trying to replicate the random noise, it gets parameterized and removed, clean frames encoded, and then new random grain of a similar texture gets added after decode. This requires solving several hard problems simultaneous, and I'e not seen any implementations reliable enough to leave on by default. And some real-world AV1 players have defects in their implementation which can yield inconsistent final experiences. But it is fundamentally a sound approach I anticipate will be an assumed part of encoders and decoders 10 years from now.
The FGS decoder module is out of loop of the AV1 decoder, and it appears lots of hardware will make the AV1 FGS it available for other codecs, like VVC. And I think there are some that can do it for HEVC using SEI metadata.
Yeah, ToS has good digital grain for its age, but it's not the same as real film grain. Nor is its grain particularly heavy.
I can't think of any broadly available very grainy sources, alas. The lack of good grainy sources in standard test libraries is probably one reason issues in grain handling aren't more directly addressed in reference encoders.
Yeah, it's some good nerdy fun. Less fun when trying to deliver the highest quality at the lowest bitrate to millions of customers, but it's certainly a bracing challenge, and a fabulous feeling when executed well.
I encoded just 30 seconds of Django 4K, one at rd=6 and one at rd=4 and yeah...You're right, rd=4 is better than rd=6. That's so weird. It's not by much, but the fact that it's better is certainly odd. I wonder if this is just a bug in the x265 software? Because typically, asking for higher quality should yield higher quality, but the quality goes down?
I running some other tests like rd=3 just to see if rd=4 is the limit of quality when it comes to film grain. Then I might see if there's any other odd interactions with some other settings later.
I wonder if this is something that the software guys could actually update in a future build?
benwaggoner
12th December 2022, 21:55
I encoded just 30 seconds of Django 4K, one at rd=6 and one at rd=4 and yeah...You're right, rd=4 is better than rd=6. That's so weird. It's not by much, but the fact that it's better is certainly odd. I wonder if this is just a bug in the x265 software? Because typically, asking for higher quality should yield higher quality, but the quality goes down?
Yeah, it's non-intuitive for sure! I presume that some of the extra stuff in other tools that get activated by rd=6 are causing some suboptimal behaviors. I'm also sure that some of the things in rd=6 are beneficial too, so it should be better than 4 if the pathological behaviors could be disabled somehow.
I running some other tests like rd=3 just to see if rd=4 is the limit of quality when it comes to film grain. Then I might see if there's any other odd interactions with some other settings later.
Please do report back!
I wonder if this is something that the software guys could actually update in a future build?
It almost certainly can be fixed. If the tools that are causing problems in rd=6. can be identified, hopefully they can be tweaked to not respond badly to grain. Or at least just deactivated with heavy grain. --tune grain is ancient and terrible, and a refactor that worked well with the modern x265 infrastructure would be welcome, even if it had to be activated manually.
Since the problems are by far the worst at 4K, I suspect there's some threshold of noise detail versus content detail where things go wonky if crossed.
excellentswordfight
13th December 2022, 16:58
Finally got a Epyc milan to play with, and its a nice upgrade over Rome.
Amd Epyc "Milan" 7543P - 6,4fps
Amd EPyc "Rome" 7402P - 4,86fps
Intel Xeon "Cascade Lake Refresh" 2x6226R - 3,97fps
All 1u Dell servers in performance mode, doing a 2160p HDR10 transcode of STEM2 with preset slow, CPU utilization is about 50-75%, all system has 64 threads and 8x16GB for Epyc, and 12x8GB RAM for Intel.
HD MOVIE SOURCE
17th December 2022, 14:58
Yeah, it's non-intuitive for sure! I presume that some of the extra stuff in other tools that get activated by rd=6 are causing some suboptimal behaviors. I'm also sure that some of the things in rd=6 are beneficial too, so it should be better than 4 if the pathological behaviors could be disabled somehow.
Please do report back!
It almost certainly can be fixed. If the tools that are causing problems in rd=6. can be identified, hopefully they can be tweaked to not respond badly to grain. Or at least just deactivated with heavy grain. --tune grain is ancient and terrible, and a refactor that worked well with the modern x265 infrastructure would be welcome, even if it had to be activated manually.
Since the problems are by far the worst at 4K, I suspect there's some threshold of noise detail versus content detail where things go wonky if crossed.
Tested some rd=3 vs rd=4 on film grain material, and rd3 and rd4 gave exactly the same results right down to the QP values. I then realized my testing method is definitely flawed. I tested on Medium, and many of the good motion and quality settings don't kick in until you passed medium. This is probably why I got exactly the same numbers.
Now, as soon as I went to rd=2 then the quality dropped.
So, once I get a chance I'm going to retest on very slow, then just switch the rd setting from 6 to 4 and see what happens. That way most settings are turned on regardless of rd setting, and that way we can see how rd scales properly.
However, when encoding Big Buck Bunny even on Medium, resulted in better quality using rd=6 vs rd=4. Which I found interesting.
Boulder
17th December 2022, 16:55
Rd 3 and 4 are the exact same, as are 5 and 6. This is stated in the docs.
benwaggoner
19th December 2022, 02:15
Yeah. They left some extra room in case they'd want to add intermediate steps in the future.
I'd love to see an --rd 5 which would be "all the stuff from --rd 6 that doesn't mess up with high resolution film grain."
A1
22nd December 2022, 08:35
Good questions. After some rumination, I have a new working theory. The big problem with noisy --rd 6 is increasingly visible differences between CUs in QP and shape. It may not be --rd 6 itself that is doing it, but features of other modes that kick in at --rd >4.
--amp, --no-amp
Enable analysis of asymmetric motion partitions (75/25 splits, four directions). At RD levels 0 through 4, AMP partitions are only considered at TU sizes 32x32 and below. At RD levels 5 and 6, it will only consider AMP partitions as merge candidates (no motion search) at 64x64, and as merge or inter candidates below 64x64.
According to your inference, as long as the maximum size of ctu is limited, the quality of noise and other high-frequency details can be guaranteed when rd6 is turned on.
DMD
9th January 2023, 22:01
Good morning.
To understand and have a reference with the hardware, I did a test with sample video @ 23.976fps with the current PC (Intel Core i5 9600K \ GeForce RTX 2070).
I set the StaxRip preset to Slower, the process is 0.1 FPS.
MISSION IMPOSSIBLE!!! :(
https://i.postimg.cc/tRVz0PsG/Screenshot-19-01-2022-08-34-42.png
Making a prediction of around 2 FPS (I hope) with a Ryzen 9 5900X, I made this reasoning, if that's correct.
In the duration of a second of movie there are about 24fps and the processor processes 2 FPS, it means that the ratio is 1:12, so a movie with a duration of 2 hours needs 24 hours of processing, is that correct?
If necessary, it is possible to divide the source file into several parts, in order to do the processing at different times and then carry out the final union?
Thank you
Good evening.
This is the first test with Ryzen 9 7950X preset "Slow", I hope that with this upgrade the coding time will be more acceptable.
https://i.postimg.cc/9QBJ0Ldn/Cattura.png (https://postimages.org/)
benwaggoner
9th January 2023, 22:11
Those are apples and kumquats comparisons; different thread counts, presets (slower v. slow), etc.
DMD
10th January 2023, 09:09
Yes. You are right, I did not use the same preset.
With "Slower" it is 0.88 fps with the same processor.
At this point I think using the "Slow" preset is a good compromise between performance and final quality, I don't nthink "Slower" significantly affects the quality.
If I haven't miscalculated, an estimated process difference of about 37h40min between the two presets, is that much time difference worth it?
Example: 2 hours of movie @24fps
0.88fps>55h
2.75fps>17h 45min
https://i.postimg.cc/wvcqKny0/Cattura-slower.png
benwaggoner
10th January 2023, 16:43
Yes. You are right, I did not use the same preset.
With "Slower" it is 0.88 fps with the same processor.
At this point I think using the "Slow" preset is a good compromise between performance and final quality, I don't nthink "Slower" significantly affects the quality.
If I haven't miscalculated, an estimated process difference of about 37h40min between the two presets, is that much time difference worth it?
Example: 2 hours of movie @24fps
0.88fps>55h
2.75fps>17h 45min
Slower is the fastest preset where some of of HEVC's more modern features kick in, like relatively deep TU recursion, weighted b-frame prediction, and B-intra encoding. It's the setting I start with by default, and iterate from. It has somewhat reduced parallelism (lookahead-slices 1 instead of 4), so might not be as optimal for benchmarking with many cores available.
Apples-to-apples comparisons can't rely on just presets, however. The number of frame threads can have a big impact on perf and a smaller impact on quality, and the default number of frame threads is based on how many cores are available. Thus comparing two processors with different core counts can see the processor with more cores running with more frame threads, improving encoding speed but potentially reducing quality. So not quite apples-to-apples.
Benchmarking is hard to do in a broadly applicable way, because there are so many encoding scenarios that can impact relative performance. Comparing at slow with default frame threads is certainly a scenario that will matter to plenty of people. For me, comparing with --preset slower --frame-threads 1 would have the most relevance. Benchmarking for realtime encoding would be very different, as predictable worst-case encoding time becomes essential. Plenty of benchmarks just compare with stock default settings.
I see you are comparing with --pmode (makes good sense if you have a lot of cores relative to frame size, but can slow things down if there aren't enough cores) and --pme (which is a net negative unless you have a whole lot of cores encoding sub-HD resolutions).
I consider it "fair" to use --pmode if it is only used when it increases throughput in a given configuration, and turned off when it doesn't. As --pmode doesn't decrease quality (and can theoretically increase it a bit).
The same can apply to using --pme selectively, although the cores needed to make it a net positive are a lot higher. But for 480p with 64 cores or something, it probably would help. I personally rarely test with more than 18/36 available for any given encoder instance. Although with all the ARM patches, Graviton2/3 with 64 cores deserves some benchmarking as well.
DMD
11th January 2023, 12:43
Slower is the fastest preset where some of of HEVC's more modern features kick in, like relatively deep TU recursion, weighted b-frame prediction, and B-intra encoding. .......
Thanks for the explanation, I still have a lot to learn from this technique.
benwaggoner
12th January 2023, 00:37
Thanks for the explanation, I still have a lot to learn from this technique.
Trust me, everyone has a lot to learn when it comes to benchmarking!
First rule is to define as specific a question as possible, and then figure out the optimal benchmarking approach for that.
DTL
8th February 2023, 20:19
Good evening.
This is the first test with Ryzen 9 7950X preset "Slow", I hope that with this upgrade the coding time will be more acceptable.
Screenshot shows it even not uses AVX512 functions of x265.
It looks 'auto' asm is still up to AVX2 and AVX512 still need to be enabled manually.
Also for getting more performance of software_for_task at different architectures it is better to use optimized builds for each architecture (making equal output work). In other case it is mostly testing different architectures for executing some single build of software (may be not optimal to any of tested architectures).
About AVX512 capable chips and x265 usage:
1. It may be better to use build for AVX512 architecture executable. May be better to use Intel C compiler mastered by Intel for 10+ years for AVX512 architecture even for AMD AVX512-compatible chips. Also having all-project (multi-file) interprocedural optimization features.
When compiler build for selected architecture it can use additional features like larger register file for params storage and special instructions. The disadvantage - the executable can only run at the target (or higher compatible) chip and will cause 'illegal instruction' crash at lower architectures. So for users of AVX512 chips it may be better to found or ask developers or self build the AVX512-targeted build for work.
So 'universal' run-about-everywhere like from SSE2 and higher architectures builds of x265 are not optimized for AVX512 architecture by compiler (and may have degraded performance).
2. In the 'old' github sources from 2020 https://github.com/videolan/x265 usage of optional handcrafted AVX512 codepaths is not enabled by default options set. So it must be activated by command line option --asm avx512 (also may be good to found all other SIMD string options somewhere). May be it is the same for newer versions for https://bitbucket.org/multicoreware/x265_git/downloads/?tab=downloads site.
benwaggoner
10th February 2023, 22:47
Screenshot shows it even not uses AVX512 functions of x265.
It looks 'auto' asm is still up to AVX2 and AVX512 still need to be enabled manually.
This is because using AVX512 causes a net performance decrease in almost all x265 use cases.
Also for getting more performance of software_for_task at different architectures it is better to use optimized builds for each architecture (making equal output work). In other case it is mostly testing different architectures for executing some single build of software (may be not optimal to any of tested architectures).
The nice thing about x265 and open source in general is that we can all compile it optimally ourselves. I think it's most useful to compare processors using the compiler and settings most optimal for that processor. Profile-guided optimizations are fair too in my opinion.
About AVX512 capable chips and x265 usage:
1. It may be better to use build for AVX512 architecture executable. May be better to use Intel C compiler mastered by Intel for 10+ years for AVX512 architecture even for AMD AVX512-compatible chips. Also having all-project (multi-file) interprocedural optimization features.
Yep, if testing AVX512 everything that can be optimized for AVX512 should be.
When compiler build for selected architecture it can use additional features like larger register file for params storage and special instructions. The disadvantage - the executable can only run at the target (or higher compatible) chip and will cause 'illegal instruction' crash at lower architectures. So for users of AVX512 chips it may be better to found or ask developers or self build the AVX512-targeted build for work.
So 'universal' run-about-everywhere like from SSE2 and higher architectures builds of x265 are not optimized for AVX512 architecture by compiler (and may have degraded performance).
ASM instructions not compatible with the current hardware don't get used, with everything up to AVX2 being used by default if available. You are 100% right about one-size-fits-all binaries not being optimal for anyone.
2. In the 'old' github sources from 2020 https://github.com/videolan/x265 usage of optional handcrafted AVX512 codepaths is not enabled by default options set. So it must be activated by command line option --asm avx512 (also may be good to found all other SIMD string options somewhere). May be it is the same for newer versions for https://bitbucket.org/multicoreware/x265_git/downloads/?tab=downloads site.
All the other ones are automatic by default. Per https://x265.readthedocs.io/en/master/cli.html#performance-options:
--asm <integer:false:string>, --no-asm
x265 will use all detected CPU SIMD architectures by default. You can disable all assembly by using --no-asm or you can specify a comma separated list of SIMD architectures to use, matching these strings: MMX2, SSE, SSE2, SSE3, SSSE3, SSE4, SSE4.1, SSE4.2, AVX, XOP, FMA4, AVX2, FMA3
Some higher architectures imply lower ones being present, this is handled implicitly.
One may also directly supply the CPU capability bitmap as an integer.
Note that by specifying this option you are overriding x265’s CPU detection and it is possible to do this wrong. You can cause encoder crashes by specifying SIMD architectures which are not supported on your CPU.
Default: auto-detected SIMD architectures
DTL
11th February 2023, 01:28
"This is because using AVX512 causes a net performance decrease in almost all x265 use cases."
I make quick test with my build with VisualStudio 2019 sources from github - at i5-11600 intel chip enabling AVX512 with command line option give about 4.8% benefit with FullHD encoding with --placebo profile over 'auto' SIMD. But it may be rare intel chip without frequency decrease at AVX512 usage. (low core number - only 6 cores).
Emulgator
11th February 2023, 20:16
An i9-11900K sees AVX-512 profit too here, I guess I can bind that to the production node of Rocket Lake.
x265 2.8 introduces --asm avx512 11900K: Speed 0,35fps -> 0,7fps yay !
(Please do not nail these numbers of mine to your door, these were just from remembering progress bars.
These 14nm chips seem to be able to pull through that workload without too much thermal penalty. (as the predecessors did)
Until AVX-512 got to be fixed/tucked/fused away for the following families by Intel themselves, as I heard...
benwaggoner
12th February 2023, 04:28
Exciting news! Too bad the latest Intel consumer chips dropped AVX512. Comparing with CPU-specific and profile-driven optimizations would be neat to see. And now we have quite a bit of ARM SIMD and some good ARM performance CPUs in Macs and Graviton, we can compare architectures too.
ReinerSchweinlin
13th February 2023, 10:34
I wonder how one of the later x64 compatible Xeon PHI would perform with x265. The cores themselfes where weaker atom cores, but if I recall correctly, they had a bunch of AVX512 features already on board and up to 64 cores per "CPU", often hosting 4 CPUs in one server (and 4 x hyperthreading), giving an insane number of 256 Threads per CPU, 1024 for one complete Server... If something like ripbot would distribute the encoding over the cores, cleverly avoiding too much bottlenecking between ram-swapping.. Might be interesting :)
I know that the cores themslelfes are way behing what modern cores could do - but if "the rest" of the core is enough to feed the AVX512 unit efficiently, maybe the number of cores makes up for the lack of frequency and cache ?
benwaggoner
13th February 2023, 21:35
I wonder how one of the later x64 compatible Xeon PHI would perform with x265. The cores themselfes where weaker atom cores, but if I recall correctly, they had a bunch of AVX512 features already on board and up to 64 cores per "CPU", often hosting 4 CPUs in one server (and 4 x hyperthreading), giving an insane number of 256 Threads per CPU, 1024 for one complete Server... If something like ripbot would distribute the encoding over the cores, cleverly avoiding too much bottlenecking between ram-swapping.. Might be interesting :)
I know that the cores themslelfes are way behing what modern cores could do - but if "the rest" of the core is enough to feed the AVX512 unit efficiently, maybe the number of cores makes up for the lack of frequency and cache ?
Modern video coding is a pretty punishing mix of single-threaded arithmetic processing (CABAC), lots of low-latency branchy mode decisions, and SIMD processing. It needs a pretty balanced and beefy CPU to do quickly; this is a big reason why CUDA-style GPU acceleration has largely vanished with modern codecs.
I expect the Atom cores would bottleneck on CABAC pretty hard. Some kind of hybrid NUMA approach with some performance cores and slower-but-SIMD cores could work. The key to a hybrid model like that would be low-latency access to shared L3 cache. That kind of mixed-core architecture is becoming pretty standard.
wyliec2
14th February 2023, 18:26
I consider it "fair" to use --pmode if it is only used when it increases throughput in a given configuration, and turned off when it doesn't. As --pmode doesn't decrease quality (and can theoretically increase it a bit).
I'm new to this forum but a long-time Handbrake user. Most of the HEVC/X265 info appears to be applicable.
I generally use a 5950X (16 core/32 thread) for encoding and find that using H.265 10-bit on 4K sources winds up with 90+% CPU utilization at Slow-Slower-Very Slow presets.
On BD sources, CPU utilization is more in the 40-50% range.
I have tried applying pmode in Handbrake (pmode=1) and found it to only be beneficial with BD encodes at Very Slow preset.
In testing with three movies at RF 19, Very Slow, pmode=1 reduces encode time 15-20% and increases CPU utilization roughly 20%.
The output files with pmode=1 are generally slightly smaller than an encode with same settings and no options.
One movie, somewhat grainy, showed a significant size reduction - 4247 MB with no options and 3474 MB with the pmode=1 option. I repeated this encode and confirmed the results. FWIW - No option took 13:46 (hours:minutes) to encode and pmode=1 took 10:31 to encode.
Based on what I'd read here, I wasn't expecting significant size change from pmode....wondering about any thoughts from experts..??
tormento
16th February 2023, 16:32
It would be nice to see newer AMD AVX2 vs AVX512 performances when encoding.
ReinerSchweinlin
17th February 2023, 16:07
Modern video coding is a pretty punishing mix of single-threaded arithmetic processing (CABAC), lots of low-latency branchy mode decisions, and SIMD processing. It needs a pretty balanced and beefy CPU to do quickly; this is a big reason why CUDA-style GPU acceleration has largely vanished with modern codecs.
I expect the Atom cores would bottleneck on CABAC pretty hard. Some kind of hybrid NUMA approach with some performance cores and slower-but-SIMD cores could work. The key to a hybrid model like that would be low-latency access to shared L3 cache. That kind of mixed-core architecture is becoming pretty standard.
Thanx for your estimation.
The atom cores on these later Xeon PHi indeed are pretty weak in integer and floating point. They do have a 16GB internal Cache though which is accessable by all the cores - maybe that helps.
But I just checked, the widely available nights Landing Chips only offer a very small subset of AVX 512 Instructions, only the latest ones have a few more modern subsets on board. Is there any ressource to x265 where the used AVX512 modi are listed?
benwaggoner
17th February 2023, 19:36
Thanx for your estimation.
The atom cores on these later Xeon PHi indeed are pretty weak in integer and floating point. They do have a 16GB internal Cache though which is accessable by all the cores - maybe that helps.
16 GB? That can fit a lot of frames.
But I just checked, the widely available nights Landing Chips only offer a very small subset of AVX 512 Instructions, only the latest ones have a few more modern subsets on board. Is there any ressource to x265 where the used AVX512 modi are listed?
The source code would be the definitive resource. There may be a higher level doc somewhere, but I couldn't find one with a quick search. But "a very small subset" is likely not compatible.
The lack of strong single-threaded perf would be the big bottleneck anyway.
Although, I just recalled that WPP might allow some WPP parallelization; nominally 1 thread per 64 pixels high, although probably only 2x better given overhead. WPP certainly allows for decoder parallelization. Even still, an Atom core is many times slower slower for CABAC-like operations than a modern Xeon core, so that's already factored into comparisons.
Modern video encoding is stressful in pretty much every way, so Amdahl's Law prevents any big improvement in one area from helping all that much.
As I've mentioned before, some years back Intel discovered that x265 pushed Xeon thermals hotter than Intel's on internal thermal test tool's theoretical worst case.
The flip side of this is that encoding benefits some from most improvements; when a new processor says it's "X%-Y%" faster, encoding is always close to the higher Y% value. We get to spend orders of magnitude more MIPS/pixel today than when I started doing compression.
Circa 1996, it took about 80 minutes to encode 1 minute of 320x240p15 on my then rocket-fast PowerMac 8100/80 workstation. I was able to charge $80/minute for a tape-to-file conversion with a $20/min surcharge for VHS (mainly to encourage the client to find the Beta SP master).
ReinerSchweinlin
18th February 2023, 13:37
16 GB? That can fit a lot of frames.
Yes, 16GB. But its not a "normal" L3 Cache, its referd to "remote L2 Cache". Its bandwith is higher than the 8 Lane DDR4 access, but not as fast as modern L3 Cache. It can be configured to act as a normal transparent Cache (like a L3 Cache), but also accessed with a seperate driver (or in a hybrid mode). too bad there are no motherboards in Europe for these Xeons. I know that its probally not really worth it, but for a small amount of money, I`d satisfy my curiosity and get one :)
The source code would be the definitive resource. There may be a higher level doc somewhere, but I couldn't find one with a quick search. But "a very small subset" is likely not compatible.
Thanx for checking though. I am not deep enough into all this to simply look up the source code and get my answer.
On a side note: When I was tinkering with CPU feature sets yesterday on an 1950x, I found odd performance differences in different runs, depending, turning AVX2 off seemed to speed things up... Seems there is some potential in individually tweaked binary compiles, taylored to a CPU (of course not worth if one wants to distribute it publicly, but tweaking a personal encoding server this way would be fun), so I probably will have to learn to compile stuff like this properly after all...
ok, back to topic...
The lack of strong single-threaded perf would be the big bottleneck anyway.
I think so, too.... These atom cores really are weak... Even a core2duo has more ooompf per core :)
Although, I just recalled that WPP might allow some WPP parallelization; nominally 1 thread per 64 pixels high, although probably only 2x better given overhead. WPP certainly allows for decoder parallelization. Even still, an Atom core is many times slower slower for CABAC-like operations than a modern Xeon core, so that's already factored into comparisons.
I remember quality penalties from too much parallelization - is it worth thinking about it or are we talking a few percent difference in efficiency here?
Modern video encoding is stressful in pretty much every way, so Amdahl's Law prevents any big improvement in one area from helping all that much.
Maybe at this point its worth mentioning that getting a XEON Phi of course is pure for academic research and interest, tinkering with old stuff, etc.... For anyone reading along - simply getting a modern Desktop CPU is a much better idea :)
As I've mentioned before, some years back Intel discovered that x265 pushed Xeon thermals hotter than Intel's on internal thermal test tool's theoretical worst case.
Haha... Whenever something like this happened back in my days at university - people reverted to introducing the "factor correction"... simply mutliply the whole equation by something that sounds reasonable (I hope Intel does better and to be fair - This was one user group of.... not so scientific members...)..
Circa 1996, it took about 80 minutes to encode 1 minute of 320x240p15 on my then rocket-fast PowerMac 8100/80 workstation. I was able to charge $80/minute for a tape-to-file conversion with a $20/min surcharge for VHS (mainly to encourage the client to find the Beta SP master).
AH, I remember these machines, I did some service on them back then... Good times... I still remember some encoding adventures - good old Abit BP6 with two P3 celeron Tualatin CPUs was able to do realtime MPEG2 for SVCD Encoding..
Edit:
Phoronix has some CPU-Infos which might be interesting:
processor : 0
vendor_id : GenuineIntel
cpu family : 6
model : 87
model name : Intel(R) Xeon Phi(TM) CPU 7210 @ 1.30GHz
stepping : 1
microcode : 0x1b0
cpu MHz : 1168.239
cache size : 1024 KB
physical id : 0
siblings : 256
core id : 0
cpu cores : 64
apicid : 0
initial apicid : 0
fpu : yes
fpu_exception : yes
cpuid level : 13
wp : yes
flags : fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush dts acpi mmx fxsr sse sse2 ss ht tm pbe syscall nx pdpe1gb rdtscp lm constant_tsc arch_perfmon pebs bts rep_good nopl xtopology nonstop_tsc cpuid aperfmperf pni pclmulqdq dtes64 monitor ds_cpl est tm2 ssse3 fma cx16 xtpr pdcm sse4_1 sse4_2 x2apic movbe popcnt tsc_deadline_timer aes xsave avx f16c rdrand lahf_lm abm 3dnowprefetch ring3mwait cpuid_fault epb pti fsgsbase tsc_adjust bmi1 avx2 smep bmi2 erms avx512f rdseed adx avx512pf avx512er avx512cd xsaveopt dtherm ida arat pln pts
bugs : cpu_meltdown spectre_v1 spectre_v2 mds msbds_only
bogomips : 2600.01
clflush size : 64
cache_alignment : 64
address sizes : 46 bits physical, 48 bits virtual
power management:
rchitecture: x86_64
CPU op-mode(s): 32-bit, 64-bit
Byte Order: Little Endian
Address sizes: 46 bits physical, 48 bits virtual
CPU(s): 256
On-line CPU(s) list: 0-255
Thread(s) per core: 4
Core(s) per socket: 64
Socket(s): 1
NUMA node(s): 1
Vendor ID: GenuineIntel
CPU family: 6
Model: 87
Model name: Intel(R) Xeon Phi(TM) CPU 7210 @ 1.30GHz
Stepping: 1
CPU MHz: 1192.466
CPU max MHz: 1500.0000
CPU min MHz: 1000.0000
BogoMIPS: 2600.01
L1d cache: 2 MiB
L1i cache: 2 MiB
L2 cache: 32 MiB
NUMA node0 CPU(s): 0-255
Vulnerability Itlb multihit: Not affected
Vulnerability L1tf: Not affected
Vulnerability Mds: Vulnerable: Clear CPU buffers attempted, no microcode; SMT mitigated
Vulnerability Meltdown: Mitigation; PTI
Vulnerability Spec store bypass: Not affected
Vulnerability Spectre v1: Mitigation; usercopy/swapgs barriers and __user pointer sanitization
Vulnerability Spectre v2: Mitigation; Full generic retpoline, STIBP disabled, RSB filling
Vulnerability Tsx async abort: Not affected
Flags: fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush dts acpi mmx fxsr sse sse2 ss ht tm pbe syscall nx pdpe1gb rdtscp lm constant_tsc arch_perfmon pebs bts rep_good nopl xtopology nonstop_tsc cpuid aperfmperf pni pclmulqdq dtes64 monitor ds_cpl est tm2 ssse3 fma cx16 xtpr pdcm sse4_1 sse4_2 x2apic movbe popcnt tsc_deadline_timer aes xsave avx f16c rdrand lahf_lm abm 3dnowprefetch ring3mwait cpuid_fault epb pti fsgsbase tsc_adjust bmi1 avx2 smep bmi2 erms avx512f rdseed adx avx512pf avx512er avx512cd xsaveopt dtherm ida arat pln pts
wyliec2
18th February 2023, 18:32
I consider it "fair" to use --pmode if it is only used when it increases throughput in a given configuration, and turned off when it doesn't. As --pmode doesn't decrease quality (and can theoretically increase it a bit).
I have been experimenting with with pmode on my 5950X platform.
I typically encode in Slow, Slower or Very Slow.
For all 4K encodes, the CPU will run at 90+% utilization and pmode causes encodes to take longer.
For BD encodes at Slow or Slower, the encodes take longer with pmode.
For BD encodes at Very Slow, pmode does reduce encode time - in one example an encode took 13 hours at Very Slow and 10 hours at Very Slow with pmode. It also seems to increase CPU utilization around 20% (from mid 40% to mid 60%).
I've only tested on 3 files and while two showed slightly smaller output file size, one showed a significant output size reduction (4247 MB without pmode and 3474 MB with pmode).
Nothing in documentation or what I've read here, lead me to expect this result....wondering if there are any thoughts/comments on this result..??
DMD
18th February 2023, 23:07
Slower is the fastest preset where some of of HEVC's more modern features kick in, like relatively deep TU recursion, weighted b-frame prediction, and B-intra encoding. It's the setting I start with by default, and iterate from. It has somewhat reduced parallelism (lookahead-slices 1 instead of 4), so might not be as optimal for benchmarking with many cores available.
Apples-to-apples comparisons can't rely on just presets, however. The number of frame threads can have a big impact on perf and a smaller impact on quality, and the default number of frame threads is based on how many cores are available. Thus comparing two processors with different core counts can see the processor with more cores running with more frame threads, improving encoding speed but potentially reducing quality. So not quite apples-to-apples.
Benchmarking is hard to do in a broadly applicable way, because there are so many encoding scenarios that can impact relative performance. Comparing at slow with default frame threads is certainly a scenario that will matter to plenty of people. For me, comparing with --preset slower --frame-threads 1 would have the most relevance. Benchmarking for realtime encoding would be very different, as predictable worst-case encoding time becomes essential. Plenty of benchmarks just compare with stock default settings.
I see you are comparing with --pmode (makes good sense if you have a lot of cores relative to frame size, but can slow things down if there aren't enough cores) and --pme (which is a net negative unless you have a whole lot of cores encoding sub-HD resolutions).
I consider it "fair" to use --pmode if it is only used when it increases throughput in a given configuration, and turned off when it doesn't. As --pmode doesn't decrease quality (and can theoretically increase it a bit).
The same can apply to using --pme selectively, although the cores needed to make it a net positive are a lot higher. But for 480p with 64 cores or something, it probably would help. I personally rarely test with more than 18/36 available for any given encoder instance. Although with all the ARM patches, Graviton2/3 with 64 cores deserves some benchmarking as well.
After your comprehensive answer, I apologize for this inexperienced question.
Taking into account that my CPU (Ryzen 7950) has 16C/32T, to perform x265 encoding of 4K HDR files, I disabled "pmode" should I also disable "pme" from my script to avoid long encoding time or how could I improve my script?
Thank you very much
----------------------------------------------------------------------------------------------------------------------------------------------------------------------------
--crf 16 --preset slower --output-depth 10 --profile main10 --level-idc 5.1 --rd-refine --vbv-bufsize 100000 --vbv-maxrate 100000 --hme-search umh,umh,star --hme --min-keyint 1 --keyint 24 --no-open-gop --pme --master-display "G(8500,39850)B(6550,2300)R(35400,14600)WP(15635,16450)L(10000000,1)" --colorprim bt2020 --colormatrix bt2020nc --transfer smpte2084 --range limited --max-cll "1000,400" --sar 1:1 --no-info --repeat-headers --aud --hrd --uhd-bd
----------------------------------------------------------------------------------------------------------------------------------------------------------------------------
guest
22nd February 2023, 04:17
After your comprehensive answer, I apologize for this inexperienced question.
Taking into account that my CPU (Ryzen 7950) has 16C/32T, to perform x265 encoding of 4K HDR files, I disabled "pmode" should I also disable "pme" from my script to avoid long encoding time or how could I improve my script?
Thank you very much
----------------------------------------------------------------------------------------------------------------------------------------------------------------------------
--crf 16 --preset slower --output-depth 10 --profile main10 --level-idc 5.1 --rd-refine --vbv-bufsize 100000 --vbv-maxrate 100000 --hme-search umh,umh,star --hme --min-keyint 1 --keyint 24 --no-open-gop --pme --master-display "G(8500,39850)B(6550,2300)R(35400,14600)WP(15635,16450)L(10000000,1)" --colorprim bt2020 --colormatrix bt2020nc --transfer smpte2084 --range limited --max-cll "1000,400" --sar 1:1 --no-info --repeat-headers --aud --hrd --uhd-bd
----------------------------------------------------------------------------------------------------------------------------------------------------------------------------
Hello yet again, DMD,
Like I said in the Staxrip thread, I use RipBot264, but the Pauly Dunne builds, which have so much more to offer, than the standard one...but I digress.
Several "power" users have commented how the 16 core Ryzens, "fall off a cliff" when encoding certain X265 video's, but "we" have come up with a "fix" that is part of the encoders command's that really gets them to do the job they're supposed to do, as well as custom x265 command's as well.
I am VERY happy with the way my 3950X, 5950X & the 7950X are performing, as well as the interloper, the 13900KF :)
I must admit that my 5950X was being bested by the 5900X with almost everything, but I changed some basic BIOS setting's and it's working better that ever before :D
DMD
22nd February 2023, 08:29
Hello yet again, DMD,
Like I said in the Staxrip thread, I use RipBot264, but the Pauly Dunne builds, which have so much more to offer, than the standard one...but I digress.
Several "power" users have commented how the 16 core Ryzens, "fall off a cliff" when encoding certain X265 video's, but "we" have come up with a "fix" that is part of the encoders command's that really gets them to do the job they're supposed to do, as well as custom x265 command's as well.
I am VERY happy with the way my 3950X, 5950X & the 7950X are performing, as well as the interloper, the 13900KF :)
I must admit that my 5950X was being bested by the 5900X with almost everything, but I changed some basic BIOS setting's and it's working better that ever before :D
I didn't know that, and I'm very surprised that the Ryzens have this problem with x265 encoding.
But I am also very happy that a solution has been found to make them work at their maximum performance.
I don't know how to apply the "fix" and sari happy to know how to do it.
As for my bios ( ASUS ROG Strix X670E-F Gaming WiFi) I only performed optimization for RAM and fast boot.
Thank you very much
DTL
23rd February 2023, 12:17
Taking into account that my CPU (Ryzen 7950) how could I improve my script?
With AMD 7xxx you may try to force usage of AVX512 with --asm avx512 . If it will not cause overheating clock trottling (https://www.hwcooling.net/en/intel-avx-512-tested-in-x265-how-to-enable-it-and-does-it-help/ ) you may got some performance benefit.
DMD
23rd February 2023, 18:50
With AMD 7xxx you may try to force usage of AVX512 with --asm avx512 . If it will not cause overheating clock trottling (https://www.hwcooling.net/en/intel-avx-512-tested-in-x265-how-to-enable-it-and-does-it-help/ ) you may got some performance benefit.
Many thanks for the suggestion.
ReinerSchweinlin
27th February 2023, 10:53
I didn't know that, and I'm very surprised that the Ryzens have this problem with x265 encoding.
But I am also very happy that a solution has been found to make them work at their maximum performance.
I don't know how to apply the "fix" and sari happy to know how to do it.
As for my bios ( ASUS ROG Strix X670E-F Gaming WiFi) I only performed optimization for RAM and fast boot.
Thank you very much
Since I also have a 1950x lying around, IŽd be happy to read more about that "fix" - could someone point me in the right direction with a link please ?
Boulder
27th February 2023, 11:50
It would be interesting to hear since I have a 5950X and have zero issues with getting the CPU work at 80-100% usage level when encoding with x265.
benwaggoner
4th March 2023, 23:16
After your comprehensive answer, I apologize for this inexperienced question.
Taking into account that my CPU (Ryzen 7950) has 16C/32T, to perform x265 encoding of 4K HDR files, I disabled "pmode" should I also disable "pme" from my script to avoid long encoding time or how could I improve my script?
You definitely want to have --pme off; I've never seen it boost throughput with anything above 480p. You should get a >2x speed improvement turning it off.
--pmode has a much bigger chance to be helpful, I'd say it's likely useful above 20 threads for 4K if using --frame-threads 1. The more modes being evaluated, the more parallelization for --pmode to take advantage of.
Using only a single frame thread can improve quality, but limits parallelization a lot, and combining it with --pmode can get some of that perf back if you have enough cores.
Looking at the reset of your command line:
--rd-refine doesn't do anything in a single pass, which your encode is.
Is --no-open-gop still required for BD compatibility with x265 (they are certainly supported by the BD format itself). With 24 frame GOPs, open GOP can provide some real benefit. If you're stuck with --no-open-gop, you could try --radl 2 to get some of the same benefit.
I don't know that --hme has proven to be that helpful. You should try with it off to see if it provides any benefit with your content.
If there is much grain in the source --rd 4 can both improve quality and throughput.
--crf 16 --preset slower --output-depth 10 --profile main10 --level-idc 5.1 --rd-refine --vbv-bufsize 100000 --vbv-maxrate 100000 --hme-search umh,umh,star --hme --min-keyint 1 --keyint 24 --no-open-gop --pme --master-display "G(8500,39850)B(6550,2300)R(35400,14600)WP(15635,16450)L(10000000,1)" --colorprim bt2020 --colormatrix bt2020nc --transfer smpte2084 --range limited --max-cll "1000,400" --sar 1:1 --no-info --repeat-headers --aud --hrd --uhd-bd
Boulder
5th March 2023, 13:47
--rd-refine doesn't do anything in a single pass, which your encode is.
It does work on CRF encodes, it doesn't need a stats file or anything. I've been trying to figure out what it actually does or what the use case is but I have no clue.
You definitely want to have --pme off; I've never seen it boost throughput with anything above 480p. You should get a >2x speed improvement turning it off.
--pmode has a much bigger chance to be helpful, I'd say it's likely useful above 20 threads for 4K if using --frame-threads 1. The more modes being evaluated, the more parallelization for --pmode to take advantage of.
Using only a single frame thread can improve quality, but limits parallelization a lot, and combining it with --pmode can get some of that perf back if you have enough cores.
Looking at the reset of your command line:
--rd-refine doesn't do anything in a single pass, which your encode is.
Is --no-open-gop still required for BD compatibility with x265 (they are certainly supported by the BD format itself). With 24 frame GOPs, open GOP can provide some real benefit. If you're stuck with --no-open-gop, you could try --radl 2 to get some of the same benefit.
I don't know that --hme has proven to be that helpful. You should try with it off to see if it provides any benefit with your content.
If there is much grain in the source --rd 4 can both improve quality and throughput.
Thank you very much for the advice, I will do some tests for a better result.
Using StaxRip I had a chance to do some tests with "number of parallel process" and "Chuncks", and I noticed that by setting the maximum value (16) for both parallel processes and Chunks, I got higher process speed, but also missing video frames.
In my personal configuration with a setting of 3-3 I was able to get a slight speed increase without any side effects, using the commands I included in the previous post.
benwaggoner
6th March 2023, 06:22
It does work on CRF encodes, it doesn't need a stats file or anything. I've been trying to figure out what it actually does or what the use case is but I have no clue.
You're right; was thinking of a different parameter.
--rd-refine, --no-rd-refine
For each analysed CU, calculate R-D cost on the best partition mode for a range of QP values, to find the optimal rounding effect. Default disabled.
Only effective at RD levels 5 and 6
It should offer a slight overall compression efficiency improvement.
Boulder
18th March 2023, 15:07
--pmode has a much bigger chance to be helpful, I'd say it's likely useful above 20 threads for 4K if using --frame-threads 1. The more modes being evaluated, the more parallelization for --pmode to take advantage of.
Using only a single frame thread can improve quality, but limits parallelization a lot, and combining it with --pmode can get some of that perf back if you have enough cores.
Just for fun, I tested frame-threads from 5 (the default for my CPU) to 1 and then with pmode on my 5950X (16C/32T). The effect of pmode on the compression efficiency is much bigger than I anticipated. The speed increase was weird because my CPU usage is already around 90-100% when encoding with the default frame-threads and no pmode.
I ran this test on a 720p encode, normal setup and settings for my 1080p->720p encodes to the media library. I do use some uncommon parameters like --no-limit-modes and --rskip 0 which probably affect the results compared to standard presets.
I seriously need to test the 4K encodes as well.
F 5 - 5718.31 kbps - 7.11 fps
F 4 - 5713.13 kbps - 6.93 fps
F 3 - 5708.73 kbps - 6.74 fps
F 2 - 5715.94 kbps - 6.23 fps (odd that the size went up..)
F 1 - 5695.93 kbps - 4.50 fps
F 1 + pmode - 5490.78 kbps - 5.88 fps
F 2 + pmode - 5521.68 kbps - 7.43 fps
F 3 + pmode - 5515.12 kbps - 7.72 fps
F 4 + pmode - 5519.36 kbps - 7.83 fps
F 5 + pmode - 5521.10 kbps - 8.01 fps
vBulletin® v3.8.11, Copyright ©2000-2026, vBulletin Solutions Inc.