View Full Version : New advanced Benchmark


Sagittaire
7th March 2017, 13:18
Here the source:
http://jfl1974.free.fr/Benchmark/Benchmark.zip

1) x264 1080p 8 bits bt.709 with BD FHD H264 source
2) x265 2160p 10 bits HDR bt.2020 with ~BD UHD HEVC HDR source
3) decoding benchmark with libavcodec and ~BD UHD HEVC HDR source
4) Advanced x265 test for complete SIMD MMX, SSE, SSE2, SSE3, SSE4, AVX, FMA3, AVX2, FMA4, XOP

Update:
-14/04/2017: new automatic benchmark

try and report your result:

|----------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| CPU (Ghz) | x264 | x265 | LAVC | auto | MMX2 | SSE | SSE2 | SSE3 | SSE4 | AVX | AVX2 | All |
|----------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| i5-3550@3.50 | 7.12 | 1.04 | 48.0 | 0.84 | 0.35 | 0.35 | 0.54 | 0.58 | 0.86 | 0.86 | N/A | N/A |
| i7-5960X@4.40 | 24.03 | 4.38 | 150.0 | 3.50 | 1.01 | 1.01 | 1.64 | 1.78 | 2.78 | 2.80 | 3.43 | N/A |
| R7 1700@3.75 | 21.66 | 3.45 | 110.0 | 2.74 | 1.00 | 1.01 | 1.67 | 1.87 | 2.61 | 2.60 | 2.67 | N/A |
| x5670@4.00 | 11.33 | 1.58 | 69.0 | 1.32 | 0.56 | 0.56 | 0.88 | 0.94 | N/A | N/A | N/A | N/A |
| i7-4770K@4.50 | 10.89 | 2.11 | 72.0 | 1.75 | 0.49 | 0.49 | 0.80 | 0.86 | 1.34 | 1.37 | 1.62 | N/A |
| i7-2600K@4.20 | 9.27 | 1.33 | 58.0 | 1.10 | 0.44 | 0.44 | 0.70 | 0.75 | 1.10 | 1.11 | N/A | N/A |
| i5-2500K@4.50 | 6.95 | 1.15 | 52.0 | 0.98 | 0.41 | 0.41 | 0.61 | 0.66 | 0.96 | 0.95 | N/A | N/A |
|----------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| R7 1700@stock | 18.51 | 2.98 | 103.0 | 2.32 | 0.96 | 0.95 | 1.41 | 1.54 | 2.11 | 2.05 | 2.08 | N/A |
| R7 1700X@stock | 20.61 | | | | | | | | | | | |
| R7 1800X@stock | 21.28 | 3.42 | 118.0 | 2.72 | 1.13 | 1.13 | 1.71 | 1.83 | 2.52 | 2.53 | 2.65 | N/A |
| i7-6700K@stock | 11.95 | 2.28 | 82.0 | 1.89 | 0.56 | 0.55 | 0.88 | 0.97 | 1.47 | 1.50 | 1.85 | N/A |
| i7-7700K@stock | 14.02 | | | | | | | | | | | |
| i7-5960X@stock | 18.18 | 3.30 | 116.0 | 2.66 | 0.76 | 0.76 | 1.23 | 1.33 | 2.09 | 2.09 | 2.58 | N/A |
| i7-6900K@stock | 19.83 | | | | | | | | | | | |
| i7-6950K@stock | 22.10 | | | | | | | | | | | |
| i7-4770@stock | 10.30 | 1.89 | 67.0 | 1.55 | 0.43 | 0.44 | 0.71 | 0.76 | 1.21 | 1.21 | 1.51 | N/A |
| E5-2670@stock | 12.64 | 1.80 | 79.0 | 1.50 | 0.62 | 0.62 | 0.94 | 1.02 | 1.46 | 1.50 | N/A | N/A |
|----------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|

x265 is new benchmark; You can find result for x264 here:

http://www.hardware.fr/getgraphimg.php?id=506&n=9

Sagittaire
2nd April 2017, 23:30
update

Brazil2
3rd April 2017, 10:15
The archive is broken.

Benchmark.zip.007 is only 20 MB and with 7-zip the archive test fails at 225 files 88%: Data error: Benchmark\Sample\Sample1080p.ts

And once extracted Sample1080p.ts is only 22 882 817 bytes large (199 159 304 expected).

Sagittaire
3rd April 2017, 21:02
I see that ... I (re)upload next weekend.

Sagittaire
8th April 2017, 16:26
Update ...

Try and report your result please ...

Danielcz
12th April 2017, 12:44
some error did not complete:

Motenai Yoda
12th April 2017, 21:07
Can you capture fps values and put them into a txt file?
also maybe adding -v quite to ffmpeg cli can help.

Selur
14th April 2017, 11:35
using a R7 1800X@stock I got the same problem as Danilecz -> https://pastebin.com/xEd81pxZ
ffmpeg\ffmpeg.exe -i Sample\Exodus_UHD_HDR_Exodus_draft.mp4 -an -f rawvideo - | x265\x265.exe --input-res 3840x2160 --fps 23.976 - -o Output\x265_2160p.265 --input-depth 10 --output-depth 10 --crf 24 --preset medium --tune grain --ssim --psnr --asm AVX,FMA3,FMA4,LZCNT,BMI1,BMI2,AVX2 --frames 100
aborts with:
frame= 105 fps=3.3 q=-0.0 size= 2551500kB time=00:00:04.37 bitrate=4777574.4kbiError writing trailer of pipe:: Broken pipe
frame= 105 fps=2.5 q=-0.0 Lsize= 2551500kB time=00:00:04.37 bitrate=4777574.4kbits/s speed=0.104x
video:2551500kB audio:0kB subtitle:0kB other streams:0kB global headers:0kB muxing overhead: 0.000000%
Conversion failed!

Cu Selur

Sagittaire
14th April 2017, 13:41
broken pipe with ffmpeg is normal because i use --frames 100 for SIMD test with x265.

benchmark work well with R7 1700@stock and R7 1700@3.75 ghz for 2 samples here in France.

I make new automatised benchmark: downlaod, retry and report your result. Thanks.

Selur
14th April 2017, 15:41
I'm still getting the:
av_interleaved_write_frame(): Invalid argument
Error writing trailer of pipe:: Invalid argument
but this time a results.log was created:
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| CPU | x264 | x265 | LAVC | auto | MMX2 | SSE | SSE2 | SSE3 | SSE4 | AVX | AVX2 | All |
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| Ryzen 7 1800X Eight-Core Processor | 21.28 | 3.42 | 118 | 2.72 | 1.13 | 1.13 | 1.71 | 1.83 | 2.52 | 2.53 | 2.65 | N/A |
(I ran the Benchmark_auto.bat)

---

also ran Bechmark.bat an it runs through without a problem now

Danielcz
14th April 2017, 17:46
i7 4770 stock (non-K), R9 290X, 32GB RAM, Win 7 utimate x64

|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| CPU | x264 | x265 | LAVC | auto | MMX2 | SSE | SSE2 | SSE3 | SSE4 | AVX | AVX2 | All |
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| Core i7-4770 | 10.30 | 1.89 | 67 | 1.55 | 0.43 | 0.44 | 0.71 | 0.76 | 1.21 | 1.21 | 1.51 | N/A |

Sagittaire
14th April 2017, 20:22
I'm still getting the:
av_interleaved_write_frame(): Invalid argument
Error writing trailer of pipe:: Invalid argument
but this time a results.log was created:
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| CPU | x264 | x265 | LAVC | auto | MMX2 | SSE | SSE2 | SSE3 | SSE4 | AVX | AVX2 | All |
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| Ryzen 7 1800X Eight-Core Processor | 21.28 | 3.42 | 118 | 2.72 | 1.13 | 1.13 | 1.71 | 1.83 | 2.52 | 2.53 | 2.65 | N/A |
(I ran the Benchmark_auto.bat)

---

also ran Bechmark.bat an it runs through without a problem now

yes ... really powerfull CPU this Rysen 7.

Ma
14th April 2017, 23:03
On my i5 3450S (with enhanced turbo) x265 4K encoding speed was 1.06.
For curiosity I changed x265 to VS 2017 AVX version and the speed was 1.11. Then I used VS 2017 AVX PGO version and the speed was 1.15 (I added only '-v warning' to ffmpeg part to cleaner output).

Then I go to the decoding -- with included in benchmark ffmpeg the speed was 48 to 49, with Zeranoe ffmpeg 2017-04-11 it was 64, with compiled by GCC 7 today snapshot the speed was 66.

The speed difference in x265 encoding is normal, but it is surprising big difference in decoding speed.

RanmaCanada
15th April 2017, 08:07
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| CPU | x264 | x265 | LAVC | auto | MMX2 | SSE | SSE2 | SSE3 | SSE4 | AVX | AVX2 | All |
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| Xeon E5-2670 0 | 12.64 | 1.80 | 79 | 1.50 | 0.62 | 0.62 | 0.94 | 1.02 | 1.46 | 1.50 | N/A | N/A |

Sagittaire
15th April 2017, 12:29
On my i5 3450S (with enhanced turbo) x265 4K encoding speed was 1.06.
For curiosity I changed x265 to VS 2017 AVX version and the speed was 1.11. Then I used VS 2017 AVX PGO version and the speed was 1.15 (I added only '-v warning' to ffmpeg part to cleaner output).

Then I go to the decoding -- with included in benchmark ffmpeg the speed was 48 to 49, with Zeranoe ffmpeg 2017-04-11 it was 64, with compiled by GCC 7 today snapshot the speed was 66.

The speed difference in x265 encoding is normal, but it is surprising big difference in decoding speed.

Yes i know for x265, I will try to introduce script to use best x265 compilation for each CPU ... ;-)

Anyway your ffmpeg result is really big surprise. You have link for ffmpeg 64 bit with GCC7 compilation?

Ma
15th April 2017, 15:31
ffmpeg (without any libs) compiled by GCC 7 from snapshot 2017-04-14:
www.msystem.waw.pl/x265/ffmpeg-2017-04-14.7z

The ffmpeg2.exe is the same with '-O2' optimize option instead of default '-O3 -fno-tree-vectorize'. For my i5 3450S this ffmpeg2.exe is slightly faster.

It looks like ffmpeg is faster in decoding HEVC than one month before.

Sagittaire
16th April 2017, 16:31
ffmpeg (without any libs) compiled by GCC 7 from snapshot 2017-04-14:
www.msystem.waw.pl/x265/ffmpeg-2017-04-14.7z

The ffmpeg2.exe is the same with '-O2' optimize option instead of default '-O3 -fno-tree-vectorize'. For my i5 3450S this ffmpeg2.exe is slightly faster.

It looks like ffmpeg is faster in decoding HEVC than one month before.

yes ... really higher speed.

THX for the ffmpeg build.

shinchiro
16th April 2017, 16:43
It looks like ffmpeg is faster in decoding HEVC than one month before.
Yeah..lots of hevc asm landed in ffmpeg's upstream recently, like:
http://git.videolan.org/?p=ffmpeg.git;a=commit;h=947230837cb6d64323590650554dad7abaf9a93f

NikosD
18th April 2017, 10:31
Hey,

I have made a few corrections to your "run.sh" file in order to be more consistent to what is running and what is being displayed.

For example, there is no "SSE3" instruction set tested, it's "SSSE3" and so on.

Also, the "All" setting is actually like "auto" so we don't need that.

I added also --no-asm option and I named it "No SIMD" in the final text, in order to see better the speedups of various SIMD sets compared to "No SIMD" setting

All of my changes are here:
http://txt.do/drlqm

SquallMX
18th April 2017, 19:00
6700K Stock 4.0 Ghz

|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| CPU | x264 | x265 | LAVC | auto | MMX2 | SSE | SSE2 | SSE3 | SSE4 | AVX | AVX2 | All |
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| Core i7-6700K | 11.95 | 2.28 | 82 | 1.89 | 0.56 | 0.55 | 0.88 | 0.97 | 1.47 | 1.50 | 1.85 | N/A |

Sagittaire
19th April 2017, 16:17
Hey,

I have made a few corrections to your "run.sh" file in order to be more consistent to what is running and what is being displayed.

For example, there is no "SSE3" instruction set tested, it's "SSSE3" and so on.

Also, the "All" setting is actually like "auto" so we don't need that.

I added also --no-asm option and I named it "No SIMD" in the final text, in order to see better the speedups of various SIMD sets compared to "No SIMD" setting

All of my changes are here:
http://txt.do/drlqm

yes, I have little bug for "all SIMD" bench.
I will make new script for choose the best x264 and x265 build between VS2017, GCC7 and ICC17

Sagittaire
19th April 2017, 16:19
6700K Stock 4.0 Ghz

|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| CPU | x264 | x265 | LAVC | auto | MMX2 | SSE | SSE2 | SSE3 | SSE4 | AVX | AVX2 | All |
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| Core i7-6700K | 11.95 | 2.28 | 82 | 1.89 | 0.56 | 0.55 | 0.88 | 0.97 | 1.47 | 1.50 | 1.85 | N/A |

the most interessing result ... thx

you can (re)make benchmark at your max overclocking?

Stock 4.0 Ghz

in fact stock is at 4.1 Ghz because turbo max for all core is 4.1 Ghz
You choose imposed 4.0 Ghz for all core or your i7-6700K is really in default stock frequency?

NikosD
19th April 2017, 16:59
yes, I have little bug for "all SIMD" bench.
I will make new script for choose the best x264 and x265 build between VS2017, GCC7 and ICC17
Don't forget to add the --no-asm option.

The "auto" is just fine for "all" SIMD sets.

Also, as I had told you in PM, the ffmpeg LAV video decoder is a lot faster in HEVC decoding than the script, probably due to more threads involved.

But using that script we can compare different CPUs to each other on the same script and not to absolutely fastest decoding.

SquallMX
19th April 2017, 18:09
the most interesting result ... thx

you can (re)make benchmark at your max overclocking?



in fact stock is at 4.1 Ghz because turbo max for all core is 4.1 Ghz
You choose imposed 4.0 Ghz for all core or your i7-6700K is really in default stock frequency?

i7-6700K stock clock for all cores at full is 4.0 GHz. Unfortunately is a DTR so no overclocking because of the cooling conditions:(.

Motenai Yoda
20th April 2017, 12:19
here my results.log
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| CPU | x264 | x265 | LAVC | auto | MMX2 | SSE | SSE2 | SSE3 | SSE4 | AVX | AVX2 | All |
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| Core i7 920 @3.8 | 7.90 | 1.14 | 49 | 0.96 | 0.37 | 0.37 | 0.60 | 0.65 | 0.96 | N/A | N/A | N/A |

RanmaCanada
23rd April 2017, 17:40
Well it looks like according to these benches I need to upgrade to at least R1700.

Thank you for the comprehensive tables.

Fador
25th April 2017, 08:29
Xeon E5-2699 v4 @ Stock

|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| CPU | x264 | x265 | LAVC | auto | MMX2 | SSE | SSE2 | SSE3 | SSE4 | AVX | AVX2 | All |
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| Xeon E5-2699 v4 | 31.04 | 6.79 | 135 | 4.95 | 1.72 | 1.73 | 2.66 | 2.88 | 4.21 | 4.22 | 4.87 | N/A |

Yanak
27th April 2017, 15:58
Hello, CPU runs @4.2Ghz
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| CPU | x264 | x265 | LAVC | auto | MMX2 | SSE | SSE2 | SSE3 | SSE4 | AVX | AVX2 | All |
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| Core i7-3770K | 10.35 | 1.46 | 72 | 1.22 | 0.47 | 0.47 | 0.76 | 0.81 | 1.22 | 1.23 | N/A | N/A |

shh
30th April 2017, 09:51
2x Xeon E5-2699 v4 @2.20GHz, 128GB RAM, 2xThreads: 22/22 (HT: off)
https://ark.intel.com/de/products/91317/Intel-Xeon-Processor-E5-2699-v4-55M-Cache-2_20-GHz
2x22 full cores oviously are a little better than just one E5-2699v4 with Hyper-Threading on. :)

|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| CPU | x264 | x265 | LAVC | auto | MMX2 | SSE | SSE2 | SSE3 | SSE4 | AVX | AVX2 | All |
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| 2x Xeon E5-2699 v4 | 37.19 | 9.36 | 118 | 5.48 | 2.13 | 2.12 | 3.25 | 3.48 | 4.90 | 4.89 | 5.28 | N/A |

Selur
3rd June 2017, 08:06
I overclocked my system to 4 GHz to see how this would affect the benchmark, and got some mixed results.
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| CPU | x264 | x265 | LAVC | auto | MMX2 | SSE | SSE2 | SSE3 | SSE4 | AVX | AVX2 | All |
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| Ryzen 7 1800X Eight-Core Processor | 21.91 | 3.39 | 74 | 2.59 | 1.19 | 1.14 | 1.66 | 1.76 | 2.05 | 2.18 | 2.12 | N/A |4GHz (Mainboard default)
| Ryzen 7 1800X Eight-Core Processor | 22.14 | 3.43 | 73 | 2.52 | 1.21 | 1.15 | 1.64 | 1.73 | 2.22 | 2.24 | 2.32 | N/A |4GHz (only changed the cpu mult)
| Ryzen 7 1800X Eight-Core Processor | 22.35 | 3.41 | 74 | 2.59 | 1.14 | 1.14 | 1.56 | 1.75 | 2.17 | 2.32 | 2.37 | N/A | 4GHz (only changed the cpu mult) + MalwareBytes and Windows Defender disabled
I'm wondering why the LAVC benchmark is so low compared to the old unoverclocked result I posted before (https://forum.doom9.org/showthread.php?p=1803733#post1803733):
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| CPU | x264 | x265 | LAVC | auto | MMX2 | SSE | SSE2 | SSE3 | SSE4 | AVX | AVX2 | All |
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| Ryzen 7 1800X Eight-Core Processor | 21.28 | 3.42 | 118 | 2.72 | 1.13 | 1.13 | 1.71 | 1.83 | 2.52 | 2.53 | 2.65 | N/A |@Stock speed

Not aware that I changed anything else on my system aside from updating Windows, since the last benchmark.

Cu Selur

Sagittaire
3rd June 2017, 20:04
Hi Selur

AMD make new compilator optimized for Ryzen (linux only?)
http://developer.amd.com/tools-and-sdks/cpu-development/amd-optimizing-cc-compiler/

You can try to make compilation for x264 and x265, make test with this benchmark and post your result?

THX

jd17
29th July 2017, 15:32
If anyone is interested, I tested my i5-7500 as well, thanks for compiling the benchmark! :)

|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| CPU | x264 | x265 | LAVC | auto | MMX2 | SSE | SSE2 | SSE3 | SSE4 | AVX | AVX2 | All |
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| Core i5-7500 | 9.21 | 1.89 | 64 | 1.55 | 0.45 | 0.45 | 0.72 | 0.78 | 1.19 | 1.20 | 1.52 | N/A |

sacd
5th August 2017, 10:24
Interesting benchmark, measures all the right applications.
This is my new editing PC running the benchmark_auto from the first post dated April 14 2017 (CPU OC 4.5GHz):


|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| CPU | x264 | x265 | LAVC | auto | MMX2 | SSE | SSE2 | SSE3 | SSE4 | AVX | AVX2 | All |
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| Core i9-7900X | 32.97 | 6.24 | 157 | 4.70 | 1.57 | 1.58 | 2.50 | 2.74 | 4.07 | 3.98 | 4.82 | N/A |

drizzit
6th August 2017, 17:07
My results with the auto bat, manually overclocked to 4.4GHz

|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| CPU | x264 | x265 | LAVC | auto | MMX2 | SSE | SSE2 | SSE3 | SSE4 | AVX | AVX2 | All |
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| Core i7-4790K | 11.70 | 2.10 | 78 | 1.74 | 0.49 | 0.49 | 0.82 | 0.89 | 1.39 | 1.40 | 1.75 | N/A |

Sagittaire
10th August 2017, 15:34
you can say if you have real full CPU charge at 100% for x264 test with 8C/16T CPU and higher (if someone have 16C/32T).

thx

Clare
10th August 2017, 21:45
I had to modify the script to run in Linux.


|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| CPU | x264 | x265 | LAVC | auto | MMX2 | SSE | SSE2 | SSE3 | SSE4 | AVX | AVX2 | All |
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| Core i7-6700HQ | 8.96 | 1.75 | 88 | 1.43 | 0.27 | 0.28 | 0.49 | 0.54 | 1.05 | 1.12 | 1.45 | N/A |


Here is the link to the modified script for Linux: https://gist.github.com/WyohKnott/72a7f35d28062bf48610a4aea92af788
It depends on python-cpuinfo, and of course, you need to install ffmpeg, x264 and x265 from your distribution repository.

Balthazar2k4
22nd August 2017, 22:02
Here you go:
|------------------------------------------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| CPU | x264 | x265 | LAVC | auto | MMX2 | SSE | SSE2 | SSE3 | SSE4 | AVX | AVX2 | All |
|------------------------------------------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| Ryzen Threadripper 1950X 16-Core Processor | 37.59 | 6.50 | 136 | 4.65 | 2.02 | 2.00 | 2.95 | 3.10 | 3.89 | 4.10 | 4.26 | N/A |

Edit: x264 does NOT utilize all 32 threads. During that portion of the test I was seeing ~80% utilization.

kolak
24th August 2017, 21:01
2x Xeon E5-2699 v4 @2.20GHz, 128GB RAM, 2xThreads: 22/22 (HT: off)
https://ark.intel.com/de/products/91317/Intel-Xeon-Processor-E5-2699-v4-55M-Cache-2_20-GHz
2x22 full cores oviously are a little better than just one E5-2699v4 with Hyper-Threading on. :)

|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| CPU | x264 | x265 | LAVC | auto | MMX2 | SSE | SSE2 | SSE3 | SSE4 | AVX | AVX2 | All |
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| 2x Xeon E5-2699 v4 | 37.19 | 9.36 | 118 | 5.48 | 2.13 | 2.12 | 3.25 | 3.48 | 4.90 | 4.89 | 5.28 | N/A |

Far beyond reasonable scaling. Such a waste of processing power :) This is only good for running many encodes at the same time or splitting long encodes into chunks.
Can you test it with 1 core, so we can see how all this cores are wasted :)

kolak
24th August 2017, 21:05
Interesting benchmark, measures all the right applications.
This is my new editing PC running the benchmark_auto from the first post dated April 14 2017 (CPU OC 4.5GHz):


|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| CPU | x264 | x265 | LAVC | auto | MMX2 | SSE | SSE2 | SSE3 | SSE4 | AVX | AVX2 | All |
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| Core i9-7900X | 32.97 | 6.24 | 157 | 4.70 | 1.57 | 1.58 | 2.50 | 2.74 | 4.07 | 3.98 | 4.82 | N/A |

Just shows how good clock makes things fast.

bladerunner1982
27th August 2017, 20:06
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| CPU | x264 | x265 | LAVC | auto | MMX2 | SSE | SSE2 | SSE3 | SSE4 | AVX | AVX2 | All |
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| Pentium G4560 | 5.02 | 0.77 | 37 | 0.66 | 0.25 | 0.25 | 0.40 | 0.44 | 0.67 | N/A | N/A | N/A |

zurv
10th October 2017, 19:15
Here is the auto benchmark run on (windows 10 x64)
i9-7980xe @ 4.7 (with AVX offset of 12 and avx512 of 15)

|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| CPU | x264 | x265 | LAVC | auto | MMX2 | SSE | SSE2 | SSE3 | SSE4 | AVX | AVX2 | All |
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| Core i9-7980XE | 44.89 | 9.34 | 186 | 6.84 | 2.48 | 2.46 | 3.69 | 3.86 | 5.25 | 5.75 | 6.71 | N/A |


I'm still playing around with the offsets. But 18 cores + avx. stuff gets HOT (mainly the VRMs)

NikosD
15th October 2017, 11:39
Here is the auto benchmark run on (windows 10 x64)
i9-7980xe @ 4.7 (with AVX offset of 12 and avx512 of 15)
[
I'm still playing around with the offsets. But 18 cores + avx. stuff gets HOT (mainly the VRMs)

With an all core turbo clock@4.7 GHz you are approaching a 600W power consumption.

Have you delidded the CPU ?

What is the translation of AVX offset of 12 in clock speed ?

I'm assuming that you have a good liquid cooler in order to go that high.

nevcairiel
15th October 2017, 12:06
What is the translation of AVX offset of 12 in clock speed ?

Every offset reduces the clock by 100MHz, so thats 1200 less in AVX2 mode.

zurv
15th October 2017, 15:46
oddly the 7980 is ez'r to cool than the 7900 (i have that too), the die is much larger on the 7980. Yes, i delided but it wasn't as helpful as with other CPUs (maybe 5C drop)

The issue with the AVX stuff (and 512 even more) is the power draw. Cooling the CPU isn't a problem, but so much power is going through the CPU package and VRM.. ugh. (I have a mono block too.) (Mesh is a problem too as the jacks up power usage.. normally i'd want to run it at 32 (default is 24.) But AVX + higher mesh is monster power/heat.
Also, quick stuff (even this benchmark) is fine. The hard part is something that will be stable for a long time. Doing a 4k video (game capture lossless) for 30min in x264 is still 2-3 hours. (hrmm.. video that no one really look at.. and YT totally re-encodes to pooo quality.. so maybe i shouldn't be using "very slow"... )

I wonder how much faster avx is vs not using it. Ie, over 18 cores I'm losing over 18gigs of speed when using avx. I'm assuming that it is worth it.

(also, the default speed for the 7980 is 2.6ghz. It OCs really well. Most people are getting 4.5-.49 on the OC. AVX is a problem, but most don't use that. That said, the new Time Spy 4k 3dmark test does.)

Atak_Snajpera
15th October 2017, 17:27
(also, the default speed for the 7980 is 2.6ghz. It OCs really well. Most people are getting 4.5-.49 on the OC. AVX is a problem, but most don't use that. That said, the new Time Spy 4k 3dmark test does.)
If you do not care about AVX performance then maybe you should just have bought ThreadRipper 1950x for 2 times less. In non-AVX tasks difference between those two is not big.

zurv
15th October 2017, 17:35
that is a silly reply. One could just have the encoder not use avx. clock to clock intel is faster and there are more core on the 7980. (also the 7980 OCs much higher than the threadripper too. So more cores, clock to clock faster and faster clocks.) perf matters more than cost. Other than encoding nothing i do uses avx. So even for single core usages this CPU is better for me than threadripper. (the 7980 OCs just as well as the 7900 (which i also have) ... well.. other than avx :) )
AVX 512 is a big deal (or so i'm told :) ) the i9s are the only desktop CPU with it. I'm not fully clear by that is :) but i haven't looked into it much.

Atak_Snajpera
15th October 2017, 18:06
clock to clock intel is faster
AMD 1 core+SMT=Intel 1 core+HT

http://i.cubeupload.com/Tsc1ox.png

AVX 512 is a big deal (or so i'm told ) the i9s are the only desktop CPU with it.
You will have to wait another 10 years to see decent support in apps for this SIMD instructions. AVX-512 for now is just a placebo.

NikosD
15th October 2017, 18:16
AVX 512 is a big deal (or so i'm told :) ) the i9s are the only desktop CPU with it. I'm not fully clear by that is :) but i haven't looked into it much.

AVX512 is simply not existent and as Atak_Snajpera said in the previous post, don't expect it soon.

Anandtech.com is looking for apps leveraging AVX512 since i9 7900X launch and they didn't find any to test it.

Also, as i asked you before, do you use an AIO liquid cooler for the 18core ?

littleD
27th December 2017, 07:38
Not really impressive but good to know
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| CPU | x264 | x265 | LAVC | auto | MMX2 | SSE | SSE2 | SSE3 | SSE4 | AVX | AVX2 | All |
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| Core i5-7200U | 4.79 | 0.94 | 35 | 0.78 | 0.23 | 0.23 | 0.36 | 0.39 | 0.61 | 0.61 | 0.77 | N/A |

Boulder
3rd January 2018, 18:41
So what is the current optimal bang for buck CPU to get if 95% of all use is encoding with x265 and processing with Vapoursynth? I'm unsure about the current reasonable prices because Ryzens and i7's are often quite close?

JMX
7th January 2018, 17:37
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| CPU | x264 | x265 | LAVC | auto | MMX2 | SSE | SSE2 | SSE3 | SSE4 | AVX | AVX2 | All |
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| Core i9-7900X | 29.37 | 5.47 | 145 | 4.38 | 1.35 | 1.35 | 2.14 | 2.32 | 3.49 | 3.50 | 4.30 | N/A |



no overclocking of any kind - just a stock computer
some error messages while the auto script ran but I didn't really payed attention.

RanmaCanada
26th April 2018, 00:03
Anyone with a Ryzen+ want to run this bench so we can compare apples to apples? It would be interesting to see it benched properly as the review sites aren't doing x265 benchmarks justice.

TEB
26th April 2018, 14:12
av_interleaved_write_frame(): Invalid argument
Error writing trailer of pipe:: Invalid argument00:00:04.37 bitrate=4777574.4kbits/s speed=0.123x
frame= 105 fps=2.3 q=-0.0 Lsize= 2551500kB time=00:00:04.37 bitrate=4777574.4kbits/s speed=0.094x
video:2551500kB audio:0kB subtitle:0kB other streams:0kB global headers:0kB muxing overhead: 0.000000%
Conversion failed!
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| CPU | x264 | x265 | LAVC | auto | MMX2 | SSE | SSE2 | SSE3 | SSE4 | AVX | AVX2 | All |
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| Ryzen 7 1800X | 19.96 | 2.85 | 67 | 2.08 | 1.02 | 1.03 | 1.52 | 1.58 | 1.91 | 2.09 | 2.16 | N/A |

dex8472
17th June 2018, 04:59
Hi guys, here's my benchmark. Stock config.

Ryzen 7 2700x @ 3.70 GHz
i7-4770k @ 3.50 GHz
i7-4790k @ 4 GHz

|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| CPU | x264 | x265 | LAVC | auto | MMX2 | SSE | SSE2 | SSE3 | SSE4 | AVX | AVX2 | All |
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| Ryzen 7 2700X | 21.62 | 3.51 | 124 | 2.78 | 1.14 | 1.15 | 1.75 | 1.90 | 2.72 | 2.74 | 2.84 | N/A |
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| Core i7-4770K | 10.03 | 1.84 | 64 | 1.50 | 0.42 | 0.42 | 0.69 | 0.75 | 1.16 | 1.19 | 1.47 | N/A |
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| Core i7-4790K | 10.76 | 2.03 | 65 | 1.59 | 0.47 | 0.47 | 0.77 | 0.84 | 1.29 | 1.31 | 1.58 | N/A |

RanmaCanada
21st July 2018, 06:06
Is there any way we can get the benchmark updated, as several people with Epyc machines have tried to run it, and it will not run.

NikosD
21st July 2018, 08:49
If it doesn't run properly for EPYC it won't run for Threadripper 2, too.

DotJun
23rd July 2018, 05:28
Will you be adding avx512 to the benchmark?

Atak_Snajpera
23rd July 2018, 09:34
If it doesn't run properly for EPYC it won't run for Threadripper 2, too.

Why???

Atak_Snajpera
23rd July 2018, 09:35
Will you be adding avx512 to the benchmark?

avx512 is almost useless
https://networkbuilders.intel.com/docs/accelerating-x265-the-hevc-encoder-with-intel-advanced-vector-extensions-512.pdf

NikosD
23rd July 2018, 12:13
Why???Due to the similar many core architecture I guess.

Unless of course EPYC has a different problem.

Atak_Snajpera
23rd July 2018, 13:21
Is there any way we can get the benchmark updated, as several people with Epyc machines have tried to run it, and it will not run.

I doubt that this benchmark is designed to saturate all logical processors on EPYC 1(64threads)/2(96threads) anyway.

2160p is only enough up to 32 threads. You will have to run 2 instances to get 100% usage in this case on EPYC 1 and 3 for EPYC 2.

RanmaCanada
24th July 2018, 22:00
I doubt that this benchmark is designed to saturate all logical processors on EPYC 1(64threads)/2(96threads) anyway.

2160p is only enough up to 32 threads. You will have to run 2 instances to get 100% usage in this case on EPYC 1 and 3 for EPYC 2.

True, but the problem is that people are reporting that it won't run on Windows Server SR2. As I don't have an EPYC myself, I sadly can't run the benchmark, and the people are only posting in the threads on reddit, and not creating accounts and posting here to report the issues. :( People have asked them to come here, but to no avail, so it's like a long game of telephone. Hopefully when TR2 is released we will get some good benchmarks :)

Atak_Snajpera
25th July 2018, 11:59
tell them to run x265 FHD Benchmark
https://s22.postimg.cc/t2k45na75/Untitled-1.png

Ps. Overclocked 2990x@4GHz should give you ~90fps in 1080p with default settings.

RanmaCanada
23rd August 2018, 00:12
And even you are having troubles with your benchmark haha :P I guess we need some new software to abuse these multi-core beasts?

Atak_Snajpera
23rd August 2018, 10:30
And even you are having troubles with your benchmark haha :P I guess we need some new software to abuse these multi-core beasts?

Because numa nodes were not set for each encoder. Who would have thought that lack of direct access to memory by DIE1 and DIE3 will be so devastating for x265 video encoding. Relax. New version will assign each encoder to own NUMA node. This should reduce data traffic on infinity nodes and hence improve speed.

Forteen88
13th December 2019, 00:46
I got "av_interleaved_write_frame(): Invalid argument" too, using "Benchmark_auto.bat" file. But the results.log file showed,
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| CPU | x264 | x265 | LAVC | auto | MMX2 | SSE | SSE2 | SSE3 | SSE4 | AVX | AVX2 | All |
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| Ryzen 5 3600 | 20.82 | 3.86 | 118 | 3.13 | 0.97 | 1.00 | 1.55 | 1.69 | 2.51 | 2.55 | 3.16 | N/A |

AMD Ryzen 3600 with DDR4-3200@3466.

RanmaCanada
18th April 2020, 08:15
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| CPU | x264 | x265 LAVC | auto | MMX2 | SSE | SSE2 | SSE3| SSE4 | AVX | AVX2 | All |
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| Ryzen 7 2700 | 17.78 | 2.97 | 99 | 2.34 | 0.96 | 0.90 | 1.47 | 1.59 | 2.22 | 2.31 | 2.34 | N/A |

In comparison to my E5-2670 that I posted exactly 3 years ago!

|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| CPU | x264 | x265 | LAVC | auto | MMX2 | SSE | SSE2 | SSE3 | SSE4 | AVX | AVX2 | All |
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| Xeon E5-2670 0 | 12.64 | 1.80 | 79 | 1.50 | 0.62 | 0.62 | 0.94 | 1.02 | 1.46 | 1.50 | N/A | N/A |

Is there anyway we can get the benchmark updated with newer builds of software? I have tried to do it myself, but I failed, badly.

edit: why can't I get the formating to stick...

Zebulon84
18th April 2020, 12:13
As there is no CPU like mine yet, I've tried the benchmark. I had some errors, but it seems I'm not the only one so here is the result of my 11 years old i5 750:

|---------------------|---------|---------|--------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| CPU | x264 | x265 | LAVC | auto | MMX2 | SSE | SSE2 | SSE3 | SSE4 | AVX | AVX2 | All |
|---------------------|---------|---------|--------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| Core i5 750 | 4.70 | 0.71 | 31 | 0.59 | 0.24 | 0.24 | 0.37 | 0.39 | 0.57 | N/A | N/A | N/A |


To have a better view of different CPU, I gathered all results posted in this thread, Intel then AMD sorted by x265 value:

|---------------------|---------|---------|--------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| CPU | x264 | x265 | LAVC | auto | MMX2 | SSE | SSE2 | SSE3 | SSE4 | AVX | AVX2 | All |
|---------------------|---------|---------|--------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| 2x Xeon E5-2699 v4 | 37.19 | 9.36 | 118 | 5.48 | 2.13 | 2.12 | 3.25 | 3.48 | 4.90 | 4.89 | 5.28 | N/A |
| Core i9-7980XE @4.7| 44.89 | 9.34 | 186 | 6.84 | 2.48 | 2.46 | 3.69 | 3.86 | 5.25 | 5.75 | 6.71 | N/A |
| Xeon E5-2699 v4 | 31.04 | 6.79 | 135 | 4.95 | 1.72 | 1.73 | 2.66 | 2.88 | 4.21 | 4.22 | 4.87 | N/A |
| Core i9-7900X @4.5 | 32.97 | 6.24 | 157 | 4.70 | 1.57 | 1.58 | 2.50 | 2.74 | 4.07 | 3.98 | 4.82 | N/A |
| Core i9-7900X | 29.37 | 5.47 | 145 | 4.38 | 1.35 | 1.35 | 2.14 | 2.32 | 3.49 | 3.50 | 4.30 | N/A |
| Core i7-8750H | 13.14 | 2.33 | 79 | 1.90 | 0.64 | 0.64 | 1.02 | 1.07 | 1.64 | 1.55 | 1.85 | N/A |
| Core i7-5960X | 18.18 | 3.30 | 116 | 2.66 | 0.76 | 0.76 | 1.23 | 1.33 | 2.09 | 2.09 | 2.58 | N/A |
| Core i7-6700K | 11.95 | 2.28 | 82 | 1.89 | 0.56 | 0.55 | 0.88 | 0.97 | 1.47 | 1.50 | 1.85 | N/A |
| Core i7-4790K @4.4 | 11.70 | 2.10 | 78 | 1.74 | 0.49 | 0.49 | 0.82 | 0.89 | 1.39 | 1.40 | 1.75 | N/A |
| Core i7-4790K | 10.76 | 2.03 | 65 | 1.59 | 0.47 | 0.47 | 0.77 | 0.84 | 1.29 | 1.31 | 1.58 | N/A |
| Core i7-4770 | 10.30 | 1.89 | 67 | 1.55 | 0.43 | 0.44 | 0.71 | 0.76 | 1.21 | 1.21 | 1.51 | N/A |
| Core i5-7500 | 9.21 | 1.89 | 64 | 1.55 | 0.45 | 0.45 | 0.72 | 0.78 | 1.19 | 1.20 | 1.52 | N/A |
| Core i7-4770K | 10.03 | 1.84 | 64 | 1.50 | 0.42 | 0.42 | 0.69 | 0.75 | 1.16 | 1.19 | 1.47 | N/A |
| Xeon E5-2670 | 12.64 | 1.80 | 79 | 1.50 | 0.62 | 0.62 | 0.94 | 1.02 | 1.46 | 1.50 | N/A | N/A |
| Core i7-6700HQ | 8.96 | 1.75 | 88 | 1.43 | 0.27 | 0.28 | 0.49 | 0.54 | 1.05 | 1.12 | 1.45 | N/A |
| Core i7-3770K @4.2 | 10.35 | 1.46 | 72 | 1.22 | 0.47 | 0.47 | 0.76 | 0.81 | 1.22 | 1.23 | N/A | N/A |
| Core i7-2600K @4.5 | 9.43 | 1.39 | 62 | 1.13 | 0.47 | 0.46 | 0.72 | 0.78 | 1.14 | 1.18 | N/A | N/A |
| Core i7 920 @3.8 | 7.90 | 1.14 | 49 | 0.96 | 0.37 | 0.37 | 0.60 | 0.65 | 0.96 | N/A | N/A | N/A |
| Core i5-7200U | 4.79 | 0.94 | 35 | 0.78 | 0.23 | 0.23 | 0.36 | 0.39 | 0.61 | 0.61 | 0.77 | N/A |
| Pentium G4560 | 5.02 | 0.77 | 37 | 0.66 | 0.25 | 0.25 | 0.40 | 0.44 | 0.67 | N/A | N/A | N/A |
| Core i5 750 | 4.70 | 0.71 | 31 | 0.59 | 0.24 | 0.24 | 0.37 | 0.39 | 0.57 | N/A | N/A | N/A |
|---------------------|---------|---------|--------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| Ryzen 9 3900X | 39.30 | 7.25 | 144 | 5.43 | 1.96 | 1.96 | 3.02 | 3.26 | 4.58 | 4.60 | 5.51 | N/A |
| Threadripper 1950X | 37.59 | 6.50 | 136 | 4.65 | 2.02 | 2.00 | 2.95 | 3.10 | 3.89 | 4.10 | 4.26 | N/A |
| Ryzen 7 3700X | 24.25 | 4.58 | 130 | 3.76 | 1.26 | 1.26 | 1.94 | 2.03 | 2.75 | 2.97 | 3.71 | N/A |
| Ryzen 5 3600 | 20.82 | 3.86 | 118 | 3.13 | 0.97 | 1.00 | 1.55 | 1.69 | 2.51 | 2.55 | 3.16 | N/A |
| Ryzen 7 2700X | 21.62 | 3.51 | 124 | 2.78 | 1.14 | 1.15 | 1.75 | 1.90 | 2.72 | 2.74 | 2.84 | N/A |
| Ryzen 7 1800X | 21.28 | 3.42 | 118 | 2.72 | 1.13 | 1.13 | 1.71 | 1.83 | 2.52 | 2.53 | 2.65 | N/A |
| Ryzen 7 1800X @4 | 22.35 | 3.41 | 74 | 2.59 | 1.14 | 1.14 | 1.56 | 1.75 | 2.17 | 2.32 | 2.37 | N/A |
| Ryzen 1700 | 18.51 | 2.98 | 103 | 2.32 | 0.96 | 0.95 | 1.41 | 1.54 | 2.11 | 2.05 | 2.08 | N/A |
| Ryzen 7 2700 | 17.78 | 2.97 | 99 | 2.34 | 0.96 | 0.90 | 1.47 | 1.59 | 2.22 | 2.31 | 2.34 | N/A |
| Ryzen 7 1800X (TEB)| 19.96 | 2.85 | 67 | 2.08 | 1.02 | 1.03 | 1.52 | 1.58 | 1.91 | 2.09 | 2.16 | N/A |
|---------------------|---------|---------|--------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
Note that Selur and TEB have the same CPU (Ryzen 1800X) but with quite different results.

tormento
27th April 2020, 16:51
My results:
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| CPU | x264 | x265 | LAVC | auto | MMX2 | SSE | SSE2 | SSE3 | SSE4 | AVX | AVX2 | All |
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| Core i7-2600K @4.5 | 9.43 | 1.39 | 62 | 1.13 | 0.47 | 0.46 | 0.72 | 0.78 | 1.14 | 1.18 | N/A | N/A |
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|

masterkivat
27th April 2020, 19:37
My results:

(Ryzen 3900x stock, 16gb DDR4@3000MHz)

https://puu.sh/FDetX/1d23de58e8.png

(Sorry for the lazyness in not formatting to text here :D

benwaggoner
27th April 2020, 23:17
Because numa nodes were not set for each encoder. Who would have thought that lack of direct access to memory by DIE1 and DIE3 will be so devastating for x265 video encoding. Relax. New version will assign each encoder to own NUMA node. This should reduce data traffic on infinity nodes and hence improve speed.
I've not really found much content below 8K where a single encode materially benefits from multiple sockets(at least not with my dual "Intel(R) Xeon(R) Gold 6240 CPU @ 2.60GHz, 2594 Mhz, 18 Core(s), 36 Logical Processor(s)").

Greenhorn
7th May 2020, 01:04
Ryzen 3700x, base clocks:
x264: 24.25
x265: 4.58
LAVC: 130
auto: 3.76
MMX2: 1.26
SSE: 1.26
SSE2: 1.94
SSE3: 2.03
SSE4: 2.75
AVX: 2.97
AVX2: 3.71
All: N/A

RanmaCanada
23rd May 2020, 18:01
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| CPU | x264 | x265 | LAVC | auto | MMX2 | SSE | SSE2 | SSE3 | SSE4 | AVX | AVX2 | All |
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| Core i7-8750H | 13.14 | 2.33 | 79 | 1.90 | 0.64 | 0.64 | 1.02 | 1.07 | 1.64 | 1.55 | 1.85 | N/A |

RanmaCanada
7th August 2020, 00:18
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| CPU | x264 | x265 | LAVC | auto | MMX2 SSE | SSE2 | SSE3 | SSE4 | AVX | AVX2 | All |
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| Core i5-9300H | 10.68 | 1.86 | 57 | 1.49 | 0.51 | 0.50 | 0.80 | 0.86 | 1.29 | 1.20 | 1.47 | N/A |

RanmaCanada
5th April 2021, 21:02
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| CPU Ryzen 7 5800X | x264 | x265 | LAVC | auto | MMX2 | SSE | SSE2 | SSE3 | SSE4 | AVX | AVX2 | All |
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
|
| 8-Core Processor | 32.44 | 5.67 | 103 | 4.34 | 1.61 | 1.65 | 2.45 | 2.67 | 3.71 | 3.69 | 4.39 | N/A |

TEB
9th April 2021, 09:56
Does a linux console version of this benchmark exist? I have alot of new amd servers at work as well as alot of intel servers i can benchmark.. All with red hat ..

benwaggoner
9th April 2021, 17:55
Ryzen 3700x, base clocks:
x264: 24.25
x265: 4.58
LAVC: 130
auto: 3.76
MMX2: 1.26
SSE: 1.26
SSE2: 1.94
SSE3: 2.03
SSE4: 2.75
AVX: 2.97
AVX2: 3.71
All: N/A
Is it time to add an AVX-512 test as well? I know that won't work on lots of systems, and probably makes things slower in general for most use cases (on my Cascade Lake I only see speedups on 8K or 4K around veryslow). But perhaps things are different with the new Intel designs?

It'd be nice to have the test track that to see how AVX-512 perf improves over time.

RanmaCanada
10th April 2021, 00:19
Does a linux console version of this benchmark exist? I have alot of new amd servers at work as well as alot of intel servers i can benchmark.. All with red hat ..

AFAIK, no. We also need the exe files updated as they are ancient. I've attempted to do it on my own, but I have no clue what I am doing :D

Sagittaire
12th April 2021, 14:43
Is it time to add an AVX-512 test as well? I know that won't work on lots of systems, and probably makes things slower in general for most use cases (on my Cascade Lake I only see speedups on 8K or 4K around veryslow). But perhaps things are different with the new Intel designs?

It'd be nice to have the test track that to see how AVX-512 perf improves over time.

I will try to update that.

Anyway Zen3 is incredible update if you compare at Zen2 or Zen1 (All CPU are 8C/16T here): AVX2 produce terrible boost
I will try to use AV1 benchmark with 4K source.

|---------------------|---------|---------|--------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| CPU | x264 | x265 | LAVC | auto | MMX2 | SSE | SSE2 | SSE3 | SSE4 | AVX | AVX2 | All |
|---------------------|---------|---------|--------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| Ryzen 7 5800X | 32.44 | 5.67 | 103 | 4.34 | 1.61 | 1.65 | 2.45 | 2.67 | 3.71 | 3.69 | 4.39 | N/A |
| Ryzen 7 3700X | 24.25 | 4.58 | 130 | 3.76 | 1.26 | 1.26 | 1.94 | 2.03 | 2.75 | 2.97 | 3.71 | N/A |
| Ryzen 7 2700X | 21.62 | 3.51 | 124 | 2.78 | 1.14 | 1.15 | 1.75 | 1.90 | 2.72 | 2.74 | 2.84 | N/A |
| Ryzen 7 1800X | 21.28 | 3.42 | 118 | 2.72 | 1.13 | 1.13 | 1.71 | 1.83 | 2.52 | 2.53 | 2.65 | N/A |
|---------------------|---------|---------|--------|---------|---------|---------|---------|---------|---------|---------|---------|---------|

jriker1
22nd June 2021, 15:51
If we are collecting these still, below is a run thru of my old dual Xeon server. Note it did fail at one point:

x265 [info]: frame B: 35, Avg QP:23.89 kb/s: 17745.37 PSNR Mean: Y:49.108 U:52.469 V:56.906 SSIM Mean: 0.991256 (20.583dB)
x265 [info]: Weighted P-Frames: Y:0.0% UV:0.0%av_interleaved_write_frame(): Invalid argument
x265 [info]: consecutive B-frames: 61.5% 26.2% 9.2% 3.1% 0.0% frame= 105 fps=2.6 q=-0.0 size= 2551500kB time=00:00:04.37 bitrate=4777574.4kbits/s speed=0.11x
Error writing trailer of pipe:: Invalid argument
video:2551500kB audio:0kB subtitle:0kB other streams:0kB global headers:0kB muxing overhead:SIM Mean Y: 0.9897323 (19.885 dB)frame= 105 fps=1.9 q=-0.0 Lsize= 2551500kB time=00:00:04.37 bitrate=4777574.4kbits/s 0.000000%87x
Conversion failed!


|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| CPU | x264 | x265 | LAVC | auto | MMX2 | SSE | SSE2 | SSE3 | SSE4 | AVX | AVX2 | All |
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| Xeon X5680 | 18.03 | 2.29 | 74 | 1.81 | 0.76 | 0.76 | 1.21 | 1.30 | 1.80 | N/A | N/A | N/A |

Also isn't some of this dependent on the speed of the drive/ssd the data's being read/written to?

excellentswordfight
22nd June 2021, 20:16
If we are collecting these still, below is a run thru of my old dual Xeon server. Note it did fail at one point:




|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| CPU | x264 | x265 | LAVC | auto | MMX2 | SSE | SSE2 | SSE3 | SSE4 | AVX | AVX2 | All |
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| Xeon X5680 | 18.03 | 2.29 | 74 | 1.81 | 0.76 | 0.76 | 1.21 | 1.30 | 1.80 | N/A | N/A | N/A |

Also isn't some of this dependent on the speed of the drive/ssd the data's being read/written to?
I havnt looked at this benchmark specifically, but I assume that the sources are rather compressed, if thats the case then I doubt that it will have much effect as long as the drive is in a somewhat idle state. For example, if you have UHD-bluray soruce at 60Mbps, thats only 7,5MB/s reads if you can encode at realtime (24fps) and most people cannot even reach that for 2160p with hevc, and usually with x264/x265 you encode to something smaller so writes shoudlnt be an issue either.

jriker1
22nd June 2021, 22:53
is it just me or is my system keeping up with i7 CPU's from some of the consolidated info I have? I also have an i7-7700k so guessing that system would output at a similar speed even though my server has 24 threads and my i7 has 8 threads. Haven't tested the i7 yet but seems like similar i7's get similar results in the x265 space.

benwaggoner
22nd June 2021, 23:23
I will try to update that.

Anyway Zen3 is incredible update if you compare at Zen2 or Zen1 (All CPU are 8C/16T here): AVX2 produce terrible boost
I will try to use AV1 benchmark with 4K source.

|---------------------|---------|---------|--------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| CPU | x264 | x265 | LAVC | auto | MMX2 | SSE | SSE2 | SSE3 | SSE4 | AVX | AVX2 | All |
|---------------------|---------|---------|--------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| Ryzen 7 5800X | 32.44 | 5.67 | 103 | 4.34 | 1.61 | 1.65 | 2.45 | 2.67 | 3.71 | 3.69 | 4.39 | N/A |
| Ryzen 7 3700X | 24.25 | 4.58 | 130 | 3.76 | 1.26 | 1.26 | 1.94 | 2.03 | 2.75 | 2.97 | 3.71 | N/A |
| Ryzen 7 2700X | 21.62 | 3.51 | 124 | 2.78 | 1.14 | 1.15 | 1.75 | 1.90 | 2.72 | 2.74 | 2.84 | N/A |
| Ryzen 7 1800X | 21.28 | 3.42 | 118 | 2.72 | 1.13 | 1.13 | 1.71 | 1.83 | 2.52 | 2.53 | 2.65 | N/A |
|---------------------|---------|---------|--------|---------|---------|---------|---------|---------|---------|---------|---------|---------|

Do I read this as AVX2 being SLOWER than the ALL? What is the difference?

benwaggoner
22nd June 2021, 23:25
I havnt looked at this benchmark specifically, but I assume that the sources are rather compressed, if thats the case then I doubt that it will have much effect as long as the drive is in a somewhat idle state. For example, if you have UHD-bluray soruce at 60Mbps, thats only 7,5MB/s reads if you can encode at realtime (24fps) and most people cannot even reach that for 2160p with hevc, and usually with x264/x265 you encode to something smaller so writes shoudlnt be an issue either.
The only time I can recall where read speed was a bottleneck was having to do some uncompressed 4K over USB 2.0 some years back. As slow as x265 4K was at the time, the USB bottleneck still made it three times slower.

Nico8583
27th August 2021, 14:14
Hi,
Here are my results :
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| CPU | x264 | x265 | LAVC | auto | MMX2 | SSE | SSE2 | SSE3 | SSE4 | AVX | AVX2 | All |
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| Ryzen 7 3700X | 26.99 | 5.16 | 147 | 4.10 | 1.38 | 1.36 | 1.91 | 2.30 | 3.35 | 3.40 | 4.10 | N/A |

|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| CPU | x264 | x265 | LAVC | auto | MMX2 | SSE | SSE2 | SSE3 | SSE4 | AVX | AVX2 | All |
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| Core i7-8700 | 19.49 | 3.73 | 119 | 3.03 | 0.91 | 0.92 | 1.45 | 1.57 | 2.38 | 2.41 | 2.99 | N/A |

All CPU @ stock, RAM 2x8GB 3200Mhz for 3700X and 2x8GB 2666Mhz for 8700

Edit : The same benchmark with slow preset instead of medium for x265 :

|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| CPU | x264 | x265 | LAVC | auto | MMX2 | SSE | SSE2 | SSE3 | SSE4 | AVX | AVX2 | All |
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| Ryzen 7 3700X | 28.24 | 2.07 | 140 | 1.35 | 0.49 | 0.49 | 0.83 | 0.91 | 1.07 | 1.07 | 1.37 | N/A |

Loomes
10th November 2021, 14:30
Here are the results for my Ryzen 9 3900x, 16GB DDR4 3200Mhz CL16:

|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| CPU | x264 | x265 | LAVC | auto | MMX2 | SSE | SSE2 | SSE3 | SSE4 | AVX | AVX2 | All |
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| Ryzen 9 3900X | 40.68 | 7.73 | 189 | 5.81 | 2.04 | 2.04 | 3.12 | 3.41 | 4.67 | 4.74 | 5.66 | N/A |

...not too shabby but I was thinking about swapping my RAM for better 3600MHz CL 14 -- would there be significant improvements?

Boulder
10th November 2021, 18:14
Memory speed doesn't matter much. You could also probably overclock to 3600MHz anyway, and I think the sweet spot for that chip's FCLK is 1800MHz so you should definitely test it. Any extensive filtering and 4K sources may require upping the Avisynth+ default cache size a lot so 16GB might not cut it in all cases. I recall a few times when I ran out of memory and had to restrict the cache size which then lowered the performance in multithreaded mode. With 32GB, it's easy to just allow max 20GB for the cache (or which Avisynth+ uses around 10-12GB most of the time).

RanmaCanada
11th November 2021, 01:53
I would have to agree with Boulder. I swapped from 3200MHz regular ram to 2400MHz ECC and I did not see a measurable difference in regards to encode speeds on my 5800x. The only thing you can do is just get more ram to ensure you don't run out haha, which I've had happen more than once! If anyone has good leads on 2x32GB DDR4 ECC for desktops (unbuffered I believe), it would be appreciated haha.

Maybe it's time to ping Sagittaire again and ask them kindly if they can update this with more modern builds. haha.

excellentswordfight
11th November 2021, 09:25
On that subject, it's a bit intresting that with Alder Lake, video encoding is one of the areas that has a performance uplift when comparing ddr4 to ddr5.

https://images.anandtech.com/graphs/graph17047/126998.png
https://tpucdn.com/review/intel-core-i9-12900k-alder-lake-ddr4-vs-ddr5/images/premiere-pro.png
https://cdn.sweclockers.com/artikel/diagram/24897?key=6802cbedbbdb26b1dfdfb39297ec46a8

But here for techpowerups x265 test there isnt much difference

https://tpucdn.com/review/intel-core-i9-12900k-alder-lake-ddr4-vs-ddr5/images/encode-h265.png

I suspect that the faster you encode the more impact it will have, so when doing sub real time encodes there isnt enough traffic to RAM to make a difference.

And on the Alderlake subject, 12700k is quite a bit faster than 5800X in x265 it seems, and rather close to 5900X. Looks like a reasonable option for x265.

tormento
11th November 2021, 12:10
On that subject, it's a bit intresting that with Alder Lake, video encoding is one of the areas that has a performance uplift when comparing ddr4 to ddr5.
I'd really like to see x265 performance with disabled E-cores, thus activating AVX-512.

RanmaCanada
11th November 2021, 14:56
On that subject, it's a bit intresting that with Alder Lake, video encoding is one of the areas that has a performance uplift when comparing ddr4 to ddr5.

https://images.anandtech.com/graphs/graph17047/126998.png
https://tpucdn.com/review/intel-core-i9-12900k-alder-lake-ddr4-vs-ddr5/images/premiere-pro.png
https://cdn.sweclockers.com/artikel/diagram/24897?key=6802cbedbbdb26b1dfdfb39297ec46a8

But here for techpowerups x265 test there isnt much difference

https://tpucdn.com/review/intel-core-i9-12900k-alder-lake-ddr4-vs-ddr5/images/encode-h265.png

I suspect that the faster you encode the more impact it will have, so when doing sub real time encodes there isnt enough traffic to RAM to make a difference.

And on the Alderlake subject, 12700k is quite a bit faster than 5800X in x265 it seems, and rather close to 5900X. Looks like a reasonable option for x265.

Techpowerup uses the slow preset where as everyone else usually uses fast or worse, because "bigger numbers are better!" Anandtech has also been garbage ever since they sold out to Intel, when they were bought by Purch (Intel) back in the day (2014).

excellentswordfight
11th November 2021, 15:22
I'd really like to see x265 performance with disabled E-cores, thus activating AVX-512.
I would strongly assume that you would loose performance, by a lot. You are disabling eight cores with similar IPC as skylake @ 3,7GHz in a load that actually uses those. For what? A few percent gain from avx512?

All xeons I've tested have had negative performance when using avx512 just cause of a few hundreds megahertz downclock.

Techpowerup uses the slow preset where as everyone else usually uses fast or worse, because "bigger numbers are better!"
I'm aware, hence my remark of encoding speed.

Could be absolutely valid testing faster encoders/presets, depends what you wanna test. In this case I found it rather nice, as it demonstrates that very fast encoders can benefit from higher memory bandwidth.

Anandtech has also been garbage ever since they sold out to Intel, when they were bought by Purch (Intel) back in the day (2014).
Ian still writes the best in-depth articles of any media outlet though. Not that this has anything to do with the subject.

tormento
11th November 2021, 15:24
You are disabling eight cores with similar IPC as skylake @ 3,7GHz in a load that actually uses those.
We don't know exactly how Alder Lake behaves with high and continuous loads. If it uses P-cores only (as I suspect), perhaps some AVX-512 could be useful.

excellentswordfight
11th November 2021, 17:37
We don't know exactly how Alder Lake behaves with high and continuous loads. If it uses P-cores only (as I suspect), perhaps some AVX-512 could be useful.
The techpowerup test actually had tests with e-cores disabled, and it makes ut quite a bit slower, I doubt avx512 will make up for a 30% performance loss.

https://tpucdn.com/review/intel-core-i9-12900k-alder-lake-12th-gen/images/encode-h265.png

Selur
12th November 2021, 22:58
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| CPU | x264 | x265 | LAVC | auto | MMX2 | SSE | SSE2 | SSE3 | SSE4 | AVX | AVX2 | All |
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| Ryzen 9 3950X 16-Core Processor | 44.76 | 7.77 | 118 | 5.67 | 2.46 | 2.47 | 3.61 | 3.86 | 5.01 | 5.01 | 5.58 | N/A |

Cu Selur

Loomes
12th November 2021, 23:05
The Ryzen 9 3950X has 25% more cores (16) than the Ryzen 9 3900X (12) and is only about 10% faster. Hm, what to make of that regarding to x265's multithreading?

Selur
13th November 2021, 07:36
My guess is that the higher the core count you will need a higher resolution content to saturate the cores fully.
Also I'm on Windows 11 so that might play into it.

Boulder
13th November 2021, 08:33
My guess is that the higher the core count you will need a higher resolution content to saturate the cores fully.

This is the most likely reason. With a 1080p source and the default CTU 64, the cores might not get fully utilized. With a 4K source, it's a very different story (or by switching to CTU size 32).

kolak
20th November 2021, 21:19
My guess is that the higher the core count you will need a higher resolution content to saturate the cores fully.
Also I'm on Windows 11 so that might play into it.

I noticed that scaling is faaaar from linear with higher cores number. At some point it's basically pointless to try to use so more cores even for UHD. You should rather switch into some chunked encoding and use eg. 16 cores per chunk.

Loomes
21st November 2021, 13:16
You should rather switch into some chunked encoding and use eg. 16 cores per chunk.
Sounds interesting. Could you please give a bit more information / sources about chunk encoding? I get the idea but how to approach this technically?

Atak_Snajpera
21st November 2021, 18:12
Sounds interesting. Could you please give a bit more information / sources about chunk encoding? I get the idea but how to approach this technically?

It works like this
https://i.postimg.cc/Vk7yzLZN/client.png

You basically divide your video clip virtually in avisynth script with Trim function and then encode all chunks at once. Then you combine all .265 files into one and mux. With this approach i can easily saturate even Threadripper 64C/128T using just 720p source.

kolak
21st November 2021, 21:41
There are problems with it as well: VBV buffer flow, VBR efficiency with very short chunks, but in most cases it works fine. You also should use stitching option in encoders and make sure headers are set correctly, specially for the last chunk.
SPS/PPS should also be set properly (for ts muxing) as some hardware boxes don't like if those change during one video.
You should also put some attention where you place chunk points, eg. best are scene changes. In case of streaming need of fixed I frame distance chunks points should be rather aligned to them.
For perfect streams it's not as easy as simple divide to n parts and encode them simultaneously, but it can be done. It all depends what is the end target.

sebastian1
24th November 2021, 17:35
5900X@PBO/+50Mhz/142W
||---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| CPU ....................| x264 | x265 | LAVC | auto | MMX2 | SSE | SSE2 | SSE3 | SSE4 | AVX | AVX2 | All |
|----------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| Ryzen 9 5900X......| 48.29 | 8.88 | 224 | 6.94 | 2.39 | 2.40 | 3.67 | 3.93 | 5.66 | 5.72 | 6.86 | N/A |

tormento
26th November 2021, 19:28
Yet nobody with Alder Lake? :)

RanmaCanada
26th November 2021, 21:33
Yet nobody with Alder Lake? :)

We would honestly need to get the benchmark updated for anything in Alder Lake to be relevant.

GEfS
1st January 2022, 06:17
R5 2600@4.1ghz 1.35V
2666mhz CL16-18-18-38 RAM.


|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| CPU | x264 | x265 | LAVC | auto | MMX2 | SSE | SSE2 | SSE3 | SSE4 | AVX | AVX2 | All |
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| Ryzen 5 2600 | 16.42 | 2.75 | 97 | 2.26 | 0.90 | 0.84 | 1.30 | 1.44 | 2.09 | 2.12 | 2.18 | N/A |

Emulgator
2nd January 2022, 17:40
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| CPU | x264 | x265 | LAVC | auto | MMX2 | SSE | SSE2 | SSE3 | SSE4 | AVX | AVX2 | All |
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| 11th Gen Core i9-11900K | 29.92 | 5.16 | 169 | 4.07 | 1.34 | 1.29 | 2.09 | 2.23 | 2.92 | 3.51 | 4.05 | N/A |

Corei9-11900K, Notebook-MB Z590
No CPU/GPU OC, Busy Clock was between 3,6 and 4,5GHz, idling around 5,1..5,3 GHz,
3200MHz RAM underclocked to 2933 (CR 2T 1467MHz 21-21-21-47)
Consumption depending on passes
CPU cores ate min.80W..max.145W
CPU Package ate min.90W..max.165W
Total System ate min.150W..max.262W

GEfS
6th January 2022, 08:44
i5-12600@stock with ID-Cooling SE-207-XT on Gigabyte Z690 Gaming X DDR4
2133mhz RAM as Corsair stock (3600C18 kit, Micron E-die) and I forgot to turn on XMP. And I don't have time for a second run cuz the machine gonna be shipped real soon.
Benchmark suite running on a NVMe Pcie 4.0 Samsung PM9A1 512GB

|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| CPU | x264 | x265 | LAVC | auto | MMX2 | SSE | SSE2 | SSE3 | SSE4 | AVX | AVX2 | All |
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| Core i5-12600K | 33.75 | 6.34 | 163 | 4.69 | 1.62 | 1.54 | 2.49 | 2.77 | 3.91 | 4.10 | 4.55 | N/A |


Will update with a proper setup (with at least 3200mhz RAM) whenever I get in touch on Alder Lake system.

rwill
7th January 2022, 18:48
|----------------------------------------------|---------|----------|---------|---------|---------|----------|---------|---------|---------|---------|---------|---------|
| CPU | x264 | x265 | LAVC | auto | MMX2 | SSE | SSE2 | SSE3 | SSE4 | AVX | AVX2 | All |
|----------------------------------------------|---------|----------|---------|---------|---------|----------|---------|---------|---------|---------|---------|---------|
| Ryzen Threadripper 3970X 32-Core Processor | 57.04 | 15.53 | 209 | 9.92 | 3.64 | 3.64 | 5.37 | 5.78 | 8.34 | 8.49 | 10.14 | N/A |

CPU utilization wasnt even close to 99% though.

rwill
7th January 2022, 18:53
By the way:

https://www.igorslab.de/en/intel-deactivated-avx-512-on-alder-lake-but-fully-questionable-interpretation-of-efficiency-news-editorial/

benwaggoner
7th January 2022, 23:18
By the way:

https://www.igorslab.de/en/intel-deactivated-avx-512-on-alder-lake-but-fully-questionable-interpretation-of-efficiency-news-editorial/
The thermal throttling caused by the current AVX512 implementation tends to counteract the potential speed gains in going from 256 to 512 bit SIMD. With x265, turning off AVX512 yields faster encoding below something like 4K at --preset veryslow.

I can see Intel very much wanting to keep AVX512 out of consumer systems where it would actually cause performance regressions most of the time. It's great news if this means they'll provide a better, cooler AVX512 soon that could be of more practical use.

AVX2 had a similar course, where the thermal throttling of the initial implementation made it only a small boost, but a later revision had better thermal characteristics and made AVX2 a lot more valuable to use.

nevcairiel
7th January 2022, 23:40
The thermal throttling caused by the current AVX512 implementation tends to counteract the potential speed gains in going from 256 to 512 bit SIMD.

Did you actually test on Alder Lake, and not some previous implementation?
Because between the new process node as well as general improvements, the downsides should be much reduced.

tonemapped
10th January 2022, 07:48
The thermal throttling caused by the current AVX512 implementation tends to counteract the potential speed gains in going from 256 to 512 bit SIMD. With x265, turning off AVX512 yields faster encoding below something like 4K at --preset veryslow.

I can see Intel very much wanting to keep AVX512 out of consumer systems where it would actually cause performance regressions most of the time. It's great news if this means they'll provide a better, cooler AVX512 soon that could be of more practical use.

AVX2 had a similar course, where the thermal throttling of the initial implementation made it only a small boost, but a later revision had better thermal characteristics and made AVX2 a lot more valuable to use.

The problem with Intel's f**kery with 'accidentally' leaving AVX512 not fused off, therefore allowing motherboard manufacturers to enable in via BIOS (which Intel would have noticed before launch), is that it gave an estimated uplift during the initial review cycle (first week) of ~9% in many applications/benchmarks. I don't see that as a coincidence.

On top of that, I know Intel added AVX2 support to the Alder Lake's Atom cores, sorry I mean "efficiency cores", so that people could utilise both the proper cores and the Atom cores, which makes total sense, but Intel stated about six months before 12 Gen. released that AVX512 would be fused off as the Atom cores can't support it.

Then there's the power issue with disabling AVX512. The Alder Lake 12900K already uses ~80% more power for ~21% more performance, when compared to a CPU with fewer cores but the same threads (5900X), and it would appear disabling AVX512 increases power use.


IgorsLab conducted some tests on the 12900K with power measured via the EPS 12V, so a reliable reading, with the following configurations:

12900K 8+8(24) / no AVX512 due to Atom cores
12900K (P-core only) / AVX512 enabled
12900K (P-core only) / AVX512 disabled


Now the interesting part is the power consumption:

12900K 8+8(24) / no AVX512 due to Atom cores: 255W
12900K (P-core only) / AVX512 enabled: 307W
12900K (P-core only) / AVX512 disabled: 328W


I'd really hoped Alder Lake would be more power efficient, but it seems Intel's taken Nvidia's approach and given up with power efficiency. My 5900X can encode two instances of x265 -slower @ ~10.6 fps (1.14V, 139W, 4.35 GHz). The 12900K is certainly impressive, but I don't see how people can justify ~255W for 21% more performance.

And in case people think I'm noting the above because I have a 5900X - the same applies to the 3080 (which I have) vs the 2080 Ti. The former offers ~17% more performance for an additional ~120W. It's obscene.

RanmaCanada
11th January 2022, 20:50
One could say Intel left AVX512 in on purpose to give idiots the false sense that their chips are superior. Now that it is being taken away, all that people will remember is that when the chips were launched, they were faster, and they won't notice, or care that their chips are now actually slower, and use far more power.

Intel breaking the law again with false advertising? Nah, they would never break the law, ever.

benwaggoner
11th January 2022, 23:22
One could say Intel left AVX512 in on purpose to give idiots the false sense that their chips are superior. Now that it is being taken away, all that people will remember is that when the chips were launched, they were faster, and they won't notice, or care that their chips are now actually slower, and use far more power.

The thing is, customer aren't going to see any actual real-world regressions without AVX512 with the implementations to date, outside of maybe some specific scientific computing applications, or 4K preset veryslow x265 encoding.

nevcairiel
12th January 2022, 00:11
Intel breaking the law again with false advertising? Nah, they would never break the law, ever.

Intel never advertised AVX512 for Alder lake, and even directly said to media that AVX512 is (supposed to be) disabled on those chips. Any conclusions drawn otherwise are entirely on the side of the media.

The vast majority of benchmarks or real-world use-cases won't even benefit from AVX512, and the fact that you have to disable the E-cores to make use of it also would only ever show any uplift on extremely heavy AVX512 work-loads, which would be capable of offsetting the reduction in cores.
For x265 for example, I would blindly claim that enabling the E-cores is extremely likely to be faster then making use of AVX512.

RanmaCanada
12th January 2022, 16:25
Intel never advertised AVX512 for Alder lake, and even directly said to media that AVX512 is (supposed to be) disabled on those chips. Any conclusions drawn otherwise are entirely on the side of the media.

The vast majority of benchmarks or real-world use-cases won't even benefit from AVX512, and the fact that you have to disable the E-cores to make use of it also would only ever show any uplift on extremely heavy AVX512 work-loads, which would be capable of offsetting the reduction in cores.
For x265 for example, I would blindly claim that enabling the E-cores is extremely likely to be faster then making use of AVX512.

https://videocardz.com/newz/intel-confirms-alder-lake-p-h-mobile-specs-publishes-hybrid-architecture-optimization-guide-for-developers

I hate to use the site, but Intel released documentation that stated all you would need to do was disable the E-Cores to enable AVX512. Intel then back peddled and claimed their official documentation was "wrong".

Atak_Snajpera
12th January 2022, 16:48
Intel never advertised AVX512 for Alder lake, and even directly said to media that AVX512 is (supposed to be) disabled on those chips. Any conclusions drawn otherwise are entirely on the side of the media.

The vast majority of benchmarks or real-world use-cases won't even benefit from AVX512, and the fact that you have to disable the E-cores to make use of it also would only ever show any uplift on extremely heavy AVX512 work-loads, which would be capable of offsetting the reduction in cores.
For x265 for example, I would blindly claim that enabling the E-cores is extremely likely to be faster then making use of AVX512.

E-Cores are very weak in floating point calculations. If you take into account lower clock then 4 E-Cores will be equal to 1 P-Core.
https://i.postimg.cc/GdGk8kQj/Flops-CPUv2.png

Locked clocks at 4 GHz (P-CORE) and 3 GHz (E-CORE) for testing purposes.
https://abload.de/img/flops61ojmu.png

nevcairiel
12th January 2022, 17:14
E-Cores are very weak in floating point calculations. If you take into account lower clock then 4 E-Cores will be equal to 1 P-Core.

Luckily video encoding doesn't use much floating point, if any at all. Its all integer math.

Its also all just theory, a proper comparison would be the only interesting part. Even if your assumption is right, there is 8 E cores on the big CPUs, if they make up 25% extra performance (eg. 2 extra cores), thats likely going to beat or match AVX512.

benwaggoner
12th January 2022, 17:16
Wow, I was not expecting a big jump in x87 performance this decade!

Atak_Snajpera
12th January 2022, 17:37
Wow, I was not expecting a big jump in x87 performance this decade!

Intel CPUs will again crush AMD CPUs in QUAKE 2 game (software mode)!

rwill
12th January 2022, 18:13
Intel CPUs will again crush AMD CPUs in QUAKE 2 game (software mode)!

You are joking but there still seem to be some middleware libraries for games that are using x87. I think the last prominent one that got called out was in Skyrim.

Atak_Snajpera
13th January 2022, 16:49
Luckily video encoding doesn't use much floating point, if any at all. Its all integer math.

Its also all just theory, a proper comparison would be the only interesting part. Even if your assumption is right, there is 8 E cores on the big CPUs, if they make up 25% extra performance (eg. 2 extra cores), thats likely going to beat or match AVX512.

AVX-512 supports integer as well
https://en.wikipedia.org/wiki/AVX-512
Introduced with Cannon Lake.[4]

AVX-512 Integer Fused Multiply Add (IFMA) - fused multiply add of integers using 52-bit precision.

GEfS
17th January 2022, 08:05
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| CPU | x264 | x265 | LAVC | auto | MMX2 | SSE | SSE2 | SSE3 | SSE4 | AVX | AVX2 | All |
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| Core i7-11700 | 28.40 | 5.55 | 168 | 4.39 | 1.36 | 1.37 | 2.14 | 2.32 | 3.18 | 3.10 | 3.97 | N/A |

B560M Aorus Elite, i7 11700, Tower cooler rated at 180W.
4.4 all core while full load. (near 90C, ambient 20C)
RAM 2x8GB bus 3200 C16-20
SSD Seagate Q5 500GB while running the test.

i5-12600@stock with ID-Cooling SE-207-XT on Gigabyte Z690 Gaming X DDR4
2133mhz RAM as Corsair stock (3600C18 kit, Micron E-die) and I forgot to turn on XMP. And I don't have time for a second run cuz the machine gonna be shipped real soon.
Benchmark suite running on a NVMe Pcie 4.0 Samsung PM9A1 512GB

|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| CPU | x264 | x265 | LAVC | auto | MMX2 | SSE | SSE2 | SSE3 | SSE4 | AVX | AVX2 | All |
|---------------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| Core i5-12600K | 33.75 | 6.34 | 163 | 4.69 | 1.62 | 1.54 | 2.49 | 2.77 | 3.91 | 4.10 | 4.55 | N/A |


Will update with a proper setup (with at least 3200mhz RAM) whenever I get in touch on Alder Lake system.

And even the worst setup for 12600K Alder Lake is still better than i7 11700.