View Full Version : Sharing Ryzen 1700 vs i7-6700 result


cojj
10th March 2017, 10:21
I7 rig - all stock setting
R7 rig - o.c'd to 3.7ghz + running at 1.24v (under-volted) + 24c idle/70c after 2days of non-stop 100% usage, cheap evo 212 air cooled) + $650nzd ~= $450usd for cpu+mobo

Encoded 720p and 1080p video with veryslow preset + 19.9 crf.

x265 [info]: HEVC encoder version 2.3+17-6e348252e902
x265 [info]: build info [Windows][GCC 6.3.0][64 bit] 10bit
x265 [info]: using cpu capabilities: MMX2 SSE2Fast SSSE3 SSE4.2 AVX AVX2 FMA3 LZCNT BMI2
x265 [info]: Main 10 profile, Level-4 (Main tier)
x265 [info]: Thread pool created using 8 threads
x265 [info]: Slices : 1
x265 [info]: frame threads / pool features : 3 / wpp(17 rows)
x265 [info]: Coding QT: max CU size, min CU size : 64 / 8
x265 [info]: Residual QT: max TU size, max depth : 32 / 3 inter / 3 intra
x265 [info]: ME / range / subpel / merge : star / 57 / 4 / 4
x265 [info]: Keyframe min / max / scenecut / bias: 23 / 250 / 40 / 5.00
x265 [info]: Lookahead / bframes / badapt : 40 / 8 / 2
x265 [info]: b-pyramid / weightp / weightb : 1 / 1 / 1
x265 [info]: References / ref-limit cu / depth : 5 / off / on
x265 [info]: AQ: mode / str / qg-size / cu-tree : 1 / 1.0 / 32 / 1
x265 [info]: Rate Control / qCompress : CRF-19.9 / 0.60
x265 [info]: VBV/HRD buffer / max-rate / init : 12000 / 12000 / 0.900
x265 [info]: tools: rect amp limit-modes rd=6 psy-rd=2.00 rdoq=2 psy-rdoq=1.00
x265 [info]: tools: rskip signhide tmvp b-intra strong-intra-smoothing deblock
x265 [info]: tools: sao


1080p:
i7: encoded 2208 frames in 2345.17s (0.94 fps), 3186.52 kb/s, Avg QP:22.74
r7: encoded 2208 frames in 1479.63s (1.49 fps), 3186.52 kb/s, Avg QP:22.74

720p:
i7: encoded 2208 frames in 1209.50s (1.83 fps), 1404.07 kb/s, Avg QP:22.15
r7: encoded 2208 frames in 859.33s (2.57 fps), 1404.45 kb/s, Avg QP:22.15


Attached are screenshots of cpu usage when encoding 1080.

ChaosKing
10th March 2017, 10:39
Could you also add a second ryzen test with smt disabled? It should be a bit faster.

Atak_Snajpera
10th March 2017, 10:55
With this you will have 100% cpu utilisation
http://forum.pclab.pl/topic/1184884-x265-FHD-Benchmark/

CruNcher
10th March 2017, 12:07
Interesting seeing that most reviews compare vs i7-7700

though Intels 6 core mainstream answer is already ready for departure if rumors are correct and that should obviously reach those values and go slightly above even @ 6C/12T, but yeah for Mainstream still insane 3 FPS even 4 or 5 would be unusable for AVG users :D

but still nice how close AMD gets now pretty unoptimized :)

cant wait the results with Ataks suggestion utilizing the threads more efficient on both and how that looks then under both 100% utilization target :)

i guess you'll endup fully utilized somewhere at the I7-6600K result in that Benchmark database

cojj
10th March 2017, 20:47
As requested - from this result seems like ryzen is ~10% slower than Haswell-e/ ~22% faster than Sandy Brdge, clock to clock/core to core/thread to thread . I was getting near full utilization (occasional small dips)

Interesting seeing that most reviews compare vs i7-7700
Maybe it's because they are in the same price range. As for me, I only have mutiple i7-6700 and a r7-1700 xD

Atak_Snajpera
10th March 2017, 21:54
Run also this http://forum.pclab.pl/topic/1105978-FlopsCPU-klasyczny-benchmark-z-1992-roku-w-nowej-oprawie/ on this overclocked Ryzen.

cojj
10th March 2017, 22:09
Run also this http://forum.pclab.pl/topic/1105978-FlopsCPU-klasyczny-benchmark-z-1992-roku-w-nowej-oprawie/ on this overclocked Ryzen.

Here you go. Please evaluate these numbers and share your conclusion :)

CruNcher
10th March 2017, 23:30
The new Intel 6 Core Mainstream (6C/12T) gonna smoke it, now Intel just has to set the price right ;)

But nothing really has changed Intel is still 1 Generation ahead and smokes AMDs Core architecture with less core count and same 14 nm target ;)

But yeah if AMD would never made it we would never have seen Intel adapting prices to it but surely the 6C/12T will still be a small overall amount more expensive when it hits soon.

Thus why Raven Ridge will be so interesting Intels GPU is still not on the same level overall to compete on the level like AMDs CPU now competes with it :)

The Time was overall not enough for Intel to advance it even if they made very huge steps on the GPU side as well, sponsored mainly by their shrinking time to market advantage that slowly vanishes thx to AMDs corporation with Samsungs Semiconductor part.

microchip8
10th March 2017, 23:38
The new Intel 6 Core Mainstream (6C/12T) gonna smoke it, now Intel just has to set the price right ;)

Got something to back that up?

CruNcher
10th March 2017, 23:58
Wait and see ;)

You could also look at those benchmarks and understand that Intel is in that Workload more advanced with it's instruction currently and that they just hold the price high till now ;)

Bloax
11th March 2017, 00:53
Could you also add a second ryzen test with smt disabled? It should be a bit faster.

Disabling SMT only helps in limited threading scenarios (e.g. games), not things like video encoding which utilize all available threads.

burfadel
11th March 2017, 01:02
The new Intel 6 Core Mainstream (6C/12T) gonna smoke it, now Intel just has to set the price right ;)

It's gonna smoke? So it will have thermal issues? :sly:

Got something to back that up?

The 6-core Coffee Lake CPU, which is a modified Kaby Lake, is expected to be released in February 2018. Its release roughly coincides with the second generation Zen, known as Zen+. Sure, you can compare the performance of 6-core Coffee Lake to Zen, but you might as well compare Zen to Haswell and not Skylake.

Also keep in mind we might not have seen exactly how well Ryzen performs yet, seeing as there's the scheduler bug, a performance processor driver on the way (allegedly), bios optimisations, as well as optimisations for specific titles.

cojj
11th March 2017, 01:28
Run also this http://forum.pclab.pl/topic/1105978-FlopsCPU-klasyczny-benchmark-z-1992-roku-w-nowej-oprawie/ on this overclocked Ryzen.

Here you go. Would appriciate if you could sharing your evaluation on these results.

microchip8
11th March 2017, 01:32
It's gonna smoke? So it will have thermal issues? :sly:



The 6-core Coffee Lake CPU, which is a modified Kaby Lake, is expected to be released in February 2018. Its release roughly coincides with the second generation Zen, known as Zen+. Sure, you can compare the performance of 6-core Coffee Lake to Zen, but you might as well compare Zen to Haswell and not Skylake.

Also keep in mind we might not have seen exactly how well Ryzen performs yet, seeing as there's the scheduler bug, a performance processor driver on the way (allegedly), bios optimisations, as well as optimisations for specific titles.

How much of an improvement will Coffee Lake be over Kaby Lake? If it's the same (read: virtually nothing) as Skylake -> Kaby Lake, then I doubt it will gone "smoke" anything

cojj
11th March 2017, 01:57
Hello everyone. Let's stay on topic and try not to sepculate the future - maybe we can do that on another thread.

To help with the discussion => Currrently (2017-03-11), with above results, is r7-1700 the best processor to get for pure x265 encoding machine? I think its pretty decent perf/$?


Run also this http://forum.pclab.pl/topic/1105978-FlopsCPU-klasyczny-benchmark-z-1992-roku-w-nowej-oprawie/ on this overclocked Ryzen.

I uploaded the result but doom9 forum says it will post it after a moderator reviews it. In the meantime, a txt version of the result is as follows:

test/single/muti/ratio
x86/52.1M/593M/11.4
x87/3.34G/28.3G/8.5
SSE2/8.87G/87.2G/9.8
AVX/11.1G/111G/10.0
AVX2/18.0G/153G/8.5

@Atak_Snajpera => I would appriciate your evaluation of these results as I find your knowledge very insightful.

Bloax
11th March 2017, 04:19
The r7-1700 is the best processor in terms of price/high-performance to get for video encoding at the moment if you manually overclock it to 3.7-3.9 Ghz, yes.

hajj_3
11th March 2017, 09:26
But nothing really has changed Intel is still 1 Generation ahead and smokes AMDs Core architecture with less core count and same 14 nm target ;)

Intel's 14nm FF+ is WAY better than the 14nm that AMD is using, it is more like 12nm in comparison. Therefore intel has a power and temperature advantage over AMD. If AMD has the same process they could increase clock speeds to make their chips more competitive in games.

Intel's motherboards for their 6, 8 and 10 core processors are far more expensive than amd's boards. Intel still hasn't reduced the prices for their chips yet either so i can't see intel's 6 core being a best seller due to the cost.

p.s i hope someone can compare 2400mhz ram vs 3200mhz ram using ryzen and kaby lake with x264 and x265 so that we can see what effect it has performance on intel and amd as previous amd architectures gave significant performance increases with faster memory, much more so than intel.

Atak_Snajpera
11th March 2017, 12:01
test/single/muti/ratio
x86/52.1M/593M/11.4
x87/3.34G/28.3G/8.5
SSE2/8.87G/87.2G/9.8
AVX/11.1G/111G/10.0
AVX2/18.0G/153G/8.5

@Atak_Snajpera => I would appriciate your evaluation of these results as I find your knowledge very insightful.
Can you post screenshot as well? Upload images to http://imgsafe.org/ or https://cubeupload.com/

Updated Flops table
http://i.imgsafe.org/41ba1e2822.png

for comparison KabyLake@5.1GHz
http://fs5.directupload.net/images/170211/doyk2oew.jpg

NikosD
11th March 2017, 18:18
R7 rig - o.c'd to 3.7ghz + running at 1.24v (under-volted) + 24c idle/70c after 2days of non-stop 100% usage, cheap evo 212 air cooled) + $650nzd ~= $450usd for cpu+mobo


Hey,

Is it possible to run both apps and post your results - FlopsCPU (single core) and x265 benchmark (GUI) - but by disabling from your BIOS:

a) Precision Turbo
b) SMT ("hyper-threading")
c) XFR

also disable 6 cores (if possible) in order to have only 2 active and make sure your CPU runs at 3.0GHz and stays there.

So, you will end up to a RyZen 2C/2T@3.0GHz in order to directly compare it with Sandy and Haswell.

Here is my updated post:
https://forum.doom9.org/showthread.php?p=1800292#post1800292

It would be extremely useful if you keep the log files (txt) created by FlopsCPU, in order to further analyze the FPU of RyZen and upload them too as a zip file.

ShogoXT
12th March 2017, 06:32
Did this in a hurry. OCed to 3.6ghz, but didnt reformat yet, so its not optimal. Cant disable cores, made sure power is on high performance mode. Dont have my noctua yet, so will have to OC more later.
http://i.imgur.com/TpqPXcr.jpg
http://i.imgur.com/HNdZO5z.jpg
http://i.imgur.com/EZlIW5K.jpg

NikosD
12th March 2017, 06:43
There is a huge fluctuation in your results regarding to multicore performance of FlopsCPU, which is a secure indicator that precision turbo and XFR were having a party.

In order to compare small differences you have to at least disable those two, if disabling cores is impossible, in order to keep your clock stable during tests.

ShogoXT
12th March 2017, 07:23
Actually I had ran that with discord open after streaming on youtube with Ryzen, playing World of Warcraft. The FPU program actually froze a couple times getting those results. If I ever get around to reformatting it should clear up. Pretty sure XFR and Turbo disable automatically when overclocking. At least one of those does. XFR only increases by 50mhz.

NikosD
12th March 2017, 07:38
test/single/muti/ratio
x86/52.1M/593M/11.4
x87/3.34G/28.3G/8.5
SSE2/8.87G/87.2G/9.8
AVX/11.1G/111G/10.0
AVX2/18.0G/153G/8.5


There is an issue here.

It seems that optimized Intel FP SIMD SSE2/AVX/AVX2-FMA3 code has a weird behavior running on AMD RyZen.

The logical steps for RyZen's SIMD HW according to Intel's hardware and compiler optimizations should be:

Ryzen SSE2 about 2.3x faster than x87

Ryzen AVX about 2x faster than SSE2

Ryzen AVX2-FMA3 about unknown times faster than AVX

We now get:

RyZen SSE2 about 2.7x faster than x87 (close , a little faster)

RyZen AVX about 25% faster than SSE2 (extremely slow)

RyZen AVX2-FMA3 about 62% faster than AVX (too fast)

RyZen's SSE2 performance is beyond Intel's 2.27-2.35x jump over scalar x87, it's around 2.65x, but they are close.

The first step could mean not so optimized scalar code x87 for RyZen, slower x87 performance of RyZen or faster SSE2 implementation of RyZen vs Intel.

But the next steps are very difficult to explain.

From SSE2 to AVX, Intel gains ~100% performance, RyZen gains 25%

This is completely unexplained, even more than low DX12 performance in gaming.

RyZen seems to gain 100% over SSE2, using AVX2-FMA3 and 62% over AVX.

Completely unpredictable behavior, possibly due to Intel compiler optimizations regarding to RyZen's HW SIMD implementation.

I don't know if there is any difference running that test using Win 7.

It shouldn't for Single Core results.

The above results explain the very low performance running AVX single precision in Tom's hardware scientific tests here:
http://www.tomshardware.com/reviews/amd-ryzen-7-1800x-cpu,4951-10.html

RyZen when running Intel optimized FP SIMD code, gains a lot leveraging AVX2-FMA3 code compared to SSE2 and AVX and looses a lot leveraging AVX code compared to Intel.

NikosD
12th March 2017, 11:32
OK...

I might have an explanation of that almost 1/2 performance of Ryzen compared to Intel for AVX1 (FP32/FP64) which can be seen as only 25% gain over SSE2 code for Ryzen.

The Intel compiler probably expects only 2 execution units for AVX, because Intel's HW has only 2.

So, it gives 1x256bit FADD AVX instruction to 1x128bit FADD AVX instead of 2x128bit FADD AVX that Ryzen has, because it thinks there is only 1 FADD inside.

So instead of splitting the 256bit instruction to 2x128bit registers in order to work simultaneously, now the 128bit register has to process two times the same 256bit instruction.

The same of course goes for FMUL, but I think the compiler can give two FMULs at the same time due to Intel's HW capability, so that's why we don't have exactly half performance but a little lesser.

That would explain the performance loss and I think is feasible for Intel to change/patch that behavior of its compiler or for AMD to do that automatically maybe with a microcode update (?).

I'm not sure about that.

These are wild guesses of course, for that huge surprise of AVX1 FP32/FP64 performance.

Atak_Snajpera
12th March 2017, 13:33
I always thought that AMD cpus are smart enough to automatically split 1x256 instruction into 2x128 bit. It looks like since AMD FX nothing has been changed in this matter so I wouldn't wait for some magic microcode update.

Sagittaire
12th March 2017, 20:32
well I think that x265fhd benchmark can't work really good with Rysen because x265fhd use multiple instance and Rysen have specific limitation in L3 cache.

Moreover, 5 instance for x265 in same time is not reallistic situation: nobody make encoding like that, and particulary with x265.

@cojj

use simply "--pme" command to optimize thread charge for rysen in 720p and 1080p

If you want really compare speed at CPU charge 100% for Rysen vs Intel, use simply 4K source with --pme command.

NikosD
12th March 2017, 20:40
We have results from two users only, which are very contradicting because although they have small differences in clock, the performance difference is big.

The users are not professional reviewers or IT professionals and we don't know the exact conditions of systems tested.

x265 FHD benchmark scales better than any other I have seen, but maybe we need 4K samples for better scaling.

Atak_Snajpera
12th March 2017, 20:42
Well I've already told you that for 16T cpu you must use atleast two instances of x265. Your magic commands won't change anything. I've tested mutiple switches and always 2 instances were overall faster. I think you are expecting too much from AMD's 8 core "SandyBridge+" cpu. ALU is only ~13% faster than in K10 architecture from 2010.

NikosD
12th March 2017, 20:47
RyZen is not Sandy+.

Its IPC starts from just little lower than Sandy and goes up to Kabylake for some workloads.

RyZen R7 is the best 8C/16T cpu you can buy nowadays and will stay that way for a long time.

Sagittaire
12th March 2017, 20:52
Well I've already told you that for 16T cpu you must use atleast two instances of x265. Your magic commands won't change anything. I've tested mutiple switches and always 2 instances were overall faster. I think you are expecting too much from AMD's 8 core "SandyBridge+" cpu. ALU is only ~13% faster than in K10 architecture from 2010.

anyway your x264fhd test don't produce coherent result, it's like that ... ;-)

in cojj test you have:

1080p:
i7: encoded 2208 frames in 2345.17s (0.94 fps), 3186.52 kb/s, Avg QP:22.74 with CPU at 100%
r7: encoded 2208 frames in 1479.63s (1.49 fps), 3186.52 kb/s, Avg QP:22.74 with CPU at 50%

Atak_Snajpera
12th March 2017, 20:57
Results do not lie. x86 is between sandybridge and ivybridge. AVX+FMA is a disaster. Only SSE2 looks good but still not near SkyLake. AMD once again fight with Intel using old method called "MoAr CoRes!".

Sagittaire
12th March 2017, 20:59
Results do not lie. x86 is between sandybridge and ivybridge. AVX+FMA is a disaster. Only SSE2 looks good but still not near SkyLake. AMD once again fight with Intel using old method called "MoAr CoRes!".

it's true, Result do not lie ... ;-)

http://www.hardware.fr/getgraphimg.php?id=446&n=10

and in serious test you have R7 1800X on par with i7-6900K.

Atak_Snajpera
12th March 2017, 21:04
it's true, Result do not lie ... ;-)

http://www.hardware.fr/getgraphimg.php?id=446&n=10

and in serious test you have R7 1800X on par with i7-6900K.

Show cpu usage in those tests and then we will talk. I guarantee that x265 wasn't saturating 10c/20t in 100%. Average usage in 1080p for 16T is around 75%.

NikosD
12th March 2017, 21:05
We have already said too many times, that FP SSE2, AVX and FMA3 have nothing to do with x265 application.

Only integer performance matters here and it's about x64, SSEx and AVX2 integer.

x64 is faster on RyZen, you can't see that from FlopsCPU but from the latencies and thoughput of instructions.

The same is true for SSEx.
RyZen is at least as fast as Kabylake.

The only performance drop is when using integer AVX2, but as you said it has double cores so...

RyZen is faster than the fastest 4C/8T from Intel using x265.

Sagittaire
12th March 2017, 21:08
Show cpu usage in those tests and then we will talk. I guarantee that x265 wasn't saturating 10c/20t in 100%. Average usage in 1080p for 16T is around 75%.

CPU charge with R7 is comparable to i7-6900K. You have the same limitation here (8C/16T).

Anyway it's simple to saturate thread for 8C/16T: use simply 4K source with --pme command. It's in x265 documentation.

Atak_Snajpera
12th March 2017, 21:11
Anyway it's simple to saturate thread for 8C/16T: use simply 4K source with --pme command. It's in x265 documentation.
Too synthetic test for me. Most people work with 1080p sources (Blu-ray discs and so on). Trust me.

RyZen is faster than the fastest 4C/8T from Intel using x265.
Let's be honest difference is not huge.

Sagittaire
12th March 2017, 21:12
Too synthetic test for me. Most people work with 1080p sources (Blu-ray discs and so on). Trust me.

seriousely, multiple instance for x265 is by far more synthetic test ... :rolleyes:

and HEVC codec is the official codec for UHD BluDisk ... not H264.

HEVC is particulary designed for UHD encoding.

NikosD
12th March 2017, 21:19
x265 is an extremely optimized AVX2 integer application.

Haswell using x265 optimized for AVX2 is 70% faster than Sandy at the same clock.

So, it's like a miracle that RyZen is faster on x265 than Kabylake which has faster clocks and full AVX2 integer implementation.

@Atak

Is it possible to upgrade your benchmark to 4K HEVC encoding ?

Atak_Snajpera
12th March 2017, 21:20
seriousely, multiple instance for x265 is by far more synthetic test ...
No it is not. I constantly run two instances on my Xeon using Distributed Encoding.

and HEVC codec is the official codec for UHD BluDisk ... not H264.
See torrent sites and tell me how many 4k rips you have there.

HEVC is particulary designed for UHD encoding.
And so what. Most people use x265 for 1080p blu-ray rips. Maybe you should tell them to stop using x265 in this case?

So, it's like a miracle that RyZen is faster on x265 than Kabylake which has faster clocks and full AVX2 integer implementation.
I can reverse your sentence. It is miracle that despite 8 threads and FMAC256 (more heat) SkyLAke/KAbyLake can overclock to ~5GHz and still compete with double cores.

Sagittaire
12th March 2017, 21:27
Is it possible to upgrade your benchmark to 4K HEVC encoding ?

I have benchmark in preparation with x264 (1080p) and x265 (2160p) encoding.

If you use --pme mode with 1080p source (with higher preset than medium) you have 75% CPU charge with 8C/16T.

Unfortunaly --pmode is not usefull because you desactivate reflist option (usefull for speed in x265).


I can reverse your sentence. It is miracle that despite 8 threads and FMAC256 (more heat) SkyLAke/KAbyLake can overclock to ~5GHz and still compete with double cores.

no, simply because R7 CPU charge is not at 100% but more at 50% with 8C/16T for 1080p source and default x265 setting.

and you have same theoric IPC for CPU 4C/8T at 5.0 Ghz than 8C/16T at 2.5 Ghz (for same achitectural CPU)

CruNcher
12th March 2017, 21:31
clocks are dead threading workload efficiency is the future of chip and core design.

Atak_Snajpera
12th March 2017, 21:32
no, simply because R7 CPU charge is not at 100% but more at 50% with 8C/16T for 1080p source and default x265 setting.
Not in my benchmark...

Sagittaire
12th March 2017, 21:36
Not in my benchmark...

your benchmark don't produce usuable result ... it's like that ... :search:

Atak_Snajpera
12th March 2017, 21:36
clocks are dead threading workload efficiency is the future of chip and core design.

Less cores + more IPC + Higher clocks is always better than
MoAr CoReS + less IPC + Lower clocks

https://i.stack.imgur.com/HTTmC.png

Sagittaire
12th March 2017, 21:41
Less cores + more IPC + Higher clocks is always better than
MoAr CoReS + less IPC + Lower clocks

https://i.stack.imgur.com/HTTmC.png

seem don't work for GPU ... or calculator station.

or even with x264:

http://www.hardware.fr/getgraphimg.php?id=446&n=9

Atak_Snajpera
12th March 2017, 21:42
your benchmark don't produce usuable result ... it's like that ... :search:

Oh really?
RyZen 1700@3.7GHz = 25.5 fps
5960X@4.4GHz = 33.9 fps

If we could overclock RyZen to 4.4GHz we would get
4.4/3.7*25.5=30.3 fps

33.9/30.3=1.12

Haswell architecture in x265 encoding is 12% more efficient than Zen. According to flops results (ALU) haswell is 14% faster than Zen. Pretty accurate don't you think?

Sagittaire
12th March 2017, 21:44
Oh really?
RyZen 1700@3.7GHz = 25.5 fps
5960X@4.4GHz = 33.9 fps

If we could overclock RyZen to 4.4GHz we would get
4.4/3.7*25.5=30.3 fps

33.9/30.3=1.12

Haswell architecture in x265 encoding is 12% more efficient than Zen.

no ... your bench don't work. It's simple to prove that.

Sagittaire
12th March 2017, 21:49
advantage for Broadwell-E is only 8% if you compare at Rysen for one Thread (CPU charge at 100%) at 3 Ghz for x265.

http://www.hardware.fr/getgraphimg.php?id=438&n=1

And for 8C/16T at 100%, Rysen at 3.6 Ghz will be on par with i7-6900K at 3.2 Ghz.

it's clear now ... ?

CruNcher
12th March 2017, 21:50
Less cores + more IPC + Higher clocks is always better than
MoAr CoReS + less IPC + Lower clocks

https://i.stack.imgur.com/HTTmC.png

https://upload.wikimedia.org/wikipedia/commons/d/d7/Gustafson.png

Yes and you ignore Gustafson great

Workloads complexity rises also

Atak_Snajpera
12th March 2017, 21:50
no ... your bench don't work. It's simple to prove that.
Yeah paste again that graph because I've already forgotten how it looks like. Seriously you are trolling now. I'm out...

NikosD
12th March 2017, 21:56
Come on guys...We are all useful here.

Take it easy and not personal.

All the arguments have been put on the table.

Let's take a rest and wait for more and better RyZen results.

Sagittaire
12th March 2017, 21:58
Yeah paste again that graph because I've already forgotten how it looks like. Seriously you are trolling now. I'm out...

well i will post benchmark in 4K with optimized x265 thread profil ...

you will see if I am troll ... :rolleyes:

cojj
12th March 2017, 22:44
Let's embrace intellectual conversation and stay away from slandering/taking things too personally.

From what I read above...
Sagittaire: And for 8C/16T at 100%, Ryzen at 3.6 Ghz will be on par with i7-6900K at 3.2 Ghz. => which means Broadwell is 12.5% more efficient than Zen
Atak_Snajpera: Haswell architecture in x265 encoding is 12% more efficient than Zen

From my understanding, Haswell=>Broadwell has very little performance gain so to me, it seems like your conclusions are the same (within the margin of error).


My Ryzen tests were done in clean-install of windows 10 with minimal processes running in the background (default windows crap only). I did not do any tuning as advised by other people (or benchmarking guides provided by AMD), it was simply out-of-box + slight o.c.

Sagittaire
12th March 2017, 22:54
My Ryzen tests were done in clean-install of windows 10 with minimal processes running in the background (default windows crap only). I did not do any tuning as advised by other people (or benchmarking guides provided by AMD), it was simply out-of-box + slight o.c.

if you want better thread perf, try with --pme command line (particulary good with high quality profil ... ;-)

--pmode command will work too with "veryslow" preset

try "--bframes 3 --frame-thread 8" for better threading

I will post 4K benchmark in few minute if you want compare real perf between i7-6700K and R7 1700

cojj
12th March 2017, 23:00
I will post 4K benchmark in few minute if you want compare real perf between i7-6700K and R7 1700

Looking forward to it :)
Could you still do 1080 benchmark too? That's my primary use-case at the moment.

NikosD
12th March 2017, 23:04
My Ryzen tests were done in clean-install of windows 10 with minimal processes running in the background (default windows crap only). I did not do any tuning as advised by other people (or benchmarking guides provided by AMD), it was simply out-of-box + slight o.c.

Your clock is 2.7% faster but your x265 result is more than 10% faster than the other RyZen.

So, he needs to retest.

BTW, can you put your clock to 3.0GHz and run again with 2 active cores (I'm not sure if you can disable 6 cores)

Can you upload in a server your FlopsCPU logs ?

Thanks.

Sagittaire
12th March 2017, 23:05
Looking forward to it :)
Could you still do 1080 benchmark too? That's my primary use-case at the moment.

will be 1080p for x264 and 2160p for x265.

But you can try 1080p with these command

--pme : work for all preset
--pmode : work for veryslow and placebo preset (else desactive reflist and produce really lower speed for other preset)
--bframes 3 --frame-thread 8: produce better threading for GOP

and try combinaison for these command for 1080p

Sagittaire
12th March 2017, 23:11
and here benchmark:
jfl1974.free.fr/Benchmark.zip

1080p for x264
2160p for x265

try and report your result on 8C/16T CPU:
-speed for x264 and x265
-CPU charge for x264 and x265

ShogoXT
12th March 2017, 23:15
I upgraded from a i7 920 so I dont mind too much on taking risks, but I didnt want another quad core for the third time.

So far I had 6-7 fps on my DVD -> 1080P remaster project on my OCed 920, but now its around 20fps. So im happy. I just hope games dont run too horribly.

I might wait to reformat when the next major windows 10 update comes out. They tend to be a bit messy.

I will also re benchmark and OC more when I get my Noctua installed.

In the meantime livestreaming 1080p 60fps faster profile on OBS gaming works pretty well.

cojj
12th March 2017, 23:17
Your clock is 2.7% faster but your x265 result is more than 10% faster than the other RyZen.

So, he needs to retest.

BTW, can you put your clock to 3.0GHz and run again with 2 active cores (I'm not sure if you can disable 6 cores)

Can you upload in a server your FlopsCPU logs ?

Thanks.

I think the other guy already said he was running some stuff when running the benchmark. source below:

Actually I had ran that with discord open after streaming on youtube with Ryzen, playing World of Warcraft. The FPU program actually froze a couple times getting those results.

I will try to the 2core test when I can. It's currently queued up with tasks which I cannot cancel easily :(

Sagittaire
12th March 2017, 23:20
So im happy. I just hope games dont run too horribly.

Rysen 7 will be between i7-4770K and i7-4790K for game.

If you don't have titan X, it's a good CPU for game in 1080p.

You can't expect really better perf with i7-7700K in 1080p if you have GTX 1060 or RX 480 GPU

CruNcher
13th March 2017, 00:16
Though Benchmarks are mostly done wrong also just FPS based benchmarks are mostly conducted wrong depending on the task to compute and the target to achieve neither FPS are always the correct way to benchmark at all.

Render times and Render Presentation times also play a big role in overall Perception and they even play a bigger role when we talking about VR, latency is not only FPS.

Still conducting FPS based Benchmarks on 3D Engine Realtime Workloads and many others is foolish no developer really judges anything based on the pure FPS counter these "Professional Benchmark Reviews" show mostly still to the Public.

NikosD
13th March 2017, 13:30
and here benchmark:
jfl1974.free.fr/Benchmark.zip

1080p for x264
2160p for x265

try and report your result on 8C/16T CPU:
-speed for x264 and x265
-CPU charge for x264 and x265

OK.

Here are the results of four systems:

Sandybridge Core i5 2400 6MB Cache,
Haswell Core i3 4170 3MB Cache,
Skylake Core i7 6700K 8MB Cache
Kabylake Core i7 7700K 8MB Cache and
RyZen R7 1700 8MB + 8MB Cache *

For all the systems, I tried to find out the IPC of x264 & x265 workload using your settings, eliminating the memory bottleneck.

So, all systems share these features:

a) Win 10 x64 system
b) SMT/ Hyperthreading disabled
c) Turbo disabled (+XFR disabled for RyZen)
d) Underclocked to 3.0GHz
e) Only 2 cores active

* Ryzen results are dual.
One running the apps on one CCX (2+0) and the other result running the apps on two CCXs (1+1)

And here comes the absolute surprise!

The 1+1 system is a little faster than 2+0 (!)

RyZen results have sent to me by Rigaya.

One possible explanation is the access to double L3 cache 16MB (1+1) vs 8MB (2+0)

I used also newer version of x264 (r2762) and x265 (v2.3+18 MS 2017 AVX/AVX2) than those versions included in the zip file to find out possible differences, on my personal systems (Sandybridge & Haswell)


Results x264: (r2744 - Default)


Skylake -2C/2T@3.0GHz -> 3.87 fps

Ryzen (1+1)-2C/2T@3.0GHz -> 3.49 fps

Haswell-2C/2T@3.0GHz -> 3.46 fps

Ryzen (2+0)-2C/2T@3.0GHz -> 3.42 fps

Sandy-2C/2T@3.0GHz -> 2.42 fps




Results x264: (r2762) newer version


Kabylake -2C/2T@3.0GHz -> 3.97 fps

Haswell-2C/2T@3.0GHz -> 3.37 fps

Sandy-2C/2T@3.0GHz -> 2.56 fps


Well, the results for default r2744 version give better fps for Haswell and much better gain at 43% against Sandybridge at the same clock.

Using a newer version r2762 of x264 the gain is smaller 32%

I did an analysis using the --no-asm and --asm SIMD instructions, where SIMD means MMX2, SSE2Fast, SSSE3, SSE4.2, AVX, FMA3, AVX2, LZCNT, BMI2.

The results gave me a 10% faster Haswell using --no-asm and better gains for MMX2 and SSE2Fast than Sandybridge at the same configuration 2C/2T@3.0GHz

AVX2 adds only 6% on Haswell, but there seems to be a little inconsistency in the results of x264 sometimes, because they can give a lot different results when running the same test again.

The larger performance advantage is by far the MMX2 and SSE2Fast instruction sets.

MMX2 gives x3 the performance of --no-asm for Sandy and x3.2 for Haswell.

SSE2Fast gives another 50% for Sandy over MMX2 and 66% for Haswell.

The rest of SIMD instruction sets (SSSE3, SSE4.2, AVX, FMA3, AVX2, LZCNT, BMI2) give from 0% (FMA3, LZCNT, BMI2) on Haswell to 8% (SSSE3) on both Haswell and Sandy.

Using data from the same analysis of RyZen and Kabylake systems, it seems that they follow exactly the same pattern.

RyZen and Kabylake have both a x3.3 MMX2 gain over --no-asm and a 60% gain of SSE2Fast for RyZen over MMX2.
Kaby gains a little more on the same SIMD SSE2fast which is 70%.

SSSE3 is 8% for both and AVX2 is only 4% for Kabylake.

But AVX2 is a little slower for RyZen than AVX and unfortunately LZCNT and BMI2 substract even more performance from RyZen.

So, it's better to add --asm AVX, in order to disable higher SIMD sets, when using RyZen and x264 to maintain maximum performance.

From the results above I can see that Skylake is faster than Haswell ~12% due to better SIMD architecture.

According to this table http://rigaya34589.blog135.fc2.com/blog-entry-916.html?sp it seems that Skylake has 50% faster integer add/sub and 100% integer mul/shifts than Haswell.

Compared to RyZen, some instructions are 4x faster, others 2x, a few less than 100% and for very few, Skylake is a little slower than RyZen.

Generally speaking, integer SIMD architecture of Skylake is faster than FP SIMD architecture compared to RyZen.

RyZen is very close to Haswell and the difference is inside the margin of error.



Now, using your version of x265 and the above four systems, I got these results:


Results x265: (v2.3+7)


Skylake-2C/2T@3.0GHz -> 0.86 fps

RyZen (1+1)-2C/2T@3.0GHz -> 0.71 fps

RyZen (2+0)-2C/2T@3.0GHz -> 0.70 fps

Haswell-2C/2T@3.0GHz -> 0.67 fps

Sandy-2C/2T@3.0GHz -> 0.44 fps

There is a good 52% gain for Haswell over Sandybridge, but less than x265 FHD benchmark of Atak that gave me a 71% gain.

Is it your version, your settings or the 4K sample ?
I don't know.

Skylake is faster than Haswell ~28% due to better architecture and faster AVX2 implementation (?), a lot better than x264 difference so I think that AVX2 implementation plays a significant role here.

Skylake is 95% (!) faster than Sandybridge at 3.0GHz.

But the real surprise here is RyZen that manages to overcome Haswell.
It seems that 4K x265 encoding, although it has more AVX2 integer optimizations, has better IPC for RyZen than 1080p x264 encoding compared to older Intel architectures (Sandybridge & Haswell) but not Skylake

Or maybe AVX2 optimized x264 encoding is actually slower for RyZen, than non-AVX2 settings.


I then tried optimized latest versions of x265 using MS VS 2017 for AVX and AVX2 sets from here http://msystem.waw.pl/x265/


Results x265: (v2.3+18 MS 2017 AVX/AVX2)

Haswell-2C/2T@3.0GHz -> 0.78 fps

Sandy-2C/2T@3.0GHz -> 0.48 fps


First of all those VS 2017 optimized version are 10% faster than yours for Sandybridge and 16% for Haswell.

Now, the gain for Haswell goes to 62.5% which is closer to x265 FullHD benchmark results from Atak.


One last thing.
Inside the zip file there are your encoding results after running the benchmark which are ~100MB.
You could delete them and re-archive the benchmark.zip without them in order to save 100MB.

P.S

Full speed (CPU 100% utilization) at stock and overclocked settings of Skylake Core i7 6700K 4C/8T,
Kabylake Core i7 7700K 4C/8T and RyZen R7 1700 8C/16T using x264 & x265 apps.


x264 (r2744):


Ryzen R7 1700@3.85GHz r2762 (--asm AVX) 23.46 fps

Ryzen R7 1700@stock 19.07 fps

Kabylake 7700K@4.8GHz 15.36 fps

Skylake 6700K@4.7GHz 14.82 fps

Skylake 6700K@stock 13.27 fps



x265(2.3+7):


Ryzen R7 1700@stock 3.19 fps

Skylake 6700K@4.7GHz 3.17 fps

Skylake 6700K@stock 2.86 fps

Sagittaire
13th March 2017, 20:24
OK.

I tried your benchmark and I had some weird results using that x264 version included.
The two CPUs are Core i3-4170 & Core i5-2400.

Your version is not the latest, but a version before latest.

It's r2744 and the latest version is r2762 as you can see here:
http://download.videolan.org/x264/binaries/win64/


v2744 is really recent version. r2762 will not change result for speed.


During testing the screen reported an fps ~3.5-4.5 fps IIRC but when finished the final line gave me 58.8 fps which is completely wrong obviously.
I then disabled HT and underclocked to 3.0GHz and gave me 9.44fps which is still huge and obviously wrong.
On the other hand, Sandybridge result at 3.0GHz with only 2 cores active gave me 1.84 fps which is very low I think.

it's simply because there are two different speed: ffmpeg frame server and x264 encoder itself. Initial speed for frame server is really high simply because x264 encoder have initial lookahead frame buffer.


Results x264: (r2762)
Haswell-2C/2T@3.0GHz -> 3.37 fps
Sandy-2C/2T@3.0GHz -> 2.56 fps

I was thinking that x264 is not too much AVX2 integer optimized, but I was wrong obviously, because I had seen a post of the developer a few years ago comparing Haswell vs Ivybridge with only 5% gain for Haswell at the same clock.

yes there are big AVX/AVX2 optimisation in x264. Certainely less than in x265 but really not negligible.


Are there any weird switches in x264 settings that are causing problems with benchmark tests ?

There no problem. It's really common command line for 1080p encoding: --preset slower --tune grain --crf 20



I then tried optimized latest versions of x265 using MS VS 2017 for AVX and AVX2 sets from here http://msystem.waw.pl/x265/

Results x265: (v2.3+18 MS 2017 AVX/AVX2)
Haswell-2C/2T@3.0GHz -> 0.78 fps
Sandy-2C/2T@3.0GHz -> 0.48 fps

I use gcc version in the benchmark. Be carrefull here, the author report than AVX version will be higher speed for AVX CPU:

All binaries do the same, so it is only about encoding speed. My recomendations are: for AVX2-CPU the fastest should be VS 2017 AVX2 version, for AVX-CPU – VS 2017 AVX version, for SSE4-CPU – VS 2017 none or GCC none version, for SSSE3-CPU – GCC SSSE3 version, for CPU without even SSSE3 – GCC none version. You can determine fastest version by comparing encoding time on the same short sample.


Anyway, it's really interessing and complete report ... ;-)

NikosD
13th March 2017, 20:42
v2744 is really recent version. r2762 will not change result for speed.


I think it will.

If you see change log file between the two versions is full of AVX2 optimizations.

So, for recent CPUs there will be difference.


it's simply because there are two different speed: ffmpeg frame server and x264 encoder itself. Initial speed for frame server is really high simply because x264 encoder have initial lookahead frame buffer.

OK.

So, where can I find the right x264 speed after running the benchmark ?

Because using your version and reading the last line of the CLI, the results are unreasonable.

Can you provide a screenshot for sure ?



yes there are big AVX/AVX2 optimisation in x264. Certainely less than in x265 but really not negligible.

Last version has surely a lot of AVX2 optimizations.

If I can find a way to measure your version according to my previous paragraph in this post, I will tell you about that too.


I use gcc version in the benchmark.

It's a little slower as you can see from Microsoft's Visual Studio 2017 AVX/AVX2 optimizations.

Sagittaire
13th March 2017, 20:50
OK.

So, where can I find the right x264 speed after running the benchmark ?

Because using your version and reading the last line of the CLI, the results are unreasonable.

Can you provide a screenshot for sure ?

at the encoding end, the speed from frame serveur and x264/x265 converge. You can find final speed at the end of encoding in x264/x265 log information (with file size, psnr, ssim ... etc etc)

NikosD
13th March 2017, 20:53
Is there somewhere a log file that I didn't see ?

I'm not in front of my PC right now and I don't remember a log file.

Sagittaire
13th March 2017, 21:00
Is there somewhere a log file that I didn't see ?

I'm not in front of my PC right now and I don't remember a log file.

Just in the end of command line windows at the end of encoding. Don't worry, you report the good information: ffmpeg has just x.x fps precision and x264/x265 have x.xx precision. Moreover the fps for framserver and x264/x265 will be strickly the same at the end of encoding.

NikosD
13th March 2017, 23:02
I have updated my previous post, adding x264 r2744 results and Skylake results from a friend.

Sagittaire
14th March 2017, 13:35
@cojj

You can test benchmark on rysen?

cojj
15th March 2017, 01:54
@Sagittaire
Sorry for late reply - I'm very busy this week + my ryzen rig is doing a lot of heavy work.
I will try to get it done over the weekend.

On the side note, I'm going to France next month for business trip. I hope people are nice there :)

mandarinka
15th March 2017, 12:51
I already posted it here (https://forum.doom9.org/showpost.php?p=1800971&postcount=123), but apparently encoding performance can be different (lower) in Balanced power plan under Windows 10, compared to High Performance plan.

http://abload.de/img/ryzen_coreparking6lkdn.png

This difference (and maybe also the difference when HPET is on/off) might be worth testing if you can.

Sagittaire
15th March 2017, 13:45
I already posted it here (https://forum.doom9.org/showpost.php?p=1800971&postcount=123), but apparently encoding performance can be different (lower) in Balanced power plan under Windows 10, compared to High Performance plan.

http://abload.de/img/ryzen_coreparking6lkdn.png

This difference (and maybe also the difference when HPET is on/off) might be worth testing if you can.

yes hardware.fr french preview discovered (cocorico!!!) this problem:
http://www.hardware.fr/articles/956-22/indices-performance.html

anyway, if you have CPU charge at 100% then it's not a problem like for x264.

however with application at charge CPU with less than 100%, you can have up to 10% improvement (x265 or game for exemple).

Atak_Snajpera
15th March 2017, 15:34
Updated Flops table
https://i.imgsafe.org/6da60ecb0f.png

Sagittaire
15th March 2017, 18:11
Updated Flops table
https://i.imgsafe.org/6da60ecb0f.png

and ... ???

it's just synthetic test. In real life application, Rysen outperform Haswell, and by far. And in most case, is on par with Broadwell-E.

Atak_Snajpera
15th March 2017, 19:11
It shows interesting data about RyZen.
Up to 8 threads RyZen 7 in integer calculation acts like SandyBridge+. When you put more load on SMTs then efficiency increases beyond SkyLake/KabyLake.
Similar story with SSE2. If you force all SMT's to work then efficiency goes above Haswell. RyZen is very uneven CPU for sure unlike SkyLake/KabyLake. Not to mention about CCX issues.

Sagittaire
15th March 2017, 20:11
It shows interesting data about RyZen.
Up to 8 threads RyZen 7 in integer calculation acts like SandyBridge+. When you put more load on SMTs then efficiency increases beyond SkyLake/KabyLake.
Similar story with SSE2. If you force all SMT's to work then efficiency goes above Haswell. RyZen is very uneven CPU for sure unlike SkyLake/KabyLake. Not to mention about CCX issues.

well it's not true:

1) the real efficiency is not flops/hz but flop/watt. With a calculation like that, all our phone would have intel soc and not arm soc.

2) one more time, in real application like x264 (we are on doom9 after all), rysen outperfom sandybridge, IvyBridge and Haswell and by far.

http://www.hardware.fr/getgraphimg.php?id=464&n=1

and for x264 efficiency (fps/watt), the best CPU in the area is R7 1700 and by far.

Atak_Snajpera
15th March 2017, 20:25
Once again you haven't understood my message. Nothing new here...

Sagittaire
15th March 2017, 20:34
Once again you haven't understood my message. Nothing new here...

well we are on doom9, I don't care about synthetic result in ALU/Hz and FPU/Hz. It is not even a correct measure of efficiency.

What is important is the practical results: and better ALU/Hz or FPU/hz seem not really usefull to Intel in x264 and x265 encoding, isn't it?

I will even generalize: This does not seem to be really useful in many applications.

Atak_Snajpera
15th March 2017, 20:40
Without this synthetic benchmark you would be still living in dream land regarding AVX2 performance. Hint: x265 likes AVX2 (256bit) alooooot.

By the way how is your x265 benchmark. Have you finally optimized it for RyZen with those crazy settings? ;)

Sagittaire
15th March 2017, 21:37
Without this synthetic benchmark you would be still living in dream land regarding AVX2 performance. Hint: x265 likes AVX2 (256bit) alooooot.


Certainely ... but i7-6900K and Rysen 7 1800X will be on par for x265 encoding. It's like that even if it's hard to understand for you.


By the way how is your x265 benchmark. Have you finally optimized it for RyZen with those crazy settings? ;)

Well it's really simple: 2160p encoding with --pme command line for 8C/16T

Anyway you can make 1080p encoding with 4C/8T to compare Rysen (4+0 configuration) with Intel 4C/8T CPU. It's really simple to make that too ... :devil:

ShogoXT
15th March 2017, 22:17
I like more data no matter what, even if it's a synthetic result.

There is news about Ryzen having a bug with FMA3 workloads. I wonder if that's why when I ran the fpu benchmarks it froze up a few times?

Also when I turn on the computer after it being off for a while, it doesn't even reach the post screen. I have to turn it off again to get it to boot. I suspect my overclock or the motherboard....

Sagittaire
15th March 2017, 22:53
and here benchmark:
jfl1974.free.fr/Benchmark.zip

1080p for x264
2160p for x265

try and report your result on 8C/16T CPU:
-speed for x264 and x265
-CPU charge for x264 and x265

and here result with R7 1700 8C/16T @stock 3.0/3.2/3.7 GHz:
x265 4K with CPU charge at 100%:
encoded 1007 frames in 315.68s (3.19 fps), 15638.93 kb/s

X264 2K with CPU charge at 100%:
encoded 1649 frames, 19.07 fps, 8510.48 kb/s


Result with i5 3550 4C/4T at 3.5 Ghz:
x265 4K with CPU charge at 100%:
1.14 fps

X264 2K with CPU charge at 100%:
6.57 fps

NikosD
15th March 2017, 22:56
and here result with R7 1700 8C/16T:


Stock clocks ?

Or overclocked ?

The default/stock clock of R7 1700 8C/16T is 3.1GHz.

In x265 has exactly the same speed of 6700K@4.7GHz, but in x264 RyZen is definitely faster.

Sagittaire
15th March 2017, 22:59
Stock clocks ?

Or overclocked ?

The default/stock clock of R7 1700 8C/16T is 3.1GHz.

at stock base/turbomin/turbomax 3.0/3.2/3.7 GHz

Sagittaire
15th March 2017, 23:10
Stock clocks ?

Or overclocked ?

The default/stock clock of R7 1700 8C/16T is 3.1GHz.

In x265 has exactly the same speed of 6700K@4.7GHz, but in x264 RyZen is definitely faster.

you can extrapole result easily:

result with R7 1700 8C/16T:
x265 4K with CPU charge at 100%:
encoded 1007 frames in 315.68s (3.19 fps), 15638.93 kb/s

X264 2K with CPU charge at 100%:
encoded 1649 frames, 19.07 fps, 8510.48 kb/s


result with R7 1700X 8C/16T (estimation for x265):
x265 4K with CPU charge at 100%:
3.44 fps

X264 2K with CPU charge at 100%:
encoded 1649 frames, 20.61 fps, 8510.48 kb/s


result with R7 1800X 8C/16T (estimation for x265):
x265 4K with CPU charge at 100%:
3.61 fps

X264 2K with CPU charge at 100%:
encoded 1649 frames, 21.62 fps, 8510.48 kb/s

Motenai Yoda
16th March 2017, 17:08
well it's not true:

1) the real efficiency is not flops/hz but flop/watt. With a calculation like that, all our phone would have intel soc and not arm soc.

[...]

and for x264 efficiency (fps/watt), the best CPU in the area is R7 1700 and by far.

yep but you have to take into account the real power consuption, nor tdp, intels never go over their tdp, amds a lot.

Well it's really simple: 2160p encoding with --pme command line for 8C/16T
why pme over pmode?

microchip8
16th March 2017, 20:40
yep but you have to take into account the real power consuption, nor tdp, intels never go over their tdp, amds a lot.


why pme over pmode?

I don't know where you get that from. The power sensors on my FX8350 CPU reports 124.5W when I run an encode using all cores. The FX8350 is a 125W CPU so what the sensor reports is accurate.

Sagittaire
16th March 2017, 21:26
why pme over pmode?

pmode desactive reflist option and reflist is really powerfull option to optimisize speed for preset < veryslow. You can use pmode only with preset veryslow or placebo.

anyway --preset veryslow --pme --pmode is certainely really good to optimisize encoding speed for 1080p source.

evilr00t
16th March 2017, 23:42
pmode desactive reflist option and reflist is really powerfull option to optimisize speed for preset < veryslow. You can use pmode only with preset veryslow or placebo.

anyway --preset veryslow --pme --pmode is certainely really good to optimisize encoding speed for 1080p source.

Just... no.

CPU: E5-2695v3
1080p60 Dark Souls 3 capture:
Common flags: --preset veryslow --qg-size 8 --crf 24.5 --lambda-file (new file from x265 thread), --frames 500 --threads 28
x265 stock : 2.58 fps (100%)
x265 --pmode : 2.43 fps (94.2%, slower)
x265 --pme : 2.54 fps (98.4%, slower)
x265 --pmode --pme: 2.29 fps (88.8%, slower)

CPU use graph:
http://imgur.com/a/V18RK

pmode+pme can't even saturate this CPU, but as you can see, it slows down the encode. More work doesn't mean it runs faster, especially on SMT systems where going over 50% means you are slowing something else down.

Here's a retest without any custom settings:

Z:\>a:\avs4x265 -P z:\x265_main.exe --preset veryslow --frames 750 --pools 28 [--pme] [--pmode] --crf 24.5 -o NUL DS3.avs

stock : encoded 750 frames in 305.89s (2.45 fps), 6178.28 kb/s, Avg QP:32.15 - 100%
--pmode : encoded 750 frames in 335.43s (2.24 fps), 5985.25 kb/s, Avg QP:32.19 - 91.4%
--pme : encoded 750 frames in 311.13s (2.41 fps), 6178.28 kb/s, Avg QP:32.15 - 98.3%
--pmode --pme: encoded 750 frames in 354.57s (2.12 fps), 5985.25 kb/s, Avg QP:32.19 - 86.5%

x265 [info]: HEVC encoder version 2.3+17-6e348252e902
x265 [info]: build info [Windows][GCC 6.3.0][64 bit] 8bit

mandarinka
17th March 2017, 15:15
I don't know where you get that from. The power sensors on my FX8350 CPU reports 124.5W when I run an encode using all cores. The FX8350 is a 125W CPU so what the sensor reports is accurate.

Indeed. If you meassure 12V CPU line power, you can confirm this. It doesn't go over TDP at default settings (unless you disable TDP limits in bios and so on). The wattage/ampers on the power supply line will be somewhat higher than 125 W when you measure, but that is because you measure before voltage regulators converting from 12V to vcore. That creates extra power consumption (Intel or AMD platform) as evidenced by the VRMs generating substantial heat. The losses here can be 15-20 % added on top of the TDP of the CPU. And of course, losses on motherboard VRM don't count towards CPU's TDP, neither with Intel not AMD.

NikosD
21st March 2017, 12:50
Updated Flops table


Judging by that extremely optimized application for finding the maximum actual flops as close as possible to max theoretical flops that I posted here https://forum.doom9.org/showthread.php?p=1801309#post1801309, your table based on Intel optimized Flops.c might not be accurate.

The application is famous now as the "FMA3 bug" of RyZen, but TheStilt has already managed to run it on RyZen 1700 and posted the results here:
http://forum.hwbot.org/showpost.php?p=480934&postcount=34

According to the results of that app, the FMA3 implementation of RyZen is extremely efficient (~100% of theoretical flops) with Haswell optimized binary for 128bit SSE/SSE2 and 256bit AVX/FMA3 using FP32 and FP64, compiled by MS Visual Studio 2015.

Now, if you see the FP64 results, the FADD and FMUL tests have the same performance for both 128bit and 256bit (SSE2 vs AVX1) using that app.

BUT using FADD+FMUL (not FMA which is something different), RyZen almost doubles (~77%) the performance of FP64 compared to FADD, reaching the results of FMA3 - only ~28% difference.

So, my suggestion is to recompile flops.c using MS Studio 2017 for SSE2/AVX/AVX2-FMA3 instructions and run it again on RyZen.

I don't know if it will be faster than ICC in absolute numbers, but I think we will see different results in relative numbers between different instruction sets SSE2/AVX/AVX2-FMA3.

I will post my binaries of flops.c compiled by MS Studio 2013 and your GUI, in order to be tested by a RyZen user, although a more recent version of MS Studio like 2017 could make a difference.

NikosD
21st March 2017, 13:36
@All RyZen owners.

OK, so this is the GUI of Atak but with my compilations of a little old MS Visual Studio 2013.

You can get it here:
https://mega.nz/#!s9cFXA6a!q-co5cemP-ZaeLFdxcNgVGtgPPUVSczBOLUZRVLEXak

After running the app, please go to menu "Main" - > "Save screenshot" and upload the image of your results.

It should look something like that:
https://s11.postimg.org/wablgq6df/Core_i3_4170_3_0_GHz_Cache_3_0_GHz_HT_OFF.jpg

Sagittaire
22nd March 2017, 01:14
for R7 1700 ...

http://reho.st/self/976028f757a85dedaf36b369bc991f96df523975.jpg

burfadel
22nd March 2017, 02:16
for R7 1700 ...

http://reho.st/self/976028f757a85dedaf36b369bc991f96df523975.jpg

Interesting how the mutli-core on Ryzen in all cases is more than 8 times more power than single core. This suggests that not all of the single core was saturated resulting in the SMT threads being beneficial when running the multi-core test.

If you scale back the results from multi-core to what single core should read, if the whole core was saturated:

x86: 64.13 MFLOPS
x87: 3.06 GFLOPS
SSE2: 4.91 GFLOPS
AVX: 4.99 GFLOPS
AVX2: 7.69 GFLOPS

It suggests something isn't quite right somewhere :)

Atak_Snajpera
22nd March 2017, 13:37
for comparision binaries compiled by Intel Compiler 15
Ryzen@4GHz and RyZen@3.7GHz
http://i.imgur.com/a44Bmnn.jpg http://i.imgur.com/UHgQYiL.jpg

Intel compiler once again generates the fastest code for AMD :) This clearly debunks any conspiracy theories that Intel compiler favours intel cpus.

Sagittaire
22nd March 2017, 14:05
for comparision binaries compiled by Intel Compiler 15
Ryzen@4GHz and RyZen@3.7GHz
http://i.imgur.com/a44Bmnn.jpg http://i.imgur.com/UHgQYiL.jpg

Intel compiler once again generates the fastest code for AMD :) This clearly debunks any conspiracy theories that Intel compiler favours intel cpus.

http://img4.hostingpics.net/pics/904693kabini.png

not so far than Jaguar or i3-4170 ... :devil:

This benchmark seem dont work really well: in practice, R7-1800X is on par with i7-6900K for all heavy application (x264, x265, 3D-calculation ... etc etc)

NikosD
22nd March 2017, 14:53
for R7 1700 ...


Intel compiler once again generates the fastest code for AMD :) This clearly debunks any conspiracy theories that Intel compiler favours intel cpus.

Not exactly the results that I was expecting from MS VC 2013.
I get similar gains with Sandy and Haswell like RyZen.

Probably MS VC 2013 compiler is too old to vectorize flops.c for SSE2/AVX/AVX2-FMA3.

Based on the fact that MS VC 2017 makes the fastest executable for x265, even faster than Intel's compiler and GCC, someone with MS Visual Studio 2017 should try to make a fast executable of flops.c in order for us to test any differences in relative numbers.

It seems that flops.c is better optimized by Intel's compiler and the autovectorizer.

Now, regarding that myth of Intel's compiler, it wasn't a myth.

Probably Intel after paying a multi-million $ penalty to AMD, after cheating and manipulating in various ways compilers, OEMs etc and grabbing a CPU share that is worth a lot more than the penalty itself, could now play fair (at least more than in the past)

But we don't compare in this example different compilers to find the faster one, but only to find specific differences in instructions sets.

The truth is that RyZen mimics a lot Intel's HW and can manage to run fast enough executables optimized for Intel, BUT that doesn't mean that it is optimized for RyZen.

The other guy with that Haswell optimized "FMA3 bug" executable had managed to optimize 128bit/256bit FP32 and FP64 flops app for Bulldozer and Piledriver compiled by MS Visual Studio 2015, having a lot of asm optimizations for those architectures.

I think we have to wait for a RyZen optimized FLOPS app by him.

burfadel
22nd March 2017, 14:54
http://img4.hostingpics.net/pics/904693kabini.png

not so far than Jaguar or i3-4170 ... :devil:

This benchmark seem dont work really well: in practice, R7-1800X is on par with i7-6900K for all heavy application (x264, x265, 3D-calculation ... etc etc)

Yeah something is bung there. 8 core should essentially be 8 times the performance of single core, the SMT threads don't count because they allow a parellel process for unused core power. If the core is saturated, they serve no purpose.

Your CPU scales correctly, the 4.2 is probably because there was some other process that momentarily ran during the single core test.

Atak_Snajpera
22nd March 2017, 16:01
Based on the fact that MS VC 2017 makes the fastest executable for x265, even faster than Intel's compiler and GCC, someone with MS Visual Studio 2017 should try to make a fast executable of flops.c in order for us to test any differences in relative numbers.
Why don't you check yourself?
https://www.visualstudio.com/downloads/

NikosD
22nd March 2017, 16:11
Why don't you check yourself?
https://www.visualstudio.com/downloads/
I have completely abandoned any thoughts of putting again developer tools in my system.

I remember you looking for some info of different compilers than Intel's to test.
I offered you one!

Nevermind, if you are not in the mood just forget it.

Sagittaire
22nd March 2017, 17:38
One more time, i don't think that this bench work correctly:

if you compare i3-4170 vs Rysen 7 in single core mode, i3-4170 produce better result and by far than Rysen 7 for x86, x87, SSE2, AVX or AVX2.

Anyway for single thread in x264 bench you have:

Rysen 7 1800X: 0.6175 fps/ghz
Core i7-6900K: 0.5756 fps/ghz
Core i7-5960X: 0.5857 fps/ghz
Core i7-7700K: 0.6866 fps/ghz
Core i7-4790K: 0.6363 fps/ghz
Core i7-3770K: 0.5512 fps/ghz
Core i7-2600K: 0.5052 fps/ghz

How Rysen can be really catastrophic in all synthetic test for all intructions and fight with i7-6900K in real application like x264 or x265?

http://www.hardware.fr/getgraphimg.php?id=440&n=1

NikosD
22nd March 2017, 17:49
One more time, i don't think that this bench work correctly:

if you compare i3-4170 vs Rysen 7 in single core mode, i3-4170 produce better result and by far than Rysen 7 for x86, x87, SSE2, AVX or AVX2.

Anyway for single thread in x264 bench you have:

Rysen 7 1800X: 0.6175 fps/ghz
Core i7-6900K: 0.5756 fps/ghz
Core i7-5960X: 0.5857 fps/ghz
Core i7-7700K: 0.6866 fps/ghz
Core i7-4790K: 0.6363 fps/ghz
Core i7-3770K: 0.5512 fps/ghz
Core i7-2600K: 0.5052 fps/ghz

How Rysen can be really catastrophic in all synthetic test for all intructions and fight with i7-6900K in real application like x264 or x265?

http://www.hardware.fr/getgraphimg.php?id=440&n=1

We must not compare apples with oranges.

The synthetic test measures theoretical FPU performance using double precision FP64 numbers in FADD/FSUB/FMUL operations of legacy x87 and vector processing SIMD like SSE2/AVX/AVX2-FMA3

None of them is used by apps like x264 or x265.

So, think of x264 & x265 apps as a measurement of integer SSE2/AVX/AVX2 performance and FlopsCPU as a measurement of FP64 SIMD SSE2/AVX/AVX2 performance.

That x86 test of FlopsCPU is irrelevant.

Now, if you try to find out FP64 real world apps that behave like FlopsCPU benchmark, you have to dig in the world of HPC, scientific, 3D, rendering apps although they are not so much optimized for FP64 SIMD like that synthetic-theoretical app.

Sagittaire
22nd March 2017, 19:38
Now, if you try to find out FP64 real world apps that behave like FlopsCPU benchmark, you have to dig in the world of HPC, scientific, 3D, rendering apps although they are not so much optimized for FP64 SIMD like that synthetic-theoretical app.

Rysen 7 seem really make good work in all 3D rendering apps too for exemple ...

There is in fact only a very rare real application where Rysen does not give a good result. Why test SIMD that are never used in practice?

CruNcher
22nd March 2017, 21:57
Intel put the ultra efficient Mobile versions on the table and KabyLake is now Flooding the Laptop Market we gonna see soon how well Ryzen will really do as well how well AMD will GPU wise do vs the Nvidia + Intel combination ;)

At least Intel + AMDs GPU combination will follow very soon RX 5xx (560M/570M) updated version against mostly the GTX 1050 TI (14nm) updated VPU/GTX 1060 (16nm).

And intels 4 Cores/4T do really good so far :)

https://www.youtube.com/watch?v=mwiY2OOz1sY&list=UUGtSdR7VYd3U9J4O99vl_Cg

Efficiency vs the Base (PS4 Pro) seems not perfect but not that bad at all :)

https://www.youtube.com/watch?v=PVEbJzIEABg

So soon well see what AMDs 4Cores/4T and 8T brings on the table vs it at perf/watt price overall ;)

More data to crunch is incoming

https://www.youtube.com/watch?v=TxYVbg2VMuw&list=UU7D8-hkY0PrS5BNgl_HBXNg

i7-7300HQ
i7-7700HQ

are damn efficient @ 45W with integrated GPU

Zen needs to at least stay competitive in that Performance range Mobile

Bristol Ridge currently tumbles somehwere @ 65W (not efficient enough).

It seems currently AMD needs to counter attack a price range of 1000 on the lowest entry with rather better or at least equal performance to a Intel Kaby Lake (14nm) /Nvidia GTX 1050 TI (14nm) combo Platform (with both superior VPU Software/Hardware state), that fight will be a tad harder then the Desktop one every watt inefficiency vs price counts now (Software/Hardware) ;)

Biggest problem for AMD Vega is not ready neither Desktop nor Mobile and they have to go without very important improvements into this now with a updated Polaris before Raven Ridge is even ready with it's Vega Core and could gain some traction with the MultiGPU scaling.

zub35
23rd March 2017, 02:40
Ryzen7 1700, avx2 on/off
http://rigaya34589.blog135.fc2.com/blog-entry-909.html

--

Owners of ryzen, test x265 on different compilers GCC, VS, ICC
With AVX2 enabled and disabled in the settings. To find out the fastest speed.

GCC / VS: http://x265.ru/builds
ICC: http://msystem.waw.pl/x265

burfadel
23rd March 2017, 03:55
Ryzen7 1700, avx2 on/off
http://rigaya34589.blog135.fc2.com/blog-entry-909.html

--

Owners of ryzen, test x265 on different compilers GCC, VS, ICC
With AVX2 enabled and disabled in the settings. To find out the fastest speed.

GCC / VS: http://x265.ru/builds
ICC: http://msystem.waw.pl/x265

It doesn't show too badly for Ryzen. An important consideration is RAM speed, something that they're improving with Ryzen and hopefully they'll have resolved in the next couple of months. Ryzen apparently loves faster memory, that test was done with 2400 MHz memory on Ryzen and 3600 MHz memory on the i7-7700K. The i7 was also clocked at the maximum overclock you should consider when not using high CPU load continuously on a regular basis (which is closer to 4.6 GHz). The 1700X was at stock speeds.

I think bios and software updates would make that graph show something different in a good way for the 1700X in 6 months time, assuming they have fully updated and running the fastest realistic RAM they can (maybe like 3600 MHz).

CruNcher
23rd March 2017, 18:21
It indeed becomes interesting comparing clock for clock with a very fast pooling time :)

https://www.youtube.com/watch?v=KDZ1dVu80as


unfortunately no one really does synced Power tests at the same time but obviously the Goal reaching 45W like KabyLake could be reached now :)

Without doubt it is one of the best AMD processors since Athlon 64 :)

And really excited to the see the Mobile results in the not so distant future vs Intel of Zen :)

ShogoXT
24th March 2017, 03:39
Ryzen7 1700, avx2 on/off
http://rigaya34589.blog135.fc2.com/blog-entry-909.html

--

Owners of ryzen, test x265 on different compilers GCC, VS, ICC
With AVX2 enabled and disabled in the settings. To find out the fastest speed.

GCC / VS: http://x265.ru/builds
ICC: http://msystem.waw.pl/x265

If I read that correctly, so it IS faster with avx2 off. Would avx be better than sse4.2?

Turns out my ram was only running at 2133 before. The mobo reduced it automatically instead of telling me the settings failed. Got new ones and it's running at 3200 for real this time. I was actually seeing frame lag in some games before I managed to eek it up to 2400 before changing the ram.

Ryzen loves fast ram , remember that.

sneaker_ger
24th March 2017, 10:57
If I read that correctly, so it IS faster with avx2 off.
For some of the tests at least. (x264 8bit and x265 8 bit but not x265 10 bit)

CruNcher
25th March 2017, 08:31
Seems slowly we see several Studios optimizing intensively for AMD already based on their PS4 work :)

http://media.bestofmicro.com/ext/aHR0cDovL21lZGlhLmJlc3RvZm1pY3JvLmNvbS8wLzQvNjU5NjY4L29yaWdpbmFsL2ltYWdlMDUyLnBuZw==/r_600x450.png

http://media.bestofmicro.com/ext/aHR0cDovL21lZGlhLmJlc3RvZm1pY3JvLmNvbS8wL0IvNjU5Njc1L29yaWdpbmFsL2ltYWdlMDUyLnBuZw==/r_600x450.png

http://media.bestofmicro.com/ext/aHR0cDovL21lZGlhLmJlc3RvZm1pY3JvLmNvbS8xLzMvNjU5NzAzL29yaWdpbmFsL2ltYWdlMDUyLnBuZw==/r_600x450.png



Eidos new Glacier2 based Engine goes a complete different way then anything else currently.


Even with its GPU internally Energy Efficiency still on Intels side for system idle.

https://img.purch.com/w/711/aHR0cDovL21lZGlhLmJlc3RvZm1pY3JvLmNvbS8zLzYvNjU4NDgyL29yaWdpbmFsLzAxLUlkbGUucG5n

But if we look at Workloads it changes rapidly

https://img.purch.com/w/711/aHR0cDovL21lZGlhLmJlc3RvZm1pY3JvLmNvbS8zLzgvNjU4NDg0L29yaWdpbmFsLzAzLUdhbWluZy1NZXRyby1MTC1GSEQtSGVhdnktV29ya2xvYWQtLnBuZw==

https://img.purch.com/o/aHR0cDovL21lZGlhLmJlc3RvZm1pY3JvLmNvbS8zLzMvNjU4NDc5L29yaWdpbmFsLzA0LVRvcnR1cmUtTG9vcC5wbmc=

https://img.purch.com/w/711/aHR0cDovL21lZGlhLmJlc3RvZm1pY3JvLmNvbS84L0MvNjU4NjY4L29yaWdpbmFsLzExLVRlbXBlcmF0dXJlcy1Ub3J0dXJlLnBuZw==


Pretty Amazing results if the data is not wrong measured and i trust Igor much more then any other Reviewer on this Entire Planet :)


That's some crazy Preview to what we might gonna see Mobile next AMD might gonna take the Lead with Raven Ridge based Mobile Systems.

Though Intel has the drawback here of the Internal GPU that shouldn't be forgotten as well ;)

And overall some of AMDs R&D still lacks in the whole package ;)

Mobile though it's already pretty clear that Intel has no chance and also the Combination of Nvidia + Intel on the lowest power possible might not save it this time it looks very good for AMD taking the Mobile Lead in overall Performance and Consumption in front of Intel + Nvidia ;)

Scorpio will be yummy :)

NikosD
25th March 2017, 09:27
I have updated my x264 & x265 benchmark results of this thread here:
https://forum.doom9.org/showthread.php?p=1800781#post1800781

Take a look.
It seems that RyZen has better IPC for x265 than x264 using the default settings of the benchmark, compared to older Intel architectures (Sandybridge & Haswell) but not for Skylake.

CruNcher
25th March 2017, 10:07
~22W more power efficient then Skylake with GPU ;)
and still ~8W then Kaby Lake also with GPU

though we have no x264/x265 based Workload Power Efficiency Data compared with AVX/AVX2.

but you could partly relate it to existing data looking at temperature results.

which are overall more efficient measurable with the OS latency inside the OS then Power ;)

we see roughly 8W higher efficiency then i7-7700 @ 3.8 GHz without any GPU core on the Gaming test

But the torture loop becomes more interesting which also displays more what x264 and x265 are ;)

and again wee see ~22W higher efficiency @ same clock then Skylake.

Though we also see that temperature wise KabyLake still leads by ~-2 °C and at 3.8 Ghz even a whoping ~-14 °C which is clear if you look @ the Power difference of 85W vs 142W almost a ~60W difference now we could calculate ~30W per real core ;)

and Ryzen over Skylake by that exact ~-2 °C

Also a thing i hate about Reviews needing to search for important data on zillion of pages instead of getting it displayed efficiently in 1 view per tested workload

All these Zillion of Pages make no real sense instead of all data per workload combined in 1 view.

Now it becomes clearer why Raven Ridge will be 4 cores only @ first even with the better yields 6 cores could become problematic overhead wise vs Intel Mobile ;)

So overall that Mobile lead could be faster over for AMD then thought with Coffe Lake ;)

NikosD
25th March 2017, 10:23
AMD could have a bigger problem with the new HEDT platform of Intel using X299 chipset and Skylake X/ Kabylake X CPUs than Broadwell-E HEDT.

Kabylake X is a quad core meaningless decision.

But Skylake X will go up to 10C/20T and will be much faster than the Broadwell-E HEDT CPU.

Thank god this new HEDT platform from AMD will use 16C/32T CPUs and if the rumor is true, it will eat Skylake X for breakfast.

http://www.guru3d.com/news-story/amd-x390-and-x399-chipsets-diagrams-reveal-hedt-information.html

CruNcher
25th March 2017, 12:00
Skylake is not Kaby Lake and Intel Plans allready to attack AMD with 6 Cores Mainstream including GPU with better Energy Efficiency then AMDs 8 Cores without GPU

The most efficient Zen combination we gonna see in Scorpio a really ambitious platform project Hardware/Software :)

Even Raven Ridge will be only a Glimpse of it's overall Efficiency.

What normal AVG PC users get sold currently is only the trash stuff ;)

NikosD
25th March 2017, 12:09
Nobody cares for energy efficiency on desktop or HEDT as long as the power consumption and temperatures are in sane levels.

AMD is on par with Intel on that and it has faster CPUs.

Kabylake CPU is a higher clocked Skylake CPU just to somehow compete with Ryzen 7.

That 6core CPU from Intel only exists because of Ryzen 8C/16T.

There would be no 6core CPU for mainstream desktop from Intel, if AMD hadn't released Ryzen.

And by the time that 6core Intel CPU will be released, Zen 2 will be released also.

6core Intel will have to fight Zen 2.

CruNcher
25th March 2017, 12:28
There would be no 6core CPU for mainstream desktop from Intel, if AMD hadn't released Ryzen.

Of curse that's called competition no one is making you a present of more performance lower consumption for the same or lower cost it all comes through competition :)

And yes overall Ryzen is nice as i said best AMD processor since Athlon 64 after that i left AMD for Intel same as i left ATI after the Radeon 9000 series for Nvidia ;)

and in both cases now AMD gave partly up on their own paths and adapting it's Technology and gains traction again in both fields chasing it's competition now based on mostly the same Cores with still slightly different approaches though which both combine into 1 unique Ecosystem which Alpha test we saw with the Consoles starting ;)

Atak_Snajpera
25th March 2017, 12:38
Leaked Ryzen 5 results
https://www.purepc.pl/image/news/2017/03/24_procesory_amd_ryzen_5_wyciekly_do_sprzedazy_za_granica_2.jpg

It shows that Ryzen in SSE2 is only slightly faster than Sandy/Ivy Bridge. I see similar performance in flops. Coincidence? I do not think so.

Sagittaire
25th March 2017, 17:20
Leaked Ryzen 5 results

It shows that Ryzen in SSE2 is only slightly faster than Sandy/Ivy Bridge. I see similar performance in flops. Coincidence? I do not think so.

absoluty not, one more time: bad information or alternative fact ... :devil:

http://www.guru3d.com/index.php?ct=articles&action=file&id=28906&admin=0a8fcaad6b03da6a6895d1ada2e171002a287bc1

On single thread R7 1800X is on par with i7-7700K. If you make benchmark in score/frequency you have:
i7-7700K: 510 unit/ghz
R7 1800X: 540 unit/ghz

On single thread, R5 1600 with turbo at 3.6 Ghz (single thread) produce the same and coherent result in your CPUz benchmark:
R5 1600: 548 unit/ghz

http://www.guru3d.com/index.php?ct=articles&action=file&id=28961&admin=0a8fcaad6b03da6a6895d1ada2e171002a287bc1


For CPUz benchmark in MT, Rysen is simply a Intel Killer: R5 1600 6C/12T produce simply better result than i7-6900K 8C/16T here:

http://www.guru3d.com/index.php?ct=articles&action=file&id=28907&admin=0a8fcaad6b03da6a6895d1ada2e171002a287bc1


source: Guru3d (http://www.guru3d.com/articles-pages/amd-ryzen-7-1800x-processor-review,9.html)

Atak_Snajpera
25th March 2017, 18:25
I don't care about cpu-z. This bench is a mistery. We know zero about nature of calculations. How many integer and how many float tests are there? How does it calculate those scores? Does it use SIMD and so on...

Look at cinebench (SSE2).

CruNcher
25th March 2017, 18:30
absoluty not, one more time: bad information or alternative fact ... :devil:

http://www.guru3d.com/index.php?ct=articles&action=file&id=28906&admin=0a8fcaad6b03da6a6895d1ada2e171002a287bc1

On single thread R7 1800X is on par with i7-7700K. If you make benchmark in score/frequency you have:
i7-7700K: 510 unit/ghz
R7 1800X: 540 unit/ghz

On single thread, R5 1600 with turbo at 3.6 Ghz (single thread) produce the same and coherent result in your CPUz benchmark:
R5 1600: 548 unit/ghz

http://www.guru3d.com/index.php?ct=articles&action=file&id=28961&admin=0a8fcaad6b03da6a6895d1ada2e171002a287bc1


For CPUz benchmark in MT, Rysen is simply a Intel Killer: R5 1600 6C/12T produce simply better result than i7-6900K 8C/16T here:

http://www.guru3d.com/index.php?ct=articles&action=file&id=28907&admin=0a8fcaad6b03da6a6895d1ada2e171002a287bc1


source: Guru3d (http://www.guru3d.com/articles-pages/amd-ryzen-7-1800x-processor-review,9.html)

Keeping power requirements for Workloads under the hood is very unimpresive and Frequency out of the sweetspot causes massive power.

And only the fewest Reviewer can measure power requirements in some way efficient guru3d is not one of them.

Sagittaire
25th March 2017, 18:31
I don't care about cpu-z. This bench is a mistery. We know zero about nature of calculations. How many integer and how many float tests are there? How does it calculate those scores? Does it use SIMD and so on...

Look at cinebench (SSE2).

Well Rysen produce certainely best result in this bench: AMD use cinebench for these technical demonstration ... :D

http://www.guru3d.com/articles_pages/amd_ryzen_7_1800x_processor_review,10.html

One more time: bad information ... :devil:

NikosD
25th March 2017, 18:35
Cinebench is probably the best SW of ALL out there to demonstrate RyZen's multithreaded performance.

It's like RyZen was built for Cinebench, the favorite benchmark of Intel in previous years (before RyZen's release)

It should be a shock for Intel fan boys to loose so badly in their favorite benchmark.

Sagittaire
25th March 2017, 18:40
Cinebench is probably the best SW of ALL out there to demonstrate RyZen's multithreaded performance.

It's like RyZen was built for Cinebench, the favorite benchmark of Intel in previous years (before RyZen's release)

It should be a shock for Intel fan boys to loose so badly in their favorite benchmark.

yes ... and Atak_Snajpera show simply that little R5 1600 will be better than i7-7700K in this bench

Atak_Snajpera
25th March 2017, 18:42
Well Rysen produce certainely best result in this bench: AMD use cinebench for these technical demonstration ... :D

http://www.guru3d.com/articles_pages/amd_ryzen_7_1800x_processor_review,10.html

One more time: bad information ... :devil:

What is your problem? Can't you do simple math?

Ryzen 1700@4GHz = 1596 CB
https://youtu.be/Dan36jGqmzE?t=8m35s

1596 * 0.75 = 1197 cb for Ryzen 5 (6C/12T) @ 4GHz

CruNcher
25th March 2017, 19:04
Amd does this because Cinebench is Intel optimized from bottom to ground ;)

https://www.youtube.com/watch?v=PzLxCo5qofo

And not because that is so bad but the entire difference since Ryzen even unoptimized ;)

also these tests are mostly flawwed because Ryzen can never holdup 4 Ghz under full Core Load

http://i.cubeupload.com/vCYY13.jpg

Sagittaire
25th March 2017, 19:08
What is your problem? Can't you do simple math?

Ryzen 1700@4GHz = 1596 CB
https://youtu.be/Dan36jGqmzE?t=8m35s

1596 * 0.75 = 1197 cb for Ryzen 5 (6C/12T) @ 4GHz

yes ... but your initial affirmation is:

It shows that Ryzen in SSE2 is only slightly faster than Sandy/Ivy Bridge. I see similar performance in flops. Coincidence? I do not think so.

... and it's completely false ... :devil:

Rysen is just slightly faster than Broadwell-E in cinebench ... ;-)

http://www.guru3d.com/index.php?ct=articles&action=file&id=28968&admin=0a8fcaad6b03da6a6895d1ada2e171002a287bc1

http://www.guru3d.com/index.php?ct=articles&action=file&id=28969&admin=0a8fcaad6b03da6a6895d1ada2e171002a287bc1

CruNcher
25th March 2017, 19:52
It is crazy that wee see reviewers only measuring Encoding thereby Decoding is much more time critical and overall more stressful for the System in terms of overall stability ;)

And such 8 cores 16 threads perfect for 4K (UHD) High Bitrate Decoding efficiency tests also of instructions and overall optimization and it can be much better correlated where problems occur in near realtime by just listening to the Fan behavior and it's overall clock switching efficiency on the OS side and scheduling made hearable under it's full Power output.

you stressing a real scenario also on the Driver side :)

http://i1.sendpic.org/i/kb/kbOlJXnoiYyTezTj0OXLKzykIhp.png

Atak_Snajpera
25th March 2017, 20:28
yes ... but your initial affirmation is:

Quote:
It shows that Ryzen in SSE2 is only slightly faster than Sandy/Ivy Bridge. I see similar performance in flops. Coincidence? I do not think so.
... and it's completely false ...

Problem Or you mad bro?
Knowing that Ryzen 5 (6C/12T) @ 4GHz reaches 1200 CB we can calculate average clock frequency during test.
1100 CB / 1200 CB * 4GHz = 3.7 GHz
Max turbo for 3930k is 3.8 GHz

http://i.imgsafe.org/6c3dcbe5a4.png
http://i.imgsafe.org/6c4a9ed938.png

CruNcher
25th March 2017, 21:12
It interesting to see how the price tumbled down and slightly up again :D

http://www.cpubenchmark.net/cpu.php?cpu=Intel+Core+i7-7700K+%40+4.20GHz&id=2874

http://www.cpubenchmark.net/compare.php?cmp[]=2874&cmp[]=2970&cmp[]=2905


hehe but thats even more funny

http://www.cpubenchmark.net/cpu.php?cpu=AMD+Ryzen+7+1700&id=2970

up down up down


most probably the SMT News :D


I do believe that Skylake levels of performance are out of reach for Zen, and Intel will have their new CPU coming out around the same time as Zen, so Zen will be two generations behind in performance compared to Intel - which is a far cry better than the six generations they are currently behind.

Zen+'s promised 15% boost, though, will best Skylake, so if they get that out in a year's time after Zen, then they may practically catch up to Intel.

He wasn't really right which is good for everyone it's only 1 Generation before time to market now for AVG user starting when Coffee Lake Hits :)

And 1 Generation is not really that much for a CPU when the Price for the competition is acceptable :)

it looks very sane on the CPU side now though very insane on the GPU with Nvidia getting each generation more crazy and sooner or late falling on it's nose ;)


Though i really would like to see some more tests on the Decoding side of Ryzen that CCX interconect latency issues could become more problematic here :)


https://community.amd.com/community/gaming/blog/2017/03/14/tips-for-building-a-better-amd-ryzen-system


WOW you rarely see that Disclaimer

2. AMD processors, including chipsets, CPUs, APUs and GPUs (collectively and individually "AMD processor"), are intended to be operated only within their associated specifications and factory settings. Operating your AMD processor outside of official AMD specifications or outside of factory settings, including but not limited to the conducting of overclocking using the Ryzen Master overclocking software, may damage your processor, affect the operation of your processor or the security features therein and/or lead to other problems, including but not limited to damage to your system components (including your motherboard and components thereon (e.g., memory)), system instabilities (e.g., data loss and corrupted images), reduction in system performance, shortened processor, system component and/or system life, and in extreme cases, total system failure. It is recommended that you save any important data before using the tool. AMD does not provide support or service for issues or damages related to use of an AMD processor outside of official AMD specifications or outside of factory settings. You may also not receive support or service from your board or system manufacturer. Please make sure you have saved all important data before using this overclocking software.

They explain to you that Netflix could stop working ;)

What is uber funny is how they tell users hey overclock your memory and you get better Render results ehhhh note to AMD you would see the same on Intel ;)

https://www.youtube.com/watch?v=qksXthUcbiQ
https://www.youtube.com/watch?v=TId-OrXWuOE
https://www.youtube.com/watch?v=HZPr-gNWdvI

Sagittaire
26th March 2017, 10:14
Problem Or you mad bro?
Knowing that Ryzen 5 (6C/12T) @ 4GHz reaches 1200 CB we can calculate average clock frequency during test.
1100 CB / 1200 CB * 4GHz = 3.7 GHz
Max turbo for 3930k is 3.8 GHz


well it's enough. Cinebench is certainely the best test in area to show Ryzen power.

Here the unit/ghz for cinebench R15 in single thread test:
Core i7-7700K: 42.44 (Kaby Lake)
Core i7-6700K: 42.14 (Kaby Lake)
Core i7-6900K: 41.62 (Broadwell E)
Core i7-4790K: 39.31 (Haswell)
Rysen 7 1800: 39.00
Core i7-5960X: 38.85 (Haswell-E)
Core i7-2600K: 32.63 (Sandy Bridge)


Here the unit/ghz for cinebench R15 in multi thread test:
Rysen 7 1800: 425.78
Core i7-6900K: 418.64 (Broadwell E)
Core i7-5960X: 408.75 (Haswell-E)
Rysen 5 1600: 323.52
Core i7-7700K: 218.18 (Kaby Lake)
Core i7-6700K: 219.75 (Kaby Lake)
Core i7-4790K: 209.00 (Haswell)
Core i7-2600K: 176.17 (Sandy Bridge)


Here the unit/ghz/core for cinebench R15 in multi thread test:
Core i7-7700K: 54.54 (Kaby Lake)
Core i7-6700K: 54.93 (Kaby Lake)
Rysen 5 1600: 53.92
Rysen 7 1800: 53.22
Core i7-6900K: 52.33 (Broadwell E)
Core i7-4790K: 52.25 (Haswell)
Core i7-5960X: 51.09 (Haswell-E)
Core i7-2600K: 44.04 (Sandy Bridge)


You understand now that Rysen 7 is by far more performant for cinebench R15 than Sandy Bridge or Ivy Bridge:
- Ryzen is on par with Haswell for single thread
- Ryzen is on par with Broadwell E for Multi thread

You can note that SMT for Rysen is more efficient than Hyperthreading for cinebench.

It's clear now?

NikosD
26th March 2017, 11:11
I have updated my x264 & x265 benchmark results of this thread here:
https://forum.doom9.org/showthread.php?p=1800781#post1800781

Take a look.
It seems that RyZen has better IPC for x265 than x264 using the default settings of the benchmark, compared to older Intel architectures (Sandybridge & Haswell) but not for Skylake.
I have updated again my results with accurate RyZen tests and with a huge surprise

Ryzen running on 1+1 is faster than 2+0 (!)

Read more on my above post.

Sagittaire
26th March 2017, 11:19
and we are on doom9 and result is even better (but coherent) for x264:

http://www.hardware.fr/getgraphimg.php?id=465&n=9

it's clear now?

Sagittaire
26th March 2017, 11:28
I have updated again my results with accurate RyZen tests and with a huge surprise

Ryzen running on 1+1 is faster than 2+0 (!)

Read more on my above post.

yes ... because you have really higher cache in 1+1 than 2+0 mode. In fact twice for 1+1 vs 2+0.

and be carefull, SMT for ryzen in more efficient than Hyperthreading. Certainely that 2C/4T test will be even better for Ryzen.

http://www.hardware.fr/getgraphimg.php?id=437&n=1

Atak_Snajpera
26th March 2017, 11:36
You do realize that x264 mainly rely on performance of integer units in cpu,right? Cinebench is different. Rendering almost exclusively runs on FPU. For example Cinebench only uses SSE2 SIMD.
Clock vs Clock SSE2 unit in RyZen is only slightly faster than in SandyBridge-E. Deal with it.

Now I see that you are an another corporate fanboy. Numbers below clearly show that RyZen = SandyBridge-E in floating-point calculations.
http://i.imgsafe.org/6c3dcbe5a4.png

and be carefull, SMT for ryzen in more efficient than Hyperthreading. Certainely that 2C/4T test will be even better for Ryzen.
It is more efficient because it starts from lower level!
http://i.imgsafe.org/79a616c1ff.png

Sagittaire
26th March 2017, 11:43
You do realize that x264 mainly rely on performance of integer units in cpu,right? Cinebench is different. Rendering almost exclusively runs on FPU. For example Cinebench only uses SSE2 SIMD.
Clock vs Clock SSE2 unit in RyZen is only slightly faster than in SandyBridge-E. Deal with it.

Now I see that you are an another corporate fanboy. Numbers below clearly show that RyZen = SandyBridge-E in floating-point calculations.


No ... you are fanboy corporate ...

I post clearly all the result that you can find in all the preview: and conclusion is really good for Rysen in cinebench. Cinebench R15 is the test that AMD itself use for make technical demonstration for Rysen.

I don't know were you find your result. But I post the result that you can find in all independant preview:

Here the unit/ghz for cinebench R15 in single thread test:
Core i7-7700K: 42.44 (Kaby Lake)
Core i7-6700K: 42.14 (Kaby Lake)
Core i7-6900K: 41.62 (Broadwell E)
Core i7-4790K: 39.31 (Haswell)
Rysen 7 1800: 39.00
Core i7-5960X: 38.85 (Haswell-E)
Core i7-2600K: 32.63 (Sandy Bridge)


Here the unit/ghz for cinebench R15 in multi thread test:
Rysen 7 1800: 425.78
Core i7-6900K: 418.64 (Broadwell E)
Core i7-5960X: 408.75 (Haswell-E)
Rysen 5 1600: 323.52
Core i7-7700K: 218.18 (Kaby Lake)
Core i7-6700K: 219.75 (Kaby Lake)
Core i7-4790K: 209.00 (Haswell)
Core i7-2600K: 176.17 (Sandy Bridge)


Here the unit/ghz/core for cinebench R15 in multi thread test:
Core i7-7700K: 54.54 (Kaby Lake)
Core i7-6700K: 54.93 (Kaby Lake)
Rysen 5 1600: 53.92
Rysen 7 1800: 53.22
Core i7-6900K: 52.33 (Broadwell E)
Core i7-4790K: 52.25 (Haswell)
Core i7-5960X: 51.09 (Haswell-E)
Core i7-2600K: 44.04 (Sandy Bridge)

It's clear that you make trolling now ...

Atak_Snajpera
26th March 2017, 11:53
I don't know were you find your result. But I post the result that you can find in all independant preview:
Results are hardcoded in Cinebench app you dumb ass! Have you ever run cinebench r15 on your PC?!?
http://i.imgsafe.org/79d8a4a197.png

CruNcher
26th March 2017, 11:57
I have updated again my results with accurate RyZen tests and with a huge surprise

Ryzen running on 1+1 is faster than 2+0 (!)

Read more on my above post.

What is the overall result for 3+3 ?

Results are hardcoded in Cinebench app you dumb ass! Have you ever run cinebench r15 on your PC?!?
http://i.imgsafe.org/79d8a4a197.png

please dont forget they will slightly differ by Windows OS


Passmarks testing is also overall much more reliable imho and represents very accurately CPU Performance and you have access to the entire database of test and they have a sample test algorithm with error margin

NikosD
26th March 2017, 11:59
You mean 4+4 and it's in the end of my post above.

CruNcher
26th March 2017, 12:08
no i mean 3+3 actually preview what Ryzen 5 will bring on the table :)


6 cores will be the new Mainstream target 6/12

8 cores full utilization will take some time for every Software Developer to achieve properly Low Level and get good scaling results mostly driven by Scorpio ;)

if i had the decisions to efficiently upgrade i would rather put my money into 6 Cores and save the rest for the GPU side instead of 2 additional cores which performance delta isn't optimal yet due to it's interconnect and also save energy cost until Software is fully ready and most probably get the internal IGPU for free too ;)

Why should i just go for 2 cores if i can use the internal IGPU to lower Workload pressure by much more then the 2 cores can help higher Performance, that would be dumb ;)

Sagittaire
26th March 2017, 12:09
Results are hardcoded in Cinebench app you dumb ass! Have you ever run cinebench r15 on your PC?!?


No, Cinebench don't indicate the real frequency, just stock frequency. R5 1600 is not disponible for official test and you don't know how work really turbo: certainely 3.2/3.4/3.6 for Stock/Turbo max all core/turbo max one core. Moreover You don't know what are the real frequency for i7-3930K: stock is 3.2 Ghz and max turbo for one core is 3.8 Ghz and reall frequency for cinebench MT will be certainely between 3.2 and 3.8 Ghz but you can't say more.

I post independant result from Guru3d and I can post many other result from Ryzen preview if you want. And in all these preview, Ryzen is an Intel Killer: A simple R7 1700 at 329$ will be on par with i7-5960X or 17-6900K at 1000$.

http://www.anandtech.com/show/11170/the-amd-zen-and-ryzen-7-review-a-deep-dive-on-1800x-1700x-and-1700/18

http://images.anandtech.com/graphs/graph11170/85881.png

CruNcher
26th March 2017, 12:32
I wonder how surprised you will be Sagittaire if you see Intels 6 Cores beating that 8 Cores end of 2017 and already beating it in the Lab ;)

My i5 2400 btw is in that range of the i3 Kaby Lake by now

http://www.cpubenchmark.net/compare.php?cmp[]=2930&cmp[]=793&cmp[]=2970

But yeah you cant really deny Ryzen 7 1700 is hell of a good deal and Future investment with that Scorpio Backup coming :)

And that Glimpse shows the potential, Eidos Labs and Nixxes already do a pretty good job with the ground up work on the Dawn Engine

http://media.bestofmicro.com/ext/aHR0cDovL21lZGlhLmJlc3RvZm1pY3JvLmNvbS8wLzQvNjU5NjY4L29yaWdpbmFsL2ltYWdlMDUyLnBuZw==/r_600x450.png

Sagittaire
26th March 2017, 12:54
I wonder how surprised you will be Sagittaire if you see Intels 6 Cores beating that 8 Cores end of 2017 and already beating it in the Lab ;)

My i5 2400 btw is in that range of the i3 Kaby Lake by now

perhaps ... but it's not actually the case ... ;-)

Actually R7 1800X at 4.0 Ghz produce twice more performance than i7-6700K at 4.0 ghz in cinebench R15 MT.

If you want make that with 6C/12T Kaby Lake, you must have something like 5.0 Ghz for stock frequency. I seriousely doubt that Intel make that in 2017.
i7-7700K have actually low power efficiency and Intel is certainely at max clock for Kaby Lake for i7-7700K for correct power efficiency.

Power Efficacity (100*fps/watt) for x264 Encoding (higher is better)
http://www.hardware.fr/getgraphimg.php?id=464&n=1

CruNcher
26th March 2017, 13:22
Yeah actually i wonder also how they could achieve it especially with the IGPU overhead on practical the same node but it's Intel ;)

If they don't make it AMD looks really in the lead at least without IGPU overhead.


Raven Ridge is becoming more and more exciting to look out for for Laptops the more Zen data spreads vs i7-7700HQ :D

http://www.cpubenchmark.net/compare.php?cmp[]=2906&cmp[]=2937&cmp[]=2970


Did somebody do a 2+2 test yet and looked at the power consumption ?

though if you just calculate with 2 you roughly get @ 7000 not enough to beat Intel Mobile on the CPU side of course it would be pretty enough to beat Intel on the complete Platform with Vega inside this time no doubt about that even with slightly lower IPC Intel would see no sunlight this time (with their shrinking advantage made pooof) in most Workloads competitive and in 3D unbeatable.

Jensen surely doesn't sleeps well currently as well hoping to sell some GTX 1050 bricks fast ;)

Sagittaire
26th March 2017, 13:32
Yeah actually i wonder also how they could achieve it especially with the IGPU overhead on practical the same node but it's Intel ;)

If they don't make it AMD looks really in the lead at least without IGPU overhead.


Raven Ridge is becoming more and more exciting to look out for for Laptops the more Zen data spreads :D

be carefull in these previews, IGPU for intel is in off mode (configuration use GTX 1080 in most case). Moreover in idle mode, iGPU certainely don't use more than 5 watts and will not change efficiency result for CPU.

CruNcher
26th March 2017, 14:50
I think there is a big possibility that Raven Ridge is gonna land here

http://www.cpubenchmark.net/cpu.php?cpu=Intel+Core+i3-7350K+%40+4.20GHz&id=2930

but with much much better GPU Performance of course :)

Also Sagittaire that 5 GHz are reachable for Intel already it surely takes not so much of fine tuning and optimization to reach it stable for the 6 Coffe Lake Cores surely you cant expect them to reach 5 GHz with full threading though

https://youtu.be/dxgrb0YR_fA?t=309

seeing such experimental results means they working hard in the lab to tighten them ;)

i wouldn't be surprised if they would dethrone Raven Ridge again 2018 with that 6 Core Mobile im actually pretty sure they will this is how the cycle works if you 1 Generation behind.

But at least the Prediction of AMD 2 Generations behind i wouldn't share it's 1 Generation 1 Cycle pretty perfect for staying alive and share the market :)


Richards next FCAT Compares are out looking at the imho super nice Ryzen 7 1700

https://www.youtube.com/watch?v=-RRt5WkVxuk

AMD made a good decision putting two on top that have slightly higher performance but aren't worth it to catch zillion of more buyers for the Ryzen 7 1700 that is a great Strategy that's gonna succeed :)

everyone will see the Value of the Ryzen 7 1700 and understand how fair it is


Especially how adaptive you can use that system and save even more resources not going all the time with the full 8 cores but switching between 6 and 8 cores depending on the Workload for now until the 8 cores get throughout used ;)


i would actually pretty much disable 2 cores and save the energy and use only the 6 cores at first for most Workloads and use those 2 extra cores only as backup till they widely low level get used on the consumer side of things ;)

https://youtu.be/-RRt5WkVxuk?t=592

He should know better that Intel is already preparing the answer and it's running in the Lab and might be even sooner released then first anticipated, intel will look closely at its market share tumbling now and might push a lot of money and scale their schedule based on that for a pretty fast release ;)

It's funny to see Reviewers that don't understand the Game they Playing in ;)

also these Video Encoding compares on the Software side with those archiving targets i find them economically blatantly idiotic for avg consumer and even for semi professional content producer but that would be worth another topic.

This is also a reason i find it pretty unfair how the R&D inside of Kaby Lake and Balancing decision gets judged by most "Pro Reviewers" overall it's so blatantly wrong and it gets pretty easy to sell those Gamer target something which is technically very questionable overall even if others want to make you believe the investment is worth it with so blatantly wrong arguments ;)

Especially if those Reviewer use Workloads from the last 20 century and don't understand Platform Balance overall.

Sagittaire
26th March 2017, 17:05
Also Sagittaire that 5 GHz are reachable for Intel allready it surely takes not so much of fine tuning and optimization to reach it stable for the 6 Coffe Lake Cores surely you cant expect them to reach 5 GHz with full threading though

Certainely for extreme overclocking but not for stock frequency.

- If you want make that with 6C/12T Kaby Lake at 5 Ghz stock, you have certainely something like 160-180 watt power consumption (and perhaps higher).
- If Intel make 6C/12T Kaby Lake at 5 Ghz stock, then the power efficacity will be really low, really lower than Sandy Bridge at stock frequency (i7-2600K and i7-7700K are on par for x264 power efficiency)
- For each new architecture, Intel annonce generaly 10% more IPC and not more.

I don't think that Intel can make quickly a new 6C/12T CPU better than R7 1800X. Anyway Intel can make certainely really quickly a 8C/16T Kaby Lake CPU (i7-7900K at 4.0/4.2/4.5 with TDP at 140 watts for exemple) certainely really faster than R7 1800X in all case.

NikosD
26th March 2017, 17:09
Intel can't do anything right now.

They can only sit down and watch their new HEDT Skylake X platform getting destroyed by AMD 16C/32T HEDT CPU and that 6C/12T is going to face Zen 2.

So...

Sagittaire
26th March 2017, 17:24
They can only sit down and watch their new HEDT Skylake X platform getting destroyed by AMD 16C/32T HEDT CPU and that 6C/12T is going to face Zen 2.


there are rumor for R7 1950X 16C/32T on desktop at 1000$ ... :D

samples circulate for validation in professional area

CruNcher
26th March 2017, 18:05
It's a Dual Socket System rarely any AVG Consumer Target though some Gamers are surely crazy even if most Numa nodes would idle all the time and consuming power for nothing some Gamers would surely buy those believing they get something from it ;)

i have to say i understand more Gamers that buy overpriced laptops then these Systems for Gaming purposes hell even would tell them buy some all color glowing Razer shit instead and be happier seeing that colors flow instead of seeing your nodes idle :D

Which we though currently see very often with Ryzen as well, according to youtube data ;)

Even with Semi Pro targeted Video Editing you will still experience that ;)

NikosD
26th March 2017, 19:25
No.

AMD HEDT is not a dual socket system.

It's a single and dual socket system.

You can go single if you want to.

A dual socket system with two 16C/32T CPUs means 64 Threads.

That's a crazy number of threads even for professional/workstation use.

It's in the territory of a serious server system.

CruNcher
26th March 2017, 19:38
No.

AMD HEDT is not a dual socket system.

It's a single and dual socket system.

You can go single if you want to.

A dual socket system with two 16C/32T CPUs means 64 Threads.

That's a crazy number of threads even for professional/workstation use.

It's in the territory of a serious server system.

Oh come on that's ridiculous i buy a car for ferrari price to run it at 80 mph constantly :D

Even as backup Redundancy for the Future that would be crazy for most AVG user.

Btw

https://youtu.be/79-s8byUkxk?t=1508

hehe Ryan tries to nail Petersen here because of VEGA and Petersen is doing his typical smiling PR bla bla ;)

NikosD
26th March 2017, 19:44
Sorry, but I can't see how a High End Desktop System should be used only on dual socket, otherwise it could be considered as slow ?

Has Intel released a dual socket system for HEDT ?

You think a 16C/32T on a single socket system is like a Fiat instead of Ferrari ?

BTW, Ryan is an Nvidiot.

I rarely pay attention to what he is saying and posting in this thread that video is completely off topic.

I will not loose my time on that.

CruNcher
26th March 2017, 20:09
High End Desktop is useless market that should have never come into existance neither should have overclocking though nowadays finally each overclocker get kicked into the butt like smokers these days by Nvidia themselves and people slowly wake up ;)

It was funny to get Petersen his coment the Kids dont really seem to understand that Vbios times are over and they never gonna get this deep into the system anymore.

Sagittaire
26th March 2017, 20:27
Well Ryzen 16C/32T is stock at 3.1 Ghz and turbo at 3.6 Ghz for 180 Watt with quad chanel RAM. You will have same performance than R7 1700 in single thread mode. You will have at least same performance than r7 1700 in game. You will have real time encoding for x264 in veryslow preset for 1080p24 and really correct speed for x265 in medium preset for 2160p24 source. For the same price than i7-6900K or i7-5960K, I prefer Rysen 16C/32T and by far.

Don't forget that you have game witch like 10C/20T CPU and more:
https://www.computerbase.de/2017-03/amd-ryzen-1800x-1700x-1700-test/4/#diagramm-watch-dogs-2-fps

you can find in computer base test (in deutch ... ;-) all the more threaded game in the area.

NikosD
28th March 2017, 18:33
Rigaya was very kind to compile flops.c under MS VC 2017, which like MS VC 2013 is incapable to produce vectorized code, only scalar.

But the FPU SIMD performance of Ryzen is still under investigation, for me at least.

I haven't managed to find proper answers to my questions yet.

NikosD
30th March 2017, 07:30
It shows interesting data about RyZen.
Up to 8 threads RyZen 7 in integer calculation acts like SandyBridge+. When you put more load on SMTs then efficiency increases beyond S.

Without this synthetic benchmark you would be still living in dream land regarding AVX2 performance. Hint: x265 likes AVX2 (256bit) alooooot.



2) one more time, in real application like x264 (we are on doom9 after all), rysen outperfom sandybridge, IvyBridge and Haswell and by far.


What do you think about RyZen's IPC in x264 & x265 apps ?

Haswell's IPC is 30%-40% better than Sandy in x264 and 50%-70% better in x265 and RyZen manages to catch Haswell's IPC on both apps.

I can't believe that RyZen has equally implemented AVX2 integer SIMD unit with Haswell, because AVX2 integer for Haswell is 256bit and RyZen's AVX2 integer is 128bit (should be)

So, what's going on ?

NikosD
2nd April 2017, 10:14
So, what's going on ?

I have to reply to myself with new data on my original post.

I did an analysis based on the --asm SIMD parameter of x264 app and the results are surprising (as often happens lately)

AVX2 gives ~5% on average for both Haswell and Kabylake, so it's actually insignificant.

SSSE3 for example gives ~8% for all platforms tested.

The huge speed up comes from MMX2 (about x3 for Sandy and x3.3 for RyZen and Kaby) and SSE2fast (from 50% to 70% over MMX2)

All the other SIMD sets do not play an important role and some of them have regressions e.g AVX2, LZCNT, BMI2 on RyZen.

Anyway, read my new data and analysis here:
https://forum.doom9.org/showthread.php?p=1800781#post1800781

Sagittaire
2nd April 2017, 13:39
I have to reply to myself with new data on my original post.

I did an analysis based on the --asm SIMD parameter of x264 app and the results are surprising (as often happens lately)

AVX2 gives ~5% on average for both Haswell and Kabylake, so it's actually insignificant.

SSSE3 for example gives ~8% for all platforms tested.

The huge speed up comes from MMX2 (about x3 for Sandy and x3.3 for RyZen and Kaby) and SSE2fast (from 50% to 70% over MMX2)

All the other SIMD sets do not play an important role and some of them have regressions e.g AVX2, LZCNT, BMI2 on RyZen.

Anyway, read my new data and analysis here:
https://forum.doom9.org/showthread.php?p=1800781#post1800781

you compare SIMD cumulation? Like this?

ffmpeg\ffmpeg.exe -i Sample\Exodus_UHD_HDR_Exodus_draft.mp4 -an -f rawvideo - | x265\x265.exe --input-res 3840x2160 --fps 23.976 - -o Output\x265_2160p.265 --input-depth 10 --output-depth 10 --crf 24 --preset medium --tune grain --ssim --psnr --asm MMX2 --frames 100

ffmpeg\ffmpeg.exe -i Sample\Exodus_UHD_HDR_Exodus_draft.mp4 -an -f rawvideo - | x265\x265.exe --input-res 3840x2160 --fps 23.976 - -o Output\x265_2160p.265 --input-depth 10 --output-depth 10 --crf 24 --preset medium --tune grain --ssim --psnr --asm MMX2,SSE,SSE2,SSE2fast --frames 100

ffmpeg\ffmpeg.exe -i Sample\Exodus_UHD_HDR_Exodus_draft.mp4 -an -f rawvideo - | x265\x265.exe --input-res 3840x2160 --fps 23.976 - -o Output\x265_2160p.265 --input-depth 10 --output-depth 10 --crf 24 --preset medium --tune grain --ssim --psnr --asm MMX2,SSE,SSE2,SSE3,SSSE3 --frames 100

ffmpeg\ffmpeg.exe -i Sample\Exodus_UHD_HDR_Exodus_draft.mp4 -an -f rawvideo - | x265\x265.exe --input-res 3840x2160 --fps 23.976 - -o Output\x265_2160p.265 --input-depth 10 --output-depth 10 --crf 24 --preset medium --tune grain --ssim --psnr --asm MMX2,SSE,SSE2,SSE3,SSSE3,SSE4,SSE4.1,SSE4.2 --frames 100

ffmpeg\ffmpeg.exe -i Sample\Exodus_UHD_HDR_Exodus_draft.mp4 -an -f rawvideo - | x265\x265.exe --input-res 3840x2160 --fps 23.976 - -o Output\x265_2160p.265 --input-depth 10 --output-depth 10 --crf 24 --preset medium --tune grain --ssim --psnr --asm MMX2,SSE,SSE2,SSE3,SSSE3,SSE4,SSE4.1,SSE4.2,AVX --frames 100

ffmpeg\ffmpeg.exe -i Sample\Exodus_UHD_HDR_Exodus_draft.mp4 -an -f rawvideo - | x265\x265.exe --input-res 3840x2160 --fps 23.976 - -o Output\x265_2160p.265 --input-depth 10 --output-depth 10 --crf 24 --preset medium --tune grain --ssim --psnr --asm MMX2,SSE,SSE2,SSE3,SSSE3,SSE4,SSE4.1,SSE4.2,AVX,AVX2 --frames 100

ffmpeg\ffmpeg.exe -i Sample\Exodus_UHD_HDR_Exodus_draft.mp4 -an -f rawvideo - | x265\x265.exe --input-res 3840x2160 --fps 23.976 - -o Output\x265_2160p.265 --input-depth 10 --output-depth 10 --crf 24 --preset medium --tune grain --ssim --psnr --asm MMX2,SSE,SSE2,SSE3,SSSE3,SSE4,SSE4.1,SSE4.2,AVX,AVX2,FMA3 --frames 100

ffmpeg\ffmpeg.exe -i Sample\Exodus_UHD_HDR_Exodus_draft.mp4 -an -f rawvideo - | x265\x265.exe --input-res 3840x2160 --fps 23.976 - -o Output\x265_2160p.265 --input-depth 10 --output-depth 10 --crf 24 --preset medium --tune grain --ssim --psnr --asm MMX2,SSE,SSE2,SSE3,SSSE3,SSE4,SSE4.1,SSE4.2,AVX,AVX2,FMA3,FMA4 --frames 100

ffmpeg\ffmpeg.exe -i Sample\Exodus_UHD_HDR_Exodus_draft.mp4 -an -f rawvideo - | x265\x265.exe --input-res 3840x2160 --fps 23.976 - -o Output\x265_2160p.265 --input-depth 10 --output-depth 10 --crf 24 --preset medium --tune grain --ssim --psnr --asm MMX2,SSE,SSE2,SSE3,SSSE3,SSE4,SSE4.1,SSE4.2,AVX,AVX2,FMA3,FMA4,XOP --frames 100

NikosD
2nd April 2017, 13:44
Yes the idea is to use --asm but not with all those architectures.

I did that only for x264, not x265 and I didn't see all of them in latest versions for my Sandy & Haswell systems.

Only those I mentioned above and you don't need to write them all with commas.

For example, when you write --asm avx2, that by default means MMX2, SSE2FAST, SSSE3, SSE4.2, AVX, FMA3, AVX2

The app automatically adds all those SIMD sets

ShogoXT
5th April 2017, 06:21
I had to change motherboards to fix my problems. Also finally reformatted so this is on a fresh install.
http://i.imgur.com/pBnqewc.jpg
Asrock x370 Taichi BIOS 1.94A BETA
AMD Ryzen 1700 @ 3.8ghz 1.25 voltage
Corsair LPX 16gb 3200mhz Hynix
MSI RX 480 Gaming X

Check out this image I found on reddit. Power consumption from voltage makes a HUGE difference. From this I tried to do 1.25 or 1.275 @ 3.9ghz but it was a no go for me. I wanted the highest I could get at 1.2ish range voltage. That 1 volt though... :D
https://www.overclockers.ru/images/lab/2017/03/25/1/03.png

Finally I also tested by playing battlefield 1 on OBS 1080p60. Cpu preset is very high, clearly I should try Faster. WARNING: I suck.
https://www.youtube.com/watch?v=Zi9lk06HEjU

I can bench more if needed.

CruNcher
5th April 2017, 21:14
Predicatively the new Intel 6 Core 6/12 should reach ~28 fps it should be only slightly lower then the X above Ryzen

NikosD
6th April 2017, 12:20
There is no 6core Intel.

There will be no 6core Intel for 2017.

Andouille
6th April 2017, 15:28
There is no 6core Intel.

http://processors.specout.com/d/m/Hexa-.-Core

NikosD
6th April 2017, 15:36
http://processors.specout.com/d/m/Hexa-.-Core
You haven't been following this thread obviously or you are trying to make wrong impressions

We are not talking about HEDT platform/CPUs, but mainstream desktop.

Cruncher wrote the "new" Intel 6core

burfadel
7th April 2017, 11:41
You haven't been following this thread obviously or you are trying to make wrong impressions

We are not talking about HEDT platform/CPUs, but mainstream desktop.

Cruncher wrote the "new" Intel 6core

Probably referring to the future Coffee Lake CPU's. X265 on Ryzen could have noticeable gains in the future with better memory support, new bioses, performance driver built in to Windows (or part of a driver package, and Ryzen optimised code.

The performance driver will likely be part of a Windows update, where CPU drivers are usually now part of. A large part of the driver will simply achieve the same as doing this now:
https://community.amd.com/community/gaming/blog/2017/04/06/amd-ryzen-community-update-3

Try the new power plan :).

Sagittaire
9th April 2017, 01:53
some result on my new benchmark


|----------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| CPU (Ghz) | x264 | x265 | LAVC | auto | MMX2 | SSE | SSE2 | SSE3 | SSE4 | AVX | AVX2 | All |
|----------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| i5-3550@3.50 | 7.12 | 1.04 | 48.0 | 0.84 | 0.35 | 0.35 | 0.54 | 0.58 | 0.86 | 0.86 | N/A | N/A |
| i7-5960X@4.40 | 24.03 | 4.38 | 150.0 | 3.50 | 1.01 | 1.01 | 1.64 | 1.78 | 2.78 | 2.80 | 3.43 | N/A |
| R7 1700@3.75 | 21.62 | 3.39 | 109.0 | 2.67 | 1.12 | 1.08 | 1.63 | 1.75 | 2.34 | 2.26 | 2.37 | N/A |
|----------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| R7 1700@stock | 18.51 | 2.98 | 103.0 | 2.32 | 0.96 | 0.95 | 1.41 | 1.54 | 2.11 | 2.05 | 2.08 | N/A |
| R7 1700X@stock | 20.61 | | | | | | | | | | | |
| R7 1800X@stock | 21.62 | | | | | | | | | | | |
| i7-7700K@stock | 14.02 | | | | | | | | | | | |
| i7-5960X@stock | 17.70 | 3.24 | 112.0 | 2.55 | 0.75 | 0.75 | 1.20 | 1.32 | 2.06 | 2.09 | 2.52 | N/A |
| i7-6900K@stock | 19.83 | | | | | | | | | | | |
| i7-6950K@stock | 22.10 | | | | | | | | | | | |
|----------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|


You can see that MMX, SSE, SSE2, SSE3 and SSE4 have really high speed with ryzen (higher than intel at same clock)
AVX2 is really powerfull with intel and makes it possible to close the gap with Ryzen

cojj
10th April 2017, 07:20
Sisoftware recently did a good benchmark review with instruction-sets:

http://www.sisoftware.eu/2017/04/05/amd-ryzen-review-and-benchmarks-cpu/

NikosD
10th April 2017, 09:24
A really bad comparison of a 8C/16T RyZen with a previous generation 4C/8T Skylake and a 6C/12T Haswell.

They should put at least some IPC figures.

Στάλθηκε από το SM-N910C μου χρησιμοποιώντας Tapatalk

NikosD
12th April 2017, 13:30
According to this table http://rigaya34589.blog135.fc2.com/blog-entry-916.html?sp it seems that Skylake has 50% faster integer add/sub and 100% integer mul/shifts than Haswell.

Compared to RyZen, some instructions are 4x faster, others 2x, a few less than 100% and for very few, Skylake is a little slower than RyZen.

Generally speaking, integer SIMD architecture of Skylake is faster than FP SIMD architecture compared to RyZen.


Στάλθηκε από το SM-N910C μου χρησιμοποιώντας Tapatalk

CruNcher
26th April 2017, 22:55
some result on my new benchmark


|----------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| CPU (Ghz) | x264 | x265 | LAVC | auto | MMX2 | SSE | SSE2 | SSE3 | SSE4 | AVX | AVX2 | All |
|----------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| i5-3550@3.50 | 7.12 | 1.04 | 48.0 | 0.84 | 0.35 | 0.35 | 0.54 | 0.58 | 0.86 | 0.86 | N/A | N/A |
| i7-5960X@4.40 | 24.03 | 4.38 | 150.0 | 3.50 | 1.01 | 1.01 | 1.64 | 1.78 | 2.78 | 2.80 | 3.43 | N/A |
| R7 1700@3.75 | 21.62 | 3.39 | 109.0 | 2.67 | 1.12 | 1.08 | 1.63 | 1.75 | 2.34 | 2.26 | 2.37 | N/A |
|----------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|
| R7 1700@stock | 18.51 | 2.98 | 103.0 | 2.32 | 0.96 | 0.95 | 1.41 | 1.54 | 2.11 | 2.05 | 2.08 | N/A |
| R7 1700X@stock | 20.61 | | | | | | | | | | | |
| R7 1800X@stock | 21.62 | | | | | | | | | | | |
| i7-7700K@stock | 14.02 | | | | | | | | | | | |
| i7-5960X@stock | 17.70 | 3.24 | 112.0 | 2.55 | 0.75 | 0.75 | 1.20 | 1.32 | 2.06 | 2.09 | 2.52 | N/A |
| i7-6900K@stock | 19.83 | | | | | | | | | | | |
| i7-6950K@stock | 22.10 | | | | | | | | | | | |
|----------------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|---------|


You can see that MMX, SSE, SSE2, SSE3 and SSE4 have really high speed with ryzen (higher than intel at same clock)
AVX2 is really powerfull with intel and makes it possible to close the gap with Ryzen

Its not much worth first you need to understand Intels CPU is bottlenecked by a integrated GPU if Intel wouldn't have it it would be even way more efficient.

It's AMD that are inefficient because they left it away investing those resources into other parts saving the transistors physical space ;)

And Intel will next beat AMD with 6 Cores + IGPU again ;)

While AMD will only have 4 Cores + IGPU by then though still overall a more efficient GPU Core ;)

but then you have to look at the whole System architecture and understand that AMD has something very efficient if they combine resources modular.

4 Cores + IGPU + External AMD GPU all interfacing nicely with each other that is Powerfull in the overall System Architecture when workloads are efficiently distributed async and throughout the hardware scheduling it itself.

Sagittaire
28th April 2017, 20:25
And Intel will next beat AMD with 6 Cores + IGPU again ;)


calculation is really simple because scaling is in practice perfect with frequency and core number for x264 speed:

i7-7700K 4C/8T Kaby Lake:
4.3 Ghz for max turbo on all cores
14.02 fps in my test for x264 encoding
110 W for power in this test

R7 1800X 8C/16T Ryzen:
3.7 Ghz for max turbo on all cores
21.62 fps in my test for x264 encoding
128 W for power in this test

if you want find i7-7800K 6C/12T potential performance it's really simple:

"i7-7800K 6C/12T" Kaby Lake:
4.3 Ghz for max turbo on all cores
21.03 fps in my test for x264 encoding
165 W for power in this test

this "i7-7800K" is just on par with R7-1800X but never Intel can make this CPU at this frequency because power is by far too high. For have acceptable power, "i7-7800K" must have 140 W in this test, and it's more 4.0 Ghz for frequency in this case.

"i7-7800K 6C/12T" Kaby Lake:
4.0 Ghz for max turbo on all cores
19.56 fps in my test for x264 encoding
140 W for power in this test

aegisofrime
16th May 2017, 07:20
Hi all,

AMD just released a new compiler, AMD Optimizing C/C++ Compiler (AOCC). Can someone try compiling x265 with it, see how it performs?

http://developer.amd.com/tools-and-sdks/cpu-development/amd-optimizing-cc-compiler/

NikosD
16th May 2017, 12:03
Hi all,

AMD just released a new compiler, AMD Optimizing C/C++ Compiler (AOCC). Can someone try compiling x265 with it, see how it performs?

http://developer.amd.com/tools-and-sdks/cpu-development/amd-optimizing-cc-compiler/

Probably you should post that info to x265 encoder thread.

Andouille
31st May 2017, 01:54
There will be no 6core Intel for 2017.

http://ark.intel.com/products/123589/Intel-Core-i7-7800X-Processor-8_25M-Cache-up-to-4_30-GHz

NikosD
31st May 2017, 03:25
http://ark.intel.com/products/123589/Intel-Core-i7-7800X-Processor-8_25M-Cache-up-to-4_30-GHz
I don't know how many times I have to reply to you the same thing



You haven't been following this thread obviously or you are trying to make wrong impressions

We are not talking about HEDT platform/CPUs, but mainstream desktop.

Cruncher wrote the "new" Intel 6core

NikosD
5th June 2017, 19:10
Delayed to February 2018

http://wccftech.com/intel-coffee-lake-delayed-2018-8th-gen-kaby-lake-refresh/

mandarinka
5th June 2017, 20:27
Delayed to February 2018

http://wccftech.com/intel-coffee-lake-delayed-2018-8th-gen-kaby-lake-refresh/

That was a mistake, on saturday taht incriminating post was already edited to say "later this year". The Raja guy probably misremmebered/misunderstood something and somebody with NDA knowledge requested him to edit it.

Also Intel's fact sheet says "coming soon", so I'm inclined to believe they will manage to bring them out in august or september.

plonk420
10th August 2017, 11:51
to be honest, intel says pricing for their 6c12t is "$383.00 - $389.00"

ryzen 7 1700X is $350, 1800X is $460

i really AM hoping they put the heat on Intel, tho.

ShogoXT
3rd October 2017, 18:19
https://www.phoronix.com/scan.php?page=news_item&px=Ryzen-Segv-Response

Has anyone seen much about the Ryzen segfault issue? Id like to hear your opinions, seems to mostly effect GCC on Linux if im reading correctly, but it could be more.

I have had a lot of issues with my platform and have tried different motherboards and RAM. I have Samsung B die ram now, but I CANNOT get it up there, at least without using the external clock generator, and have it boot. It might be the memory controller on the CPU itself.

My 1700 also requires too much voltage past 3.9ghz so im kinda stuck on 3.8 @ 1.36v . It generates a TON of heat while streaming and encoding. Was getting up to 95c streaming PUBG at 1080p 60 fps with x264 faster setting on OBS. Replaced the cooler and its down a little bit, but sometimes reaches 85c still.

Considering the new AMD compiler, let me know if you want me to try other benchmarks when I try to tweak this further. I may RMA my 1700 as well to fix the segfault and get better OC results. As people have said they have noticed much greater headroom after RMA.

burfadel
4th October 2017, 02:34
You're not gong to be able to go much higher than that regardless. It's not a RMA issue because overclocking doesn't count, and faults like that are common in processors, including Intel, most you are unaware. They can be resolved with microcode updates. Needless RMA'ing just means they'll recover the cost through increased prices. This is especially true as your argument is you just can't overclock as much as you like. Even more to the point, it's a 1700 not 1700X, so even though it's not locked unlike non-K Intel CPU's there is no guarantee that you can run it above rated speed. The difference between 3.9 and 3.8 is a fraction over 2 percent, it's not worth quibbling over. Also sounds like you were using the stock cooler, that thing is literally only rated to stock CPU speeds like Intel coolers. RMA is only intended to replace truly faulty parts, not any damage caused by the user (such as running at 95 C when overclocking).

They're releasing Zen+ which will be available in about 5 months along with a new 400 series chipset (X470). It will have a higher headroom and apparently be on an improved process and higher clocks. If you want faster upgrade then :).

ShogoXT
4th October 2017, 13:31
Yes they might do a bios fix at some point, but so far they DO indeed RMA for this specific issue. I've simply been reading other people's reports over the last month.

The results are if your build date is before 1725, you have the segfault issue likely. After that date you probably don't have the issue. The epyc and threadripper platforms are on newer process it seems and do not have the issue as well.

Reports note that these do versions do have better results on ram support and the voltage wall. One person had as much as 4ghz @ 1.25v which is quite impressive. Others show similar improvements.

This is what AMD offers willingly if you provide proof them and work with their support. They will ask what settings you run at and to fiddle with bios settings. I personally would prefer to avoid possible long term file corruption, so why not do it?

ShogoXT
5th October 2017, 16:56
http://www.overclockersclub.com/reviews/intel_core_i7_8700k__core_i5_8400/7.htm

Equal to ryzen on x264 and better by a bit on x265. Will depend on price, but count on possible problems due to it being rushed and requiring power changes.

Atak_Snajpera
5th October 2017, 18:36
8700k@4.3 GHz is about 25% faster than 1700@3.7GHz
http://i.cubeupload.com/c6s6ss.png

Source -> https://youtu.be/v5EioeVC_QY?t=2m17s

ShogoXT
5th October 2017, 19:33
Yea it was one of my fears. I knew ryzen was never going to be equal to intel on core performance due to AVX2 usage in x265, but I havent been doing real encoding projects lately. Just encoding and gaming on Twitch using x264.

I will feel salty though if they suddenly support streaming with x265...

On the other hand Z370 is pretty much a dead end platform. They actually had to change the pin layout so much on the socket that you CANNOT use Skylake and Kaby Lake on Z370, do to the power changes. They should have standardized a new socket, but clearly they were rushed from Ryzen.

Next year will be the Coffee Lake 8 core CPU on Z390 chipset, likely with similar changes.

Ryzen AM4 socket will be usable until 2020. My hope is because of 12NM LP or 7NM LP fab, high clocks plus a greater core count through 6 core CCXs will be possible on the same AM4 socket. Just conjecture though...

Atak_Snajpera
6th October 2017, 12:35
However if you also use some filtering in AviSynth (QTMC / MDegrian) then this difference will be much smaller because most filters use only SSE2.

LigH
6th October 2017, 13:34
Next chapter of Threadripper competitors: the German IT magazine c't tested Core i9-7960X (16 cores) and even i9-7980XE (18 cores!). Skat, anyone?!

Regarding Cinebench 15, there was only little advantage over R-TR 1950X. Despite twice the price. But it may be a different result for more AVX* heavy duties.

Atak_Snajpera
6th October 2017, 13:40
i9-7980XE@4GHz should achieve 80 fps in my benchmark ;) More than the fastest EPYC 32C/64T.

ShogoXT
6th October 2017, 17:10
https://www.youtube.com/watch?v=f7BqAjC4ZCc

Warning if you consider that 18 core and overclock it at all. The release motherboards are completely inadequate for handling it. You need to make sure the motherboard has more than the standard 8 pin CPU connector. Ideal 8+4 or 8+8. Choose your psu with that in mind.

Vrm cooling is an issue too as a lot of them would hit 85c then throttle.

Mobo vendors like ASRock released new versions like the x299 taichi XE a week ago.
https://www.asrock.com/mb/Intel/X299%20Taichi%20XE/index.us.asp

nevcairiel
6th October 2017, 17:44
That video is from June, a lot has happened since then, and plenty new boards were released - plenty with more then a single 8-pin. Also, the VRM topic was blown quite a bit out of proportion. Certainly its far more important on those CPUs then ever before (including ThreadRipper, it also uses a lot of power), but if you OC you should always be aware to cool all important components adequately - that includes the VRM. Putting a small active fan there to cool them is usually plenty to avoid any throttling, for example, or using a full-cover water block.

Asmodian
6th October 2017, 23:31
i9-7980XE@4GHz should achieve 80 fps in my benchmark ;) More than the fastest EPYC 32C/64T.

It should be much better than that even, I am getting 103 fps with a 7900X @ 4.7 GHz.

Actually clocks are a bit more complicated; 4.8 GHz on two cores, 4.7 GHz on the rest, and -200 MHz for AVX heavy workloads. During a benchmark run the cores spend a lot of time at both 4.8/4.7 and 4.6/4.5 GHz. Still, a 7980XE@4.0GHz should be really good for x264 and x265. :)

Atak_Snajpera
7th October 2017, 12:20
It should be much better than that even, I am getting 103 fps with a 7900X @ 4.7 GHz.
With 1920x1080 and default medium x265 profile?!?! I DON'T think so!

NikosD
7th October 2017, 13:02
From the first reviews, it seems that 18 core Core i9 is a TDP 200W processor at default speeds and with a toothpaste between the CPU and IHS.

Overclocking that power hungry processor at 5.5GHz leads to 800W (!!) power consumption.

Not a lot people would go that high, but generally speaking OC of Core i9 above 12 cores is a difficult and probably dangerous task.

It needs a very good water cooling solution even at default speeds, you can imagine what it needs for overclocking.

sneaker_ger
7th October 2017, 13:29
It needs a very good water cooling solution even at default speeds
Says who? Not Intel.

Asmodian
7th October 2017, 16:25
With 1920x1080 and default medium x265 profile?!?! I DON'T think so!

Sorry, by "my benchmark" I thought you meant your x264 one because I couldn't find your x265 benchmark in your signature link.

Of course, I only get 54.3 fps with your x265 benchmark. :o

http://i1222.photobucket.com/albums/dd496/asmodian3/x265%20Benchmark_zpsahdotluw.png

ShogoXT
23rd November 2017, 13:15
Hey has the newer versions of x265 gotten any better on Ryzen in performance?

I was just curious if any optimizations have been made or if it's still around Intel 4c/8t range.

I don't expect much because of avx2, but was hoping even a little boost vs March reviews.

LigH
23rd November 2017, 13:23
I don't find any x265 patches (https://bitbucket.org/multicoreware/x265/commits/branch/default) which mention "Ryzen" in their descriptions.

Also I do not remember anyone explaining which kind of NUMA thread pooling would be optimal to circumvent the design flaws.

burfadel
23rd November 2017, 16:14
There is a new microcode that will be available across most Ryzen boards shortly, it should give a nice little boost to those speeds and hopefully better support for high speed memory :). Such to the extent that there will be a very noticeable difference between using DDR4-2400 RAM and DDR4-3333 RAM, much more than the i7-8700K. Also keep in mind the i7-8700K costs considerably more than the Ryzen 1700.

ShogoXT
23rd November 2017, 18:21
I don't find any x265 patches (https://bitbucket.org/multicoreware/x265/commits/branch/default) which mention "Ryzen" in their descriptions.

Also I do not remember anyone explaining which kind of NUMA thread pooling would be optimal to circumvent the design flaws.

http://www.tomshardware.com/reviews/amd-ryzen-threadripper-1950x-game-performance,5207-2.html

Here is what I know.
Ryzen cores are layed out in CCXs. They are up to 4 cores depending on if cores are disabled.
Ryzen dies are made up of 2 CCXs for up to 8 cores.
6 core 1600s are CCXs with 1 core disabled in each CCX. 3+3
Ryzen 1200 is 2+2.
CCXs are connected with infinity fabric that runs at half memory speed usually maxing out in the 3200 ram speed range of effectiveness.
2133 and 2400 really penalizes system performance.
Raven ridge is 1 CCX+ Vega.

Threadripper is 4 desktop ryzen dies connected like epyc, but with 2 dead dies.
According to the link above, the more you move around data between dies the more it really slows down.
Note that "game mode" completely disables a die so this does not happen.

If latency matters a lot (I don't know) for x265 then you want to keep most of the work between the 4 cores inside each CCX.

I don't know how you identify real cores from logical threads let alone cores from their different CCXs though in software.

ShogoXT
23rd November 2017, 18:23
There is a new microcode that will be available across most Ryzen boards shortly, it should give a nice little boost to those speeds and hopefully better support for high speed memory :). Such to the extent that there will be a very noticeable difference between using DDR4-2400 RAM and DDR4-3333 RAM, much more than the i7-8700K. Also keep in mind the i7-8700K costs considerably more than the Ryzen 1700.

I did hear they are completely redesigning the microcode from the ground up for Raven ridge and future CPUs for the sake of being modular.

Asmodian
24th November 2017, 04:48
If latency matters a lot (I don't know) for x265 then you want to keep most of the work between the 4 cores inside each CCX.

I don't know how you identify real cores from logical threads let alone cores from their different CCXs though in software.

This is very interesting for future CPU designs from both companies, going to smaller modular dies and architectures is more efficient (cost and material wise) so more layers of NUMA awareness might be very useful. Instead of two NUMA nodes for the 1950X you could have four "NUMA level 1" nodes, one for each CCX, then two "NUMA level 2" nodes, one for each die. A dual socket system would have two "NUMA level 3" nodes, one for each CPU.

This way the OS and/or x265 could schedule threads that share data on the lowest NUMA level available or put independent threads as far away as possible. Does anyone know if something like this is already possible or planned? It seems key for future performance with the wide interest in multiple dies and modular architectures.