View Full Version : Computex 2018: AMD TR 32 core CPU , Intel 28 core 5ghz CPU announced


ShogoXT
6th June 2018, 20:15
https://www.anandtech.com/show/12906/amd-reveals-threadripper-2-up-to-32-cores-250w-x399-refresh

https://www.anandtech.com/show/12907/we-got-a-sneak-peak-on-intels-28core-all-you-need-to-know

Looks like Intel lost the core advantage for their x299 platform, but they used some kind of phase change water cooling for it. I imagine it requires a lot of power and cooling to keep 5ghz.

AMD on the other hand looks like they finally got rid of the dummy dies and went nearly full Epyc, besides the PCIE lanes and memory channels. The slide mentions it has the Zen+ improvements which include XFR2 and Precision Boost 2, both have which increased single core performance for Ryzen quite a bit. Works on existing socket.

Only issue still is preventing programs and threads from crossing the glue too much.

ShogoXT
6th June 2018, 20:30
Also there is currently a debate whether or not the Intel 28 core 5ghz is even feasible, there is a good chance they caught wind of the 32 core TR shortly beforehand and panic pushed a demo rig together with hidden chillers just to take some of the hype away from AMD.

Will there be a product to ship?

Atak_Snajpera
7th June 2018, 10:37
Gamer Nexus nicely summed up this fake 5GHz on all 28 cores
https://youtu.be/tRH0-QwhvVQ

foxyshadis
8th June 2018, 05:41
The Anand article is honestly more informative, despite being rushed to production. And it's not fake, it's just a total custom jobbo (even the motherboard!) with unobtanium, to the point that they were able to choose the top five overclockable cpus off the factory line, pair them with a custom mobo and a top-end watercooler.... and still only one could hit it. That really said it all. They say a lot of vague things, hope everyone salivates and builds the product up in their head.

It is an exciting demonstration of the limits of technology, if nothing else.

I just hope nobody actually thinks they'll be able to buy this, save for mortgaging their house to buy this specific box.

Atak_Snajpera
8th June 2018, 12:17
Cinebench (SSE2) test runs on that machine only for few seconds so maintaining 5GHz during that period wasn't that difficult. CPU basically didn't have enough time to warm up ;)
I bet that 28 core @ 5GHz would not survive few hours of encoding in x265 (4k + preset veryslow + AVX2). That 5GHz would most likely drop to base level in couple of minutes ;)
Whole system could even crash with forced AVX-512. Yes. So for me 5GHz on that CPU is fake. Extremely desperate move in reaction to Threadripper 32core.

NikosD
8th June 2018, 12:36
This is by far the most explanatory video of what really happened during that "28core 5.0GHz" demo.

ROFL
https://www.youtube.com/embed/ozcEel1rNKM?feature=oembed

foxyshadis
11th June 2018, 09:04
This is by far the most explanatory video of what really happened during that "28core 5.0GHz" demo.

ROFL
https://www.youtube.com/embed/ozcEel1rNKM?feature=oembed

If posts could be stickied, this would be it.

ShogoXT
19th June 2018, 20:33
http://translate.google.com/translate?hl=en&sl=auto&tl=en&u=https%3A%2F%2Fwww.hkepc.com%2F16912%2F%25E7%258D%25A8%25E5%25AE%25B6__32%25E6%25A0%25B8%25E5%25BF%258364%25E7%25B7%259A%25E7%25A8%258B4GHz_Boost_AMD_Ryzen_Threadripper_2990X%25E7%258D%25A8%25E5%25AE%25B6%25E6%259B%259D%25E5%2585%2589

6300 in Cinebench.

Atak_Snajpera
27th June 2018, 12:33
Rumours say that 2990x will start from just 1500 euro! (18 core intel costs almost 2000 euro!)
https://pclab.pl/zdjecia/artykuly/blind/2018/06/ryzen/ryzen582.jpg

Looks like we will have affordable video encoding beast ;)

NikosD
27th June 2018, 15:59
This is simply suicidal for AMD.

I think it's going to be above 2000$ for the 32C/64T SKU.

ShogoXT
28th June 2018, 10:41
Overclockers have noted that it is unwise to overclock it on a normal motherboard as most vrms of x399 are insufficient beyond stock clocks.

https://youtu.be/aKe7CnZT9ZE

Haven't had a chance to watch it yet, but here's one of their opinions.

Atak_Snajpera
28th June 2018, 10:59
Do you really have to overclock 32 cores??? 3.4 GHz on all cores is enough.

NikosD
4th August 2018, 09:14
Well, it seems that official release of The Monster (TR2) is very close, August 13th.

The Beast of 32C/64T will run at 4.0GHz on ALL cores using air-cooling - a huge cooler from Cooler Master and AMD.

It's going to reach more than 6000 running Cinebench.

Intel's response leaked that is going to be announced around October and released on November-December but it's not going to add more cores - up to 18C/36T just with better clocks.

So, with a performance of an overclocked EPYC 32C/64T and without competition from Intel, AMD doesn't really sell TR2 at 1800$

AMD donates TR2 to the people!

https://videocardz.com/77031/amd-ryzen-threadripper-2990wx-2970wx-2950x-and-2920x-specs-and-pricing-leaked

huhn
4th August 2018, 13:16
double the price of the 16 core variant so the margin should be the same thanks to AMD glue there is no different in yield. that's how you make money the price is transparent.

NikosD
4th August 2018, 16:47
You don't seem to understand the concept.

If it was Intel, the price should be 10000$.

They sell Xeon 28C/56T at that price.

Cary Knoop
4th August 2018, 18:25
32 cores for $1800, and more lanes than Intel, another game changer.
AMD did it again!

huhn
4th August 2018, 21:09
because there are CPU with unreasonable prices doesn't mean AMD have to do this right now.

of cause they could go intel extreme they did this before but this is how to take market share and the TR socket never got "over" priced CPU yet.

i don't even want to know how expesive this 28 core xeon is in production compared to the glue AMD CPUs.

NikosD
7th August 2018, 08:24
You are unsuccessfully trying to cover Intel's and Nvidia's madness regarding profit from their products.

They also share a mad product segmentation and unethical tactics for decades, in order to milk more profit.

Yes, we all know that AMD is a company too but I doubt they would go that way even if they were in the position of Nvidia and Intel.

A Xeon at 10.000$ doesn't cost 10.000$ because of its production cost, but because it says "Intel" in front of Xeon.

huhn
7th August 2018, 15:33
where did i said it will cost 10k i only said is far more expensive than a TR and that's a given not even defending intel here that would be a bad try anyway...

and to be fair this over priced "thing" is a server CPU and epyc is not know to be priced much better the flag ship there cost 4-5k.

TR should be compared with 2066 like AMD is doing right now.
the flag ship Intel Core i9-7980XE get's stomped and is more expensive than a TR2.

these high end epyc and xeon CPU are quite different just alone the multi CPU support...

what so ever what i take from this is that using glue is a far better way to get high core counts than huge die sizes.

NikosD
9th August 2018, 08:44
In order to have an equal overclocking performance of Intel's 28C@5.0GHz and TR2 32C@5.0GHz using Cinebench, some people managed to push TR2 at this frequency using liquid cooler (nitrogen)

The results as you can see are impressive, TR2 is faster even than Intel's "secret weapon with an unknown CPU" hiding in the closet.

TR2 could come up with a Cinebench score around 8000 vs 7330 of Intel.

https://www.tweaktown.com/articles/8699/amds-threadripper-unboxing-overview/index3.html

huhn
9th August 2018, 14:12
these liquid nitrogen garbage overclocks are just a waist of time. totally not practical just people having fun wasting power.

the stock speeds are impressive.

NikosD
9th August 2018, 15:33
The liquid nitrogen based overclocking is very impressive to see, it's like a show on its own.

Also, it's a kind of extreme overclocking to reach the limits of the architecture and remove any thermal performance limitations.

Intel started a war by water-cooling with a huge forbidden cooler in many countries their CPU in order to break a record in Cinebench and AMD just replied.

With their new aircooler CoolerMaster Wraith they can push 32C/64T at 4.0GHz all-core which is a huge achievement at 250W.

The rest is just for the performance crown.

And don't forget TR2 is a real CPU that you can order right now if you want.

Intel's CPU looks like an illusion...

FranceBB
9th August 2018, 23:38
Preface:


Intel actually has the 98% of the server market share, mostly because Xeon CPUs have been fast and reliable for quite some time.
Besides, even though AMD released its "new" server products, Intel Xeon CPUs are still faster, mostly because instructions, memory handling and load distributions are implemented in a way that they handle certain type of loads faster in a multithreading environment and also in a single thread one.
For instance, afaik the fastest AMD CPU is the Epyc 7601 which has 32 core, 64 threads, a maximum frequency of 3.2 GHz, 64 MB of cache L3 and instructions 'till AVX2.
The Intel opponent is the Intel Xeon Platinum 8176 which has 28 core, 56 threads, a maximum frequency of 3.80 GHz, 38.5 MB of cache L3 and instructions 'till AVX-512.
Please note that the lack of AVX-512 instructions in AMD CPUs can also play an important role in favor of Intel for certain type of calculations.


Demonstration:


In order to demonstrate that Intel is faster than AMD, I report the results from Cinebench, the famous benchmarking program that many people use (including Linus Sebastian) and other benchmarking tools:

Cinebench R11.5, 64bit (Single-Core):

https://i.imgur.com/rM20PLo.png

As you can see, single thread calculations show that Intel Xeon is way faster than AMD Epyc when it comes to using cores singularly.
This is kinda expected, because of the way the CPU is implemented, but you'll be surprised to find out what the multi-thread result shows.

Cinebench R11.5, 64bit (Multi-Core):

https://i.imgur.com/mDIutr8.png

Despite having less shared L3 cache, less cores and less threads, Intel Xeon is still slightly faster than AMD Epyc.

Passmark CPU Mark:

https://i.imgur.com/bYpGfS2.png

As you can see, even Passmark shows that Intel Xeon is faster than the AMD Epyc in multi-threading calculations.

Geekbench 3, 64bit (Single-Core):

https://i.imgur.com/ordvM6x.png

Once again, AMD is noticeably slower than Intel on single-core calculations.

Geekbench 3, 64bit (Multi-Core):

https://i.imgur.com/DGBfeyU.png

And Geekbench shows that Intel is slightly faster than AMD on multi-threading calculations as well.


Final Conclusion:


In other words, even though Intel CPUs are kinda overpriced, they perform better than the AMD ones, due to their internal design, their implementation and also because of AVX-512 instructions set support that lacks in AMD CPUs.
Intel has always provided good, fast and reliable hardware for servers and workstations, that's why pretty much every company decided to pick Intel CPUs for their workloads.
Xeon CPUs don't come cheap, but you get what you pay for.
Unfortunately, though, the consumer side doesn't reflect the server-side scenario.
Intel CPUs for consumers (i3, i5, i7) are not as good as the Xeon ones, they didn't have many cores in the past and they have never been cheap.
This is mostly because Intel didn't really have any competitor for years, that's why we started to think about i3 as 2c/4th, i5 as 4c/4th and i7 as 4c/8th.
AMD CPUs didn't quite manage to reach the same performances of Intel CPUs since the days of dual core.


A bit of history:


Back when CPUs were mono-core, AMD had a particular way of handling calculations and memory management in their chips and they managed to get more calculations per cycle than Intel CPUs.
Having more calculations per cycle granted AMD a significant part of the market 'cause many people decided to get an AMD CPU.
As result, to keep the peace against AMD, Intel increased the frequency of their CPUs, but this led to an increase in voltages as well, which led to an increase in temperatures 'cause air coolers didn't quite manage to keep them very cool.
On the other hand, AMD had the same performances with less frequency, less voltage and temperatures were fine.
They also started naming their CPUs with the Intel equivalent frequency number; for instance, an AMD Athlon 3200+ at 2.2GHz was "the same" as an Intel processor with the same characteristics, but running at 3.2GHz.
Back in the days (2003-2004 if I recall correctly), a $200 Athlon 64 3200+ chip performed pretty much the same as an $800 Pentium 4 3.2 GHz, but it was cheaper and it spent less energy, so it was a "go for it".
Once again, this was because AMD managed to get more calculations done per CPU cycle.
When Intel understood that raising the frequency wasn't the right answer, they came up with the "multi-core" idea and they released their first multi-core CPU.
At the very beginning, it was slower than a normal single thread CPU in common daily scenarios, 'cause programs weren't aware of the second core, so they used Core0 only, while Core1 was sitting there in idle, doing pretty much nothing.
Anyway, it turned out to be good to do many tasks at the same time, due to the fact that users were able to divide the workload and it somehow found its way to the public.
Eventually, programmers started implementing programs in a different way to make use of the second additional core and this gave a great speed boost.
In the meantime, AMD released other single-thread CPUs, like the AMD Athlon 3700+, which were fine except for the fact that once the developers made their programs aware of the second core, AMD single-thread CPUs couldn't cope with the Intel ones.
This forced AMD to make its own dual core CPU, but because of the way AMD CPUs were implemented, they didn't work well in a dual core scenario.
The peculiar production way that granted them success back when they were mono-core allowing more calculations per cycle somehow worked against them when they had to implement dual core.
Later on, Intel continued developing their own CPUs and 4cores CPUs came out (Dual Core Duo) and this led them to develop their multi-core architecture further and further implementing multi-threading and naming their CPUs i3, i5, i7.
AMD tried to fight Intel in many ways with a failure after another, like the Intel Phenom II x6 (a 6core 6 thread CPU that had lower performances than an Intel i5 4c/4th), the famous AMD Bulldozer FX that worked very badly with multi-thread and didn't divide the work-load really well, having really bad performances in pretty much every common scenario and so on.
Enthusiast were more and more prone to buy Intel and the company raised prices without introducing many features in the consumer CPU.
They did develop other features, but they didn't release them by purpose 'cause there was no reason to: they had the whole market for themselves, so they just focused on the enterprise-side which demanded faster and better CPUs.
After many failures to reach Intel and after deluding enthusiast consumers year after year, AMD decided to dedicate their new CPUs to the entry-level consumer market, including a decent GPU inside their processors called APU.
APU found a positive response by customers that didn't want to spend much and couldn't afford to buy an Intel CPU and a dedicated GPU.


Nowadays:


Nowadays, AMD finally dropped their "peculiar" way of making CPUs in favor of an Intel-like implementation, featuring what recalls the Intel multi-threading; that's why Ryzen and Epyc came out.
However, this whole way of making CPUs is new to AMD engineers while Intel engineers have done it for years and years, so it's "normal" that Intel CPUs are more fine-tuned than the AMD ones, however this is good for the market because now that Intel finally has a rival, it's forced to either release new products or lower their prices.


My "story" as consumer:
I've been running Intel CPUs in the 90s 'till the new millennium, when I moved to AMD. I've been a satisfied AMD customer for four years, buying a CPU after another and enjoying their great success for single-thread CPUs, but then after 2004 I got pissed off year after year, especially after I bought the Phenom II x6 (6c/6th) thinking that it was going to compete with the Intel i7 CPUs, but I ended up encoding things slower than people with an i5 CPU. After that, I moved to Intel; I've been an Intel customer ever since and I'm not planning to go back to AMD anytime soon.

NikosD
10th August 2018, 08:03
...After that, I moved to Intel; I've been an Intel customer ever since and I'm not planning to go back to AMD anytime soon.

It seems that you moved to Intel for good, I think you could actually work for them in a way.

Where to start from ?

Intel has 98% of server market when last year had 99+%

EPYC managed to gain almost 2% which is not that big, but even Intel's CEO (former) said that is struggling to keep AMD below 15% - 20% which is close to the biggest share AMD ever had during Opteron days - 25%

Intel's 28Cores processor costs 10.000$ while AMD's 32Cores processor costs around 4.000$

AMD supports 2 TB RAM per socket versus 768 GB RAM for Intel.

The memory advantage is huge for AMD.

Also Intel supports 48 PCIe lanes in single socket versus 128 PCIe lanes for AMD.

The PCIe lanes advantage is huge for AMD.

The spec_int and spec_float benchmarks which are more important for server market than Cinebench and passmark (!) favor once again AMD.

The second half of 2018 will bring AMD to 5% share at the end of the year for server market.

Server market is a difficult market and needs time to validate and trust new platforms.

Due to the broken Intel's 10nm process which will not probably be fixed even in latest 2019, AMD will have no rival in server market for 2019 when it will release its new EPYC 2 server CPU in Q1 at 7nm and 48C/96T initially and 64C/128T later.

It's a 28 cores vs 64 cores battle with no luck for Intel.

AMD, as former Intel's CEO has already said, will reach more than 20% at the end of 2019 in server market because big companies have already samples of EPYC 2 using 7nm in their hands and can see the difference in speed compared to ancient and slower Xeons in 14nm.

And the price is always much better for AMD.

Now regarding AVX512, there is simply no market for these instructions yet.

Intel and other reviewers are struggling to find benchmarks and real apps to leverage AVX512 but with no luck, besides one or maybe two.

You seem to skip HEDT category (high end desktop) which is our topic here for a reason.

On Monday August 13th, AMD will release officially Threadripper 2 or Threadripper 2000 series with 32C/64T and 4.0GHz all-core speed using air-cooling.

Intel simply doesn't have an answer to this category because its best CPU is just a 18C/36T and will stay this way till the end of the year and probably next year too.

On August 31th, AMD will release a new 16C/32 CPU and on October 12C/24T & 24C/48T

These AMD Threadripper processor have simply no rival from Intel and will dominate the HEDT category.

And of course AMD has the PCIe lanes advantage for Threadripper, exactly like EPYC vs Xeon.

Finally, regarding mainstream desktop CPUs, AMD has the core number advantage of 8C/16T vs 6C/12T for Intel which can only win at 1080p gaming for 6% (which doesn't really matter) and some light threaded apps.

All multithreaded-aware apps favor AMD which is the trend nowadays and of course 1440p or 4K gaming is exactly the same between Intel and AMD.

For the next year Ryzen 2 is coming based on the same Zen 2 architecture using 7nm like EPYC 2 and rumors already are talking about 12 cores or even 16 cores for mainstream desktop.

I see no good future for Intel at least for 2019, like most people actually.

Blue_MiSfit
10th August 2018, 08:47
I'm super impressed with the new 32 core Threadripper systems.

I don't do a ton of encoding / highly threaded loads at home, but if I did I'd be looking at a Threadripper of some sort!

As it stands my most performance sensitive application is Lightroom, mainly when working with high megapixel (often stitched) images from my Nikon D810 DSLR. Lightroom heavily favors single threaded performance, especially from Intel, so I have a watercooled i7-7700k.

That being said, the 8 core boost frequency on the new Threadripper is pretty cool - when you're in "game mode" it clocks 8 cores up quite high so you get good single threaded performance. Regardless, I bet the Intel offerings will still be faster for Lightroom :devil:

I absolutely love how AMD has been coming up with such innovative, disruptive products, though. I definitely can see them taking a nice chunk of the cloud server market with Epyc, and maybe a decent chunk of the HEDT market with Threadripper.

Atak_Snajpera
10th August 2018, 11:49
Let's be honest even Intel knows that they won't be able to compete with upcoming Epycs in 7nm (64C/128T)
Interesting read
https://semiaccurate.com/2018/08/07/intel-has-no-chance-in-servers-and-they-know-it/

In other words, even though Intel CPUs are kinda overpriced, they perform better than the AMD ones, due to their internal design, their implementation and also because of AVX-512 instructions set support that lacks in AMD CPUs.
Forget about avx-512 because it overheats CPU like crazy!
https://networkbuilders.intel.com/docs/accelerating-x265-the-hevc-encoder-with-intel-advanced-vector-extensions-512.pdf

StvG
10th August 2018, 11:50
... Nowadays, AMD finally dropped their "peculiar" way of making CPUs in favor of an Intel-like implementation, featuring what recalls the Intel multi-threading;...

Both companies have different implementations? Zen is CCX based, Intel is monolithic.
With this (https://en.wikipedia.org/wiki/Intel_UltraPath_Interconnect) Intel could make something Zen-like (up to 40 cores?).

Groucho2004
10th August 2018, 12:05
As it stands my most performance sensitive application is Lightroom, mainly when working with high megapixel (often stitched) images from my Nikon D810 DSLR. Lightroom heavily favors single threaded performance, especially from Intel, so I have a watercooled i7-7700k.
Odd. I use Adobe Camera Raw (through Photoshop) for my Nikon raw files which is basically the same as Lightroom and it clearly uses all 4 cores of my i5 2500K, whether I make adjustments or convert a bunch of raw files to tif/jpeg.

ShogoXT
12th August 2018, 12:48
I always thought that AVX 512 was the future, but not so much anymore. Fab density is increasing and this has forced more cooling. There is no way Intel can keep pushing x86, high clocks, and complex instruction sets in a world that is going toward parallelization.

nevcairiel
12th August 2018, 12:57
Its the same as it was with AVX2, the initial implementation produces too much heat, the next ones will be refined, especially with a die shrink, and it'll be much more valuable.
AVX512 is also not only about 512-bit instructions, it also adds a whole lot of useful features to 128-bit and 256-bit instructions that should allow faster and smarter code in the future, without the heat penality.

But it'll take some more time for AVX512 support to be a bit more wide-spread and less heat-dependent and software to fully understand to properly use it without incurring any of the penalities for it to become effective.

NikosD
12th August 2018, 13:40
AVX2 had no such problems in the first implementation of Haswell.

The turbo modes of AVX2 were lower than any other mode, but this hasn't changed since then.
It's not that it was improved.

AVX512 market is largely covered by other more parallelized hardware like GPUs.

Intel added 4 special instructions regarding AI called DLBoost in next generation Cascade Lake server CPUs as a superset of AVX512, but we all know that AI and ML/DL is a market for GPUs.

The GP part of GPGPU is becoming more and more general and this is a trend that looks like is not going to stop for HPC and all the other aspects of AVX512 that could be useful.

huhn
12th August 2018, 22:10
is there even a heat problem with AVX512?

seriously does it even matter that the CPU is using a lower clock as long as the program runs faster? so be it...
getting a cooler that will keep it working even with AV512 and full boost clocks so not a hard too and trivial with deliding.

and if there is any real problem with heat on intel right now it's there thermal "paste".

nevcairiel
12th August 2018, 22:34
AVX2 had no such problems in the first implementation of Haswell.

The turbo modes of AVX2 were lower than any other mode, but this hasn't changed since then.
It's not that it was improved.

Actually it was improved. In Haswell the entire CPU (ie. all cores) would clock down if any core used AVX2. Since Broadwell this is no longer the case, and only the core actually running AVX2 is clocking down if needed.
This resulted in a steep performance penality when using AVX2 in Haswell, which had people equally hesitant at first. But now AVX2 is quite useful in a multitude of applications, including Video.

Broadwell was also Intels first CPU to be made on Intels tri-gate transistor process, which helped efficiency and helped to reduce the power/heat requirements of the AVX2 units.

The really dense SIMD compute areas is where process improvements give the most benefit, since they produce the most heat on a small area.

I fully expect those clock offsets to improve with Ice Lake, whenever that comes out. Maybe by then developers have also figured out how to properly utilize it.

NikosD
12th August 2018, 23:24
But if your code doesn't run on all cores, how are you going to gain performance ?

Multithreading is necessary for any modern code, especially for HPC.

So, once again the clock will go down for AVXx.

And Broadwell is actually a non existent CPU for desktop.

Anyway, I think GPUs are nowadays more suitable for these kind of calculations like AVX512 and Top500 supercomputer list, clearly shows this exactly.

huhn
13th August 2018, 00:47
that's the point it's supposed to be faster when correctly used even with the lower clocks. there is lot's lot's of stuff that can't be paralysed
why even waist your time on creating a 32 core CPU if a 1080 has 2560 "cores" for the same reason and much more.

AVX512 is used in x265 do GPU do this better?

where does this total blind assumption come from that it is something GPU will be better at.

i wouldn't be shocked if zen2 will support it..

Atak_Snajpera
13th August 2018, 11:30
Its the same as it was with AVX2, the initial implementation produces too much heat, the next ones will be refined, especially with a die shrink, and it'll be much more valuable.
after 2020 maybe ;)

i wouldn't be shocked if zen2 will support it..
If we are lucky zen2 will have 4xFMAC128 instead of 2xFMAC128. This means that AVX-512 will require 2 cycles like in SkyLake-X 7800x

seriously does it even matter that the CPU is using a lower clock as long as the program runs faster? so be it...
getting a cooler that will keep it working even with AV512 and full boost clocks so not a hard too and trivial with deliding.
The issue is that it may not run faster ;) Even Intel in his document does not recommend using avx-512 in x265.

They had to disable 24 and 20 cores respectively in order to show some gains. Yeah that makes sense. You buy $10k CPU and use 4 or 8 cores in video encoding.
https://s22.postimg.cc/h3nyhv5wx/Untitled-1.png

NikosD
13th August 2018, 12:18
@huhn

Oh man.

How many unreasonable phrases in one post.

AVX512 is a specialized instruction set used in very very very rare cases, especially in an efficient way.

For x265 even Intel suggests to enable AVX512 under certain circumstances or else it's not worth it or even worse it's slower than AVX2

AVX512 is just a special case of Intel's propaganda due to the fact that it can't produce more general purpose cores.

So, it can't win the match fair in general purpose silicon and it has already started to make things useless like AVX512 look as something important and necessary.

Unfortunately, this kind of propaganda works as we see people in x265 thread to buy Intel CPUs for AVX512 to use it with x265!

Poor guys...

The reason that AMD delivered many cores is because obviously there are many cases that general purpose computing can be parallelized and run on many threads.

But obviously this is not the case for AVX512.

Of course AMD will provide support for these kind of instruction sets, like always

But it is more important and necessary to implement fast SSEx and AVX/AVX2 than AVX512

So, simple.

huhn
13th August 2018, 12:21
so we just ignore the benefits on the 10 core CPU?

and again you can workaround the clockspeed with a simple bios setting and proper cooling.

From Figures 1 and 2 we can make the following inferences:
• For desktop and workstation SKUs (like the Intel Core i9-
7900X processor that we tested), Intel AVX-512 kernels
can be enabled for all encoder configurations, because the
reduction in CPU clock frequency is rather low.
• For server SKUs (like the Intel Xeon Platinum 8180
processor on which we tested), the frequency dip is higher
and increases with more cores being active. Therefore,
Intel AVX-512 should only be enabled when the amount
of computation per pixel is high, because only then is the
clock-cycle benefit able to balance out the frequency
penalty and result in performance gains for the encoder.
Specifically, we recommend enabling Intel AVX-512 only
when encoding 4K content using the slower or veryslow
preset in the main10 profile. We do not recommend
enabling Intel AVX-512 kernels for other settings
(resolutions, profiles, or presets), because unexpected
inversions with respect to using the Intel AVX2 kernels
may result.
Conclusion
This paper described our experience with accelerating
x265 and open-source HEVC encoders, with Intel AVX-512
instructions. From our experience we recommend that
for workstation and client CPUs that have Intel AVX-512
instructions, the kernels may be used across all profiles of
x265. However, for server-grade CPUs that have the Intel
AVX-512 instructions, this acceleration should be used only
for certain profiles of x265 that focus on encoding high
resolution video (4K and higher) in the main10 profile using
the slower or veryslow presets due to the impact that the
Intel AVX-512 instructions have to clock frequency. For other
profiles on server CPUs with Intel AVX-512 instructions,
enabling these kernels is not recommended.

so this read like a not recommend for you?

Atak_Snajpera
13th August 2018, 12:31
so we just ignore the benefits on the 10 core CPU?
Yes because We are living now in 16-32 core HEDT world. These days 10 core HEDT CPU for $1k is sooo last age. ;) Ryzen 3700 will have more cores than that old 7900x.

huhn
13th August 2018, 12:42
@NikosD

yes 512 is not that important right now like AVX2 wasn't important when it was released.

if you would just read why AVX512 looses speed on the 8180 you would instantly see that AVX512 can be used with as many cores as present the problem is just a design flaw from intel.

but calling it propaganda is such much easier...

ShogoXT
14th August 2018, 10:51
https://www.overclock3d.net/reviews/cpu_mainboard/amd_threadripper_2950x_and_2990wx_review/11

Reviews are out. 2990wx is worse than the 2950x even in x264 and x265. In nearly every other workload it beats the core i9.

Edit: Actually I was hasty. Seems hit and miss. 2950 looks good though. People are blaming memory usage or thread utilization on Windows. Linux is a bit better.


https://www.phoronix.com/scan.php?page=article&item=amd-linux-2990wx&num=5

Atak_Snajpera
14th August 2018, 12:04
Well you need 4 instances of x265 to saturate all those 64 threads. Those amateurs just run handbreak and expect to have full utilization.

ShogoXT
14th August 2018, 17:26
Well you need 4 instances of x265 to saturate all those 64 threads. Those amateurs just run handbreak and expect to have full utilization.

Does that x265 benchmark run multiple instances? Or is it just ripbot that does that? I post on a few tech forums so I could get at least one to consider updating their test setup.

Atak_Snajpera
14th August 2018, 18:01
My x265 FHD Benchmark (http://forum.pclab.pl/topic/1184884-x265-FHD-Benchmark/) runs 5 instances at once.

Blue_MiSfit
14th August 2018, 18:32
Odd. I use Adobe Camera Raw (through Photoshop) for my Nikon raw files which is basically the same as Lightroom and it clearly uses all 4 cores of my i5 2500K, whether I make adjustments or convert a bunch of raw files to tif/jpeg.

Bulk exports, sure. Interactive editing is still heavily dependent on single threaded performance. This has been getting better, but it's still a big deal.

Groucho2004
15th August 2018, 03:50
Bulk exports, sure. Interactive editing is still heavily dependent on single threaded performance. This has been getting better, but it's still a big deal.
As I mentioned, whatever I do in ACR, be it adjusting exposure, sharpening/noise reduction, etc., all of the 4 cores are being used, quite evident by looking at the CPU usage in Task Manager.
Disabling all but 1 core in the BIOS makes interactive editing very slow and choppy, confirming the advantage of multiple cores.

However, after some research it seems that Lightroom/ACR fails to efficiently utilise more than 4 cores which means that the new, shiny ThreadRipper won't be of much use in this context.

Atak_Snajpera
15th August 2018, 14:20
Looks like there are some serious issues with scheduler in windows. 5 instances of x265 running and 2990WX just chokes
https://p.xfastest.com/~sinchen/GIGABYTE-X399-AORUS-XTREME/GIGABYTE-X399-AORUS-XTREME-66.jpg

Source -> https://www.xfastest.com/thread-221870-1-1.html

ShogoXT
15th August 2018, 14:22
My x265 FHD Benchmark (http://forum.pclab.pl/topic/1184884-x265-FHD-Benchmark/) runs 5 instances at once.

Im trying to bug some reviewers about it right now.

Question: How many does Ripbot264 usually do?

I hope to see more benchmarks made out of VPX and whenever AV1 encoders are finally stable. H265 patent problems have really slowed down the video compression world and we will have to catch up quick after AV1 gets around.

Atak_Snajpera
15th August 2018, 14:30
Question: How many does Ripbot264 usually do?
It is up to the user. You can run even 16 instances if you are working with some low resolution footage (DVD)
http://i.cubeupload.com/e8JoyV.png

ShogoXT
18th August 2018, 03:22
Is that x264 benchmark still around? Does it have multiple instances as well?

For comparing AMD to Intel streaming power x264 is usually used. So im usually curious about that too.

NikosD
18th August 2018, 09:09
A few comments regarding the 32C/64T release of TR2.

It seems that is a niche CPU, focused on specific workloads like rendering, encoding, compiling even more than 16C release.

Not only because of doubling the number of cores but also due to the specific implementation.

The compromise of drop-in compatibility to the existing X399 motherboards, led to reduced Infinity Fabric bandwidth and increased latency of compute dies with a performance impact on memory intensive workloads.

I saw a huge performance advantage over the new 2950X 16C CPU (72%) and 18C Intel's CPU using the new Blender Beta benchmark similar to Corona, POV-ray, Cinema4D etc but also saw a performance drop below 2950X (!) level in certain tasks.

Windows scheduler makes things worse but fortunately (and unfortunately because MS is involved) I read this paragraph in a review:
AMD continues working with Microsoft to route threads to the die with direct-attached memory first, and then spill remaining threads over to the compute dies.

Unfortunately, the scheduler currently treats all dies as equal, operating in Round Robin mode.

As a result, even moderately-threaded applications can suffer at the hands of high memory latency and low throughput.

This is further complicated by thread migration.

According to AMD, Microsoft has not committed to a timeline for updating its scheduler

32C/64T desktop performance for 1800$ is unforseen, but it needs to pay attention to the workload.

But who can beat 16C/32T of 2950X at 900$ ?

Atak_Snajpera
18th August 2018, 12:33
Is that x264 benchmark still around? Does it have multiple instances as well?
no just single instance. I'm planning to update x265 FHD Benchmark so each instance (except 5th) will run on specific NUMA node.

NikosD
19th August 2018, 17:54
Redmond, we've got a problem...
(unless the problem is the poor developer's NUMA implementation)

7zip compression benchmark running on 2990WX (32C/64T)

@Atak_Snajpera We definitely need this NUMA aware version of your x265 benchmark.

https://s22.postimg.cc/4yxav5g4x/7zip.png

Cary Knoop
19th August 2018, 18:03
Redmond, we've got a problem...
(unless the problem is the poor developer's NUMA implementation)

7zip compression benchmark running on 2990WX (32C/64T)

@Atak_Snajpera We definitely need this NUMA aware version of your x265 benchmark.

https://s22.postimg.cc/4yxav5g4x/7zip.png
The question is whether Redmond and other software vendors were notified in a timely fashion by AMD.

NikosD
20th August 2018, 06:25
The platform of Windows hasn't been called "Wintel" by accident.

Linux distributions didn't have previous notice in order to implement a proper NUMA aware OS.

And Ryzen processors based on Zen architecture have been released since March 2017 and Threadripper processors since August 2017.

It's been a year already.

huhn
20th August 2018, 07:46
what has windows to gain from intentionally slowing down AMD CPUs...

this CPU is a niche product with low priority on a consumer grade OS.

linux/unix OS are used on super computer and mainframes where windows has no market share to speak of and this is a typical CPU used there and NUMa is quite important there.
if this has anything todo with this result. because numa is supported since win7.

NikosD
20th August 2018, 15:14
what has windows to gain from intentionally slowing down AMD CPUs...

Money from Intel, I wrote it above.
Intel is trying to buy (literally) some time by all means (literally) in order to fix their 10nm broken process and to be competitive again with a new architecture.

It's not the first time for Intel and Nvidia, they are full in unethical moves in order to dominate their markets.

How do you think they become this big ?


if this has anything todo with this result. because numa is supported since win7.

The problem is obviously in the implementation of NUMA and Windows scheduler.

Windows are always ready for Intel processors, in a very strange way.

Like x265, in a strange way too.

Atak_Snajpera
20th August 2018, 15:31
It is not really a x265 fault that by default it uses all numa nodes. The issue is that AMD released another odd cpu after AMD FX.
I do not know if you remember but before patch for win7 FX-8150 was being detected as 8C/8T CPU. This meant that all integer units were treated as full cores. After patch AMD FX 8xxx were "downgraded" to 4C/8T.

https://www.anandtech.com/show/5448/the-bulldozer-scheduling-patch-tested

M$ should implement better way of detecting topology of those newer "non-standard" cpus. Currently scheduler is clearly not aware what NUMA has direct access to memory.

amichaelt
20th August 2018, 16:00
Money from Intel, I wrote it above.

Which makes no sense. Microsoft pulls in twice as much revenue as Intel as has at least 10 times as much cash on hand. They aren't exactly struggling to make ends meet.

NikosD
20th August 2018, 16:18
It is not really a x265 fault that by default it uses all numa nodes. The issue is that AMD released another odd cpu after AMD FX.
I do not know if you remember but before patch for win7 FX-8150 was being detected as 8C/8T CPU. This meant that all integer units were treated as full cores. After patch AMD FX 8xxx were "downgraded" to 4C/8T.

https://www.anandtech.com/show/5448/the-bulldozer-scheduling-patch-tested

M$ should implement better way of detecting topology of those newer "non-standard" cpus. Currently scheduler is clearly not aware what NUMA has direct access to memory.

Yes, AMD is the underdog and needs sometimes to make "odd" CPUs.

But, in a very strange way, every "odd" CPU from AMD is odd for MS and SW vendors but every odd CPU from Intel is like an everyday CPU.

Look at the collaboration of Intel and x265 developers.
They worked together to optimize it for AVX 512.

But the same piece of software is not optimized for Ryzen, we are still waiting for it one and a half year after the release.

With Threadripper and its NUMA architecture is worse of course.

The FX processor was once again a lot faster in Linux out of the box versus Windows.

Intel makes the rules for everything on CPUs, but if it wasn't AMD we would stay to Pentiums and x86 architecture another 10 years at least.

Like the scheme of Core i3 2C/4T, Core i5 4C/4T and Core i7 4C/8T.

That scheme lasted from 2011 until Ryzen.

Oh and the HEDT category.

You could buy a 10C/20T Core i7 6950X for 1800$ before Threadripper, which is the cost for the 32C/64T from AMD right now.

Think about it.

Intel is the most ridiculous tech company along with Nvidia of course.

huhn
20th August 2018, 18:35
what do you want intel to do not optimise for there architecture
?
what should x265 do ignore help should they ignore there biggest customer base?
ryzen supports AVX2 and it is slower than the intel version and that'S the newest instruction.

you are aware that intel cpu are not rarely run better on linux too just like amd CPUs...

there is a lot shady stuff happening and most can't be seen here. but this is not shady.

what new CPU feature should they focus on? AVX512 was already in CPUs in 2015. amd even retired there 3Dnow!.

NikosD
9th October 2018, 20:47
So...The mythical monster chip that Intel was hiding in the closet is just a poor old Skylake X/Cascade Lake (?) 28C/56T CPU, unlocked and overclocked in order to justify its skyrocketed price and power consumption.

Intel will release it on December though, so we have to wait a little more to see who is going to justify again Intel's decisions regarding new CPUs.

huhn
7th November 2018, 14:55
zen 2 looks like it is fixing quite some issues that zen has.

up to 8 core per die 256 bit AVX.

so it's no longer a dream that we my end up with a CPU with a single CCX with no ram latency limitation.

to bad they only showed rome.

Atak_Snajpera
7th November 2018, 17:34
But you will get 16C (2 dies + IO die) for the price of 9900k. Regarding AVX. Yes x265 will work noticeable faster now!

huhn
7th November 2018, 18:55
i'm pretty sure i'll be fine with 8 cores.

nevcairiel
7th November 2018, 19:26
Mainstream just doesn't need 16 cores at this time. It barely needs 8. It would be rather annoying if the only offering they have is two CPU dies, and lower core counts being half-disabled. I rather have an optimized 8-core in one die with no inter-die communication troubles.
IMHO Ryzen should stay 8-cores and anyone that wants 16 can get ThreadRipper.

huhn
7th November 2018, 22:48
if i understood it correctly PCIe is still on the cores so two dies are not welcome to me.

Atak_Snajpera
8th November 2018, 12:23
Mainstream just doesn't need 16 cores at this time. It barely needs 8. It would be rather annoying if the only offering they have is two CPU dies, and lower core counts being half-disabled. I rather have an optimized 8-core in one die with no inter-die communication troubles.
IMHO Ryzen should stay 8-cores and anyone that wants 16 can get ThreadRipper.

Who cares what mainstream needs! Streamers will love 16c ryzen.
Even overclocked 9900k is too slow for 1080p60fps with preset medium. BTW. Few years ago people were also saying mainstream does not need 16 thread cpu. Without AMD you would be stuck in never ending 4 core era. AMD may release 16c ryzen just to show middle finger to intels 9900k.

INTEL: Look AMD we also have 16 thread cpu!
AMD: HOLD MY BEER!

AMD did similar move with 2990WX (32C/64T) this year. In this case middle finger to Intels 18C/36T HEDT.

nevcairiel
8th November 2018, 13:03
Who cares what mainstream needs!

99% of everyone, thats who. We already have 16c CPUs, its called ThreadRipper.

High-end gaming CPUs is where AMD needs to catch up right now, and more cores is not going to do it anymore, not in 2018/19. A 16c Ryzen would be no competition to a 9900k, not for gaming, not unless they can also close the ST gap.
An 8c with higher IPC and higher clock, that would show up a 9900k. Cores is easy, and cores is also not everything.

I'm sure they could easily release a 16c Ryzen if they wanted to. I'm just saying thats not what I want to buy. I want a really fast 8c with no compromises (ie. no half-disabled dies, no inter-die latency problems, etc), with supreme ST performance as well as decent MT, because a lot of things I do are still ST bound.

I don't care about stupid fights who has the most cores or whatever. I care about a CPU that makes sense for my needs, which align with the majority of high-end gamers as it happens.

Atak_Snajpera
8th November 2018, 13:56
Threadripper has 4 channels while Ryzen only 2. So segmentation is maintained. I do not understand why you complain about performance in games. 2700x with improved turbo boost work very well in games. Performance is slower due to weaker 128bit avx unit.(many newer games use AVX). Zen2 fixes that.

huhn
8th November 2018, 15:17
if you bottleneck the GPU ryzen 2 is fine but if you want to run very high FPS like 240/144 hz stable it simply can't compete the latency issues are simply killing it.

and newer games don't use AVX in a way that it is the main speed concern the issue for AMD currently is ram latency and the clockspeed. pre sandy bridge CPU are still performing as expected in new games.

or do you really think OC ram in CPU limited scenarios are magical making AVX much faster on AMD but not on intel and just look at threadripper when it is used with all cores it so unbelievable slow in games.

Atak_Snajpera
8th November 2018, 15:21
and newer games don't use AVX in a way that it is the main speed concern the issue for AMD currently is ram latency and the clockspeed. pre sandy bridge CPU are still performing as expected in new games.

Frostbite engine (for example BFV strongly uses AVX) and Unreal Engine 4 use AVX.
https://software.intel.com/en-us/articles/intel-software-engineers-assist-with-unreal-engine-419-optimizations

huhn
8th November 2018, 16:08
this was march 2018 so at an even later date this engine is used in games and can you remember a game that updated there unity engine and got a major performance boost that can be pointed at AVX? me neither.

how can you claim BFV is using AVX >heavily< if it isn't even out yet.

Atak_Snajpera
8th November 2018, 16:39
this was march 2018 so at an even later date this engine is used in games and can you remember a game that updated there unity engine and got a major performance boost that can be pointed at AVX? me neither.

how can you claim BFV is using AVX >heavily< if it isn't even out yet.

Have you ever heard about open beta? Some guy analyzed BFV in Vtune (if I remember correctly) and noticed that AVX instructions are used a lot more than in BF1. I wouldn't be surprised if new tomb raider was also optimized for avx. Some claim that Project Cars 2 also uses avx.

huhn
8th November 2018, 17:03
the open beta start tomorrow.

and again i'm not saying AVX isn't used i'm just questioning if it a major concern.

battlefield one runs as expected on a sandy bridge CPU which doesn't have AVX2. i'm not aware of a any prove that performance differences between generations is from AVX differences not clockspeed/IPC.

Atak_Snajpera
8th November 2018, 18:40
the open beta start tomorrow.

and again i'm not saying AVX isn't used i'm just questioning if it a major concern.

battlefield one runs as expected on a sandy bridge CPU which doesn't have AVX2. i'm not aware of a any prove that performance differences between generations is from AVX differences not clockspeed/IPC.

You must be living in alternative universe
https://www.ea.com/games/battlefield/news/battlefield-5-open-beta-announcement

https://www.youtube.com/watch?v=cAsyo8gIyys&feature=em-comments

https://i.postimg.cc/y8HXdKs5/Untitled-1.png

huhn
9th November 2018, 02:43
ok sorry this game is not on my rader and never will be on it be i seriously see a big problem in this random comment

so let's assume everything is 100 % AVX code in new games.
because there is no word about AVX2 the problem for zen is AVX2 not AVX 128 bit is not good but fine or at least on par with jaguar.

the next thing... battlefield 5 is a console game what do consoles use as a CPU AMD jaguar what can an AMD jaguar do native AVX not AVX2.

and the PC gaming margin is only 10% maybe 20% it so meaningless compared to console and they will optimise for the console. they will look into the PC port in there spare time.

than we have tests like this and not random comments: https://www.youtube.com/watch?v=U4k_ErEg-FU

and they doesn't even try to get high framerates totally ignoring the possibility to play at lower graphic settings to get your target framerate.

nevcairiel
9th November 2018, 07:30
so let's assume everything is 100 % AVX code in new games.
because there is no word about AVX2 the problem for zen is AVX2 not AVX 128 bit is not good but fine or at least on par with jaguar.

Both AVX and AVX2 are 256-bit. AVX is floating point, AVX2 is integer and FMA. In simplified terms.

Atak_Snajpera
9th November 2018, 13:28
ok sorry this game is not on my rader and never will be on it be i seriously see a big problem in this random comment

so let's assume everything is 100 % AVX code in new games.
because there is no word about AVX2 the problem for zen is AVX2 not AVX 128 bit is not good but fine or at least on par with jaguar.

the next thing... battlefield 5 is a console game what do consoles use as a CPU AMD jaguar what can an AMD jaguar do native AVX not AVX2.

and the PC gaming margin is only 10% maybe 20% it so meaningless compared to console and they will optimise for the console. they will look into the PC port in there spare time.

than we have tests like this and not random comments: https://www.youtube.com/watch?v=U4k_ErEg-FU

and they doesn't even try to get high framerates totally ignoring the possibility to play at lower graphic settings to get your target framerate.

Even AVX1 code runs noticeable slower on zen
https://i.postimg.cc/4xNL0x99/Flops_CPU.png

huhn
9th November 2018, 16:38
Both AVX and AVX2 are 256-bit. AVX is floating point, AVX2 is integer and FMA. In simplified terms.

because you really know this stuff can you elevate the use of the 128 bit version of AVX which is unofficially called AVX 128.

or is there little to nothing you can gain from doing it this way?

nevcairiel
9th November 2018, 18:34
AVX with 128-bit registers has small minor benefits, usually because there happen to be some instrustrictons that may do the job better, and because AVX allows more efficient use of some other instructions (ie. distinct destination registers for some instructions, instead of an implied one). The differences to SSE1/2/3/4.1 versions of the same code (which also are 128-bit) is relatively small, maybe 5-10% (for a single given function), while doubling the register size to 256 can in an ideal case of course be up to twice as fast.

NikosD
10th November 2018, 09:35
Intel's strong PR and marketing teams are trying really hard with tons of money behind them to persuade people that more cores are useless or at least not necessary.

It's the same old story for the dark ages of 2011 - 2017 era, when the same teams were telling us that the Core i3 2C/4T - Core i5 4C/4T - Core i7 4C/8T scheme is the best that we could ever had.

No more than 4 cores for mainstream desktop, ordered Intel.

If you wanted more, you should pay 2000$ to buy 10 cores of Broadwell-E 6950X in 2017.

On March 2017 a world revolution took place.

The revolution of Zen architecture and the first implementation of RyZen processor for mainstream desktop, using 8C/16T in a very affordable price of 350$ - 500$ (initially)

Nowadays, you can buy 32C/64T for 1800$ from AMD, when Intel still sells 18C/36T for 2000$ - even this means 80% more cores for the same price just a year later for Intel.

Cascade-AP, a one-off processor from Intel, mimicks the Zen first generation architecture for servers (EPYC) and rises the sum of cores to 48 from 28 of Skylake-EP using GLUE (as Intel called Infinity Fabric of Zen architecture) to add two Skylake 24C processors to one Cascade 48C.

Intel is trying hard to be as less humiliated as possible compared to the absolute monster of EPYC 2, a multi-die (8+1) CPU of 64C/128T adding PCIe v4.0 and Infinity Fabric 2 to the equation.

Intel guys, please relax...

If you can't avoid something, just enjoy it.

It's coming!

Racer
11th November 2018, 14:30
Looks like there are some serious issues with scheduler in windows. 5 instances of x265 running and 2990WX just chokes
https://p.xfastest.com/~sinchen/GIGABYTE-X399-AORUS-XTREME/GIGABYTE-X399-AORUS-XTREME-66.jpg

Source -> https://www.xfastest.com/thread-221870-1-1.html

Is there actually an updated benchmark with adjusted numa pools or did anybody check if this gives an performance improvement for the 2990WX?

Instance 1 = --numa-pools "+,-,-,-"
Instance 2 = --numa-pools "-,+,-,-"
Instance 3 = --numa-pools "-,-,+,-"
Instance 4 = --numa-pools "-,-,-,+"

Atak_Snajpera
11th November 2018, 18:37
Do you have access to 2990wx?