Log in

View Full Version : x264 and Intel's Media Engine / Sandy Bridge


Pages : 1 2 3 4 5 [6] 7 8 9

kolak
14th January 2011, 13:12
@SeeManRun: thanks for your test results. I think though that for the most optimal results you should upgrade your x264 to latest revision, r1867 that is.
You also didn't specify where the resizing occurs: Avisynth? If yes then which method was used - Spline, Bicubic, Lanczos? Last but not least: which source filter have you used? FFMS2, DGDecNV, DirectShow?
All these conditions may have big influence on overall encoding performance observed...
Please post your full AVS script and Avisynth version used and all above questions will be answered.

It is very funny when people do speed test, by feeding encoders through big avisynth scripts- juts render script into yv12 file :)

Andrew

SeeManRun
14th January 2011, 16:53
@SeeManRun: thanks for your test results. I think though that for the most optimal results you should upgrade your x264 to latest revision, r1867 that is.
You also didn't specify where the resizing occurs: Avisynth? If yes then which method was used - Spline, Bicubic, Lanczos? Last but not least: which source filter have you used? FFMS2, DGDecNV, DirectShow?
All these conditions may have big influence on overall encoding performance observed...
Please post your full AVS script and Avisynth version used and all above questions will be answered.

With an upgraded version (x264 core:112 r1867 22bfd31) I got the following log file. Please take a look and let me know thoughts. Also should note that my CPU is only around 50% used during the test. When I use my own quality settings for blu-ray movies it is 100% on the second pass (yes I use 2 pass encoding).

Let me know what else I can adjust.

poisondeathray
14th January 2011, 17:21
attachments have to be approved , if you copy & paste text, or use a 3rd party free hosting site (e.g. mediafire.com, sendspace.com) there is no wait

SeeManRun
14th January 2011, 19:12
attachments have to be approved , if you copy & paste text, or use a 3rd party free hosting site (e.g. mediafire.com, sendspace.com) there is no wait

Thanks for the info. Do you prefer 3rd party posting or giant messages opposed to approving attachments by chance?

easyfab
14th January 2011, 19:45
Hi, here is a little test ( in french) :x264 VS MediaEspresso et Quick Sync Video with the new SB i5 2300 and i7 2600k

http://www.pcinpact.com/articles/intel-qsv-mediaespresso-x264-video/415-1.htm

shon3i
14th January 2011, 20:10
Here's my little contribution

source: park_joy_1080p.y4m
avs script:
RawSource("park_joy_1080p.y4m",1920,1080,"YV12")
AssumeFPS(24000,1001)
Spline36Resize(1280,720)
Encoders: Mainconcept H264 using CUDA, x264 r1834
Settings: Mainconcept
http://img12.imageshack.us/img12/1313/mcsettings1.th.jpg (http://img12.imageshack.us/i/mcsettings1.jpg/)http://img214.imageshack.us/img214/2465/mcsettings2.th.jpg (http://img214.imageshack.us/i/mcsettings2.jpg/)http://img211.imageshack.us/img211/7046/mcsettings3.th.jpg (http://img211.imageshack.us/i/mcsettings3.jpg/)http://img211.imageshack.us/img211/7046/mcsettings3.th.jpg (http://img211.imageshack.us/i/mcsettings3.jpg/)http://img843.imageshack.us/img843/8770/mcsettings4.th.jpg (http://img843.imageshack.us/i/mcsettings4.jpg/)http://img684.imageshack.us/img684/1676/mcsettings5.th.jpg (http://img684.imageshack.us/i/mcsettings5.jpg/)

x264: preset: superfast, tunings: none, ABR
cabac=1 / ref=1 / deblock=1:0:0 / analyse=0x3:0x3 / me=dia / subme=1 / psy=1 / psy_rd=0.00:0.00 / mixed_ref=0 / me_range=16 / chroma_me=1 / trellis=0 / 8x8dct=1 / cqm=0 / deadzone=21,11 / fast_pskip=1 / chroma_qp_offset=0 / threads=9 / sliced_threads=0 / nr=0 / decimate=1 / interlaced=0 / constrained_intra=0 / bframes=3 / b_pyramid=2 / b_adapt=1 / b_bias=0 / direct=1 / weightb=1 / open_gop=0 / weightp=1 / keyint=250 / keyint_min=23 / scenecut=40 / intra_refresh=0 / rc=abr / mbtree=0 / bitrate=6000 / ratetol=1.0 / qcomp=0.60 / qpmin=0 / qpmax=51 / qpstep=4 / ip_ratio=1.40 / pb_ratio=1.30 / aq=1:1.00

Target bitrate: 6mbps
Encoding logs: Mainconcept
Input:
---------------------
Video source: C:/Users/shon3i/Desktop/test.avs (DirectShow Import)
Output settings:
---------------------
Filename: C:/Users/shon3i/Desktop/test_out.264
VIDEO: H.264/AVC CUDA, NTSC, 1280x720pixels, 23.976pfps, VBR, 6000kbps (max N/A)
multiplexing: Elementary
---------------------

H.264 Encoding Done.
--------------------------------------------------------------------------------
Frames: 500 incoming, 500 encoded
Avg. Bitrate: 5999.40 kbits per second
Time elapsed: 12.35 seconds
--------------------------------------------------------------------------------
Transcoding results:
------------------
Done: C:/Users/shon3i/Desktop/test.avs -> C:/Users/shon3i/Desktop/test_out.264 (0h 00m 12.523s, 39.9 FPS)
-------------------

x264:
x264 [info]: frame I:2 Avg QP:25.30 size:178487
x264 [info]: frame P:254 Avg QP:31.24 size: 49476
x264 [info]: frame B:244 Avg QP:34.03 size: 9280
x264 [info]: consecutive B-frames: 2.8% 94.8% 2.4% 0.0%
x264 [info]: mb I I16..4: 6.9% 40.9% 52.2%
x264 [info]: mb P I16..4: 0.7% 1.7% 0.9% P16..4: 76.6% 0.0% 0.0% 0.0% 0.0% skip:20.0%
x264 [info]: mb B I16..4: 0.2% 0.2% 0.0% B16..8: 11.4% 0.0% 0.0% direct:21.9% skip:66.4% L0:24.9% L1:43.2% BI:31.9%
x264 [info]: final ratefactor: 26.27
x264 [info]: 8x8 transform intra:49.0% inter:29.5%
x264 [info]: coded y,uvDC,uvAC intra: 83.7% 71.6% 54.9% inter: 32.0% 12.9% 3.3%
x264 [info]: i16 v,h,dc,p: 32% 14% 42% 12%
x264 [info]: i8 v,h,dc,ddl,ddr,vr,hd,vl,hu: 16% 14% 30% 5% 7% 6% 6% 6% 9%
x264 [info]: i4 v,h,dc,ddl,ddr,vr,hd,vl,hu: 14% 17% 23% 5% 10% 6% 7% 7% 11%
x264 [info]: i8c dc,h,v,p: 49% 18% 24% 8%
x264 [info]: Weighted P-Frames: Y:2.4% UV:0.8%
x264 [info]: kb/s:5826.40
encoded 500 frames, 41.20 fps, 5826.40 kb/s

Output:
Mainconcept - http://www.mediafire.com/?5c9a4geg809bnyu
x264 - http://www.mediafire.com/?95zzd5338q2rnnx

Enjoy.

poisondeathray
14th January 2011, 20:37
Hi, here is a little test ( in french) :x264 VS MediaEspresso et Quick Sync Video with the new SB i5 2300 and i7 2600k

http://www.pcinpact.com/articles/intel-qsv-mediaespresso-x264-video/415-1.htm



Thanks, Maybe something was lost in the Google translation, but it looks on their tests x264 was actually faster on the "ultra fast" preset, and about the same speed on "very fast preset ?" But it was 1/2 the filesize with better quality ?

But - the rollover screenshots aren't of the same frame. They have a zip file, I'll have a look at that

poisondeathray
14th January 2011, 20:49
With an upgraded version (x264 core:112 r1867 22bfd31) I got the following log file. Please take a look and let me know thoughts. Also should note that my CPU is only around 50% used during the test. When I use my own quality settings for blu-ray movies it is 100% on the second pass (yes I use 2 pass encoding).

Let me know what else I can adjust.

For some reason it looks like you're deinterlacing a progressive source and doing a crop & resize



FFVideoSource("movies\00588 temp files\00588.mkv", cachefile="movies\00588 temp files\00588.ffindex")
AssumeFPS(23.976)
Crop(0,0, -Width % 8,-Height % 8)
ConvertToYV12()
Yadif()
Crop(0,4,-0,-4)
BicubicResize(1920,1072,0,0.5)

poisondeathray
14th January 2011, 21:26
Here's my little contribution

<snip>


Mainconcept - http://www.mediafire.com/?5c9a4geg809bnyu
x264 - http://www.mediafire.com/?95zzd5338q2rnnx


Thanks,

No surprise, not looking too good for the Mainconcept CUDA encoder...

What were your system specs shon3i ? CPU and GPU ?

mp3dom
14th January 2011, 21:38
Here's my little contribution
source: park_joy_1080p.y4m

The settings for both encoders are near the same?
Comparing the CUDA settings vs. x264 settings seems that CUDA uses only 1 B-frames (3 on x264), that CUDA can potentially have I frame every frame (x264 at least every 23), CUDA cannot reference B frames and doesn't have pyramid Bframe (both options enabled in x264). Probably doesn't change anything at the end anyway.

SeeManRun
14th January 2011, 22:09
For some reason it looks like you're deinterlacing a progressive source and doing a crop & resize

These options seem to be in the default. I will have to take a look if removing them gets me some higher speed. I never resize videos, only crop, so this is interesting.

poisondeathray
14th January 2011, 22:17
Those aren't defaults for x264 . You mean the defaults GUI you're using (staxrip ?)

I would say yadif is your main bottleneck. Not only will it be slower, it will be lower quality

If you want to test speed , use x264.exe to remove other bottlenecks

If you're using single threaded ffms , that could be a bottleneck as well (ffms-mt is the multithreaded version , I don't know which one you're using in staxrip)

LoRd_MuldeR
14th January 2011, 23:48
I would say yadif is your main bottleneck. Not only will it be slower, it will be lower quality

You will hardly find a (software) deinterlacer that is as fast as Yadif and still gives decent quality.

For an encoder test/comparison I would recommend to do the deinterlacing beforehand and feed all encoders with the identical progressive (deinterlaced) source.

poisondeathray
14th January 2011, 23:50
You will hardly find a (software) deinterlacer that is as fast as Yadif and still gives decent quality.


I agree.

But I still wouldn't use it on a 23.976 progressive source. :)

deadrats
15th January 2011, 00:38
My upgraded machine is a Core i7 2600k with lots of RAM on an ASUS P8P67, all at stock but this thing is turboing up to 3.8 ghz all the time.


how about you do a QS test with tmpg express 5 for us so we can put this thread to bed.

shon3i
15th January 2011, 02:24
Thanks,

No surprise, not looking too good for the Mainconcept CUDA encoder...

What were your system specs shon3i ? CPU and GPU ?
GeForce 9600GT 512MB (stock), Phenom X4 @ 2.2 (stock). Btw to my eyes, Mainconcept is better than any GPU encoder these days. Badaboom (elemental) is even worse

The settings for both encoders are near the same?
Comparing the CUDA settings vs. x264 settings seems that CUDA uses only 1 B-frames (3 on x264), that CUDA can potentially have I frame every frame (x264 at least every 23), CUDA cannot reference B frames and doesn't have pyramid Bframe (both options enabled in x264). Probably doesn't change anything at the end anyway. What i am trying to prove here is what we talking about last 2-3 pages back, that A encoder can be more efficient thatn B encoder and produce significant better quality for same and even faster speed. Nothing else. I off course can match all x264 settings to decrease quality and make same quality as GPU encoder, i will probably get significant more speed with x264 nothing else. I intended to use slowest possible settings in GPU encoder and use same-speed settings in x264, and as i said i can even decrease x264 settings from defaults and i will probably get nothing than more speed.

And about I frames, both encodes have only 2 I frames, you can check streams through Elecard Stream Eye

SeeManRun
15th January 2011, 04:57
how about you do a QS test with tmpg express 5 for us so we can put this thread to bed.

Will work on that later, but as others have posted you may need a h67 motherboard for that, which I don't have.

SeeManRun
15th January 2011, 04:59
Those aren't defaults for x264 . You mean the defaults GUI you're using (staxrip ?)

I would say yadif is your main bottleneck. Not only will it be slower, it will be lower quality

If you want to test speed , use x264.exe to remove other bottlenecks

If you're using single threaded ffms , that could be a bottleneck as well (ffms-mt is the multithreaded version , I don't know which one you're using in staxrip)

Yes, defaults in the GUI. Removing de-interlacing certainly speeds it up, and I never resize but it seems screwing with it made that happen. Fixing that up and using ultrafast i was able to hit over 135 fps and it looks much better than I would have expected.. Will have to play with this more later.

SeeManRun
15th January 2011, 06:51
how about you do a QS test with tmpg express 5 for us so we can put this thread to bed.

I can't seem to find version 5, and since I don't speak Japanese, I don't really want to go through the pain of trying to figure out what to run when it will likely be incorrect.

frenchfries
15th January 2011, 13:38
how about you do a QS test with tmpg express 5 for us so we can put this thread to bed.

I too would like to see someone with a SB run a comparison but unfortunately he cannot as he said he has a P67, which does not allow the use of the "QuickSync" bits yet

My encodes taken at a CRF of 18 and using slower settings seem to come out at around 10Mb/s with very little perceivable loss over the original bluray. If the source is grainy, as it so often seems to be these days, I can notice the difference more easily but tuning for film boosts the file size significantly. A trade off I am unwilling to make.

As mentioned by another poster, if you don't care about diskspace why do you transcode at all? I would simply decrypt the content and leave it at that, noty even remux, if that were a viable position for me.

OT - Dark Shikari reminds of Linus Torvalds in forums, comes across as knowing what he is on about but usually very quick to get angry. Neuron2 also fits into this camp. That being said, you also have on several occasions snapped and "argued the person not the topic"

deadrats
15th January 2011, 14:52
As mentioned by another poster, if you don't care about diskspace why do you transcode at all? I would simply decrypt the content and leave it at that, noty even remux, if that were a viable position for me.

if you go back and reread what i wrote i routinely advise people not to bother transcoding, the only time i do any re-encoding is if i have a tv capture of a sporting event or tv show and i want to cut out the commercials, otherwise i will only keep the video stream as it is, keep only the highest quality audio stream and remux.

the people that transcode a commercial BD from 30+ mb/s to 10 mb/s are obviously those that intend to pirate the content: they wish to save as much bandwidth as they can sharing it or they intend to return the BD back to the video store and wish to use as little space as possible "backing up" the content.

OT - Dark Shikari reminds of Linus Torvalds in forums, comes across as knowing what he is on about but usually very quick to get angry. Neuron2 also fits into this camp. That being said, you also have on several occasions snapped and "argued the person not the topic"

you hit the nail on the head, the pseudo-religious fervor displayed by the x264 faithful is prevalent within the entire open source community; there was a time i did linux reviews in the old ET forums and said reviews used to be quite popular, they used to regularly get hundred's of views and comments and my email would get flooded.

the funny thing was that if i ripped into that distro or linux in general, the faithful would spend post after post telling me i was wrong, i was incompetent, i didn't know what i was talking about and/or that i was a microsoft employee out to spread FUD in order to destroy linux.

if i liked a particular distro, at times i was accused of being an employee of that company (such as red hat) or at times my positive review got linked on the home page of that distro as a testimonial.

if i have snapped at times it's because DS is single handedly responsible for all the anti-gpu acceleration FUD you're likely to run into: people with no programming experience, that think java is a type of bean, fill forum after forum with some of the most retarded BS you are likely to find and it has DS' signature all over it.

some of the most stupid:

1) gpu's are only good at simple calculations but the calculations performed by x264, and required by video encoding in general, are too complex.

2) gpu's can't do video encoding, that's why cuda sucks.

3) gpu's are only good at tasks that are highly parallel in nature but video encoding is inherently a linear task.

there's a number of others but these are the ones that really annoy me.

LoRd_MuldeR
15th January 2011, 15:25
some of the most stupid:

1) gpu's are only good at simple calculations but the calculations performed by x264, and required by video encoding in general, are too complex.

Indeed GPU cores are very simple compared to CPU cores. Initially GPU cores have been designed for 3D rendering and for nothing else. Starting with the GPGPU hype the GPU vendors have integrated more and more "CPU-like" features into their GPU cores. These features actually hurt 3D performance (because they occupy transistors that could have been used for stuff that helps 3D rendering otherwise), but they allow for more efficient/comfortable GPGPU calculations. Still, even with all the GPGPU extensions, GPU cores are lacking many features that you are used from CPU programming. Especially getting the memory management and access patterns right on the GPU is a nightmare! If you claim the opposite, you probably have never implemented a non-trivial algorithm on the GPU (I have, so I know what I'm talking about). The advantage of GPU cores is that they often can "hide" their disadvantages, simply because there are so many cores in a GPU. You can think of the GPU as a highly parallel co-processor. Still only problems that scale to thousands of threads will run on the GPU efficiently. Problems that are inherently sequential will never run efficiently on the GPU, because all those beautiful cores will be idle most of the time...

2) gpu's can't do video encoding, that's why cuda sucks.

CUDA is just an interface for GPGPU programming. There are many problems (mostly scientific ones) that are extremely parallel and thus fit perfectly well onto the GPU. However "video encoding" is a problem which is not that suitable for GPGPU (and thus for CUDA or for OpenCL or for whatever GPGPU interface you use). And that's mainly because of point 3.

3) gpu's are only good at tasks that are highly parallel in nature but video encoding is inherently a linear task.

That's basically true! However video encoding is not inherently sequential in its whole, but there are important parts in the encoding pipeline that are inherently sequential. So if at all, only certain parts of the video encoding pipline will work efficiently on the GPU. Unfortunately moving only parts of the video encoder to the GPU while keeping all the rest on the CPU isn't easy to do either. That's because transferring data between the GPU (device memory) and the CPU (host memory) causes an enormous delay which can easily kill all your nice speed-up (believe me, I know what I'm talking about ^^). The existing "CUDA encoders" workaround that problem by using very simple algorithms for those "problematic" parts of the encoding pipeline. And that's one of the reasons why they suck (quality-wise) that much...

nm
15th January 2011, 16:25
My encodes taken at a CRF of 18 and using slower settings seem to come out at around 10Mb/s with very little perceivable loss over the original bluray. If the source is grainy, as it so often seems to be these days, I can notice the difference more easily but tuning for film boosts the file size significantly. A trade off I am unwilling to make.

You should generally always tune for film. Raise CRF until you get the bitrates that you are after.

mariush
15th January 2011, 17:26
deadrats, since you think programming in Cuda is so easy, I think you should read this paper from nVidia, which shows just how many things you have to keep in mind if you wish to have efficient and fast code on gpu:

http://developer.download.nvidia.com/compute/cuda/1_1/Website/projects/reduction/doc/reduction.pdf

And this above is a simple reduction, not multihexagon search or something more complex...

What Lord_Mulder says above is correct... if you were to implement the algorithms that are now in x264 that give quality to the encodings on the gpu, it will take so much time to transfer the data back and forth to the GPU and arrange it in a way the algorithms would run in parallel on the video card efficiently... and after you do this the code will run just a bit faster on the GPU... both together would just make the whole encoding slower.

later edit: You can also see this 29 page slideshow, which discusses using Cuda in x264: https://docs.google.com/present/view?id=df3rvqmk_1gsxgzgts&pli=1

The slideshow's project page is this one: https://sites.google.com/site/x264cuda/

poisondeathray
15th January 2011, 17:33
I'm a non programmer, but I've seen Lord Mulder explain this before and it seems like very reasonable explanations to me

@deadrats
1) So why do you think - in programming terms - the current GPU encoders "suck" ? Surely the big companies have enough resources to hire good programmers and develop software? It's not in the infancy stage anymore - we're on 3rd and 4th generation GPU encoders now.

2) More importantly, how could you improve them? or how could you integrate a GPU-CPU encoder etc... ie. How can you make something that - me - an end user would want to use?

dj_tjerk
15th January 2011, 17:42
if you go back and reread what i wrote i routinely advise people not to bother transcoding, the only time i do any re-encoding is if i have a tv capture of a sporting event or tv show and i want to cut out the commercials, otherwise i will only keep the video stream as it is, keep only the highest quality audio stream and remux.
Blasphemy! Why don't you cut without re-encoding?


the people that transcode a commercial BD from 30+ mb/s to 10 mb/s are obviously those that intend to pirate the content: they wish to save as much bandwidth as they can sharing it or they intend to return the BD back to the video store and wish to use as little space as possible "backing up" the content.

Or those that, well, want to actually play a movie on a device that can't handle such bitrates/resolutions, or can't store that amount of data. Also, nice fallacy there.


you hit the nail on the head, the pseudo-religious fervor displayed by the x264 faithful is prevalent within the entire open source community; there was a time i did linux reviews in the old ET forums and said reviews used to be quite popular, they used to regularly get hundred's of views and comments and my email would get flooded.
So you're saying you used to be an arrogant and allegedly worshipped, though contributing, user? I.e. what you claim DS to be?


the funny thing was that if i ripped into that distro or linux in general, the faithful would spend post after post telling me i was wrong, i was incompetent, i didn't know what i was talking about and/or that i was a microsoft employee out to spread FUD in order to destroy linux.
Can't say you're wrong about the distro stuff, as I haven't seen any of your reviews or the mentioned posts. What I can say however is that you don't know how video compression works, based on your sole argument that "bits per pixel" is always and directly correlated with (perceived) quality.


if i have snapped at times it's because DS is single handedly responsible for all the anti-gpu acceleration FUD you're likely to run into: people with no programming experience, that think java is a type of bean, fill forum after forum with some of the most retarded BS you are likely to find and it has DS' signature all over it. DS has explained over and over how to approach offloading some parts of the encoding process to the GPU to people asking willing to try it. However, every single one of those persons has given up, because it isn't as trivial to implement as they thought. I wouldn't expect anyone to not become a bit curt when the topic of "gpu encoding" surfaces once more, and people are proclaiming it's the next best thing and x264 should use it.


some of the most stupid:

See LoRd_MuldeR's post. The GPU is good at what it's good at (duh), but one often has to resort to highly parallel algorithms that are much worse than a serial algorithm.

CruNcher
15th January 2011, 18:40
GeForce 9600GT 512MB (stock), Phenom X4 @ 2.2 (stock). Btw to my eyes, Mainconcept is better than any GPU encoder these days. Badaboom (elemental) is even worse

What i am trying to prove here is what we talking about last 2-3 pages back, that A encoder can be more efficient thatn B encoder and produce significant better quality for same and even faster speed. Nothing else. I off course can match all x264 settings to decrease quality and make same quality as GPU encoder, i will probably get significant more speed with x264 nothing else. I intended to use slowest possible settings in GPU encoder and use same-speed settings in x264, and as i said i can even decrease x264 settings from defaults and i will probably get nothing than more speed.

And about I frames, both encodes have only 2 I frames, you can check streams through Elecard Stream Eye

It's better then Elementals current badaboom encoder no doub't but 2.0 has to bee seen also its not as good as Nvidias Encoder yet but the Q1 version should beat it (also feature wise).
So in your case 8 SMs (G94) are slightly slower as 4 AMD cores @ subme 1 in Mainconcepts Encoder :) i can update this with G92 results @ 14 SMs :) if we can get someone with a Fermi based G100 card and a AMD system it would be perfect :)

Though it would be much easier to create a automated benchmark system based on Nvidias Encoder then on Mainconcepts for easier result gathering in a external thread :) (with x264 as a system reference bench for both 720p/1080p)

shon3i
15th January 2011, 19:23
@CruNcher, is there trial of nvidia encoder or in some application? i willing to try. Btw after i do test on older Athlon 4800 @ 2.4GHz i get speed like ~25-30fps with x264, aslo i tryed on friends GTX 460, i didn't get noticeable speed few fps more nothing marginal

poisondeathray
15th January 2011, 19:36
Btw after i do test on older Athlon 4800 @ 2.4GHz i get speed like ~25-30fps with x264, aslo i tryed on friends GTX 460, i didn't get noticeable speed few fps more nothing marginal

shon3i - mainconcept's marketing slides show big improvement , but larger improvement with more cores, say GTS250 vs. GTX470 on the GPU encoder, (more on hexcore, but even on dual core setup, speedup is significant). Maybe that older Athlon core was bottleneck ?

http://www.mainconcept.com/fileadmin/user_upload/download/product_sheets/CUDA-Sheets_06-2010.pdf

deadrats
15th January 2011, 21:48
...

wow.

1) gpgpu features kill 3d performance? is that why each successive generation of gpu gets faster and faster? gpgpu capabilities have been available since dx8 class cards, when shaders were first introduced, 3d performance has increased many orders of magnitude since the gf3 days, clearly the additional transistors haven't hurt performance all that much.

2) re: memory transfer, a similar claim could be made with regards to general cpu programming - you have to load the data from the hdd to ram and write the results back, if your app doesn't handle memory management internally then the OS does it for you, but either way it takes place. you don't seem to have a problem with the performance implications of that, so why the objection that gpgpu programming has a similar requirement? furthermore, the ram on graphics cards is faster than system ram (gddr3 or 5 vs ddr2 or 3), so the performance penalty is significantly less.

but this approach fails to utilize the full power of modern gpu's: i know that DS tried to get a SAD function coded in cuda (i've seen his posts over at the nvidia developers forum) and i know he was successful in coding it but was disappointed with the performance versus the one used in x264.

what he failed to account for is that cuda functions, unlike c/c++ functions, can be called numerous times, the proper way to use a gpu powered SAD function is to compile the "kernel" and call it repeatedly to perform the SAD calculations in parallel (obviously on different frames). you could code the gpu SAD to store the values in an array that can then be read from the encoder as they are needed.

you don't have to code each individual part of an encoder so that each part runs faster than it's x86 counterpart, all you have to do is code it so all the parts together have a lower running time than all the x86 parts together.

test after test with tmpg express 5 shows that the cuda encoder runs faster than the included x264, using only 20-25% of the 128 cores my gts250 has, and the performance disparity increases as the bit rate increases, at true BD bit rates, x264 just can't keep up.

3) the main reason cuda encoders "suck" is because quite a few of them use the reference nvidia h264 encoder which was meant as a template, a learning tool, it was never meant to be used in a commercial product.

the reality is this: gpgpu programming in not easy, it requires a completely different way of thinking about problems and their solutions, something x86 programmers are not used to. add to that the fact that nvidia hasn't made cuda programming easy for those that cut their teeth doing procedural programming, they have sponsored cuda courses only on the graduate level and it's no wonder cuda encoders suffer a bit.

but that doesn't mean the technology sucks, just that people don't know how to use it yet.

deadrats
15th January 2011, 21:52
I'm a non programmer, but I've seen Lord Mulder explain this before and it seems like very reasonable explanations to me

@deadrats
1) So why do you think - in programming terms - the current GPU encoders "suck" ? Surely the big companies have enough resources to hire good programmers and develop software? It's not in the infancy stage anymore - we're on 3rd and 4th generation GPU encoders now.

2) More importantly, how could you improve them? or how could you integrate a GPU-CPU encoder etc... ie. How can you make something that - me - an end user would want to use?

see my response to "mulder".

LoRd_MuldeR
15th January 2011, 22:32
1) gpgpu features kill 3d performance? is that why each successive generation of gpu gets faster and faster? gpgpu capabilities have been available since dx8 class cards, when shaders were first introduced, 3d performance has increased many orders of magnitude since the gf3 days, clearly the additional transistors haven't hurt performance all that much.

As a matter of fact, the number of transistors that are available in a GPU is limited. And all transistors that the designers spend for improving GPGPU-specific features cannot be used for 3D Rendering-specific features anymore. Modern GPU's definitely contain various design decisions that have been made in favor of GPGPU performance rather than 3D Rendering performance. Of course the overall 3D performance still has improved in each GPU generation, but it surely could have been improved a lot more of if they did optimize for the "raw" 3D performance only...

2) re: memory transfer, a similar claim could be made with regards to general cpu programming - you have to load the data from the hdd to ram and write the results back, if your app doesn't handle memory management internally then the OS does it for you, but either way it takes place. you don't seem to have a problem with the performance implications of that, so why the objection that gpgpu programming has a similar requirement? furthermore, the ram on graphics cards is faster than system ram (gddr3 or 5 vs ddr2 or 3), so the performance penalty is significantly less.

You obviously don't have the slightest idea what you are talking about. First of all, on the GPU you must "upload" all input data from the host memory (CPU) to the device memory (GPU) first and later you must "download" all the results from device memory back to host memory. This is in addition to loading/saving the data from/to the HDD. Moreover memory access to the GPU memory from a GPU kernel is very different compared to access to the main memory from a CPU program. Access to the "global" GPU memory is very slow and (by default) it's not cached at all! Therefore each group of GPU cores has a fast (but small!) "shared" memory attached to it (some of that "shared" memory can be used as cache in the Fermi generation). Moreover even access to the "shared" memory must be organized very carefully, because it is split into several memory banks. Different GPU cores can access memory located in different banks simultaneously, but access to memory in the same bank must be serialized, which costs a lot of performance. And so on...

but this approach fails to utilize the full power of modern gpu's: i know that DS tried to get a SAD function coded in cuda (i've seen his posts over at the nvidia developers forum) and i know he was successful in coding it but was disappointed with the performance versus the one used in x264.

what he failed to account for is that cuda functions, unlike c/c++ functions, can be called numerous times, the proper way to use a gpu powered SAD function is to compile the "kernel" and call it repeatedly to perform the SAD calculations in parallel (obviously on different frames). you could code the gpu SAD to store the values in an array that can then be read from the encoder as they are needed.

you don't have to code each individual part of an encoder so that each part runs faster than it's x86 counterpart, all you have to do is code it so all the parts together have a lower running time than all the x86 parts together.

Again you have not the slightes idea what you are talking about :rolleyes:

3) the main reason cuda encoders "suck" is because quite a few of them use the reference nvidia h264 encoder which was meant as a template, a learning tool, it was never meant to be used in a commercial product.

the reality is this: gpgpu programming in not easy, it requires a completely different way of thinking about problems and their solutions, something x86 programmers are not used to. add to that the fact that nvidia hasn't made cuda programming easy for those that cut their teeth doing procedural programming, they have sponsored cuda courses only on the graduate level and it's no wonder cuda encoders suffer a bit.

but that doesn't mean the technology sucks, just that people don't know how to use it yet.

CUDA (and GPGPU in general) is available for several years now and various "big" companies, including NVidia, have been pushing GPGPU aggressively since then. Still all the GPGPU encoders that we have seen on the market until now aren't anywhere near the state-of-the-art software encoders. If not all programmers/companies working on GPGPU encoders are completely incompetent, it is time to realize that the currently available GPUGPU platforms are NOT as suitable for video encoding as the GPU vendors and their marketing departments try to make us believe all the time...


Okay, deadrats. I have been watching you trolling this thread for quite some time now. Whenever somebody explains something to you, you respond immediately and claim the opposite, although from your posts it often quite obvious that you don't know the technical background very well. So I kindly ask you stop at this point and focus on the original topic again!

shon3i
15th January 2011, 22:52
@CruNcher, is there trial of nvidia encoder or in some application? i willing to try. Btw after i do test on older Athlon 4800 @ 2.4GHz i get speed like ~25-30fps with x264, aslo i tryed on friends GTX 460, i didn't get noticeable speed few fps more nothing marginal
Infact you been right on Athlon machine i can't propretly (without bottleneck) playback avs script. HDD will blow up.

Well that cards will be probably more expensive than an good six or eighth core processor.

frenchfries
16th January 2011, 02:12
wow.

1) gpgpu features kill 3d performance? is that why each successive generation of gpu gets faster and faster? gpgpu capabilities have been available since dx8 class cards, when shaders were first introduced, 3d performance has increased many orders of magnitude since the gf3 days, clearly the additional transistors haven't hurt performance all that much.


Have you noticed that GPU's performance increase in each successive generation has been tapering off? At least a small part of this is due to the designers being unable to use all the transistors in the budget for raw 3d work.

I don't appreciate your implication that I am a pirate merely because I don't have unlimited disk space. In fact if I was a pirate I would have considerably more money to spend on disk space and could then simply download BD images and not have to transcode.

I mainly transcode for ease of use, not having to go find the disc is good. If I can make something smaller with no perceivable difference, to me at least, why wouldn't I? It only costs me a little time and my CPU a lot of time :P

Finally, I was quite optimistic when this talk of GPU accelerated transcoding started and it was only after I saw the output that I became jaded. I am so glad the badaboom had a trial or I would have asked for my money back.

deadrats
16th January 2011, 03:44
I don't appreciate your implication that I am a pirate merely because I don't have unlimited disk space. In fact if I was a pirate I would have considerably more money to spend on disk space and could then simply download BD images and not have to transcode.

I mainly transcode for ease of use, not having to go find the disc is good. If I can make something smaller with no perceivable difference, to me at least, why wouldn't I? It only costs me a little time and my CPU a lot of time :P

it's not an implication, it's a logical conclusion. first of all who watches a BD on their pc, don't you prefer watching it on a hdtv with a BD player?

second, to me the term "back up" means just that; use a BD burner to make a duplicate disk, not transcode it down to 10 mb/s (or less).

as for being jaded, so long as you use enough bit rate, the quality differences are nill.

furthermore all the companies that buy per seat licenses to main concept's cuda encoder or adobe's "mercury" engine would disagree. another company that would disagree with you is microsoft, seeing how they filed for, and were granted patents for, gpu powered encoding.

the thought that a handful of programmers that give their software away for free knows more about gpgpu than the people making millions of it is laughable.

aegisofrime
16th January 2011, 03:51
it's not an implication, it's a logical conclusion. first of all who watches a BD on their pc, don't you prefer watching it on a hdtv with a BD player?

Maybe I just want to watch porn in my own room instead of in the living room with my family.

poisondeathray
16th January 2011, 03:55
Maybe I just want to watch porn in my own room instead of in the living room with my family.

Best reply EVER :p

cacepi
16th January 2011, 04:16
furthermore all the companies that buy per seat licenses to main concept's cuda encoder or adobe's "mercury" engine would disagree. another company that would disagree with you is microsoft, seeing how they filed for, and were granted patents for, gpu powered encoding.

<sarcasm>Yes, because it's not like Mainconcept wants to charge you $1200 for an encoder, or Adobe $800 for their editor suite ($2600 for Master Collection), or Microsoft patent GPU encoding on their platforms so they can charge Adobe and Mainconcept licensing fees.</sarcasm>

You're certainly not that naïve, are you?

And yes, I understand the point you're trying to make, but what does that really have to do with how superior, or inferior, x264 is to OpenCL/CUDA encoding? And be honest here.

frenchfries
16th January 2011, 07:57
it's not an implication, it's a logical conclusion. first of all who watches a BD on their pc, don't you prefer watching it on a hdtv with a BD player?


Its called a HTPC or media centre.
http://en.wikipedia.org/wiki/Home_theater_PC
I prefer being able to easily browse my movies from the couch rather than getting up and having to look at the dvd storage unit and attempt to read the titles from the spine.
http://img504.imageshack.us/img504/3333/previewxf5.png

deadrats
16th January 2011, 11:41
And yes, I understand the point you're trying to make, but what does that really have to do with how superior, or inferior, x264 is to OpenCL/CUDA encoding? And be honest here.

what does it have to do with it?

as i have said, people are willing to spend thousands on per seat licenses, or out right pirate the software, rather than use the legally free alternative, i think that speaks volumes as to how x264 is viewed compared to other encoders.

x264 has an unfair competitive advantage by virtue of being free, if they started charging $2500 for a single seat license, how many people do you think would still choose it over main concept gpu powered encoder?

there seems to be this ridiculous mythos that sprouts up around any open source project, linux being a prime example, the thought being that a bunch of hobbyists knows more about OS design than the people making billions off it.

absurd.

in all honesty, i have seen numerous encodes where x264 fell apart, it happens with all encoders once you start bit rate starving an encode.

i think we can all agree that given the same source, if we encoded to 1080p with a very reasonable 15 mb/s or 720p with a reasonable 10 mb/s, more than likely, we wouldn't be able to tell the difference between encodes.

once we apply some professional standards to our encodes x264's advantages quickly disappear.

Mixer73
16th January 2011, 11:53
as i have said, people are willing to spend thousands on per seat licenses, or out right pirate the software, rather than use the legally free alternative, i think that speaks volumes as to how x264 is viewed compared to other encoders.

Do not underestimate the impact of nepotism and gladhanding.

Sharktooth
16th January 2011, 12:10
@deadrats: compression algos are made to COMPRESS. so the more they compress (retaining image or audio quality for lossy encoders) the better they are. that means, the efficiency of an encoder is the ability to produce the lowest bitrate for an expected image quality. in this case x264 is unbeatable. plus it is very configurable.
your idea about setting a bitrate for a resolution is WRONG. 1st coz bitrate depends on source compressibility and compressibility depends on the encoder. sure you can shoot a super high bitrate and whatever encoder you use the result will look the same. but the principle is you want to COMPRESS that material... otherwise why use a high efficiency compressor? you can just stick to MPEG-2 or MPEG-4 ASP and set, let's say, 25mbps for 1080p.
so, the whole point of this discussion, is not "the bitrate". we care about the COMPRESSION and picture/sound quality. something that oversimplified encoders running on GPUs are not capable to deliver together.

iwod
16th January 2011, 12:50
Can we keep things just SB related..... not another GPU Encoding Thread,

the_corona
16th January 2011, 14:17
Oh my, I was so excited when I saw how many news posts this thread got, only to discover it turned into a GPU encoder argument :-(

deadrats, it seems you are quite happy with throwing enough bits at a problem so pretty much any encoder will do fine. That's great, so stick with it and stop saying everyone who disagress is a pirate and a mere idiot. What is the point of your posts?

You keep making all this claims, yet have nothing to show for it. You don't seem to understand the first thing about Cuda or OpenCL programming (see your "understanding" of memory).

Regarding SB, DS has said in the past it could potentially be useful, but Intel, who have in the past promised to provide a patch by them for x264, just dissapeared (why? did they find its not helping afterall? Maybe its politics, who knows) and does not adequatly provide documentation for the new low level calls.

I humbly suggest you stop making all these claims and accusations. Or in the very least return to talking about SB and not CUDA because this thread is about the former after all (you do know there is a huge difference don't you?)

LoRd_MuldeR
16th January 2011, 14:26
@everybody:
As the discussion in this thread is definitely going into the wrong direction for some time now, I urge everybody to focus on the original topic again.
This means, if you don't have to add anything new to the discussion about Sandy Bridge (QuickSync) please do not post in this thread!

(BTW: I already moved two of deadrats's profanity/no-use replies to moderation. And I will do so with other off-topic posts, in the hope we can resurrect the thread)

SeeManRun
16th January 2011, 20:17
@everybody:
As the discussion in this thread is definitely going into the wrong direction for some time now, I urge everybody to focus on the original topic again.
This means, if you don't have to add anything new to the discussion about Sandy Bridge (QuickSync) please do not post in this thread!

Agreed. It seems I am the only one in this thread that I have seen yet with a Sandy Bridge (that I bought specially for x264, not for QS though). I had one private message from someone about running some tests, but so far they are proving too much work for me to do (installing diff versions of avisynth) and I don't really want to corrupt my currently functioning work flow.

How can I help you guys in some way by running some tests? I can offer this: I have a 3.4 ghz SB CPU that will run 8 threads and is fully unlocked. I have a Core 2 Quad 6600 running at 2.4 ghz. By disabling hyperthreading and lowering the speed of my SB, I can compare 4 SB cores at 2.4 ghz to 4 core2duo cores (2+2 of course) running at 2.4 ghz and post the results (unfortunately RAM will be 4 vs 16 gigs, but from what I have seen x264 doesn't usually take more than 1 in my existing work flows).

What I need though is a fairly self contained test so I can more or less just download the package and run it and post the results (I can supply BD or DVD content from my collection).

As for QuickSync, I have a p67 motherboard, so last I heard I cannot tap into the QS functions, though that does seem like an artificial limitation, unless the QS functionality is part of the on board video card which is disabled on p67.

deadrats
16th January 2011, 23:41
@SeeManRun

you could download the japanese version of tmpg express and run the media sdk encoder in software mode. if you don't wish to go through that hassle, download gom encoder, make sure to install the IPP encoder, and run some visual quality tests comparing x264 with IPP. you choose which encoder is used in the settings.

as a side note, main concept has jumped on the QS bandwagon:

http://www.mainconcept.com/products/partner-products/intel/h264avc-encoder-sdk.html

evidently intel has decided to share access to it's low level api with them :rolleyes:

LoRd_MuldeR
17th January 2011, 00:04
http://www.mainconcept.com/products/partner-products/intel/h264avc-encoder-sdk.html

evidently intel has decided to share access to it's low level api with them :rolleyes:

"The MainConcept™ H.264/AVC Encoder SDK for Intel® Quick Sync Video (QSV) acts as wrapper for the Intel Media SDK 2.0."

:rolleyes:

Sharktooth
17th January 2011, 00:06
ok, so then it will be crappy as the others...

aegisofrime
17th January 2011, 03:22
From what I read on Anandtech, Quick Sync is a special dedicated piece of silicon, separate from the GPU. I'm not an engineer, but to my mind there's no reason why Quick Sync needs Intel FDI to work. I'm hoping one of the more adventurous motherboard manufacturers (I'm looking at you ASRock) will be able to get around this limitation.

As for Mainconcept, they are just jumping onto the bandwagon for marketing purposes. "Hey lookie, we are among the first to support Quick Sync!" I expect as time goes by there will be more refined implementations.