View Full Version : CUDA H.264 vs x264 Speed and Image Quality Benchmarks, discussion


St Devious
10th July 2009, 07:54
Hardware


Intel Q9450 @ 3.2 Ghz
4GB RAM @ 800 MHz 5-4-4-12 2T
GTS 250 512 MB @ 738/1836/1100 MHz (Core/Shader/Memory)


http://i27.tinypic.com/20z7l21.jpg

Software

Nvidia GeForce 186.18 WHQL Drivers
MeGUI 0.3.1.1047
x264 r1178 Jeeb's Build
Mediacoder 0.7.1.4475 for CUDA GPU encoding



Source Videos
VforVendetta 2000Kbps 1280x720 Clip (http://mirror05.x264.nl/Dark/force.php?file=./x264clips/VForVendetta.mkv)

Mediainfo on Source
Video
ID : 2
Format : AVC
Format/Info : Advanced Video Codec
Format profile : High@L4.0
Format settings, CABAC : Yes
Format settings, ReFrames : 4 frames
Muxing mode : Container profile=Unknown@4.0
Codec ID : V_MPEG4/ISO/AVC
Duration : 1mn 52s
Nominal bit rate : 2 000 Kbps
Width : 1 280 pixels
Height : 720 pixels
Display aspect ratio : 2.35
Frame rate : 23.976 fps
Resolution : 24 bits
Colorimetry : 4:2:0
Scan type : Progressive
Bits/(Pixel*Frame) : 0.091
Writing library : x264 core 65 r999+1 eb3ef1b
Encoding settings : cabac=1 / ref=4 / deblock=1:-1:-1 / analyse=0x3:0x113 / me=tesa / subme=9 / psy_rd=1.0:0.0 /
mixed_ref=1 / me_range=32 / chroma_me=1 /trellis=0 / 8x8dct=1 / cqm=0 / deadzone=4,4 / chroma_qp_offset=-2 / threads=3 /
nr=0 / decimate=0 / mbaff=0 / bframes=3 / b_pyramid=1 / b_adapt=2 / b_bias=0 / direct=3 / wpredb=1 / keyint=250 / keyint_min=25 /
scenecut=40(pre) / rc=2pass / bitrate=2000 / ratetol=1.0 / qcomp=0.60 / qpmin=10 / qpmax=51 / qpstep=4 / cplxblur=20.0 / qblur=0.5 / ip_ratio=1.40
/ pb_ratio=1.30 / aq=1:1.00


Encoding Settings
CUDA GPU

http://i31.tinypic.com/2hxr3nd.jpg

MeGUI x264
Preset ultrafast used here
program --bitrate 800 --no-mixed-refs --bframes 1 --no-weightb --direct temporal --nf --no-cabac --subme 1 --partitions none --scenecut 0 --me dia
--threads auto --thread-input --aq-mode 0 --output "output" "input" --subme 0 --preset ultrafast

Speed Results


CUDA GPU H.264 - 23.5s 114.4 FPS
x264 preset ultrafast - 30s 90.8 FPS


Image comparison

Source
http://i29.tinypic.com/1e8dc1.jpg

x264 @ 800 Kbps
http://i27.tinypic.com/dnhl4g.png

CUDA GPU @ 800 Kbps
http://i28.tinypic.com/dfanhf.png

Output File Mediainfo

x264 @ 800 Kbps
Video
ID : 1
Format : AVC
Format/Info : Advanced Video Codec
Format profile : Main@L3.1
Format settings, CABAC : No
Format settings, ReFrames : 2 frames
Codec ID : avc1
Codec ID/Info : Advanced Video Coding
Duration : 1mn 52s
Bit rate mode : Variable
Bit rate : 900 Kbps
Nominal bit rate : 800 Kbps
Maximum bit rate : 2 229 Kbps
Width : 1 280 pixels
Height : 720 pixels
Display aspect ratio : 16/9
Frame rate mode : Constant
Frame rate : 23.976 fps
Resolution : 24 bits
Colorimetry : 4:2:0
Scan type : Progressive
Bits/(Pixel*Frame) : 0.041
Stream size : 12.1 MiB (100%)
Writing library : x264 core 68 r1179M 96e2229
Encoding settings : cabac=0 / ref=1 / deblock=0:0:0 / analyse=0:0 / me=dia / subme=0 / psy_rd=0.0:0.0 / mixed_ref=0 /
me_range=16 / chroma_me=1 / trellis=0 / 8x8dct=0 / cqm=0 / deadzone=21,11 / chroma_qp_offset=0 / threads=6 / nr=0 / decimate=1 / mbaff=0 /
bframes=1 / b_pyramid=0 / b_adapt=1 / b_bias=0 / direct=1 / wpredb=0 / keyint=250 / keyint_min=25 / scenecut=0 / rc=abr / bitrate=800 /
ratetol=1.0 / qcomp=0.60 / qpmin=10 / qpmax=51 / qpstep=4 / ip_ratio=1.40 / pb_ratio=1.30 / aq=0



CUDA GPU @ 800 Kbps

Video
ID : 1
Format : AVC
Format/Info : Advanced Video Codec
Format profile : High@L5.1
Format settings, CABAC : Yes
Format settings, ReFrames : 2 frames
Codec ID : avc1
Codec ID/Info : Advanced Video Coding
Duration : 1mn 52s
Bit rate mode : Variable
Bit rate : 869 Kbps
Maximum bit rate : 1 776 Kbps
Width : 1 280 pixels
Height : 720 pixels
Display aspect ratio : 16/9
Frame rate mode : Constant
Frame rate : 23.976 fps
Resolution : 24 bits
Colorimetry : 4:2:0
Scan type : Progressive
Bits/(Pixel*Frame) : 0.039
Stream size : 11.6 MiB (100%)


http://i32.tinypic.com/beggah.jpg

Shows that GPU temperature rose by 6 C when encoding. CPU Usage was almost 95% on all 4 cores during the encode.


More soon...

Suggest settings and comparisons you would like to see.

EDIT: Thank you for you suggestion guys.

As I said in my OP, that there is more to come with different sources at different resolutions.

I'm not trying to advertise anyone here. Just feeding my curiosity to see what kind of encoding does the free CUDA encoder do and probably helping others feeling the same way in the process.

The settings in this encode were used to test the pure speed of x264, to see if it could be as fast as GPU encode at similar or better quality. Since that didn't happen, I will try to match the quality and see what is the difference in performance.

As to the hardware used, this is the best thing I have access to right now. Also as others said GTS 250 is a last generation GPU based on the G92 chip used in 8800GTS 512 MB, 9800GTX, 9800 GTX+.

Also I may use Badaboom and MediaShow Espresso in future if time permits.

Also I plan on using this video

1080p VBR Video Quality Test Streams
Sony HDW-F900 footage, 1080p@25, 18 Mbps average, 30 Mbps peak in a 35 Mbps Transport Stream (259,534,064 bytes)

on this page http://www.w6rz.net/ as one of the sources. Please let me know if there is another uncompressed source you would like me to try.

I'm not too sure about the questions regarding the decoder, I only have ffdshow+haali media splitter installed on my system.
I would really appreciate if you let me know If i need to change something with the decoders.

roozhou
10th July 2009, 08:45
MeGUI uses avisynth to frameserve but mediacoder uses mencoder to decode and pipes raw data to encoders.
With "ultrafast" settings decoding may become a bottleneck for x264.

kumi
10th July 2009, 08:56
I wonder what settings MediaCoder used to arrive at these benchmarks (http://blog.mediacoderhq.com/benchmarks-cuda-h-264-vs-x264/)?

stanleyhuang
10th July 2009, 08:59
Absolutely. Though the decoding is done in the separate process and there is a large ring-buffer to connect decoder and encoder, the decoding is still a bottleneck on a fast multi-core processor. Fortunately mplayer-mt/ffmpeg-mt has multi-threaded H.264 decoding.

MeGUI uses avisynth to frameserve but mediacoder uses mencoder to decode and pipes raw data to encoders.
With "ultrafast" settings decoding may become a bottleneck for x264.

Dark Shikari
10th July 2009, 09:03
Why did you use such retardedly low quality settings with x264?

At a minimum your goal should be to match the quality of the two; it makes no sense to test two encoders against each other in terms of speed at vastly different compression settings.

Also, if you're going to compare two encoders, use raw video input, not a highly compressed H.264 stream whose decoding method differs between the two encoders.

stanleyhuang
10th July 2009, 09:06
I think Q9450 and GTS250 is not the hardware of the same level, at least not the same price. ;-)

Dark Shikari
10th July 2009, 09:13
I think Q9450 and GTS250 is not the hardware of the same level, at least not the same price. ;-)This as well. It's rather disingenuous to pick a last-generation CPU and compare to a current GPU (and then question why the former is slower).

stanleyhuang
10th July 2009, 09:14
Actually x264 do generate better quality when it is configured for maximum quality, but that will also make the transcoding extremely slow.

Fr4nz
10th July 2009, 09:18
This is a totally borked comparison...

Dark Shikari
10th July 2009, 09:19
Actually x264 do generate better quality when it is configured for maximum quality, but that will also make the transcoding extremely slow.You hardly need "maximum quality"; even the (reasonable) faster settings beat the crappy CUDA encoder easily quality-wise, as has been tested dozens of time before with exactly this encoder.

Of course, it doesn't help that he even turned off subpixel motion vectors though... his settings are completely ridiculous.This is a totally b0rked comparison...Yes, and he posted it in a very official-looking fashion despite the entire thing being done in a completely idiotic and haphazard manner.

I mean seriously:

1. Pick a GPU that's more costly and faster than the CPU.
2. Go out of the way to pick the worst possible settings for x264 (literally!)
3. Use two different decoders to feed the different encoders.
4. Show how the GPU encoder looks so much better than x264.

I'm not going to bother with this anymore as it's clear that this guy is here solely to try to advertise crappy encoders by performing intentionally bad tests.

stanleyhuang
10th July 2009, 09:30
I've published the x264 options in my benchmark (http://blog.mediacoderhq.com/benchmarks-cuda-h-264-vs-x264/). Under this configuration, both encoders have near (x264 is slightly better) output quality. The CPU I used costs about US$ 195, the GPU (display adapter with 896MB onboard GDDR3) I used costs about US$ 235.

PS1: For serious encoding, I myself use x264.
PS2: I don't know and have no relationship with St Devious and I don't think he can benefit anything by advertising the crappy encoder. I just saw this post by a back-link to my blog.

Fr4nz
10th July 2009, 10:04
2. Go out of the way to pick the worst possible settings for x264 (literally!)

I think this is the most important point that invalidates the comparison: you have to use *possibily* the same settings in both encoders in order to make a credible comparison.

stanleyhuang
10th July 2009, 10:13
He might just want x264 to work as fast as it can.

roozhou
10th July 2009, 10:52
He might just want x264 to work as fast as it can.

I wonder why x264 ran so slowly with --preset ultrafast on a Quad-core.

slavickas
10th July 2009, 11:19
This as well. It's rather disingenuous to pick a last-generation CPU and compare to a current GPU (and then question why the former is slower).

err no GTS 250 = 9800GTX+ ~= 8800 GTS

Reimar
10th July 2009, 11:44
I wonder why x264 ran so slowly with --preset ultrafast on a Quad-core.

Probably due to decoding speed. Unfortunately it is unclear which decoders were used. If either the source was uncompressed or at least DXVA+readback was used for x264 or the source used some format that the GPU can't accelerate it might make some sense as an encoder comparison so far the main conclusions is: A fast GPU can decode H.264 a lot faster than a slow CPU with some random (probably single-threaded) decoder. Not exactly news. And not in any way related to x264.

roozhou
10th July 2009, 11:53
Probably due to decoding speed. Unfortunately it is unclear which decoders were used. If either the source was uncompressed or at least DXVA+readback was used for x264 or the source used some format that the GPU can't accelerate it might make some sense as an encoder comparison so far the main conclusions is: A fast GPU can decode H.264 a lot faster than a slow CPU with some random (probably single-threaded) decoder. Not exactly news. And not in any way related to x264.

How can one perform DXVA+readback? Is there any opensource implementation available?

ajp_anton
10th July 2009, 12:06
This as well. It's rather disingenuous to pick a last-generation CPU and compare to a current GPU (and then question why the former is slower).He DID pick a last-generation GPU.

And about picking the "worst possible settings", he's trying to match the speeds, not the quality.

However, it doesn't say if x264 was able to use all cores. It says "95% on all 4 cores", was this during the x264 encode? And why only show a picture of the last core? There's a nice picture of the task manager with all 4, separately or combined.
Not to mention what decoder was used (use uncompressed), and why not use the source of the source?

Reimar
10th July 2009, 13:07
How can one perform DXVA+readback? Is there any opensource implementation available?

I actually didn't mean to imply it is possibly, I don't know (I know it is possible on Linux with VDPAU), but given that IDirectXVideoDecoderService::CreateVideoDecoder takes IDirect3DSurface9 as render target I'd expect you should be able to read back from those surfaces.
That is DXVA2 only though...

tph
10th July 2009, 13:19
Any video encoder comparison needs to use raw video as input, otherwise you're benchmarking decoder performance.

LoRd_MuldeR
10th July 2009, 13:27
Any video encoder comparison needs to use raw video as input, otherwise you're benchmarking decoder performance.

Usually the decoder should be orders of magnitude faster than the encoder. So the time for decoding should be negligible.

Reading "raw" data from the HDD may actually become a bottleneck, especially for HD content...

CruNcher
10th July 2009, 13:35
as the others already said a lot here is flawed settings wise comparing main vs high then the subme is totally off 0 for x264 is laughable keyint 250 vs 15 is also a good joke ;) also would have been nice if you could add Elemental Technologies Badaboom to the test (Main Profile) :)

@Lord_Mulder
not entirely true especially when you decode more complex streams +higher resolution (Full HD) on Nvidias Hardware Decoder (VP2/VP3) it can be more beneficial speed wise in a encoding chain (especially if you add other GPU effects additionaly as denoising or deinterlacing to it) it heavily depends on your usage scenario Badaboom for example does this quiet good by decoding all the inputs it supports on the VP2 same scenario can be simulated for x264 with Donald Graft his Nvidia API Decoder :)

And it depends on a lot more for example you have a older CPU and want to encode faster now buying a complete new system maybe means to make a complete architecture change (you calculate the costs your budget explodes) now you have a old GFX card inside and decide to buy say a 8800 GT you only need to change the GFX card get future OS gimick support , 3D Games support, Video Encoding and Enhancing including a very powerful Hardware HD Decoder all in one (that is a pretty good deal) not to say all the other CUDA/OpenCL able applications that are still to come :)

you can say what you want but the ION platform shows here what is possible in a very small form factor with this combination putting a i7 quadcore in that case would melt it and every efficiency be wasted ;)

And the extremest case what energy usage is worth you which encoding quality ? (though this one mostly doesn't apply for end consumers as most of the times they don't care about this question)

LoRd_MuldeR
10th July 2009, 13:56
If I get you right, you are saying that we can speed-up the decoding part even more by using GPU-accelerated decoders.

So this supports my point that using "raw" (uncompressed) input may not be the best idea.

I'd rather use a fast lossless compressor, such as HuffYUV, for the source instead of uncompressed YUV/RGB data...

CruNcher
10th July 2009, 14:00
ah sorry misunderstood you meant decoder in general not CPU Decoding here :) yep you right then of course raw is problematic in terms of IO and the test should be done on separate harddrives (in/out) and best with separate hardware controllers (to lower CPU utilization on a consumer system) or directly from RAM :)

Elemental Technologies developed based on all these requirements their own GPU powered Encoder rack :) http://www.elementaltechnologies.com/products/server
combining all the features of the GPU with the Power of the CPU in a Energy Efficient way (sure not yielding the best quality but for the bitrate target of it a very efficient quality HVS wise with a very good energy output).

St Devious
10th July 2009, 14:48
Thank you for you suggestion guys.

As I said in my OP, that there is more to come with different sources at different resolutions.

I'm not trying to advertise anyone here. Just feeding my curiosity to see what kind of encoding does the free CUDA encoder do and probably helping others feeling the same way in the process.

The settings in this encode were used to test the pure speed of x264, to see if it could be as fast as GPU encode at similar or better quality. Since that didn't happen, I will try to match the quality and see what is the difference in performance.

As to the hardware used, this is the best thing I have access to right now. Also as others said GTS 250 is a last generation GPU based on the G92 chip used in 8800GTS 512 MB, 9800GTX, 9800 GTX+.

Also I may use Badaboom and MediaShow Espresso in future if time permits.

Also I plan on using this video

1080p VBR Video Quality Test Streams
Sony HDW-F900 footage, 1080p@25, 18 Mbps average, 30 Mbps peak in a 35 Mbps Transport Stream (259,534,064 bytes)

on this page http://www.w6rz.net/ as one of the sources. Please let me know if there is another uncompressed source you would like me to try.

I'm not too sure about the questions regarding the decoder, I only have ffdshow+haali media splitter installed on my system.
I would really appreciate if you let me know If i need to change something with the decoders.

@ ajp_anton - thanks for pointing that out. I'll put the x264 CPU Usage in the next encode. The image that I put up and 95% on all 4 cores is actually the CPU Usage during the GPU encode. And the image shows the usage of all 4 cores. If you look closely, each one of them is color coded and all of them are in there. It just so happens that they are hidden behind the last one as the usage is about the same on all.

CruNcher
10th July 2009, 15:04
you can leave MediaShow Espresso away it will only let Nvidias Encoder look worse as they don't have any idea of what they doing with Nvidias API in their GUI over there @ Cyberlink i have the feeling ;)
Stans Cli Encoder Interface to cuvidenc.dll is currently the most purest way to make usage of Nvidias Encoder API via cuvidenc.dll :) the other do to much insane stuff in their input Logic (most of them are optimized for Mobile Encoding and buggy as hell) for advanced usage ;)
You only need to use GPU Decoding in a x264 chain if you compare vs Badaboom and only if the input is VP2 Decodable in your case Stans encoder is only a interface to Nvidias API so no Decoding going on there :)

Just to make it clear there are only 2 Nvidia Capable Cuda Encoder Cores currently (Nvidia,Elemental Technologies)

Nvidia = cuvidenc.dll coming with Nvidias driver (used by almost all 3rd party Cuda advertised products Mediashow Espresso, Nero MoveIt, Loiloscape, PowerDirector 7, Vreveal HD, Stans Cli Encoder (Mediacoder))

Elemental Technologies = Badaboom (Consumer Main Profile only),Elemental Accelerator (Professional High Profile),Elemental Server (Professional High Profile)

St Devious
10th July 2009, 15:08
you can leave MediaShow Espresso away it will only let Nvidias Encoder look worse as they don't have any idea of what they doing with Nvidias API in their GUI over there @ Cyberlink i have the feeling ;)
Stans Encoder is currently the most purest way to make usage of Nvidias Encoder API via cuvidenc.dll :) the other do to much insane stuff in their input Logic (most of them are optimized for Mobile Encoding and buggy as hell) for advanced usage ;)
You only need to use GPU Decoding in a x264 chain if you compare vs Badaboom and only if the input is VP2 Decodable in your case Stans encoder is only a interface to Nvidias API so no Decoding going on there :)

ok, mediashow espresso is out for this round.

So If I use the video I suggested in my previous post, do i need to worry about decoding holding back the encoding ? Do I need to change anything around with the decoders ?

CruNcher
10th July 2009, 15:20
Only if you compare vs Badaboom as it will decode the Mpeg-2 on your GPU :) best to avoid any IO differences would be decoding from a Ramdisk and additionally as output for the speed estimation into NUL.

St Devious
10th July 2009, 15:26
Only if you compare vs Badaboom as it will decode the Mpeg-2 on your GPU :)

what about x264 vs stan's CUDA H.264 encoder with that mpeg-2 sample ? do i need to change anything ?

CruNcher
10th July 2009, 15:40
CABAC on
Subme 3
partitions default
scenecut 1
bframes 2
8x8dct 1

Stans Cli Encoder
-idrp 250

not quiet sure about the partitioning though but none is definitely not reflecting Nvidias Encoder


PS: I practically have the same card as you 8800 GT (G92) but another CPU also Dual Core Athlon XP 64 Toledo :)

i would also add my results here your Core 2 Quad should be a lot more efficient at least in the CPU Encoding part :)

LoRd_MuldeR
10th July 2009, 16:04
Just to make this clear:

It's not Stan's H.264 encoder. It's just Stan's interface to NVIDIA's H.264 encoder. The very same encoder we have seen producing crap quality in other applications ;)

Also that "encoder" was removed from the MediaCoder package due to licensing issues. At least that was the case when I last checked...

CruNcher
10th July 2009, 16:08
Yes and it's clear that Nvidia didn't like that ;) especialy the 3rd parties hehe ;D

@Lord_Mulder
anyway it was problematic with those 3rd party apps to get on the spot target results either they where crippled or buggy :P

St Devious
10th July 2009, 16:15
Just to make this clear:

It's not Stan's H.264 encoder. It's just Stan's interface to NVIDIA's H.264 encoder. The very same encoder we have seen producing crap quality in other applications ;)

Also that "encoder" was removed from the MediaCoder package due to licensing issues. At least that was the case when I last checked...

Alright, let's call it Mediacoder CUDA encoder ?

And it is back in the latest release.

Yes and it's clear that Nvidia didn't like that ;) especialy the 3rd parties hehe ;D

@Lord_Mulder
anyway it was problematic with those 3rd party apps to get on the spot target results either they where crippled or buggy :P

I think Mediacoder has more options to configure the encoder than Badaboom, unless I'm wrong. Don't recall seeing High profile in Badaboom

@CruNcher - sure, I'll add your results too

CruNcher
10th July 2009, 16:18
But Badaboom is bad to compare it's a unique core compared to the 3rd party apps and yes the consumer version has yet no High Profile support but there are sightings in the last version of High Profile and AQ support which could come soon :)

Oh its back :) so Stan payed his royalities ;)

dj_tjerk
10th July 2009, 17:22
First of all I have to say it's quite useless testing at your decoding/reading/(bitstream writing?) maximum (In your case, 114fps, in my case it was 118 fps). But I can't make nvidia's encoder go any slower too (as in.. give better quality) so it can't be helped.
Secondly, you're probably showing like one of the first frames for the output files, which is nice but doesn't really work for single pass --bitrate encodes. NVidia's encoder is apparantly better at that though.

If you wanna have x264 encode at the max speed too (as was the case on my system here), you might wanna go for sth like --bframes 0 --subme 5 and remove the --no-cabac. That caused no slowdown for me whatsoever in comparison to ultrafast's defaults.

Also, I looked at the output of the file encoded with mediacoder/cuda, and leaving aside that the colors are totally off, it looks a lot like abusing ttempsmooth(). Surfaces are moving nice and steady, but every couple frames almost the entire frame changes.

stanleyhuang
10th July 2009, 17:32
I'm not trying to advertise anyone here. Just feeding my curiosity to see what kind of encoding does the free CUDA encoder do and probably helping others feeling the same way in the process.
Hoping some people can free themselves from prejudices.

As to the hardware used, this is the best thing I have access to right now. Also as others said GTS 250 is a last generation GPU based on the G92 chip used in 8800GTS 512 MB, 9800GTX, 9800 GTX+.
What I was thinking is that you are using a more advanced CPU and a less advanced GPU.

BTW: I think you can use MediaCoder to do benchmark for both x264 and CUDA H.264 encoder. This will eliminate the difference in decoding.

stanleyhuang
10th July 2009, 17:35
Also that "encoder" was removed from the MediaCoder package due to licensing issues. At least that was the case when I last checked...

The encoder is re-added (since 0.7.1.4470) now after we signed a formal license with nvidia. We are currently working on our own cuda-based video filtering features including down-scaling, de-interlacing, 3D-denoising and pull-up.

CruNcher
10th July 2009, 17:42
Stanley does Nvidia provide the API for their Cuda Encoder for free in the CUda 2.2 SDK does the usage of it needs to be licensed (payed) i guess so i mean you virtualy a competitor now to all those 3rd party applications (Cyberlink,Nero,Loilo) that make money from selling this :D

stanleyhuang
10th July 2009, 17:46
I am sorry I don't think I can disclose this due to the NDA we signed.

CruNcher
10th July 2009, 17:48
Hehe Cyberlink and Co will be pissed if that goes around (though they have the nicer GUI) ;) but anyway Nvidia sells their cards and is happy :D
How you 3rd parties fight against each other is not their thing :D
Yeah sure understand that confidential stuff highly business critical lol ;)

Manao
10th July 2009, 18:04
If you wanna have x264 encode at the max speed too (as was the case on my system here), you might wanna go for sth like --bframes 0 --subme 5 and remove the --no-cabac. That caused no slowdown for me whatsoever in comparison to ultrafast's defaults.You'd better keep bframes too. They usually reduce bitrate by 20% at the same quality (cabac is "only" 10%, deblocking is 10% also).

St Devious
10th July 2009, 18:05
Hoping some people can free themselves from prejudices.

not sure what you mean. But I know that CPU encoding is where its at, and GPU encoding is still in its infancy and can't compete with x264 in its current form. When I see slides from Nvidia and ATI and review sites that GPU encoding is 5 times faster than CPU encoding and there is no mention of image quality, that makes me kinda angry.

So I'm just trying to provide proof.

What I was thinking is that you are using a more advanced CPU and a less advanced GPU.

Not sure how you would measure that.

BTW: I think you can use MediaCoder to do benchmark for both x264 and CUDA H.264 encoder. This will eliminate the difference in decoding.

That's a good idea but as in business they say, you can't make everyone happy, I fear I would get flamed for that too for trying to advertise mediacoder

dj_tjerk
10th July 2009, 18:25
You'd better keep bframes too. They usually reduce bitrate by 20% at the same quality (cabac is "only" 10%, deblocking is 10% also).
Maybe so, but in my case --bframes 1 with --no-cabac and --subme 0 caused it to slow down to about 108 fps (whereas setting --bframes 0 and --subme 5 and removing --no-cabac caused no slowdown whatsoever from my maximum). So I got subme and cabac for "free".

St Devious however chose to set bframes 1 before changing anything else, eventhough with his goal of comparing quality at the same speed (i.e. maximum possible speed), he should've enabled cabac and subpixel refinements, since those would've most likely caused x264 to encode at the same speed as cuda (assuming his system behaves the same as mine does).

St Devious
10th July 2009, 18:49
St Devious however chose to set bframes 1 before changing anything else, eventhough with his goal of comparing quality at the same speed (i.e. maximum possible speed), he should've enabled cabac and subpixel refinements, since those would've most likely caused x264 to encode at the same speed as cuda (assuming his system behaves the same as mine does).

I am under the impression that presets override any settings. is that the case ?

LoRd_MuldeR
10th July 2009, 18:52
I am under the impression that presets override any settings. is that the case ?

Nope. Presets don't overwrite anything! If a preset is selected, it is applied first. Then all explicit options are applied (and may overwrite what the preset has set).

In other words: Instead of starting from the defaults, we start from the selected preset. But that's all.

Maybe you are thinking of the new "--profile" option (which is applied after the explicit options). That option does overwrite your settings, as it enforces a certain level!

St Devious
10th July 2009, 19:08
Presets don't overwrite anything! If a preset is selected, it is applied first. Then all explicit options are applied (and may overwrite what the preset has set).

ok then I will try to set the same option as the preset ultrafast in the MeGUI x264 settings GUI

stanleyhuang
11th July 2009, 05:00
not sure what you mean.
I mean the people claiming that you intend to make these benchmarks to advertise something. I hope they can be a bit open-minded.

That's a good idea but as in business they say, you can't make everyone happy, I fear I would get flamed for that too for trying to advertise mediacoder
That's likely. ;-)
So every word should be neutral here.

St Devious
11th July 2009, 06:03
I'm comparing now with best quality that each encoder can provide at a certain bitrate.

Using this video

1080p VBR Video Quality Test Streams
Sony HDW-F900 footage, 1080p@25, 18 Mbps average, 30 Mbps peak in a 35 Mbps Transport Stream (259,534,064 bytes)

on this page http://www.w6rz.net/

Encoding both encoders to 4Mbps and 6 Mbps.

Using 2 pass unrestricted HQ profile in Megui, would extra quality provide some better quality ?

And here are the settings for CUDA encoder in Mediacoder. Please let me know what to set to compare to match x264

http://i26.tinypic.com/2gv1vsl.jpg

stanleyhuang
11th July 2009, 06:12
The CUDA H.264 encoder in current release doesn't support 2-pass mode yet.

St Devious
11th July 2009, 06:29
The CUDA H.264 encoder in current release doesn't support 2-pass mode yet.

i meant used 2 pass unrestricted HQ for x264.

But How can i configure the CUDA encoder to provide maximum quality ?

Here is an update, click on the images to get full size 1920x1080 image

x264 @ 6 Mbps
http://www4.picturepush.com/photo/a/1958117/220/1958117.png (http://www.picturepush.com/public/1958117)

CUDA @ 6 Mbps
http://www4.picturepush.com/photo/a/1958117/220/1958117.png (http://www.picturepush.com/public/1958118)


Source
http://www1.picturepush.com/photo/a/1958119/220/1958119.png (http://www.picturepush.com/public/1958119)

Mediainfo

Source
Video
ID : 49 (0x31)
Menu ID : 1 (0x1)
Format : MPEG Video
Format version : Version 2
Format profile : Main@High
Format settings, Matrix : Default
Duration : 2mn 7s
Bit rate mode : Variable
Bit rate : 32.5 Mbps
Nominal bit rate : 30.0 Mbps
Width : 1 920 pixels
Height : 1 080 pixels
Display aspect ratio : 16/9
Frame rate : 25.000 fps
Colorimetry : 4:2:0
Scan type : Progressive
Bits/(Pixel*Frame) : 0.628
Stream size : 493 MiB (92%)


x264
Video
ID : 1
Format : AVC
Format/Info : Advanced Video Codec
Format profile : High@L5.1
Format settings, CABAC : Yes
Format settings, ReFrames : 5 frames
Codec ID : avc1
Codec ID/Info : Advanced Video Coding
Duration : 2mn 7s
Bit rate mode : Variable
Bit rate : 6 000 Kbps
Maximum bit rate : 16.6 Mbps
Width : 1 920 pixels
Height : 1 080 pixels
Display aspect ratio : 16/9
Frame rate mode : Constant
Frame rate : 25.000 fps
Resolution : 24 bits
Colorimetry : 4:2:0
Scan type : Progressive
Bits/(Pixel*Frame) : 0.116
Stream size : 90.9 MiB (100%)
Writing library : x264 core 68 r1179M 96e2229
Encoding settings : cabac=1 / ref=5 / deblock=0:0:0 / analyse=0x3:0x133 / me=umh / subme=7 / psy_rd=1.0:0.0 / mixed_ref=1 / me_range=16 / chroma_me=1 / trellis=2 / 8x8dct=1 / cqm=0 / deadzone=21,11 /
chroma_qp_offset=-2 / threads=6 / nr=0 / decimate=1 / mbaff=0 / bframes=3 / b_pyramid=1 / b_adapt=2 / b_bias=0 / direct=3 / wpredb=1 / keyint=250 / keyint_min=25 / scenecut=40 / rc=2pass / bitrate=6000 / ratetol=1.0 /
qcomp=0.60 / qpmin=10 / qpmax=51 / qpstep=4 / cplxblur=20.0 / qblur=0.5 / ip_ratio=1.40 / pb_ratio=1.30 / aq=1:1.00


CUDA
Video
ID : 1
Format : AVC
Format/Info : Advanced Video Codec
Format profile : High@L5.1
Format settings, CABAC : Yes
Format settings, ReFrames : 2 frames
Codec ID : avc1
Codec ID/Info : Advanced Video Coding
Duration : 2mn 7s
Bit rate mode : Variable
Bit rate : 6 035 Kbps
Maximum bit rate : 42.0 Mbps
Width : 1 920 pixels
Height : 1 080 pixels
Display aspect ratio : 4/3
Frame rate mode : Constant
Frame rate : 25.000 fps
Resolution : 24 bits
Colorimetry : 4:2:0
Scan type : Progressive
Bits/(Pixel*Frame) : 0.116
Stream size : 91.4 MiB (100%)

Selur
11th July 2009, 07:28
@St Devious: your x264 and your CUDA Screenshot link to the same picture,... (CUDA should be: http://www.picturepush.com/public/1958118)

RunningSkittle
11th July 2009, 08:17
Levels are different in the x264 image as well.

St Devious
11th July 2009, 14:04
@St Devious: your x264 and your CUDA Screenshot link to the same picture,... (CUDA should be: http://www.picturepush.com/public/1958118)


thanks, corrected it now.

Levels are different in the x264 image as well.

what do you mean ?

LoRd_MuldeR
11th July 2009, 14:20
what do you mean ?

He talks about luminance levels, I think!

http://forum.doom9.org/showthread.php?t=143689

But I'm not so sure whether this is a level problem, because the "black" intensity looks identical in both screenshots.
Typically you would get "black" vs. "dark gray" when comparing pictures of different luminance levels, but that's not the case here.
Looks more like something is wrong with the color intensities. Look at the red shirt and the cyan bag!

buba king
11th July 2009, 14:21
what do you mean ?

The colors are different in the x264 screen shot... probably a bad color conversion somewhere.

St Devious
11th July 2009, 14:26
MeGUI had me add ConvertToYV12() at the end of script, is that a problem ?

LoRd_MuldeR
11th July 2009, 14:54
MeGUI had me add ConvertToYV12() at the end of script, is that a problem ?

Nope. That's because x264 doesn't accept anything but YV12() data. But you should use the very same source script for both encoders ;)

St Devious
11th July 2009, 15:12
Nope. That's because x264 doesn't accept anything but YV12() data. But you should use the very same source script for both encoders ;)

hmm.. didn't think that way. Guess that may have invalidated the above images.

Will encode with an avisynth script when i get home tomm.

btw what about Stan's suggestion of using Mediacoder to encode with both CUDA and x264 ? that wa decoders should be the same.

LoRd_MuldeR
11th July 2009, 15:23
btw what about Stan's suggestion of using Mediacoder to encode with both CUDA and x264 ? that wa decoders should be the same.

Given that MediaCoder really uses the same decoders, processing and colorspace conversion for both encoders ;)

I'd go with Avisynth to be 100% sure that the source data is identical for both encoders!

The less steps there are between your source and the encoder, the less things can corrupt your results.

Directly feeding the source into the encoder from Avisynth (using the identical script) is the most reliable method.

St Devious
11th July 2009, 15:30
Given that MediaCoder really uses the same decoders, processing and colorspace conversion for both encoders ;)

I'd go with Avisynth to be 100% sure that the source data is identical for both encoders!

The less steps are between your source and the encoder, the less things can influence/corrupt your results.

Directly feeding the source into the encoder from Avisynth (using the identical script) is the most reliable method...

thanks.

How does the decoding happen ?

Like what will Mediacoder and MeGUI use to decode the avisynth file ?

And how does the avisynth file decode the .ts source file ?

LoRd_MuldeR
11th July 2009, 15:38
Like what will Mediacoder and MeGUI use to decode the avisynth file ?

If you use Avisynth input, they won't decode anything! They'll receive uncompressed video data from Avisynth ;)

Which doesn't exclude that MediaCoder applies color conversions or other processing on that data...


And how does the avisynth file decode the .ts source file ?

Depends on your Avisynth script. What XYZSource() filter did you use in your script?

St Devious
11th July 2009, 16:08
Depends on your Avisynth script. What XYZSource() filter did you use in your script?

don't remember atm, but whatever MeGUI Avisynth Script creator uses when you feed .d2v to it

LoRd_MuldeR
11th July 2009, 16:15
don't remember atm, but whatever MeGUI Avisynth Script creator uses when you feed .d2v to it

MPEG2Source() respectively DGDecode :D

St Devious
11th July 2009, 16:24
MPEG2Source() respectively DGDecode :D

alright, so avisynth decodes the MPEG2 Source, and then feeds uncompressed data to any encoder and that way it eliminates any change in color or anything.

Learned something new today ! :D

How the resource usage by the decoding ? like some were saying earlier in the thread that decoding might have caused encoding to slow down.

Sharktooth
11th July 2009, 16:30
DGDecode uses a single threaded reference MPEG2 decoder... so, for better results, i'd suggest to use DGNVTools... or at least DGMpegDecNV...

St Devious
11th July 2009, 16:47
DGDecode uses a single threaded reference MPEG2 decoder... so, for better results, i'd suggest to use DGNVTools... or at least DGMpegDecNV...

do you mean instead of using MPEG2Source() in the avisynth script, I use DGMpegDecNV ?

Haven't used it before. Isn't that neuron's software and you have to pay to use it ?

Sharktooth
11th July 2009, 16:48
yep it is. but if you own and nvidia card it is a MUST.

St Devious
11th July 2009, 17:00
yep it is. but if you own and nvidia card it is a MUST.

I could do that, but my worry is that people might not be able to reproduce my results, since everyone wouldn't be able to buy that software.

LoRd_MuldeR
11th July 2009, 17:06
I could do that, but my worry is that people might not be able to reproduce my results, since everyone wouldn't be able to buy that software.

If people can't use DGMpegDecNV, they can't use NVIDIA's CUDA encoder anyway :p

But be aware that DG's NV tools are payware...

St Devious
11th July 2009, 17:25
If people can't use DGMpegDecNV, they can't use NVIDIA's CUDA encoder anyway :p

But be aware that DG's NV tools are payware...

That's what I meant, they can't use it because it's payware.

Sagekilla
11th July 2009, 17:34
@LoRd_MuldeR: Would using the GPU for decoding then using it for encoding create any noticeable impact or is it essentially "free" decoding?

St Devious
11th July 2009, 17:36
@LoRd_MuldeR: Would using the GPU for decoding then using it for encoding create any noticeable impact or is it essentially "free" decoding?

I think it should be free decoding, since the encoding part is not using 100% GPU, at least that's what I have found on ATI Radeons. But I imagine nvidia's CUDA can't be that efficient at this stage to be able to use 100% of GPU.

LoRd_MuldeR
11th July 2009, 17:48
@LoRd_MuldeR: Would using the GPU for decoding then using it for encoding create any noticeable impact or is it essentially "free" decoding?

Isn't the decoding happening on a specialized chip (e.g. VP2) anyway? So the actual GPU shouldn't be involved at all.

But I imagine nvidia's CUDA can't be that efficient at this stage to be able to use 100% of GPU.

It's more a question whether your CUDA application (kernel) is able to keep all the multiprocessors of the GPU busy all the time ;)

Most applications simply don't scale well enough (because the tasks aren't independent/parallelizable) or they suffer from memory bottlenecks (bank conflicts).

CUDA simply isn't the perfect platform for video encoding. Or in other words: Video encoding isn't the perfect application for CUDA.

If CUDA was as good for video encoding as they try to make us believe, we would have seen a competitive CUDA encoder until now. But there is none...

CruNcher
11th July 2009, 17:50
The Decoding doesn't utilize the GPU in anyway VP2 is it's own processor inside the GPU that runs with 400 Mhz and is ARM based (it is it's own logic) same as with UVD (ATI) :)
doing anything on it wont slowdown the GPU part of the encoding as that is done via CUDA

You can watch a Blu-Ray Movie and @ the side of it Play Crysis it wont impact the game alot :P (also energy wise the Playback part would roughly take 3-5 Watts the Game part alot more like 100 watts depends on the GPU and Game Settings)

Think of it like your Standalone Player integrated inside the GPU fully independent of the rest (only the actuall bus transfer could make problems Playing Crysis and a Blu-Ray + AACS Decryption @ the same time ;))

What this makes possible Nvidia showed in a Hard test on the ION platform doing multiple tasks at once without majorly slowing down anything :) (Encoding,Playback,Editing) @ the same time.

Though for this to work properly you need a OS (and driver Model) that is optimized for a GPU that is where Windows XP lacks it was not designed to manage GPU resources (when utilized by multiple tasks) very efficiently and where Vista and Win7 Shines :)

St Devious
11th July 2009, 18:00
Isn't the decoding happening on a specialized chip (e.g. VP2) anyway? So the actual GPU shouldn't be involved at all.


The Decoding doesn't utilize the GPU in anyway VP2 is it's own processor inside the GPU that runs with 400 Mhz and is ARM based (it is it's own logic) same as with UVD (ATI) :)
doing anything on it wont slowdown the GPU part of the encoding as that is done via CUDA

You can watch a Blu-Ray Movie and @ the side of it Play Crysis it wont impact the game alot :P (also energy wise the Playback part would roughly take 5 Watts the Game part alot more like 100 watts depends on the GPU and Game Settings)

Think of it like your Standalone Player integrated inside the GPU fully independent of the rest (only the actuall bus transfer could make problems Playing Crysis and a Blu-Ray + AACS Decryption @ the same time ;))

What this makes possible Nvidia showed in a Hard test on the ION platform doing multiple tasks at once without majorly slowing down anything :) (Encoding,Playback,Editing) @ the same time.

Thanks, got it !

STaRGaZeR
11th July 2009, 20:57
The Decoding doesn't utilize the GPU in anyway VP2 is it's own processor inside the GPU that runs with 400 Mhz and is ARM based (it is it's own logic) same as with UVD (ATI) :)

UVD runs at core frequency in ATI cards.

blubberbirne
11th July 2009, 22:09
MediaCoder Cuda H264 Encoder don't work with H264 TS Source Files for me.
I encodet a 45.2sec TS Stream. The Result is a 1min 40sec clip. Frame Rate is in Source and Encoded File 25fps

stanleyhuang
12th July 2009, 04:09
Our next step is to use hardware for H.264 decoding, possible by using API of VP2.

stanleyhuang
12th July 2009, 04:10
MediaCoder Cuda H264 Encoder don't work with H264 TS Source Files for me.
I encodet a 45.2sec TS Stream. The Result is a 1min 40sec clip. Frame Rate is in Source and Encoded File 25fps

Have you set a deinterlancer?
You can discuss this on mediacoder forum.

blubberbirne
12th July 2009, 08:50
Sorry, no deinterlacer set.

CiNcH
12th July 2009, 09:35
The Decoding doesn't utilize the GPU in anyway VP2 is it's own processor inside the GPU that runs with 400 Mhz and is ARM based (it is it's own logic) same as with UVD (ATI)
H.264 HD should be decodeable in software on a 400 MHz ARM? Is there some special hardware on the ARM?

CruNcher
12th July 2009, 11:29
It seems so look @ Tegra what that little thing is capable off soon first devices will start the era of mobile 1080p it works since the G92 was released soon it's market time (devices where already announced for end of 2008 but seems the crisis killed those plans) :) now most probably Microsoft will be the first to release it to market in form of the Zune HD Mobile Device.


The most interesting Videos about it (though this is compared to the decoder logic only like in the G92 a whole ARM CPU)


http://www.youtube.com/watch?v=AXinHWiat2s
http://www.youtube.com/watch?v=2QJ-ETt3kMk (shows a workaround to in browser Flash Hardware Decoding they don't get it running in the browser yet (see Adobe announcement) ;) )
http://www.youtube.com/watch?v=f5sEIf5-FJs
http://www.youtube.com/watch?v=qWqlKBp9qQ0
http://www.youtube.com/watch?v=lTEaWfTO-zE
http://www.youtube.com/watch?v=RChFjrcBsx4 (Future Roadmap of Tegra)

708145
12th July 2009, 12:57
H.264 HD should be decodeable in software on a 400 MHz ARM? Is there some special hardware on the ARM?

The ARM is just controlling the special purpose hardware blocks used for bitstream processing (BSP), decoding and "image enhancements" (VP2/3).

bis besser,
T0B1A5

stanleyhuang
12th July 2009, 13:03
H.264 HD should be decodeable in software on a 400 MHz ARM? Is there some special hardware on the ARM?

Quite likely with DSP extension.

St Devious
12th July 2009, 16:32
can sombody help me tune the CUDA encoder for best quality

http://i26.tinypic.com/2gv1vsl.jpg

ChronoCross
12th July 2009, 17:35
Doesn't MediaCoder Violate the GPL? While not relevant to the discussion here in this thread I figure I'd mention it as no one should be using software that violates a license, it's much like how we don't discuss pirated material.

http://roundup.ffmpeg.org/roundup/ffmpeg/issue1162

RunningSkittle
12th July 2009, 21:32
Doesn't MediaCoder Violate the GPL? While not relevant to the discussion here in this thread I figure I'd mention it as no one should be using software that violates a license, it's much like how we don't discuss pirated material.

http://roundup.ffmpeg.org/roundup/ffmpeg/issue1162

The encoder is re-added (since 0.7.1.4470) now after we signed a formal license with nvidia. We are currently working on our own cuda-based video filtering features including down-scaling, de-interlacing, 3D-denoising and pull-up.

ChronoCross
13th July 2009, 03:18
Originally Posted by stanleyhuang
The encoder is re-added (since 0.7.1.4470) now after we signed a formal license with nvidia. We are currently working on our own cuda-based video filtering features including down-scaling, de-interlacing, 3D-denoising and pull-up.


yeah they fixed the nvidia issue however they still violate the GPL and LGPL. Not to mention the donation button and the ads on the site which is quite shameful.....

St Devious
13th July 2009, 04:13
Here's another comparison. This time both x264 and CUDA were encoded from the same avisynth script.

Avisynth Script as made by MeGUI AVISynth Script Creator
DGDecode_mpeg2source("D:\Videos\1080p25.d2v", info=3)
ColorMatrix(hints=true, threads=0)
#deinterlace
#crop
#resize
#denoise


Tried to use best quality settings for x264. Used the Unrestricted 2 pass Extra Quality profile in MeGUI. Changed keyframe interval size to 50 to match the CUDA encoder.

Click on Images for 1920x1080 size image

x264 6Mbps - 1007s
http://www1.picturepush.com/photo/a/1968734/1024/Picture-Box/x264-6Mbps.png (http://www.picturepush.com/public/1968734)

CUDA 6 Mbps - 93s
http://www4.picturepush.com/photo/a/1968732/1024/Picture-Box/CUDA-6Mbps.png (http://www.picturepush.com/public/1968732)

Source
http://www5.picturepush.com/photo/a/1968733/1024/Picture-Box/Source.png (http://www.picturepush.com/public/1968733)

x264 Mediainfo
Video
ID : 1
Format : AVC
Format/Info : Advanced Video Codec
Format profile : High@L5.0
Format settings, CABAC : Yes
Format settings, ReFrames : 8 frames
Codec ID : avc1
Codec ID/Info : Advanced Video Coding
Duration : 2mn 7s
Bit rate mode : Variable
Bit rate : 6 000 Kbps
Maximum bit rate : 16.4 Mbps
Width : 1 920 pixels
Height : 1 080 pixels
Display aspect ratio : 16/9
Frame rate mode : Constant
Frame rate : 25.000 fps
Resolution : 24 bits
Colorimetry : 4:2:0
Scan type : Progressive
Bits/(Pixel*Frame) : 0.116
Stream size : 91.0 MiB (100%)
Writing library : x264 core 68 r1181M 49bf767
Encoding settings : cabac=1 / ref=8 / deblock=1:-1:-1 / analyse=0x3:0x133 / me=umh / subme=8 /
psy_rd=1.0:0.0 / mixed_ref=1 / me_range=16 / chroma_me=1 /
trellis=2 / 8x8dct=1 / cqm=0 / deadzone=21,11 /
chroma_qp_offset=-2 / threads=6 / nr=0 / decimate=1 / mbaff=0
/ bframes=3 / b_pyramid=1 / b_adapt=2 / b_bias=0 / direct=3 /
wpredb=1 / keyint=50 / keyint_min=25 / scenecut=40 /
rc=2pass / bitrate=6000 / ratetol=1.0 / qcomp=0.60 / qpmin=10
/ qpmax=51 / qpstep=4 / cplxblur=20.0 / qblur=0.5 /
ip_ratio=1.40 / pb_ratio=1.30 / aq=1:1.00

Mediainfo on CUDA file
Video
ID : 1
Format : AVC
Format/Info : Advanced Video Codec
Format profile : High@L5.1
Format settings, CABAC : Yes
Format settings, ReFrames : 2 frames
Codec ID : avc1
Codec ID/Info : Advanced Video Coding
Duration : 2mn 7s
Bit rate mode : Variable
Bit rate : 6 043 Kbps
Maximum bit rate : 44.3 Mbps
Width : 1 920 pixels
Height : 1 080 pixels
Display aspect ratio : 4/3
Frame rate mode : Constant
Frame rate : 25.000 fps
Resolution : 24 bits
Colorimetry : 4:2:0
Scan type : Progressive
Bits/(Pixel*Frame) : 0.117
Stream size : 91.6 MiB (100%)

CUDA Settings
http://i25.tinypic.com/2vubtpi.jpg

Here's more from above files

x264 6Mbps
http://www1.picturepush.com/photo/a/1968739/220/1968739.png (http://www.picturepush.com/public/1968739)

CUDA 6Mbps
http://www4.picturepush.com/photo/a/1968737/220/1968737.png (http://www.picturepush.com/public/1968737)

Source
http://www5.picturepush.com/photo/a/1968738/220/1968738.png (http://www.picturepush.com/public/1968738)

roozhou
13th July 2009, 04:22
Tried to use best quality settings for x264. Used the Unrestricted 2 pass Extra Quality profile in MeGUI. Changed keyframe interval size to 50 to match the CUDA encoder.


Your comparison is meaningless because x264 w/ best quality settings will beat CUDA encoder undoubtfully. Everyone here knows that.

St Devious
13th July 2009, 04:26
Your comparison is meaningless because x264 w/ best quality settings will beat CUDA encoder undoubtfully. Everyone here knows that.

By how much ?

Take a look at the second set of pics. I'm a little surprised that there is not much difference. The difference that there is, might not be noticeable when the video is in motion.

6Mbps might have made a difference here, will do a 4Mbps encode now.

Anther round of images. CUDA loses badly here, look at the far trees

CUDA 6Mbps
http://www5.picturepush.com/photo/a/1968743/220/1968743.png (http://www.picturepush.com/public/1968743)

x264 6Mbps
http://www1.picturepush.com/photo/a/1968744/220/1968744.png (http://www.picturepush.com/public/1968744)

Source
http://www2.picturepush.com/photo/a/1968745/220/1968745.png (http://www.picturepush.com/public/1968745)

roozhou
13th July 2009, 04:42
By how much ?
Take a look at the second set of pics. I'm a little surprised that there is not much difference. The difference that there is, might not be noticeable when the video is in motion.
[/IMG][/URL]

Why not upload your encoded clip? A video is more convincing than a single screenshot.

St Devious
13th July 2009, 04:48
Why not upload your encoded clip? A video is more convincing than a single screenshot.

ok, doing that.

In the meantime here's some more images

CUDA 6 Mbps
http://www1.picturepush.com/photo/a/1968749/220/1968749.png (http://www.picturepush.com/public/1968749)

x264 6Mbps
http://www2.picturepush.com/photo/a/1968750/220/1968750.png (http://www.picturepush.com/public/1968750)

Source
http://www3.picturepush.com/photo/a/1968751/220/1968751.png (http://www.picturepush.com/public/1968751)

stanleyhuang
13th July 2009, 04:56
Doesn't MediaCoder Violate the GPL? While not relevant to the discussion here in this thread I figure I'd mention it as no one should be using software that violates a license, it's much like how we don't discuss pirated material.

http://roundup.ffmpeg.org/roundup/ffmpeg/issue1162

Dude, please DO NOT mess this any longer. MediaCoder is not GPLed and it does not contain any GPL code and don't wag your tongue too freely.

stanleyhuang
13th July 2009, 05:02
Your comparison is meaningless because x264 w/ best quality settings will beat CUDA encoder undoubtfully. Everyone here knows that.

But some people still care about the balance between speed and quality in some cases. I think these comparison is a good work from some aspect.

stanleyhuang
13th July 2009, 05:09
can sombody help me tune the CUDA encoder for best quality

http://i26.tinypic.com/2gv1vsl.jpg

The defaults are almost the settings with the best quality. You can still tune-up number of b-frames and peak bitrate options.

stanleyhuang
13th July 2009, 05:13
yeah they fixed the nvidia issue however they still violate the GPL and LGPL. Not to mention the donation button and the ads on the site which is quite shameful.....

So you think only GPLed software can have a donation button and have ads on the software's website right? Lots and lots of freewares as well as free softwares in this world are supported by Google Adsense now.

St Devious
13th July 2009, 05:20
Anybod who wants to download the latest files from the encode mentioned in post #89 http://forum.doom9.org/showpost.php?p=1304901&postcount=89

CUDA 6Mbps
http://www.mediafire.com/file/vxjhznimklz/1080p25 CUDA 6Mbps.mp4

x264 6Mbps
http://www.mediafire.com/file/dzmmzynz2lz/1080p25 x264 6 Mbps.mp4

Word to the wise: You might have to mess around with frame numbers in AvsP in order to get same frame on both the files. Don't know why that is happening.

stanleyhuang
13th July 2009, 05:22
Word to the wise: Mediacoder messes up on the AR. Don't know if its a problem with the CUDA encoder or mediacoder in general.

Have you tried manually setting an aspect ratio for output?

roozhou
13th July 2009, 05:24
Dude, please DO NOT mess this any longer. MediaCoder is not GPLed and it does not contain any GPL code and don't wag your tongue too freely.

MC bundles MODIFIED version of GPLed software, e.g. mplayer/mencoder and x264. But we cannot find the source codes or related patches anywhere. This is a violation of GPL.

St Devious
13th July 2009, 05:27
Have you tried manually setting an aspect ratio for output?

you have a PM.

no, just saw that option. now it works fine. I'll Upload the new file.

stanleyhuang
13th July 2009, 05:59
MC bundles MODIFIED version of GPLed software, e.g. mplayer/mencoder and x264. But we cannot find the source codes or related patches anywhere. This is a violation of GPL.

The patch of MPlayer is freely available here (http://www.mediacoder.cn/dl/mplayer_patch_by_stanley.zip). Welcome to improve it.
Yes the bundled x264 is modified, but not modified by me. Actually it's from http://www.x264.nl as I found it's faster than my own build.
In addition, the bundled FFmpeg is from here (http://oss.netfarm.it/mplayer-win32.php) and here (http://ffmpeg.arrozcru.org/builds/)

Ramir Gonzales
13th July 2009, 14:42
After looking at the comparison pics and the amount of time of the encoding to get to this quality, I finally understand why DS doesn't want to react to this thread anymore...:rolleyes:

St Devious
13th July 2009, 14:47
After looking at the comparison pics and the amount of time of the encoding to get to this quality, I finally understand why DS doesn't want to react to this thread anymore...:rolleyes:

Why is that ?

nm
13th July 2009, 14:53
After looking at the comparison pics and the amount of time of the encoding to get to this quality, I finally understand why DS doesn't want to react to this thread anymore...:rolleyes:
Well, we have yet to see a speed comparison at the same level of quality, as suggested by Dark Shikari. Since NVIDIA's encoder doesn't have settings to tune the encoding quality, this should be pretty simple: choose a reasonable bitrate and find the fastest x264 settings that match the quality of NVIDIA's encoder at that bitrate. I'd suggest trying the x264 presets first. I'm sure people will help you tune the settings further once you have found the closest preset.

Sharktooth
13th July 2009, 14:59
also 6mbps is an overkill try a set of lower bitrates like 1,500, 2,500, 4,500...

St Devious
13th July 2009, 15:02
also 6mbps is an overkill try a set of lower bitrates like 1,500, 2,500, 4,500...

ok will do. just covering my bases with all the bitrates :D

Ramir Gonzales
13th July 2009, 15:30
Why is that ?

Of the fact that x264 never had the power to convince the crowd as xvid did, and it's now crystal clear that it never will.

You see the believers (mostly) in this same forum here trying every possible method for us to convince us that x264 is better than the rest, "try this", "try that", "no! that setting is too low", "no! that setting is too high", "no, you still cannot configure the other one as you wish", etc...etc...

Yet the facts are everywhere, if you need quality from x264 it will take you way too much time to get it. Competition will get us what we need, just like the fact that xvid STILL is able to get the same quality (or even better) in the same encoding timeframe.

Sharktooth
13th July 2009, 15:36
@Ramir Gonzales: what you smoked? there's NO WAY on earth and on the universe xvid can deliver the same quality as x264 in the same encoding timeframe. THERE IS NO WAY... at least in the current state of xvid.
Also, x264 new presets are just too easy to use AND there are even several GUIs that are way too easy as well.

St Devious
13th July 2009, 15:37
Of the fact that x264 never had the power to convince the crowd as xvid did, and it's now crystal clear that it never will.

You see the believers (mostly) in this same forum here trying every possible method for us to convince us that x264 is better than the rest, "try this", "try that", "no! that setting is too low", "no! that setting is too high", "no, you still cannot configure the other one as you wish", etc...etc...

Yet the facts are everywhere, if you need quality from x264 it will take you way too much time to get it. Competition will get us what we need, just like the fact that xvid STILL is able to get the same quality (or even better) in the same encoding timeframe.

I'm not sure I agree with you. I'm a firm believer in x264's quality, especially at low bitrates when compared to Xvid. I would rather have the encoding take more time if I can display the final output at a better quality than a encoder that takes less time.

Even though I'm a firm believer in x264, I'm still willing to give the competition a try.

x264 might not have convinced you, but it just might if it was ported to GPU encoding. I'm sure with so many talented people working on it, the day is not far when a quality free encoder works on GPU.

Afterall GPU computing is still in its infancy, with interesting developments CUDA, OpenCL, DirectX Compute happening.

Sharktooth
13th July 2009, 15:40
x264 is faster than ANY cuda encoder if you build 2 PCs with the same budget AND when comparing the encoders using the same features or similar quality (otherwise wont be a fair comparison).
so please, go ahead with tests... you will notice that as well.

Dark Eiri
13th July 2009, 17:04
I've tried encoding a 1080p video (live performance in a stadium with lots of high wide angles and sparks everywhere) with the CUDA Encoder, using MediaCoder, at "Xbox 360" settings (3 B-Frames, Level 4.1, everything else at default, I have no idea if it's possible to get more than 1 ref-frame on this thing), 10 Mbps, just as I encode everyday with x264 (and get awesome quality, by the way), and guess what? It doesn't completely suck!

I mean, it's still highly humiliated by x264 (though x264 encoded @ 2.7 fps, probably bottlenecked by DGDecode or Avisynth/Yadif, and CUDA @ 18 fps, I don't care about speed, quality is my thing), but it's highly watchable. It doesn't have as much fine detail, and some areas are a little blocky, but for the average user, I think it would look fine.

CUDA Deinterlacer (Advanced options) does absolutely nothing. I guess it's not a complete feature yet. It would be really nice to see CUDA decoding and deinterlacing in a free encoder too.

It's still really crappy for SD content, though. x264 @ 2 Mbps looks transparent for the most cases, CUDA @ 2 Mbps looks just as bad as Youtube. I guess the smaller the frame size, higher is the complexity (since the bitrate needed for "transparency" doesn't really scale).

roozhou
13th July 2009, 17:05
@St Devious
If you need fast encoding for x264, you should first try to speed up your decoding
1) Use a fast decoder, e.g. libmpeg2 in ffmpeg/mencoder for MPEG2 or CoreAVC CUDA for H264. Do not use slow decoders like DGmpgdec.
2) No pre-processing, no colorspace conversion. If you have to, there must be something wrong with your decoder or source.

The Sony HDW-F900 sample used in your test decodes at ~140fps on my 3.0G E8400 with ffdshow using only a single core. It won't be surprised to see x264 encoding 1080P footage at 100+fps on your Q9450.

stanleyhuang
13th July 2009, 17:05
But if we take the GPU-based video filtering into account, the results may also differ, as GPU is far better at jobs like down-scaling, deinterlacing and other visual enhancements than doing encoding.

stanleyhuang
13th July 2009, 17:09
@St Devious
If you need fast encoding for x264, you should first try to speed up your decoding

This is only true when converting a 1080 movie to a low-res (say 480x272) H.264, in such case, decoding becomes bottleneck.
If you doing 1080 to 1080 transcoding, decoding will never be the bottleneck on a quad-core.

roozhou
13th July 2009, 17:10
But if we take the GPU-based video filtering into account, the results may also differ, as GPU is far better at jobs like down-scaling, deinterlacing and other visual enhancements than doing encoding.

Yes, but GPU doesn't have to do everything. Let GPU do decoding, scaling and deinterlacing while keeping CPU doing encoding.

stanleyhuang
13th July 2009, 17:13
CUDA Deinterlacer (Advanced options) does absolutely nothing. I guess it's not a complete feature yet. It would be really nice to see CUDA decoding and deinterlacing in a free encoder too.

CUDA-based deinterlacer and other video filters will soon be available. We are just working on this part. These will not only limited to CUDA encoder. You will be able to use the filters with any encoders including x264.

roozhou
13th July 2009, 17:18
CUDA-based deinterlacer and other video filters will soon be available. We are just working on this part.

Will it work in DGAVCDecNV's way or accept unprocessed frames from main memory and output processed frames to main memory.

Personally I prefer the second way because we can insert the filter into any place of the filter chain.

stanleyhuang
13th July 2009, 17:33
Will it work in DGAVCDecNV's way or accept unprocessed frames from main memory and output processed frames to main memory.

Personally I prefer the second way because we can insert the filter into any place of the filter chain.

The process is quite straightforward. Uncompressed frames (currently only YV12 or i420) are transferred from main memory to GPU memory, a certain CUDA kernel downloaded to GPU and processes the frames (in dozens, for better parallelism and less overhead), and the processed ones are read back to main memory. As a CUDA kernel can't do really much, this process may repeat if several filters are to be applied.

Dark Eiri
13th July 2009, 17:35
CUDA-based deinterlacer and other video filters will soon be available. We are just working on this part. These will not only limited to CUDA encoder. You will be able to use the filters with any encoders including x264.

Now that's interesting! Will it be nVidia's deinterlacer or a custom made one?

stanleyhuang
13th July 2009, 17:37
Yes, but GPU doesn't have to do everything. Let GPU do decoding, scaling and deinterlacing while keeping CPU doing encoding.

Porting decoder to CUDA is not less sophisticated as porting an encoder, not mentioning there are so many different formats to decode.

roozhou
13th July 2009, 17:40
The process is quite straightforward. Uncompressed frames (currently only YV12 or i420) are transferred from main memory to GPU memory, a certain CUDA kernel downloaded to GPU and processes the frames (in dozens, for better parallelism and less overhead), and the processed ones are read back to main memory. As a CUDA kernel can't do really much, this process may repeat to if several filters are to be applied.

Anyway, it's awesome. AFAIK this will be the first stand-alone CUDA video filter. I guess a lot of MeGUI users will switch to MediaCoder if you finish this:)

stanleyhuang
13th July 2009, 17:40
Now that's interesting! Will it be nVidia's deinterlacer or a custom made one?

We are implementing the deinterlacer and some other filters.

roozhou
13th July 2009, 17:41
Porting decoder to CUDA is not less sophisticated as porting an encoder, not mentioning there are so many different formats to decode.

Not CUDA but VP2.

ChronoCross
13th July 2009, 17:47
So you think only GPLed software can have a donation button and have ads on the software's website right? Lots and lots of freewares as well as free softwares in this world are supported by Google Adsense now.

No but I feel that crooks who violate the GPL and LGPL (which you still are) don't deserve to make money off their illegal activities. Besides you made a gui....while it no doubt took a little bit to code a majority of the useful features are provided by third party programs.....so don't try to pretend your some great coding god who deserves any respect.

stanleyhuang
13th July 2009, 17:56
No but I feel that crooks who violate the GPL and LGPL (which you still are) don't deserve to make money off their illegal activities. Besides you made a gui....while it no doubt took a little bit to code a majority of the useful features are provided by third party programs.....so don't try to pretend your some great coding god who deserves any respect.

VERY NARROW-MINDED.
I am never pretending anything or your god and I state clearly the things behind the GUI.
Besides, MediaCoder is not just a GUI which can be written with a little coding and we are implementing something never existed before.

St Devious
13th July 2009, 17:57
@St Devious
If you need fast encoding for x264, you should first try to speed up your decoding
1) Use a fast decoder, e.g. libmpeg2 in ffmpeg/mencoder for MPEG2 or CoreAVC CUDA for H264. Do not use slow decoders like DGmpgdec.
2) No pre-processing, no colorspace conversion. If you have to, there must be something wrong with your decoder or source.


How do I implement libmpeg2 in the avisynth script I had a few posts back ?

This is only true when converting a 1080 movie to a low-res (say 480x272) H.264, in such case, decoding becomes bottleneck.
If you doing 1080 to 1080 transcoding, decoding will never be the bottleneck on a quad-core.

So would the decoding slow down the encoding in the Avisynth script I had ?

Will it work in DGAVCDecNV's way or accept unprocessed frames from main memory and output processed frames to main memory.

Personally I prefer the second way because we can insert the filter into any place of the filter chain.

What is this process we are talking about ?

Anyway, it's awesome. AFAIK this will be the first stand-alone CUDA video filter. I guess a lot of MeGUI users will switch to MediaCoder if you finish this:)

Continued from the above question, Why is this ?

And, How resource intensive is deinterlacing ?

stanleyhuang
13th July 2009, 17:58
Not CUDA but VP2.

At the moment, nvidia still does not allow programming the VP2 but only allows APIs for read-back the decoded frames.

stanleyhuang
13th July 2009, 18:01
So would the decoding slow down the encoding in the Avisynth script I had ?
I don't think it will noticeably slow down encoding on your powerful processor.

And, How resource intensive is deinterlacing ?
I roughly estimate that it's comparable to decoding.

roozhou
13th July 2009, 18:17
How do I implement libmpeg2 in the avisynth script I had a few posts back ?

Use ffdshow + DSS or just use mediacoder and add "-vc mpeg12," to mencoder's commandline. I highly doubt avisynth itself adds additional overhead to decoding. Try playing something in MPC-HC and playing a simple DirectShowSource("xxx") avs script in the same player. You will notice significant higher CPU load with avs script.

So would the decoding slow down the encoding in the Avisynth script I had ?
Yes if your x264 is encoding as fast as CUDA encoder.

What is this process we are talking about ?
DGAVDDecNV's deinterlacing is a part of its decoder. That means you cannot use it separately.

Continued from the above question, Why is this ?
And, How resource intensive is deinterlacing ?
Try nnedi2/eeedi/TDeint in avisynth or mcdeint in mplayer. You will see how slow a good software deinterlacer is.
Currently the best deinterlacer in MediaCoder is yadif. There is an option for mcdeint but MC is not using it correctly.

St Devious
13th July 2009, 18:21
Use ffdshow + DSS or just use mediacoder and add "-vc mpeg12," to mencoder's commandline. I highly doubt avisynth itself adds additional overhead to decoding. Try playing something in MPC-HC and playing a simple DirectShowSource("xxx") avs script in the same player. You will notice significant higher CPU use with avs script.

That 1080p source doesn't play for me in MPC-HC, but plays in WMP.

What does "-vc mpeg12" do ?

Still not sure how to use ffdshow + DSS (what's DSS ?) while encoding ?

LoRd_MuldeR
13th July 2009, 18:25
What does "-vc mpeg12" do ?

Force the decoder to MPEG-1/MPEG-2.

Still not sure how to use ffdshow + DSS (what's DSS ?) while encoding ?

DirectShowSource()

BTW: DirectShowSource() will use ffdshow if it's the only suitable DirectShow decoder available or if it has the highest Merit of all suitable decoders.

roozhou
13th July 2009, 18:27
What does "-vc mpeg12" do ?

To set libmpeg2 as the preferred decoder since libavcodec's mpeg2 decoder, which is slower, has higher priority in mplayer/mencoder.

Still not sure how to use ffdshow + DSS (what's DSS ?) while encoding ?
DSS = DirectShowSource

St Devious
13th July 2009, 18:38
Force the decoder to MPEG-1/MPEG-2.



DirectShowSource()

BTW: DirectShowSource() will use ffdshow if it's the only suitable DirectShow decoder available or if it has the highest Merit of all suitable decoders.

To set libmpeg2 as the preferred decoder since libavcodec's mpeg2 decoder, which is slower, has higher priority in mplayer/mencoder.

DSS = DirectShowSource

I see, thank you.

So in my avisynth script
DGDecode_mpeg2source("D:\Videos\1080p25.d2v", info=3)
ColorMatrix(hints=true, threads=0)
#deinterlace
#crop
#resize
#denoise

I replace DGDecode_mpeg2source with DirectShowSource() and the encoder will use ffdshow as If i have set ffdshow to decode MPEG2 ?

Is libmpeg2 faster than libavcodec in there ? And are these two options faster than DGDecode_mpeg2source ?

roozhou
13th July 2009, 18:48
Is libmpeg2 faster than libavcodec in there ?
Yes, I tested on C2D and K8. Libavcodec supports multithreading but it consumes more CPU time than libmpeg2, leaving less available CPU time for x264.
And are these two options faster than DGDecode_mpeg2source ?
Yes, I am quite sure of that, but don't expect too much.

LoRd_MuldeR
13th July 2009, 18:49
I replace DGDecode_mpeg2source with DirectShowSource() and the encoder will use ffdshow as If i have set ffdshow to decode MPEG2 ?

Nope. DirectShowSource must point to the MEPG file (PS or TS), not to the .d2v project file.

Also in addition to an MPEG-2 decoder (e.g. ffdshow), a suitable DirectShow Splitter will be needed for MPEG PS/TS files. I think Haali Media Splitter can do that.

roozhou
13th July 2009, 18:51
Nope. DirectShowSource must point to the MEPG file (PS or TS), not to the .d2v project file.

Also in addition to an MPEG-2 decoder (e.g. ffdshow), a suitable DirectShow Splitter will be needed for MPEG PS/TS files. I think Haali Media Splitter can do that.

MPC's mpeg splitter is just fine, and faster than Haali.

PS. I usually remux PS/TS test samples into mkv.

LoRd_MuldeR
13th July 2009, 18:53
MPC's mpeg splitter is just fine, and faster than Haali.

But MPC's "internal" filters aren't available to DirectShowSource, unless the "external" (standalone) versions are installed/registered.
So he either needs to install Haali's Splitter, MPC's Splitter or a similar one...

(BTW: Another option would be using FFmpegSource instead of DirectShowSource, which saves you from any DirectShow hassle)

roozhou
13th July 2009, 18:57
But MPC's "internal" filters aren't available to DirectShowSource, unless the "external" versions are installed/registered.

There are stand-alone filters (http://www.xvidvideo.ru/component/option,com_docman/task,cat_view/Itemid,11/gid,19/orderby,dmdate_published/) on xvidvideo.ru.

St Devious
13th July 2009, 18:59
Nope. DirectShowSource must point to the MEPG file (PS or TS), not to the .d2v project file.

Also in addition to an MPEG-2 decoder (e.g. ffdshow), a suitable DirectShow Splitter will be needed for MPEG PS/TS files. I think Haali Media Splitter can do that.

thanks, so no need for DGIndex anymore ?

I have Haali set to decode MPEG-TS. IF I don't set it, then what splitter does the MPEG-TS ?

MPC's mpeg splitter is just fine, and faster than Haali.

PS. I usually remux PS/TS test samples into mkv.

That's interesting. Do you use AVIDemux to remux ?

But MPC's "internal" filters aren't available to DirectShowSource, unless the "external" versions are installed/registered.

(BTW: Another option would be using FFmpegSource instead of DirectShowSource, which saves you from any DirectShow hassle)

What's DirectShow hassle ?

What do i need for FFmpegSource ? Does this point to .d2v or the .ts file ?

LoRd_MuldeR
13th July 2009, 19:03
thanks, so no need for DGIndex anymore ?

Nope. DGIndex and DGDecode/MPEG2Source() belong together. DirectShowSource() is unrelated to DGIndex.

I have Haali set to decode MPEG-TS. IF I don't set it, then what splitter does the MPEG-TS ?

DirectShow will use whatever DirectShow Splitter for MPEG-TS is installed on your system (if there's any).

If there's more than one suitable Splitters installed, the one with the highest Merit will be picked.

If no suitable DirectShow Splitter is installed on your system, then DirectShow will simply fail to build the graph.

What's DirectShow hassle ?

With DirectShowSource() you need all the required DirectShow filters installed on your system. At least a suitable DirectShow splitter plus a suitable decoder.

FFmpegSource() doesn't need that. It doesn't use any "external" decoders or splitters. Instead it runs "out of the box". It's still "beta" though...

What do i need for FFmpegSource ? Does this point to .d2v or the .ts file ?

TS file. FFmpegSource() is unrelated to DGIndex.

roozhou
13th July 2009, 19:22
FFmpegSource() doesn't need that. It doesn't use any "external" decoders or splitters, it runs "out of the box". It's still "beta" though...


As I mentioned in previous post, libavcodec's mpeg2 decoder used in FFMS is slower than libmpeg2 used in ffdshow.

St Devious
13th July 2009, 19:32
If there's more than one suitable Splitters installed, the one with the highest Merit will be picked.

Is there any way to check the merit of filters installed on the system ?

If no suitable DirectShow Splitter is installed on your system, then DirectShow will simply fail to build the graph.

What does building the graph mean ?

FFmpegSource() doesn't need that. It doesn't use any "external" decoders or splitters. Instead it runs "out of the box". It's still "beta" though...

I meant, do you need plugin or something for FFmpegSource() to work or some .exe file in avisynth folder ?

LoRd_MuldeR
13th July 2009, 19:37
Is there any way to check the merit of filters installed on the system ?

http://www.softella.com/dsfm/index.en.htm

What does building the graph mean ?

DirectShow connects the source (e.g. AVI, PS or TS file) with the sink (e.g. Video Renderer, Audio Device or Avisynth/DSS) by putting the required Filters in-between.

This is called a DirectShow Graph:

http://img222.imageshack.us/img222/5712/dsgraph.th.png (http://img222.imageshack.us/img222/5712/dsgraph.png)

If DirectShow cannot construct a chain of Filters that connects the Source to the Sink, then building the graph failed.

I meant, do you need plugin or something for FFmpegSource() to work or some .exe file in avisynth folder ?

Nothing is needed, except for the "FFMS2.dll" file, which should be located in your Aviynth/Plugins folder.

That's it. All the required splitters and decoders are "built-in", thanks to libavcodec :)

Myrsloik
13th July 2009, 23:07
That's it. All the required splitters and decoders are "built-in", thanks to libavcodec :)

Not completely true. Haali's splitter will be used (or more exactly the parser part, directly through COM) for mpeg ps/ts and ogg/ogm contents. So make sure to have the latest version installed.

CruNcher
14th July 2009, 00:08
@stanleyhuang
are there any differences between a Cuda 2.2 and 2.3 driver in nvcuvenc.dll ? does your cli application interfaces with both correctly ?

St Devious
14th July 2009, 01:55
Slight problem while playing the .ts file with DirectShowSource() in Avisynth

http://i30.tinypic.com/11rexy9.jpg

Video plays after I press close

LoRd_MuldeR
14th July 2009, 20:10
Looks like a suitable Audio decoder is missing! Try DirectShowSource("C:\My File.ts", audio=false) ;)

St Devious
14th July 2009, 21:12
Looks like a suitable Audio decoder is missing! Try DirectShowSource("C:\My File.ts", audio=false) ;)

that plays, thank you.

what decoder could be missing on the file ?

LoRd_MuldeR
14th July 2009, 21:57
that plays, thank you.

What audio format does your source file contain? If you don't know, ask MediaInfo.

St Devious
14th July 2009, 22:51
What audio format does your source file contain? If you don't know, ask MediaInfo.

it's AC3. Ffdshow is set to decode AC3 with liba52

roozhou
15th July 2009, 03:42
it's AC3. Ffdshow is set to decode AC3 with liba52

I guess this is the problem of the ts file itself. None of my player can play that ts with sound.

St Devious
15th July 2009, 03:46
I guess this is the problem of the ts file itself. None of my player can play that ts with sound.

do you get any errors ?

i don't think that clip is supposed to have sound, just a second odd of audio somewhere.

edison
15th July 2009, 10:29
my test:
http://www.pcinlife.com/article/graphics/2009-07-15/1247632564d831.html

is there any idea to tweak the image quality for CUDA encoding on MediaCoder ? :)

dj_tjerk
15th July 2009, 11:50
Nope, I don't think there's any way to make it look better. I guess a 2-pass encode would look better, but there's no option for that yet afaik.

CruNcher
15th July 2009, 14:54
x264
12:480 sec = --preset ultrafast
13:915 sec = --preset veryfast --profile baseline --scenecut -1 --nf --partitions i8x8,p8x8 --aq-mode 0
14:204 sec = --preset veryfast --profile baseline --scenecut -1 --nf --partitions i8x8,p8x8


cuda
15:000 sec = ultrafast

3 mbits target 1280x720 25fps

so @ low bitrate targets cuda (in form of Nvidias Encoder) is pretty much useless compared to a old CPU architecture in this case AMD Athlon(tm) 64 X2 Dual Core Processor 4200+ (1603Mhz)

the visual quality in this case is also on the side of x264 compared vs a G92 8800GT 112 Stream Processors

X264 X2 vs G92 112 Stream Processors = X264 wins @ Baseline 3 mbit target (speed/quality)

stanleyhuang
15th July 2009, 15:54
is there any idea to tweak the image quality for CUDA encoding on MediaCoder ? :)

An update (http://www.mediacoderhq.com/dlupdate.htm) is available which adds a new option named "slice count". Increasing this option will bring possibly better quality.

CruNcher
15th July 2009, 16:00
how should the slice count that is by default currently @ 4 (2 cores) change anything visualy ?
changed slicecount to 1 though i can see no visual improvement only a slight bitrate behaviour change (hits target more accurate now) :)

@stanleyhuang
cudaH264Enc.exe output is missing all the processing data utilization (cpu/kernel), time, fps arent shown anymore :(


Video stream: 276480.000 kbit/s (34560000 B/s) size: 675993600 bytes 19.560 s
ecs 489 frames

NvEncodeTestAPI returned with return value = 0


PS: I would say from the current results i gathered so far a G92 with 112 Stream Processors (Shaderspeed 1.7 GHz) is currently as powerfull (if compared to x264 optimizations) as a Dualcore K8 @ 2.7 GHz (Toledo Core) in terms of Video Encodindg, though still have todo cabac, high bitrate, fullhd resolution tests


Cabac 720p (Main Profile) 3mbit:

x264 = --preset veryfast --profile main -b 0 --scenecut -1 --nf --partitions i8x8,i4x4,p8x8
cuda = -iw 1280 -ih 720 -fpsnum 25 -fpsden 1 -profile 1 -preset -1 -idrp 250 -qp 25 -qpp 28 -qpb 31 -slicecount 1 -pinterval 1 -darw 16 -darh 9 -rc 1

x264 = 15.9 sec
cuda = 17.5 sec

http://mirror05.x264.nl/CruNcher/force.php?file=./GPU-Encoding/x264-veryfast-aq-abr-6000-cabac-main.mp4
http://mirror05.x264.nl/CruNcher/force.php?file=./GPU-Encoding/cuda-ultrafast-6000-cabac-main.mp4

Cabac 1080p (High Profile) 6 mbit:

x264 = --preset veryfast --profile high -b 0 --scenecut -1 --nf --partitions i8x8,i4x4,p8x8
cuda = -iw 1920 -ih 1080 -fpsnum 25 -fpsden 1 -profile 2 -preset -1 -idrp 250 -qp 25 -qpp 28 -qpb 31 -slicecount 1 -pinterval 1 -darw 16 -darh 9 -rc 1

x264 = 247 sec
cuda = 253 sec

http://mirror05.x264.nl/CruNcher/force.php?file=./GPU-Encoding/x264-veryfast-aq-abr-1080p-6000-cabac-high.mp4
http://mirror05.x264.nl/CruNcher/force.php?file=./GPU-Encoding/cuda-ultrafast-1080p-6000-cabac-high.mp4

Result = A Geforce 8800 GT with 112 Stream Processors clocked @ 1.7 GHz has roughly the same Encoding Power as x264 on a Dualcore K8 @ 2.7 GHz, though still x264 can't be beat in Visual Quality results

St Devious
16th July 2009, 00:41
An update (http://www.mediacoderhq.com/dlupdate.htm) is available which adds a new option named "slice count". Increasing this option will bring possibly better quality.

i see that b-frames max are 16 now. How does increasing them from 8 to 16 affect quality ?

Is there a wiki somewhere explaining the meaning of all the options in CUDA encoder and how their values affect quality ?

LoRd_MuldeR
16th July 2009, 00:53
i see that b-frames max are 16 now. How does increasing them from 8 to 16 affect quality ?

We know from x264 that more than ~3-4 consecutive b-frames are only helpful in very rare cases. I doubt this is fundamentally different for the CUDA encoder.

Unless the CUDA encoder adaptively decides the optimal number of b-frames (like x264 does) and the option only limits the maximum, raising b-frames too much may even hurt quality.

St Devious
16th July 2009, 02:51
x264 = --preset veryfast --profile high -b 0 --scenecut -1 --nf --partitions i8x8,i4x4,p8x8
cuda = -iw 1920 -ih 1080 -fpsnum 25 -fpsden 1 -profile 2 -preset -1 -idrp 250 -qp 25 -qpp 28 -qpb 31 -slicecount 1 -pinterval 1 -darw 16 -darh 9 -rc 1


I'll try those CMD on my source and report back.

and why are you setting pinterval (bframes) to 0 ? Were you testing fastest for CUDA ?

Based on my testing, setting the bframes from 3 to 0 increases the time taken by about 5s, while setting it to 16 increases it by only 1s or even in some case none.

btw, how are you encoding CUDA ? I don't see those options in my commandline in mediacoder. i see something like

# ".\codecs\cudaH264Enc.exe" -i "$(SourceFile)" -o "$(DestFile)" -profile 2 -preset -1 -level 51 -idrp 250 -qp 25 -qpp 28 -qpb 31 -gop -deblock -forceintra -forceidr -pinterval 4 -darw 16 -darh 9 -rc 1 -abit $(VideoBitrate)

We know from x264 that more than ~3-4 consecutive b-frames are only helpful in very rare cases. I doubt this is fundamentally different for the CUDA encoder.

Unless the CUDA encoder adaptively decides the optimal number of b-frames (like x264 does) and the option only limits the maximum, raising b-frames too much may even hurt quality.

Thanks, just tested it for myself, 3 vs 16. no difference in quality, but 1s difference in encoding time.

CruNcher
16th July 2009, 10:51
i was testing lowest complexity possible, theres only 1 thing that is strange in the test with the Test Sequence combination @ BlueSky (1st Sequence) X264 is heavily reacting creating a massive peak over time that slowly vanishes not sure but might be caused by the AQ in this case.

http://s12.directupload.net/images/090716/3oadcya4.png

a lot of black frames and then RC seems to explode for several frames :(

it isn't such a big problem Nvidias Hardware Decoder has no problem Decoding this spike but in Software like ffh264 you can see that it slowsdown a little when this heavy bitrate curve hits the decoder.

Also it seems that Nvidia is using some automaticly calculated VBV setting (to avoid such peaks) seeing that it declares a bitrate in the stream.

dj_tjerk
16th July 2009, 11:24
Would you mind testing for funzies how the bitrate distribution looks with crf (play with the value of crf till you end up near the desired bitrate)?

CruNcher
16th July 2009, 12:19
As CRF is almost a 2pass the result isn't really surprising :)

http://mirror05.x264.nl/CruNcher/force.php?file=./GPU-Encoding/x264-veryfast-aq-crf-1080p-6000-cabac-high.mp4

http://s12.directupload.net/images/090716/wklv4i7o.png

especialy riverbed which imho is also the most complex scene motion estimation wise of all gets visualy out much better then with ABR or Nvidias RC :)

Nvidia Encoder (Cuda)
http://s12b.directupload.net/images/090716/temp/e7udo6dg.png (http://s12b.directupload.net/images/090716/e7udo6dg.png)
x264 abr
http://s12.directupload.net/images/090716/temp/rz4y2yw3.png (http://s12.directupload.net/images/090716/rz4y2yw3.png)
x264 crf
http://s12.directupload.net/images/090716/temp/qo7k2yzv.png (http://s12.directupload.net/images/090716/qo7k2yzv.png)

@stanleyhuang
there is also a bug in cudah264enc or better MediaCoders logic when setting -rc 0 (quality based) the user needs to provide a peak bitrate or the encoder errors out with


NVVE_AVG_BITRATE Failed

Set Encoder Params failed

stanleyhuang
16th July 2009, 13:49
there is also a bug in cudah264enc or better MediaCoders logic when setting -rc 0 (quality based) the user needs to provide a peak bitrate or the encoder errors out with

Thanks, will fix this.

CruNcher
16th July 2009, 15:31
Hmm i get allways the same bitrate result with quality based encoding no matter what the -rc 0 xx or -pbit settings are ?

PS: A new Cuda 2.3 Driver has been released by Nvidia 190.38, lets see if they changed something on the Video Encoder :)

DiKey
6th October 2009, 01:52
The last post is at 16.07.2009. There was no news?

seemees
21st October 2009, 05:53
stanleyhuang
When you make a stable version of MediaCoder with Cuda 2.3 support on h264 and may be on Dirac? I can't make an AVI file with Dirac or with h264. Only mkv with mkvmerge.exe (h264). What we need to produce play-able AVI with Dirac and with h264 - what codec or CLI program, please help.
with best regards, seemees

Audionut
21st October 2009, 12:09
Why on earth would you want to put h.264 in avi!! The only way to do so is with a hack.

Use a container that natively supports h.264.

seemees
21st October 2009, 12:49
Audionut
Why not? :-)
AVI container is in options - isn't it? And with virtual dub and ffdshow or h264vfw - AVI maked easy. But in Media codec - NOT.
AVI support for h264 ready:
http://en.wikipedia.org/wiki/Comparison_of_container_formats#cite_note-H.264_in_AVI-3
Through an updated x264/ffdshow filter it is possible to view H.264 in an AVI file.
With best regards to you, seemees

roozhou
21st October 2009, 12:56
Why on earth would you want to put h.264 in avi!! The only way to do so is with a hack.

Use a container that natively supports h.264.

What hack? And why do you call it a hack? I wrote a program putting H264 into avi in the same way of other codecs. It plays fine. I have even seen DV store 1920x1080 AVC + AAC in avi on a SD Card.

LoRd_MuldeR
21st October 2009, 13:25
Please, no AVI flamewar ;)

Once again: You must distinguish between limitations in AVI and limitations in VFW (Video for Windows).

While most AVI's are created using VFW Codecs (e.g. VirtualDub), there are other, better ways to create AVI's (MEncoder, Avidemux, etc).

As far as I know the main problem with AVI itself is that it doesn't know about B-Frames. Frames in AVI are either makred as Key-Frames (which equals IDR in H.264) or as Delta-Frames (which equals P-Frames in H.264). Also AVI doesn't support "progressive download", as the index block is stored at the end of the file. The problem with VFW (and not with AVI itself!) is that it has a strict "one frame in, one frame out" behaviour. So the encoder cannot access "future" frames without returning the current frame first, which makes it impossible to code B-Frames.

As far as I know, the workaround (or "hack") that DivX came up with to do B-Frames with VFW is: Return dummy Null-Frames and then later pack the B- and P-Frame into one frame.

That is commonly referred to as the "Packed Bitstream" hack. But it's specific to VFW. It's not required for AVI in general...

roozhou
21st October 2009, 13:38
As far as I know, the workaround (or "hack") that DivX came up with to do B-Frames with VFW is: Return dummy Null-Frames and then later pack the B- and P-Frame into one frame.

That is commonly referred to as the "Packed Bitstream" hack. But it's specific to VFW. It's not required for AVI in general...

AVC in avi does NOT use "Packed Bitstream". The H264 frames are sequentially stored in avi in its decoding order, which is the same way MKV and MP4 do. Windows' built-in avi splitter and all H264 decoders can handle it and generate correct frames and PTS.

LoRd_MuldeR
21st October 2009, 13:46
AVC in avi does NOT use "Packed Bitstream". The H264 frames are sequentially stored in avi in its decoding order, which is the same way MKV and MP4 do. Windows' built-in avi splitter and all H264 decoders can handle it and generate correct frames and PTS.

I didn't imply that "packed bitstream" is used with H.264 in AVI. I just said that "hack" exists to workaround a limitation in VFW related to creating B-Frames. However if one would like to create a H.264 AVI file through the VFW interface and use B-Frames, a similar "workaround" would be needed, I think. Tools that create AVI files not through VFW (e.g. Avidemux) don't need it for sure...

Audionut
21st October 2009, 14:12
Well, fwiw, using a 17 year old container for a 6 year old standard is ....

<pengvado> <Dark_Shikari> if a fansubber or DVD rip group uploaded an H.264-in-AVI file they'd get laughed off the internet
<Dark_Shikari> Probably
<pengvado> but asp-in-avi wouldn't be laughed at
<Dark_Shikari> of course not, ASP in AVI is normal unless you want softsubs
<pengvado> and asp-in-avi requires exactly the same ugly hacks
<Dark_Shikari> Probably because its been that way so long that everyone is accustomed to it
<pengvado> which just goes to show that our campaign to use a new codec as an excuse to tell people to upgrade their container is working

LoRd_MuldeR
21st October 2009, 14:26
Well, fwiw, using a 17 year old container for a 6 year old standard is ....

...completely unrelated to whether that container is suitable or not. If at all, we need technical arguments against using AVI, instead of polemic. And I tried to sum them up here (http://forum.doom9.org/showthread.php?p=1336612#post1336612).

BTW: Isn't this getting off-topic? :p

seemees
21st October 2009, 15:18
For respectable All
I have a dozen AVI with h264, I have a some (a little) mkv with it. Here the big differencies (on PC of course):
1.AVI seek very fast.
2.MKV seek very SLOW. more then 1-2 seconds some time.
3.AVI errors easy to recover (AVI reconst and setra).
4.MKV reconst - bad idea. No one really worked product for do this (I know i bad searcher in inet).
5.AVI smoller size and fast downloads from inet (cause a smaller bitrate used < 2000 kbps).
6.MKV huge 4-30G size and hard to downloads from inet (bitrate > 4000 kbps).
7.AND ANY of MKV 100% has a 1-50 frame error with bad artifacts and so on (lost lines, squares, color dismatches...) It is an entropy of BIG size...
Thus The AVI need us. The developers MUST implement him in his products. Isn't it?

One thing really RIGHT worked with MKV is Haali media splitter. And other applications have many errors and mismatchs for this format...

MP4 is same as MKV. New and no handful, AVI is really-really old but GOOD friend of us.

With best regards, seemees

Guest
21st October 2009, 15:20
Lord_Mulder is right, this is OT. Please take this to an appropriate place.

CpT
21st October 2009, 17:52
@CruNcher I get the same thing with -pbit + the latest mediacoder version. Doesn't seem to do anything.

Something I've been toying around with:
qp level I 20
qp level P 22
qp level B 22

The main issue I'm having with it is it doesn't caculate aspect correctly for anamorphic. And the resizer seems borked.

nakTT
14th December 2009, 16:15
1) Any new development on this?
2) What is the quality (at the same bitrate) comparison now?
3) Is CUDA getting closer to x264 than ever before or x264 is pulling away?

Please share info.

Firebird
14th December 2009, 16:52
Is CUDA getting closer to x264
No. It will never be as good as x264 is.

LoRd_MuldeR
14th December 2009, 17:01
Is CUDA getting closer to x264 No. It will never be as good as x264 is.

That statement doesn't make sense. CUDA is a platform technology while x264 is one specific software. So you are comparing apples and oranges ;)

So the real question is: Will GPGPU-based (CUDA, Stream, OpenCL, etc.) H.264 encoders eventually beat CPU-only encoders performance-wise and quality-wise?

Well, currently it doesn't look like this will happen soon. But this may change with upcoming GPU generation...

Cyber-Mav
14th December 2009, 17:04
cuda is getting closer now in quality.

nakTT
14th December 2009, 17:12
That statement doesn't make sense. CUDA is a platform technology while x264 is one specific software. So you are comparing apples and oranges ;)

So the real question is: Will GPGPU-based (CUDA, Stream, OpenCL, etc.) H.264 encoders eventually beat CPU-only encoders performance-wise and quality-wise?

Well, currently it doesn't look like this will happen soon. But this may change with upcoming GPU generation...
Thanks Loard, that's what I meant.

Quality wise, what the upcoming GPU generation got to do with it? It is just the hardware, I can understand if it speed wise. Please shed some light on this.

:thanks:

roozhou
14th December 2009, 17:29
Thanks Loard, that's what I meant.

Quality wise, what the upcoming GPU generation got to do with it? It is just the hardware, I can understand if it speed wise. Please shed some light on this.

:thanks:

There is no video encoding chip on any GPU so the encoding quality has nothing to do with GPU generation or CPU models. You get same quality between a P3 and a i7 with x264 if you use the same settings.

Quality is only determined by the algorithm that the encoder uses. This applies to both CUDA and x264.

nakTT
14th December 2009, 17:42
There is no video encoding chip on any GPU so the encoding quality has nothing to do with GPU generation or CPU models. You get same quality between a P3 and a i7 with x264 if you use the same settings.

Quality is only determined by the algorithm that the encoder uses. This applies to both CUDA and x264.
That is my understanding that I'm trying to share on my previous post. Perhaps we are wrong/right?

nakTT
14th December 2009, 18:16
Thanks for your thought Stephen. It doesn't occur to me until you point it out. Thanks again.

LoRd_MuldeR
14th December 2009, 19:24
There is no video encoding chip on any GPU...

Not yet. But there is dedicated decoder hardware on any modern graphics card already. Also there are encoding solutions available that ship with a dedicated encoder hardware/stick.

Therefore it's not completely absurd to think about adding dedicated encoder chips to future GPU generations...

...so the encoding quality has nothing to do with GPU generation or CPU models.

Well, the capabilities of the first GPGPU-enabled GPU generation were pretty limited. Since then the GPU manufactures have added new GPGPU-specific capabilities with each generation.

Certain encoding algorithms, that can not be implemented (efficiently) on the current GPU generation, may be implementable on future GPU generations.

So we may see improved GPGPU encoders on future GPU generations indeed! Especially since the development of GPU's currently is rapid, while the development of CPU's is slowing down.

Of course there is no automatism. We'll have to wait and see whether future GPU generations will be more suitable for video encoder than the current GPU's.

nakTT
15th December 2009, 03:45
Well, the capabilities of the first GPGPU-enabled GPU generation were pretty limited. Since then the GPU manufactures have added new GPGPU-specific capabilities with each generation.

Certain encoding algorithms, that can not be implemented (efficiently) on the current GPU generation, may be implementable on future GPU generations.

So we may see improved GPGPU encoders on future GPU generations indeed! Especially since the development of GPU's currently is rapid, while the development of CPU's is slowing down.

Of course there is no automatism. We'll have to wait and see whether future GPU generations will be more suitable for video encoder than the current GPU's.
So is my understanding correct to say that CPU encoding give programmer a huge of room to maneuver as appose to GPGPU?

Thanks for a very newbie friendly explanation. I have always like the way you treat newbies. Keep it up.


:thanks:

kidjan
16th December 2009, 04:37
ok, doing that.

In the meantime here's some more images

CUDA 6 Mbps
http://www1.picturepush.com/photo/a/1968749/220/1968749.png (http://www.picturepush.com/public/1968749)

x264 6Mbps
http://www2.picturepush.com/photo/a/1968750/220/1968750.png (http://www.picturepush.com/public/1968750)

Source
http://www3.picturepush.com/photo/a/1968751/220/1968751.png (http://www.picturepush.com/public/1968751)

IMO, quality comparisons like this would be a lot more useful with SSIM (and possibly PSNR) measurements. My $.02, possibly wrong. It's a lot easier to encode the same video to equal bitrates and then see how it fares with an objective measurement than posting screenshots.

Puncakes
16th December 2009, 09:19
IMO, quality comparisons like this would be a lot more useful with SSIM (and possibly PSNR) measurements. My $.02, possibly wrong. It's a lot easier to encode the same video to equal bitrates and then see how it fares with an objective measurement than posting screenshots.

I don't know about you, but considering the fact that those measurements are useless for comparing actual visual quality, I think I'd rather have screenshots.

LoRd_MuldeR
16th December 2009, 10:52
So is my understanding correct to say that CPU encoding give programmer a huge of room to maneuver as appose to GPGPU?

You must think of the GPU as a massively parallel processor. So GPGPU (CUDA, Stream, OpenCL, etc) gives the programmer access to a massively parallel co-processor.

And we are not talking about four or eights threads here. We are talking about hundreds or even better thousands of threads that need to run on the GPU!

So if you want to leverage the theoretical processing power of a GPU, your problem must be highly parallelizeable and new algorithms are needed that scale to hundreds/thousands of threads.

Therefore not any problem is suitable for the GPU. There are inherently sequential problems that don't fit on the GPU at all!


The GPU cores are many, but they are very limited. Especially memory access to the (global) GPU memory is extremely slow, because it's not cached at all (except for texture memory).

Thus we must try to "hide" slow memory access with calculations, which means that we need much more GPU threads than we have GPU cores.

Well, each group/block of GPU cores has its own local "shared" memory that is fast, but the size of that per-block shared memory is small. Way too small for many things!

Also we can't sync the shared memories of different blocks, so whenever threads from different blocks need to "communicate", this needs to be done through the slow "global" memory.

Even organizing/synchronizing the threads within a block is a though task, because "bad" memory access patterns can slow down your GPU program significantly!


Last but not least the GPU cannot access the main/host memory at all. Hence the CPU program needs to upload all input data to the graphic's device first and later download all the results.

That "host <-> device" data transfer is a serious bottleneck and means that you cannot run "small" functions on the GPU, even if they are a lot faster there.

What worth is it to complete a calculation in 1 ms instead of 10 ms, but it takes 20 ms to upload/download the data to/from the graphic's device? Yes, it's completely useless!

So if we move parts of our program to the GPU, this must be significant parts with enough "work" to justify the communication delay. It's not trivial to find such parts in your software.

Remember: Those parts must also be highly parallelizable and efficient parallel algorithms for the individual problem must exists (or must be developed).


See also:
http://developer.download.nvidia.com/compute/cuda/2_0/docs/NVIDIA_CUDA_Programming_Guide_2.0.pdf

Dark Shikari
16th December 2009, 10:58
And we are not talking about four or eights threads here. We are talking about hundreds or even better thousands of threads that run on the GPU.More like 20,000.

nakTT
16th December 2009, 11:30
Thanks again LoRd_MuldeR, for your informative posting. I really enjoy reading the info.


:thanks:

Limit
16th December 2009, 14:53
I wonder if the next generation CPU/GPU combo chips like Llano/Fusion would make any difference. The latencies should be significant lower although they are still conected over PCIe. What do you think, would such a APU be useful for x264?

nakTT
16th December 2009, 15:21
I wonder if the next generation CPU/GPU combo chips like Llano/Fusion would make any difference. The latencies should be significant lower although they are still conected over PCIe. What do you think, would such a APU be useful for x264?
IMHO integrated GPU (be it just on the same packaging or on the same silicone) will be nowhere near the power of a high-end discreet GPU. Please note that those kind of CPUwithGPU are targeted towards notebooks and other lightly demanding graphic usage like office PC and others.

LoRd_MuldeR
16th December 2009, 15:30
I wonder if the next generation CPU/GPU combo chips like Llano/Fusion would make any difference. The latencies should be significant lower although they are still conected over PCIe. What do you think, would such a APU be useful for x264?

Well, it may make the bottleneck less critical, but certainly doesn't remove it, as the basic architecture still is the same. Unless they use more PCIe lanes for the internal interconnect than they used for the "external" PCIe bus, there won't be much difference. And even if there is a difference, the way we call GPU kernels/programs is still the same: Upload input data from the host to the device, invoke the GPU kernel, wait for completion (while maybe doing other things on the CPU) and finally download the results from the device back to the host. Also I doubt that the combined CPU/GPU chip packages will contain very powerful GPU's. It will be more like what he have as "on board" graphics chips now. Not anywhere near high-end GPU's.

However with NVidia's new "Fermi" GPU generation there will be significant improvements for the per-block "shared" GPU memory: It's now much larger and it can (optionally) be used to cache accesses to the global GPU memory. This may (or may not) significantly help for specific problems. Also this is one example for what I said before: Future GPU generations may be more suitable for implementing video compression algorithms than the current generation. In the case of Fermi I cannot tell you whether it helps video encoding or not. The Codec gurus need to decide ^^

Limit
16th December 2009, 16:14
Well, it may make the bottleneck less critical, but certainly doesn't remove it, as the basic architecture still is the same. Unless they use more PCIe lanes for the internal interconnect than they used for the "external" PCIe bus, there won't be much difference.

PCIe is a point-to-point link so it should be possible to run the GPU's link with a much higher clock rate. Standard PCIe runs at 100MHz. With CPU, GPU and PCIe Controller on the same die a much higher frequency for the GPU PCIe link should be a small problem. For example if you get it running with 1GHz, you increase the bandwidth and decrease the latency by a factor of 10.

And even if there is a difference, the way we call GPU kernels/programs is still the same: Upload input data from the host to the device, call the kernel, wait for completion (while maybe doing other things on the CPU) and finally download the results from the device back to the host.

Afaik the integrated GPUs has no own memory besides the small caches. So there is no need to copy data from host memory to device memory because it is the same memory.

Also I doubt that the combined CPU/GPU chip packages will contain very powerful GPU's. It will be more like what he have as "on board" graphics chips now. Not anywhere near high-end GPU's.

That is clear. The last rumours I heard speak of 240 shader units for AMD/ATIs first generation Fusion APU. That is far from high-end but its computing power is still higher then any avaible CPU's.

LoRd_MuldeR
16th December 2009, 16:28
Afaik the integrated GPUs has no own memory besides the small caches. So there is no need to copy data from host memory to device memory because it is the same memory.

Well, it then "shares" the RAM modules with the CPU - not to be confused with the on-chip shared memory. But this doesn't mean that the GPU can directly access the same memory locations that the CPU uses. We don't know it yet, but I would assume they simply "lock" a certain range of the physical main memory address space for the GPU. So we'd still have to copy the input data from the "regular" memory area (used by the CPU) over to some place in the memory area reserved for GPU - and the results need to copied back the same way.

Also we are talking about Intel CPU's here, the upcoming "Arrandale" to be precise. So far Intel doesn't offer any GPGPU API for their GPU's. Until Intel does so (probably by making their GPU's accessible through OpenCL), we cannot use those combined CPU/GPU chips for anything but graphics output or video decoding at all! And if you look at the OpenCL API, it is defined similar to the CUDA API. In particular there is "host" memory that OpenCL kernels explicitly cannot access! And there's the "global" (device) memory, which all OpenCL kernels can access.

PCIe is a point-to-point link so it should be possible to run the GPU's link with a much higher clock rate. Standard PCIe runs at 100MHz. With CPU, GPU and PCIe Controller on the same die a much higher frequency for the GPU PCIe link should be a small problem. For example if you get it running with 1GHz, you increase the bandwidth and decrease the latency by a factor of 10.

That sounds like pure speculation. Unless there are some facts, I will assume that the "internal" PCIe-based interconnect will be roughly at the same level as "external" PCIe 2.0 is nowadays...

ajp_anton
16th December 2009, 16:39
Also we are talking about Intel CPU's here, the upcoming "Arrandale" to be precise. So far Intel doesn't offer any GPGPU API for their GPU's. Until Intel does so (probably by making their GPU's accessible through OpenCL), we cannot use those combined CPU/GPU chips for anything but graphics output or video decoding at all!No to mention that when we say "the integrated GPUs aren't powerful", we mean the low end of AMD and Nvidia, which is far from what Intel has to offer =)

edison
16th December 2009, 19:12
Fermi have 128KB L2 cache (per memory controller,total 768KB) that can use to speed up thread block sync.

and, MCP7X/GT200 and the later gpus can use system memory without copying to dedicated (video) memory for significant perf improvement, NVIDIA call this "zero copy accesss".

LoRd_MuldeR
16th December 2009, 20:48
Yup, that's one of the new features introduced with CUDA 2.2, don't know if OpenCL will have a similar functionality. But this doesn't solve the fundamental problem! All that "zero copy accesss" does is: The GPU can now fetch memory directly from the host process' memory space. This memory fetch is done across PCIe and it bypasses the "global" device memory. But still the speed for accessing the main ("host") memory from the GPU isn't anywhere near the CPU, as all data still must go through the PCIe bus - this means limited bandwidth as well as additional latency! You may be able to "hide" the latency by doing "stream" processing: Upload the next data element to the device, while the current data element is being processed. So when the current data element is finished processing, the next element has already arrived at the device and thus the device doesn't need to wait - it can continue with processing immediately. But it highly depends on the individual application/problem whether stream processing is feasible or not.

Last but not least it should be mentioned that the "zero copy accesss" is only supported by the Geforce-200 series and later, which currently excludes most CUDA-enabled devices!

slavickas
16th December 2009, 21:59
>Last but not least it should be mentioned that the "zero copy accesss" is only supported by the Geforce-200 series and later
minus GTS 250 thanks to renaming :)

kidjan
17th December 2009, 03:13
I don't know about you, but considering the fact that those measurements are useless for comparing actual visual quality, I think I'd rather have screenshots.

You're not comparing "actual visual quality"--you're comparing a reference image (i.e. input to the encoder) with an outputted image (i.e. decoded output). If you were comparing actual visual quality, we'd all be scrutinizing the quality of the input material.

Furthermore, I find a "screenshot" completely inadequate, given that any singular image from encoded video may not be indicative of the overall quality. SSIM or some other objective measurement should be used over the duration of video, and in the perfect world you'd return a box-plot (http://en.wikipedia.org/wiki/Box_plot) of cumulative SSIM/PSNR/whatever scores. That box plot would communicate A) outliers (which correspond with particularly bad frames), B) median SSIM, and C) how consistent image quality is based on the quartiles/whiskers.

I digress, and I'm open to the possibility that I'm being a pedantic jerk. But I find this method of posting screen shots of singular frames to be almost completely devoid of merit.

nm
17th December 2009, 16:30
You're not comparing "actual visual quality"--you're comparing a reference image (i.e. input to the encoder) with an outputted image (i.e. decoded output). If you were comparing actual visual quality, we'd all be scrutinizing the quality of the input material.

He means the actual (subjective) visual quality compared to the reference, of course.

Furthermore, I find a "screenshot" completely inadequate, given that any singular image from encoded video may not be indicative of the overall quality.

I agree, although not completely. I think screenshots are more useful than SSIM/PSNR scores or graphs when comparing different encoders.

SSIM or some other objective measurement should be used over the duration of video, and in the perfect world you'd return a box-plot (http://en.wikipedia.org/wiki/Box_plot) of cumulative SSIM/PSNR/whatever scores. That box plot would communicate A) outliers (which correspond with particularly bad frames), B) median SSIM, and C) how consistent image quality is based on the quartiles/whiskers.

I digress, and I'm open to the possibility that I'm being a pedantic jerk. But I find this method of posting screen shots of singular frames to be almost completely devoid of merit.
"Objective" measures such as PSNR or SSIM do not represent visual quality very well, so I'd argue that posting them as box-plots or graphs is also almost completely devoid of merit. I prefer full videos and randomly selected screenshots (with matched frametypes).

kidjan
17th December 2009, 22:43
He means the actual (subjective) visual quality compared to the reference, of course.

No, he does not mean the "visual quality" compared to the reference. He means the "visual similarity" compared to the reference. Quality is the wrong word.


"Objective" measures such as PSNR or SSIM do not represent visual quality very well, so I'd argue that posting them as box-plots or graphs is also almost completely devoid of merit. I prefer full videos and randomly selected screenshots (with matched frametypes).

Of course they don't represent "visual quality"; that isn't what SSIM or PSNR are used for, nor is it what the comparisons here are really about. Again: the comparisons are between a reference image and a decoded output. That does not need to be a "qualitative" comparison, nor should it be. For this sort of task, an objective meaure is clearly preferable to improperly conducting ad-hoc comparisons with singular frames because it allows us to quantify how well of a job the encoder did matching the input.

Lastly, if someone is going to do the Mean-Opinion-Score approach being done in this thread, stop labeling the encoder output. They're biasing results since observers know which encoder produced what output. Assuming the MOS approach is preferable (again, I disagree), then this thread is a wonderful example of what not to do in experimental design.

And now I'm definitely a pedantic jerk. Sorry. Will stop posting now.

nm
18th December 2009, 00:28
No, he does not mean the "visual quality" compared to the reference. He means the "visual similarity" compared to the reference. Quality is the wrong word.

You can replace the word "quality" with "similarity" if it pleases you. I doubt you'll get other people to use it though, since the term "quality" is pretty established in this context. See http://en.wikipedia.org/wiki/Video_quality and some of the references listed there, for example. "Similarity" is used in content-based retrieval and pattern matching research.

Of course they don't represent "visual quality"; that isn't what SSIM or PSNR are used for, nor is it what the comparisons here are really about. Again: the comparisons are between a reference image and a decoded output.

Yes, we all know what these comparisons are about. You don't need to get that worked up about terminology. It's besides the point and completely irrelevant.

That does not need to be a "qualitative" comparison, nor should it be. For this sort of task, an objective meaure is clearly preferable to improperly conducting ad-hoc comparisons with singular frames because it allows us to quantify how well of a job the encoder did matching the input.

The issue is that not all encoders try to optimize for PSNR or SSIM. Some have other "psy" optimizations that produce significantly worse SSIM scores but higher visual similarity.

I'd only use PSNR and SSIM for measuring the effect of certain (non-psy) parameters of an encoder, against itself.

kidjan
18th December 2009, 08:42
Some have other "psy" optimizations that produce significantly worse SSIM scores but higher visual similarity.

I'm skeptical of this claim. On what information do you base this assertion?

Regardless, it doesn't change the fact that the methods used in this thread are completely without any scientific merit. If they're going to do MOS, at least do it with some semblence of rigor.

LoRd_MuldeR
18th December 2009, 11:10
I'm skeptical of this claim. On what information do you base this assertion?

x264 development! When Psy-Optimizations were added, the subjective quality was improved significantly, as agreed by most users. At the same time SSIM metric did drop significantly.

You can try it out yourself: Encode the same source clip with "--ssim --tune film" and with "--ssim --tune ssim" (the latter includes "--no-psy"). Then you'll see...

(Side note: I'm currently developing a SSIM-based metric for a specific application and I'm having a hard time to measure a certain "effect" - one that is clearly visible to human viewers)

kidjan
23rd December 2009, 02:41
x264 development! When Psy-Optimizations were added, the subjective quality was improved significantly, as agreed by most users. At the same time SSIM metric did drop significantly.

You can try it out yourself: Encode the same source clip with "--ssim --tune film" and with "--ssim --tune ssim" (the latter includes "--no-psy"). Then you'll see...

(Side note: I'm currently developing a SSIM-based metric for a specific application and I'm having a hard time to measure a certain "effect" - one that is clearly visible to human viewers)

Thanks Mulder.

I'd be interested to see how you arrived at these results (i.e. what is "significant," how did you determine MOS dropped, etc.), but maybe in a different thread, I've already disrupted this one enough. Thanks for the feedback.

sethk
30th December 2009, 02:24
Yup, that's one of the new features introduced with CUDA 2.2, don't know if OpenCL will have a similar functionality. But this doesn't solve the fundamental problem! All that "zero copy accesss" does is: The GPU can now fetch memory directly from the host process' memory space. This memory fetch is done across PCIe and it bypasses the "global" device memory. But still the speed for accessing the main ("host") memory from the GPU isn't anywhere near the CPU, as all data still must go through the PCIe bus - this means limited bandwidth as well as additional latency! You may be able to "hide" the latency by doing "stream" processing: Upload the next data element to the device, while the current data element is being processed. So when the current data element is finished processing, the next element has already arrived at the device and thus the device doesn't need to wait - it can continue with processing immediately. But it highly depends on the individual application/problem whether stream processing is feasible or not.

Last but not least it should be mentioned that the "zero copy accesss" is only supported by the Geforce-200 series and later, which currently excludes most CUDA-enabled devices!

Zero-copy access is also available on the MCP79x series integrated GPUs, where the integrated GPU is also the memory controller (northbridge) for the CPU and my understanding is it doesn't have much of a penalty in use on this family of chips. Data accessed using this command does not need to be copied to another area before being accessed, although that is an option. In fact the whole point of zero-copy access is to .. have no copying going on. (See post here (http://forums.nvidia.com/index.php?showtopic=92290&view=findpost&p=519529)).

The GT200 on the other hand is a more complicated scenario, but it sounds like it might still be helpful in some encoding situations.

LoRd_MuldeR
30th December 2009, 02:39
Zero-copy access is also available on the MCP79x series integrated GPUs, where the integrated GPU is also the memory controller (northbridge) for the CPU and my understanding is it doesn't have much of a penalty in use on this family of chips. Data accessed using this command does not need to be copied to another area before being accessed, although that is an option. In fact the whole point of zero-copy access is to .. have no copying going on. (See post here (http://forums.nvidia.com/index.php?showtopic=92290&view=findpost&p=519529)).

And that's a low-performance "on board" chip, which isn't anywhere near NVidia's "high end" boards. I wouldn't expect great performance from that one ;)

Sure, it will beat everything that Intel can deliver currently, but unfortunately that doesn't mean much...

The GT200 on the other hand is a more complicated scenario, but it sounds like it might still be helpful in some encoding situations.

No, it doesn't solve the fundamental problem at all. All data still needs to go through the "slow" PCIe bus, twice! That means a serious bottleneck, bandwidth-wise and delay-wise.

All that "zero-copy access" does is: Data doesn't need to be uploaded to "global" memory before it can be copied to the "shared" memory of the individual block. It can go to the "shared" memory of the destination block immediately now. That avoids one indirection in some cases, but certainly not in all cases. In cases where more than one single thread block needs to access the data, we still need to store it in "global" memory. Also the data still needs to be uploaded via PCIe, which IMO is the most important problem! For a project I was working on we had to move calculations to the GPU, although they were slower(!) on the GPU than on the CPU. But downloading all the intermediate data to the CPU (host memory) was so slow, that we had to do everything on the GPU (global device memory) and only downloaded the final result. So believe me: The bottleneck of having to upload/download all the data through the PCIe bus (which currently is the only way on all competitive GPU's) is not just some "theoretical" thing. It's something you will encounter in reality! And it's something that can easily kill all the "nice" speedup you did expect by GPGPU processing. People are even trying to transfer data from one GPU board to another GPU board via DVI to avoid PCIe ^^

sethk
1st January 2010, 01:27
And that's a low-performance "on board" chip, which isn't anywhere near NVidia's "high end" boards. I wouldn't expect great performance from that one ;)

Sure, it will beat everything that Intel can deliver currently, but unfortunately that doesn't mean much...



No, it doesn't solve the fundamental problem at all. All data still needs to go through the "slow" PCIe bus, twice! That means a serious bottleneck, bandwidth-wise and delay-wise.

All that "zero-copy access" does is: Data doesn't need to be uploaded to "global" memory before it can be copied to the "shared" memory of the individual block. It can go to the "shared" memory of the destination block immediately now. That avoids one indirection in some cases, but certainly not in all cases. In cases where more than one single thread block needs to access the data, we still need to store it in "global" memory. Also the data still needs to be uploaded via PCIe, which IMO is the most important problem! For a project I was working on we had to move calculations to the GPU, although they were slower(!) on the GPU than on the CPU. But downloading all the intermediate data to the CPU (host memory) was so slow, that we had to do everything on the GPU (global device memory) and only downloaded the final result. So believe me: The bottleneck of having to upload/download all the data through the PCIe bus (which currently is the only way on all competitive GPU's) is not just some "theoretical" thing. It's something you will encounter in reality! And it's something that can easily kill all the "nice" speedup you did expect by GPGPU processing. People are even trying to transfer data from one GPU board to another GPU board via DVI to avoid PCIe ^^

Definitely true about the MCP being a slow solution - didn't mean to imply otherwise, but I mention it for two reasons - one, GT200 is not the only current solution supporting zero-copy and two, it did not have the same latency issues with zero-copy access.

As you say, this is all academic because of the low shader pipeline count on this solution.

For the GT200, even with the high latency induced by accessing main memory through the PCIe bus, it has enough bandwidth (8GB / sec) to main memory that I imagine it could be used in creative ways in conjunction with local memory.

Now the GT200 has 1GB of local memory, so you could certain use the local graphics memory (GDDR) to buffer large groups of uncompressed frames, and have the main CPU handle decompression and streaming of the uncompressed frames to GPU memory (which could store a large enough buffer of frames to handle backwards and forwards frame access), and write the compressed data back to main memory without hitting a PCIe memory bottleneck. Since most of the work would need to be done on uncompressed frames and that could be happening in GPU memory instead of main memory, the PCIe latency may not be as big a deal as if you tried to do everything in main memory.

I haven't attempted any of this, and I may be off-base in my thinking, but it was what I pictured as a reasonable approach.

yuvi
1st January 2010, 19:52
As stated in CUDA2.2PinnedMemoryAPIs.pdf, mapped pinned memory is always beneficial for integrated GPUs (and afaik supported by all CUDA-capable integrated GPUs since it's only needed that the driver map the memory the GPU uses into the application's VM space.) Discrete GPUs are completely different: mapped pinned memory is pretty much only beneficial over copying the data to device memory if the data is only accessed once and all accesses are coalesced. Thus, for most kernels zero-copy is slower for discrete GPUs.

DiKey
16th September 2010, 22:46
Is anything new with Fermi capabilities? Some new programs or algorithms?

Pakmenu
2nd January 2012, 11:09
If anybody noticed the bitrate in x264 peaks at 2229kbps! to make an average of 900kbps...
The cuda clip peaks at 1776kbps, which makes for less variability, and thus MORE bitrate at low motion, low contracst scenes.
Just think: to have scenes at 2230kbps and yet have an average of 900kbps, you have to have a lot of scenes having only let's say only 300kbps. to have a peak at 1780 you can have low mation scenes at maybe 500kbps.

Anyway: comparing a FRAME and not knowing the FRAMESIZE the encoder used.... is pointless! a fair comarison would look at one whole GOP structure (I frame and dependent b/p frames) and have both encoders set to encode the same size for all the frames in the GOP, then look at the quality.

The MORE variability a bitrate has, the more will be allocated at high contrast, high movement scenes and LESS bitrate to low contrast, low motion scenes. It looks like the picture grabbed is extremely dark, low contrast and probably not a lot of mation.
This means x264 would have allocated MORE bitrate to better lighted and higher motion scenes.
since it's not known how the rest of the clip looks, the rest of the clip might look better for x264!

Anyway when looking only at FRAME comparrison, the SIZE of THE FRAME should be taken EQUAL to be fair in comparrison!
(the framesize can be looked up when enabeling OSD in FFDshow.)

Indeed a pointless and unfair comparison!

Didée
2nd January 2012, 14:49
Ok, then let's look at full video streams (http://www.mediafire.com/?d6xb7m46ua4l6iv). No change, CUDA encode looks just crap.

CruNcher
2nd January 2012, 14:53
i can speak about experience with the Nvidia Encoder and it was the most interesting one (i analyzed most every change Nvidia Engineers did in the past on it (saw lot of bugs go by, was impressed it being the first one to support FreXt), by driver revision) until Quicksync came up which is rapidly enhancing :) and got another Quality boost

Though i would really like to see how ORBX http://us.download.nvidia.com/downloads/GTC_Videos/flvs/4008_GTC2010..wmv currently compares to H.264 on the GPU being desinged entirely for the GPU from the groundup :)
Jules brought some powerfull stuff together the last years with bringing Paul Debevec's Lightstage HDR Rendering Research on board of his cloud vision and his Engine work and doing his GPU Video Codec work :D

Though it could be that AMD is coming back also with their far from good Performance in the last years (Research wise) it seems they have some interesting stuff in the cradle out of labs for this year (the Realtime Deshaking introduction surprised a little and now the improvement of it v 2.0 in their coming GPUs which also could give their General Video Encoding a big boost)

hajj_3
7th January 2012, 16:57
@dark shikari, i don't suppose you might want to write an updated article about h.265/hevc as the complete draft of the spec is due in february.