View Full Version : ffdshow tryouts project: Discussion & Development
clsid
19th December 2008, 20:24
Can't say that the "normal" branch of ffdshow became significant faster between Beta-5 (r2033) and Beta-6 (r2527) on my system...Since beta5 there were also some fixes regarding compliance to the H.264 specification. Some of those had a negative performance impact.
r2527 is 4% faster on my system than builds from last week. CPUs with SSE2 should get bigger gains.
Your results show a smaller increase in performance because you tested with 4 threads. Test the non-MT builds with 1 thread.
clsid
19th December 2008, 20:27
if you can provide the generic/GCC/ICL 10.1 SSE/ICL 10.1 SSE2 builds of the same rev, I can run compares on h264 HD decoding + sharpening in RGB32HQ on an o/c Q6600(I'm also a Remoulade beta tester)
Generic and ICL10 builds are online. GCC build is no longer officially supported.
LoRd_MuldeR
19th December 2008, 20:30
You results show a smaller increase in performance because you tested with 4 threads. Test the non-MT builds with 1 thread.
I don't get this. If the latest revision really runs faster than the older revision with one single thread, why should this speed-up (difference) suddenly go away with several threads?
Especially when both versions compared use the very same "old" multi-threading implementation...
laserfan
19th December 2008, 20:38
Is there any way to identify the version of ffdshow (libavcodec in particular) via a .cmd line?
fastplayer
19th December 2008, 20:43
I don't get this. If the latest revision really runs faster than the older revision with one single thread, why should this speed-up go away with several threads?
Especially when both versions compared use the very same "old" multi-threading implementation...
My guess:
Thread creating, spawning, switching etc. are all operations that require themselves a bunch of CPU cycles, cache, latency etc. Apparently, these "costs" are so high on quad core that they totally negate any performance gain.
leeperry
19th December 2008, 20:47
Generic and ICL10 builds are online. GCC build is no longer officially supported.
ok, but I guess the idea is to benchmark beta6 ?
so if you wouldn't mind to build it in icl10/generic, I'll be happy to try them :)
there used to be separate sse1/sse2 ICL versions, but that doesn't existe in ICL 10.1 anymore ? when yesgrey3 built ICL10.1 versions of the Reclock resampler, some of them were much faster than others...I could ask him for the best settings.
also, that'd be good to take audio PP in account, because that's where the ICL10 versions shine IMO...and timecodec can't benchmark that..
LoRd_MuldeR
19th December 2008, 20:53
Apparently, these "costs" are so high on quad core that they totally negate any performance gain.
If that was the case, then multiple threads would run slower than one thread (in the none-MT version) on my Machine.
But in fact the none-MT version never uses uses more than two threads for H.264. And, although it can't keep up with the MT version, the two thread patch gives some speed-up!
I don't see why this should no longer be the case (and even turn to the opposite), after further optimizations have arrived...
roozhou
19th December 2008, 21:02
If that was the case, then multiple threads would run slower than one thread (in the none-MT version).
But in fact the none-MT version never uses uses more than two threads for H.264. And, although it can't keep up with the MT version, the two thread patch gives some speed-up!
I don't see why this should no longer be the case (and even turn to the opposite), after further optimizations have arrived...
What CPU are you using? Recently Dark Shikari has merged some of x264's SSE2 codes into ffmpeg. If you are using non-Phenom AMD CPU, it won't give you significant speedup.
LoRd_MuldeR
19th December 2008, 21:04
What CPU are you using? Recently Dark Shikari has merged some of x264's SSE2 codes into ffmpeg. If you are using non-Phenom AMD CPU, it won't give you significant speedup.
See "My Specs" in my signature. SSE2 is supported by my CPU.
clsid
19th December 2008, 21:08
@leeperry,
I am not going to make any more builds. Four is enough for today. SSE2 showed no gain in the past over SSE with the ICL builds, so I am not going to waste my time on that again.
Use the trunk builds.
@LoRd_MuldeR,
I never said that there won't be any speedup with >1 threads. I just said that testing with 1 thread gives a better picture that isn't clouded by the effect of a crappy MT implementation.
LoRd_MuldeR
19th December 2008, 21:11
@LoRd_MuldeR,
I never said that there won't be any speedup with >1 threads. I just said that testing with 1 thread gives a better picture that isn't clouded by the effect of a crappy MT implementation.
Okay. I will do another test later and I will explicitly enforce one single thread. Right now a capture is in progress...
fastplayer
19th December 2008, 21:12
If that was the case, then multiple threads would run slower than one thread (in the none-MT version).
I think you misunderstood me. In this particular situation the performance gains that have been achieved by the FFmpeg guys, are not enough to overcome the cost that is associated with spawning a 2nd, 3rd or 4th thread. Either that or the changes they made are just less SMP-friendly...
LoRd_MuldeR
19th December 2008, 21:22
I think you misunderstood me. In this particular situation the performance gains that have been achieved by the FFmpeg guys, are not enough to overcome the cost that is associated with spawning a 2nd, 3rd or 4th thread. Either that or the changes they made are just less SMP-friendly...
Sure. But if two threads already run faster than one thread, which is the case on my system (even with the none-MT version), then I hardly can imagine how further optimizations can make the "one thread" variant run faster, while there is no noticeable speed-up for the "two threads" variant. I'd rather assume that these optimizations don't help my Core2 as much as other CPUs. But as said before, I will do more tests later.
fastplayer
19th December 2008, 21:27
Generic and ICL10 builds are online. GCC build is no longer officially supported.
You mean GCC is not used anymore for compiling ffdshow.ax, correct?
yesgrey
19th December 2008, 21:31
when yesgrey3 built ICL10.1 versions of the Reclock resampler, some of them were much faster than others...I could ask him for the best settings.
The best settings are always application dependent...
For the resampler the best settings were to not use any Intel extensions...
Maybe with the new code the result is different, but I don't believe it.
tetsuo55
19th December 2008, 21:40
A few people have reported an issue with the x64 builds of ffdshow where ffdshow uses an unusually high amount of CPU cycles, resulting in bad playback (stuttering, frame drops). This issue is present for a long time now, so not related to any recent changes.
Weird thing is that the CPU usage returns to normal when the OSD is enabled in ffdshow video decoder.
Does anyone have an idea what might cause this problem, and why the OSD makes a difference?
I myself am unable to reproduce the issue on a clean install of Vista x64.
Just based on the description this seems to be the inversion of the bug i reported earlier. OSD enabled causes it to show 100% CPU when combined with auto-post processing.
I think the OSD is broken in more ways than one resulting in unexpected behaviour in several parts of ffdshow
In the case you mention it helps, it the case i mentioned it hurts.
Imho it should get a higher priority than it has at the moment
fastplayer
19th December 2008, 21:46
I'd rather assume that these optimizations don't help my Core2 as much as other CPUs.
M. Niedermayer made quite a few H.264-related commits in the past few days and he's running a Merom CPU (castrated C2D):
http://lists.mplayerhq.hu/pipermail/ffmpeg-cvslog/2008-December/018415.html
tetsuo55
19th December 2008, 22:00
Yeah all those h264 commits look great, all those 0.? speedups have to combine into a nice ?.? somewhere.
According to the changelog the speedups where measured on a pentium dual. Also a lot of unneeded checks and calculations where dropped.
This means on a per sample and per system basis the increase in FPS can be pretty high(probably never more than 10% though)
Also there seems to be more work done on realmedia 30 and 40, hopefully there should be less problems with it now and i hope that its soon fully able to replace realplayer.
Snowknight26
19th December 2008, 23:16
As for the interlacing flags, are you sure that the stream has correct flags?
I don't know how to find that out, but here (http://stfcc.org/misc/00000.cut2.m2ts) is a sample of where it changes from progressive to interlaced and vice versa.
STaRGaZeR
20th December 2008, 00:10
I don't know how to find that out, but here (http://stfcc.org/misc/00000.cut2.m2ts) is a sample of where it changes from progressive to interlaced and vice versa.
Your sample is MBAFF. That means it's interlaced at macroblock level. In the same frame it may be macroblocks coded as interlaced and others as progressive. In order to view this correctly the entire stream has to be deinterlaced, just like it is now.
BTW Tiesto rules :p
haruhiko_yamagata
20th December 2008, 00:46
Some new numbers:
[...]
Can't say that the "normal" branch of ffdshow became significant faster between Beta-5 (r2033) and Beta-6 (r2527) on my system...
Beta6 is 5 - 12% faster for me. Please make sure you get the new pre-beta6.
Revision 16239 - Directory Listing
Modified Fri Dec 19 13:45:13 2008 UTC (10 hours, 15 minutes ago) by darkshikari
Port x264 deblocking code to libavcodec. This includes SSE2 luma deblocking code and both MMXEXT and SSE2 luma intra deblocking code for H.264 decoding. This assembly is available under --enable-gpl and speeds decoding of Cathedral by 7%.
Snowknight26
20th December 2008, 01:07
Your sample is MBAFF. That means it's interlaced at macroblock level. In the same frame it may be macroblocks coded as interlaced and others as progressive. In order to view this correctly the entire stream has to be deinterlaced, just like it is now.
Since thats the case, when I enable the deinterlacer, it shouldn't be deinterlacing the progressive macroblocks... but it does anyway.
haruhiko_yamagata
20th December 2008, 01:11
when I enable the deinterlacer, it shouldn't be deinterlacing the progressive macroblocksThis is wrong. A macroblock encoded using progressive algorithm may require deinterlacing.
fastplayer
20th December 2008, 01:26
300-tlr2_h1080p.mov | terminatorsalvation-tlr2_h1080p.mov | Madagascar.avi
2033: 44.7 | 43.2 | 198.2
2527: 45.4 | 44.9 | 202.5
2527 ICL: 45.4 | 45.0 | 202.4
- All results in dfps
- First 2 trailers: 1920x800 H.264, 3rd one: 1280x720 DX50
- CPU: Athlon64 3500+
Snowknight26
20th December 2008, 01:29
This is wrong. A macroblock encoded using progressive algorithm may require deinterlacing.
May require. If it doesn't require it, does it still get deinterlaced? Might I remind you of the screenshots I previously posted (#5721 (http://forum.doom9.org/showpost.php?p=1225749&postcount=5721)).
haruhiko_yamagata
20th December 2008, 01:40
May require. If it doesn't require it, does it still get deinterlaced? Might I remind you of the screenshots I previously posted (#5721 (http://forum.doom9.org/showpost.php?p=1225749&postcount=5721)).Anyway, the decoder flagged the frame correctly.
Deinterlace or not is choice of deinterlacers. If you use a good deinterlacer, you will have satisfactory results.
LoRd_MuldeR
20th December 2008, 02:27
Beta6 is 5 - 12% faster for me. Please make sure you get the new pre-beta6.
I use the latest version. Re-downloaded, just to be sure. File from 2008-12-19, 18:01.
This time I ran the comparison with only one single thread, so the multi-threading code can't have any impact.
However there still is no remarkable difference between Beta-5 and preBeta-6. See:
E:\HD\freedom EP1 sample.mkv, 1920x1080, High@L4.1
[ffdshow, rev2033, Beta-5, 2008-07-05, 1 thread]
User: 31s, kernel: 0s, total: 31s, real: 30s, fps: 19.8, dfps: 20.0
User: 30s, kernel: 0s, total: 30s, real: 30s, fps: 20.0, dfps: 20.0
User: 30s, kernel: 0s, total: 30s, real: 30s, fps: 20.0, dfps: 19.9
[ffdshow, rev2527, Pre-Beta 6, 2008-12-19, 1 thread]
User: 30s, kernel: 0s, total: 30s, real: 30s, fps: 20.0, dfps: 20.0
User: 30s, kernel: 0s, total: 30s, real: 30s, fps: 19.9, dfps: 20.0
User: 31s, kernel: 0s, total: 31s, real: 30s, fps: 19.7, dfps: 19.9
[ffdshow-MT, rev2525, 2008-12-20, 1 thread]
User: 30s, kernel: 0s, total: 30s, real: 30s, fps: 20.1, dfps: 20.0
User: 31s, kernel: 0s, total: 31s, real: 31s, fps: 19.6, dfps: 19.5
User: 31s, kernel: 0s, total: 31s, real: 31s, fps: 19.5, dfps: 19.5
[ffdshow-MT, rev2525, 2008-12-20, 4 threads]
User: 2s, kernel: 0s, total: 2s, real: 8s, fps: 244.9, dfps: 71.5
User: 2s, kernel: 0s, total: 2s, real: 8s, fps: 236.1, dfps: 71.3
User: 2s, kernel: 0s, total: 2s, real: 8s, fps: 244.9, dfps: 71.0
[CoreAVC, Version 1.8.5]
User: 0s, kernel: 0s, total: 0s, real: 7s, fps: 625.8, dfps: 83.2
User: 0s, kernel: 0s, total: 0s, real: 7s, fps: 773.0, dfps: 83.0
User: 1s, kernel: 0s, total: 1s, real: 7s, fps: 588.4, dfps: 82.3
[DivX H.264 Decoder, Beta-3]
User: 1s, kernel: 0s, total: 1s, real: 6s, fps: 499.0, dfps: 89.4
User: 1s, kernel: 0s, total: 1s, real: 6s, fps: 433.2, dfps: 89.0
User: 1s, kernel: 0s, total: 1s, real: 7s, fps: 486.7, dfps: 88.0
Another sample, just to be sure. But same result:
E:\HD\Crowd Run 2160p UHD CRF22 x264-CtrlHD.mkv
[ffdshow, rev2033, Beta-5, 2008-07-05, 1 thread]
User: 121s, kernel: 0s, total: 121s, real: 121s, fps: 4.1, dfps: 4.1
User: 121s, kernel: 0s, total: 121s, real: 121s, fps: 4.1, dfps: 4.1
User: 121s, kernel: 0s, total: 121s, real: 121s, fps: 4.1, dfps: 4.1
[ffdshow, rev2527, Pre-Beta 6, 2008-12-19, 1 thread]
User: 120s, kernel: 0s, total: 120s, real: 120s, fps: 4.1, dfps: 4.1
User: 120s, kernel: 0s, total: 120s, real: 120s, fps: 4.1, dfps: 4.1
User: 121s, kernel: 0s, total: 121s, real: 120s, fps: 4.1, dfps: 4.1
[ffdshow-MT, rev2525, 2008-12-20, 4 threads]
User: 8s, kernel: 0s, total: 8s, real: 35s, fps: 61.1, dfps: 14.2
User: 7s, kernel: 0s, total: 8s, real: 35s, fps: 62.4, dfps: 14.2
User: 7s, kernel: 0s, total: 7s, real: 35s, fps: 62.9, dfps: 14.2
[DivX H.264 Decoder, Beta-3]
User: 3s, kernel: 0s, total: 4s, real: 28s, fps: 121.7, dfps: 17.5
User: 4s, kernel: 0s, total: 4s, real: 28s, fps: 104.6, dfps: 17.5
User: 4s, kernel: 0s, total: 4s, real: 28s, fps: 118.5, dfps: 17.4
Snowknight26
20th December 2008, 06:24
Deinterlace or not is choice of deinterlacers. If you use a good deinterlacer, you will have satisfactory results.
Which deinterlacer that comes with ffdshow would qualify as being 'good?'
Dark Shikari
20th December 2008, 06:33
I use the latest version. Re-downloaded, just to be sure. File from 2008-12-19, 18:01.
This time I ran the comparison with only one single thread, so the multi-threading code can't have any impact.
However there still is no remarkable difference between Beta-5 and preBeta-6. See:Are you sure the person who compiled your copy of ffdshow had yasm installed? Otherwise, none of the new assembly code will get used... :rolleyes:
(it also requires --enable-gpl...)
haruhiko_yamagata
20th December 2008, 07:29
@LoRd_MuldeR
It may dependent on samples. Please try premiere-paff.ts or bbc-japan_1080p.mov.
fastplayer
20th December 2008, 11:34
Here are more results on my single-core Athlon64, this time with various filters applied:
http://i40.tinypic.com/2pyamhk.pnghttp://i41.tinypic.com/30vbpxe.png
http://i40.tinypic.com/2ep4l1x.pnghttp://i39.tinypic.com/9scor4.png
Looks like the MSVC compilers have caught up with ICL. Or ICL just doesn't like AMD CPU's :D
H.264 decoding performance has increased by 1.5-4% when comparing Beta 5 and pre-Beta 6.
clsid
20th December 2008, 13:08
Are you sure the person who compiled your copy of ffdshow had yasm installed? Otherwise, none of the new assembly code will get used... :rolleyes:
(it also requires --enable-gpl...)
Yasm 0.7.2, so don;t worry ;)
Beta5 is from before Michael's major H.264 compliance fixes. Which have had some impact on performance. So that is why prebeta6 should not be compared with it, but with a build from say last week.
I have seen speedups varying from 1% to 11% on a Core2.
clsid
20th December 2008, 13:13
Looks like the MSVC compilers have caught up with ICL. Or ICL just doesn't like AMD CPU's :D
I use patched ICL libs. Without that performance would even be worse on AMD.
ICL only shows benefit in just a few filters, like xsharpen. Perhaps you could test some more processing filters to see if there are more that benefit?
haruhiko_yamagata
20th December 2008, 13:53
I use patched ICL libs. Without that performance would even be worse on AMD.
ICL only shows benefit in just a few filters, like xsharpen. Perhaps you could test some more processing filters to see if there are more that benefit?As for xsharpen, MSVC9 does better job than expected.
Another problem was equalizer. Is it better now?
leeperry
20th December 2008, 14:14
FF.MKV (1080p h264)
720p spline resize/unsharp masking/LSF/ConvertToRGB32()
ffdshow_rev2447_20081208_clsid_sse_icl10.exe :
User: 1s, kernel: 0s, total: 1s, real: 3s, fps: 46.0, dfps: 25.9
User: 1s, kernel: 0s, total: 2s, real: 3s, fps: 43.5, dfps: 25.9
User: 1s, kernel: 0s, total: 1s, real: 3s, fps: 43.8, dfps: 26.0
ffdshow_rev2527_20081219_clsid_sse_icl10.exe :
User: 1s, kernel: 0s, total: 1s, real: 3s, fps: 44.5, dfps: 26.5
User: 1s, kernel: 0s, total: 1s, real: 3s, fps: 44.2, dfps: 26.3
User: 2s, kernel: 0s, total: 2s, real: 3s, fps: 43.5, dfps: 26.4
ffdshow_rev2527_20081219_clsid.exe :
User: 1s, kernel: 0s, total: 1s, real: 3s, fps: 45.3, dfps: 26.5
User: 1s, kernel: 0s, total: 1s, real: 3s, fps: 43.8, dfps: 26.5
User: 1s, kernel: 0s, total: 2s, real: 3s, fps: 43.5, dfps: 26.5
ffdshow_rev2488_20081213_xxl_mt.exe :
User: 3s, kernel: 0s, total: 3s, real: 3s, fps: 27.4, dfps: 22.9
User: 3s, kernel: 0s, total: 3s, real: 3s, fps: 26.9, dfps: 22.9
User: 3s, kernel: 0s, total: 3s, real: 3s, fps: 27.4, dfps: 23.0
ffdshow_rev2527_20081219_clsid_sse_icl10.exe + CoreAVC 1.8.5 :
User: 1s, kernel: 0s, total: 1s, real: 3s, fps: 47.2, dfps: 28.0
User: 1s, kernel: 0s, total: 1s, real: 3s, fps: 47.2, dfps: 28.0
User: 1s, kernel: 0s, total: 1s, real: 3s, fps: 45.6, dfps: 27.8
ffdshow_rev2527_20081219_clsid_sse_icl10.exe + Remoulade beta3 :
User: 1s, kernel: 0s, total: 1s, real: 3s, fps: 44.2, dfps: 27.6
User: 2s, kernel: 0s, total: 2s, real: 3s, fps: 40.9, dfps: 27.7
User: 1s, kernel: 0s, total: 1s, real: 3s, fps: 44.2, dfps: 27.7
Death.and.Life.of.Bobby.Z.mkv (720p 2.35 h264)
unsharp masking/LSF/ConvertToRGB32()
ffdshow_rev2447_20081208_clsid_sse_icl10.exe :
User: 8s, kernel: 2s, total: 11s, real: 14s, fps: 44.3, dfps: 34.2
User: 8s, kernel: 2s, total: 10s, real: 14s, fps: 46.2, dfps: 35.4
User: 8s, kernel: 2s, total: 10s, real: 14s, fps: 46.9, dfps: 35.3
ffdshow_rev2527_20081219_clsid_sse_icl10.exe :
User: 8s, kernel: 2s, total: 11s, real: 14s, fps: 45.4, dfps: 35.5
User: 8s, kernel: 2s, total: 10s, real: 14s, fps: 46.3, dfps: 35.5
User: 8s, kernel: 2s, total: 10s, real: 14s, fps: 46.5, dfps: 35.4
ffdshow_rev2527_20081219_clsid.exe :
User: 8s, kernel: 2s, total: 11s, real: 14s, fps: 44.8, dfps: 34.2
User: 8s, kernel: 2s, total: 11s, real: 14s, fps: 44.2, dfps: 34.2
User: 8s, kernel: 2s, total: 11s, real: 14s, fps: 44.1, dfps: 34.2
ffdshow_rev2488_20081213_xxl_mt.exe :
User: 8s, kernel: 2s, total: 10s, real: 13s, fps: 48.4, dfps: 37.1
User: 8s, kernel: 2s, total: 10s, real: 13s, fps: 47.1, dfps: 36.9
User: 8s, kernel: 1s, total: 10s, real: 13s, fps: 46.2, dfps: 36.4
ffdshow_rev2527_20081219_clsid_sse_icl10.exe + CoreAVC 1.8.5 :
User: 8s, kernel: 1s, total: 10s, real: 13s, fps: 49.2, dfps: 37.8
User: 8s, kernel: 1s, total: 10s, real: 13s, fps: 48.7, dfps: 37.8
User: 8s, kernel: 1s, total: 10s, real: 13s, fps: 48.4, dfps: 38.0
ffdshow_rev2527_20081219_clsid_sse_icl10.exe + Remoulade beta3 :
User: 8s, kernel: 2s, total: 10s, real: 12s, fps: 46.9, dfps: 38.7
User: 9s, kernel: 1s, total: 11s, real: 13s, fps: 44.3, dfps: 38.1
User: 8s, kernel: 2s, total: 10s, real: 12s, fps: 46.1, dfps: 38.9
Dila_720p.mkv (720p 1.78 h264)
unsharp masking/LSF/ConvertToRGB32()
ffdshow_rev2447_20081208_clsid_sse_icl10.exe :
User: 1s, kernel: 0s, total: 2s, real: 2s, fps: 21.8, dfps: 19.0
User: 1s, kernel: 0s, total: 2s, real: 2s, fps: 22.5, dfps: 18.9
User: 1s, kernel: 0s, total: 2s, real: 2s, fps: 22.2, dfps: 19.2
ffdshow_rev2527_20081219_clsid_sse_icl10.exe :
User: 1s, kernel: 0s, total: 2s, real: 2s, fps: 22.2, dfps: 19.8
User: 1s, kernel: 0s, total: 2s, real: 2s, fps: 23.9, dfps: 19.8
User: 1s, kernel: 0s, total: 2s, real: 2s, fps: 22.7, dfps: 19.9
ffdshow_rev2527_20081219_clsid.exe :
User: 1s, kernel: 0s, total: 2s, real: 2s, fps: 22.1, dfps: 19.8
User: 1s, kernel: 0s, total: 2s, real: 2s, fps: 21.9, dfps: 19.8
User: 1s, kernel: 0s, total: 2s, real: 2s, fps: 23.7, dfps: 19.6
ffdshow_rev2488_20081213_xxl_mt.exe :
User: 1s, kernel: 0s, total: 2s, real: 2s, fps: 22.7, dfps: 20.3
User: 1s, kernel: 0s, total: 2s, real: 2s, fps: 23.4, dfps: 20.1
User: 1s, kernel: 0s, total: 2s, real: 2s, fps: 22.7, dfps: 20.1
ffdshow_rev2527_20081219_clsid_sse_icl10.exe + CoreAVC 1.8.5 :
User: 1s, kernel: 0s, total: 2s, real: 2s, fps: 23.9, dfps: 20.3
User: 1s, kernel: 0s, total: 2s, real: 2s, fps: 23.9, dfps: 20.5
User: 1s, kernel: 0s, total: 2s, real: 2s, fps: 23.4, dfps: 20.3
ffdshow_rev2527_20081219_clsid_sse_icl10.exe + Remoulade beta3 :
User: 1s, kernel: 0s, total: 2s, real: 2s, fps: 22.7, dfps: 20.1
User: 1s, kernel: 0s, total: 2s, real: 2s, fps: 22.7, dfps: 20.1
User: 1s, kernel: 0s, total: 2s, real: 2s, fps: 23.5, dfps: 20.1
notes :
1)-test were conducted on a 3.3Ghz Q6600(8*415)
2)-the ffdshow MT version failed with the FF.mkv 1080p sample, I've uploaded it here :
http://www.megaupload.com/?d=YGJA8GE2
3)-even though Remoulade sometimes has higher dfps, its fps is lower than CoreAVC...and CoreAVC seems to offer better realtime performance, so I'm not sure that the dfps figure is the only one to care for?
4)-timecodec doesn't measure ffdshow audio performance(which is higher w/ ICL10 builds)
5)-I've set timecodec to the highest priority on 4 cores w/ as little background processes as possible(in low priority on single cores) + short h264 samples to measure the decoding speed, not the computer throughput
fastplayer
20th December 2008, 14:48
By the way, don't browse the page with IE6... :D
Christmas is around corner, I'm in a generous mood. So here's one last gift to you IE6-diehards:
ffdshow's new homepage (http://ffdshow-tryout.sourceforge.net/index2.php) now displays correctly on your "browser"! :D
Well, it's readable, but i would like it a bit bigger. I think it is a click smaller than normal. I am on 1680x1050 20" monitor (maybe on 21"-22" there's no problem).
I've made some improvements to the homepage which should please - not everyone - but most of us:
Helvetica and Trebuchet MS are used for improved readability
Font size and line height have been increased:
Should look sexier now on high-res displays. On resolutions of 1024x768 and less it looks a bit too bulky but I'm just too lazy to let font size adjust itself dynamically based on screen resolution... :p
haruhiko_yamagata
20th December 2008, 14:54
Got a similar problem with recent builds (normal and MT, but not x64) when playing interlaced MPEG-2 and using Yadif.
With Yadif enabled (internal version -or- Avisynth version) I get "slow motion" playback, which disappears when the OSD is enabled!
Had a lot of private discussion with haruhiko_yamagata about that problem already, but no solution yet.
It came down to the following result:
* The time ffdshow spends in Yadif is always okay.
* When the slow motion happens, then ffdshow spends a very long time (much too long) in the "convert" function!
* The "convert" time is back to normal with OSD enabled.
Note that there was no such problem in rev2347, seems it started around rev2391 ...The slow down is caused by very slow V-RAM access. It's a bug of video driver. ffdshow is just triggering it.
Does "High quality YV12 to RGB conversion" matter?
STaRGaZeR
20th December 2008, 16:02
Just to confirm this issue (http://forum.doom9.org/showpost.php?p=1225914&postcount=5744) with the MT branch, Snowknight26's Tiesto sample (http://forum.doom9.org/showpost.php?p=1226004&postcount=5770) has the sample problem in at least frame 538 or surrounding frames using MPC's internal MPEG PS/TS/PVA splitter, same freeze, same behaviour. Using only 1 thread fixes it again.
LoRd_MuldeR
20th December 2008, 16:09
The slow down is caused by very slow V-RAM access. It's a bug of video driver. ffdshow is just triggering it.
I wonder how ffdshow can be effected by V-RAM access. It doesn't access the graphics card directly, it just sends over the decoded frame to the next filter in graph, right? So I would think that the "convert" time only measures the time to convert the frame internally (in ffdshow) and send it to the next filter (most likely a renderer), but nothing more. Even if it had to wait for a "slow" renderer, that problem would disappear as soon as queuing is on. But it doesn't. Also as mentioned before, the problem isn't there in older revision and it reproducible appears as soon as I install a newer revision. And last but not least: How should the renderer detect that Yadif is used and then decide to do a "slow" V-RAM access now? While it does "fast" V-RAM access with KernelBob in use. Crazy, isn't it?
Does "High quality YV12 to RGB conversion" matter?
Nope. Makes no difference.
It's only "OSD" that for some reason makes a difference :confused:
Are you sure the person who compiled your copy of ffdshow had yasm installed? Otherwise, none of the new assembly code will get used... :rolleyes:
I took clsid's build from ffdshow's sourceforge site. And I think he knows what he does :)
LoRd_MuldeR
20th December 2008, 17:38
@LoRd_MuldeR
It may dependent on samples. Please try premiere-paff.ts or bbc-japan_1080p.mov.
Okay, tried again with the "premiere-paff.ts" sample. But still the difference between Beta-5 (r2033) and preBeta-6 (r2527) is negligible:
E:\HD\premiere-paff.ts
[ffdshow, rev2033, Beta-5, 2008-07-05, 1 thread]
User: 29s, kernel: 0s, total: 30s, real: 30s, fps: 39.8, dfps: 39.4
User: 29s, kernel: 0s, total: 29s, real: 30s, fps: 39.9, dfps: 39.4
User: 30s, kernel: 0s, total: 30s, real: 30s, fps: 39.7, dfps: 39.2
[ffdshow, rev2033, Beta-5, 2008-07-05, 4 threads]
User: 3s, kernel: 0s, total: 3s, real: 13s, fps: 341.1, dfps: 87.5
User: 3s, kernel: 0s, total: 3s, real: 13s, fps: 315.8, dfps: 87.3
User: 3s, kernel: 0s, total: 3s, real: 13s, fps: 335.2, dfps: 87.2
[ffdshow, rev2527, Pre-Beta 6, 2008-12-19, 1 thread]
User: 29s, kernel: 0s, total: 29s, real: 30s, fps: 40.0, dfps: 39.7
User: 29s, kernel: 0s, total: 29s, real: 30s, fps: 40.2, dfps: 39.7
User: 29s, kernel: 0s, total: 29s, real: 30s, fps: 40.2, dfps: 39.7
[ffdshow, rev2527, Pre-Beta 6, 2008-12-19, 4 threads]
User: 3s, kernel: 0s, total: 3s, real: 13s, fps: 317.1, dfps: 87.7
User: 3s, kernel: 0s, total: 3s, real: 13s, fps: 353.8, dfps: 87.3
User: 3s, kernel: 0s, total: 3s, real: 13s, fps: 321.1, dfps: 87.0
[ffdshow-MT, rev2525, 2008-12-20, 1 thread]
User: 30s, kernel: 0s, total: 30s, real: 30s, fps: 39.4, dfps: 39.1
User: 30s, kernel: 0s, total: 30s, real: 30s, fps: 39.5, dfps: 39.0
User: 30s, kernel: 0s, total: 30s, real: 30s, fps: 39.6, dfps: 39.0
[ffdshow-MT, rev2525, 2008-12-20, 4 threads]
User: 2s, kernel: 0s, total: 2s, real: 9s, fps: 457.6, dfps: 130.2
User: 2s, kernel: 0s, total: 2s, real: 9s, fps: 446.9, dfps: 130.0
User: 2s, kernel: 0s, total: 2s, real: 9s, fps: 404.3, dfps: 129.3
[CoreAVC Decoder, v1.8.5]
User: 1s, kernel: 0s, total: 1s, real: 7s, fps: 1032.6, dfps: 160.2
User: 1s, kernel: 0s, total: 1s, real: 7s, fps: 899.0, dfps: 159.2
User: 1s, kernel: 0s, total: 1s, real: 7s, fps: 1091.7, dfps: 158.9
[DivX H.264 Decoder, Beta-3]
User: 1s, kernel: 0s, total: 1s, real: 7s, fps: 858.6, dfps: 169.8
User: 1s, kernel: 0s, total: 1s, real: 7s, fps: 771.9, dfps: 169.0
User: 1s, kernel: 0s, total: 1s, real: 7s, fps: 734.8, dfps: 168.7
ash925
20th December 2008, 18:43
With versions 2503 and 2527 I am getting a black or green output in virtualdub(1.8.6)though the videos run fine in mediaplayer classic(patched build).Version 2033 runs video fine, I want to know is there any way to get proper output in virtualdub(any settings that need to be changed). The problem exists with h264 videos in avi or mp4 container(mp4 plugin is used of course).
LoRd_MuldeR
20th December 2008, 18:47
With versions 2503 and 2527 I am getting a black or green output in virtualdub(1.8.6)though the videos run fine in mediaplayer classic(patched build).Version 2033 runs video fine, I want to know is there any way to get proper output in virtualdub(any settings that need to be changed).
Media players access ffdshow through the DirectShow interface, while VirtualDub uses the VfW interface only.
Goto "ffdshow" -> "VFW configuration" and make sure all required video decoders (Codecs) are enabled on the "Decoder" tab.
Also make sure the desired color format(s) are checked on the "Output" page.
BTW: About what video formats we are talking here?
ash925
20th December 2008, 18:58
Was trying to update the Format thing but for some reason the site became tooo slow.
H264 videos in avi or mp4 videos.
I had the h264 option in vfw enabled each time.
LoRd_MuldeR
20th December 2008, 19:02
H264 videos in avi or mp4 videos.
VirtualDub doesn't support MP4 files. Unless a new MP4 input plugin was released recently...
I had the h264 option in vfw enabled each time.
Then it should work. At least it does here.
clsid
20th December 2008, 19:20
Try pressing F9 (show input pane).
ash925
20th December 2008, 19:39
@LoRd_MuldeR:"VirtualDub doesn't support MP4 files. Unless a new MP4 input plugin was released recently..."
Well I found a plugin and it seems to work(at least upto now).
@clsid:Input and ouput pane both are enabled.
It seems strange that MPC can use it, but virtualdub can't.
LoRd_MuldeR
20th December 2008, 19:42
It seems strange that MPC can use it, but virtualdub can't.
MPC is a DirectShow-baes player, VirtualDub uses VFW Codecs. These are two completely different things!
The only reason why ffdshow works in VDub at all is that it provides a special VfW interface for such "legacy" applications.
ash925
20th December 2008, 19:50
Thanks for the clarification LoRd_MuldeR.
BTW problem solved, I accidentally put a matroska file in Virtualdub and the video shows just fine, must have something to do with Haali Splitter.
LoRd_MuldeR
20th December 2008, 20:19
...must have something to do with Haali Splitter.
Impossible. As said before, VirtualDub doesn't use DirectShow. Hence it doesn't use any DirectShow filters, such as Haali Splitter.
VirtualDub only supports VfW Codecs and it's limited to AVI files, using it's own internal AVI splitter.
Other containers than AVI can only be opened in VDub through special VDub Input Plugins and even that feature was only added recently.
squid_80
21st December 2008, 00:11
There is a directshow input plugin...
vBulletin® v3.8.11, Copyright ©2000-2026, vBulletin Solutions Inc.