Log in

View Full Version : Optimization of Script


paul_slocum
7th February 2009, 21:45
I have a complex script that I want to optimize because it takes about an hour to render on my computer. I'm pasting the script below to see if anyone has any ideas, but I was also wondering if there is any documentation on how AVIsynth works internally and how it uses memory so I can optimize scripts myself. It appears that the layer function is what's really eating up most of the processing time. Thanks!

-paul

# Set up video/audio sources
clip1=DirectShowSource("crazyguy.avi").converttoRGB32
clip1=assumeFPS( clip1, 30)
music=WAVSource("hemthis4.wav")

# Blank clips to be used for spacers
blS = trim(BlankClip(clip1, color=$000000),0,14)
blL = trim(BlankClip(clip1, color=$000000),0,15)

# Video track sequences
t0S = trim(clip1,933,947)
t0L = trim(clip1,933,948)
track0=t0S+
t0L+t0S+t0L+t0L+t0S+t0L+t0L+t0S+t0L+blL+blS+blL+blL+blS+blL+
t0L+t0S+t0L+t0L+t0S+t0L+t0L+t0S+t0L+t0S+t0L+t0L+t0S+blL+blL+
blS+blL+blL+blS+blL+blL+blS+blL+blL+t0S+t0L+t0L+t0S+t0L+t0L+
t0S+t0L+t0S+t0L+t0L+t0S+t0L+t0L+t0S+t0L+t0L+t0S+t0L+t0L+t0S+
t0L+t0L+t0S+t0L+t0L+t0S+t0L+t0L+t0S+t0L+t0S+t0L+t0L+t0S+t0L+
t0L+t0S+t0L+t0L+t0S+t0L+t0L+t0S+t0L+t0L

t1S = trim(clip1,743,757)
t1L = trim(clip1,743,758)
track1=t1S+
t1L+t1S+t1L+t1L+t1S+t1L+t1L+t1S+t1L+t1L+t1S+t1L+t1L+t1S+t1L+
blL+blS+blL+blL+blS+blL+blL+blS+blL+blS+blL+blL+blS+t1L+t1L+
t1S+t1L+t1L+t1S+t1L+t1L+t1S+t1L+t1L+blS+blL+blL+blS+blL+blL+
blS+blL+blS+blL+blL+blS+blL+t1L+t1S+t1L+t1L+t1S+t1L+t1L+t1S+
t1L+t1L+t1S+t1L+t1L+t1S+t1L+t1L+t1S+t1L+t1S+t1L+t1L+t1S+t1L+
t1L+t1S+t1L+t1L+t1S+t1L+t1L+t1S+t1L+t1L

t2S = trim(clip1,262,276)
t2L = trim(clip1,262,277)
track2=t2S+
t2L+t2S+t2L+t2L+t2S+t2L+t2L+t2S+t2L+t2L+t2S+t2L+t2L+t2S+t2L+
t2L+t2S+t2L+t2L+t2S+t2L+t2L+t2S+t2L+t2S+t2L+t2L+t2S+t2L+t2L+
t2S+t2L+t2L+t2S+t2L+t2L+t2S+t2L+t2L+t2S+t2L+t2L+t2S+t2L+t2L+
t2S+t2L+t2S+t2L+t2L+t2S+t2L+t2L+t2S+t2L+t2L+t2S+blL+blL+blS+
blL+blL+blS+blL+blL+blS+blL+blL+blS+blL+blS+blL+blL+blS+blL+
blL+blS+t2L+t2L+t2S+t2L+t2L+t2S+t2L+t2L

t3S = trim(clip1,529,543)
t3L = trim(clip1,529,544)
track3=t3S+
t3L+t3S+t3L+t3L+t3S+t3L+t3L+t3S+t3L+t3L+t3S+t3L+t3L+t3S+t3L+
t3L+t3S+t3L+t3L+t3S+t3L+t3L+t3S+t3L+t3S+t3L+t3L+t3S+t3L+t3L+
t3S+t3L+t3L+t3S+t3L+t3L+t3S+t3L+t3L+t3S+t3L+t3L+t3S+t3L+t3L+
t3S+t3L+t3S+t3L+t3L+t3S+t3L+t3L+t3S+t3L+t3L+t3S+blL+blL+blS+
blL+t3L+t3S+t3L+t3L+t3S+t3L+t3L+blS+blL+blS+blL+blL+blS+blL+
blL+blS+t3L+t3L+t3S+t3L+t3L+t3S+t3L+t3L

t4S = trim(clip1,700,714)
t4L = trim(clip1,700,715)
track4=t4S+
t4L+t4S+t4L+t4L+t4S+t4L+t4L+t4S+t4L+t4L+t4S+t4L+t4L+t4S+t4L+
t4L+t4S+t4L+t4L+t4S+t4L+t4L+t4S+t4L+t4S+t4L+t4L+t4S+t4L+t4L+
t4S+t4L+t4L+t4S+t4L+t4L+t4S+t4L+t4L+t4S+t4L+t4L+t4S+t4L+t4L+
t4S+t4L+t4S+t4L+t4L+t4S+t4L+t4L+t4S+t4L+t4L+t4S+t4L+t4L+t4S+
t4L+blL+blS+blL+blL+blS+blL+blL+t4S+t4L+t4S+t4L+t4L+t4S+t4L+
t4L+t4S+t4L+t4L+t4S+t4L+t4L+t4S+t4L+t4L

t8S = trim(clip1,508,522)
t8L = trim(clip1,508,523)
track8=t8S+
t8L+t8S+t8L+t8L+t8S+t8L+t8L+t8S+t8L+t8L+t8S+t8L+t8L+t8S+t8L+
blL+blS+blL+blL+blS+t8L+t8L+t8S+blL+blS+blL+blL+blS+blL+blL+
blS+t8L+t8L+t8S+t8L+blL+blS+blL+blL+blS+blL+blL+blS+blL+blL+
t8S+t8L+t8S+blL+blL+blS+blL+t8L+t8S+t8L+t8L+t8S+t8L+t8L+t8S+
t8L+t8L+t8S+t8L+t8L+t8S+t8L+t8L+t8S+t8L+t8S+t8L+t8L+t8S+t8L+
t8L+t8S+t8L+t8L+t8S+t8L+t8L+t8S+t8L+t8L

t9S = trim(clip1,752,766)
t9L = trim(clip1,752,767)
track9=t9S+
t9L+t9S+t9L+t9L+t9S+t9L+t9L+t9S+t9L+t9L+t9S+t9L+t9L+t9S+t9L+
t9L+t9S+t9L+t9L+t9S+blL+blL+blS+t9L+t9S+t9L+t9L+t9S+t9L+t9L+
t9S+blL+blL+blS+blL+t9L+t9S+t9L+t9L+t9S+t9L+t9L+t9S+t9L+t9L+
blS+blL+blS+t9L+t9L+t9S+t9L+t9L+t9S+t9L+t9L+t9S+t9L+t9L+t9S+
t9L+t9L+t9S+t9L+t9L+t9S+t9L+t9L+t9S+t9L+t9S+t9L+t9L+t9S+t9L+
t9L+t9S+t9L+t9L+t9S+t9L+t9L+t9S+t9L+t9L

t14S = trim(clip1,256,270)
t14L = trim(clip1,256,271)
track14=blS+
blL+blS+blL+blL+blS+blL+blL+blS+blL+blL+blS+blL+blL+blS+blL+
blL+blS+blL+blL+blS+blL+blL+blS+blL+blS+blL+blL+blS+blL+blL+
blS+blL+blL+blS+blL+blL+blS+blL+blL+blS+blL+blL+blS+blL+blL+
blS+blL+blS+blL+blL+blS+blL+blL+blS+blL+blL+blS+blL+blL+blS+
blL+blL+blS+blL+blL+blS+blL+blL+blS+blL+blS+blL+blL+blS+blL+
blL+blS+blL+blL+blS+blL+blL+blS+blL+blL

t15S = trim(clip1,256,270)
t15L = trim(clip1,256,271)
track15=blS+
blL+blS+blL+blL+blS+blL+blL+blS+blL+blL+blS+blL+blL+blS+blL+
blL+blS+blL+blL+blS+blL+blL+blS+blL+blS+blL+blL+blS+blL+blL+
blS+blL+blL+blS+blL+blL+blS+blL+blL+blS+blL+blL+blS+blL+blL+
blS+blL+blS+blL+blL+blS+blL+blL+blS+blL+blL+blS+blL+blL+blS+
blL+blL+blS+blL+blL+blS+blL+blL+blS+blL+blS+blL+blL+blS+blL+
blL+blS+blL+blL+blS+blL+blL+blS+blL+blL

t16S = trim(clip1,119,133)
t16L = trim(clip1,119,134)
track16=t16S+
t16L+t16S+t16L+t16L+t16S+t16L+t16L+t16S+t16L+t16L+t16S+t16L+t16L+t16S+t16L+
t16L+t16S+t16L+t16L+t16S+t16L+t16L+t16S+t16L+t16S+t16L+t16L+t16S+t16L+t16L+
t16S+t16L+t16L+t16S+t16L+t16L+t16S+t16L+t16L+blS+blL+blL+blS+blL+blL+
blS+blL+blS+blL+blL+blS+blL+blL+blS+blL+blL+blS+blL+blL+blS+
blL+blL+blS+blL+blL+blS+blL+blL+blS+blL+blS+blL+blL+t16S+t16L+
t16L+t16S+blL+blL+blS+blL+blL+blS+blL+blL

t17S = trim(clip1,711,725)
t17L = trim(clip1,711,726)
track17=t17S+
t17L+t17S+t17L+t17L+t17S+t17L+t17L+t17S+t17L+t17L+t17S+t17L+t17L+t17S+t17L+
t17L+t17S+t17L+t17L+blS+blL+blL+blS+blL+blS+blL+blL+blS+t17L+t17L+
t17S+blL+blL+blS+blL+blL+blS+blL+blL+t17S+t17L+t17L+t17S+t17L+t17L+
t17S+t17L+t17S+t17L+t17L+t17S+t17L+t17L+t17S+t17L+t17L+t17S+t17L+t17L+t17S+
t17L+t17L+t17S+t17L+t17L+t17S+t17L+t17L+t17S+t17L+t17S+t17L+t17L+t17S+t17L+
t17L+t17S+t17L+t17L+t17S+t17L+t17L+t17S+t17L+t17L

# Overlay all video tracks into final clip
final = BlankClip(track1, color=$000000)
final = Layer(final, track0, "fast")
final = Layer(final, track1, "fast")
final = Layer(final, track2, "fast")
final = Layer(final, track3, "fast")
final = Layer(final, track4, "fast")
final = Layer(final, track8, "fast")
final = Layer(final, track9, "fast")
final = Layer(final, track14, "fast")
final = Layer(final, track15, "fast")
final = Layer(final, track16, "fast")
final = Layer(final, track17, "fast")

# Dub music and return final clip
final = audiodub( final, music )
return(final)

Sagekilla
7th February 2009, 23:38
Well it's no wonder it takes forever to render. Your splicing pattern is utterly ridiculous and no "patterns" that you can exploit.. You can try using loop() to reduce something like t0s+t0l+t0s+t0l to (t0s+t0l).loop(2)

IanB
7th February 2009, 23:46
Yes the 11 consecutive Layer operations are your problem. All the trims and joins are zero processing cost, they are built in the compile phase only.

You could replace the Layer(... "Fast") with Merge(). Merge() works with all pixel formats so you could remove the ConvertToRGB32() and work in the faster YV12 colour space.

Do you really want ((((((((((t0/2+t1)/2+t2)/2+t3)/2+t4)/2+t8)/2+t9)/2+t14)/2+t15)/2+t16)/2+t17)/2 It seems a little unusual to me.

I remember someone writing a multiple mix plugin for doing the weighted addition of many clips. Maybe that would work better. :search:

paul_slocum
7th February 2009, 23:49
The splicing doesn't seem to be an issue. I can render a single "track" in realtime. It's the layering that seems to be the issue. It seems odd to me that the layering is so slow.

The only solution I've thought of so far is to layer first and then splice. I can layer one segment of each permutation of layering, and then splice all the resulting layered loops at the end. It's just a bit more difficult to write it to work this way.

The script is generated by a C++ program that reads song data from music software that I wrote, and then converts it into an AVIsynth script with loops that are perfectly in sync with the song's tempo. That's why the splicing seems kinda random.

(update)

Ian, I wrote this before I saw your post. Thanks, I'll try the merge function and look for the multi-mix function.

(update2)

Merge doesn't seem to shave off any significant amount of time. Guess it's either write the permutation version, or just deal with the long render time. Also tried the "average" plugin which merges multiple videos, but it wasn't much faster either.

Leak
8th February 2009, 11:24
You could try splitting your file into several files (one per track) then using VirtualDub or avs2avi to encode them into AVI files using some lossless codec and using those, or you could try opening the separate AVS files directly with AVISource...

I bet that's going to speed things up a bit...

(Just curious - how long is your audio file, anyway?)

np: Fennesz - Vacuum (Black Sea)

IanB
8th February 2009, 12:27
Do you really want ((((((((((t0/2+t1)/2+t2)/2+t3)/2+t4)/2+t8)/2+t9)/2+t14)/2+t15)/2+t16)/2+t17)/2Layer(... "Fast") runs at a reasonable speed, Merge(... 0.5) is only marginally faster but it supports all colour spaces.

YV12 should be about 2.6 times faster than RGB32, just because of the reduced data size. YUY2 would be about 2 times.

But think a little harder about the 11 consecutive Layer operations, were dealing with 8 bit data, I doubt if data before track4 (1/128th) or track8 (1/64th) will be even visible in the result.

paul_slocum
8th February 2009, 23:18
Ian: You're right about the odd x/2/2/2 merging technique and that some of the layers may barely be visible. I'm going to try to use a method that eliminates that problem.

Leak: The test song I'm using is just an excerpt and it's only 45 seconds. Splitting it up is a good idea, although for my purposes it'll probably work better to split it up into temporal sections rather than sections of tracks. And if you like Fenessz you might actually be into what I'm doing.

I decided I'm going to rewrite the software to make it do the merge first. I've thought about this method a lot, and I think it will be dramatically more efficient.

Thanks for all the help guys, really appreciate it. I've been doing this stuff on my own for a while -- good to know there's a friendly and active AVIsynth community. I'll post my efficiency results and some video output examples soon.

*.mp4 guy
8th February 2009, 23:53
Just a quick thought, looking at all of those calls, this might be IO bound, not cpu bound. I'm probably completely wrong, but its worth checking.

IanB
9th February 2009, 00:57
@paul_slocum,

What are you really trying to achieve with each of your tracks? Perhaps a still of an output frame might help explain things.

Maybe this is the style you are looking for ...
T01=Merge(track0, track1)
T23=Merge(track2, track3)
T0123=Merge(T01, T23)

T48=Merge(track4, track8)
T914=Merge(track9, track14)
T48914=Merge(T48, T914)

T012348914=Merge(T0123, T48913)

T1516=Merge(track15, track16)
T151617=Merge(T11516, track17, 0.3333)

final=Merge(T012348913, T151617, 0.1875)
...

paul_slocum
9th February 2009, 19:21
Okay, I rewrote the software and now it generates a much more efficient script that effectively does the same thing. It renders each segment of layering first, then builds the sequence. It did cut the render time in half, but it should be doing a lot better than that.

The problem is that AVIsynth seems to be re-rendering the "average" filters each time I use them in the final sequence. I've tried making a version of the script below that only includes each "permutation" of the averaging in the final output once, and it renders in about 8 minutes. But when I make the final sequence longer (like it is in the script below) even though it just reuses the same clips, it takes about 35 minutes. It seems to me that AVIsynth should remember that it already rendered that stuff when it encounters the same clip again and just copy the results of the first render. Is there any way to make it do that? (...other than rendering the averages to an AVI and then using a second script to sequence them)

# Set up video/audio sources
clip1=DirectShowSource("crazyguy.avi").converttoRGB32
clip1=assumeFPS( clip1, 30)
music=WAVSource("hemthis4.wav")

# Video tracks
t0 = trim(clip1,218,233)
t1 = trim(clip1,553,568)
t2 = trim(clip1,781,796)
t3 = trim(clip1,797,812)
t4 = trim(clip1,870,885)
t8 = trim(clip1,316,331)
t9 = trim(clip1,997,1012)
t14 = trim(clip1,111,126)
t15 = trim(clip1,257,272)
t16 = trim(clip1,333,348)
t17 = trim(clip1,951,966)

# Merge permutations
pL0 = average(t14, 0.5, t15, 0.5)
pS0 = trim( pL0, 0, 14)
pL1 = average(t0, 0.333333, t14, 0.333333, t15, 0.333333)
pS1 = trim( pL1, 0, 14)
pL2 = average(t1, 0.25, t8, 0.25, t14, 0.25, t15, 0.25)
pS2 = trim( pL2, 0, 14)
pL3 = average(t1, 0.2, t8, 0.2, t14, 0.2, t15, 0.2, t17, 0.2)
pS3 = trim( pL3, 0, 14)
pL4 = average(t1, 0.2, t9, 0.2, t14, 0.2, t15, 0.2, t17, 0.2)
pS4 = trim( pL4, 0, 14)
pL5 = average(t0, 0.25, t8, 0.25, t14, 0.25, t15, 0.25)
pS5 = trim( pL5, 0, 14)
pL6 = average(t0, 0.2, t9, 0.2, t14, 0.2, t15, 0.2, t17, 0.2)
pS6 = trim( pL6, 0, 14)
pL7 = average(t0, 0.2, t8, 0.2, t14, 0.2, t15, 0.2, t17, 0.2)
pS7 = trim( pL7, 0, 14)
pL8 = average(t1, 0.2, t8, 0.2, t14, 0.2, t15, 0.2, t16, 0.2)
pS8 = trim( pL8, 0, 14)
pL9 = average(t1, 0.2, t9, 0.2, t14, 0.2, t15, 0.2, t16, 0.2)
pS9 = trim( pL9, 0, 14)
pL10 = average(t14, 0.333333, t15, 0.333333, t16, 0.333333)
pS10 = trim( pL10, 0, 14)
pL11 = average(t2, 0.2, t3, 0.2, t14, 0.2, t15, 0.2, t16, 0.2)
pS11 = trim( pL11, 0, 14)
pL12 = average(t2, 0.2, t4, 0.2, t14, 0.2, t15, 0.2, t16, 0.2)
pS12 = trim( pL12, 0, 14)
pL13 = average(t2, 0.25, t3, 0.25, t14, 0.25, t15, 0.25)
pS13 = trim( pL13, 0, 14)

# Produce final sequence
final =pS0+pL0+pS0+pL0+pL0+pS0+pL0+pL0+pS0+pL0+pL1+pS1+pL1+pL1+pS1+pL1+\
pL2+pS2+pL2+pL2+pS3+pL4+pL4+pS4+pL3+pS3+pL3+pL3+pS3+pL5+pL5+\
pS5+pL6+pL6+pS6+pL6+pL7+pS7+pL7+pL7+pS8+pL8+pL8+pS8+pL8+pL8+\
pS9+pL9+pS9+pL8+pL8+pS8+pL8+pL10+pS10+pL10+pL10+pS10+pL11+pL11+pS11+\
pL11+pL12+pS12+pL12+pL12+pS12+pL12+pL12+pS11+pL11+pS11+pL11+pL11+pS13+pL13+\
pL13+pS13+pL10+pL10+pS10+pL10+pL10+pS10+pL10+pL10

# Dub music and return final clip
final = audiodub( final, music )
return(final)

IanB
9th February 2009, 22:06
Ah you found the Average plugin ;)

What format is crazyguy.avi, DirectShowSource does not seek well at the best of times and is truly abysmal for keyframe sparse formats e.g. Xvid with a keyframe every 300.

You can try to mitigate this by ramping up the Avisynth cache with SetMemoryMax(1024)

If that still does not help transcode to an all keyframe format like Huffyuv and use AviSource().

Also the Average plugin by mg262 reverts to slow C code for odd numbers of clips. It may be worth duplicating 1 clip just to get an even clip count to allow the use of the iSSE code. i.e.pL12 = average(t2, 0.1, t2, 0.1, t4, 0.2, t14, 0.2, t15, 0.2, t16, 0.2)

Also loose the ConvertToRGB32()! For this application a YUV format will be 2 times as fast. If you do need RGB output do it last.

paul_slocum
10th February 2009, 06:59
I did everything you said, and it now renders in about 2 minutes, which is great. It will probably take quite a bit longer with a full song, but it's still very reasonable. Converting to Huffyuv made the biggest difference.

It was an XVID file, and I'm sorry if that was a dumb mistake to be using MPEG4. I'm not sure what the keyframes are on the video, it was from my digital camera. I just figured that it'd be okay since it can render the XVID file in realtime and probably caches everything it renders. But I don't really know much about the internals of the software. I probably should just look at the source code sometime.

Thanks so much for your help. I'll post some examples of the output from all this soon.

Leak
10th February 2009, 15:39
I just figured that it'd be okay since it can render the XVID file in realtime and probably caches everything it renders.
It probably worked in realtime with one of your "tracks" alone because all decoded frames of the XviD file fit into AviSynth's cache, but all together probably didn't.

And having to decode at the last keyframe repeatedly can't be good for performance - hence the speedup with HuffYUV where every frame is a keyframe...

paul_slocum
10th February 2009, 20:17
Okay, here's an example of the output. This is totally rough and unfinished, just an example of the kind of thing you can do with this software.

(~11MB)
http://www.qotile.net/temp/vconv_demo.avi