View Full Version : Dolby Pro Logic II encoding test (90° vs 180° phase shift)
3dsnar
7th June 2006, 08:33
Here are two clips,
http://forum.videohelp.com/images/guides/p1514165/180vs90.7z
And the reference (original 6 channel input) clip.
http://forum.videohelp.com/images/guides/p1514168/fifthelem_6chnl.7z
Each downmixed with:
Lt = FL{0°} + 0.7071 C{0°} + 0.7071 LFE{0°} + 0.866 SL{180°} + 0.5 SR{180°}
Rt = FR{0°} + 0.7071 C{0°} + 0.7071 LFE{0°} + 0.5 SL{0°} + 0.866 SR{0°}
and
Lt = FL{0°} + 0.7071 C{0°} + 0.7071 LFE{0°} + 0.866 SL{90°} + 0.5 SR{-90°}
Rt = FR{0°} + 0.7071 C{0°} + 0.7071 LFE{0°} + 0.5 SL{-90°} + 0.866 SR{+90°}
Please test them with your dolby certified decoders
and decide which ones sounds better (eg. A vs B version) with reference to the original sound.
The poll question is related to fifthelem_A.mp3 and fifthelem_B.mp3,
while speech_A.mp3 and speech_B.mp3 should sound identical
(to see that both downmixing methods are compatible with DPLII decoders)
scharfis_brain
7th June 2006, 14:17
I cannot tell for sure (tendency goes to clip A) which is the 180 or 90 degree clip because I curently have no possibility to play back the AC3 in direct 5.1 off my PC.
but anyways this is NOT a fair comparision.
fair would have been:
Each downmixed with:
Lt = FL{0°} + 0.7071 C{0°} + 0.7071 LFE{0°} + 0.866 SL{180°} + 0.5 SR{180°}
Rt = FR{0°} + 0.7071 C{0°} + 0.7071 LFE{0°} + 0.5 SL{0°} + 0.866 SR{0°}
and
Lt = FL{0°} + 0.7071 C{0°} + 0.7071 LFE{0°} + 0.866 SL{+90°} + 0.5 SR{+90°}
Rt = FR{0°} + 0.7071 C{0°} + 0.7071 LFE{0°} + 0.5 SL{-90°} + 0.866 SR{-90°}
or Rockarias style:
Lt = FL{0°} + 0.7071 C{0°} + 0.7071 LFE{0°} + 0.866 SL{0°} + 0.5 SR{180°}
Rt = FR{0°} + 0.7071 C{0°} + 0.7071 LFE{0°} + 0.5 SL{180°} + 0.866 SR{0°}
and
Lt = FL{0°} + 0.7071 C{0°} + 0.7071 LFE{0°} + 0.866 SL{+90°} + 0.5 SR{-90°}
Rt = FR{0°} + 0.7071 C{0°} + 0.7071 LFE{0°} + 0.5 SL{-90°} + 0.866 SR{+90°}
but you are testing two types of four concurring types.
the traditional mixing with 180°
and rockarias inverse mixing with 90°
And I think this is perfectly audible when you can hear "weapons loaded" at the very beginning of the sample.
sample A lets it sound from the center.
sample B lets it sound from the surround.
But this is not due to the different phases but more due to the reason that one of the surround channels (in relation to each other) become inverted with your 90° matrix. (IMO)
(this is the thing I always tried to explain as center surround issue to Rockaria. But it seems to be a slightly different effect here)
so would you redo the samples, please?
Rockaria
7th June 2006, 15:02
http://forum.doom9.org/showpost.php?p=783211&postcount=42
IMO this matrix is useless.
Just imagine the MonoSurround-condition: Rear left and rear right are carrying the same signal.
With this matrix you'll succesfully eliminate the mono surround out of the downmix.
This means to me, that
Dolby Pro Logic Left Right Center Rear Left Rear Right
Left Total 1.000 0.000 0.7071 0.866 -0.5
Right Total 0.000 1.000 0.7071 -0.5 0.866
is a derivation from the faulty matrix above and should not be used.
You may experience wider sounding surround channels, because every middle information (mono information) is weakened in the surround downmix.
or Rockarias style:
Lt = FL{0°} + 0.7071 C{0°} + 0.7071 LFE{0°} + 0.866 SL{0°} + 0.5 SR{180°}
Rt = FR{0°} + 0.7071 C{0°} + 0.7071 LFE{0°} + 0.5 SL{180°} + 0.866 SR{0°}
and
Lt = FL{0°} + 0.7071 C{0°} + 0.7071 LFE{0°} + 0.866 SL{+90°} + 0.5 SR{-90°}
Rt = FR{0°} + 0.7071 C{0°} + 0.7071 LFE{0°} + 0.5 SL{-90°} + 0.866 SR{+90°}
but you are testing two types of four concurring types.
the traditional mixing with 180°
and rockarias inverse mixing with 90°
(this is the thing I always tried to explain as center surround issue to Rockaria. But it seems to be a slightly different effect here)
I found, that Rockaria's personal matrix has problems with centered surround effects that are meant to be reproduced by both surround speakers (DPLII) or the surround-back speaker (DPLIIx).
Inverting the phase of one of the Surround channels like this modified matrix does can enhance the perceived width and separation of both channel because it is simply something like crosstalk reduction. But At the cost of centered surround sounds. That's why I prefer the unaltered matrix.
.....
....
@scharfis, do you see your posts are 'arbitrary itself' ?
Would you stop posting against anybody with no basis & no understanding at all? That does not help anything.
traditional mixing with 180° :: what the traditional means? your myths?, your Gods' words?
inverse mixing with 90° : youself admitted you have no clear picture of the DPL II model already?
If you continue posting like this, I will regard them as personal attacks, only to depreciate any resonable conversations.
scharfis_brain
7th June 2006, 15:23
I don't see a contradiction to myself here. I think that these quotes still enhance the statements of my first post in this thread.
Also it is not a personal attack as I said many times in past, too.
If you are interpreting my answers as personal attacks it is your problem. Not mine!
traditional mixing with 180° :: what the traditional means? your myths?, your Gods' words?The wikipedia style of matrix. The Besweet style of matrix. The AC3filter style of matrix. The ffdshow style of matrix. Is this enough traditional style?
inverse mixing with 90° : youself admitted you have no clear picture of the DPL II model already?
Of course not, as I already did some manual mixes myself, build some active Dolby surround mixing ciruits etc.
If you should find irony in the last sentence, keep it for yourself. Thanks.
Actually I think you are the person that is not able or willing to do a objective conversation without personal attacks.
I never did attack you. But you do!
tebasuna51
7th June 2006, 19:03
Test with a SONY STR-DE495 receiver, DPL II Movie mode.
There are a clear difference between A and B in "weapons loaded" (?) and "Yes, Sir" (3 sec.).
There are presents in FL, FR, BL, BR original audio, and not at Center channel.
In A "Yes, Sir" is basically only at Center channel.
In B "Yes, Sir" is present in all five channels, then is not perfect but better than A.
Then my vote for B.
I tried also a C option (matrix 3 from original thread) with indistinguishable differences with B (at least for my ears).
Rockaria
7th June 2006, 19:10
@scharfis, I am pretty much impressed by your convenient logic adaptible to any situations.:(
In my previous quotes :
The contradiction in the 1st and 3rd quoted messages from you looks to me clearly an irony. (and the 2nd one just a justification)
The 'inverse 90 degree' is actually the +-90 degree phase shifts, to be correct, based on my best reasoning & understaning which describes the DPL II model from Dolby's unclear documents.
Look at the 1st quote.
Now do you distinguish which is useless? : my useless model vs your invaluable objections
scharfis, I just want you not to play with my id at all.:cool:
@3dsnar,
The two fifth_elm clips showed a distinct difference in the first part, the _b clip sounded some wider fronts
But I have no idea which is with better seperation fidelity without the original 6ch clip(in any form such as mp4) as I mentioned in other thread.
[edit] Yeah, I see tebasuna has the same opinion.
Now with the original 6ch AC3, I see the fifth_elm_B has the better seperation fidelity. voted.
3dsnar
8th June 2006, 08:41
OK, thanks all of you for votes.
BTW. The original signal is provided too.
Yes, B is 180 deg phase shift (the simpler approach),
A is 90 deg. phase shift.
The sign variations are not important, because they result
in phase invertion in the decoded signal and do not affect the separation (as Tebasuna already noticed). Therefore A and B are enough to distinguish between 90 and 180 phase shifts.
I agree with Scharfis that the 90 deg phase shift should have been prepared with the same sign style, to be fully consistent with the 180 deg. shift. Maybe next time ;)
(early) conclusions.
1) There is a noticable difference between the 90 deg phase shift and 180 deg phase shift
2) 90 deg seem to produce significantly worst results, thus probably the DPLII downmixing equation is based on 180 deg. shifts (simple sign change).
---
BTW. Please test Aud-X DSfilter DPLII decoder (or rather DPLII decoder simulation),
by sending the output as AC3 through SPDIF to your HT amps.
I am curious of your opinions and especially your thoughts regarding its quality vs certified dolby decoders quality
:thanks: 3d
Rockaria
8th June 2006, 09:39
OK, thanks all of you for votes.
BTW. The original signal is provided too.
But after my post in other thread, without any notice or acknowledgement.
http://forum.doom9.org/showpost.php?p=837679&postcount=63
You might want to add two more models to fully reflect the discussions in the related threads:
Lt = FL{0°} + 0.7071 C{0°} + 0.7071 LFE{0°} + 0.866 BL{-90°} + 0.5 BR{-90°}
Rt = FR{0°} + 0.7071 C{0°} + 0.7071 LFE{0°} + 0.5 BL{90°} + 0.866 BR{90°}
and
Lt = FL{0°} + 0.7071 C{0°} + 0.7071 LFE{0°} + 0.866 SL{180°} + 0.5 SR{0°}
Rt = FR{0°} + 0.7071 C{0°} + 0.7071 LFE{0°} + 0.5 SL{0°} + 0.866 SR{180°}
...
You might also want to test with original 6ch(mp4 format) vs dpl II mix(music + speaker test clip), to compare reasonably.
The FFDShow can switch between PCM/AC3 in digital out mode on the fly making it easy to compare.
As I said many times, the speaker test clip is most generous on any models(more than 95% of the seperation quality).
Also the matrices values(Ls1,Ls2,Rs1,Rs2....) might affect the seperation quality by the phase shift degree change.
Comparing the Wikipedia matrix with the current one would be reasonable(also making it reasonable using the variable reference than the values).
So far, the fifthelem_B showed noticeable better seperation, but the test is not setup to compare the seperation fidelity(original vs DPL II).
Yes, A is 180 deg phase shift (the simpler approach),
B is 90 deg. phase shift.
Can you provide the proof in a source format(no dll)?
The sign variations are not important, because they result
in phase invertion in the decoded signal and do not affect the separation (as Tebasuna already noticed). Therefore A and B are enough to distinguish between 90 and 180 phase shifts.
Again, it's just your very dangerouse assumption.
You can test it with FFDShow which can adjust the matrix value with the sign.
It shows no difference only with the simplest speaker test file.
Also prove exactly which Tebasuna's observations concure your assumptions.
I agree with Scharfis that the 90 deg phase shift should have been prepared with the same sign style, to be fully consistent with the 180 deg. shift. Maybe next time
also check my quote above. indeed, the polls feels like something public.
(early) conclusions.
1) There is a noticable difference between the 90 deg phase shift and 180 deg phase shift : agreed.
2) 90 deg seem to produce significantly worst results, thus probably the DPLII downmixing equation is based on 180 deg. shifts (simple sign change).
I agree again we are using totally differernt languages. Apparently it seems what you wanna see.
---
BTW. Please test Aud-X DSfilter DPLII decoder (or rather DPLII decoder simulation),
by sending the output as AC3 through SPDIF to your HT amps.
I am curious of your opinions and especially your thoughts regarding its quality vs certified dolby decoders quality My independant Ad. :
Sorry, unfortunately, I am very much satisfied with the FFDShow with probably 1000% of proven better features and stability.
I also believe they will provide the DPL II encoding with 90 deg. phase shift very soon(with current adjustable matrix).
Rockaria
8th June 2006, 09:51
Was it a blind test to fool people by an anonymous person?
Here are two clips,
http://forum.videohelp.com/images/gu...165/180vs90.7z
And the reference (original 6 channel input) clip.
http://forum.videohelp.com/images/gu...helem_6chnl.7z
Each downmixed with:
Lt = FL{0°} + 0.7071 C{0°} + 0.7071 LFE{0°} + 0.866 SL{180°} + 0.5 SR{180°}
Rt = FR{0°} + 0.7071 C{0°} + 0.7071 LFE{0°} + 0.5 SL{0°} + 0.866 SR{0°}
and
Lt = FL{0°} + 0.7071 C{0°} + 0.7071 LFE{0°} + 0.866 SL{90°} + 0.5 SR{-90°}
Rt = FR{0°} + 0.7071 C{0°} + 0.7071 LFE{0°} + 0.5 SL{-90°} + 0.866 SR{+90°}
Yes, A is 180 deg phase shift (the simpler approach),
B is 90 deg. phase shift.
Yes, B is 180 deg phase shift (the simpler approach),
A is 90 deg. phase shift.
You definitely need to provide the encoding source of your test poll.
3dsnar
8th June 2006, 10:18
I have informed you of the input included (after you suggested so).
You should read my posts more carefully.
-----------
I am asking to verify my DPLII simmulation decoding algorithm,
not the entire DSfilter.
-----------
Here is the matlab code that I used for downmixing.
function [y, y2]=DPLIIdownmix(s)
FL=s(:,1);
FR=s(:,2);
C=s(:,3);
LFE=s(:,4);
SL=s(:,5);
SR=s(:,6);
y=zeros(length(FL),2)
y2=zeros(length(FL),2);
wgF=1;
wgC=sqrt(0.5);
wgA=sqrt(0.75);
wgB=sqrt(0.25);
Nrm=wgF + wgC + wgC + wgA + wgB;
wgF=wgF/Nrm;
wgC=wgC/Nrm;
wgA=wgA/Nrm;
wgB=wgB/Nrm;
%180 deg phase shifts
y(:,1) = FL*wgF + C*wgC + LFE*wgC - wgA*SL - wgB*SR;
y(:,2) = FR*wgF + C*wgC + LFE*wgC + wgB*SL + wgA*SR;
%90 deg phase shifts
y2(:,1) = FL*wgF + C*wgC + LFE*wgC + wgA*imag(hilbert(SL)) - wgB*imag(hilbert(SR));
y2(:,2) = FR*wgF + C*wgC + LFE*wgC - wgB*imag(hilbert(SL)) + wgA*imag(hilbert(SR));
Rockaria
8th June 2006, 10:31
I have informed you of the input included (after you suggested so).
You should read my posts more carefully.
Now I see at the last part....
So you feel OK performing the blind test poll without any prior permission?
Anyway, I want to believe it can be forgiven if the test routine proves used correctly.
Anybody know where I can get the matlab?
3dsnar
8th June 2006, 10:55
So you feel OK performing the blind test poll without any prior permission?
Yes :D
tebasuna51
8th June 2006, 11:00
BTW. Please test Aud-X DSfilter DPLII decoder (or rather DPLII decoder simulation),
by sending the output as AC3 through SPDIF to your HT amps.
I am curious of your opinions and especially your thoughts regarding its quality vs certified dolby decoders quality
Sorry, I have attached my PC to a old amp with only stereo input (with DPL I capability).
To test 5.1 or DPL II I need to burn a CD and go to other room with a DVD/DivX/mp3 player attached to the SONY receiver.
I tried, to make some kind of test, use your Aud-X DSF in GraphEdit but don't accept WAVDEST or DUMP output. Only accept DirectSound output (like you say in other thread).
For my configuration your Aud-X DSF is unusable.
Rockaria
8th June 2006, 11:17
@3dsnar, you seem to be very confident about ...what?
Anyway, I myself won't engage in such flip-flop games any more.
http://en.wikipedia.org/wiki/MATLAB : something that costs money for what I do not owe.
http://en.wikipedia.org/wiki/Hilbert_transform : seems to be the logic for transforming a square wave form to a strange curve with 90 deg. phase shifts
Possibly we can find the corresponding function in any forms of avisynth/sox plugin..
It's gonna take a long while to prepare the verification test.... oh well, nothing to believe.
It seems it has just lengthened the route further, forcing me to continue my own approach.
At least, I believe I have proved the reference model 4 better than 2 practically, and 1/3 better than 2/4 theoretically. period.
/fooled
3dsnar
8th June 2006, 11:37
You can believe the test, or not.
You can prepare your own downmixes as well.
Freedom of choice.
-------
Hilbert transform shifts all of the signal (represented as Fourier series) sinusoidal components by Pi/2 (90 deg). Hence, such operation may result in waveform change.
Due to long discussions related to creating DPLII downmixes,
and especially 90 dg. vs 180 dg. phase shift myths and speculations, I hope the test has its value for reasonable people.
Rockaria
8th June 2006, 17:32
Freedom of choice.
Oh yeah, my favorest word!:rolleyes: But when abused it becomes what we have just experienced.
So you are proud of playing with the POLL!
Do you even realize you abused the public POLL for your personal purpose?
Also do you understand I tried my best to prevent this kind of misleading, in the first post in other thread about your POLL?
http://forum.doom9.org/showpost.php?p=837679&postcount=63
What I am doubting most is the ethic of an anonymous identity, who wants to move people dramitically with minimum honest efforts.
For an example, your faulty DSFilter unable to connect the out pin to other filters is what I pointed out 5 month ago, still not fixed.
Your gesture is something like identifying the bugs with all the smooth generic promising words, but in fact no honest consideration & fix at all. That's why I can't trust yours at all to be honest, including all the other hidden invaluable un-disclosable know-wheres(no know-hows).
Your 'DPL II decoder' is another good example of the misleading. aud-x? .oh well..I forgot it..
Why don't you just put more time to fix the bugs silently in the one-person-owned-WE .com, for your future, instead of playing the waste games?
Even if I become to verify by myself your doubtful process and result, there's no doubt you just(afterall) created the minimum result(who knows if it's even intented mystypings) that you wanna see in a very deceiving way and you seem to be satisfied you proved something, moved people, realizing no ethical & logical problem at all. Another perfect example of Public Misleading.
I believe you will have to ACTIVELY prove your QUESTIONED HONESTY with proper perfectly OPEN processes and full considerations as I described in other thread, realizing you just created a bigger problem.
Until then, I believe all the other issues are none of your concern.
Thanks anyway for your little event(a true surprise never experienced before) that makes me think of something very different.:eek:
3dsnar
8th June 2006, 18:40
IMHO you've got a serious problem man.
Sorry, but conversation with you is waste of time.
And I am probably not alone with this sad conclusion.
This is the last time I replied to your attacks.
Rockaria
8th June 2006, 18:45
Will see.
Rockaria
9th June 2006, 04:26
Well, I don't have the matlab, nor could find a corresponding plugin for avisynth yet.
Which is why I want to share my idea looking for positive contributions from anybody talented and equipped.
Below is my draft plan to figure out the closest DPL II model for s/w emulations, possibly in a very OPEN, independant, objective, fair and fully reflected but economic way.
Now some more tests are included such as 'the verification of the Hilbert() algo' as well as many other essential considerations to go through the proper approach & procedure.
If the minimum test environment is not setup soon, I am going to add the results in my original thread when available.
1. prepare proper tools & organize the test.
. s/w tools : matlab or avisynth/sox with proper plugin(esp. for 90 deg phase shift)
. resource orgarnizer : an independant user with proper DPL II understanding, s/w tools, h/w equips, resources.
. refined model : the orgarnizer refines the test procedures & models
. prepare, test, report and share the results
2. prepare a 6ch mixed original source for full channel complexity and easy identification
The DPL II encoding with avisynth is listed in the below thread.(for models m12, m22)
http://forum.doom9.org/showpost.php?p=786690&postcount=82
<avs 6ch clips mixing example>
spk=DirectShowSource("SSWAV06.m4a")
muz=DirectShowSource("6chmusic.m4a")
...
mxx=MixAudio(spk, muz, 0.4, 1)
..
dpl2Enc(mxx, 0.7071, 0.7071, 0.866, -0.5, -0.5, 0.866)
...
3. identify the weights of orginal matrices : including the coefficients on rears to be shared on each channel, with no phase shifts(passive)
<mxA> : widely used for s/w encoding
DPL II Lf Rf C Ls Rs
Lt 1.000 0.000 0.7071 0.866 0.5
Rt 0.000 1.000 0.7071 0.5 0.866
<mxB> : Wikipedia one, possibly from Dolby
DPL II Lf Rf C Ls Rs
Lt 1.000 0.000 0.707 0.8165 0.5774
Rt 0.000 1.000 0.707 0.5774 0.8165
4. identify target DPL II models(formula) to compare
<m11>
Lt = mix(Lf.0°, C.0°, Ls1.-90°, Rs2.-90°) == mix(mix(Lf, C), mix(Ls1, Rs2).-90°)
Rt = mix(Rf.0°, C.0°, Ls2.+90°, Rs1.+90°) == mix(mix(Lf, C), mix(Ls2, Rs1).+90°)
<m12> rears(m11).-90°
Lt = mix(Lf.0°, C.0°, Ls1.-180°, Rs2.-180°) == mix(mix(Lf, C),-mix(Ls1, Rs2))
Rt = mix(Rf.0°., C.0°., Ls2.0°, Rs1.0°} == mix(Rf, C, Ls2, Rs1)
<m13> m11.-90° == m11?
Lt = mix(Lf.-90°, C.-90°, Ls1.-180°, Rs2.-180°) == mix(mix(Lf, C).-90°,-mix(Ls1, Rs2))
Rt = mix(Rf.-90°, C.-90°, Ls2.0°, Rs1.0°) == mix(mix(Rf, C).-90°, mix(Ls2, Rs1))
...
<m21>
Lt = mix(Lf.0°, C.0°, Ls1.+90°, Rs2.-90°) == mix(mix(Lf, C), Ls1.+90°, Rs2.-90°)
Rt = mix(Rf.0°, C.0°, Ls2.-90°, Rs1.+90°) == mix(mix(Rf, C), Ls2.-90°, Rs1.+90°)
<m22> rears(m21).-90°
Lt = mix(Lf.0°, C.0°, Ls1.0°, Rs2.-180°) == mix(mix(Lf, C, Ls1),-Rs2)
Rt = mix(Rf.0°, C.0°, Ls2.-180°, Rs1.0°) == mix(mix(Rf, C, Rs1),-Ls2)
...
When -180° = -, +90° = Hilbert(), -90° = -180°.+90° = -.+90°
5. perform the verification of Hilbert() algo if any useful, with m12 * mxA
<m12>
Lt = mix(Lf.0°, C.0°, Ls1.-180°, Rs2.-180°) == mix(mix(Lf, C), -mix(Ls1, Rs2)})
Rt = mix(Rf.0°, C.0°, Ls2.0°, Rs1.0°) == mix(Rf, C, Ls2, Rs1)
<-180° = -(+90°.+90°) == -Hilbert().Hilbert()>
Lt = mix(Lf.0°, C.0°, Ls1.-90°.-90°, Rs2.-90°.-90° == mix(mix(Lf, C), -mix(Ls1, Rs2).+90°.+90°)
Rt = mix(Rf.0°, C.0°, Ls2.0°, Rs1.0°)
6. verify if m13 == m11 with mxA
.discard m13 if identical or refine the 'phase shift' concept
7. verify if different phase shift degrees makes any difference with (m11, m12) * mxA
.discard the worse model
8. prepare & perform the seperation fidelity comparison with remaining candidates combination(matrices, models)
m11 * mxA : <reference1>
m11 * mxB :
m12 * mxA : <reference2>
m13 * mxA :
m21 * mxA : <reference3>
m21 * mxB :
m22 * mxA : <reference4>
9. conclusions
. the test orgarnizer share & report the test results
. users evaluate & verify the results(polls and/or posts)
Thanks,
ursamtl
9th June 2006, 13:39
If you have Plogue Bidule, why don't you just model the Dolby encoding using it? Bidule has a built-in Hilbert Transform filter (although in my experiments, it seems to produce a -90° phase shift, so for the +90° shift, you'll need to invert the signal). Another option for the Hilbert is to use Christian Budde's excellent Phasebug VST plugin (freeware). This gives you the possibility of any phase shift in a 360° circle.
If you don't have Plogue Bidule, try Audomulch.
I haven't done much with DPLII since I've never been impressed with its results, but I can tell you that in my experiments with Ambisonics, the 90° phase shift used in some encoding designs is essential. For example, in the superstereo circuits I've modelled using Plogue Bidule and Phasebug, the 90° phase shift balances the ambience nicely across the surround speakers, but moving it to 0° or 180° forces the ambience all to one surround speaker on another. Again, the application is different from DPL, but I've read a lot of surround sound documentation over the past couple of years and the Hilbert Transform is everywhere!
Finally, I know you guys are passionate about your points of view, but let's relax and enjoy this experimentation.
Regards,
Steve.
scharfis_brain
9th June 2006, 13:59
one question: how was the hilbert transform performed nearly 20 years ago when Dolby introduced their "Dolby Surround" ?
I know, that bandpass filters alter the phase. But the phase deviation is not constant with changes in frequency.
Also I think that 25 years ago computers weren't fast enough to do such calculations in realtime.
Do you have an idea how dolby 'could' have made the frequency independent phase shift?
ursamtl
9th June 2006, 14:15
The Hilbert Transform was done electronically. You can find quite a few circuits around on the net with a bit of judicious googling. :) It's sometimes known as a "dome filter."
scharfis_brain
9th June 2006, 14:43
Oh! That's nice stuff. Chaining some filters to achieve a near to constant phase shift over a defined range of frequencies.
scharfis_brain
9th June 2006, 15:36
I just did a small test with simple 180 degree mixing how the sign affects the surround image:
The Normal mixing:
http://home.arcor.de/scharfis_brain/samples/traditional_matrix_180_degree.png
And the inverted matrix:
http://home.arcor.de/scharfis_brain/samples/inverted_matrix_180_degree.png
samples can be downloaded here:
http://home.arcor.de/scharfis_brain/samples/DPL2-inverse-vs-traditional.rar
3dsnar: now compare my inverted sample to your 90 degree sample. both show the same behavior: "weapons loaded" comes from the center.
And that is why your 90 degree sample also is inverted.
I hope that it now should be clear what I meant all the time that Rockarias personal (inverted) Matrix destroys some of the surround information.
Rockaria
9th June 2006, 16:55
I will also have to agree that 'Rockarias personal (inverted) Matrix' destroyed scharis brain cells totally.
Sorry now brainless, I feel the responsibility. So what can I do for you?
@ursamtl, I appreciate your invaluable unbiased information.
I will look into the mentioned tools and try to find a way in my environment to apply the DPL II 90 degree phase shifts by the spec in the resonable models.
By the way, I seem to have made some annoying useless neighbors unavoidably.
My wife says she's gonna move right now although I want to give them some more opportunities to make themselves good citizons until this weekend.
The apatment is wooden, transfering the earthqakes every once in a while and the alu. shutters shap ultra sonic waves poking awaken my ears, soul and body more than 30 times a day. Their languages are nothing but noises to me, but they seems proud of keeping differnt languages in every different villlage, just 5 miles away.
My best idea atm is turning up the VOA(the foreign language to them) out from the window around the same origin(balcony) to neutralize. And it seems working now. But my wife says it's gonna go back tomorrow like yesterday saying there are such memories that don't last a day.
scharfis_brain
9th June 2006, 17:20
LOL. You are still taking it personal.
You were the person swapping the signs, weren't you?
So I refer this kind of matrix to you.
but anyways. did you listen to the samples?
Rockaria
9th June 2006, 18:38
Just imagine, what happens to your person when you replay whatever sweet music more than 10 times everyday.
Even the most beautiful aria(unfortunately, yours infact was some annoying discrete cracklings) becomes ugly.....at least to me.
You will get disappointed if I say I have already tested the active matrix DPL II encoding with all the possible combinations of the sign and rear coefs with FFDShow, Ac3Filter and Avisynth MixAudio() explained in the original thread.
So those things are no news to me, unfortunately making me feel no need to repeat. But you seem to be still repeating the same old song.
please remind, that arbitrary phase shifting is dependant to frequency.
I think it doesn't make any difference to the decoder whether the encoder shifts the surround +90° and -90° or whether it shifts 0° and 180°.
It is the phase difference (180°) between Lt and Rt that counts here.
where can I find these documents? I was unsuccessful searching them on the dolby.com site.
The contents of the document you linked to me I already knew. I hoped it contained some more specific information about how DPL2 is encoded. But it just was like some marketing talk (for sure, this is not your fault)
slnc = i.amplify(0).invert()
Sure, it's not your fault either(I'm unavoidably repeating). Btw, why the fault directed to Rockaria' among all the reasonable smooth words?.
But if you continue to repeat it blindedly like this, It becomes clearly your faults, even evil. Please maintain the HISTORY to avoid the WW III.
By the way, 'the inverting' reminds me of the reflecting on the mirror in the same phase region, while shifting kinda sliding through the phase regions. So I believe they are different and the 'phase shift' is correct in this case
3dsnar
9th June 2006, 18:51
3dsnar: now compare my inverted sample to your 90 degree sample. both show the same behavior: "weapons loaded" comes from the center.
And that is why your 90 degree sample also is inverted.
I hope that it now should be clear what I meant all the time that Rockarias personal (inverted) Matrix destroys some of the surround information.
Scharfis, I agree with you (partly).
Yes, the inverted version causes wrong position of some sounds in the surround panorama (in case of the 180 phase shifts).
The 90 deg. phase shift in general destroys much more significantly the complete sound (i.e. the inverted 180 deg shift differs from the inverted 90 deg shift. The 90 deg is even worst.).
So far it seems that the the best results are obtained with this equation:
Lt = FL{0°} + 0.7071 C{0°} + 0.7071 LFE{0°} + 0.866 SL{180°} + 0.5 SR{180°}
Rt = FR{0°} + 0.7071 C{0°} + 0.7071 LFE{0°} + 0.5 SL{0°} + 0.866 SR{0°}
Cheers, 3d.
Rockaria
9th June 2006, 19:34
It's gonna be my last wasteful discussion on the unclear theory until I get the resonable fully considered test results to verify the models, unless Rockaria is pointed directly or indirectly hereafter.
4. identify target DPL II models(formula) to compare
<m11>
Lt = mix(Lf.0°, C.0°, Ls1.-90°, Rs2.-90°) == mix(mix(Lf, C), mix(Ls1, Rs2).-90°)
Rt = mix(Rf.0°, C.0°, Ls2.+90°, Rs1.+90°) == mix(mix(Lf, C), mix(Ls2, Rs1).+90°)
<m12> rears(m11).-90°
Lt = mix(Lf.0°, C.0°, Ls1.-180°, Rs2.-180°) == mix(mix(Lf, C),-mix(Ls1, Rs2))
Rt = mix(Rf.0°., C.0°., Ls2.0°, Rs1.0°} == mix(Rf, C, Ls2, Rs1)
<m13> m11.-90° == m11?
Lt = mix(Lf.-90°, C.-90°, Ls1.-180°, Rs2.-180°) == mix(mix(Lf, C).-90°,-mix(Ls1, Rs2))
Rt = mix(Rf.-90°, C.-90°, Ls2.0°, Rs1.0°) == mix(mix(Rf, C).-90°, mix(Ls2, Rs1))
...
<m21>
Lt = mix(Lf.0°, C.0°, Ls1.+90°, Rs2.-90°) == mix(mix(Lf, C), Ls1.+90°, Rs2.-90°)
Rt = mix(Rf.0°, C.0°, Ls2.-90°, Rs1.+90°) == mix(mix(Rf, C), Ls2.-90°, Rs1.+90°)
<m22> rears(m21).-90°
Lt = mix(Lf.0°, C.0°, Ls1.0°, Rs2.-180°) == mix(mix(Lf, C, Ls1),-Rs2)
Rt = mix(Rf.0°, C.0°, Ls2.-180°, Rs1.0°) == mix(mix(Rf, C, Rs1),-Ls2)
...
When -180° = -, +90° = Hilbert(), -90° = -180°.+90° = -.+90°
As it explaines, the Rt in <m12>, has the image mixed with all 4 channels in the same phase region 0, making it the worst choice for the seperations theoretically.(it's a very simple math). It seems concuring my test results with the prevailing <m12> having Lf<Rf, if you read my original thread.
Now the seperation efficiency more clearly looks like m21>m22>m11>m12, which of course needs some simple listening tests.
If I were to explain why the rears have the coefs shared on different channels based on the <m21>, <m22> model is :
. there may exists some overlappings of the images spanning on the different phase regions.
. by comparing the two identical but somewhat overlapped rear coef images, it would be possible to get the closer original channel image enhanced by the servo feedback.
/OPEN minded & mutual respects to the anonymous...
DarkAvenger
9th June 2006, 23:41
Your test is probably flawed if I understand correctly. A properly mastered 5.1 ac3 track already contains 90° shifted surround sounds. That's why a simple downmix should be enough.
ursamtl
9th June 2006, 23:48
Actually this brings up a point I was going to make earlier in this thread. The Dolby documentation talks about + or - 90° phase shifts during the encoding phase only.
scharfis_brain
10th June 2006, 00:09
hmmm. since every DVD-Standalone player has DPL downmixing built in (some also offer DPL2!), what about recording the downmixed & outputted audio of such a device?
Then one can tell for sure, whether ther is +/-90 or 0/180 degree downmixing.
ursamtl
10th June 2006, 00:34
To "downmix" means taking the a 6-channel input and mixing it appropriately for stereo playback on non-multichannel systems. I'm not sure that a downmixed stereo output is matrixed in such a way that it can be subsequently "unmatrixed" into 6 channels so that you could do this comparison. As I understand it, the downmix process is only to present fairly listenable audio on a stereo system.
To be able to determine phase shifts as you suggest would require taking the downmixed stereo file and running it through a decoder and then comparing the phase of the resulting channels to the original 6 individual source channels before the downmix. From what I've read, the matrix process is such that complete recovery of original source channels is not possible.
I wonder if all the attention to whether or not the Hilbert transform /90° phase shift is necessary isn't because it's much easier for casual programmers to implement a 0/180° algorithm. For those who want to try, my understanding is that you need to implement a FIR filter with certain cooefficients. Here's a routine for calculating the cooefficients: http://www.musicdsp.org/archive.php?classid=3#195.
Rockaria
10th June 2006, 08:31
I have identified one(or two) more DPL II candidate from VideoHelp.
http://www.videohelp.com/forum/archive/t294463.html
This model appears to be similar to what ursamtl mentioned in his very helpful posts for the quest.
So I am adding this to the candidates as :
<m31>
Lt = mix(Lf.0°, C.0°, Ls1.-90°, *Rs2.-90°) == mix(mix(Lf, C), mix(Ls1, *Rs2).-90°)
Rt = mix(Rf.0°, C.0°, *Ls2.+90°, Rs1.+90°) == mix(mix(Rf, C), mix(*Ls2, Rs1).+90°)
<m32>
Lt = mix(Lf.0°, C.0°, Ls1.+90°, *Rs2.+90°) == mix(mix(Lf, C), mix(Ls1, *Rs2).+90°)
Rt = mix(Rf.0°, C.0°, *Ls2.-90°, Rs1.-90°) == mix(mix(Rf, C), mix(*Ls2, Rs1).-90°)
When * = invert()
Ok, here's an ANSWER finally...
This won't give you EXACT Dolby Surround Encoding (without a true encoding plugin), but it'll come mighty close.
Put your 6 waves into Audition.
Pan them Like this:
LF --> 100%L
C --> 50%L + 50%R
RF --> 100%R
LS -->(50%L + PhaseInverted 50%R) w/ +90º PhaseShift (if you know how to do that, otherwise ignore the phaseshift)
RS -->(50%R + PhaseInverted 50%L) w/ -90º PhaseShift (same as above)
LFE --> (Already LowPassFiltered up to ~80Hz) ?? 25%L+25%R
Mix down to a stereo WAVE file (making sure not to overload the mixer itself--bring everything down equally if you have to)
Then, you can convert to AC3 2.0, with the "Dolby Surround Indicated" checked.
Good luck,
Scott
(Static) Phase shifting is where the waveform's SINE wavefronts are shifted by a constant Phase Angle. E.G. a textbook "Sine Wave" will become a "Cosine Wave" when shifted 90º. It's fairly easy to do with analog Electronics, but not so easy to do digitally.
Why?
Because a shift by degrees is related in phase, not in time. Shifting a single frequency like above is equivalent to a delay of the same SINE--but only for that frequency. That means that to duplicate using standard delay techniques, you have to have a frequency-dependent time delay.
Thankfully, there is a VST plugin that can do the same thing easily:
http://www.savioursofsoul.de/Christian/VST/PhaseBug.zip
Like I said, if this is too much extra work, you can try skipping the phase shift.
Scott
The PhaseBug also works with Audacity with the help of its VST Enabler, for the rear phase shifts.
And the Audacity has the required effects such as Amplify() and Invert().
It will accept 6ch wav and save to any supported codec containing DPL II stereo signal inside.
I also located the free matlab clone scilab-4.0.exe, which has the very intrinsic hlib function for the Hilbert Transform.
Yeah, no such builtin Hilbert() which certainly require a custom function like what ursamtl addressed again.
But I couldn't find any avisynth plugin(HilbertShift, InvertAudio) yet, for the DPL II automation when the the best s/w emulation model is finally established.
Now all the requird basic functionalities seem to be ready with Audacity : 6ch read, mix, amplify, shift(PhaseBug VST), invert, 2ch save
I will update my original plan to save the space when I finish verifying the all required functionalities.
BTW, an unresolved issue on the PhaseBugMono(in the below links) :
. it has the effect scale input range as 0.000~1.000..
. the document has different gui and scale : in degree
. anybody can kindly explain how to map the effect scale to phase degree?
http://audacity.sourceforge.net/download/windows
http://audacity.sourceforge.net/help/faq?s=install&i=vst-enabler
http://www.savioursofsoul.de/Christian/Plugins.htm
3dsnar
10th June 2006, 08:50
Your test is probably flawed if I understand correctly. A properly mastered 5.1 ac3 track already contains 90° shifted surround sounds. That's why a simple downmix should be enough.
Such statement cannot be actually true.
Because, in case of reach 5.1 mixes there is alot going on,
and some sounds in the surround are (roughly) 90 deg shifted against fronts, some contain similar phase, and some are only present in the surrounds (or fronts). So it is difficult to talk about phase relations.
Anyway, as I understand we are interested in the best DPLII downmixing equation applicapable for 5.1 sources. Since the sources practically in most of the cases come from DVDs (and we do not have the raw tracks used for creating the 5.1 signals), thus we are interested how to convert the DVD 5.1 audio signals to DPLII stereo.
3dsnar
10th June 2006, 08:56
To be able to determine phase shifts as you suggest would require taking the downmixed stereo file and running it through a decoder and then comparing the phase of the resulting channels to the original 6 individual source channels before the downmix. From what I've read, the matrix process is such that complete recovery of original source channels is not possible.
I wonder if all the attention to whether or not the Hilbert transform /90° phase shift is necessary isn't because it's much easier for casual programmers to implement a 0/180° algorithm. For those who want to try, my understanding is that you need to implement a FIR filter with certain cooefficients. Here's a routine for calculating the cooefficients: http://www.musicdsp.org/archive.php?classid=3#195.
Yes, you've got the point. It is not so easy to implement hilbert transform.
In practice, the hilbert transform is applied via spectrum operations, not FIR filtering apprach, because convolutive filtering is a bit computationally heavy.
You have to calculate a complex spectrum (with FFT), manipulate the spectrum bins, and go back to time domain with IFFT. This requires overalapp-add aproach (to preserve signal continuity on the frame edges). If someone is really interested in that, I could help (please PM me. I cannot directly share the source code - sorry).
There is an alternative way to see how the downmix should look like.
Generate an appropriate test signal (let's say square and sinusoidal signals). Each playing
separately in each channel. Than use DPLII certified decoder and see the phase changes
in the decoded multichannel output.
I have no possibility to capture such signal.
Can someone do that?
(even capturing it throuth analog output/input would do)
I prepared the signals.
http://forum.videohelp.com/images/guides/p1514184/squares.7z
1) Just in case, the original
2) The 180 deg. downmix
Lt = FL{0°} + 0.7071 C{0°} + 0.7071 LFE{0°} + 0.866 SL{180°} + 0.5 SR{180°}
Rt = FR{0°} + 0.7071 C{0°} + 0.7071 LFE{0°} + 0.5 SL{0°} + 0.866 SR{0°}
3)
Lt = FL{0°} + 0.7071 C{0°} + 0.7071 LFE{0°} + 0.866 SL{-90°} + 0.5 SR{-90°}
Rt = FR{0°} + 0.7071 C{0°} + 0.7071 LFE{0°} + 0.5 SL{+90°} + 0.866 SR{+90°}
(so this time with the same sign style - i.e. no sign inversion)
DarkAvenger
10th June 2006, 09:16
Such statement cannot be actually true.
Knowing is better than believing. Refer to page 3-6, 4-12ff, 4-19, D-1 of document: http://www.dolby.com/assets/pdf/tech_library/46_DDEncodingGuidelines.pdf
3dsnar
10th June 2006, 09:21
I know this document...
But, what dolby writes is one thing, and how it is done, it could be a different story.
Ofcourse I am talkin about creating DPLII downmixes (not preparing material for 5.1 dolby digital).
But we do not have to argue about this,
but can simply find out.
Please read my previous post.
DarkAvenger
10th June 2006, 09:34
Yes, and I wrote "properly mastered"... and that you are using unsuitable sources for testing.
3dsnar
10th June 2006, 09:37
Uhm, you could be right :o
It would be better to analyze
the waveform changes and see if the DPLII decoder adjusts the (assumed) phase shifts.
---
On the other hand, since some sources may have already adjusted phase,
and some may not, this means that there is no universal solution for the downmixing equation...
And dependingly on the material, you have to use 180 deg or 90 deg. (to achieve what dolby suggests). But will see the decoded downmixes (if someone will be able to perform the decoding
and record the results), to be certain.
ursamtl
10th June 2006, 13:59
Yes, you've got the point. It is not so easy to implement hilbert transform.
In practice, the hilbert transform is applied via spectrum operations, not FIR filtering apprach, because convolutive filtering is a bit computationally heavy.
That's not what Dolby says about its approach. I think they are clear in their explanation of why and how for the Hilbert. They designed DPLII, etc., so I hope they know what they're talking about ;)
(from Dolby Digital Professional Encoding Guidelines, Appendix D, page 1)
Purpose
The 90-Degree Phase Shift filter provides a means for an encoding engineer to create a multichannel Dolby Digital bitstream that can be downmixed to a Dolby Surroundcompatible Lt/Rt output. Without this filter, point-source elements panned from Surround to Center in the multichannel mix would seem to pan from Surround to Left and then to Center when downmixed to Lt/Rt and reproduced using a Dolby Surround Pro Logic decoder.
This filter should generally be used whenever encoding a multichannel signal unless it is known that the 5.1-channel source does not contain point-source element pans. For example, if the source was recorded using five discrete microphones placed in the corners of an auditorium, there is no panning between channels and the filter could be safely disabled. If in doubt, use a DP562 to downmix the 5.1-channel program to Lt/Rt, Dolby Surround Pro Logic decode the Lt/Rt signals, and then set the filter to the setting that sounds best.
Description
The 90-degree phase-shift is created using a very long FIR filter. Since this filter introduces a significant time delay, the other four channels are delayed using a PCM delay line so that all six channels are kept in sample alignment. This filter has exactly
90-degree phase shift at all frequencies. The magnitude response is flat across most of the spectrum, rolling off at the lower edge of the audio band (-3 dB below 30 Hz).
I don't know how much clearer we could get than this! Dolby clearly states how they do it: a very long FIR filter. The source code for such a filter is available around the net, either at the audiodsp.org link I provided yesterday or at a couple of other filter designer sites. Dolby also mentions that this introduces a time delay (also known as latency), so the channels not passed through the Hilbert need to be delayed to compensate for this. Maybe instead of investing all this time and energy in trying to prove one side of the argument or the other, it'd make more sense to try coding it the way Dolby describes it and see how the results turned out? Just a suggestion.:)
Regards,
Steve.
tebasuna51
10th June 2006, 14:06
Knowing is better than believing. Refer to page 3-6, 4-12ff, 4-19, D-1 of document: http://www.dolby.com/assets/pdf/tech_library/
46_DDEncodingGuidelines.pdf
Reading the document:
1) The info is for Dolby Surround Pro Logic (DPL I), nothing about DPL II.
2) Using the info in this DPL II discussion:
a) The phase shift (and the related FIR filters), when needed, is always a process in encoder ac3 5.1 phase, in downmix process only a inverter is used to add the surround mix to the Lt channel.
b) The phase shift is not needed "if the source was recorded using five discrete microphones placed in the corners of an auditorium".
c) The downmix of a Dolby Digital 5.1 compliant signal (with or without phase shift included) only need a inverter. Matrix 1 (simple downmix).
d) The use of Matrix 1 downmix can't guarantee a correct DPL mix if the 5.1 source is not Dolby Digital compliant, but none can guarantee this if the method to obtain the 5.1 signal is unknown.
BTW, there are some contradictions in Dolby documentation, for instance:
At Authoring Dolby Digital and Dolby E Bitstreams (2002), 3.6 Downmixing (page 17), http://forum.doom9.org/showthread.php?p=332259#post332259
"The Lt/Rt downmix sums the surround channels and adds them in phase to the left channel and out of phase to the right channel.
This allow a Dolby Surround Pro Logic decoder to reconstruct the L/C/R/S channels for a Pro Logic home theater."
If this is true we must use a matrix 3 style, with correct (not inverted) SL, SR channels using software DPL II Cyberlink decoder.
3dsnar
10th June 2006, 16:30
@Ursamtl,
I am just talking about en efficient implementation.
The result should be very similar (i.e. spectral multiplications are equivelent to convolution in the time domain).
@Tebasuna. All I was just going to write you did :)
All I can say is EXACTLY. 100% agree.
To summarize, the appropriate phase relations in the 5.1 input should be as close as possible to the Dolby specs. (e.g. application of appropriate phase shifts, etc).
The downmixing process can be viewed as a separate thing.
Cheers, 3d
Rockaria
10th June 2006, 20:21
@tebasuna51
I do not understand clearly what you mean by 'in phase', 'out of phase', 'inverted', 'matrix 3' and the use of 'software DPL II Cyberlink decoder' here.
--------------------------
A more detailed DPL II doc with diagrams of .. DPL1 which I linked to scharfis in my original thread.(look @ around P-3)
http://www.dolby.com/assets/pdf/tech_library/209_Dolby_Surround_Pro_Logic_II_Decoder_Principles_of_Operation.pdf
About the issue on DD5.1(ac3) the possibility of occasionally having the rears shifted :
. I don't want to believe the decoded already discrete 5.1 ch stream needs any phase shift effects on the rears for further channel seperations unless it's EX or similar encoded
. there are contradictions in many related docs : i.e. this doc only mentions the 90 deg phase shift in DPL (II) encoding
. so the target should be limited to normal(no phase shifted) 5.1ch sources for the DPL II encoding emulation.
My interpretation on the DPL I is now adjusted as ursamtl addressed :
Lt = mix(Lf.0°, C.0°, *S.+/-90°)
Rt = mix(Rf.0°, C.0°, S.+/-90°)
. when * = invert(), +/- = + or - phase shift(prolly +)
. no 180° is mentioned here at all : I initially misinterpreted the - polarity(invert) as - phase shift, 180 deg off relations as results.
--------------------------
Now the issue is clearly what reasonable DPL II rear stereo seperation plans we can imagine(in combination with existing matrices) :
<m11> adjusted : no relations between the coefs, polarity between the channel : simple coefs mix
Lt = mix(Lf.0°, C.0°, *Ls1.+/-90°, *Rs2.+/-90°) == mix(Lf, C, *mix(Ls1, Rs2).+/-90°)
Rt = mix(Rf.0°, C.0°, Ls2.+/-90°, Rs1.+/-90°) == mix(Rf, C, mix(Ls2, Ls1).+/-90°)
<m12> the prevailing s/w version : we clearly see the Lf<Rf effect
<m13> to be adjusted
...
<m21> : adjusted : phase between the coefs, polarity between the channels
Lt = mix(Lf.0°, C.0°, *Ls1.+/-90°, *Rs2.-/+90°) == mix(mix(Lf, C), *mix(Ls1.+/-90°, Rs2.-/+90°))
Rt = mix(Rf.0°, C.0°, Ls2.-/+90°, Rs1.+/-90°) == mix(mix(Rf, C), mix(Ls2.-/+90°, Rs1.+/-90°))
<m22> my adjusted s/w emulation that has around 20% better seperations than <m12> : both 180 deg shifts
...
<m31> : adjusted : polarity between the coefs, phase between the channels
Lt = mix(Lf.0°, C.0°, *Ls1.+/-90°, Rs2.+/-90°) == mix(mix(Lf, C), mix(*Ls1, Rs2).+/-90°)
Rt = mix(Rf.0°, C.0°, Ls2.-/+90°, *Rs1.-/+90°) == mix(mix(Rf, C), mix(Ls2, *Rs1).-/+90°)
<m32> : replaced : polarity between the coefs & channels, phase between the channels
Lt = mix(Lf.0°, C.0°, *Ls1.+/-90°, Rs2.+/-90°) == mix(mix(Lf, C), mix(*Ls1, Rs2).+/-90°)
Rt = mix(Rf.0°, C.0°, *Ls2.-/+90°, Rs1.-/+90°) == mix(mix(Rf, C), mix(*Ls2, Rs1).-/+90°)
<m33> : added : polarity between the coefs & channels
Lt = mix(Lf.0°, C.0°, *Ls1.+/-90°, Rs2.+/-90°) == mix(mix(Lf, C), mix(*Ls1, Rs2).+/-90°)
Rt = mix(Rf.0°, C.0°, Ls2.+/-90°, *Rs1.+/-90°) == mix(mix(Rf, C), mix(Ls2, *Rs1).+/-90°)
<m34> : added : polarity between the coefs
Lt = mix(Lf.0°, C.0°, *Ls1.+/-90°, Rs2.+/-90°) == mix(mix(Lf, C), mix(*Ls1, Rs2).+/-90°)
Rt = mix(Rf.0°, C.0°, *Ls2.+/-90°, Rs1.+/-90°) == mix(mix(Rf, C), mix(*Ls2, Rs1).+/-90°)
...
. when * = invert(), +/- = + then - phase shift(a), -/+ = - then + phase shift(b) : i.e. <m21a>,, <m32b>
. we are not certain on Dolby's 90° whether - or +, so applying on the models respectively
. not still sure why there need the relations to be established between the channels, but will evaluate them.
. these DPL II model candidates adjustments will be reflected on the original plan.
tebasuna51
11th June 2006, 00:58
@tebasuna51
I do not understand clearly what you mean by 'in phase', 'out of phase', 'inverted', 'matrix 3' and the use of 'software DPL II Cyberlink decoder' here.
- 'in phase' and 'out of phase' are literal text in document:
Authoring Dolby Digital and Dolby E Bitstreams
Of course is ambiguous, but can be interpreted like the inverted (or {180}) surround signal mus be added to Rt instead to Lt.
- 'inverted', 'matrix 3' and 'software DPL II Cyberlink decoder' are quick references to my test in previous threads. Here a brief resume (with simplified coeficients and results to see the relevant questions):
Purpose: check the relation between input and output channels in the process
6channelwav -> DPL II downmix -> DPL II upmix -> 6channelwav'
Where:
- 6channelwav is a Channel_Test (channels separated in time)
- DPL II upmix was made with Cyberlink PowerDVD 6, Audio Effect dsf (software decode).
- DPL II downmix was made with BeHappy using:
Matrix 1 (like BeSweet/Azid)
LT = L + 0.7 C - 0.8 SL - 0.5 SR
RT = R + 0.7 C + 0.5 SL + 0.8 SR
Matrix 3 (with inverted signs for back channels)
LT = L + 0.7 C + 0.8 SL + 0.5 SR
RT = R + 0.7 C - 0.5 SL - 0.8 SR
Results (L, R, ... input channels, L', R'... output channels) :
M1 decoded in Movie mode
L' = 0.7 L
R' = 0.7 R
C' = 0.6 C
SL' = - 0.7 SL
SR' = - 0.7 SR
M3 decoded in Movie mode
L' = 0.7 L
R' = 0.7 R
C' = 0.6 C
SL' = 0.7 SL
SR' = 0.7 SR
But I don't know if all soft/hard decoders have similar behaviour.
Rockaria
11th June 2006, 03:19
can be interpreted like the inverted (or {180}) Yea, literally you are right.
And with the above linked (recentest maybe) doc's mentoned 'polarity', I believe it's a plain invert(mirroring) on the same phase region.
I also tested your mentioned matrix 3 with FFDShow, with various clips, ac3/dts/aac/wav, 6ch/5ch, varying complexity, h/w DPL II decoder in movie mode.
I found it's indeed better(closer to ffdshow's ac3 out mode) with some types of clips.
I also found the matrix 2(mine) becomes extremely bad with winamp5 bundled 6ch music, forcing me to give up messing with these 180 deg. models any more.
But they will still be included in the tests with some more content types. Just no more pinky expectations.
So I conclude it(the s/w 180 deg. phase shift emulations) heavily depends on the contents. The possible reasons are :
. the phase shifts originally imbeded in the sources as DarkAvenger said.(it will also affect 90 deg models)
. the number of source channels without LFE and/or Center.
. channel complexity of the clips
. no enough/accurate channel seperation plan(algo, esp. Hilbert() & Invert())
...
Now I have encoded a 6ch speaker test 90 deg. phase shifted DPL II based on the <m31> model.
It plays perfectly as expected(most generous no-concurrent-ch-clip).
But what is important now is I assure the Audacity with PhaseBug is working perfectly on my current slow P4, you can use the same procedure & env.
. read a 6ch source wav or ogg..
. duplicate each rear for invert
. -3dB on center, -3dB/-9dB for coef1/coef2 on each rears.(sorta 1:3, minimum residuals)
. PhaseBugMono +-90deg phase shifts on each 4 rear coef. (*check edit : 0.75 : 90 deg., 1.0 : 180, 0.25 -90, 0.00 : -180 )
. align the channels to correct positions L-C-R
. save to 2ch DPL II ogg. done.
I am changing to XP-M soundstorm. I need some significant performance...
Let's see if these 90 deg. phase shifts and inverts gonna make any differnces....;)
[edit]
.The phaseBug input scale seems working strangely : it changes the degree but advances a lot and no negative direction ??
raquete
11th June 2006, 16:48
- 'in phase' and 'out of phase' are literal text in document:
Authoring Dolby Digital and Dolby E Bitstreams
Of course is ambiguous, but can be interpreted like the inverted (or {180}) surround signal mus be added to Rt instead to Lt.
a good lecture about phase: http://en.wikipedia.org/wiki/Phase_%28waves%29
It is common to speak of inverting the polarity of a wave as "flipping the phase" or "shifting the phase by 180 degrees". These are not completely equivalent, though, since a 180 degree phase shift of all signal frequencies would also delay the signal. Inverting the signal is instantaneous.
i have one .pdf from Dolby.Inc showing about LT/RT and respectives phases,later i post the link to download this file(i need to find it)
very interesting thread guys,i'm loving it,go ahead!
;)
tebasuna51
11th June 2006, 19:15
It is common to speak of inverting the polarity of a wave as "flipping the phase" or "shifting the phase by 180 degrees". These are not completely equivalent, though, since a 180 degree phase shift of all signal frequencies would also delay the signal. Inverting the signal is instantaneous.
Not at all. Any delay is a undesired result, if exist any delay then not all frecuencies have the same phase shift.
Invert the polarity of a wave is the perfect method to shift the phase by 180 degrees to all frecuencies.
In matrix equations the terms SL{180} and -SL are absolutely equivalents.
Rockaria
11th June 2006, 22:47
In fact, that's a fundamental issue in the digital world.
Suppose a best visual assimilating situation : throw a stone in a pond and recode the event.
. you will see the tree-age-circles(wave) spereding outside forever until the origin loses it's energy
. it's a continuous analogic wave circles spreading to all directions(360 degree)
. now the issue is how to capture and compact the phenomenon in a digital way from that abundant of analogic data.
. suppose you are far outside from the center of the event and trying to recode the WAVE on the TIME AXIS.
...
My best imagination is a STRING of COIL, which circulates 0~360 degree, but also advances on the time axis, making any degree (i.e. +-3601~ deg) of phase shift possible, spanning any number of samples, definately causing exact amount of time delay.
With more circular scanings(sample rate), the more accurate/bigger data will be attained on each second(bit rate).
Also with more recodings from multi directions, you are getting more accurate spatial informations(5.1~ch, I assume the minimum is 8 mics/speakers for real 3d recoding).
Also if you straighten the COIL wire on the more granular phase*time linear domain, it appears similar to the diagram on Wikipedia.
Any correction is welcomed.
raquete
11th June 2006, 23:46
@ Rockaria.... :goodpost:
My best imagination is a STRING of COIL, which circulates 0~360 degree, but also advances on the time axis,..
something like this?
http://upload.wikimedia.org/wikipedia/en/e/e1/Wave-chirplet-wave-only.png
more lecture : (Complex sinusoids)
http://en.wikipedia.org/wiki/Negative_frequency
Rockaria
12th June 2006, 02:55
Yes, in a microphone's point of view, standing upright, scanning 360 deg. horizantally, 48000times a second(48khz/sec), but mostly recoding the waves from the origin direction [* check edit] on the zero phase degree zone when properly mastered[/*].
The WAV recoded, should be regarded as a bunch of all the waves(coils) scanned having different freq(hz), phases(circles), varying volumes(diameter)..
This might cause a problem synchronizing all the inner waves causing a significant delay in the phase shifting.
The inverting(flipping, mirroring) is a simple negative movement of the wav images on the same time(phase) axis of the grand wav(the captured wave), having no phase advance.
But the changing the polarity might have two different possible interpretations by the axis on the target waves : when the each inner wave is considered, the phase shifting will also make sense, but when the grand wav is considered, the inverting will be more reasonable, provided & proved that the phase shifting on the grand wave will result in different phase shifts on the inner waves because of each differnt frequencies(asking an analogic process or synchroniztions of all the channels when a DSP is used..).
But considering Dolby also included the '90deg phase shifting' in the diagrams, I believe they used the 'polarity' for inverting.
@raquete, thanks for the link that best visually explaining the coil(wave), as well as the related topics. :thanks:
[edit] The recoded mono wave might having the direction(phase shifted) is based on the fact that Dolby's doc on DD saying 'the simple downmixing the properly mastered waves from mics will be enough for DPL(II)...", just a simple assumption.
. some modification on the polarity change
Rockaria
12th June 2006, 23:00
A verification of avisynth's Amplify(-n) or MixAudio() when n=0.000~1 == invert()
not -180 deg phase shift as explained in the doc
<shift360.avs>
a=DirectShowSource("SSWAV06.wav")
a1=GetChannel(a,1)
a2=a1.Amplify(-1)
a3=a2.Amplify(-1)
a4=a3.Amplify(-1).Amplify(-1)
a5=a4.Amplify(-1).Amplify(-1)
a6=a5.Amplify(-1).Amplify(-1)
MergeChannels(a1,a2,a3,a4,a5,a6)
<capwav.cmd>
bepipe.exe --script "import(^shift360.avs^)">6ch380sft.wav
Audacity in max magnified scale on 48khz 6~sec 6ch wav clip
. 0.00001 sec increment view :: 0.00001 * 48000hz :: 0.48sample or cycle
. no advance with all 6ch(except a2 exactly inverted), identical
So all the - signed mix in similar scripts/language might also be the simple invert(), no advance in the phase*time domain.
If this is true, all the existing s/w DPL II emulations are using invert() only : partial implementations : no 90 deg phase shifts
Also note that the Hilbert() used in the quoted matlab script is not a built-in function implying the test result having no basis at all.
I believe that's enough for the verification.
I need to find a way to use the PhaseBugMono correctly :
. assuming it being the allpass phase shift filter
http://www.harmony-central.com/Effects/Articles/Phase_Shifting/
http://en.wikipedia.org/wiki/Phase_shifting
. mapping the input scale(0.000~1.000) to +-90 & +-180 deg scales : it's almost random (to me)
. synchronizing(aligning) the filtered channels to the exact phase*time domain of the clip
. or find other filters to do it more efficiently & effectively.
. or use other environment(tools)
It's taking time...
raquete
13th June 2006, 06:25
from Phase_Shifting article:
"To create the notches, we only need to mix the allpass filters output with the input....So when we mix a delayed copy of the signal with the original, there will be notches at equally spaced frequencies. This is exactly how the flanger operates, and thus, the flanger is merely one type of a phase shifter."
i can "proove" it! :cool:
in 1981 i did one miraculous phaser for guitar(inside analog sintesizer) that had 8 all pass filters turning like spiral using 8 FETs(2n3819 if i right remember) and 8 operational amplifiers(741) !
good link and another cool post Rockaria, thanks.
Rockaria
14th June 2006, 05:49
The invert(), -90 deg phase shift(PhaseBug, allpass filter) and the coefficients did the tric!:)
The earlier mentioned simplest model with -90 phase shift played by far the closest to the AC3 6ch out mode.
<m11> adjusted : no relations between the coefs, polarity between the channel : simple coefs mix
Lt = mix(Lf.0°, C.0°, *Ls1.-90°, *Rs2.-90°) == mix(Lf, C, *mix(Ls1, Rs2).-90°)
Rt = mix(Rf.0°, C.0°, Ls2.-90°, Rs1.-90°) == mix(Rf, C, mix(Ls2, Rs1).-90°)
Firstly over 95 % of the rear stereo seperation leaded the rest channels seperation further around 85%, considering the center/lfe are inherently unable to be seperated perfectly(simple mix). Now we can say the 90deg shift model reconstructs the channels over 90%.
<PhaseBug>
I had to mess around with the PhaseBugMono input scales to best map to the degrees and the amount of the delay to be adjusted.
. 48k/44.1k PhaseBug value 0.322 shifts -90deg
. 360 deg(4time) advances 0.256 sec
. -90 deg advances 0.064 sec
. +-180(2times of 90deg shifts) deg advances 0.128 sec
. +90(270 : 3times of -90deg shifts) deg advances 0.192 sec
. somebody might be able to theorize and establish a more accurate formula
<6ch source>
I chose the previously mentioned winamp5 bundled aac 'Hey Buddy.....aac' because it's open and by far one of the best clip having variety/clear channel harmony.
. transcoded to ogg -q5 6ch with foobar to be imported in the Audacity.
. the channel order might be different depending on the ogg cli encoder versions. but can be easily verified within Audacity.
<fronts stereo mix>
I made the front mix in a seperate project first because the rear can be simply mixed with the front mix again to make a complete DPL II
. copied L,R,C,LFE to a new project named front. L & R already set to correct channel attributes when imported.
. center -3dB, LFE as it is(kinda low initially)
. performed the quickMix() & played to verify : OK
<rear stereo mix>
I made a rear mix to make a complete mix with the prior front mix
. copied the rears to a new project named rear ans set the channel names/attributes to Ls1 & Rs1 : -3dB
. duplicated the Ls1 to Ls2 having the channel attribute to Right, Rs1 to Rs2 attributed to Left : -9dB
. performed invert() on Ls1 & Rs2
. performed the PhaseBugMono 90deg shifts on all the coefficients(0.322)
. performed the cut() of the beginning 0.064sec(align() also can do it) on all the coefs
. performed the quickMix() & played to verify : it played clearly on each rear speakers.
<complete mix>
I made a complete mix and exported to complete.ogg
. copy the front streo mix to the rear & played to verify : almost identical to the original AC3 out mode
. perform the quickMix() & played to verify : same , now safe to export to complete.ogg.
<remaining issues>
. the worldcup
. refine the models for further test : well, I think the coefs did the 'rear stereo channel seperation algo' perfectly.
. OPENing the conclusion & result : the best way is 'DIY' but can share the clips by email requests.
. Theorizing(formulating) the process and automation(programming) : will discuss some time later
@raquete, thanks for the confirmation(allpass filter) and I want to let you guys know that all the cricial issues are hinted by ursamtl in just few words. :thanks:
We did some research, reasoning and tried to adapt to the facts and tools....
Also there might be some mistakes I couldn't catch. Any corrections will be apprciated.
Thanks.
[edit] flawed! the rear cancellation occured(in the beginning on Ls). check below details.
[edit jun14] I meant it that we cannot conclude the best model, not flawed in the approach.
[edit Jun17] verified PhaseBugMono(0.322) = -90 deg shift, as said in a Dolby's doc.
[edit Jun18) typo on the model
Rockaria
14th June 2006, 17:33
Well, it seems too earlier to draw the conclusions. I admit above tests contains some flaws, not enough considerations.
The major problem is the CANCELLATIONS occured when merging the front and rear.inv.90deg-shfted stereos, killing the surround images.(especially the Ls in the beginning). We still need to refine the models and procedures.
. when reviewing the rear channels while phase shfting, I see the image spinning on the time-axis.
. the 90 phase shifted images also contains somewhat altered informations which might be CANCELLED/ALTERED when merged with other channel images which also contains informations on the same phase zones.
. if I understand it correctly, all the channel images reviewed have some directional information(not totally round or flat on the time-aixs, phase shifts imbeded initially), as what DarkAvenger introduced, which might lead to the perfect rear seperation impossible inherently, without any proper mastering maybe.
I's gonna take far more time for me, especially these days....step by step....well, It couldn't be that simple...:(
3dsnar
14th June 2006, 20:14
Also note that the Hilbert() used in the quoted matlab script is not a built-in function implying the test result having no basis at all.
1) hilbert(...) one of the functions from signal processing toolbox
http://www.mathworks.com/access/helpdesk/help/toolbox/signal/hilbert.html
2) you can view phase shift as a time delay, but also as an instantaneous operation. I.e. cosinusoid can be viewed as a time shifted sinusoid. Since in digital spectrum, signal is represented as a set of sinusoids/cosinusoids, thus appropriate time shift of each component is equal to instantaneous operation. Hence both are equivelent.
Rockaria
14th June 2006, 21:19
hilbert(...) one of the functions from signal processing toolbox
OK. Now I see, it is included in a related mathwotk package implying it should perform 90deg shifts correctly on all the inner waves(frequencies) with some delay.
you can view phase shift as a time delay, but also as an instantaneous operation
I agree we can have different interpretations and applications depending on the point of views, if we can ignore the time(frequency cycle) advance.
But what matters here is, how Dolby used the definitions for the DPL(II) encoding.
. how did you interpret the dolby's DPL I diagram in your formula?
. the - sign in the formula, interpreted as -180deg shift in your formula, actually works as invert() or -180 deg shift with the mathlab script?
. the 90 deg shift as +90 deg shift withs *1/4 cycle advance or something else?
(*1/4cycle : minimum sync cycles +1/4 cycle of the lowest freq.)
3dsnar
14th June 2006, 21:45
As I understand Dolby recommendations (which identical to Tebasuna's interpretation) that the 90 deg. shift is necessary as some preprocessing operation. I.e. it depends on the signal, and may or may not be applied to achieve the best separation during DPLII decoding. For example, probably in the example that I provided, the 5.1 input had already properly shifted FL and FR signals. But in some cases it may be necessary.
If I am correct, such speculation probably implies that during the decoding phase, the inversion of 90 deg. shift is not applied. Since the decoder cannot assume what was the original signal.
It requires some tests. But I think we are pretty close to find out how is it done.
- 180 deg. shift is identical (from mathematical point of view) to sign change. For example if you would apply the hilbert transform twice, than you would obtain something identical to the sign change (with some numerical error differences, which can be ommited).
imag(hilbert(s)) is like converting sinus to minus cosinus. And it is applied to all sinusoidal components of the input signal.
This is frequency dependent operation, which means that with the increase of frequency, the time shift gets smaller (since period duration decreases). I think that dolby is talking about the same thing. Especially because they refer to the hilbert transoform (which is applied in time domain as convolutive filtration, but this is the same to the FFT based operation used in matlab).
Rockaria
15th June 2006, 00:48
Matrix 1 (like BeSweet/Azid)
LT = L + 0.7 C - 0.8 SL - 0.5 SR
RT = R + 0.7 C + 0.5 SL + 0.8 SR
Matrix 3 (with inverted signs for back channels)
LT = L + 0.7 C + 0.8 SL + 0.5 SR
RT = R + 0.7 C - 0.5 SL - 0.8 SR
Isn't it too early to say/believe something/somebody is right or wrong?
I believe it just tends to lead to the simple justifications making things ambiguous again..
We are just guessing(interpreting) and testing to figure out the best model for DPL II decoders.
What matters here is the correct OPEN approach to figure out the uncertain matters.
We may not be sure if those rear shifts are meant for the channel seperations in 5.1 DD(I strongly doubt it).
But they may just have some directions by the mic positions(unintended side effects), but certainly not intended for the DPL II channel seperations.
But if you regard the - sign above(180deg phase shift initially to easy implement & replace the 90deg shifts) as the invert() now,
isn't it too dangerous to genralize the formula(without the 90deg shifts) to be used for the general type of clips?
Another irony is Tebasuna seems to have experienced the Matrix3 generally better than the same Matrix1(yours)...
Anyway, if we have to align the shifted channels(like I did), it might have a similar effect as the instantanuous half-flipping.
Maybe I need some more research and tests to understand & implement 'how the 90deg phase shifts works in DPL II with the coefs'.
moved to http://forum.doom9.org/showpost.php?p=841158&postcount=65
[edit]
The quoted type of models with LS.invert() only(without the 90deg shifts) actually produces unbalanced volumes on the fronts as I theoretically proved before.
Now I am clearly excluding them from the candidate list.
The Matrix1 showed Lf < Rf while the Matrix3 showed Lf > Rf.
The closest-Dolby model <m11> showed no such unbalances.
Also my 90deg version <m21> showed no such unbalance but showed not enough seperation of the rear : rears mixed to fronts.
I want to temporarily conclude the <m11> is the most faithful & simplest model so far, until we get any progress.
<m11> no other relations between the coefs, polarity between the channel, coefs +90 deg off from the fronts
Lt = mix(Lf.0°, C.0°, *Ls1.-90°, *Rs2.-90°) == mix(Lf, C, *mix(Ls1, Rs2).-90°)
Rt = mix(Rf.0°, C.0°, Ls2.-90°, Rs1.-90°) == mix(Rf, C, mix(Ls2, Rs1).-90°)
When * == invert(), -90° == 1/4 phase shift
[edit Jun17] verified PhaseBugMono(0.322) = -90 deg shift, as said in a Dolby's doc.
[eidt Jun18] typo on the model
raquete
15th June 2006, 04:29
i can't do good questions/answers in english :o then...
the colors in this screenshot are showing the details that are relevants for me, i will this help the thread:
http://img153.imageshack.us/img153/4018/surroundchannelprocessing2aa.png
i understand "can be" as "we can" and not as " (we) have to".
please,comments and corrections in my "colored details".
no answers:
and reading lots of .pds from Dolby.Inc i can't find answer for this question:
is -90 or +90 degrees?
thanks.
Rockaria
15th June 2006, 05:35
i understand "can be" as "we can" and not as " (we) have to".
please,comments and corrections in my "colored details".
no answers:
and reading lots of .pds from Dolby.Inc i can't find answer for this question:
is -90 or +90 degrees?
thanks.
There are two different modes in the downmixing from the Dolby DD 5.1 sources:
. simple downmix for a stereo speaker mode : don't need the invert() and .90 deg phase shift. But you can apply them expecting almost no difference for the stereo speaker.
. DPL (II) encoding for 5.1ch speaker mode through a DPL II decoder : as well as the invert() & phase shifts, the volume controls are necessary(as seen in many matrix form and the -3dBs in Dolby's documents)
I am going to test the +90 deg shift, but I think it's at best same as the -90 phase shift.
[result]
same result with <m11> having all coefs -90deg shifted.
it seems with the + or - 90 deg shifts, the rear images are crossed on the front phase that the decoder counts.
[edit Jun17] verified PhaseBugMono(0.322) = -90 deg shift, as said in a Dolby's doc.
raquete
15th June 2006, 11:37
I am going to test the -90 deg shift, but I think it's at best same as the +90 phase shift.
[result]
same result with <m11> having all coefs +90deg shifted.
it seems with the + or - 90 deg shifts, the rear images are crossed on the front phase that the decoder counts.
as ask don't hurt (but can bore...lol):
did you could listen differences from -90 to +90 ? some type of "cancelations" ?
thanks so much Rockaria, (you are very dedicated and interested) .
:-)
Rockaria
15th June 2006, 17:53
did you could listen differences from -90 to +90 ? some type of "cancelations" ?
That's a very important point which I was about to verify if the 1/2 cycle offs between the channels or coefs makes any differences on the cross(1/2 : 90-270 degree) phase pane shifts.
The rear channels will be 1/2 cycle(from 1/4 cycle or phase) advanced-mixed to the fronts, creating different mixtures.
When I was doing the above test, I didn't feel any noticeable differences between them, especially on the beginning part(as I mentioned before).
I will report back anytime soon. :thanks:
as ask don't hurtActually you are helping ourselves in a very positive manners!
you are very dedicated and interested
So you are in your best;) contributing way!
3dsnar
16th June 2006, 08:53
There is a simple way to find out how exactly the DPLII downmix should look like. And we also should be certain about the result.
In order to find out we need to
* know the original 5.1 source
* The source has to be a test file, where at the same time only one channel is playing and all channels are used through the sample
* the source waveform should be easy to analyze (i.e. periodic with easy to determine amplitude)
---------------------
* The source can be downmixed with any proposed matrix.
* We need to record output of the certfied Dolby decoder.
As we already know, for such a test sample,
the decoder should nearly perfectly recreate the source (no leakage between channels).
By analyzing how the produced 5.1 signal looks like, we will be able to determine the amplitude of each channel (i.e. what weights should be used) and phase.
I have provided such a test sample (the source and just in case two downmixes).
http://forum.doom9.org/showthread.php?t=112122&page=2
I guess the most problematic thing is to record the output of the certified decoder.
If someone will be able to produce such signal, I will do all the analysis in Matlab, and provide
all pictures (presenting input signal against the decoded one), etc.
Rockaria
16th June 2006, 11:55
In my previous test :
I just proved invert(+90 deg phase shift) not = -90deg phase shift.
-90.align()==+90.+90.+90.align() <> +90.invert() : looks similar but plays differently.
so +90.align() <> half-flipping : +90.+90.align() == +180.align() <> invert()
It turned out having some critical flaw in the method align() because of the gui nature of the Audacity.
I used the 'finding the muting point' for the alignNew() when exactly inverted. Now it has proved :
+180.alignNew() == invert(), when PhaseBugMono(1.000) = 180 Deg Shift
It also proved :
-90 == +90.alignNew().invert()
The effects this introduces are :
. the coefs' alignment must be strictly controlled. ( I am not sure if it affects the cross channels).
. the accuracy of the 90 deg shift seems not affecting much.
...
So far <m11> has proved itself being the most provable model for DPL II, best seperations with no unbalance between the front as seen on existing models.
The issue on the originally shifted rears still remains :
. Is it a mastering problem(neutralizing the position) or need some flattening processing while phase shifting?
Rockaria
16th June 2006, 11:58
Some interpretations :
%180 deg phase shifts
y(:,1) = FL*wgF + C*wgC + LFE*wgC - wgA*SL - wgB*SR;
y(:,2) = FR*wgF + C*wgC + LFE*wgC + wgB*SL + wgA*SR;
%90 deg phase shifts
y2(:,1) = FL*wgF + C*wgC + LFE*wgC + wgA*imag(hilbert(SL)) - wgB*imag(hilbert(SR));
y2(:,2) = FR*wgF + C*wgC + LFE*wgC - wgB*imag(hilbert(SL)) + wgA*imag(hilbert(SR));
<formula 1>
Lt = mix(mix(Lf, C, LFE), mix(Ls1, Rs2).invert())
Rt = mix(Rf, C, LFE, Ls2, Rs1)
<formula 2>
Lt = mix(mix(Lf, C, LFE), mix(Ls1.hilbert(), Rs2.hilbert().invert()))
Rt = mix(mix(Rf, C, LFE), mix(Ls2.hilbert().invert(), Rs1.hilbert()))
<formula 3>
Lt = mix(mix(Lf, C, LFE), mix(Ls1, Rs2).hilbert().invert())
Rt = mix(mix(Rf, C, LFE), mix(Ls2, Rs1).hilbert()))
- when L,R,C,LFE,Ls1,Ls2,Rs1,Rs2 are attained by any matrix including the Dolby dB format.
@3dsnar,
You can verify the hilbert() by checking if hilbert().hilbert() = 180 deg phase shift becomes invert() with the formula 1.
The formula 3 or similar must have been a better choice for a reanoable comparison than the formula 2.
3dsnar
16th June 2006, 16:52
The matlab function produces -90 phase shift (it is like converting
cosine to sine). The function produces returns the real part of the signal (which is the input) and its imaginary part (the phase shifted signal).
So hilbert().hilbert() prduces -180 phase shift (equal to sign change).
Here is a picture:
http://forum.videohelp.com/images/guides/p1514168/hilberthilbert.jpg
cheers, 3d
Rockaria
16th June 2006, 17:25
%90 deg phase shifts
y2(:,1) = FL*wgF + C*wgC + LFE*wgC + wgA*imag(hilbert(SL)) - wgB*imag(hilbert(SR));
y2(:,2) = FR*wgF + C*wgC + LFE*wgC - wgB*imag(hilbert(SL)) + wgA*imag(hilbert(SR));
..
The matlab function produces -90 phase shift
Then you tested with a brand new formula.
Some other points :
. how did you output the graph, through matlab scripts or photoshop(no help in this case)?
. is the imagepart any delayed ? 0/4, 1/4, 3/4 or any reasonable length?
. when I did the 90deg Shift with PhaseBug with a sine, it altered the signal(lower) as the phase shifts repeats. with a complex image, I couldn't notice such alteration.
. can you output the graph with a complex image through matlab script?
3dsnar
16th June 2006, 17:33
Hmm, I did not test the formulas. I have prepared some test signals, according to the formulas, so maybe someone will be able to use it with a DPLII decoder and record results.
They may be incorrect, it is not important. By analysing the produced signal it should be possible to see what the decoder produces, and thefore to guess how the correct formula should look like. I explained that here:
http://forum.doom9.org/showthread.php?p=841101#post841101
-------------------------
Yes, in Matlab. you can draw any signal in Matlab, and then export is as JPG for example (this is what I did).
When you copy-paste these instructions to the Matlab command window:
s=sin((1:44100)/1000);
figure;plot(s); hold on; plot(imag(hilbert(s)),'r'); hold on; plot(imag(hilbert(imag(hilbert(s)))),'g')
legend('sine','imag(hilbert(sine))', 'imag(hilbert(imag(hilbert(sine))))')
You will obtain the same picture.
Rockaria
16th June 2006, 17:58
Hmm, I did not test the formulas. I have prepared some test signals, according to the formulas, so maybe someone will be able to use it with a DPLII decoder and record results.
They may be incorrect, it is not important....
..... That seems to be the major difference between us...
No consensus & progress can be reasonably made with this different approach.
Anyway, what I can say here is :
. prove first and open to the public for evaluations : now you can test freely with my approach.
. also I make corrections ASAP whenever I find flaws in my approach to prevent any misleading : now you see I am also a human.
I will stop the activity on this issue for a while until we get any significant progress than the model <m11>.
/DoItYourself
3dsnar
16th June 2006, 18:15
Ineed, my conclusions were made to early, as I admitted here
http://forum.doom9.org/showthread.php?t=112122&page=2
Now I am trying to find out the truth, based on analysing the DPLII output signal, but I am not able to record a decoded signal (outputed by a DPLII decoder), but I hope that someone will be able to provide it.
If someone will be able to, it will be possible to determine precisely what should be the phase shifts, as well as weights. Without it I cannot make any progress.
ursamtl
16th June 2006, 20:22
Ok here are a couple of curve balls to get you guys out of the Hilbert rut ;)
Perhaps a bit of background reading might help you guys in your quest. Check out an archive of Michael Gerzon's papers at http://www.audiosignal.co.uk/Gerzon%20archive.html. Although he doesn't discuss Pro Logic directly, he does discuss a lot of the fundamentals behind modern surround sound systems. When I recently posted a question on the necessity of a 90deg phase shift to the experts on the sursound mailing list, one of them suggested the answer could be found in one of these papers. I found no obvious quick answer, but there are a lot of fascinating aspects here.
Another issue worth investigating is the whole notion of steering logic, which is quite fuzzy to me. I understand the purpose, but not the implemention. According to Dolby, steering logic is the big difference between DPL versions I and II. Therefore, it follows that, if you want to successfully model a DPLII decoder in software, you have to be able to model the steering logic circuit.
Regards,
Steve.
Rockaria
16th June 2006, 21:53
@raqueue,
This time I got somewhat different/detailed results than before, probably because of the more accurate new align method with more attention.
The (check edit)+90 model(m11a) was worse than -90 shift producing some right-shifted sounds in the beginning part.
<test condition>
. AN7(soundstorm) spdif pcm out mode, Onkyo TX-L5 in DPL II movie mode.
. 6ch source clip : winamp5 bundled 'Hey Buddy.aac' encoded to ogg using foobar2k v0.9 for Audacity input
. Source AC3 Play : FFDShow Ac3 6ch 640kbps out mode
. Audacity + PhaseBugMono : mix processing, DPL II preview. : system mixer spdif pcm out mode
. DPL II ogg play : FFDShow AC3 out mode for 2ch ogg DPL II input.
<tested models> 1:3 -dB ratio on Ls1:Ls2, Rs1:Rs2, ch-coefs are cross-ch mixed
<m10> : (Ls1,Rs2).invert() : 80%, not enough seperation on Rs, Lf < Rf
<m10a> : (Rs1,Ls2).invert() : 80%, not enough seperation on Ls, Lf > Rf
<m11> : (Ls1,Rs2).invert(), rear.hilbert() : 90% to ac3 6ch out
<m11a> : m11.rear.invert() == (Ls1,Rs2).invert(), rear.hilbert().invert() : 85%, the beginning part(airplain) a bit shifted to the Rf
<m21> : (Ls1, Ls2).hilbert(), (Rs1, Rs2).hilbert().invert() : 80%, actively balanced(L-R, F-S), not enough rear seperation
<m21a> : similar to m21
<m22> : (Ls1, Rs1).hilbert(), (Ls2, Rs2).hilbert().invert() : ...
...
Best Candidate : <m11> 3:1 ratio relations between the coefs, polarity between the channel, coefs +90 deg off from the fronts
Lt = mix(Lf.0°, C.0°, LFE.0°, *Ls1.-90°, *Rs2.-90°) == mix(mix(Lf, C, LFE), mix(Ls1, Rs2).hilbert().invert())
Rt = mix(Rf.0°, C.0°, LFE.0°, Ls2.-90°, Rs1.-90°) == mix(mix(Rf, C, LFE), mix(Ls2, Rs1).hilbert())
- when * == invert(), -90° == hilbert() == PhaseBugMono(0.322),
channels are properly attenuated with panning(weights on F-S & L-R) & ratio(center, lfe, coefs)
<conclusion>
. rear.hilbert() seems to affect the rear seperations from the fronts.
. invert() seems to affact the seperation of mixed rear channels.
. the cross mixed rear coefs(3:1 ratio on Ls1:Ls2, Rs1:Rs2) seems to affect the stereo seperation
. the <m11> seperation plan seems to be enough for the DPL II decoder
. other combinations tested so far, produced at best negative effects
<Audacity Issues>
. easily align the shifted channels with time-shift-tool mode.
. can preview the different models easily with undo/redo, quickmix,duplicate.. features
<other remaining Issues>
. treating the originally imbeded positioned ch image
. test,evaluating & gathering the reasonably working candidates.
. automation(programming) with Avisynth, mathlab or other scriptbased(batch) gui tools.
...
@3dsnar,
I hope somebody can provide some help for you to figure out the DPL II decoder functionality.
I believe some better Dolby certified S/W DPL II decoder may provide some ways to capture at any point of the stream processing.
@ursamtl, thanks for another pinpointing lecture(clues). It's gonna be some good brain-soothing time for me.
Hope anybody can test with my setup and find any better models...
I will also report back when I find any better models.... Thanks.
[edit]
http://www.dolby.com/assets/pdf/tech_library/44_SuroundMixing.pdf
Enforced basis by an another document explaining DPL II in a DPL I's point of view.
1. encoder models(look @P7-1) : several types
a. polarity inversion : like most existing s/w invert() only encoder
This method lets you place sounds at each of the four cardinal points: Left, Center, Right and Surround. This is an approximation of Dolby Surround and is limited in its capability. ... However, it does not allow for either interior sounds or center-to-surround pans. To achieve these effects, phase encoding is needed.
...
2. h/w encoder(@P8-1) : it clearly mentions all-pass network, lag(delay), large-frequence.
The all-pass networks are designed so that over the range of the surround bandpass filter, the phase shift of the surround path output always lags that of the left and right by as close to 90 degrees as pratically possible. All-pass networks with this property have large frequencydependent phase lag.
...
3. rear stereo seperation in DPL II : where can I find a Dolby's doc explaining in terms of coefs?
[edit Jun17] verified PhaseBugMono(0.322) = -90 deg shift, as said in a Dolby's doc.
above attached Dolby's doc says it's -90deg shift, so I tested it with a sine signal.
The PhasBugMono(0.322) shifted backward, which is -90 deg. and concures the behaviors of other tools.
Thanks for reminding me. I will adjust the models I tested.
[edit Jun18] about the volume control & phase-shift/invert mix
What need to be kept in the coef stereo seperation is the ratio between them to minimize the residuals and possibly to best seperate the rear channels by cross-ch-embeded-coef-image comparisons.
. weights : controls the PANNING between F-S & L-R. also controls the relative volumes on C & LFE
. coefs : the ratio approximating 3:1 on coef1:coef2, -3dB vs -9~-12dB was best for me with minimum residuals
. clipping control : Amplify(-3~-6dB) on the related channels(fronts or rears) will be proper prior to any mixing, depending on the original volumes.
. gain setting on each channel will be proper to control the weights and coefs ratio prior to any mixing or exporting.
. normalizing should be used on the mixed clips : T(Lt,Rt) or Ft(front)/St(rear)
. removing any possible DC-offsets in the normalize() option will be useful prior to any hilbert() or PhaseBugMono(0.322)
. be reminded the dB is the ratio on the logarithmic line. the resulting dB will be multiplied(not added) in the same sign region. http://en.wikipedia.org/wiki/Decibel
[edit Jun 22] typos on the tested models
bleo
19th June 2006, 03:31
Hi everyone. I authored the thread A better Dolby Pro Logic II downmix (http://forum.doom9.org/showthread.php?t=57988) almost three years ago, and I'm glad to see that there's still interest in Dolby Pro Logic II! :)
I see that the current debate has several different arguments rolled in together so I would like to address them separately.
The first is Rockaria's proposal that I shall describe in my own words as "inverting SR with respect to SL". Specifically, I am referring to the matrices:
<m21>
Lt = FL + 0.7071 C + 0.7071 LFE + 0.866 SL {+90} + 0.5 SR {-90}
Rt = FR + 0.7071 C + 0.7071 LFE + 0.5 SL {-90} + 0.866 SR {+90}
<m22>
Lt = FL + 0.7071 C + 0.7071 LFE + 0.866 SL {0} + 0.5 SR {+180}
Rt = FR + 0.7071 C + 0.7071 LFE + 0.5 SL {+180} + 0.866 SR {0}
I believe these are wrong. Take the case that SL and SR are identical, that is, the audio is intended to sound from between the surround speakers. When encoded through <m22>, Lt and Rt will each contain S at half power and phase {0} (the SL/SR has been dropped since I have defined SL and SR as identical). Thus when decoded, S will wrongly sound from C' (I shall use the apostrophe ['] to denote decoded sound channels).
Similarly for <m21>, Lt and Rt will each contain S at half power and phase {+90}. Thus when decoded, S will wrongly sound from C' but also with {+90} phase shift with respect to the original S.
I have not done 3dsnar's test in post 1 of this thread since I have no convenient internet or audio equipment where I am currently. However, going on the word of others, tebasuna51: "'weapons loaded' and 'Yes, Sir' are presents in FL, FR, BL, BR original audio, and not at Center channel". 3dsnar later reveals that version A uses <m21>. The results of scharfis_brain: "sample A lets it sound from the center" and tebasuna51: "In A "Yes, Sir" is basically only at Center channel" agree with my explanation above.
In conclusion, <m21> and <m22> will sound fine when only one channel is active at a time. Even SL and SR will decode to the correct channel since the {180} phase difference between Lt and Rt is maintained. However, the matrices will be incorrect when multiple channels are active, such as in the case where SL and SR are identical as I have described, but also when they are different, the decoded SR' will be {180} out of phase with SL'. The result will be similar to listening to a stereo CD with reversed polarity on one speaker. I can elaborate on this if anyone requests but I wish to keep this post concise.
The second debate is whether we should use {90} or {180} phase shifts. I shall refer to the following matrices:
<m11>
Lt = FL + 0.7071 C + 0.7071 LFE + 0.866 SL {-90} + 0.5 SR {-90}
Rt = FR + 0.7071 C + 0.7071 LFE + 0.5 SL {+90} + 0.866 SR {+90}
<m12>
Lt = FL + 0.7071 C + 0.7071 LFE + 0.866 SL {+180} + 0.5 SR {+180}
Rt = FR + 0.7071 C + 0.7071 LFE + 0.5 SL {0} + 0.866 SR {0}
Again, both matrices will sound fine when only one channel is active at a time. However, take the case that C and S are identical (SL/SR omitted for brevity), that is, the audio is intended to sound from between the center and surround speakers, such as in the case of a front to rear pan. When encoded through <m12>, S will cancel out C in Lt and only Rt will contain the sound, so it will be wrongly decoded to R'.
This will not occur in <m11> because S is {90} out of phase with respect to C. The {180} phase difference is also maintained for S between Lt and Rt so the decoder will be able to decode it to S'.
This explanation seems to agree with what ursamtl has quoted from the Dolby literature in post 41 of this thread. Dolby seems to imply that downmixers will use simple {180} phase shifts like <m12>. The original audio engineer should apply the {90} phase shift to the original SL and SR if they contain "point-source elements panned from S to C". Thus <m12> will be transformed into <m11> and we will not have the problem described above.
My conclusion is that it is too difficult to do the {90} phase shift ourselves. I would stick to <m12> and hope that there are no C/S pans or that the original audio engineer already applied the {90} phase shift if required.
The third debate is whether Lt or Rt should have SL and SR mixed (+) or (-). I shall refer to the following matrices:
<m12>
Lt = FL + 0.7071 C + 0.7071 LFE + 0.866 SL {+180} + 0.5 SR {+180}
Rt = FR + 0.7071 C + 0.7071 LFE + 0.5 SL {0} + 0.866 SR {0}
<m12a>
Lt = FL + 0.7071 C + 0.7071 LFE + 0.866 SL {0} + 0.5 SR {0}
Rt = FR + 0.7071 C + 0.7071 LFE + 0.5 SL {+180} + 0.866 SR {+180}
It is plain to see that SL' and SR' in the decoded output will both be {180} phase shifted with respect to the original SL and SR when using the incorrect matrix. The results of tebasuna51's tests in Software Dolby Pro-Logic II decoder?, post 41 (http://forum.doom9.org/showpost.php?p=836023&postcount=41), show that <m12a> may be the correct matrix. However, do note that when using a DPL1, or DPL2 decoder in Movie mode, SL' and SR' are delayed by 10msec which will have a frequency-dependent effect on your phase measurements. tebasuna51's results do in fact show that there are phase differences between DPL2 Movie and Music modes, even when using the same downmix matrix. My conclusion is that further tests are required to determine the behavior of different hardware and software DPL2 decoders.
Rockaria
19th June 2006, 07:49
Hi bleo, I felt I need to put some comments on your post.
I also read your guide long ago and tried to adapt myself to the formula <m12> but have experienced some problems that made me seek other models :
. the rear seperation is not applied fully that it causes the unbalances in normal music play : Lf < Rf for m12 and Lf > Rf for m12a.
. the m2~ models which I tested did not make such unbalances and actually produced better seperations on some/most types of clips (to me).
. to be precise, only one faithful model(by the spec) is correct, others are just partially implemented or incorrect.
But I finally concluded that the 180 deg shift(actually invert by dolby's spec) models are lacking something crucial, which later turned out being the C-S panning which will also affect the front seperation(Lf <> Rf).
You can easily notice the unbalanced playback when you switch between <m12> & <m12a>(Lf <> Rf), unless anybody want to be satisfied with the approximated implementation, as Dolby themselves admitted(as linked in the edit part above in my post).
As you see, I am now settled down to the <m11> model which is adopting the invert & -90 degree shift(as a result +-90 deg shifted & aligned). And it's not that hard to encode, just not automated yet.
I have tested with various clips so far and in fact got astonished by the seperation quality it might have been able to provide earlier to the DPL II users who still prefer the advantages than any other discrete formats can do : compact size, multi platforms ...
The hilbert transform seems not have been any new algorithm since very long ago.
The only issue I am having is that I am looking for a theoretical background which explains the rear stereo seperation with the cross-channel-imbeded-rear-coefficients reasonably. It proves to be correct through my ears & heart but yet through my brain and eyes. Everything else(invert, 90 phase shift) is clearly explained(and linked here) by their DPL I explanations and diagrams.
One more thing I want to point out is that the PowerDvd DPL II decoder might not be complete yet.
I believe there is no reason the music mode to have inverted channel seperations than the movie mode in normal DPL II decoding.
Anyway, I know I got the correct results finally, even though I am not relying on the DPL II encoding that much on 6ch play. (once experienced, I won't go back:)
I just want it treated as a plain fact. As you said it has been partial implementation. The difficulty of the hilbert() implementation is another thing.
But it's also my pleasure exchanging different opinions especially when we read & understand each other's views clearly.;)
bleo
19th June 2006, 19:24
Hi Rockaria. I welcome your debate and especially your test results. In fact, I hope that you will see that your results agree with my theories!
In reply to your concerns:
[Rockaria]: the rear seperation is not applied fully that it causes the unbalances in normal music play : Lf < Rf for m12 and Lf > Rf for m12a.
For <m12>:
Lt = FL + 0.7071 C + 0.7071 LFE + 0.866 SL {+180} + 0.5 SR {+180}
Rt = FR + 0.7071 C + 0.7071 LFE + 0.5 SL {0} + 0.866 SR {0}
let us take the case that FL and S are identical (SL/SR omitted for brevity). When encoded through <m12>, S will cancel out half of FL in Lt, and be present in Rt at half power and phase {0}. Thus, the sound will be wrongly moved towards R' and in this case be decoded to C'. Even when FL and S are only partially identical, that is, they contain only some identical waveforms, these waveforms will cancel each other out in Lt and be added to Rt, resulting in an overall output where FL' < FR' as you have seen.
<m12> is a compromise for easier downmixing and the example illustrates the difficulties of matrix encoding multiple concurrent channels.
[Rockaria]: the m2~ models which I tested did not make such unbalances and actually produced better seperations on some/most types of clips (to me).
Perhaps this is true, but you have created another problem for the case where SL and SR are identical (a mono surround) which I have already described. I believe that this may be common in movie soundtracks, which is very relevant to Doom9 Forum readers!
Furthermore, the better separation that you hear is probably the result of SL' being {180} out of phase with SR', again as I have already described.
[Rockaria]: to be precise, only one faithful model(by the spec) is correct, others are just partially implemented or incorrect.
The correct matrix is <m11>. However, <m12> is much easier computationally. How long does it take to {90} phase shift the two surround channels of a 2 hour movie soundtrack containing some 700 million samples?? Also, if the original soundtrack engineer already applied the {90} phase shift as per the Dolby specifications, then <m12> is the correct matrix to use.
[Rockaria]: I believe there is no reason the music mode to have inverted channel seperations than the movie mode in normal DPL II decoding.
A DPL1, or DPL2 decoder in Movie mode (but not in Music mode) delays SL' and SR' by 10msec which will have a frequency-dependent effect on your phase measurements.
Rockaria
19th June 2006, 21:13
It seems you are starting the debate when I want to save the words.:(
I think I understood your points clearly but unfortunately you seems not because I see you repeating the formula again.
Actually I found some issues with your formula style because of several reasons I described before.
. The + & - signs in the equations : I've already seen them misleading some people to do arithmetic operations on the wave images.
. the varibles regarded as constants : we have seen already too many matrices having different combinations when Dolby only suggested -3dBs in th spec. Your constants does not consider the clipping control, normalizations and custom pannings. I believe what is important is the ratio(1:3) between the rear coefs to minimize the residuals and best seperate the stereo images from the cross mixed 4 coefs images.
. does not faithfully reflect the Dolby's spec : Dolby says (Ls1, Rs2).invert then apply hilbert() to all the rear coefs resulting +-90 shifts on the rear coeffs. But your formula interpreted and simplified it as '180 deg off only' and applied the invert(-) only.
The simplification can be used only for analysis or optimization of the logic(definitely with a full formula remarked).
Now, some people will still regard the 180deg and the matrix as the magic numbers when they are just simple relative ratios and forget the invert(), hilbert() and any other necessary gain contols by your insistent contributions.
If you read my linked Dolby's doc again, the invert only matrix is considered for game surrounds, not for the professional music.
The movie sounds may be actually similar to game sounds that I agree it might suit for casual movie users but certainly not for music users.
The channels might having the directions is also another variables. I believe there is no constant such as already 90 deg shifted channels if the original engineer mastered the clips for general purpose use(especially for DD5.1). Please provide the links to generalize your concerns. I believe the positioned channels are just some minimul unintended side effects.
And the 10ms rear delay in the PowerDVD in movie mode play... Do you really believe if you shift your listening position to 1 m or so backward, the SL is played inverted? I think the DAC & speaker set is another layer..and the DPL II process is limited to encode & decode layer.
It is also possible for some people to try to use anyting in hand to backup the theory. And when they are abused, we call it a justification.
And 'how long does it take to phase shift?'.... I'd say it depends on the experience. I can't make people understood who just try to debate without any personal tests(verifications) and formal basis(link please!). It took me just 1.5 times of the invert().
Please read the linked Dolby's doc & my previous posts(to understand my views) and test it for yourself to verify if you want to continue the conversation. I've tried to keep the consistency but corrected anything I found wrong ASAP. Because I believe they are just simple facts, nothing related to human as far as they keep the correct approach(conversation).
bleo
20th June 2006, 04:40
The correct matrix is <m11> or <m11a>. (The one with {+/-90} phase shifts). I don't know how much clearer I can be.
Now, according to ursamtl's quote in post 41 of this thread from the "Dolby Digital Professional Encoding Guidelines", the encoding engineer of a DD5.1 soundtrack may or may not apply the {90} phase shift to the surround channels. If it is applied, then since:
{90}*<m12> = <m11>
thus <m12> would be the correct matrix to use when downmixing this particular soundtrack to DPL2, and in these cases only.
We are not discussing the coefficients here (0.7071, 0.866, 0.5); they were already discussed three years ago. If you wish, you may normalize them to prevent clipping.
And Rockaria, you need to go and learn how a time delay will cause a frequency-dependent phase shift...
Rockaria
20th June 2006, 05:58
The correct matrix is <m11> or <m11a>. (The one with {+/-90} phase shifts). I don't know how much clearer I can be.
Did I say <m11> incorrect? And have you ever mentioned the <m11> model in all through your post before we started to verify the models? Why do you appear in the last stage of the research and use my <m11> for a justification of your severe mistakes on <m12>?
Now, according to ursamtl's quote in post 41 of this thread from the "Dolby Digital Professional Encoding Guidelines", the encoding engineer of a DD5.1 soundtrack may or may not apply the {90} phase shift to the surround channels. If it is applied, then since:
...{90}*<m12> = <m11>
thus <m12> would be the correct matrix to use when downmixing this particular soundtrack to DPL2, and in these cases only.
I know ursamtl is a person with inspirations, experiences and no-biased opinions.
But quoting his message for your justifications won't be welcomed by anybody and it does not make yourself another ursamtl especially without the correct understandings.
If you read the message again, there is no proof that Dolby is using the positioned channels for the DD5.1 encoding.
The originally positioned mic sources might be useful for simple DPL II mix, like HRTF. It may not even need the invert(), a very special case especially customed to bleo.
We are not discussing the coefficients here (0.7071, 0.866, 0.5); they were already discussed three years ago. If you wish, you may normalize them to prevent clipping. Then we wasted 3 years because of your misleading simplified formula.
And Rockaria, you need to go and learn how a time delay will cause a frequency-dependent phase shift...
What is this?:angry: Do you want to believe you understand better than me? Simply because you joined this forum 2 years earlier?
You definitely need to prove you are not accusing & depreciating people. Or I might feel we need to correct your immature attitudes onto people and matters.
You are still not providing any resonable or formal basis(yeah, a simple blinded debate).:cool:
specise_8472
21st June 2006, 10:30
Here is a good reson for 90 degree phase shift!!
The surround channel should ideally be processed so that the
components of the signal in the two encoded Left Total (Lt) and
Right Total (Rt) channels are at +90 degrees and -90 degrees
phase shifts from the original surround signal. This is because of
the need for dynamic panning effects, for example causing a
sound to sweep over the heads of the audience, front to rear. If no
plus and minus 90 degree phase shift is applied, then the
Surround channel would have to be panned fully left, and then
phase inverted and panned fully right. Then the following would
occur during a front to rear sweep:
When the signal is in the Centre channel, it consists of
L + R. When the signal is in the Surround (i.e. rear) channel, it
consists of L - R. For a moving sweep from front to rear, the
level must be reduced from full to zero level over time in the
Centre channel, and increased from zero level to full level over
the same amount of time in the Surround channel. However , the
problem occurs when the level of the Centre channel is the same
as the level of the Surround channel. The ‘+R’ component in the
Centre channel and the ‘-R’ component in the Surround channel
will cancel each other out completely, leaving the signal only in
the Left reproduction channel. Thus the sound will appear to
sweep from front to the left and then to the rear, which is not the
desired effect.
In order to alleviate this effect of leftwards sweeping,
the plus and minus phase shifts are added in order to de-correlate
the -R component of the panned sound in the Surround channel
from the +R component of the panned sound in the Centre
channel. Therefore the component of the sound in the +R part of
the Centre channel is not the exact inverse of the sound in the -R
part of the Surround channel, and thus the sound is not cancelled
in the Right reproduction channel when the levels of the Centre
and Surround channels are the same during the front to rear
panning. The actual way to implement these phase shifts was
examined in some detail, however, the solution used was to delay
the surround signal before sending it and its inverse to the Left
Total and Right Total channels respectively.
If the Surround signal is delayed by a time interval
equal to 90 degrees of the wavelength of a frequency which is in
the centre of the bandwidth occupied by the surround signal, then
the phase shift is 90 degrees at that frequency, and deviates from
90 degrees as the frequency gets higher and lower than the centre
frequency. At this point the signal can be approximated as having
been phase shifted by 90 degrees. The delayed signal can be
lowered by 3 dB, and sent to the Left Total channel, then the
signal can be inverted and sent to the Right Total channel. In this
way, there is an approximate phase shift of plus and minus 90
degrees, and the signal is encoded differentially onto the two
transmission channels Left Total and Right Total. The time delay
causes the surround channel’s -R component to be de-correlated
from the centre channel’s +R component, in a similar fashion to
the way the pure plus and minus 90 degree phase shifts decorrelate
these signal components. 800Hz seems to be a good
mid-frequency point on the 100Hz to 7kHz bandwidth ( taking
into account the logarithmic frequency law). At this frequency, a
full wavelength phase shift is achieved by a time delay of 1/800
seconds. This is equal to 1.25 milliseconds. Thus a 1/4
wavelength phase shift is equal to 0.3125ms. The lowest
frequency in the reproduced surround channel is 100Hz. This has
a period of 10 ms, and thus a 1/4 wavelength period of 2.5 ms.
The time delay value chosen should then be a multiple of the
value for 1/4 wavelength of 800Hz , that is it should be a multiple
of 0.3125ms. It should also be greater than 2.5ms , in order to
effectively de-correlate the surround signal effectively at the
lower frequencies. A time delay of 7.8 ms was found to be most
effective during algorithm testing. This is the equivalent to a
delay of 6.25 wavelengths at 800Hz.
ursamtl
21st June 2006, 13:25
Here is a good reson for 90 degree phase shift!!
Interesting reading, specise. What's your source for this?
specise_8472
21st June 2006, 19:23
Interesting reading, specise. What's your source for this?
http://profs.sci.univr.it/~dafx/Final-Papers/pdf/Carugo.pdf
bleo
21st June 2006, 20:36
Thank you specise_8472. That explanation was what I was trying to say when I was describing the deficiencies of <m12> (the matrix with {180} phase shifts for the surrounds in Lt). The only difference being in <m12>, a front to rear pan wrongly goes from C' to R' to S'. Your explanation uses a slightly different matrix that is actually <m12a>--{180} phase shifts for the surrounds in Rt). So as described, a front to rear pan wrongly goes from C' to L' to S'.
Rockaria
21st June 2006, 20:55
That seems to be a fairly well described resoning on DPL 1, although we are not sure of the source.
. rear.hilbert() seems to affect the rear seperations from the fronts.
. invert() seems to affact the seperation of mixed rear channels.
. the cross mixed rear coefs(3:1 ratio on Ls1:Ls2, Rs1:Rs2) seems to affect the stereo seperation
The major difference between the DPL 1 and DPL II is of cource the third step of algo : rear stereo seperation.
The DPL II decoder steering logic(we have no spec yet, ursamtl) refined by the servo feed back may perform the three stacked seperation plans almost concurrently that any approximated results from each step will affect each other instantly.
The all-ch mixed Rt of the invert-only versions will undoubtedly confuse the DPL II decoder to some degree.(theorectically & practically proved)
Even when we play these invert-only encoded clips in stereo mode, I mostly experience the Lt <Rt making me assume :
. the playback volume of the inverted image not identical to the original image.
. the cross phase stacked(90 deg shifted) image has generally lower altitude(volume) differenthiating it a little from the simple downmix.
The limitation of the 'mid-frequency of 100Hz to 7kHz bandwidth' seems to be excluded from the DPL II version, iirc, correct me if wrong.
And I agree there are two possible approaches in the phase shift :
a. 1/4 cycle time delay on all frequencies(inner waves) and align(sync) them(look at Valex's post below) : depending on the granularity of the frequency band(to perform the shift together), the amount of delay and phase shift degree will be differently approximated. http://forum.doom9.org/showpost.php?p=360179&postcount=20
b. The alteration of the wave signal on all the frequencies(hilbert...) : basically same as the upper methods but might alter the signal resulting in some side effects.
Anyway, we need to theorize the two steps of the align() for any automations depending on the algo used for the 90deg phase shifting & target accuracies. I need some more research and analysis for this to formulate.
Thanks for the decent quote specise who has been providing quite unbiased opinions on the issues.
raquete
21st June 2006, 22:23
That seems to be a fairly well described resoning on DPL 1... right!
The limitation of the 'mid-frequency of 100Hz to 7kHz bandwidth' seems to be excluded from the DPL II version, iirc, correct me if wrong. .:goodpost: no corrections,you're completely right
DPL II channels are full range(less LFE of course)
bleo
21st June 2006, 23:14
Out of interest, I reread the Dolby Surround Mixing Manual (http://www.dolby.com/assets/pdf/tech_library/44_SuroundMixing.pdf) to investigate the concept in the 2nd part of specise_8472's quote: using a time-delay to cause a frequency-dependent phase shift. Note that this discussion is only relevant to DPL ONE encoders. In Chapter 8: Theory of Operation, 8.1: Encoder:
All processing [-3dB, 100Hz-7KHz bandpass filter, and Dolby B noise reduction] in the surround path contributes to the total degree of phase shift for that channel. The all-pass networks are designed so that over the range of the surround bandpass filter [100Hz-7KHz], the phase shift of the surround path output always lags that of the left and right by as close to 90 degrees as pratically possible. All-pass networks with this property have large frequencydependent phase lag. Thus for instance at 1 kHz, the left and right paths through the Dolby Model SEU4 give phase shifts of roughly –550 degrees, while the surround path, measured at the right total output, has about 90 degrees more lag (approximately –640 degrees total).
After doing some calculations, I take this to mean that processing of the front channels causes a delay of 1.528msec, and the surround channel by 1.778msec, so that the surround channel is delayed 0.25msec relative to the fronts. This will cause the following phase shifts in the surround channel relative to the fronts:
100Hz -9°
1KHz -90°
7KHz -630°
So it seems that DPL 1 does not use fixed phase shifts like +/-90° or 0/180°. This method should still be enough to decorrelate the surround channel from the fronts, but I suspect that this then requires the 100Hz-7KHz bandwidth limitation on the surround channel. Following this reasoning, if DPL2 is not bandwidth limited on the surround channels, then it probably does not use a time-delay phase shift, but instead I guess that the technology became available to do a constant 90° all-frequency phase shift.
specise_8472
22nd June 2006, 01:27
This article is good at explaining more - that a sample delay of 3 gives the needed results.
http://www.gamasutra.com/features/sound_and_music/19981218/surround_01.htm
specise_8472
22nd June 2006, 01:43
The major difference between the DPL 1 and DPL II is of cource the third step of algo : rear stereo seperation.
The DPL II decoder steering logic(we have no spec yet, ursamtl) refined by the servo feed back may perform the three stacked seperation plans almost concurrently that any approximated results from each step will affect each other instantly.
Yes but to my thinking, DPLII could be mixed exactly the same way, just that you mix LS into the left and RS into the right. Then you should be able to extract these back out by the fact that they have a phase difference. IE you can recover DPLI you should be able to do exactly the same with each individual channel.
Also DPLIIx would be the same, but with the CS mixed ito the Rear surrounds at -3db and easily recoverable out of the recovered rears.
I could be wrong here, just my thinking.
I did find last night some docs on the steering logic used in similar decoders. Based upon Right/Left and Left/Right.
I will play around with the DPLII encoder tonight and see what gives. It also offers decoding as well.
bleo
22nd June 2006, 07:32
I've done some more reading of Mixing information of Dolby Pro Logic II (http://www.dolby.com/assets/pdf/tech_library/214_Mixing with Dolby Pro Logic II Technology.pdf) and in the section Music Mode vs Cinema Mode Decoding it says:
Cinema mode is very close to the standard Dolby Surround decoder, with the exception of the full-bandwidth stereo surrounds. Also, a polarity inversion of one of the channels is used to spread the signal in the room. So in fact, the inversion of SR' has nothing to do with the 10+msec delay of the surround channels.
Now, I have done some tests encoding stereo music with matrix <m11a> to L'/C'/R'; R'/SR'; and SL'/SR'. I used Cool Edit to do the 90° phase shifts, WinDVD's DPL2 decoder, Sonic Foundry Soft Encode DD5.1 encoder and finally WinDVD Dolby Headphone to listen to the result since I don't have 5.1 speakers where I am currently.
As an aside, the default option in Soft Encode is to 90° phase shift the surround channels... but that's another debate and more about DPL ONE compatibility anyway...
The main result was that I found SL' and SR' were 180° out of phase in Music mode, and in-phase in Movie mode. This agrees with tebasuna51's earlier results, but seems contrary to what the Dolby specs say.
anyway, it's very late and I'm a bit stumped right now... so does anyone have any clue what this is supposed to mean??
specise_8472
22nd June 2006, 08:36
A snippet from the Surcode DPLII Encoder Manual.
9.4 Invert Right Surround
In Pro Logic II, the Ls and Rs outputs are intended to be out of phase by 180 degrees for surround steered inputs. However, if the decoder is set to Movie mode, the polarity of the Rs output is inverted in order for the Ls and Rs outputs to be in phase. This improves phantom rear center imaging.
ursamtl
22nd June 2006, 13:04
A snippet from the Surcode DPLII Encoder Manual.
This would make sense since DPLII music mode is generally promoted by Dolby as intended for simulating 5-channel surround when listening to regular stereo music. As such, they don't want accurate phantom center imaging in the rear; they want diffuse ambience.
scharfis_brain
22nd June 2006, 13:10
also I want to add that there won't be any phantom center surround anymore when inverting one of the surround channels before downmixing. The centered surround will be sent to the front center instead.
So this must be a mistake in the SurCode Manual or even in the SurCode program itself.
Rockaria
22nd June 2006, 18:38
If I understand it correctly, the Music mode is designed for non-DPL(ii) encoded music, which will generally require some aggressive sparial effect through fronts and rears.
Also the Movie mode which is designed for DPL II encoded clips(it has the Ls inverted during the encoding time) will generally require the correct reconstructions of the original channels.
The problems we have now to verify prior to any conclusions seem to be :
. what happens if we decode the DPL II clip in Music mode? will it play the Ls inverted?
. is the Difussion/rear-phantom-center effect used after or during of the decoding process?
. for the DPL II Movie mode play, it is also possible to interpret as by inverting one of the channels again : 1. both out of phase : more difussion, 2. both in phase : rear phantom center, closer to the original
. what is the effect of one rear channel only inverted through the speakers? : unbalanced/directional playback?
I did find last night some docs on the steering logic used in similar decoders. Based upon Right/Left and Left/Right.
I will play around with the DPLII encoder tonight and see what gives. It also offers decoding as well.
It would be great to have the theoretical basis of the rear-stereo-seperation through the cross-mixed-rear-channels steering decoder logic. :thanks:
When not cross mixed, I got one rear channel playing in the front.
When not mixed in correct ratio(1:3), I got residuals(no full seperation) on both rears.
Of course, these experiences are limited to my DPL II decoders(onkyo, sony) and my s/w DPL II encoding experiments..
bleo
22nd June 2006, 18:47
also I want to add that there won't be any phantom center surround anymore when inverting one of the surround channels before downmixing. The centered surround will be sent to the front center instead.
So this must be a mistake in the SurCode Manual or even in the SurCode program itself.
Hi scharfis_brain! Yes this is true if one of the surround channels is 180° out of phase with the other before or during downmixing. This is the case with <m21> and <m22> which we already agree on.
But I think the SurCode manual is referring to inverting the SR' output after decoding. The WinDVD DPL2 decoder also has this option.
This concurs with my test results. The only thing that remains to be interpreted is when the Dolby manual says: " a polarity inversion of one of the channels is used to...", does "spread the signal in the room" mean the [I]final SL' and SR' are in-phase or 180° out of phase? It would seem that the answer is in-phase.
Rockaria
22nd June 2006, 19:33
also I want to add that there won't be any phantom center surround anymore when inverting one of the surround channels before downmixing. The centered surround will be sent to the front center instead.
Yes this is true if one of the surround channels is 180° out of phase with the other before or during downmixing. This is the case with <m21> and <m22> which we already agree on.
No, unfortunately it seems you are mistaking(in agreeing & quoting). If you read the model revision history :
I corrected the '180 deg phase shift 'to invert(), because they are misleading by misunderstanding.
I excluded all the invert-only models from the candidate list : to m10~
I revised all the models to have invert & 90 deg phase shift included : the model variations are basically for figuring out the best rear-stereo-seperation plan which some people here never have considered seriously.
Also I have repeatedly said, <m11> is the most faithful-simplest DPL II model so far, other combinations I've tested gave negative effects at best.
So as I said, it's basical now to understand the rear-stereo-seperation plan to have the complete skeleton of the DPL II architecture, by understanding the DPL II steering logic maybe.
bleo
22nd June 2006, 20:18
one of the surround channels is 180° out of phase with the other before or during downmixing. This is the case with <m21> and <m22>
Translated to your notation, I meant:
if Ls = Rs, such as the case of a phantom center surround, then in <m21>, Ls1 = (Rs2).invert() and Ls2 = (Rs1).invert()
thus they will cancel each other out in Lt and Rt.
Rockaria
22nd June 2006, 20:45
..<m21>, Ls1 = (Rs2).invert() and Ls2 = (Rs1).invert()
thus they will cancel each other out in Lt and Rt.
Then by your theory, any invert mixed channels will cancel each other (completely)!:scared:
Thus by your notation of 180 deg shift which is infact invert().delay(2/4cycle) :
Ls=0.85xxSL{180} + 0.5SR{180}, Rs=0.85xxSR{0} + 0.5SL{0}
Lt=L+C+LFE-Ls
should (completely) cancel each other, yeah no lefts..
<m2~> are kinda reasonable plans, but proved not complying the DPL II decoder mechanism by tests. Thus I can agree it need to be removed from the candidate list for this reason only.
Another thing : how were you able to connect his 'invert in one rear channel' to <m21>, <m22>?
Have you got any prej- information before decoding?
..inverting one of the surround channels before downmixing.
Also I want to remind you that without including the rear-stereo-seperation logic, affecting concurrently each other, I can only rely on the playbacks through the actual 5.1 speakers.
bleo
22nd June 2006, 21:20
Then by your theory, any invert mixed channels will cancel each other (completely)!:scared:
Thus by your notation of 180 deg shift which is infact invert().delay(2/4cycle) :Actually I meant 180° phase shift = just invert() with no delay. Sorry for the misunderstanding.
Ls=0.85xxSL{180} + 0.5SR{180}, Rs=0.85xxSR{0} + 0.5SL{0}
Lt=L+C+LFE-Ls
should (completely) cancel each other, yeah no lefts..Yes! I think we are finally on the same wavelength! (pun intended! :)) For <m12> using 180° phase shifts, Ls will cancel out L in Lt if they are identical. This is the problem that both specise_8472 and I were describing. That is why we need to use 90° phase shifts. Then all of the channels will be nicely separated by being in different channels--Lt or Rt, or by phase--must not be 180° out of phase with another channel in the Lt or Rt mix.
Rockaria
22nd June 2006, 21:40
Actually I meant 180° phase shift = just invert() with no delay. Sorry for the misunderstanding.
By the way, whose misunderstanding you meant? mine or yours?:rolleyes:
While you say 180 degrees, you actually use '-' in the formula, which is invert().
The 'out of phase' also is actually -180 ~ 0 phase region, not simply invert(-) or -180 deg.
Repeating my points, any correct description based on DPL 1 decoder behavior, might not be useful that much when the rears have STEREO channels in DPL II.
Anyway, I don't want to be trapped by so many exceptional what ifs based on the legacy DPL 1....;)
bleo
22nd June 2006, 21:56
I believe there's a small typographical error in post #73:
<m21> : (Ls1,Rs2).invert(), (Ls1, Ls2).hilbert(), (Ls1, Ls2).hilbert().invert()should read:
<m21> : (Ls1,Rs2).invert(), (Ls1, Ls2).hilbert(), (Rs1, Rs2).hilbert().invert()
Another thing : how were you able to connect his 'invert in one rear channel' to <m21>, <m22>?Because <m21> = <m11a> with (Rs1, Rs2)invert(), you are effectively inverting Rs during downmixing.
Also I want to remind you that without including the rear-stereo-seperation logic, affecting concurrently each other, I can only rely on the playbacks through the actual 5.1 speakers.This sounds like the original work from 2002: BeSweet v1.4b12 & Dolby Surround II Matrix (http://forum.doom9.org/showthread.php?s=&threadid=27936). My thread A better Dolby Pro Logic II downmix (http://forum.doom9.org/showthread.php?t=57988) is a continuation of this work. I would like to add that the original 'surround2' downmix had the coefficients:
Lt = L + 0.7071 C - 0.8165 BL - 0.5774 BR
Rt = R + 0.7071 C + 0.5774 BL + 0.8165 BR
but all normalized by multiplying by 0.3225.
After those threads, the 'rear-stereo-separation logic' still has not been fully explained. But the results are still good enough for me ;)
Rockaria
22nd June 2006, 22:18
OK. thanks for pointing the typos instead of answering my Qs.
(Actually I stopped the further model refining after the reults of the <m11> and thinking the seperation plan is optimized for the DPL II decoder. Of course, I will have to reagrrage the models any time soon.)
You seem to be really proud of your invert-only s/w encoder efforts, three years ago. I wonder if Valex' has implemented the 90 deg shifts by delaying the frequencies accordingly since then.
But I want to make it very cleare that WE are interested in the 90 deg phase shifts implementation(no ommissions from the Dolby's spec).
If you are not interested in the enhancements, please at least stop advertising or over-protecting your partial implementation. No music users will go back, once tasted the full implementations. Nobody here is depreciating your efforts also but yourself.:cool:
bleo
23rd June 2006, 04:08
Ls' and Rs' are separated by their different amplitudes (the 0.866 and 0.5 coefficients) in Lt and Rt.
I have tested the matrix <m11a>:
Lt = L + 0.7071 C - 0.866 Ls (-90°) - 0.5 Rs (-90°)
Rt = R + 0.7071 C + 0.5 Ls (-90°) + 0.866 Rs (-90°)
I made a 6-channel sample with 10sec of pink noise in L, then C, R, Rs, Ls. I then downmixed to DPL2 using <m11a> in Cool Edit. I then decoded back to 6-channels using WinDVD's DPL2 filter. The output channels were all very well separated with very little leakage.
I will upload the files once I organize some webspace.
Rockaria
23rd June 2006, 05:11
I have tested the matrix <m11a>:
Lt = L + 0.7071 C - 0.866 Ls (-90°) - 0.5 Rs (-90°)
Rt = R + 0.7071 C + 0.5 Ls (-90°) + 0.866 Rs (-90°)
It's actually <m11>...
Thanks for the first verification(besides me) of the seperation quality of the 90° shift model.:thanks:
The model from the specise's linked DPL 1 document(now I can open it) seems actually based on <m11a>.
I am not familiar with the CoolEdit's PhaseShift(which direction actually it shifts if we (can) ignore the align()), but I don't think the sign will make not so big differences, although it may actually create the different mixure of the images.
The constants used seems reasonable(1:3 on coef dB) and I believe can be used for the default values before any custom pannings(based on Dolby's doc I linked recently, I read we sometimes need to adjust the panning by the ch-signal types used).
My rough interpretation(just to match) might be helpful for dB unit users(based on DPL 1+) :
. -3dB == 0.5 ~= 0.7071
. (-3dB, -9dB) == (0.5, 1/8) ~= (sqrt(3/4), sqrt(1/4)) == (0.866 , 0.5)
http://en.wikipedia.org/wiki/Decibel
raquete
23rd June 2006, 05:41
hy boys,
i will this help the thread:
from "Dolby Surround Pro Logic II Decoder Principles of Operation.pdf (Dolby.Inc)"
(is too big but very clever)
please pay atention in the green retangle(and red lines) and in the text about "antiphase signal".
http://img471.imageshack.us/img471/8374/018by.png http://img217.imageshack.us/img217/6011/026oy.png http://img217.imageshack.us/img217/2186/034zi.png http://img69.imageshack.us/img69/6267/049zw.png
thanks.
Rockaria
23rd June 2006, 06:31
Yes, invert() == '-' in the bleo's formula! Everybody seems to agree now.
That must be a good visual guide, better than the links(several times here).:D
One thing to remind is that it's for DPL 1. For DPL 2, I found the 1:3~4 ratio on the coef 1: coef 2 is optimized.
It is similar to the -4.5dB or -6dB attenuation ratio used by bleo(btw, what is the source?).
Another thing is that depending on the mix process(mix all together or by steps) and algo used, the resulting ratio will be changed.
One last thing is that there are actually some more formal/personal matrices used beside the Dolby's -3dBs.
I think at least we need to evaluate them depending on the mix processes, effects & tools we choose. Thanks.
3dsnar
23rd June 2006, 06:42
Ls' and Rs' are separated by their different amplitudes (the 0.866 and 0.5 coefficients) in Lt and Rt.
I have tested the matrix <m11a>:
Lt = L + 0.7071 C - 0.866 Ls (-90°) - 0.5 Rs (-90°)
Rt = R + 0.7071 C + 0.5 Ls (-90°) + 0.866 Rs (-90°)
I made a 6-channel sample with 10sec of pink noise in L, then C, R, Rs, Ls. I then downmixed to DPL2 using <m11a> in Cool Edit. I then decoded back to 6-channels using WinDVD's DPL2 filter. The output channels were all very well separated with very little leakage.
I will upload the files once I organize some webspace.
Bleo, can you do the same with the signals from this post:
http://forum.doom9.org/showthread.php?p=838650#post838650
It will be also possible to determine phase changes after separation (to see wether the decoder compensates the 90 deg. phase shift), etc.
I will provide all the pictures and analysis results
which I will perform in Matlab.
Please let me know,
3d
raquete
23rd June 2006, 07:08
Yes, invert() == '-' in the bleo's formula! Everybody seems to agree now. this is not what i understood. :stupid:
the S input is also reduced by 3dB,but before being divided equally between Lt and Rt,the signal has 90-degree phase shift applied relative to L,C, and R. Finally, the S signals are carried in Lt/Rt with opposite polarities
correct me please if i'm wrong but what i understood about "opposite polarity" is:
S(-90) to Lt
S(+90) to Rt
:thanks:
Rockaria
23rd June 2006, 07:50
Well, suppose if you have only -90 deg shift logic.
If we can ignore the delay(cycle advance) :
+90 == -90.invert() == -90 - 180 == -90 + 180 : any routine is ok
If we count the delay :
+90 == -90.invert() == (-90 + 180).align()... : align() is introduced
So :
S(-90).invert() == S(+90), same result but the process included to get the final result is differernt
The reason why I do not want to optimize the process(formula) is to respect(identify) the original encoding & decoding process.
hilbert(). invert() instead of simple +90
Mathmatically, it looks stupid. But sometimes it's very useful to identify the hidden process/logic.
Also If you continue to optimze them from person to person any further for any reasons, you might even get the 180 deg(invert-only) model
or to an extreme, the simple downmix, progressively detached from the reality.
ursamtl
23rd June 2006, 13:52
hy boys,
i will this help the thread:
from "Dolby Surround Pro Logic II Decoder Principles of Operation.pdf (Dolby.Inc)"
(is too big but very clever)
please pay atention in the green retangle(and red lines) and in the text about "antiphase signal".
Exactly!
No offense to the participants, but seriously this thread has been going on for days now and dancing around the same point. Dolby says that it encodes the surround signal with a -90deg phase shift for the Lt channel and a +90deg phase shift for the Rt channel. This applies to DPLII as well as DPLI.
The difference between DPLI and DPLII are (as I understand them) the following and only the following:
DPLI: surrounds limited 100Hz to 7kHz with special Dolby B noise reduction
DPLII: surrounds full range
DPLI: one mode of operation with 20ms delay on surround to take advantage of the Haas effect so that sounds appear to be coming from the front channels.
DPLII: two modes of operations. Movie Mode, which features the same 20ms delay as DPLI and an active Center. Music Mode, which does not have a delay between front and rear. IIRC, it also does not have a center speaker output but relies on phantom imaging for the front Center.
DPLI: single channel surround output (may be spread across two speakers but still mono).
DPLII: stereo surround output using a steering logic to ensure that the matrixed surround channel is separated into two distinct channels.
Notice that there doesn't seem to be any real difference in the encoding method. There may be commercial encoding software available for DPLII, but Dolby's documentation seems to indicate that the differences between DPLI and DPLII are all on the decoding side.
Now, if you look carefully at all of the decoding differences, there are only two that are not quite clear or relatively easy to duplicate in hobby/project software: the Dolby B noise reduction and the steering logic. As for the noise reduction, if one is interested more in DPLII, it's not very crucial. The steering logic, however, IS essential.
Therefore, IMHO, you folks would more quickly get some good results if you simply accepted Dolby's recommendation to encode the surrounds with the -90deg and +90deg phase shift and then focus on the steering logic. You'll notice that Bleo used the +-90 in his most recent post and the results were very good.
Sorry if this message seems a bit harsh, but it seems this thread at times is getting bogged down or going around in circles with arguments over whether a -- is the + or whether inverting a channel is the same as a -, etc., etc.
So here's the real challenge for you folks. Figure this out:
(thanks to Raquete for the image)
http://img69.imageshack.us/img69/6267/049zw.png
Regards,
Steve.
3dsnar
23rd June 2006, 14:18
Therefore, IMHO, you folks would more quickly get some good results if you simply accepted Dolby's recommendation to encode the surrounds with the -90deg and +90deg phase shift and then focus on the steering logic. You'll notice that Bleo used the +-90 in his most recent post and the results were very good.
Exactly.
However do we all agree, that if the surround channels are already +-90 deg shifted the simple downmixing method (with only sign inverting) is correct?
So the best automated way would be to scan the whole signal and determine phase relation between fronts and surrounds (and adjust it if necessary) and then proceed with downmixing.
Rockaria
23rd June 2006, 19:11
Boys go dancing?
Sorry if this message seems a bit harsh, but it seems this thread at times is getting bogged down or going around in circles with arguments over whether a -- is the + or whether inverting a channel is the same as a -, etc., etc.
<<it seems to produce a -90° phase shift, so for the +90° shift, you'll need to invert the signal
...
Notice that there doesn't seem to be any real difference in the encoding method.
It's more than harsh. What a supid boy try to make another +90° shift logic when they already have -90°?
Raquete, was my answer above enough for your inquiry or too enough(at least it's my 20 years old system analysis rule)?
ursamtl, so you decided to completely forget the cross channel-coefficient-imbeding which is the major different feature for the DPL II decoder steering logic?
Nobody has explained the coef ratio(1:3) but started to use it since 3 years ago. Keeping the silence is a good virtue on legacy thing?
If you don't understand the black box clearly, at least have a correct approach to figure out the uncertain things prior to any attempt of misleading conclusions. That's the basic virtue of any participants.
Why nobody except me and now bleo test the new models and verify the seperation quality it produces? Do you think you can prove it with the mouth only?
ursamtl
23rd June 2006, 19:34
Exactly.
However do we all agree, that if the surround channels are already +-90 deg shifted the simple downmixing method (with only sign inverting) is correct?
So the best automated way would be to scan the whole signal and determine phase relation between fronts and surrounds (and adjust it if necessary) and then proceed with downmixing.
I agree that the +-90deg shift should occur in the encoding phase only since that's what Dolby describes in their documentation. I also agree with the decoder using the surround and the inverted surround since this is also documented by Dolby.
What I don't get is the need to scan for phase relationships and adjust. Simply downmix using the +-90deg phase shifts, since downmixing is in essence the same thing as encoding in the context we're discussing.
ursamtl
23rd June 2006, 19:52
Rockaria,
There were two reasons I mostly stayed silent on this debate. Number one, I have very little interest in Dolby Pro Logic I or II. Number two, the tone of this thread has often been quite negative and combative. I tried to point out earlier that we can all be passionate about our positions, but let's relax and enjoy the experimentation. However, you seem to take this far too personally. As long as what I said seemed to back up your "side" of the argument, you referred to me as being unbiased. If, however, I write something you don't like or that doesn't back your "side", you suddenly seem to get angry! Dude, relax! Don't take this as a personal attack on you. If you do, then that's your problem. Don't turn the thread into your personal battle.
If you read what I wrote, you will see that I pointed out that steering logic is what you should be focusing on. Obviously I didn't forget it! If I don't join in testing all the different coefficients it's because I don't want or need to. I played around with Dolby Pro Logic I and II over two years ago and have no further desire to. I jumped into this discussion mostly to nudge it in a more positive direction. That seems to be impossible if anyone disagrees with you because you take it far too personally! If you do, it just makes the debate unpleasnt for everyone concerned. Plus, if the thread gets too negative, the moderators will simply close it. That certainly won't help anyone gain a better understanding!
Boys go dancing?
It's more than harsh. What a supid boy try to make another +90° shift logic when they already have -90°?
Raquete, was my answer above enough for your inquiry or too enough(at least it's my 20 years old system analysis rule)?
ursamtl, so you decided to completely forget the cross channel-coefficient-imbeding which is the major different feature for the DPL II decoder steering logic?
Nobody has explained the coef ratio(1:3) but started to use it since 3 years ago. Keeping the silence is a good virtue on legacy thing?
If you don't understand the black box clearly, at least have a correct approach to figure out the uncertain things prior to any attempt of misleading conclusions. That's the basic virtue of any participants.
Why nobody except me and now bleo test the new models and verify the seperation quality it produces? Do you think you can prove it with the mouth only?
Rockaria
23rd June 2006, 20:52
I'm beginning to agree with 32, ****************
I also start to change my view on your approches or D****, ****IS COM****L.
Before mentioning the 'like, dislike, personal, angry, jump on', what you need to understand is we must keep the mutual respects to the anonymous.
If you have backuped any of my views, it should be because you think it's right, nothing else. Otherwise, it's just a child game.
As you admitted in the beginning, you may not have enough knowledge on DPL II because of several reasons, compared to your other experiences.
I liked the unbiased opinions from your experiences on this audio area. But as for the system analysis on the uncertain area is what I am majored in.
You started to have & expose the WILL to drive. To my point of view, by asserting the DPL II encoding SAME as DPL I and excluding the rear-stereo seperation plan, even though I explaind so many times about the rear stereo-coefs, you started dancing.
It's more than dangerous in my views, almost same as regarding the 90 deg shift concept to 180 deg one.
It's like trying to explain the water with H2 only without considering the Oxygen.
At least, I skimmed the Dolby's docs twice before. Why do you think I didn't know the steering logic? simply because I aknowledged you by 'pin-pointing' & 'no spec yet,ursamtl' ? What if I say I't was my way of confirmation?
Don't look down on people who might have better experiences on certain area, which is also the reason we can help each other.
As far as you keep the unbiased opinions, I will acknowledge you. But if you ever attempt to mislead in the approach, I think it's my duty to point them out. That's all. So you can think you are warned by me Dude.
About the 'debate, combat,,,,' : Have you seen myself accusing people first & at all? I just wanted to protect myself & VIEW and promote the possible debates to conversations. This thread started to prove the legacy is better than the Dolby's spec, yeah over-protecting the VIEWS related, not in a scientific way at all. And I think we have almost calmed down it to the correct direction, until you started it suddenly.
scharfis_brain
23rd June 2006, 21:20
One question: how 'unbiased' (as you understand it) are YOU?
ursamtl
23rd June 2006, 21:24
One question: how 'unbiased' (as you understand it) are YOU?
Hehe. Good one. Reminds me of that line in Orwell's Animal Farm about all animals being equal, except that some are more equal than others!
Rockaria
23rd June 2006, 21:26
One question: how 'unbiased' (as you understand it) are YOU?What is the reason asking me? Would you elaborate some more on it, followed by your own confession?
Rockaria
23rd June 2006, 21:27
Hehe. Good one. Reminds me of that line in Orwell's Animal Farm about all animals being equal, except that some are more equal than others!
Boy, I totally changed now. Play the animal kids game here.
bleo
23rd June 2006, 22:05
However do we all agree, that if the surround channels are already +-90 deg shifted the simple downmixing method (with only sign inverting) is correct?
So the best automated way would be to scan the whole signal and determine phase relation between fronts and surrounds (and adjust it if necessary) and then proceed with downmixing.
I agree. The 6-channel materials that we are downmixing to DPL2 are mostly DD5.1 movie soundtracks. The Sonic Foundry Soft Encode DD5.1 encoder by default applies 90° phase shift to the surround channels. So if we were to downmix such a DD5.1 soundtrack to DPL2, the input channels will be: L, R, C, Ls (90°), Rs (90°); and the matrix that we actually use will be:
Lt = L + 0.7071 C - 0.866 Ls - 0.5 Rs
Rt = R + 0.7071 C + 0.5 Ls + 0.866 Rs
Substitute the input channels into the matrix and the total result will be:
Lt = L + 0.7071 C - 0.866 Ls (90°) - 0.5 Rs (90°)
Rt = R + 0.7071 C + 0.5 Ls (90°) + 0.866 Rs (90°)
which is exactly what we want :) and coincidently, what we've been doing for the past several years, blissfully unaware of 90° phase shifts... :)
Now the question is how do we know if the surrounds of a DD5.1 soundtrack are already 90° phase shifted? I don't think it's easy without having the original 6-channel source from before DD5.1 encoding. Otherwise, we don't know the correct phases of the surround channels, or even what relationship they are supposed to have with the fronts (they may contain completely different material).
bleo
23rd June 2006, 22:16
@raquete: I'm still trying to work out which should be + or - 90° in the matrix. It is a bit more complicated because one of the surrounds gets inverted in the output depending on whether you are using Movie or Music mode.
@3dsnar: I downloaded your test samples, but there's a problem since the WinDVD DPL2 decoder silences the first 2 seconds of its output. I will create a similar test sample myself, so please tell me if there are any specific parameters that I should include.
bleo
23rd June 2006, 23:48
I made a 6-channel test sample that contained 3 sec silence then 3 sec sine waves in the L, then C, R, Ls and Rs channels. I downmixed it to DPL2 using Cool Edit and the matrix:
Lt = L + 0.7071 C - 0.866 Ls (+90°) - 0.5 Rs (+90°)
Rt = R + 0.7071 C + 0.5 Ls (+90°) + 0.866 Rs (+90°)
I then upmixed it using the WinDVD DPL2 decoder in Movie mode.
Results:
- The entire output was delayed by 5msec.
- The surround output channels were delayed by an additional 15msec.
- The surround output channels were both -90° phase shifted relative to the original inputs.
In a second experiment, I encoded a 6-channel test sample that contained sine waves to DD5.1 using Sonic Foundry Soft Encode. It phase shifted the surround channels by +90°
Rockaria
23rd June 2006, 23:54
bleo, thanks for the detailed test results.
I agree. The 6-channel materials that we are downmixing to DPL2 are mostly DD5.1 movie soundtracks...
Indeed, I read the SoftEncode has the '90 deg phase shift option' for DD5.1 rear channels, which later can be invert-downmixed to DPL or DPL II(cross coefs mixed).
This filter should generally be used whenever encoding a multichannel signal unless
it is known that the 5.1-channel source does not contain point-source element pans.
should use-unless-does not contain == should use-if not-not-contains == should use-if contains?
However here arise some issues:
. rears shifted for later possible DPL invert-downmix only from DD5.1 having the point-source element pans
. does the DD.5.1 decoder detect the 'rears shifted' info anyhow(meta tag or stream) and play the rears in normal phase?
. if not, doesn't it have any negative effects in volume and panning in AC3 decoding itself?
I believe it's still an exceptional case limited to DD5.1 encoding & decoding/downmix and the solution(if any) should be sought there.
The full DPL II solution is also good for PC 6ch music play through the analog-2ch-DPL II receiver(FFDShow), let alone the multi platform, compact size use, with relatively decent seperation fidelity & overall decent quality. And it at teast does not prove the Legacy is better than the Dolby's spec.
My simple idea for the detection of the shifted image is : it has generally low altitude(volumes) than the non shifted one, which can be verified by shfting it visually.
So the general consensus is the FULL implementation of the DPL II spec and possibly have the option to detect the originally 90 deg. shifted rears and suggest not to apply again?
Rockaria
24th June 2006, 02:36
@Raqueue and anybody still misunderstanding the +-90 shifts on the Dolby's doc,
The S input is also reduced by 3 dB, but before being divided equally between Lt and Rt,
the signal has 90-degree phase shift applied relative to L, C, and R.
Finally, the S signals are carried in Lt/Rt with opposite polarities
(note the "-" sign in the summing stage feeding the Lt output). Clearly, it describes there are practically two steps which you missed by just looking at the result : diagram.
When the stream is finally mixed to the Rt(or Lt, one side), it will be invert-mixed.
Also the cost of invert logic(circuit) is relatavely ignorable than the all-pass filter, it's not desirable at all to develop a redundant +90 shift logic, which is exactly what Dolby described in the document.
You didn't need to use the misleading 'boys enforced by Exactly recursively' on this small discovery which I evaluated already.
Rockaria
24th June 2006, 03:14
A snippet from the Surcode DPLII Encoder Manual.
9.4 Invert Right Surround
In Pro Logic II, the Ls and Rs outputs are intended to be out of phase by
180 degrees for surround steered inputs. However, if the decoder is set to
Movie mode, the polarity of the Rs output is inverted in order for the Ls
and Rs outputs to be in phase. This improves phantom rear center
imaging.
A snippet from the Surcode DPLII Encoder Manual.
This would make sense since DPLII music mode is generally promoted by
Dolby as intended for simulating 5-channel surround when listening to
regular stereo music. As such, they don't want accurate phantom center
imaging in the rear; they want diffuse ambience.
also I want to add that there won't be any phantom center surround anymore
when inverting one of the surround channels before downmixing. The centered
surround will be sent to the front center instead.
So this must be a mistake in the SurCode Manual or even in the SurCode
program itself.
Did I read it wrong? I wonder why the movie mode issue suddenly changed to the music mode?
Does anybody think this kind of posts any helpful to promote the positive contributions?
I actually started to worry from here.... The Good Will seems to have decided to stay as a gentleman as before.
bleo
24th June 2006, 04:29
This filter should generally be used whenever encoding a multichannel signal unless it is known that the 5.1-channel source does not contain point-source element pans.This is confusing English. It means:
- If the source contains point-source element pans, you must use 90° phase shifts on the DD5.1 surround channels.
- If the source definitely does not contain point-source element pans, you can choose to turn off the 90° phase shift filter. (It should not hurt to leave it on).
- By default, the 90° phase shift filter is on.
- Note that most of the time, you don't know whether the source contains point-source element pans.
Also in section 4.7.5: Surround Channel 90-Degree Phase-Shift it says:
The 90-Degree Phase-Shift parameter should always be left enabled except under specific conditions. These include, but are not necessarily limited to, system calibration, encoding of certain test signals, and in the extremely rare case when the discrete playback of highly coherent program material may be compromised.
I take "highly coherent program material" to mean that the surround channels must keep the same phase relationship with the fronts. There is an example of this back in section D.1:
if the source was recorded using five discrete microphones placed in the corners of an auditorium, there is no panning between channels and the filter could be safely disabled.The microphones are already setup in positions matching DD5.1 playback speakers. The phase relationships should not be altered.
Note that this entire post is about 90° phase shifting the surround channels when encoding DD5.1.
raquete
24th June 2006, 04:48
@Raqueue and anybody still misunderstanding the +-90 shifts on the Dolby's doc,...
i'm not misunderstanding anything because i don't know anything about DPL I or II.
i don't know "nothing" (and who don't know nothing don't understand anything and can't misunderstand anything :p hey,i'm philos-off :p )
Rockaria trust me,i'm learning here with you all,this thread is like a "school" for me.
Rockaria
24th June 2006, 06:48
Yeah, that's what I finally figured out by excluding the two 'NOT's. I don't like this style.(two negatives==positive==invert().invert()) : severe waste of my CPU cycles
- If the source definitely does not contain point-source element pans, you can choose to turn off the 90° phase shift filter. (It should not hurt to leave it on).
This type of corner-positioned mic sources are considered to have (invert)phase shifted rear channels. So there might still necessary some more inverting process to cross mix the rear channels for later DPL II downmixes.
Also if we consider the processing of the 90 phase shifts on the rears, without any mastering, the rears might have 0 & 180 phases, which may require some verifying before deciding to enable. Also I am not sure if it's OK(fidelity) for the DD5.1 playback.
The Dolby family codecs seems asking a strongly forced compatibility as using all 'should's & default-on option.
But what is certain is that we can safely regard it as a simple suggestion only when we consider the later optional DPL downmixing.
. The 90-Degree Phase Shift filter provides a means for an encoding engineer to create a multichannel Dolby Digital bitstream that can be downmixed to a Dolby Surroundcompatible Lt/Rt output.
. The SoftEncode has the default-on option for 90-Degree Phase Shift
. But I strongly believe we SHOULD disable the feature for a wider compatibility
Personally, I think it's an undesirable forced feature for either DD5.1 playback or no-custom-downmix., I will choose 'disabled' if I have to DD encode.
The issue of "highly coherent program material" implies there can be some unexpected side effects(fidelity loss) by the unrecovered DD5.1 playback from the optional feature. I will not choose the 'should be enabled option for some rare later DPL downmix opportunity' for this reason too.
Thanks for your analysis bleo, it helped me a lot for my evaluation.
Rockaria
24th June 2006, 06:56
What we've learnt from the 'phase shifts' :
i don't know "nothing" == you know everything ( you know everything when you are interested in, in the end or now if aligned)
positive+positive = 45+45|155 = 90|200 = 90|-160 = positive|negative = +-positive : so no right(x).right(y) pls. it's confusing..
@raquete, what I count is you love music, since long ago. that's not nothing, believe me:cool:
3dsnar
24th June 2006, 09:19
I made a 6-channel test sample that contained 3 sec silence then 3 sec sine waves in the L, then C, R, Ls and Rs channels. I downmixed it to DPL2 using Cool Edit and the matrix:
Lt = L + 0.7071 C - 0.866 Ls (+90°) - 0.5 Rs (+90°)
Rt = R + 0.7071 C + 0.5 Ls (+90°) + 0.866 Rs (+90°)
I then upmixed it using the WinDVD DPL2 decoder in Movie mode.
Results:
- The entire output was delayed by 5msec.
- The surround output channels were delayed by an additional 15msec.
- The surround output channels were both -90° phase shifted relative to the original inputs.
In a second experiment, I encoded a 6-channel test sample that contained sine waves to DD5.1 using Sonic Foundry Soft Encode. It phase shifted the surround channels by +90°
OK, Bleo, thanks for the test.
1) I assume that the sinusoids (in all channels) had the same amplitudes before downmixing. Did they have the same amplitudes (in all channels) after decoding (if yes, this implies the weights were correct, because they must be an arbitrary values assumed by the decoder in order to recreate the original amplitude of all channels).
2) The fact that the decoded signals were -90 deg shifted means that the downmixing equations should look like
Lt = L + 0.7071 C + 0.866 Ls (+90°) + 0.5 Rs (+90°)
Rt = R + 0.7071 C - 0.5 Ls (+90°) - 0.866 Rs (+90°)
(???)
3) It would be the best to use for example square waves (instead of sinusoids) because it will be possible to distinguish between time delay and phase shift applied to all sinusoidal partials (i.e. phase shift does not affect the waveform shape)
Also, my observation (I am curious if any of you agree) is that
since the decoder does not compensate the 90 deg. phase shift
it means that the process of preparing the audio material is something separate than the downmixing (it is like a recommendation how to prepare the audio material for downmixing)
So I think that NOT including the 90 phase shifts in the downmixing equation is fully correct, because the inverse process does not compensate that (assuming that the inverse process is supposed to attempt to recreate the original). The 90 deg. shift should be viewed as some audio preprocessing to assure the best possible performance. However treating it as a part of the equations implies using them, which is not true for all cases.
So the DPLII material preparation should be rather described as a two stage process
1) Audio preprocessing (to assure correct phase relations between fronts and surrounds)
2) Downmixing with the use of equations with only sign reversing.
But I guess this is a bit matter of taste...
------
And some short remarks to some other posts
A) It is not possible to determine whether the signal is, or is not 90 deg shifted (without knowing the original). I.e. any methodology will work in some rare cases, and in others will not
(consider for example pure sinusoids or noise)
B) It is possible to roughly calculate phase relation between pair of channels (e.g. between FL and SL or FR and SR). It is also possible to change the phase. It could be done for the total signal, or in overlapping frames (they would have to be pretty long though, e.g. 1 sec. or so in order not to produce artifacts by phase cancallation of overlapping segments, etc)
So it would be possible to scan through the signal and determine global phase relation between the fronts are surrounds, and correct it if necessary for DPLII downmixing. This would be a bit computationally expensive.
Also, it could be done locally (because phase can fluctuate - because the surrounds and fronts may contain different audio material, etc.) in 1 sec. overlapping segments. This would be a bit less computationally complex to the global phase adjustments (in order to determine phase you have to compute FFTs for each short time frame - e.g 2048 samples for 48000 Hz signal, with 50% overlapp).
But is it worth it to implement such complex algorithm? Especially because probably over 90% of movie soundtracs have already properly (for DPLII downmixing) phase shifted?
bleo
24th June 2006, 16:44
Hi 3dsnar, thank you for your comments.
1) I assume that the sinusoids (in all channels) had the same amplitudes before downmixing. Did they have the same amplitudes (in all channels) after decoding (if yes, this implies the weights were correct, because they must be an arbitrary values assumed by the decoder in order to recreate the original amplitude of all channels).The sine waves in the input were all -3.2dB. The sine waves in the output were as follows:
L' C' R' Ls' Rs'
L -6.2 -43.3 -69.5 -49.2 -47.8
C -43.2 -6.2 -43.2 -42.9 -42.9
R -69.5 -43.3 -6.2 -47.8 -49.2
Ls ~-50.5 ~-62 ~-53 -6.3 -35
Rs ~-57.5 -61.7 -50.8 -34.3 -6.3Looking along the main diagonal, we see that the outputs of L', C' and R' were decreased by -3dB, and the outputs of Ls' and Rs' were decreased by -3.1dB.
All other values off the main diagonal indicate channel leakage. Some are inevitable due to the steering logic. For example, L and R were pure sine waves present solely in their respective channels, yet they still leaked by a small amount to C', Ls' and Rs'. Others may be due to rounding errors in our downmix matrix coefficients and +90° phase shift. For example, there is a fair amount of leakage of Ls and Rs to the other surround output. Perhaps we can refine the matrix to reduce this.2) The fact that the decoded signals were -90 deg shifted means that the downmixing equations should look like
Lt = L + 0.7071 C + 0.866 Ls (+90°) + 0.5 Rs (+90°)
Rt = R + 0.7071 C - 0.5 Ls (+90°) - 0.866 Rs (+90°)
(???)I have not tested this yet, but I theorize that both surround outputs will then be +90° phase shifted relative to the original inputs (when using DPL2 Movie mode).3) It would be the best to use for example square waves (instead of sinusoids) because it will be possible to distinguish between time delay and phase shift applied to all sinusoidal partials (i.e. phase shift does not affect the waveform shape)I will use square waves in my future tests.
I will also respond to the rest of your (interesting) discussion when I have time... :)
raquete
24th June 2006, 18:59
I will use square waves in my future tests.
what about differents frequences between channel ? !
something like:
L = 440Hz -3.2dB
R = 1Khz -3.2db
and 2Khz -15dB inside L and R channels (as center).
:thanks: for your tests bleo.
3dsnar
24th June 2006, 19:26
OK, Bleo. Great!
So we have found out one thing for certain:
The weights values that you used in the equations are correct.
Also if we would use the following downmixing eqs:
(no 90 deg phase shifts)
Lt = L + 0.7071 C + 0.866 Ls + 0.5 Rs
Rt = R + 0.7071 C - 0.5 Ls - 0.866 Rs
we should obtain (roughly) the input 6 channel signal.
If you'll find a bit of time, it would be great to perform the
test with the above eqs, and based on square waves.
---
Good idea Raquete, although it is better not to use sine waves.
So square waves with different fundamental frequencies for different channels.
Rockaria
25th June 2006, 00:52
As discussed in the related posts, we are concerned if there are any commercial DVDs supporting this optional feature.
http://forum.doom9.org/showpost.php?p=844338&postcount=119 : The 6-channel materials that we are downmixing to DPL2 are mostly DD5.1 movie soundtracks
http://forum.doom9.org/showpost.php?p=844361&postcount=122 : I believe it's still an exceptional case limited to DD5.1 encoding & decoding/downmix and the solution(if any) should be sought there.
http://forum.doom9.org/showpost.php?p=844442&postcount=125
http://forum.doom9.org/showpost.php?p=844466&postcount=127 : The issue of "highly coherent program material" implies there can be some unexpected side effects(fidelity loss) by the unrecovered DD5.1
The capability of seperating this two processes is backed-up by identifying the two seperate steps of the DPL encoding :
http://forum.doom9.org/showpost.php?p=844418&postcount=123 : (Ls,Rs).hilbert() => (Ls.invert.mix, Rs.mix).
I believe the initial DPL II threads participants were fully aware of the limitation of the invert-only version.
The formula only had some potentials to be misinterpreted, forcing us to waste enormous energies even to get down to discuss the full implementation.
It must be a big discovery, if it ever turnes out that all/most existing DD5.1s are ready for the invert-only downmix.
They might be identified in the real world as :
. By the DPL compatibility mark on the movie jacket, not with the 2ch DPL but with 5.1ch DD : has anybody seen them?
. By receivers for 2ch speakers : I see only passive matrix downmix(cross mixed) or straight mix(not cross mixed)
. By dvd players for DPL II receivers : it's mostly provable if any such dvds exist.
. By licence permission for general transcode on PCs or other platform : definitely only on their certified devices.
. By the acknowigements by the s/w DPL II(invert-only) encoder developers : The SoftEncode suggested 90 deg shift issue is just recently introduced here
.
.
Anyway, I believe there should be no issues in concluding the final DPL II stream must contain the FULL DPL II encoding(invert & 90deg phase shifts). if we can prove the below assumption is INCORECT by resolving the issue 2(below).
http://forum.doom9.org/showpost.php?p=844489&postcount=129 : So I think that NOT including the 90 phase shifts in the downmixing equation is fully correct, because the inverse process does not compensate that....
<history>
[Jun24] looking for any evidences.
Rockaria
25th June 2006, 01:03
The process is of course : 2ch coefs-cross-mix(Ls,Rs)-> 90 deg shift(hilbert()) -> +-polarity mix(simple mix & invert mix) on Lt & Rt
Some facts from Dolby's doc :
1. the result should be : Ls(in phase : +90), Rs(out phase : -90)
"The Lt/Rt downmix sums the surround channels and adds them in phase to the left channel and out of phase to the right channel.
http://web.archive.org/web/20031206104650/http://dolby.com/pro/digaudio/pa.ma.1102.Standards.S.pdf P18 middle
http://forum.doom9.org/showpost.php?p=838744&postcount=42
2. The diagrams say to apply invert() to the Ls when mixing to Lt : Ls.hilbert().invert() == in phase(+90), so Ls.hilbert() == -90 deg shift.
http://www.dolby.com/assets/pdf/tech_library/44_SuroundMixing.pdf P8-1
Also it concurs the Dolby description in the same page :
Thus for instance at 1 kHz, the left and right paths through the Dolby
Model SEU4 give phase shifts of roughly –550 degrees, while the surround path, measured at the
right total output, has about 90 degrees more lag (approximately –640 degrees total).
-640 - (-550) = -90 for the 90 degrees shift mentioned
3. it also mostly concurs other tools' hilbert() directions
. mathlab : hilbert() = -90
. PhaseBugMono(0.322) = -90 : it advances on the time-axis but when aligned it start from the -90 phase image.
. CoolEdit : PhaseShift(90) = -90?
. Soft Encode DD5.1 rear shifted: ?? -> +90 decoded
. Surcode assumed : -90 => +90 decoded, *check below 4. Surcode by specise.
4. analysis of the decoder outputs.
. PowerDvd(invert-only) by tebasuna : Avisynth, BeSweet, Azid
http://forum.doom9.org/showpost.php?p=838939&postcount=45
. Windvd by bleo
http://forum.doom9.org/showpost.php?p=844358&postcount=121
CoolEdit, Windvd movie mode.
Lt = L + 0.7071 C - 0.866 Ls (+90°) - 0.5 Rs (+90°) => -90° decoded
Rt = R + 0.7071 C + 0.5 Ls (+90°) + 0.866 Rs (+90°) => -90° decoded
==>change all (+90°)s to (-90°)s==> CoolEdit : PhaseShift(90) = -90?
. Surcode by specise.
http://forum.doom9.org/showpost.php?p=843644&postcount=90
9.4 Invert Right Surround
In Pro Logic II, the Ls and Rs outputs are intended to be out of phase by 180 degrees for surround steered inputs. However, if the decoder is set to Movie mode, the polarity of the Rs output is inverted in order for the Ls and Rs outputs to be in phase. This improves phantom rear center imaging.
-Ls(-90), Rs(-90) -> Ls(+90), Rs(-90).invert() ==> Ls(+90), Rs(+90)
<issues>
. no reverse phase shift logic in the decoder? : 90 -> 0
<history>
[Jun24] more inputs to be updated until cleared out.
Rockaria
25th June 2006, 20:52
Dolby simply say -3dB for the C-mix and another -3dB for the S-mix in the DPL 1 docs.
If we include the LFE and rear STEREO to the formula, it will become somewhat more complicated.
Also depending on the encoding tools and mixing steps(stacking mix) the resulting ratio will be changed.
This will also include some of the essential effects such as normalizing(mostly on the final mix), the DRC(enough bit size, clipping control by pre-gaining), stacked mixing(grouping) and the custom panning if the default matrix does not yield the desired harmony.
1. 1:2 SoundPressure(p) ratio on the coefs :
http://forum.doom9.org/showpost.php?p=145721&postcount=1 , same one as Wikipedia
Lt = 1.0*L + 0.707*C + 0.707*LFE - 0.8165*Ls - 0.5774*Rs
Rt = 1.0*R + 0.707*C + 0.707*LFE + 0.5774*Ls + 0.8165*Rs
. residuals
. resulting max linear ratio = 3.808
. max(safe) pre-gaining 1/3.808 = 0.2626 == -11.6dB == 10log(0.2626^2)
2. 1:3 SoundPressure(p) ratio between the coefs by bleo :
http://forum.doom9.org/showthread.php?t=57988 : good so far for existing s/w encoders
Lt = L + 0.7071 C + 0.7071 LFE - 0.866 BL - 0.5 BR
Rt = R + 0.7071 C + 0.7071 LFE + 0.5 BL + 0.866 BR
. 0.7071 == -3dB, -0.866 == -1.25dB, -0.5 == -6dB
. resulting max linear ratio = 3.78
. max(safe) pre-gaining 1/3.78 = 0.2646 == -11.5dB == 10log(0.2646^2)
3. Dolby's way (-3dBs & 1:3 SoundPressure(p)) :
simple assigning of the 3:1 ratio to the coef dB value : (-3dB, -9dB)
http://forum.doom9.org/showpost.php?...4&postcount=73 bottom : good results
Lt = (L, -3dB(C, LFE), (-3dB Ls, -9dB Rs)<*)
Rt = (R, -3dB(C, LFE), (-9dB Ls, -3dB Rs)<)
changing it to the linear gain ratio format would look like :
Lt = (L, 0.7071(C, LFE), (0.7071 Ls, 0.3536 Rs)<*)
Rt = (R, 0.7071(C, LFE), (0.3536 Ls, 0.7071 Rs)<)
. when < : -90 deg shift, * : invert()
using the coef which satisfies the two conditions to have the effective -3dB attnuation
http://forum.doom9.org/showpost.php?p=846551&postcount=140
. 1:3 ratio on the Sound Pressures
. -3dB result by the opposite polarity
. it will have +3dB gain effect by reverse inverting the Ls DPL II decoding time
in dB format :
Lt = (L, -3dB(C, LFE), (-1.25dB(Ls), -6.02dB(Rs))<*)
Rt = (R, -3dB(C, LFE), (-6.02dB(Ls), -1.25dB(Rs))<)
in Sound Power format(P) : same as 2.
Lt = (1L, 0.7071(C, LFE), (0.866Ls, 0.5Rs)<*)
Rt = (1R, 0.7071(C, LFE), (0.5Ls, 0.866Rs)<)
. resulting max linear ratio = 3.78 == 11.55dB(10log(3.78^2)), max(safe) pre-gaining -11.55dB
. the custom panning can be applied to each channel(variable) and groups(Lf-Rf, Ls-Rs, F-S,...)
. the coef ratio also can control the F-S & Ls-Rs panning but not desirable because it also imples not enough seperation/cancellations
. the seperation plans are stacked by <, *, (1, 0.7071 + phantom center, 0.866, 0.5) for the input to the DPL II decoder
Also we can use this 'redirecting the LFE to the SubWoofer(instead of the center)' method :
http://forum.doom9.org/showpost.php?p=846578&postcount=141
4. stacking mix(channel grouping) : will be mostly useful when using the gui tools
. the FRONT, C, LFE is one logical group : FRONT, and the coefs-cross-mixed-rears is another group REAR
. These groups ideally can be mixed, tuned before the final mix with normalization.
. the gains will be changed in a stacked way
. pre gaining methods : averaging or peak
. batch encoding(mixing) will be useful with : default attenuations, peak-pre-gaining, and final normalizing
http://en.wikipedia.org/wiki/Decibel
http://www.phys.unsw.edu.au/~jw/dB.html
Decibel.10^(dB/10) == SoundPressure.p^(1/2) == SoundPower.P || Decibel.10^(dB/20) == SoundPower.P
SoundPower.P^2 == SoundPressure.10Log(p) == Decibel.dB || SoundPower.10log(P^2) == Decibel.dB
http://en.wikipedia.org/wiki/Root_mean_square
5. analysis of the decoder out amplitudes
. need to verify all the candidates on both(90deg shift, invert-only) models.
. the rears can be affected by the un recovered 90 deg shifts of the image
. it will be strongly affected by the seperation plans(models)
. especially by the seperation/cancellation order(rear->center+LFE->front), dynamically by the steering logic enhanced by the servo feedbacks.
. steering logic might be allocating the ratios from the predefined gains(+3dBs) to the channels regardless of the input attenuations.
it turned out that the decoder does not control the decoded channel gains. if you give the ratio 0 on a channel, it produces no sound on the channel(except the phantom center), making the gain control independant from the decoder(transparent).
. In this case, analysing the decoded volumes might not be that useful, relying on the speakers would have the higher priority.
a. a result from bleo http://forum.doom9.org/showthread.php?p=844619#post844619
. tested with a sign wave, <m11a>
. may require some more tests with fully complex streams :
<issues>
<history>
[Jun25] some more refinings of the gains, stacked mixing, formula, analysis of the decoder output
[Jun25] modified 3.a., daft coef formula by gain & coef dB ratio 1:3, typos
[Jun26] refined the model <m11>, it's attenuation variables & coef constants
[Jun27] the gain control on the channels is tranparent to the decoder(except the phantom center)
[Jun29] final(?) modifications on the models
bleo
26th June 2006, 02:47
My findings on Issue #3:
If the difference between the coefficients for Ls and Rs is:
< 4.77dB then Ls will leak into Rs' and Rs into Ls', so the surround channel separation will be poor.
> 4.77dB then Ls will leak into L' and Rs into R', so the surround channels will spread too wide and 'wrap around' to the fronts.
= 4.77dB then Ls' and Rs' will have maximum channel separation without leaking into L' and R'.
Secondly, if the coefficients for Ls^2 + Rs^2 are:
> 1 then the volume of Ls' + Rs' will be louder than Ls + Rs.
< 1 then the volume of Ls' + Rs' will be quieter than Ls + Rs.
= 1 then the volume of Ls' + Rs' will be equal to Ls + Rs.
So the coefficients should be:
0.866, 0.5 = SQRT(3/4), SQRT (1/4) = -1.25dB, -6.02dB
Rockaria
26th June 2006, 03:57
Ignoring the phase & the polarity and counting one channel only,
a. If we give 0dB(1) attenuation to the rears :
T = mix(F, -3dB C,-3dB LFE, S), when S= (coef1, coef2)
= mix(F, 0.7071 C, 0.7071 LFE, S)
. resulting max linear ratio = 3.4142 == 10.665779 dB
. if we exclude rears : 2.4142 == 7.6554649... dB, so the difference(if we can do -) = 3.0103...dB
b. If we allocate linear ratio 1 S to the coeffs in 1:3 ratio : (sqrt(3/4), sqrt(1/4) = (0.8660, 0.5)
T = mix(F, 0.7071 C, 0.7071 LFE, 0.866 S1, 0.5 S2)
. resulting max linear ratio = 3.4142 - 1 + 0.5 + 0.866 = 3.7802 == 11.550295555...dB
if we exclude the rears : 2.4142 == 7.6554649... dB, the difference(if we can do -) = 3.8945...dB
if we count the coefs only : 0.866 + 0.5 = 1.366 == 2.7090 dB
I was able to get your 4.77dB by : 6.02dB - 1.25dB = 4.77, although I am not sure what it means.
But as you see, it keeps the 1:3 coef ratio(that's why I said reasonable before), but assign & allocationg 0dB(1) between the coefs.
I just didn't want to duplicate it in the group 3 to respect the original efforts.
As far as the rear coef ratio(allocating) keeps 1 : 3~ dB, it seperates well between the rear stereo(no residuals).
And the assigning of the attenuations on the rears(in your case 0dB == 1) is another issue(F-S panning maybe)...
specise_8472
26th June 2006, 11:00
At the end of the day, the downmixes are almost there.
BUT
Dolby Prologic 1/2 specifications are to have phase shifted +-90 degree (=180 degree) rear. So no amount of figure twiddling and tweeking will ever make spec files. If it was that easy to do this way, I think Dolby Laboratories and their technicians would have figured it out by now, and patented it.
A good starter patent and pointers to other patents is US 6,760,448. ALso includes steering and 3 rear channels.
Just my 2c worth.
Rockaria
26th June 2006, 12:00
Hmm, one of my role is straightening the looking-complicated stuffs in easy words & logics.
But if it looked too easy(except specise ;) ) then there's something wrong. So the brand new almost-FINAL full version follows.
if use -3dBs for the attenuations by the Dolby's spec
Lt = mix(L, -3dB(C, LFE), -3dB(Ls, 0.5774 Rs)<*)
Rt = mix(R, -3dB(C, LFE), -3dB(0.5774 Ls, Rs)<)
. when < : -90 deg phase shifts, * : invert(), so both rears are 180 deg off polarity on the same time axis.
changing it to the linear gain ratio format would look like :
Lt = mix(L, 0.7071mix( C, LFE), 0.7071mix(Ls, 0.5774 Rs)<*)
Rt = mix(R, 0.7071mix( C, LFE), 0.7071mix(0.5774 Ls, Rs)<)
. resulting max linear ratio = 3.53. max(safe) pre-gaining 1/3.53 = 0.2833 == -11.0dB == 10log(0.2833^2)
. might need to apply 0.7071mix( 0.7071C, 0.7071LFE) ==0.5mix(C, LFE)
the attenuations are variables to allow any custom pannings from the defaults, the rear coefs are constants for the REAR-STEREO seperation plan. so the seperation plans are stacked by <, *, (1, 0.5774) for the input to a ever-fading DPL II decoder to be reconstructed almost identical. I also feel full of coins in my pocket with some various convincing results on this useless stuff, but might also be useful for IIx or III.
for anybody interested in the automation, the -90deg phase shift logic must be fully emulated, not by the simple some freq delay, but by hilbert() or full ranges freq delays aligned dynamically in the buffer, thinking it's gonna be issue #4, some time later, possibly followed by the steering logic of the decoder.
/steering logic specs
Rockaria
29th June 2006, 08:32
bleo, I finally figured out the reason why the rear coeff ratios rule need to be your (0.866, 0.5).
Besides the mentioned rear seperation plan(1:3 on the coef attenuation),
there is one more condition that the result of the channel gain difference must be -3dB by the opposite polarity of the duplication imbeded on the other channels.
It will cancel the extra sound pressures in the main coef(also when played in the stereo mode) resulting in effectively -3dB attenuations.
<condition1 : 1:3 ratio on the Sound Pressures>
XX=3*YY (x= SQRT(3)y)
<condition2 : -3dB result by the opposite polarity>
0.7071^2(-3dB)=XX - YY
So :
3YY -YY=-.5
YY=.25, y=.5
XX=.75, x=.866
The Dolby's spec seems to have been changed in DPL II by the STEREO implementations on the rears.
I will update the model including an advanced LFE mixing method to be effectively played in the SUB Woofer(not to the center in reduced volume).
Also note that there are differences in the two decoding mode(passive in stereo mode & active in DPL II decoding mode)
By reverse changing the polarity in Ls (in the movie mode, although I don't notice the rear volume difference than the music mode), the total gain will be 0dB(3/4 + 1/4 p) on the rears(creating the rear phantom center, no extra +3dB gaining) unlike the center(the decoder will add 3dB).
But the passive decoder(in stereo mode) will still have the effective -3dB attenuations on the center & rears.
[Jun30] By some tests in the post http://forum.doom9.org/showpost.php?p=847043&postcount=149 :
. the stereo speaker out by the opposite polarity will somehow modify the overall volume
. so the output volume will be adjusted creating some diffussion by the opposite polarity.
. so the (0.866, 0.5) is not proved as the only correct coefs ratio yet, maybe some more test with both output(decoder, speaker) analysis are necessary
[Jun30] Considering the DPL II is also designed for the stereo playback, edstimating its total output volume would be useful to define the proper pre-mix attenuations. The DIFFUSION effect by the cross channel imbeded coef2 volume is considered having -1.5 ~ +1.5 P(sound power) by the distances inbetween making the overall stereo outputs -3DB ~0 dB attenuations from the original level.
The DPL II rear decoding would require some more analysis by measuring the decoder output levels by changing the attenuation values. So far (0.866, 0.5) seems to be the optimized value.
Rockaria
29th June 2006, 10:21
...
DPL II channels are full range(less LFE of course)
He threw up another meaningful words again, I was thinking about that too though, and finally had some time to messup with some richer harmony.
One of my receivers I am testing is not directing the mixed LFE to the Subwoofer. The center speaker is set to 'large'(for more rich bass) so setting it to small/medium might help to direct it to the subwoofer.
The fact that if I mix the LFE to Lf or Rf, it goes to the subwoofer, gave me a hint to mix the LFE to both Lf and Rf in invert & 90phase shifted way.
There can be several ways to achieve the LFE to be redirected to the SUBWoofer by :
. equaly positioning the LFE to the fronts(like center) : no redirection(in my setup), but when changed the center to 'small', it redirected
. panning the LFE to some unbalanced positions to the front(Lf-Rf) : some portion will go to the SUB, will also affect the unbalanced normal stereo playback
. phaseshift(Lf -90), invert(Rf) of the LFE : complicated, might have some side effects depending on the existing fronts LFE portion(cancellation)
.other combinations of these (normal, shift, invert) gave me less SUB Woofer volumes.
. mixing the LFE to the rears before the rear shift & invert process
Some changes from the latest model :
Lt = (L, LFE<, -3dB(C), (-1.25dB(Ls), -6.02dB(Rs))<*)
Rt = (R, LFE*, -3dB(C), (-6.02dB(Ls), -1.25dB(Rs))<)
changing it to the linear gain ratio format would look like :
Lt = (1L, 1LFE<, 0.7071C, (0.866Ls, 0.5Rs)<*)
Rt = (1R, 1LFE*, 0.7071C, (0.5Ls, 0.866Rs)<)
. resulting max linear ratio = 4.07 == 12.20dB(10log(4.07^2)), max(safe) pre-gaining -12.20dB
. this way gave richer LFE directed to the subwoofer
. might affect the existing LFE portions of the fronts(cancellation)
. it seems giving better rear seperations surprisingly(?)
. the PhantomCenter(of the LFE) might stay in the center though
. also it has very similar volume in the stereo mode playback as before(not too loud).
PreMixing the LFE to the Rears :
Ls = (Ls, LFE), Rs = (Rs, LFE)
dbFormat ratio :
Lt = (L, -3dB(C), (-1.25dB(Ls), -6.02dB(Rs))<*)
Rt = (R, -3dB(C), (-6.02dB(Ls), -1.25dB(Rs))<)
Sound Power ratio :
Lt = (1L, 0.7071C, (0.866Ls, 0.5Rs)<*)
Rt = (1R, 0.7071C, (0.5Ls, 0.866Rs)<)
. no negative effect, the safe-pregaining would be similar to above.
<conclusion>
. if the receiver supports the speaker size setting(or center-LFE-SUB redirecting) or portable use, pre-mix it to the center
. if the receiver supports rear-LFE-SUB redirect, pre-mix it to the rears.
. otherwise, test the front invert-shift-mix
<history>
[Jun 29] initial version tested with some 6ch music. : good results
[Jun 29] added a method premixing to the rears, conclusions.
bleo
30th June 2006, 02:29
I made a 6-channel test sample in Cool Edit that contained the following:
0-3 sec, silence
3-6 sec, L 300Hz square wave
6-9 sec, C 400Hz square wave
9-12 sec, R 500Hz square wave
12-15 sec, Ls 600Hz square wave
15-18 sec, Rs 700Hz square wave
I then downmixed it to DPL2 in Cool Edit using the matrix:
Lt = L + 0.7071 C + 0.866 Ls + 0.5 Rs
Rt = R + 0.7071 C - 0.5 Ls - 0.866 Rs
I then upmixed it back to 6 channels using WinDVD DPL2 decoder in Movie mode.
Results:
- The square waves were intact in the output indicating that no phase shift occurred.
- There was a 5msec delay of the entire output, and an additional 15msec delay of Ls' & Rs'.
- After accounting for this delay, the phases of all output channels were identical to the inputs, and there was no inverting either.
- The amplitude of the square waves fluctuated a bit at the channel switchover points (3, 6, 9, 12, 15 secs). This was probably an artifact of the steering logic.
Discussion:
This matrix produced output that was closest to the original input. The phases were the same and the relative channel volumes were the same with an overall reduction of -3dB. However this is NOT a DPL2 spec. compliant downmix. A discussion on 90° phase shifts will follow...
bleo
30th June 2006, 04:01
Let's say we forget about DPL2 and DD5.1 for the moment. Then we take a 5-channel soundtrack and phase shift both rears by 90°. What is the result to the listener?... In most cases, nothing! Especially for movie soundtracks, this is because there is usually no phase correlation required between the fronts and the rears. Furthermore, it would be difficult to keep such a phase correlation intact because the listener does not necessarily sit in the exact center of all 5 channels, and the rears are usually time delayed as well. (This difficulty may have been why quadraphonic sound did not become popular).
Now for DPL2, the rears MUST be 90° phase shifted to DEcorrelate them from the fronts in the downmix. Otherwise, if C and S are identical, they will cancel each other out on either Lt or Rt of the downmix. Once they cancel out in the downmix, you cannot get them back--this is the disadvantage of matrix encoding compared to discrete multi-channels.
So you 90° phase shift the input rears to make sure they are not cancelled out in the DPL2 downmix. After you decode back to 5 channels, the output rears are still 90° phase shifted. What is the consequence to the listener? Again, nothing! The input & output rears do not need to have any phase correlation with the input & output fronts.
Now for DD5.1, why do the Dolby specs say to 90° phase shift the rears? Firstly because it has no consequence for a listener using 5.1 channels--the rears do not need to have any phase correlation with the fronts. Secondly (and more importantly for our discussion), when downmixing DD5.1 to DPL 1 or 2, PRE- 90° phase shifting the rears allows the DD5.1 DEcoder to downmix using a very simple arithmetic matrix:
Lt = L + 0.7071 C - 0.866 Ls - 0.5 Rs
Rt = R + 0.7071 C + 0.5 Ls + 0.866 Rs
or
Lt = L + 0.7071 C + 0.866 Ls + 0.5 Rs
Rt = R + 0.7071 C - 0.5 Ls - 0.866 Rs
(signs to be determined by further experiments)
The final DPL2 decoded rears will of course be 90° phase shifted compared to the original sountrack before DD5.1 encoding, but have the same phase as the DD5.1 output. What is the consequence to the listener? You guessed it--nothing!
So when downmixing a DD5.1 soundtrack to DPL2, use one of the above matrices. DO NOT 90° phase shift the rears--it should have already been done when encoding to DD5.1.
When SHOULD you apply 90° phase shifts to the rears when downmixing to DPL2? These are mostly unique cases:
- If your multichannel material was never encoded to DD5.1 (e.g. you created it yourself).
- If you want to experiment...
raquete
30th June 2006, 04:52
Now for DD5.1, why do the Dolby specs say to 90° phase shift the rears? Firstly because it has no consequence for a listener using 5.1 channels--the rears do not need to have any phase correlation with the fronts.
sorry,i don't agree and it's nothing have to do with DD5.1,follow me:
http://img425.imageshack.us/img425/9361/controlroom2bf.png
think that you listen the sound and not decoding
q1
what happen if the surround left and right have the same phase as front left and right?
a1- cancellation!
q2
what if the sL and sR have +90 degrees?
a2- cancellation!
q3
what if the sL and sR have -90 degrees?
a3- cancelation!
q4
what happen if sL have +90 and sR -90 degrees?
a4-
the sL is 90 degrees from the L channel.
the sR is 90 degrees from the R channel.
the sL and sR are now 180 degrees
no more cancellations and:
as show the picture,the surrounds are placed round 90 degrees from the frontal channels this is why they are turned 90 degrees and one of them with +90 and other with -90 to don't have cancellation because one speaker in front other sounding with the same phase will give.... cancellation.
this is what i understood(with experience) why they used 180 degrees between surrounds and at the same time 90 degrees from the fronts!
3dsnar
30th June 2006, 06:50
Raquette,
we are talking about 90 deg phase shift of the waveform
(or shifting by PI/2 in radians). It has not much to do with angles between speakers ;)
raquete
30th June 2006, 07:17
@ 3dsnar
It has not much to do with angles between speakers
but have.
maybe i was unclear,let me try to explain better:
has no consequence for a listener using 5.1 channels--the rears do not need to have any phase correlation with the fronts.
if the surround channels have the same phase(no matter if phase 0,+90 or -90) and one speaker is in front of the other,you get cancellation.
as the surrounds are placed in the left and right sides,they turn his phases in (+ and -)90 degrees in the audio too.this result in differents phases between the surrounds and no cancellations will happen.
this not means that the sound with turned phase sounds better
then,if it not sounds better,we have to questions:
1-why they turn the phases?
because one speaker in front other playing the sound with the same phase give cancellation
2- and why -90 and +90?
because the surrounds are nearly (or perfectly) 90 degrees from the central channel(and can't have the same phase)
(of course,poor english here)
later i(will if needed)post another picture showing how and why the surrounds have to stay round 90 degrees from the frontals and 180 degrees(between sL/sR)
Rockaria
30th June 2006, 11:31
I made a 6-channel test sample in Cool Edit that contained the following:
0-3 sec, silence
3-6 sec, L 300Hz square wave
6-9 sec, C 400Hz square wave
9-12 sec, R 500Hz square wave
12-15 sec, Ls 600Hz square wave
15-18 sec, Rs 700Hz square wave
I then downmixed it to DPL2 in Cool Edit using the matrix:
Lt = L + 0.7071 C + 0.866 Ls + 0.5 Rs
Rt = R + 0.7071 C - 0.5 Ls - 0.866 Rs
I then upmixed it back to 6 channels using WinDVD DPL2 decoder in Movie mode.
Results:
- The square waves were intact in the output indicating that no phase shift occurred.
- There was a 5msec delay of the entire output, and an additional 15msec delay of Ls' & Rs'.
- After accounting for this delay, the phases of all output channels were identical to the inputs, and there was no inverting either.
- The amplitude of the square waves fluctuated a bit at the channel switchover points (3, 6, 9, 12, 15 secs). This was probably an artifact of the steering logic.
Discussion:
This matrix produced output that was closest to the original input. The phases were the same and the relative channel volumes were the same with an overall reduction of -3dB. However this is NOT a DPL2 spec. compliant downmix. A discussion on 90° phase shifts will follow...
Well, whatever is the purpose of this test, there always are another eyes trying to see what it means actually..
I try to figure out conditions not described clearly :
. purpose : to prove if the DPL II decoders work correctly without the 90 deg shift process with non-pre-shifted clips
. conditions : very similar to tebasuna's powerdvd test before where he recommended the matrix3 to be used
- Windvd should represent the standard of the DPL II decoders
- the input streams should represent the normal combinations
The conclusion implied in the test result : although nothing specified explicitly
. windvd or powerdvd are de facto standard : whether complient to the DPL II spec or simply Dolby..
. the 90 phase shift is not necessary, matrix3 is the rule
. not regaining the -3dB on rears is correct
The conclusion I could derive from the result : although the previous test with 90 deg shift looked more reliable
. the decoder performs invert on one channel, does not perform 'reverse phase shift' as expected
. not sure if it is complient or not with the two test results : does not meet the conditions to conclude
. the -3dB rear output attenuation says not correct volume regaining by either wrong input or logic
. for the invert-only versions(and these s/w dpl II decoders), the matrix3 is the way to go!
/believing is personal freedom
tebasuna51
30th June 2006, 12:27
I made a 6-channel test sample in Cool Edit that contained the following:
0-3 sec, silence
3-6 sec, L 300Hz square wave
6-9 sec, C 400Hz square wave
9-12 sec, R 500Hz square wave
12-15 sec, Ls 600Hz square wave
15-18 sec, Rs 700Hz square wave
I then downmixed it to DPL2 in Cool Edit using the matrix:
Lt = L + 0.7071 C + 0.866 Ls + 0.5 Rs
Rt = R + 0.7071 C - 0.5 Ls - 0.866 Rs
I then upmixed it back to 6 channels using WinDVD DPL2 decoder in Movie mode.
Same test using PowerDVD6 DPL2 decoder in Movie mode. Exact results, only:
- There was a 5msec delay of the entire output, and an additional 15msec delay of Ls' & Rs'.
I get a 15msec total delay of Ls' & Rs' (5 global + 10 additional)
Rockaria
30th June 2006, 15:10
@raqueue, there are three possible cases that can happen by mixing the two identical signals before & after the DAC process.
Suppose you are playing the identical stereo streams in stereo mode(DPL II will aggressively decode them in surrounds) :
combinations : (0,0),(0, -180),(0, 90), (0,-90), (90,90), (90,-90), (-90,-90)....
a. mixing the streams in one channel before the DAC :
- the stream will completely cancel each other : (90,-90), (0, -180)...
- increased volumes : (0,0), (90,90), (-90,-90) > (0,+-90)...
b. phantom center speaker effect : (0,0), (90, 90), (-90,-90) > (0,+-90)..
c. difussion effect : (0, -180), (90,-90) > (0,+-90)...
@ bleo, at least we seem to agree the final DPL II stream must contain +-90 cross stacked mixes for the better seperation quality.
Some problems I am noticing is that the conditions to be proved are confused :
a. 90 shifts are nessary for normal 5.1 streams for better seperations : clearly explained in the Dolby's doc, doesn't need to prove anything else.
b. 90 shifts are not necessary at all except for some experiments : need to prove lots of things before claiming
- the DPL II encoding is mostly used only for DD5.1 movie clips
- all the DD5.1 movie clips rears are already 90 deg phase shifted
- there are actually some certified equipments for normal users which can do the invert-only downmix
- the 90 deg shifts has no negative effects at all for movie clips in the correlations between the channels maybe
- the MOVIE clips are always different from the MUSIC in the correlations between the channels maybe
- all the non-pre-90 shifted clips are just for hobby
Not an easy job .. impossible. Why not just let the methods be chosen by those who know what they are doing?
3dsnar
30th June 2006, 15:49
sorry,i don't agree and it's nothing have to do with DD5.1,follow me:
http://img425.imageshack.us/img425/9361/controlroom2bf.png
think that you listen the sound and not decoding
q1
what happen if the surround left and right have the same phase as front left and right?
a1- cancellation!
q2
what if the sL and sR have +90 degrees?
a2- cancellation!
q3
what if the sL and sR have -90 degrees?
a3- cancelation!
q4
what happen if sL have +90 and sR -90 degrees?
a4-
the sL is 90 degrees from the L channel.
the sR is 90 degrees from the R channel.
the sL and sR are now 180 degrees
no more cancellations and:
as show the picture,the surrounds are placed round 90 degrees from the frontal channels this is why they are turned 90 degrees and one of them with +90 and other with -90 to don't have cancellation because one speaker in front other sounding with the same phase will give.... cancellation.
this is what i understood(with experience) why they used 180 degrees between surrounds and at the same time 90 degrees from the fronts!
Ofcourse the speakers position are of importance. But this is a completely different issue. Here we are talking about the problems related to phase properties for downmixing and decoding (before the decoded sound gets outside the decoder).
So it is not that I disagree / agree with you :)
I just wanted to clear things out.
raquete
30th June 2006, 17:22
@ 3dsnar
So it is not that I disagree / agree with you
of course,i (we all) know that. ;)
I just wanted to clear things out. no need because you are always very clear in your posts.
remember,you are one "creator",i'm one single "listener".
@ Rockaria
c. difussion effect : (0, -180), (90,-90) > (0,+-90)...
difussion effect takes to listening effects when the audio is up or downmixed.(this is why i posted about speakers positions and respectives audio phases too)
all i can say about your (complete) last post is:
perfect (clever explanations,word by word) :cool: :cool:
edit: typos
Rockaria
1st July 2006, 19:08
The DIFFUSION effect looks like another expression(transformed energy) of the cancellation.
Like throwing another stone in the pond, the area between them will have some counter-wave effect making the power weaker.
If you change the phase from 0 - +-90 - +-180 between the two speakers, you will FEEL the center sound becoming weaker(scattered).
Not the entire cancellation(like digitally when mixing) but some type of transformed(to different type of energy : speaker) interference..
So I agree by including the digital phase altering mixing process(& delay,,,), the speaker diffusion/phantom center effects would make a full consideration of this environment.:)
raquete
1st July 2006, 19:42
ok.:goodpost: (one more)
one single test for "cancellations" :
1-create(generate tones) 100Hz -30 seconds stereo wave-same phase in L/R.
listen in one stereo amplifier(or in pc with good sound)where the speakers are in front of you(of course,left and right speakers)
you will listen one very clean sound.
2- now place one speaker in front the other.(face to face)
L --> <-- R
see(or better,listen) how much cancellation happen.
3- most important: i will that you understand what i mean :p
Rockaria
1st July 2006, 20:24
I see the sound becoming almost muted with extremely close faced speakers & identical stereo playback with a ch invert phased.;)
The speaker's transformed energy consists of several factors :
- direction : it has 360 degree but more directed to the front of the speaker
- carrier : air(thiner density of the matriels than water or circuit in digital/analog formats), it will be delayed, weakened by the distance.
- set up : 5.1(3 fronts and 2 rears facing each other but with some enough distance), angle & distance(&delay),,, fators are crucial
- room(ambience) : interferences, reflections(matriels), decay, echos...
Thanks for providing the excellent extreme example.:goodpost:
[edit]
The same DIFFUSION or cancellation effect will also happen when changing the one speaker polarity(swapping the red & black wires).
Rockaria
20th July 2006, 20:36
. for the invert-only versions(and these s/w dpl II decoders), the matrix3 is the way to go!
http://forum.doom9.org/showpost.php?p=844727&postcount=134
Based on the Dolby's and surecode's docs, the invert only versions should use the tebasuna's matrix3(in avs or similar scripts) :
- the DPL II is expected to have the out of phase signal in Rs : the matrix1 will have (-Ls, +Rs) encoding time
- the decoder movie mode inverts the Rs(as explained in the surecode doc) : the matrix1 will have (-Ls, -Rs) both inverted
- so the matrix3 will have (+Ls, -Rs) encoding time and (+Ls, +Rs) decoding time.
- but these models do not fully seperate the channels(unbalanced), although will never be agreed or admitted..
Because most 90 phase deg shift routines perform -90 deg shifts effectively, invert mixing the Ls(coeffs) will have (+Ls, -Rs) encoding time and (+Ls, +Rs) decoding time.
I have tried to find a plugin or ways to implement/utilize the 90 deg phase shift in avisynth with no success yet. Developing the plugin also does not have the proper environment(esp. time) for me. But as I found the softEncode surround 90 deg phase shift option performs correctly for the DD downmix encoding/play, I conclude it's the ac3 encoder's role to do the optional 90deg shift + downmix + (DD2.0 / WAV encode/STDOUT).
http://forum.doom9.org/showpost.php?p=850563&postcount=15
http://forum.doom9.org/showpost.php?p=853856&postcount=50
And the complete background of the DPL II downmixing process in DD encoding(rear 90deg phase shifts option) and decoding(auto downmix mode selection based on annex->dmixmod meta info) : http://forum.doom9.org/showpost.php?p=856250&postcount=98
I think the softEncode's batch mode will still be useful for the DPL II encoding automation if we don't care about the temporary WAV AC3 files.
As for the Dynamic DPL II encoding for the ds filter use(FFDShow...), there are some conditions I can think of for now :
. the freq-delay method must be applied to the all-pass(granular frequency bands) and aligned
. the buffer must be large enough to sync(align) the delays between the freq-bands as well as other channels.
. the accuracy of the phase shifts degree does not affect the quality that much
...
This technology is already fading even without being utilized of the full potentials and advantages.. that's too bad..
/.
[edit] added a reference in the same thread
vBulletin® v3.8.11, Copyright ©2000-2026, vBulletin Solutions Inc.