View Full Version : New B-VOP code
Tommy Carrot
10th July 2004, 12:39
As i checked the CVS, i noticed there are massive changes in the B-VOP motion estimation. Syskin, could you explain what should we expect from it? Faster speed, better mode-decisions, RDO?
SeeMoreDigital
10th July 2004, 12:46
Hi Tommy,
Is this with XviD v1.0.1 ?
If so, this might explain why XviD's 1B-VOP works differently in hardware to DivX and Nero's 1B-VOP.
Cheers
Tommy Carrot
10th July 2004, 12:50
It's in the HEAD branch (i guess it's the same as 1.01). The changes are 2 days old.
IMO the root of the hardware problems is that xvid places the b-frames adaptively, while the others place them in fixed pattern.
SeeMoreDigital
10th July 2004, 13:09
Originally posted by Tommy Carrot
IMO the root of the hardware problems is that xvid places the b-frames adaptively, while the others place them in fixed pattern. Well even though I don't confess to understand computer code this explanation makes sense to me!
I wonder why the guys at XviD have not provided this explanation in the past (or have I missed it).
Is there a work around within any of XviD's settings to simulate 'fixed pattern' B-VOP? Or could a third party tool be developed to do it. Such as Moitah's excellent MPEG4 Modifier?
Cheers
Manao
10th July 2004, 13:18
AFAIK, fixed frame-type pattern has nothing to do with the MPEG-4 norm, so I don't think XviD's b-frames issues with hardware are related to that issue. Moreover, a pattern is fixed up to a certain limit : I-frames at scene changes must be inserted, and that certainly disrupt the pattern.
And non fixed pattern can't be made fixed after compression : it needs at complete recompression, by the codec itself, because doing so, you'll change reference frames.
sysKin
10th July 2004, 14:39
Originally posted by Tommy Carrot
As i checked the CVS, i noticed there are massive changes in the B-VOP motion estimation. Syskin, could you explain what should we expect from it? Faster speed, better mode-decisions, RDO?
As for now, you should expect more speed (especially with qpel) and more quality (but not much). And new bugs of course, one of them was fixed an hour ago.
In the future you should expect VHQ - this code was made ready for it.
Oh, and CVS HEAD is not 1.0, it's 1.1.
Radek
Tommy Carrot
10th July 2004, 15:20
Thanks, good to hear it.
virus
10th July 2004, 17:42
Originally posted by sysKin
In the future you should expect VHQ - this code was made ready for it.hi Radek,
just a question: will it be turned on by default at the same level than for the P-frames? Or maybe there will be two different GUI controls for it? (allowing, for example: VHQ 4 for P-frames and VHQ 1 for B-frames)
cheers :)
virus
lordadmira
10th July 2004, 20:29
How exactly will the VHQ work with B frames Syskin? And also could u contribute to my QP/BVOP thread. ;)
sysKin
11th July 2004, 04:13
Originally posted by virus
just a question: will it be turned on by default at the same level than for the P-frames? Or maybe there will be two different GUI controls for it? (allowing, for example: VHQ 4 for P-frames and VHQ 1 for B-frames)Most of that depends on possible PSNR gain. I'm thinking that vhq for b-frames should be VHQ5, but if psnr gain is big, it will be active before that.
Anyway it's just a GUI choice.
There might also exist VHQ6 (maybe later), which makes an extra RD-based motion search in b-frames. Again, that depends if there is any psnr to win from it or not.
(Oh and don't worry about the fact I'm using psnr - R-D search is, kinda by definition, a psnr-improving technique).
@lordadmira: I have nothing to say in that thread, your theories there are a complete failure you know ;)
I don't like to speculate anyways.
VHQ in b-frames will work in exactly the same way as they do in p-frmaes.
virus
11th July 2004, 04:36
Originally posted by sysKin
Most of that depends on possible PSNR gain. I'm thinking that vhq for b-frames should be VHQ5, but if psnr gain is big, it will be active before that.
Anyway it's just a GUI choice.Thx for the explanation :)
I have the impression that choosing the "right" level of VHQ could be pretty difficult. The max number of B-frames also plays an important role, no? If you aim at high quality in a reasonable time and have 60-65% of B-frames (typical for 2/1.50/1.00), you may want to spend more time on them - example: VHQ4 for B-vops but just VHQ1/2 for P-vops.
But with settings like 1/1.50/1.00 you may want to spend the same time on both B- and P-vops... example choose VHQ3 for both. A lot of possibilities are open if you have two dropdown lists to play with ;)
Also, can you estimate how much slower the VHQ5 mode is/will be compared to the actual VHQ4? (this is my last question, promised! :D)
[EDIT] mmmhhh... re-reading the posts, I think we're actually speaking about 2 different things... you're talking about a VHQ5 which is "VHQ4 + SAD-based mode decision for B-frames" and VHQ6 as "VHQ4 + R/D-based mode decision for B-frames", right? I was talking about two independent settings, let's call them P-VHQ and B-VHQ, meant to choose the amount of additional work for each frametype. Hope I understood well... :D
cheers
virus
lordadmira
13th July 2004, 01:38
Originally posted by sysKin
@lordadmira: I have nothing to say in that thread, your theories there are a complete failure you know ;) :o What theories? I'm not pushing any theories there. One person has warped my post into something like that but I'm not. I'm trying to figure out the worth of QP in relation to BVOPs. It's a question. Crusty's FAQ implies that QP is no good without BVOPs.
LA
Manao
13th July 2004, 05:48
Sorry, but : Crusty's FAQ implies that QP is no good without BVOPs.The FAQ doesn't imply that, you're the only one to make that deduction from Crusty's explanations. Crusty doesn't even speak of b-frames when speaking of Qpel.
Teegedeck
13th July 2004, 06:51
Exactly. Please; no spreading of false information here!
Splashdriver
13th July 2004, 08:17
Originally posted by Manao
Sorry, but : The FAQ doesn't imply that, you're the only one to make that deduction from Crusty's explanations. Crusty doesn't even speak of b-frames when speaking of Qpel.
True. I've just reread Crusty's FAQ and really, Crusty's FAQ doesn't speak of enabling Qpel when B-frames are activated in order to increase compressibility (where previous FAQ's DID).
Greetings,
Splashdriver
Koepi
13th July 2004, 09:13
what is going on here?
bframes and qpel _are_ unrelated.
qpel _can_ increase compressability further together with bframes.
(there's nothing to discuss about. that's fact.)
SeeMoreDigital
13th July 2004, 09:44
Koepi
Sorry to keep pushing you XviD guys on this. But will one of you ever be able to explain why XviD's 1B-VOP works differently in hardware to DivX's and Nero's 1B-VOP.
If you don't want to block up a thread talking about it, please PM me?
Many thanks
Sagittaire
13th July 2004, 10:21
Good news ... very good news ...
actual bframe quant decision mode isn't really good: the local quality differences are really too important sometimes (with defaut setting 2/1.50/1.00).
lordadmira
13th July 2004, 20:08
Originally posted by Koepi
what is going on here?
bframes and qpel _are_ unrelated.
qpel _can_ increase compressability further together with bframes.
(there's nothing to discuss about. that's fact.) I understand the scenario where QP helps together with B frames. However I don't see how it can help without B frames. And no one has been able to give an explanation one way or the other.
Crusty's FAQ does imply what I stated. He uses examples where motion is caught over multiple frames. This example requires B frames. Without B frames motions are only caught over 1 frame, from one P frame to the next. I don't see any value to catching motion at subpixel resolution between adjacent frames. This is because it's impossible for a pixel/macroblock to move less than 1 pixel from one frame to the next, due to the digital nature of the image. If anybody can truly explain why this is wrong, please do so.
LA
stephanV
13th July 2004, 20:31
I'm no expert but i don't think it works exactly like this... how motion is "caught" using QPEL (or HPEL(?) for that matter) has got nothing to do with the frame type. Although P-frames only use information from the previous I/P-frame, this doesnt mean a block can't be predicted for several frames. And i think it is in these cases (which are quite common) QPEL could be useful. QPEL is looking forward 4 frames to do motion estimation and thus results in more accurate motion vectors for each block.
Crusty's FAQ (chapter C3e) (http://www.vslcatena.nl/~ronald/docs/xvidfaq.html#C3e) even mentions an example with P-frames only:
If we assume that there is one vector for every macroblock (there might be 4 or 0), at the resolution of 640x272 and 24fps and P-frames only, two bits for every macroblock take 40 x 17 x 2 x 24 = 32640 bits or 32.5 kbps.
Also note, that the word B-frame is not mentioned at all in the explanation of QPEL.
Manao
13th July 2004, 20:36
This is because it's impossible for a pixel/macroblock to move less than 1 pixel from one frame to the next, due to the digital nature of the imageAlright. Imagine a black picture ( pixels at 0 ), with a white square in the middle ( pixels at 255 ). We read the value of the pixels on an horizontal line and we find :
0 0 255 255 255 0 0
Now, the same picture, but this time with the square moved by half a pixel on the right. If we read the value of the pixels on the same line, we'll read :
0 0 128 255 255 128 0
Now, if we try to find the motion vector, using a block based approach, with SAD distorsion, with a pel precision, we will find that either 0 or 1 are good candidates.
But, we could also try to find a motion vector with a half pel precision. To do so, we have to make an interpolation of the first picture. Let's say we made a bilinear interpolation, the picture becomes :
0 0 0 128 255 255 255 255 255 128 0 0 0
And now, for searching the motion vector, we have to set of values ( odd and even ones ), which correspond to pel vectors and half pel one :
0 0 255 255 255 0 0
Which is the original one, and which correspond to pel vectors, and this one :
0 128 255 255 128 0
Which is for the half pel vectors. We can see that a vector 1/2 gives an exact match, distorsion-wise. Hence, half pel motion is possible between two consecutive frames.
Now, since you seem to think quartel pel is kind of different, let remake the experiment, with moving the block by a quarter of a pixel on the right. The picture becomes :
0 0 192 255 255 64 0
Now, we can make a bilinear interpolation to enlarge the first picture by 4 :
0 0 0 0 0 64 128 192 255 255 255 255 255 255 255 255 255 192 128 64 0 0 0 0 0
We can build four set of values, which are :
0 0 255 255 255 0 0
0 64 255 255 192 0
0 128 255 255 128 0
0 192 255 255 64 0
And we find a perfect match ( once again ), this time with a 1/4 vector.
You didn't understand my previous explanations on the other thread, I hope this one will show you the light.
Now, with a more logic reasoning, if quarter pixel movement between two adjacent frames was impossible, as you're claiming, how could it be possible between two frames not adjacent ? Those two frames are linked by adjacent frames, and if only pel and helf pel movement were possible between them, then the total movement between the two frames would be the sum of the motion vectors of the adjacent frames, and so would still be half pel at most.
lordadmira
13th July 2004, 20:55
Ah, what u have done is included the idea of texture. If ur example is indeed correct then Crusty needs a new Qpel section.
Keopi? Is this right? What can u add?
LA
PS: I wasn't "claiming" anything. U like to rewrite my words...
virus
13th July 2004, 20:57
to put it simply: a digital image results from the sampling of a continuous scene (example: for a DVD movie, that's the original filmed scene with the actors, the objects, the background and so on). And in a continuous scene, movements can be of any length, which correspond to an arbitrary distance (in pixels) in the sampled image, even stuff like 0.31624, or 1.018, or 2.21 pixels. Simple as that. Subpixel motion compensation allows to reduce the impact of such effects on the compression.
virus
Manao
13th July 2004, 21:09
Ah, what u have done is included the idea of textureUh ? No, I only included the idea of image, that's all.Crusty needs a new Qpel section.Not at all, you're the only one to see in it something that isn't. Reread it, you'll see that Crusty is _perfectly_ right. And I don't understand your fixation on Qpel : half pel works the same, and is used whatever the settings are, and that doesn't bother you ( which may be by the way the reason you don't understand Crusty explanations )
lordadmira
13th July 2004, 21:29
U need the idea of texture in order to say that the 0 0 192 255 255 64 0 block is the decendant of the 0 0 255 255 255 0 0 block. Without texture the 255 pixel has to be matched against another 255 pixel somewhere in the future. I wouldn't use the word perfect on Crusty's QP section. Even he himself calls it a crude example. The same issue I was raising with QP also applies to HP. For the sake of brevity I didn't say anything about it. That would really have thrown u into a conniption. :D There isn't a glossary around here so if my use of the word texture is less than cannonical, just gonna have to deal.
LA
RadicalEd
13th July 2004, 22:25
Originally posted by lordadmira
U need the idea of texture in order to say that the 0 0 192 255 255 64 0 block is the decendant of the 0 0 255 255 255 0 0 block. Without texture the 255 pixel has to be matched against another 255 pixel somewhere in the future.
That's the whole point of h/qpel. The encoder can interpolate the image data and estimate that so that the new block can be directly compensated from the original block without storing any texture data.
Originally posted by lordadmira
And no one has been able to give an explanation one way or the other.
I thought my post in the original thread (http://forum.doom9.org/showthread.php?s=&threadid=79377) was pretty proficient...
Originally posted by RadicalEd
As far as I understand it, the codec will literally interpolate the macroblock up to subpel depth in an attempt to get a better match. The plane of pixels in an image isn't continuous, but the actual objects in the image are/were before being digitized. So it's possible for something to shift a half pixel to the right, and in doing so merely cause the whole pixel to change slightly. Unless I'm totally off base, this is what half/quarter pel is designed to account for.
lordadmira
13th July 2004, 23:03
Originally posted by RadicalEd
That's the whole point of h/qpel. The encoder can interpolate the image data and estimate that so that the new block can be directly compensated from the original block without storing any texture data. I guess one problem is that we all have a different understanding of the word "interpolate". Reading ur old explanation now, I can see how it makes some sense. The way u used interpolate didn't make any sense to me. Interpolation means filling in interveneing values from two distant points (inter - between). A better word to use would be upsample. However I probably still wouldn't have known what u were talking about. The texture idea (non-canonical usage) makes much more sense to me.
LA
sysKin
14th July 2004, 06:20
Originally posted by lordadmira
I guess one problem is that we all have a different understanding of the word "interpolate". Reading ur old explanation now, I can see how it makes some sense. The way u used interpolate didn't make any sense to me. lordadmira...
After two long threads of hpel/qpel discussion, it turns out that you don't know what sub-pixel interpolation is?? Please, I'm not telling you you should understand anything, but you seem to be very knowledgable about details of mpeg-4 without having any knowladge about how mpeg compression works.
Everything Manao, virus, RadicalEd, Crusty and others said is exactly and accurately true. Because of that, I suggest you understand that, instead of (/before) trying to push your theories.
You might also want to google for mpeg-1 compression explaination and familiarize yourself with the terminalogy used. Yes, interpolation means "something in the middle" - and in mpeg, you've got spatial interpolation (in sub-pixel filters) and temporal interpolation (in b-frames), and you should simply know which of the two is meant each time, instead of being surprised about this term.
Your use of "texture" is also wrong - just find out what it means in terms of mpeg before using it.
Regards,
Radek
PS. You're free to ask me about any details, I don't mind helping you with that.
lordadmira
14th July 2004, 07:15
Originally posted by sysKin
PS. You're free to ask me about any details, I don't mind helping you with that. It's interesting how hard it is to wring information out of people..... I don't equate that with pushing a theory. When things aren't documented, that's the only avenue left. And after all, that's all we've got to go on: bits and pieces written down here and there. So when somebody writes something down trying to convey x but people come away with y, there's a problem. And that's something that people already in the know can't judge. Because they can't evaluate a text in the same way that a first timer does. Ur own knowledge is applied in the process. It's the same phenomenon of why u miss ur own spelling mistakes, it's already correct in ur mind. U need an outsider to proof read what u write. So then someone like me comes along and makes statements based off what has been written. Then there's all kinds of commotion because of the difference in perspective. As if said person is just pulling things out of their arse. So if there's any problem here it lies with the documentation that's out there. Hell that's why I came here in the first place! To try and find something out! I don't just accept what's put in front of me like a good litte comrade. Denying that their could be anything wrong or ambiguous about one's own works is rather arrogant. If I'm explaining something to somebody and they don't understand or get the wrong idea, that's my fault, not their fault. I understand what ur saying about jargon but common sense has to apply also. Perl misuses the word interpolate too, doesn't make it right. Where jargon deviates from normal meanings there needs to be asterisks and glossaries. I don't see any glossary associated with anything Xvid. Maybe I missed it. And I've only seen thin incomplete Mpeg-4 glossaries. I only used texture as a stopgap word, I never suggested it was "the" meaning and indicated such. Atleast Crusty made the effort to write a FAQ.
This is not a flame. This is flame retardant and pre-emptive retardant for future reference. So with that, let this thread be locked.
LA
malkion
14th July 2004, 07:28
Ah, the tautology, lol.
@sysKin, you've mentioned the quality gain from using vhq-5 may be only a little. As I currently use vhq-4 with all encodes, do you think vhq-5 will become the new standard, per say?
sysKin
14th July 2004, 08:17
Originally posted by malkion
@sysKin, you've mentioned the quality gain from using vhq-5 may be only a little. As I currently use vhq-4 with all encodes, do you think vhq-5 will become the new standard, per say? I hate speculations - so the answer is: no idea, and I can't tell until it's ready.
virus
14th July 2004, 11:40
Originally posted by sysKin
I hate speculations - so the answer is: no idea, and I can't tell until it's ready. hi Radek,
please make sure to yell when the code is ready so we can help you with the tests. No matter what you need: R/D curves, fixed quantizer w/ all possible settings, visual impressions (look, we even do blind tests here :D), whatever. Just say what you need and we will help, no problem!
In this way at least we'll have a chance to give something back and let you concentrate on the actual coding. The more XviD gets refined, the more difficult the assessment about the improvements will be. So, don't waste your time doing everything by yourself... just ask ;)
cheers :)
virus
virus
14th July 2004, 11:58
Originally posted by lordadmira
It's interesting how hard it is to wring information out of people.....
I won't say so. Many people tried to correct your wrong statements but you always answered "uh? what ur saying doesn't make any sense" or "u don't know what ur talking about" or simply blaming someone other (like Crusty) of your own misconceptions.
And when your false claims are finally torn in pieces by 5-6 different people, you resort to standard trolling techniques like saying "it was not a claim, just a question" or "oh well, but we're off-topic" or "don't explain this to me, but to xyz which don't understand this stuff" or "nobody gave an answer" when, as a matter of fact, 5 people gave the same answer and you've simply ignored them.
So as you can see, it is not a matter of pure knowledge, but a matter of respect for the people which waste their time to clear up all the FUD you've spread on this board. If you wanna an explanation, we're here. But then listen to the explanations and trust people who definitely know what they're talking about. I won't try to teach you japanese, because I assume your knowledge of it is vastly superior to mine. But then, please do not try to teach sysKin about the "true power of the B-frame" because in this field, sysKin knows everything and you should simply shut up and thank him for the explanation. Well, sorry to seem harsh, but I think you get the idea ;)
virus
lordadmira
14th July 2004, 21:27
So Xvid's B frames didn't work the way I thought they did. Big deal. My own tests showed that and I changed my opinion (after consulting with Teeg I might add). The only claim I made was that about B frames. Everything else was a question whether u accept that or not. How can u fault me for not understanding a particular person's particular explanation of something? Saying something over and over again doesn't help. That's childish. Believe it or not I'm hanging around here to both learn and contribute.
LA
virus
14th July 2004, 21:59
Originally posted by lordadmira
How can u fault me for not understanding a particular person's particular explanation of something?
No, unfortunately you didn't get the idea.
Never talked about your understanding capabilities. Please re-read my post. I talked about how you usually react to the explanations. You're missing the point of my words completely. Let me repeat it again: it's not a matter of understanding things, but a matter of correct behaviour with other people. Please re-read what I've written. There's nothing wrong if someone doesn't know everything about a specific topic. (more disturbing would be to *pretend to know* what you don't know, anyway)
virus
lordadmira
14th July 2004, 22:14
I think I have behaved properly. I'm always conscious to restrict my comments and give people the benefit of the doubt. But it seems like people aren't so apt to give that to me. Probably because I'm a boat rocker. :D But still that's no excuse.
Teegedeck
14th July 2004, 23:06
To summ it up; we all had to type a lot. Nothing's broken and we can get on with business as usual. My thanks go out to those who have submitted such splendid explanations of the inner workings of MPEG-4! This forum is a treasure trove (I've always wanted to use that expression ;)) of knowledge, thanks to you!
vBulletin® v3.8.11, Copyright ©2000-2026, vBulletin Solutions Inc.