View Full Version : Any news about a potential successor to VVC / H.266?
kurkosdr
19th August 2022, 18:15
I know it's a bit early to ask, but with Leonardo Chiariglione having announced the "death of MPEG", I thought there is no harm in asking:
Has anyone heard any news about a potential successor to VVC / H.266?
Now, don't get me wrong, VVC is pretty advanced and all, but it still can't compress a UHD stream to the bitrate of an FHD H.264 stream with the same quality, while at the same time terrestrial bands are shrinking worldwide. That's why I am wondering whether the 10-year cadence (a new format every 10 years or so) will be kept alive despite the recent restructuring of MPEG or VVC / H.266 will be the end of the line.
Also, I know AV1 exists and will likely keep evolving, but it's not an ISO or ETSI standard, and free-to-air and free-sat broadcasters can only use ISO and ETSI standards (and ATSC standards for the US). In other words, ISO standards are still important if free-to-air and free-sat are to keep up with internet distribution in terms of quality.
nhw_pulsar
19th August 2022, 19:44
There is currently MPEG VVC ECM: Enhanced Compression Model (beyond VVC) which presents some impressive improvement over VVC, something like 20-30% if I remember correctly.MPEG is also studying machine learning next-generation video coding tools, notably super-resolution and neural-network-based loop filters. For more info, details are available on their website at standard explorations.
A drawback is that it is really slower notably to decode...
Also the ultra-impressive ECM is however (as always) PSNR- and SSIM- driven, which don't correlate well with human visual system.For example, on the clic learned image compression challenge website, I have seen an image at high compression with BPG and ECM, and I visually preferred the BPG image... As also very quickly with VVC, I don't know if my VTM binaries are right (I also use the default command line and slowest preset), but I find that VVC lacks of neatness, with a comparison with my codec, I find that NHW has more neatness (but less precision), but visually I find that neatness is more pleasant than precision... Don't know if ECM will correct neatness? But of course this is my very personal opinion, for example the industry absolutely doesn't share it and completely sticks with PSNR, SSIM and precision...
Cheers,
Raphael
Jamaika
20th August 2022, 09:30
What are BPG photos and who uses them? After all we have the HEIF/AVIF standard.
The purchase of equipment is also very interesting nowadays. After the promotion, Orange offers smartphones for X EURO, but it can buy equipment for 75% cheaper on amazon through an intermediary who keeps 15% commission.
As for the vision of 8K LTM, LCEVC, EVC, AV2, JXS. They are paid and only on the SAT.
The Webb telescope has problems, and it is not known what the photos of the 16K universe will be in. A good topic for a clairvoyant.
https://www.v-nova.com/try-v-nova-video-compression-technologies/
https://cloud.qencode.com/lcevc-video-codec
https://www.lcevc.org/how-lcevc-works/
https://www.mpeg.org/
Blue_MiSfit
21st August 2022, 03:56
LC-EVC is a big step forward for a lot of use cases, especially when paired with next gen formats like VVC and AV1.
It's not a slam dunk for every use case though :)
nhw_pulsar
21st August 2022, 19:25
Hello,
Just a quick correction, sorry for my misleading information, BPG is not better than VVC intra.On my image test sets, VVC VTM intra is a little visually better than BPG for me, but it seems rather slight (BPG -m 1 also takes 78ms to encode, whereas VTM intra takes 75seconds to encode...), I also don't test at extreme compression.But the difference really appears for video coding I think, because I don't test video sequences, but I have seen few comparisons between HEVC and VVC at the same bitrate, and VVC was then really visually better in the video case.
Cheers,
Raphael
rwill
22nd August 2022, 06:05
MOM !
People are comparing shitty (reference) encoder implementations against each other again !!!
ksec
22nd August 2022, 14:32
It is progressing at a surprisingly fast rate. ECM 4 + EE1.2 managed ~30% BD-Rate with 4K compared to VTM 11. ( VVC Reference Encoder, although VTM 16 do outperform VTM 11 by about ~10% so the actual difference isn't as big )
We should have some news about ECM 6 soon. We are not far off from the 40-50% BD Rate compared to VVC.
SeeMoreDigital
22nd August 2022, 15:06
I know it's a bit early to ask, but with Leonardo Chiariglione having announced the "death of MPEG", I thought there is no harm in asking...
Hmmm...
Leonardo Chiariglione's announcement was made over two years ago (Sat 06 Jun 2020) over on the MPEG Home Page (https://www.mpeg.org/) when his chairmanship of the group came to an end after 32 years!
I suspect he was more than a little pissed off... Was he pushed out? Probably ;)
benwaggoner
24th August 2022, 18:45
AV2 is also in development, and certainly will be available before ECM. It's too early to say if it'll be as good as VVC for real-world content, let alone better.
ksec
25th August 2022, 14:42
I would go as far as to say ECM is actually ahead of AV2 in terms of development. At this rate we could have next gen ( H.267? ) spec ready by 2025.
It will be interesting because ECM is looking like the first Video Codec which would requires dedicated Hardware Decoder to work. We are looking at ~5x the Decoding Complexity compared to VVC.
FranceBB
29th August 2022, 14:26
But most importantly... where the heck is x266?! :(
Honestly, is there any plan from Multicoreware to release what they have done and make it open source once and for all? 'cause honestly, I've been doing all my tests with Fraunhofer's VVEnc which is somewhat more usable than the VTM reference encoder, but still, I'd like to have x266 to play with and start the integration with our systems... :(
P.s looks like they're gonna be at this year's IBC, so maybe I can ask them in person https://multicorewareinc.com/ibc2022/
What do you say, Ben? Shall we get there together on behalf of Doom9 xD?
EDIT: Ok, I've actually booked with them on Monday, September 12. I'll ask about x266 and I'll come back here with some more info. Hopefully it's gonna be good news. Stay tuned.
terrestrial bands are shrinking worldwide. That's why I am wondering if the 10-year cadence (a new format every 10 years or so) will be kept alive despite the recent restructuring of MPEG or if VVC / H.266 will be the end of the line.
Don't worry about terrestrial, I'm sure that in 10 years time there are still gonna be channels in MPEG-2 720x576 yv12 25i TFF BT601 alive xD
Blue_MiSfit
29th August 2022, 20:02
Let us know what MCW says - I haven't spoken to them in quite a long time (since before Pradeep left).
benwaggoner
30th August 2022, 18:31
What do you say, Ben? Shall we get there together on behalf of Doom9 xD?
Yeah. We should have a doom9 meetup even if there's a critical mass going.
kurkosdr
2nd September 2022, 21:29
But most importantly... where the heck is x266?! :(
Honestly, is there any plan from Multicoreware to release what they have done and make it open source once and for all? 'cause honestly, I've been doing all my tests with Fraunhofer's VVEnc which is somewhat more usable than the VTM reference encoder, but still, I'd like to have x266 to play with and start the integration with our systems... :(
P.s looks like they're gonna be at this year's IBC, so maybe I can ask them in person https://multicorewareinc.com/ibc2022/
What do you say, Ben? Shall we get there together on behalf of Doom9 xD?
EDIT: Ok, I've actually booked with them on Monday, September 12. I'll ask about x266 and I'll come back here with some more info. Hopefully it's gonna be good news. Stay tuned.
To be honest, with the development of open-source encoders for ISO standards having slowed down to a standstill (for example: no high-quality AAC-LC encoder, or HE-AAC encoder, or MPEG-H 3D Audio encoder), I am glad Multicoreware has stepped up to create an open-source VVC encoder, so at least that thing is covered. It's the same deal with EVC, we had to wait for Samsung to donate an open-source encoder (xeve) to have one. So, let's give MultiCoreware some time, they have delivered well enough with x265 to deserve our trust IMO.
FranceBB
2nd September 2022, 22:34
Absolutely, I totally trust them, but there are two things that worry me now and didn't worry me in 2013 in the x265 days:
- They don't have an account where they post regularly on Doom9 anymore (and they deleted the x265_Project account from which the devs used to post)
- Pradeep Ramachandran has left Multicoreware (June 2015 - May 2021) and he was the manager of the video-related development team who was responsible for x265 and I have no idea who picked up the task after him
which is why I'm going to meet with them and ask them how things are going.
Speaking of which, I'm gonna open a new topic on monday as I'd like to collect all the questions that the community might wanna ask them and report those to them so it would be as if you all came to their boot at IBC with me. ;)
kurkosdr
3rd September 2022, 16:42
LC-EVC is a big step forward for a lot of use cases, especially when paired with next gen formats like VVC and AV1.
It's not a slam dunk for every use case though :)
Is it any good when layered over HEVC? HEVC has emerged as the defacto format for UHD 4K video in terrestrial and satellite, so anything that could be done to reduce the bitrate without breaking compatibility is always welcome. With this arrangement, the current HEVC decoders could potentially decode a lower-bitrate HEVC stream while any newer receivers could also decode the LC-EVC enhancement to get the equivalent quality of a higher bitrate HEVC-only stream. Provided LC-EVC over HEVC is worthwhile, of course.
To be honest, I am a bit disappointed by the current state of UHD in broadcasting. We are taking state-of-the-art DVB-T2 muxes being able to broadcast a grand total of 3 channels (https://www.digitalbitrate.com/dtv.php?mux=HEVC&liste=1&live=1&lang=en). With VVC it would increase to a grand total of 5 (meanwhile with H.264 FHD you can fit 6). That's why I think that UHD on broadcasting won't be much of a success unless some further improvement is made on compression.
hajj_3
3rd September 2022, 20:07
To be honest, I am a bit disappointed by the current state of UHD in broadcasting. We are taking state-of-the-art DVB-T2 muxes being able to broadcast a grand total of 3 channels (https://www.digitalbitrate.com/dtv.php?mux=HEVC&liste=1&live=1&lang=en). With VVC it would increase to a grant total of 5 (meanwhile with H.264 FHD you can fit 6). That's why I think that UHD on broadcasting won't be much of a success unless some further improvement is made on compression.
iptv is becoming more popular so dvb-t2 will become less relevant. Plenty of bandwidth available with DVB-S2 and DVB-C.
FranceBB
3rd September 2022, 22:23
To be honest, I am a bit disappointed by the current state of UHD in broadcasting. We are taking state-of-the-art DVB-T2 muxes being able to broadcast a grand total of 3 channels (https://www.digitalbitrate.com/dtv.php?mux=HEVC&liste=1&live=1&lang=en). With VVC it would increase to a grant total of 5 (meanwhile with H.264 FHD you can fit 6). That's why I think that UHD on broadcasting won't be much of a success unless some further improvement is made on compression.
That's true for DTT, but don't think Satellite is any better 'cause HotBird is overcrowded and bandwidth is expensive af. :(
iptv is becoming more popular so dvb-t2 will become less relevant. Plenty of bandwidth available with DVB-S2 and DVB-C.
Not really, it's becoming more popular in cities only.
If you think about rural communities, they can easily get a DTT feed or a Satellite feed, but they can't really watch the telly over the internet.
To make a quick example, during the 2nd heatwave in Italy I didn't wanna stay in Milan and sweat my balls off on a 104°F heatwave, so I went to a mountain village at 5000 feet which was far better with as little as 61°F. Thanks God they had a Satellite dish and I could watch the game quite happily in UHD at 25 Mbit/s H.265 HEVC 4:2:0 HLG HDR 10bit planar, 'cause otherwise I would have had to miss it (yes, miss it). Why? Well, 'cause even though the hotel had Wi-Fi, their Wi-Fi was hooked up to an ADSL modem 'cause nothing else was available or ever got to that village. Same goes for mobile connection which was limited to 3G only. Until internet is gonna be everywhere at decent speeds and with modern technologies, IPTV will never be a thing. If you use satellite, however, you have pretty much the same quality everywhere and it doesn't matter where people live, they almost certainly will be able to install a dish and watch their favorite programs at high quality. Internet is fine if you have things like OTT where you wanna watch a movie or a TV Series and it can load all the time you want as buffering doesn't matter, but if you have a live event like a Football game, then latency matters and in that case a satellite feed will always be the only option for lots of people, especially for those living in rural communities. Speaking of internet-only based providers, this is mainly the reason why companies like Netflix and Amazon who offer on demand movies and tv series are succeeding in Italy, but those who only offer Sports like DAZN are not (and never will without satellite).
p.s I also work for a company who airs via satellite, so I'm biased xD
ksec
4th September 2022, 06:42
ECM 4.0 Tagged, ECM 5.0 expected in October.
ksec
27th December 2022, 14:04
ECM 4.0 Tagged, ECM 5.0 expected in October.
Now they tested ECM 6.0 vs VTM 11
The results of the two tests reveal congruent results for the performance of the ECM and the VTM. Both, the laboratory test conducted with naïve viewers and the on-site test with experts demonstrate a clear visual benefit of the ECM when compared to the VTM for a significant number of cases. For the laboratory results, reported BD rate savings indicate a benefit of about 38% for the UHD test sequences and about 32-33% for the HD test sequences on the given test set.
The 38% average was with 3 test doing 45% reduction and two test with 30% only.
Pretty damn impressive if you ask me. This is excluding the EE2 work using Neural Network. And they still have plans for ECM 7.
hajj_3
5th March 2023, 14:59
ECM 8.0 is out: https://vcgit.hhi.fraunhofer.de/ecm/ECM/-/tree/ECM-8.0
ksec
5th March 2023, 18:32
ECM 8.0 is out: https://vcgit.hhi.fraunhofer.de/ecm/ECM/-/tree/ECM-8.0
Nice. We will have to wait for the next JVET meeting in April to know its results. It is progressing rather quickly.
DTL
21st March 2023, 13:18
What is really required to make some real progressive step over classic h.264 is significant or complete move to object-oriented encoding. For any type of scene (including any natural scene).
Currenly as supplement to 'old-block-oriented' MPEGs I making simple pre-processing engine (pmode=1 for MDegrainN at mvtools2 software pack for avisynth environment) and it shows good benefit already (about 60 -> 90% of skipped blocks in x264 at crf=18 encoding and about 40..50% less bitrate at mostly static scene and about 5% less bitrate at mixed-scenes movie). And it still not support 'forward' motion compensation for moving scenes so total mixed-scenes title encoding benefit will be higher after full idea implementation. Example of processing software at https://forum.doom9.org/showthread.php?p=1984527#post1984527 .
But it is still very poor amature-level software design and from still not very died current civilization industry of video codecs design expected more complete and professional solution:
Video codec must extract from incoming frame sequence for encode the scene content (textures, objects, lighting, motion data) and create standard compatible bitstream for enduser decoder to decode 2D frames sequence to feed to standard enduser 2D pixel display.
Currently implemented only very small part of this process: The MDegrainN pmode=1 search for best texture view of the small image patch over current tr-scoped frames pool (+tr frame around current, including 'backward' motion compensation to perform dissimilarity metric analysis) and duplicate this texture in output frame sequence to MPEG coder. So it is 100%/full temporal denoising. MPEG coder understand it as scene texture element and not encode it in each frame (skip-block) and only reference it as element from scene textures dataset (some ref frame in h.264 standard) and apply some simple transform (motion) data (zero-MV currently) to create each output time-sampled frame to display. So FHD-sized frame of the about static scene with speaking talent can be encoded to about 500 bytes sized B-frames of h.264 standard (compression ratio of about 6000:1). So object-oriented MPEG codecs expected easily to reach >10000:1 compression ratios for natural movies too.
So that I want from industry-desigen next-gen motion pictures (MPxx) compression standard features:
1. Understand and compensate for many possible transforms (not 2D translate-only as in current MPEGs), but also
- rotate transform (3 axis)
- skew transform
- scale transform
- lighting/shading transform
- may be other possible base geometric transforms of 2D projections of textured+lit 3D objects
2. Have more advanced scene analysis engine to understand physically non-changed scene textures (only damaged by photon-shot noise and/or changed by supported by current MPEG-generation list of transforms). So the encoded data is consist of scene textures dataset and transforms for each part of the scene elements (areas). Also motion-compensation should be expanded in naming to Transforms Compensating meaning not only 2D translate transform can be compensated but much more advanced list of possible geometric and lighting transforms.
Really it is also a part of future denoise engine of mvtools2 development too. So both MPEG compression and temporal denoising still have large part of equal processing. So in some future 'perfect world' expected some final combining of temporal denosie + MPEG encoder into single engine so users no more need to apply MPEG-pre processing with temporal denoise before final MPEG encoder.
Are there any existing solutions moving to this way ?
" announced the "death of MPEG""
MPEG is simply motion pictures experts group. It will dead about at the end of current civilization (really may be soon enough). But after this residual creatures may be not very busy with motion pictures processing. For the next long 'dark ages'.
"I am a bit disappointed by the current state of UHD in broadcasting. We are taking state-of-the-art DVB-T2 muxes being able to broadcast a grand total of 3 channels. With VVC it would increase to a grand total of 5 (meanwhile with H.264 FHD you can fit 6). That's why I think that UHD on broadcasting won't be much of a success unless some further improvement is made on compression."
The tech solution is very simple - do not broadcast unlimited number of junk channels. Not count number of channels as an advantage. Simply put 1 UHD per 8 MHz physical bandwidth. But good designed.
"iptv is becoming more popular so dvb-t2 will become less relevant."
Some creatures really live simply at the surface of the planet but not in the cities. And in the current quickly dying civilization we already have disabled unlimited-traffic internet tariffs of wireless 3G/4G providers a few months ago and last week my pre-paid internet 3G is almost died and support tried to fix it several days long without any good success. So mostly poor text-based internet left only. And no wire or optical provider can make wire/optical connection for poor creatures living far from city and even from multi-room multi-levels buildings in a small self-built homes. So only DVB-T or DVB-S is really only high enough bandwidth way to got some movie content in such places.
kurkosdr
24th March 2023, 16:23
... from still not very died current civilization industry of video codecs design ... will dead about at the end of current civilization (really may be soon enough) ... And in the current quickly dying civilization ...
Dude, can you stop doing this? Doom9 is a technical forum, not a place to live your eschatological fantasies. If you want to do this, there are websites run by cranks who believe in this kind of stuff. They will even help you buy gold and prepper supplies from the affiliate links.
DTL
24th March 2023, 19:43
We as visual tech industry engineers see the internals of the processes at about half a century interval. Though still trying to make this visual industry a bit better. But it is good to correct the efforts with some propositions to the close enough future (about 1/3..1/4 of a century or less).
As about object-oriented encoding I remember I read some promises at time of h.264 developement may be 10..15 years ago. But it looks all hype about object-orienting video compression is now dead ? I see close to nothing move to this really benefitical direction. May be it was found it can not build well universal codec for both broadcasting (requiremnt of a very low channel switching time before decoder can start to feed output decoded frames sequence) and other ways of content delivery ? I understand the torrents-way full-file form of content delivery via IP network before starting of playback is somewhere at the lowset priority at the main industry video codecs developers. Though it may be main way of content compressing and delivery at some communities and they are mostly interested in high quality video compression at the smallest file size per title (low network transfer cost, low storage cost, high number of happy seeds, longlive of release in a network and so on). So the requirement of startup time before decoder can collect all required scene data may be removed in file-based content delivery scenarious and fast enough random file acess playback storages (like HDD or better flash).
So may be future video codecs may divided to broadcast and streaming-fiendly with lower quality and/or higher average bitrate required and title-per-file codecs for private CD/DVD/BD or online-purchased files backup into smallest file size. While the 'local cutscene bitrate' may be as high as required per selected cutscenes and may be much higher strict requirements for broadcast-friendly video codecs.
benwaggoner
24th March 2023, 22:16
Object-oriented compression either requires the content be authored as objects (ala Atmos and MPEG-H) or a good way to extract objects from existing basement media. One can imagine a codec which is basically a stream of GPU instructions to draw the video. However, video is full of things that look like objects but aren't, or don't act like objects, or turn into different kinds of objects. And they all exist under different lighting, and get color timed to get a look that may not have been possible to do in real life. It's not hard to come up with a promising demo like this. RealMedia blew our minds when they remade a South Park episodes as animated vector graphics with a soundtrack. Standard definition that could stream in real time over a 56K modem!
Demos of actual film and video content getting automatically converted into objects and then reconstructed accurately? That I've never seen. Maybe possible with some crazy machine learning system, to some degree, but reconstruction to something that feels shot on 35mm film at 24p with a 180 degree shutter is a couple orders of magnitude beyond what we know how to do today. Getting something understandable, sure. But something that feels like the original seems very challenging.
Of course, codecs have been getting all kind of features that make encoding objects in the source a lot more efficient. All the different CU sizes (including rectangular) and asymetric motion partitions lets us pretty accurately make an object get motion prediction distinct from the background. Intra-frame prediction lets us reuse repeated visual elements (great with text). But since it is fundamentally pixels in and out, stuff that can't be parsed as objects comes through just fine.
I also get nervous about too much ML compression. Bitstream corruption could go from annoying green blocks appearing to different dialog being generated and characters wearing different clothing.
birdie
25th March 2023, 08:53
I also get nervous about too much ML compression. Bitstream corruption could go from annoying green blocks appearing to different dialog being generated and characters wearing different clothing.
Or saying the wrong words since ML is now applied to audio as well. :D
DTL
25th March 2023, 15:49
"Object-oriented compression either requires the content be authored as objects (ala Atmos and MPEG-H) or a good way to extract objects from existing basement media."
It is should be already 'very simple' task for todays 'neural networks' and other 'machine learning'. They can even create some images from sort of several bytes text description.
As I see in practice the total imaging system way to the enduser may go in a bit wrong direction when engineers try to compare video-compression codec with metrics like 'as clear as source' (with typical PSNR/SSIM/VIF and other same metrics). The final goal of the movie pictures content consumption is really the feeling in the head. And it may be not required not only lossless encode the noise part of scene image transfer but also may be some significant part of 'real' scene not greatly add to the total feeling of the movie. It is real challenging task for 'very cool future codecs': analyse total movie content (standard runtime of about 1,5 hours) and create some low-bytes encoded description so at player side the 'decoder' can reproduce very close to this movie frame sequence (may be in different resolution and so on) may be used not very same objects and textures. But giving close to the watching 'initial source' movie feeling in the consumer head. The tested by 'equality metrics' PSNR between source frame sequence and decoder output may be very low. But customer experience and satisfaction will be close to perfect. Depending on the movie description format it may create 'much more 10000:1 compression ratio' and also render any required by enduser resolution.
"One can imagine a codec which is basically a stream of GPU instructions to draw the video."
Currently MPEGs and temporal denoisers works close to 'simple 2D rendering' operating with small patches of an each frame called 'blocks' and either assume the block is the same (100% motion compensation in MPEGs and 'skip' block status) or encode 'residual error' that can not be compensated (really described as a sequence of small sized commands to renderer) with current system supported transforms (or not really geometrically transform at all).
"And they all exist under different lighting,"
The lighting is also from 'transforms' family and very easy to describe in small byte sequence. For example when some part of the scene is shaded by moving object - the shaded blocks only got 'lighting' transform data changed if other geometry and light sources are static. The task for 'temporal denoisers' is equal - detect and 'back compensate' as many transforms as possible to 'distillate' as clean as possible original texture view. After this the 'denoised' frames may be restored as 1 only clean texture database source +number of transforms applied in each output frame (translation, scaling, rotation, lighting and so on). The only difference between temporal denosier and video codec : temporal denoiser uses all this data internally for 'regenerate clean frames view' and output same RAW frames as input. And video codec can output clean textures +transform data in 'smaller bytesize' form and pass to 'decoder' to restore image frames sequence in full size for displaying. Same as we have now with MPEGs using I-frames and ref-frames (?) as short-time texture database and MVs data to describe texture translate-transform only compensation with each output frame render. All other transforms not supported by current MPEGs are treated as 'unsupported' and encoded as residual error imposible to compensate and increase output bitrate and filesize.
I think to add support of 'lignting' transform analysis and 'compensation' into mvtools software in some not very far future - it is easy enough. Also the 'rotate' and 'scale' transform also may be useful in many real cutscenes. So all 'base' SRT (Scale Rotate Translate) transforms will be covered +lighting. Currently as old enough MPEGs it support only 'translate' transform analysis and compensation. But can reuse Translate analysis data ('standard blocks MVs') from hardware MPEG encoder ASIC if it present in the system and have required API via drivers. So if some new version of hardware MPEG encoder ASIC will provide more analysis transforms data (scale, rotate and other) - it can be also easily enough reused for temporal degraining too. But currently Microsoft hardware accelerators API for motion pictures analysis only support Translate transform.
Yes - making motion search engine for analysis of 'rotate' transform in 2 or better all 3 axises will make significant performance penalty to current transtale-only search engine in mvtools so offloading this also simple but compute-loaded work to some ASIC is very 'nice to have' feature. Also if this engine will be reused in MPEG coder it may make significant saving of investments for both denoise and compression tasks.
"Of course, codecs have been getting all kind of features that make encoding objects in the source a lot more efficient. All the different CU sizes (including rectangular) and asymetric motion partitions lets us pretty accurately make an object get motion prediction distinct from the background."
If we currently have all MPEGs from MPEG-1 to may be h.266 only support of analysis and compensation of translate-tranform only it looks the real engineering progress in video codecs design is not very big.
"But since it is fundamentally pixels in and out, stuff that can't be parsed as objects comes through just fine."
There is already sort of partial object-orienting in may be any MPEG - it is block-based compression. Block is small object (can be treated as small texture) and it can be or can not be supplemented with transform data to try to reach better compression in frames sequence.
So encoder can have 2 datapaths:
1. Standard compression of block as samples-array without Transform Compensation (simple M-JPEG)
2. Attempt to analyse full list of supported transforms and create encoding using Transforms Compensation
Next is compare compressed dataset size from path1 and path2 and select the lowest if the decoded image difference with input source is below threshold (quality level of encode param).
So only if block can not reach better compressability after attempt to transform-compensated compression it is encoded as unknown transform or impossible to compensate 'standard JPEG'. Also if some transforms compensation shows benefit - the intermediate path is checked if the partial transform compensation + residual error JPEG-like compression can give less compressed datasize. All other blocks are encoded with lower bits count if Transform Compensation shows close to ideal compensation and residual error close to zero (or below quality-threashold) and the total per timeslice or per file size is reduced. Example of ideally compensated by TC blocks are simply static blocks.
So for the typical cutscene: talking talent at fixed background - all blocks of fixed background are encoded as 100% TC with 'skipped' blocks.
The facial animation of talking talent is created from skin/face texture + transforms dataset (scale, rotate, translate, skew,...) . So also can be better encoded using simple static face single texture database + set of transforms data for each frame.
Also for codecs the hierarchy of database data may be given:
1. Group of frames textures database
2. Cutscene (group of group o frames) textures database.
3. Group of cutscenes texture database.
4. Total movie textures database.
With using more and more higher level of textuure database - more and more data can be skipped from lower levels for the total movie to file compression. For example single talent typically frequently present in many cutscenes of the movie and single texture database may be used. Textures database may be sort of 'macro-used ref frames' for MPEG encoder with 'unlimited' number of ref frames (computed after first pass of analysis of full movie to encode in multi-pass encoding mode). And decoder can have fast random access to total movie textures (ref frames) database for any output frame decoding.
benwaggoner
27th March 2023, 16:34
Yes, ML can generate images from a short text description, which is amazing. But it can't take an image, describe it in some high-level pseudo-semenatic way, and then regenerate the generally same image it started with. Nor is it feasible to have a sequence of images, and have ML code the slight differences between consecutive frames in order to reconstruct the original frames.
Maybe for something simple and of consistent style. For example, a Simpsons episode. There's a huge number of episodes to train on, with a somewhat limited number of art styles, characters and locations frequently on the screen. It's all fundamentally vector animation, which can be expressed quite concisely. I suspect having a cluster of ML models that could generate a movie from scratch would be a lot easier than one that can take existing content, compress it to a much lower bitrate than feasible today, and then reconstruct the same image. When a ML can make a movie, it'll know to only try to make the kind of movies it knows how to make. An input video sequence can be almost anything.
With a lot of handwaving, $1B, and five years I can imagine coming up with some sort of efficient ML system that can make a generic Simpsons episode. But that would still be a generic episode. If they did a modern 3D episode with ray tracing, the ML would have no way to express that accurately. There would still need to be a fallback to more traditional pixels-to-pixels encoding for novel stuff the ML wasn't trained on.
And a technology that can just take moving images and turn them into animated vectors in a pleasing way just doesn't exist, and that's the easier first stage of what you're talking about, ignoring reconstruction.
ML can be very powerful for predicting a right result for cases that land within the range of training data it was given. But trying to go out of that range can result in bizarre results and even "hallucinations" where it provides something plausible but wrong. Training a ML to handle any sort of video that's been made or might be made is an impossible task. Traditional codecs have their downsides, but they are robust. It's nigh impossible to find content that just breaks a codec anymore; at reasonable bitrates, the content is always recognizable as itself, even if degraded.
In a classic information theory sense, we can think of the ML model itself as the codebook that takes the compressed signal that is reconstructed into something similar enough to the original. The more variety and specificity needed, the more training and the bigger ML model that needs to be downloaded and stored is.
ML is able to interpolate missing data in a more analog-like fashion, providing a verisimilitude without the artifacts of traditional encoding. But that verisimilitude means that bitstream or other errors can result in something wrong-but-plausible instead of an obvious artifact. Not good if a different actor's face is used, or text on signage says something else. It's a feature that traditional codecs will lose detail, but not make it up in ways that would change the narrative.
I think the MPAI approach to have something like a traditional codec at the core, enhanced by ML when ML can enhance useful, makes a lot more sense. ML would be great at rendering a pebbled street, since the location of the individual pebbles don't matter. But classic motion compensation would be needed to make sure the pebbles don't randomize when the camera pans. ML has also showed promising results for better deblocking/deringing without detail loss.
DTL
27th March 2023, 20:01
" If they did a modern 3D episode with ray tracing, the ML would have no way to express that accurately."
The 'accuracy measurement' is really quality setting of compression. If user want lower sized compressed file it accept lower accuracy metric. Very close to todays video codecs with PSNR and other really accuracy metrics of how decoded result differs from input. And the final goal of 'highly abstracted' video compression method is to create more or less accurate (similar) feeling in mind in compare with watching original non-compressed content. So it have right to significantly step away from original frames samples values, object positions and other 'look'.
"An input video sequence can be almost anything."
It may be not completely correct. The human-oriented video codec may act close to human vision system - it create some (very simple) 3D scene model using 1 or 2 2D projections and it also learnable (as small babe start to watch surround world it learn how to reconstruct the limited world model in mind using limited size 2D projections from eyes to brain internals). If place this planet trained mature human to significantly different environment it may lost ability to 'space orient' using eyes only and need to re-learn again.
So the typical movie scenes are limited in possible 'samples array values' (not pure white random noise) and describe not infinitely big number of 'typical 3D' or even 'simple 2D' scenes. It is really a sort of shame for current civilization but it was found the number of plots for movies is also not only not infinitely big but also _very_ limited. So it is now very hard for content writers to write something significantly new and good for general public. Sort of quick exhausting crypto valid values from close to infinitely big count of even integer numbers.
"a technology that can just take moving images and turn them into animated vectors in a pleasing way just doesn't exist,"
In simple example it is working way of almost all current 'video compression codecs' - turn input into possible to transform-compensate array of textures (blocks) smaller in size in compare with input frame sequence and try to encode only residual error after transform-compensating. The transform dataset (currently only motion vectors for translate-transform compensation) expected to be much smaller in compare with input block sequence in input frames sequence. Sadly most of MPEGs from beginning only support translate-transform analysis and compensation. From new codecs expected more supported transforms.
And yes - if we increase number of supported transforms - the datasize of transforms set will be somehow bigger. So it is the task for encoder to select best way for each block - either encode it as MJPEG only (no-temporal compression) or make transform-compensation of 1 or more supported transforms +encode residual error and check the smallest possible data output while keeping selected quality level for residual decoding error.
stax76
28th March 2023, 00:34
A related news article that recently appeared in my news feed:
Apple acquired a startup using AI to compress videos. (https://techcrunch.com/2023/03/27/apple-acquired-a-startup-using-ai-to-compress-videos/?guccounter=1&guce_referrer=aHR0cHM6Ly9hcHAucmFpbmRyb3AuaW8v&guce_referrer_sig=AQAAAH2e4_hBg74yy7PXsXnfnGU64uhOeO_ce4a-Rhgl18wBUtwIZmTr481PXrlLTvn1qVRyd_uM0E4WA28FqZibr2CILSnfpntkBJ4x8h5eXYE-borEy0OppE4u_aohk3fWOMuBFcQQ2L9iKg3xBUrhzteoe47mbHxuwUJysn0TkRpJ)
benwaggoner
28th March 2023, 02:39
So it have right to significantly step away from original frames samples values, object positions and other 'look'.
Oh my, we must work in very different markets. In professional content, a slight change in hue or contrast can mean a week of my time figuring out what went wrong and getting a fix implemented. Very senior executives start emailing me when a creative complains that the audio dynamic range is off. We can't change color, dynamic range, cropping, or anything else that would modify creative intent.
Changing the position of objects and the look of the title? Contracts would be updated to bar the use of compression technologies that could result in that.
Full stop for security cameras, news gathering, or anything else where accuracy is more important than verisimilitude. Courts were originally were iffy on using MPEG-2 video in evidence because "it could be changed from what happened."
benwaggoner
28th March 2023, 02:45
A related news article that recently appeared in my news feed:
Apple acquired a startup using AI to compress videos. (https://techcrunch.com/2023/03/27/apple-acquired-a-startup-using-ai-to-compress-videos/?guccounter=1&guce_referrer=aHR0cHM6Ly9hcHAucmFpbmRyb3AuaW8v&guce_referrer_sig=AQAAAH2e4_hBg74yy7PXsXnfnGU64uhOeO_ce4a-Rhgl18wBUtwIZmTr481PXrlLTvn1qVRyd_uM0E4WA28FqZibr2CILSnfpntkBJ4x8h5eXYE-borEy0OppE4u_aohk3fWOMuBFcQQ2L9iKg3xBUrhzteoe47mbHxuwUJysn0TkRpJ)
That's basically ML driven preprocessing, softening less important detail to make it easier to encode. It's not clear if they're doing it baseband only, or with integration in the the encoder itself; noise-reduction enhancement like that works better if it is quantization-aware.
This is a pretty powerful class of techniques to improve quality at low bitrates while maintaining compatibility with existing decoders. The new experimental --mctf filter in x265 is another stab at the same concept, but probably less sophisticated.
ML integration into the decoder side could help restore some of the missing detail in those "less important" regions, as the feel of the texture is more important the the specific pattern of the foliage. That's a kind of loss that creative aren't likely to notice or care about.
DTL
28th March 2023, 18:02
"Oh my, we must work in very different markets. In professional content, a slight change in hue or contrast can mean a week of my time figuring out what went wrong and getting a fix implemented. Very senior executives start emailing me when a creative complains that the audio dynamic range is off. We can't change color, dynamic range, cropping, or anything else that would modify creative intent."
Yes - 'Hi-Fi' and 'Hi-End' markets quality requirements are much higher in compare with 'general public / mass market' . So it may be also not good to attempt to use 'single MPEG version for decade' in so different quality groups. For general public usage and mass-market much more 'scene abstractive' video codecs may be used. Same as MP3 128 kBit/s for general public is much more interesting in compare with HD-Audio with lossless 192 kHz/24bit. But both marketing applications exist.
Also the progressive video codec can be 'receiver-optimized' and use simple fact of 'human vision receive bitrate' - it is really much lower GBit/s so even uncompressed FHD at SDI 1+ Gbit/s is really extra redundant. So if codec is highly-adaptive to enduser visual system it also can have significant benefit while keeping 'good quality feeling' at customer's mind and customer can successfully pay per service. Though the 'internal tech/engineered measured quality of service' will be far from 'Hi-Fi' and 'Hi-End' markets.
I work at general public state broadcasting and I see how awful is 2 Mbit/s h.264 SD (realtime MPEG compression with small GOP size) airing in compare with studio uncompressed FHD SDI source but general public typically make zero complaints on quality (over a years and decades). So it work for years and decades very well.
"Changing the position of objects and the look of the title? Contracts would be updated to bar the use of compression technologies that could result in that."
There is some real secret - if enduser not see the ground-truth source he may not notice any difference if total 'video compression system' is 'smart enough'. Typically in most of broadcasting tasks endusers never see the real source of footage. The typical target marketing goal is some final feeling in mind/head of the customer. And customer generally pay for it. It is not about exact samples/pixels values and very precise scene elements locations.
As I now work on sort of artistic images creation I see how really _most_ of real imaging and ofcourse broadcast imaging typically have _very_poor_ artistic design (close to zero) and even some significant changes can not more ruine the typically absent any artistic value.
It is not about specially designed artistic titles of several years of production with per-scene and per-frame careful and costly fine-tuning and cost about $300 000 000 per 1.5 hours (5400 seconds) of runtime. It is about 99+% of typical content for video compression.
If we can create most of broadcast content priced about $50000 per second of runtime (about $500000 per scenecut) so it may require better video compression.
benwaggoner
28th March 2023, 20:48
The name of the game in my business is "preservation of creative intent." We try to deliver the experience the creators made and intended to be seen as accurately as possible. The ultimate goal is something that would pass an A/B test with the mezzanine (Rings of Power hits that goal). It is accepted that the representation of creative intent will fall short of small devices and low bandwidth. But we never, never, never want to do anything that would change the creative intent to something else. Loss of detail is vastly more acceptable than the introduction of false detail.
Taking your MP3 example, a ML-enabled compression technology would be able to reproduce the same sound in fewer bits by offering better was to reconstruct a complex signal from a compact compressed representation. But the object-based AI you're talking about would be more like taking a .WAV file and then converting it to MIDI. Sure, that could work well for some content, and deliver extremely low bitrates. But if there are novel instruments or effects in there, or lyrics, a MIDI version would express a radically different creative intent.
Of course there will be convergence and gray areas over time. But in general, the more a compression technology bends towards synthesis versus reproduction, the bigger risks to maintaining creative intent there will be. And there will always need to be fallbacks to classic compression approaches for elements that the ML doesn't know how to synthesize accurately.
benwaggoner
28th March 2023, 21:17
One can think of the difference between what can be synthesized and what can't as the residual.
DTL
28th March 2023, 22:21
" But the object-based AI you're talking about would be more like taking a .WAV file and then converting it to MIDI. Sure, that could work well for some content, and deliver extremely low bitrates. But if there are novel instruments or effects in there, or lyrics, a MIDI version would express a radically different creative intent."
Yes - it is very close to sound model: Take RAW WAV at the AI-coder input and it will output MIDI-like dataset for distribution. Though depending on the requested by customer file size (and resulted quality) the used instruments set may be 'standard library' or _partial_ new instruments description may be included into compressed dataset. And the quality of compression (adjustable by user) is the amount of 'real instruments/ textures' used - more or less. Now we have it with current MPEGs - the lower requested bitrate the lower number of texture details and other scene elements enduser got to watch.
"And there will always need to be fallbacks to classic compression approaches for elements that the ML doesn't know how to synthesize accurately."
Yes - the initial implementations of 'highly abstractive' video compressors may be very poor. But expected to become better in quality using 'learning'. As was mentioned current civilization already about exhaust all possible movie plots and accumulated a good library of 'classic movies' over half a century in colours. So it now simply a task to machine-learning compression creators to scan over this classic-movies database and decrease mean residual error of compression. Using as much as possible existing movies in library. Also a 'typical basic database of textures' for decoder (downloadable once) may be created. It expected to be not very large.
Same as MIDI standard instruments library looks like updated once at the OS install or with drivers for playback card or software. And yes - the quality of MIDI playback depends on quality of player hardware (cost). So it simple marketing parameter - the more user pay - the more quality it got at movie (content) playback.
Also it is easily scalable for quality over a range of playback devices from poor smartphones with low power and cost to home-cinema premium playback devices of high compute power. It is also very marketing-friendly.
DTL
29th March 2023, 09:33
"Courts were originally were iffy on using MPEG-2 video in evidence because "it could be changed from what happened.""
It is simply very different applications cases. Because human visual system is _very_ limited in possible to 'understand' datarate the following 2 cases of getting visual information from frame possible:
1. If human creature have very long time for each frame and each frame area analysis - it can get most of visual information presented in a frame using long time sequential frame scan by different (located and sized) areas. It is typical use case of Large Format photoframe consumption - user take hours of viewing of total frame and of different frame areas to consume all data from 16K or much larger resolution frame. If the frame is designed by good artist and have many 'levels' to view.
2. If frame sequence presenting is externally clocked at very high framerate (like 25 frames per second or even more) - the human creature have only chances of capturing some common feeling of the total cutscene presenting and may be some very limited areas of the scene in more detailed form. The general public broadcasting and many other use cases of moving pictures systems are about this case. So wisely designed video compression system can use this feature of enduser to decrease datarate and or filesize.
The limitation of 2. can also harm the user's life in some real life use cases like mission-critical viewing tasks like driving a car at high speed - if human visual system overloaded with high input datarate make skip of some really important part of input scene the serious incident may happen. Though the view was clear and all scene content was presented to the eyes in clear non-distorted form.
FranceBB
6th April 2023, 17:11
The name of the game in my business is "preservation of creative intent." We try to deliver the experience the creators made and intended to be seen as accurately as possible. The ultimate goal is something that would pass an A/B test with the mezzanine (Rings of Power hits that goal).
I would assume that's because Prime has almost exclusively TV Series and movies (i.e creative content in general).
You're lucky, in a way, 'cause the way you get the master is the way you output it.
I mean, from your use case it probably doesn't matter if it's PQ at 23,976, if it's BT709 SDR at 25p etc while for linear broadcasting it's always a matter of performing the "best conversion possible" to respect the airing standard which are 50p BT2020 HLG for UHD, 25i TFF BT709 SDR for FULL HD and 25i TFF BT601 SDR for SD.
Back in my streaming services days (Crunchyroll and Viewster, 2013-2015), I was in the same safe boat as you guys as I didn't really have to change the framerate or the transfer or the primaries etc, but ever since I moved to a TV and started working in linear broadcasting (January 6th, 2016) I had to actively start doing those conversions in the least painful way possible.
By the way, the thing DTL is talking about would kinda work for a news channel, which almost every linear broadcaster has anyway. ;)
I work at general public state broadcasting and I see how awful is 2 Mbit/s h.264 SD (realtime MPEG compression with small GOP size) airing in compare with studio uncompressed FHD SDI source but general public typically make zero complaints on quality (over a years and decades). So it work for years and decades very well.
I feel you. Sky TG24 (the news channel) is still in SD on terrestrial (DTV), so seeing what happens to a perfectly good FULL HD 25i BT709 SDI uncompressed 1 Gbit/s feed from the video mixer is really painful... Luckily, on the satellite it is in FULL HD at decent bitrate, but that only covers around 4 million people and taking into account that the Italian population is 55 million people, one can quickly realize that the sad truth is that the overall majority watches news in SD... :(
benwaggoner
6th April 2023, 17:25
I would assume that's because Prime has almost exclusively TV Series and movies (i.e creative content in general).
You're lucky, in a way, 'cause the way you get the master is the way you output it.
Prime Video is doing a whole lot of sports content and live channels now as well. But yeah, for scripted stuff the job is to reproduce the mezzanine as accurately as possible.
I mean, from your use case it probably doesn't matter if it's PQ at 23,976, if it's BT709 SDR at 25p etc while for linear broadcasting it's always a matter of performing the "best conversion possible" to respect the airing standard which are 50p BT2020 HLG for UHD, 25i TFF BT709 SDR for FULL HD and 25i TFF BT601 SDR for SD.
Yeah, IP delivery makes things much simpler by allowing for more complex output options. And we don't ever have to deliver interlaced or HLG for anything!
Back in my streaming services days (Crunchyroll and Viewster, 2013-2015), I was in the same safe boat as you guys as I didn't really have to change the framerate or the transfer or the primaries etc, but ever since I moved to a TV and started working in linear broadcasting (January 6th, 2016) I had to actively start doing those conversions in the least painful way possible.
I started out in good old analog 480i production, so had sort of an inverse path to yours. And it sure was a breath of fresh air once I'd proved to everyone that we didn't need to do frame rate conversions etcetera.
I feel you. Sky TG24 (the news channel) is still in SD on terrestrial (DTV), so seeing what happens to a perfectly good FULL HD 25i BT709 SDI uncompressed 1 Gbit/s feed from the video mixer is really painful... Luckily, on the satellite it is in FULL HD at decent bitrate, but that only covers around 4 million people and taking into account that the Italian population is 55 million people, one can quickly realize that the sad truth is that the overall majority watches news in SD... :(
Ugh. It was that way in the USA for a long time as well. Hopefully your year-on-year numbers are moving in the right direction. I heard in another meeting today that 50% of EU eyeball hours are OTT already.
It's funny to recall how much of my career was trying to reach a high quality standard def, first on CD-ROM, and then on the web. It's fun to have HD SDR be the fallback now, with creators increasingly considering the 4K HDR the "real" version of the content.
ksec
17th July 2023, 14:52
And ECM 9.1 is out. https://vcgit.hhi.fraunhofer.de/ecm/ECM/-/tree/ECM-9.1
According to the latest test from JVET Geneva meeting, ECM 9.1 finally achieved *overall* 30% BD-Rate vs VTM 11 in Random Access.
benwaggoner
18th July 2023, 17:38
And ECM 9.1 is out. https://vcgit.hhi.fraunhofer.de/ecm/ECM/-/tree/ECM-9.1
According to the latest test from JVET Geneva meeting, ECM 9.1 finally achieved *overall* 30% BD-Rate vs VTM 11 in Random Access.
30% is impressive! Hopefully the irreducible decode complexity increase isn't >>2x to get those gains.
nevcairiel
18th July 2023, 17:59
30% sounds about average, its a similar ballpark as VVC claimed over HEVC at the slower speeds. Much less and adopting a new codec is barely worth it.
ksec
22nd October 2023, 08:19
And ECM 10 is out in the usual place
According to the latest test from JVET meeting, ECM 10 achieved overall 33% BD-Rate vs VTM 11 in Random Access. With lots of improvement going in Low-Delay.
The most impressive thing is we get ~44% of BD-Rate in Text and Graphics in Motion Category. ( The overall result exclude this Category )
30% is impressive! Hopefully the irreducible decode complexity increase isn't >>2x to get those gains.
Compared to VTM we are looking at 8x increase in Encoding Complexity ( Fair but far from perfect ) and 8x increase in decoding. ( oouh )
I am not entirely sure how they intend to achieve 60% BD-Rate ( Yes 60%, not 50% as they are usually aiming at ). We might have hit the end of the S curve in terms of how we could further compress video.
benwaggoner
23rd October 2023, 21:23
And ECM 10 is out in the usual place
According to the latest test from JVET meeting, ECM 10 achieved overall 33% BD-Rate vs VTM 11 in Random Access. With lots of improvement going in Low-Delay.
Yeah, we got the same update at the SMPTE Media Technology Summit last week. Quite impressive progress for just three years since standardization
The most impressive thing is we get ~44% of BD-Rate in Text and Graphics in Motion Category. ( The overall result exclude this Category )
Oh, that is good. Not having screen content extensions be mandatory in Main in HEVC is one of my biggest frustrations from it. Those would make game streaming a lot better.
Compared to VTM we are looking at 8x increase in Encoding Complexity ( Fair but far from perfect ) and 8x increase in decoding. ( oouh )
Yeah, 8x decoder complexity is going to be a hard sell; we've never seen close to that big a jump in successive codecs. Although AV2's ML extensions might mean an atypical jump on the AOM side as well. Certainly the $ and mm^2 for a current-get video decoder are WAY down compared to MPEG-2. Even if we had an 8x jump circa 2026. Moore's Law keeps on giving.
I am not entirely sure how they intend to achieve 60% BD-Rate ( Yes 60%, not 50% as they are usually aiming at ). We might have hit the end of the S curve in terms of how we could further compress video.
At least not without some serious ML on the encoder side to make some more advanced transforms feasible to encode. For example, a spherical warp.
Getting a good film grain synthesis feature built-in is probably the biggest lower-hanging fruit left, as so much of the energy of the most difficult to encode content is grain or other temporally and spatially random noise.
Comcast and others are pushing for AOM's AVFG1 proposal, which would allow any codec to trigger the AV1 FGS post filter via a SEI message. AV1 FGS isn't great, but it certainly can be made better with better parameterization (SVT-AV1 sells it way short). The biggest gap is that the grain synthesis is defined relative to display resolution, not source content resolution, so you get cases where grain particle sizes vary on different displays. Since grain is a physical property, a given piece of synthetic grain should take up a specified percentage of the content, so a low resolution screen should show proportionally smaller grain.
ksec
10th February 2024, 15:54
ECM 11 is out, tiny improvement made on top of EMC 10, very small encoding and decoding efficiency improvement. Looking at possibly merging work on Neural Network to further push down BD-Rate.
Tommy Carrot
10th February 2024, 20:10
Anyone care to share a Windows build of this encoder? I'm very curious how it performs.
Jamaika
11th February 2024, 23:30
https://www.sendspace.com/file/ymwwdd
Tommy Carrot
12th February 2024, 16:22
Thank you very much!
In random access mode unfortunately it crashes at the 2nd frame every time, but i could get it work in allintra mode, so i can at least examine how does it perform for still image compression.
Jamaika
12th February 2024, 18:44
I didn't check codecs. I created encoder only in AVX. I had no time.
ksec
13th February 2024, 10:15
Thank you very much!
In random access mode unfortunately it crashes at the 2nd frame every time, but i could get it work in allintra mode, so i can at least examine how does it perform for still image compression.
Last time I checked on it ( ECM4 I think in late 2022 or early 2023? ) It was ridiculously good. To the point I think we might as well forget VVC and move forward with FVC / H.267.
hajj_3
13th February 2024, 11:17
ecm paper: https://arxiv.org/pdf/2401.02145.pdf
Tommy Carrot
13th February 2024, 12:38
I didn't check codecs. I created encoder only in AVX. I had no time.
No worries mate, i'm grateful for your build anyway.
Last time I checked on it ( ECM4 I think in late 2022 or early 2023? ) It was ridiculously good. To the point I think we might as well forget VVC and move forward with FVC / H.267.
Well, i could not test random access mode yet, but allintra is definitely very promising, still image quality is very noticably better than in VVC, and compared to HEIF, it's a completely different class.
kurkosdr
13th February 2024, 20:41
Last time I checked on it ( ECM4 I think in late 2022 or early 2023? ) It was ridiculously good. To the point I think we might as well forget VVC and move forward with FVC / H.267.
Personally, I still can't believe technology was able to improve significantly beyond HEVC. I had a university course that covered the basics up to AVC in 2014, I have read about the neat tricks HEVC employs to go beyond AVC, but VVC and beyond are a mystery to me. HEVC was already supposed at the limits of entropy and the limits of post-processing.
Also, are the gains for ECM consistent across the resolution range? I mean, AVC can already do 720x576p25 at 2mbps with decent quality, so does this mean that with ECM we'll be able to get 2*0.6*0.6*0.7= 0.5Mbps SD video with decent quality?
benwaggoner
14th February 2024, 20:04
Last time I checked on it ( ECM4 I think in late 2022 or early 2023? ) It was ridiculously good. To the point I think we might as well forget VVC and move forward with FVC / H.267.
VVC is a completed standard starting to go into silicon. H.267 may be the best codec to use in 2030, but I'd rather not have to stick to HEVC until then.
benwaggoner
14th February 2024, 20:14
Personally, I still can't believe technology was able to improve significantly beyond HEVC. I had a university course that covered the basics up to AVC in 2014, I have read about the neat tricks HEVC employs to go beyond AVC, but VVC and beyond are a mystery to me. HEVC was already supposed at the limits of entropy and the limits of post-processing.
Whomever said that had a limited imagination! There are tons of encoding techniques that haven't been implemented into a codec yet. Complex 3D warping. More elaborate forms of texture synthesis than just film grain.
A speculative codec from 25 years now could be a ML kernel that turns into an optimized entropy decoder for other ML kernels that synthesize a movie (including actors, acting, lighting, sets, dialog), with residuals as other ML kernels that correct the first approximation. A future codec could be a superset of H.267, Unreal Engine, ChatGPT, and three other things we've never considered. A video codec is just a stream of bits that turn into a sequence of images.
Also, are the gains for ECM consistent across the resolution range? I mean, AVC can already do 720x576p25 at 2mbps with decent quality, so does this mean that with ECM we'll be able to get 2*0.6*0.6*0.7= 0.5Mbps SD video with decent quality?
Older codecs weren't as well optimized for higher resolutions, so you'll typically see bigger gains there. AVC High Profile, which added 8x8 blocks was to make HD encoding more efficient. AVC Main was originally tuned much more for SD. HEVC can do 32x32 texture units, which is a whole lot better. VVC has some advantages over HEVC for 8K. Reference encoders are SLOW, so there has been a natural tendency to focus on lower resolutions as experiments can be run much faster. 4K is 24x the pixels of 720x480!
We're likely past the point of diminishing returns for improving higher resolutions more than lower ones. For moving images, it's really hard to resolve more than 4K anyway, and pretty much impossible for 24p with 1/48th sec motion blur.
kurkosdr
15th February 2024, 17:24
Whomever said that had a limited imagination! There are tons of encoding techniques that haven't been implemented into a codec yet. Complex 3D warping. More elaborate forms of texture synthesis than just film grain.
A speculative codec from 25 years now could be a ML kernel that turns into an optimized entropy decoder for other ML kernels that synthesize a movie (including actors, acting, lighting, sets, dialog), with residuals as other ML kernels that correct the first approximation. A future codec could be a superset of H.267, Unreal Engine, ChatGPT, and three other things we've never considered. A video codec is just a stream of bits that turn into a sequence of images.
Personally, I'd never use an ML encoder, since it can create things that aren't there. I've made my peace with the concept of lossy compression by considering it a form of clever downscaling in the frequency domain (the weights of the higher frequencies are recorded with less accuracy, that's what the whole "integer-divide the DCT'ed block by the quantization table" does). But ML could create things that were never there. For example, how such a video would be admissible as evidence, and how can we claim any kind of creative intent is maintained? And let's be real, in the real world no residual will be recorded in the name of bitrate reduction, much like both AVC and HEVC are typically pushed to their limits by content providers/broadcasters. But imagine instead of loss of detail and blockiness you get rogue details.
Anyway, back on topic, any good info on what ECM does to achieve its compression gains over VVC (and VVC over HEVC)?
benwaggoner
15th February 2024, 20:16
Personally, I'd never use an ML encoder, since it can create things that aren't there. I've made my peace with the concept of lossy compression by considering it a form of clever downscaling in the frequency domain (the weights of the higher frequencies are recorded with less accuracy, that's what the whole "integer-divide the DCT'ed block by the quantization table" does).
Good description of pretty much all important codecs from JPEG on. It's really down to a perceptually optimized multidimensional scaling at heart.
But ML could create things that were never there. For example, how such a video would be admissible as evidence, and how can we claim any kind of creative intent is maintained? And let's be real, in the real world no residual will be recorded in the name of bitrate reduction, much like both AVC and HEVC are typically pushed to their limits by content providers/broadcasters. But imagine instead of loss of detail and blockiness you get rogue details.
Yeah, "hallucinations" become an increasingly big risk with ML. And even with non-ML techniques to some degree, as compression ratios get more and more complex. Film Grain Synthesis can create a different grain texture. Complex error concealment, deranging, or deblocking can interpolate detail that wasn't there originally instead of just losing detail.
In general principle, the more advanced compression gets, the more plausible the results of a bitstream error will be, as everything but entropy gets squeezed out. Bits spent that can indicate a later bit is wrong is a failure of arithmetic encoding.
One way around this would be deterministic ML - if we have models that always respond the same way to the same input, we can ensure that pixels we see in QA are the same pixels that will be seen anywhere. We can think about this like the transition from MPEG-2 using floating point to iDCT in modern codecs. We don't need future codecs to be compatible with general purpose ML engines for decode; we can continue to specify particular behavior in decoders.
ksec
18th February 2024, 06:48
VVC is a completed standard starting to go into silicon. H.267 may be the best codec to use in 2030, but I'd rather not have to stick to HEVC until then.
VVC 1.0 was standardised in June 2020 at the time it was VTM 9. But in reality it was 2.0 in 2022 April with all compliance and standard test ready. Which at the time it was VTM 16.
I think FVC / ECM 11 right now is much closer to VTM 9 / VVC 1.0 stage. And in terms of Next Gen Codec, this is by far the earliest reach by historic standard. Part of me think that is due to COVID and researcher got insane productivity while being stuck at home.
I still remember I was impressed by VVC during VTM development, well turns out we squeeze another ~30% BD-Rate out of it within 2-3 years time.
But on the other hand it will likely be the first codec that requires hardware decoder to work.
Jamaika
18th February 2024, 16:42
Test VVC encoder
vvencapp: VVenC, the Fraunhofer H.266/VVC Encoder, version 1.6.1-c5da1b5 [Windows][GCC 14.0.0][64 bit][SIMD=SSE42]
Encoder always converts videos to 10bit. Currently frame rate isn't read in libavcodec.
VVEncoderApp_avx.exe -i "113.yuv" -o "output_vvenc_10bit.vvc" -v 5 -t 4 -s 1280x720 --fps 3000/1001 -c yuv420 -ip 256 --internal-bitdepth 10 --bitrate 3Mbps -f 200 --passes 2 --preset medium --level auto --tier main
mp4box.exe -new -add output_vvenc_10bit.vvc output_vvenc_10bit.mp4
Track Importing VVC - Width 1280 Height 720 FPS 3000/1001
VVC Import results: 200 samples (377 NALUs) - Slices: 2 I 0 P 198 B - 0 SEI - 1 IDR - 0 CRA
VVC Stream uses forward prediction - stream CTS offset: 31 frames
[vvc @ 0000019cf30e4ef0] Intra Block Copy is not implemented. Update your FFmpeg version to the newest one from Git. If the problem still occurs, it means that your file has a feature which has not been implemented.
[vvc @ 0000019cf30e4ef0] frame 3, P( 0, 0) failed with -1163346256
[vvc @ 0000019cf30e4ef0] Intra Block Copy is not implemented. Update your FFmpeg version to the newest one from Git. If the problem still occurs, it means that your file has a feature which has not been implemented.
[vvc @ 0000019cf30e4ef0] frame 4, P( 5, 0) failed with -1163346256
[vvc @ 0000019cf30e4ef0] Intra Block Copy is not implemented. Update your FFmpeg version to the newest one from Git. If the problem still occurs, it means that your file has a feature which has not been implemented.
[vvc @ 0000019cf30e4ef0] Intra Block Copy is not implemented. Update your FFmpeg version to the newest one from Git. If the problem still occurs, it means that your file has a feature which has not been implemented.
[vvc @ 0000019cf30e4ef0] frame 6, P( 7, 0) failed with -1163346256
[vvc @ 0000019cf30e4ef0] Intra Block Copy is not implemented. Update your FFmpeg version to the newest one from Git. If the problem still occurs, it means that your file has a feature which has not been implemented.
[vvc @ 0000019cf30e4ef0] frame 8, P( 0, 1) failed with -1163346256
[vvc @ 0000019cf30e4ef0] frame 5, P( 2, 0) failed with -1163346256
VVCSoftware: VTM Encoder Version 23.1-b4dd0e7 [Windows][GCC 14.0.0][64 bit] [SIMD=AVX]
ffvvc doesn't like VTM. Video doesn't play well.
VTMEncoderApp_avx.exe --SummaryVerboseness -c "encoder_randomaccess_vtm.cfg" --InputFile=114.yuv --BitstreamFile=output_vtm_10bit.vvc --SourceWidth=1280 --SourceHeight=720 --FrameRate=29.970 --InputBitDepth=10 --InternalBitDepth=10 --OutputBitDepth=10 --MSBExtendedBitDepth=10 --InputChromaFormat=420 --ChromaFormatIDC=420 --ConformanceWindowMode=1 --FramesToBeEncoded=200 --MatrixCoefficients=1 --InputColorPrimaries=-1 --LMCSSignalType=0 --Level=4 --BDPCM=1 --Tier=main --HashME=1 --IBC=1 --MaxCUWidth=16 --MaxCUHeight=16 --CTUSize=32 --QP=32 --RateControl=1 --TargetBitrate=3000000 --MaxBTLumaISlice=32 --MaxBTChromaISlice=32 --MaxBTNonISlice=32 --MaxTTLumaISlice=32 --MaxTTChromaISlice=32 --MaxTTNonISlice=32 --ColorTransform=0 --VideoFullRange=0 --InputSampleRange=0 --AspectRatioInfoPresent=1 --ChromaLocInfoPresent=1 --OverscanInfoPresent=1 --Log2MaxTbSize=5 --VirtualBoundariesPresentInSPSFlag=1 --EnableDecodingCapabilityInformation=1 --DecodingRefreshType=1 --HrdParametersPresent=0
mp4box.exe -new -add output_vtm_10bit.vvc output_vvenc_10bit.mp4
Track Importing VVC - Width 1280 Height 720 FPS 25000/1000
OpenGOP detected - adjusting file brand
VVC Import results: 48 samples (95 NALUs) - Slices: 3 I 0 P 45 B - 0 SEI - 1 IDR - 2 CRA
VVC Stream uses forward prediction - stream CTS offset: 31 frames
VVCSoftware: ECM Encoder Version 11.0-ce78934 (VTM-10.0-45dfe06) [Windows][GCC 14.0.0][64 bit] [SIMD=AVX]
ECM isn't probably VVC codec. Non-standard.
ECMEncoderApp_avx.exe --SummaryVerboseness -c "encoder_randomaccess_ecm.cfg" --InputFile=114.yuv --BitstreamFile=output_ecm_10bit.vvc --SourceWidth=1280 --SourceHeight=720 --FrameRate=29.970 --InputBitDepth=10 --InternalBitDepth=10 --OutputBitDepth=10 --MSBExtendedBitDepth=10 --InputChromaFormat=420 --ChromaFormatIDC=420 --ConformanceWindowMode=1 --FramesToBeEncoded=200 --MatrixCoefficients=1 --InputColorPrimaries=-1 --LMCSSignalType=0 --Level=4 --BDPCM=1 --Tier=main --HashME=1 --IBC=1 --MaxCUWidth=16 --MaxCUHeight=16 --CTUSize=32 --QP=32 --RateControl=1 --TargetBitrate=3000000 --MaxBTLumaISlice=32 --MaxBTChromaISlice=32 --MaxBTNonISlice=32 --MaxTTLumaISlice=32 --MaxTTChromaISlice=32 --MaxTTNonISlice=32 --ColorTransform=0 --VideoFullRange=0 --InputSampleRange=0 --AspectRatioInfoPresent=1 --ChromaLocInfoPresent=1 --OverscanInfoPresent=1 --Log2MaxTbSize=5 --VirtualBoundariesPresentInSPSFlag=1 --EnableDecodingCapabilityInformation=1 --DecodingRefreshType=1 --HrdParametersPresent=0
mp4box.exe -new -add output_ecm_10bit.vvc output_ecm_10bit.mp4
[VVC] wrong num tile columns 207 in PPS
[VVC] Error parsing NAL unit type 16
[VVC] Error parsing Picture Param Set
[vvc @ 000001692f8cc2b0] sps_delta_qp_in_val_minus1[i][j] out of range: 610, but must be in [0,255].
[vvc @ 000001692f8cc2b0] Failed to read unit 1 (type 15).
[vvc @ 000001692f8cc2b0] Failed to parse picture unit.
[vvc @ 000001692f8c1760] Could not find codec parameters for stream 0 (Video: vvc, none): unspecified size
Consider increasing the value for the 'analyzeduration' (0) and 'probesize' (5000000) options
https://www.sendspace.com/file/ar9enm
benwaggoner
19th February 2024, 21:17
VVC 1.0 was standardised in June 2020 at the time it was VTM 9. But in reality it was 2.0 in 2022 April with all compliance and standard test ready. Which at the time it was VTM 16.
I think FVC / ECM 11 right now is much closer to VTM 9 / VVC 1.0 stage. And in terms of Next Gen Codec, this is by far the earliest reach by historic standard. Part of me think that is due to COVID and researcher got insane productivity while being stuck at home.
I still remember I was impressed by VVC during VTM development, well turns out we squeeze another ~30% BD-Rate out of it within 2-3 years time.
Codec developers have been worried we're running out of ideas for decades. But it never quite seems to happen, thank goodness. And thank Moore's Law allowing us to keep on increasing decode compute requirements some and encode compute requirements a bunch.
But on the other hand it will likely be the first codec that requires hardware decoder to work.
You think? Not having a software decoder for testing would add a lot of friction. Hopefully at least it could be done in a GPU accelerated implementation so we can get work done before HW decoders are available.
Jamaika
23rd February 2024, 18:55
Intra Block Copy VVC errors have been corrected. Videos can be played.
Converting VTM videos requires adding framerate (-r). Video still stutters bit.
https://github.com/ffvvc/FFmpeg/pull/198
In function 'prepare_intra_edge_params_8',
inlined from 'intra_pred_8' at extra/vvc_intra_template.c:627:5:
extra/vvc_intra_template.c:535:21: warning: writing 16 bytes into a region of size 0 [-Wstringop-overflow=]
535 | left[i] = top[i] = left[0];
| ~~~~~~~~^~~~~~~~~~~~~~~~~~
extra/vvc_intra_template.c: In function 'intra_pred_8':
extra/vvc_intra_template.c:625:21: note: at offset -105 into destination object 'edge' of size 3136
625 | IntraEdgeParams edge;
| ^~~~
https://www.sendspace.com/file/f2zj4b
ksec
17th March 2024, 06:41
You think? Not having a software decoder for testing would add a lot of friction. Hopefully at least it could be done in a GPU accelerated implementation so we can get work done before HW decoders are available.
It is definitely the largest increase in terms of decoder complexity. At least not a viable options for vast majority of machines, especially Mobile Phones or Laptops. I dont think we will ever get a dav1d level decoder for VVC or FVC / H.267 ( FVC not being an official name ). And even if we assume we do get someone to write the insane amount of hand written assembly of dav1d. If we consider VVC being similar level of complexity as AV1. And we are looking at 8-10x decoding complexity of AV1 / VVC to FVC. You will need 8x more powerful computer to decode an FVC file with dav1d level of optimisation.
We are looking at an Apple M3 / or current Top Tier CPU from x86 to barely decode 1080P FVC files at 30fps with 100% CPU usage with an dav1d level of decoder ( Which we likely wont get ). I dont expect this level of CPU performance to filter through to the low end any time soon or even within next 10 years. And we haven't even talked about 4K.
So realistically any usage of FVC would requires dedicated hardware acceleration.
But Of course I hope I am utterly wrong.
From April's Meetings.
The rate reduction for natural sequences over VTM 11 in RA configuration for {Y, U, V} increased from ECM-11.0’s {-22.56%, -31.91%, -33.67%} to ECM-12.0’s {-24.01%, -33.20%, -35.34%}.
Tommy Carrot
29th July 2024, 08:19
I've been curious about ECM / H.267 for a while, and i could not find a working windows build of this codec anywhere, and i dont want to beg for builds either, so with some reluctance, i set up visual studio on my work PC (btw, it's cool and all, but i feel ~40 GB for it is a bit excessive.). To my surprise, i could compile ECM encoder without a hitch, no crashes, everything works as it should.
So i could finally try it myself. It is very slow, even slower than the AV2 encoder, so much so that testing it is kinda difficult. (a 100 frames long 720p clip took about 14 hours to finish). But the quality/efficiency is very impressive indeed. Nothing really comes close to it, especially at lower bitrates. It easily outclasses AV2 in it's current stage.
birdie
1st August 2024, 14:24
I've been curious about ECM / H.267 for a while, and i could not find a working windows build of this codec anywhere, and i dont want to beg for builds either, so with some reluctance, i set up visual studio on my work PC (btw, it's cool and all, but i feel ~40 GB for it is a bit excessive.). To my surprise, i could compile ECM encoder without a hitch, no crashes, everything works as it should.
So i could finally try it myself. It is very slow, even slower than the AV2 encoder, so much so that testing it is kinda difficult. (a 100 frames long 720p clip took about 14 hours to finish). But the quality/efficiency is very impressive indeed. Nothing really comes close to it, especially at lower bitrates. It easily outclasses AV2 in it's current stage.
Sounds great except VVenc loses badly to x265 under certain conditions and x266 is nowhere to be found and you're already talking about H.267.
https://github.com/fraunhoferhhi/vvenc/discussions/389
I'm not a fan of how VVC has been deployed so far.
GeoffreyA
1st August 2024, 16:53
Sounds great except VVenc loses badly x265 under certain conditions and x266 is nowhere to be found and you're already talking about H.267.
https://github.com/fraunhoferhhi/vvenc/discussions/389
I'm not a fan of how VVC has been deployed so far.
I think the problem is that VVenC denoises the video during encoding, leading to a softer picture. Recently, working with anime carrying artifacts and quantisation noise, I saw evidence of this. I found that libaom wiped the video clean, leading to an acceptable if soft picture. With VVC, I saw the same effect: the picture was cleaned up, though AV1 did a better job. (At low bitrates, too, it seems to inherit HEVC's, or x265's, characteristic artifacts along lines.) I don't know if VVenC exposes the option to disable denoising, but if it did, video might be sharper.
birdie
1st August 2024, 17:49
I think the problem is that VVenC denoises the video during encoding, leading to a softer picture. Recently, working with anime carrying artifacts and quantisation noise, I saw evidence of this. I found that libaom wiped the video clean, leading to an acceptable if soft picture. With VVC, I saw the same effect: the picture was cleaned up, though AV1 did a better job. (At low bitrates, too, it seems to inherit HEVC's, or x265's, characteristic artifacts along lines.) I don't know if VVenC exposes the option to disable denoising, but if it did, video might be sharper.
There are options for that but they invariably blow up the bitrate:
https://github.com/fraunhoferhhi/vvenc/discussions/388
benwaggoner
1st August 2024, 18:08
I think the problem is that VVenC denoises the video during encoding, leading to a softer picture. Recently, working with anime carrying artifacts and quantisation noise, I saw evidence of this. I found that libaom wiped the video clean, leading to an acceptable if soft picture. With VVC, I saw the same effect: the picture was cleaned up, though AV1 did a better job. (At low bitrates, too, it seems to inherit HEVC's, or x265's, characteristic artifacts along lines.) I don't know if VVenC exposes the option to disable denoising, but if it did, video might be sharper.
I don't think it is denoising specifically, versus getting some higher QPs but with a codec that does a much better job of concealing block and particularly inter block artifacts.
Are you comparing AV1 and VVC at the same bitrates? I'd expect VVC to do somewhat better than AV1 at this in general. Of course, AV1 encoders are a lot more mature at this point.
Reencoding from source that already has video encoding artifacts is always a tricky challenge. Generally the more simlilar the codecs are, the better; reencoding from H.264 to HEVC is generally cleaner than from H.264 to VP9 as HEVC can fall back to pretty much symmetrically encoding the input pixels. Similarly, I'd expect VP9 to reencode better to AV1 than VVC, as they share so much common architecture (but I've not tested that).
As a best practice, cleaning up artifacts before reencoding is preferred. Garbage in is always at least as much garbage out, and often worse than that. The bitrates required to not come out worse than a source already encoded for distribution are often as high or higher than the original bitrate anyway. Reference and enterprise encoders are always tested and tuned on uncompressed sources, as that's what premium content has. Rest assured CrunchyRoll's sources don't have those kind of artifacts!
Given Google's heavy involvement in AV1 encoder development and use via YouTube, it would make absolute sense that libaom was tuned to handle artifacts typical in user-generated content, not just professional mezzanines or uncompressed test sources.
GeoffreyA
2nd August 2024, 12:48
Memory wasn't the best, so I did a quick test again. I would say that, at low bitrates, AV1 and VVC are on the same footing, but could be trading one artifact for another, and that will be subjective. What's certain, though, is that the bad source was "cleaned up," leading to better picture. At higher bitrates, both are preserving the artifacts of the source, leading to worse picture. So, this supports your suggestion that higher QPs, along with concealment, are saving the day. (It could also be that denoising, if present, is being cut down with more bitrate.)
The source is x264-encoded, 5 Mbps, 480p anime supposedly taken straight from the Blu-ray; it appears to be from the same master used for the DVD releases. It is not in the best shape: soft, along with subtle ringing, I'd say. I agree that one should clean up artifacts before encoding, but in this case, AV1 killed two birds with one stone: more compression and better picture (to a limit)!
Generally, though, I find that VVC is slightly ahead of AV1 and a touch sharper.
GeoffreyA
3rd August 2024, 20:04
Looks like there is some sort of denoising after all: MCTF. See birdie's link above as well as the following, p. 16:
https://ieeexplore.ieee.org/stamp/stamp.jsp?arnumber=9503377
FranceBB
3rd August 2024, 20:15
Rest assured CrunchyRoll's sources don't have those kind of artifacts!
Yep, I worked there in 2013 and we used to get Apple ProRes HQ files at 23,976p. Despite not being lossless, they were high quality mezzanine files. The whole licensing thing for Japanese Anime was a very big issue, though, which meant that according to which agreement you had, you could end up either getting a proper master or getting the same "TX Ready" master sent to Japanese Broadcasters and that was unfortunately XDCAM-50 (i.e a 50 Mbit/s Long GOP MPEG-2 stream at 29.970 interlaced TFF with 3:2 pulldown). Luckily it was still 1920x1080, albeit 8bit and with lots of banding. With anime being anime you also had the hard question of "what the heck do I do know" when some shows were clearly 23,976 with 3:2 pulldown BUT had either the opening or the ending 29.970 progressive. Do I deinterlace everything to 29.970p to preserve the opening/ending? Do I just IVTC everything to 23,976p (thus decimating the opening/ending)? The guideline back then was to always encode at 23,976p in those situations which is what we did. One of the worst mezzanine files we used to receive was Naruto Shippuden (I also always found odd that they kept using the same numbering as if the first series and the shippuden were the same thing, so it was like 220 + whatever episode number it was. Like episode 392 of the shippuden was 612). Why I'm mentioning Naruto in particular? Well, because we didn't have access to the TX file it was coming from and the guys at TV Tokyo were creating an high bitrate H.264 progressive file for all the streaming platforms licensing it for simulcast, but they were so not used to it that they didn't quite always encode it right, so you could end up with a 23,976fps progressive file in which some scenes had repeated frames as they were inverse telecined incorrectly. If anything, I think they were deinterlacing to 29,970 and then decimating to 23,976p so sometimes, if the pattern was right after they trimmed the clock, the colorbars etc, it matched and some other times it didn't if it was off by a few frames. :(
Anyway, from there our own mezzanine file was created (an H.264 1920x1080 level 4.1 4:2:0 8bit + AAC in mp4 with very very very high bitrate and almost always 23,976p) and then sent to the distribution encoder (along with the .ass) that would create the renditions automatically for the various resolutions and bitrates to populate the CDN. At that point, the publishing team would test those on the various devices to make sure everything was fine before the content was "unlocked" automatically at the scheduled time.
It feels like a lifetime ago, considering that I've abandoned the streaming sector a long time ago to dedicate to linear broadcasting and I've been working at Sky for 8 and a half years.
Given Google's heavy involvement in AV1 encoder development and use via YouTube, it would make absolute sense that libaom was tuned to handle artifacts typical in user-generated content, not just professional mezzanines or uncompressed test sources.
Yep and we can see this a lot in YouTube videos encoded AV1. While VP9 often presents plenty of artifacts, AV1 streams are generally much softer. This is because the AV1 approach is to blur and average out details rather than showing blocking while being bit-starved, so if the source already had artifacts it's very much possible that those ended up being "cleaned" by the same averaging out process. Whether that was intentional or not I'm not sure, but something tells me that AV1 was probably just trying to average out / smooth out problematic uncorrelated high frequency components and due to more luck than anything compression artifacts from low bitrate sources actually usually end up in that category.
Generally, though, I find that VVC is slightly ahead of AV1 and a touch sharper.
I would be surprised if it wasn't.
The goal by MPEG for H.266 VVC was to at the very least be 35% more efficient than H.265 HEVC (with 40% being the target) and in general VVEnc tests show it to be around 13% more efficient than AV1. Mind you, VVEnc is just 4 years old.
GeoffreyA
5th August 2024, 15:31
Yep, I worked there in 2013 and we used to get Apple ProRes HQ files at 23,976p.
It's interesting hearing about your experience in the industry. Most properly-encoded anime is almost always 23.976. Regarding that poor source I mentioned, namely "Sorcerer Hunters," there were other issues too. Inverse telecining in FFmpeg, with fieldmatch, decimate, and others, worked well for the main 26 episodes; but the OVA proved troublesome and I was left with combing in at least one scene. Same story in VapourSynth. So I've left it for now till I learn more.
Regarding AV1 and softness, I think future codecs will increasingly take this path because that is what is viewed, seemingly by many, as high quality these days. For my part, I find it regrettable.
nevcairiel
5th August 2024, 23:08
Regarding AV1 and softness, I think future codecs will increasingly take this path because that is what is viewed, seemingly by many, as high quality these days. For my part, I find it regrettable.
This is not new, it started with HEVC already, in contrast to H264.
Of course the alternative to softness is blocking artifacts, like H264 was famous for. I rather have softness then blocking artifacts.
Thats why in many cases people instantly preferred HEVC at very low bitrates, since a soft image is more watchable then a blocky H264 encode
The real solution is more bitrate, but we aren't getting that.
Of course thats the same reason new codecs may appear a bit sharper again - at the same bitrate you get a bit more quality, less need to reduce details, eg. less soft.
ksec
6th August 2024, 12:28
I've been curious about ECM / H.267 for a while, and i could not find a working windows build of this codec anywhere, and i dont want to beg for builds either, so with some reluctance, i set up visual studio on my work PC (btw, it's cool and all, but i feel ~40 GB for it is a bit excessive.). To my surprise, i could compile ECM encoder without a hitch, no crashes, everything works as it should.
So i could finally try it myself. It is very slow, even slower than the AV2 encoder, so much so that testing it is kinda difficult. (a 100 frames long 720p clip took about 14 hours to finish). But the quality/efficiency is very impressive indeed. Nothing really comes close to it, especially at lower bitrates. It easily outclasses AV2 in it's current stage.
H.267 ECM ( I believe the original name for it was Future Video Codec aka FVC but I dont see it being used anywhere any more) has roughly 8 - 10x the encoding complexity compared to VVC's VTM.
Right now ECM is showing about 25% BD-R compared to VTM. And up to 50% for Text and Graphics with Motions.
It also increase decoding complexity by 8x, the highest we have seen in any codec generation. I am just not entirely sure how this will work on Mobile. May be we could use it with LCEVC at 720P to reconstruct 1080P files?
Using VVC + LCEVC, which is what the Brazil TV 3.0 are going to be using in 2025, is proving to be high quality and extremely resource efficient. I remember earlier this year Brazil Government and University published a paper showing anywhere from -10% ( using more Bit Rate ) to 60% reduction [1] compared to VVC alone. And this is with early stage VVC and LCEVC encoder.
Hopefully we will have other encoder to play with soon. I hope there will be a Beamr for VVC.
[1] Somewhat interesting is that LCEVC is developed by V-Nova, a British Company and being true to British fashion they absolutely down play the potential of the tech. Which is quite amusing to me. XD
GeoffreyA
6th August 2024, 17:23
This is not new, it started with HEVC already, in contrast to H264.
Of course the alternative to softness is blocking artifacts, like H264 was famous for. I rather have softness then blocking artifacts.
Thats why in many cases people instantly preferred HEVC at very low bitrates, since a soft image is more watchable then a blocky H264 encode
The real solution is more bitrate, but we aren't getting that.
Of course thats the same reason new codecs may appear a bit sharper again - at the same bitrate you get a bit more quality, less need to reduce details, eg. less soft.
I fully agee---there's no contest between the soft image and blocky one---and this has been happening since HEVC. But one does get the feeling, if strictly incorrect, that transparency is harder to achieve than before.
benwaggoner
8th August 2024, 21:15
This is not new, it started with HEVC already, in contrast to H264.
Of course the alternative to softness is blocking artifacts, like H264 was famous for. I rather have softness then blocking artifacts.
Thats why in many cases people instantly preferred HEVC at very low bitrates, since a soft image is more watchable then a blocky H264 encode
The real solution is more bitrate, but we aren't getting that.
Of course thats the same reason new codecs may appear a bit sharper again - at the same bitrate you get a bit more quality, less need to reduce details, eg. less soft.
I don't think there's anything keeping modern codecs from having as much detail as older ones, and in some cases they have tools that allow even greater detail than before. HEVC's transform skip and lossless CU options available in all profiles mean complex graphics with very sharp edges can be encoded perfectly, which isn't always possible in any frequency transform only codec.
It's more that the more advanced the codec, the better in-loop artifact reduction tools it has. Before in-loop deblocking, once a reference frame hit too high a QP, quality was trashed for the rest of the GOP, and you had shorter max GOP durations. Once you could get soft instead of blocky, future frames that referenced a high QP frame could spend bits adding detail instead of trying to erase erroneous detail of blocks.
HEVC added SAO, which does similar stuff for ringing artifacts. VVC is able to do much better with motion vector artifacts than prior codecs, allowing high QP prediction not look as artificial.
These sorts of tools all shift codecs towards having high QP result in just loss of detail instead of detail loss with introduction of wrong detail (ala MPEG-2, where 8x8 block patterns often became painfully visible). With a good encoder, that means it takes fewer bits to hit a certain level of high quality. For noisy content, it can also mean that it's not feasible to save all that many bits over prior codecs without losing some detail.
It also means that bitrates can be pushed much lower without introducing distracting artifacts. If a broadcaster is using fixed RF bandwidth, adding more, softer channels makes compelling economic sense, as you can get away with a lot more compression before customer start complaining. And even for IP streaming, bandwidth costs are a big part of the total cost of the business, so reducing them can help the bottom line a lot.
Still, it's not like we were getting artifact free high detail from those sectors before; it's pretty much premium content delivered over IP where the economics made sense for using enough bits to look consistently good. I'll take softer over soft-with-blocks any day. Although the psychovisual factors there can get complex; the added high frequencies of DCT artifacts in MPEG-2 offered a certain "sizzle" that some customers actually preferred over the uncompressed source.
Another reason we can see early-development encoders tend towards softness is that PSNR is intrinsically biased in that direction with per-frame QP. Encoders need psychovisually optimized adaptive quantization to lower QP in flatter areas to provide a more subjectively balanced encode. SSIM and VMAF are only a little better at properly accounting for the value of preserving detail more in low-detail areas.
ksec
23rd January 2025, 15:09
INSIGHT: Future of Video Compression ITU/ISO Workshop
MPEG and ITU organized a workshop in Geneva titled “Future Video Coding – Advanced Signal Processing, AI, and Standards” as part of preparations for defining the H.267 codec.
I cover here the requirements, with presentations made by Samsung, Amazon, China Mobile, and MainConcept. https://bit.ly/3PClCSq
Samsung
Samsung highlighted the deployment of different codecs across various markets: the MPEG family is used for broadcast/pay TV/OTT, while AOM codecs are prevalent in OTT/social media platforms. The company emphasized that codec efficiency is no longer the highest priority due to advancements in network capabilities (e.g., 5G, Fiber). Samsung's key priorities are:
- Licensing cost (Samsung has implemented all existing codecs: MPEG-2, VP8, VP9, AVC, HEVC, AV1).
- Low complexity (important for encoding on smartphones, smart glasses, and high-end TVs).
Samsung also discussed the evolution of traditional codecs: moving from pre-H.266 codecs based on a collection of algorithmic tools to hybrid deterministic/ML-based tools for H.267, and eventually, to fully end-to-end ML-based encoders (autoencoders) by the 2030 timeframe.
Amazon
Amazon presented various points of innovation and improvement:
- Film grain synthesis (introduced in AV1 and later adopted by VVC).
- Subjective evaluation of new codecs, considering all levels of the ABR ladder.
- Native support for multiview video.
- Error resilience in UDP transport.
- Enhanced handling of banding and low-light conditions.
- Increased use of subjective metrics, such as VMAF.
- Advocacy for AI-based compression that can be updated even after the standard is published.
China Mobile
China Mobile focused on UHD (4K/8K), 5G, and emphasized:
- Lower latency and increased parallel processing.
- UHD capture using portable devices.
- Real-time adaptation of compression parameters based on varying network conditions.
- Balancing compression efficiency with reduced complexity (challenging with AI).
- Including both peak and average bitrates in benchmarks.
- Adaptive compression based on image content, network conditions, device type, and unicast demands.
- Proper handling of AI-generated video.
- Using subjective quality measurements for evaluations.
MainConcept
MainConcept’s presentation was highly grounded, reminding the audience that a codec takes 10–15 years to progress from standardization to large-scale deployment. This means H.267 may only see widespread adoption around 2035–2040.
Their recommendations for H.267 included:
- No additional compression gain compared to VVC.
- Prioritizing resolutions like 1080p and 4K.
- Encoding tools focused on improving quality rather than reducing bitrate.
- Ultra-low latency (sub-frame).
- CPU-friendly decoding.
- Parallel processing support.
- Shorter time-to-market cycles.
https://www.linkedin.com/feed/update/urn:li:activity:7286232368439341056/
Pretty much echoing what I have been saying about compression efficiency, complexity, 5G and quality / bitrate issues for years.
kurkosdr
23rd January 2025, 16:59
Samsung also discussed the evolution of traditional codecs: moving from pre-H.266 codecs based on a collection of algorithmic tools to hybrid deterministic/ML-based tools for H.267, and eventually, to fully end-to-end ML-based encoders (autoencoders) by the 2030 timeframe.
Yay! I can't wait for decoders that hallucinate details that weren't there in the first place. Also, I can't wait for the inevitable scandal when some EBU broadcaster re-broadcasts Euronews or BBC World terrestrially at a low bitrate and the decoder hallucinates details that were never there over some politically important footage and the re-broadcaster gets blamed for messing with the footage. And yes, some EBU broadcasters re-broadcast Euronews or other EBU broadcasters terrestrially (for example, Greece's ERT re-broadcasts BBC World terrestrially).
At least artifacts are recognizable as artifacts, they don't hallucinate things that were never there, not plausibly at least.
FranceBB
23rd January 2025, 22:06
Yeah... I'm also a bit skeptical on the introduction of machine learning in H.267 encoders, basically what Samsung wants to do. The main reason is that they would be inevitably subject to hallucinations and while the current compression artifacts are easily recognizable by everyone, adding machine learning to an encoder could lead to hallucinations which are absolutely critical to avoid, especially in high importance contexts like news and media archival. Even nowadays we have all kind of trickeries implemented on the consumer devices like linear interpolation to increase the framerate, various upscaling and post processing techniques implemented by the various TVs etc, but they're on the consumer devices and they can always be disabled. If we put those things in the encoders, then there's gonna be no escape from it, for everyone, from mezzanine files to distribution files, anything can hallucinate and obviously the lower the bitrate the greater the chances of it happening. If they wanna introduce them, fine, but they should be encoder specific, it should be possible to disable them and they should only allow for things like helping with predictions for a better motion-compensation etc.
By the way, there's also something that wasn't included in the report which I think is important: H.267 expectations are for it to provide a 25% bitrate reduction compared to H.266, which just shows how difficult it's getting to create new more efficient codecs after H.264.
MPEG-2 is 30% more efficient than MPEG-1.
xvid is 10% more efficient than MPEG-2.
H.264 is 40% more efficient than xvid.
H.265 is 35% more efficient than H.264.
H.266 is 30% more efficient than H.265.
H.267 will be 25% more efficient than H.266.
Codec Efficiency Gain (vs. Predecessor)
MPEG-1 ~0% (no predecessor)
MPEG-2 ~30%
Xvid ~10%
H.264 ~40%
H.265 ~35%
H.266 ~30%
H.267 ~25% (estimated)
Are we just going towards a saturation point? Perhaps that would explain why they wanna integrate machine learning stuff (which I'm against obviously).
rwill
23rd January 2025, 23:14
Can't wait for Xerox Moments (https://www.dkriesel.com/en/blog/2013/0802_xerox-workcentres_are_switching_written_numbers_when_scanning) when encoding...
Otherwise just a note: Mpeg-1 and Mpeg-2 have around the same efficiency. It is just that Mpeg-2 supports interlaced.
Mpeg-1 is quite superior to H.261 (50%+ reduction ?) though because H.261 has very limited motion compensation. It is H.261 that has no predecessor
benwaggoner
24th January 2025, 20:18
Can't wait for Xerox Moments (https://www.dkriesel.com/en/blog/2013/0802_xerox-workcentres_are_switching_written_numbers_when_scanning) when encoding...
Otherwise just a note: Mpeg-1 and Mpeg-2 have around the same efficiency. It is just that Mpeg-2 supports interlaced.
Mpeg-1 is quite superior to H.261 (50%+ reduction ?) though because H.261 has very limited motion compensation. It is H.261 that has no predecessor
MPEG-2 also added half-per motion compensation to MPEG-1, which was a significant efficiency improvement.
rwill
24th January 2025, 20:39
I am pretty positive that Mpeg-1 already had Half Pel MC, having implemented Mpeg-1/2 encoders, decoders and the like.
benwaggoner
27th January 2025, 18:44
I am pretty positive that Mpeg-1 already had Half Pel MC, having implemented Mpeg-1/2 encoders, decoders and the like.
You are correct.
And sheesh, that was so long ago now!
modus-ms325c
28th January 2025, 02:03
maybe it's about time we brought MPEG-1/2 back from the ashes...
kurkosdr
28th January 2025, 18:10
maybe it's about time we brought MPEG-1/2 back from the ashes...
What do you mean by "bring MPEG-1/2 back from the ashes"? MPEG-1 or MPEG-2 are still used when compatibility with old equipment or old formats (VCD and DVD-Video accordingly) is desired, they never went away. Especially not MPEG-2.
But use them for modern things like HD? Hell no. Even if you want something royalty-free and ISO standard, MPEG-4 Part 2 is reasonably well-supported and is royalty-free except in Brazil (and will be fully royalty-free sometime in 2026). MPEG-2 is royalty-free except in Malaysia (and won't be royalty-free until sometime in 2035 due to submarine patents in that country). There is no reason to use MPEG-1 or MPEG-2 other than compatibility with old equipment or old formats.
FranceBB
28th January 2025, 23:38
Well, as someone who encodes MPEG-2 stuff regularly, I wouldn't mind if someone got the old encoders polished (HC-ENC, x262, lavc). Ideally it would be nice to have a properly multi thread and numa node aware open source MPEG-2 encoder. I mean, sure, the encoding complexity of MPEG-2 is very low for modern CPUs, but the problem is that most of the time we're limited in the amount of resources used, not to mention that most encoders are either plain C / C++ only or have only old assembly optimizations like SSE2 and could benefit from multithreading and assembly optimizations up to AVX512, but I guess no one will ever spend time doing that...
rwill
29th January 2025, 05:09
Don't think that Mpeg-2 can benefit from anything above SSE2. 16x16 Macroblocks and all.
Regarding good Mpeg-2 encoders, I have done multiple tests over the years and I am still looking for some free encoder to beat my y262 encoder in metrics/subjective quality... but one does not simply switch horses at the end of the race right?
modus-ms325c
29th January 2025, 16:30
Regarding good Mpeg-2 encoders, I have done multiple tests over the years and I am still looking for some free encoder to beat my y262 encoder in metrics/subjective quality... but one does not simply switch horses at the end of the race right?
idk man, maybe work on y262's missing features first since no one else will or is unable to do so for you
rwill
29th January 2025, 18:19
idk man, maybe work on y262's missing features first since no one else will or is unable to do so for you
So... whats missing really?
benwaggoner
29th January 2025, 21:12
What do you mean by "bring MPEG-1/2 back from the ashes"? MPEG-1 or MPEG-2 are still used when compatibility with old equipment or old formats (VCD and DVD-Video accordingly) is desired, they never went away. Especially not MPEG-2.
Are you aware of anyone still making VCDs this decade? If so, do you know why? I would think that those old pre-DVD VCD players used for movie piracy back in the day would have all broken down ages ago.
It's delightful to see ancient tech still being used!
Last I heard, the US Navy was still using VC-1 on submarines.
rwill
30th January 2025, 06:16
High Flyer – Etihad’s inflight entertainment (https://www.broadcastprome.com/case-studies/high-flyer-etihads-inflight-entertainment/)
Thales and Panasonic have different audio and video encoding requirements. Panasonic requires its video files to be encoded at MPEG 4 1.5mbps with 16:9 aspect ratio, and its audio files in mp3 at 128kbps. Thales requires Hollywood movies to be encoded in MPEG 1 at 1.5mbps with an aspect ratio of 4:3 and MPEG 2 in 16:9. All other video content for Thales is encoded in MPEG 1 at 1.5mbps aspect ratio 4:3, while audio files have the same format as Panasonic.
Article is from 2017 and airplanes do not get upgrades often?
modus-ms325c
30th January 2025, 16:01
So... whats missing really?
https://files.catbox.moe/kvqsdx.png
rwill
30th January 2025, 17:27
https://files.catbox.moe/kvqsdx.png
It does not have Frame Threading because it has Slice Threading. Slice Threading is the more sane choice for Mpeg-2.
About Dual Prime and Field Pictures, do you know what these are and where they have a real world use case?
modus-ms325c
31st January 2025, 01:40
I don't know what they are, haven't seen what they do in practice, and was most definitely not the writer of a README that consists of these phrases.
rwill
31st January 2025, 06:53
I don't know what they are, haven't seen what they do in practice, and was most definitely not the writer of a README that consists of these phrases.
Then please be more careful with statements like these:
idk man, maybe work on y262's missing features first since no one else will or is unable to do so for you
modus-ms325c
31st January 2025, 15:06
*** edited ***
avih
31st January 2025, 19:56
@modus-ms325c next time it's a ban. Please behave.
benwaggoner
3rd February 2025, 18:59
High Flyer – Etihad’s inflight entertainment (https://www.broadcastprome.com/case-studies/high-flyer-etihads-inflight-entertainment/)
Article is from 2017 and airplanes do not get upgrades often?
The in-seat entertainment gets updated more frequently than you might think, as newer generations can save a lot of weight, thus fuel, and thus operating expenses. Some of the older systems were quite a few kilos per seat, including wiring.
Some airlines, like Alaska, have ditched in-seat entertainment entirely and just give everyone free WiFi access to the preloaded content library on the plane. That lets them use H.264 + AAC-LC at the minimum, and it would probably work to make it HEVC today. That said, I don't think compression efficiency is directly tied to cost savings anymore, so they'll emphasize compatibility over that. Once you've got the WiFi to handle one codec, a more efficient one doesn't help that much. And while better compression could store more titles, storage is also getting cheaper on its own. I expect that they're more limited by licensing than storage anyway.
kurkosdr
4th February 2025, 00:54
It does not have Frame Threading because it has Slice Threading. Slice Threading is the more sane choice for Mpeg-2.
About Dual Prime and Field Pictures, do you know what these are and where they have a real world use case?
If memory serves me well, a "field-picture" encodes a pair of interlaced fields (top and bottom field). Basically, in MPEG-2, in interlaced mode, you can either have a picture encoded as a progressive frame or as a pair of fields. The first type of picture is useful for "fake interlace" (needed to get a 25p video on DVD-Video because DVD-Video doesn't support progressive mode for example) or for scenes with zero motion in true interlaced video, and the other is for scenes with motion in true interlaced video.
Note: when I say "true interlaced video" I mean the kind of interlaced video that would give you those annoying comb artifacts during motion if you tried to weave/overlay its field-pairs into frames.
If so, does it mean y262 can't encode true interlaced video?
rwill
4th February 2025, 07:44
If memory serves me well, a "field-picture" encodes a pair of interlaced fields (top and bottom field). Basically, in MPEG-2, in interlaced mode, you can either have a picture encoded as a progressive frame or as a pair of fields. The first type of picture is useful for "fake interlace" (needed to get a 25p video on DVD-Video because DVD-Video doesn't support progressive mode for example) or for scenes with zero motion in true interlaced video, and the other is for scenes with motion in true interlaced video.
Note: when I say "true interlaced video" I mean the kind of interlaced video that would give you those annoying comb artifacts during motion if you tried to weave/overlay its field-pairs into frames.
If so, does it mean y262 can't encode true interlaced video?
You have not understood interlaced support in Mpeg-2.
y262 can encode interlaced content just fine.
kurkosdr
4th February 2025, 11:46
You have not understood interlaced support in Mpeg-2.
y262 can encode interlaced content just fine.
Then what are "field pictures"? Any hint?
rwill
4th February 2025, 17:12
Then what are "field pictures"? Any hint?
Well maybe I am the wrong one to answer that. As far as I know a field picture in the context of Mpeg-2 is a coded picture where picture_structure is equal to "Top field" or "Bottom field".
If you need more information: The whole internet is at your fingertips.
hajj_3
4th February 2025, 17:16
This thread should be integrated into this thread: https://forum.doom9.org/showthread.php?t=185662
kurkosdr
4th February 2025, 17:29
Well maybe I am the wrong one to answer that. As far as I know a field picture in the context of Mpeg-2 is a coded picture where picture_structure is equal to "Top field" or "Bottom field".
If you need more information: The whole internet is at your fingertips.
That's what the internet tells me, a "field picture" is basically a field. If you have a 720x576i50 video, a field picture will be a 720x288 picture to encode a field (either top or bottom). Because if you have motion in that particulate field-pair and try to weave the two fields into a 720x576 frame and encode it as a "frame picture", you'll get combing artifacts that will need lots of bitrate to compress without massive artifacting. So, better encode that field-pair it as two pictures 720x288 each in "field picture" mode, not "frame mode".
Anyway, have you tried true interlaced (with combing artifacts when you weave/overlay) with y262? How well does it perform, say at 4mbps for 720x576 during motion scenes?
rwill
4th February 2025, 18:25
Anyway, have you tried true interlaced (with combing artifacts when you weave/overlay) with y262? How well does it perform, say at 4mbps for 720x576 during motion scenes?
Well of course I have. It performs. I do not know how well because I had skipped the stuff with the reference.
Asmodian
5th February 2025, 03:59
Anyway, have you tried true interlaced (with combing artifacts when you weave/overlay) with y262?
I would not call that 'true interlaced'. I would call that treating interlaced video as progressive. It is never good. If a codec cannot encode as fields it is better to bob deinterlace than encode the two fields from different timepoints in the same frame.
rwill
5th February 2025, 07:51
I would call that treating interlaced video as progressive.
What kurkosdr described, with is own words, appears to be interlaced content. I do not know why you suddenly come along with treating it as progressive content. No one talked about that. Can you explain further?
Asmodian
6th February 2025, 22:42
I mean encoding an interlaced frame as a single picture with combing artifacts (treating it as a single progressive frame) results in much worse quality than encoding the two fields as seperate pictures and weaving during decode.
The combing artifacts are very sharp high frequency detail that does not work well with our lossy frequency domain compression methods. This is even worse when considering 4:2:0 chroma sampling.
rwill
7th February 2025, 07:10
I know this thread is originally about a H.266 successor but I just have to ask, do you guys have ever heard about field-based prediction in a frame picture in the context of Mpeg-2 video ?
I mean it all started with modus-ms325c where I had hopes he can tell me what features to add to y262 so it might see wider use, but I got disappointed.
Then kurkosdr and you came along and, if I understood you right, tried to argue that you need field pictures to encode interlaced content. But this is just not true, at least for Mpeg-2 video. I got disappointed again.
benwaggoner
7th February 2025, 19:49
The combing artifacts are very sharp high frequency detail that does not work well with our lossy frequency domain compression methods. This is even worse when considering 4:2:0 chroma sampling.
Yeah, interlaced-as-progressive is very common in the worst video on the internet. Very early YouTube, along with anamorphic-as-square and limited-range-as-full.
Interlaced was the whole point of 4:2:2, really. I'm not sure why it remains such a big deal in production and post for progressive-only titles. It seems like 4:2:0 or 4:4:4 should be used.
Asmodian
10th February 2025, 00:04
do you guys have ever heard about field-based prediction in a frame picture in the context of Mpeg-2 video ?
You mean adaptive field/frame compression? Changing between field or frame based encoding per-macroblock? My experiance (a long time ago now) was that this does not work as well for interlaced video compared to always using field pictures. I always got some blocks that should have been encoded field based, but were encoded frame based instead.
It might not be a big deal for y262 adoption, hopefully no one is encoding interlaced video today, but it is a feature I would want for encoding camcoder VHS tapes into MPEG2.
benwaggoner
13th February 2025, 22:53
You mean adaptive field/frame compression? Changing between field or frame based encoding per-macroblock? My experiance (a long time ago now) was that this does not work as well for interlaced video compared to always using field pictures. I always got some blocks that should have been encoded field based, but were encoded frame based instead.
I mean interlaced, period. Scripted entertainment has been exclusively progressive for years now. We still see some interlaced production for broadcast and sports, but that's declining, and mainly for legacy stuff that doesn't support HEVC anyway.
It might not be a big deal for y262 adoption, hopefully no one is encoding interlaced video today, but it is a feature I would want for encoding camcoder VHS tapes into MPEG2.
Yeah, having 4:2:2 is definitely essential for archiving interlaced sources.
kurkosdr
16th February 2025, 22:11
I mean interlaced, period. Scripted entertainment has been exclusively progressive for years now. We still see some interlaced production for broadcast and sports, but that's declining, and mainly for legacy stuff that doesn't support HEVC anyway.
Personally, I'd prefer 50p/60p at 1080p resolution was universally supported even in H.264 hardware. Let the broadcaster/encoder side decide how they want to de-interlace the legacy content, not the de-interlace filter in my TV that might or might not mess it up. At least we got it with HEVC, since 50p/60p is supported across the resolution range in pretty much all HEVC hardware, even at 2160p resolution.
IMO, allowing interlaced on anything other than 480i and 576i was a huge mistake. If my old CRT TV can't handle the new HD resolution, why allow a mode (interlaced) that's tailor-made for my old CRT TV for those higher resolutions? You have to convert anyway.
Z2697
17th February 2025, 06:25
You just confused the relationship of "macro block level interlacing" and "field pictures".
Which is nothing beyond "they both used for interlaced content".
The former is what you are really talking about. Which is supported in y262.
Field pictures is like the way (and the only way) HEVC handles interlaced contents today.
FranceBB
17th February 2025, 10:15
allowing interlaced on anything other than 480i and 576i was a huge mistake.
Yes... and you know what's the most ironic thing in all this? We're not even shooting interlaced anymore, we're shooting 2160p at 50fps, the whole thing is produced progressively even on live events that go through video mixers etc and then the matrix/transfer/primaries conversion + downscaling + dithering + fields division to create the 25i TFF FULL HD version and the 25i TFF SD version occurs at the very end of the chain... So yeah, it could have been 50p from beginning to the end if only they allowed us to do that, but unfortunately the choice at the time was between HD (1280x720 50p) and FULL HD (1920x1080 25i TFF) and the overwhelming majority of broadcasters picked the latter with just a few exceptions like NRK (the Norwegian public broadcaster) which went to HD 50p.
At least we got it with HEVC, since 50p/60p is supported across the resolution range in pretty much all HEVC hardware, even at 2160p resolution.
Yeah... The great advantage of H.265 for me was moving from interlaced to progressive 50p and from 8bit planar to 10bit planar. That on its own would have been a good enough reason to move to H.265, aside from the whole UHD resolution bump, BT2020 wide color gamut and later also HDR PQ/HLG transfers.
benwaggoner
24th February 2025, 17:43
Yes... and you know what's the most ironic thing in all this? We're not even shooting interlaced anymore, we're shooting 2160p at 50fps, the whole thing is produced progressively even on live events that go through video mixers etc and then the matrix/transfer/primaries conversion + downscaling + dithering + fields division to create the 25i TFF FULL HD version and the 25i TFF SD version occurs at the very end of the chain... So yeah, it could have been 50p from beginning to the end if only they allowed us to do that, but unfortunately the choice at the time was between HD (1280x720 50p) and FULL HD (1920x1080 25i TFF) and the overwhelming majority of broadcasters picked the latter with just a few exceptions like NRK (the Norwegian public broadcaster) which went to HD 50p.
We almost killed interlaced with ATSC 1.0. The idea of playing broadcast on computers was really exciting back then, and Intel was adamant that interlaced just couldn't be done well on the PC. Broadcasters fought for interlaced, but were about to cave to Intel when some Intel engineer found a lousy-but-functional way to do it, and here we are today.
(as I remember from Joel Brinkley's history of HD Defining Vision, the greatest video standards true crime page-turner thriller of our era. Highly recommended for anyone interested in video standards or history. I hope they make a TV series of it.)
Yeah... The great advantage of H.265 for me was moving from interlaced to progressive 50p and from 8bit planar to 10bit planar. That on its own would have been a good enough reason to move to H.265, aside from the whole UHD resolution bump, BT2020 wide color gamut and later also HDR PQ/HLG transfers.
10-bit was really the killer feature for HEVC, as it enabled HDR-10m with compression efficiency at UHD resolutions #2. Those drove very quick and broad implementation for the codec when it was only a few years old (although both 10-bit and HDR were pretty late additions, and there wasn't PQ optimization built in). It's universal in mobile devices and near universal in living room devices. Most new devices had HEVC decoders by the time the standard was as old as VVC is now.
vBulletin® v3.8.11, Copyright ©2000-2026, vBulletin Solutions Inc.