View Full Version : Alliance for Open Media codecs
hydra3333
22nd September 2018, 02:39
Excelent post :thanks: And +1
Tommy Carrot
22nd September 2018, 04:00
If anyone is curious how the encoding speed improved in the last half year or so, here's a little comparison between the different versions:
version | enc time | filesize
0.1.0-9348 | 152 | 183494
0.1.0-9559 | 119 | 182391
0.1.0-9658 | 109 | 182118
1.0.0-6 | 94 | 182856
1.0.0-82 | 86 | 182952
1.0.0-181 | 61 | 184139
1.0.0-245 | 59 | 189154
1.0.0-399 | 51 | 180221
1.0.0-541 | 22 | 184505
1.0.0-629 | 15 | 184616
I used cpu-used=1 in all cases, and tried to match the bitrates as close as possible in constant quality mode. The encoding speed is still very slow, but as you can see, it finally started to improve at a higher pace lately (although the last 2 builds used CONFIG_LOWBITDEPTH=1, so they are not really comparable). Aomenc is still way too slow for any productive purpose, but at least it's possible to test it now.
Nintendo Maniac 64
22nd September 2018, 04:46
Excelent post :thanks:
And +2
fight between Rider and Saber for whoever knows what I'm talking about
The original visual novel (https://vndb.org/v11) is better. :p
(and that's not all all because I've had a hand with some of the media stuff in the latest and upcoming unofficial English PC versions of "Fate/stay night Realta Nua" (http://forums.nrvnqsr.com/showthread.php/4745-Fate-Stay-Night-Realta-Nua-PC-version-Mirror-Moon-TL-insertion-project/page585), nope no way what 'choo talkin' bout Willis you crazy)
NikosD
22nd September 2018, 08:33
If anyone is curious how the encoding speed improved in the last half year or so, here's a little comparison between the different versions:
version | enc time | filesize
0.1.0-9348 | 152 | 183494
.
.
1.0.0-629 | 15 | 184616
It seems it's an order of magnitude faster.
10 times faster is a lot, but obviously it has further big margins to improve.
I wonder how long is going to take for another 10 times improvement.
SmilingWolf
22nd September 2018, 09:20
The original visual novel (https://vndb.org/v11) is better. :p
(and that's not all all because I've had a hand with some of the media stuff in the latest and upcoming unofficial English PC versions of "Fate/stay night Realta Nua" (http://forums.nrvnqsr.com/showthread.php/4745-Fate-Stay-Night-Realta-Nua-PC-version-Mirror-Moon-TL-insertion-project/page585), nope no way what 'choo talkin' bout Willis you crazy)
I'd have you fight me IRL for dissing anything ufotable, but that evened out after you gave me a reason to install FSN again :D
(And for being on the first line in the works, thanks to y'all for what you do!)
Also, the pre-edit "the original visual novel from 2004" was just fine, I didn't choose FSN just because the clips look good ;)
I wonder how long is going to take for another 10 times improvement.
That's an interesting question.
The 0.1.0-9348 (https://forum.doom9.org/showthread.php?p=1840615#post1840615) build is from April, so that's 146 days (https://www.timeanddate.com/date/durationresult.html?d1=29&m1=4&y1=2018&d2=21&m2=9&y2=2018&ti=on) between the first and the last tested build. Now, of course these things don't really follow a linear time-optimization correlation, buuuuut I wonder where we'll be in another 4 months :)
Maybe they'll get cpu-used=4 on par (speed wise) with x265's placebo-and-then-slowed-down-some-more, using the no-wpp etc. etc. settings benwaggoner suggested a few pages ago (https://forum.doom9.org/showthread.php?p=1849998#post1849998), which in my limited testing came around at half aomenc's time a week or two ago.
Also rav1e might become a real game changer for av1 encoding in the same timeframe. I surely hope so.
And a small heads up, Video Dev Days (https://www.videolan.org/videolan/events/vdd18/) have officially begun, and today they're going to showoff dav1d, their very own decoder!
Exciting times ahead!
Selur
22nd September 2018, 09:35
And a small heads up, VideoLAN Developer Days have officially began, and today they're going to showoff dav1d, their very own encoder!
decoder not encoder
according to: https://www.videolan.org/videolan/events/vdd18/#saturday
SmilingWolf
22nd September 2018, 09:40
Aye, caffeinated typo slipped in. Corrected, thanks!
EDIT: in the meantime, the presentation has finished. Does anyone know if there's a livestream or at least a youtube channel where they upload stuff as they go?
nevcairiel
22nd September 2018, 09:48
EDIT: in the meantime, the presentation has finished. Does anyone know if there's a livestream or at least a youtube channel where they upload stuff as they go?
Unfortunately, I don't think so. Its rather hilarious that a open-source multimedia conference can't figure out streaming or at least on-demand videos afterwards. :)
Mr_Khyron
22nd September 2018, 14:49
dav1d source code
https://code.videolan.org/videolan/dav1d
:)
SmilingWolf
22nd September 2018, 15:14
dav1d source code
https://code.videolan.org/videolan/dav1d
:)
Uwheee!
dav1d-0.0.1-7-bb521ef9: https://mega.nz/#!w9xlxZpT!_KfDw7rp9LWMosjNIdSZeuqhgOr7PaHkG8eJgoHLwrg
EDIT: botched on Windows because of the usual open(file, "r") bug. Needs to be open(file, "rb") for binary objects, lest files be treated as text and the input mangled.
Will submit a patch ASAP
EDIT2: https://code.videolan.org/videolan/dav1d/merge_requests/12/diffs
Zebulon84
22nd September 2018, 15:48
Previous Video Dev Days conferences are available on Vimeo (https://vimeo.com/videolan/collections) so I guess it will be the same for 2018. But videos for 2017 Video Dev Days were posted last may, so we may have to wait quite a while.
hajj_3
22nd September 2018, 19:00
Previous Video Dev Days conferences are available on Vimeo (https://vimeo.com/videolan/collections) so I guess it will be the same for 2018. But videos for 2017 Video Dev Days were posted last may, so we may have to wait quite a while.
possibly, the 2016 ones were posted on their youtube channel in a timely manner: https://www.youtube.com/user/VideoLANorg/videos
Nintendo Maniac 64
23rd September 2018, 04:28
Now the wait begins for dav1d to be implemented into LAVfilters and for my low multi-threaded utilization woes to be gone once and for all!
...please tell me I'm not being overly optimistic by saying that? The way I see it, a Xeon x3470 shouldn't be all too different from the likes of an i7-8550U seeing as the Xeon has a ~33% IPC deficit while the i7 has a ~33% clockrate deficit, so the end result should be pretty similar (not counting AVX anyway).
I mean, if its too slow on an i7-8550U, then it'll be too slow for almost all laptops in existence except for maybe those 6core i9 laptops and those crazy gamer/professional laptops that have full-fat desktop CPUs like that Asus with a Ryzen 1700.
you gave me a reason to install FSN again :DJust as long as it's that upcoming "merge project" mentioned on the last couple of pages at the end of the linked thread, otherwise I'll never forgive you. :p
(to clarify, the original visual novel only renders at 800x600 while this "merge project" will have HD visuals and high-res text - that alone should make it a no-brainer)
Also, the pre-edit "the original visual novel from 2004" was just fineI only removed the "2004" text because I thought using "original" would be clearer to convey that the visual novel came before any think else Fate-related.
SmilingWolf
23rd September 2018, 10:02
Now the wait begins for dav1d to be implemented into LAVfilters and for my low multi-threaded utilization woes to be gone once and for all!
...please tell me I'm not being overly optimistic by saying that? The way I see it, a Xeon x3470 shouldn't be all too different from the likes of an i7-8550U seeing as the Xeon has a ~33% IPC deficit while the i7 has a ~33% clockrate deficit, so the end result should be pretty similar (not counting AVX anyway).
I mean, if its too slow on an i7-8550U, then it'll be too slow for almost all laptops in existence except for maybe those 6core i9 laptops and those crazy gamer/professional laptops that have full-fat desktop CPUs like that Asus with a Ryzen 1700.
I'm afraid we're not quite there yet. Using my most CPU-intensive clip (PresageFlowerFight, 1080 encoded with as many tiles as possible) I have only been able to observe a 45% max utilization with --framethreads 8
My CPU is an i7-4770, with 4/8 cores
Some timings:
# time ./dav1d.exe --framethreads 8 -o /dev/null --muxer yuv4mpeg2 -i Fight.cq20.1080p.ivf 2> /dev/null
real 0m19,460s
user 0m0,000s
sys 0m0,000s
# time aomdec.exe --threads=8 -o /dev/null Fight.cq20.1080p.ivf
real 0m5,170s
user 0m0,000s
sys 0m0,000s
Right now dav1d is implemented in pure C and of course it shows. It'll be more interesting after ASM optimizations start trickling in.
EDIT: after playing a bit with the numbers I've been able to push the times down a bit further and CPU util higher (50-53% range):
# time ./dav1d.exe --framethreads 6 --tilethreads 2 -o /dev/null --muxer yuv4mpeg2 -i Fight.cq20.1080p.ivf 2> /dev/null
real 0m17,800s
user 0m0,000s
sys 0m0,000s
I only removed the "2004" text because I thought using "original" would be clearer to convey that the visual novel came before any think else Fate-related.
I see, I see (https://www.youtube.com/watch?v=R8y1GILBUcc)
hajj_3
23rd September 2018, 10:13
MPC-BE 1.5.2.3976 Beta has been released which can play my 1920x800 551kbps bitrate AV1 file that i have smoothly, used up to 46% cpu on core i3-7100U 2.4ghz
Link: https://sourceforge.net/projects/mpcbe/files/latest/download
LigH
23rd September 2018, 17:41
New uploads: (MSYS2; MinGW32: GCC 7.3.0 / MinGW64: GCC 8.2.0)
AOM v1.0.0-643-gaf3e5cc66 (https://www.mediafire.com/file/z697gkw8877yhqu/aom_v1.0.0-643-gaf3e5cc66.7z)
rav1e 0.1.0 (d330de0 / 2018-09-23) (https://www.mediafire.com/file/u5wrgqg77ehukm0/rav1e_0.1.0_2018-09-23_d330de0.7z)
dav1d 0.0.1 (5e05e65 / 2018-09-23) (https://www.mediafire.com/file/djxkwy12x2cv3wq/dav1d_0.0.1_2018-09-23_5e05e65.7z)
marcomsousa
24th September 2018, 12:24
AV1 already have 3 encoders and/or decoders, it's a good start..
olduser217
25th September 2018, 04:33
I heard that Ateme demonstrated 4k60 AV1 decoding on TV (seems like LG TV) during the recent IBC Show.
Anybody attended the show and saw it?
marcomsousa
25th September 2018, 10:03
I heard that Ateme demonstrated 4k60 AV1 decoding on TV (seems like LG TV) during the recent IBC Show.
Anybody attended the show and saw it?
At NAB Conference:
ATEME: Video (https://www.youtube.com/watch?v=7_731UnjMuo)
Other AOMedia: Video1 (https://www.youtube.com/watch?v=CVFWZbolHhE) Video2 (https://www.youtube.com/watch?v=k6WSqb4O2x0) Video3 (https://www.youtube.com/watch?v=NW3h0s2E4EQ&list=PL97T7zfqOOF108asipLcqHAdTOkCGsFsa&index=3) Video4 (https://www.youtube.com/watch?v=aAayA_2jCb8)
SmilingWolf
25th September 2018, 14:17
32/64bits binaries:
av1-1.0.0-654-gd0076f507: https://mega.nz/#!kto20KoR!XbcrlXv7QZFRzks38Jqm8oss8P3IDlZx3X0tfeVwJx4
64bits binaries only (I don't have a 32bits toolchain in my MSYS2 env):
dav1d-0.0.1-37-9075f0e: https://mega.nz/#!co5wFQyQ!EZhG33K4zBp6MHbEqfFdNc3p2qvIKmdYdPPtDJqCWBw
The dav1d build has been patched so that the input/output files are opened in binary mode, for the reasons previously explained. Band-aid solution until the PR is accepted upstream.
savage747
25th September 2018, 21:18
I dusted off my own automated codec testing setup so I can track AV1-development. Of course, objective metrics do not replace subjective testing, but this should still be okay to get some very rough estimates.
The following graph takes about 45 minutes to generate on my Ryzen 2700 (8 cores, 16 threads). To keep things sane, I'm using --cpu-used=4, which is a somwhat "fast" setting.
The modern codecs may perform even better on HD resolutions.
benwaggoner
26th September 2018, 00:56
I dusted off my own automated codec testing setup so I can track AV1-development. Of course, objective metrics do not replace subjective testing, but this should still be okay to get some very rough estimates.
It can be VERY rough. It would be pretty typical for a new psychovisual feature to result in improved quality AND reduced scores with objective metrics. That is pretty much the defining feature of a psychovisual optimization :).
The following graph takes about 45 minutes to generate on my Ryzen 2700 (8 cores, 16 threads). To keep things sane, I'm using --cpu-used=4, which is a somwhat "fast" setting.
The modern codecs may perform even better on HD resolutions.
HEVC certainly has a bigger advantage over H.264 at UHD resolutions.
It's hard to say with AV1 since the encoders are so slow that there isn't a substantial corpus of high-quality HD AV1 encodes to evaluate yet. In theory it should also scale well, but today's encoders obviously haven't been able to get a lot of resolution-tuned psychovisual optimization yet.
A whole lot of AV1's discussed capabilities are more informed speculation than anything based on real-world demonstrations of real-world scenarios.
IgorC
26th September 2018, 01:36
Would it make sense just drop a bitdepth of 8 bits already? Especially now when AV1 was just realesed. All right, decoder optimizations and hardware support will come anyway.
Yes, 8 bits is faster to encode/decode but banding ruins a major part of quality gains at quite large range of bitrates.
VP9 8-bits vs 10 bits is a day and night difference.
https://sonnati.files.wordpress.com/2016/06/10bit2.png
https://sonnati.wordpress.com/2016/06/17/does-vp9-deserve-attention-part-ii/
Even VP9/HEVC 10-12bit are already so advanced to produce block-free video even at quite low bitrates.
I wouldn't be surprised if AV1 8 bits would look worse than VP9/HEVC 10-12 bits (or just comparable)
olduser217
26th September 2018, 03:28
At NAB Conference:
ATEME: Video (https://www.youtube.com/watch?v=7_731UnjMuo)
Other AOMedia: Video1 (https://www.youtube.com/watch?v=CVFWZbolHhE) Video2 (https://www.youtube.com/watch?v=k6WSqb4O2x0) Video3 (https://www.youtube.com/watch?v=NW3h0s2E4EQ&list=PL97T7zfqOOF108asipLcqHAdTOkCGsFsa&index=3) Video4 (https://www.youtube.com/watch?v=aAayA_2jCb8)
Thanks for the video link.
Just curious whether the Ateme demo was done with existing TV hardware (assuming software decoding with ARM based SOC), or the TV was just been used as a display for decoding done with PC.
Blue_MiSfit
26th September 2018, 08:01
Well, they were using the player in the TV (at least, it looks like the standard LG player UI) - so my guess is a special engineering prototype TV with an FPGA or something in it :devil:
benwaggoner
26th September 2018, 16:57
Would it make sense just drop a bitdepth of 8 bits already? Especially now when AV1 was just realesed. All right, decoder optimizations and hardware support will come anyway.
Yes, 8 bits is faster to encode/decode but banding ruins a major part of quality gains at quite large range of bitrates.
VP9 8-bits vs 10 bits is a day and night difference.
https://sonnati.files.wordpress.com/2016/06/10bit2.png
https://sonnati.wordpress.com/2016/06/17/does-vp9-deserve-attention-part-ii/
Even VP9/HEVC 10-12bit are already so advanced to produce block-free video even at quite low bitrates.
I wouldn't be surprised if AV1 8 bits would look worse than VP9/HEVC 10-12 bits (or just comparable)
HEVC eliminates most of the 10-bit advantage over 8-bit that H.264 had. If the source doesn’t have banding, you don’t get much new banding even at lower bitrates. I think AV1 should have ballpark similar improvements.
But yeah, it would be great if the “Main” profile for future codecs always supported at least 10-bit. That’s required for HDR, which is quickly becoming mainstream. It’s not like the SoC or GOU vendors are developing 8-bit only decoders anymore, even if some display pipelines are 8-bit RGB. But 10-bit 64-960 4:2:0 Y’CbCr makes for better 0-255 RGB 4:4:4 anyway.
benwaggoner
26th September 2018, 17:01
Well, they were using the player in the TV (at least, it looks like the standard LG player UI) - so my guess is a special engineering prototype TV with an FPGA or something in it :devil:
If it’s a fast 8-core ARM or something, software decode should be feasible, especially if the GPU was leveraged for some operations. Decoder optimization is a LOT easier and faster than encoder optimization, since there is only one right answer for decoding. We’ll have a good idea of the potential performance of AV1 decoders way before we will for encoders.
blurred
28th September 2018, 09:52
Interesting discussion: https://www.reddit.com/r/linux/comments/9i4h7y/av1_video_samples_now_available_on_youtube_netflix/
It points some FSF licensing issues from 2010:
- https://www.fsf.org/blogs/community/google-free-on2-vp8-for-youtube :"Until we move to free formats, the threat of patent lawsuits and licensing fees hangs over every software developer, video creator, hardware maker, web site and corporation -- including you."
- https://www.fsf.org/blogs/licensing/googles-updated-webm-license : "Unfortunately, the interaction between the copyright license and the patent license made the result GPL-incompatible. Based on the concerns of developers writing GPL-covered software, Google publicly stated that they would take some time to review the WebM license and try to address the community's concerns. Today, they released a revised license, and it is GPL-compatible."
"The most important part of the change is that Google has separated the patent license from the copyright license. Now the copyright license on the software is a totally standard three-clause BSD license, which is clearly compatible with the GPL. The patent license, in turn, provides distributors with permission to exercise all the rights, and meet all the conditions, in the GPL, as required by GPLv2 section 7; and those permissions are consistent with the ones provided by the patent grant in GPLv3 section 11. All this means that developers distributing GPL-covered software can take advantage of the patent license without running afoul of the GPL's conditions, whether they're using GPLv2 or GPLv3."
Does it still apply to AV1?
SmilingWolf
29th September 2018, 06:43
MSYS2/GCC 8.2 builds:
av1-1.0.0-692-g16c9affcb: https://mega.nz/#!U1AVFI5R!jPhrtngqfp0uNnVUo-GLfZMDT-h8Yl8p05ACp9Fn4Ok
dav1d-0.0.1-103-deab253: https://mega.nz/#!xpAHSYwB!hA4Unc81jOTbhxgKrCCRTMd0qKqaX3oODqiV3cs3ZP8
dav1d builds are also available here: https://code.videolan.org/videolan/dav1d/pipelines
Cloud button on the right
The first ASM routines for dav1d have been checked in! Time to benchmark again :)
TEB
29th September 2018, 19:11
I heard that Ateme demonstrated 4k60 AV1 decoding on TV (seems like LG TV) during the recent IBC Show.
Anybody attended the show and saw it?
Yupp saw it, but infact it was a HEVC encoded output from a AV1 source (high bitrate) as far as i know..
Mr_Khyron
30th September 2018, 00:32
https://www.twoorioles.com/
Offering unprecedented compression gains, high video quality and broad industry support, AV1 promises to be the next universal video compression standard. With 20 to 30 percent better quality than VP9 and HEVC, AV1 has buy-in from such technology leaders as Google, Apple, Microsoft, Netflix, Amazon, Intel, Cisco, Facebook and Mozilla. We’re helping to lead the charge with the industry's first commercial encoder for AV1. As part of our EVE (Efficient Video Encoder) family, EVE for AV1 is in beta testing with several industry partners. With our 16 years of video compression innovation, we’re making the promise of AV1 a reality.
https://www.twoorioles.com/eve-for-av1/
Mr_Khyron
30th September 2018, 19:28
Alibaba Cloud, the cloud computing arm of Alibaba Group, has joined the Alliance for Open Media (AOMedia) as a Promoter member. Alibaba Cloud will collaborate with AOMedia and other industry leaders in pursuit of an open and royalty-free AOMedia video codec, AV1.
https://aomedia.org/alibaba-cloud-joins-the-alliance-for-open-media/
benwaggoner
1st October 2018, 00:01
https://www.twoorioles.com/
https://www.twoorioles.com/eve-for-av1/
I really wish PR and such would stop throwing around this "20-30% better than HEVC and VP9" stuff. There is no way to know if that's achievable or not yet, without mature AV1 encoders. And we know that the best HEVC encoders can outperform the best VP9 encoders by easily 20-30% for real-world scenarios, so a quote assuming VP9 and HEVC are equivalent just invalidates the AV1 comparison.
Honestly, finding real-world scenarios where libvpx can deliver better quality @ bitrate @ perf than x264 is a challenge, although I suspect that's primarily due to x264's massively greater psychovisual tuning and perf optimization efforts, not differences in the bitstream standards.
NikosD
1st October 2018, 06:56
I really wish PR and such would stop throwing around this "20-30% better than HEVC and VP9" stuff. There is no way to know if that's achievable or not yet, without mature AV1 encoders...
Honestly, finding real-world scenarios where libvpx can deliver better quality @ bitrate @ perf than x264 is a challenge, although I suspect that's primarily due to x264's massively greater psychovisual tuning and perf optimization efforts, not differences in the bitstream standards.
I think the same could be said for HEVC too compared to AVC.
I remember the days when x265 was claiming 50% more compression for the same quality compared to x264, but you have to dig too much to find such a case, if any.
olduser217
1st October 2018, 07:41
Yupp saw it, but infact it was a HEVC encoded output from a AV1 source (high bitrate) as far as i know..
Do you mean that the AV1 bitstream was actually transcoded to HEVC stream and then playback using the LG TV?
If this is the case, it seems like the purpose was to showcase the AV1 encoding/transcoding capability rather than decoding of AV1 on existing SOC hardware.
LigH
1st October 2018, 07:47
@olduser217: I understood that the same way...
marcomsousa
1st October 2018, 08:35
Do you mean that the AV1 bitstream was actually transcoded to HEVC stream and then playback using the LG TV?
If this is the case, it seems like the purpose was to showcase the AV1 encoding/transcoding capability rather than decoding of AV1 on existing SOC hardware.
@olduser217: I understood that the same way...
They are a software company. They have a software solution that permit encoding and transcoding to BIG companies.
So, a big company can (in 2019) store in AV1 format and transcoding to ANY format LIVE.
Or encoding from ANY format to AV1.
At this moment, they are changing their solution to be Public Cloud Server friendly. Changing to microservices architecture capable to scale any number of servers automatically.
PS: They don’t have anything magically, so encoding it’s still slow and expensive.
olduser217
1st October 2018, 09:38
They are a software company. They have a software solution that permit encoding and transcoding to BIG companies.
So, a big company can (in 2019) store in AV1 format and transcoding to ANY format LIVE.
Or encoding from ANY format to AV1.
At this moment, they are changing their solution to be Public Cloud Server friendly. Changing to microservices architecture capable to scale any number of servers automatically.
PS: They don’t have anything magically, soy encoding it’s still slow and expensive.
@marcomsousa @LigH
Thanks for the explanation.
So, for the exhibition, the AV1 bitstream (I heard during the demo, the AV1 source was from an USB thumbdrive which was plugged into the TV) was uploaded to their server and downloaded as HEVC bitstream, then playback from the TV which support HEVC decoding?
hajj_3
1st October 2018, 15:16
Video Dev Days 2018 videos:
https://www.youtube.com/watch?v=1bSsP0Wi46E - AV1: in the end, what got in?
https://www.youtube.com/watch?v=ytsRYKQc6kQ - rav1e: the best rust AV1 encoder
https://www.youtube.com/watch?v=UhIgBdrKyNM - Dav1d: a fast new AV1 decoder
marcomsousa
1st October 2018, 15:41
...
Thanks, just add a new dav1d video and labels
Google AOM
Chrome (Q4 2018)
WebRTC Integration (2019)
Android Q (Q3 2019)
SOC Hardware (2020)
Beelzebubu
1st October 2018, 16:17
And we know that the best HEVC encoders can outperform the best VP9 encoders by easily 20-30% for real-world scenarios, so a quote assuming VP9 and HEVC are equivalent just invalidates the AV1 comparison.
[..]
libvpx
I'm all for pointing out PR for what it is, but to equate "libvpx" with "best VP9 encoders" is inherently unfair as an argument against an actually-good VP9 encoder.
LigH
1st October 2018, 17:44
The media-autobuild_suite already allows using libdav1d in ffmpeg; unfortunately, from time to time, other libraries may break it, so consider to exclude what you don't really need...
benwaggoner
1st October 2018, 19:24
I think the same could be said for HEVC too compared to AVC.
I remember the days when x265 was claiming 50% more compression for the same quality compared to x264, but you have to dig too much to find such a case, if any.
I have seen 50% reduction at very low bitrates and UHD resolutions. But for moderately grainy SD/HD content, x264 versus x265 is more like 30%.
Grain/noise parameterization and reconstruction is probably the single biggest next step in encoding performance, since random noise is intrinsically uncompressible.
benwaggoner
1st October 2018, 19:37
I'm all for pointing out PR for what it is, but to equate "libvpx" with "best VP9 encoders" is inherently unfair as an argument against an actually-good VP9 encoder.
I wasn't specifically calling out libvpx here, but I've not been able to get a real-world comparison of anything substantially better in terms of potential quality (not even quality @ perf).
I'm quite curious about what Eve can really do! But I've struggled to find samples for which source is available so I can try a real apples-to-apples. Their website seems to just have frames, not even any actual video examples.
If anyone can point me towards any, I'd really appreciate it, and could probably make some comparative HEVC encodes available.
For example, here's a recent Tears of Steel x265 test I did (1.5 Mbps average 4 Mbps peak, max 5 sec GOP, unlimited encoding time)
https://1drv.ms/v/s!AlvIQZWsyeO-kKplp2EQ8-Q4bCNVZw
It is pretty awesome that we can deliver a pretty good 1080p experience at VideoCD bitrates! 20x more pixels, and I think better per-pixel quality.
I've been doing 1, 1.5, and 2 Mbps ToS encodes using x265, x264, xvid, WMV VC-1, and VC-1 adaptive resolution Smooth Streaming. Hoping to kick off a VP9 set today, but it's been surprisingly hard to get good detailed parameter tuning documentation for libvpx.
I really hope libaom will get something as good as x265.readthedocs.io! That's really the gold standard of encoder documentation to date.
TD-Linux
1st October 2018, 22:43
"The most important part of the change is that Google has separated the patent license from the copyright license. Now the copyright license on the software is a totally standard three-clause BSD license, which is clearly compatible with the GPL. The patent license, in turn, provides distributors with permission to exercise all the rights, and meet all the conditions, in the GPL, as required by GPLv2 section 7; and those permissions are consistent with the ones provided by the patent grant in GPLv3 section 11. All this means that developers distributing GPL-covered software can take advantage of the patent license without running afoul of the GPL's conditions, whether they're using GPLv2 or GPLv3."
Does it still apply to AV1?
Yes, it does.
Blue_MiSfit
2nd October 2018, 00:35
All the improvements in dav1d is fabulous. Up to 50fps for 8 bit 1080p with 4 threads AND NO ASM YET is pretty amazing.
Mr_Khyron
2nd October 2018, 00:37
http://www.jbkempf.com/blog/post/2018/Introducing-dav1d
LigH
2nd October 2018, 07:14
And wiiaboo + schmidthubert made ffmpeg build again. So build your favourite kind of ffmpeg with dav1d. — Sorry, too early, API may have changed, some exports are not found suddenly.
@benwaggoner: PM?
Adonisds
2nd October 2018, 15:14
I just did some performance testing with the 1080p 30fps AV1 encode of the Gus Kenworthy & Tom Wallisch X Games Slopestyle GoPro Preview (https://www.youtube.com/watch?v=_fAOe8oz8qM) video in MPC-HC v1.8.1 x64 with its built-in LAVfilters; I originally tried the Halo video but I found the X Games video to be much more demanding (but also much more motion-sick inducing, especially when playing at slower than real-time).
With my 4c/8t Nehalem Xeon x3470 I was only seeing ~25% CPU utilization at maximum even though I was unable to play back the video in real-time (it was somewhere between 16fps and 20fps). Mathematically that should mean that it's only using 2 threads, but disabling SMT and setting my BIOS to only enable 2 cores resulted in noticably worse performance, yet setting the BIOS to enable 3 cores without SMT resulted in the same 16-20fps performance I was originally seeing yet at only ~67% CPU utilization.
At least with the LAVfilters bundled with MPC-HC v1.8.1 x64, it would seem that the AV1 decoder can only utilize 3 cores and no SMT, yet even then the 2 less loaded cores are only hitting around half of their according core's available utilization.
And for reference, the 720p 30fps AV1 encode of that same video played back without a hitch on my Xeon - heck it left enough headroom that I could turn on a bunch of motion interpolation which greatly helped alleviate the motion sickness I got from watching the 1080p AV1 encode playback at sub-20fps frame rates (it's times like this that I thank the devs over at Nintendo for making F-Zero X and F-Zero GX native 60fps games).
Now I'm a bit out-of-the-loop, but I couldn't help but notice that YouTube-DL was using the .MP4 extension for AV1 downloads - is that in fact correct behavior? (and no, I don't mean AVC1, otherwise my PC would have been playing back the videos easy-peasy).
How do you turn on motion interpolation?
Mr_Khyron
2nd October 2018, 18:44
https://itpeernetwork.intel.com/open-source-visual-cloud/
Next Gen CODECs
Streaming media is the foundation for visual cloud workloads. Our open source projects will also advance the next generation of CODECs, specifically Scalable Video Technology (SVT) for HEVC and AV1.
Additionally, Intel is contributing to the SVT-HEVC Encoder core to enable high performance, quality, and scalability of HEVC video encoding under a highly permissive BSD and patent license. This HEVC-compliant encoder library core achieves excellent density-quality tradeoffs and is highly optimized for Intel® Xeon® Scalable and Intel® Xeon® D processors.
We are also forming a new SVT-AV1 open source encoder project as an enhancement to the Alliance for Open Media (AOM) to provide a cleaner, easier to use codebase. Community support is critical to open source innovation, and Intel welcomes contributions to the SVT-AV1 project. Register for email updates at 01.org.
As video consumption and generation continues to grow, so will the number of industries that must deliver high-bandwidth, low-latency video at scale. Open source software is fundamental to meeting the demands, and Intel is committed to collaborating with industry leaders to grow the community.
NikosD
2nd October 2018, 23:07
https://www.youtube.com/watch?v=UhIgBdrKyNM - Dav1d: a fast new AV1 decoderThis is a very fast decoder.
Using just C with almost no ASM and no SIMD at all, they managed to be 60% faster than libaom v1.0.0.0 and capable of 1080p decoding of a 4Mbps clip at 50fps using 4 cores.
In the presentation they said that they are expecting the SIMD version to hit a 4x decoding acceleration, meaning 200 fps for 1080p using 4 cores.
That would be probably the biggest step up in software decoding SIMD acceleration.
NikosD
3rd October 2018, 08:16
If someone could build and provide latest libaom and dav1d decoders in DirectShow filter form, I'm interested in opening a new thread for AV1 decoder's evaluation.
I have various CPUs from Core 2 Duo to Haswell, so I can follow the improvements in ASM, MultiThreading and SIMD acceleration for both AV1 decoders.
BTW, are there any other AV1 decoders ?
But I need them to be registered as DirectShow filters.
foxyshadis
3rd October 2018, 08:23
I have seen 50% reduction at very low bitrates and UHD resolutions. But for moderately grainy SD/HD content, x264 versus x265 is more like 30%.
Grain/noise parameterization and reconstruction is probably the single biggest next step in encoding performance, since random noise is intrinsically uncompressible.
I have a feeling that video standards are eventually going to have a similar come-to-jesus moment that graphic card makers faced when they were finally called out for totally ignoring minimum FPS. That led to some pretty big driver changes within a year, but video standards have a much longer lead time....
NikosD
3rd October 2018, 08:54
...I did however notice that, out of the three cores being utilized on my Xeon x3470, one of them is pretty much fully pegged and the other two are about half utilization - this would then equal the ~25% utilization I'm seeing.
Now normally I would let this all go as simply a case of "not having fast enough single-threaded performance" and be done with it, but the fact that you and others were seeing over 60% utilization on your own 4c/8t CPUs (which implies it was balancing the load across more than 4 threads) makes me thing something still isn't quite right here - I mean, weaker single-threaded performance shouldn't result in fewer threads being utilized, and if anything you'd want it to be the opposite, no?
My preliminary results for LAV x64 0.72.0-15 and MPC-HC v1.8.2 using libaom v1.0.0-552 show me only ~75% CPU utilization of a 2C/4T Haswell Core i3-4170@3.7GHz, meaning that the current status of libaom decoder is an up to 3 threads SMT implementation.
The above situation leads to a ceiling in performance of course, meaning that a lot of 1080p AV1 samples out there, are simply not real-time decoded by libaom.
Probably even the no SIMD version of dav1d as it is right now, due to the 4 threads implementation IIRC, could do it better.
But generally speaking, "we need a bigger boat"
clsid
3rd October 2018, 15:19
The number of threads for libaom depends on the number of tiles in the video.
Newer libaom build is already 10-15% faster. You should wait with tests until dav1d gets into FFmpeg/LAV. Then libaom as decoder will become obsolete.
NikosD
3rd October 2018, 20:03
The number of threads for libaom depends on the number of tiles in the video.
The number of tiles is a property of the video created by the encoder or is it something that is been created by the decoder ?
LigH
3rd October 2018, 20:08
The encoder decides that, usually based on the resolution of the video, especially the height, if it was not specified expliticly as parameter.
NikosD
3rd October 2018, 21:21
By watching the presentation of dav1d, he gave me the impression that a tile-based decoder decides the number, the scheme, the orientation etc of the tiles, but without saying too many details in the presentation.
Wrong impression obviously.
Beelzebubu
3rd October 2018, 22:37
Sorry if that created confusion; tiles (and thus the amount of tile threads a decoder can use) are set by encoder; frame threads is selectable by decoder without requiring special encoder settings.
user1085
3rd October 2018, 22:48
Sorry if that created confusion; tiles (and thus the amount of tile threads a decoder can use) are set by encoder; frame threads is selectable by decoder without requiring special encoder settings.For comparison with other decoders for H.265, H.264, do we know how many threads they use?
Beelzebubu
3rd October 2018, 23:03
Depends on the implementation. FFmpeg's decoders for H264/HEVC typically use only frame threads in the default configuration, the exact number depends on the number of cores on your system. OpenHEVC allows you to combine frame and wave-front threading (similar to x265). Frame+Tile - like Frame+WFP - allows better scaling at ultra-high resolutions or very low-end systems, but is not typically necessary for real-time playback on normal systems.
olduser217
4th October 2018, 03:25
Sorry if that created confusion; tiles (and thus the amount of tile threads a decoder can use) are set by encoder; frame threads is selectable by decoder without requiring special encoder settings.
For AV1, decoder frame threading is related to encoder (bitstream encoded) too.
CDF tables can be updated (or not updated) at the end of decoding for a frame, depending on a frame header syntax element in the bitstream.
If the CDF tables need to be updated at the end of the decoding for a frame, then frame threading should be impossible.
jonatans
4th October 2018, 05:08
- https://www.fsf.org/blogs/licensing/googles-updated-webm-license : "Unfortunately, the interaction between the copyright license and the patent license made the result GPL-incompatible. Based on the concerns of developers writing GPL-covered software, Google publicly stated that they would take some time to review the WebM license and try to address the community's concerns. Today, they released a revised license, and it is GPL-compatible."
"The most important part of the change is that Google has separated the patent license from the copyright license. Now the copyright license on the software is a totally standard three-clause BSD license, which is clearly compatible with the GPL. The patent license, in turn, provides distributors with permission to exercise all the rights, and meet all the conditions, in the GPL, as required by GPLv2 section 7; and those permissions are consistent with the ones provided by the patent grant in GPLv3 section 11. All this means that developers distributing GPL-covered software can take advantage of the patent license without running afoul of the GPL's conditions, whether they're using GPLv2 or GPLv3."
Does it still apply to AV1?
Yes, it does.
Well.. the problem still applies to AV1. But the solution does not ;)
In WebM, the software license and patent license are clearly separated https://www.webmproject.org/license/ "The WebM codec source code and specification are licensed differently."
In AOM, the software license and patent license are clearly combined https://aomedia.org/license/ "Software released by the Alliance for Open Media is made available under a combination of the following licenses"
Nintendo Maniac 64
4th October 2018, 06:06
How do you turn on motion interpolation?
I'm using third party software for that - it is not part of MPC-HC, LAVfilters, nor AV1.
I was just making a point about how much CPU headroom I still had is all.
NikosD
4th October 2018, 07:21
Sorry if that created confusion; tiles (and thus the amount of tile threads a decoder can use) are set by encoder; frame threads is selectable by decoder without requiring special encoder settings.
For AV1, decoder frame threading is related to encoder (bitstream encoded) tooThank you for your comments.
Then it must be a great coincidence that all three AV1 clips I tried to decode had 75% CPU utilization on a four threaded CPU.
Is it some kind of default settings of the encoder to produce such clips ?
Can you post a sample that libaom decoder can use 4 threads or more during decoding ?
Thank you.
Beelzebubu
4th October 2018, 11:24
If the CDF tables need to be updated at the end of the decoding for a frame, then frame threading should be impossible.
The VDD presentation explains how we still accomplish frame threading in this scenario. Efficient frame threading is possible, even with CDF table dependencies. We did this in ffvp9 also, there is nothing new about this approach.
olduser217
4th October 2018, 12:30
The VDD presentation explains how we still accomplish frame threading in this scenario. Efficient frame threading is possible, even with CDF table dependencies. We did this in ffvp9 also, there is nothing new about this approach.
Thanks for the information.
Do you mind to point out the link to the presentation if it is available?
Beelzebubu
4th October 2018, 12:53
Thanks for the information.
Do you mind to point out the link to the presentation if it is available?
https://www.youtube.com/watch?v=UhIgBdrKyNM
Adonisds
4th October 2018, 15:39
I'm using third party software for that - it is not part of MPC-HC, LAVfilters, nor AV1.
I was just making a point about how much CPU headroom I still had is all.
I understood your point. But I'd like to start using motion interpolation for all my videos. What software do you use?
Adonisds
4th October 2018, 15:41
Would a modern 8 core/16 threads processor be able to software decode av1 4k60 hdr video once the decoder is more optimized?
Blue_MiSfit
4th October 2018, 17:25
I'd better hope so! If not, AV1 is going to have a hard time.
mzso
4th October 2018, 19:20
Honestly, finding real-world scenarios where libvpx can deliver better quality @ bitrate @ perf than x264 is a challenge, although I suspect that's primarily due to x264's massively greater psychovisual tuning and perf optimization efforts, not differences in the bitstream standards.
When was the last time a now format's encoder could beat the predecessor in bitrate and performance at the same time? Those days are gone it seems to me.
alex1399
4th October 2018, 20:44
The HEVC have tried very hard to not let its decode complexity overwhelm the decode complexity of AVC too much. In modern days, AV1 could try something better with the cost of decode complexity.
Nintendo Maniac 64
4th October 2018, 23:14
What software do you use?
SVP (http://svp-team.com/).
Note that there are several variants of it:
the old SVP 3.1.7 is free but only works on Windows and with programs like MPC-HC; also it might only work with 32bit media players.
the full-featured version of SVP 4 is only free on Linux and works with VLC and mpv.
the basic yet free version of SVP 4 which only works on Windows and with programs like MPC-HC much like v3.1.7, but has considerably reduced configuration options compared to both the full-featured version and the old 3.1.7 version.
the paid full-featured "Pro" version on Windows and Mac is largely identical to the free version of SVP 4 on Linux, though on Windows it also works with MPC-HC as well as VLC and/or mpv.
One thing to keep in mind is that interpolating to refresh rates that are exact multiples of the source framerate will provide a smoother result with fewer artifacts - e.g. 30fps interpolated to 120Hz (4x) is better than 30fps interpolated to 144Hz (4.8x) - this is most easily accomplished with something like like MPC-HC's or madVR's built-in automatic resolution changer which can be used to change your refresh rate depending on a given video frame rate (though you may need to create a custom resolution in order to access certain refresh rates on your display).
Blue_MiSfit
5th October 2018, 05:13
Yikes, no thanks. Looks like the motion interpolation on every TV these days :(
Nintendo Maniac 64
5th October 2018, 08:26
Yikes, no thanks. Looks like the motion interpolation on every TV these days :(
...what were you expecting? That's simply what motion interpolation is.
If you don't like high framerate video, then no amount of motion interpolation is going to look good - that's the goal of motion interpolation after all.
If you simply don't like interpolation artifacts, then certainly a piece of software designed to run at fullspeed with cranked settings on even a 10-year old quad core CPU isn't going to deliver a cleaner image than dedication interpolation hardware in modern TVs.
It is worth mentioning however that the Linux version, Pro versions, and the old v3.1.7 do have "2m (min artifacts)" and "1.5m (less artifacts)" settings as well as a "32 px. Large 0" setting that are particularly ideal for people that dislike high motion interpolation but still want improved motion resolution.
Additionally, it can make a big difference how your source content was originally recorded - my main use is for motorsports that are actually natively recorded and broadcast at 50fps but are then downsampled to 25fps for internet streaming (I'm looking at you Formula E). These sorts of 50fps --to-> 25fps content retains the faster camera shutter speed of the original 50fps recording which means that applying motion interpolation will make it look much more akin to native HFR content than what you'd get if you interpolated 24fps movie content recorded with a slow camera shutter speed (as is typically used in cinema).
And in my opinion, applying motion interpolation with cranked settings for native 50fps motorsports on a 100Hz CRT is glorious - the sense of speed you get from the cars that way is just unmatched!
________EDIT________
I just don't care for motion interpolation at all.
But you didn't say what aspect it is that you don't like, so I felt it necessary to try to answer every angle I knew of.
Now I know a lot of people don't like the artifacts, but that would practically require something like a Threadripper 2990WX + 64GB RAM + a GPU with Nvidia's A.I. Tensor cores in order to have truly artifactless interpolation that looks like real native HFR.
Keep in mind however that the higher the native frame rate of the video, the less artifacts there are: 50fps --to-> 100Hz has quite a bit fewer artifacts than 25fps --to-> 100Hz, and since 50fps is natively HFR anyway the overall "feeling" isn't exactly Earth-shatteringly different either when using interpolation.
Also interpolation in general can be a god-send for low framerate content like 15fps (which is what my father's smartphone camera uses for videos recorded in low-light situations) that can otherwise look really choppy (such as a recently recorded fireworks video he took).
Let's get back to AV1 :)
Hey now, you were the one that decided to not ignore the subject and "let my post be". :p It was even at the end of a page which would have allowed it to have easily been ignored.
Blue_MiSfit
5th October 2018, 08:28
So we're getting OT here, but if you like this stuff - great!
I just don't care for motion interpolation at all.
Let's get back to AV1 :)
Nintendo Maniac 64
7th October 2018, 03:02
It seems that LAVfilters v0.73 / MPC-HC v1.8.3 improves AV1 performance a tad (~20% faster), but it's multi-threaded utilization is still poor (only ~26% utilization with 4c/8t Nehalem).
The biggest benefit I found is that only using dual core no longer completely tanks performance.
Utilizing a Nehalem Xeon x3470, I had the following performance when trying to decode the 1080p AV1 video-only stream from Gus Kenworthy & Tom Wallisch X Games Slopestyle GoPro Preview (https://www.youtube.com/watch?v=_fAOe8oz8qM) in MPC-HC v1.8.3 64bit:
4c/8t @ 2.93GHz: ~20fps
4c/8t @ 3.20GHz: ~22fps
3c/3t @ 3.20GHz: ~22fps
2c/4t @ 3.46GHz: ~20fps
2c/2t @ 3.46GHz: ~20fps
Anything above 3c/3t still sees no benefit, and SMT still sees no utilization (though Nehalem's implementation of SMT is certainly going to be weaker than more modern implementations e.g. Ryzen).
benwaggoner
7th October 2018, 20:36
Would a modern 8 core/16 threads processor be able to software decode av1 4k60 hdr video once the decoder is more optimized?
Can a modern system do that for HEVC? HEVC decode is going to be more inherently parallelizable due to WPP. And I'm not aware of any software decoders that can do a realtime 2160p60 HEVC on any hardware I've looked at.
That kind of pixel fill rate is normally the domain of hardware decoders, or at least hybrid CPU-GPU implementations (like the Xbox 360 H.264 decoder).
For commercial content, the DRM hardware requirements to play UHD on a PC has always come on systems that have a 2160p60 HW decoder anyway. Since there is already a large installed base with has the DRM support bu not HW AV1, I would expect HEVC to remain the dominant codec for delivering premium UHD content for years to come.
But I believe Profile @ Level for AV1 requires some degree of tiling for UHD resolutions. Even if not required, I imagine some de facto guidelines about tiling to improve decode perf would become standard.
Parallelizing decode of Golden/IDR, I, P, B, and b frames can also be a useful technique if plenty of memory is available. You'd just do a lookahead to encode and buffer the referenced frames in tier order. That was nigh impossible with VP9, but I think is feasible for AV1.
LigH
7th October 2018, 22:08
New uploads: (MSYS2; MinGW32: GCC 7.3.0 / MinGW64: GCC 8.2.0)
AOM v1.0.0-735-g9b21428c8 (https://www.mediafire.com/file/y9981oqyft1987f/aom_v1.0.0-735-g9b21428c8.7z)
rav1e 0.1.0 (4d185f7 / 2018-10-04) (https://www.mediafire.com/file/r05hl811nnlm15n/rav1e_0.1.0_2018-10-04_4d185f7.7z)
dav1d 0.0.1 (c6788ed / 2018-10-04) (https://www.mediafire.com/file/6j3ip3u28u9n1ab/dav1d_0.0.1_2018-10-04_c6788ed.7z)
Nintendo Maniac 64
8th October 2018, 11:31
Can a modern system do that for HEVC? HEVC decode is going to be more inherently parallelizable due to WPP. And I'm not aware of any software decoders that can do a realtime 2160p60 HEVC on any hardware I've looked at.
I just did some tests with my trusty 4c/8t Nehalem Xeon x3470 @ 2.93GHz, and I was able to play back a particularly intensive 40fps (not a typo) 3840x2160 20Mbps HEVC video clip without issue, though it seems just barely as even 41fps resulted in stuttering.
So at least for 60fps 4k HEVC at the same 20Mbps bitrate, you would only need at most 50% more CPU horsepower, though if it was 60fps 30Mbps then you might need 100% more CPU horsepower.
...however, even 100% more performance should be relatively easy to achieve when you consider the following:
Modern CPU architectures (Haswell and Zen1) are ~50% faster clock-for-clock than Nehalem (and Sky/Kaby/Coffee Lake has even faster IPC as will Zen2)
Having 50% more CPU threads (6c/12t) is quite common nowadays and can now even be had in high-end laptops (not to mention that 8c/16t is also a thing)
50% higher clockspeed (4.4GHz) is right around the max turbo speed for modern higher-end desktop CPUs on both Intel and AMD (and ~5GHz overclocks are not at all uncommon anymore)
Combine all three aspects and you should not only be able to play 4k HEVC @ 60fps without issue, but you could probably even do it at 90fps if not 120fps.
nevcairiel
8th October 2018, 11:50
Consumer bitrate HEVC content can definitely software decode through ffmpeg/avcodec on a high-end'ish machine at 4K 10-bit 60 FPS (ie. like those 60 FPS UHD Blu-rays). WPP isn't really needed for that either, since its a bitstream feature thats not mandatory. Frame-based threading is fine.
benwaggoner
8th October 2018, 17:15
I just did some tests with my trusty 4c/8t Nehalem Xeon x3470 @ 2.93GHz, and I was able to play back a particularly intensive 40fps (not a typo) 3840x2160 20Mbps HEVC video clip without issue, though it seems just barely as even 41fps resulted in stuttering.
So at least for 60fps 4k HEVC at the same 20Mbps bitrate, you would only need at most 50% more CPU horsepower, though if it was 60fps 30Mbps then you might need 100% more CPU horsepower.
...however, even 100% more performance should be relatively easy to achieve when you consider the following:
Modern CPU architectures (Haswell and Zen1) are ~50% faster clock-for-clock than Nehalem (and Sky/Kaby/Coffee Lake has even faster IPC as will Zen2)
Having 50% more CPU threads (6c/12t) is quite common nowadays and can now even be had in high-end laptops (not to mention that 8c/16t is also a thing)
50% higher clockspeed (4.4GHz) is right around the max turbo speed for modern higher-end desktop CPUs on both Intel and AMD (and ~5GHz overclocks are not at all uncommon anymore)
Combine all three aspects and you should not only be able to play 4k HEVC @ 60fps without issue, but you could probably even do it at 90fps if not 120fps.
Good analysis. Thanks!
This suggests that common consumer system isn’t likely to be able to do SW 2160p60 decode for some time to come. It seems likely to me that the installed base of computers that can decode 2160p60 HEVC in HW is many times higher than that can do it in SW.
Also, 60p isn’t THAT common in professional content; it’s mainly seen with sports. Entertainment seems stubbornly stuck at 24p, which is a lot easier to decode.
Most PCs don’t have UHD displays, of course, and many of the devices that do have a pixel size so small that delivering UHD resolutions is kind of pointless (on a 15.4” 3840x2160 display, a well-encoded 1080p 24-60p isn’t going to be obviously degraded compared to a 2160p on the same screen).
But for cases where UHD playback, particularly at high frame rates, matters, HEVC is going to have a much bigger installed base beyond 2020.
dapperdan
8th October 2018, 20:28
Mozilla Research grants available to work on AV1 related projects:
https://mozilla-research.forms.fm/mozilla-research-grants-2018h2/forms/5348
Core Web Technologies: Mozilla has been deeply involved in creating and releasing AV1, an open, royalty-free video encoding format. We are looking for someone to port the rate control model from Theora/Daala to AV1, and to explore how best to use novel AV1 features like alt-refs and frame super-resolution to optimize the result.
MoSal
8th October 2018, 20:37
i7-7700K is capable of software-decoding HEVC 4K@60fps (10bit, >50mb/s) in real time using ffhevc. But a faster CPU (or a faster decoder) is probably required for smooth day-to-day usage.
The better optimized ffvp9 decoder is way way faster than ffhevc. I know the theoretical complexity of HEVC and VP9 decoders is not the same. But the gap here is really largely attributed to better optimizations.
Assuming dav1d will reach a sufficiently-optimized state in months, I think we can conservatively predict that above mid-range devices sold in 2019 should be capable of software-decoding AV1 4K@60 fps just fine.
Blue_MiSfit
8th October 2018, 20:57
My 6700k at work struggles to decode 4kp24 VP9 in software (in latest Chrome playing 4k YouTube content). Switch on hardware decoding, however, and my little GTX 960 does it without breaking out a sweat (windows reports approx 30% decode utilization).
Software playback of 4kp60 AV1 on 8 modern cores seems like a reasonable target to me, but it's kind of a moot point. Your average consumer watching 4k videos on YouTube probably doesn't have an 8 core CPU. I certainly don't, and I'm a hardcore tech / video nerd :)
I think YouTube will be using VP9 for 4kp60 VOD for quite some time, since there's lots of hardware decoders out there by now, and Google takes a very dogmatic position against HEVC
Nintendo Maniac 64
9th October 2018, 06:05
My 6700k at work struggles to decode 4kp24 VP9 in software
That's strange as I can play a 2160p30 VP9 YouTube video at 47fps (therefore has an effective bitrate of 28Mbps) in MPC-HC v1.8.3 x64 on my Xeon x3470 which is basically just a 1st gen i7, so surely a 6th gen i7 should have no issue with 60fps 4k VP9 let alone 24fps.
in latest Chrome
...oh, that might be the problem.
If you play the exact same VP9 video stream in MPC-HC then it should be a total cakewalk for your 6700k; I believe this is because Chrome uses the libvpx decoder while MPC-HC and LAVfilters uses ffvp9.
It's my impression that Firefox also utilizes ffvp9, so you may be able to alternatively use that instead of Chrome in order to get performance similar to what is seen with MPC-HC/LAVfilters.
Clare
9th October 2018, 14:36
If you want to patch ffmpeg to use libdav1d, I have this patch: https://gist.github.com/WyohKnott/09e84e4fc9f67fc1b8190a9855119ba9
To test with mpv, use mpv --vd=libdav1d
Some Rav1e data: https://wyohknott.github.io/video-formats-comparison/rav1e.html
Speed is there but the quality is way way below x264 for now.
NikosD
9th October 2018, 16:11
Would a modern 8 core/16 threads processor be able to software decode av1 4k60 hdr video once the decoder is more optimized?I'd better hope so! If not, AV1 is going to have a hard time.It seems that LAVfilters v0.73 / MPC-HC v1.8.3 improves AV1 performance a tad (~20% faster), but it's multi-threaded utilization is still poor (only ~26% utilization with 4c/8t Nehalem).
Anything above 3c/3t still sees no benefit, and SMT still sees no utilization (though Nehalem's implementation of SMT is certainly going to be weaker than more modern implementations e.g. Ryzen).Can a modern system do that for HEVC? HEVC decode is going to be more inherently parallelizable due to WPP. And I'm not aware of any software decoders that can do a realtime 2160p60 HEVC on any hardware I've looked at.Consumer bitrate HEVC content can definitely software decode through ffmpeg/avcodec on a high-end'ish machine at 4K 10-bit 60 FPS (ie. like those 60 FPS UHD Blu-rays). WPP isn't really needed for that either, since its a bitstream feature thats not mandatory. Frame-based threading is fine.i7-7700K is capable of software-decoding HEVC 4K@60fps (10bit, >50mb/s) in real time using ffhevc.If you want to patch ffmpeg to use libdav1d, I have this patch.
In your analysis following the initial question, you missed the single most important parameter regarding SW decoding speed which is not resolution (4K), frame rate speed (60fps) or quality (10bit).
It's bitrate/ bandwidth.
The UHD Blu-rays mentioned above, are 4K60fps HDR10 at 140Mbps HEVC video stream.
Please, find me a non HEDT CPU that could decode such stream in real-time.
I mean I doubt if Ryzen 2700X could do it or the Core i9 9900K that announced yesterday.
So, regarding decoding of 4K60fps AV1 at similar bitrates of 140Mbps is surely out of question.
Now, regarding AV1 decoding progress, latest LAV 0.73 using libaom is around 13% faster than previous libaom version of LAV 0.72.x which is really impressive and this is something I mentioned at LAV filters thread, the progress of libaom.
Still, it seems to me as a 3 threaded decoder.
Eventually, it looks like dav1d is not the only AV1 decoder progressing and the battle for the fastest AV1 decoder is still alive.
clsid
9th October 2018, 17:21
If you want to patch ffmpeg to use libdav1d, I have this patch: https://gist.github.com/WyohKnott/09e84e4fc9f67fc1b8190a9855119ba9Is this compatible with latest dav1d git head? I can successfully build dav1d and the ffmpeg libs (for LAV Filters) but it crashes at playback start.
dav1d.exe is working ok, but annoying that it requires outputting to a (huge) file.
Clare
9th October 2018, 17:53
Is this compatible with latest dav1d git head? I can successfully build dav1d and the ffmpeg libs (for LAV Filters) but it crashes at playback start.
dav1d.exe is working ok, but annoying that it requires outputting to a (huge) file.
Yeah you need a recent GIT, I used da97ba3f from Oct, the 4th.
Mr_Khyron
9th October 2018, 18:05
https://www.youtube.com/watch?v=rMgUk002JrU
Nintendo Maniac 64
9th October 2018, 21:11
It's bitrate/ bandwidth.
The UHD Blu-rays mentioned above, are 4K60fps HDR10 at 140Mbps HEVC video stream.
To my credit, I've started mentioning bitrate in my more recent posts talking about HEVC and VP9 decoding performance for this very reason. I also linked to the AV1-encoded YouTube video clip that I was testing with, so one could always find out the bitrate themselves if they absolutely needed to know what it was.
Nevertheless, I personally think that 4k AV1 encoded at anything higher than 50Mbps will be extremely unlikely to exist in the wild even at 120fps. The major reason for this is because it's my impression that current video streaming services already only use ~20Mbps for their 4k 24fps content with codecs that are less efficient than AV1, and higher framerate do not require a linear increase in bitrate (last I checked, YouTube used only 33% to 50% more bits for 60fps compared to 30fps, yet the 60fps encode always looked better).
Now while that only covers internet streaming and not the likes of disc media, I see it even more unlikely for AV1 to be adopted for such a thing due to a case of "not invented here" syndrome with the MPEG-LA and the various companies more interested in traditional broadcast media and the like. The one exception to this might be game consoles, but they require relatively modern and/or high-end GPUs anyway and therefore are extremely likely to have AV1 hardware decoding...and even then, as graphics become better and better, there becomes less and less of a need to utilize any sort of pre-recorded video (which takes up more storage capacity) rather than just rendering the scene in real-time.
mzso
9th October 2018, 21:15
In your analysis following the initial question, you missed the single most important parameter regarding SW decoding speed which is not resolution (4K), frame rate speed (60fps) or quality (10bit).
It's bitrate/ bandwidth.
The UHD Blu-rays mentioned above, are 4K60fps HDR10 at 140Mbps HEVC video stream.
Please, find me a non HEDT CPU that could decode such stream in real-time.
I mean I doubt if Ryzen 2700X could do it or the Core i9 9900K that announced yesterday.
So, regarding decoding of 4K60fps AV1 at similar bitrates of 140Mbps is surely out of question.
Why bother using a more efficient format if you're hellbent on wasting bandwidth anyway? Use 50Mbps max or stick with HEVC/AVC. (Blu-rays throw bandwidth at the problem instead of using efficient encoders to begin with.) Typical FullHD movies are (or can be if encoded efficiently) essentially perfect as far as human perception is concerned, so even with AVC you would only need at most 100mbps to match it at UHD@60p. And HEVC is supposed to be more efficient than AVC.
Mr_Khyron
9th October 2018, 22:56
There are some 4K AV1 samples to download here
https://www.elecard.com/videos :cool:
I can play them with MPC-HC 1.8.3 and LAV Filters 0.73.0-1
NikosD
9th October 2018, 23:19
I can play them with MPC-HC 1.8.3 and LAV Filters 0.73.0-1
And your CPU is ?
benwaggoner
10th October 2018, 05:08
There are some 4K AV1 samples to download here
https://www.elecard.com/videos :cool:
I can play them with MPC-HC 1.8.3 and LAV Filters 0.73.0-1
Including the UHD clips?
Also, does anyone know if the sources for those clips are available? This would be a great chance to do some apples-to-apples comparisons!
Mr_Khyron
10th October 2018, 08:22
And your CPU is ?
Including the UHD clips?
Also, does anyone know if the sources for those clips are available? This would be a great chance to do some apples-to-apples comparisons!
My Cpu is a Ryzen 1700 and when i play the UHD clips all cores are at %40 and i get between 8-16 fps
benwaggoner
10th October 2018, 08:26
My Cpu is a Ryzen 1700 and when i play the UHD clips all cores are at %40 and i get between 8-16 fps
Ah! Not at real-time then.
nevcairiel
10th October 2018, 08:42
For the record my i9-7900X@4.5G can play those Elecard UHD clips in realtime with libaom through LAV Filters (achieves ~33 fps on the 22mbit clip, the clip is 25 fps). Considering the state of the multi-threading in libaom, the 10 cores aren't really doing me any good, so it must be the relatively high clock I have it running at, and full AVX2.
Judging by these results, and the limited number of threads it uses now, with proper threading and more optimizations, I have no doubt that such clips would run fine on 2018/19 mainstream CPUs (ie. decently clocked quadcores with AVX2 and above), once the decoders are fully done. Of course the highest clip they had was 22mbit only, but considering AV1 is more likely to become a web format then a optical disc format, extremely high bitrates are probably going to remain very rare.
hydra3333
10th October 2018, 11:32
My Cpu is a Ryzen 1700 and when i play the UHD clips all cores are at %40 and i get between 8-16 fps
A dummy's question - may one enquire : I have mpc-hc, how to I tell the fps ? I see it stutters at the higher bitrates but am unsure how you measure "i get between xx-yy fps" ? (yes I have "View Statistics" ticked)
NikosD
10th October 2018, 17:15
https://ark.intel.com/products/186605/Intel-Core-i9-9900K-Processor-16M-Cache-up-to-5-00-GHz-
Hopefully someone can test the 9900K soon.
What did you mean ?
Intel has already paid a company for misleading benchmarks.
Crooks...They should pay a huge penalty for these dirty, low tricks.
https://www.techspot.com/article/1722-misleading-core-i9-9900k-benchmarks/
Nintendo Maniac 64
11th October 2018, 00:33
A dummy's question - may one enquire : I have mpc-hc, how to I tell the fps ? I see it stutters at the higher bitrates but am unsure how you measure "i get between xx-yy fps" ? (yes I have "View Statistics" ticked)
To be honest, I don't really know how to do so either, so I do it the manual way via trial and error - I import the videostream into mkvtoolnix, set the frame rate to something (like 20p), export to an mkv with the new framerate, and then see if can playback this new mkv smoothly without any stutter.
If it does stutter, then I do the exact same process again but with a lower framerate (like 15p).
If it plays without any stutter, then I still do the exact same process again but with a higher framerate.
I keep doing this trial-and-error process until I find the highest framerate that doesn't result in any stuttering (note that I only use integer fps values though, like 24p, 25p, 26p, etc - testing fractional framerates as well would be WAY too time consuming with this method).
This method does have the benefit however of working on any media player or browser or the like.
And yes, technically without using a variable refresh rate display or a bajillion different custom resolutions, you're going to always have some visual stutter due to many of the the tested framerates not being an exact multiple of your display's refresh rate, but the sort of stuttering caused by inadequate video decoder performance tends to be way worse than any sort of telecine judder or the like.
olduser217
11th October 2018, 02:34
https://code.fb.com/video-engineering/facebook-video-adds-av1-support/
Facebook is working on to add AV1 support too.
The browsers that supported AV1 is able to play the AV1 version of embedded video (but highest resolution only 360p).
benwaggoner
11th October 2018, 03:50
To be honest, I don't really know how to do so either, so I do it the manual way via trial and error - I import the videostream into mkvtoolnix, set the frame rate to something (like 20p), export to an mkv with the new framerate, and then see if can playback this new mkv smoothly without any stutter.
If it does stutter, then I do the exact same process again but with a lower framerate (like 15p).
If it plays without any stutter, then I still do the exact same process again but with a higher framerate.
That's a pretty good process given the tools available today.
I do note that it might somewhat underestimate SW decoder performance requirements for real-world content significantly.
Slowing down 60p to 30p will result in a stream that may be easier to decode than the same content natively captured at 30p. This is because twice as much motion happens between 30p frames, so there's more prediction and motion vectors per frame to process. Also, the bitrate will drop by half in a 60-30 conversion, when real-world a 30p might be 70-80% the bitrate of a 60p for the same spatial quality (since twice as much change per frame is being captured).
Of course, if real-world decoder characteristics are understood, rate control techniques like VBV can cap the worst-case decoding times, although at the potential risk of capping maximum quality for difficult segments. We got a little spoiled from the last decade-ish of relatively ubiquitous H.264 HW decoding :).
marcomsousa
11th October 2018, 09:45
ffmpeg -benchmark -i Stream2_AV1_4K_22.7mbps.webm -f null -
Video: wrapped_avframe, yuv420p, 3840x2160 [SAR 1:1 DAR 16:9], q=2-31, 200 kb/s, 25 fps, 25 tbn, 25 tbc (default)
frame= 3604 fps= 16 q=-0.0 Lsize=N/A time=00:02:24.16 bitrate=N/A speed=0.622x
bench: utime=1069.000s stime=12.891s rtime=231.970s
bench: maxrss=856880kB
Since the video was 25 fps, in benchmark give that my PC is only capable to decode at 15-16 fps (speed=0.622x) at with this 22.7mbps video.
CPU: Intel Core i7-8550U
Decoder: ffmpeg-20181007-0a41a8b-win64 - libaom-av1 1.0.0-691-gbb8157b89
MoSal
11th October 2018, 11:51
frame= 3604 fps= 26 q=-0.0 Lsize=N/A time=00:02:24.16 bitrate=N/A speed=1.02x
CPU: Intel Core i7-7700k (60-65% utilization).
Decoder: libaom (1.0.0.r749.g955242e6a6, -DCONFIG_LOWBITDEPTH=1).
hydra3333
11th October 2018, 12:26
Cough,
frame= 3604 fps= 14 q=-0.0 Lsize=N/A time=00:02:24.16 bitrate=N/A speed=0.577x
bench: utime=1170.531s stime=6.344s rtime=249.925s
bench: maxrss=852156kB
CPU: Intel i7-i3820
Decoder: ffmpeg version a day or two old, N-92147-gf85fa100db ; libaom 1.0.0-708-gdf7131064 commit df7131064bf37fb5c7ee427ba564c31a2ed8bbbe (the one before it conflicts with libvpx) without DCONFIG_LOWBITDEPTH
Commandline: "ffmpeg.exe" -benchmark -i Stream2_AV1_4K_22.7mbps.webm -f null -
marcomsousa
11th October 2018, 13:07
Youtube already support AV1 when upload videos.
Then I uploaded (av1 video) to YouTube, the results:
YouTube successfully recognized the video and re-encoded it to h.264 and VP9.
YouTube did not display the original in AV1 and dit also not re-encode to AV1.
Source (https://www.reddit.com/r/AV1/comments/9n0h7o/i_uploaded_an_4k_60fps_av1_video_to_youtube_this/)
Pushman
11th October 2018, 13:58
frame= 3604 fps= 11 q=-0.0 Lsize=N/A time=00:02:24.16 bitrate=N/A speed=0.435x
ffmpeg version N-92132-g0a41a8bf29 Copyright (c) 2000-2018 the FFmpeg developers
[libaom-av1 @ 000001d804fecb00] 1.0.0-691-gbb8157b89
CPU: Intel i3-4170
clsid
11th October 2018, 14:17
GraphStudioNext has a performance test feature which can be used to measure how many fps a DirectShow decoder can deliver.
Clare
11th October 2018, 15:24
Does aomenc support multithread yet? I have tile-columns=4 and row-mt=1 but it still only using 1 core.
marcomsousa
11th October 2018, 15:38
Does aomenc support multithread yet? I have tile-columns=4 and row-mt=1 but it still only using 1 core.
you forget --threads=8?
aomenc -v -w 1920 -h 1080 --cpu-used=0 --target-bitrate=1500 --threads=8 --profile=0 --aq-mode=0 --lag-in-frames=25 --auto-alt-ref=1 --tile-columns=4 --row-mt=1 -o test15.webm test1.y4m
This use all CPU.
Tune to you logical cores.
mandarinka
11th October 2018, 16:57
Can a modern system do that for HEVC? HEVC decode is going to be more inherently parallelizable due to WPP. And I'm not aware of any software decoders that can do a realtime 2160p60 HEVC on any hardware I've looked at.
I think it was recently mentioned here that FFmpeg doesn't do WPP simultaneously in addition to frame threading. I'm also aware of it not scaling very well, maybe this is the reason. (Doesn't OpenHEVC support doing this?)
But since Nevcariel says it works on some PCs, I guess it is a matter of CPU and RAM bandwidth. And perhaps single-thread per-core performance. Many slower cores might not cut it due to bw/scaling issues, but fewer ones on 4,0-4,5 GHz like those Kaby Lake/Coffee Lake chips could?
FFmpeg's HEVC decoder isn't yet/atm optimised thoroughly, there is some intrinsics optimizations from openhevc missing (LAV Video decoder has them though) and there is probably some other pickable fruit too. There just wasn't motivation on the side of devs it seems (preferences for the google/royalty-free formats etc).
Clare
11th October 2018, 16:57
you forget --threads=8?
aomenc -v -w 1920 -h 1080 --cpu-used=0 --target-bitrate=1500 --threads=8 --profile=0 --aq-mode=0 --lag-in-frames=25 --auto-alt-ref=1 --tile-columns=4 --row-mt=1 -o test15.webm test1.y4m
This use all CPU.
Tune to you logical cores.
Ooops thanks now it's working.
aomenc --threads=8 --cpu-used=4 --tile-columns=4 --row-mt=1 --passes=2 --pass=2 --bit-depth=10 --input-bit-depth=10 --end-usage=q --cq-level=28 --fpf=Chimera_DCI4k2398p_HDR_P3PQ.log -o Chimera_DCI4k2398p_HDR_P3PQ.ivf Chimera_DCI4k2398p_HDR_P3PQ.y4m
Edit: it bursted on all core for 30 seconds but went back to one core afterwards :(
easyfab
11th October 2018, 17:20
for --row-mt=1 I think you need to wait that https://aomedia-review.googlesource.com/c/aom/+/72801 is merged.
Clare
11th October 2018, 17:46
for --row-mt=1 I think you need to wait that https://aomedia-review.googlesource.com/c/aom/+/72801 is merged.
I already patched my build with this. It seems that as soon as the first frame is finished rendering, it drops back to one core.
Nintendo Maniac 64
11th October 2018, 20:51
Slowing down 60p to 30p will result in a stream that may be easier to decode than the same content natively captured at 30p. This is because twice as much motion happens between 30p frames, so there's more prediction and motion vectors per frame to process. Also, the bitrate will drop by half in a 60-30 conversion, when real-world a 30p might be 70-80% the bitrate of a 60p for the same spatial quality (since twice as much change per frame is being captured).
Yep indeed, this is why I specifically use YouTube's 30fps encodes when dealing with a frame rate that's less than 60fps as it's better to under-estimate performance and then be pleasantly surprised to find out that real performance is better.
benwaggoner
12th October 2018, 01:23
I think it was recently mentioned here that FFmpeg doesn't do WPP simultaneously in addition to frame threading
That is going to be way more dependent on the encoder used than ffmpeg itself. x265 uses WPP AND frame-threads by default if you have multiple cores, and I doubt ffmpeg would disable that. Turning off WPP silently would be a big problem, as WPP also has significant decoder impact as well, particularly with multithreaded software decode.
It's an ongoing challenge with all encoder to make sure that the right flags are allowed by products that incorporate them. Making sure that commonly used tools like ffmpeg integrate libaom (or a superior alternative) well is pretty darn important, as that's what lots of reviewers and evaluators will use.
Getting good, actionable documentation into ffmpeg and in general is also important. Listing options without explaining why one might want to use it and its pros/cons isn't really documentation. x265.readthedocs.io is the gold standard here, and I don't even have a runner up.
benwaggoner
12th October 2018, 01:42
Yep indeed, this is why I specifically use YouTube's 30fps encodes when dealing with a frame rate that's less than 60fps as it's better to under-estimate performance and then be pleasantly surprised to find out that real performance is better.
Yeah, your approach is the best I can think of until it becomes more feasible to personally encode test content that makes use of a realistic array of AV1 features.
Testing fast encodes risks skipping features, making decoding simpler for a SW decoder than real-world competitive quality AV1 encodes would be. Some examples from past codecs where faster encoder modes simplify impact decoder performance that can be turned off for encoder performance include:
Weighted prediction
In-loop deblocking or SAO
Number of reference frames
Number of B-frames
Kurosu
12th October 2018, 12:15
There just wasn't motivation on the side of devs it seems (preferences for the google/royalty-free formats etc).
I'd rather say that most capable devs had enough of (big) corps freeloading and/or are paid to do something else. In a sense, it is actually a good thing that they learnt such a lesson, as it indicates a maturing and more professional community. And if AoM is ready to play fair in that regard, then it's "not-AoM" loss.
As for HEVC, the mentioned WPP+frame threading could be available today. All in all, there's likely 30% speed-up left.
Source: said capable devs.
Monarc
12th October 2018, 16:30
ffmpeg -benchmark -i Stream2_AV1_4K_22.7mbps.webm -f null -
Video: wrapped_avframe, yuv420p, 3840x2160 [SAR 1:1 DAR 16:9], q=2-31, 200 kb/s, 25 fps, 25 tbn, 25 tbc (default)
frame= 3604 fps= 16 q=-0.0 Lsize=N/A time=00:02:24.16 bitrate=N/A speed=0.622x
bench: utime=1069.000s stime=12.891s rtime=231.970s
bench: maxrss=856880kB
Since the video was 25 fps, in benchmark give that my PC is only capable to decode at 15-16 fps (speed=0.622x) at with this 22.7mbps video.
CPU: Intel Core i7-8550U
Decoder: ffmpeg-20181007-0a41a8b-win64 - libaom-av1 1.0.0-691-gbb8157b89
on my Intel(R) Core(TM) i5-3550 CPU @ 3.30GHz
ffmpeg -benchmark -i Stream2_AV1_4K_22.7mbps.webm -f null -
ffmpeg version n4.0.2
[libaom-av1 @ 0x55e28fe6f180] 1.0.0-759-g90a15f4f28
frame= 3604 fps=6.2 q=-0.0 Lsize=N/A time=00:02:24.16 bitrate=N/A speed=0.246x
video:1886kB audio:0kB subtitle:0kB other streams:0kB global headers:0kB muxing overhead: unknown
bench: utime=579.607s
bench: maxrss=602748kB
ffmpeg -threads 4 -benchmark -i Stream2_AV1_4K_22.7mbps.webm -f null -
frame= 3604 fps= 16 q=-0.0 Lsize=N/A time=00:02:24.16 bitrate=N/A speed=0.643x
video:1886kB audio:0kB subtitle:0kB other streams:0kB global headers:0kB muxing overhead: unknown
bench: utime=656.087s
bench: maxrss=737080kB
benwaggoner
13th October 2018, 00:26
I'd rather say that most capable devs had enough of (big) corps freeloading and/or are paid to do something else. In a sense, it is actually a good thing that they learnt such a lesson, as it indicates a maturing and more professional community. And if AoM is ready to play fair in that regard, then it's "not-AoM" loss.
Also, multithreaded decode speed isn't THAT important a feature for ffmpeg, which is mainly for transcoding. Lots of operations that include decode are going to be encoder-bound. And the other operations likely have work to do on unused cores. Multithreaded decode uses more memory and sometimes CPU, so an export-bound operation may actually be a little faster with a single-threaded decoder.
You are right that lots of ffmpeg and other open-source encoder/tool development is corporate funded, and so a lot of its development is driven by what companies want enough to pay for. And HEVC source is still pretty uncommon outside of some high end contribution streams from live events.
As for HEVC, the mentioned WPP+frame threading could be available today. All in all, there's likely 30% speed-up left.
Source: said capable devs.
That's reasonable. Decoders hit the point of asymptotic speed increases a lot sooner than encoders as there is only one "right" answer. AV1 is going to have >>30% decoder perf headroom still. I am curious how the greater variety of tools will impact optimal decoder performance in AV1 versus HEVC. AV1 is MUCH better designed for both HW and multithreaded SW decoders than VP9 and earlier, which had limitations in the loop filter and frame signaling that made them almost single-threaded. Which was generally fine for a PC web browser of the era, but had some real limitations for battery operated devices or those with lots of slower cores.
We'll probably have a good sense of the fundamental real-world decoder perf differences between HEVC and AV1 in 2019.
Kurosu
13th October 2018, 06:30
Also, multithreaded decode speed isn't THAT important a feature for ffmpeg, which is mainly for transcoding
True, the decoding speed increase is no longer useful for a lot of companies. Security is often more important.
AV1 is MUCH better designed for both HW and multithreaded SW decoders than VP9 and earlier, which had limitations in the loop filter and frame signaling that made them almost single-threaded.
One of those bottlenecks in thread scaling, what I'd call entropic state sharing across frames, is still there in AV1. This is to me something that was decided irrespective of SW decoding. Another sign that indeed the interest in it is only transitory.
We'll probably have a good sense of the fundamental real-world decoder perf differences between HEVC and AV1 in 2019.
And in 2020, hopefully, new products will all have HW decoding of AV1, so that would then no longer matter.
Mjpeg
13th October 2018, 16:29
New article from Streaming Media:
http://www.streamingmedia.com/Articles/Editorial/Featured-Articles/HEVC-VP9-AV1-and-VVC-Presenting-a-Codec-Update-in-11-Charts-127956.aspx
On the second page there is an intriguing slide from Bitmovin that with a pure software encoder version of libaom, they attained a speedup for cpu-used=0 from 1000x of VP9 in the early days down to 40x today with improvements ongoing. The claim from AOM was always that there was lots of room for optimization work once the spec was settled and ... well maybe that's happening. The proof will be when those speedups are in mainline FFmpeg and show up in tests on this very thread!
The slide also quotes software decode of 3.4x of VP9 single-threaded.
utack
13th October 2018, 18:27
New article from Streaming Media:
http://www.streamingmedia.com/Articles/Editorial/Featured-Articles/HEVC-VP9-AV1-and-VVC-Presenting-a-Codec-Update-in-11-Charts-127956.aspx
Regarding the "poor results for AV1".
The BBC paper used single pass encode
The parameters “--passes=1” and “--lag-in-frames=0” were set to run AV1 in single pass mode
without the possibility of looking ahead in the video sequence before encoding
The Harmonic paper used a fixed GOP of less than a single second of video material
--min-gf-interval=16 --max-gf-interval=16
The Ateme paper does not mention any configuration for the encoder
That is basically three results that are only good for academic publication count and the gutter in the real world
On another note, there is a 5 minute 720p clip with super diverse content out there now on youtube where VP9 and AV1 version have basically the same file size:
https://www.youtube.com/watch?v=WaWnLiffxJ4
You can do some direct comparison of quality
marcomsousa
14th October 2018, 08:25
I already patched my build with this. It seems that as soon as the first frame is finished rendering, it drops back to one core.
Another fix was merged
Fix allocation of workers for enc row-mt
Workers allocated for row based multi-threading of encoder are
now evaluated as minimum of num of threads and total number of
sb rows in the frame to be encoded.
Change-Id: I07501f43514f1ee45dd6637fe56411432930396c
dapperdan
14th October 2018, 09:41
New article from Streaming Media:
http://www.streamingmedia.com/Articles/Editorial/Featured-Articles/HEVC-VP9-AV1-and-VVC-Presenting-a-Codec-Update-in-11-Charts-127956.aspx
On the second page there is an intriguing slide from Bitmovin that with a pure software encoder version of libaom, they attained a speedup for cpu-used=0 from 1000x of VP9 in the early days down to 40x today with improvements ongoing.
Might be worth noting that the slide is from a Google employee, the Bitmovin employee was just in the audience.
I'm intrigued to hear more about the AI strategies they are using and wonder how portable those are once the "trick" has been revealed by the AI.
Tommy Carrot
14th October 2018, 11:29
On the second page there is an intriguing slide from Bitmovin that with a pure software encoder version of libaom, they attained a speedup for cpu-used=0 from 1000x of VP9 in the early days down to 40x today with improvements ongoing.
The speed improvements are nice, but many of those optimizations came at the cost of the quality. I'd estimate the compression efficiency dropped about 2-3% since the earlier versions.
soresu
14th October 2018, 19:55
Personally, considering that BBC is directly involved in development of VVC, and given that they joined the AV1 party relatively late in the development cycle, I'm inclined to find their so called research analysis somewhat suspect.
The Streaming Media article mentions their credibility on the issue, coming from a member of AOMedia, without taking into account just how late they joined.
At the very least it seems a conflict of interest to post such a negative paper so early in the post standard development of AV1, even more so considering that VVC is 2 years away from even being standardised.
Does anyone here have any idea what level of patent/proposals BBC have in the VVC working group?
Edit: Sorry if this is on the wrong thread, it concerns both AV1 and VVC so I wasnt sure which to drop it in.
Clare
15th October 2018, 16:53
AV1 Image File Format (AVIF) https://people.xiph.org/~negge/AVIF2018.pdf
Mile-High Video Workshop videos http://mile-high.video/files/mhv2018/
with topics such as:
Video Encoding and HEVC
Into the Depths: The Technical Details behind AV1
AV1 vs. HEVC: Perceptual Evaluation of Video Encoders
Codec Comparison from TCO and Compression Efficiency Perspective
VVC – The Next-Generation Video Standard of the Joint Video Experts Team
Pushing Encoding Quality and Speed with x265
Massively Parallel Encoding
benwaggoner
15th October 2018, 19:32
The speed improvements are nice, but many of those optimizations came at the cost of the quality. I'd estimate the compression efficiency dropped about 2-3% since the earlier versions.
A 2.5x perf improvement with only a 2-3% efficiency tradeoff is a pretty good optimization option. That's ballpark to a x26? preset ratio in perf and quality.
And things have to get LOTS faster to do the psychovisual and rate control tuning to get AV1 into something that can be practically compared to other encoders/bitstreams. Quality won't matter within an order of magnitude of the current speeds. It's not like anyone is actually delivering in volume 1080p encoded with --preset placebo either!
Quality @ Bitrate @ Perf!
SmilingWolf
16th October 2018, 11:05
According to my tests on my F.Y.C clip (1280x720, 480 frames) aomenc right now has got roughly the same encoding quality @ bpp of 1 month ago, with a deviation within the 1%, but is also somewhere between 14-25% faster than 2 months ago for --cpu-used=4
Which means that clip now encodes, in single thread at CQ 20, in 66 minutes rather than 87 on my i7-4770.
aomenc 1.0.0-359-g1bc580401:
# aomenc --frame-parallel=0 --tile-columns=0 --auto-alt-ref=2 --cpu-used=4 --passes=2 --threads=1 --lag-in-frames=25 --end-usage=q --cq-level=20 -o test.av1.cq20.webm orig.i420.y4m
Pass 1/2 frame 480/481 92352B 1539b/f 36899b/s 21265 ms (22.57 fps)
Pass 2/2 frame 480/480 5995490B 99924b/f 2395780b/s 5245119 ms (0.09 fps)
aomenc 1.0.0-775-g7f492af02:
# aomenc --frame-parallel=0 --tile-columns=6 --auto-alt-ref=1 --cpu-used=4 --tune=psnr --passes=2 --threads=1 --end-usage=q --cq-level=20 -o test.av1.cq20.webm orig.i420.y4m
Pass 1/2 frame 480/481 92352B 1539b/f 36899b/s 17764 ms (27.02 fps)
Pass 2/2 frame 480/480 6092708B 101545b/f 2434645b/s 3969928 ms (0.12 fps)
mzso
16th October 2018, 12:11
AV1 Image File Format (AVIF) https://people.xiph.org/~negge/AVIF2018.pdf
Funnily enough there's barely any mention of AVIF. Like 8 pages.
benwaggoner
16th October 2018, 16:53
Funnily enough there's barely any mention of AVIF. Like 8 pages.
Well, everything about intra coding is relevant to AVIF.
It's interesting to see the use of VMAF for still images. I question its relevance, as VMAF was only calibrated on >=24p moving images. Has anyone published any data suggesting VMAF is useful for measuring still images?
Also, the image comparisons versus HEVC don't show the command line. HEVC doesn't have --preset stillimage like x264, but I've tried an adaption of those parameters in x265. I'd expect it would improve the quality of the HEVC (HEIF?) image significantly. I saw some of the new Adobe tools that came out the last few days also include an HEIF exporter, which would be interesting to compare. Looks like fixed QP is being used for both AV1 and HEVC, which is going to be suboptimal for both codecs. Particularly for anything that includes mixed natural/synthetic imagery.
It amused me that it ended with a frames-per-minute graph that doesn't reference frame size, nor whether this is for still only or for temporal encoding. The idea of a still image taking 38 seconds to encode is going to horrify anyone using an image server.
marcomsousa
16th October 2018, 17:08
AV1 Image File Format (AVIF) https://people.xiph.org/~negge/AVIF2018.pdf
Presentation explained
https://www.youtube.com/watch?v=On9VOnIBSEs
marcomsousa
16th October 2018, 19:55
Google released today Chrome 70
Chrome 70 adds an AV1 decoder (it's enabled by default) to Chrome Desktop x86-64 based on the official bitstream specification. At this time, support is limited to “Main” profile 0 (https://aomediacodec.github.io/av1-spec/#profiles) and does not include encoding capabilities. The supported container is MP4 (ISO-BMFF (https://aomediacodec.github.io/av1-isobmff)) (see From raw video to web ready for a brief explanation of containers (https://developers.google.com/web/fundamentals/media/manipulating/files#how_are_media_files_put_together)).
To try AV1:
Update your stable Chome just type in the url: chrome://settings/help and restart Chrome
Go to the YouTube TestTube page (https://www.youtube.com/testtube).
Select "Prefer AV1 for SD" or "Always Prefer AV1" to get the desired AV1 resolution. Note that at higher resolutions, AV1 is more likely to experience playback performance issues on some devices.
Try playing YouTube clips from the AV1 Beta Launch Playlist (https://www.youtube.com/playlist?list=PLyqf6gJt7KuHBmeVzZteZUlNUQAVLwrZS).
Confirm the codec av01 in "Stats for nerds".
dapperdan
17th October 2018, 08:19
And things have to get LOTS faster to do the psychovisual and rate control tuning to get AV1 into something that can be practically compared to other encoders/bitstreams. Quality won't matter within an order of magnitude of the current speeds. It's not like anyone is actually delivering in volume 1080p encoded with --preset placebo either!
Netflix already use VP9, which is lacking good rate control and psychovisual tuning, but still delivering higher quality (VMAF) at lower bitrates than their other streams (according to their blog posts) and they have said they'd roll out AV1 when it was within 4-10x slower than VP9. Presumably based on getting higher quality that makes that tradeoff worthwhile for them at that point.
Bitmovin claims their latest encoder release is twice as fast as stock AV1 and 20x slower than VP9 so things don't seem to be that far off AV1 being delivered on a large scale.
They also have the extra bonus of being able to encode their in-house stuff without film grain and add it later, a feature they were keen to have added to AV1.
mzso
17th October 2018, 11:44
Netflix already use VP9, which is lacking good rate control
Is that still true? My ipression these days from youtube are that the bitrates are quite stable. (Might be wrong)
They also have the extra bonus of being able to encode their in-house stuff without film grain and add it later, a feature they were keen to have added to AV1.
Stupidest feature ever. They should have implemented an optional "hellyeahIwantnoiseLOL" decoder feature instead.
LigH
17th October 2018, 12:26
Don't miss the detail that bitrate optimization for streaming mainly means ABR — and the smaller the assumed decoding buffer and allowed preload time, the closer it gets to CBR, which will cause quality fluctuations. Oh, the joy of the VBV model.
Adonisds
17th October 2018, 13:37
SVP (http://svp-team.com/).
Note that there are several variants of it:
the old SVP 3.1.7 is free but only works on Windows and with programs like MPC-HC; also it might only work with 32bit media players.
the full-featured version of SVP 4 is only free on Linux and works with VLC and mpv.
the basic yet free version of SVP 4 which only works on Windows and with programs like MPC-HC much like v3.1.7, but has considerably reduced configuration options compared to both the full-featured version and the old 3.1.7 version.
the paid full-featured "Pro" version on Windows and Mac is largely identical to the free version of SVP 4 on Linux, though on Windows it also works with MPC-HC as well as VLC and/or mpv.
One thing to keep in mind is that interpolating to refresh rates that are exact multiples of the source framerate will provide a smoother result with fewer artifacts - e.g. 30fps interpolated to 120Hz (4x) is better than 30fps interpolated to 144Hz (4.8x) - this is most easily accomplished with something like like MPC-HC's or madVR's built-in automatic resolution changer which can be used to change your refresh rate depending on a given video frame rate (though you may need to create a custom resolution in order to access certain refresh rates on your display).
Thanks
Phanton_13
17th October 2018, 17:32
Is that still true? My ipression these days from youtube are that the bitrates are quite stable. (Might be wrong) Yes, you have to take in consideration that even if the bitrate is quite stable you can can have a bad distribution of it. I found that in vp9 if you limit the maximun quantification between 24-38 the overall quality is boosted without much variation on final bitrate/size (max value to be used depends of bitrate and resolution), basically you prevent the encoding rate control to be fucked sometimes with its corresponding drop in quality.
benwaggoner
17th October 2018, 19:51
Netflix already use VP9, which is lacking good rate control and psychovisual tuning, but still delivering higher quality (VMAF) at lower bitrates than their other streams (according to their blog posts) and they have said they'd roll out AV1 when it was within 4-10x slower than VP9. Presumably based on getting higher quality that makes that tradeoff worthwhile for them at that point.
VMAF has a much better subjective correlation than SSIM, but it is pretty far from perfectly subjectively correlated with double-blind subjective measurements. The risk of any new metric is over-optimizing an encoder for metric scores instead of actual subjective experience. I saw plenty of places where it got things wrong in the previous version; I've not worked extensively with the latest update, though, which has doubtless improved.
Netflix is also going from H.264 to VP9/AV1, so they don't need to beat or even meet HEVC quality to get a worthwhile improvement, of course.
Netflix IS using HEVC for all UHD and HDR encoding AFAIK.
I'm not sure where VP9 versus H.264 quality stands today. I'm running an encoding challenge, and would love to have someone provide best-effort VP9 and even AV1 samples for comparison.
https://forum.doom9.org/showthread.php?t=175776
Bitmovin claims their latest encoder release is twice as fast as stock AV1 and 20x slower than VP9 so things don't seem to be that far off AV1 being delivered on a large scale.
That would suggest that stock AV1 is only 40x slower than VP9, which is not my understanding at all. Perhaps they are talking the highest speed mode of their encoder, which would involve some quality degradation. Quality @ Perf is the important thing.
Getting good multithreading into AV1 is going to be really important, since time to market matters for a lot of content. Being able to encode something 8x faster on 16 cores probably doesn't matter for Netflix or YouTube given their chunking and lack of day-after-broadcast content. But it's critical for other markets, and essential for live encoding. Some of the VP9 and AV1 comparisons with x264 and x265 were artificially limited to 1-2 cores. Which makes sense if optimizing for absolute volume of minutes encoded, but understates the speed advantages of x26? for latency-critical tasks.
They also have the extra bonus of being able to encode their in-house stuff without film grain and add it later, a feature they were keen to have added to AV1.
That is a pretty huge feature! It was optional for H.264 (only required in HD-DVD decoders, but I don't know of anything authored with it). Random noise is mathematically uncompressible, so this kind of noise synthesis is an extremely promising way to improve quality and compressibility of the most challenging content.
...and will also reveal how much of the apparent detail in film comes from the grain. Older films, particularly, look really soft without the grain, and often just don't have much spatial detail.
But don't underestimate the challenge of the removing grain part; parameterizing it and then reconstructing it on playback are the easy parts. It's way more feasible now than 12 years ago, but it isn't trivial of something that can run 100% automated without messing up sometimes.
Unfortunately production workflows put in film grain much earlier than the encoding stage, so it's already baked in way before it gets to an encoder.
Mr_Khyron
17th October 2018, 22:29
http://www.socionext.com/en/pr/sn_pr20180831_01e.pdf
AV1 Hardware Accelerated Encoder Solution
There will be a demonstration of the world’s first hardware-based implementation of an AV1 encoding system.
AV1 is a new, advanced video data compression standard established by the Alliance for Open Media, a consortium
that includes Socionext, and is capable of compressing data about 30% more effectively than HEVC, whilst keeping the same image quality.
A hardware encoder already?! :cool:
utack
17th October 2018, 23:16
Random noise [...] will also reveal how much of the apparent detail in film comes from the grain.
Same frame with (from libaom) and without grain (from dav1d)
http://screenshotcomparison.com/comparison/122531
Biggest difference is in the river if you ask me
It is probably even more important when pushing the bitrate even lower in a scene like this
Audionut
17th October 2018, 23:55
Unfortunately production workflows put in film grain much earlier than the encoding stage, so it's already baked in way before it gets to an encoder.
Sounds like production workflows need overhauling. :devil:
I can't recall the specifics, but I do recall some sort of overview on bandwidth statistics, showing that some large percentage of worldwide bandwidth is Netflix.
Grain = bandwidth!
benwaggoner
18th October 2018, 01:08
Sounds like production workflows need overhauling. :devil:
Yeah, I'll get on that once I wean the industry off 24p :).
I can't recall the specifics, but I do recall some sort of overview on bandwidth statistics, showing that some large percentage of worldwide bandwidth is Netflix.
Yeah, the majority of internet traffic is now some sort of video. The public numbers use a wide variety of methodology, and have been known to get it wrong by a factor of 2x.
This is one reason we need to forever push the boundaries of compression efficiency. Once we get to transparent compression quality, we need to keep making the required bandwidth to reach that ever lower. Over my 25 years doing digital video, we've seen ~20% improvement every year in the bandwidth required to achieve a given level of quality. Decent 1080p today takes fewer bits than ugly 320x176 did back in 1995.
benwaggoner
18th October 2018, 01:13
http://www.socionext.com/en/pr/sn_pr20180831_01e.pdf
A hardware encoder already?! :cool:
It's a hardware appliance. From the description, it's a fast CPU plus some acceleration. Probably source decode, preprocessing, coarse motion search, frame type selection, weighted prediction, that kind of thing.
ASICs and even GPU-based compression was left behind with HEVC; there are so many different tools a modern codec can use (>2x more in AV1 than HEVC), and GPU's aren't great at the tight inner loops and rapid mode determination required to get good use out of a modern bitstream. Fast individual cores with strong SIMD and L1/2/3 cache run rings around hardware that best works with a one way waterfall cascade of operations.
And CABAC is all about having a few very fast cores. That needs to be done on CPU.
Heck, lots of HW-on-GPU and ASIC encoders don't even have good B-frame support.
benwaggoner
18th October 2018, 01:17
Same frame with (from libaom) and without grain (from dav1d)
http://screenshotcomparison.com/comparison/122531
Biggest difference is in the river if you ask me
It is probably even more important when pushing the bitrate even lower in a scene like this
Good example. And also a good illustration of the challenges with film grain removal. The degrained version also took out the ripples from the water, which were NOT grain. Small details that move stochastically can be very hard to discriminate from grain. It requires some flavor of bidirectional optical flow analysis. Turbulence can be a lot like grain mathematically, but not at all visually.
Selur
20th October 2018, 17:37
According to the Wiki (https://en.wikipedia.org/wiki/AV1#Profiles)
av1 high profile should support 4:2:0 with 8bit and 10bit, but:
ffmpeg -y -loglevel fatal -threads 8 -i "F:\TestClips&Co\files\test.avi" -map 0:0 -an -sn -vf zscale=rangein=tv:range=tv -pix_fmt yuv420p -vsync 0 -f yuv4mpegpipe - | aomenc --passes=1 --pass=1 --target-bitrate=1500 --end-usage=vbr --profile=1 --cpu-used=3 --undershoot-pct=0 --overshoot-pct=0 --buf-sz=6 --buf-initial-sz=4 --buf-optimal-sz=5 --drop-frame=0 --kf-min-dist=0 --kf-max-dist=250 --auto-alt-ref=1 --arnr-maxframes=7 --arnr-strength=5 --noise-sensitivity=0 --sharpness=0 --static-thresh=0 --tune-content=default --tile-columns=0 --tile-rows=0 --min-gf-interval=0 --max-gf-interval=0 --threads=2 --width=640 --height=352 --i420 --input-bit-depth=8 --bit-depth=8 --row-mt=0 --cdf-update-mode=1 -o "E:\Temp\18_13_01_6710_01.ivf" -
ffmpeg -y -loglevel fatal -threads 8 -i "F:\TestClips&Co\files\test.avi" -map 0:0 -an -sn -vf zscale=rangein=tv:range=tv -pix_fmt yuv420p10le -strict -1 -vsync 0 -f yuv4mpegpipe - | aomenc --passes=1 --pass=1 --target-bitrate=1500 --end-usage=vbr --profile=1 --cpu-used=3 --undershoot-pct=0 --overshoot-pct=0 --buf-sz=6 --buf-initial-sz=4 --buf-optimal-sz=5 --drop-frame=0 --kf-min-dist=0 --kf-max-dist=250 --auto-alt-ref=1 --arnr-maxframes=7 --arnr-strength=5 --noise-sensitivity=0 --sharpness=0 --static-thresh=0 --tune-content=default --tile-columns=0 --tile-rows=0 --min-gf-interval=0 --max-gf-interval=0 --threads=2 --width=640 --height=352 --i420 --input-bit-depth=10 --bit-depth=10 --row-mt=0 --cdf-update-mode=1 -o "E:\Temp\18_14_27_4610_01.ivf" -
give me:
Profile 1 requires 4:4:4 color format
also av1 professional profile should support 4:2:0 with 8bit, 10bit, 12bit, but:
ffmpeg -y -loglevel fatal -threads 8 -i "F:\TestClips&Co\files\test.avi" -map 0:0 -an -sn -vf zscale=rangein=tv:range=tv -pix_fmt yuv420p -vsync 0 -f yuv4mpegpipe - | aomenc --passes=1 --pass=1 --target-bitrate=1500 --end-usage=vbr --profile=2 --cpu-used=3 --undershoot-pct=0 --overshoot-pct=0 --buf-sz=6 --buf-initial-sz=4 --buf-optimal-sz=5 --drop-frame=0 --kf-min-dist=0 --kf-max-dist=250 --auto-alt-ref=1 --arnr-maxframes=7 --arnr-strength=5 --noise-sensitivity=0 --sharpness=0 --static-thresh=0 --tune-content=default --tile-columns=0 --tile-rows=0 --min-gf-interval=0 --max-gf-interval=0 --threads=2 --width=640 --height=352 --i420 --input-bit-depth=8 --bit-depth=8 --row-mt=0 --cdf-update-mode=1 -o "E:\Temp\18_16_01_5510_01.ivf" -
and
ffmpeg -y -loglevel fatal -threads 8 -i "F:\TestClips&Co\files\test.avi" -map 0:0 -an -sn -vf zscale=rangein=tv:range=tv -pix_fmt yuv420p10le -strict -1 -vsync 0 -f yuv4mpegpipe - | aomenc --passes=1 --pass=1 --target-bitrate=1500 --end-usage=vbr --profile=2 --cpu-used=3 --undershoot-pct=0 --overshoot-pct=0 --buf-sz=6 --buf-initial-sz=4 --buf-optimal-sz=5 --drop-frame=0 --kf-min-dist=0 --kf-max-dist=250 --auto-alt-ref=1 --arnr-maxframes=7 --arnr-strength=5 --noise-sensitivity=0 --sharpness=0 --static-thresh=0 --tune-content=default --tile-columns=0 --tile-rows=0 --min-gf-interval=0 --max-gf-interval=0 --threads=2 --width=640 --height=352 --i420 --input-bit-depth=10 --bit-depth=10 --row-mt=0 --cdf-update-mode=1 -o "E:\Temp\18_16_44_1810_01.ivf" -
give me:
Profile 2 bit-depth < 10 requires 4:2:2 color format
From the looks of it:
Main at least supports:
420 with 8 and 10 bit
High at least supports
444 with 8 and 10 bit
Professional at least supports
420 with 12 bit
422 with 8,10,12 bit
444 with 8,10,12 bit
-> is the wiki wrong of is this a missing feature or bug in aomenc?
Some reliably info would be nice.
(using: av1 - AOMedia Project AV1 Encoder 1.0.0-810-gc9c806a80)
Cu Selur
nevcairiel
20th October 2018, 17:50
The spec is slightly vague when it comes to this, but even if it would be technically allowed to encode a lower format in a higher profile, you should never do that. There is absolutely no reason to. It'll just screw everything over.
The only thing the profile controls is chroma/bitdepth, there are no other variables, so pick the one appropriate for your content and don't do anything else.
A very strict reading of the spec might even forbid this, ie. Section 6.4.1 (General sequence header OBU semantics) has a table with allowed features per profile, and it has no "backwards" notes.
So I'll go with that. Its not allowed.
Selur
20th October 2018, 18:29
Looking at https://aomediacodec.github.io/av1-spec/av1-spec.pdf
The Main profile supports YUV 4:2:0 or monochrome bitstreams with bit depth equal to 8 or 10.
The High profile further adds support for 4:4:4 bitstreams with the same bit depth constraints.
Finally, the Professional profile extends support over the High profile to also bitstreams with bit depth equal to 12, and also adds support for the 4:2:2 video format.
at page 635 of 665
Thus how I understand it:
Main should support:
4:2:0 at 8bit and 10bit
High should support:
4:2:0 at 8bit and 10bit
4:4:4 at 8bit and 10bit
and Professional should support
4:2:0 at 8bit, 10bit and 12bit
4:2:2 at 8bit, 10bit and 12bit
4:4:4 at 8bit, 10bit and 12bit
Where as aomenc reports:
Profile 1 requires 4:4:4 color format
and
Profile 2 bit-depth < 10 requires 4:2:2 color format
from the looks of it aomenc has a bug here.
Looking at 6.4.1 I agree with you it should be:
Main:
4:0:0 with 8bit
4:2:0 with 8 and 10 bit
High:
4:4:4 with 8 and 10 bit
Professional:
4:0:0 with 8, 10, 12 bit
4:2:0 with 12bit
4:2:2 with 8, 10, 12 bit
4:4:4 with 8, 10, 12 bit
Cu Selur
sneaker_ger
20th October 2018, 20:17
I think they don't want to have one format in different profiles. If I go by what you say I would have the choice to encode 4:2:0 10 bit as either Main or Professional.
So I think this is correct:
Main should support:
Monochrome at 8bit and 10bit
4:2:0 at 8bit and 10bit
High should support:
4:4:4 at 8bit and 10bit
and Professional should support
Monochrome at 12 bit
4:2:0 at 12 bit
4:2:2 at 8bit, 10bit and 12bit
4:4:4 at 12bit
sneaker_ger
20th October 2018, 20:30
Btw: MkvToolNix (https://forum.doom9.org/showpost.php?p=1854553&postcount=131) 28.0.0 now with finalized support for AV1 reading/writing in mkv/webm.
nevcairiel
20th October 2018, 21:22
Looking at https://aomediacodec.github.io/av1-spec/av1-spec.pdf
at page 635 of 665
Annex A is not really restricting the bitstream, its just naming the profiles.
The actual rules are in Section 6.4.1 like I said in my previous post (currently page 112)
6.4.1. General sequence header OBU semantics
seq_profile specifies the features that can be used in the coded video sequence.
seq_profile Bit depth Monochrome support Chroma subsampling
0 8 or 10 Yes YUV 4:2:0
1 8 or 10 No YUV 4:4:4
2 8 or 10 Yes YUV 4:2:2
2 12 Yes YUV 4:2:0, YUV 4:2:2, YUV 4:4:4
This table defines the bitstream rules. As stated there, its not allowed.
aomenc seems to behave just fine. Wikipedia is wrong (what else is new?)
And as said before, there is no sane reason to ever want to do that.
sneaker_ger
20th October 2018, 21:58
Wikipedia is wrong
Liar! :devil:
nevcairiel
20th October 2018, 23:21
That just makes the table look very weird on Wikipedia though, and it also doesn't communicate the limitations of encoding 4:2:0 or 4:4:4 in the professional profile correctly (although the text before it does, but reading pfff)
sneaker_ger
21st October 2018, 09:54
Ok, table corrected (again :( ) and cleaned. Hope it sticks...
Though I'm still wondering about the decoder levels that are allegedly required. Can't find anything about that in the specs. Level 2.2 doesn't even seem to be defined yet. I removed it for now.
SmilingWolf
21st October 2018, 10:05
32/64bits binaries:
av1-1.0.0-811-g68baec84b: https://mega.nz/#!w4Y2AawR!T4zzPbckJmJI02cEKsfxQ5hUhbFg0MhmHr536z_btCM
IgorC
23rd October 2018, 00:11
HEVC eliminates most of the 10-bit advantage over 8-bit that H.264 had. If the source doesn’t have banding, you don’t get much new banding even at lower bitrates. I think AV1 should have ballpark similar improvements.
But yeah, it would be great if the “Main” profile for future codecs always supported at least 10-bit. That’s required for HDR, which is quickly becoming mainstream. It’s not like the SoC or GOU vendors are developing 8-bit only decoders anymore, even if some display pipelines are 8-bit RGB. But 10-bit 64-960 4:2:0 Y’CbCr makes for better 0-255 RGB 4:4:4 anyway.
I guess You're right. x265 doesn't suffer that much from banding as VP9 does.
https://mattgadient.com/x264-vs-x265-vs-vp8-vs-vp9-examples/
https://mattgadient.com/results-encoding-8-bit-video-at-81012-bit-in-handbrake-x264x265/
Interesting news. it's astonishing to see Machine Learning in coding field.
New version of Opus audio codec uses ML. Quality gains are more than significant. Now speech/audio detector has a human-level intelligence. :eek:
https://hub.packtpub.com/opus-1-3-a-popular-foss-audio-codec-with-machine-learning-and-vr-support-is-now-generally-available/
https://opus-codec.org/
I can't imagine what can be done with AV1 in future.
marcomsousa
23rd October 2018, 10:28
Mozilla released today Firefox 63
Firefox 63 adds an AV1 decoder (it's disabled by default).
To try AV1:
Update your stable Firefox (installer (http://ftp.mozilla.org/pub/firefox/releases/63.0/))
Enable AV1: about:config?filter=media.av1.enabled pass the alert message and enable av1 then restart Firefox
Go to the YouTube TestTube page (https://www.youtube.com/testtube).
Select "Prefer AV1 for SD" or "Always Prefer AV1" to get the desired AV1 resolution. Note that at higher resolutions, AV1 is more likely to experience playback performance issues on some devices.
Try playing YouTube clips from the AV1 Beta Launch Playlist (https://www.youtube.com/playlist?list=PLyqf6gJt7KuHBmeVzZteZUlNUQAVLwrZS).
Confirm the codec av01 in "Stats for nerds".
hajj_3
23rd October 2018, 11:51
Firefox 64 Beta is out, which i think enables AV1 by default.
UPDATE: I don't think it is enabled by default, this video won't play: http://video.1ko.ch/codec-comparison/videos/av1-2018-06_550.webm
Pushman
23rd October 2018, 14:10
https://www.bunkus.org/blog/
Here is MKVToolNix v28.0.0: the first release to support AV1 in its finalized form. mkvmerge can read it from Matroska/WebM, MP4, IVF container files and from raw OBU streams. mkvextract will extract it to IVF files. Apart from that a couple of bugs were fixed and usabitliy enhancements made.
New features and enhancements for AV1:
mkvmerge: AV1 parser: updated the code for the finalized AV1 bitstream specification. Part of the implementation of #2261.
mkvmerge: AV1 packetizer: updated the code for the finalized AV1-in-Matroska & WebM mapping specification. Part of the implementation of #2261.
mkvmerge: AV1 support: the `--engage enable_av1` option has been removed again. Part of the implementation of #2261.
mkvmerge: MP4 reader: added support for AV1. Part of the implementation of #2261.
mkvextract: added support for extracting AV1 to IVF. Part of the implementation of #2261.
Mosu
23rd October 2018, 14:18
Btw: MkvToolNix (https://forum.doom9.org/showpost.php?p=1854553&postcount=131) 28.0.0 now with finalized support for AV1 reading/writing in mkv/webm.
https://www.bunkus.org/blog/
Note that the sequence header parser code in mkvmerge v28 was buggy (https://gitlab.com/mbunkus/mkvtoolnix/issues/2410). Those issues have been fixed since. If you want to use mkvmerge/MKVToolNix for AV1, you should use the latest continuous builds (Windows (https://mkvtoolnix.download/windows/continuous/), Linux AppImage (https://mkvtoolnix.download/appimage/continuous/)) instead.
PatchWorKs
24th October 2018, 12:25
Hi everyone, I'm trying to understand how to obtain the better quality/speed/size for FHD/3 (aka 360p or, more precisely, 640x*) resolution videos @ 1 Mbps.
Since AV1 (but other codecs too) supports multiple bitdepth do you suggest to use 12 ?
About chroma: choosing 444 instead of 420 could help to obtain better quality ?
Thanks in advice for anyone can help.
alex1399
24th October 2018, 15:52
Apply moderate De-grain, De-noise, Sharpen and more would help the visual quality. Some encode enhancer (forum.videohelp.com/threads/389555) have friendly GUI to work around if you have video card.
M4ST3R
24th October 2018, 17:14
Mozilla released today Firefox 63
Firefox 63 adds an AV1 decoder (it's disabled by default).
To try AV1:
Update your stable Firefox (installer (http://ftp.mozilla.org/pub/firefox/releases/63.0/))
Enable AV1: about:config?filter=media.av1.enabled pass the alert message and enable av1 then restart Firefox
Go to the YouTube TestTube page (https://www.youtube.com/testtube).
Select "Prefer AV1 for SD" or "Always Prefer AV1" to get the desired AV1 resolution. Note that at higher resolutions, AV1 is more likely to experience playback performance issues on some devices.
Try playing YouTube clips from the AV1 Beta Launch Playlist (https://www.youtube.com/playlist?list=PLyqf6gJt7KuHBmeVzZteZUlNUQAVLwrZS).
Confirm the codec av01 in "Stats for nerds".
It doesn't work me. The youtube testtube page says the AV1 codec is not available in my browser yet.
I use firefox 63 and I have enabled media.av1.enabled
benwaggoner
24th October 2018, 17:54
Hi everyone, I'm trying to understand how to obtain the better quality/speed/size for FHD/3 (aka 360p or, more precisely, 640x*) resolution videos @ 1 Mbps.
Since AV1 (but other codecs too) supports multiple bitdepth do you suggest to use 12 ?
Higher will make encoding and decoding slower. I wouldn’t use >8-bit unless working with a source that is >8-bit with >8-bit detail.
About chroma: choosing 444 instead of 420 could help to obtain better quality ?
444 gives you 24-bit per pixel instead of 12-bit. Making some naive assumptions, one could expect that to double encoding and decoding time. And most progressive content is totally fine with 4:2:0 sub sampling. The exception is if you have very sharp edges between saturated colors. Like sharply rendered red text on a green background. Natural images are almost always great with 4:2:0.
Also, VBR 1 Mbps at 640x360 isn’t a very challenging bitrate; H.264 can do quite well for lots of content at that bitrate, and HEVC can do well at significantly lower. Heck, I could do a nice WMV VC-1 in those constraints a decade ago. Doing encoding tuning at a bitrate where things look good is a lot harder since subtle improvements or degradations might not be visible. Using a bitrate where it’s never going to look great makes the impact of changes more visible and more meaningful. Making mediocre quality significantly better is where differences in encoders really matter.
FWIW. I’m able to get mediocre 1920x800 out of HEVC at 1 Mbps. That’s 1/9th the bits per pixel as 1 Mbps at 640x360.
An “interesting” bitrate to compare 640x360 natural image content at would be more in the 200-500 Kbps range for VBR. If you’re doing a constant bitrate encode, higher can make sense. I’m not sure how accurate CBR rate control is in AV1, though. VP3-9 weren’t ever any good at VBV compliance.
Sent from my iPad using Tapatalk
lvqcl
24th October 2018, 18:36
It doesn't work me. The youtube testtube page says the AV1 codec is not available in my browser yet.
I use firefox 63 and I have enabled media.av1.enabled
32-bit Firefox or 64-bit one?
Mystery Keeper
24th October 2018, 18:54
Works for me. Sad thing is: the uploaded videos are uploaded in inferior formats so far. Re-coding into AV1 doesn't make them better. The first video I watched had color banding. But no DCT artifacts.
marcomsousa
24th October 2018, 21:29
The commit libaom-v1.0.0-820-g4c118dc5e enables by default CONFIG_LOWBITDEPTH=1, 8bit content optimized codepaths.
PatchWorKs
25th October 2018, 08:11
Higher will make encoding and decoding slower. I wouldnÂ’t use >8-bit unless working with a source that is >8-bit with >8-bit detail.
Yes, but - as you probably know - more bits results in better compression...
Anyway source is AVC 264 FHD 8bits 4:2:0
444 gives you 24-bit per pixel instead of 12-bit. Making some naive assumptions, one could expect that to double encoding and decoding time. And most progressive content is totally fine with 4:2:0 sub sampling. The exception is if you have very sharp edges between saturated colors. Like sharply rendered red text on a green background. Natural images are almost always great with 4:2:0.
Yes, I made some tests @444 and encoded videos seems - to me- less contrasted/colorful (more natural ?)...
Also, VBR 1 Mbps at 640x360 isnÂ’t a very challenging bitrate; H.264 can do quite well for lots of content at that bitrate, and HEVC can do well at significantly lower. Heck, I could do a nice WMV VC-1 in those constraints a decade ago. Doing encoding tuning at a bitrate where things look good is a lot harder since subtle improvements or degradations might not be visible. Using a bitrate where itÂ’s never going to look great makes the impact of changes more visible and more meaningful. Making mediocre quality significantly better is where differences in encoders really matter.
Well, the challenge is to stay under 1Gb for 2hrs of stream (audio+video) preserving quality.
Anyway I'll try @ 750Kbps to understand if it's possible to obtain acceptable results.
FWIW. IÂ’m able to get mediocre 1920x800 out of HEVC at 1 Mbps. ThatÂ’s 1/9th the bits per pixel as 1 Mbps at 640x360.
An “interesting” bitrate to compare 640x360 natural image content at would be more in the 200-500 Kbps range for VBR. If you’re doing a constant bitrate encode, higher can make sense. I’m not sure how accurate CBR rate control is in AV1, though. VP3-9 weren’t ever any good at VBV compliance.
I'm testing "realtime" quality encoding, so VBR is not allowed.
Here's the FFMPEG commandline I've used to test VP9:
ffmpeg.exe
-y
-hwaccel auto
-i "<source_filename>"
-c:v libvpx-vp9
-pix_fmt yuv444p12le
-cpu-used 5
-r 30
-g 90
-quality realtime
-speed 7
-threads 4
-row-mt 1
-tile-columns 1
-frame-parallel 0
-qmin 4
-qmax 48
-b:v 1M
-maxrate 1M
-bufsize 1M
-sws_flags lanczos
-vf "crop=<parm>,scale=iw/3:ih/3"
-sn
-an
-f webm "<output>_VP9.webm"
This performs between 30 and 50fps (depending on source complexity) on i3-4160 testing machine.
utack
28th October 2018, 13:44
the past ~2 weeks libaom has been producing files for me it can't decode any more, a few frames into the stream it dies and never recovers
Failed to decode frame: Corrupt frame detected
Additional information: Failed to decode tile data
dav1d is not dying, but shows a lot of blocky colorful artifarcts in many GOPs
Is that happening to everyone else or did i find a rare bug in my build proccess or encoder settings?
Selur
28th October 2018, 14:52
Is that happening to everyone else or did i find a rare bug in my build proccess or encoder settings?
Encoded a bunch of files last week with with up-to-date builds of aomenc and rav1e and had no problem.
SmilingWolf
29th October 2018, 08:41
the past ~2 weeks libaom has been producing files for me it can't decode any more, a few frames into the stream it dies and never recovers
dav1d is not dying, but shows a lot of blocky colorful artifarcts in many GOPs
Is that happening to everyone else or did i find a rare bug in my build proccess or encoder settings?
If you're using more than one slice and more than one thread, you've just met my old friend BUG 2054 (https://bugs.chromium.org/p/aomedia/issues/detail?id=2054)
So far I have reached the conclusion it happens (but not always!) when forcing the encoder to use columns with the width of 2 superblocks (128 pixels). Doesn't happen when the columns have a width of 256 pixels or more on the same file with all the other settings being the same.
I have been doing single threaded encodes for weeks because of this.
hajj_3
29th October 2018, 10:59
MPC-BE beta 1.5.3 v4106 has been released which supports AV1: https://sourceforge.net/projects/mpcbe/files/MPC-BE/Release%20builds/1.5.2/MPC-BE.1.5.2.x64-installer.zip/download
utack
29th October 2018, 11:50
If you're using more than one slice and more than one thread, you've just met my old friend BUG 2054 (https://bugs.chromium.org/p/aomedia/issues/detail?id=2054)
That might be it. Or the 10bit pipeline.
I switched to 8bit and no "threads" option and that works for now.
benwaggoner
29th October 2018, 18:48
Yes, but - as you probably know - more bits results in better compression...
Anyway source is AVC 264 FHD 8bits 4:2:0
Yes, I made some tests @444 and encoded videos seems - to me- less contrasted/colorful (more natural ?)...
There REALLY shouldn’t be any value in using more chroma sub samples than the source! Unless you are downscaling in the same operation. A 720p 4:2:0 source’s chroma samples would be 4:4:4 at 360p.
You might want to try some kind of blind test, though. You shouldn’t see more color or more contrast with more samples. You might see sharper edges between areas of strong color. For example, with ClearType text, which plays all sorts of RGB 444 subpixel tricks that don’t translate well to even other LCD panels, let along different color spaces.
(Pro tip for doing screen recording; disable chroma subpixel rendering for text. Otherwise it can look unpredictably funky on different displays).
Revan654
30th October 2018, 21:30
1. With Rav1e anyone know what type -r Reconstruction takes (Boolean, String, Int, etc...) and what it actually does. I can not find any info on it and the help menu is blank for that option.
2. This has been likely asked a dozen times over, Whats the main reason why aomenc is slower compared to other encoders out there?
I'm just starting to research AV1.
marcomsousa
30th October 2018, 22:56
2. This has been likely asked a dozen times over, Whats the main reason why aomenc is slower compared to other encoders out there?
libaom it was developed for research purposes during AV1 design.
Isn't optimized for speed, only to be a good reference encoder/decoder
Almost all other encoder, the reference encoder isn't fast enough.
That's why there are 3º parties implementation of the same codecs dav1d (https://code.videolan.org/videolan/dav1d) and rav1e (https://github.com/xiph/rav1e)
Now that the AV1 design is finished, the reference encoder (https://aomedia.googlesource.com/aom/+log/master) is now optimizing there code for speed.
We need to wait until there are hardware acceleration (for decoding 4k/8k UHD) (2, 3 years to be mainstream)
New codecs are (always) more complex that the codecs before.
New codecs that are releasing now are too much complex for todays CPU, but they are designed to be using mainstream in future CPU 3 years from now.
Before AV1 became mainstream, AOM was to begin working in AV2, that will me more complex that AV1, and will be design to CPU released 7 or 8 years from now.
Increase the complex is the only way to produce better quality with less space. And that is ok if the hardware continues to improve over time..
LigH
2nd November 2018, 14:19
New uploads: (MSYS2; MinGW32: GCC 7.3.0 / MinGW64: GCC 8.2.0)
AOM v1.0.0-864-g351711076 (https://www.mediafire.com/file/lpgb984xqh79733/aom_v1.0.0-864-g351711076.7z)
now with TPL model (RDO modulation based on frame temporal dependency) and block based denoiser
rav1e 0.1.0 (7492fc5 / 2018-11-01) (https://www.mediafire.com/file/p33yn56lvruc2nz/rav1e_0.1.0_2018-11-01_7492fc5.7z)
dav1d 0.0.1 (287ba91 / 2018-11-02) (https://www.mediafire.com/file/k8tzk89944xvmx9/dav1d_0.0.1_2018-11-02_287ba91.7z)
Mjpeg
2nd November 2018, 18:10
Report from a 1 day meeting on future codecs, h264 through VVC.
http://www.streamingmedia.com/Articles/Editorial/Featured-Articles/At-the-Battle-of-the-Codecs-Answers-on-AV1-HEVC-and-VP9-128213.aspx
The most interesting quote is from a Youtube encoding engineer that AV1 encoding time is down to 16x slower vs. VP9, so that's a nice performance trend (no doubt giving up some % of quality). What's important (as benwaggoner always says) is what quality@perf tradeoffs are available.
benwaggoner
2nd November 2018, 20:39
Report from a 1 day meeting on future codecs, h264 through VVC.
http://www.streamingmedia.com/Articles/Editorial/Featured-Articles/At-the-Battle-of-the-Codecs-Answers-on-AV1-HEVC-and-VP9-128213.aspx
The most interesting quote is from a Youtube encoding engineer that AV1 encoding time is down to 16x slower vs. VP9, so that's a nice performance trend (no doubt giving up some % of quality). What's important (as benwaggoner always says) is what quality@perf tradeoffs are available.
Yeah, YouTube is sort of a special case for encoding. Given how much content they get and the average views/upload, economically they’re going to spend fewer MIPS/pixel than for premium content. And they do their stuff (last I heard) single-threaded on unused-at-that-moment Google servers, ala a spot instance. Which is why libvpx never got a lot of multithreading, nor were the VPx series of bitstream vetted for parallelizability.
Flip side is no one expects spectacular quality from YouTube. It’s free. So even though video game captures always look terrible (lots of high frequency sharp edges...), that’s what people are used to and so they don’t really think about it any more. So YouTube can experiment a lot at how to make good-enough quality fast, which is a different direction from a lot of other folks, and a very useful one in encoder development.
It can be a good starting point for live encoders, although a YouTube can handle some content taking 3x longer to encode than other content due to complexity.
HEVC and AV1 have great tools for making text and video game footage a LOT better. But they are also pretty expensive to add to normal mode detection, so lot of heuristics to figure out when to use them are important.
Mjpeg
2nd November 2018, 22:26
I'm all for AV1 encoders getting faster, but of course on an absolute scale it's pretty slow still, spending a lot of CPU for all that efficiency.
I always figured live-encoding would be quite a wait, although the official framing of AV! always mentions that live-stream/chat was an important case they had thought about.
So anyway I find a talk by Nathan Egge of Mozilla
Most of the slides look familiar but there's a claim about live encoding I had not seen before.
https://people.xiph.org/~negge/AVIF2018.pdf
page 55:
rav1e Live Encoding
Shown at IBC in Sept 2018
● 640x480 @ 30 fps
● Single tile / thread
● Simplified feature set
I guess one sign of AV1 progress will be when someone posts a link here to an AV1 webcam.
SmilingWolf
3rd November 2018, 11:01
32/64bits binaries (GCC 8.2):
av1-1.0.0-877-ge5761e020: https://mega.nz/#!1sgl3QrB!x6F6SfLzz6smB9wJlfczj3y9z1TcBvWZ-IMSzZukIWs
A long standing multithreading bug has been fixed tonight, so here's a new build
Cc @utack
v0lt
3rd November 2018, 16:11
A long standing multithreading bug has been fixed tonight, so here's a new build
I still don’t see aomenc.exe using more than one thread.
SmilingWolf
3rd November 2018, 17:44
It won't by default. Tile columns, rows and threads count are all set to 0.
You have to give at least --tile-columns=1 (and/or --tile-rows=1) --threads=2 to have it use multithreading.
LigH
3rd November 2018, 18:12
32 bit GCC 8.2? ... Does it exist for Windows? MSYS2 did not yet solve internal compiling errors, I believe.
SmilingWolf
3rd November 2018, 18:24
I'm cross compiling from a linux VM. Made it easier to switch between compiler versions back when I was investigating the optimization related bug, then the bug was worked around and the environment stuck.
Silver lining, I'm not stuck with an outdated compiler and I don't have to manage my own MSYS2 package.
Clare
3rd November 2018, 18:35
It won't by default. Tile columns, rows and threads count are all set to 0.
You have to give at least --tile-columns=1 (and/or --tile-rows=1) --threads=2 to have it use multithreading.
Just use --row-mt=1 instead of tiles, it maxes out all my threads.
SmilingWolf
3rd November 2018, 18:57
What does your command line look like? So far I've been unable to make row-mt work myself
My tries so far:
this one generates invalid bitstream:
../../bin8/aomenc --frame-parallel=0 --tile-columns=2 --tile-rows=2 --row-mt=1 --threads=4 --auto-alt-ref=1 --cpu-used=4 --tune=psnr --passes=2 --end-usage=q --cq-level=40 --test-decode=fatal -o test.av1.cq40.webm orig.i420.y4m
Pass 1/2 frame 480/481 92352B 1539b/f 36899b/s 17849 ms (26.89 fps)
Pass 2/2 frame 19/0 0B 17871 ms 1.06 fps [ETA unknown] 2423FFailed to decode frame 2 in stream 0: Corrupt frame detected
Failed to decode tile data
This one works but uses only 12-13% (one core) of my 4c/8t CPU:
../../bin8/aomenc --frame-parallel=0 --tile-columns=2 --tile-rows=2 --row-mt=1 --auto-alt-ref=1 --cpu-used=4 --tune=psnr --passes=2 --end-usage=q --cq-level=40 --test-decode=fatal -o test.av1.cq40.webm orig.i420.y4m
Pass 1/2 frame 480/481 92352B 1539b/f 36899b/s 18130 ms (26.47 fps)
Pass 2/2 frame 16/0 0B 18152 ms 52.88 fpm [ETA unknown]
v0lt
3rd November 2018, 19:30
It won't by default. Tile columns, rows and threads count are all set to 0.
You have to give at least --tile-columns=1 (and/or --tile-rows=1) --threads=2 to have it use multithreading.
I always ask 4 threads, but it never worked.
aomenc --codec=av1 --cq-level=20 --threads=4
Added:
I do not understand what are the columns and rows in this context. If a codec divides a frame into identical independent cells, it is unclear how it can effectively compress in a multi-thread mode.
SmilingWolf
3rd November 2018, 22:20
Tile columns (click to enlarge):
https://thumb.ibb.co/jYeVnf/Screenshot-2.png (https://ibb.co/jYeVnf)
The frame is divided in N (10 in my case) columns of equal width.
Each column in indipendent, so every thread can indipendently work on a tile.
Using tile columns (and/or rows) makes decoding faster too, because each tile can be decoded by a separate thread, making playback way smoother
You command line doesn't show any multithreading because you only have a single big tile, which is being worked on by thread 0, leaving threads 1,2,3 with nothing to do.
v0lt
4th November 2018, 06:54
@SmilingWolf
In this case, the video stream received in the multi-thread mode can be worse than the one-thread mode (with the same bitrate of course). Because motion prediction algorithms will not be able to work effectively.
Added:
I also noticed that some files are decoded by 2 threads, while others are always in single-threaded mode. This is strange, because It is not clear how this will affect hardware decoding.
SmilingWolf
4th November 2018, 11:04
@SmilingWolf
In this case, the video stream received in the multi-thread mode can be worse than the one-thread mode (with the same bitrate of course). Because motion prediction algorithms will not be able to work effectively.
Indeed tile columns affect compression efficiency. However it is also the only way to have threaded decoding and the only way to have smooth 1080p decoding back when I began testing (far before my registration on doom9). Well, at least it was before dav1d, which seems to be using a couple different parallelization techniques.
I'll be running a couple of simple test encodes and decodes and report back some numbers.
Added:
I also noticed that some files are decoded by 2 threads, while others are always in single-threaded mode.
Do you have any samples? Files coming from YouTube perhaps? I'd like to inspect them. I know at least some of the first videos they put online in the AV1 test playlist used a single tile column, which forced single threaded decoding in anything using libaom (e.g. Firefox, Chrome, FFMpeg)
v0lt
4th November 2018, 12:13
@SmilingWolf
As far as I remember, streams obtained using rav1e v1.0.116 were decoded in single-threaded mode. But samples from elecard.com loaded at least 2 cores.
Now it is difficult for me to recheck it, because the decoder in the player works faster than before.
SmilingWolf
4th November 2018, 12:20
Alright, rav1e doesn't support tiles yet, so every frame is a single big column
I'll inspect the Elecard samples ASAP.
Meanwhile my encodes are finishing up, so I'll post size, quality and decoding time differences when they're done
SmilingWolf
4th November 2018, 18:23
The clip used is the F.Y.C one I described some pages ago (http://forum.doom9.org/showthread.php?p=1852449#post1852449)
aomenc/aomdec: 1.0.0-877-ge5761e020
dav1d: 0.0.1 e0c3186
Quality and sizes:
Sizes:
test.av1.cq20.tc0.ivf: 5956739
test.av1.cq20.tc2.ivf: 6001827 +0.75%
test.av1.cq20.tc6.ivf: 6091937 +2.22%
PSNR-HVS-M:
test.av1.cq20.tc0.ivf: 43.192
test.av1.cq20.tc2.ivf: 43.1736 -0.04%
test.av1.cq20.tc6.ivf: 43.1489 -0.10%
MS-SSIM:
test.av1.cq20.tc0.ivf: 26.5095
test.av1.cq20.tc2.ivf: 26.4895 -0.07%
test.av1.cq20.tc6.ivf: 26.467 -0.15%
Decoding:
# aomdec --threads=8 --progress -o /dev/null test.av1.cq20.tc0.ivf
480 decoded frames in 4660361 us (103.00 fps)
# aomdec --threads=8 --progress -o /dev/null test.av1.cq20.tc2.ivf
480 decoded frames in 3365067 us (142.64 fps) +27,79%
# aomdec --threads=8 --progress -o /dev/null test.av1.cq20.tc6.ivf
480 decoded frames in 3267103 us (146.92 fps) +29,89%
# time dav1d -i test.av1.cq20.tc0.ivf -o /dev/null --muxer yuv4mpeg2 -q --framethreads 8 --tilethreads 4
480 decoded frames in 1997 ms (240,36 fps)
# time dav1d -i test.av1.cq20.tc2.ivf -o /dev/null --muxer yuv4mpeg2 -q --framethreads 8 --tilethreads 4
480 decoded frames in 1747 ms (274,75 fps) +12,51%
# time dav1d -i test.av1.cq20.tc6.ivf -o /dev/null --muxer yuv4mpeg2 -q --framethreads 8 --tilethreads 4
480 decoded frames in 1763 ms (272,26 fps) +11,71%
TC 0 means one single column (whole frame)
TC 2 generates 4 columns
TC 6 generates 10 columns
I write they "generate" N columns because there's an upper limit to how many columns fit in a given horizontal resolution, TC 6 implies an actual max of 2^6=64 columns and I use it as a catch all to generate as many columns as possible for my clips.
So the take aways from all this:
for very negligible quality and size differences you can get up to 30% faster decoding performances on 720p. I'd expect it to be even more noticeable on higher resolutions;
dav1d is now faster than libaom;
interestingly, dav1d gave slightly better results with less columns on this particular clip. This might warrant more thorough investigation in the future.
Also, RE: Elecard clips:
they use tile columns (5 for the 720p clips, 10 for the HD clips), so that's why the player used more than one core on those
Mr_Khyron
6th November 2018, 19:15
ffmpeg -hide_banner -t 60 -c:v libdav1d -threads 16 -tilethreads 4 -i Stream2_AV1_4K_22.7mbps.webm -benchmark -f null -
[libdav1d @ 000001ec27984180] libdav1d bd747b1
Input #0, matroska,webm, from 'Stream2_AV1_4K_22.7mbps.webm':
Metadata:
encoder : libwebm-0.2.1.0
Duration: 00:02:24.12, start: 0.000000, bitrate: 22728 kb/s
Stream #0:0(eng): Video: av1 (Main), yuv420p(tv), 3840x2160, SAR 1:1 DAR 16:9, 25 fps, 25 tbr, 1k tbn, 1k tbc (default)
[libdav1d @ 000001ec27a66b80] libdav1d bd747b1
Stream mapping:
Stream #0:0 -> #0:0 (av1 (libdav1d) -> wrapped_avframe (native))
Press [q] to stop, [?] for help
Output #0, null, to 'pipe:':
Metadata:
encoder : Lavf58.22.100
Stream #0:0(eng): Video: wrapped_avframe, yuv420p, 3840x2160 [SAR 1:1 DAR 16:9], q=2-31, 200 kb/s, 25 fps, 25 tbn, 25 tbc (default)
Metadata:
encoder : Lavc58.39.100 wrapped_avframe
frame= 1500 fps= 77 q=-0.0 Lsize=N/A time=00:01:00.00 bitrate=N/A speed=3.09x
video:785kB audio:0kB subtitle:0kB other streams:0kB global headers:0kB muxing overhead: unknown
bench: utime=228.359s stime=38.609s rtime=19.687s
bench: maxrss=2776212kB
I tried bencmarking with ffmpeg 4.2
from 16fps with libaom to 77fps with Dav1d
Clare
9th November 2018, 19:02
What does your command line look like? So far I've been unable to make row-mt work myself
My tries so far:
this one generates invalid bitstream:
../../bin8/aomenc --frame-parallel=0 --tile-columns=2 --tile-rows=2 --row-mt=1 --threads=4 --auto-alt-ref=1 --cpu-used=4 --tune=psnr --passes=2 --end-usage=q --cq-level=40 --test-decode=fatal -o test.av1.cq40.webm orig.i420.y4m
Pass 1/2 frame 480/481 92352B 1539b/f 36899b/s 17849 ms (26.89 fps)
Pass 2/2 frame 19/0 0B 17871 ms 1.06 fps [ETA unknown] 2423FFailed to decode frame 2 in stream 0: Corrupt frame detected
Failed to decode tile data
This one works but uses only 12-13% (one core) of my 4c/8t CPU:
../../bin8/aomenc --frame-parallel=0 --tile-columns=2 --tile-rows=2 --row-mt=1 --auto-alt-ref=1 --cpu-used=4 --tune=psnr --passes=2 --end-usage=q --cq-level=40 --test-decode=fatal -o test.av1.cq40.webm orig.i420.y4m
Pass 1/2 frame 480/481 92352B 1539b/f 36899b/s 18130 ms (26.47 fps)
Pass 2/2 frame 16/0 0B 18152 ms 52.88 fpm [ETA unknown]
aomenc -v --threads=8 --cpu-used=4 --row-mt=1 --lag-in-frames=25 --auto-alt-ref=1--passes=2 --pass=2 --bit-depth=10 --input-bit-depth=10 --end-usage=q --cq-level=28 -o Chimera_DCI4k2398p_HDR_P3PQ.ivf Chimera_DCI4k2398p_HDR_P3PQ.y4m
Mr_Khyron
9th November 2018, 20:16
https://mspoweruser.com/microsoft-release-av1-video-codec-for-windows-10/
Microsoft has released support for the new AV1 royalty-free video codec for Windows 10 via the Microsoft Store.
AOMedia Video 1 (AV1), is an open, royalty-free video coding format designed for video transmissions over the Internet. It is being developed by the Alliance for Open Media (AOMedia) and is meant to be a successor to VP9 without relying on any MPEG patents.
The AV1 extension in the Microsoft Store is an early beta version of the AV1 software decoder. Since this is an early release, users may see some performance issues when playing AV1 videos.
Microsoft says they will be regularly updating the codec via automatic store updates.
Find the new codec in the Microsoft Store here.
hydra3333
10th November 2018, 10:08
he he, clicked on "Get" 30 times in the microsoft store and nothing happens ... that may be saying something about quality.
v0lt
10th November 2018, 13:51
I tried bencmarking with ffmpeg 4.2
from 16fps with libaom to 77fps with Dav1d
Where can I download ffmpeg with libdav1d library?
lvqcl
10th November 2018, 14:22
from 16fps with libaom to 77fps with Dav1d
AFAICS dav1d has only x86-64 AVX2 assembly code, right?
I wonder what's their plans about older hardware...
SmilingWolf
10th November 2018, 15:04
aomenc -v --threads=8 --cpu-used=4 --row-mt=1 --lag-in-frames=25 --auto-alt-ref=1--passes=2 --pass=2 --bit-depth=10 --input-bit-depth=10 --end-usage=q --cq-level=28 -o Chimera_DCI4k2398p_HDR_P3PQ.ivf Chimera_DCI4k2398p_HDR_P3PQ.y4m
That worked, thanks!
Where can I download ffmpeg with libdav1d library?
Win64 GCC 8.2 static build:
ffmpeg-4.2-92396-g55e021f39b: https://mega.nz/#!IgAAVayA!jpzHzBaE6hZpmCb4_1Fdj-es2oRV-FnbjR-ruOD8lCI
- libaom 1.0.0-902-g03d8ebedc
- libdav1d 58fc516
NikosD
10th November 2018, 18:58
In order to GET the new MS AV1 codec from MS Store, you need to install the forbidden (banned) Windows October 2018 Update.
Test:
MS Windows October x64
Core i3 4170
DXVA Checker (new beta version)
Sample:
Chimera AV1 1080p 8bit (Netflix free sample)
LAV x64 0.73.1 vs MS MFT AV1
LAV x64 19/34/144 (min/avg/max fps) CPU Usage: 57/70/83 (%)
MS MFT AV1 15/26/156 CPU Usage: 50/68/81
It seems that AOM AV1 codec is ~30% faster than MS MFT AV1 on average fps
v0lt
10th November 2018, 19:16
@SmilingWolf
Thank.
But my results are different from those that were announced here.
I ran the following tests:
ffmpeg -hide_banner -t 10 -c:v libaom-av1 -i Stream2_AV1_4K_22.7mbps.webm -benchmark -f null -
ffmpeg -hide_banner -t 10 -c:v libdav1d -i Stream2_AV1_4K_22.7mbps.webm -benchmark -f null -
ffmpeg -hide_banner -t 10 -c:v libdav1d -threads 4 -tilethreads 4 -i Stream2_AV1_4K_22.7mbps.webm -benchmark -f null -
And got the following results:
libaom-av1 - max 14 fps
libdav1d - max 7.4 fps
libdav1d -threads 4 -tilethreads 4 - max 9.7 fps
Added:
Intel i5-3570k, Windows 7 Sp1 x64.
richardpl
10th November 2018, 20:20
Probably because you are not using right arch and CPU combo.
Wolfberry
11th November 2018, 04:36
I ran the following tests:
ffmpeg -hide_banner -t 10 -c:v libaom-av1 -i Stream2_AV1_4K_22.7mbps.webm -benchmark -f null -
ffmpeg -hide_banner -t 10 -c:v libdav1d -i Stream2_AV1_4K_22.7mbps.webm -benchmark -f null -
ffmpeg -hide_banner -t 10 -c:v libdav1d -threads 4 -tilethreads 4 -i Stream2_AV1_4K_22.7mbps.webm -benchmark -f null -
I ran the same test as above and get 16/38/46 fps.
What is the CPU you use for testing?
It might be related to the AVX2 code used in dav1d.
v0lt
11th November 2018, 05:12
it might be related to the avx2 code used in dav1d.
sse2, sse4.1?
Aleksoid1978
11th November 2018, 07:37
Very "good" optimisation dav1d - much slower on my system...
Nintendo Maniac 64
11th November 2018, 09:04
AFAICS dav1d has only x86-64 AVX2 assembly code, right?
I wonder what's their plans about older hardware...
Don't forgot that Pentiums and Celerons don't support AVX, and this includes the models that use full-fat Sky/Kaby/Coffee cores such as the ever-popular 2c/4t G4560 and its successor the G5400 (as well as the variants with the faster iGPU like the G4600 and G5500).
And of course, it's those very same AVX-lacking Celerons and Pentiums and such that would stand to gain the biggest benefit from any such software decoder optimizations because those processors simply lack the raw "moar cores!" computational grunt that their i7 and Ryzen brethren have for brute-forcing their way through.
So needless to say, it'd be pretty disappointing to me if dav1d pretty much required having an AVX-capable CPU in order to have any benefit.
Very "good" optimisation dav1d - mush slower on my system...
...that's not a Fernando Alonso reference (https://www.redditmedia.com/mediaembed/6ar8fb), is it?
Mystery Keeper
11th November 2018, 13:14
I wish aomenc/vpxenc had GOP-level parallelism. When each thread is encoding one GOP, and then they are stitched together. That would make use of all CPU power without compromising quality/compression.
Selur
11th November 2018, 13:25
I wish aomenc/vpxenc had GOP-level parallelism.
Which would require 2pass encoding and a fixed gop structue (in regard to the gop sizes), iirc 2nd pass normally should be able to overwrite GOP to archive vbv limits (not totally sure).
Gravitator
11th November 2018, 13:49
ffmpeg -hide_banner -t 10 -c:v libaom-av1 -i 1.mp4 -benchmark -f null - (43 fps)
ffmpeg -hide_banner -t 10 -c:v libdav1d -i 1.mp4 -benchmark -f null - (52 fps)
ffmpeg -hide_banner -t 10 -c:v libdav1d -threads 1 -tilethreads 2 -i 1.mp4 -benchmark -f null - (61 fps)
ffmpeg -hide_banner -t 10 -c:v libdav1d -threads 2 -tilethreads 2 -i 1.mp4 -benchmark -f null - (65 fps)
lvqcl
11th November 2018, 14:08
sse2, sse4.1?
It seems that one of dav1d developers said: "we don't care about mmx/sse2 support anyway" (link (http://lists.ffmpeg.org/pipermail/ffmpeg-devel-irc/2018-October/005348.html)). Have no idea about sse4.1.
SmilingWolf
11th November 2018, 15:12
It seems that one of dav1d developers said: "we don't care about mmx/sse2 support anyway" (link (http://lists.ffmpeg.org/pipermail/ffmpeg-devel-irc/2018-October/005348.html)). Have no idea about sse4.1.
BBB is part of TwoOrioles, so it might have been referred to the company based on its userbase.
Still, MMX is hardly relevant nowadays. SSE4.1 as the lowest bar doesn't sound too unreasonable
Also relevant: https://code.videolan.org/videolan/dav1d/issues/15#note_22262
NikosD
11th November 2018, 19:06
It seems that one of dav1d developers said: "we don't care about mmx/sse2 support anyway"
Have no idea about sse4.1.Still, MMX is hardly relevant nowadays. SSE4.1 as the lowest bar doesn't sound too unreasonable.MMX is too old and not that beneficial as it can reach only 64bits (maybe 80bits max)
SSEx should be the base as it is 128bit with very fast implementation on all CPUs of the last 10 years.
Especially SSE2 is mandatory for x64 architecture.
From the last link it's obvious that dav1d developers targeted AVX2 for 256bit acceleration using ASM, but not exclusively.
They are going to optimise for SSEx later.
So no worries, I think.
marcomsousa
11th November 2018, 20:50
if they want to go with 4k and 8k videos they have to use AVX2.
Nintendo Maniac 64
12th November 2018, 00:30
Especially SSE2 is mandatory for x64 architecture.
You can also usually safely target SSE3 (no, not SSSE3) as well since it's supported on all DDR2-capable 64bit x86 CPUs and newer.
(the only 64bit x86 CPUs that don't support SSE3 are some socket 754 and 939 Athlon 64s which used DDR1)
Mystery Keeper
12th November 2018, 05:22
Which would require 2pass encoding and a fixed gop structue (in regard to the gop sizes), iirc 2nd pass normally should be able to overwrite GOP to archive vbv limits (not totally sure).
I'm totally fine with that. I usually use 2pass anyway. And, of course, I meant I wish they had it as an option.
LigH
12th November 2018, 12:47
@ Nintendo Maniac 64:
Even AMD Athlon64/Phenom (K8-K10 arch.) support some SSE3; but x264/x265 does not use it, considers their implementation as "too slow", I believe.
marcomsousa
12th November 2018, 23:16
SSE3-optimised av1_nn_predict
https://aomedia.googlesource.com/aom/+/486cc9894b7e76b09b4ee37dff6f313f27b1c501
I have developed a SIMD-optimised neural network implementation using
SSE3. I have also added functional equivalence tests between this and
the original implementation. I added aom_clear_system_state() to a few
places where FPU operations are used after av1_nn_predict.
Speed-ups over the original C implementation for various network shapes:
10x64x16: 1.72x
12x12x1: 2.72x
12x24x1: 2.35x
12x32x1: 3.34x
18x24x4: 0.94x
18x32x4: 0.93x
4x16x1: 2.01x
8x16x1: 1.89x
8x16x4: 2.02x
8x24x1: 2.77x
8x32x1: 2.98x
8x64x1: 3.76x
9x32x3: 1.08x
4x8x4: 1.66x
A few awkwardly-shaped networks are slightly slower: these could be
padded to more convenient sizes to use the SIMD kernels.
I also wrote an AVX/AVX2 implementation but on these relatively small
networks it was barely faster than the SSE3 code.
Nintendo Maniac 64
12th November 2018, 23:23
Even AMD Athlon64/Phenom (K8-K10 arch.) support some SSE3
...but this is exactly what I alluded to?
Athlon 64 CPUs are available on socket 754, 939, and AM2; 754 and 939 used DDR1 memory while AM2 used DDR2, and all AM2 CPUs support SSE3.
(there are some socket 754 and 939 CPUs that support SSE3, though it's kind of hit and miss).
Phenom for reference requires at least DDR2.
LigH
13th November 2018, 08:50
I'm sorry, I don't know socket numbers... :o - so we looked at the same topic from different angles. :D
v0lt
13th November 2018, 19:26
I ran the same test as above and get 16/38/46 fps.
What is the CPU you use for testing?
It might be related to the AVX2 code used in dav1d.
Intel i5-3570k (SSE4.1, SSE4.2, AVX), Windows 7 Sp1 x64.
SmilingWolf
15th November 2018, 00:29
Status report!
Previous edition: http://forum.doom9.org/showthread.php?p=1852449#post1852449
Whatever paragraph I don't repeat here can be assumed to be the same as in the aforementioned post
First of all: graphs! Click to enlarge
Y axis: chosen metric
X axis: bits per pixel
720p:
https://thumb.ibb.co/hDPSs0/msssim-720.png (https://ibb.co/hDPSs0)https://thumb.ibb.co/j8ObkL/psnrhvsm-720.png (https://ibb.co/j8ObkL)
1080p:
https://thumb.ibb.co/it4XQL/msssim-1080.png (https://ibb.co/it4XQL)https://thumb.ibb.co/izFvef/psnrhvsm-1080.png (https://ibb.co/izFvef)
BD rates for 720p:
x264 -> rav1e (yeah you read that right!)
RATE (%) DSNR (dB)
MSSSIM -0.736889 0.0375593
PSNRHVS -5.5274 0.375081
rav1e -> x265
RATE (%) DSNR (dB)
MSSSIM -26.5291 1.29942
PSNRHVS -27.1134 1.70509
x265 -> libaom
RATE (%) DSNR (dB)
MSSSIM -18.9088 0.7852
PSNRHVS -15.3123 0.761791
BD rates for 1080p:
x264 -> rav1e (yeah you read that right again!)
RATE (%) DSNR (dB)
MSSSIM -4.92009 0.235151
PSNRHVS -7.23088 0.473125
rav1e -> x265
RATE (%) DSNR (dB)
MSSSIM -26.7063 1.16103
PSNRHVS -28.0007 1.53902
x265 -> libaom
RATE (%) DSNR (dB)
MSSSIM -26.486 0.938124
PSNRHVS -21.7431 0.905916
Encoders:
x264 157-2935-545de2f
x265 2.9-4-471726d3a046
rav1e 0.1.0-702-ab4d23e2
libaom 1.0.0-908-g3a607f7b0
Cmdlines:
x264 --preset veryslow --tune ssim --crf 16 -o test.x264.crf16.264 orig.i420.y4m
x265 --preset veryslow --tune ssim --crf 16 -o test.x265.crf16.hevc orig.i420.y4m
rav1e --low_latency false -o test.rav1e.cq80.ivf --quantizer 80 -s 2 --tune psnr orig.i420.y4m
aomenc --frame-parallel=0 --tile-columns=3 --auto-alt-ref=1 --cpu-used=4 --tune=psnr --passes=2 --threads=2 --end-usage=q --cq-level=20 --test-decode=fatal -o test.av1.cq20.webm orig.i420.y4m
Notes:
So as you can see, the rav1e and aomenc cmdlines have been slightly adjusted to take advantage of the bugfixes and updates from the last months.
In particular, rav1e has been gifted by Frank Bossen the ability to create a B-pyramid, which almost single handedly decreed rav1e's advantage over x264.
A word of warning on this last point: it's still kind of a mixed bag. In very flat, static scenes like PresageFlowerWalk x264 still rules by quite a margin, while rav1e takes the crown in clips like F.Y.C and PresageFlowerFight
F.Y.C, x264 -> rav1e:
RATE (%) DSNR (dB)
MSSSIM -18.451 1.01281
PSNRHVS -25.7463 2.03419
PresageFlowerFight, x264 -> rav1e:
RATE (%) DSNR (dB)
MSSSIM -31.4953 1.80761
PSNRHVS -31.0827 2.27546
PresageFlowerWalk, x264 -> rav1e:
RATE (%) DSNR (dB)
MSSSIM 66.2264 -1.70084
PSNRHVS 70.8208 -2.28853
(as always, a negative BD rate means improvement, positive means regression)
Considerations about times with libaom:
I'm using my desktop PC to run all the encodes. It is also my main study/work PC, so the times can come quite off. Plus, I run multiple encodes in parallel, which further messes up timings.
HOWEVER, between annoying bugs and a lot of stuff, the first report did cost me nearly a week of time (this includes having to re-run some encodes because sh*t happened) ONLY to encode with libaom.
Taking advantage of the recent bugfixes and improvements I have been able to rework my workflow and bring down that time to a couple days only, WITHOUT having to touch the --cpu-used parameter and no night time encoding.
All in all, I am pretty satisfied.
This concludes my (bi-monthly?) report.
As always, I'm open to any kind of feedback to improve my comparisons and my encodes.
benwaggoner
16th November 2018, 19:53
So, what's everyone's favorite AV1 decoder app on Windows? Chrome looks to be not converting from video to PC range correctly (blacks are washed out, contrast is low, etcetera). Is there a nightly of something that does AV! correctly for apples-apples?
SmilingWolf
16th November 2018, 22:45
So, what's everyone's favorite AV1 decoder app on Windows? Chrome looks to be not converting from video to PC range correctly (blacks are washed out, contrast is low, etcetera). Is there a nightly of something that does AV! correctly for apples-apples?
VLC 3.0.5 (Nightly). I fixed my nVidia settings just today because I had that same problem while playing back the ToS fragment I use for the tests. Plays out correctly now.
In alternative, ffplay for quick stuff when I already have a bunch of command prompts open in the right path.
LigH
16th November 2018, 23:27
I use almost only MPC-HC. Which uses LAV Filters with a direct API. It was able to play AV1 clips from the YouTube beta playlist and some tiny own encodes (I don't have powerful CPU's available). So, only a limited experience, yet, but it appears to work.
SmilingWolf
18th November 2018, 12:40
32/64bits binaries (GCC 9.0):
av1-1.0.0-941-gd2a592e1c: https://mega.nz/#!F5Am2KyK!9aQ6_7mM2pDJMsW11CM01Jjsa1R7S6_OaZahvKCHPWQ
mandarinka
19th November 2018, 10:51
I wish aomenc/vpxenc had GOP-level parallelism. When each thread is encoding one GOP, and then they are stitched together. That would make use of all CPU power without compromising quality/compression.
You could get the same results by splitting manually into X parts end encode them separately at once. I'm not sure how much does libvpx/libaom count with that. It works great with x264 and x265 (using raw output at least).
mandarinka
19th November 2018, 10:58
@ Nintendo Maniac 64:
Even AMD Athlon64/Phenom (K8-K10 arch.) support some SSE3; but x264/x265 does not use it, considers their implementation as "too slow", I believe.
SSE3 is not particularly useful for multimedia and it's just a few instructions introduced in Presscot P4 and Venice 90nm K8.
You probably mean SSSE3 (SSS instead of SS) aka "Suplemental SSE3" which is a confusing and dumb name. Probably should have been SSE4 but got renamed for marketing reasons. Or SSE3 was not supposed to be SSE3 originally.
SSSE3 is very useful for encoding and decoding, but only comes on Core 2 chips, and Bobcat/Bulldozer and later cores from AMD. K10 and K8 end at the not-so-important SSE3.
(Note that x265 actually needs SSSE3 + SSE4 to be useful, you are barred from most of assembly optimization if you only have SSSE3, like with 65nm Core 2s or pre-Sandy Bridge Pentium/Celeron).
LigH
19th November 2018, 13:09
Thanks, mandarinka, that explains a bit. I meant SSE3 of 2004 (a.k.a. "Prescott New Instructions" PNI, according to Wikipedia), originally. SSSE3 of 2006 did not arrive in AMD CPUs before the "Cat" (Fusion APU) and "Heavy Equipment" series, so Athlon64/Phenom are clearly out of business.
benwaggoner
20th November 2018, 01:08
You could get the same results by splitting manually into X parts end encode them separately at once. I'm not sure how much does libvpx/libaom count with that. It works great with x264 and x265 (using raw output at least).
Naïve Split-and-stich risks violating VBV at the stitch boundaries and/or reducing quality at those boundaries in order to ensure VBV.
Not that VBV is being used in any AV1 testing I've seen so far.
utack
20th November 2018, 04:13
So I am not entirely sure about what the stats file from first pass includes.
When using pure "q" mode for constant quality, is there a benefit to doing a first pass, or does the first pass only determine how to distribute bitrate when a target bitrate and vbr is specified?
marcomsousa
20th November 2018, 16:37
Building Modern Web Media Experiences: AV1 (Chrome Dev Summit 2018)
https://youtu.be/iTC3mfe0DwE?t=612
VP9 vs H.264
AV1 is 30% smaller in size that VP9
Support in companies
Support in browsers, WebRTC, web
Switch Codecs and Containers in MSE (AV1,VP9,H.264)
DEMO Switch codecs MSE http://storage.googleapis.com/change_type/index.html
uneedme
21st November 2018, 11:24
Hi all
Still, anywhere could find the detail explained parameter functions and arguments range?
forgive my poor wording...
high-end spree means nothing...
utack
22nd November 2018, 00:09
dav1d is doing well
http://www.jbkempf.com/blog/post/2018/dav1d-toward-the-first-release
Wolfberry
22nd November 2018, 11:48
64-bit GCC 8.2.0 binaries: av1-1.0.0-962-1468e60d7 (https://drive.google.com/drive/folders/1xZQABtoaSFgGu11YstmHKYLzO3elemlC)
AVX2 ver of highbd dr predictions Z1,Z3
perfromance increase 1.22x-20x depending on input params
NikosD
22nd November 2018, 18:47
dav1d is doing well
http://www.jbkempf.com/blog/post/2018/dav1d-toward-the-first-release Dav1d is very fast indeed and although is optimized for AVX2, RyZen manages to be a lot faster than Haswell, albeit Haswell has twice as fast AVX2 implementation.
Scaling to more threads and better hyperthreading implementation along with better clocks (?) for the specific SKUs, probably gave RyZen the clear lead.
mandarinka
22nd November 2018, 19:52
The Haswell chips they use for testing is a mobile 4C/8T quadcore which probably runs with low clocks (probably some macbook, so...) and the other is a 4C/4T lower-price desktop SKU which is why it will have lower performance than Ryzen. BTW that Ryzen is a hexacore 6C/12T anyway (yay for AMD!).
littleD
22nd November 2018, 20:49
Wonder how they ran six/eight thread benchmark on 4core/4 thread cpu. If they did, that means decoder has internal switch for thread count. And whats more, single core is underutilized since more CPU threads gives more performance. And since benchmark on 6 core zen gives better results than on 4 thread haswell means AOM decoder they compare to, is highly single threaded. Both decoders have still room to improve anyway.
And. If they compare speed on 12 thread zen this means they compare threading. Global Comparison benchmark shows it.
utack
26th November 2018, 00:17
I encoded the full 4096x1714 4K version of Tears of Steel with libaom.
Bitrate of the final file is about 3.5mbit/s, and quality is definitely on a scale of at least "very good".
Have fun looking what aom can do with that little bitrate, or at benchmarking:
https://drive.google.com/drive/folders/1qp4SvIxzitLipFiudG3iJDautldl0luL?usp=sharing
benwaggoner
26th November 2018, 18:47
I encoded the full 4096x1714 4K version of Tears of Steel with libaom.
Bitrate of the final file is about 3.5mbit/s, and quality is definitely on a scale of at least "very good".
Have fun looking what aom can do with that little bitrate, or at benchmarking:
https://drive.google.com/drive/folders/1qp4SvIxzitLipFiudG3iJDautldl0luL?usp=sharing
Wow, what was the encoding time like?
Can you share the command line you used?
easyfab
26th November 2018, 19:22
@utack
with my AMD 2700x and libdav1d : 98 fps.
frame=17620 fps= 98 q=-0.0 Lsize=N/A time=00:12:14.16 bitrate=N/A speed=4.09x
video:9223kB audio:412879kB subtitle:0kB other streams:0kB global headers:0kB muxing overhead: unknown
bench: utime=1003.859s stime=182.547s rtime=179.705s
bench: maxrss=2521828kB
Did you encode with some tiles ? because It only use 40-50% of the CPU.
You should, it encode faster ( use more threads ) and decode also faster .
SmilingWolf
26th November 2018, 20:39
Did you encode with some tiles ? because It only use 40-50% of the CPU.
You should, it encode faster ( use more threads ) and decode also faster .
He used row-mt and no tiles. At the moment trying to use both makes aomenc sputter invalid bitstreams.
row-mt can maximise CPU usage at encode time, but as you said the lack of tiles reduces playback performance
utack
26th November 2018, 21:47
Wow, what was the encoding time like?
A little over 3 days
Can you share the command line you used?
He used row-mt and no tiles.
It is in the google drive folder
xz -dc tearsofsteel-4k.y4m.xz | aomenc --cpu-used=4 --row-mt=1 --threads=8 --kf-max-dist=250 --bias-pct=75 --webm -o sk.webm --end-usage=q --aq-mode=1 --cq-level=44 --codec=av1 --passes=2 --pass=2 --fpf=/tmp/fpf -
At the moment trying to use both makes aomenc sputter invalid bitstreams.
It seems to do so in the outro, aom craps out at frame 16xxx and I reported it:
https://bugs.chromium.org/p/aomedia/issues/detail?id=2262
row-mt can maximise CPU usage at encode time, but as you said the lack of tiles reduces playback performance
Using dav1d for playback with frame parallel decoding works great, and also does not crash when it reaches the "corrupt" frame
easyfab
26th November 2018, 21:55
My try with the first 1500 frames @1000kb/s with row-mt + tiles
My 2 pass command line :
7z.exe" x "ToS_1920x800_xdither.7z" -so | aomenc.exe --cpu-used=4 --row-mt=1 --threads=16 --tile-columns=4 --tile-rows=2 --kf-max-dist=250 --bias-pct=75 --webm -o tos.webm --target-bitrate=1000 --codec=av1 --passes=2 --pass=2 --fpf=fpf --limit=1500 -
Pass 2/2 frame 1500/1481 8290530B 2633147 ms 34.18 fpm [ETA unknown]
Pass 2/2 frame 1500/1500 8314991B 44346b/f 1064304b/s 2656175 ms (0.56 fps)
And the result file : https://www.sendspace.com/file/dcf6ii
It seems to decode fine for me.
SmilingWolf
26th November 2018, 22:01
Using dav1d for playback with frame parallel decoding works great, and also does not crash when it reaches the "corrupt" frame
Still, using dav1d and having tiles makes a difference between 30-50% faster decoding on my system using Elecard's Holi Festival 4K clip (tested with ffmpeg, latest libdav1d, -tilethreads 1/2/4).
No reason to give that up IMO
BTW, just to be clear, using row-mt does NOT automatically introduce tiling. Using AOMAnalyzer to take a look at the clip confirms that every frame is just one big tile. There are no columns nor rows
@easyfab:
That's interesting, that it decodes properly. It used to croak in the first couple frames for me (http://forum.doom9.org/showthread.php?p=1856831#post1856831)
Maybe it has been fixed and it flew under my radar. Guess I'll check now
EDIT: Holysmoly you're right, mixing tiles, threads and row-mt together has been fixed!
easyfab
26th November 2018, 22:19
for me ( 16 threads cpu ) without tiles, row-mt alone is more than 2x slower ( only use 4/6 threads / cpu usage 10-20 % ) less than 20 fpm.
with row-mt + tiles it use all threads ( cpu usage 50-60% ) and give 34 fpm on TOS 1500 first frames . I'm ok to loose a bit quality with tiles to gain 2x speed.
SmilingWolf
26th November 2018, 22:22
We share the same opinion. The tradeoffs are vastly in favour (http://forum.doom9.org/showthread.php?p=1856939#post1856939) of tiling
julius666
28th November 2018, 00:30
My try with the first 1500 frames @1000kb/s with row-mt + tiles
My 2 pass command line :
7z.exe" x "ToS_1920x800_xdither.7z" -so | aomenc.exe --cpu-used=4 --row-mt=1 --threads=16 --tile-columns=4 --tile-rows=2 --kf-max-dist=250 --bias-pct=75 --webm -o tos.webm --target-bitrate=1000 --codec=av1 --passes=2 --pass=2 --fpf=fpf --limit=1500 -
Pass 2/2 frame 1500/1481 8290530B 2633147 ms 34.18 fpm [ETA unknown]
Pass 2/2 frame 1500/1500 8314991B 44346b/f 1064304b/s 2656175 ms (0.56 fps)
And the result file : https://www.sendspace.com/file/dcf6ii
It seems to decode fine for me.
Wow, this looks incredibly good! And it plays smoothly on my old Thinkpad X230, which is actually the worst (AV1-capable) HW I have access to atm.
This is almost unbelievable given the codec's infancy.
Nintendo Maniac 64
28th November 2018, 10:34
it plays smoothly on my old Thinkpad X230
Are you playing this back with dav1d or libaom?
julius666
29th November 2018, 22:10
Are you playing this back with dav1d or libaom?
I tried it with the latest mpv on Arch Linux. It uses libaom by default:
[vd] Codec list:
[vd] libaom-av1 (av1) - libaom AV1
[vd] Opening decoder libaom-av1
[vd] No hardware decoding requested.
[vd] Using software decoding.
[vd] Detected 4 logical cores.
[vd] Requesting 5 threads for decoding.
[ffmpeg/video] libaom-av1: 1.0.0
[vd] Selected codec: libaom-av1 (libaom AV1)
mandarinka
6th December 2018, 19:57
I see nobody talking about it here, yet, but it seems Google or some other members decided to "unfreeze" the bitstream, there will by the look of it be an incompatible AV1.1.0 revision.
Supposedly it's mostly driven by requests of hardware implementers which wanted to restrict the format a bit in some places.
It looks like hardware might only implement AV1.1.0 mostly, not AV1.0? (This is a bit speculative, maybe there will be exceptions).
Current AV1.0 decoders will support the more restricted AV1.1.0 streams but new (hardware) decoders targetting AV1.1.0 won't be compatible with AV1.0 video. Rather messy, IMHO they should have just delayed the codec finalization and do it properly the first time around, instead of this, but I guess politics prevailed.
There's remarkable silence about it, but I guess it's not the best news so they don't want to brag about it too much until it's final (not sure if the decision is set in stone already).
utack
6th December 2018, 20:33
Supposedly it's mostly driven by requests of hardware implementers which wanted to restrict the format a bit in some places.
.
Even worse that it seems to be about tiles.
These should have never been implemented in a new codec the first place, it is just a lazy workaround because libaom sucks at frame parallel encoding and decoding.
rav1e and dav1d won't need them
SmilingWolf
6th December 2018, 21:34
Ronald Bultje commented on frame parallelism being a bad thing for VP9 (https://blogs.gnome.org/rbultje/2014/02/22/the-worlds-fastest-vp9-decoder-ffvp9/), so not much of a surprise it was turned off by default in AV1/libaom
Tiles help with decoding, do I have to re-link my tests every time this stuff is brought up?
The proposed changes to tile width management only had to be set in stone, all I know is that they have been the de-facto standard in the libaom codebase since at least this summer, when I first began encoding 1080p clips.
EDIT: looks like I was thinking about something else in these regards, nevermind
And since we're talking about this, why at least not link the to the source of the news?
https://www.reddit.com/r/AV1/comments/a1038a/av11_is_launching_soon_and_will_include_breaking/
nevcairiel
6th December 2018, 22:31
Tiles help with decoding, do I have to re-link my tests every time this stuff is brought up?
That may be true, but you can make the same argument for a lot of things. That alone does not necessarily justify a feature thats generally rather annoying, and generally considered a remnant of the past.
In any case, at the current point in time, AV 1.0 bitstreams will just eventually vanish, and AV 1.1 will be the actual standard people use. Outside of tech demos, AV1 is generally unused still.
And with knowing this change is here, noone is going to adopt it now until this is cleared up and "final" again.
mandarinka
8th December 2018, 01:08
I can't tell how correct it is, but this was an interesting read: https://codecs.multimedia.cx/2018/12/why-i-am-sceptical-about-av1/
Author is a former libav/ffmpeg developer if you don't remember his name.
nevcairiel
8th December 2018, 11:13
He can ramble about hardware influence all he wants, but if you don't get the hardware people onboard, your codec is DOA anyway.
utack
8th December 2018, 15:34
I can't tell how correct it is, but this was an interesting read: https://codecs.multimedia.cx/2018/12/why-i-am-sceptical-about-av1/
Author is a former libav/ffmpeg developer if you don't remember his name.
That the daala people who provided most of the new ideas startet their own rav1e encoder from scratch is supporting this blog post.
I also don't get how hardware seems to play a big role, VP9 was not made with hardware in mind, and Qualcomm Samsung Nvidia AMD Intel, as well as some random Chinese SOC vendors still managed to make a hardware decoder. So hardware designers can't be scared off that easily it seems?
Mjpeg
8th December 2018, 17:11
He can ramble about hardware influence all he wants, but if you don't get the hardware people onboard, your codec is DOA anyway.
I agree. I'm just a lurker, but the HEVC licensing debacle opened up a window for a few years, so AV1 needed to jump in quickly to have a chance, which leads to the not-so-radical design that annoys him. I think what he misses is that if AV1 is can succeed, then we'll get AV2, AV3 etc.
It's super clear to me that hardware is crucial because of playback. A codec cannot succeed if "OMG Youtube Is Killing My Battery" is a big reddit thread!
SmilingWolf
8th December 2018, 17:39
That the daala people who provided most of the new ideas startet their own rav1e encoder from scratch is supporting this blog post.
rav1e was started to provide a minimal and fast encoder as early as possible, it was not some kind of statement.
There are even some slides about why working on the existing libvpx-then-branched-libaom codebase made things hard in some xiph.org user folder, I'll edit if I can find them again.
It mainly had to do with libaom being big, old and full of experiments anyway
---
Alright found a couple right off the bat:
https://people.xiph.org/~tdaede/rav1e_vdd_2017.pdf
Started as a reimplementation of AV1 in order to find bitstream and specification bugs
● Could do an encoder or decoder:
– Decoder (especially fuzzed) more useful to find mismatches
– Encoder doesn’t need all features implemented to work
● Algorithmic improvements over libaom
https://people.xiph.org/~tdaede/rav1e_vdd_2018.pdf
Background on libaom:
Derived from libvpx codebase
● Reference implementation, “sort of usable”
● Much encoder behavior is inherited from previous VPx codecs
– multiple frame passes
– weird rate control
As for hardware, we can speculate how Google had to drag some vendors on board with who knows what promises or deals as a last ditch effort to promote VP9. It was not designed with hardware in mind, hardware support came very late, and whoa look at the amount of people willing to make VP9 encodes around outside of Google/Youtube! /s
One would think they learnt from that experience. Again, we can only speculate, but that would provide a good explanation for the veto power of hardware vendors this time around.
That may be true, but you can make the same argument for a lot of things. That alone does not necessarily justify a feature thats generally rather annoying, and generally considered a remnant of the past.
But it's a matter of fact that we're stuck with it. Considering it's not that hard to understand how to use it from an encoder (the person, not the software) POV and the low overhead, discouraging its use will only have people creating streams that can't be reproduced on a very large amount of systems. Next thing you know, plenty of people will be screaming bloody murder because their 16 core system stutters on a 1080p clip with mild to heavy motion and the CPU is still underused.
What are the modern alternatives to tiling? I remember reading somewhere WPP is patented, so what would be the next best alternative?
alex1399
9th December 2018, 11:37
Oh no, Youtube coupled the high define resolution with the 60fps in all most every video. It uses a trivial motion interpolation that simply duplicates frame from previous and converts native frame-rate into 60fps. What a jerk. Thats why it works so hard on some low end PC with high speed Internet.
Nintendo Maniac 64
9th December 2018, 22:23
It seems that at some point YouTube actually implemented AV1 "in the wild". Much like how VP9 was rolled out, it seems that the larger the view count then the more likely you'll get an AV1 encode...but not entirely - this is most obvious on Linus Tech Tips where this LTT video with ~1.6M views (https://www.youtube.com/watch?v=U5dJn7V_4Bk) has AV1 but this other LTT video posted just 1 day before with ~2.8M views (https://www.youtube.com/watch?v=VWHlPH23P-w) does not...
Unless YouTube flipped an AV1 switch sometime specifically on December 3rd only for newly-posted video, making the first-linked video quality while the second-linked video did not?
Anyway, I just finished a bunch of CPU decoding performance testing for one of SmilingWolf personal videos that he encoded in AV1, and while the absolute performance numbers are relatively meaningless since they largely depend on things like bitrate, framerate, and resolution, the relative decoding performance numbers should still be of interest.
The way I measured included (but was not limited to) having a video clip play at 17fps and then seeing what the lowest clockrate required was for a given CPU architecture and thread configuration with a 45nm Core 2 Duo @ 3.5GHz as the baseline. I also tested within a single architecture for performance scaling at various clockrate, and save for Wolfdale (possibly due to the lack of an integrated memory controller), the performance scaling was for all intents and purposes identical between Nehalem and Haswell (that is, the percentage amount of extra clockrate necessary to play back 24fps vs 17fps was darned-near exactly the same)
All of this was tested with MPC-HC 1.8.3 x64 (LAVfilters 0.73, libaom) and only on CPUs that supported SSE 4.1 as CPUs lacking this instruction set would have needed a 10+ GHz overclock (I'm not kidding) such as the Phenom II and the 65nm Conroe-based Core 2 Duo (even though the later supports SSSE3 and not just SSE3).
And to clarify, the percentage below is simply how much faster a given CPU should be if it had the same clockrate as the baseline 2c/2t Wolfdale (which itself was clocked at 3.5GHz).
100% - Wolfdale 2c/2t
119% - Nehalem 2c/2t
130% - Wolfdale 4c/4t
138% - Nehalem 2c/4t
167% - Haswell 2c/2c
175% - Nehalem 4c/4t
175% - Nehalem 4c/8t (not a typo)
And from some of the other tests I did, I was able to extrapolate the performance of Haswell CPUs configured at 2c/4t, 4c/4t, and 4c/8t (again, relative to a 2c/2t Wolfdale) as I do not have access to Haswell CPUs with thread configurations greater than 2c/2t:
199% - Haswell 2c/4t
241% - Haswell 4c/4t
241% - Haswell 4c/8t (not a typo)
For those that don't know their CPU architectures...
Wolfdale = second generation desktop Core 2 Duo/Quad, 45nm die-shrink
Nehalem = 45nm, first generation of Intel CPUs that use the Core i5/i7 branding (i3 didn't come along until the 32nm Westmere die-shrink...which was still considered "1st gen" oddly enough)
Haswell = 22nm, fourth generation of Intel CPUs that use the Core i3/i5/i7 branding
Zebulon84
10th December 2018, 00:18
It seems that at some point YouTube actually implemented AV1 "in the wild". Much like how VP9 was rolled out, it seems that the larger the view count then the more likely you'll get an AV1 encode...but not entirely - this is most obvious on Linus Tech Tips where this LTT video with ~1.6M views (https://www.youtube.com/watch?v=U5dJn7V_4Bk) has AV1 but this other LTT video posted just 1 day before with ~2.8M views (https://www.youtube.com/watch?v=VWHlPH23P-w) does not...
Unless YouTube flipped an AV1 switch sometime specifically on December 3rd only for newly-posted video, making the first-linked video quality while the second-linked video did not?
The one with more views without AV1 is also longer (18 min vs 11 min), so it may be still encoding, or length is taken into account when choosing which video is worth spending hours of encoding time for a few AV1 views.
Thanks for the benchmarks, it proves it's possible to do software decoding of AV1 on desktop, but it's quite heavy, and will be hard on laptop battery and probably too much for any smartphones. Do you plan to add ...lake, Ryzen or dav1d ?
Nintendo Maniac 64
10th December 2018, 01:03
it proves it's possible to do software decoding of AV1 on desktop, but it's quite heavy, and will be hard on laptop battery and probably too much for any smartphones
Remember that this depends heavily on video resolution, framerate, and bitrate.
In particular, the video I was using I had manually slowed down to 17fps because that was the highest framerate my Nehalem x3470 could handle when configured as 2c/2t and turbo disabled (2.93GHz), which is why I did not provide any absolute performance numbers and focused on relative performance.
Do you plan to add ...lake, Ryzen or dav1d ?
I don't have any such CPUs, and I've no idea how to benchmark dav1d since coding and such is totally not my specialty (my expertise is much more in hardware).
However, Zen-based CPUs should have per-GHz performance similar to Haswell while Sky/Kaby/Coffee lake will have slightly better per-GHz performance than Haswell, so for the most part you can just use the Haswell relative performance numbers (not the clockrate!) as a reference for those architectures.
hajj_3
11th December 2018, 17:11
dav1d v0.1 has been released: http://www.jbkempf.com/blog/post/2018/First-release-of-dav1d
sneaker_ger
11th December 2018, 17:52
And, we've been experimenting with shaders, notably for the Film Grain feature.
shaders = GPU?
v0lt
11th December 2018, 17:52
dav1d v0.1 has been released: http://www.jbkempf.com/blog/post/2018/First-release-of-dav1d
Good news.
I will wait for the ffmpeg build with both libaom and libda1d libraries. I want to compare the speed of work in the same conditions.
SmilingWolf
11th December 2018, 20:07
64bits, GCC 8.2:
ffmpeg 4.2-92673-g876ed08b0d: https://mega.nz/#!QxpinIyQ!HBtUEzFObdc5RDFEc3UrzOdaRo8QxNxABGMlPcvgSWA
- libaom 1.0.0-1024-g5b8f393fe
- libdav1d 0.1.0 c0501f1
sneaker_ger
11th December 2018, 20:48
Thx. Seems dav1d is still easily 40% slower on an i5-2500K (AVX but no AVX2) compared to libaom. Only ~50% CPU utilization on both.
nevcairiel
11th December 2018, 21:14
SSE* code is still actively being worked on, and is actively coming in right now, so its getting faster day by day on those systems. AVX1 doesn't help a codec like this much, since AVX instructions are primarily floating point, and only AVX2 adds the required integer instructions.
SmilingWolf
11th December 2018, 21:25
To follow the progress of SSSE3 implementation: https://code.videolan.org/videolan/dav1d/issues/216
Same thing for NEON: https://code.videolan.org/videolan/dav1d/issues/215
An article on dav1d 0.1.0 by the same guy who's been doing most of the benchmarks that appeared in the official blogposts: https://medium.com/@ewoutterhoeven/dav1d-0-1-0-release-the-first-benchmarks-5404360e44e3
v0lt
12th December 2018, 04:03
SSE* code is still actively being worked on, and is actively coming in right now, so its getting faster day by day on those systems. AVX1 doesn't help a codec like this much, since AVX instructions are primarily floating point, and only AVX2 adds the required integer instructions.
They claim that dav1d is always faster than libaom. They say that there are problems only in single-threaded mode. This lie breaks. :mad:
On modern desktop, dav1d is very fast, compared to other decoders:
Pentium G5600 not modern?
But, since the previous blogpost, we've added more assembly for desktop, and we've merged some assembly for ARMv8, and for older machines (SSSE3).
We're now as fast as libaom, in single-thread, on ARMv8, and faster with more threads.
My tests:
ffmpeg -t 10 -c:v libaom-av1 -i Stream2_AV1_4K_22.7mbps.webm -benchmark -f null -
ffmpeg -t 10 -c:v libdav1d -i Stream2_AV1_4K_22.7mbps.webm -benchmark -f null -
ffmpeg -t 10 -c:v libdav1d -threads 4 -tilethreads 4 -i Stream2_AV1_4K_22.7mbps.webm -benchmark -f null -
Result:
libaom-av1 - 14 fps
libdav1d - max 7.1 fps
libdav1d -threads 4 -tilethreads 4 - max 9.6 fps
I got the exact same result a month ago (https://forum.doom9.org/showthread.php?p=1857317#post1857317).
Wolfberry
12th December 2018, 04:45
This commit (https://code.videolan.org/videolan/dav1d/commit/02312cae6c45a58d1b275ad80eb6c41271415c3b) may help.
Not sure if it is CLI only or can be used in ffmpeg.
Nintendo Maniac 64
12th December 2018, 05:04
Pentium G5600 not modern?.
Unfortunately, many people do not realize that Intel Pentiums do not support AVX at all.
At least going forward there's now an AMD alternative in the form of the Athlon 200GE which does support AVX2 (in addition to having a better iGPU and actual sane prices in lieu of Intel's 14nm shortage), but that processor was only just released a couple months ago.
MoSal
12th December 2018, 05:48
libaom-av1 - 14 fps
libdav1d - max 7.1 fps
libdav1d -threads 4 -tilethreads 4 - max 9.6 fps
I got the exact same result a month ago (https://forum.doom9.org/showthread.php?p=1857317#post1857317).
Can you try -threads 8 -tilethreads 1?
marcomsousa
12th December 2018, 10:55
ffmpeg-4.2-92681-0e833f6
- libaom 1.0.0-1028-78e6b2c
- libdav1d 0.1.0 73067e5
ffmpeg -t 10 -c:v libaom-av1 -i Stream2_AV1_4K_22.7mbps.webm -benchmark -f null -
ffmpeg -t 10 -c:v libdav1d -i Stream2_AV1_4K_22.7mbps.webm -benchmark -f null -
ffmpeg -t 10 -c:v libdav1d -threads 4 -tilethreads 4 -i Stream2_AV1_4K_22.7mbps.webm -benchmark -f null -
Result:
Code:
libaom-av1 - 21 fps 0.780x speed
libdav1d - 41 fps 1.65x speed
libdav1d -threads 4 -tilethreads 4 - 58 fps 2.31x speed
CPU: Intel i7 8550U (MMX, SSE, SSE2, SSE3, SSSE3, SSE4.1, SSE4.2, EM64T, AES, AVX, AVX2, FMA3)
Gravitator
12th December 2018, 12:02
ffmpeg-4.2-92396-g55e021f39b (https://forum.doom9.org/showthread.php?p=1857294#post1857294)
- libaom 1.0.0-902-g03d8ebedc
- libdav1d 58fc516
ffmpeg -hide_banner -t 10 -c:v libaom-av1 -i 1.mp4 -benchmark -f null - (43 fps)
ffmpeg -hide_banner -t 10 -c:v libdav1d -i 1.mp4 -benchmark -f null - (52 fps)
ffmpeg -hide_banner -t 10 -c:v libdav1d -threads 1 -tilethreads 2 -i 1.mp4 -benchmark -f null - (61 fps)
ffmpeg -hide_banner -t 10 -c:v libdav1d -threads 2 -tilethreads 2 -i 1.mp4 -benchmark -f null - (65 fps)
ffmpeg-4.2-92681-0e833f6 (https://forum.doom9.org/showthread.php?p=1859780#post1859780)
- libaom 1.0.0-1028-78e6b2c
- libdav1d 0.1.0 73067e5
ffmpeg -hide_banner -t 10 -c:v libaom-av1 -i 1.mp4 -benchmark -f null - (45 fps)
ffmpeg -hide_banner -t 10 -c:v libdav1d -i 1.mp4 -benchmark -f null - (51 fps)
ffmpeg -hide_banner -t 10 -c:v libdav1d -threads 1 -tilethreads 2 -i 1.mp4 -benchmark -f null - (58 fps)
ffmpeg -hide_banner -t 10 -c:v libdav1d -threads 2 -tilethreads 2 -i 1.mp4 -benchmark -f null - (63 fps)
nevcairiel
12th December 2018, 12:50
They claim that dav1d is always faster than libaom. They say that there are problems only in single-threaded mode. This lie breaks. :mad:
Actuall it says that it will soon be faster then other decoders on all platforms. "soon" and not now.
If you don't have AVX2, the decoder is still being bottlenecked quite heavily, and also won't thread quite as nicely because reference frames take too long to decode, for example. The SSSE3 work is still at early stages - if you look at the ticket linked above, only a small part of assembly has been covered in SSSE3 yet.
NikosD
12th December 2018, 14:39
SSSE3 code base is fundamental because it's the first instruction test supported by all Core 2 Duo and above (not Pentium 4) and also it's very useful for decoding (at least on previous codecs like H.264/H.265)
But I don't know if they want to go back to even older instruction sets and CPUs like SSE2.
We'll see.
sneaker_ger
12th December 2018, 15:15
Shame for AMD users. K10 (like Phenom II) and similar don't have SSSE3. Produced up to 2012. Of course it will be years until AV1 is de-facto required (if ever) so by then...
SmilingWolf
12th December 2018, 15:24
But I don't know if they want to go back to even older instruction sets and CPUs like SSE2.
https://code.videolan.org/videolan/dav1d/issues/207#note_24056
nevcairiel
12th December 2018, 15:30
Shame for AMD users. K10 (like Phenom II) and similar don't have SSSE3. Produced up to 2012. Of course it will be years until AV1 is de-facto required (if ever) so by then...
The marketshare of non-SSSE3 desktop CPUs is so small that noone is really going to bother with that, particularly because many of those CPUs are often times going to be too slow for any real use anyway.
And in all honesty, if you bought a K10 in 2012 or anywhere near to that, you just did it wrong, even on the low-end market.
Intel introduced SSSE3 all the way back in 2006, afterall. Its hardly "new" even in 2012.
Ultimately its up to the developers how they want to spend their time, but as mentioned in the ticket linked above, pure SSE2 is often a lot more painful to write then using SSSE3 enhancements.
NikosD
12th December 2018, 15:38
https://code.videolan.org/videolan/dav1d/issues/207#note_24056 Thank you.
So, SSSE3 is the minimum.
Little pity for AMD CPUs.
sneaker_ger
12th December 2018, 15:39
Steam HW Survey says 3% don't have SSSE3, only 0.01% don't have SSE3.
https://store.steampowered.com/hwsurvey
nevcairiel
12th December 2018, 15:43
SSE3 (without the third S) is mostly useless for video. Its primarly floating-point.
For video, which needs integer instructions, you only have a few meaningful steps: (everything left out is mostly floating point or otherwise not related, like SSE3, AVX1, etc).
- MMX
- SSE2
- SSSE3
- SSE4.1
- AVX2
- AVX512
Obviously noone cares about MMX anymore. SSE4.1 is only useful in special cases. And obviously AVX512 is not rolled out and perhaps even understood widely enough yet, maybe in a few years.
So, by and large, that leaves SSE2, SSSE3, AVX2. The difference between SSE2 and SSSE3 is not gigantic, same 128-bit registers afterall, SSSE3 only adds a bunch of new instructions - but some of those are really useful and make code much simpler and easier to write.
clsid
12th December 2018, 16:33
The optimizations in Dav1d are currently mostly for 8-bit only. So for 10-bit libaom may still be faster.
Development pace in Dav1d is pretty high, so we will have a fast decoder long before there is actual widespread AV1 content (beyond the current demo files and a few Youtube videos).
v0lt
12th December 2018, 16:34
Can you try -threads 8 -tilethreads 1?
I test again. i5-3570K.
libaom-av1 - max 14 fps
libdav1d - max 7.2 fps
libdav1d -threads 4 -tilethreads 4 - max 9.7 fps
libdav1d -threads 8 -tilethreads 1 - max 10 fps
Actuall it says that it will soon be faster then other decoders on all platforms. "soon" and not now.
I carefully read their "press releases". I did not see them writing about slow speed without AVX2. But they know exactly about this. This happens the second time. "Press releases" write for sponsors?
I'm waiting for the dav1d to be faster on my processor. I want to see truthful information, not PR.
SmilingWolf
12th December 2018, 16:53
I carefully read their "press releases". I did not see them writing about slow speed without AVX2. But they know exactly about this. This happens the second time. "Press releases" write for sponsors?
I'm waiting for the dav1d to be faster on my processor. I want to see truthful information, not PR.
No you didn't.
The blogpost links twice to this previous one for detailed perf reports: http://www.jbkempf.com/blog/post/2018/dav1d-toward-the-first-release
Today, dav1d is very fast on AVX2 processors, which should cover a bit more than 50% of the CPUs used on the desktop. We wrote 95% of the code needed for AVX2, but there is still a bit more achievable.
We're readying the SSE and the ARM optimizations, to do the same. They will be very fast too, in the next weeks.
It's clearly stated that dav1d is the fastest on AVX2.
Then the same post you claim to have read very carefully states that work on SSSE3 has only just begun.
Since the Pentium G5600 only supports extensions up to SSE4.2 it's clear you'll have to wait some more.
Spare the rage and read some more
v0lt
12th December 2018, 17:18
No you didn't.
The blogpost links twice to this previous one for detailed perf reports: http://www.jbkempf.com/blog/post/2018/dav1d-toward-the-first-release
Please quote the text where it is written that dav1d without AVX2 will run slower.
This information I could find only in the discussion of beta testing.
SmilingWolf
12th December 2018, 17:29
Please quote the text where it is written that dav1d without AVX2 will run slower.
This information I could find only in the discussion of beta testing.
A certain extension provides a speedup. It follows, w/o said extension things will be slower.
You have the wunderbar vector extensions: you have the speedup these provide.
You can't use the vector extensions: you're going to run on C code, which is gonna be slower. Which is the reason these multimedia extensions exist in the first place.
Doesn't really take a degree to understand.
I got it, everyone around here got it, it seems you're the only one left out. Wonder where the problem lies?
easyfab
12th December 2018, 18:06
And if you want the latest info for SIMD you should look :
AVX2 https://code.videolan.org/videolan/dav1d/issues/78
SSSE3 https://code.videolan.org/videolan/dav1d/issues/216
ARM / NEON https://code.videolan.org/videolan/dav1d/issues/215
As you can see for AVX2 it's pretty much done, but only a few for others. And that only for 8bit if i'm correct.
Nintendo Maniac 64
12th December 2018, 19:40
many of those CPUs are often times going to be too slow for any real use anyway.
And in all honesty, if you bought a K10 in 2012 or anywhere near to that, you just did it wrong, even on the low-end market.
Keep in mind that even the Llano 1st gen APUs lacked SSSE3 due to their K10-derived CPU architecture.
As someone with both a Phenom II x4 and a Core 2 Quad (actually a Phenom II x2 unlocked to x4 and a quad Wolfdale Xeon), I find that the latter has pretty sub-par multicore scaling in video workloads - yes it's faster than a Core 2 Duo, but not quite at the level that you'd expect as I showed in my post two pages back (https://forum.doom9.org/showthread.php?p=1859536#post1859536) (if Wolfdale had the same scaling from 2c/2t to 4c/4t as Nehalem, then 4c/4t Wolfdale would've only needed ~2.4GHz, not 2.7GHz)
This then commonly results in the Phenom actually performing similar to if not better than the Core 2 Quad on a per-GHz basis assuming the tested code isn't heavily relying on SSSE3 or SSE4.1 (as is obviously the case currently with AV1 decoding), and the Phenom not only tended to have higher stock clocks but even came in 6 core variants as well.
Similarly, I've also previously documented that the Phenom II is faster than Core 2 Quad clock-for-clock in SVP video interpolation (http://www.svp-team.com/forum/viewtopic.php?pid=68658) (which is a task that loves "moar cores!" and SMT threads).
benwaggoner
12th December 2018, 20:06
And if you want the latest info for SIMD you should look :
AVX2 https://code.videolan.org/videolan/dav1d/issues/78
SSSE3 https://code.videolan.org/videolan/dav1d/issues/216
ARM / NEON https://code.videolan.org/videolan/dav1d/issues/215
As you can see for AVX2 it's pretty much done, but only a few for others. And that only for 8bit if i'm correct.
Do we have numbers for the installed base of AVX2 capable PCs? They've been in all new mainstream systems for several years now. I'd guess it's >50% already.
Nintendo Maniac 64
12th December 2018, 21:22
2Do we have numbers for the installed base of AVX2 capable PCs? They've been in all new mainstream systems for several years now.
I realize I sound like a broken record at this point, but the newest Pentiums and Celerons still do not support AVX, and this even applies to the models that use the full-fat Sky/Kaby/Coffeelake cores (though with smaller cache size) such as the ever-popular 2c/4t Pentium G4560 and its direct successor the G5400.
(and again, going forward the Athlon 200GE is a wiser choice of CPU, but that's only been on the market for a couple months now)
mzso
13th December 2018, 12:08
Hi!
On the decoder sides Dav1d and libAOM are the only two options? I see Firefox has a Dav1d option, which doesn't work too well, because it freezes on the bitmovin demo. (I guess the other is libaom.) The default decoder plays the video completely smoothly now on my computer.
PS:
By the way, can I download these streams?
The player is pretty trashy, the quality always resets and doesn't want to change unless I seek.
mzso
13th December 2018, 12:22
Even worse that it seems to be about tiles.
These should have never been implemented in a new codec the first place, it is just a lazy workaround because libaom sucks at frame parallel encoding and decoding.
rav1e and dav1d won't need them
Why shouldn't we like tiled encoding?
utack
13th December 2018, 14:28
Why shouldn't we like tiled encoding?
They make compression efficiency worse.The current implementation splits the frame into equal parts, and most of the times you get a split right in the center of the picture where most action takes place.
dav1d demonstrates pretty well that frame parallel decoding works fairly well, other encoders managed to get perfect frame parallel encoding done, so it just seems a lazy solution until libaom gets row_mt running well.
LigH
13th December 2018, 16:34
New uploads: (MSYS2; MinGW32: GCC 7.4.0 / MinGW64: GCC 8.2.1)
AOM v1.0.0-1030-g7ac3eb1bb (https://www.mediafire.com/file/ro61f2rjhrw5y76/aom_v1.0.0-1030-g7ac3eb1bb.7z)
New parameters:
--enable-dual-filter=<arg> Enable dual filter (0: false, 1: true (default))
--enable-order-hint=<arg> Enable order hint (0: false, 1: true (default))
--enable-dist-wtd-comp=<arg Enable distance-weighted compound (0: false, 1: true (default))
--enable-masked-comp=<arg> Enable masked (wedge/diff-wtd) compound (0: false, 1: true (default))
--enable-interintra-comp=<a Enable interintra compound (0: false, 1: true (default))
--enable-diff-wtd-comp=<arg Enable difference-weighted compound (0: false, 1: true (default))
--enable-interinter-wedge=< Enable interinter wedge compound (0: false, 1: true (default))
--enable-interintra-wedge=< Enable interintra wedge compound (0: false, 1: true (default))
--enable-global-motion=<arg Enable global motion (0: false, 1: true (default))
--enable-warped-motion=<arg Enable local warped motion (0: false, 1: true (default))
--enable-obmc=<arg> Enable OBMC (0: false, 1: true (default))
rav1e 0.1.0 (64b9f50 / 2018-12-13) (https://www.mediafire.com/file/up68lvmt9j8m9e3/rav1e_0.1.0_2018-12-13_64b9f50.7z)
dav1d 0.1.0 (e5bca59 / 2018-12-13) (https://www.mediafire.com/file/v8d90s8ykco32zl/dav1d_0.1.0_2018-12-13_e5bca59.7z)
SmilingWolf
13th December 2018, 16:52
They make compression efficiency worse.
In x265, WPP hurts efficiency too. Should we stop using it?
The clip used is the F.Y.C one I described some pages ago (http://forum.doom9.org/showthread.php?p=1852449#post1852449)
Cmdlines:
x265 --preset veryslow --tune ssim --crf 20 -F 1 --no-wpp -o test.x265.crf20.1F.00WPP.hevc orig.i420.y4m
x265 --preset veryslow --tune ssim --crf 20 -F 1 -o test.x265.crf20.1F.12WPP.hevc orig.i420.y4m
Sizes:
test.x265.crf20.1F.00WPP.hevc: 5566953
test.x265.crf20.1F.12WPP.hevc: 5612446 (+0.81%)
PSNR-HVS-M:
test.x265.crf20.1F.00WPP.hevc: 42.9368
test.x265.crf20.1F.12WPP.hevc: 42.9299 (-0.02%)
MS-SSIM:
test.x265.crf20.1F.00WPP.hevc: 26.3172
test.x265.crf20.1F.12WPP.hevc: 26.3112 (-0.02%)
With libaom the compression efficiency loss is very very low with an acceptable amount of tiles (in this case, 4 on a 720p clip).
I have already measured it: http://forum.doom9.org/showthread.php?p=1856939#post1856939.
That's -0.75% space efficiency with 0.0X% loss in quality. It's even comparable to x265's WPP!
On the other hand, libaom's --frame-parallel=1 exhibits a 6% overhead. Just so that we're clear, libaom's --frame-parallel has got nothing to do with libdav1d's decoding option with the similar name, which doesn't depend on any optional characteristic of the bitstream.
You can already have row-mt WITH tiles which should work decently. Maybe combine it with chunked encoding for better overall performance.
So again, no excuses to not use tiles.
marcomsousa
13th December 2018, 16:57
I realize I sound like a broken record at this point, but the newest Pentiums and Celerons still do not support AVX, and this even applies to the models that use the full-fat Sky/Kaby/Coffeelake cores (though with smaller cache size) such as the ever-popular 2c/4t Pentium G4560 and its direct successor the G5400.
Dav1d is already optimize AVX2 (~50% market share)
Now they will begin optimizing for SSS3 and SSE4.1 that all CPU have.
They not know if it will work fine with just this two extensions...
1) You must think that this codec will not be mainstream until there are some HW encoders (6 months to 1 more year for Big Companies have a custom HW encoder)
2) If a Celerons can't decode 1080p with dav1d, the player have two options: serve another codec, or serve the same codec with less resolution.
If you are big enough like youtube you can serve H264 to that HW and save bandwidth with the majority
Or, it's just fine to serve AV1 720p videos to Celerons, and more to the others. (they shouldn't be too picky).
For sure 8k video will be only be serve with AV1 in Youtube, like today VP9 is for >1080p.
Beelzebubu
13th December 2018, 18:18
Ronald Bultje commented on frame parallelism being a bad thing for VP9 (https://blogs.gnome.org/rbultje/2014/02/22/the-worlds-fastest-vp9-decoder-ffvp9/), so not much of a surprise it was turned off by default in AV1/libaom
No, that's a mis-interpretation. Frame parallelism is great.
For encoding, the speedup is slightly better than for tile parallelism, and the quality loss per added thread is less than for tile threading. For example, in my experiments, frame-multithreading in Eve/VP9 costs 0.0% BDRATE loss for a 1.8x speedup going from 1 to 2 threads, but tile threading only gives a 1.7x speedup and has a BDRATE quality loss of around 0.5%. This pattern holds for more threads, and tends to be true across multiple codecs and encoders. Now, obviously, libaom/vpx have no frame threaded encoding so not much to be said there. But in x264, my experience from many years ago is that they switches from slice to frame multi-threaded encoding for the same reason: better scaling *and* less BDRATE quality loss. So far, so good.
OK, next, decoding. This is trickier. For ffh264, for example, we classically found that frame-multithreaded decoding gives a higher speedup than slice-multi-threaded decoding per added thread. Given this pattern of frame multithreading scaling better *and* having less quality loss than within-frame alternatives in a variety of codecs, you'd expect everything to be good, right? Well, not exactly. It holds true, but only to some extend.
The problem in decoding of vp9/av1 is that frames depend on entropy output of the previous frame. For h264/5, cabac state resets in each frame, but this is not true for vp9/av1. So, for frame-multithreading, you need to split decoding in 2 passes, and pass 1 of the next dependent frame can only start when the previous frame finished it's pass 1 and started its pass 2. So, vp9/av1 *decoding* scale less well than h264 *decoding* when using frame multithreading. Fortunately, the system load doesn't go up either, so really what it means is that you need more threads to fully saturate a system. It's even better if you combine frame and tile threading, like what dav1d does.
Wait, you're asking now, what about that statement that frame parallelism is bad in libvpx? Well, it's not what you think it is. --frame-parallel in vpxenc has nothing to do with frame multi-threading in the encoder. It's a header bit that removes the entropy dependency I just talked about. So now, it scales better when using frame multi-threading, which is why this bit is called the "frame parallelism" bit, but it also costs you all backwards entropy, incurring >1% BDRATE quality loss. However, there is no reason to do this. Hardware is not allowed to support higher resolutions with vs. without this feature, and there is no software decoder that implements frame multithreading with but not without entropy dependencies disabled. And if entropy dependencies are present, you can saturate system load anyway by simply using more threads. So the whole thing is kind of silly. Why give up quality for no gain whatsoever?
I can't tell how correct it is, but this was an interesting read: https://codecs.multimedia.cx/2018/12/why-i-am-sceptical-about-av1/
Author is a former libav/ffmpeg developer if you don't remember his name.
Kostya Shishkov.
Beelzebubu
13th December 2018, 18:29
Do we have numbers for the installed base of AVX2 capable PCs? They've been in all new mainstream systems for several years now. I'd guess it's >50% already.
It's around 50%, depending on what statistics you look at. So, I think some people have already tried to address the dav1d performance metrics, so to summarize:
single-threaded, playing 8-bits/component content on >=Haswell (i.e. AVX2=1) will give a 40-80% FPS increase when using dav1d compared to libaom;
multi-threaded, when using the right combinations of frame and tile threading (or just really large numbers) you can get several times higher FPS using dav1d compared to libaom when playing back 8-bit content on Hawell or newer (i.e. AVX2=1);
pre-Haswell (e.g. SSSE3 (https://code.videolan.org/videolan/dav1d/issues/216)), non-x86 (e.g. Neon (https://code.videolan.org/videolan/dav1d/issues/215)), 32bit (the AVX2 assembly is 64-bit only) and 10-bits/component are not yet done. They will not be faster ATM, and possibly significantly slower. We're working on it;
Firefox has a problem (https://bugzilla.mozilla.org/show_bug.cgi?id=1512462) integrating nasm so their version of dav1d has all assembly disabled ATM.
benwaggoner
13th December 2018, 18:48
No, that's a mis-interpretation. Frame parallelism is great.
For encoding, the speedup is slightly better than for tile parallelism, and the quality loss per added thread is less than for tile threading. For example, in my experiments, frame-multithreading in Eve/VP9 costs 0.0% BDRATE loss for a 1.8x speedup going from 1 to 2 threads, but tile threading only gives a 1.7x speedup and has a BDRATE quality loss of around 0.5%. This pattern holds for more threads, and tends to be true across multiple codecs and encoders. Now, obviously, libaom/vpx have no frame threaded encoding so not much to be said there. But in x264, my experience from many years ago is that they switches from slice to frame multi-threaded encoding for the same reason: better scaling *and* less BDRATE quality loss. So far, so good.
Also, content may not be encoded with slices/tiles, but almost certainly will be encoded with hierarchically structured reference frames (like a I P B b structure) where the majority of frames aren't reference frames (e.g. all non-ref b-frames can be decoded in parallel as long as their reference frames are already decoded). So a performant decoder needs to have frame level parallelism, even if it also has slice/tile level as well.
The problem in decoding of vp9/av1 is that frames depend on entropy output of the previous frame. For h264/5, cabac state resets in each frame, but this is not true for vp9/av1. So, for frame-multithreading, you need to split decoding in 2 passes, and pass 1 of the next dependent frame can only start when the previous frame finished it's pass 1 and started its pass 2. So, vp9/av1 *decoding* scale less well than h264 *decoding* when using frame multithreading. Fortunately, the system load doesn't go up either, so really what it means is that you need more threads to fully saturate a system. It's even better if you combine frame and tile threading, like what dav1d does.
So, decoding will be limited by serial decoding of entropy decoding? Do non-reference frames still update and thus serialize the entropy state? If decoding the "bbbb" in an IbbbbBbbbbP" sequence is serialized, that'll really impact decoder parallelization. but if all the non-ref b frames inherit the CABAC state of the most recently decoded reference frame, than it'll be a lot easier.
Beelzebubu
13th December 2018, 19:50
So, decoding will be limited by serial decoding of entropy decoding? Do non-reference frames still update and thus serialize the entropy state? If decoding the "bbbb" in an IbbbbBbbbbP" sequence is serialized, that'll really impact decoder parallelization. but if all the non-ref b frames inherit the CABAC state of the most recently decoded reference frame, than it'll be a lot easier.
Frames with a "similar entropy" reference each other, so a high-level P might use the previous P (which is coded 16 frames back) as its entropy reference, and a non-reference inner B frame (which might not be a reference picture at all for pixel purposes) may actually use the previous inner B-frame (which may well be the one directly before this, or usually 2 and sometimes 3 frames back) as its reference. So this certainly influences how well frame-multithreading scales, not in the worst possible way but not ideal either.
And that's why you see weird things where using 256 instead of 128 threads (I think this is 32/16 frame threads x 8 tile threads) on a 32 core leads to pretty significant speedups (like this (https://medium.com/@ewoutterhoeven/dav1d-0-1-0-release-the-first-benchmarks-5404360e44e3)).
benwaggoner
13th December 2018, 20:00
Frames with a "similar entropy" reference each other, so a high-level P might use the previous P (which is coded 16 frames back) as its entropy reference, and a non-reference inner B frame (which might not be a reference picture at all for pixel purposes) may actually use the previous inner B-frame (which may well be the one directly before this, or usually 2 and sometimes 3 frames back) as its reference. So this certainly influences how well frame-multithreading scales, not in the worst possible way but not ideal either.
And that's why you see weird things where using 256 instead of 128 threads (I think this is 32/16 frame threads x 8 tile threads) on a 32 core leads to pretty significant speedups (like this (https://medium.com/@ewoutterhoeven/dav1d-0-1-0-release-the-first-benchmarks-5404360e44e3)).
Great analysis, thanks!
And huh, I can just imagine the tears of people trying to implement low-cost HW decoders for this. I can see how interframe entropy could provide a percent or two of compression efficiency, though.
I would rather have per-frame entropy and no slice requirement if I had a choice.
Beelzebubu
13th December 2018, 20:05
Great analysis, thanks!
And huh, I can just imagine the tears of people trying to implement low-cost HW decoders for this. I can see how interframe entropy could provide a percent or two of compression efficiency, though.
I would rather have per-frame entropy and no slice requirement if I had a choice.
TBH, from what I understand from people in the relevant committees, this was proposed for HEVC also. The reason they didn't do it had nothing to do with HW, though, but was simply to keep the VoD and RTC use cases technically more similar. (Entropy dependencies are obviously disabled for RTC use cases.)
benwaggoner
13th December 2018, 21:16
TBH, from what I understand from people in the relevant committees, this was proposed for HEVC also. The reason they didn't do it had nothing to do with HW, though, but was simply to keep the VoD and RTC use cases technically more similar. (Entropy dependencies are obviously disabled for RTC use cases.)
For RTC you could have backwards entropy states just fine, I think. So IPPPPPP could have each P reference the entropy state of the previous P. Error correction for lost packets would require trickiness. AV1 RTC would have the same issues.
Limiting entropy state reference to reference frames/tiles would be a lot more robust, but of reduce value. A bunch of non-ref b frames referencing the same frames probably have a lot more in common than any do to the ref-B/P/I frames they reference...
Random access would also be slowed by interframe entropy coding; it's essentially adding another layer of reference dependencies. Entropy is easier to decode, but getting to an arbitrary frame in a long GOP could require decoding the entropy state of a lot more frames than it would with a traditional IbBbP with inter-frame entropy only. With 8 b-frames, getting to an arbitrary frame of H.264/HEVC requires decoding about 1/8th of frames between the IDR and the target frame. Seems like it could be a lot worse in AV1, if I am understanding correctly.
Beelzebubu
13th December 2018, 21:20
Random access would also be slowed by interframe entropy coding; it's essentially adding another layer of reference dependencies. Entropy is easier to decode, but getting to an arbitrary frame in a long GOP could require decoding the entropy state of a lot more frames than it would with a traditional IbBbP with inter-frame entropy only. With 8 b-frames, getting to an arbitrary frame of H.264/HEVC requires decoding about 1/8th of frames between the IDR and the target frame. Seems like it could be a lot worse in AV1, if I am understanding correctly.
Yes, you're correct, random access (seeking) is going to be slower because of this.
mandarinka
14th December 2018, 00:31
In x265, WPP hurts efficiency too. Should we stop using it?
Why do you think the bestest encoders haven't? :cool: Enlightened ones have stropped using frame threading. :devil:
Do we have numbers for the installed base of AVX2 capable PCs? They've been in all new mainstream systems for several years now. I'd guess it's >50% already.
Steam is probably one of the largest datasets available but it is probably quite skewed. It covers disproportionate number of gaming-used computers, but likely almost no HTPCs or office-usage PCs. And all of those are going to watch AV1 video in browsers, even if it is just video ads. So real AVX2 penetration is likely worse than Steam shows, because of the Pentiums/Celerons and the like.
For illustration, look for example at the difference in Windows 10 versus Windows 7 usage shown by general browsing-based statistics sources and by Steam. The former show ~45% for W10 while Steam gives it over 60 %.
SmilingWolf
14th December 2018, 08:35
Why do you think the bestest encoders haven't? :cool: Enlightened ones have stropped using frame threading. :devil:
I am unsure of the meaning of this.
My point was that there is no point in not using either frame threading, WPP (for x265) or tiling (for libaom) when the overhead is not only so low, but even very similar between the two.
Yet I have never seen WPP get the same amount of flack tiling gets, especially considering tile-threading in libdav1d can contribute up to +108% of the decoding performance on its own: https://docs.google.com/spreadsheets/d/1AO3lDZnpC8pNJffOknY1rIxXwLog_ISwHhO_sv3Xlhg/edit#gid=1238661928
Kurosu
14th December 2018, 12:34
Tiles will cause a coding efficiency loss, even if negligible in the big picture. But it is not such a boon either, except for encoders with particular limits, or software decoders. Same for WPP, which really is more a software decoder thing. Contrary to dav1d, your regular HEVC software decoder does not exploit the combined "threadability" of frames and tiles/WPP.
nevcairiel
14th December 2018, 12:39
In the long run, features that allow faster software decoding are really just wasted coding efficiency. When a codec goes mainstream, you'll have a full stack of hardware decoders, which usually don't care that much about these things.
On top of that, if you look at frame threading numbers, the advantage from tile threading shrinks extremely rapidly. Comparing its speed advantage without frame threading is really only a very limited picture.
SmilingWolf
14th December 2018, 13:17
In the long run, features that allow faster software decoding are really just wasted coding efficiency. When a codec goes mainstream, you'll have a full stack of hardware decoders, which usually don't care that much about these things.
On top of that, if you look at frame threading numbers, the advantage from tile threading shrinks extremely rapidly. Comparing its speed advantage without frame threading is really only a very limited picture.
True, and true. I don't even have a retort to that.
I still think that we can care about removing tiling from a libaom encoding workflow whenever the hardware goes mainstream and makes 4K decodable even on budget CPUs like v0lt's Pentium G5600, which should be 2-3 years (?), but I'm ok with the above. Hopefully in the same time rav1e will get proper psy-RD and frame-parallel encoding, too, so we won't have to care about it anyway.
My main heat for the whole tiling debate comes from excluding from early adoption (i.e. right about now) a lot of low-medium tier systems with "inappropriate" encoding settings. In my early tests libdav1d could scale much better on my processor if combined with tiling rather than simply incresing the frame-threads above a certain threshold. Hard to justify a 4MB difference in 1GB of video when said video can't be decoded in real time at all.
Still, the spreadsheet I quoted makes me think I should run the numbers again for dav1d. It has been a couple of months after all.
Mierastor
14th December 2018, 18:37
"Intel: AV1 support not yet in Gen11 Graphics, but coming soon after"
https://www.reddit.com/r/AV1/comments/a5ufft/intel_av1_support_not_yet_in_gen11_graphics_but/
Meaning late 2020, if Intel as usual introduces new CPU generations late in the year?
Since these introductions have often only been paper launches, large-scale availability will only occur in 2021?
nevcairiel
14th December 2018, 19:45
Thats about the time frame most here would expect hardware support. Maybe in 2020, or thereabouts.
Nintendo Maniac 64
14th December 2018, 21:36
But lets be honest here - with AMD finally being a viable alternative again, who is really buying Intel for their graphics capabilities? :p
Motenai Yoda
14th December 2018, 22:14
the ones that don't care about gpu capabilities and still get a display without need a discrete card
nevcairiel
14th December 2018, 22:19
But lets be honest here - with AMD finally being a viable alternative again, who is really buying Intel for their graphics capabilities? :p
Gen11 is also supposed to be significantly faster. And Intel has among the best media capabilities today already, while AMD has the worst.
So for a small form factor media PC, there would be no competition for me.
Nintendo Maniac 64
15th December 2018, 04:16
the ones that don't care about gpu capabilities and still get a display without need a discrete card
Uhhhh... (https://en.wikipedia.org/wiki/List_of_AMD_accelerated_processing_unit_microprocessors)
huhn
17th December 2018, 13:43
you are aware that amd is still missing VP9 for hardware decoding even on vega.
while AMD is very competitive in the CPU market there GPU's are currently at an all time low. vega is using a lot of power needs a huge die and is really slow if you take size into consideration and the hardware decoder is pretty much worse than the nvidia cards that are getting 4 years old.
Mr_Khyron
17th December 2018, 18:53
you are aware that amd is still missing VP9 for hardware decoding even on vega.
while AMD is very competitive in the CPU market there GPU's are currently at an all time low. vega is using a lot of power needs a huge die and is really slow if you take size into consideration and the hardware decoder is pretty much worse than the nvidia cards that are getting 4 years old.
https://en.wikichip.org/w/images/a/a1/vega-whitepaper.pdf
on page 14
Vega” can also decode the VP9 format at resolutions up to 3840x2160 using a hybrid approach where the video and shader engines collaborate to offload work from the CPU.
sneaker_ger
17th December 2018, 19:56
AFAIK the newer AMD ones like Ryzen 5 2500U (Raven Ridge/ Vega 8) now have VP9 10 bit ASIC decoding.
But yeah, they are late. Don't expect it to be different for AV1.
huhn
17th December 2018, 21:39
hybrid decoding has nothing todo with hardware decoding.
letting a CPU and GPU core do the work of an ASIC is simply not the same.
it would be nice if the newest APU have asic decoder but not that trust worth test i found told me the vega 11 is hybrid.
alex1399
18th December 2018, 08:54
No sense, the amd hybrid decoding of vp9 just works as the way that intel do on the hybrid decoding of hevc in their 6th generation processor skylake. They ARE hardware decode.
Blue_MiSfit
18th December 2018, 19:33
Keep things on topic, please.
SmilingWolf
18th December 2018, 22:31
Status report, redux
"rav1e is doing well" edition
1st edition: http://forum.doom9.org/showthread.php?p=1852449#post1852449
2nd edition: http://forum.doom9.org/showthread.php?p=1857587#post1857587
Whatever paragraph I don't repeat here can be assumed to be the same as in the aforementioned posts
First of all: graphs! Click to enlarge
Y axis: chosen metric
X axis: bits per pixel
720p:
https://i.ibb.co/ZJ8GZVw/msssim-720.png (https://ibb.co/ZJ8GZVw) https://i.ibb.co/fVWNjjy/psnrhvsm-720.png (https://ibb.co/fVWNjjy) https://i.ibb.co/6ZQ6sTn/hvmaf-720.png (https://ibb.co/6ZQ6sTn)
1080p:
https://i.ibb.co/3dKr5f5/msssim-1080.png (https://ibb.co/3dKr5f5) https://i.ibb.co/4NNHxCC/psnrhvsm-1080.png (https://ibb.co/4NNHxCC) https://i.ibb.co/nDTK5y5/hvmaf-1080.png (https://ibb.co/nDTK5y5)
Encoders improvement over time:
720p:
https://i.ibb.co/R0dQbtc/msssim-overtime-720.png (https://ibb.co/R0dQbtc) https://i.ibb.co/9rJQXbQ/psnrhvsm-overtime-720.png (https://ibb.co/9rJQXbQ)
1080p:
https://i.ibb.co/Y76p6Tp/msssim-overtime-1080.png (https://ibb.co/Y76p6Tp) https://i.ibb.co/WBhV3rN/psnrhvsm-overtime-1080.png (https://ibb.co/WBhV3rN)
BD rates for 720p:
Codecs ladder: | x264 relative:
x264 -> rav1e | x264 -> rav1e
RATE (%) DSNR (dB) | RATE (%) DSNR (dB)
MSSSIM -10.7345 0.541324 | MSSSIM -10.7345 0.541324
PSNRHVS -15.3271 1.07245 | PSNRHVS -15.3271 1.07245
HVMAF -7.40703 2.07138 | HVMAF -7.40703 2.07138
----------------------------|-----------------------------
rav1e -> vp9 | x264 -> vp9
RATE (%) DSNR (dB) | RATE (%) DSNR (dB)
MSSSIM -15.1057 0.68453 | MSSSIM -21.81 1.08927
PSNRHVS -11.2436 0.654976 | PSNRHVS -22.9586 1.51188
HVMAF -19.2883 2.57019 | HVMAF -24.1102 3.58993
----------------------------|-----------------------------
vp9 -> x265 | x264 -> x265
RATE (%) DSNR (dB) | RATE (%) DSNR (dB)
MSSSIM -4.25723 0.169151 | MSSSIM -25.6195 1.21115
PSNRHVS -8.19042 0.41409 | PSNRHVS -29.8289 1.83058
HVMAF -10.6714 0.708441 | HVMAF -31.2046 4.45371
----------------------------|-----------------------------
x265 -> av1 | x264 -> av1
RATE (%) DSNR (dB) | RATE (%) DSNR (dB)
MSSSIM -18.9088 0.7852 | MSSSIM -38.0511 1.97653
PSNRHVS -15.3123 0.761791 | PSNRHVS -38.6659 2.56119
HVMAF -18.0023 1.0489 | HVMAF -44.0411 4.3982
BD rates for 1080p:
Codecs ladder: | x264 relative:
x264 -> rav1e | x264 -> rav1e
RATE (%) DSNR (dB) | RATE (%) DSNR (dB)
MSSSIM -20.0683 1.00011 | MSSSIM -20.0683 1.00011
PSNRHVS -21.9935 1.47903 | PSNRHVS -21.9935 1.47903
HVMAF -18.4773 3.96202 | HVMAF -18.4773 3.96202
------------------------------|-----------------------------
rav1e -> vp9 | x264 -> vp9
RATE (%) DSNR (dB) | RATE (%) DSNR (dB)
MSSSIM -18.2653 0.729489 | MSSSIM -31.1754 1.49143
PSNRHVS -14.922 0.755605 | PSNRHVS -30.1275 1.87845
HVMAF -20.1645 2.31195 | HVMAF -32.4505 4.72978
------------------------------|-----------------------------
vp9 -> x265 | x264 -> x265
RATE (%) DSNR (dB) | RATE (%) DSNR (dB)
MSSSIM 5.13717 -0.177855 | MSSSIM -28.9956 1.18206
PSNRHVS -0.096748 -0.0123981 | PSNRHVS -31.474 1.63676
HVMAF -3.78107 0.0881882 | HVMAF -34.6185 4.22357
------------------------------|-----------------------------
x265 -> av1 | x264 -> av1
RATE (%) DSNR (dB) | RATE (%) DSNR (dB)
MSSSIM -26.486 0.938124 | MSSSIM -45.4959 2.12535
PSNRHVS -21.7431 0.905916 | PSNRHVS -43.4792 2.56047
HVMAF -22.7091 1.17861 | HVMAF -48.0404 4.69582
Encoders:
x264 157-2935-545de2f
x265 2.9-4-471726d3a046
rav1e 0.1.0-977-64b9f501
libaom 1.0.0-908-g3a607f7b0
libvpx 1.7.0-1352-gea57f9acd
Cmdlines:
x264 --preset veryslow --tune ssim --crf 16 -o test.x264.crf16.264 orig.i420.y4m
x265 --preset veryslow --tune ssim --crf 16 -o test.x265.crf16.hevc orig.i420.y4m
rav1e --low_latency false -o test.rav1e.cq80.ivf --quantizer 80 -s 2 --tune psnr orig.i420.y4m
aomenc --frame-parallel=0 --tile-columns=3 --auto-alt-ref=1 --cpu-used=4 --tune=psnr --passes=2 --threads=2 --end-usage=q --cq-level=20 --test-decode=fatal -o test.av1.cq20.webm orig.i420.y4m
vpxenc --codec=vp9 --frame-parallel=0 --tile-columns=2 --good --cpu-used=0 --tune=psnr --passes=2 --threads=2 --end-usage=q --cq-level=20 --test-decode=fatal --ivf -o test.vp9.cq20.ivf orig.i420.y4m
Quality settings:
x264: CRF 16-24 step 1, and 24-34 step 2
x265: CRF 16-24 step 1, and 24-34 step 2
rav1e: CQ 80-160 step 16
aomenc: CQ 20-40 step 4
vpxenc: CQ 20-48 step 4
VMAF: model used: nflxall_vmafv4, pooling: harmonic_mean
Notes:
Revisiting the sequences from the previous report:
F.Y.C, x264 -> rav1e
RATE (%) DSNR (dB)
MSSSIM -21.8223 1.2258
PSNRHVS -28.408 2.29056
HVMAF -18.7682 3.89782
PresageFlowerFight, x264 -> rav1e
RATE (%) DSNR (dB)
MSSSIM -33.8353 1.95641
PSNRHVS -33.5774 2.47118
HVMAF -29.2079 7.17318
PresageFlowerWalk, x264 -> rav1e
RATE (%) DSNR (dB)
MSSSIM -6.98635 0.333085
PSNRHVS -3.78743 0.244148
HVMAF 2.23865 -0.150589
MAJOR improvements in static/very low motion scenes thanks to adaptive keyframe selection (https://github.com/xiph/rav1e/commit/869fef7002e7a7595bf6891228bd58d70a26a670), marginal improvements in the others.
Phanton_13
19th December 2018, 00:36
SmilingWolf, can you also indicates the encoding speed, it don't need to be in fps as it can be in relation of a specific encoder, is more to see the variation in the speed of aoemc and rav1e due to optimizations or other improvements.
utack
19th December 2018, 05:26
The cpu-used heuristics in libaom seems to be poorly tuned in lossless mode.
Tested with a digitally animated gif (https://mir-s3-cdn-cf.behance.net/project_modules/max_1200/90d83d73226609.5c0277698d016.gif)
hitting 50% size increase at 1/4 encode time saved compared to cpu-used=0, and double the size at 60% the encode time of cpu-used=0
each dot presets one cpu-used value
https://i.imgur.com/RR3IcEb.png
Nintendo Maniac 64
19th December 2018, 07:31
digitally animated gif
If you ever want to test higher-quality animation examples (read: not limited to 256 colors), then perhaps try an animated PNG.
SmilingWolf
19th December 2018, 11:06
SmilingWolf, can you also indicates the encoding speed, it don't need to be in fps as it can be in relation of a specific encoder, is more to see the variation in the speed of aoemc and rav1e due to optimizations or other improvements.
I use my PC under various loads while encoding, so I can't reliably measure times, which is why I haven't included anything time specific so far.
A quick test on the F.Y.C clip however has displayed an improvement, on average, of 12% over the last month (1.0.0-1058-g8547359cf vs 1.0.0-908-g3a607f7b0).
At the same time, quality as measured by MS-SSIM and PSNR-HVS-M has gone down ~2.5%, circa to the levels of 3 months ago (1.0.0-577-g8ae39302e).
As for rav1e, the times have only grown worse because there have been many new additions but almost no early breakout (or none at all?) strategies have been implemented yet.
And as of right now, some ASM optimizations are disabled (https://github.com/xiph/rav1e/blob/f0e759175f52a1dae4233ff528557fcf0e3c8319/src/me.rs#L10) on Windows (https://github.com/xiph/rav1e/blob/f0e759175f52a1dae4233ff528557fcf0e3c8319/src/predict.rs#L157), so I guess I'm basically running on Rust code. So not much of a speedup on that front either (yet).
Phanton_13
20th December 2018, 14:52
Thanks SmilingWolf for the info.
sneaker_ger
22nd December 2018, 11:31
Simply crashes on my system even with a simple "ffmpeg" (no parameters). Windows 7 x64, i5-2500K (SSE4.2 and AVX).
sneaker_ger
22nd December 2018, 17:47
Still crashes.
lvqcl
22nd December 2018, 18:40
My CPU doesn't support SSE4.2, only 4.1, but I still tried to run it.
ffmpeg.exe crashes at shlx instruction, which is part of BMI2 (Haswell/Excavator).
nevcairiel
22nd December 2018, 19:11
Its usually not a good idea to build such a restricted binary, nor does it really give a meaningful speed enhancement. Just use a better build.
I often intentionally restrict what I allow compilers to do, because even GCC 8 is still terrible at using advanced instruction set instructions, and it can and will cause issues left and right.
This is especially true in a large and old code base like ffmpeg, which has code that was written 15 years ago, and code that as written just now, code that was painstakingly hand-optimized, and whatnot which can result in quite varying stress on the compiler.
Wolfberry
22nd December 2018, 22:01
I think the crash comes from that I configured gcc to have skylake by default, forget to use the --cpu switch, will rebuild them later.
Thanks nevcairiel.
SmilingWolf
23rd December 2018, 01:26
GCC 8.2.1 20181221, static build
ffmpeg-4.2-92779-g8b53d1322f: https://mega.nz/#!5kJxjA7I!dHKMhcYcjQPVZZWEBVTkzgNpOxT5hyylOEHTDtESTYM
- libaom-av1 1.0.0-1103-g9a48f9ca5
- libdav1d 0.1.0-38-g1703f21
SmilingWolf
23rd December 2018, 13:22
My AVIF toolkit: https://mega.nz/#!5oQE2Sob!STZHdk4ob4ptHknMvNcB4JxbCt9xdu3WUKkg7iyh2EM
Needs MSYS2. It's mostly an hack job.
Don't mess with the directory structure. Images to convert to AVIF in "images", AVIF to convert to PNG in "contained".
Due to stuff right now the script assumes all the AVIF images follow the BT709 matrix for YUV -> RGB conversion.
Usage:
encode.MT.sh takes in input a quality for aomenc
encode.ST.sh takes in input an image path and a quality setting
decode.MT.sh takes no arguments
decode.ST.sh takes the AVIF file path
ST and MT are for Single Thread and Multi Thread, respectively. Would be more correct to say multi-process but you get the idea. The default is 6 processes, you can tweak it by modifying xargs' -P option in the MT script.
The defaults are for high quality (--cpu-used=0 for extra overkill), so it might take a while to convert everything depending on your machine. You can tweak aomenc's options in encode.ST.sh
Also included a python script to gather the metrics' weighted average from a given stats file (auto generated by encode.MT.sh)
SmilingWolf
23rd December 2018, 17:24
zscale (libzimg) should be preferred (personal opinion)
While I do use libzimg from time to time, it has a tendency to randomly crash (esp. when downscaling video), so I don't want it into an automated workflow.
For the mod2 thing: setting YUV444 just for that looks way too wasteful unless you already have to preserve zones of high contrast like in the DDMC.png example. So maybe something like scale=-2:ih to set the width to a multiple of 2?
lvqcl
23rd December 2018, 17:47
I decided to test how my old CPU (Intel Core2 Quad Q9300, SSE4.1) decodes AV1 using ffmpeg from SmilingWolf build.
Test video: https://www.youtube.com/watch?v=PiWyCQV52h0 , 1280x720p.
-c:v libaom-av1: 1.89x realtime, utime = 94 sec.
-c:v libdav1d: 1.30x realtime, utime = 158 sec.
-c:v libdav1d -threads 4 -tilethreads 4: 1.67x realtime, utime = 159 sec.
-c:v libdav1d -threads 8 -tilethreads 8: 1.74x realtime, utime = 163 sec.
-c:v libdav1d -threads 16 -tilethreads 16: 2.02x realtime, utime = 161 sec.
(I hava no idea what threads and tilethreads options do, so I just tested various values for them)
So, on my system dav1d requires ~160/94=1.7 times more CPU time than aom.
MoSal
24th December 2018, 12:16
(I hava no idea what threads and tilethreads options do, so I just tested various values for them)
Try with tilethreads set to 1.
lvqcl
24th December 2018, 18:01
-c:v libdav1d: 1.31x realtime, utime = 156 sec.
-c:v libdav1d -threads 4 -tilethreads 1: 1.31x realtime, utime = 157 sec.
-c:v libdav1d -threads 8 -tilethreads 1: 1.61x realtime, utime = 158 sec.
-c:v libdav1d -threads 16 -tilethreads 1: 1.80x realtime, utime = 159 sec.
-c:v libdav1d -threads 32 -tilethreads 1: 1.98x realtime, utime = 160 sec.
v0lt
24th December 2018, 18:22
@lvqcl
What is "utime"? This is not like decoding time.
NikosD
24th December 2018, 20:20
What's the progress of dav1d leveraging SSSE3 assembly optimizations ?
Are we still based on AVX2 only for dav1d ?
lvqcl
24th December 2018, 20:37
ffmpeg prints something like this:
bench: utime=178.762s stime=2.839s rtime=58.362s
IIUC:
utime = total time spent on user code (across all CPU cores)
stime = total time spent on system code
rtime = "real time" aka wall time
So: it took 58.362 seconds to decode a video, but CPU time spent on decoding was 178.762+2.839 sec.
That is, (178.762+2.839)/58.362 = 3.1 cores were active (on average) during decoding.
Wolfberry
25th December 2018, 01:55
@NikosD
SSSE3: issue #216 (https://code.videolan.org/videolan/dav1d/issues/216) (7 / 28)
AVX2: issue #78 (https://code.videolan.org/videolan/dav1d/issues/78) (9 / 52)
NikosD
25th December 2018, 08:46
So, AVX2 is missing only 4:4:4 and SVC but SSSE3 is missing everything (almost) for 8bit.
Thank you!
utack
25th December 2018, 19:43
Did a quick test for some typical "sent by phone video". Shot on a phone, 30s, medium resolution and bitrate
x264 crf 26 and "placebo" to get a bitrate estimate for medium-poor quality, libaom cpu-used=3 in 2pass mode to match the bitrate and compare
Turns out for this medium resolution (720p), and with a lot of "high frequency" motion (water waves and grass) x264 is still extremely competitive, and imho even beats libaom here in 1/3 screenshots
http://screenshotcomparison.com/comparison/126513
hajj_3
29th December 2018, 11:28
It looks like some new SSSE3 optimisations for Dav1d have been submitted: https://code.videolan.org/videolan/dav1d/commit/9ea56386dee2706d94f3c2dac1720bcf4961aaba
hajj_3
30th December 2018, 12:04
what AV1 decoder does the latest MPC-BE x64 use (v1.5.3 4246 beta)? When playing a 720p 25fps AV1 video it uses up to 61% cpu on my kabylake i3-7100u, using vlc player 3.0.5 (which uses dav1d) it uses up to 35% playing the same video.
v0lt
30th December 2018, 20:12
@hajj_3
MPC-BE used libaom git-v1.0.0-748-g8048e8c0b.
https://sourceforge.net/p/mpcbe/code/HEAD/tree/trunk/lib64/
marcomsousa
3rd January 2019, 09:16
@hajj_3
MPC-BE used libaom git-v1.0.0-748-g8048e8c0b.
https://sourceforge.net/p/mpcbe/code/HEAD/tree/trunk/lib64/
Just update to libaom git-v1.0.0-1116-g00c80e6b5 (3 hours ago) next build will be updated.
benwaggoner
4th January 2019, 20:45
Did a quick test for some typical "sent by phone video". Shot on a phone, 30s, medium resolution and bitrate
x264 crf 26 and "placebo" to get a bitrate estimate for medium-poor quality, libaom cpu-used=3 in 2pass mode to match the bitrate and compare
Turns out for this medium resolution (720p), and with a lot of "high frequency" motion (water waves and grass) x264 is still extremely competitive, and imho even beats libaom here
Water waves and grass are really hard to encode, and classic per-frame PNSR or SAD style optimization don't yield good results. There's a lot of psychovisual tuning to keep the motion looking natural without getting block-based basis pattern leaking in. And a lot of rate control to keep a part of a frame with that content looking good without sucking all the bits away from the rest of the frame and making them look bad.
That's the kind of stuff that comes from a mature encoder with lots of psychovisual tweaks. Which defines x264 in spades, and which x265 inherited a lot of. The real-world performance of those encoders has more to do with the foundational legacy of loving obsessive attention from quality @ bitrate obsessed video pirates than any particular underlying bitstream features. I bet an x262 could have outperformed any MPEG-2 encoder for anime DVDs, for example.
utack
4th January 2019, 23:13
Water waves and grass are really hard to encode, and classic per-frame PNSR or SAD style optimization don't yield good results. There's a lot of psychovisual tuning to keep the motion looking natural without getting block-based basis pattern leaking in. And a lot of rate control to keep a part of a frame with that content looking good without sucking all the bits away from the rest of the frame and making them look bad.
That's the kind of stuff that comes from a mature encoder with lots of psychovisual tweaks. Which defines x264 in spades, and which x265 inherited a lot of. The real-world performance of those encoders has more to do with the foundational legacy of loving obsessive attention from quality @ bitrate obsessed video pirates than any particular underlying bitstream features. I bet an x262 could have outperformed any MPEG-2 encoder for anime DVDs, for example.
Thanks for your insight.
Would you attribute this mostly to excellent psychovisual tuning or are video streams of small dimensions with a lot of motion and 4x4 blocks areas where AV1 might always be much better even in a theoretical best case scenario?
benwaggoner
5th January 2019, 21:10
Thanks for your insight.
Would you attribute this mostly to excellent psychovisual tuning or are video streams of small dimensions with a lot of motion and 4x4 blocks areas where AV1 might always be much better even in a theoretical best case scenario?
H.264 and HEVC both have 4x4 blocks as well, so that feature alone isn’t going to be make-or-break.
As for comparing formats, the codec specs are like what you have in your fridge. The encoder is like the cook. A great cook can make simple ingredients into something wonderful, and a terrible cook can make a disaster out of the best ingredients. A great cook with a wide variety of great ingrediants is what gives the optimal results.
In comparing codecs, all we really can compare is the dishes that come out of the kitchen, though. Is a meal great or bad due to the cook or the ingredients? It’s hard to say and involves a lot of educated guesses and speculation.
For example, x264 with —preset placebo —tune film is probably going to produce better quality @ bitrate with typical content than libaom at its absolute fastest settings. It’s really quality @ perf @ bitrate, and that’s controlled by encoder optimization even more than the bitstream format. Stuff like AVX2 optimization will produce better quality within a given bitrate @ perf, because more options get tried and tools get used. And that’s with absolutely no change to psychovisual tuning or bitstream. It’s just the same results, faster.
Of course even that can be impacted by bitstream details. The bigger block sizes of HEVC mean that AVX2 and AVX512 offer bigger gains than with x264. Even choice of processor can change relative quality @ perf @ bitrate, as differenr encoders make better or worse use of lots of cores or more advanced SIMD.
We can only really know how “good” HEVC or AV1 or VVC are based on the best avaialable encoder for a given use case. And that can be hard to predict. Certainly cable MPEG-2 is a lot more efficient today than anyone predicted or could demonstrate when MPEG-2’s spec was finished.
hajj_3
7th January 2019, 13:08
http://www.streamingmedia.com/Articles/News/Online-Video-News/Unified-Patents-Challenges-Velos-Media-Patent-128870.aspx
LigH
7th January 2019, 13:18
I was afraid to click an uncommented URL ... but did anyway.
This article is about HEVC licensing. Only marginally related to AOM.
TomV
16th January 2019, 01:42
Hey AV1 experts... can you share your command line for your highest subjective quality encodes? In other words, what's your recommended settings for the equivalent of --preset veryslow?
Tommy Carrot
16th January 2019, 12:57
The preset equivalent for aomenc is --cpu-used. Contrary to what would be logical, it has nothing to do with threading. --cpu-used=0 is the equivalent of preset placebo, while 8 is fastest (more precisely the least slow) preset. I would not go over 6, after that the quality regression is really noticeable.
utack
16th January 2019, 13:24
I would not go over 6, after that the quality regression is really noticeable.
What kind of SSIM difference is noticeable?
Speed 1, 5 and 8 (https://www.arewecompressedyet.com/?job=av1_sp1_010819%402019-01-11T15%3A41%3A56.697Z&job=av1_sp5_011019%402019-01-11T20%3A49%3A41.789Z&job=av1_sp8_011019%402019-01-11T19%3A26%3A13.348Z)
Tommy Carrot
16th January 2019, 16:11
I'm talking about visual quality. --cpu-used 7 and 8 are looking significantly worse than 6, while the encoding speed doesn't improve that much.
benwaggoner
16th January 2019, 18:51
Hey AV1 experts... can you share your command line for your highest subjective quality encodes? In other words, what's your recommended settings for the equivalent of --preset veryslow?
Probably more like —preset very slow —tune film (or grain, or animation). The psychovisual tuning that’s implicit in x264/5’s RF model, and explicit in —tune and other parameters, are an important component.
benwaggoner
16th January 2019, 18:55
What kind of SSIM difference is noticeable?
Speed 1, 5 and 8 (https://www.arewecompressedyet.com/?job=av1_sp1_010819%402019-01-11T15%3A41%3A56.697Z&job=av1_sp5_011019%402019-01-11T20%3A49%3A41.789Z&job=av1_sp8_011019%402019-01-11T19%3A26%3A13.348Z)
Mean of per frame SSIM is not a great metric, honestly. The mean of any per-frame metric isn’t going to catch variability of quality, and any spatial-only metric will be bad at catching frame strobing or other visual discontinuities.
The latest VMAF is the least-bad metric available, by a good margin. But even it isn’t very useful at comparing between high quality encodes and some kinds of artifacts.
Encoding history is littered with encoders with promising PSNR and SSIM scores that just didn’t look very good.
TomV
17th January 2019, 17:14
The preset equivalent for aomenc is --cpu-used. Contrary to what would be logical, it has nothing to do with threading. --cpu-used=0 is the equivalent of preset placebo, while 8 is fastest (more precisely the least slow) preset. I would not go over 6, after that the quality regression is really noticeable.
Yes, I'm aware of that. I've been using --cpu-used=1 Thanks.
TomV
17th January 2019, 17:20
Probably more like —preset very slow —tune film (or grain, or animation). The psychovisual tuning that’s implicit in x264/5’s RF model, and explicit in —tune and other parameters, are an important component.
--preset is not an option (I get an error message)
There is a --tune option (psnr, ssim, cdef-dist, daala-dist), and a --tune-content option (default, screen)
Here's the help file. It seems like every option that would improve visual quality is on by default.
C:\Test>aomenc --help
Usage: aomenc <options> -o dst_filename src_filename
Options:
--help Show usage options and exit
-c <arg>, --cfg=<arg> Config file to use
-D, --debug Debug mode (makes output deterministic)
-o <arg>, --output=<arg> Output filename
--codec=<arg> Codec to use
-p <arg>, --passes=<arg> Number of passes (1/2)
--pass=<arg> Pass to execute (1/2)
--fpf=<arg> First pass statistics file name
--limit=<arg> Stop encoding after n input frames
--skip=<arg> Skip the first n input frames
--good Use Good Quality Deadline
-q, --quiet Do not print encode progress
-v, --verbose Show encoder parameters
--psnr Show PSNR in status line
--webm Output WebM (default when WebM IO is enabled)
--ivf Output IVF
--obu Output OBU
--q-hist=<arg> Show quantizer histogram (n-buckets)
--rate-hist=<arg> Show rate histogram (n-buckets)
--disable-warnings Disable warnings about potentially incorrect encode settings.
-y, --disable-warning-prompt Display warnings, but do not prompt user to continue.
--test-decode=<arg> Test encode/decode mismatch
off, fatal, warn
Encoder Global Options:
--yv12 Input file is YV12
--i420 Input file is I420 (default)
--i422 Input file is I422
--i444 Input file is I444
-u <arg>, --usage=<arg> Usage profile number to use
-t <arg>, --threads=<arg> Max number of threads to use
--profile=<arg> Bitstream profile number to use
-w <arg>, --width=<arg> Frame width
-h <arg>, --height=<arg> Frame height
--forced_max_frame_width Maximum frame width value to force
--forced_max_frame_height Maximum frame height value to force
--stereo-mode=<arg> Stereo 3D video format
mono, left-right, bottom-top, top-bottom, right-left
--timebase=<arg> Output timestamp precision (fractional seconds)
--fps=<arg> Stream frame rate (rate/scale)
--global-error-resilient=< Enable global error resiliency features
-b <arg>, --bit-depth=<arg> Bit depth for codec (8 for version <=1, 10 or 12 for version 2)
8, 10, 12
--lag-in-frames=<arg> Max number of frames to lag
--large-scale-tile=<arg> Large scale tile coding (0: off (default), 1: on)
--monochrome Monochrome video (no chroma planes)
--full-still-picture-hdr Use full header for still picture
Rate Control Options:
--drop-frame=<arg> Temporal resampling threshold (buf %)
--resize-mode=<arg> Frame resize mode
--resize-denominator=<arg> Frame resize denominator
--resize-kf-denominator=<a Frame resize keyframe denominator
--superres-mode=<arg> Frame super-resolution mode
--superres-denominator=<ar Frame super-resolution denominator
--superres-kf-denominator= Frame super-resolution keyframe denominator
--superres-qthresh=<arg> Frame super-resolution qindex threshold
--superres-kf-qthresh=<arg Frame super-resolution keyframe qindex threshold
--end-usage=<arg> Rate control mode
vbr, cbr, cq, q
--target-bitrate=<arg> Bitrate (kbps)
--min-q=<arg> Minimum (best) quantizer
--max-q=<arg> Maximum (worst) quantizer
--undershoot-pct=<arg> Datarate undershoot (min) target (%)
--overshoot-pct=<arg> Datarate overshoot (max) target (%)
--buf-sz=<arg> Client buffer size (ms)
--buf-initial-sz=<arg> Client initial buffer size (ms)
--buf-optimal-sz=<arg> Client optimal buffer size (ms)
Twopass Rate Control Options:
--bias-pct=<arg> CBR/VBR bias (0=CBR, 100=VBR)
--minsection-pct=<arg> GOP min bitrate (% of target)
--maxsection-pct=<arg> GOP max bitrate (% of target)
Keyframe Placement Options:
--enable-fwd-kf=<arg> Enable forward reference keyframes
--kf-min-dist=<arg> Minimum keyframe interval (frames)
--kf-max-dist=<arg> Maximum keyframe interval (frames)
--disable-kf Disable keyframe placement
AV1 Specific Options:
--cpu-used=<arg> CPU Used (0..8)
--auto-alt-ref=<arg> Enable automatic alt reference frames
--sharpness=<arg> Loop filter sharpness (0..7)
--static-thresh=<arg> Motion detection threshold
--row-mt=<arg> Enable row based multi-threading (0: off, 1: on (default))
--tile-columns=<arg> Number of tile columns to use, log2
--tile-rows=<arg> Number of tile rows to use, log2
--enable-tpl-model=<arg> RDO modulation based on frame temporal dependency
--arnr-maxframes=<arg> AltRef max frames (0..15)
--arnr-strength=<arg> AltRef filter strength (0..6)
--tune=<arg> Distortion metric tuned with
psnr, ssim, cdef-dist, daala-dist
--cq-level=<arg> Constant/Constrained Quality level
--max-intra-rate=<arg> Max I-frame bitrate (pct)
--max-inter-rate=<arg> Max P-frame bitrate (pct)
--gf-cbr-boost=<arg> Boost for Golden Frame in CBR mode (pct)
--lossless=<arg> Lossless mode (0: false (default), 1: true)
--enable-cdef=<arg> Enable the constrained directional enhancement filter (0: false, 1: true (default))
--enable-restoration=<arg> Enable the loop restoration filter (0: false, 1: true (default))
--enable-rect-partitions=< Enable rectangular partitions (0: false, 1: true (default))
--enable-dual-filter=<arg> Enable dual filter (0: false, 1: true (default))
--enable-intra-edge-filter Enable intra edge filtering (0: false, 1: true (default))
--enable-order-hint=<arg> Enable order hint (0: false, 1: true (default))
--enable-tx64=<arg> Enable 64-pt transform (0: false, 1: true (default))
--enable-dist-wtd-comp=<ar Enable distance-weighted compound (0: false, 1: true (default))
--enable-masked-comp=<arg> Enable masked (wedge/diff-wtd) compound (0: false, 1: true (default))
--enable-interintra-comp=< Enable interintra compound (0: false, 1: true (default))
--enable-smooth-interintra Enable smooth interintra mode (0: false, 1: true (default))
--enable-diff-wtd-comp=<ar Enable difference-weighted compound (0: false, 1: true (default))
--enable-interinter-wedge= Enable interinter wedge compound (0: false, 1: true (default))
--enable-interintra-wedge= Enable interintra wedge compound (0: false, 1: true (default))
--enable-global-motion=<ar Enable global motion (0: false, 1: true (default))
--enable-warped-motion=<ar Enable local warped motion (0: false, 1: true (default))
--enable-filter-intra=<arg Enable filter intra prediction mode (0: false, 1: true (default))
--enable-smooth-intra=<arg Enable smooth intra prediction modes (0: false, 1: true (default))
--enable-paeth-intra=<arg> Enable Paeth intra prediction mode (0: false, 1: true (default))
--enable-cfl-intra=<arg> Enable chroma from luma intra prediction mode (0: false, 1: true (default))
--enable-obmc=<arg> Enable OBMC (0: false, 1: true (default))
--enable-palette=<arg> Enable palette prediction mode (0: false, 1: true (default))
--enable-intrabc=<arg> Enable intra block copy prediction mode (0: false, 1: true (default))
--enable-angle-delta=<arg> Enable intra angle delta (0: false, 1: true (default))
--disable-trellis-quant=<a Disable trellis optimization of quantized coefficients (0: false (default) 1: true)
--enable-qm=<arg> Enable quantisation matrices (0: false (default), 1: true)
--qm-min=<arg> Min quant matrix flatness (0..15), default is 8
--qm-max=<arg> Max quant matrix flatness (0..15), default is 15
--reduced-tx-type-set=<arg Use reduced set of transform types
--frame-parallel=<arg> Enable frame parallel decodability features (0: false (default), 1: true)
--error-resilient=<arg> Enable error resilient features (0: false (default), 1: true)
--aq-mode=<arg> Adaptive quantization mode (0: off (default), 1: variance 2: complexity, 3: cyclic refresh)
--deltaq-mode=<arg> Delta qindex mode (0: off (default), 1: deltaq 2: deltaq + deltalf)
--frame-boost=<arg> Enable frame periodic boost (0: off (default), 1: on)
--noise-sensitivity=<arg> Noise sensitivity (frames to blur)
--tune-content=<arg> Tune content type
default, screen
--cdf-update-mode=<arg> CDF update mode for entropy coding (0: no CDF update; 1: update CDF on all frames(default); 2: selectively update CDF on some frames
--color-primaries=<arg> Color primaries (CICP) of input content:
bt709, unspecified, bt601, bt470m, bt470bg, smpte240, film, bt2020, xyz, smpte431, smpte432, ebu3213
--transfer-characteristics Transfer characteristics (CICP) of input content:
unspecified, bt709, bt470m, bt470bg, bt601, smpte240, lin, log100, log100sq10, iec61966, bt1361, srgb, bt2020-10bit, bt2020-12bit, smpte2084, hlg, smpte428
--matrix-coefficients=<arg Matrix coefficients (CICP) of input content:
identity, bt709, unspecified, fcc73, bt470bg, bt601, smpte240, ycgco, bt2020ncl, bt2020cl, smpte2085, chromncl, chromcl, ictcp
--chroma-sample-position=< The chroma sample position when chroma 4:2:0 is signaled:
unknown, vertical, colocated
--min-gf-interval=<arg> min gf/arf frame interval (default 0, indicating in-built behavior)
--max-gf-interval=<arg> max gf/arf frame interval (default 0, indicating in-built behavior)
--gf-max-pyr-height=<arg> maximum height for GF group pyramid structure (1 to 4 (default))
--sb-size=<arg> Superblock size to use
dynamic, 64, 128
--num-tile-groups=<arg> Maximum number of tile groups, default is 1
--mtu-size=<arg> MTU size for a tile group, default is 0 (no MTU targeting), overrides maximum number of tile groups
--timing-info=<arg> Signal timing info in the bitstream (model unly works for no hidden frames, no super-res yet):
unspecified, constant, model
--film-grain-test=<arg> Film grain test vectors (0: none (default), 1: test-1 2: test-2, ... 16: test-16)
--film-grain-table=<arg> Path to file containing film grain parameters
--denoise-noise-level=<arg Amount of noise (from 0 = don't denoise, to 50)
--denoise-block-size=<arg> Denoise block size (default = 32)
--enable-ref-frame-mvs=<ar Enable temporal mv prediction (default is 1)
-b <arg>, --bit-depth=<arg> Bit depth for codec (8 for version <=1, 10 or 12 for version 2)
8, 10, 12
--input-bit-depth=<arg> Bit depth of input
--input-chroma-subsampling chroma subsampling x value.
--input-chroma-subsampling chroma subsampling y value.
--sframe-dist=<arg> S-Frame interval (frames)
--sframe-mode=<arg> S-Frame insertion mode (1..2)
--annexb=<arg> Save as Annex-B
Stream timebase (--timebase):
The desired precision of timestamps in the output, expressed
in fractional seconds. Default is 1/1000.
Included encoders:
av1 - AOMedia Project AV1 Encoder 0.1.0-11038-g437d957f8 (default)
Use --codec to switch to a non-default encoder.
benwaggoner
17th January 2019, 19:43
--preset is not an option (I get an error message)
There is a --tune option (psnr, ssim, cdef-dist, daala-dist), and a --tune-content option (default, screen)
Sorry, I wasn't clear I was offering x264/5 syntax for what you want an AV1 equivalent to. Which does not (yet?) exist in libaom.
Libaom has very little psychovisual optimization compared to the x26? codecs, or psychovisual tuning options. Libaom is mare like a speed-optimized reference encoder right now. Production AV1 encoders will need to have much more psychovisual tuning.
I'm kinda surprised no one has started making a xAV1 based on x264 like x265 started. Although so many fundamental structures derived from the VPx series would make it a lot harder to start. HEVC was enough of a H.264 subset that getting an x265 that did SOMETHING wasn't THAT hard. AV1-as-an-ecosystem has a serious disadvantage in not having a well-tuned, production-grade, open-source VPx encoder to start from.
Libvpx never got the obsessive focus from thousands of quality/efficiency/speed obsessed video pirates that was the foundation of what x264 is today.
Just look at the H.264, or even the HEVC, forums here, and all the posts over many years. That's the kind of community focus that makes for a great encoder.
Look at the "Posts" column:
https://forum.doom9.org/forumdisplay.php?f=17
TomV
17th January 2019, 21:25
Sorry, I wasn't clear I was offering x264/5 syntax for what you want an AV1 equivalent to. Which does not (yet?) exist in libaom.
Oh. That was my original question... "what's your recommended settings for the equivalent of --preset veryslow?" Of course I'm intimately familiar with x265 syntax. I defined a fair amount of it. :)
zub35
17th January 2019, 23:22
benwaggoner If the codec will always be so slow to encode on x86 CPUs, then this encoder will be exclusively for large companies that are capable of acquiring a HW encoder.
AV1 it is a medal from two sides:
1. free and high quality 2. the need to purchase special HW encoders.
Therefore, what is the point of the community to develop it in of quality improvement.
AV1 in its current form (x86), is simply a way of advertising and popularization, nothing more.
p.s. AV1 - "free" (need buy HW) codec for only youtube...
hajj_3
18th January 2019, 00:31
benwaggoner If the codec will always be so slow to encode on x86 CPUs, then this encoder will be exclusively for large companies that are capable of acquiring a HW encoder.
AV1 it is a medal from two sides:
1. free and high quality 2. the need to purchase special HW encoders.
Therefore, what is the point of the community to develop it in of quality improvement.
AV1 in its current form (x86), is simply a way of advertising and popularization, nothing more.
p.s. AV1 - "free" (need buy HW) codec for only youtube...
rav1e is a reasonably fast encoder, it can do several fps and will get faster with time.
Also youtube won't be the only ones adopting this. Facebook, bbc iplayer, netflix etc and many others will be adopting this.
foxyshadis
19th January 2019, 10:18
benwaggoner If the codec will always be so slow to encode on x86 CPUs, then this encoder will be exclusively for large companies that are capable of acquiring a HW encoder.
It's not completely set in stone, but I really believe H.264/AVC might be the last codec easily encodable in pure x86/x64. lntel and AMD have agreed on some extensions to make H.265/HEVC and AV1 not suck quite as much, but the obvious direction is in GPU or fixed-function encoding.
hajj_3
19th January 2019, 11:01
lntel and AMD have agreed on some extensions to make H.265/HEVC and AV1 not suck quite as much
source?
alex1399
20th January 2019, 03:20
I think that is the reason H.264/AVC ditched some complicated features whether they can be implemented by hardware decoder or not. In some extensions, heavy taxing on cpu loading from software encoding perspective is quite possible.
Gravitator
20th January 2019, 08:23
How do I beat a pulsating noise?
I need to save/delete it across the entire sample (using the AOM settings).
> Encoded sample (https://files.videohelp.com/u/227452/AOM%20test%20sp6.mkv)
> Original sample (https://files.videohelp.com/u/227452/SW2.mkv)
aomenc --passes=2 --pass=1 --target-bitrate=800 --end-usage=vbr --fpf="PATH TO THE .stats FILE" --profile=0 --cpu-used=6 --min-q=0 --max-q=63 --bias-pct=70 --minsection-pct=15 --maxsection-pct=10000 --lag-in-frames=25 --drop-frame=0 --undershoot-pct=0 --overshoot-pct=0 --buf-sz=6 --buf-initial-sz=4 --buf-optimal-sz=5 --drop-frame=0 --kf-min-dist=0 --kf-max-dist=250 --auto-alt-ref=1 --arnr-maxframes=7 --arnr-strength=5 --noise-sensitivity=0 --sharpness=0 --static-thresh=0 --tune-content=default --tile-columns=0 --tile-rows=0 --aq-mode=0 --min-gf-interval=0 --max-gf-interval=0 --threads=2 --width=1920 --height=816 --i420 --input-bit-depth=10 --bit-depth=10 --row-mt=0 --cdf-update-mode=1 -o NUL -
aomenc --passes=2 --pass=2 --target-bitrate=800 --end-usage=vbr --fpf="PATH TO THE .stats FILE" --profile=0 --cpu-used=6 --min-q=0 --max-q=63 --bias-pct=70 --minsection-pct=15 --maxsection-pct=10000 --lag-in-frames=25 --drop-frame=0 --undershoot-pct=0 --overshoot-pct=0 --buf-sz=6 --buf-initial-sz=4 --buf-optimal-sz=5 --drop-frame=0 --kf-min-dist=0 --kf-max-dist=250 --auto-alt-ref=1 --arnr-maxframes=7 --arnr-strength=5 --noise-sensitivity=0 --sharpness=0 --static-thresh=0 --tune-content=default --tile-columns=0 --tile-rows=0 --aq-mode=0 --min-gf-interval=0 --max-gf-interval=0 --threads=2 --width=1920 --height=816 --i420 --input-bit-depth=10 --bit-depth=10 --row-mt=0 --cdf-update-mode=1 -o OUTPUTFILE -
benwaggoner
21st January 2019, 19:26
It's not completely set in stone, but I really believe H.264/AVC might be the last codec easily encodable in pure x86/x64. lntel and AMD have agreed on some extensions to make H.265/HEVC and AV1 not suck quite as much, but the obvious direction is in GPU or fixed-function encoding.
HEVC is certainly encodable on x64, although 32-bit becomes impractical at high quality at very high frame sizes. The trend for even live encoding has been towards highly multithreaded CPU encoding. Fixed-function ASIC style hardware is too inflexible when there are SO many options for how to encode every block, with lots of psychovisual tuning to be done. The more complex codecs get, the more high quality encoders are on CPU. And arithmetic entropy coding really benefits from peak single-thread performance. I think there is hope for hybrid CPU/GPU/ASIC/FPGA models, but I don’t see professional quality encoding not to heavily use CPU anytime soon.
I don’t think there is anything intrinsically hard about AV1 for software encoding. If anything SW encoding could be easier than HW encoding due to parallelization limitations. The bigger issue is that VPx hasn’t had a truly competitive quality @ perf encoder in YEARS. And a reference encoder isn’t a great starting point, which libaom sort of is. If one wanted to build a quality @ perf optimized encoder from scratch, especially for low latency live encoding, I’d start by just implementing the mandatory features of a bitstream and then add features incrementally, seeing what their quality @ perf is. Starting with an encoder that HAS to use ALL codec features like a reference encoder does can be harder than going front the ground up.
Sent from my iPhone using Tapatalk
benwaggoner
21st January 2019, 19:30
rav1e is a reasonably fast encoder, it can do several fps and will get faster with time.
Also youtube won't be the only ones adopting this. Facebook, bbc iplayer, netflix etc and many others will be adopting this.
Do we have data on quality @ perf @ bitrate?
It’s pretty easy to make a fast encoder. But making one that is fast and produces competitive quality at a given bitrate is a lot harder.
AV1 is new enough that I don’t expect encoders to give us a clear sense of what the potential quality @ perf for the bitstream is yet. There is a lot of quality and perf optimization to be in the ballpark to compare.
Sent from my iPhone using Tapatalk
utack
24th January 2019, 23:13
A new Speech on AV1 by Tim Terriberry
https://www.youtube.com/watch?v=qubPzBcYCTw
Nintendo Maniac 64
25th January 2019, 02:22
A new Speech on AV1 by Tim Terriberry
https://www.youtube.com/watch?v=qubPzBcYCTw
Kind of amusing that, for something about new fancy-pants video codecs, the video itself has telecine judder (it's been telecine'd from 25fps to 30fps).
TomV
25th January 2019, 08:07
A new Speech on AV1 by Tim Terriberry
https://www.youtube.com/watch?v=qubPzBcYCTw
Tim is a smart engineer, but engineers typically aren't well equipped to present legal opinions / advice (from 2 min to 9 min). Nobody pays for patents twice. If you license from both MPEG LA and HEVC Advance, the companies that are in both only get paid once. The patent chart Tim is using is out of date and inaccurate. Fraunhofer sold their patents to GE, which is why GE has HEVC patents. Canon licenses their patents through MPEG LA. Velos Media has told hundreds of companies what they charge, and they've signed many companies to their license program.
The truth is that every competitive device that supports video now supports HEVC in hardware. Billions of devices, with billions more sold each year. Most every TV, smartphone, tablet or connected set-top box (including Google Chromecast Ultra). If the patent situation were really untenable, Apple, Samsung, LG, Sony, Amazon, Google, GoPro, Roku and hundreds of other device OEMs wouldn't be incorporating HEVC in their devices. And we wouldn't be watching 4K HDR HEVC movies from Netflix, Amazon, Hulu, Vudu, and Apple. If you want another perspective, I gave a talk at the SF Video meetup on the topic... https://www.youtube.com/watch?v=vgE8-4rcXl0
The Alliance for Open Media isn't the only group working to deal with the difficulties of licensing patents for industry standards. MPEG is dealing with it. The Media Coding Industry Forum is dealing with it. And outside firms like Unified Patents, with their Video Codec Zone (specifically focused on HEVC) are dealing with it. It's an ongoing challenge (both for HEVC and for new standards in development), but it's being dealt with.
On the other hand, I'm blazing along at less than 1 frame per minute of 1080P with aomenc cpu-used=1 on a fast Core i7-7820X (with hyperthreading disabled, for the fastest possible single-threaded performance). The resulting videos are roughly on par with my HEVC encodes at identical bit rates (sometimes better, but very often worse). They're clean, but soft and lacking detail.
TD-Linux
26th January 2019, 22:10
Do we have data on quality @ perf @ bitrate
Here's a link to AWCY as shown in Tim's presentation:
https://beta.arewecompressedyet.com/?job=x264-veryslow%402018-11-04T00%3A40%3A25.690Z&job=vp9_Sept-06-19%402018-09-06T14%3A41%3A38.675Z&job=master-s1-high-latency-525f981376bd
tl;dr it is better than x264 at every bitrate, but still worse that libvpx VP9. It is also currently about 10x slower than x264, which is blazing fast compared to libaom but still has a lot of room for improvement.
benwaggoner
26th January 2019, 22:56
AWCY doesn’t include any metrics that are well demonstrated to be able to finely discriminate between quality of different encoders and codecs. VMAF is the least-bad we’ve ever had, but can still be off quite a bit for individual clips, especially if the use codec features or psychovisual optimizations that weren’t included in their test clips. For example, VMAF is bad at rating effectiveness of low-Luna adaptive quant, I speculate because they didn’t include any clips that used different ways to do that in their testing. VMAF is a very impressive effort, but it is not magic. Like all machine learning aystems, it tried to predict what a human would answer given complex input, based on a. large set of example inputs and answers. But I’d it doesn’t have human input for some kinds of inputs, the validity of its predicted ratings for those inputs is unpredictable at best.
Also, the value of mean or even harmonic mean of per-frame scores is limited for clips much more than 10 seconds. A movie encoded in CBR and a VBR encode at the same ABR might up with the same mean score per frame, but the VBR would be strongly preferred by viewers as it offers consistent quality, with the worse sections being a lot better than the worst in a CBR encode.
Comparing psychovisual optimizations, rate control, and significantly different encoders tools requires subjective double-one testing before any confidence in objective meassures’ applicability.
Net-net: you can’t know how good video looks without real people looking at it when techniques are used that weren’t incorporated in an objective metric. If we see a high correlation between MOS and VMAF for a new technique/codec, then we can start trusting that metric.
utack
27th January 2019, 15:30
Does anyone know what the deal with Qualcomm is?
Is the YouTube rollout and Netfix talk pushing them towards making a hardware decoder, or do they have some interests in MPEG doing well and will try to delay AV1 support in phones for a while?
Djfe
28th January 2019, 18:25
A rendered 8k video on YouTube with lots of flickering (epilepsy warning), details like rain drops on a helmet etc.
https://youtu.be/fOWsamMv_v4
Maybe good for comparing av1 to vp9 on YouTube (up to 480p currently)
Once YouTube gets better encoders anyways
obvious already: more blurred but definitely less blocky and less obvious artifacts)
1:30min into the video is probably the best place to compare (and the hardest part for their av1 implementation so far at that bitrate)
benwaggoner
28th January 2019, 20:37
A rendered 8k video on YouTube with lots of flickering (epilepsy warning), details like rain drops on a helmet etc.
https://youtu.be/fOWsamMv_v4
Maybe good for comparing av1 to vp9 on YouTube (up to 480p currently)
Once YouTube gets better encoders anyways
obvious already: more blurred but definitely less blocky and less obvious artifacts)
1:30min into the video is probably the best place to compare (and the hardest part for their av1 implementation so far at that bitrate)
Are you getting at AV1 encode at 480p and below somehow? It shows as VP9 for me at every bitrate (using Chrome).
That is a very interesting clip from a compression perspective. It'll really stress weighted prediction (all those strobes) and adaptive quantization (intense variation in frequency distribution). Tons of value from intraframe prediction.
The VP9 encode is not doing well; at 8K scaled down to my 4K monitor there's lots of blocking and banding issues on the guy. I don't have an immediate intuition for how much is limitations in VP9 versus libvpx. That's a kind of content not in the standard libraries of clips encoders get tuned against.
I would expect libaom to do pretty well against it at a slow preset, as libaom is doing a pretty broad mode search with its myriad tools. So it might find lots of oddball methods that work well with this clip. Probably a big gap between slower and faster modes.
Nintendo Maniac 64
28th January 2019, 20:51
Are you getting at AV1 encode at 480p and below somehow? It shows as VP9 for me at every bitrate (using Chrome).
Don't you have to opt into using AV1 on YouTube?
Anyway, I can definitely confirm via youtube-dl that AV1 (listed as av01) encodes do in fact exist for that video at 480p resolution and lower:
https://imgoat.com/uploads/aa1883c641/190367.png
benwaggoner
28th January 2019, 20:57
Don't you have to opt into using AV1 on YouTube?
Anyway, I can definitely confirm via youtube-dl that AV1 (listed as av01) encodes do in fact exist for that video at 480p resolution and lower:
Yes, it can be set here: https://www.youtube.com/testtube
Now I need to figure out how to do side/by/side in different codecs. Worth comparing to the x264 encodes as well.
benwaggoner
28th January 2019, 21:51
Does anyone know what the deal with Qualcomm is?
Is the YouTube rollout and Netfix talk pushing them towards making a hardware decoder, or do they have some interests in MPEG doing well and will try to delay AV1 support in phones for a while?
I don’t know anything about Qualcomm specifically, but it can take quite a while to go from final spec to design to tape-out to samples to full-scale fab to products launching with a new SoC.
It was being generally discussed in the industry that AV1’s bitstream finalization delays caused chipmakers to miss the 2019 product design window. A HW accelerated decoder is a lot more flexible, but fixed-function decoder needs to be RIGHT. Small product flaws can wind up impact the entire industry for years. And the combination of video decode and DRM is complex with very high functional requirements.
And I’ve heard that implementing AV1 in hardware is more complex than anticipated, due to relatively low parallelism opportunities and how many discreet tools can get applied to any given pixel. Getting a decoder running on a low-power chip is quite diffeeenr than with 1-2 very fast x64 threads. Say what you will about the MPEG process, but it is good at constraining decode complexity for software and hardware.
Nintendo Maniac 64
28th January 2019, 23:39
Now I need to figure out how to do side/by/side in different codecs.
Download each individual video stream via youtube-dl and then play them back in their own video player program window?
alex1399
29th January 2019, 18:28
libavfilter could be utilized to perform some sort of [0:v]crop=in_w/2:in_h:0:0[VL];[1:v]crop=in_w/2:in_h:in_w/2:0[VR];[VL][VR]hstack stuff
mandarinka
29th January 2019, 23:39
AWCY doesn’t include any metrics that are well demonstrated to be able to finely discriminate between quality of different encoders and codecs. VMAF is the least-bad we’ve ever had, but can still be off quite a bit for individual clips, especially if the use codec features or psychovisual optimizations that weren’t included in their test clips.
Isn't there also the possibility that there's a sort of implicit "training" for this metric included in one codec/encoder and not the other?
I don't know whether VMAF was used in some x265 tuning, but given the age of all the significant parts of x264 codebase, I am fairly sure that there has been no attempts to do this.
Meanwhile VMAF was IIRC used during development of Daala and AV1 and maybe AOMenc/Rav1e? In that case, there could be some inherent bias in the metric towards those codecs that would then add some imaginary advantage above their real compression quality into the numbers, when measured by VMAF. Simply because their output was implicitly tuned to get better VMAF, because VMAF was used to test new tools/analysis/RDO and so on.
benwaggoner
30th January 2019, 00:20
Isn't there also the possibility that there's a sort of implicit "training" for this metric included in one codec/encoder and not the other?
It is an inevitability. Generally the utility of a metric goes down once it is codified, because people start optimizing for that metric instead of the subjective quality that metric approximates. So the correlation of the metric with subjective ratings becomes weaker, as metric-specific tricks get implemented.
I don't know whether VMAF was used in some x265 tuning, but given the age of all the significant parts of x264 codebase, I am fairly sure that there has been no attempts to do this.
x265 was around long before VMAF, and a VMAF useful for UHD has only been around a few months. Libaom seems have been getting a lot more tuning-by-VMAF.
Meanwhile VMAF was IIRC used during development of Daala and AV1 and maybe AOMenc/Rav1e? In that case, there could be some inherent bias in the metric towards those codecs that would then add some imaginary advantage above their real compression quality into the numbers, when measured by VMAF. Simply because their output was implicitly tuned to get better VMAF, because VMAF was used to test new tools/analysis/RDO and so on.
VMAF wasn't around for most/all of Daala work. AV1 is really the first bitstream to have its practical implementations start in the VMAF era. This is one reason I'm suspicious of VMAF scores for AV1. Good analysis.
In particular I worry that VMAF is insufficiently sensitive to temporal shifts in video quality. A VMAF of 70, 65, 60, 60, 65, 60, 60 might come out as a nice "VMAF=65.3" but be a annoying to watch. Frame strobing was a weakness of libvpx.
TD-Linux
30th January 2019, 01:00
The patent chart Tim is using is out of date and inaccurate. Fraunhofer sold their patents to GE, which is why GE has HEVC patents. Canon licenses their patents through MPEG LA.
The chart is the one used in Leonardo's blog (http://blog.chiariglione.org/a-crisis-the-causes-and-a-solution/), though I think it's originally from streamingmedia.com. I've attached an updated version for future presentations.
https://people.xiph.org/~tdaede/HEVC.png
jonatans
30th January 2019, 02:14
I originally created the figure and presented it for the first time during Streaming Tech Sweden 2017. There have been some changes since then.
Here is an updated figure:
https://www.divideon.com/images/HevcPatentHolders190130.png
Please note that the figure is only based on public information available from ISO/IEC/ITU and from the patent pools. Please also note that not all of the MPEG LA patent holders are shown in the figure.
hajj_3
30th January 2019, 02:32
I originally created the figure and presented it for the first time during Streaming Tech Sweden 2017. There have been some changes since then.
Here is an updated figure:
https://www.divideon.com/images/HevcPatentHolders190130.png
Please note that the figure is only based on public information available from ISO/IEC/ITU and from the patent pools. Please also note that not all of the MPEG LA patent holders are shown in the figure.
I think i read that Franhaufer sold their HEVC patents to General Electric, if true you should remove Franhaufer from your diagram.
TomV
30th January 2019, 02:51
The chart is the one used in Leonardo's blog (http://blog.chiariglione.org/a-crisis-the-causes-and-a-solution/), though I think it's originally from streamingmedia.com. I've attached an updated version for future presentations.
https://people.xiph.org/~tdaede/HEVC.png
It's originally from Jonatan Samuelsson, a.k.a. jonatans (https://forum.doom9.org/member.php?u=227153)
Even with the updates, the main problem with this chart is that it's a bit misleading.. for 2 reasons. First, I think most people in the video industry assume that AVC patent licensing is and was perfectly clean and simple... that all necessary standard-essential patents were licenseable in the MPEG LA patent pool. That's not true. Nokia, Qualcomm, Broadcomm, Blackberry, Texas Instruments, MIT all hold standard-essential AVC patents outside the MPEG LA pool (although Qualcomm messed up and a judge ruled they can't assert them for AVC). Multiple legal battles have been fought over AVC patents, including some pretty big cases... Microsoft v Motorola, and Apple v Nokia. Today, everyone can agree that the patent licensing situation for AVC is much better than it is for HEVC. But it didn't start out that way, and it took some time for the situation to settle. Also, there are quite a few more patent holders in some of these HEVC pools than shown in this chart.
In his talk, Tim mentioned that patents are issued that may not be valid (https://youtu.be/qubPzBcYCTw?t=498), and then said "and you could go around and try to invalidate them all, but they're really expensive to do that, and there's a lot of them". Well, if you're a multi-billion dollar company (Apple, Samsung, Google, Amazon, etc.), you have a lot of lawyers, and that's what they're paid to do. If you're being asked to pay tens or hundreds of millions of dollars a year in patent license fees, you have all the motivation in the world to spend whatever it takes on legal fees to right-size the problem. When multiple multi-billion dollar companies have this issue, collectively there is a lot of motivation. It turns out that when challenged in court, most patents don't hold up. They can be invalidated for many reasons... prior art, unpatentable claims, obviousness, the invention was anticipated, etc. This type of effort is being undertaken by Unified Patents (as a service to many large tech companies), and there is a relatively new law called the America Invents Act that provides a faster, less expensive way to get rid of bad patents, called an Inter Partes Review (IPR). Unified already filed an IPR against Velos Media, and you can expect more such filings under their Video Codec domain. But you don't even have to invalidate patents in order not to pay a fortune.
Again, keep in mind that no one is asking for patent license fees for content distribution (streaming, etc.), except for UHD-Blu-ray disc (a small per-disc fee to HEVC advance). Only hardware device manufacturers need to license HEVC patents, and they are dealing with that issue and they continue to support HEVC in every device they make that supports video. For video services, HEVC is free. Now that the majority of active end-user devices support HEVC, it makes a lot of financial sense for video services to make their VOD catalog, or the majority of their live channels available in both AVC and HEVC (not just 4K and HDR content... all content). The bandwidth savings and customer experience improvement far outweigh the additional cost of encoding and CDN storage.
TomV
30th January 2019, 02:58
I think i read that Franhaufer sold their HEVC patents to General Electric, if true you should remove Franhaufer from your diagram.
That's true. I don't think it was ever announced, but I can assure you that I got confirmation from a very reliable source.
TomV
30th January 2019, 03:02
I originally created the figure and presented it for the first time during Streaming Tech Sweden 2017. There have been some changes since then.
Here is an updated figure:
https://www.divideon.com/images/HevcPatentHolders190130.png
Please note that the figure is only based on public information available from ISO/IEC/ITU and from the patent pools. Please also note that not all of the MPEG LA patent holders are shown in the figure.
Hey... we were both responding at the same time (cross posting). Thanks for posting an update Jonatan. Your efforts on this, and in the MC-IF are really appreciated.
mandarinka
30th January 2019, 04:39
Interesting that some of the champions supporting or helping AOM are in the problematic(?) group of unpooled HEVC licensors... you would say these companies support AV1 because they hated that. :devil:
kuchikirukia
30th January 2019, 05:55
Here's a link to AWCY as shown in Tim's presentation:
https://beta.arewecompressedyet.com/?job=x264-veryslow%402018-11-04T00%3A40%3A25.690Z&job=vp9_Sept-06-19%402018-09-06T14%3A41%3A38.675Z&job=master-s1-high-latency-525f981376bd
tl;dr it is better than x264 at every bitrate, but still worse that libvpx VP9. It is also currently about 10x slower than x264, which is blazing fast compared to libaom but still has a lot of room for improvement.
Taking x264 PSNR and SSIM values without setting --tune PSNR and SSIM is fail.
nevcairiel
30th January 2019, 10:35
Taking x264 PSNR and SSIM values without setting --tune PSNR and SSIM is fail.
You get one encode to compare, do you really want that to be one tuned for PSNR? Because that would be the real fail.
One encode, several metrics. Not re-encoding targeted for metrics.
TD-Linux
30th January 2019, 10:42
Taking x264 PSNR and SSIM values without setting --tune PSNR and SSIM is fail.
There are more metrics than those at the link, but fair. I compared against x264 --tune PSNR as well and it still beats x264 at PSNR:
https://beta.arewecompressedyet.com/?job=x264-veryslow-tune-psnr%402019-01-30T08%3A44%3A59.051Z&job=master-s1-high-latency-525f981376bd
MoSal
30th January 2019, 12:29
In particular I worry that VMAF is insufficiently sensitive to temporal shifts in video quality. A VMAF of 70, 65, 60, 60, 65, 60, 60 might come out as a nice "VMAF=65.3" but be a annoying to watch. Frame strobing was a weakness of libvpx.
That's not really a problem. Per-frame data is available (wrote this (https://github.com/MoSal/vmaf-plot) mostly in a couple of hours). It's even available for multiple metrics, which is nice.
The real problem is that VMAF is not that good. It, for example, spectacularly fails with samples that greatly benefit from AQ (yes, I know you already hinted at this).
jonatans
30th January 2019, 13:57
Thanks for posting an update Jonatan. Your efforts on this, and in the MC-IF are really appreciated.
Thank you Tom. And thanks for providing additional context to this interesting and complicated matter.
I think i read that Franhaufer sold their HEVC patents to General Electric, if true you should remove Franhaufer from your diagram.
This is correct. But my understanding is that Fraunhofer did not sell all their HEVC patents. They are listed as licensor in HEVC Advance. In the latest patent list from HEVC Advance there are two Fraunhofer patents listed: https://www.hevcadvance.com/pdfnew/HEVC-Patent-List-January-2019.pdf
Beelzebubu
30th January 2019, 17:56
In particular I worry that VMAF is insufficiently sensitive to temporal shifts in video quality. A VMAF of 70, 65, 60, 60, 65, 60, 60 might come out as a nice "VMAF=65.3" but be a annoying to watch. Frame strobing was a weakness of libvpx.
First and foremost: yes! It's great to see some technical & independent thinking of how good VMAF really is.
Netflix uses "hVMAF" as official notation in their charts. "h" means "harmonic", which means it uses harmonic (https://en.wikipedia.org/wiki/Harmonic_mean) means, which bias towards the least favourable. So in your example, the harmonic mean would be 62.66, whereas the average would be 62.86. Neither of these is 65.3. So I'm personally not as concerned about the averaging mechanism aspect of your concern. (In the CLI, use --pool harmonic_mean or something similar, depending on which exact tool you use.) On the other hand, I don't believe that VMAF uses temporal consistency in the reconstruction (the "motion" component is calculated from the source), so that particular concern ("frame throbbing" - i.e. keyframe pulsing or grain/textured-background tearing) I agree with.
Actually, I have to hedge a little here, since I'm not 100% sure VIF (another VMAF component) has a temporal component to it. I don't think it does but I'm not 100% sure.
Since we're on the subject, here's some more of my personal concerns about VMAF:
* it's luma-only;
* AQ (x264/5) or SAO (x265) appear to have a negative impact on vmaf score, which is inconsistent with the reported visual results. I do have more detailed thoughts on this but let's leave that for some other time;
* the actual MOS/VMAF correlation depends very strongly on the viewing environment and therefore on the used model file, but most poeple simply use the default model without knowing what viewing environment it represents.
Just to be clear, I'm not trying to talk badly about VMAF, I think it's a great tool, it's better than the alternatives and it's fantastic that they opensourced the library as well as the models so that we can learn and understand how it works and constructively critique it. Hopefully, over time, that will make it even better, which should be the ultimate goal.
Separately, I also do agree with you that in the end, we should probably make a distinction between codecs optimized using VMAF vs. those that did not. This isn't an excuse to suck at writing encoders or to not use VMAF when writing encoders, but at the end of the day, we have to acknowledge that as in any metric, we're assuming a perfect correlation between our metric-of-the-day and the visual experience (or MOS score). That correlation will in practice always be imperfect, and therefore tuning towards/using that metric needs to be done with care and with visual confirmation (otherwise queue up the incoming VMAF artifacts - I wonder what they will look like?).
MoSal
30th January 2019, 21:57
@Beelzebubu
What's really funny, Netflix will not be using VMAF on their published AOM content as is. Why? Because of film grain synthesis ;)
benwaggoner
31st January 2019, 06:00
@Beelzebubu
What's really funny, Netflix will not be using VMAF on their published AOM content as is. Why? Because of film grain synthesis ;)
Well, if VMAF is used on the reconstructed video, it should be as good as VMAF is at dealing with film grain.
If VMAF is bad at dealing with film grain, they need to address that.
The nice thing about VMAF is that it's really a machine learning framework. They can keep on adding new clips and kinds of encodings and training it to rate those. The big expenses is getting the subjective ratings to use as ground-truth data. But VMAF itself can always be as good as the ground truth data from subjective testing.
TD-Linux
31st January 2019, 06:38
The nice thing about VMAF is that it's really a machine learning framework. They can keep on adding new clips and kinds of encodings and training it to rate those. The big expenses is getting the subjective ratings to use as ground-truth data. But VMAF itself can always be as good as the ground truth data from subjective testing.
The inputs to VMAF itself are the outputs of a bunch of simpler metrics. In that way, the machine-learned part is sort of a "meta-metric". That also means that if the input metrics all respond poorly to film grain, no amount of machine learning is going to be able to make sense of it. I think more work on the input metrics will be needed before VMAF can be used to make film grain decisions.
I don't know what Netflix currently does, but if I were them I would filter the grain from the video, run the VMAF-targeting dynamic optimizer to produce the rate controlled stream, and then add the noise parameters back as a final step.
LigH
31st January 2019, 08:55
Franhaufer
Fraunhofer
Frau = woman
Hof = yard
TomV
31st January 2019, 19:27
HEVC Advance standard essential patent owned by GE challenged as likely invalid (https://www.unifiedpatents.com/news/2019/1/31/tx7k046x6yvail8s3jz4v8tes3x4m7)
benwaggoner
31st January 2019, 20:33
The inputs to VMAF itself are the outputs of a bunch of simpler metrics. In that way, the machine-learned part is sort of a "meta-metric". That also means that if the input metrics all respond poorly to film grain, no amount of machine learning is going to be able to make sense of it. I think more work on the input metrics will be needed before VMAF can be used to make film grain decisions.
I don't know what Netflix currently does, but if I were them I would filter the grain from the video, run the VMAF-targeting dynamic optimizer to produce the rate controlled stream, and then add the noise parameters back as a final step.
Good point on the underlying metrics. In particular I think the temporal metric was quite weak. They changed it for the most recent VMAF, but I'm not confident it'll catch all common kinds of visible temporal distortions. Two frames can look equally "good" but switching between them can be terribly jarring. Open GOP and RADL exist in large part to smooth inter-GOP transitions. And that still requires some cleverness around GOP boundaries to do well.
Mr_Khyron
1st February 2019, 20:13
https://github.com/OpenVisualCloud/SVT-AV1
Welcome to the GitHub repo for the SVT-AV1 encoder! To see a list of feature request and view what is planned for the SVT-AV1 encoder, visit our Trello page: http://bit.ly/SVT-AV1 Help us grow the community by subscribing to our SVT-AV1 mailing list
:cool:
Hardware
The SVT-AV1 Encoder library supports the x86 architecture
CPU Requirements
In order to achieve the performance targeted by the SVT-AV1 Encoder, the specific CPU model listed above would need to be used when running the encoder. Otherwise, the encoder runs on any 5th Generation Intel® Core™ processor, (Intel® Xeon® CPUs, E5-v4 or newer).
RAM Requirements
In order to run the highest resolution supported by the SVT-AV1 Encoder, at least 48GB of RAM is required to run a 4k 10bit stream multi-threading on a 112 logical core system. The SVT-AV1 Encoder application will display an error if the system does not have enough RAM to support this. The following table shows the minimum amount of RAM required for some standard resolutions of 10bit video per stream:
Resolution Minimum Footprint (GB)
4k 48gb
1080p 16gb
720p 8gb
480p 4gb
Selur
1st February 2019, 20:17
so an Intel only encoder?
nevcairiel
1st February 2019, 20:20
so an Intel only encoder?
It should be able to run on any AVX2 CPU.
But the entire series of SVT encoders (SVT-HEVC is also a thing) is designed specifically for a use-case of running them on powerful datacenter systems with loads of memory and CPU cores.
benwaggoner
1st February 2019, 20:47
https://github.com/OpenVisualCloud/SVT-AV1
:cool:
Wow, that's a LOT of RAM for 4K. But if it's somewhat proportional to number of cores, no biggie. Any 112 logical core system is going to have >> 48 GiB RAM. The biggest c5 instance today is 72 logical threads and 144 GiB RAM.
I don't think there's ever been an encoder that could usefully use anything like 112 cores except via GOP-level parallelism. But hey, it's Intel.
I've not been able to find much detailed documentation about the SVT HEVC or AV1 projects. Do they mean "Scalable Video" ala SVC and SHVC with enhancement layers, mainly used in videoconferencing? Or scalable in the sense of scaling with hardware?
Leveraging the new low-level encoder SDK from Intel offers some interesting potential for very fast initial estimates for encoding, leaving the CPU to focus more on refinement. There isn't an AV1 encoder in the current Intel CPUs, obviously, but perhaps some VP9 functionality added in Kaby/Coffee Lake can be leveraged. Certainly things like weighted prediction and coarse motion vectors could be reused to some degree. SVT HEVC has a full 8-bit HEVC encoder implementation to leverage in Skylake-S+ and 10-bit in Kaby/Coffee.
Unfortunately there aren't any Xeon processors with VP9 encoding yet. The best available is the 8/16 core i9-9900K. I don't see any public roadmap for when AV1 might be added. Ice Lake? I see that has an all new HEVC encoder at least. Although given tape-out schedules and how recent the AV1 bitstream was finalized, a full fixed-function implementation might not be there before Tiger Lake. (all just personal speculation fueled by Wikipedia).
I am very curious to see what comes out of the next generation of GPU-assisted software-defined encoding. Having it all on-die instead avoid the PCI bus latency challenges of past GPU+CPU implementations.
nevcairiel
1st February 2019, 20:50
I've not been able to find much detailed documentation about the SVT HEVC or AV1 projects. Do they mean "Scalable Video" ala SVC and SHVC with enhancement layers, mainly used in videoconferencing? Or scalable in the sense of scaling with hardware?
Hardware. Its not producing "scalable video".
TomV
2nd February 2019, 00:26
I've not been able to find much detailed documentation about the SVT HEVC or AV1 projects. Do they mean "Scalable Video" ala SVC and SHVC with enhancement layers, mainly used in videoconferencing? Or scalable in the sense of scaling with hardware?
No. Intel bought eBrisk (they already owned a good chunk of eBrisk, thanks to the Altera acquisition, because Altera had invested in eBrisk), and then open sourced their HEVC encoder. Then they started focusing on AV1, and now they've open sourced that encoder. The HEVC encoder is fast, but the video quality is not competitive. I'm not sure if it can beat x264 under equal conditions. It certainly can't beat x265 or Beamr 5 under any conditions.
TomV
2nd February 2019, 00:26
Hardware. Its not producing "scalable video".
Not hardware. Software.
nevcairiel
2nd February 2019, 00:43
Not hardware. Software.
You should read the context of the question that answer was to. ;-)
To make sure its not lost again, let me paraphrase: :p
Q: Scalable Video, or Scaling with Hardware?
A: Hardware.
TomV
2nd February 2019, 02:47
You should read the context of the question that answer was to. ;-)
To make sure its not lost again, let me paraphrase: :p
Q: Scalable Video, or Scaling with Hardware?
A: Hardware.
OK... I see. Exactly what they were thinking when they used this acronym, which, as Ben points out, is confusingly similar to Scalable Video Coding and Scalable HEVC Video Coding, I don't know. Nothing to see here... move along.
hajj_3
3rd February 2019, 17:47
Intel SVT-AV1 benchmarks: https://twitter.com/fg118942/status/1092045469981671424
soresu
3rd February 2019, 20:54
Is it just me or is rav1e actually pulling out ahead of VP9 at some bitrates on the graph? If so thats a nice milestone for rav1e, given the timeframe.
benwaggoner
4th February 2019, 19:41
Intel SVT-AV1 benchmarks: https://twitter.com/fg118942/status/1092045469981671424
Is there more documentation on what's actually being tested and graphed here? Based on the parameters, it doesn't appear to be controlled for encoding speed. And odd to use --tune ssim for x264/x265 for VMAF, which is a superior objective metric than SSIM.
I wish tests would provide the actual per-frame VMAF scores instead of just a mean. For real-world duration stuff, variability of quality can hurt subjective quality in a way that VMAF itself won't capture. Keyframe strobing on one frame every 5 seconds doesn't drag down the mean much, but it can be a very annoying artifact viewers can clap along to.
Nintendo Maniac 64
4th February 2019, 20:31
I wish tests would provide the actual per-frame VMAF scores instead of just a mean.
I presume you mean (pun not intended) in a manner similar to frame-time graphs used in GPU performance benchmarks, or at least just also showing a 1% low? For a similar reason, they came about since showing the average frame rate hides any uneven frame delivery which is much more important to game playability.
The only thing is that such graphs would tend to be limited to having a single bitrate or quality setting since the bottom axis in such a situation would be time rather than bitrate.
...which might very well be why people don't do it - because they want to show a single graph with various differing bitrates rather than a really detailed graph but only at a single bitrate or quality setting.
EDIt: A 1% low graph would at least let you do this, but it would still also require making a second graph (unless you don't even care about the mean at all, in which case you could just graph a 1% low and call it a day).
nevcairiel
4th February 2019, 20:37
We get SSIM graphs with per-frame curves, so its not that of a "new" idea to also do that for VMAF or the likes.
benwaggoner
5th February 2019, 00:56
I presume you mean (pun not intended) in a manner similar to frame-time graphs used in GPU performance benchmarks, or at least just also showing a 1% low? For a similar reason, they came about since showing the average frame rate hides any uneven frame delivery which is much more important to game playability.
The only thing is that such graphs would tend to be limited to having a single bitrate or quality setting since the bottom axis in such a situation would be time rather than bitrate.
...which might very well be why people don't do it - because they want to show a single graph with various differing bitrates rather than a really detailed graph but only at a single bitrate or quality setting.
EDIt: A 1% low graph would at least let you do this, but it would still also require making a second graph (unless you don't even care about the mean at all, in which case you could just graph a 1% low and call it a day).
Having a harmonic mean of the worst 0.1%, 1%, 10% would be quite useful, yes.
But the actual VMAF output is just per-frame scores, so anyone publishing a mean VMAF already has the data. Even if they don't want to plot the data, they could still make the log files available for download.
fg118942
5th February 2019, 03:24
Having a harmonic mean of the worst 0.1%, 1%, 10% would be quite useful, yes.
But the actual VMAF output is just per-frame scores, so anyone publishing a mean VMAF already has the data. Even if they don't want to plot the data, they could still make the log files available for download.
Log files and encoded videos are here.
https://www.dropbox.com/s/nbnlsicvslptt2c/vidyo1_720p_60fps.7z?dl=0
I am encoding it with tune ssim because I followed the instructions in this article.
https://www.streamingmedia.com/Articles/Editorial/Featured-Articles/AV1-A-First-Look-127133.aspx
I may not be able to answer difficult questions as I am not good at English.
TomV
5th February 2019, 16:57
Intel SVT-AV1 benchmarks: https://twitter.com/fg118942/status/1092045469981671424
x264 and x265 preset slower is not the right preset to use versus aomenc --cpu-used = 0 and SVT-AV1 enc-mode 0. This test should compare with x264, x265 --preset placebo. Better yet, forget objective metrics. Just show us the video, so we can judge for ourselves the bit rates that produce matching subjective quality.
kanaka
6th February 2019, 10:03
My AVIF toolkit: https://mega.nz/#!5oQE2Sob!STZHdk4ob4ptHknMvNcB4JxbCt9xdu3WUKkg7iyh2EM
I tested avif format with this photo
https://personal.sron.nl/~pault/images/colourvisiontest_small.png
Avif file was different from source... (text wasn't readable), so i removed
--color-primaries=bt709 --transfer-characteristics=bt709 --matrix-coefficients=bt709
and result was ok. Why did you put this color profile?
I'm thinking about conversion my 12bit raw photos to avif. Is is possible? What pix_format shoul I use?
fg118942
6th February 2019, 11:46
x264 and x265 preset slower is not the right preset to use versus aomenc --cpu-used = 0 and SVT-AV1 enc-mode 0. This test should compare with x264, x265 --preset placebo. Better yet, forget objective metrics. Just show us the video, so we can judge for ourselves the bit rates that produce matching subjective quality.
I thought that the point was right so I added placebo data.
https://i.imgur.com/V1WH6GJ.png
Also, the video encoded with SVT-AV1 has been uploaded here.
https://www.dropbox.com/s/nbnlsicvslptt2c/vidyo1_720p_60fps.7z?dl=0
LigH
6th February 2019, 18:27
New uploads: (MSYS2; MinGW32: GCC 7.4.0 / MinGW64: GCC 8.2.1)
AOM v1.0.0-1299-g54eabb5c8 (https://www.mediafire.com/file/q1x1a8akjgtqu8c/aom_v1.0.0-1299-g54eabb5c8.7z)
rav1e 0.1.0 (2cec0f9 / 2019-02-06) (https://www.mediafire.com/file/w1o3g5wdye8w8o5/rav1e_0.1.0_2019-02-06_2cec0f9.7z)
dav1d 0.1.1 (caca572 / 2019-02-06) (https://www.mediafire.com/file/aym9cb9e1ct5vi7/dav1d_0.1.1_2019-02-06_caca572.7z)
benwaggoner
6th February 2019, 21:05
Log files and encoded videos are here.
https://www.dropbox.com/s/nbnlsicvslptt2c/vidyo1_720p_60fps.7z?dl=0
I am encoding it with tune ssim because I followed the instructions in this article.
https://www.streamingmedia.com/Articles/Editorial/Featured-Articles/AV1-A-First-Look-127133.aspx
I may not be able to answer difficult questions as I am not good at English.
Thank you!
benwaggoner
6th February 2019, 21:07
I tested avif format with this photo
https://personal.sron.nl/~pault/images/colourvisiontest_small.png
Avif file was different from source... (text wasn't readable), so i removed
--color-primaries=bt709 --transfer-characteristics=bt709 --matrix-coefficients=bt709
and result was ok. Why did you put this color profile?
I'm thinking about conversion my 12bit raw photos to avif. Is is possible? What pix_format shoul I use?
709==sRGB, so I am surprised it made a difference. Perhaps a 0-255 versus 16-235 luma range conversion? Making text unreadable would be a weird result, though.
benwaggoner
6th February 2019, 21:09
x264 and x265 preset slower is not the right preset to use versus aomenc --cpu-used = 0 and SVT-AV1 enc-mode 0. This test should compare with x264, x265 --preset placebo. Better yet, forget objective metrics. Just show us the video, so we can judge for ourselves the bit rates that produce matching subjective quality.
If we are comparing to very slower encoders, I recommend adding --tskip to x265 as well. That can help efficiency with text, cel animation, and other content with synthetically sharp edges.
kanaka
7th February 2019, 08:48
709==sRGB, so I am surprised it made a difference. Perhaps a 0-255 versus 16-235 luma range conversion? Making text unreadable would be a weird result, though.
check yourself http://screenshotcomparison.com/comparison/129605
TD-Linux
7th February 2019, 11:07
check yourself http://screenshotcomparison.com/comparison/129605
Aha, looks like 601 vs 709 matrix. Although JPEG is normally sRGB primaries, it uses what is basically a full-range 601 matrix. So if your sources are JPEG, a 601 matrix makes the most sense.
kanaka
7th February 2019, 11:23
Aha, looks like 601 vs 709 matrix. Although JPEG is normally sRGB primaries, it uses what is basically a full-range 601 matrix. So if your sources are JPEG, a 601 matrix makes the most sense.
Source is png (https://personal.sron.nl/~pault/images/colourvisiontest_small.png)
and there is commands from encode.ST.sh
./bins/ffmpeg -r 1 -y -hide_banner -loglevel fatal -i "$1" -vf scale=out_color_matrix=bt709:flags=lanczos+accurate_rnd+bitexact+full_chroma_int+full_chroma_inp,format=yuv420p10le -strict -1 "temp/orig_$filename.y4m"
./bins/aomenc --threads=4 -v --cpu-used=4 --end-usage=q --cq-level=$quality --sharpness=7 --bit-depth=10 --full-still-picture-hdr --color-primaries=bt709 --transfer-characteristics=bt709 --matrix-coefficients=bt709 --ivf -o "encoded/$filename.ivf" "temp/orig_$filename.y4m"
//edit: I had older version of toolkit. New toolkit use yuv420p and works ok.
benwaggoner
7th February 2019, 19:57
Aha, looks like 601 vs 709 matrix. Although JPEG is normally sRGB primaries, it uses what is basically a full-range 601 matrix. So if your sources are JPEG, a 601 matrix makes the most sense.
sRGB uses 709, which itself is the average of the 601 PAL (EBU 3213) and NTSC (SMPTE C) primaries. As an industry, we should probably stop talking about "601 primaries" since there are actually two different ones, unless we use it as shorthand for "the primaries used by the original SD video format."
709 was the compromise for HD to make it "international" - as the average of the two, if 601 gets treated as 709 or vise versa, that minimizes the worst-case error compares to 601 PAL <> 601 NTSC.
https://en.wikipedia.org/wiki/Rec._709#Primary_chromaticities
As a parochial American, I thought 709 was dumb when it came out, but I have since gained the wisdom to appreciate its simple brilliance.
soresu
7th February 2019, 22:44
benwaggoner you just mentioned AV2 several times in the EVC thread on February 1st. Do you know where current work on AV2 is being committed to if it is public yet? The googlesource.com git site leaves something to be desired as far as usability and search is concerned.
benwaggoner
8th February 2019, 21:45
benwaggoner you just mentioned AV2 several times in the EVC thread on February 1st. Do you know where current work on AV2 is being committed to if it is public yet? The googlesource.com git site leaves something to be desired as far as usability and search is concerned.
I've heard from some people that they are doing some initial work on it, and the hope is that it will be a relatively quick turnaround.
One potential wrinkle to the VPx and AVx codecs is that they know in advance when essential patents are going to expire, so tools could be designed in advance and only deployed when IP is cleared up. So there could be stuff that was too early for AV1 that could be reused. That's just my own personal speculation, though. But that could speed some things up.
People looking at AV2 have also been a lot more optimistic about it than AV1, which didn't get enough attention to HW decoder design optimization, or getting tools to work together orthogonally. One comment I heard is that one tool might be sharpening a pixel while another is smoothing it.
2020 should be an interesting year in the codec space, with AV2, VVC, and EVC all potentially being far enough along to evaluate, and H.264, HEVC, and AV1 still competing for current deployments. After UHD, HDR, HFR, and object-based audio all launching in 2014-2016, it's been a little dull around new media technologies. So I'm pretty amped by all the exciting fun 2020-2022 is going to be for codecs! It'll be an interesting mix of technical, business, and legal factors, and I really don't have a guess yet about what the codec world will look like in five years*!
And audio stuff is heating up with xHE-AAC, AC-4 with Atmos, and MPEG-H all going mainstream.
* Well, I bet we'll still be using MP4 as a container format.
soresu
8th February 2019, 22:47
Is there any word on Daala techniques like PVQ and Activity Masking, and also ANS going into AV2?
Though from the direction of VVC and Google's own encoding research priorities - I could see a more than healthy dose of machine learning put to use in AV2 aswell. ML seems to be affording some very significant complexity/efficiency gains in the area of Path Tracing, and I'm sure all of the involved AOM parties would cheer improvements in encoding complexity.
IgorC
10th February 2019, 18:10
2020 should be an interesting year in the codec space, with AV2, VVC, and EVC all potentially being far enough along to evaluate, and H.264, HEVC, and AV1 still competing for current deployments. After UHD, HDR, HFR, and object-based audio all launching in 2014-2016, it's been a little dull around new media technologies. So I'm pretty amped by all the exciting fun 2020-2022 is going to be for codecs! It'll be an interesting mix of technical, business, and legal factors, and I really don't have a guess yet about what the codec world will look like in five years*!
This statement is beyond of a healthy optimism.
The market of video codecs is cooling down. There are few reasons for that. Royalty free formats start to gain share and some external factors like a big improvement of network bandwidth especially during last years.
And audio stuff is heating up with xHE-AAC, AC-4 with Atmos, and MPEG-H all going mainstream.
:(
I don't know where You get this information from but this is not what happens with audio codecs lately.
xHE-AAC has nothing to do with mainstream. It's a low bitrate codec and companies adopt it only where bandwidth is very scarce. xHE-AAC/AC4 has no advantage over AAC (22 years old format) at 96-128+ kbps. Audio formats are mature at this point.
AC3 patents have expired in 2017 while LC-AAC's will be expired during 2019-2020. It will be imposible to force some company to use new codec when there are AC3 and LC-AAC with expired patents. Giant streaming platforms, Netflix and Youtube, use AAC and Opus. They don't plan to use any new audio codecs in near future.
Plus there is no one single developer team working on xHE-AAC, MPEG-H or AC4 audio codec. And xHE-AAC isn't actually a new format. It's a standard since 2012. Where its development? Adoption?
Blue_MiSfit
10th February 2019, 23:21
Ultra low bitrate is highly desirable for a few specific use cases for companies delivering video:
1) Countries with extremely poor (~2G, to maybe 3G at best) cellular connectivity. Delivering even good quality SD video is totally acceptable here. Total bit budget is often like 200 - 300 Kbps though, so you really do need to use the lowest bitrate audio you can possibly use. 96 Kbps for stereo AAC is not feasible. Opus is great here, but it doesn't have universal support, so more development into other formats is absolutely welcome.
2) Download / offline playback. Imagine you're at the airport about to board a flight. You forgot to download something to watch on your phone / tablet during the flight! You want to be able to download a movie or a couple episodes of a series as quickly, probably using over-crowded WiFi or cellular connectivity. See above.
soresu
11th February 2019, 03:14
IgorC,
I dont know about XHE-AAC, but I heard that UK FreeSAT chose AC-4 as the audio format for its next evolution. Link here (https://dolbyac4.com/uk/).
Other platforms supporting it are shown, aswell as multiple hardware partners (Broadcom, Cadence, HiSilicon,
Mediatek, MStar Semiconductor,
Novatek and Realtek).
IgorC
11th February 2019, 04:14
1) Countries with extremely poor (~2G, to maybe 3G at best) cellular connectivity. Delivering even good quality SD video is totally acceptable here. Total bit budget is often like 200 - 300 Kbps though, so you really do need to use the lowest bitrate audio you can possibly use. 96 Kbps for stereo AAC is not feasible.
This is not true.
Look at the report https://opensignal.com/reports-data/global/data-2018-11/state_of_wifi_vs_mobile_OpenSignal_201811.pdf
The modest mobile and/or fixed connections are about ~2-3 Mbps (in Algeria). Far from yours 200-300 kbps.
Generally people have a wrong idea that every county in Africa, Asia and Latin America (where I live actually) has very bad internet connection.
Here in Latin America I get 10+ Mbps on 4g/LTE+/4G+.
And Indians are mad about their "slow" 4G connection. It's "just" 6 Mpbs! https://www.indiatimes.com/technology/news/india-has-over-86-percent-4g-availability-but-the-worst-data-speed-in-the-world-at-6-07-mbps-340086.html
https://ispspeedindex.netflix.com/country/india/
Do You still think "96 Kbps for stereo AAC is not feasible" and "India is so 2G", right?
Ultra low bitrate is highly desirable for a few specific use cases for companies delivering video:
Yes, corner cases. Not mainstream as Ben claims.
Opus is great here, but it doesn't have universal support, so more development into other formats is absolutely welcome.
Opus is used in Youtube, an endless number of VoIP and telephone clients including Cisco corporate solutions like Webex, Skype etc. And it is supported by large number of platflorms including Android and iOS https://caniuse.com/#search=opus
So are You sugesting to use something better like xHE-AAC which doesn't even has one single available encoder? Oh, nice. That will do.
2) Download / offline playback.
What is wrong with current VP9, H.264, H.265, Opus and HE/AAC codecs?
xHE-AAC isn't any better than HE/AAC, Opus at 96 kbps, which is already low bitrate. https://www.ietf.org/lib/dt/documents/LIAISON/file1298.doc
IgorC,
I dont know about XHE-AAC, but I heard that UK FreeSAT chose AC-4 as the audio format for its next evolution. Link here (https://dolbyac4.com/uk/).
Other platforms supporting it are shown, aswell as multiple hardware partners (Broadcom, Cadence, HiSilicon,
Mediatek, MStar Semiconductor,
Novatek and Realtek).
Great. Both xHE-AAC and AC4 have similar quality as they have the same/similar compression tools. So AC4 has an advantage but only on low bitrate as well. It makes sense to use it where BW is expensive like digital radio DRM but I won't expect it to see on internet platforms like Netflix, YouTube (Opus AAC), Spotify (AAC 128-256k, Vorbis 96/160/320l), Apple Music (AAC 256k), Tidal (96-256 kbps AAC and lossless FLAC) etc.
soresu
11th February 2019, 04:45
As someone who comes from a village in northern England, I can tell you that it only just got upgraded to VDSL from the 3 mbps ADSL 2 it had been at for 5-8 years.
Thankfully it is only 1.5-2 miles from the closest exchange, but many rural communities are much further out than that and still lack the FTTC/VDSL upgrades that have existed near the exchanges for over half a decade (therefore stuck with ultra low ADSL data rates). Expensive 4G mobile broadband data is sadly a bad option if you plan to consume any significant amount of video per month.
All this adds up to the fact that low/ultra low bitrate video is far from corner case, even in first world countries - mainly because rural areas being lower population density are treated like third world countries by BT/Open Reach.
It wouldnt surprise me to find out that many rural places in Europe and the US suffer from similarly slow uptake of landline fibre based broadband technologies.
TomV
11th February 2019, 07:30
This is not true.
Look at the report https://opensignal.com/reports-data/global/data-2018-11/state_of_wifi_vs_mobile_OpenSignal_201811.pdf
Igor, you're arguing with 2 technical professionals who are key members of their respective Tier 1 companies... Amazon and Disney. They have access to much better insights on end-user bandwidth and client device capabilities than you or I. These companies will license proprietary codecs like Dolby AC-4 or xe-AAC if and when that makes sense. Software audio decoding is certainly feasible on most devices, especially at very low bit rates (when audio is also likely mixed to one channel).
I'm glad you have decent bandwidth in Latin America. In many developing areas of the world, bandwidth is still scarce and expensive. And even if mobile networks have been upgraded, that end-customer that a video streaming service is trying to take care of may still have an older device.
iwod
11th February 2019, 07:44
As someone who comes from a village in northern England, I can tell you that it only just got upgraded to VDSL from the 3 mbps ADSL 2 it had been at for 5-8 years.
Thankfully it is only 1.5-2 miles from the closest exchange, but many rural communities are much further out than that and still lack the FTTC/VDSL upgrades that have existed near the exchanges for over half a decade (therefore stuck with ultra low ADSL data rates). Expensive 4G mobile broadband data is sadly a bad option if you plan to consume any significant amount of video per month.
All this adds up to the fact that low/ultra low bitrate video is far from corner case, even in first world countries - mainly because rural areas being lower population density are treated like third world countries by BT/Open Reach.
It wouldnt surprise me to find out that many rural places in Europe and the US suffer from similarly slow uptake of landline fibre based broadband technologies.
I know this is slightly off topic, but I can assure you, comparatively speaking BT isn't doing such a bad job at rural areas. They are actively investing into G.Fast and VDSL 35b. One of the earliest implementor of ADSL2+, ( That is up to 5000M from exchange ). The future is that once 5G matures, setting up Gigabits wireless network using Microwave as backbone should be way cheaper than layering out fibre. So I am optimistic in rural area's broadband.
But yes, ultra low bitrate ( Sub 1Mbps ) is still required in many places, especially if you are doing video which is hogging a lot of the capacity. There is a huge capacity difference between constantly hanging on to a 1Mbps Data stream than doing once in a while 6Mbps Speed test.
So hopefully as both Network Technologies improves and Compression improves, the long tail of world population can all enjoy online streaming video within the next decade. I just hope future codec focus more on sub 2-4Mbps bitrate,
hajj_3
11th February 2019, 12:20
They are actively investing into G.Fast and VDSL 35b
BT are using VDSL2-17A Annex B not 35b.
IgorC
11th February 2019, 13:32
Igor, you're arguing with 2 technical professionals who are key members of their respective Tier 1 companies... Amazon and Disney. .
That's really good. Then they should know that 200-300 kbps is a far from reality.
I have provided a study with real numbers of world bandwidth and I'm myself network specialist who has an information what's going on with ISPs and how xHE-AAC (ultra low bitrate audio format) will change very little if anything as it has already happened with MPEG Surround (standard since 2007). All those professionals were very positive how this MPEG Surround will save bandwidth to million people. Has it?
Phanton_13
11th February 2019, 16:01
I don't think that 200-300 kbps is that far from reality, specially in a mobile phone as even through my phone speed is typilally +30mbps is not thar rare the occasion when it drops down to 1000-500 kbps, basicallly drops in signal quality or being in a overpopulate cell. On the other part more than bandwidth what it save is data usage and as most mobile connections are billed by data usage not by bandwidth this has economical advantage for the user.
I also admit that xhe-aac is mainly a letdown that have seen no real adoption beyond DRM and is understandable for more than one reason.
TomV
11th February 2019, 19:23
Again, the people who run worldwide video streaming services know exactly how many customers are bandwidth limited, and they know which of the many Adaptive Bit Rate renditions (tiers) their customers are streaming. When customers can't sustain higher bit rates, the client player application requests the lowest bit rate rendition from the server. The video service provider has logs of all of this activity, so they know what % of streams are at the lowest bit rate tier, and where/when this happens. Some have tiers as low as 100 kbps (including about 10 kbps for mono audio). With HEVC, this tier is watchable, even if the picture size is reduced to 320x240. It's not possible at all with AVC. For those in rural parts of the world where the best they can get is a 2G mobile connection, they're thrilled to get super-low bit rate video, as long as they can tell what's happening, and they don't have to wait for buffering too often. It sure beats what they used to have, which was no video, or extremely long buffering times.
IgorC
11th February 2019, 19:53
Again, the people who run worldwide video streaming services know exactly how many customers are bandwidth limited,
So do network engineers know (in fact even better)
Again, where is xHE-AAC adoption? It's as old as HEVC. HEVC was adopted, xHE-AAC wasn't.
If ultra low bitrates are so important why nobody hurries to adopt this codec?
Name me just one relatively large broadcasting company who has adopted it.
Name me just one developer team (not Fraunhofer themselves) who actually developing xHE-AAC encoder in this moment.
Can You do it?
Because if You can't I don't see a reason to keep this dicussion.
The video service provider has logs of all of this activity, so they know what % of streams are at the lowest bit rate tier, and where/when this happens. Some have tiers as low as 100 kbps... (including about 10 kbps for mono audio).
Are You sure this number isn't very small? 0.1%, 1-2% ?
Can You present any statistics?
Nintendo Maniac 64
11th February 2019, 20:43
So um, as someone on the outside, I've got to ask - outside of DRM and satisfying "not invented here" syndrome, what benefit does xHE-AAC provide over something like Opus?
benwaggoner
11th February 2019, 21:25
So um, as someone on the outside, I've got to ask - outside of DRM and satisfying "not invented here" syndrome, what benefit does xHE-AAC provide over something like Opus?
It supposedly offers somewhat better quality at very low bitrates. I've not seen a detailed double-blind listening test to validate that, though.
I think xHE-AAC is going to become more broadly supported on platforms, out of momentum. AAC licensees now get access to xHE-AAC for free so it's a trivial effort to roll in xHE-AAC support with platforms updates. And use of AAC in MPEG-4 streams is broadly understood and implemented. Opus does have a mapping, but I've not seem much use of it. MPEG-4 as a file format is more dominant than H.264 was; the Matroska based container formats aren't a significant player for commercially distributed content.
The biggest recent news is that Android Pie has xHE-AAC built in.
xHE-AAC also offers gapless switching between bitrates without having to do overlap decoding. Does Opus support that. I consider this a significant advantage of xHE-AAC over HE v1/v2 and LC, which can only do gapless switching within v1 or v2, but not between. xHE-AAC can switch from very low bitrates to very high quality, which wasn't feasible before.
And don't knock how critical DRM is for premium content. It's a Very Big Deal. And there are SoC reasons why mixing encrypted video and unencrypted audio can be problematic.
Blue_MiSfit
11th February 2019, 21:48
^ DRM is absolutely positively mandatory - no two ways about it. You simply will not get the rights to distribute content if you don't have approved DRM implementations, and this is extremely specific e.g hardware implementations of PlayReady, Widevine with progressive restrictions to unlock HD or UHD content.
I don't love DRM, but it's just table stakes when you're delivering premium content. There's no way around it, so the best we can do is make it as unobtrusive as possible!
Regarding the 200 - 300 Kbps scenario - another use case would be rural customers with satellite or very poor cell service. I grew up in a very small town and many of my friends live outside the city limits where there simply is no broadband. Not even DSL, though you could maybe get an ISDN or T1 line if you're a masochist :D
You might get one bar of LTE (or two if you stand in exactly the right place), and you're sharing this one solitary tower with your neighbors, so during peak times you're lucky to get 1 Mbps, assuming you're not over your data cap, at which point you drop down to under 500 Kbps. Satellite can be fast, but generally is very over-sold in these areas, and cannot keep up with demand during peak times. It also has extreme data caps, and the throttled speed is extremely slow.
Anyway - hoping we can get back on topic - AOM codec discussion :)
Maybe viewing through the ultra low bitrate lens, has anyone done very low bitrate 2 pass VBR tests with AV1? I'd be interested to see how that might fit into the above scenario regarding downloading content for offline playback. This is a neat scenario because you get the bonus of being able to skip all the compromises one must make when encoding for adaptive bitrate delivery and can use very large buffers (vbv-maxrate + vbv-bufsize) and longer adaptive keyframe intervals.
benwaggoner
11th February 2019, 22:29
^ DRM is absolutely positively mandatory - no two ways about it. You simply will not get the rights to distribute content if you don't have approved DRM implementations, and this is extremely specific e.g hardware implementations of PlayReady, Widevine with progressive restrictions to unlock HD or UHD content.Yeah. It is table stakes for anything that isn't piracy or user-generated content. If AV1 matters outside of social networks, it'll be because it gets 1st class DRM support, which requires HW decoders and other deep SoC integration. Weird little SoC DRM design decisions have kept a lot of amazing things from happening.
Maybe viewing through the ultra low bitrate lens, has anyone done very low bitrate 2 pass VBR tests with AV1? I'd be interested to see how that might fit into the above scenario regarding downloading content for offline playback. This is a neat scenario because you get the bonus of being able to skip all the compromises one must make when encoding for adaptive bitrate delivery and can use very large buffers (vbv-maxrate + vbv-bufsize) and longer adaptive keyframe intervals.
It's not so much a long buffer window, but having a maxrate>>bitrate. Maxrate and bufsize can be kept at the profile @ level maximums while ABR can be way way lower. Doing maximum 10 sec Open GOP is pretty reasonable.
I worry that AV1's interframe CABAC dependencies will impair random access enough to make the practical maximum GOP duration a lot smaller. Some of that could probably be addressed via encoder tweaks, at the loss of a little efficiency.
Phanton_13
11th February 2019, 22:37
I think that I misled wen I reference DRM, I was referencing to Digital Radio Mondiale where xhe-aac is one of the codecs used.
The lowest that I went with AV1 is q50 for sd anime and is surprisingly watchable (resulting in about 100 to 200 kbps of video bitrate, the audio was opus at 32kbps). Definitely CDEF is a great tool. Around 30MB per episode... I wish that this quality-bitrate ratio was available 20 years ago.
IgorC
12th February 2019, 00:24
So um, as someone on the outside, I've got to ask - outside of DRM and satisfying "not invented here" syndrome, what benefit does xHE-AAC provide over something like Opus?
Both Opus and xHE-AAC are hybrid music/speech formats and have essentially same quality. xHE-AAC is slightly better than Opus at 16-32 kbps, quality at 48-64 kbps is the same for both and Opus is slightly better than all family LC-AAC/HE-AAC/xHE-AAC at higher bitrates.
Many people try to market a feature of xHE-AAC as switching between bitrates as something outstanding and so on. But in reality it's not a premium feature and Opus supports it from the very beginning. More here https://wiki.hydrogenaud.io/index.php?title=Opus
Also xHE-AAC is a high delay codec and it's not suitable for real time communications like VOIP calls etc. Company who wants real time communication should also adopt low-delay xHE-AAC fork (EVS). And if You want stereo EVS then this is just another extension. So we got mutltiple codecs and/or extensions:
1.xHE-AAC
2.low delay EVS
3.EVS stereo extension (aka IVAS)
...
It's sort of LC-AAC, HE-AAC, HE-AACv2, Low delay AAC (LD-AAC), Enhanced LD-AAC ( low delay HE-AAC) aka ELD-AAC, ELD-AAC v2 (HE-AACv2 low delay), xHE-AAC, EVS ( low delay xHE-AAC), IVAS (low delay stereo xHE-AAC)...
While Opus is just one single format for everything: high-,low-delay, stereo and multichannel codec.
So I'm not surprised that Opus is popular in internet community while xHE-AAC support is non-existent. Android 9 will get an xHE-AAC decoder but there is still no any encoder
Tommy Carrot
12th February 2019, 00:58
Maybe viewing through the ultra low bitrate lens, has anyone done very low bitrate 2 pass VBR tests with AV1? I'd be interested to see how that might fit into the above scenario regarding downloading content for offline playback.
Very low bitrate is definitely the strong point of AV1 (more precisely aomenc). At around crf 28-30 compression rates, it's starting to get better than x265, and the lower the bitrate, the bigger the advantage. XVC, and especially VVC are still better though, the VVC reference encoder is seriously amazing at ultra low bitrate scenarios.
benwaggoner
12th February 2019, 01:21
Very low bitrate is definitely the strong point of AV1 (more precisely aomenc). At around crf 28-30 compression rates, it's starting to get better than x265, and the lower the bitrate, the bigger the advantage. XVC, and especially VVC are still better though, the VVC reference encoder is seriously amazing at ultra low bitrate scenarios.
Ooh, do you have any examples or command lines for the >x265 for very low bitrates? Low bitrate, low resolution video enables MUCH faster encode/evaluate/reencode cycles and would be a great place to evaluate AV1's potential strengths.
Tommy Carrot
12th February 2019, 01:41
Ooh, do you have any examples or command lines for the >x265 for very low bitrates? Low bitrate, low resolution video enables MUCH faster encode/evaluate/reencode cycles and would be a great place to evaluate AV1's potential strengths.
Nothing special, i pretty much use the default settings.
aomenc --end-usage=q --cq-level=xx --cpu-used=1 --kf-max-dist=250 -v -o av1.webm test.y4m
I haven't tried 2-pass bitrate mode though, only CQ mode with using the quantizer with the closest bitrate.
Mr.Radar
14th February 2019, 21:17
I think that I misled wen I reference DRM, I was referencing to Digital Radio Mondiale where xhe-aac is one of the codecs used.
For those who are unaware, Digital Radio Mondiale is the digital broadcasting standard used on the shortware radio bands. Due to the narrow bandwidth of shortwave broadcast channels and the high amount of error correction required to avoid dropouts the usable bitrates are very low. Here is a 2017 comparison between xHE-AAC and Opus at those bitrates (https://www.youtube.com/watch?v=sEUwScTX-R4). More recent versions of libopus should produce significantly more competitive results in the speech-focused samples (libopus 1.3 enabled wideband audio down to 9 kbps) but on music samples xHE-AAC would probably retain the edge since Opus at very low bitrates is effectively operating mostly as a speech codec in its SILK mode.
Mr_Khyron
15th February 2019, 00:28
https://github.com/OpenVisualCloud/SVT-AV1/pull/55
Over 50% reduction in memory requirements
:)
soresu
15th February 2019, 08:45
It says "low core count (e.g. 4-core machines)" though, does that mean it doesnt apply at all to higher core counts, or that they alreeady have a similar optimisation implemented or in the pipe?
nevcairiel
15th February 2019, 09:23
It probably means that memory usage didn't scale properly to different core counts. The more cores you run, the more memory its going to need, that is not unexpected from any encoder. :)
Mr_Khyron
16th February 2019, 01:42
https://phoronix.com/scan.php?page=news_item&px=SVT-AV1-Speed-Progress
It was just a few weeks ago that Intel open-sourced the SVT-AV1 project as a CPU-based AV1 video encoder. In the short time since publishing it, there's already been some significant performance improvements.
Since the start of the month, SVT-AV1 has added multi-threaded CDEF search, more AVX optimizations, and other improvements to this fast evolving AV1 encoder. With having updated the test profile against the latest state as of today, here's a quick look at the performance of this Intel open-source AV1 video encoder.
Nintendo Maniac 64
16th February 2019, 03:14
"low core count (e.g. 4-core machines)"
That's certainly amusing considering that Intel seemed to love branding quad cores as i7 up until recently. :p
Speaking of quad core i7 CPUs...
https://phoronix.com/scan.php?page=news_item&px=SVT-AV1-Speed-Progress
I'm going to go out on a limb and guess that, on the SVT-AV1 v2019-02-15 results, the lack of performance delta seen between the i7-7740X and Ryzen 2700X is due to the former's considerably faster AVX2 implementation even though the Ryzen has twice as many cores and threads as the i7?
Selur
16th February 2019, 07:38
btw. https://ci.appveyor.com/project/OpenVisualCloud/SVT-AV1/build/artifacts offers Windows binaries of the stv-av1 encoder (SvtAv1EncApp.exe and SvtAv1EncApp.dll are needed)
Blue_MiSfit
16th February 2019, 21:33
Very cool. I'm playing with this now.
Default settings are quite fast - it did 12 fps for 480p on my i7 7700k with 50% CPU usage and 1.2 GB RAM usage.
Here's the basic usage guide:
https://github.com/OpenVisualCloud/SVT-AV1/blob/master/Docs/svt-av1_encoder_user_guide.md
A sample command:
.\SvtAv1EncApp.exe -i .\beauty_480p.yuv -b out.ivf -w 848 -h 480 -fps 24
Yes, I had to encode at 848x480 and not 854 - comically this encoder requires mod 8 input :)
Results aren't too bad - the default CQ 50 setting produced a 623 Kbps file that looks marginally okay, and CQ 40 produced a 1.2 Mbps file that looks a lot better. There are some odd artifacts, almost like the edges of objects wiggle a bit every other frame.
However, I'm getting very jerky playback for some reason. I've tried a LAV nightly build in MPC-HC, latest VLC, and ffplay, and they all show the issue. I confirmed my input YUV is 24 fps and the output webm is 24 fps. When I step through it frame by frame it's all there, but for some reason it's jerky on playback. I don't recall seeing this with aomenc encodes, and the decoder isn't close to maxing out one core, so I'm not sure what's up with that...
tnti
18th February 2019, 02:20
Very cool. I'm playing with this now.
Default settings are quite fast - it did 12 fps for 480p on my i7 7700k with 50% CPU usage and 1.2 GB RAM usage.
Here's the basic usage guide:
https://github.com/OpenVisualCloud/SVT-AV1/blob/master/Docs/svt-av1_encoder_user_guide.md
A sample command:
.\SvtAv1EncApp.exe -i .\beauty_480p.yuv -b out.ivf -w 848 -h 480 -fps 24
Yes, I had to encode at 848x480 and not 854 - comically this encoder requires mod 8 input :)
Results aren't too bad - the default CQ 50 setting produced a 623 Kbps file that looks marginally okay, and CQ 40 produced a 1.2 Mbps file that looks a lot better. There are some odd artifacts, almost like the edges of objects wiggle a bit every other frame.
However, I'm getting very jerky playback for some reason. I've tried a LAV nightly build in MPC-HC, latest VLC, and ffplay, and they all show the issue. I confirmed my input YUV is 24 fps and the output webm is 24 fps. When I step through it frame by frame it's all there, but for some reason it's jerky on playback. I don't recall seeing this with aomenc encodes, and the decoder isn't close to maxing out one core, so I'm not sure what's up with that...
https://github.com/OpenVisualCloud/SVT-AV1/issues/33
TD-Linux
21st February 2019, 23:16
I worry that AV1's interframe CABAC dependencies will impair random access enough to make the practical maximum GOP duration a lot smaller. Some of that could probably be addressed via encoder tweaks, at the loss of a little efficiency.
You don't need to worry - AV1 probability dependencies can only come from one of the reference frames, so it doesn't place any additional impairment on seekability.
(also note that CABAC is a misnomer, as it's not binary. The spec doesn't give it an acronym, but dav1d uses MSAC).
Beelzebubu
22nd February 2019, 03:01
You don't need to worry - AV1 probability dependencies can only come from one of the reference frames, so it doesn't place any additional impairment on seekability.
I think his concern that a decoder (or a stupid player, which is 99.9% of them) don't know this. A good container format (like mp4) can represent the reference structure in its atoms, and then a good decoder + good container + good encoder can do the right thing. But if any one of them fails, you'll have a worse seeking experience if you want to do frame-exact user experience. It's up to all devs to make sure that doesn't happen, and like I said, this is multi-factorial so it's easy to forget and screw up.
benwaggoner
22nd February 2019, 21:41
I think his concern that a decoder (or a stupid player, which is 99.9% of them) don't know this. A good container format (like mp4) can represent the reference structure in its atoms, and then a good decoder + good container + good encoder can do the right thing. But if any one of them fails, you'll have a worse seeking experience if you want to do frame-exact user experience. It's up to all devs to make sure that doesn't happen, and like I said, this is multi-factorial so it's easy to forget and screw up.
Yeah, the goal is for the decoder to determine the minimum sequence of frames required to decode a particular frame. With a IbbbBbbbPbbbBbbbPbbbbBbbbI kind of structure, decoding the last "b" frame in an Open GOP should require just six frames (IPPBIb) in decode order typically. But that requires the reference list because sometimes a b could reference two B frames back and things like that. With multiple reference frames it's impossible to reliably know the hierarchy without knowing what each frame references.
And there are patterns that can be spec-legal but that existing encoders might not do. And then better encoders add those to improve quality.
TD-Linux
22nd February 2019, 22:37
Yeah, the goal is for the decoder to determine the minimum sequence of frames required to decode a particular frame. With a IbbbBbbbPbbbBbbbPbbbbBbbbI kind of structure, decoding the last "b" frame in an Open GOP should require just six frames (IPPBIb) in decode order typically. But that requires the reference list because sometimes a b could reference two B frames back and things like that. With multiple reference frames it's impossible to reliably know the hierarchy without knowing what each frame references.
Yeah, if you want to do that you'll need a reference list parser. But that's been true for a very long time - even x264 produces streams that require you to do this.
But regardless, whatever structure you pick, the CDFs always follow that same structure, so they are "free" from a seekability point of view.
benwaggoner
23rd February 2019, 00:31
Yeah, if you want to do that you'll need a reference list parser. But that's been true for a very long time - even x264 produces streams that require you to do this.
But regardless, whatever structure you pick, the CDFs always follow that same structure, so they are "free" from a seekability point of view.
That's good news. I had heard suggestions otherwise.
sneaker_ger
24th February 2019, 23:29
Seems like dav1d is now like 40% faster than libaom on my i5-2500K (no AVX2). And there's a pull request for additional SSSE3 optimizations with like another 50% speedup.
https://code.videolan.org/videolan/dav1d/merge_requests/599#note_29705
nevcairiel
25th February 2019, 00:06
Once the CDEF patch is merged, it should definitely beat AOM in nearly all situations, in 8-bit decoding anyway. 10-bit/12-bit hasn't been worked on much yet, performance wise.
sneaker_ger
26th February 2019, 15:32
CDEF patch is merged now.
nevcairiel
26th February 2019, 15:53
Here is an up-to-date performance comparison of aomdec vs dav1d (I believe some of the tests are still running as I write this)
https://docs.google.com/spreadsheets/d/1rkPMHgy7cXEsT9KeYF-NVZQNiEGEkvrLnLo2FaLWiwA/edit#gid=1835354908
Single-Threading, it still loses in some cases on SSE 4.1 CPUs (aom implements mostly SSE4.1 for pre-AVX), but Multi-Threading it makes up for that.
One missing part is prep_8tap (https://code.videolan.org/videolan/dav1d/merge_requests/604), which can have a decent impact.
Their goal for releasing dav1d 0.2.0 is to be faster then aomdec in 8-bit in all targeted scenarios.
sneaker_ger
26th February 2019, 17:03
There are some user benchmarks on an iPhone (https://code.videolan.org/videolan/dav1d/issues/15#note_29792) with crazy results.
https://i.imgur.com/tqR6ttB.png
1080p with 46 to 74 fps. Years ago we told people they would never be able to watch new codecs on a phone if it didn't have ASIC for the codec ...
nevcairiel
26th February 2019, 17:27
Years ago we told people they would never be able to watch new codecs on a phone if it didn't have ASIC for the codec ...
They really still shouldn't want to/have to, since it'll make the phone hot and the battery empty. :)
benwaggoner
26th February 2019, 18:00
There are some user benchmarks on an iPhone (https://code.videolan.org/videolan/dav1d/issues/15#note_29792) with crazy results.
https://i.imgur.com/tqR6ttB.png
1080p with 46 to 74 fps. Years ago we told people they would never be able to watch new codecs on a phone if it didn't have ASIC for the codec ...
Fun way to amaze yourself - figure out how long ago you first got a computer that was more powerful than your current phone (benchmark/metric of your choice). A modern iPhone has lots of fast cores with SIMD instructions and a pretty darn powerful GPU. It would smoke any hot gaming rig of 10 years ago, and a typical laptop of 5 years ago.
Generally each new codec generation aims to be more than twice as complex to decode as the prior generation. With Moore’s Law gains, computing devices get a LOT more than 2x faster in that time.
The bigger challenge with software codecs is getting them integrated into hardware DRM required to play premium content above 480p.
It’s exciting to see these perf gains with AV1. Hopefully we’ll get to a reasonably “done” version of dav1d later this year so we can ballpark decoder requirements for AV1 versus other codecs. I’m particularly interested in how much parallelism is possible. VP9 was nearly single-threaded, which was problematic
A big question for AV1’s viability is how many extra transistors (and thus how much extra die size and SoC cost) full HW decode will take. I’ve not heard any details from anyone who has taped out an implementation yet. Anyone else?
NikosD
26th February 2019, 18:16
Single-Threading, it still loses in some cases on SSE 4.1 CPUs (aom implements mostly SSE4.1 for pre-AVX), but Multi-Threading it makes up for that. Unfortunately, in single-threading mode, dav1d looses even using SSSE3 in Dua Lipa clip.
It needs further optimization for pre-AVX2 SIMD, but the multi-threading performance especially for AVX2 is impressive.
EwoutH
26th February 2019, 20:10
@nevcairiel You're fast :)
That bench included !604 (prep_8tap), so this is most likely the performance as it's going to look like at 0.2.0 release. Not all prep_8tap functions are converted to ssse3 yet, so there is potentially another 5-12% speedup possible for ssse3, highly depending on content.
iwod
28th February 2019, 08:34
https://aomediacodec.github.io/av1-avif/
v1.0.0, 19 February 2019 but page still says it is a draft.
That is the Image format AVIF you are looking at. Not sure if that is what you want.
Q3CPMA
28th February 2019, 10:45
Personally, I find AVIF a lot more interesting than AV1 seeing as JPEG is technologically much more obsolete than AVC (if you don't consider 4K, of course). I just hope they don't force chroma subsampling like on webp.
hajj_3
28th February 2019, 13:01
Personally, I find AVIF a lot more interesting than AV1 seeing as JPEG is technologically much more obsolete than AVC (if you don't consider 4K, of course). I just hope they don't force chroma subsampling like on webp.
Jpeg XL will be ratified later this year, that should be much nicer than AVIF.
Rumbah
28th February 2019, 22:05
Jpeg XL will be ratified later this year, that should be much nicer than AVIF.
Will it be royalty free?
Zebulon84
1st March 2019, 05:40
In this document (Final Call for Proposals for a Next-Generation Image Coding Standard (JPEG XL)) (https://jpeg.org/downloads/jpegxl/jpegxl-cfp.pdf) issued in last April, it is stated page 3: This new JPEG activity aims to develop a new image coding standard that provides state-of-the-art image compression performance, and that addresses shortcomings in current standards. To encourage widespread adoption, an important goal for this standard is to support a royalty-free baseline
So probably partially. But will il be enough to have wider adoption than previous standards (JPEG 2000, JPEG XR...) ?
mzso
1st March 2019, 23:27
They really still shouldn't want to/have to, since it'll make the phone hot and the battery empty. :)
There's a solution. :)
https://cdn.vox-cdn.com/thumbor/1HzHTQhjVUrhVDHQC8k1hdM9HPg=/0x0:2040x1360/1200x800/filters:focal(913x433:1239x759)/cdn.vox-cdn.com/uploads/chorus_image/image/63124099/energizer_unit_vsavov7.0.jpg
Jpeg XL will be ratified later this year, that should be much nicer than AVIF.
Yay! Two more formats to sink into obscurity...
hajj_3
2nd March 2019, 02:08
It looks like the dav1d decoder v0.2.0 has been released: https://code.videolan.org/videolan/dav1d/commit/0e55d462957323119b9e90b203f534b9fe6f169a
marcomsousa
2nd March 2019, 07:31
It looks like the dav1d decoder v0.2.0 has been released: https://code.videolan.org/videolan/dav1d/commit/0e55d462957323119b9e90b203f534b9fe6f169a
Isn't release yet, just a change for the next release.
Wait until there is a 0.2.0 tag and/or release.
https://code.videolan.org/videolan/dav1d/tags
marcomsousa
2nd March 2019, 08:07
The version has already bumped to 0.2.0 (https://code.videolan.org/videolan/dav1d/commit/e811c4767d0c698fc674e9603e1e63d15acff16e)
Are you sure that 0.2.0 is finished?
There isn't a tag neither a release version.
https://code.videolan.org/videolan/dav1d/tags
https://code.videolan.org/videolan/dav1d/releases
That change was in 25/02 (https://code.videolan.org/videolan/dav1d/commit/e811c4767d0c698fc674e9603e1e63d15acff16e) and the other was 13 hours ago (https://code.videolan.org/videolan/dav1d/commit/0e55d462957323119b9e90b203f534b9fe6f169a).
I think that the final version of 0.2.0 isn't release yet, but it's seems that will be release in the follow hours/days.
But I confirm that this is a litle strange. They have a version without being finished. That it seems that is the way before, they don't use -dev neither -snapshot.
nevcairiel
2nd March 2019, 09:14
The release is scheduled for Monday currently. The version bump means nothing, its only in preparation for the release, only tags really matter. You would typically do that a few days earlier to ensure everything comes out as expected.
Its not really that strange either, you guys are just overthinking it.
Wolfberry
2nd March 2019, 09:46
The version bump means nothing.
Just like this commit (https://code.videolan.org/videolan/dav1d/commit/0c173fd19326435e575d884b6f40ca135aed1885)
It seems like 0.1.1 is just a transition from 0.1.0 to 0.2.0 and will never get released. But I can be wrong, anyway.
nevcairiel
2nd March 2019, 10:35
0.1.1 was never released indeed, the changelog entries being created for it have been recycled into the 0.2.0 changelog.
sneaker_ger
2nd March 2019, 12:03
The version has already bumped to 0.2.0 (https://code.videolan.org/videolan/dav1d/commit/e811c4767d0c698fc674e9603e1e63d15acff16e)
libaom 1.0.0-1402-g442f429de
libdav1d 0.2.0-493155af
(https://drive.google.com/open?id=1TrR10EL8nDD4WO2nHhHvaSKVhNaWTmcC)
Now dav1d is like 55% faster on i5-2500K (AVX, no AVX2) than libaom on Blackmagic Pocket Cinema "Nature" sample 1080p24. :cool:
v0lt
3rd March 2019, 17:03
ffmpeg-4.2-N-93276-g3b23eb283a-win64-static (https://drive.google.com/open?id=11-RIGTnNB5jjHrO4puotRHsNu_YdGMKr)
Test on this build. i5-3570K.
-t 10 Stream2_AV1_4K_22.7mbps.webm -benchmark -f null -
libaom-av1 - max 15 fps
libdav1d - max 13 fps
libdav1d -threads 4 -tilethreads 4 - max 18 fps
libdav1d -threads 8 -tilethreads 1 - max 18 fps
I see progress compared to last time (https://forum.doom9.org/showthread.php?p=1859851#post1859851).
Please give me a link to a good 10-bit AV1 sample.
sneaker_ger
3rd March 2019, 17:08
Please give me a link to a good 10-bit AV1 sample.
http://download.opencontent.netflix.com/?prefix=AV1/Chimera/
nevcairiel
3rd March 2019, 18:18
Note that 10-bit is not an optimization target at all yet, since its not in use in the real world at all yet either.
v0lt
3rd March 2019, 19:00
@nevcairiel
Why does libdav1d give a speed less than libaom-av1 with default settings? 8 and 10-bit. Poorly selected default settings?
i5-3570K, AV1 10-бит.
-ss 5 -t 20 Chimera-AV1-10bit-1920x1080-6191kbps.mp4 -benchmark -f null -
libaom-av1 - max 25 fps
libdav1d - max 21 fps
libdav1d -threads 4 -tilethreads 4 - max 25 fps
libdav1d -threads 8 -tilethreads 1 - max 26 fps
soresu
3rd March 2019, 20:25
Dunno if anyone noticed the CNN restoration experiment commited to the "experimental" branch at the AOMedia git repo.
Link here (https://aomedia.googlesource.com/aom/+log/refs/heads/experimental).
EwoutH
3rd March 2019, 22:03
@v0lt Which version of dav1d are you using? You can use `dav1d --version` to check.
A big factor is because your CPU doesn't support AVX2 instructions, where dav1d is most optimized for (Haswell (4000-series) and newer support it). Your CPU i5-3570K falls back on SSSE3, which recently got a lot faster (https://www.reddit.com/r/AV1/comments/awxvks/ssse3_performance_improvement_of_dav1d_blogpost/) with dav1d.
Also can you try -tilethreads 4 without a threads option?
Anyway there is at least one patch incoming (https://code.videolan.org/videolan/dav1d/merge_requests/622) that's not yet in FFmpeg that will improve performance with 10 to 30% percent, depending on content.
dapperdan
3rd March 2019, 22:21
I interpreted volts question as being "if I can get faster results from dav1d by passing some settings, why can't it figure those out by default".
Presumably, if those patches land the default might beat libaom, but there would still be extra speed a ilavle as long as you knew the right options to pass.
There may well be a good reason for these not being the defaults though.
Gravitator
4th March 2019, 10:43
Through the synthetic test mode ffmpeg, I get higher rates of DAV1D than AOM. But in real conditions I get worse playing than AOM (via ffplay). We are waiting for optimization for SSSE3.
sneaker_ger
4th March 2019, 15:15
Also can you try -tilethreads 4 without a threads option?
On i5-2500K with ffmpeg (binary from one post above) -tiletreads 4 (alone) is some 15% slower than -threads 6 -tilethreads 2. ("Blackmagic Pocket Cinema Camera 4K ‘Nature’-oAbB4dQOz4I" 1080p24)
v0lt
4th March 2019, 17:21
@v0lt Which version of dav1d are you using? You can use `dav1d --version` to check.
ffmpeg-4.2-N-93276-g3b23eb283a-win64-static provided by Wolfberry. But the links are gone.
ffmpeg-4.2-N-93290-g88d0be1c0e-win64-static (https://drive.google.com/open?id=1jpnMkcW8j1UPD4StPDt9ZwFtwNTXnD1C)
Thank. This build works better.
i5-3570K
ffmpeg -hide_banner -t 10 -c:v <codec¶meters> -i Stream2_AV1_4K_22.7mbps.webm -benchmark -f null -
libaom-av1 - max 15 fps
libdav1d - max 16 fps
libdav1d -threads 4 -tilethreads 4 - max 24 fps
libdav1d -threads 8 -tilethreads 1 - max 21 fps
ffmpeg -hide_banner -ss 5 -t 20 -c:v <codec¶meters> -i Chimera-AV1-8bit-1920x1080-6736kbps.mp4 -benchmark -f null -
libaom-av1 - max 60 fps
libdav1d - max 64 fps
libdav1d -threads 4 -tilethreads 4 - max 105 fps
libdav1d -threads 8 -tilethreads 1 - max 87 fps
ffmpeg -hide_banner -ss 5 -t 20 -c:v <codec¶meters> -i Chimera-AV1-10bit-1920x1080-6191kbps.mp4 -benchmark -f null -
libaom-av1 - max 25 fps
libdav1d - max 21 fps
libdav1d -threads 4 -tilethreads 4 - max 26 fps
libdav1d -threads 8 -tilethreads 1 - max 27 fps
benwaggoner
4th March 2019, 18:58
There's a solution. :)
https://cdn.vox-cdn.com/thumbor/1HzHTQhjVUrhVDHQC8k1hdM9HPg=/0x0:2040x1360/1200x800/filters:focal(913x433:1239x759)/cdn.vox-cdn.com/uploads/chorus_image/image/63124099/energizer_unit_vsavov7.0.jpg
Yay! Two more formats to sink into obscurity...
The time **IS** ripe for a new replacement for JPEG/GIF/PNG. We're looking at >>2x compression efficiency improvements. And HEIF is getting good traction in the Apple ecosystem.
I would prefer an AV1 in HEIF over a new file format, though. HEIF has a lot of great features as a container format.
nevcairiel
4th March 2019, 22:50
HEIFs tiling is pure cancer. It's a typical format designed by committee, overly complex for its own good.
sneaker_ger
5th March 2019, 10:34
dav1d 0.2.0 is final.
https://medium.com/@ewoutterhoeven/dav1d-0-2-0-covering-all-pcs-including-mobile-eac3e43868c2
hajj_3
5th March 2019, 10:34
DAV1D 0.2.0 final is out: https://code.videolan.org/videolan/dav1d/tags
Hopefully a new vlc player will be released soon with this integrated.
utack
5th March 2019, 15:16
DAV1D 0.2.0 final is out: https://code.videolan.org/videolan/dav1d/tags
Hopefully a new vlc player will be released soon with this integrated.
Not yet
You can check here
https://git.videolan.org/?p=vlc.git;a=blob;f=contrib/src/dav1d/rules.mak;hb=HEAD
http://downloads.videolan.org/pub/videolan/dav1d/
hin12
7th March 2019, 03:45
Is 480p the highest quality option for all AV1 videos on YouTube? I'm seeing that a lot of popular videos are being encoded to AV1, but only at the 144p, 240p, 360p and 480p resolutions.
sneaker_ger
7th March 2019, 10:03
Yeah. But even Gangnam Style (3.3 Billion views) is only available in 720p with AV1 (1080p is H.264+VP9 only.), Despacito (6 Billion views) is limited to 480p AV1. On the AV1 beta playlist it seems 1080p is max for AV1, 1440p and 2160p are reserved for VP9.
dapperdan
7th March 2019, 13:13
When VP9 rolled out they talked about how delivering video to people with decent spec machines but terrible connectivity was a sweet spot where VP9 increased the amount of video watched. I'd guess they've run the numbers and they get more benefit from encoding the lower sizes in AV1 so that's what they prioritize.
lvqcl
7th March 2019, 16:08
Some time ago I was able to download "Childish Gambino - Feels Like Summer" (https://www.youtube.com/watch?v=F1B9Fk_SgI0) in 1080p AV1 format; now its max. available resolution for AV1 is 720p.
Maybe 1080p AV1 is just too heavy to decode.
sneaker_ger
7th March 2019, 16:11
Or maybe they are re-encoding with new encoder version/settings and it will take some weeks before it's finished. :devil:
benwaggoner
7th March 2019, 17:57
Some time ago I was able to download "Childish Gambino - Feels Like Summer" (https://www.youtube.com/watch?v=F1B9Fk_SgI0) in 1080p AV1 format; now its max. available resolution for AV1 is 720p.
Maybe 1080p AV1 is just too heavy to decode.
It would have been in many cases. Perhaps it'll come back when Chrome has a highly optimized decoder.
Also, the CPU requirements for encoding 1080p are 2.5.x that for 720p. Perhaps quality compromises were made to get publishing time to be reasonable that reduced competitiveness, or there were limits on available encoder capacity?
A 1080p H.264 encode takes >>100x less compute than a 1080p AV1. Even Google only has so much free CPU capacity in a day.
Encoder performance improvements plus better quality/speed tradeoff modes will change things dramatically. With continued strong development at the current pace, we could potentially have encoders that can provide better quality than x264 at the same bitrate and encoding time by late 2020. That encoder wouldn't even try to do 99.9% of the stuff that libaom tries to do, but real encoders don't try exhaustive searches, but use lots of heuristics and early exits.
mandarinka
8th March 2019, 19:56
I ran a new test of Rav1e on the same source as in this past post last year: https://forum.doom9.org/showpost.php?p=1850675&postcount=882
(the motivation was a claim that Rav1e supposedly started to beat x264 "at any bitrate" so I thought I could as well repeat my highly specific test - note that this is by no means supposed to be authoritative. Unless you care about this type of content which I do.)
I used the last Rav1e release available at the time (https://github.com/xiph/rav1e/releases/tag/20190219), at quality 20 and speed 0, tune psy. There are no other tuning options, so I could not do any tweaks - not my fault, that's the developers' policy. Resulting bitrate was 14526kbps. Encoding speed was 0.005 fps on Ryzen 3 2200G (~1 core used).
My x264 commandline is kind of tuned for this content - I copied what I last used on a similar bluray (but I dropped --qpmax which probably worsens efficiency). Note that three pass was used to get closer to Rav1e's bitrate but due to the short length probably, x264 still undershot (14337 kbps). Using three pass should not give quality boost.
The encoding speed was about 0.05 fps due to extremely placebo settings and 1 thread used (same as Rav1e). You could probably get 0.3 without much if any damage to quality.
Here are images from the encode: http://imgbox.com/g/yrCWrTYtF4
In the dust storm scene that uses higher bitrate than the rest and is at the start of the 910frame clip, Rav1e does well and it seems to be very close - I am not totally convinced it is as good as x264 as I think there are some signs of kinda low-passing the noisy blocks, but I am not completely confident it is worse either, although I am inclined to say x264 is in fact better here.
The rest of the clip has Rav1e clearly deficient despite the large bitrate though. It constantly smoothes the solid/flat areas and generally can't keep the milder grain and texture (which is a flaw). So its psychovisual decisions probably still aren't ready for high quality transparent encoding. This is a general problem with any non x264/x265 encoder probably, even x265 would drop texture detail everywhere before it got aq and psyrdo.
I'm trying to upload the source and streams (not fun to reproduce the Rav1e one hah), but uloz.to seems to fail on me. One thing that I should perhaps note about source - it has its chroma temporally denoised (not luma). In case that particularly matters for rav1e rate control... it only has constant quantizer mode though, afaik.
Commandlines:
rav1e.exe --speed 0 --quantizer 20 --tune Psychovisual n:\test.y4m -o psy-q20-speed0.ivf
x264-2935-64.exe n:\etr-testLL.mkv --qcomp 0.60 --aq-mode 1 --pass 3 --bitrate 14526 --no-mbtree --min-keyint 5 --keyint 240 --b-pyramid normal --deblock 0:0 --psy-rd 1.0:0.0 --aq-strength 0.8 --preset placebo --bframes 9 --direct auto --me tesa --merange 64 --subme 11 --threads 1 --colormatrix bt709 --sar 1/1 --chromaloc 0 --input-range tv --range tv --force-cfr --fps 24000/1001 -o n:\test-etr-x264-10bit-noqpmax.mkv --output-depth 10 --stats n:\noqpm.txt
Atak_Snajpera
9th March 2019, 13:44
What is the point of using AV1 with such insanely (15Mbps) high bitrate for anime? Why not use XviD or even MPEG-2 if you have such high bitrate budget?
mandarinka
9th March 2019, 14:12
What is the point of using AV1 with such insanely (15Mbps) high bitrate for anime? Why not use XviD or even MPEG-2 if you have such high bitrate budget?
Transparent quality? I don't know what you expect from mpeg2 or xvid here but they would not give it to you. After all, the pictures clearly show that even 14 megabits isn't enough when encoder doesn't do the right decisions.
In any case, the point of it is stated in the post - it was claimed by certain people (and it was not just internet randoms) that Rav1e beats x264 at any bitrate, already. I wanted to test that claim.
(It's also a realistic case for me, but I would not use unfinished/untuned encoders normally for that, of course.)
Atak_Snajpera
9th March 2019, 14:14
Transparent quality? I don't think what you expect frommpeg2 or xvid here but they would not give it to you.
In any case, the point of it is stated in the post - it was claimed by certain people (not just internet randoms) that Rav1e beats x264 at any bitrate already. I wanted to test that claim.
(It's also a realistic case for me, but I would not use new/untuned encoders normally for that, of course.)
OK. I will have to check how bad is this codec in parkjoy at 5Mbps. I'm expecting similar disaster to that crappy SVT-HEVC encoder. ;)
mandarinka
9th March 2019, 14:18
Parkjoy might benefit from the compression strength advantage of the format, same as it usually shows in low bitrate tests. Rav1e has no AQ yet though so that will probably be a big disadvantage because parkjoy iirc benefited a lot from it?
But you would not run into this "can't have transparent dirt on a flat color area in cel anime" issue, that's for sure.
Wolfberry
9th March 2019, 14:20
I recommend to use the official AppVeyor builds here: https://ci.appveyor.com/project/tdaede/rav1e/history
mandarinka
9th March 2019, 14:41
I recommend to use the official AppVeyor builds here: https://ci.appveyor.com/project/tdaede/rav1e/history
At the time I went to #aomedia to ask for these binaries (I recall I could not find the artifact button that leads to the binary there/it was not working atm... not sure now) and somebody there recommended to get the release one. So I did that.
Wolfberry
9th March 2019, 14:50
The pre-release is just a weekly snapshot and the latest one is already kinda old, so...
mandarinka
9th March 2019, 14:52
Yeah, obviously. At the time I ran this test though, it was just 9 or 10 days old. You have to remember I used speed 0, so just getting 910 frames encoded took three or four days (PC hibernated overnight). The following week I was kind of busy which added more delay to this post. So it wouldn't really be a big difference if I used appveyor then.
Wolfberry
9th March 2019, 15:16
AFAIK, assembly is disabled on windows at the moment, so you are basically running on rust code.
I am recommending AppVeyor builds in case someone want to try it out now.
Atak_Snajpera
9th March 2019, 16:03
OH MY GOD! The encoding speed is just INSANELY SLOOOOOOOOOOOOOOOOOOOOOOOW. 90 minutes have passed and my 8C/16T managed to encode only 350 frames from just first pass.
How can you even test any codec if encoding takes so much time?
mandarinka
9th March 2019, 16:17
AFAIK, assembly is disabled on windows at the moment, so you are basically running on rust code.
Are you sure? I was asking about the speed in the channel too, exactly for this reason, but didn't get this information there. Maybe it would be a good idea if the encoder displayed ASM used like x264 does.
OH MY GOD! The encoding speed is just INSANELY SLOOOOOOOOOOOOOOOOOOOOOOOW. 90 minutes have passed and my 8C/16T managed to encode only 350 frames from just first pass.
How can you even test any codec if encoding takes so much time?
I don't think Rav1e even threads much. That's why I ran x264 at 1 thread too, to get even ground in the multithreading quality trade-offs (with the idea that Rav1e doesn't make them yet but will in the future when it threads).
Edit: Or not - looks like threading as well as some other options were added after my testing.
Atak_Snajpera
9th March 2019, 16:22
https://i.imgsafe.org/3d/3da34b7c3e.png
Wolfberry
9th March 2019, 16:23
Take a look at the source code: https://github.com/xiph/rav1e/blob/master/build.rs#L25
So yes, nasm is explicitly disabled on windows.
mandarinka
9th March 2019, 17:16
Seems threads and other new options got introduced in after the build I used. Though that should not matter for quality, I used the slowest possible one before the speed levels rebalancing.
Does anybody have any explanation about what the --train-rdo flag does?
And also, what does bottom-up encoding mean?
Atak_Snajpera
9th March 2019, 19:52
Why rav1e ignored my 5Mbps bitrate and gave me 10Mbps ?!?
https://i.imgsafe.org/40/40b77cd424.png
When I was watching encoded file I was quite impressed by quality but then I realized that I was comparing 10Mbps AV1 with 5Mbps x264 :(
comparision
AV1 10Mbps vs x264 10Mbps
https://www.mediafire.com/file/t3hwslldv5vzlk2/AV1-10Mbps.mkv/file
https://www.mediafire.com/file/bhyulmmn5hn122n/x264-10Mbps.mkv/file
quick screenshots
AV1 -> http://atak-snajpera.5v.pl/images/AV1-10Mbps.mkv_snapshot_00.03.000.png
x264 -> http://atak-snajpera.5v.pl/images/x264-10Mbps.mkv_snapshot_00.03.000.png
spoiler AV1 is still in (blurry) woods...
mandarinka
9th March 2019, 20:29
Rav1e rate control is rather basic I think, they only just landed a first attempt on it. You generally want to encode AV1 (even with 1pass cqp) and then run 2-pass x264 to match, like I did :)
Edit: looks like it doesn't do too well on PJ either, the grassy bank is obvious but also the people are still artifacty. x264 has sharp artifacty edges on the figures while Rav1e has them blurrer which might be more pleasant, but it still artifacts a lot there. (and I wonder what's with the superstrong artifacting in the first few frames).
Wolfberry
10th March 2019, 03:21
Information about train-rdo: https://github.com/xiph/rav1e/pull/948
It is currently disabled at all speed levels since it is not usable yet.
mandarinka
10th March 2019, 21:20
Or so it is approximate (worse) RDO that uses cost estimates based on a trained model instead of actual values.
(So a speed option, not something improving quality).
I guess I could have figured that out from the name, should have put more thinking into it.
IgorC
11th March 2019, 19:08
DAV1D 0.2.0 final is out: https://code.videolan.org/videolan/dav1d/tags
Hopefully a new vlc player will be released soon with this integrated.
MPC-BE 1.5.3 (build 4455) beta (https://forum.doom9.org/showthread.php?p=1867953#post1867953) has switched from libaom to dav1d 0.2.x
hajj_3
11th March 2019, 22:46
MPC-BE 1.5.3 (build 4455) beta (https://forum.doom9.org/showthread.php?p=1867953#post1867953) has switched from libaom to dav1d 0.2.x
great news, i just tested it, it uses far less cpu playing av1 files than before :)
Nintendo Maniac 64
12th March 2019, 05:53
great news, i just tested it, it uses far less cpu playing av1 files than before :)
I was paranoid that this would only be the case for AVX-enabled CPUs, but even on my AVX-less Haswell Pentium G3258 I was seeing ~50% faster performance than what I would get via MPC-HC v1.8.4 with its built-in LAVfilters.
sneaker_ger
12th March 2019, 10:37
Yes, the dav1d developers added a lot of SSSE3 optimizations for older CPUs. AVX is useless, AVX2 is good. See the discussions and links on the last pages.
Nintendo Maniac 64
12th March 2019, 20:57
I also just tried a very recent nightly build of LAVfilters as well (0.73.1-30 built on 2019-03-08), and unfortunately the resulting performance indicates that it's likely still using the AOM decoder.
older CPUs
Reminder that even the newest Pentiums lack AVX. Heck in terms of actual IPC and performance-per-clock, my own Pentium is only two generations old, and one of those generations were practically skipped-over for the desktop.
nevcairiel
12th March 2019, 22:47
I also just tried a very recent nightly build of LAVfilters as well (0.73.1-30 built on 2019-03-08), and unfortunately the resulting performance indicates that it's likely still using the AOM decoder.
LAV will be switching to dav1d 0.2.1 soon, which was released only minutes ago. I wanted to use 0.2.0 already but knowing that a small improvement was coming up shortly, it made sense to wait.
LigH
13th March 2019, 14:30
New uploads: (MSYS2; MinGW32: GCC 7.4.0 / MinGW64: GCC 8.3.0)
AOM v1.0.0-1457-geca009dba (https://www.mediafire.com/file/l0kwzzlkz0mk36n/aom_v1.0.0-1457-geca009dba.7z/file)
rav1e 0.1.0 (fbecf18 / 2019-03-13) (https://www.mediafire.com/file/836rpxr2cqar688/rav1e_0.1.0_2019-03-13_fbecf18.7z/file)
dav1d 0.2.1 (408d048 / 2019-03-13) (https://www.mediafire.com/file/yopfiwjye3cgye8/dav1d_0.2.1_2019-03-13_408d048.7z/file)
IgorC
14th March 2019, 00:57
Good News: AV1 Encoding Times Drop to Near-Reasonable Levels (https://www.streamingmedia.com/Articles/Editorial/Featured-Articles/Good-News-AV1-Encoding-Times-Drop-to-Near-Reasonable-Levels-130284.aspx)
https://dzceab466r34n.cloudfront.net/Images/ArticleImages/InlineImages/122183-AV1-Figure1-ORG.jpg
Tommy Carrot
14th March 2019, 03:32
Good News: AV1 Encoding Times Drop to Near-Reasonable Levels (https://www.streamingmedia.com/Articles/Editorial/Featured-Articles/Good-News-AV1-Encoding-Times-Drop-to-Near-Reasonable-Levels-130284.aspx)
Indeed, it's getting faster. Unfortunately at the cost of compression efficiency, according to my tests it has gotten worse by about 3-5% since december, and --cpu-used=0 currently barely has better quality than --cpu-used=1 had 3 months old ago, while the encoding speed is still being slower. In other word, for the same quality/efficiency, the encoding speed hasn't really improved much, but the maximum efficiency has gotten worse.
EwoutH
14th March 2019, 13:01
The VLC Dev and 3.0 branches now use dav1d 0.2.1.
Here are new Nightly builds for:
VLC Dev Windows (64-bit) (https://nightlies.videolan.org/build/win64/last/)
VLC 3.0 Windows (64-bit) (https://nightlies.videolan.org/build/win64/last-3/)
VLC Android 3.1 (64-bit) (https://nightlies.videolan.org/build/android-armv8a/)
Chimera (http://download.opencontent.netflix.com/?prefix=AV1/Chimera/) (8-bit 1080p) now plays most scenes fluently on my Galaxy S7 (Exynos 8890 with 4x Mongoose 1, 4x Cortex-A53 and 4GB LPDDR4). Summer in Tomsk 1080p (from Elecard (https://www.elecard.com/videos)) plays fluently, while Summer Nature 1080p still has a stutter here and there (bitrate is a lot higher). 720p videos are now decoded without a hitch.
foxyshadis
14th March 2019, 23:25
The rest of the clip has Rav1e clearly deficient despite the large bitrate though. It constantly smoothes the solid/flat areas and generally can't keep the milder grain and texture (which is a flaw). So its psychovisual decisions probably still aren't ready for high quality transparent encoding. This is a general problem with any non x264/x265 encoder probably, even x265 would drop texture detail everywhere before it got aq and psyrdo.
It never ceases to astonish me how every single codec seems to prize PSNR over visual quality at the start, after all these years, and bolt on psychovisual as an afterthought. Also, how many codecs are designed with only low-bitrate in mind, assuming that high bitrate will automatically be transparent, so there's no reason to waste time on that. You'd think that Google would have, well, Googled for the successes and failures in this field before surging ahead. Or at least listened to Monty, since he's The Woz of codec design.
marcomsousa
15th March 2019, 10:11
Android Q introduces support for the open source video codec AV1. This allows media providers to stream high quality video content to Android devices using less bandwidth. In addition, Android Q supports audio encoding using Opus - a codec optimized for speech and music streaming, and HDR10+ for high dynamic range video on devices that support it.
The MediaCodecInfo API introduces an easier way to determine the video rendering capabilities of an Android device. For any given codec, you can obtain a list of supported sizes and frame rates using VideoCodecCapabilities.getSupportedPerformancePoints(). This allows you to pick the best quality video content to render on any given device.
benwaggoner
16th March 2019, 00:53
It never ceases to astonish me how every single codec seems to prize PSNR over visual quality at the start, after all these years, and bolt on psychovisual as an afterthought. Also, how many codecs are designed with only low-bitrate in mind, assuming that high bitrate will automatically be transparent, so there's no reason to waste time on that. You'd think that Google would have, well, Googled for the successes and failures in this field before surging ahead. Or at least listened to Monty, since he's The Woz of codec design.
PSNR may not be good, but it is easy to calculate.
And Google REALLY likes objective metrics. They trust a number from a computer more than their own eyes.
It's one of those "looking where the light is" kinds of problems.
Nintendo Maniac 64
16th March 2019, 23:17
They trust a number from a computer more than their own eyes.
Geez, no wonder people like Elon Musk are so concerned about certain tech companies going all-in on deep A.I.
benwaggoner
19th March 2019, 01:00
Say, does anyone know if the Android Q AV1 decoder has/will have PlayReady integration? If so, what level?
Blue_MiSfit
19th March 2019, 04:41
I strongly doubt it. If there is an integration I'd be shocked if it's not software based, which would max out at SL2000.
marcomsousa
19th March 2019, 08:53
Firefox 66 - Support for AV1 codec is activated on Windows by default.
birdie
19th March 2019, 14:13
Firefox 66 - Support for AV1 codec is activated on Windows by default.
It was already enabled (https://www.mozilla.org/en-US/firefox/65.0/releasenotes/) in Firefox 65.
Nintendo Maniac 64
23rd March 2019, 02:04
It would seems that LAVFilters v0.74.1 and therefore also MPC-HC v1.8.6 are now using dav1d, resulting in similar performance gains that were seen in MPC-BE beta builds (which I previously measured as being ~50% faster than MPC-HC v1.8.4 and its included LAVFilters v0.73).
Pushman
23rd March 2019, 08:42
https://github.com/Nevcairiel/LAVFilters/releases/tag/0.74
NEW: Using the dav1d AV1 decoder for significantly improved AV1 decoding performance
Nintendo Maniac 64
23rd March 2019, 20:40
https://github.com/Nevcairiel/LAVFilters/releases/tag/0.74
Note, there's a newer 0.74.1 version of LAVFilters in case anybody is thinking that above link is the newest version (which it's not):
https://github.com/Nevcairiel/LAVFilters/releases
mzso
23rd March 2019, 22:58
Hi!
Is there anything out there that can encode valid AVIF images?
Does (dev versions of) Chrome/Firefox support viewing AVIF images? (Anything else?)
nevcairiel
24th March 2019, 00:22
Hi!
Is there anything out there that can encode valid AVIF images?
Does (dev versions of) Chrome/Firefox support viewing AVIF images? (Anything else?)
You can encode with this tool, for example:
https://github.com/Kagami/go-avif (binary: https://ci.appveyor.com/project/Kagami/go-avif/build/artifacts)
They also have a avif.js demo that can decode it using JavaScript (https://kagami.github.io/avif.js/). I don't think any browsers have native support yet, since the spec for AVIF was only officially signe-off a week ago or so.
Interestingly, the upcoming Windows 10 "19H1" seems to have native support, both in Explorer for thumbnails, as well as in Paint for importing - assuming the AV1 Video Extension is installed.
Wolfberry
24th March 2019, 00:44
AVIF is basically AV1 codec in HEIF container, so you can also use ffmpeg with libaom enabled (or ffmpeg + aomenc) to encode the pictures into ivf and then use MP4Box -add-image to convert the ivf files into avif.
nevcairiel
24th March 2019, 01:36
AVIF is basically AV1 codec in HEIF container, so you can also use ffmpeg with libaom enabled (or ffmpeg + aomenc) to encode the pictures into ivf and then use MP4Box -add-image to convert the ivf files into avif.
To be fair, any AVIF file is a HEIF file, but if your tool doesn't know what its doing, then a HEIF file with AV1 in it may not be valid AVIF (since it defines some constraints to make it easier on readers that only want to support AVIF and not the full HEIF spec).
That said, MP4Box actually supports AVIF however, but you should be passing "-brand avif" to make sure it makes a proper one.
Compatible software and some invocation commands are listed here:
https://github.com/AOMediaCodec/av1-avif/wiki
Wolfberry
24th March 2019, 01:44
Thanks for the link, SmilingWolf's script also adds -ab miaf, I guess that is optional?
nevcairiel
24th March 2019, 02:30
Thanks for the link, SmilingWolf's script also adds -ab miaf, I guess that is optional?
Yeah, sort of optional. The spec says it should be there, but its not clear to me if you need to use one of its profiles. You can have a whole list of alternate/compatible brands in such a file. The AVIF spec lists them:
https://aomediacodec.github.io/av1-avif/#profiles-constraints
"avif" from the AVIF spec itself.
"mif1" from HEIF
"miaf" from the MIAF spec
and finally "MA1B" or "MA1A" to identify the AV1 profile in use.
Or various alterations of the above for special types of images.
... this ISOBMFF image format stuff is a real mess. :)
NikosD
6th April 2019, 07:59
I did a small comparison between lib-aom and lib-dav1d earlier this week.
I used LAV Video 0.73.1-31 x64 for lib-aom and LAV Video 0.74.1-1 for lib-dav1d (no idea if LAV uses the latest versions of both libraries, but should be close to latest)
The two systems mentioned below have Win 10 October Update and I used DXVA Checker v4.20 for both of them.
The AV1 sample is Chimera - AV1 - 1080p - 8bit - 6.7Mbps file.
Skylake system:
Core i5 6500 (4C/4T) at 3.2GHz (All core turbo 3.3GHz) using 1x8GB DIMM of DDR4-2133 MHz RAM (Single channel)
Haswell system:
Core i3 4170 (2C/4T) at 3.7GHz (No Turbo) using 2x8GB DIMMs of DDR3-1600 MHz RAM (Dual channel)
Benchmark results:
(min/avg/max fps)
Skylake Core i5 6500:
libaom 30/48/149 CPU usage 65%
dav1d 77/128/282 CPU usage 91%
Haswell Core i3 4170:
libaom 21/37/153 CPU usage 71%
dav1d 53/91/240 CPU usage 90%
Comments:
For both AVX2 capable CPUs, the speedup is similar ~2.5 times faster for dAV1d than libaom (2.46 for Haswell and 2.67 for Skylake) and the CPU usage also goes to ~90% for both.
The difference is huge for multi-threaded systems.
Core i5 6500 manages to be faster from 30% (using libaom) to 41% (using dav1d) than Core i3 4170.
dapperdan
6th April 2019, 09:23
MSU HEVC/AV1/VP9 encoding test results:
http://compression.ru/video/codec_comparison/hevc_2018/
soresu
9th April 2019, 01:11
Did anyone see the SVT-AV1 announcement from Intel/Netflix.
Link here (https://www.phoronix.com/scan.php?page=news_item&px=Intel-SVT-AV1-Announced).
Still havent seen a comprehensive comparison of it compared to libaom, x265 and vp9 quality wise.
tnti
9th April 2019, 03:38
https://mp.weixin.qq.com/s?__biz=MzU1NTEzOTM5Mw==&mid=2247489793&idx=1&sn=6a148c57f2aa9885533f6b278ca3b94e&chksm=fbd9b12fccae3839d044c60fca9fa3a72dbf9b684562e1af7c3ef69ad78f1e3ad620416be03e&xtrack=1&scene=0&subscene=131&clicktime=1554775340&ascene=7&devicetype=android-27&version=2700033c&nettype=WIFI&abtest_cookie=BAABAAoACwASABMABQAjlx4AVpkeAMiZHgDZmR4A3JkeAAAA&lang=zh_CN&pass_ticket=fjjXRX43%2F5cmbhcL8uajT4nAMGLFrTTwZb7hV1Ggt4VVt5J1zgccAZAL4jdEkgDp&wx_header=1
关于对比rav1e和SVT-AV1,以下为来自Zoe Liu的回复:
“rav1e和SVT-AV1在github上的开源编码器,迄今为止还没有完成AV1标准中的最主要的编码工具集。对于任意一款AV1编码器,如果没有实现AV1标准的主要工具,只是在码流格式上符合AV1标准规范,而无法体现AV1的标准优势,其实还没有达到作为AV1编码器的基准要求。
我们对最新的SVT-AV1 github版本做了相对粗略的RD性能测试,其目前的编码效率还停留在AV1的前身VP9的大致水平。对这样开发阶段的编码器,放到AV1类别里面做编码性能与速度评估,不仅没有意义,对这样的开源项目也是不公平的。”
The following is the content of machine translation:
Regarding the comparison between rav1e and SVT-AV1, the following is a response from Zoe Liu:
"Rav1e and SVT-AV1 open source coders on GitHub have not yet completed the most important coding tool set in AV1 standard. For any AV1 encoder, if there is no main tool to implement the AV1 standard, it only conforms to the AV1 standard in bit stream format, but can not reflect the standard advantages of AV1. In fact, it has not reached the standard requirements as AV1 encoder.
We have done a relatively rough RD performance test for the latest version of SVT-AV1 github, and its coding efficiency is still at the approximate level of VP9, the predecessor of AV1. It is not only meaningless to evaluate the coding performance and speed of the coder in the AV1 category, but also unfair to such open source projects.
hajj_3
9th April 2019, 12:21
Intel to make their own AV1 software decoder:
https://www.phoronix.com/scan.php?page=news_item&px=Intel-TODO-AV1-Decode-FFmpeg
https://trello.com/c/oEZQai7L/26-av1-decoder
utack
9th April 2019, 13:19
Intel to make their own AV1 software decoder:
https://www.phoronix.com/scan.php?page=news_item&px=Intel-TODO-AV1-Decode-FFmpeg
https://trello.com/c/oEZQai7L/26-av1-decoder
Seems like they were not happy dav1d runs well on Ryzen :p
hajj_3
9th April 2019, 14:00
Seems like they were not happy dav1d runs well on Ryzen :p
Hopefully the dav1d decoder will copy any improvements that intel makes in their decoder. I have an intel kaby lake chip so i'm not complaining about intel making a new software decoder :P
nevcairiel
9th April 2019, 15:45
Seems like they were not happy dav1d runs well on Ryzen :p
Well dav1d certainly won't get slower through any of their efforts. :)
That said, dav1d puts a lot of effort into AVX2, which current-gen Ryzen doesn't even do particularly fast, so..
Additionally, that Trello "av1 decoder" task existed before dav1d was even a thing, and contrary to the Phoronix article, dav1d is actually quite decently fast on 8-bit content (work on speeding up 10-bit has started).
BUT, one should know that an encoder basically contains a decoder already, because it has to practically decode encoded frames to get accurate reference frame information, so splitting that out may always be an option.
mandarinka
9th April 2019, 16:13
Dav1d will be always better for us end users, I think that's almost certain. SVT encoders are kind enterprise code which likely won't be too receptive about various issues of users in the wild, who have weird needs like non-AVX2 CPUs, non-mod16/8/4 video (for example) and integration into various players which the coders of these server-side encoders for commercial use are not going to see as important or even be aware of.
Dav1d devs are going to be more in touch with this and generally, the ffmpeg circle sort of development has proved itself with decoders.
foxyshadis
10th April 2019, 00:43
Yeah, sort of optional. The spec says it should be there, but its not clear to me if you need to use one of its profiles. You can have a whole list of alternate/compatible brands in such a file. The AVIF spec lists them:
https://aomediacodec.github.io/av1-avif/#profiles-constraints
"avif" from the AVIF spec itself.
"mif1" from HEIF
"miaf" from the MIAF spec
and finally "MA1B" or "MA1A" to identify the AV1 profile in use.
Or various alterations of the above for special types of images.
... this ISOBMFF image format stuff is a real mess. :)
As I read it, "miaf" is required for AVIF, and defines a fair number of important constraints. I wish it was possible to view MIAF's draft standard, but if you aren't in MPEG, ISO, or IEC it looks like that's not possible at all for at least another month.
Nintendo Maniac 64
10th April 2019, 17:56
That said, dav1d puts a lot of effort into AVX2, which current-gen Ryzen doesn't even do particularly fast, so..
Though Intel does seem to like segmenting their CPU lineup with regards to AVX feature-set.
NikosD
11th April 2019, 07:47
Even the current Ryzen architecture is very fast using dAV1d probably due to the highly multi-threaded nature of both (dAV1d and Ryzen)
The multi-threaded gain seems to be more significant than the AVX2 loss for current-gen Ryzen.
soresu
14th April 2019, 17:02
Seems that rav1e is prioritising chunked/segment encoding early on to maximise parallelism, they just opened a new project list to target specific tasks. Link here (https://github.com/xiph/rav1e/projects/9).
IgorC
14th April 2019, 18:11
Xiph update on one year of AV1 (slides from NAB 2019) [PDF]
https://people.xiph.org/~negge/NAB2019.pdf
user1085
14th April 2019, 18:43
Seems that rav1e is prioritising chunked/segment encoding early on to maximise parallelism, they just opened a new project list to target specific tasks. Link here (https://github.com/xiph/rav1e/projects/9).How does chunked/segment encoding work?
soresu
14th April 2019, 22:11
Much like the name suggests, by splitting the file to be encoded by chunks/segments, each being encoded in a separate instance of the encoder, but each also having the lower parallelism of tile or frame based encoding on top of that.
Its more beneficial for systems like servers, or dual socket workstations with large core counts, where the WPP, tile and frame parallel techniques are not enough on their own to saturate all cores with work.
Apparently it has its own problems such as memory consumption and chiefly the time taken to initialise each instance.
Also the segments would have to be chosen based on detected scene change otherwise you could have some very obvious bitrate changes visible mid shot.
I dont think it is suitable for real time encoding, unless you have a significant time delay on transmission to detect scene changes before carving up a new segment to be encoded possibly?
Dyomich
15th April 2019, 16:58
MSU HEVC/AV1/VP9 encoding test results:
http://compression.ru/video/codec_comparison/hevc_2018/
dapperdan, thank you for the link. Here is more information about the results:
MSU received many requests on the results of announced AV1 participation in comparison.
AV1 is quickly being developed, but still too slow to participate in our usual fast-universal-ripping use cases.
This is why AV1 was included only in this special report (so as in 2017).
Download free PDF report: http://compression.ru/video/codec_comparison/hevc_2018/pdf/MSU_HEVC_AV1_comparison_2018_P4_HQ_encoders.pdf
This part of the comparison usually releases later because of much lower encoding speed than in other use cases (formal limit was 0.005 fps but actually unlimited).
According to only quality scores, the places of the competitors are the following:
AV1
VP9, x265 and sz265
sz264 and x264
On the speed-quality chart there are four Pareto-optimal participants: AV1, VP9, x264, and SIF Encoder:
http://compression.ru/video/codec_comparison/hevc_2018/figures/speed_quality_av1_2018.png
You can compare the latest results to the results from out similar comparison in 2017 (http://compression.ru/video/codec_comparison/hevc_2017/),
(Important note: in 2017 we used VBR mode for AV1 encoder, and in 2018 constant QP was used, as VBR was several times slower).
SmilingWolf
16th April 2019, 18:57
Status report!
"MSU preceded me" edition
1st edition: https://forum.doom9.org/showthread.php?p=1852449#post1852449
2nd edition: https://forum.doom9.org/showthread.php?p=1857587#post1857587
3rd edition: https://forum.doom9.org/showthread.php?p=1860475#post1860475
Whatever paragraph I don't repeat here can be assumed to be the same as in the aforementioned post
First of all: graphs!
Click to enlarge
Y axis: chosen metric
X axis: bits per pixel
720p:
https://i.ibb.co/Sxr5vnm/hvmaf-720.png (https://ibb.co/Sxr5vnm) https://i.ibb.co/vXJCNvM/msssim-720.png (https://ibb.co/vXJCNvM) https://i.ibb.co/XDQ82Xs/psnrhvsm-720.png (https://ibb.co/XDQ82Xs)
1080p:
https://i.ibb.co/sKXkrZV/hvmaf-1080.png (https://ibb.co/sKXkrZV) https://i.ibb.co/PTvVsvN/msssim-1080.png (https://ibb.co/PTvVsvN) https://i.ibb.co/F7mM4YX/psnrhvsm-1080.png (https://ibb.co/F7mM4YX)
BD rates for 720p:
Codecs ladder: | x264 relative:
x264 -> svtav1 | x264 -> svtav1
RATE (%) DSNR (dB) | RATE (%) DSNR (dB)
MSSSIM -3.28672 0.142426 | MSSSIM -3.28672 0.142426
PSNRHVS -4.36439 0.235114 | PSNRHVS -4.36439 0.235114
HVMAF -7.28468 0.186648 | HVMAF -7.28468 0.186648
----------------------------|-----------------------------
svtav1 -> vp9 | x264 -> vp9
RATE (%) DSNR (dB) | RATE (%) DSNR (dB)
MSSSIM -14.9778 0.624669 | MSSSIM -21.5937 1.19737
PSNRHVS -17.0162 0.881864 | PSNRHVS -23.6381 1.70941
HVMAF -17.6481 0.768638 | HVMAF -21.7739 2.36707
----------------------------|-----------------------------
vp9 -> x265 | x264 -> x265
RATE (%) DSNR (dB) | RATE (%) DSNR (dB)
MSSSIM -5.14321 0.226527 | MSSSIM -25.8745 1.38639
PSNRHVS -8.41874 0.455384 | PSNRHVS -30.4345 2.09471
HVMAF -13.8727 0.673276 | HVMAF -31.395 3.32434
----------------------------|-----------------------------
x265 -> av1 | x264 -> av1
RATE (%) DSNR (dB) | RATE (%) DSNR (dB)
MSSSIM -17.7008 0.800381 | MSSSIM -36.613 2.16722
PSNRHVS -14.1648 0.748597 | PSNRHVS -37.5985 2.80939
HVMAF -12.6967 0.474379 | HVMAF -39.7078 2.87016
BD rates for 1080p:
Codecs ladder: | x264 relative:
x264 -> svtav1 | x264 -> svtav1
RATE (%) DSNR (dB) | RATE (%) DSNR (dB)
MSSSIM -2.03138 0.058404 | MSSSIM -2.03138 0.058404
PSNRHVS 3.95915 -0.147954 | PSNRHVS 3.95915 -0.147954
HVMAF -4.05114 -0.0123967 | HVMAF -4.05114 -0.0123967
-----------------------------|------------------------------
svtav1 -> vp9 | x264 -> vp9
RATE (%) DSNR (dB) | RATE (%) DSNR (dB)
MSSSIM -29.5206 1.00448 | MSSSIM -33.7275 1.65389
PSNRHVS -32.951 1.40212 | PSNRHVS -33.0678 2.11895
HVMAF -34.7504 1.33459 | HVMAF -33.0663 3.73028
-----------------------------|------------------------------
vp9 -> x265 | x264 -> x265
RATE (%) DSNR (dB) | RATE (%) DSNR (dB)
MSSSIM 6.91124 -0.232017 | MSSSIM -30.515 1.24701
PSNRHVS 2.12517 -0.103003 | PSNRHVS -32.954 1.71647
HVMAF -5.87407 0.106315 | HVMAF -35.0935 3.13962
-----------------------------|------------------------------
x265 -> av1 | x264 -> av1
RATE (%) DSNR (dB) | RATE (%) DSNR (dB)
MSSSIM -28.0678 0.994389 | MSSSIM -47.6133 2.2674
PSNRHVS -23.2568 0.966762 | PSNRHVS -45.5839 2.73622
HVMAF -24.6825 0.769329 | HVMAF -50.6976 3.39944
Encoders:
x264 157-2935-545de2f
x265 3.0-1-ed72af837053
libvpx-vp9 1.8.0-374-gb758ba795
SVT-AV1 0.0.1-525-5a8b857e
libaom 1.0.0-1605-g64a2ffb72
Cmdlines:
x264 --preset veryslow --tune ssim --crf 16 -o test.x264.crf16.264 orig.i420.y4m
x265 --preset veryslow --tune ssim --crf 16 -o test.x265.crf16.hevc orig.i420.y4m
vpxenc --codec=vp9 --frame-parallel=0 --tile-columns=1 --auto-alt-ref=6 --good --cpu-used=0 --tune=psnr --passes=2 --threads=2 --end-usage=q --cq-level=20 --test-decode=fatal --ivf -o test.vp9.cq20.ivf orig.i420.y4m
SvtAv1EncApp.exe -i orig.i420.yuv -b test.svtav1.cq20.ivf -w 1280 -h 720 -q 20 -enc-mode 3 -fps-num 24000 -fps-denom 1001 -intra-period 23
aomenc --frame-parallel=0 --tile-columns=1 --auto-alt-ref=1 --cpu-used=4 --tune=psnr --passes=2 --threads=2 --row-mt=1 --end-usage=q --cq-level=20 --test-decode=fatal -o test.av1.cq20.webm orig.i420.y4m
VMAF: model used: vmaf_v0.6.1, pooling: harmonic_mean
Notes:
TearsOfSteel720 and TheFifthElement, two clips in the 720p category, have a vertical resolution incompatible with SvtAv1EncApp (not divisible by 8), so they have been excluded from all measurements this time.
Next time they will be padded, all the encodes re-made and quality measurements re-taken.
Meanwhile, rav1e has got a nasty bug that makes it bloat encodes, which brings up to 25% BD rate regression, so it has been excluded from this edition.
I'll try to post an amendment when the above problems have been fixed.
This concludes this report.
As always, I'm open to any kind of feedback to improve my comparisons and my encodes.
Nintendo Maniac 64
16th April 2019, 20:52
Status report!
Good to see you're not dead - you hadn't logged in for nearly 4 months. :p
Anyway quick question - since dav1d seems to be pretty much the standard go-to AV1 decoder nowadays, I don't suppose you'd be interested in the possibility of more decoder performance tests?
Basically I'm thinking how that Presage Flower video clip you had me test is going to have pretty outdated results regarding AV1's relative CPU software decoding requirements (https://forum.doom9.org/showthread.php?p=1859536#post1859536) seeing as dav1d is like 50% faster even on CPUs that don't support AVX (I've not yet tested it on CPUs without SSE4.1 and/or SSSE3 however).
singhkays
16th April 2019, 23:02
Status report!
Glad to see this! I'm working on a sub-500Kbps comparison myself for AV1/x264/VP9 (this is to replace GIF with the smallest video that'll play natively in a browser)
Do you know how these new commandline options in ffmpeg affect quality/speed for AV1?
enable_cdef, enable_global_motion, and intrabc
https://git.ffmpeg.org/gitweb/ffmpeg.git/commit/995889abbf31389aa0f002ea6599c7c89526d1d6
marcomsousa
17th April 2019, 18:00
1) Comparision of libaom, SVT-AV1 and rav1e for still-image coding on subset1 from derf's collection. JPEG files are also encoded with libjpeg for the reference.
https://github.com/Kagami/av1-bench/raw/master/graphs/still1.png
https://github.com/Kagami/av1-bench
2) go-avif comes with handy CLI utility avif. It supports encoding of JPEG and PNG files to AVIF
Static 64-bit builds for Windows, macOS and Linux are available at releases page.
https://github.com/Kagami/go-avif
3) AVIF (AV1 Still Image File Format) polyfill for the browser
https://github.com/Kagami/avif.js
demo https://kagami.github.io/avif.js/
Nintendo Maniac 64
17th April 2019, 18:15
Any comparisons for lossless AVIF?
marcomsousa
17th April 2019, 18:17
Chrome will use dav1d by default for decoding.
Chromium 74 [1] (week of 2019-Apr-23)
https://chromium.googlesource.com/chromium/src.git/+/ede4345734371ac35b4e50cea7ef3332dabd54d5%5E%21/
SmilingWolf
17th April 2019, 20:54
Any comparisons for lossless AVIF?
I've got you covered... somewhat. My comparison was made with a mix of CG wallpapers and Pixiv anime fanart. Here: https://docs.google.com/spreadsheets/d/1bG-CSTn_L79MtS7A-wR572H5_W_3kb2xX5AWagKn1ak/
The numbers are 3 months old but I don't think they have moved too much.
Yup, not ded, but busy :)
Unless dav1d has started adding 10bit asm or has made big improvements in the multithreading department I don't think there'll be much to see w.r.t. improvements against my Presage Flower encode (10bit YUV420P), but you're very welcome to try and prove me wrong of course, benchmarks are always good!
Do you know how these new commandline options in ffmpeg affect quality/speed for AV1?
https://git.ffmpeg.org/gitweb/ffmpeg.git/commit/995889abbf31389aa0f002ea6599c7c89526d1d6
I never tested any of that flags with the exception of frame_parallel, which I always set to 0 (disabled).
I'd imagine one would want to leave the CDEF filter on, but I have hardly any idea what most of the others do.
Nintendo Maniac 64
18th April 2019, 00:06
I've got you covered... somewhat. My comparison was made with a mix of CG wallpapers and Pixiv anime fanart. Here: https://docs.google.com/spreadsheets/d/1bG-CSTn_L79MtS7A-wR572H5_W_3kb2xX5AWagKn1ak/
a30.png
a_cs03.png
Gee that file name formatting sure looks familiar... *double checks the contents of "bgimage.xp3"* yup, knew it.
Also it's kind of neat that one can access the source pixiv artwork's page with nothing more than just the file name since it's the same set of numbers used in both the file name and in a given artwork's web page URL.
Unless dav1d has started adding 10bit asm or has made big improvements in the multithreading department I don't think there'll be much to see w.r.t. improvements against my Presage Flower encode (10bit YUV420P)
Oh derp, I completely forgot about it being 10bit. However, dav1d apparently does scale across multiple CPU threads extremely well.
singhkays
18th April 2019, 00:09
I never tested any of that flags with the exception of frame_parallel, which I always set to 0 (disabled).
I'd imagine one would want to leave the CDEF filter on, but I have hardly any idea what most of the others do.
Actually, after I posted I found more info in the AV1 paper here https://jmvalin.ca/papers/AV1_tools.pdf
Intra Block Copy: AV1 allows its intra coder to refer backto previously reconstructed blocks in the same frame, in a mannersimilar to how inter coder refers to blocks from previous frames.It can be very beneficial for screen content videos which typicallycontain repeated textures, patterns and characters in the same frame.Specifically, a new prediction mode named IntraBC is introduced, andwill copy a reconstructed block in the current frame as prediction. Thelocation of the reference block is specified by a displacement vector ina way similar to motion vector compression in motion compensation.Displacement vectors are in whole pixels for the luma plane, andmay refer to half-pel positions on corresponding chrominance planes,where bilinear filtering is applied for sub-pel interpolation.
Warped Motion Compensation: Warped motion models areexplored in AV1 by enabling two affine prediction modes, globaland local warped motion compensation [10]. The global motiontool is meant for handling camera motions, and allows conveyingmotion models explicitly at the frame level for the motion betweena current frame and any of its reference
Constrained Directional Enhancement Filter (CDEF): CDEFis a detail-preserving deringing filter designed to be applied afterdeblocking that works by estimating edge directions followed byapplying a non-separable non-linear low-pass directional filter of size5×5 with 12 non-zero weights [14]. To avoid extra signaling, thedecoder computes the direction per 8×8 block using a normative fastsearch algorithm that minimizes the quadratic error from a perfectdirectional pattern
I tested out these options with the following CLI however, I saw no change in the average bitrate of the output file, speed of encoding or any difference in VMAF scores of different files.
./ffmpeg -i scene3.mp4 -c:v libaom-av1 -b:v 200k -filter:v scale=720:-1 -strict experimental -cpu-used 1 -tile-rows 0 -tile-columns 2 -row-mt 1 -threads 8 -pass 1 -f mp4 -y /dev/null &&
./ffmpeg -i scene3.mp4 -pix_fmt yuv420p -movflags faststart -c:v libaom-av1 -b:v 200k -filter:v scale=720:-1 -strict experimental -cpu-used 1 -tile-rows 0 -tile-columns 2 -row-mt 1 -threads 8 -pass 2 -hide_banner ./av1/av1_200k-none.mp4
./ffmpeg -i scene3.mp4 -c:v libaom-av1 -b:v 200k -filter:v scale=720:-1 -strict experimental -cpu-used 1 -tile-rows 0 -tile-columns 2 -enable-cdef 1 -enable-global-motion 1 -enable-intrabc 1 -row-mt 1 -threads 8 -pass 1 -f mp4 -y /dev/null &&
./ffmpeg -i scene3.mp4 -pix_fmt yuv420p -movflags faststart -c:v libaom-av1 -b:v 200k -filter:v scale=720:-1 -strict experimental -cpu-used 1 -tile-rows 0 -tile-columns 2 -enable-cdef 1 -enable-global-motion 1 -enable-intrabc 1 -row-mt 1 -threads 8 -pass 2 -hide_banner ./av1/av1_200k-all.mp4
./ffmpeg -i scene3.mp4 -c:v libaom-av1 -b:v 200k -filter:v scale=720:-1 -strict experimental -cpu-used 1 -tile-rows 0 -tile-columns 2 -enable-cdef 1 -row-mt 1 -threads 8 -pass 1 -f mp4 -y /dev/null &&
./ffmpeg -i scene3.mp4 -pix_fmt yuv420p -movflags faststart -c:v libaom-av1 -b:v 200k -filter:v scale=720:-1 -strict experimental -cpu-used 1 -tile-rows 0 -tile-columns 2 -enable-cdef 1 -row-mt 1 -threads 8 -pass 2 -hide_banner ./av1/av1_200k-cdef-only.mp4
./ffmpeg -i scene3.mp4 -c:v libaom-av1 -b:v 200k -filter:v scale=720:-1 -strict experimental -cpu-used 1 -tile-rows 0 -tile-columns 2 -enable-global-motion 1 -row-mt 1 -threads 8 -pass 1 -f mp4 -y /dev/null &&
./ffmpeg -i scene3.mp4 -pix_fmt yuv420p -movflags faststart -c:v libaom-av1 -b:v 200k -filter:v scale=720:-1 -strict experimental -cpu-used 1 -tile-rows 0 -tile-columns 2 -enable-global-motion 1 -row-mt 1 -threads 8 -pass 2 -hide_banner ./av1/av1_200k-globmotion-only.mp4
./ffmpeg -i scene3.mp4 -c:v libaom-av1 -b:v 200k -filter:v scale=720:-1 -strict experimental -cpu-used 1 -tile-rows 0 -tile-columns 2 -enable-intrabc 1 -row-mt 1 -threads 8 -pass 1 -f mp4 -y /dev/null &&
./ffmpeg -i scene3.mp4 -pix_fmt yuv420p -movflags faststart -c:v libaom-av1 -b:v 200k -filter:v scale=720:-1 -strict experimental -cpu-used 1 -tile-rows 0 -tile-columns 2 -enable-intrabc 1 -row-mt 1 -threads 8 -pass 2 -hide_banner ./av1/av1_200k-intrabc-only.mp4
soresu
18th April 2019, 02:33
Wow, I suspected SVT-AV1 to have lower quality, but not so bad as this - it barely seems to pass x264 and is worse at some resolutions and bitrates.
Wasnt rav1e supposed to have passed x264 quality some time ago?
If these graphs match to VMAF too, I cant see why Netflix would announce support for SVT-AV1 when they already have a performant VP9 solution which presumably at least matches libvpx quality, if not better with EVE-VP9.
marcomsousa
18th April 2019, 15:24
I cant see why Netflix would announce support for SVT-AV1 when they already have a performant VP9 solution which presumably at least matches libvpx quality, if not better with EVE-VP9.
Netflix choose SVT-AV1 because of architecture, language and how scalable it is.
"SVT-AV1 uses parallelization at several stages of the encoding process."
"However, rav1e is written in Rust programming language, whereas an encoder written in C has a much broader base of potential developers."
The quality is a matter of time.
Introducing SVT-AV1: a scalable open-source AV1 framework by Netflix
https://medium.com/netflix-techblog/introducing-svt-av1-a-scalable-open-source-av1-framework-c726cce3103a
SmilingWolf
18th April 2019, 17:03
Wasnt rav1e supposed to have passed x264 quality some time ago?
Oh it absolutely did, as far back as 4 months ago (https://forum.doom9.org/showthread.php?p=1860475#post1860475) already.
Unfortunately there's a bug (#1081 (https://github.com/xiph/rav1e/issues/1081)) that is currently badly damaging the BD rates, so no comparison from me until that's fixed. One of the PRs (#1215 (https://github.com/xiph/rav1e/pull/1215)) seems promising in that regards, after some ironing out it should bring the encoding efficiency back on track.
mandarinka
18th April 2019, 18:17
Nice metrics in those curves are nice, but in reality, is the ability to keep the visual quality good there?
That said, I don't think it's even fair to expect miracles until there's adaptive quantization. I don't think there has been a good encoder without adaptive quantization yet (and psyrdo too possibly?) so if we're just assuming Rav1e won't need it, that might be kind of unwarranted.
benwaggoner
18th April 2019, 18:30
Much like the name suggests, by splitting the file to be encoded by chunks/segments, each being encoded in a separate instance of the encoder, but each also having the lower parallelism of tile or frame based encoding on top of that.
Its more beneficial for systems like servers, or dual socket workstations with large core counts, where the WPP, tile and frame parallel techniques are not enough on their own to saturate all cores with work.
Apparently it has its own problems such as memory consumption and chiefly the time taken to initialise each instance.
This isn't too bad if you use >>10 second chunks. Not how YouTube does it, but more common common for premium content
Also the segments would have to be chosen based on detected scene change otherwise you could have some very obvious bitrate changes visible mid shot.
This is the bigger problem. Beyond scene changes, ensuring VBV compliance at splice points requires either very conservative rate control near splice points, or risk out of comformance spikes. This isn't a huge deal with SW decoders (although a bitrate spike will likely result in dropped frames in a complex high motion scene where dropped frames are most visible). But it can be a serious issue with HW decoders; either they can't handle existing content, or the die size and cost has to increase to handle out-of-spec streams. This was a big deal ~10 years ago when x264 didn't enforce levels requirement, and a 1080p flagged as Level 4.0 could still have 9 reference frames and break decoders.
I dont think it is suitable for real time encoding, unless you have a significant time delay on transmission to detect scene changes before carving up a new segment to be encoded possibly?
Best-case, latency increases by max chunk duration. So you get into quality/latency tradeoffs. Definitely not feasible for videoconferencing!
Note that "real" encoders get around these issues by using frame-level parallelism. x265 can saturate a good number of cores at 1080p.
benwaggoner
18th April 2019, 18:33
1) Comparision of libaom, SVT-AV1 and rav1e for still-image coding on subset1 from derf's collection. JPEG files are also encoded with libjpeg for the reference.
https://github.com/Kagami/av1-bench/raw/master/graphs/still1.png
https://github.com/Kagami/av1-bench
VMAF isn't a still image metric. Has anyone run a correlation for VMAF against subjective testing for still images? If not, I wouldn't trust it as a metric for this use case. Nor PSNR or SSIM.
2) go-avif comes with handy CLI utility avif. It supports encoding of JPEG and PNG files to AVIF
Static 64-bit builds for Windows, macOS and Linux are available at releases page.
https://github.com/Kagami/go-avif
With PNG in there, wouldn't we be comparing YUV<>RGB?
benwaggoner
18th April 2019, 18:45
Nice metrics in those curves are nice, but in reality, is the ability to keep the visual quality good there?
Since libaom's encoding was tuned with VMAF, we'd expect it to have a higher VMAF at any given subjective quality than other encoders that weren't tuned against it. There have been several papers that have shown HEVC or VVC's reference encoders delivering subjectively superior quality than AV1 even when VMAF predicts AV1 will look better.
VMAF is our least-bad objective metric ever, and keeps on getting better. But like any ML system, it is fundamentally a statistics-driven system that tries to give the same answer given similar data as the data-answer pairs it was trained on. So VMAF isn't going to be great at discriminating between AV1 and VVC artifacts without a lot of double-blind subjective ratings of AV1 and VVC artifacts.
This isn't a bug in VMAF; it's just the way that systems like that work.
That said, I don't think it's even fair to expect miracles until there's adaptive quantization. I don't think there has been a good encoder without adaptive quantization yet (and psyrdo too possibly?) so if we're just assuming Rav1e won't need it, that might be kind of unwarranted.
Right, and VMAF is insensitive to lots of adaptive quant psychovisual optimizations. For example, it doesn't do a good job of picking up on gradients or banding in darker scenes, for example. I'm guessing they didn't have a lot of variations of adaptive quant modes in their test set, so VMAF didn't get trained to distinguish between them.
And even with HVMAF, keyframe strobing will be hard to notice as it is just one bad frame per GOP. I don't know how many IDR frames VMAF was tested against; lots of codec tests wind up being a single long GOP, or scene-adaptive IDR only. So cases where a shot is longer than the max GOP duration may not have been sampled much.
benwaggoner
18th April 2019, 18:53
Xiph update on one year of AV1 (slides from NAB 2019) [PDF]
https://people.xiph.org/~negge/NAB2019.pdf
There certainly has been a ton of progress in the last year!
Although I am again frustrated by the lack of a real apples-to-apples subjective quality comparison. The only one given was libaom versus HEVC HM for ultra low latency (Slide 38). I don't know that the HM is even optimized for low latency; libaom has a lot more rate control than the typical reference encoder.
And even then, we can see that while Y-PSNR a bitrate increase of 5%, subjective MOS testing showed a decrease of -4%. Metrics are not closely coupled!
The VVC JEM, conversely, showed a 32% decrease for Y-PSNR and 30% decrease for MOS; much better correlated.
benwaggoner
18th April 2019, 19:09
Cmdlines:
x264 --preset veryslow --tune ssim --crf 16 -o test.x264.crf16.264 orig.i420.y4m
x265 --preset veryslow --tune ssim --crf 16 -o test.x265.crf16.hevc orig.i420.y4m
Why --tune ssim if targeting VMAF?
We know that --tune ssim looks subjectively worse than --tune film in x264 and not using --tune at all in x265.
vpxenc --codec=vp9 --frame-parallel=0 --tile-columns=1 --auto-alt-ref=6 --good --cpu-used=0 --tune=psnr --passes=2 --threads=2 --end-usage=q --cq-level=20 --test-decode=fatal --ivf -o test.vp9.cq20.ivf orig.i420.y4m
And why a different tune, PSNR, here?
SvtAv1EncApp.exe -i orig.i420.yuv -b test.svtav1.cq20.ivf -w 1280 -h 720 -q 20 -enc-mode 3 -fps-num 24000 -fps-denom 1001 -intra-period 23
aomenc --frame-parallel=0 --tile-columns=1 --auto-alt-ref=1 --cpu-used=4 --tune=psnr --passes=2 --threads=2 --row-mt=1 --end-usage=q --cq-level=20 --test-decode=fatal -o test.av1.cq20.webm orig.i420.y4m
VMAF: model used: vmaf_v0.6.1, pooling: harmonic_mean
Also PSNR.
It seems like the same tuning should be used across all encoders! Although tuning for a given metric and then comparing with that metric is more a test of mathematical correctness of rate control than something that says much about viewer experience.
We've seen data that shows libaom underperforms HM and particularly the VVC JEM in subjective metrics versus objective metrics. I'm guessing because libaom has baked in a lot of VMAF-tuned optimizations.
The gold standard for AV1's current competitiveness would be a double-blind comparison of subjective quality at the same total encoding time.
I guess I'm unsure on what exactly the goal of these particular tests are, or how they are expected to be fruitfully applied.
Double-blind testing is a whole lot of work, but inescapably necessary at this point in the codec universe. Things are going to be crazy over the next few years with H.264, HEVC, and AV1 today and with VVC, EVC, and AV2 on the horizon. VMAF is going to need a data set with subjective tests of the "flavor" of artifacts each produces to be able to make good inter-codec quality comparisons.
It'd be nice to know the relative encoding times as well.
SmilingWolf
18th April 2019, 19:29
--tune ssim on both x264 and x265 gave the best objective metrics scores for both PSNR-HVS-M and MS-SSIM (more than --tune psnr even when measuring PSNR-HVS-M).
VMAF has not been tested because it has been added to the encoding and scoring pipeline later than when I carried out the tune tests.
--tune psnr in the libvpx and libaom cmdlines is there as more of a way to make it explicit.
You can consider the whole "--tune" thing in those encoders as either a joke or a misnomer: instead of turning some knobs like they do on the x26X encoders, they set the RDO metric used during encoding.
To add insult to injury, in libaom out of 4 tunes (psnr, ssim, cdef-dist, daala-dist) 2 of them are usable only in single-threaded builds and give terrible results, and ssim is not even implemented, leaving --tune psnr as the only available one.
And it's not a question of setting it or leaving it alone, --tune psnr is the default (https://aomedia.googlesource.com/aom/+/08ee14a076e0a9f4405d14cd5d301bfa52f4cdf7/av1/av1_cx_iface.c#164) and there is no way to change or unset it. Whatever you do, whether you know or not, if you encode with libaom you're using --tune psnr.
Relative encoding times are unavailable (or, rather, unrealiable) because the machine comes under various loads because I use it while the encodes are running, and have set the pipeline to leave me at least a couple of free cores at all times.
sneaker_ger
19th April 2019, 00:18
Short decoding speed test on 10 year old Intel T3400 (2C2T laptop CPU, SSSE3, no SSE4)(both zeranoe's ffmpeg 20190417-8a3ed5a-win64-static, dav1d 20190410-44d0de4 -threads 4 -tilethreads 2), Chimera 720p 8 bit:
libaom: 749.241 (12 fps)
dav1d: 293.281 (30 fps)
Nintendo Maniac 64
19th April 2019, 00:37
Short decoding speed test:
libaom: 749.241 (12 fps)
dav1d: 293.281 (30 fps)
OK, just how are you going about benchmarking this?
Last time I inquired about this, the best way was pretty much just trial and error by using something like mkvtoolnix to set a given frame rate and then use madvr's OSD to see if there were any dropped frames while playing it back.
10 year old Intel T3400 (2C2T laptop CPU, SSSE3, no SSE4)
Kind of odd that CPU released in late 2008 when it uses the same architecture (Merom) as the original Conroe/Merom (65nm) Core 2 Duo from 2006 (though being a Pentium it has less L2 cache).
Weirder yet considering that the Wolfdale/Penryn 45nm Intel CPUs were available by then, and Nehalem was even available on desktop.
sneaker_ger
19th April 2019, 00:39
OK, just how are you going about benchmarking this?
Last time I inquired about this, the best way was pretty much just trial and error by using something like mkvtoolnix to set a given frame rate and then use madvr's OSD to see if there were any dropped frames while playing it back.
This is just using ffmpeg -benchmark and the fps values are averages. I didn't test for framedrops during difficult scenes.
soresu
19th April 2019, 02:45
Hmmm, the commit here (https://aomedia.googlesource.com/aom/+/d85d2477239d9b9bf36d94daccec79f48ab3784d) on the libaom experimental branch has the title "Add comparison between cnn and cdef/restoration."
I wonder if this means they are targetting an ML tool to replace CDEF, which wouldnt surprise me considering how Tim Terriberry mentioned CDEF being evaluated for a more efficient replacement during the latter stages of AV1 development.
Nintendo Maniac 64
19th April 2019, 05:36
This is just using ffmpeg -benchmark
I'll be honest, I'm actually completely unfamiliar with using ffmpeg...I am at least familiar with how to use command line, but I've no idea what to actually input to get ffmpeg's benchmark argument to actually function.
Could you perhaps share the exact entire command you used? From there I should be able to figure out how to get things going over here.
(software is a bit of a weak point for me - hardware is much more of my specialty)
nevcairiel
19th April 2019, 08:55
Could you perhaps share the exact entire command you used? From there I should be able to figure out how to get things going over here.
If you want to benchmark solely decoding, something like this:
ffmpeg -benchmark -i file.mp4 -f null -
Fill in the filename, of course, but don't move its position in the command line. :)
If you want to benchmark DirectShow on Windows, a far better option then your madVR hack is to use GraphStudioNext, which has View -> Performance Test, which lets you specify a file and a decoder, and it'll run only that decoder, without rendering involved.
hajj_3
19th April 2019, 10:46
DAV1D decoder v0.2.2 has been released, here are the changes:
- Large improvement on MSAC decoding with SSE, bringing 4-6% speed increase. The impact is important on SSSE3, SSE4 and AVX-2 cpus
- SSSE3 optimizations for all blocks size in itx
- SSSE3 optimizations for ipred_paeth and ipref_cfl (420, 422 and 444)
- Speed improvements on CDEF for SSE4 CPUs
- NEON optimizations for SGR and loop filter
- Minor crashes, improvements and build changes
dapperdan
19th April 2019, 12:48
Double-blind testing is a whole lot of work, but inescapably necessary at this point in the codec universe. Things are going to be crazy over the next few years with H.264, HEVC, and AV1 today and with VVC, EVC, and AV2 on the horizon. VMAF is going to need a data set with subjective tests of the "flavor" of artifacts each produces to be able to make good inter-codec quality comparisons.
It'd be nice to know the relative encoding times as well.
MSU did some subjective testing with their subjectify.us platform for their recent HEVC tests. Interestingly VP9 improved more than x265 when you compare SSIM to the subjective scores, though two other HEVC encoders beat both.
I think Netflix did a talk about how to use machine learning to reduce the number of comparisons that the real humans needed to do, making this kind of thing more efficient.
Nintendo Maniac 64
19th April 2019, 20:26
ffmpeg -benchmark -i file.mp4 -f null -
Yep that's exactly what I needed, and things are working now!
...except that the ffmpeg build I used seems to use the AOMedia AV1 decoder rather than dav1d. So now the question is where are you getting your ffmpeg builds so that they actually use dav1d?
Beelzebubu
19th April 2019, 20:44
Yep that's exactly what I needed, and things are working now!
...except that the ffmpeg build I used seems to use the AOMedia AV1 decoder rather than dav1d. So now the question is where are you getting your ffmpeg builds so that they actually use dav1d?
They probably use both, but prefer aom. To use dav1d, try -c:v libdav1d before -i.
foxyshadis
20th April 2019, 13:14
I'd like to solicit opinions on splitting this thread up, especially into aom, rav1e, dav1d, still image (avif) news, as well as solicitations to get the best quality command lines. I'd like to create a separate AV1 forum entirely at this point, but one megathread does not a forum make.
SmilingWolf
20th April 2019, 14:33
VMAF isn't a still image metric. Has anyone run a correlation for VMAF against subjective testing for still images?
Here you go, based on the TID2013 dataset (http://www.ponomarenko.info/tid2013.htm):
Actual profile:
Spearman: | Kendall:
PSNRHA 0.938 | PSNRHA 0.787
PSNRHMA 0.934 | PSNRHMA 0.777
PSNRHVS 0.926 | PSNRHVS 0.766
PSNRHVSM 0.917 | PSNRHVSM 0.749
FSIMc 0.915 | FSIMc 0.742
FSIM 0.911 | FSIM 0.736
WSNR 0.897 | WSNR 0.718
MSSIM 0.887 | MSSIM 0.697
VSNR 0.882 | VSNR 0.690
VMAF_v0.6.1 0.863 | VMAF_v0.6.1 0.675
VMAF_rb_v0.6.3 0.862 | VMAF_rb_v0.6.3 0.674
NQM 0.857 | NQM 0.666
PSNR 0.825 | PSNR 0.624
VIFP 0.815 | VIFP 0.621
PSNRc 0.803 | PSNRc 0.596
SSIM 0.788 | SSIM 0.577
Simple profile:
Spearman: | Kendall:
PSNRHA 0.953 | PSNRHA 0.818
PSNRHVS 0.951 | PSNRHVS 0.809
FSIM 0.949 | FSIM 0.795
FSIMc 0.947 | FSIMc 0.792
PSNRHVSM 0.938 | PSNRHMA 0.785
PSNRHMA 0.937 | PSNRHVSM 0.780
WSNR 0.933 | WSNR 0.772
PSNR 0.913 | PSNR 0.745
VSNR 0.912 | VSNR 0.731
MSSIM 0.905 | MSSIM 0.720
VIFP 0.897 | VIFP 0.714
VMAF_rb_v0.6.3 0.891 | VMAF_rb_v0.6.3 0.698
VMAF_v0.6.1 0.889 | VMAF_v0.6.1 0.696
PSNRc 0.876 | PSNRc 0.689
NQM 0.875 | NQM 0.681
SSIM 0.837 | SSIM 0.628
Full profile:
Spearman: | Kendall:
FSIMc 0.851 | FSIMc 0.666
PSNRHA 0.819 | PSNRHA 0.643
PSNRHMA 0.813 | PSNRHMA 0.631
FSIM 0.801 | FSIM 0.629
MSSIM 0.787 | MSSIM 0.607
VMAF_rb_v0.6.3 0.749 | VMAF_rb_v0.6.3 0.564
VMAF_v0.6.1 0.748 | VMAF_v0.6.1 0.563
PSNRc 0.687 | VSNR 0.508
VSNR 0.681 | PSNRHVS 0.507
PSNRHVS 0.654 | PSNRc 0.496
PSNR 0.640 | PSNRHVSM 0.481
SSIM 0.637 | PSNR 0.470
NQM 0.635 | NQM 0.466
PSNRHVSM 0.625 | SSIM 0.463
VIFP 0.608 | VIFP 0.456
WSNR 0.580 | WSNR 0.446
All bitmap images have been converted to raw full range YUV444P with ffmpeg and then measured with the vmafossexec program.
ffmpeg.exe -i i01_01_1.bmp -vf "scale=flags=accurate_rnd+bitexact+full_chroma_int+full_chroma_inp,format=yuvj444p" i01_01_1.bmp.yuv
vmafossexec.exe yuv444p 512 384 reference_images/i01.bmp.yuv distorted_images/i01_01_1.bmp.yuv model/vmaf_v0.6.1.pkl
vmafossexec.exe yuv444p 512 384 reference_images/i01.bmp.yuv distorted_images/i01_01_1.bmp.yuv model/vmaf_rb_v0.6.3/vmaf_rb_v0.6.3.pkl --ci
I'm also attaching the raw scores, for completeness sake.
A note on how to read the numbers:
from the paper (http://www.ponomarenko.info/papers/tid2013.pdf) I get the following: a SROCC of 0.95 is considered excellent, 0.90 is good, and 0.85 is barely acceptable.
bstrobl
20th April 2019, 16:32
I'd like to solicit opinions on splitting this thread up, especially into aom, rav1e, dav1d, still image (avif) news, as well as solicitations to get the best quality command lines. I'd like to create a separate AV1 forum entirely at this point, but one megathread does not a forum make.
Seems sensible, I would welcome a couple more threads.
TomV
20th April 2019, 17:59
I'd like to solicit opinions on splitting this thread up, especially into aom, rav1e, dav1d, still image (avif) news, as well as solicitations to get the best quality command lines. I'd like to create a separate AV1 forum entirely at this point, but one megathread does not a forum make.
Makes sense. Implementations should have separate threads from the main standardization effort and aomenc. AOM/AV1 news, legal discussions, etc. can be separate threads.
NikosD
20th April 2019, 18:31
I'd like to solicit opinions on splitting this thread up, especially into aom, rav1e, dav1d, still image (avif) news, as well as solicitations to get the best quality command lines. I'd like to create a separate AV1 forum entirely at this point, but one megathread does not a forum make. Too much effort, too many double posts, too much separated information and too much overhead in general.
Probably a separation of AV1 encoding and AV1 decoding would be more than enough for AV1 codec.
Audionut
21st April 2019, 13:40
Too much effort, too many double posts, too much separated information and too much overhead in general.
Agreed. Not busy enough yet. You can come back after a couple of days and still might only have a full page to read.
dapperdan
21st April 2019, 14:42
VMAF isn't designed for still images, but they do provide the tools to create your own VMAF for specific use cases (e.g. anime on a phone screen, or video game cobtebt) so it surprises me that no one has taken the framework and applied it to still images yet.
It should in theory be able to fuse the results of those other still image tests and create something even better aligned with human reported scores than any one alone. Presumably not Netflix's main use case but you'd think they deliver enough still images to make it worthwhile since they already have the skills.
SmilingWolf
21st April 2019, 15:04
VMAF isn't designed for still images, but they do provide the tools to create your own VMAF for specific use cases (e.g. anime on a phone screen, or video game cobtebt) so it surprises me that no one has taken the framework and applied it to still images yet.
It should in theory be able to fuse the results of those other still image tests and create something even better aligned with human reported scores than any one alone. Presumably not Netflix's main use case but you'd think they deliver enough still images to make it worthwhile since they already have the skills.
I was looking into this very matter earlier today and the main problem is, as always for this kind of problems, the lack of high quality MOS datasets. In particular, the only "extensive" dataset I've found is TID2013, and even that only comprises of 2 kinds of image compression distortions, for 25 images, at 5 intensities = 250 distorted images and relative scores.
When calculating the SROCC for only the "compression" distortions (JPEG and J2K) these are the results:
--- top 33%
PSNRHA 0.9686
DSSIM -0.9683
PSNRHVS 0.9677
PSNRHMA 0.9651
PSNRHVSM 0.9603
FSIMc 0.9589
FSIM 0.9580
VMAF_rb_v0.6.3 0.9524
SSIMULACRA -0.9519
VMAF_v0.6.1 0.9505
WSNR 0.9468
--- middle
MSSIM 0.9427
VIFP 0.9380
--- low 33%
PSNRc 0.9200
CQM 0.9190
PSNR 0.9170
VSNR 0.9162
SSIM 0.9147
NQM 0.9023
I also tweeted to Jon Sneyers about the dataset they used to validate SSIMULACRA, will see if he can release it indipendently of a blogpost that now, after two years, is probably not going to happen.
hajj_3
24th April 2019, 17:41
dav1d decoder v0.3.0 is out:
Changes for 0.3.0 'Sailfish':
------------------------------
This is the final release for the numerous speed improvements of 0.3.0-rc.
It mostly:
- Fixes an annoying crash on SSSE3 that happened in the itx functions
nevcairiel
24th April 2019, 19:14
Just because someone updates the changelog doesn't mean it has been released already. You can see actual release tags here, hopefully to help avoid confusing premature announcements:
https://code.videolan.org/videolan/dav1d/tags
There is no 0.3.0 yet. There will need to be one or two additional maintenance changes before that is the case. Probably in a day or two.
Motenai Yoda
24th April 2019, 19:47
there isn't a 3.0 tag yet, but you can always git the master branch with the last commit
dapperdan
25th April 2019, 12:35
Interesting snippet here:
https://www.streamingmedia.com/Articles/Editorial/Featured-Articles/NAB-2019-NGCodec-Talks-Hardware-Based-High-Quality-Live-Video-Encoding-131160.aspx
what we believe is, it's really a variant of their VP9.
Jan Ozer: When you say variant you mean...
Oliver Gunasekara: So Intel bought a company called eBrisk which they then open-sourced and that is the team that has delivered this. And what they did for time to market was take their VP9 implementation and just remove all the functionality that is not appropriate for AV1, tweak the syntax to have a legal AV1. So the end result is, it is an AV1 encoder but it doesn't perform anywhere near like the capabilities that AV1 can deliver. That will come in the future.
Basically claims the AV1-SVT encoder has barely begun development. If that's the case should be interesting to follow it's progress.
I noticed a ticket on their tracker where people were asking them to tag a pre-release so they could begin the process of integrating with Austria etc and the Devs didn't think it was ready for even a pre-release status, then a press release came out announcing version 1.0 was ready.
clsid
25th April 2019, 15:17
- Fixes an annoying crash on SSSE3 that happened in the itx functionsThis fix is only for non-Windows systems. So not important for most of us.
Mjpeg
25th April 2019, 15:51
Really nice writeup of adding tiles to rav1e - explains what tiles are nicely:
https://blog.rom1v.com/2019/04/implementing-tile-encoding-in-rav1e/
(credit: reddit av1 channel https://www.reddit.com/r/AV1/)
Beelzebubu
25th April 2019, 16:16
This fix is only for non-Windows systems. So not important for most of us.
It depends on the build configuration (stack alignment, to be exact), but this could trigger on all systems.
[edit] removed some nonsense because I misread your reply, sorry about that.
iwod
25th April 2019, 19:12
Interesting snippet here:
https://www.streamingmedia.com/Articles/Editorial/Featured-Articles/NAB-2019-Twitch-Talks-VP9-AV1-and-its-Five-Year-Encoding-Roadmap-131163.aspx
Basically claims the AV1-SVT encoder has barely begun development. If that's the case should be interesting to follow it's progress.
I noticed a ticket on their tracker where people were asking them to tag a pre-release so they could begin the process of integrating with Austria etc and the Devs didn't think it was ready for even a pre-release status, then a press release came out announcing version 1.0 was ready.
I cant find the quoted snippet anymore.
dapperdan
25th April 2019, 21:52
Sorry, posted wrong link:
https://www.streamingmedia.com/Articles/Editorial/Featured-Articles/NAB-2019-NGCodec-Talks-Hardware-Based-High-Quality-Live-Video-Encoding-131160.aspx
Blue_MiSfit
26th April 2019, 00:11
Good interview. I spoke with Oliver from NGCodec at NAB this year and I agree with a lot that was said.
FPGA is neat and disruptive because it's cloud native now, so you can get a lot of the flexibility of a pure software solution. I think there's a span of a few years where FPGAs make a lot of sense for dense live encoding, but then eventually ASIC encoders get even better / faster / more power efficient, and software encoders continue to offer better quality.
I think offline encoding for VOD streaming will still be done in software no matter what. I thought maybe there'd be a use case for FPGA AV1 encoding in the next year or so while software encoders (and CPUs) get fast enough to make AV1 encoding practical, but Oliver didn't seem to think this was a great use case. In retrospect, I'm inclined to agree.
ShogoXT
27th April 2019, 21:09
I'd like to solicit opinions on splitting this thread up, especially into aom, rav1e, dav1d, still image (avif) news, as well as solicitations to get the best quality command lines. I'd like to create a separate AV1 forum entirely at this point, but one megathread does not a forum make.
I think for sure that AV1 needs it's whole forum section like hevc has. Within it there for separate threads for rav1e, media industry news, etc.
It's very difficult for a sporadic doom9 reader like myself to follow with what has been discussed in this thread...
VincAlastor
3rd May 2019, 08:50
dav1d 0.3.0 decodes AV1 video’s 24% faster on SSSE3, 26% on SSE4.1 and 4% on AVX2 (all PC), and 12% faster on Arm64 (mobile).
https://medium.com/@ewoutterhoeven/dav1d-0-3-0-sailfish-armed-to-the-teeth-af5bbf845a16
singhkays
6th May 2019, 15:10
https://www.singhkays.com/blog/its-time-replace-gifs-with-av1-video/
I did a quick comparison of AV1 vs x264 vs VP9 at ultra low bitrates and how it can be used to replace GIFs in the browser
https://www.singhkays.com/blog/its-time-replace-gifs-with-av1-video/
I did a quick comparison of AV1 vs x264 vs VP9 at ultra low bitrates and how it can be used to replace GIFs in the browser
"80% better than H.264".
The claimed 50% bit rate reduction of VP9 vs. AVC is not substantiated by independent studies, or in practice by anyone. Also, you can't add bit rate reductions, you have to multiply bit rate ratios. If B encodes to the same quality as A at 0.5x the bit rate, and C encodes to the same quality as B with 0.7x the bit rate, C theoretically is 0.7 x 0.5 = 0.35x the bit rate of A... a 65% reduction, not 80%. But in practice studies have shown that AV1 is roughly on par with HEVC when measured with objective metrics (PSNR, SSIM, VMAF, etc.), delivering roughly a 50% bit rate reduction over AVC. However, measured subjectively it's behind HEVC. Although all of the above measures video, and not still still image compression, I expect the results for image (I frame only) compression to be quite close. Mozilla published a study (https://research.mozilla.org/2013/10/17/studying-lossy-image-compression-efficiency/)in 2013 which confirmed the superiority of HEVC still image compression over other existing formats. Strangely, the link to the study no longer works, but I saved a copy. Maybe one of the Mozilla guys can reshare it.
I agree that content publishers and web sites should be leveraging more powerful video codecs for still image compression. They can start with AVC, as device support is ubiquitous, and it is an improvement over JPEG and GIF. If/when they add support for an advanced codec, they will want the largest range of devices to support that codec, and they will want hardware decoding (for speed and vastly reduced power consumption). They can leverage HEVC (in HEIC container files... based on the ISO Base Media File Format, the evolution of .mov and .mp4) for the majority of devices which already have hardware HEVC support.
See https://nokiatech.github.io/heif/comparison.html
Google developed the WebP standard (https://en.wikipedia.org/wiki/WebP) for still image compression, based on VP8 technology. It's only 10% more efficient than JPEG. What you're proposing would seem to be a new version of WebP.
sneaker_ger
6th May 2019, 22:27
Mozilla published a study (https://research.mozilla.org/2013/10/17/studying-lossy-image-compression-efficiency/)in 2013 which confirmed the superiority of HEVC still image compression over other existing formats. Strangely, the link to the study no longer works, but I saved a copy. Maybe one of the Mozilla guys can reshare it.
http://web.archive.org/web/20160312174628/http://people.mozilla.org/~josh/lossy_compressed_image_study_october_2013/
soresu
7th May 2019, 16:28
Possible partial GPU acceleration coming in Dav1d during this years GSoC.
I wonder how much latency is incurred for only partial GPU decode, some guy going by atomnuker discussed possible GPU AV1 at FOSSDEM last year I think.
NikosD
7th May 2019, 17:45
Support for offloading some of the dav1d AV1 video decoder's work to GPUs using compute shaders in OpenGL/Vulkan/Metal/Direct3D. https://www.phoronix.com/scan.php?page=news_item&px=Google-GSoC-2019-Projects
nevcairiel
7th May 2019, 18:19
Don't get too excited quite yet, such efforts have in the past been problematic to get truely faster. We'll have to see.
soresu
7th May 2019, 19:57
I'd be less interested in faster playback using GPU in favor of lower power for phones prior to decoder ASIC rollouts, which probably wont be until at least next year.
I think Rockchip's recently announced RK3588 SoC for 2020 has an AV1 decoder, but details were sparse in the announcement.
The thing which concerns me most is the lack of any announcement from Qualcomm in support of AV1, considering their huge market share in Android devices
sneaker_ger
7th May 2019, 21:02
Yes, hardware support is disappointing.
IIRC for HEVC the Samsung and LG TVs were the first to have hardware decoding in late 2013/early 2014 after HEVC approval by ITU in April 2013. Now AV1 finalization was in June 2018 and still no hardware in sight, really. Intel probably Tiger Lake Q2 2020 at the earliest. Nvidia with Ampere also 2020? Nothing from Qualcomm.
Looks like they all started working on it pretty late.
alex1399
8th May 2019, 16:39
However, the feature of VP8 hardware decoding was barely seen in commercial products. When does the VP9 finalization happen?
nevcairiel
8th May 2019, 19:02
When does the VP9 finalization happen?
What do you mean?
Majority of new devices has VP9 support now.
benwaggoner
8th May 2019, 19:35
I'd be less interested in faster playback using GPU in favor of lower power for phones prior to decoder ASIC rollouts, which probably wont be until at least next year.
I think Rockchip's recently announced RK3588 SoC for 2020 has an AV1 decoder, but details were sparse in the announcement.
The thing which concerns me most is the lack of any announcement from Qualcomm in support of AV1, considering their huge market share in Android devices
I've heard indications that the extra transistors required for AV1 decoding are a lot higher than anticipated, and higher than the delta for HEVC. The cost in increased die size is a lot more than the savings from not paying MPEG-LA fees.
That would indicate a trend towards AV1 decode launching in high end chipsets first, and taking longer to get into lower-cost handsets.
benwaggoner
8th May 2019, 19:37
Yes, hardware support is disappointing.
IIRC for HEVC the Samsung and LG TVs were the first to have hardware decoding in late 2013/early 2014 after HEVC approval by ITU in April 2013. Now AV1 finalization was in June 2018 and still no hardware in sight, really. Intel probably Tiger Lake Q2 2020 at the earliest. Nvidia with Ampere also 2020? Nothing from Qualcomm.
Looks like they all started working on it pretty late.
The incremental cost to add AV1 into a CPU would be a lot lower than in a SoC, because a GPU already has so many transistors.
The make or break is how many extra mm^2 the decoder takes.
marcomsousa
9th May 2019, 16:19
Yes, hardware support is disappointing.
Now AV1 finalization was in June 2018 and still no hardware in sight, really. Intel probably Tiger Lake Q2 2020 at the earliest. Nvidia with Ampere also 2020? Nothing from Qualcomm.
Looks like they all started working on it pretty late.
What? HW support in less that one year? no way...
1st August 2018, 11:02
(...)
About the AV1 roadmap
We just complete phase 1.
In 1 year we complete phase 2.
In 2 to 3 years we complete phase 3.
In 4 to 5 years we complete phase 4.
https://forum.doom9.org/attachment.php?attachmentid=16445&stc=1&d=1533115434
So, now we are in phase 2. We need to wait 1 or 2 more years to have HW support. So we are in schedule.
* Next year we will starting see some high end CPU/GPU with AV1 HW decode support.
* In 2021 HW decode for low end CPU/GPU and encode for high end CPU/GPU. (some TVs, consoles)
* And in 2022 for all cpu/gpu. (All modern TVs, consoles)
Note: The first HW with encoding support will have bad quality comparing with software. They have to mature over the years.
And if you are asking, all the nextgen consoles that will release next year will not have AV1 HW decode support (because they release with a cpu of this year).
EwoutH
9th May 2019, 16:57
I've heard indications that the extra transistors required for AV1 decoding are a lot higher than anticipated, and higher than the delta for HEVC.
Hmm, this is quite strange. From what I've heard hardware vendors had a large say at the table with AV1 design, with the purpose of minimizing the complexity of encoding and decoding hardware. Do you know which functions or aspects of the fixed function hardware takes more die size than expected?
Also some two detail from Google Stadia: At launch they will use VP9 hardware encoding and somewhere in the future they will switch to AV1, but only when hardware encoding is available.
Stadia Streaming Tech: A Deep Dive (Google I/O'19) (https://www.youtube.com/watch?v=9Htdhz6Op1I)
15:19 for VP9, 22:53 for AV1, 31:45 for hardware video encoders.
Mr_Khyron
10th May 2019, 07:33
https://www.reddit.com/r/AV1/comments/bmp5v9/allegro_dvt_announces_ale210_encoder_ip_with_av1/
Allegro DVT released the first AV1 hardware encoder IP publicly known. The AL-E210 succeeds the AL-E200 with the main addition being support for the AV1 codec. It supports Profile 0, meaning 4:2:0 chroma subsampling with 8 and 10 bit color depth.
It claims support real-time encoding up to 4K, but with multiple cores up to 8K or (/ and?) 120fps. With the AL-E200 one core could encode 4K at 30fps, so with multiple cores this could be higher
Press release:Allegro DVT Introduces the Industry First Real-Time AV1 Video Encoder Hardware IP for 4K/UHD Video Encoding Applications (http://www.allegrodvt.com/allegro-dvt-introduces-the-industry-first-real-time-av1-video-encoder-hardware-ip-for-4kuhd-video-encoding-applications/)
Product page:AL-E210 Encoder IP (http://www.allegrodvt.com/products/silicon-ips/al-e210/)
I've heard indications that the extra transistors required for AV1 decoding are a lot higher than anticipated, and higher than the delta for HEVC. The cost in increased die size is a lot more than the savings from not paying MPEG-LA fees.
now I'm wondering, whether PVQ from the daala folks would've made it smaller ^^ :p
IIRC PVQ was less complex so easier to realize by hw, but was decided against, because of needed additional development time, even though it would've increased efficiency, too.
We'll know when gen2 will be realized (in 5 or 10 years) hehe
They can start with AVC, as device support is ubiquitous, and it is an improvement over JPEG and GIF.
That was like 5 years ago hehe (for animated pictures on major sites)
https://blog.embed.ly/what-twitter-isnt-telling-you-about-gifs-e1b74068cebd
https://rigor.com/blog/optimizing-animated-gifs-with-html5-video
they do it for reduced memory footprint, pause/play, reduced size, better caching, hw decoding, partial decoding, fast first playthrough, higher possible bit depth for animations etc.
EDIT:
for regular pictures it needs to be an image format people can download and share
https://www.cnet.com/news/facebook-tries-googles-webp-image-format-users-squawk/
hevc in ISOBMFF (heic/heif) might turn out a good solution if windows users don't need to add support manually through the microsoft store and android switches to it.
then whatsapp could switch to it as well (with transcoding back to jpeg for older devices)
EDIT2:
at least Firefox finally added WEBp support (including animations) so it counts as an alternative (especially when they update their internal codec again)
https://hacks.mozilla.org/2019/01/firefox-65-webp-flexbox-inspector-new-tooling/
foxyshadis
10th May 2019, 15:49
now I'm wondering, whether PVQ from the daala folks would've made it smaller ^^ :p
IIRC PVQ was less complex so easier to realize by hw, but was decided against, because of needed additional development time, even though it would've increased efficiency, too.
We'll know when gen2 will be realized (in 5 or 10 years) hehe
PVQ is smaller purely in the frequency domain, where Daala tried to keep everything. But it couldn't stay in the frequency domain with the current state of the art, so it ended up being both less efficient after everything was combined and optimised. Of course, Monty is constantly throwing new things at the wall to see what sticks, so maybe AV2 will change that.
EwoutH
14th May 2019, 22:46
AV1 is now quickly becoming mainstream, is it time for an separate Sub-Forum for AV1 (or AOM codecs in general)? I would love separate threads for aom, dav1d, rav1e, svt-av1, hardware acceleration and applications that support AV1.
Mjpeg
15th May 2019, 05:13
AV1 is now quickly becoming mainstream, is it time for an separate Sub-Forum for AV1 (or AOM codecs in general)? I would love separate threads for aom, dav1d, rav1e, svt-av1, hardware acceleration and applications that support AV1.
My vote is it's too early for that. It's nice reading about the encoder and decoder stuff in one place for now. Or put another way .. at what average messages/day rate should it split Maybe 25? Where well short of that now.
EwoutH
15th May 2019, 11:28
I think the number of messages per day is now also limited because small things don't seem important enough for a combined thread. With split threads I would report way more often about small changes in dav1d performance, SVT-AV1's troubles with Open Source and rav1e and aom related news. Now it just doesn't seem important enough to report on in a general thread.
For example:
Henrik Gramner just made some really interesting optimizations to dav1d's msac function in MR696 (https://code.videolan.org/videolan/dav1d/merge_requests/696) and MR697 (https://code.videolan.org/videolan/dav1d/merge_requests/697), resulting in about 1.5% faster performance on all x86-64 platforms. Martin Storsjö (wbs) is working on adding the same improvements with NEON for Arm64.
dav1d just got a logo (https://code.videolan.org/videolan/dav1d/merge_requests/681)
For benchmarking purposed, dav1d now can be run (https://code.videolan.org/videolan/dav1d/merge_requests/670) at fixed frame rates
SVT-AV1 just reduced memory usage (https://github.com/OpenVisualCloud/SVT-AV1/pull/242) by 30 to 70 percent
Someone is working on a decoder (https://github.com/OpenVisualCloud/SVT-AV1/pull/246) in SVT-AV1
Some effort is made to get SVT-AV1's development process more transparant and compliant with open source practices in this issue (https://github.com/OpenVisualCloud/SVT-AV1/issues/238)
soresu
15th May 2019, 15:58
A thought I had, was to make splitting up the general thread easier, you could just take everything pre-bitstream freeze and label it as such.
Thats got to be a serious portion of the thread, and mostly about AV1 development?
benwaggoner
15th May 2019, 23:31
AV1 is now quickly becoming mainstream, is it time for an separate Sub-Forum for AV1 (or AOM codecs in general)? I would love separate threads for aom, dav1d, rav1e, svt-av1, hardware acceleration and applications that support AV1.
Certainly tools and infrastructure are getting improved and deployed. But who actually has live content in AV1 other than YouTube?
dapperdan
16th May 2019, 13:11
But who actually has live content in AV1 other than YouTube?
Facebook appear to have it all hooked up to deploy AV1 since about a year ago, but whether they actually use it for anything other than testing yet I don't know.
Anyone use Facebook to watch popular videos? Do they have a "stats for nerds" equivalent?
Facebook appear to have it all hooked up to deploy AV1 since about a year ago, but whether they actually use it for anything other than testing yet I don't know.
Anyone use Facebook to watch popular videos? Do they have a "stats for nerds" equivalent?
The likelihood is close to Zero. Although that is just a pure guess.
The latest report from Facebook shows 92% of its users are on Mobile Apps. That is up from 90% YoY and trending towards 100%. I doubt they would want to push AV1 for the sake of saving bandwidth ( even less than a rounding error for Facebook ) at the expense of user battery. ( Which will lower user engagement time, and lower ads revenue, everything Facebook is valued of )
EwoutH
19th May 2019, 19:25
SVT-AV1 v0.5.0 (https://github.com/OpenVisualCloud/SVT-AV1/releases/tag/v0.5.0) just got tagged!
Windows build (https://ci.appveyor.com/project/OpenVisualCloud/svt-av1/builds/24654459/artifacts)
EwoutH
20th May 2019, 13:31
Amphion announced the CS8142 (PDF (http://www.amphionsemi.com/wp-content/uploads/Amphion-CS8142-One-Pager.pdf)), the first AV1 hardware decoder!
SVT-AV1 just opened a huge PR with unit tests: #260 (https://github.com/OpenVisualCloud/SVT-AV1/pull/260). They also have a lot of development branches (https://github.com/hguermaz/SVT-AV1/branches) open, including one with a lot of AVX2 and AVX512 optimizations (https://github.com/hguermaz/SVT-AV1/commits/lm-opt).
dapperdan
21st May 2019, 18:50
Faster decoding in Firefox thanks to Dav1d:
https://blog.mozilla.org/blog/2019/05/21/latest-firefox-release-is-faster-than-ever/
We have seen great growth in the use of AV1 even in just a few months, with our latest figures showing that 11.8% of video playback in Firefox Beta used AV1, up from 0.85% in February and 3% in March.
sneaker_ger
21st May 2019, 19:49
11.8%? How? Did Youtube roll out AV1 big time?
dapperdan
21st May 2019, 21:33
The 80/20 rule would suggest you only need to encode 20% of the currently popular videos to get 80% of the views so getting up to 11% should be relatively straightforward if YouTube focus on the big hits.
lvqcl
25th May 2019, 16:37
11.8%? How? Did Youtube roll out AV1 big time?
It seems that Youtube now streams SD video (480p and below) in AV1 format by default.
hin12
25th May 2019, 17:52
It seems that Youtube now streams SD video (480p and below) in AV1 format by default.
All channels with >1 million subs I've seen have AV1 available in SD for videos uploaded from January onwards. I posted this on March 6th:
Is 480p the highest quality option for all AV1 videos on YouTube? I'm seeing that a lot of popular videos are being encoded to AV1, but only at the 144p, 240p, 360p and 480p resolutions.
You had to enable the setting on testtube back then, though.
lvqcl
25th May 2019, 18:40
You had to enable the setting on testtube back then, though.
And now it's unnecessary; that's the point.
stax76
28th May 2019, 09:16
What is currently the fastest AV1 encoder?
Tommy Carrot
28th May 2019, 12:35
What is currently the fastest AV1 encoder?
SVT-AV1 with the default preset. The quality is bad though.
Aomenc with --cpu-used=8 --rt is even faster, but the quality is absolutely rancid.
stax76
28th May 2019, 12:50
SVT-AV1 with the default preset. The quality is bad though.
Aomenc with --cpu-used=8 --rt is even faster, but the quality is absolutely rancid.
I'll be waiting for improved encoders then.
ChaosKing
28th May 2019, 13:26
rav1e is "fast" with ok quality.
stax76
28th May 2019, 13:36
rav1e is "fast" with ok quality.
Last time I tried rav1e 2019-04-30 and it was SLOW AS HELL, like < 1 fps.
ChaosKing
28th May 2019, 13:38
Last time I tried rav1e 2019-04-30 and it was SLOW AS HELL, like < 1 fps.
1 fps is fast for av1 :D
soresu
28th May 2019, 14:19
I'm just wondering if the quality regression in rav1e has been fixed yet, with activity masking coming from GSOC it should be in a pretty good place quality wise if nothing is still broken.
VincAlastor
29th May 2019, 06:56
What is currently the fastest AV1 encoder?
Again thank you very much for staxrip!!!
i use in staxrip a custom cmd line for ffmpeg aomenc from streaming media experts.
-c:v libaom-av1 -crf 40 -b:v 0 -strict experimental -threads 24 -tiles 8x2 -cpu-used 5 -row-mt 1
https://www.streamingmedia.com/Articles/Editorial/Featured-Articles/Good-News-AV1-Encoding-Times-Drop-to-Near-Reasonable-Levels-130284.aspx
https://www.reddit.com/r/AV1/comments/9afr5d/my_1700x_encoding_speed_results_mostly_with/
also new version of rav1e with tiles support brings up to 3 fps for fullhd content
https://blog.rom1v.com/2019/04/implementing-tile-encoding-in-rav1e/
birdie
29th May 2019, 13:35
At this time shouldn't we consider AV1 a failure and move on to newer codecs, e.g AV2?
It doesn't have fast enough decoders to decode on mobile at 1080p on most devices (>80%).
It still doesn't have encoders which are anywhere fast enough to be usable by mere mortals.
Its hardware adoption is not there - the spec was finalized almost half a year ago, and AV1 is nowhere to be seen in Zen 2.0 (Ryzen 3000), Radeon RDNA 5700 or Intel Ice Lake. No word on its decoding acceleration even in the recently announced Arm's Cortex-A77/Mali-G77.
nevcairiel
29th May 2019, 15:19
You do realize that a newer codec would end up even more complex, and thus even slower, and also slowing down hardware adoption even more?
All your points can be applied for any new codec.
- Mobile hardware, especially on the low-end, is inherently slow. The dav1d AV1 decoder can decode 1080p on mid-range and high-end mobile devices. But for mainstream roll out, thats not going to be used anyway, because it uses too much battery.
- Encoder development takes years. You must not have been around when any other codec was new. On top of that, any new codec is also always going to be slower then previous codecs. You pay for quality or compression with speed.
- Hardware development also takes years. Noone knolwedgeable ever realistically expected AV1 to show up in hardware before the end of 2020 or so. The turn-around times for hardware are really long. Hardware you see launch/announced today was long through the design process before AV1 was finished.
Any other new codec would be in the exact same situation a year or so after the spec was officially finalized. In fact this is already looking pretty good on adoption, most major browsers now include AV1 decoders, and YouTube is rolling out content.
Basically, do some more research on how codec lifetime has been in the past for any other codec.
soresu
29th May 2019, 16:50
You do realize that a newer codec would end up even more complex, and thus even slower, and also slowing down hardware adoption even more?
All your points can be applied for any new codec.
- Mobile hardware, especially on the low-end, is inherently slow. The dav1d AV1 decoder can decode 1080p on mid-range and high-end mobile devices. But for mainstream roll out, thats not going to be used anyway, because it uses too much battery.
- Encoder development takes years. You must not have been around when any other codec was new. On top of that, any new codec is also always going to be slower then previous codecs. You pay for quality or compression with speed.
- Hardware development also takes years. Noone knolwedgeable ever realistically expected AV1 to show up in hardware before the end of 2020 or so. The turn-around times for hardware are really long. Hardware you see launch/announced today was long through the design process before AV1 was finished.
Any other new codec would be in the exact same situation a year or so after the spec was officially finalized. In fact this is already looking pretty good on adoption, most major browsers now include AV1 decoders, and YouTube is rolling out content.
Basically, do some more research on how codec lifetime has been in the past for any other codec.
I wish I could upvote this, you pretty much laid it out nice and clear there.
soresu
29th May 2019, 16:54
To add to Nevcariel's reply to Birdie, I would not expect any ASIC decoders to appear on AMD CPU's as they are generally integrated into the GPU design, which also means they end up in APU's aswell - once they have an ASIC it will be on the nearest APU or GPU release from that point.
sneaker_ger
29th May 2019, 21:43
We'll have to wait and see when or if free AV1 encoders surface that can beat e.g. x265 at not more than 10x slow down. It's hard to please the doom9 crowd. VP9 never got to that point, maybe AV1 will be the same.
nevcairiel
30th May 2019, 11:10
We'll have to wait and see when or if free AV1 encoders surface that can beat e.g. x265 at not more than 10x slow down. It's hard to please the doom9 crowd. VP9 never got to that point, maybe AV1 will be the same.
Encoders are no longer being made for this crowd. The primary design goal is massive-scale cloud encoding for YouTube, Netflix, Amazon, and everyone else that fits the encode-once, download hundreds of thousands of times scenario.
In such a scenario, even the slowest encoder is acceptable if it saves enough bytes.
In that scenario, VP9 also didn't fail. It gets used for a lot of content on the web.
birdie
30th May 2019, 11:32
Encoders are no longer being made for this crowd. The primary design goal is massive-scale cloud encoding for YouTube, Netflix, Amazon, and everyone else that fits the encode-once, download hundreds of thousands of times scenario.
In such a scenario, even the slowest encoder is acceptable if it saves enough bytes.
In that scenario, VP9 also didn't fail. It gets used for a lot of content on the web.
This is what I was talking about.
AV1 is not a codec for masses. It's a very special codec for content delivery. That's it. And that makes it and its discussion kinda worthless.
And VVC is already miles better/faster/more effective than AV1.
richardpl
30th May 2019, 11:36
And VVC is already miles better/faster/more effective than AV1.
Whatever you think.
dapperdan
30th May 2019, 12:25
I'm not sure it's fair to say VP9 and AV1 were designed purely for those use cases, it's just that that's one of the easiest niches to win if you're a next-gen codec where encoding time is traded for smaller size.
That lets them use it profitably from day one while expanding further into other niches. It probably has impacts on how much effort goes into multithreading or other features that this use case doesn't need.
Libvpx seems to equal x265, subjectively, objectively and in encoding time in the recent MSU study (and both are near the head of the pack) The argument now seems to be that it's the rate control that makes it unsuitable for many users despite good showing in test scenarios. But libvpx having bad rate control is a rather different claim than the VP9 format being 10x slower than it should be and therefore a disaster.
sneaker_ger
30th May 2019, 12:35
Doom9 crowd can't exactly live on the potential of a spec. It needs (free/cheap) access to a well-rounded encoder implementation with a sweet spot on bitrate distribution/AQ/speed. x264 and x265 meet those demands. libvpx? Not so much.
lvqcl
30th May 2019, 13:02
(free/cheap) access to a well-rounded encoder implementation with a sweet spot on bitrate distribution/AQ/speed.
I suspect that we won't see such encoders for VVC.
nevcairiel
30th May 2019, 13:41
I'm not sure it's fair to say VP9 and AV1 were designed purely for those use cases, it's just that that's one of the easiest niches to win if you're a next-gen codec where encoding time is traded for smaller size.
Its not a design target for the codec itself, because the codec really doesn't care. I'm talking about encoders. The huge open-source push that made x264 as great as it is for "personal" encodes is unlikely to repeat itself. Companies driving encoder development do not target doom9ers. You can already see this on x265 where the community involvement is pretty low, and this will only get worse as the computational complexity of codecs goes up and the "personal use" usecases get less attractive.
This is what I was talking about.
AV1 is not a codec for masses. It's a very special codec for content delivery. That's it. And that makes it and its discussion kinda worthless.
This will not change with any future codec. Not with AV2, not with VVC, or anything that follows. The computational complexity increase in all those future codecs just makes it impractical for "hobbyist" use.
And VVC is already miles better/faster/more effective than AV1.
Dream on. A MPEG reference encoder has never won any price in any of those categories.
Doom9 crowd can't exactly live on the potential of a spec. It needs (free/cheap) access to a well-rounded encoder implementation with a sweet spot on bitrate distribution/AQ/speed. x264 and x265 meet those demands. libvpx? Not so much.
And unless the "Doom9 crowd" is going to develop their own encoder, noone is going to do that for them. Thats how x264 was ultimately born, it was made by video enthusiasts, not a company.
---
It really all comes down to the inherent complexity of newer codecs. As it gets more and more impractical to encode them due to the speed, less and less people and smaller companies are going to use them, and the entire ecosystem shifts over to only the bigger players that have the volume to host huge encoding farms. Even if there was a perfect free AV1 encoder, it would still be slow. There is no going fast without sacrificing quality or compression, at which point you eventually cross into the domain of already established codecs, and you lose your reason to even use the newer codec in the first place. Hence, development is no longer targeting individuals or small companies.
sneaker_ger
30th May 2019, 14:01
And unless the "Doom9 crowd" is going to develop their own encoder, noone is going to do that for them. Thats how x264 was ultimately born, it was made by video enthusiasts, not a company.
I'm not judging, just saying how it is. People see the advertisement for the new shiny codec and can't wait to profit from "50% less bitrate". They need to reduce their hopes. As you say AV1 today isn't for them and maybe never will. Of course I wouldn't mind to be proved wrong in the future. :)
benwaggoner
30th May 2019, 22:20
Encoders are no longer being made for this crowd. The primary design goal is massive-scale cloud encoding for YouTube, Netflix, Amazon, and everyone else that fits the encode-once, download hundreds of thousands of times scenario.
In such a scenario, even the slowest encoder is acceptable if it saves enough bytes.
Oh, these things are always down to cost benefit. Google is just willing to subsidize VPx/AV1 to a huge degree, and likely has a lot of spot-idle capacity in data centers to do the encoding.
But high quality encoding doesn't work with chunks of a few seconds. Encoding longer sequences allows for IDRs and shot changes and more aggressive VBV use. YouTube can have quite a bit of keyframe strobing with difficult content for these reasons. YouTube quality wouldn't be acceptable for lots of premium content.
There has never been a VP9 encoder that offers sufficient quality OR performance for premium content, and there isn' one for AV1 yet either. I don't think this is because the VP9 bitstream wasn't capable of it, it's just that no one wrote an encoder with good psychovisual tuning, intra-frame parallelism, and other stuff.
Encoders are a real chicken-egg problem. There needs to be enough companies willing to pay for better encoders to create a competitive market so that companies work hard to make better encoders than each other. That market never emerged for VP9, so libvpx never saw the kind of quality and performance improvement of, say, the H.264 or HEVC reference encoders to the best available commercial encoders.
There is clearly more interest in AV1 than there ever was for VP9, and more quality innovation already than VP9 has had to date. Which is very promising.
In that scenario, VP9 also didn't fail. It gets used for a lot of content on the web.
Other than YouTube?
VP9/s niche is user-generated non-DRM social media content. And in practice, H.264 would have offered at least equivalent quality at equal encoding time due to faster and more psychovisually tuned encoders. An x264 running at veryslow speed is going to be the same speed as a quite low-complexity VP9 encoder, especially on high-core systems. The quality comparison for high volume use are done at quality @ bitrate @ time.
benwaggoner
30th May 2019, 22:28
I'm not judging, just saying how it is. People see the advertisement for the new shiny codec and can't wait to profit from "50% less bitrate". They need to reduce their hopes. As you say AV1 today isn't for them and maybe never will. Of course I wouldn't mind to be proved wrong in the future. :)
We have seen 50% improvements generation-to-generation with the MPEG codecs. But that's at the same point of development. An HEVC enocoder with seven years of refinement isn't going to be only half as good as a VVC reference implementation. But comparing HEVC encoders seven years after spec freeze to VVC seven years after code freeze probably will show ~50% bitrate savings. Some content will even be more. Film grain synthesis could result in 75% reduction in some of the most difficult content, for example.
AV1 wasn't ever promised to be more than 20% better, and even that was in mean PSNR, not psychovisually. AV1 encoders haven't demonstrated any advantage over HEVC with subjective quality, even with the reference encoders.
The VPx code base started VERY heavily PSNR-tuned, which may be way it does well with that today. I don't have any reason to think that is a limitation of the AV1 bitstream versus just the libaom encoder.
But the general case of "AV1 can deliver the same subjective quality at meaningfully lower bitrates than HEVC" has yet to be demonstrated.
nevcairiel
30th May 2019, 23:19
Other than YouTube?
Netflix uses it for mobile devices, with the EVE encoder I believe, which they have found to be equal to x265 quality at the time, if not slightly better in some cases (last blog on that was from December)
TD-Linux
30th May 2019, 23:46
Last time I tried rav1e 2019-04-30 and it was SLOW AS HELL, like < 1 fps.
You can try speeding it up with the --tile-cols-log2 and --tile-rows-log2 option if you have a lot of cores. I'm working on making these more automatic.
IgorC
31st May 2019, 02:10
But the general case of "AV1 can deliver the same subjective quality at meaningfully lower bitrates than HEVC" has yet to be demonstrated.
Unfortunately this can be the case.
Beamr HEVC 2 Mbps does visually better than AV1 3 Mbps. Both VP9 and AV1 are heavily optimizied for specific metrics.
Beamr 2 Mbps (https://drive.google.com/open?id=12rOGjYJI2pmKUVZYgGiQD8E1awI7e75l)
AV1 3 Mbps (https://demo.bitmovin.com/public/firefox/av1/)
Asilurr
31st May 2019, 09:55
Libvpx seems to equal x265, subjectively, objectively and in encoding time in the recent MSU study (and both are near the head of the pack).At its worst, a contemporary version of libvpx will encode 8-bit 4:2:0 content as slowly as a contemporary version of x265. Often libvpx is noticeably faster than x265, given similar approaches to encoding complexity.
Here's a quick&dirty test, using an admittedly peculiar content source. Lena_std.tif from lenna.org, RGB24 converted to 8-bit 4:2:0, both downscaled and then upscaled two times each (i.e. 256x256, 1024x1024) into 250 frames long videos. All encoders are the latest versions available today on Wolfberry's public GDrive (https://drive.google.com/drive/folders/1xZQABtoaSFgGu11YstmHKYLzO3elemlC). All encoding times are reported by Win10's PowerShell: Measure-Command {start-process process -argumentlist "args" -Wait}, and expressed in milliseconds.
x264 common settings: --crf 15 --preset veryslow --tune stillimage
405 --threads 1: 03130.76 || 03142.22 || 03119.85 (03130.94)
405 --threads 2: 02101.52 || 02105.54 || 02105.00 (02104.02)
420 --threads 1: 26487.85 || 27508.10 || 26479.85 (26825.27)
420 --threads 2: 16336.86 || 15308.29 || 15327.17 (15657.44)
x265 common settings: --crf 15 --preset veryslow --frame-threads 1 --lookahead-slices 1
505 --no-wpp: 05160.82 || 05161.18 || 05154.55 (05158.85)
505 --wpp: 04131.71 || 04142.67 || 04131.78 (04135.39)
520 --no-wpp: 68123.55 || 68134.36 || 67103.71 (67787.21)
520 --wpp: 36649.91 || 37674.62 || 34628.27 (36317.60)
x265 common settings: --crf 15 --preset placebo --cu-lossless --rd-refine --tskip -qg-size 64 --ref 6 --bframes 16 --me sea --subme 7 --frame-threads 1 --lookahead-slices 1
505 --no-wpp: 063053.81 || 061024.00 || 062028.34 (062035.38)
505 --wpp: 039684.32 || 038664.65 || 038684.41 (039011.13)
520 --no-wpp: 796345.10 || 796345.10 || 796345.10 (796345.10)
520 --wpp: 235701.00 || 235701.00 || 235701.00 (235701.00)
libvpx common settings: --lag-in-frames=25 --passes=2 --end-usage=q --cq-level=20 --good --cpu-used=0 --kf-max-dist=250 --auto-alt-ref=6 --tile-rows=0 --enable-tpl=1 --frame-parallel=0 --ivf
905 --tile-columns=0 --row-mt=0 --threads=1: 05167.92 || 05164.62 || 05167.18 (05166.57)
905 --tile-columns=5 --row-mt=1 --threads=2: 04133.40 || 04143.00 || 04147.51 (04141.44)
920 --tile-columns=0 --row-mt=0 --threads=1: 50871.23 || 49853.74 || 49845.72 (50190.23)
920 --tile-columns=5 --row-mt=1 --threads=2: 31568.60 || 31571.60 || 28518.96 (30553.05)
XAB .. output type
X .. 4 for x264, 5 for x265, 9 for libvpx
AB .. 05 for 256x256, 20 for 1024x1024
Only one x265 beyondplacebo encode at each resolution due to appalling performance.
The average values are reported in brackets, for each output type and encoding settings.
As some sort of hardware footprint is inevitable (HDD versus SSD, memory specs and amount, platform, processor specs), the above targeted the lowest common denominator: parallelism disabled, then parallelism crippled to check the scaling. Conclusions based on this small-scale test:
1. x264 is at least twice faster than either x265 or lipvpx, at comparable encoding complexity.
2. x265 parallelizes better than libvpx as encoding complexity increases.
3. x265 versylow is at best comparable with the slowest libvpx settings, lagging behind as the resolution increases.
4. x265 beyondplacebo is significantly slower than the slowest libvpx settings.
As mentioned above, this is just one possible comparison. The enthusiasts should definitely do their own, and only then report whichever encoder to be the speed demon. It's very hard to conceive that chunk-based libvpx (the obvious caveats aside: logistical hassle and quality issues) is ever slower than x265 regardless of the hardware footprint, each of them at its highest encoding complexity. But instead of testing themselves (with whatever source they please, at whatever resolution and bit depth, on whatever encoding machine), people prefer to reiterate ad absurdum that libvpx is always slower than x265. And it was, indeed, years ago.
mandarinka
31st May 2019, 13:30
Last time I tried rav1e 2019-04-30 and it was SLOW AS HELL, like < 1 fps.
If it was on Windows, it is possible you downloaded a build with assembly disabled. Last time I ran test, it was the case (official windows build from the project), and I only got that information like two weeks later after spending days watching the atrociously slow FPS counter. I guess devs expect everybody interested to be on Linux or something. Naturally it would help if the encoder signalled assembly being used like good old x264/x265 do (good idea for linux distros too, because there used to be clueless packagers that disabled ASM accidentaly) or warned when it's not, but I think nobody got around to code that yet.
Encoders are no longer being made for this crowd. The primary design goal is massive-scale cloud encoding for YouTube, Netflix, Amazon, and everyone else that fits the encode-once, download hundreds of thousands of times scenario.
In such a scenario, even the slowest encoder is acceptable if it saves enough bytes.
In that scenario, VP9 also didn't fail. It gets used for a lot of content on the web.
Sadly it's not just about speed, but about quality too. They seem to just care about some metrics on low bitrate content, actual high quality encoding with transparent quality gets a finger.
Its not a design target for the codec itself, because the codec really doesn't care. I'm talking about encoders. The huge open-source push that made x264 as great as it is for "personal" encodes is unlikely to repeat itself. Companies driving encoder development do not target doom9ers. You can already see this on x265 where the community involvement is pretty low, and this will only get worse as the computational complexity of codecs goes up and the "personal use" usecases get less attractive.
This will not change with any future codec. Not with AV2, not with VVC, or anything that follows. The computational complexity increase in all those future codecs just makes it impractical for "hobbyist" use.
There might also be another factor at play: Google and friends siphon away those enthusiasts to work on decoders/encoders for whatever formats they come with. I guess the developers are happy doing that, but I can't help thinking that resource/creative power was kind of wasted polishing me too project like VP9, which never got good anyway. Perhaps it will be a waste with AV1 too, we shall see (hopefully not but track record from predecessors isn't good).
sneaker_ger
31st May 2019, 13:50
But high quality encoding doesn't work with chunks of a few seconds. Encoding longer sequences allows for IDRs and shot changes and more aggressive VBV use. YouTube can have quite a bit of keyframe strobing with difficult content for these reasons. YouTube quality wouldn't be acceptable for lots of premium content.
If only Youtube had like .. multiple .. videos coming in daily. Then they could encode them simultaneously on a single CPU each. :devil: (And serve AVC or fast setting VP9/AV1 until they are done.)
You could do a fast first pass for scenechange detection and vbv estimation and send the chunks along with that info.
nevcairiel
31st May 2019, 16:37
actual high quality encoding with transparent quality gets a finger.
This is not something any streaming service does though, so why should they invest into developing stuff for a goal they don't even need?
Its just how it goes. And for UHD Blu-ray discs, they can just throw massive bitrates at it to solve any such issues.
As said above, it all comes back to the same thing: If you want a codec for a use-case that noone else focuses on, then do the work, instead of the complaining. :)
mandarinka
31st May 2019, 19:26
This is not something any streaming service does though, so why should they invest into developing stuff for a goal they don't even need?
Its just how it goes. And for UHD Blu-ray discs, they can just throw massive bitrates at it to solve any such issues.
As said above, it all comes back to the same thing: If you want a codec for a use-case that noone else focuses on, then do the work, instead of the complaining. :)
That's okay response to complains, but irrelevant response to criticism of the technicals. BTW, I'm not sure all the free software developers contributing are paid by those streaming companies, yet they contribute to this software that is directed at the companies interest only as you say. Perhaps there is some exploitation of volunteer goodwill?
Also you can't really expect users to jump to programming and so that shouldn't really be thought of as a solution. Even if with the open source principles, you are actually entitled to it. But, it's not like, easy to do. Second, users want to do the using, not switching roles to open source programmers.
It's probably easier to just not use such software (and use x264, x265, whatever is better). Though with this "open formats" movement/fandom/advocacy, there is this curious anomaly that people are so strong subscribers to the concept that they want to use it even if it "is not for them" at all... well there are lots of weird layers to these debates.
I think this superstrong mindshare that bends people's views is another reason why pointing out the technical problems should keep being done. And not brushed aside by "it's not for you" arguments. The enthusiasts al over the internets seem to think "it" is for them, by the look of it.
Maybe sometimes they also think all those winning compression tests and netflix blogs are also for them (while perhaps they also aren't?).
soresu
31st May 2019, 21:39
There might also be another factor at play: Google and friends siphon away those enthusiasts to work on decoders/encoders for whatever formats they come with. I guess the developers are happy doing that, but I can't help thinking that resource/creative power was kind of wasted polishing me too project like VP9, which never got good anyway. Perhaps it will be a waste with AV1 too, we shall see (hopefully not but track record from predecessors isn't good).
If codec engineers are being siphoned away from anywhere it is other codec projects.
The guy under the name xiph_mont was working on a new audio codec called Ghost when Daala started up in earnest and it was abandoned - and he later moved to AV1 development with most of those that worked on Daala.
dapperdan
1st June 2019, 07:39
Xiphmont was employed by Red Hat and Mozilla, both companies that have a business model (and a mission) incompatible with codec licenses and therefore strong motivation for "me-too" codecs that they can integrate properly.
(Which reminds me, VP9 has been the default codec choice for webRTC in Firefox and Chrome for a couple of years. I can't quickly find any stats on what kind of usage this gets. Originally the browsers agreed a compromise of supporting both VP8 and H.264 baseline and Firefox shipped that via a licencing hack where Cisco provide the binary blob and it doesn't cost them anything in licence fees because they were already at the annual cap.)
Tommy Carrot
2nd June 2019, 13:56
What's the difference between deltaq and AQ in aomenc? Deltaq changes the quantizer of the frames, and AQ changes the quantizers of the blocks within the frames, or am i completely wrong?
Also, what is tpl-model, what does it do?
TD-Linux
3rd June 2019, 19:22
What's the difference between deltaq and AQ in aomenc? Deltaq changes the quantizer of the frames, and AQ changes the quantizers of the blocks within the frames, or am i completely wrong?
Deltaq is at superblock granularity, whereas "AQ" uses segment support which is down to 4x4 granularity. The two bitstream features are sorta redundant but their coding is optimized for different uses - Deltaq was designed for sub-frame rate targeting, and segments are more for mbtree/psy purposes.
Tommy Carrot
4th June 2019, 12:05
Deltaq is at superblock granularity, whereas "AQ" uses segment support which is down to 4x4 granularity. The two bitstream features are sorta redundant but their coding is optimized for different uses - Deltaq was designed for sub-frame rate targeting, and segments are more for mbtree/psy purposes.
Thanks for the explanation.
hajj_3
6th June 2019, 08:39
vlc 3.0.7 is out, not sure how to check what version of dav1d it uses though.
birdie
6th June 2019, 11:03
vlc 3.0.7 is out, not sure how to check what version of dav1d it uses though.
It's been using dav1d since 3.0.5.
hajj_3
6th June 2019, 11:04
It's been using dav1d since 3.0.5.
yes, but i wanted to know what version of dav1d v3.0.7 is using.
birdie
6th June 2019, 11:26
yes, but i wanted to know what version of dav1d v3.0.7 is using.
0.3.1
vidschlub
9th June 2019, 11:53
At this time shouldn't we consider AV1 a failure and move on to newer codecs, e.g AV2?
It doesn't have fast enough decoders to decode on mobile at 1080p on most devices (>80%).
It still doesn't have encoders which are anywhere fast enough to be usable by mere mortals.
Its hardware adoption is not there - the spec was finalized almost half a year ago, and AV1 is nowhere to be seen in Zen 2.0 (Ryzen 3000), Radeon RDNA 5700 or Intel Ice Lake. No word on its decoding acceleration even in the recently announced Arm's Cortex-A77/Mali-G77.
I find your comment extremely surprising considering your join date of 2006. I would expect such a comment from a newbie.
Unless the entirety of the developer and encoding community are lying to me, what is occurring with AV1 is absolutely par the norm for new codecs.
265 is still a dog to encode without a reasonable monster of a PC. How will an even more compressed, newer codec, that's open (and therefore can't break terrible patents) begin to compete only 6 months in?
The only thing that AV1 has going for it, is it's openness and the hope that /so many players/ throwing themselves at the problem, will slowly address the performance issues.
None the less, it's been 6 months. You're not going to see this getting hardware acceleration for at least 6 more months on any devices.
I suspect it'll be ubiquitous at best case scenario of 3 years. (more experienced members, welcome to correct me)
vidschlub
9th June 2019, 12:06
Encoders are no longer being made for this crowd. The primary design goal is massive-scale cloud encoding for YouTube, Netflix, Amazon, and everyone else that fits the encode-once, download hundreds of thousands of times scenario.
In such a scenario, even the slowest encoder is acceptable if it saves enough bytes.
In that scenario, VP9 also didn't fail. It gets used for a lot of content on the web.
It never even began to occur to me that one day, my personal local library, would never be in AV1 format.
Yet here you are outlining exactly why it is very unlikely to and it kind of blows me away, you're totally correct.
AV1 /at scale/ when a video is being watched upwards of 500 times a week, makes so much more sense. Those encode times will eventually pay for themselves.
sneaker_ger
9th June 2019, 12:29
Probability of seeing AV1 decoding in Turing refresh?
nevcairiel
9th June 2019, 13:06
Probability of seeing AV1 decoding in Turing refresh?
None.
NikosD
9th June 2019, 16:49
nVidia and AMD (possibly Intel for Gen11 iGPUs) could add in drivers a hybrid decoding approach of AV1 using the GPU itself (shaders) but not ASIC yet.
birdie
9th June 2019, 19:52
I find your comment extremely surprising considering your join date of 2006. I would expect such a comment from a newbie.
When I was writing that comment I was thinking about VP9. Aside from YouTube/Netflix you'd consider this codec a failure. The scene doesn't use it. Doom9 users don't really use it. It's become a great codec for content delivery. It's not really used anywhere else.
It's kinda strange we have projects like x264/x265 for patent encumbered H.264/H.265 codecs, yet nothing like that for VP9/AV1.
dapperdan
9th June 2019, 21:35
I believe that is the niche that rav1e is aiming for.
And if Apple adopts AV1 and it becomes ubiquitous then then I think it'll have been a good move for the focus to have shifted from VP9, even if it means VP9 becomes a bit of a lost generation.
There's been some suggestion that SVT-AV1 has already passed libvpx for the "encode a single video on a single machine" case, while still having room to improve further.
Mr_Khyron
11th June 2019, 15:30
https://www.singhkays.com/blog/av1-ecosystem-update-may-2019/
Table of Contents
SVT-AV1 is making strides!
Android Q gets AV1 support
Firefox 67 release makes AV1 decoding default on all desktop platforms
Visionular Aurora AV1 codec claims it’s faster and better than x265
BBC compares AV1 & VVC
Amphion Semiconductor Hardware decoder
Mystery Keeper
11th June 2019, 16:35
When I was writing that comment I was thinking about VP9. Aside from YouTube/Netflix you'd consider this codec a failure. The scene doesn't use it. Doom9 users don't really use it. It's become a great codec for content delivery. It's not really used anywhere else.
It's kinda strange we have projects like x264/x265 for patent encumbered H.264/H.265 codecs, yet nothing like that for VP9/AV1.
Are you kidding? I'm using VP9 and loving it!
benwaggoner
11th June 2019, 17:17
If only Youtube had like .. multiple .. videos coming in daily. Then they could encode them simultaneously on a single CPU each. :devil: (And serve AVC or fast setting VP9/AV1 until they are done.)
You could do a fast first pass for scenechange detection and vbv estimation and send the chunks along with that info.
Yeah, something like that would work. It's more overhead, of course, and reduces total throughput. But that's what I'd do if I was trying to get good quality out of what are essentially spot instances.
This is not something any streaming service does though, so why should they invest into developing stuff for a goal they don't even need?
Its just how it goes. And for UHD Blu-ray discs, they can just throw massive bitrates at it to solve any such issues.
Streaming services for premium content do target really high quality, and most of the time at the top bitrate you won't see visible artifacts.
It's the user-generated content world where you see a visible quality ceiling. The sources aren't as good, and the economics for how many MIPS/pixel and how many Mbps to spend yield more conservative choice.
Also there is a big political motivation to use of non-MPEG codecs by some of the biggest UGC platforms, even when it doesn't make strict economic sense.
AV1 /at scale/ when a video is being watched upwards of 500 times a week, makes so much more sense. Those encode times will eventually pay for themselves.
That's the hope of AV1. At this point H.264 has very mature encoders, so quality @ bitrate @ speed of AV1 and VP9 really don't offer any substantial improvements, and there are quality regressions versus x264 for some content types.
When I was writing that comment I was thinking about VP9. Aside from YouTube/Netflix you'd consider this codec a failure. The scene doesn't use it. Doom9 users don't really use it. It's become a great codec for content delivery. It's not really used anywhere else.
It's kinda strange we have projects like x264/x265 for patent encumbered H.264/H.265 codecs, yet nothing like that for VP9/AV1.
Yeah, it's a chicken-and-egg thing. Because the market assumes that MPEG codecs are going to be widely used, a lot of people start building commercial codecs while the spec is still being finalized.
Also, the MPEG reference encoders just aren't useful for production due to speed and features. The vp* and AV1 series get a sort of hybrid reference/production encoder. It's "good enough" so people haven't bothered with ground-up new encoders. And specs haven't been close to MPEG quality before AV1.
And we can't discount the unique impact of x264. Legions of video pirates competing on making the best looking files as small as possible as quickly as possible to post to torrent sites meant lots of eyeballs on a very wide range of source content; much more diverse than typical encoder test content libraries. Dozens of people deep diving on tunings instead of a handful. Lots of eyeballs on every new beta to see what's different.
x264 just got good in ways that might be impossible to ever replicate. HEVC is close enough to H.264 that things like CRF and psychovisual tuning worked well enough to refine from. And x264 set a high bar that commercial encoder vendors had to strive to beat.
VP9 never had that kind of interest. AV1 is certainly showing much more competition in commercial encoders already than any vp* codec ever did.
soresu
11th June 2019, 17:41
It's kinda strange we have projects like x264/x265 for patent encumbered H.264/H.265 codecs, yet nothing like that for VP9/AV1.
rav1e is basically the x264 of AV1 in terms of a community or free software driven project, time will tell if it gets even remotely as much traction as x264 did early on.
Its also as much of a successor to the Theora project, which also had On2 lineage stemming from VP3 being open sourced. For such a limited base codec, they managed to get a lot out of Theora (Ptalabvorm) before VP8 made it redundant for web video.
x265 on the other hand was never truly a successor to x264 in terms of community from what I've seen - it was driven by MultiCoreWare from the get go, and controlled by them rather than community (don't quote me there).
I'd also say rav1e is also kind of a test to see if a production codec can be viable if written in Rust, I think its the first?
soresu
11th June 2019, 17:47
VP9 never had that kind of interest. AV1 is certainly showing much more competition in commercial encoders already than any vp* codec ever did.
I think the open development process has alot to do with that interest and competition, it helps to get the whole thing going during the standardisation part, and not after.
The fact that it doesn't belong to one singular company helps too I think (like AC3/AC4/DTS). For all the reach of Youtube, noone wants to suffer with their bottom line because Google decided to make a change in codecs.
I think even H264 and H265 would not have prevailed so well without a similar development process - albeit one more encumbered with patent jockeying and so forth.
`Orum
12th June 2019, 02:54
I'd also say rav1e is also kind of a test to see if a production codec can be viable if written in Rust, I think its the first?
Large parts of it are in assembly (https://github.com/xiph/rav1e/tree/master/src/x86), and I can only see that trend continuing as they improve compression efficiency while trying to avoid sacrificing encoding speed.
While I'd love to see this become the "next x264," I'm not going to hold my breath. Far too many of the best encoders now are not free and especially not open source.
birdie
12th June 2019, 10:50
Weird results from BBC: https://www.bbc.co.uk/rd/blog/2019-05-av1-codec-streaming-processing-hevc-vvc: no source videos, no codecs version, nothing.
soresu
12th June 2019, 13:11
Weird results from BBC: https://www.bbc.co.uk/rd/blog/2019-05-av1-codec-streaming-processing-hevc-vvc: no source videos, no codecs version, nothing.
It's not weird, they have a stake in VVC - so unfortunately they seem to be pushing a FUD angle against the nascent, standardised AV1 in an attempt to make VVC look better.
I'd hoped for better from the BBC considering the quality of their iPlayer platform. As it is, from these barely veiled attacks I'm in doubt that they will use AV1 at all.
Has anyone tried using SVT-AV1 on AMD hardware yet, especially Threadripper 16-32 cores? I'm curious to see how far Intel specific optimisations gimp AMD performance.
`Orum
12th June 2019, 18:09
Has anyone tried using SVT-AV1 on AMD hardware yet, especially Threadripper 16-32 cores? I'm curious to see how far Intel specific optimisations gimp AMD performance.
I don't have a threadripper but I can test on a R7 1700 when I get home, if that interests you.
As for the "Intel specific" optimizations, that could help or hurt on AMD, but at the very least they seem to use Visual Studio instead of icl (which is really more "AMD specific degradation" than "Intel specific optimization").
Edit: Does anyone have a binary available? It seems to require VS 2017 or 2019, and I won't install those while I have 2015 installed. Alternatively I can use their release, but it's several weeks old and a very active project.
`Orum
13th June 2019, 04:20
Alright, here are some numbers on my R7 1700. Source was a 1080p BD I had handy. Options were: -q 30 -n 1000 -i stdin -w 1920 -h 1080 -enc-mode 4
Total Frames Frame Rate Byte Count Bitrate
1000 30.00 fps 5734937 1376.38 kbps
Channel 1
Average Speed: 1.853 fps
Total Encoding Time: 539574 ms
Total Execution Time: 541517 ms
Average Latency: 47668 ms
Max Latency: 64363 ms
Encoder finished
iwod
13th June 2019, 15:27
It's not weird, they have a stake in VVC - so unfortunately they seem to be pushing a FUD angle against the nascent, standardised AV1 in an attempt to make VVC look better.
I'd hoped for better from the BBC considering the quality of their iPlayer platform. As it is, from these barely veiled attacks I'm in doubt that they will use AV1 at all.
Has anyone tried using SVT-AV1 on AMD hardware yet, especially Threadripper 16-32 cores? I'm curious to see how far Intel specific optimisations gimp AMD performance.
And they are also part of the Open Media Alliance. So we now consider anything that is better than AV1 as FUD?
dapperdan
13th June 2019, 17:46
https://medium.com/vimeo-engineering-blog/behind-the-scenes-of-av1-at-vimeo-a2115973314b
Vimeo adopting AV1, specifically the rav1e encoder with an explicit wish to make it the new x264.
unpause
14th June 2019, 10:27
Alright, here are some numbers on my R7 1700. Source was a 1080p BD I had handy. Options were: -q 30 -n 1000 -i stdin -w 1920 -h 1080 -enc-mode 4
Using the same options I got the following results on an R7 2700X. Ubuntu 19.04, release mode built from HEAD (5fd69642f40655d2ec7ac6ffb8cb2c678650e1e7):
SUMMARY --------------------------------- Channel 1 --------------------------------
Total Frames Frame Rate Byte Count Bitrate
1000 30.00 fps 13116427 3147.94 kbps
Channel 1
Average Speed: 2.653 fps
Total Encoding Time: 376893 ms
Total Execution Time: 378036 ms
Average Latency: 33390 ms
Max Latency: 47422 ms
Encoder finished
soresu
14th June 2019, 18:06
And they are also part of the Open Media Alliance. So we now consider anything that is better than AV1 as FUD?
I actually live in the UK, have friends that have worked for the BBC, and I personally consider their motives to be suspect here - joining the AOM is by no means any guarantee that they did so with good intentions at that point, or at any point since, any more than nVidia's presence in the Khronos standards body guarantees their good intentions towards future efforts with OpenCL.
I'm not saying that AV1/libaom are without their failings, but the graphs in the blog articles seem to show worst case scenario numbers for AV1 from what I've seen from other sources in the past (including here), while showing only improvement for VVC - the only positive thing they can write is that AV1/libaom has gotten much faster since their last test.
The best I can say is that they are being somewhat disingenuous towards AOM's efforts thus far, though their potential patent stake in VVC causes me to lean towards a more nefarious angle on the matter.
I would add that I don't in any way believe that VVC or MPEG codecs are intrinsically bad, in fact as an avid follower of ML/AI tech in media I am quite interested to see how it performs once implemented.
It is the vortex of financial incentives that churn around MPEG efforts that has me reaching for my FUD colored glasses.
mandarinka
14th June 2019, 22:11
And you think the companies like Google with vested interest in VP9/AV1 have no incentive to astroturf or paint their product in better light and competing in worse than is fair?
Actually, it might be open source enthusiasts and evangelists that volunteer/vigilante/follow these things as a hobby and not as a job/living that are the worst offenders with FUD ("Fear, uncertainty and doubt") or untrue claims, because companies are actually somewhat afraid of being held accountable.
These folks probably often don't even get they do something dishonest or if they do, they think it's fine because "we are the good guys" (no, principles should hold for everybody.). Even if we put apart more controversial fields for sake of not starting offotpic flame... good example is for example the twitter/phoronix forums marketing of the Raptor Engineering (Power9 vendor) that routinely uses strongly dishonest FUD against x86 to sell their stuff to paranoid people as a company that does it, libreboot as non-profit people that do it and then the general fandom of this free hardware(firmware) movements as general internet people that then perpetuate it further. You could probably find a lot of that blinded hypocrisy here too.
MoSal
14th June 2019, 23:10
No need for multi-paragraph comments arguing who is FUDing who (old-style FUDing is a tired tactic anyway).
BBC R&D (not necessarily representative of everyone in the organization) provided zero info that would allow anyone to replicate their results, let alone analyzing and arguing their usefulness. Period.
soresu
15th June 2019, 18:48
And you think the companies like Google with vested interest in VP9/AV1 have no incentive to astroturf or paint their product in better light and competing in worse than is fair?
I think that's a false equivalency, AV1 and VP9 are a means to an end for Google, Netflix, Cisco etc - the end being pushing more video for a given bandwidth without being encumbered by uncertainty fostered by divergent patent pools, as happened with HEVC (and I absolutely believe it will happen again with VVC given time).
The end may not be 'just' patent royalties for all MPEG members, but it will certainly be a top consideration for most of them.
The difference is actually in the product wording you mentioned:
For Google and the other AOM content creators, the product is the content and increased access created by the codecs existence.
For MPEG members, the product is the codec itself.
Obviously that oversimplifies the matter somewhat, but I think that represents the main gist of it.
bstrobl
17th June 2019, 14:37
First SoC launched by Realtek: https://www.realtek.com/en/press-room/news-releases/item/realtek-launches-worldwide-first-4k-uhd-set-top-box-soc-rtd1311-integrating-av1-video-decoder-and-multiple-cas-functions
EwoutH
17th June 2019, 14:40
Press release: https://www.realtek.com/en/press-room/news-releases/item/realtek-launches-worldwide-first-4k-uhd-set-top-box-soc-rtd1311-integrating-av1-video-decoder-and-multiple-cas-functions (Realtek Launches Worldwide First 4K UHD Set-top Box SoC (RTD1311), Integrating AV1 Video Decoder)
soresu
17th June 2019, 20:07
Faster than I expected for a hardware ASIC release, with HDMI 2.1 support no less - though HDMI 2.1 is obviously overkill for 4K60p video, the release doesn't say anything about higher than 4K support.
hajj_3
17th June 2019, 21:04
Press release: https://www.realtek.com/en/press-room/news-releases/item/realtek-launches-worldwide-first-4k-uhd-set-top-box-soc-rtd1311-integrating-av1-video-decoder-and-multiple-cas-functions (Realtek Launches Worldwide First 4K UHD Set-top Box SoC (RTD1311), Integrating AV1 Video Decoder)
nice to see but realtek charges a lot more for their chips than amlogic, allwinner, huawei, mediatek and rockchip which is why cheap android tv boxes don't use them.
IgorC
17th June 2019, 22:25
https://www.reddit.com/r/AV1/comments/bt7l7a/svtav1_now_offers_higher_quality_than_x264_x265/
Everybody compares AV1 to VP9 on 8 bits for both.
What about VP9 10 bits?
It makes also sense to compare AV1 8 bits vs VP9 10 bits.
VP9 10 bits has several advantages over AV1 8 bits at this moment:
Significantly faster encoding
Significantly faster decoding
Hardware support
Should has ~10-20% better compression than VP9 8 bits (not that far from AV1’s 25-30%)
It makes sense to employ VP9 10 bits at least for 2-3 years more until a final jump to AV1. It will give additional time for AV1 to develop better encoders and decoders.
Well, Google and Netflix already use VP9 10 bits for their HDR content but SDR could benefit as well.
Plus VP9 benefits a LOT from 10 bits because it suffers from blocking and banding not less than H.264 as it indicates here https://sonnati.wordpress.com/2016/06/17/does-vp9-deserve-attention-part-ii/
soresu
19th June 2019, 09:24
Always interesting to hear about AI/ML tidbits related to video encoding, the Visionular speaker at the Big Apple Video conference (26th June) has this in her summary:
"Finally, we will introduce certain AI+codec techniques that could provide certain novel coding tools leveraging the use of deep learning for the next AOM standard, possibly AV2."
dapperdan
19th June 2019, 18:33
I think this paper covers one such proposal for AV2 and has the speaker as an author:
https://link.springer.com/chapter/10.1007/978-3-319-94361-9_18
Couldn't quickly find a public version, but the citations took me to this which I think was another technique discussed:
https://arxiv.org/pdf/1804.09291
foxyshadis
20th June 2019, 08:56
Reminds me of NNEDI, both in that there can be surprising visual gains and in that it will require enormous CPU & GPU power just to decode.
But isn't discussion of AV2 getting way off topic? We have an entire forum for discussing new and potential codecs, this thread is just getting more polluted and useless every week.
benwaggoner
20th June 2019, 19:40
https://www.reddit.com/r/AV1/comments/bt7l7a/svtav1_now_offers_higher_quality_than_x264_x265/
Everybody compares AV1 to VP9 on 8 bits for both.
What about VP9 10 bits?
It makes also sense to compare AV1 8 bits vs VP9 10 bits.
Why not compare AV1 10-bit vs. VP9 10-bit?
The barriers to using 10-bit seem pretty similar in both cases, namely longer encoding time, lack of 10-bit sources or processing chains (getting much better), reliance on good dithering in display system, somewhat slower SW decode, and rarer HW decode support.
AFAIK, no one is planning any 8-bit only AV1 decoders, so 10-bit might be able to be used by default more often with AV1 if HW decoders become dominant. SW decoders need more optimization for 10-bit to make it competitive for higher resolutions.
I expect 10-bit to become generally mainstream as it is required for HDR, and we're near or past the tipping point where the majority of new video consumption devices support at least HDR-10.
IgorC
20th June 2019, 23:43
Of course You can compare AV1 and VP9 both 10bits.
My main point was not very disruptive move from VP9 8 bits to VP9 10 bits as a short term strategy (2-3 years).
Let’s put some numbers. My i7 notebook uses 20-25% of CPU during Youtube 1080p@60fps (VP9 8 bits). If it was VP9 10 bits that would be 5-7% additional CPU usage. Still pretty acceptable.
Now AV1 8 bits consumes whooping 60% at that resolution and framerate (and that with the last version of dav1d). While my notebook still can play it but a fan noise and overall slowness are quite annoying.
Let alone AV1 10 bits. Dav1d hasn’t any 10 bits code yet and it will take some time to get fast 10 bits decoding and/or hardware acceleration. My notebook gets very hot and drops a few frames here and there with near 100% CPU load with AV1 10 bits on 1080p@60. Also my another notebook with Kaby Lake i7 already has VP9 8-/ 10- bits hardware acceleration. So why not?
VP9 8 bits suffers from strong blocking and banding in dark areas and tones in my experience with Youtube and mobile Netflix videos. While VP9 10 bits can handle it very well with a little extra CPU overhead.
benwaggoner
20th June 2019, 23:55
Of course You can compare AV1 and VP9 both 10bits.
My main point was not very disruptive move from VP9 8 bits to VP9 10 bits as a short term strategy (2-3 years).
Let’s put some numbers. My i7 notebook uses 20-25% of CPU during Youtube 1080p@60fps (VP9 8 bits). If it was VP9 10 bits that would be 5-7% additional CPU usage. Still pretty acceptable.
Now AV1 8 bits consumes whooping 60% at that resolution and framerate (and that with the last version of dav1d). While my notebook still can play it but a fan noise and overall slowness are quite annoying.
Let alone AV1 10 bits. Dav1d hasn’t any 10 bits code yet and it will take some time to get fast 10 bits decoding and/or hardware acceleration. My notebook gets very hot and drops a few frames here and there with near 100% CPU load with AV1 10 bits on 1080p@60. Also my another notebook with Kaby Lake i7 already has VP9 8-/ 10- bits hardware acceleration. So why not?
VP9 8 bits suffers from strong blocking and banding in dark areas and tones in my experience with Youtube and mobile Netflix videos. While VP9 10 bits can handle it very well with a little extra CPU overhead.
Well the obvious NOW strategy is to keep using H.264 or HEVC, which have broad HW decoder support, and not use CPU at all. x264 properly tuned is at least as good as VP9 for lots of real world content.
I don't see any software decoder solution becoming mainstream, since we have more than good enough HW codec options now.
Fingers crossed that all AV1 HW decoders include 10-bit support. It would be great to have a codec out there where >8-bit support is guaranteed.
IgorC
21st June 2019, 01:19
Well the obvious NOW strategy is to keep using H.264 or HEVC
Yeah, nice try. :p
But fortunately Google and Netflix don't think so.
Both make a major accent on VP9 and AV1.
Cheers.
Quikee
21st June 2019, 04:46
Fingers crossed that all AV1 HW decoders include 10-bit support. It would be great to have a codec out there where >8-bit support is guaranteed.
There are 3 AV1 profiles (Main, High and Professional) and all define a mandatory 10-bit support. From that I think it's a good chance there will be HW 10-bit support from the beginning. Of course HW manufacturers still can disappoint.
VP9 has 10-bit support only from profile 2 on, where profile 0 and 1 are 8-bit only, which is why 10-bit HW support is less common.
soresu
22nd June 2019, 07:51
AFAIK, no one is planning any 8-bit only AV1 decoders, so 10-bit might be able to be used by default more often with AV1 if HW decoders become dominant. SW decoders need more optimization for 10-bit to make it competitive for higher resolutions.
True enough, but the libaom decoder isn't nearly as well optimised as dav1d at present, and even the dav1d devs seem to be completely ignoring 10 bit optimisation while they concentrate on getting 8 bit working well across at least x86 SSSE3/SSE4/AVX2, and ARM NEON SIMD targets.
I'd say that the progress so far is pretty incredible for such a young codec, decoder and encoder wise.
From what I remember it also took a while for the OpenHEVC decoder library to get optimised, and they weren't concentrating on mobile nearly as much as the AV1 groups seem to be - though the increased core counts and IPC of current ARM implementations may have influenced that focus to some degree.
nevcairiel
22nd June 2019, 08:22
the dav1d devs seem to be completely ignoring 10 bit optimisation while they concentrate on getting 8 bit working well across at least x86 SSSE3/SSE4/AVX2, and ARM NEON SIMD targets.
10-bit isn't being "ignored", 8-bit was quite simply a much higher priority since content with that is actually available to the public since YouTube started shipping it, while 10-bit is not. And there is only so many hours in a day.
Work on 10-bit has started now, but due to the nature of the beast, SIMD stuff cannot be easily ported from 8-bit to 10-bit, so its a lot of work still.
dapperdan
27th June 2019, 09:09
Lots of interesting talks at the Big Apple Video event (that's Apple as in New York, not Mac Os X):
Probably the most interesting for this group is the second half of Ronald Bultje's talk which covers his Eve-AV1 encoder and some conparisons with other codec and encoders:
https://vimeo.com/344663992
But lots of other interesting stuff from other speakers too if you click through to the channel to see the full list.
I'll also mention this one from Cisco as it's got a kind of boring sounding title but had some interesting stuff around complexity Vs speed in AV1 after the live demos.
https://vimeo.com/344366650
mandarinka
27th June 2019, 17:03
I wonder how much does VMAF really speak about visual quality and compression efficiency while keeping detail (as opposed to the usual issue with metrics, the "blur more for maximum PSNR/SSIM" effect), seeing how in those slides, *everything* except Rav1e and x264 is shown as matching or outdoing x265. Well, I guess there's already the usual assertion/claim that x265 = lipvpx-vp9 that raises questions. :rolleyes: I always stop wondering at that point in these presentations...
dapperdan
27th June 2019, 19:12
My theory is that the people hired for the subjective tests that underly the objective stats or that vote VP9 as very slightly better than x265 in the MSU tests on subjectify.us have a different notion of quality than the kind of person who is interested in codecs for their own sake.
Like, I read a paper recently where someone was applying their grain synthesis approach to HEVC and the subjective tests they did to prove it worked showed they could get basically all the subjective benefit by just doing the noise removal step and not bothering to add the grain back in, something that could be done by any encoder, for any codec (and I'm guessing this makes up part of the secret sauce of some encoders).
But I guess someone who said they could get a massive increase in subjective quality via the Psy optimisation of basically blurring the input would get some pushback on that view in some quarters, even with subjective tests to back it up.
Link to the paper. It seems at higher qualities the people saw the added grain as a defect rather than a quality improvement (though still a statistical tie mostly).
https://arxiv.org/abs/1904.11754
soresu
28th June 2019, 16:29
Having watched some (but not all) the BAV presentations - I know that AV1 is currently not ideal for running a battery of tests at short notice, but did they really need to use such outdated versions of competing codecs?
Im pretty sure that the x265 build was from January, and the libaom build from february in one of them.
Maybe I'm missing something and those builds were picked for stability?
Blue_MiSfit
28th June 2019, 22:31
Great talk from Ronald. If only I had the time to do an evaluation of Eve_AV1
soresu
29th June 2019, 15:21
I noticed that several talks mentioned rav1e, but none directly covered it
Was I missing a video, or did the Mozilla/Xiph rav1e guys not get a talk at BAV?
birdie
30th June 2019, 11:23
Twitch's AV1 deployment roadmap (from Big Apple Video 2019)
https://i.redd.it/blmo96lxl2731.png
Beelzebubu
1st July 2019, 17:33
I noticed that several talks mentioned rav1e, but none directly covered it
Was I missing a video, or did the Mozilla/Xiph rav1e guys not get a talk at BAV?
I believe that because the conference was organized by Vimeo/Mozilla, who just announced (https://press.vimeo.com/61553-vimeo-introduces-support-for-royalty-free-video-codec-av1) a partnership around rav1e, they wanted to prevent a potential conflict of interest in talk selection and decided to not give a talk on it.
Beelzebubu
1st July 2019, 17:41
Having watched some (but not all) the BAV presentations - I know that AV1 is currently not ideal for running a battery of tests at short notice, but did they really need to use such outdated versions of competing codecs?
Im pretty sure that the x265 build was from January, and the libaom build from february in one of them.
Maybe I'm missing something and those builds were picked for stability?
The builds used in the Eve-AV1 talk (https://vimeo.com/344663992) were:
rav1e c68d68c6fa80dabf5e4ed9b379f090572eb43d96 (Mon Jun 3 2019)
libaom a385cc44e15833f56de45bbbc1cc6c474751ac9f (Wed Apr 24 2019)
x264 5493be84cdccecee613236c31b1e3227681ce428 (Thu Mar 14 2019)
x265 12522:10decf67c077 (Fri Jun 07 2019)
SVT-AV1 6fd564611bdb48a2a6d2c7b90a91b4b1bdbe74b9 (Mon Jun 10 2019)
libvpx f836d8ba87dcba437228580fe65afe151ccf7659 (Thu Apr 25 2019)
So basically - ignoring x264 for a second (which is pretty mature/stable) - some from late April and some from early June, none from January or February.
soresu
1st July 2019, 19:30
Ah, must have misread the presentation then, easier for me to read from slides than video - still the position of rav1e seems odd, it shows on graphs to be still hovering around x264 - I could have sworn it passed x264 months ago, and then nothing has been said since despite all the work commits being merged into it.
Is it still suffering that regression from awhile ago?
"I believe that because the conference was organized by Vimeo/Mozilla, who just announced a partnership around rav1e, they wanted to prevent a potential conflict of interest in talk selection and decided to not give a talk on it."
Yes that makes sense, just seemed a little odd, like going to WWDC and getting nothing from Apple - still they got a lot of mentions from the presenters in any case.
TD-Linux
1st July 2019, 21:55
The rav1e stats are correct. We still fall behind on high bitrate VMAF - most likely due to that being more sensitive to activity masking (aq) which is still in progress by s_p.
The multithreading performance is limited by a serialization point of the loop filters between frames (by far the slowest part of rav1e right now). There's some outstanding PRs to make it better, e.g. https://github.com/xiph/rav1e/pull/1396
dapperdan
1st July 2019, 22:02
The Visionular talk from Zoe Liu used builds from Jan and Feb.
I think the x265 release used was the last stable release so that doesn't seem too crazy. You could easily cry foul if someone used a non stable git commit and it performed worse than expected due to hitting a bug.
It's also worth bearing in mind that that was basically the same talk as given at the Agora.io thing, so some people tour these things around for a while, it's not ridiculous for them to reuse slides and not have something fresh for every talk they give. Fairly certain I'd seen the NGCodec talk slides before too.
benwaggoner
2nd July 2019, 17:53
The builds used:
rav1e c68d68c6fa80dabf5e4ed9b379f090572eb43d96 (Mon Jun 3 2019)
libaom a385cc44e15833f56de45bbbc1cc6c474751ac9f (Wed Apr 24 2019)
x264 5493be84cdccecee613236c31b1e3227681ce428 (Thu Mar 14 2019)
x265 12522:10decf67c077 (Fri Jun 07 2019)
SVT-AV1 6fd564611bdb48a2a6d2c7b90a91b4b1bdbe74b9 (Mon Jun 10 2019)
libvpx f836d8ba87dcba437228580fe65afe151ccf7659 (Thu Apr 25 2019)
Are the actual command line parameters used documented somewhere?
benwaggoner
2nd July 2019, 18:11
I wonder how much does VMAF really speak about visual quality and compression efficiency while keeping detail (as opposed to the usual issue with metrics, the "blur more for maximum PSNR/SSIM" effect), seeing how in those slides, *everything* except Rav1e and x264 is shown as matching or outdoing x265. Well, I guess there's already the usual assertion/claim that x265 = lipvpx-vp9 that raises questions. :rolleyes: I always stop wondering at that point in these presentations...
VMAF is the least-bad objective metric we've ever had, but it's still far from perfect.
Also, VMAF isn't static. Netflix comes out with new ML models that will give different (and more accurate) scores compared to older models. It can be estimated for mobile, 1080p, or UHD devices. The scores vary based on the resolution of comparison (720p tested at 720p will deliver higher scores than 720p tested at 1080p, compared to 1080p encoding). And it is a per-frame metric, and how to aggregate per-frame scores into an overall clip quality is an unanswered question. Using a harmonic mean helps, but even that is probably only useful for <20 second durations. A single VMAF score for a whole movie or episode could indicate highly variable quality or highly consistent quality.
None of this is a diss on VMAF. Netflix did what they set out to do well, put a huge amount of effort into it, and made reasonable design decisions. But like all metrics, it measures what it is designed to measure, not what we wish it measured :).
I've seen VMAF do a poor job of detecting:
Banding in gradients
Detail in low luma
Adaptive quantization improvements
Artifacts in encoders/formats that weren't included in the VMAF training set
Differences between two pretty high quality encodes.
And it doesn't do HDR at all. It used to not do UHD, but does now.
Another problem with a popular metric is that developers start tuning for that metric instead of what the metric is supposed to measure (subjective quality in this case). When developers start tuning for metrics over eyes, the correlation of that metric to subjective quality actually gets WORSE. So, for an encoder like libaom that got tuning based on VMAF ratings, we'd expect that its VMAF scores would be higher relative to actual subjective quality than for encoders that weren't tuned that were. But it'll be better than ones tuned for PSNR, like the vp? series.
Not that tuning for VMAF is a bad strategy. But it does result in less meaningful VMAF scores.
benwaggoner
2nd July 2019, 18:17
My theory is that the people hired for the subjective tests that underly the objective stats or that vote VP9 as very slightly better than x265 in the MSU tests on subjectify.us have a different notion of quality than the kind of person who is interested in codecs for their own sake.
These are double-blind tests; the people doing it just compare two encodes.
That said, I've not seen any study demonstrating better subjective quality from a well-tuned libpvx encode versus a well-tuned x265 encode, using the same bitrate @ time.
Like, I read a paper recently where someone was applying their grain synthesis approach to HEVC and the subjective tests they did to prove it worked showed they could get basically all the subjective benefit by just doing the noise removal step and not bothering to add the grain back in, something that could be done by any encoder, for any codec (and I'm guessing this makes up part of the secret sauce of some encoders).
This opens up the interesting question of no-reference quality versus creative intent. Someone just looking at a clip without grain might think it looks great. But if the creators meant there to be grain, than the output isn't accurate. That's something that some studies might not rate. And if customers dislike grain, they might rate the encoded version higher than the source!
But I guess someone who said they could get a massive increase in subjective quality via the Psy optimisation of basically blurring the input would get some pushback on that view in some quarters, even with subjective tests to back it up.
Well, that is what adaptive quantization is all about, really. Put the artifacts where they are less painful, and used the saved bits where they'll provide the most visible improvements.
There are similar debates about TV's default "vivid" mode. Some people claim to like it, even though what's displayed in manifestly wrong on many axes.
benwaggoner
2nd July 2019, 18:17
Great talk from Ronald. If only I had the time to do an evaluation of Eve_AV1
Is Eve available for evaluation in any way? I've never been able to get my hands on a build, or clips encoded to my specifications.
benwaggoner
2nd July 2019, 18:19
I think the x265 release used was the last stable release so that doesn't seem too crazy. You could easily cry foul if someone used a non stable git commit and it performed worse than expected due to hitting a bug.
And there weren't any substantial quality improvements inx x265 between the Jan 2019 builds before the June 3.1 release.
How the encoders got tuned is what matters. And if quality is being compared at fixed encoding time, performance improvements become quality improvements.
dapperdan
2nd July 2019, 19:31
These are double-blind tests; the people doing it just compare two encodes.
My point still stands for tests that intend to be double-blind since some people's abilities and/or preferences would effectively unblind the test.
Imagine, for example, people who believe that tube amps or vinyl is better than digital audio. In a double-blind test they would probably still vote for the tube amp or the vinyl because it has distinctive audio characteristics that can't be removed without invalidating the test. They can hear things that they prefer and associate (conciously or not) with quality.
On the other hand, they would potentially be fooled by audio that had been processed to sound like tube amps or vinyl or passed through a digital chain before output.
I considered this possibility because two recent tests that were presented as being negative for AV1 specifically mentioned that some of their test participants were video engineers. They mentioned this as evidence that it was all done properly, but it seemed like an obvious test methodology failure to me.
I think it was Monty from Xiph that said his party trick used to be identifying the encoder used just by listening to mp3s, and I bet certain bitrates and content would let people here do the same with video codecs and there's a possibility their opinion scores would differ from Joe Public as a result.
dapperdan
2nd July 2019, 19:54
That said, I've not seen any study demonstrating better subjective quality from a well-tuned libpvx encode versus a well-tuned x265 encode, using the same bitrate @ time.
Have you read the full version of the last MSU subjective comparison? I've only read the free snippet, which doesn't have enough info to say either way, but it's possible that fits the criteria or is at least in the right ballpark, potentially a statistical tie:
http://www.compression.ru/video/codec_comparison/hevc_2018/#subjective_report
On the other hand, similar to how complaints about electric cars are now "I don't like the minimalism of their touchscreen interfaces" when not too long ago you'd hear how they were physical impossibilities, I think the fact that we're now at this level of complaint for the previous generation of royalty-free codecs is a testament to how far we've come.
soresu
2nd July 2019, 21:00
Is Eve available for evaluation in any way? I've never been able to get my hands on a build, or clips encoded to my specifications.
Seems like a wonky business model if one of Amazon's principal video engineers can't get their hands on a build of it to at least do some testing.
Blue_MiSfit
2nd July 2019, 21:39
Re: Eve evaluation, I've never tried, TBH. I've been wanting to spend more time looking at Beamr 5x (fantastic so far!) but have been quite busy.
benwaggoner
3rd July 2019, 01:30
Imagine, for example, people who believe that tube amps or vinyl is better than digital audio. In a double-blind test they would probably still vote for the tube amp or the vinyl because it has distinctive audio characteristics that can't be removed without invalidating the test. They can hear things that they prefer and associate (conciously or not) with quality.
Yep, and then we're back into "Zen and the Art of Motorcycle Maintenance" style philosophical ruminations on the nature and meaning of "quality." Which is unavoidable at a certain point, which is why we try to test for something more specific than just quality. Accuracy to a source and creative intent can be quite different from a no-reference "is this clip pleasing" or "do you see anything wrong with this clip?"
On the other hand, they would potentially be fooled by audio that had been processed to sound like tube amps or vinyl or passed through a digital chain before output.
Exactly. And it's not a particularly hard thing to synthesize. It's not like the film grain in Marvel movies is ACTUALLY grain-from-film. It's digitally synthesized. Grain helps make blending in VFX a lot easier, and allows for rendering at 2K instead of 4K.
I considered this possibility because two recent tests that were presented as being negative for AV1 specifically mentioned that some of their test participants were video engineers. They mentioned this as evidence that it was all done properly, but it seemed like an obvious test methodology failure to me.
It depends on what the question they were asking was intended to be. But yeah, having a bunch of video engineers look at something is very different than having the general public look at something, and can provide different (but both useful!) answers. Video engineers are going to pick up on more subtle things, and are going to care about accuracy and creative intent more.
Generally I'll have video experts to an initial pass on something to see "is there something that can be seen here?" and then using double-blind testing with a more general population to confirm details. The second is a LOT slower and more expensive than the first, of course.
I think it was Monty from Xiph that said his party trick used to be identifying the encoder used just by listening to mp3s, and I bet certain bitrates and content would let people here do the same with video codecs and there's a possibility their opinion scores would differ from Joe Public as a result.
Oh, no doubt. I've done that party trick **many** times. x264 versus WMV3 versus VC-1 versus Main Concept versus VP9; it's generally pretty obvious if you've been in the field for a while.
Beelzebubu
3rd July 2019, 13:45
Is Eve available for evaluation in any way? I've never been able to get my hands on a build, or clips encoded to my specifications.
Have you asked?
benwaggoner
3rd July 2019, 22:55
Seems like a wonky business model if one of Amazon's principal video engineers can't get their hands on a build of it to at least do some testing.
To be clear, I speak only for myself, not Amazon, on these forums. I actually tinker with video stuff on my off hours too. I should probably get out more.
Anyway, I requested an optimal Eve encoding for My encoding challenge (https://forum.doom9.org/showthread.php?t=175776&highlight=benwaggoner), but they declined to participate.
It is common for encoder vendors who think they are doing some magic things in the bitstream to want to have the bitstream output under NDA and such. I get the impulse, but it just isn't practical for doing actual comparisons or due diligence evaluation.
dapperdan
4th July 2019, 08:12
When you say they "declined to participate" did they respond and say they didn't want to take part or did you just not hear from them after making a broad request in a forum post?
I believe the comment above yours saying ("Have you asked?") Is written by a developer of EVE, which suggests they didn't know they'd been asked, so possibly an email has got lost in a spam trap.
Ilya87
4th July 2019, 18:07
Hi guys, I've desided to make a comparison of x264, rav1e and x265 encoders with 500, 600, 700, 800, 900, 1000 kbit/s with the following settings:
rav1e -b $g --tiles 6 -s 5 --matrix BT470BG /D/sintel/sintel720.y4m --output sintel720_rav1e_s5_$g.ivf
rav1e -b $g --tiles 6 -s 3 --matrix BT470BG /D/sintel/sintel720.y4m --output sintel720_rav1e_s5_$g.ivf
x264 -t 2 -m 11 --me umh --weightp 2 --direct spatial --aq-mode 2 --b-adapt 2 -B $g -b 4 -r 6 -I 240 --b-pyramid normal --no-dct-decimate --no-fast-pskip -A all -o sintel720_x264_$g.264
x265 /D/sintel/sintel720.y4m --y4m -o "Sintel720_x265_$g.h265" --rd 3 -b 4 --b-adapt 2 --b-pyramid --ref 6 -I 240 --bitrate $g --aq-mode 2 --weightp --weightb -m 2 --no-early-skip --psy-rd 1 --me star
where $g stands for bitrate value
My OS is Arch Linux x86_64 and CPU Core i5 8600K, rav1e was build recently and for testing 1191 frames of sintel 1k 16bit (from 12987 to 14177) were taken and converted to 720x306 yuv420p. x265 and x264 are from the distro's repository.
To measure MS-SSIM and PSNR-HVS-M daala's tools were used. To measure VMAF score I used ffmpeg's VMAF filter.
Results:
https://i110.fastpic.ru/big/2019/0704/69/c8b3817496c197135466a8ee32998c69.png
https://i110.fastpic.ru/big/2019/0704/a6/93b47b0ce22c8127bab17738dfee1ea6.png
https://i110.fastpic.ru/big/2019/0704/e6/151968254c3ab19c8abded7ee4a49ae6.png
x265 is a clear winner with 50.59-61.01 fps (lowest to highest bitrate settings)
x264 80.08-102.73 fps
rav1e s3 1.603-2.187 fps
rav1e s5 3.736-4.469 fps
Average CPU utilization of rav1e was 66%-70% (and I couldn't increase it).
marcomsousa
4th July 2019, 22:44
Intel SVT-AV1 0.6 Released With AV1 Decoding, SIMD Optimizations
https://www.phoronix.com/scan.php?page=news_item&px=Intel-SVT-AV1-0.6-Released
Ilya87
4th July 2019, 23:35
Intel SVT-AV1 0.6 Released With AV1 Decoding, SIMD Optimizations
https://www.phoronix.com/scan.php?page=news_item&px=Intel-SVT-AV1-0.6-Released
Still not supported dimensions multiple 2, only 8. Still segfaults. And many other bugs.
benwaggoner
7th July 2019, 20:23
Still not supported dimensions multiple 2, only 8. Still segfaults. And many other bugs.
It is a 0.6 release. I'd expect a smaller set of limitations like that and other issues will still be in 0.7.
Nintendo Maniac 64
8th July 2019, 05:39
While now a version old, Phoronix tested (on Linux) the encoding performance of SVT-AV1 v0.5 on the new 3rd gen AMD Ryzen 8core (3700X) and 12core (3900X) chips compared to existing Intel CPUs (primarily the 8core 9900K and 16core 7960X):
https://www.phoronix.com/scan.php?page=article&item=ryzen-3700x-3900x-linux&num=4
soresu
9th July 2019, 05:46
Well there goes my bank account down the tubes after seeing those Ryzen 3000 results - roll on september so I can become poor and happy with my 3950X.
benwaggoner
10th July 2019, 17:38
While now a version old, Phoronix tested (on Linux) the encoding performance of SVT-AV1 v0.5 on the new 3rd gen AMD Ryzen 8core (3700X) and 12core (3900X) chips compared to existing Intel CPUs (primarily the 8core 9900K and 16core 7960X):
https://www.phoronix.com/scan.php?page=article&item=ryzen-3700x-3900x-linux&num=4
Wow, interesting result! And no way did Intel worry about AMD optimizations when compiling SVT :sly: I wonder what the comparison between cpu-tuned x265 and libaom would be like, which should tilt more in AMD's favor.
I note that the top Intel processor used has only half the cores as the top AMD, so this difference could easily be due to multithreading more than per-core performance improvements. But that in no way invalidates the price/performance delta.
Also, and AV1 encoder that's running only ~2.5x slower than a HEVC encoder! Of course, I have no idea if the output quality is similar. As always, the key metric is quality @ bitrate @ performance.
benwaggoner
10th July 2019, 17:45
Also, I note that the Intel processor used in comparison is from 2017. The current equivalent would probably be the i9-9980XE, which as two more cores and 7% faster clock. That would probably have similar SVT performance to the Threadripper. At more than 2x the price, though (although for an encoding workstation/instance, the CPU is typically less than half the cost).
SmilingWolf
10th July 2019, 20:18
Status report!
"Yes I keep tweaking the params" edition
1st edition: https://forum.doom9.org/showthread.php?p=1852449#post1852449
2nd edition: https://forum.doom9.org/showthread.php?p=1857587#post1857587
3rd edition: https://forum.doom9.org/showthread.php?p=1860475#post1860475
4th edition: https://forum.doom9.org/showthread.php?p=1871939#post1871939
Whatever paragraph I don't repeat here can be assumed to be the same as in the aforementioned post
First of all: graphs!
Click to enlarge
Y axis: chosen metric
X axis: bits per pixel
720p:
https://i.ibb.co/Rh3db1D/hvmaf-720.png (https://ibb.co/Rh3db1D) https://i.ibb.co/fGffmdM/msssim-720.png (https://ibb.co/fGffmdM) https://i.ibb.co/86Y4ssM/psnrhvsm-720.png (https://ibb.co/86Y4ssM)
1080p:
https://i.ibb.co/sCzgwpd/hvmaf-1080.png (https://ibb.co/sCzgwpd) https://i.ibb.co/BNj3DCR/msssim-1080.png (https://ibb.co/BNj3DCR) https://i.ibb.co/cJcN2by/psnrhvsm-1080.png (https://ibb.co/cJcN2by)
BD rates for 720p:
Codecs ladder: | x264 relative:
x264 -> svtav1 | x264 -> svtav1
RATE (%) DSNR (dB) | RATE (%) DSNR (dB)
MSSSIM -10.5381 0.426713 | MSSSIM -10.5381 0.426713
PSNRHVS -11.296 0.557542 | PSNRHVS -11.296 0.557542
HVMAF -19.6867 0.689824 | HVMAF -19.6867 0.689824
----------------------------|-----------------------------
svtav1 -> vp9 | x264 -> vp9
RATE (%) DSNR (dB) | RATE (%) DSNR (dB)
MSSSIM -12.4136 0.464516 | MSSSIM -24.2802 1.23124
PSNRHVS -13.288 0.615572 | PSNRHVS -25.1991 1.68477
HVMAF -14.5152 0.598246 | HVMAF -26.3686 2.81799
----------------------------|-----------------------------
vp9 -> x265 | x264 -> x265
RATE (%) DSNR (dB) | RATE (%) DSNR (dB)
MSSSIM -1.73618 0.0667664 | MSSSIM -26.2541 1.24552
PSNRHVS -6.07444 0.298073 | PSNRHVS -30.4815 1.87719
HVMAF -9.04578 0.359953 | HVMAF -31.4265 3.28152
----------------------------|-----------------------------
x265 -> av1 | x264 -> av1
RATE (%) DSNR (dB) | RATE (%) DSNR (dB)
MSSSIM -20.8531 0.881529 | MSSSIM -39.9238 2.1343
PSNRHVS -16.9627 0.860883 | PSNRHVS -40.3335 2.76154
HVMAF -23.5865 1.00102 | HVMAF -48.1341 3.64521
BD rates for 1080p:
Codecs ladder: | x264 relative:
x264 -> svtav1 | x264 -> svtav1
RATE (%) DSNR (dB) | RATE (%) DSNR (dB)
MSSSIM -14.3136 0.452642 | MSSSIM -14.3136 0.452642
PSNRHVS -10.1078 0.374405 | PSNRHVS -10.1078 0.374405
HVMAF -20.4048 0.58988 | HVMAF -20.4048 0.58988
----------------------------|-----------------------------
svtav1 -> vp9 | x264 -> vp9
RATE (%) DSNR (dB) | RATE (%) DSNR (dB)
MSSSIM -19.1279 0.563386 | MSSSIM -34.6951 1.70828
PSNRHVS -21.5428 0.778635 | PSNRHVS -33.6391 2.16168
HVMAF -21.4399 0.750138 | HVMAF -34.3162 3.93015
----------------------------|-----------------------------
vp9 -> x265 | x264 -> x265
RATE (%) DSNR (dB) | RATE (%) DSNR (dB)
MSSSIM 8.56339 -0.282927 | MSSSIM -30.5146 1.24699
PSNRHVS 3.02814 -0.139956 | PSNRHVS -32.9536 1.71646
HVMAF -3.70741 0.0299945 | HVMAF -35.6727 3.2304
----------------------------|-----------------------------
x265 -> av1 | x264 -> av1
RATE (%) DSNR (dB) | RATE (%) DSNR (dB)
MSSSIM -28.044 1.00637 | MSSSIM -47.6676 2.30149
PSNRHVS -23.4583 0.991831 | PSNRHVS -45.8303 2.79923
HVMAF -26.6387 0.978822 | HVMAF -51.9814 3.88658
Encoders:
x264 157-2970-5493be8
x265 3.1-4-4f6dde51a5db
libvpx-vp9 1.8.0-591-g19bda215d
SVT-AV1 0.6.0-1424-8977f443
libaom 1.0.0-2036-ge2c1d5ef8
Cmdlines:
x264 --preset veryslow --tune ssim --crf 16 -o test.x264.crf16.264 orig.i420.y4m
x265 --preset veryslow --tune ssim --crf 16 -o test.x265.crf16.hevc orig.i420.y4m
vpxenc --codec=vp9 --frame-parallel=0 --tile-columns=0 --auto-alt-ref=6 --good --cpu-used=0 --tune=psnr --passes=2 --threads=1 --end-usage=q --cq-level=20 --test-decode=fatal --ivf -o test.vp9.cq20.ivf orig.i420.y4m
SvtAv1EncApp.exe -i orig.i420.yuv -b test.svtav1.cq20.ivf -w 1280 -h 720 -q 20 -enc-mode 3 -fps-num 24000 -fps-denom 1001 -intra-period 23
aomenc --frame-parallel=0 --tile-columns=0 --auto-alt-ref=1 --cpu-used=4 --tune=psnr --passes=2 --threads=2 --row-mt=1 --end-usage=q --cq-level=20 --test-decode=fatal -o test.av1.cq20.webm orig.i420.y4m
VMAF: model used: vmaf_b_v0.6.3, pooling: harmonic_mean, bagging score (arithmetic mean of 21 models' scores)
Notes:
TearsOfSteel720 and TheFifthElement, two clips in the 720p category, had a vertical resolution incompatible with SvtAv1EncApp (not divisible by 8).
They have been padded to 1280x536, so they have been included in this round of measurements again.
Meanwhile, rav1e still has got a nasty bug that makes it bloat encodes, which brings up to 25% BD rate regression, so it has been excluded from this edition.
Again, no time infos because I use the PC while it encodes etc. etc.
If somebody REALLY wants some encoding time infos I can run a battery of encodes under ideal conditions on my favourite 1080p clip (PresageFlowerFight) and report the stats in a followup post (ping @benwaggoner)
This concludes this report.
As always, I'm open to any kind of feedback to improve my comparisons and my encodes.
Nintendo Maniac 64
10th July 2019, 20:41
I note that the top Intel processor used has only half the cores as the top AMD, so this difference could easily be due to multithreading more than per-core performance improvements.
Keep in mind that the 2990WX uses a very nontraditional CPU die topology that makes it more akin to something like a dual-socket system with two full CPUs. It's so nontraditional that you basically need to use Linux to get any semblance of good performance at all (Wendell from level1techs did a good analysis on the subject in this video here (https://www.youtube.com/watch?v=M2LOMTpCtLA)).
You can also see from the results that even normal Threadripper like the 12core 2920X (which still uses a somewhat nonstandard die configuration) is getting beaten by the 9900K and 3700X which both use a very traditional CPU core configuration by comparison (one could even argue that the separate I/O die on the 3700X is actually more traditional and is akin to the days of northbridges and external memory controllers ala Athlon XP and Core 2 Duo).
Nevertheless, there could very well be a point of diminishing returns in terms of multicore scalabilty for SVT-AV1 that 32c/64t just isn't seeing the utilization that it could otherwise, and even more-so with such the nontraditional core arrangement of the 2990WX.
Also, I note that the Intel processor used in comparison is from 2017.
While true, keep in mind that the per-GHz performance on Intel has not changed at all and won't change until their 10nm parts.
The current equivalent would probably be the i9-9980XE, which as two more cores and 7% faster clock.
The i9-7960X was not the flagship part of its generation - there was in fact an 18core 7980XE during that gen as well (albeit with a bit lower base clock).
This tells me that Phoronix wasn't actually trying to use the highest-end Intel CPU parts that are available, even within a given CPU generation.
mandarinka
12th July 2019, 23:11
When you say they "declined to participate" did they respond and say they didn't want to take part or did you just not hear from them after making a broad request in a forum post?
I believe the comment above yours saying ("Have you asked?") Is written by a developer of EVE, which suggests they didn't know they'd been asked, so possibly an email has got lost in a spam trap.
I suspect that offer is for Amazon, not for the puproses of this open forum/us plain end users.
Recently somebody asked on AOM IRC whether it would be possible for the Parkjoy encode that Two Orioles showed on the recent conference presentation to be shared/uploaded. They were rejected (https://freenode.logbot.info/aomedia/20190704) - by their words the encodes are not actually"classified", but it is "too much work" (https://freenode.logbot.info/aomedia/20190705). Take it as you will but it AFAIK there is not a signle public stream or file encoded by their software, out in the open (or am I missing something?), so it might not be a coincidence or something that is gonna change. I don't want this to sound like bashing, but perhaps NDAing all that is the business policy that they want/need/are forced to use by the nature of the field. Few years have passed and I can't see how there was no way some more transparent comparison test couldn't have been arranged one way or another, so I assume they just don't wish to do that. Sharing samples like what Beamr guys did is one way they could brag about their quality without the encoder leaving their hands...
You would probably have to get such sample from streaming/video services who are using the software, when content encoded with Eve appears via them.
Blue_MiSfit
15th July 2019, 23:37
Yeah what the heck, guys! Why can't they release some sample bitstreams of open source content? The pros among us are totally interested in commercial encoders, but only if they can compete honestly.
benwaggoner
17th July 2019, 02:31
I suspect that offer is for Amazon, not for the puproses of this open forum/us plain end users.
The shootout is a personal effort of mine, unrelated to my day job. And a direct request for an Eve sample. was made and declined by someone personally familiar with us both.
Recently somebody asked on AOM IRC whether it would be possible for the Parkjoy encode that Two Orioles showed on the recent conference presentation to be shared/uploaded. They were rejected (https://freenode.logbot.info/aomedia/20190704) - by their words the encodes are not actually"classified", but it is "too much work" (https://freenode.logbot.info/aomedia/20190705). Take it as you will but it AFAIK there is not a signle public stream or file encoded by their software, out in the open (or am I missing something?), so it might not be a coincidence or something that is gonna change. I don't want this to sound like bashing, but perhaps NDAing all that is the business policy that they want/need/are forced to use by the nature of the field. Few years have passed and I can't see how there was no way some more transparent comparison test couldn't have been arranged one way or another, so I assume they just don't wish to do that. Sharing samples like what Beamr guys did is one way they could brag about their quality without the encoder leaving their hands...
And Beamr has seen a lot of success. I'm not sure of anyone using Eve in production.
You would probably have to get such sample from streaming/video services who are using the software, when content encoded with Eve appears via them.
Do we know of any that have confirmed they are using Eve in production?
LigH
18th July 2019, 09:25
New uploads: (MSYS2; MinGW32 / MinGW64: GCC 9.1.0)
AOM v1.0.0-2084-g42451f74e (https://www.mediafire.com/file/a4w52hb8l2wbrzt/aom_v1.0.0-2084-g42451f74e.7z/file)
rav1e 0.1.0 (20190430-207-g6ac87d8) (https://www.mediafire.com/file/j1fqdmj9ye5ejm0/rav1e_0.1.0_20190430-207-g6ac87d8.7z/file) built 2019-07-17; new verbose version numbering
dav1d 0.3.1 (2019-07-17, g15a9386) (https://www.mediafire.com/file/ylq5u0k6af32flj/dav1d_0.3.1_2019-07-17_15a9386.7z/file)
IgorC
24th July 2019, 15:20
AV1 Ecosystem Update: June 2019
https://www.singhkays.com/blog/av1-ecosystem-update-june-2019/
Spyros
6th August 2019, 23:13
dav1d 0.4.0 'Cheetah' released (https://code.videolan.org/videolan/dav1d/-/tags/0.4.0)
It supports all the AV1 features and all bitdepths.
0.4.0 brings large improvements in speed on ARM64 (up to 25% speedup) and minor improvements on SSE and ARM. It also improves the RAM usage quite significantly, sometimes more than halving the RAM used.
FFmpeg 4.2 released (https://ffmpeg.org/index.html#pr4.2) with AV1 decoding support (through libdav1d) & more
LigH
8th August 2019, 08:06
Instead, MABS disabled ffmpeg support for librav1e because the previously working patch doesn't work anymore. I guess the two projects have to find a common and more stable API again.
Nintendo Maniac 64
8th August 2019, 19:10
Phononix has a new review including more SVT-AV1 v0.5 encoding performance metrics (located ~1/3 the way down the page) as well as dav1d v0.3 decoding performance metrics (located ~2/3 the way down the page); do note that all this testing was conducted on Ubuntu Linux:
https://www.phoronix.com/scan.php?page=article&item=amd-epyc-7502-7742&num=4
This time they were testing the new Zen2-based AMD Epyc 32core (Epyc 7502) and 64core (Epyc 7742) chips against existing top-end Intel Xeon and AMD Epyc CPUs in both single-socket and dual-socket configurations.
dav1d v0.3 decoding was also included on the performance-per-dollar page, though oddly enough SVT-AV1 was not (scroll down to around half way down the page):
https://www.phoronix.com/scan.php?page=article&item=amd-epyc-7502-7742&num=9
And in the according forum thread for that review, there's a post containing information for dav1d's decoding performance in fps at both 1080p and 4k as well as frames-per-dollar at 4k:
https://www.phoronix.com/forums/forum/phoronix/latest-phoronix-articles/1118516-amd-epyc-7502-epyc-7742-linux-performance-benchmarks#post1118526
soresu
24th August 2019, 12:57
Github commits on rav1e have been fairly busy recently, any chance we can get a comparative improvement since the last result on this thread?
soresu
26th August 2019, 14:11
I found a gitlab repo for the dav1d GPU acceleration GSoC, seems like SGR and CDEF have been implemented in Vulkan, and the same repo even has a GLES branch.
Link here (https://code.videolan.org/stebler/dav1d/commits/vulkan3).
It will be interesting to see if they can get weaker non ASIC SoC's running well by taking advantage of the previously untapped GPU.
benwaggoner
27th August 2019, 20:56
I found a gitlab repo for the dav1d GPU acceleration GSoC, seems like SGR and CDEF have been implemented in Vulkan, and the same repo even has a GLES branch.
Link here (https://code.videolan.org/stebler/dav1d/commits/vulkan3).
It will be interesting to see if they can get weaker non ASIC SoC's running well by taking advantage of the previously untapped GPU.
For all the attention encoding on GPU has had over the years, compression with a modern codec is actually about the worst video-related task to run on a GPU. Preprocessing, compositing, and decoding are all much more determinate and parallelizable processes than optimal encoding in complex modern codecs with so many interrelated mode decisions.
Nintendo Maniac 64
27th August 2019, 21:31
For all the attention encoding on GPU has had over the years, compression with a modern codec is actually about the worst video-related task to run on a GPU. Preprocessing, compositing, and decoding are all much more determinate and parallelizable processes than optimal encoding in complex modern codecs with so many interrelated mode decisions.
Considering that the post you quoted is referring to dav1d rather than rav1e, I would presume that they were in fact referring to GPU-accelerated decoding rather than GPU-accelerated encoding.
soresu
27th August 2019, 23:54
Considering that the post you quoted is referring to dav1d rather than rav1e, I would presume that they were in fact referring to GPU-accelerated decoding rather than GPU-accelerated encoding.
Correct, I understand there are limitations to what the average ARM SoC can do, but leaving the GPU running idle during decode seems a sad waste.
soresu
31st August 2019, 21:31
Given Qualcomm has still yet to join AOM, I wouldn't expect them to do a Hexagon DSP implementation of an AV1 decoder as they did with HEVC back in the day.
NikosD
1st September 2019, 08:43
@benwaggoner
There is no such thing as "GPU encoding" nowadays.
It's an old term referring to the old days where GPGPU processing used for video encoding.
Nowadays all GPUs from the three major vendors (Intel, nVidia, AMD) contain a fixed-function hardware unit, an ASIC, just for the purpose of video decoding/encoding/pre-post processing.
The quality of H.265 hardware encoding of Turing cards aka Turing specific ASIC for video encoding is more than good enough and its speed is out of this world compared to software encoders.
Nintendo Maniac 64
2nd September 2019, 05:56
The quality of H.265 hardware encoding of Turing cards aka Turing specific ASIC for video encoding is more than good enough and its speed is out of this world compared to software encoders.
EposVox has similarly praised both the HEVC encoder on Navi as well as Turing's AVC encoder for being very fast with very good quality as well, especially Navi's HEVC encoder (though has also mentioned that actually getting it to work is a bit of a pain).
(Navi's AVC encoder is still no better than previous AMD GPUs however, meaning you really shouldn't use it)
benwaggoner
3rd September 2019, 19:30
@benwaggoner
There is no such thing as "GPU encoding" nowadays.
It's an old term referring to the old days where GPGPU processing used for video encoding.
There are still some products and ongoing experimentation for how to leverage GPU in parallel with CPU for improved encoding. But software has certainly pulled ahead in the last five years.
This could be of particular interest for new codecs like AV1, EVC, and VVC where mature fixed-function implementations aren't yet available. The rapid iteration possible with software is a huge benefit in early-stage codec development and deployment.
Nowadays all GPUs from the three major vendors (Intel, nVidia, AMD) contain a fixed-function hardware unit, an ASIC, just for the purpose of video decoding/encoding/pre-post processing.
Of course. And they are essential for things like game streaming where "good enough" quality without taxing GPU or CPU primary processing is needed. But the efficiency is going to be a lot lower than with a good software encode (Like 30%+ higher bitrates required).
The quality of H.265 hardware encoding of Turing cards aka Turing specific ASIC for video encoding is more than good enough and its speed is out of this world compared to software encoders.
Yes, certainly. The economics might not make sense for content that gets streamed multiple times, but for personal use when file size is less of a concern than encoding time, there's a place for it.
Although we don't have any GPUs with AV1 fixed function units yet, do we?
NikosD
3rd September 2019, 19:39
This could be of particular interest for new codecs like AV1, EVC, and VVC where mature fixed-function implementations aren't yet available. The rapid iteration possible with software is a huge benefit in early-stage codec development and deployment.
Although we don't have any GPUs with AV1 fixed function units yet, do we? I think there are no HW encoders or decoders for AV1 yet.
And I'm starting to believe that due to the complexity of the codec, it could be the first time that we will not see hybrid (GPU+CPU) decoders/ encoders and we will go straight to fixed-function decoders/encoders of AV1.
Agreed with your post above.
NikosD
3rd September 2019, 20:37
I think it's already posted, but just as a reminder we could also wait for Vulkan/OpenGL GPU assisted hybrid decoding of dAV1d.
I have my doubts, but OK:
https://www.phoronix.com/scan.php?page=news_item&px=DAV1D-Vulkan-GLES-Experiment
soresu
4th September 2019, 03:44
Not sure when the GSoC finishes, but he's still posting commits on that branch including CDEF opts.
marcomsousa
4th September 2019, 19:06
I think there are no HW encoders or decoders for AV1 yet.
Realtek have one HW decoder for av1 - https://www.realtek.com/en/press-room/news-releases/item/realtek-launches-worldwide-first-4k-uhd-set-top-box-soc-rtd1311-integrating-av1-video-decoder-and-multiple-cas-functions
Android 10 is released with the support for Opus audio support and AV1 video codec support- https://www.phoronix.com/scan.php?page=news_item&px=Android-10-Released
Av1 August news - https://www.singhkays.com/blog/av1-ecosystem-update-august-2019/
NikosD
4th September 2019, 19:55
Realtek have one HW decoder for av1 - https://www.realtek.com/en/press-room/news-releases/item/realtek-launches-worldwide-first-4k-uhd-set-top-box-soc-rtd1311-integrating-av1-video-decoder-and-multiple-cas-functions Right!
I had something in mind that I had recently read, so I found it also here, in this thread: First SoC launched by Realtek: https://www.realtek.com/en/press-room/news-releases/item/realtek-launches-worldwide-first-4k-uhd-set-top-box-soc-rtd1311-integrating-av1-video-decoder-and-multiple-cas-functions Moreover, I took the chance and found out this recorded video of the chipset in action, decoding in real-time a 4K60fps AV1 clip from YouTube without dropping a frame: https://www.youtube.com/watch?v=IGtjBwMwTtE
Lastly, I even found an announcement of Realtek regarding an integrated circuit RTD2893, capable of 8K AV1 decoding: https://www.realtek.com/en/press-room/news-releases/item/realtek-wins-three-best-choice-awards-at-computex-taipei-2019-including-2-best-choice-golden-awards-copy.
Realtek is first and fast!
Blue_MiSfit
5th September 2019, 00:20
...recorded video of the chipset in action, decoding in real-time a 4K60fps AV1 clip from YouTube without dropping a frame: https://www.youtube.com/watch?v=IGtjBwMwTtE
That's awesome!
soresu
5th September 2019, 15:01
Working AV1 decoder chips are good, I'm guessing they might make it into 2020 TV's.
Of course there are also the forthcoming Rockchip SoC's coming next year - I'm pretty sure one of them (RK3530 I think) was aimed at HDTV's too, with the other (RK3588) likely to make it into cheap Android TV boxes with a bit more oomph than previously seen in that segment, given most non smartphone ARM based kit tops out at low clocked A72/A73 CPU cores.
soresu
5th September 2019, 16:13
On a side note, the experimental branch of AOM has a patch for 'NN based entropy coding' which presumably is a bit more involved than mere rate decision or motion estimation optimisations.
I wonder how it compares to the AV1 entropy coding for quality/complexity, isn't that a derivative of the Daala technique?
marcomsousa
5th September 2019, 20:48
Paint.NET adding AV1 (*.avif) decoding in v4.2.2
vidschlub
11th September 2019, 03:48
It's kinda strange we have projects like x264/x265 for patent encumbered H.264/H.265 codecs, yet nothing like that for VP9/AV1.
I am unsure if I previously responded to this but this is more info I wasn't aware of.
Isn't there 3x AV1 encoders?
Are all 3 of them closed source or only designed for mass scale installs on cloud hardware or something?
Why are the 3 AV1 encoders inferior (in some ways?) to the x264 and x265 projects?
LigH
11th September 2019, 08:05
How much time has already been spent to develop x264 and x265? And since when are "final" AV1 specifications available? Compare and ask again. :sly:
hajj_3
11th September 2019, 12:06
Possible updates about AV1 adoption over the next week: https://aomedia.org/aomedia-members-demo-av1-at-ibc2019-demos-spotlight-av1s-royalty-free-ultra-high-definition-uhd-web-video-capabilities/
soresu
11th September 2019, 18:43
I am unsure if I previously responded to this but this is more info I wasn't aware of.
Isn't there 3x AV1 encoders?
Are all 3 of them closed source or only designed for mass scale installs on cloud hardware or something?
Why are the 3 AV1 encoders inferior (in some ways?) to the x264 and x265 projects?
There are 3 open source AV1 encoders that I know of - LIBAOM (AOM reference), SVT-AV1 (Intel) and RAV1E (Xiph/Mozilla), in various stages of speed and quality optimisation.
There are other closed source, proprietary encoders like Aurora (ML focused optimisations galore), EVE-AV1 (Two Orioles/Ronald Bultje), and likely the usual customers like Ateme, Cisco and such have their own encoders in various stages at the moment.
The recent Big Apple Video conference gave a run down on most of them save RAV1E, likely due to Vimeo's conflict of interest there as a conference sponsor and a RAV1E contributor/sponsor.
The main point to take away is that x265 has 6-7 years of development, and x264 has closer to 15 years of development. They are both fairly mature, if not necessarilly the fastest options today (SVT-HEVC may change the game on x265 on speed).
Meanwhile AV1 was only standardised last year, it takes time to get these new codecs ship shape for production purposes, let alone significant maturity.
marcomsousa
13th September 2019, 04:29
Adapt to multi-codec world or die, warns Bitmovin as AV1 escalates
Observing the video ecosystem from the video developer perspective is often overlooked, overshadowed somewhat by those higher up, so it was refreshing to read a report from encoding expert Bitmovin – giving the devs a well-deserved voice. Results from Bitmovin’s third annual developer survey showed an expectantly overwhelming reliance on H.264, although interestingly one-in-five developers plan to implement AV1 in 2020 – with big ramifications for the wider video industry. Device manufacturers, browser vendors, and content distributers like Cisco, Mozilla, and YouTube have already started implementing AV1 on larger scales, leading Bitmovin to conclude that AV1 is well positioned to compete with H.265/HEVC and to succeed VP9 for open-source use cases in 2020. This goes against the majority of conversations…
https://rethinkresearch.biz/articles/adapt-to-multi-codec-world-or-die-warns-bitmovin-as-av1-escalates/
LigH
13th September 2019, 07:19
It's all "conservative politics", even in technology :sly:
soresu
13th September 2019, 09:14
What's conservative about adapting, or am I missing something in the full article?
stax76
14th September 2019, 19:04
If x265 and nvenc are the last encoders relevant for home users then I will miss building encoder GUIs, the first generations were painful but the last two generations were quite fun. :)
benwaggoner
17th September 2019, 10:22
If x265 and nvenc are the last encoders relevant for home users then I will miss building encoder GUIs, the first generations were painful but the last two generations were quite fun. :)
There isn't anything codec-specific about nvenc. NVidia supports HW decode for multiple codecs, and can add more.
Something like 85% of new phones have HEVC HW decode, and we don't even have a release date for the first phone SoC that does hardware AV1. UHD Blu-ray and ATSC 3.0 are HEVC. There is going to be a substantial market for HEVC for a decade or more, just like there is still quite a lot of MPEG-2 still being encoded and delivered, and just like there will be a lot of H.264 for years too.
Heck, there was still Windows Media PlaysForSure content being published as of 2-3 years ago.
LigH
17th September 2019, 19:51
The usual threshold of hardware encoders is the limited temporal complexity. This hurts more for more advanced codecs taking more advantage of temporal redundancies. NVEnc may be a useful realtime encoder for AVC and HEVC. But it will be beaten easily if you can encode "offline" and spend more efforts into looking ahead.
marcomsousa
20th September 2019, 19:07
There is lot of interest right now for HEVC 8k real time for the next sport events..
Nintendo Maniac 64
21st September 2019, 00:11
Phoronix has more benchmark numbers of SVT-AV1 v0.5 and dav1d v0.3, this time with the 48-core EPYC 7642:
https://www.phoronix.com/scan.php?page=article&item=amd-epyc-7642&num=3
excellentswordfight
23rd September 2019, 08:32
There is lot of interest right now for HEVC 8k real time for the next sport events..
Which is crazy to me, I work with alot of UHD sport events, and we dont use enough bandwith already to do it justice, both on contribution and distribution side. The move to 8k is just beyond me, it must be tons of money from the monitor manufacturers cause I cant see any broadcasters or productions to see benefits in it. The price/performance ratio for UHD is already awful. I dont get the res craze, if you dont wanna spend the bits, dont increase it. Heck XDCAM50 (mpeg2 1080i) can still look better then what google is doing with 4k.
I would rather see a push for 1080p50/60, full 10bit pipline and rec2020 with a modern codec and a decent bitrate to replace the norm of 1080i/720p h264, cause it still looks amazing and takes way less effort and money to upgrade to. Cause most viewers wouldnt wanna finance an multi million dollar upgrade for some tech dreams, but I guess that we can still make them.
This is from an live/broadcast perspective mind you, VOD is a different scenario.
benwaggoner
24th September 2019, 20:44
I would rather see a push for 1080p50/60, full 10bit pipline and rec2020 with a modern codec and a decent bitrate to replace the norm of 1080i/720p h264, cause it still looks amazing and takes way less effort and money to upgrade to. Cause most viewers wouldnt wanna finance an multi million dollar upgrade for some tech dreams, but I guess that we can still make them.
This is from an live/broadcast perspective mind you, VOD is a different scenario.
And that is certainly what North American sports broadcasters are focused on for the next big thing: 1080p60 10-bit HEVC HDR.
Interest in broadcast 8K is mainly in countries like Japan. South Korea, and China where there is a lot more available RF to use and where TV production is material to the national economy.
It is an interesting question for how many bits with what codec where a higher resolution pays off. At 6 Mbps CBR, 1080p60 is obviously better than 4K. But going from 1080 H.264 to 1080 HEVC worked at similar bitrates. The next gen VVC codec looks like it might offer the efficiency to do 8K VVC at the same bitrates as 4K HEVC. But we're some years out from having practical real-time VVC 8K encoders. Even 8K HEVC is still emerging tech, although certainly being done in trails and such.
Nintendo Maniac 64
25th September 2019, 03:58
Even 8K HEVC is still emerging tech, although certainly being done in trails and such.
A 64core Zen2 Epyc processor can already do realtime 8k HEVC 10bit encoding, and at 79fps to boot:
https://www.techspot.com/news/81905-amd-epyc-rome-cpu-performs-real-time-8k.html
So if you only need 30fps or 24fps, or perhaps 50fps or 25fps for 50Hz territories, then you could get away with a considerably lower CPU core count.
Heck even for 60fps you could probably get away with a 48 core Epyc since, if the multi-threaded encoding scaling was 100%, you'd be seeing 59.25fps on a 48core Epyc. However, since multi-threaded encode scaling almost never perfectly scales with core-count, and since CPUs with fewer cores tend to also have higher base clocks, it'd be quite likely that you could even see slightly above the required 60fps from a "mere" 48core Zen2 Epyc processor.
excellentswordfight
25th September 2019, 08:46
And that is certainly what North American sports broadcasters are focused on for the next big thing: 1080p60 10-bit HEVC HDR.
So i noticed, "3G" seems to be much more a thing there than here in europe. Even if a broadcaster would like to switch to 1080p/3G most productions over here dont offer that contribution anyway. I visited a brand new station/studio in the states not that long ago, everything built on 1080p60, nothing for UHD.
I'm still not sold on HDR for live content tbh, I have played around with both PQ and HLG and both comes with backwards comparability to SDR issues and complexity, and to be frank the live productions I've seen had issues even i HDR. I'm all for "HDR" as a tech, but when it comes to real world live productions it creates a lot of headache, maybe to much for it to ever become mainstream. Something that just going 10bit (which productions in most cases already are) and just increasing the colorspace doesnt (rec2020 specc), and imo is good enough, especially on a high contrast tv-set like an oled.
Interest in broadcast 8K is mainly in countries like Japan. South Korea, and China where there is a lot more available RF to use and where TV production is material to the national economy.
Yeah there is no suprice that it's the same countries that wanna push the consumer hw...
A 64core Zen2 Epyc processor can already do realtime 8k HEVC 10bit encoding, and at 79fps to boot:
https://www.techspot.com/news/81905-amd-epyc-rome-cpu-performs-real-time-8k.html
So if you only need 30fps or 24fps, or perhaps 50fps or 25fps for 50Hz territories, then you could get away with a considerably lower CPU core count.
Heck even for 60fps you could probably get away with a 48 core Epyc since, if the multi-threaded encoding scaling was 100%, you'd be seeing 59.25fps on a 48core Epyc. However, since multi-threaded encode scaling almost never perfectly scales with core-count, and since CPUs with fewer cores tend to also have higher base clocks, it'd be quite likely that you could even see slightly above the required 60fps from a "mere" 48core Zen2 Epyc processor.
Well, although it is impressive, without knowing the actual image quality and bitrate of the produced encode, it doesnt mean that much to me.
Nintendo Maniac 64
25th September 2019, 18:09
without knowing the actual image quality and bitrate of the produced encode, it doesnt mean that much to me.
If broadcasters were that concerned with quality and bitrate, then they wouldn't be pushing to broadcast 8k in the first place. :p
marcomsousa
26th September 2019, 09:56
Broadcom BCM7218X STB SoC Comes With AV1 Hardware Decoding
https://www.cnx-software.com/2019/09/25/broadcom-bcm7218x-stb-soc-av1-hardware-decoding-wifi-6/
marcomsousa
27th September 2019, 18:36
AV1 codec encoder/decoder implementation SVT-AV1 releases version 0.7.0, with more SIMD optimizations for performance, improved multithreading support, and more video features.
ChaosKing
27th September 2019, 19:04
AV1 codec encoder/decoder implementation SVT-AV1 releases version 0.7.0, with more SIMD optimizations for performance, improved multithreading support, and more video features.
I get 11fps with a 1080p BD source. CPU ryzen 2600. Not bad.
soresu
29th September 2019, 12:59
New Phoronix tests with SVT-AV1 (0.6 and 0.7)
Link here (https://www.phoronix.com/scan.php?page=news_item&px=EPYC-7742-Xeon-8280-Video-Enc).
Nintendo Maniac 64
29th September 2019, 17:33
New Phoronix tests with SVT-AV1 (0.6 and 0.7)
Link here: https://www.phoronix.com/scan.php?page=news_item&px=EPYC-7742-Xeon-8280-Video-Enc
Don't forget they included dav1d v0.4.0 results as well, and they're not only measuring performance by FPS now but are even including the max and min FPS to boot.
And they're even testing 10bit AV1 decoding in dav1d! This makes me very happy since, from previous testing, I actually found dav1d's 10bit decode performance on typical consumer CPU thread counts (2 to 8 threads) to be slower than the reference AV1 decoder, so I'll be keeping an eye out for any future performance gains for 10bit AV1 decoding.
(for reference, their most recent review (https://www.phoronix.com/scan.php?page=article&item=amd-epyc-7642&num=3) had used dav1d v0.3.0 and measured performance by seconds with 8bit only)
soresu
29th September 2019, 21:53
And they're even testing 10bit AV1 decoding in dav1d! This makes me very happy since, from previous testing, I actually found dav1d's 10bit decode performance on typical consumer CPU thread counts (2 to 8 threads) to be slower than the reference AV1 decoder, so I'll be keeping an eye out for any future performance gains for 10bit AV1 decoding.
There will be a lot to gain, the various SIMD issue/bugs on gitlab show basically nothing accelerated for 10 bit as yet, as they have been concentrating on 8 bit for all ISA currently:
AVX2 (https://code.videolan.org/videolan/dav1d/issues/78).
SSSE3 (https://code.videolan.org/videolan/dav1d/issues/216).
NEON (https://code.videolan.org/videolan/dav1d/issues/215).
Nevcariel mentioned something about 10 bit content not being available at the moment (broadcast, not test content ala Chimera), so I doubt it's a priority while there are still missing gaps in the 8 bit SIMD code, which there still is on NEON at the very least.
Nintendo Maniac 64
30th September 2019, 01:22
Nevcariel mentioned something about 10 bit content not being available at the moment
Well, non yar-har-fiddle-dee-dee content anyway. :p
benwaggoner
30th September 2019, 18:59
A 64core Zen2 Epyc processor can already do realtime 8k HEVC 10bit encoding, and at 79fps to boot:
https://www.techspot.com/news/81905-amd-epyc-rome-cpu-performs-real-time-8k.html
So if you only need 30fps or 24fps, or perhaps 50fps or 25fps for 50Hz territories, then you could get away with a considerably lower CPU core count.
Making the raw bitstream is becoming feasible, yes. But having it work at industrial scale with broadcast reliability, with good 8K monitoring solutions, end to end distribution and all that? Still at the experimental stage.
An open question would be at what bitrate an 8K signal would look better than a 4K signal. Being able to spend 4x the MIPS per pixel in 4K can be material, and it takes pretty low QPs to preserve detail in an 8K frame that wouldn't be in a 4K downconvert.
Practical 8K could well wait for the VVC codec. I've not seen much AV1 in 8K, but only 20% better than HEVC isn't likely to be sufficient. "8K in 4K bandwidth" is a good story, just like HEVC gave "4K in 1080p bandwidth" versus H.264.
benwaggoner
30th September 2019, 19:03
I'm still not sold on HDR for live content tbh, I have played around with both PQ and HLG and both comes with backwards comparability to SDR issues and complexity, and to be frank the live productions I've seen had issues even i HDR. I'm all for "HDR" as a tech, but when it comes to real world live productions it creates a lot of headache, maybe to much for it to ever become mainstream. Something that just going 10bit (which productions in most cases already are) and just increasing the colorspace doesnt (rec2020 specc), and imo is good enough, especially on a high contrast tv-set like an oled.
HLG has an impossible goal: provide a single signal that does HDR on HDR TVs and SDR on SDR TVs. HLG winds up being tuned to optimize one or the other, or just be mediocre for both. Good live HDR today is produced and distributed in PQ, and SDR versions are made in parallel or derived from the PQ. It's a big challenge since it's a rare production where ALL feeds can be PQ, so SDR needs to get upconverted for some stuff still.
But when you've got a good HDR PQ source, HDR PQ output can look great; much better than SDR at the same bitrate.
soresu
1st October 2019, 00:05
"8K in 4K bandwidth" is a good story, just like HEVC gave "4K in 1080p bandwidth" versus H.264.
Isn't that 8K at 2x 4K bandwidth considering the 50% improvement target of VVC?
Likewise HEVC is at best hitting 70-75% reduction vs AVC for 4K, and even then x265 does not seem to be quite living up to that currently given the recent MSU test results (https://www.compression.ru/video/codec_comparison/hevc_2018/#4k_report), though maybe I read them wrong?
Nintendo Maniac 64
3rd October 2019, 19:08
Even more SVT-AV1 v0.7 benchmarks from Phononix, this time all on the i9-7980XE but testing performance between different OSes and distros (including Windows 10):
https://www.phoronix.com/scan.php?page=article&item=windows-linux-creators&num=4
benwaggoner
4th October 2019, 18:21
Isn't that 8K at 2x 4K bandwidth considering the 50% improvement target of VVC?
Likewise HEVC is at best hitting 70-75% reduction vs AVC for 4K, and even then x265 does not seem to be quite living up to that currently given the recent MSU test results (https://www.compression.ru/video/codec_comparison/hevc_2018/#4k_report), though maybe I read them wrong?
In general you need fewer bits per pixel as the number of pixels goes up, as any given artifact is a lot smaller. This is probably extra true for 8K; a completely messed up 4x4 block at 8K takes up as much visual field as 1 bad pixel at 1080p. The classic rule of thumb is that bitrate should go up proportionately to the 3/4ths power of the change in frame size. Thus:
new bitrate=old bitrate * (width/oldwidth * height/oldheight)^0.75.
But with modern codecs that scale better to high resolutions, the factor is going to be lower/ maybe 2/3rds? That works out to about 2x more bits, which gets covered by the 2x efficiency improvement.
Also, newer codecs have less objectionable distortions, so PSNR underestimates subjective improvements due to error-suppression features like in-loop deblocking. AV1 has a ton of those, which is presumably why we are seeing its greatest strengths at lower bitrates.
This is all ballpark. But H.264 gave 720p at roughly 480p MPEG-2 bandwidth and HEVC gave 4K at roughly 1080p H.264 bandwidth.
marcomsousa
5th October 2019, 08:59
libgav1 new AV1 decoder from Google. Focus on android OS.
https://chromium.googlesource.com/codecs/libgav1/
Nintendo Maniac 64
5th October 2019, 22:15
libgav1 new AV1 decoder from Google. Focus on android OS.
https://chromium.googlesource.com/codecs/libgav1/
And Phoronix has already managed to benchmark it against dav1d v0.4.0 on a bunch of different AMD and Intel CPUs:
https://www.phoronix.com/scan.php?page=news_item&px=dav1d-libgav1-AV1-performance
soresu
6th October 2019, 01:06
Yeah seems libgav1 has a long way to go for x86 at least.
ARM NEON is almost fully accelerated for 8 bit video in dav1d, so until Phoronix does some tests with ARM cores I'd assume a similar result there too.
On a different note, I've been keeping an eye on the experimental AOM/AV2 code branch for a while now.
Seems like a fair amount of work has gone in to it already - though judging any current cumulative improvement is difficult without an obvious AWCY link that compares the AOM master to the experimental code.
Can anyone point me to something here?
Tommy Carrot
6th October 2019, 02:04
On a different note, I've been keeping an eye on the experimental AOM/AV2 code branch for a while now.
Seems like a fair amount of work has gone in to it already - though judging any current cumulative improvement is difficult without an obvious AWCY link that compares the AOM master to the experimental code.
Can anyone point me to something here?
According to this (https://aomedia.googlesource.com/aom/+/613bf43dfe7153973a4cd07b5f053ec6da1c6ff6), the gains are currently at around 0.55%. Not much currently, but i guess they are only starting to add new tools.
marcomsousa
8th October 2019, 16:33
AV1 is ready for prime time: SVT-AV1 beats x265 and libvpx in quality, bitrate and speed
https://medium.com/@ewoutterhoeven/av1-is-ready-for-prime-time-svt-av1-beats-x265-and-libvpx-in-quality-bitrate-and-speed-31c1960703db
benwaggoner
8th October 2019, 17:28
AV1 is ready for prime time: SVT-AV1 beats x265 and libvpx in quality, bitrate and speed
https://medium.com/@ewoutterhoeven/av1-is-ready-for-prime-time-svt-av1-beats-x265-and-libvpx-in-quality-bitrate-and-speed-31c1960703dbThe VMAF scores are too close to indicate any real subjective difference. And things get confusing given some encoders, and AV1 itself, are tuned against VMAF. The more a metric is explicitly targeted, the lower its correlation to subjective quality measurements becomes.
It does suggest quality @ bitrate are becoming ballpark similar, which itself is quite an accomplishment versus encoders that have been developed for far longer.
Sent from my SM-T837V using Tapatalk
Tommy Carrot
8th October 2019, 20:57
I'd say for visual quality SVT-AV1 isn't quite there yet. There still are some motion artifacts that are not really harmful for the metrics, but visually are quite unpleasant. The slowest preset mostly gets rid of them, that's where SVT-AV1 really starting to get the upper hand over x265, but that's just way too slow.
But still, SVT-AV1 is improving quite steadily. Aomenc is still better for maximum quality per bitrate, but it's just too slow, and the encoding speed is improving at a snail's pace.
soresu
9th October 2019, 01:24
I'm more interested in that ML augmented Aurora AV1 encoder, sounds like it's got some gains over libaom's quality, while putting on some considerable speed, though my memory of their BAV talk is a little hazy now.
RanmaCanada
9th October 2019, 02:24
AV1 is ready for prime time: SVT-AV1 beats x265 and libvpx in quality, bitrate and speed
https://medium.com/@ewoutterhoeven/av1-is-ready-for-prime-time-svt-av1-beats-x265-and-libvpx-in-quality-bitrate-and-speed-31c1960703db
If it really was ready for primetime, they would have posted their sources instead of graphs that can easily be inconsequential. We've all seen encodes that score perfectly, but look like utter garbage.
There are no pictures, no videos, nothing visual for people to look at. Which when you are dealing with encodes, is the most important thing.
soresu
9th October 2019, 02:46
If it really was ready for primetime, they would have posted their sources instead of graphs that can easily be inconsequential. We've all seen encodes that score perfectly, but look like utter garbage.
There are no pictures, no videos, nothing visual for people to look at. Which when you are dealing with encodes, is the most important thing.
The graphs are AreWeCompressedYet, the linked Medium post lists its sources at the bottom, and the AWCY source lists its test videos.
RanmaCanada
9th October 2019, 03:06
The graphs are AreWeCompressedYet, the linked Medium post lists its sources at the bottom, and the AWCY source lists its test videos.
I don't think you understood what I meant. There are are no visuals on the page. No frame grabs, no side by side video comparisons, etc. When dealing with a visual medium like codecs, you need to actually see what the end result looks like. Graphs mean nothing as you can program a codec to cheat and give every single signal that a testing program wants. It's what we do here almost every single time someone fights over what codec is better, and what implementations work best.
As I said, we've all seen encodes that score incredibly high, if not perfect, but visually look like garbage.
dapperdan
9th October 2019, 07:42
As I said, we've all seen encodes that score incredibly high, if not perfect, but visually look like garbage.
Have we? Can you point to one such example on arewecompressedyet?
I generally believe the numbers, but for the same reason that people give for not believing the numbers.
Arewecompressedyet has been churning out stats for years now, using standard test clips, standard metric code, documented command lines, all available on GitHub and publicised all the results.
Basically it's the gold standard for objective metrics.
And despite regularly reading complaints here about how it doesn't reflect subjective reality, I've never once seen a link to a specific result, video or even screenshot to back these claims up.
Not that I would put particular weight on one example even if found. It would be like reading a study about a new cancer drug and saying "What about Bob? Bob died! We should use healing crystals instead."
I have seen subjective study results that generally back up the objective results within and between codecs (and this is how they're initially tested and built, so it shouldn't really be a surprise, but it shows they've not been too gamed since)
So something odd with human psychology is going on, but I don't think it's the metrics that are the problem.
edit: I should note the one stat were I do thibk arewecompressedyet is fallible.Id not take a single encoding time number seriously, I'd check there was a consistent pattern, and double check it for multithreaded behaviour if that was relevant.
Tommy Carrot
9th October 2019, 12:12
Dapperdan, there is a simple example. X264 (and probably x265) gives much better results for nearly all metrics with --tune psnr or ssim, compared to the default settings, but visually it's almost always considerably worse.
Also, as i said, SVT-AV1 has pretty good metrics with -enc-mode 4, but visually it's nowhere near x265 for example.
nevcairiel
9th October 2019, 12:41
Metrics have a simple problem that we're seeing here - and that is that none of them truely quantify human perception. If it would accurately reflect that, then there would be no problem, not even as newer metrics come closer, they are still not there.
If you create a new metric, its usually "good" at first, because the metric compares encoders, and encoders don't do anything to cheat.
But as a metric gets more popular, encoders will try to cheat and optimize for those metrics. That means they score high in those metrics - it does not mean the image got better. For a metric to remain relevant for a long time, you would have to make sure of two things, (a) that the metric is somewhat related to human vision, and (b) that encoders are forbidden to optimize for it. Unfortunately (b) is an impossible goal.
This is why purely "objective" metric-driven comparisons are always flawed, if they are not accompanied by subjective human tests as well, and both objective and subjective tests and results are properly aligned.
VMAF got worse as a reference as encoders like SVT and x265 started to optimize for that particular metric, instead of just optimizing for a better visual quality.
dapperdan
9th October 2019, 12:55
Do you happen to have a link to subjective study on naive users (i.e. not people who know what classic artifacts look like) showing that is really the case?
This an extreme example and it may well be true, but even then how do we know for a fact that the person doing the tuning wasn't paying too much attention to one specific artifact and not subtly ruining other areas they weren't focussed on? If the answer is just "well before there was an artifact here and now it's gone then that's not really convincing to me, compared with multiple objective stats that generally correlate with user reports in lab settings.
If all the metrics across all the different test clips show the same improvement in that case, then I'd be suspicious that the metric is actually correct and the tuning is wrong.
I'll try and locate two such encodes on AreWeconpressedYet and compare them.
Found a couple of examples, x264 very slow on beta.areweconpressedyet. not a perfect match as they were done about 6 months apart but, if we ignore that:
As you might expect tune PSNR wins on PSNR Y and APNSR Y, though not on the Cb and Cr variants where they're neck and neck.
The gap similarly closes massively on PSNR HVS, with the standard tuning winning at higher bitrates. Generally most metrics don't really change other than PSNR for Y and SSIM variants.
I'd guess from looking at the results that tune PSNR just means switching off anything they did to tune for SSIM ( the Fast SSIM variant specifically) but that effected PSNR negatively. Those two metrics move in opposite directions, where other metrics (even other PSNR and SSIm variants) show no real difference. So it's somewhat ironic (assuming I'm right about that) to raise this as an obvious case of subjective being better than tuning to metrics.
I'll try to find a tune ssim example, which if I'm correct will do better than default on fast SSIm especially, while losing on PSNR and not really affecting most other metrics.
dapperdan
9th October 2019, 13:56
I couldn't find any x264 tune ssim encodes, but I did find some x265 tune PSNR that show basically the same pattern. Tuning for PSNR at the expense of Fast SSIM, while other metrics don't really change.
The major difference is that the PSNR tuning seems to help with CB and Cr for x265, not just the Y channel.
Fascinating. So it appears x264 and x265 have their own version of VMAF which is just the combination of PSNR and FastSSIM metrics.
LigH
9th October 2019, 15:03
So it appears x264 and x265 have their own version of VMAF which is just the combination of PSNR and FastSSIM metrics.
That sounds confusing to me. PSNR and SSIM are mainly frame-to-frame metrics, whereas VFAM should also consider motion, thus have a temporal window.
The main metric used in x264 and x265 is the "rate factor", if I'm not completely wrong...
dapperdan
9th October 2019, 15:36
I just meant in the sense that VMAF is a fusion of other metrics, including PSNR.
VMAF use fancy statistics to combine them all based how they agreed with subjective measurements but it looks like x264 just combined Fast SSIM and PSNR when they didn't move in opposite directions.
mzso
9th October 2019, 16:27
Will SVT-AV1 and rav1e be added to ffmpeg as encoder libraries?
Or is that not possible because of licensing?
sneaker_ger
9th October 2019, 18:55
AFAIK there are no licensing issues. Both rav1e and SVT-AV1 have open licences (BSD 2-Clause and BSD+patent). Some patches already seem to exist but it seems aren't fully ready yet (i.e. developers still need to do some work or want to wait a bit more for the respective projects to mature).
quietvoid
9th October 2019, 19:12
rav1e is still missing a stable C API for ffmpeg, IIRC.
TD-Linux
10th October 2019, 01:51
Will SVT-AV1 and rav1e be added to ffmpeg as encoder libraries?
Or is that not possible because of licensing?
For rav1e we are waiting on releasing 0.1. We're down to one blocker bug, so it should be out soon: https://github.com/xiph/rav1e/issues/1636
TD-Linux
10th October 2019, 01:54
Fascinating. So it appears x264 and x265 have their own version of VMAF which is just the combination of PSNR and FastSSIM metrics.
Unfortunately it's impossible to use VMAF directly for RDO because an encoder has to evaluate distortion on very small blocks, down to 8x8 pixels. So the best you can do is come up with something that approximates it.
LigH
10th October 2019, 07:43
Just that you mention rav1e ... this week they broke compilation for the x86 target, possibly in an attempt to add assembler optimizations. The developers seem to lack of a cross-compilation environment for x86 (where I am not sure if that specifically means 32 bit code) for automated testing. I discovered that issue during my usual sporadic runs of the media-autobuild suite.
Kirakishou
14th October 2019, 07:42
Hi guys. I’m not good at English, sorry for that. Got file [Xrip][Nekopara][OVA_Extra][GB][1080P][AV1_10bit].mp4 from nyaa. Last and several dozen previous build of mpv playback it choppy. x86 and x64 version of mpv.
But mpc-hc 1.8.4.x86 playback it smooth. Only x86 build of mpc-hc 1.8.4, x64 also playback it choppy. And next builds of mpc-hc also playback it choppy. As far as I understand there were some changes in LAV Filters which led to the worst result. Maybe this information will be useful to someone.
sneaker_ger
14th October 2019, 10:18
MPC-HC 1.8.4 (LAV v0.73.1) uses libaom for AV1 decoding, 1.8.5 (LAV v0.74) uses dav1d. dav1d isn't optimized for 10 bit AV1 yet and it seems for this particular case and your hardware that libaom is faster (for 8 bit dav1d is much faster). Probably same problem for mpv. As dav1d matures this problem will be solved.
https://code.videolan.org/videolan/dav1d/issues/216
https://code.videolan.org/videolan/dav1d/issues/78
benwaggoner
15th October 2019, 00:27
Unfortunately it's impossible to use VMAF directly for RDO because an encoder has to evaluate distortion on very small blocks, down to 8x8 pixels. So the best you can do is come up with something that approximates it.
And VMAF isn't THAT great a metric. Many encoders will do some more sophisticated things internally. particularly around maintaining temporal coherence. VMAF does at least include a lightweight interframe comparison metric, but it doesn't do anything new to figure out how the variation of quality of individual frames impacts the overall viewer experience.
benwaggoner
15th October 2019, 00:31
I couldn't find any x264 tune ssim encodes, but I did find some x265 tune PSNR that show basically the same pattern. Tuning for PSNR at the expense of Fast SSIM, while other metrics don't really change.
The major difference is that the PSNR tuning seems to help with CB and Cr for x265, not just the Y channel.
Well, "help" in the sense that PSNR metrics would improve. --tune psnr reduces the subjective quality of the content at a given bitrate BY optimizing only for improved PSNR scores.
NikosD
23rd October 2019, 13:02
First commercial AV1 hardware decoder, claimed by Chips&Media and it's called Wave510A.
Can handle 4K60fps AV1 main profile using one core@500MHz and expands to dual core@1000MHz for 8K60fps.
Supports AV1 8bit/10bit up to 8Kx8K and up to 50Mbps
More here:
https://www.anandtech.com/show/15003/chipsmedia-launches-wave510a-hardware-av1-decoder-ip
and here:
https://en.chipsnmedia.com/page/product_view/5919
soresu
23rd October 2019, 13:57
Can handle 4K60fps AV1 main profile using one core@500MHz and expands to dual core@1000MHz for 8K60fps.
I wonder what the power draw for those configurations are.
benwaggoner
24th October 2019, 16:46
First commercial AV1 hardware decoder, claimed by Chips&Media and it's called Wave510A.
Can handle 4K60fps AV1 main profile using one core@500MHz and expands to dual core@1000MHz for 8K60fps.
Supports AV1 8bit/10bit up to 8Kx8K and up to 50Mbps
More here:
https://www.anandtech.com/show/15003/chipsmedia-launches-wave510a-hardware-av1-decoder-ip
and here:
https://en.chipsnmedia.com/page/product_view/5919
Interesting!
I wish there was some hint on how many transistors this takes, so we could estimate the silicon cost of adding it to a chip. It could be bigger than normal as this is JUST an AV1 decoder, without sharing anything with H.264/HEVC/VP9/etcetera decoders. In a more mature implementation, one would expect an integrated decoder which supports multiple bitstreams. That takes a lot fewer transistors in total that having all those as independent decoders.
400/500 MHz is pretty reasonable, as it can run in a processor in a relatively lower power state for better battery life on long-term content.
I am not a deep SoC guy, so take all above with an appropriately scaled grain of salt.
I'm looking forward to seeing an announcement for the first device with HW AV1 decode. AV1 isn't relevant for premium content until a material portion of customers have devices with HW decoders with integrated HW DRM.
So much hinges on whether the additive cost of AV1 decode will be low enough to be a default in lower cost SoCs in the next year or two. I'm kinda startled how murky that still is as we approach 2020.
birdie
24th October 2019, 21:10
Considering the timeline it looks like SnapDragon 865 (or whatever comes after SnapDragon 855) won't support AV1 HW decoding which is a huge bummer.
soresu
24th October 2019, 22:17
Considering the timeline it looks like SnapDragon 865 (or whatever comes after SnapDragon 855) won't support AV1 HW decoding which is a huge bummer.
Samsung's next chip doesn't have AV1 and they have actually joined AOM, so Qualcomm's conspicuous absence from AOM makes their support for AV1 dubious at best.
hajj_3
25th October 2019, 19:19
Leaked Amlogic roadmap showing the Amlogic S905X4, S908X and S805X2 chips that will support AV1. The S905X4 looks to be shipping within the next few months:
https://www.cnx-software.com/wp-content/uploads/2019/10/Amlogic-S905X4-S905X8-S805X2-AV1-8K-Processors-Large.jpg
Adonisds
26th October 2019, 21:46
Does Youtube keep the original of every video? Will every video receive an AV1 version? Just future ones?
sneaker_ger
26th October 2019, 22:09
Does Youtube keep the original of every video?
Yes.
Will every video receive an AV1 version? Just future ones?
Possibly every video but we can't say for sure. It will cost a lot of cpu time/money to convert all videos in all resolutions (+SDR/HDR) to AV1. For the longest time Youtube did not encode all videos to VP9 either. So we don't really know how it will be for AV1.
Adonisds
26th October 2019, 22:39
Yes.
Possibly every video but we can't say for sure. It will cost a lot of cpu time/money to convert all videos in all resolutions (+SDR/HDR) to AV1. For the longest time Youtube did not encode all videos to VP9 either. So we don't really know how it will be for AV1.
Thank you. Are all videos from before VP9 available in VP9?
vidschlub
26th October 2019, 23:48
Meanwhile AV1 was only standardised last year, it takes time to get these new codecs ship shape for production purposes, let alone significant maturity.
Due to the extended interest in AV1 from such a wide group of companies, could we expect to see it gain traction faster than 265 at least?
vidschlub
26th October 2019, 23:57
[Xrip][Nekopara][OVA_Extra][GB][1080P][AV1_10bit].mp4 .
As predicted, anime teams are always first to adopting crazy new codecs as soon as possible. I love these guys, they're nuts <3
birdie
27th October 2019, 11:44
Thank you. Are all videos from before VP9 available in VP9?
No. Google encodes into VP9 only videos with more than N number of views or views pre day where N or N/day are yet to be determined.
Due to the extended interest in AV1 from such a wide group of companies, could we expect to see it gain traction faster than 265 at least?
I really doubt that considering that AV1 is up to two orders of magnitude more computationally expensive and older x86 CPUs cannot even decode FullHD 60fps videos encoded in it in real time.
dapperdan
27th October 2019, 14:18
Due to the extended interest in AV1 from such a wide group of companies, could we expect to see it gain traction faster than 265 at least?
Probably a lot depends on how you define traction in this situation.
For example, it wouldn't surprise me if the total stream watch time for AV1 is already above HEVC due to YouTube encoding low res versions of popular videos. Mozilla released some numbers that suggested AV1 Firefox video views on their nightly channel rapidly rose to 20% of all video plays just before YouTube paused their AV1 rollout. Would be interested to see where that's gone since.
Instagram already uses a software decoder for VP9 on Android ( a version of libvpx surprisingly) so switching to dav1d and AV1 for popular content isn't totally unbelievable even before hardware decoders are widespread. (I'm not sure if AV1 would be much better in terms of bitrate than VP9 for their user generated content at low bitrates, but if SVT and dav1d are sufficiently better than libvpx then it would actually save them time and energy to upgrade.
I'd love to see a real world comparison of watching something like Breaking Bad on a phone on a metered connection. What are realistic bitrates for these users, if you can get a 30% bitrate saving just from synthetic film grain and AV1's tools appear to work better at lower bitrates how much does the software decoding actually cost you? Once you factor in network savings is it actually noticeable against the baseline of having the screen on? What about in the download scenario? Is it worth the battery hit to see 5 more episodes on your monthly bandwidth allowance?
IgorC
27th October 2019, 17:38
I wonder why dav1d developers have dedicated time to optimize for SSE2. Isn't SSSE3 already old enough? AMD has catched up and implemented SSSE3 in 2011. Even outdated Core 2 Duo has SSSE3.
While 10 bits decoding has literally zero optimizations till moment.
P.S. Few years ago I have tested 10 years old laptop with Pentium T4200 (SSSE3) which now rests unused. It could barely play Youtube VP9 720p videos while still dropped some frames, leave alone AV1 with its 3x complexity.
AV1 would be actually a downgrade for this kind of hardware (from VP9 720p to AV1 360/480p). And we're talking about CPU with SSSE3.
birdie
27th October 2019, 19:02
https://code.videolan.org/videolan/dav1d/-/releases#0.5.1
http://download.opencontent.netflix.com/?prefix=AV1/Sparks/
Netflix posted new AV1 samples with and without film grain in 540p, 1080p and 2160p
Neither mpv, nor ffplay can open these *.obu files. Any ideas how one can play them?
sneaker_ger
27th October 2019, 19:20
Mux to mkv using mkvmerge first.
nevcairiel
28th October 2019, 10:43
dav1ds AVX2 is fine. If you want to properly compare SSSE3 vs AVX2, then you need to look at Single Threaded benchmarks. Multi-Threading is often limited in scaling, where such differences can "hide".
But you should also not expect twice the performance from AVX2, since once you optimize everything possible with SSSE3/AVX2, the remaining parts that cannot be optimized so easily will impact the performance the most.
Beelzebubu
28th October 2019, 12:55
Ok...So, I take a look at the single threaded performance and I see a 20% gain of AVX2 compared to SSSE3.
On what system (chipset)?
soresu
28th October 2019, 14:43
BTW, any plans for AVX-512 in near future ?
Is there any benefit on this ?
There was a merge request/issue some time ago for adding some support for it (specifically mentioned as Ice Lake), but I don't think any actual optimisations have been committed to the master yet, going by my git commit RSS/Atom feed anyway.
I'm more interested in the GPGPU work that happened over the summer, another mirror repo had some further Vulkan work that seemed like bugfixes or 'piping' as it were.
Is there any chance of getting some bench figures on that work soon?
NikosD
28th October 2019, 17:22
Few things going on there:
YMM (e.g. AVX2) functions are never exactly 2x as fast as XMM (e.g. SSSE3) functions, even in theoretical conditions; True, but to be honest I was expecting something like 60% - 70%.
That's a reasonable gain going from SSEx to AVX2
YMM upper lane use will cause CPU downclocking (but not on modern AMD CPUs, I'm being told); Using a Haswell for years I believed that too, but actually Intel has solved the issue since Skylake.
My Core i3 9100F can keep all core turbo of 4.0GHz forever using AVX2 optimized code (but not power virus like Prime95 small FFT)
certain code in SIMD functions does not use YMM upper lanes (effectively), usually because the block size is too small (width=4-8), but sometimes because we don't want a function-pointer-call overhead (multisymbol coding);
and obviously, a lot of code is not SIMD'ed at all, it's 50%-50% between SIMD and non-SIMD at best.
Together, that means the speedup is well below half of half, so 20% is not entirely unreasonable. Sucks a bit, but you can't beat reality. It's not that I don't believe you or nevcairiel, but personally judging by H.265 encoding/decoding and VP9 decoding, I think the transition from SSEx to AVX2 could be more impressive than 20%.
For me it's still unreasonable.
Will see...
I'm more interested in the GPGPU work that happened over the summer, another mirror repo had some further Vulkan work that seemed like bugfixes or 'piping' as it were.
Is there any chance of getting some bench figures on that work soon? I think pure fixed-function HW will appear for the first time in 2020 using Ampere architecture of nVidia (hopefully) and as I have said in the past, GPGPU can't be that effective with such a complex codec like AV1, in my opinion.
Regarding benchmarks, if there is a DirectShow or a MediaFoundation filter exposed via DXVA2/D3D11VA for AV1 compatible with nVidia GPUs, I would definitely try it although I really don't expect too much from GPGPU for video codecs.
Nintendo Maniac 64
28th October 2019, 19:51
1080p on SSE2 is not our goal. The goal is to have a baseline support so ~5 years (or even earlier?) from now, AV1 can be the baseline, not H.264. We don't know for sure, but this may imply some basic need for SSE2 support. So we're exploring what is possible and how much work it'd be.
Question about what constitutes a baseline - would this happen to be 8bit AV1 only?
I ask because I would have expected by now for 10bit to be the standard baseline for AV1 even for non-HDR content at sub-1080p resolutions since, at least back in the h.264 days, doing such can improve compression (though obviously no hardware 10bit h.264 decoder exists even today...but that's not an issue for newer codecs), but last I checked dav1d's 10bit decode performance was actually slower than even aomedia's reference decoder!
YMM upper lane use will cause CPU downclocking (but not on modern AMD CPUs, I'm being told)
Indeed, AVX2 workloads on Zen-based CPUs (Epyc, Ryzen, Athlon with iGPUs) do not cause any downclocking. It's for this reason that Zen2 (which has a full-width 256bit AVX2 implementation unlike Zen1/+'s half-width 128bit AVX2 implementation) can actually keep up quite well on a per-thread basis to Intel's AVX-512-equipped CPUs in several AVX-512-accelerated workloads like x265 despite no current AMD processor supporting AVX-512.
Beelzebubu
28th October 2019, 21:12
personally judging by H.265 encoding/decoding and VP9 decoding, I think the transition from SSEx to AVX2 could be more impressive than 20%.
For me it's still unreasonable.
OK, let's test your claim. HEVC/VP9/dav1d decoding using FFmpeg, recent snapshot, on my local Haswell laptop of a same-quality encoded sample of the same file (ToddlerFountain), everything single-threaded, in alphabetical order:
AV1: SSSE3 12.913s vs. AVX2 10.638s = 21.4% faster;
HEVC: SSSE3: 15.188s vs. AVX2: 10.938s = 38.9% faster;
VP9: SSSE3 7.133s vs. AVX2 6.748s = 5.7% faster;
So, that's weird, AV1 is halfway between HEVC and VP9 - what's going on here? It's simple: look at the profiles. First of all, let's do AV1, and check the top-10 functions for both runs (SSSE3 first, then AVX2):
2.25 s 18.2% 2.25 s dav1d_prep_8tap_ssse3.hv_w8_loop
1.25 s 10.1% 1.25 s motion_field_projection
700.00 ms 5.6% 700.00 ms decode_coefs
601.00 ms 4.8% 601.00 ms decode_b
475.00 ms 3.8% 475.00 ms dav1d_find_ref_mvs
391.00 ms 3.1% 391.00 ms add_tpl_ref_mv
363.00 ms 2.9% 363.00 ms dav1d_put_8tap_ssse3.hv_w8_loop
329.00 ms 2.6% 329.00 ms dav1d_prep_8tap_ssse3.h_loop
326.00 ms 2.6% 326.00 ms dav1d_cdef_filter_8x8_ssse3.k_loop
242.00 ms 1.9% 242.00 ms add_ref_mv_candidate
vs.
1.36 s 12.9% 1.36 s dav1d_prep_8tap_avx2.hv_w8_loop
1.20 s 11.4% 1.20 s motion_field_projection
741.00 ms 7.0% 741.00 ms decode_coefs
638.00 ms 6.0% 638.00 ms decode_b
479.00 ms 4.5% 479.00 ms dav1d_find_ref_mvs
433.00 ms 4.1% 433.00 ms add_tpl_ref_mv
247.00 ms 2.3% 247.00 ms add_ref_mv_candidate
241.00 ms 2.2% 241.00 ms ..@949.end
229.00 ms 2.1% 229.00 ms mc
206.00 ms 1.9% 206.00 ms dav1d_put_8tap_avx2.hv_w8_loop
What do we see? Nothing unexpected (except maybe the large percentage of time spent in ref_mvs.c functions, which we know about and is tracked in #217 (https://code.videolan.org/videolan/dav1d/issues/217)). Some time spent in Properly optimized functions, but nothing major.
OK, let's look at VP9 - again SSSE3 first, then AVX2:
2.16 s 29.5% 2.16 s decode_coeffs_8bpp
574.00 ms 7.8% 574.00 ms decode_mode
474.00 ms 6.4% 474.00 ms ff_vp9_loop_filter_h_16_16_ssse3
437.00 ms 5.9% 437.00 ms 0x1090748d0
366.00 ms 4.9% 366.00 ms ff_vp9_intra_recon_8bpp
277.00 ms 3.7% 277.00 ms ff_vp9_loop_filter_v_16_16_ssse3
229.00 ms 3.1% 229.00 ms 0x10907478f
220.00 ms 3.0% 220.00 ms ff_vp9_fill_mv
184.00 ms 2.5% 184.00 ms ff_vp9_decode_block
177.00 ms 2.4% 177.00 ms ff_vp9_loopfilter_sb
vs.
2.29 s 32.1% 2.29 s decode_coeffs_8bpp
576.00 ms 8.0% 576.00 ms decode_mode
387.00 ms 5.4% 387.00 ms ff_vp9_intra_recon_8bpp
381.00 ms 5.3% 381.00 ms ff_vp9_loop_filter_h_16_16_avx
261.00 ms 3.6% 261.00 ms 0x1075bc8d0
244.00 ms 3.4% 244.00 ms ff_vp9_loop_filter_v_16_16_avx
224.00 ms 3.1% 224.00 ms ff_vp9_decode_block
222.00 ms 3.1% 222.00 ms 0x1075bc78f
173.00 ms 2.4% 173.00 ms ff_vp9_fill_mv
168.00 ms 2.3% 168.00 ms ff_vp9_loopfilter_sb
(Sorry for the hex codes.) What you see here is simple. There is not much AVX2. There is AVX-XMM (Sandybridge), but that only helps a couple of percent at best, apparently (with three-operand instructions, and SSE4 opcodes).
Last, HEVC (SSSE3 first, then AVX2):
2.45 s 16.0% 2.45 s ff_hevc_hls_residual_coding
1.40 s 9.1% 1.40 s put_hevc_qpel_uni_w_hv_8
766.00 ms 5.0% 766.00 ms put_hevc_qpel_bi_w_hv_8
702.00 ms 4.6% 702.00 ms put_hevc_epel_uni_w_hv_8
685.00 ms 4.4% 685.00 ms put_hevc_qpel_hv_8
512.00 ms 3.3% 512.00 ms ff_hevc_hls_filter
511.00 ms 3.3% 511.00 ms ff_hevc_deblocking_boundary_strengths
498.00 ms 3.2% 498.00 ms hls_coding_quadtree
366.00 ms 2.4% 366.00 ms hls_transform_tree
332.00 ms 2.1% 332.00 ms put_hevc_epel_bi_w_hv_8
vs.
2.45 s 26.6% 2.45 s ff_hevc_hls_residual_coding
549.00 ms 5.9% 549.00 ms ff_hevc_hls_filter
474.00 ms 5.1% 474.00 ms ff_hevc_deblocking_boundary_strengths
448.00 ms 4.8% 448.00 ms hls_coding_quadtree
344.00 ms 3.7% 344.00 ms hls_transform_tree
233.00 ms 2.5% 233.00 ms pred_angular_2_8
230.00 ms 2.5% 230.00 ms hls_prediction_unit
192.00 ms 2.0% 192.00 ms intra_pred_2_8
184.00 ms 2.0% 184.00 ms 0x102e3032a
178.00 ms 1.9% 178.00 ms 0x102e30829
Aha, we have the opposite problem here: HEVC has no SSSE3 fallback for most routines, it only has AVX optimizations. No wonder the difference is so big, and even then, it's only ~40%... If I compare the SSE4.2 performance to AVX2 for the same file, I get a couple of % at best, even though ffhevc has a fair bunch of AVX2 optimizations.
I'm going to leave the decoding claim for you to re-visit if you wish, but I don't think my data supports your claim. In fact, dav1d appears to do quite well.
So, let's move over to the encoding claim: that is entirely plausible. Encoding is much more DSP heavy than decoding, and the expected speed-up is thus bigger. I would indeed expect a significantly-larger-than-20% speedup from AV1 encoding on AVX2 vs. SSEx.
benwaggoner
28th October 2019, 22:36
Question about what constitutes a baseline - would this happen to be 8bit AV1 only?
I ask because I would have expected by now for 10bit to be the standard baseline for AV1 even for non-HDR content at sub-1080p resolutions since, at least back in the h.264 days, doing such can improve compression (though obviously no hardware 10bit h.264 decoder exists even today...but that's not an issue for newer codecs), but last I checked dav1d's 10bit decode performance was actually slower than even aomedia's reference decoder!
There are plenty of 10-bit H.264 decoders out there. There certainly are some devices that can decode 10-bit HEVC but only 8-bit H.264, but plenty who can do 10-bit of both.
While H.264 did show a significant efficiency improvement from 10-bit encoding even of 8-bit sources, HEVC showed much less gain (due to improvements in 8-bit). I wouldn't assume that AV1 would see a gain similar to H.264 without significant testing.
The most important thing about 10-bit is that it's required for HDR content. And HDR is definitely on the path to become mainstream over the next five years.
NikosD
28th October 2019, 23:05
OK, let's test your claim. HEVC/VP9/dav1d decoding using FFmpeg, recent snapshot, on my local Haswell laptop of a same-quality encoded sample of the same file (ToddlerFountain)...
I'm going to leave the decoding claim for you to re-visit if you wish, but I don't think my data supports your claim. In fact, dav1d appears to do quite well. Your analysis is very interesting.
Basically you say that ffmpeg HEVC decoding has no actual SSSE3 optimizations, that's why it gains 40% comparing previous SSE vs AVX2, which is double than AV1 but not that good, while SSE4.2 vs AVX2 is almost the same.
On the other hand, VP9 has no actual AVX2 optimizations that's why SSSE3 vs AVX2 is so close.
TBH, I remembered ffvp9 to be one of the best optimized decoders ever and I thought it was due to AVX2 and not SSSE3 optimizations.
The way you presented your research, it seems that all decoders are doomed in the SSEx vs AVX2 battle.
I will search it a little better and come back if i find something interesting.
Thank you for your time!
nevcairiel
28th October 2019, 23:14
There are plenty of 10-bit H.264 decoders out there. There certainly are some devices that can decode 10-bit HEVC but only 8-bit H.264, but plenty who can do 10-bit of both.
I think you got that backwards. 10-bit H.264 hardware decoders are very rare, in consumer space anyway, while any modern HEVC decoder will handle 10-bit, so devices that can do 8-bit H.264 only, but 10-bit HEVC are ample, and growing with every new device coming out.
Nintendo Maniac 64
29th October 2019, 07:46
And this i straight Haswell, newer chipsets (Zen2, Skylake) will get more, as will encoders.
Well unless it's a Pentium since those still lack AVX support altogether even if it's a desktop Skylake-based Pentium like the G5400 and such (which are effectively what an i3 used to be with Pentiums now being 2core/4thread).
Also, I thought Haswell's implementation of AVX2 was pretty much the same as Skylake's? (and I already touched on the Zen1/+ vs Zen2 implementation of AVX2 in my previous post)...unless you were alluding to Skylake-X's support of AVX-512.
nevcairiel
29th October 2019, 08:50
Well unless it's a Pentium since those still lack AVX support altogether even if it's a desktop Skylake-based Pentium like the G5400 and such (which are effectively what an i3 used to be with Pentiums now being 2core/4thread).
Chips like this is why SSSE3 was still a primary optimization target, among other things. Can't fix the lack of AVX2, but SSSE3 will do the best that is possible on them.
Also, I thought Haswell's implementation of AVX2 was pretty much the same as Skylake's? (and I already touched on the Zen1/+ vs Zen2 implementation of AVX2 in my previous post)...unless you were alluding to Skylake-X's support of AVX-512.
AVX2 improved a bit in Skylake, the better process allows the downclock to be less aggressive, and instruction latencies were slightly improved. The clocking difference would make the biggest difference there.
Skylake is afterall a different micro-architecture then Haswell, its just that since then we didn't get anything new anymore.
soresu
29th October 2019, 12:10
I think pure fixed-function HW will appear for the first time in 2020 using Ampere architecture of nVidia (hopefully) and as I have said in the past, GPGPU can't be that effective with such a complex codec like AV1, in my opinion.
Regarding benchmarks, if there is a DirectShow or a MediaFoundation filter exposed via DXVA2/D3D11VA for AV1 compatible with nVidia GPUs, I would definitely try it although I really don't expect too much from GPGPU for video codecs.
It interests me purely for decoding on platforms that lack a more modern CPU core, but still have GPU power to divy up the decoding effort, which is otherwise going to waste.
The nVidia Shield TV is a good example of this - a relatively weak CPU by modern terms with a strong GPU.
Cortex A57 is decent but aging now, even some Amazon products have more recent ARM cores with higher IPC, and obviously phone products that currently exist are completely limited to CPU power which would strain to put out 4K24 in many cases, let alone 4K60 - which maybe Apple's flagship Axx SoC could do at the moment with dav1d using 8 bit content.
It's all about taking advantage of otherwise wasted compute power to augment decode fps, perhaps even do so more efficiently if enough can be executed on the GPU without extraneous copy/transfer overheads to the CPU.
soresu
29th October 2019, 12:17
I think you got that backwards. 10-bit H.264 hardware decoders are very rare, in consumer space anyway, while any modern HEVC decoder will handle 10-bit, so devices that can do 8-bit H.264 only, but 10-bit HEVC are ample, and growing with every new device coming out.
Unfortunately not fast enough, I bought an Amazon Fire Stick for my dad in 2017 assuming that the advertised HEVC support meant up to 10 bit, only to find out to my horror that most of the HEVC encoded videos I have did not work on it.
At least the newer FS 4K I bought this year has 10 bit capability, still the experience was somewhat disheartening considering the sheer amount of 10 bit content available at the time I bought the first Fire Stick, 5 years after HEVC was standardised.
Beelzebubu
29th October 2019, 12:42
I'm more interested in the GPGPU work that happened over the summer, another mirror repo had some further Vulkan work that seemed like bugfixes or 'piping' as it were.
Is there any chance of getting some bench figures on that work soon?
Last slide (https://www.twitch.tv/videos/498918740?t=4h40m49s) in this presentation (https://www.twitch.tv/videos/498918740?t=4h37m50s), although it was cut-off at a hard 3-minute limit. Slide shows lower [NO!]fps-per-watt[NO!] watt-per-fps when using the GPU code we have so far compared to pure CPU.
NikosD
29th October 2019, 13:03
For ffvp9, it was the other way around, we did everything-and-more in SSSE3, and then did a couple of things (some MC, some inverse transforms) in AVX2, but the smaller inverse transforms and MC, as well as the loopfilters and most intra predictors, were never done. So it's fairly incomplete. I managed to find out a few more details regarding AVX2 optimizations of ffvp9 which are good (for 10bit and 12bit) on the specific routines and with a clear distance from the extremely optimized SSSE3 version.
vp9_diag_downleft_32x32_10bpp_c: 1101.2
vp9_diag_downleft_32x32_10bpp_sse2: 145.4
vp9_diag_downleft_32x32_10bpp_ssse3: 137.5
vp9_diag_downleft_32x32_10bpp_avx: 134.8
vp9_diag_downleft_32x32_10bpp_avx2: 94.0
vp9_diag_downleft_32x32_12bpp_c: 1108.5
vp9_diag_downleft_32x32_12bpp_sse2: 145.5
vp9_diag_downleft_32x32_12bpp_ssse3: 137.3
vp9_diag_downleft_32x32_12bpp_avx: 135.2
vp9_diag_downleft_32x32_12bpp_avx2: 94.0
AVX2 version is 32% faster than SSSE3 for vp9 ipred_dl_32x32_16
vp9_diag_downleft_32x32_12bpp_c: 1534.2
vp9_diag_downleft_32x32_12bpp_sse2: 145.9
vp9_diag_downleft_32x32_12bpp_ssse3: 140.0
vp9_diag_downleft_32x32_12bpp_avx: 134.8
vp9_diag_downleft_32x32_12bpp_avx2: 78.9
AVX2 version is 44% faster than SSSE3 for ipred_dl_32x32
vp9_vert_left_16x16_12bpp_c: 273.8
vp9_vert_left_16x16_12bpp_sse2: 69.4
vp9_vert_left_16x16_12bpp_ssse3: 35.3
vp9_vert_left_16x16_12bpp_avx: 34.6
vp9_vert_left_16x16_12bpp_avx2: 22.4
AVX2 version is 37% faster than SSSE3 for ipred_vl_16x16
1.2x is nothing bad, though. And this i straight Haswell, newer chipsets (Zen2, Skylake) will get more, as will encoders. I was ready to tell you to test exactly the same things using a more modern implementation of AVX2 than Haswell, like Skylake onwards (basically it's the same old 2015 Skylake architecture for all Intel CPUs after Skylake).
I have a Haswell Core i3 4170 and a Coffee Lake Refresh Core i3 9100F and I will try latest dAV1d using LAV Video when 0.5.x version becomes embedded in the decoder (hopefully soonish)
excellentswordfight
29th October 2019, 15:55
0.2.1, the newest available at the time. You can use a nightly version (https://files.1f0.de/lavf/nightly/) which would come with 0.5.1, the newest available right now.
That won't necessarily guarantee that 2160p60 will play on a mobile U-series CPU, but it got the best chances.
Ah, ok, I missed the date on the build, I see now that it was released quite some time ago so that makes sense.
Yeah, was not expecting 60fps, cant do that with sw hevc decoder either on this cpu. I was just interested how far sw decoding has come. But I will get an nightly build then,
Just be careful not to pick the 10-bit variant of the Netflix Chimera video. 10-bit is not optimized at all yet, and its not representative of real-world content yet. YouTube for example only delivers AV1 8-bit so far.
And since there is no 8-bit 2160p variant of Chimera, thats your answer.
I was using a 10bit sample; the new sparks one that was linked on previous page.
soresu
29th October 2019, 16:08
Last slide (https://www.twitch.tv/videos/498918740?t=4h40m49s) in this presentation (https://www.twitch.tv/videos/498918740?t=4h37m50s), although it was cut-off at a hard 3-minute limit. Slide shows lower fps-per-watt when using the GPU code we have so far compared to pure CPU.
Thanks for the update, disappointing but to be expected I guess for an early effort which likely expends a lot of power and time copying data back and forth between the CPU and GPU.
Also reports of Mali GPU power efficiency (used in the Kirin 970/Huawei P20) haven't been stellar either compared to alternatives in the mobile market, especially when A73 is such an efficient CPU core design to compare against.
dapperdan
29th October 2019, 20:43
Thanks for the update, disappointing but to be expected I guess for an early effort
I think the code works. It gets more FPS at the same power usage, or uses less power for the same FPS. I think the corment about "lower fps-per-watt" was intended to be the opposite, "lower watts-per-fps" or at least that's my reading of the graph from the presentation.
Beelzebubu
29th October 2019, 20:50
I think the code works. It gets more FPS at the same power usage, or uses less power for the same FPS. I think the corment about "lower fps-per-watt" was intended to be the opposite, "lower watts-per-fps" or at least that's my reading of the graph from the presentation.
Oops, yes, sorry, my bad. I'll update my post. Thanks for noticing.
TEB
30th October 2019, 13:17
Unfortunately not fast enough, I bought an Amazon Fire Stick for my dad in 2017 assuming that the advertised HEVC support meant up to 10 bit, only to find out to my horror that most of the HEVC encoded videos I have did not work on it.
At least the newer FS 4K I bought this year has 10 bit capability, still the experience was somewhat disheartening considering the sheer amount of 10 bit content available at the time I bought the first Fire Stick, 5 years after HEVC was standardised.
I can confirm this based on my research too. "all" devices released in the last 2-3 years support HEVC, but the devil´s in the details here. 10bit support lacks quite alot (main10), and also many older chipsets, from a spec, support main10 up to 4kp60 but in reality they wont provide more than 8bit@level 3.0..
It depends on a combination of chipset, microcode, OS version etc..
Example: I was testing x265 10bit encoded content in main10 on my older Samsung G7edge that has a snapdragon 820, which from the spec supports all we need (4kp60 uhd...whatever that means in reality..), but i was not able to HW decode anything over main@l2.1 via XO player.. VLC happily decoded main10@l4.1 with a 70% cpu usage on 6 of the 8 cores.. but power drain was awefull..
On H.264, i havent seen any chipset in the last years, regardless of how good the HEVC part of the soc is, able to do anything over hp@4.2. Even the latest AV1 decode enabled chipsets cant do 10bit h.264 either.
Blue_MiSfit
30th October 2019, 18:09
^ exactly. Even though the SOC should do 4kp60, the actual implementation in $phone can only do level 3.0.
soresu
31st October 2019, 14:09
There was some sort of AOM event a couple of weeks ago apparently.
A load of presentation slides can be found at this link (https://aomedia.org/aomedia-research-symposium-2019/).
soresu
31st October 2019, 19:44
Lots of Deep, Neural and ML related stuff discussed at the AOM event it seems - I think we can guess the main direction of AV(x) codecs in the future.
benwaggoner
31st October 2019, 19:46
Unfortunately not fast enough, I bought an Amazon Fire Stick for my dad in 2017 assuming that the advertised HEVC support meant up to 10 bit, only to find out to my horror that most of the HEVC encoded videos I have did not work on it.
At least the newer FS 4K I bought this year has 10 bit capability, still the experience was somewhat disheartening considering the sheer amount of 10 bit content available at the time I bought the first Fire Stick, 5 years after HEVC was standardised.
Any HDR-capable Fire Stick/TV supports 10-bit decode. The feature did come to FireTV first, though, a couple of generations back.
benwaggoner
31st October 2019, 19:53
Lots of Deep, Neural and ML related stuff discussed at the AOM event it seems - I think we can guess the main direction of AV(x) codecs in the future.
Do you mean the encoders or the bitstream definition itself?
I get nervous about actually tuning bitstream features based on ML, because we still lack in well subjectively-correlated metrics. VMAF is the least bad ever, but is SDR only. and the whole question of how individual frame metrics get aggregated into good metrics for interframe encoding remains barely examined. A mean of individual frame values isn't that useful for a clip that is more than a few second of a single shot.
We've seen that AV1 shows better VMAF to MOS ratios than other codecs, which could be a result of this sort of curve fitting to one metric. It's generally true that the more a metric gets used, the lower its subjective correlation becomes, as encoders get increasingly tuned to the metric instead of to subjective ratings.
soresu
31st October 2019, 21:21
Do you mean the encoders or the bitstream definition itself?
I get nervous about actually tuning bitstream features based on ML, because we still lack in well subjectively-correlated metrics. VMAF is the least bad ever, but is SDR only. and the whole question of how individual frame metrics get aggregated into good metrics for interframe encoding remains barely examined. A mean of individual frame values isn't that useful for a clip that is more than a few second of a single shot.
We've seen that AV1 shows better VMAF to MOS ratios than other codecs, which could be a result of this sort of curve fitting to one metric. It's generally true that the more a metric gets used, the lower its subjective correlation becomes, as encoders get increasingly tuned to the metric instead of to subjective ratings.
I've only skimmed 2-3 of the presentation slide decks, but there is certainly an appreciation for bitstream compatibility where the ML models are concerned.
Though Google at least have been knocking on this particular door for at least a few years now, I figure they would not still be at it if they thought it could not be harnessed in a standardised way for a codec bitstream.
Just to be clear, my own level of understanding of all of this is fairly amateur compared to experts on here - I mostly have an avid interest in codecs and more recently ML too (due to various ML optimisations in CG rendering and production fields).
soresu
1st November 2019, 19:59
Found an interesting slide deck called"Adaptive Optimal Linear Estimators for Enhanced Motion Compensated Prediction".
It goes over several things, though I'm not sure if they are alternative solutions or potentially additive improvements.
Assuming they are additive, it discusses at least an average 11.5% BD rate improvement over baseline (presumably AV1 is the baseline).
Link here (https://aomedia.org/wp-content/uploads/2019/10/KenRose_UCSB.pdf).
NikosD
2nd November 2019, 11:00
TLDR:
Win7 64bits, i7-4770k, 3.40GHz (stock), improvement between 0.2.1 and 0.5.1 using only SSSE3 accelerated routines, single thread:
Chimera: 33.2%
Dua Lipa: 34.9%
Included are some AVX2 tests too, because yes. Ok, so you replied to different issues than those I questioned with my results, but still your results are interesting.
It seems that single threaded performance has increased for SSSE3 but AVX2 over SSSE3 is very tiny.
Still, both SSSE3 and AVX2 in single threaded mode are ~30 something % faster for 0.5.1 vs 0.2.1
BUT as I have already stated in my results, the CPU utilization during real-world multi-threaded decoding, eats ALL of the single threaded performance in case of Dua Lipa clip and most of the single threaded gain for the other clips for both SSSE3 and AVX2 versions.
In other words, for the real-world multi-threaded decoding the absolute sum of gain and loss between 0.2.1 vs 0.5.1 is dead zero for Dua Lipa and so small for the other clips.
Sorry, but I can't call this situation as progress after seven months, if overall multi-threaded decoding performance gain is zero.
NikosD
2nd November 2019, 11:49
FFMpeg says it's going to use 4 frame threads and 3 tile threads to decode the files, so I'll be using those numbers.
Chimera: 34%
Dua Lipa: 29.7%
Ok, now your results are even more interesting.
Where can I can I find those executables of 0.2.1 and 0.5.1 versions to run them on my systems ?
I have used the LAV filters versions posted above.
NikosD
2nd November 2019, 12:02
FFMpeg says it's going to use 4 frame threads and 3 tile threads to decode the files, so I'll be using those numbers.
Chimera: 34%
Dua Lipa: 29.7%
Also, I have to say that you are using half of your threading power meaning only 50% CPU logical utilization as you have an 8 threaded CPU and you are using only 4 threads.
I think you have to test it again with at least 8 threads in order to use hyperthreading and all of your CPU's processing power.
SmilingWolf
2nd November 2019, 14:45
Is there any particular reason DXVA Checker grays out the CPU usage line when I try to do the benchmarks? Is there some option I need to set?
NikosD
2nd November 2019, 15:06
Is there any particular reason DXVA Checker grays out the CPU usage line when I try to do the benchmarks? Is there some option I need to set? You only have to run the benchmark (I run it 3 times and take the average) and for every run you will see at the end the CPU usage.
It has also min/avg/max value, even for CPU usage.
Just leave it to finish.
SmilingWolf
2nd November 2019, 15:50
Yeah I tried doing that and it didn't work, see screenshots
https://i.ibb.co/sV30XGJ/Screenshot-1.png (https://ibb.co/sV30XGJ) https://i.ibb.co/2gPh0H6/Screenshot-2.png (https://ibb.co/2gPh0H6)
NikosD
2nd November 2019, 16:06
Yeah I tried doing that and it didn't work, see screenshots
https://i.ibb.co/sV30XGJ/Screenshot-1.png (https://ibb.co/sV30XGJ) https://i.ibb.co/2gPh0H6/Screenshot-2.png (https://ibb.co/2gPh0H6) Ok, you have to go to LAV filters settings and choose SW decoding.
It seems to me that you are using HW decoding.
SmilingWolf
2nd November 2019, 16:24
I did wonder if that was the case, but when I open LAVFilters' config panel this is what I see:
https://i.ibb.co/RcM5bCC/Screenshot-1.png (https://ibb.co/RcM5bCC)
Fresh installation, nothing touched
NikosD
2nd November 2019, 16:59
Fresh installation, nothing touched Latest DXVA Checker v4.2.1 and Connect to Renderer selected ?
Also, when you select the AV1 file, does LAV say unsupported inside DXVA Checker ?
SmilingWolf
2nd November 2019, 17:19
Latest DXVA Checker v4.2.1, system equipped with a GTX1080 (440.97).
It says Unsupported, yes.
This version does not seem to have a "Connect to Renderer" option anywhere.
https://i.ibb.co/2N34LRy/Screenshot-1.png (https://ibb.co/2N34LRy)
SmilingWolf
3rd November 2019, 07:26
Ok, but SmilingWolf and you, have tested different things than me.
Firstly, he posted single threaded performance difference and I posted multi-threaded performance difference
Conveniently forgetting about my two posts dedicated to multi threaded performance aren't we?
http://forum.doom9.org/showthread.php?p=1889274#post1889274
http://forum.doom9.org/showthread.php?p=1889289#post1889289
There isn't a DXVA Checker report yet, afternoon spent trying to make it work notwithstanding, but as I said, CPU utilization goes between 70% and 90% with the two sequences used.
littleD
3rd November 2019, 08:21
Some older versions of dxva checker shows CPU usage. But i did short test and the results of dxch was around 88% utilization while system monitor was showing 100%. Maybe thats why authors of the program turned off the feature temporarily, because of inconsistent results?
SmilingWolf
3rd November 2019, 08:50
I'd love to keep arguing, but I have to agree that won't make the board any favor, so let's let bygones be bygones.
I have already removed all sorts of config files, fresh installations, even reboots etc.
Have you run your own benches on this (4.2.1) version, or on an older one? If the latter happens to be the case, what exact version, so I can download it from the VideoHelp archive?
I don't get what you mean by "no internal commands". They are cmdline applications, just use the same cmdlines I used. I was inside an MSYS shell just so that I could use the "time" command, but I suppose PowerShell on Win10 has got something similar.
Word of advice, they only digest pure IVF files, so at least the Dua Lipa video will have to be freed of its container using "ffmpeg -i Dua_Lipa.mp4 -c:v copy Dua_Lipa.ivf"
NikosD
3rd November 2019, 09:45
I'm not a developer, so environments like MSYS , Visual Studio etc are not frequently installed on my system.
I have compiled a few apps from time to time, even my own code decades ago (!) but I'm not going to do it now setting up MSYS.
I'll give PowerShell a try of course, as I use it from time to time for my job (although I still do a lot using cmd)
But that IVF thing is another obstacle.
Regarding DXVA Checker I used v4.2.1 which of course has everything, as I told you before.
Min/avg/max for FPS and CPU utilization.
Your main problem is that you see things like Video Engine and GPU utilization and you shouldn't.
You need a cleaner OS.
Tomorrow I'll try setting LAV to single-thread mode and run the same tests with Skylake at work.
If nobody here in this forum can confirm or reject my multi-thread results using so familiar tools like LAV filters and DXVA Checker, I'll try to reproduce yours single-thread results.
P.S
DXVA Checker is a sophisticated and accurate tool and the CPU utilization refers to itself only, as a process, not general CPU utilization during its running.
nevcairiel
5th November 2019, 09:40
Comparisons between LAV 0.74.1 and later nightly versions are flawed since the threading strategy changed in FFmpeg, which resulted in 0.74.1 using more frame threads then the later nightlies, making 0.74.1 artificially faster. As such, all your results are invalidated.
This is why you should use as little software as possible to do benchmarking (ie. go as close to the core as possible), as you never know what changes might interfer with your conclusions.
I've also once again changed the thread distribution in 0.74.1-30 from last night, and while its going to use more threads again now, similar to the old logic, its not going to be identical to 0.74.1 in all cases (because I added more tile threads on high core-count CPUs)
Mr_Khyron
7th November 2019, 23:56
AOMedia Research Symposium 2019 Videos
https://www.youtube.com/playlist?list=PL97T7zfqOOF3YKvniyywewtWKpxXky8iI
utack
8th November 2019, 14:49
Lesson learnt from WebP
There are only two I can think of
despite being technically more advanced you can still lose to a decades old legacy format when your encoder is terrible
it does not matter that your format is worse than the legacy competition, if you claim that it is better often enough others will start parroting it and adopt it
Seriously the only area where it might be a tiny bit better is for ultra-high compression where it does not start falling apart as badly as jpeg, for any sane (mid ot high) image quality range the vast array of jpeg encoders are doing a significantly better job of retaining detail
dapperdan
8th November 2019, 20:04
Webp had some other benefits over JPEG outside of compressing photographic images.
JPEG XL seems like WebP's successor in this regard. It's targeted at lots of pain points that would make it a good choice to replace JPEG (and PNG and GiF) on the web and in the browser even if it didn't beat JPEG on compression, though it claims that as well. And maybe the JPEG name will help, though that doesn't seem to have benefitted anyone but the original JPEG.
Not sure there's room for AVIF and JPEG XL but maybe they have subtly different niches.
dapperdan
8th November 2019, 20:09
Ronald's slide showing 4 AV1 encoders all scaling well seems like an improvement from his slide at BIG Apple Video where only SVT seemed to be managing that, with Eve just behind and Rav1e and libaom trailing.
Not sure it it's a direct comparison to the earlier slide but if it is then things should be a lot better for AV1 when cores are available.
soresu
9th November 2019, 01:38
Ronald's slide showing 4 AV1 encoders all scaling well seems like an improvement from his slide at BIG Apple Video where only SVT seemed to be managing that, with Eve just behind and Rav1e and libaom trailing.
Not sure it it's a direct comparison to the earlier slide but if it is then things should be a lot better for AV1 when cores are available.
An interesting point Ronald made implies that AV1 has an intrinsic parallel scaling limitation due to an oversight during the encoder development (12:20 in the video), something to do with superblock boundaries.
Hopefully a lesson learned for AV2 efforts going forward.
marcomsousa
9th November 2019, 15:49
Rav1e release 0.1.0
First official release, published during the Video Dev Days 2019 in Tokyo.
Features
Intra and inter frames
64x64 superblocks
4x4 to 64x64 RDO-selected square and 2:1/1:2 rectangular blocks
DC, H, V, Paeth, smooth, and a subset of directional prediction modes
DCT, (FLIP-)ADST and identity transforms (up to 64x64, 16x16 and 32x32 respectively)
8-, 10- and 12-bit depth color
4:2:0 (full support), 4:2:2 and 4:4:4 (limited) chroma sampling
11 speed settings (0-10)
Near real-time encoding at high speed levels
Rate control (single-pass and two-pass)
Temporal RDO
Scene cut detection
CLI tool and C API
https://github.com/xiph/rav1e/releases/tag/0.1.0
mzso
9th November 2019, 16:41
Hi!
Are there any AV1 videos on youtube besides the beta playlist? So far I haven't found any.
marcomsousa
9th November 2019, 16:49
Hi!
Are there any AV1 videos on youtube besides the beta playlist? So far I haven't found any.
Almost all top videos, but only in low resolutions.
mzso
9th November 2019, 19:20
Almost all top videos, but only in low resolutions.
Thanks. Well, I guess I won't be coming across many then. I don't watch stuff like that, and it looks like a few million views are far from enough. A 4+ billion Ed Sheeran song had it up to to 2160p, but a 2+ billion Taylor swift song only has it up to 720p. I managed to find some Wired videos with a couple million views, that have AV1 though.
It seems like Firefox's (well, Waterfox's to be accurate) AV1 decoding is quite poor. MPV's (after upgrading) and LAV's seem to be a lot better, no hangs or stutter.
marcomsousa
10th November 2019, 09:29
HandBrake 1.3.0 Released
* Added support for reading AV1 via libdav1d
https://github.com/HandBrake/HandBrake/releases
dapperdan
10th November 2019, 16:25
A 4+ billion Ed Sheeran song had it up to to 2160p, but a 2+ billion Taylor swift song only has it up to 720p. I managed to find some Wired videos with a couple million views, that have AV1 though.
A YouTube engineer made a comment about people using YouTube as a radio station, but their licence requiring video, so ultra low bitrate AV1 being a kind of workaround for contractual obligations.
So it's possible that YouTube is measuring the bitrate/resolution of these views and prioritising the ones that are viewed at high quality, not just the ones that are viewed a lot. Just a guess though.
Beelzebubu
11th November 2019, 07:58
An interesting point Ronald made implies that AV1 has an intrinsic parallel scaling limitation due to an oversight during the encoder development (12:20 in the video), something to do with superblock boundaries.
The superblock boundary one to get superblock-row multithreading (an encoder-side-only version of partitions in vp8 or wavefront in hevc) is quite micro, because it is easily worked around by just ignoring the superblock edge's correctness and sacrifice your search' accuracy a tiny little bit. The frame multi-threading one is a bigger deal.
benwaggoner
11th November 2019, 22:50
A YouTube engineer made a comment about people using YouTube as a radio station, but their licence requiring video, so ultra low bitrate AV1 being a kind of workaround for contractual obligations.
So it's possible that YouTube is measuring the bitrate/resolution of these views and prioritising the ones that are viewed at high quality, not just the ones that are viewed a lot. Just a guess though.
A lot of the "TV for radio" clips have super static backgrounds or just scrolling lyrics. So they should encode down super small with more advanced codecs and long GOPs.
Adonisds
18th November 2019, 23:09
When do you think Stadia using AV1 will be available?
marcomsousa
19th November 2019, 00:00
When do you think Stadia using AV1 will be available?Stadia is use cheap client cpu, so AV1 will be in next year for high-end cpu, and the follow year to cheap devices.. So 2 to 3 year for sure.
Marco Sousa
soresu
19th November 2019, 08:25
Stadia is use cheap client cpu, so AV1 will be in next year for high-end cpu, and the follow year to cheap devices.. So 2 to 3 year for sure.
Marco Sousa
There's nothing to suggest AV1 will be in either Samsung or Qualcomm's flagship SoC next year, Qualcomm still has yet to even join AOM when last I checked.
On the other hand the likes of Amlogic have plans (https://www.cnx-software.com/2019/10/20/amlogic-s905x4-s908x-s805x2-av1-1080p-4k-8k-media-processors/)to have AV1 in lower end chips for 2020, so you appear to be wrong on both counts.
benwaggoner
19th November 2019, 18:18
Stadia is use cheap client cpu, so AV1 will be in next year for high-end cpu, and the follow year to cheap devices.. So 2 to 3 year for sure.
On the decode side, maybe in a few years. But Does Stadia have a path to a good 4Kp60 encoder with sufficient efficiency and cost?
There is a lot of untapped potential in using how the game is rendered to drive encoder optimization that could help here (the game knows your motion vectors!). Not sure if anyone's researching that with AV1 currently.
benwaggoner
19th November 2019, 18:55
There's nothing to suggest AV1 will be in either Samsung or Qualcomm's flagship SoC next year, Qualcomm still has yet to even join AOM when last I checked.
On the other hand the likes of Amlogic have plans (https://www.cnx-software.com/2019/10/20/amlogic-s905x4-s908x-s805x2-av1-1080p-4k-8k-media-processors/)to have AV1 in lower end chips for 2020, so you appear to be wrong on both counts.
So far, all the announced AV1 chips have been for living room devices, and nothing for mobile, correct?
Living room has a lot more breathing room for power consumption, thermal management, and thus process. The mobile chips are where every extra transistor costs the most.
Of course, having a working HW decoder at all is a huge milestone, even if it might take a bit for those to migrate into mobile SoCs.
huhn
20th November 2019, 03:47
i wouldn't wait for them to announce a hardware decoder they may just add them or evne ship them already without a word and without any way to access them.
as an example the "first" nvidia card with an HEVC main10 decoder was the 960 this card has a VP9 profile 0 decoder too it was not possible to access it for about a year or so.
the 960 was release in 22.01.2015
VP9 was finalised 17.06.2013
that's just 17 months
i dare to say that AV1 has far more attraction then VP9
AV1 is 19 month old so if someone really wanted to they could have added it and it wouldn't be the first time they made it public which may sound odd at first but if it can't be used anyway it may just be a better this way.
There is a lot of untapped potential in using how the game is rendered to drive encoder optimization that could help here (the game knows your motion vectors!). Not sure if anyone's researching that with AV1 currently.
using multiply frames for encoding is a very bad thing in this case because latency is very important in this task there is a reason the currently released stadia is at best a very poor joke they want money for.
this only counts for lookahead which should be completely avoided if possible.
and how would you even access these information from the game in the first place.
soresu
20th November 2019, 13:16
So far, all the announced AV1 chips have been for living room devices, and nothing for mobile, correct?
He was talking Stadia which uses a Chromecast in its founder bundle - that is a living room/wall powered device.
huhn
20th November 2019, 15:26
stadia "works" with pretty much every device. you don't need a chromecast a phone can do it so can a web browser on the PC. there is missing support for iOS and such but what ever.
hajj_3
20th November 2019, 16:24
New RAV1E build out with big performance improvements when using tiling: https://github.com/xiph/rav1e/releases/tag/20191120
huhn
20th November 2019, 21:35
i don't even see any use case for stadia and AV1 now.
even if the end device is able to decode it you still have to real time low latency encode it which is just not possible at 4K60 UHD.
Mr_Khyron
22nd November 2019, 17:20
Visionular AV1 encoder AURORA demo video
http://35.185.250.137:8080/#
https://i.imgur.com/T3N47lb.png
LigH
23rd November 2019, 19:10
Isn't it "Tears of Steel"? :o
Adonisds
23rd November 2019, 19:41
i wouldn't wait for them to announce a hardware decoder they may just add them or evne ship them already without a word and without any way to access them.
as an example the "first" nvidia card with an HEVC main10 decoder was the 960 this card has a VP9 profile 0 decoder too it was not possible to access it for about a year or so.
the 960 was release in 22.01.2015
VP9 was finalised 17.06.2013
that's just 17 months
i dare to say that AV1 has far more attraction then VP9
AV1 is 19 month old so if someone really wanted to they could have added it and it wouldn't be the first time they made it public which may sound odd at first but if it can't be used anyway it may just be a better this way.
using multiply frames for encoding is a very bad thing in this case because latency is very important in this task there is a reason the currently released stadia is at best a very poor joke they want money for.
this only counts for lookahead which should be completely avoided if possible.
and how would you even access these information from the game in the first place.
So will the GPUs releasing next year from Nvidia and AMD probably decode AV1?
nevcairiel
23rd November 2019, 20:10
So will the GPUs releasing next year from Nvidia and AMD probably decode AV1?
AMD has historically been very slow adopting new codecs, while NVIDIA has usually pushed it decently fast. But any speculation is just that.
huhn
23rd November 2019, 20:53
AV1 has a major problem i forgot.
they changed the bit stream about 10 month ago. i don't have detail if thsi affects decoding but as far as i understand yes it did back then.
maybe AMD cares more now since they are now entering the mobile market. taking recent development crashing with hardware decoding and the lying with polaris into account it will be intel or nvidia first.
it's just speculation anyway.
Beelzebubu
23rd November 2019, 23:47
AV1 has a major problem i forgot.
they changed the bit stream about 10 month ago. i don't have detail if thsi affects decoding
It does not. The errata-1 is exactly that, an errata to clarify some bitstream (mostly level) constraints. Actual decoding is not affected, it is just intended to simplify worst-case for hardware design.
Actual non-cosmetic changes that I am aware of in the errata1:
#226 (https://github.com/AOMediaCodec/av1-spec/commit/619a830b27cf8195479de51f2d2fc1598c1e2d99)
#227 (https://github.com/AOMediaCodec/av1-spec/commit/f1509dd6f7382e0a9b07dcb535cf9d021b229865)
#228 (https://github.com/AOMediaCodec/av1-spec/commit/11843e6361ddf25b67c651612646aac441d1880d)
Mr_Khyron
25th November 2019, 16:09
SVT-AV1 v0.7.5
https://github.com/OpenVisualCloud/SVT-AV1/releases/tag/v0.7.5
- RDOQ for 10-bit
- Inter Intra Class pruning at MD-Staging
- Global Motion Vector support for 8-bit and 10-bit
- Interpolation Filter Search support for 10-bit
- Palette Prediction support
- 2-pass encoding support
- ATB 10-bit support at the encode pass
- Simplified MD Staging [only 3 stages]
- Inter-Inter and Inter-Intra Compound for 10-bit
- Intra Paeth for 10-bit
- Filter Intra Prediction
- New-Near and Near-New support
- OBMC Support for 8-bit and 10-bit
- RDOQ Chroma
- ATB Support for Inter Blocks
- Temporal Filtering for 10-bit
- Eight-pel support in predictive ME
- MCTS Tiles support
- Added AVX512 Optimizations
- Added AVX2 Optimizations
Nintendo Maniac 64
26th November 2019, 02:30
As usual, a new version of SVT-AV1 is released just in time for Phoenix to benchmark the previous 0.7 version on the new i9-10980XE and Threadripper 3960X and 3970X. :p (they unfortunately didn't have a Ryzen 3950X to benchmark)
https://www.phoronix.com/scan.php?page=article&item=amd-linux-3960x-3970x&num=8
Also included are benchmarks of rav1e v0.1 (1080p), dav1d v0.5.0 (1080p, 4k, and 10bit 1080p).
AV1 performance-per-watt charts can be found on this page (scroll down):
https://www.phoronix.com/scan.php?page=article&item=amd-linux-3960x-3970x&num=12
...and AV1 performance-per-dollar charts can be found on this page (also scroll down):
https://www.phoronix.com/scan.php?page=article&item=amd-linux-3960x-3970x&num=13
Mr_Khyron
26th November 2019, 12:55
https://www.mediatek.com/products/smartphones/dimensity-1000
MediaTek Dimensity 1000
We’re redefining the flagship smartphone experience with MediaTek Dimensity mobile series. MediaTek's technological and product leadership is exhibited in our new series of 5G-integrated SoC’s (system on chip) designed for premium-to-flagship smartphones.
Through the clever combination of the most powerful and innovative technologies in a leading design, this tiny 7nm chip is the new era of mobility, where everyone, and everything is seamlessly connected.
With MediaTek Dimensity, Expect Incredible.
Blur-busting displays, fast 4K multimedia & AV1 decoding
HFR (high frame rate) displays are not just an essential feature for gamers to attack the action, but they’re also beneficial to the everyday experience with notably smoother scrolling of webpages and animations in apps. The MediaTek Dimensity 1000 brings this supreme sensation to smartphones with FullHD+ displays up to 120Hz and 2K+ up to 90Hz.
In addition to enabling hardware video encoding and decoding at 4K 60FPS, the MediaTek Dimensity 1000 is the world’s 1st mobile SoC with AV1 format support: the latest and most advanced video streaming technology supported by major global technology and content makers, helping to enable new levels of visual detail, crisper sound and higher resolutions in the latest wave of streaming media.
hajj_3
26th November 2019, 12:57
The Mediatek Dimensity 1000 chip can hardware decode 4k 60fps AV1, it can't encode in AV1 in case anyone was wondering. Still fantastic news though :)
soresu
27th November 2019, 19:03
The Mediatek Dimensity 1000 chip can hardware decode 4k 60fps AV1, it can't encode in AV1 in case anyone was wondering. Still fantastic news though :)
Encoder ASICs usually take longer to appear if memory serves, I certainly remember HEVC decoders came before encoders in mobile.
Adonisds
29th November 2019, 17:11
The Mediatek Dimensity 1000 chip can hardware decode 4k 60fps AV1, it can't encode in AV1 in case anyone was wondering. Still fantastic news though :)
These 4k60 hardware decoders, are they only capable of decoding 1/4 of that in HDR mode?
Blue_MiSfit
30th November 2019, 00:46
Why would that be the case? The decoder doesn't have to work any harder to decode HDR (assuming you're already doing 10 bit, which you should be).
soresu
30th November 2019, 00:54
Why would that be the case? The decoder doesn't have to work any harder to decode HDR (assuming you're already doing 10 bit, which you should be).
Perhaps he means tone mapping for displaying HDR content on SDR screens?
Nintendo Maniac 64
30th November 2019, 01:47
Phononix benchmarked the performance of Windows 10 vs Linux in both rav1e v0.1 and SVT-AV1 v0.7 on a Threadripper 3970X: https://www.phoronix.com/scan.php?page=article&item=3970x-windows-linux&num=5
Adonisds
30th November 2019, 01:56
Why would that be the case? The decoder doesn't have to work any harder to decode HDR (assuming you're already doing 10 bit, which you should be).
Perhaps he means tone mapping for displaying HDR content on SDR screens?
No, I'm assuming SDR is done with 8 bits and HDR with 10. Why do you assume people would use 10 bit for SDR? Youtube currently uses 8 bits for SDR AV1.
I'm also assuming that these decoders advertised as 4k60 capable are only capable of that in 8 bits. I would a bit surprised if that was not the case.
So possibly they are not really ready for HDR content. And since AV1 is a codec for a time in the future when HDR could be common, that is dissapointing. I would think that since Netflix is pushing hard for it, at least 10bit 4k24 would be supported.
Adonisds
30th November 2019, 02:00
No, I'm assuming SDR is done with 8 bits and HDR with 10. Why do you assume people would use 10 bit for SDR? Youtube currently uses 8 bits for SDR AV1.
I'm also assuming that these decoders advertised as 4k60 capable are only capable of that in 8 bits. I would a bit surprised if that was not the case.
So possibly they are not really ready for HDR content. And since AV1 is a codec for a time in the future when HDR could be common, that is dissapointing. I would think that since Netflix is pushing hard for it, at least 10bit 4k24 would be supported.
Looks like I'm wrong in my second assumption, thankfully. At least the WAVE510A decoder does decode 10 bit 4k60
Edit: looks like I also misunderstood what Blue_MiSfit said. He just said decoders should do 10bit, not SDR videos. Well, that was all pointless. But let my mistakes all be public. I'm just gonna go hide in shame.
Blue_MiSfit
1st December 2019, 00:05
All good :)
10 bit SDR is totally valid, especially in the premium case where source content is at least 10 bit 4:2:2. It's unfortunate that the first wave of hardware HEVC decoders were 8 bit, else all HEVC would be 10 bit today.
We only encode 10 bit HEVC on the service I work on (SDR, HDR10, and Dolby Vision)
NikosD
1st December 2019, 11:22
Anandtech agrees that MediaTek's Dimensity 1000 could be the first consumer mobile SoC to support hardware-decoding of AV1
Media encoding capabilities fall in at 4K60, but here the biggest surprise lies in the chipset's support for AV1 video decoding.
As far as we're aware, this make the D1000 the very first consumer mobile SoC to support the format, which is a great leap forward in terms of future-proofing the devices which are based on the new chip.
benwaggoner
4th December 2019, 23:01
Perhaps he means tone mapping for displaying HDR content on SDR screens?HDR tone mapping is generally done in the SoC or GPU. It shouldn't impact decoder performance, but might increase battery draw.
Of course, doing 1080p on a phone is nearly placebo for most TV/film content. 4K is beyond overkill for such a small screen.
Generally decoders support the same resolution/fps in all supported bit modes. It's more expensive to go beyond 8-bit, but not in a way that makes it materially cheaper to only support >8-bit at lower resolutions.
Sent from my SM-T837V using Tapatalk
huhn
5th December 2019, 01:49
isn't 10 bit HEVC still generally better then 8 bit but the difference is very very small?
and the whole story of 8bpp and 16bpp x265. wasn't that an very important part why x264 10 bit was much better thanks to 10 bit bpp instead of 8.
is the 8 bpp x265 even alive?
how does AV1 handle this is this even comparable?
isn't a 8 bit quality pipe absolutely impossible anyway? just the level/RGB conversation pretty much proves it or even the chroma upscale.
Blue_MiSfit
5th December 2019, 09:10
I didn't mean to imply that encoding 8 bit content in 10 bit AV1 would be the right thing to do.
I just meant that given premium content almost always being > = 10 bit 4:2:2 it makes sense to preserve 10 bit, especially when picky studios are highly critical of how subtle gradients are handled in their content.
Blue_MiSfit
6th December 2019, 02:03
Regarding Google working on their own AV1 decoder, I'd imagine this has to do with being the masters of their own destiny so they can use the decoder however they want on Android / anywhere else, free from any potential licensing incompatibility with dav1d.
nevcairiel
6th December 2019, 09:23
Regarding Google working on their own AV1 decoder, I'd imagine this has to do with being the masters of their own destiny so they can use the decoder however they want on Android / anywhere else, free from any potential licensing incompatibility with dav1d.
dav1d has one of the most liberal licenses anywhere (2-clause BSD), Android already uses libraries under far stricter licenses.
Nevermind that Chrome already uses dav1d as well, so its not like Google somehow doesn't know about it, or something like that.
The reason is probably some BS politics, and has no logical backing.
Blue_MiSfit
7th December 2019, 02:31
Hmm. That's unfortunate, I wonder if we'll ever know the real story here :)
benwaggoner
9th December 2019, 23:23
I didn't mean to imply that encoding 8 bit content in 10 bit AV1 would be the right thing to do.
I just meant that given premium content almost always being > = 10 bit 4:2:2 it makes sense to preserve 10 bit, especially when picky studios are highly critical of how subtle gradients are handled in their content.Premium content is NOT >=10-bit 4:2:2. 10-bit 4:2:0 is all that anyone is distributing to consumers via streaming and disc. The masters are, certainly, but it's not like those look visually better than distribution codecs at a sufficient bitrate.
Sent from my SM-T837V using Tapatalk
benwaggoner
9th December 2019, 23:25
isn't 10 bit HEVC still generally better then 8 bit but the difference is very very small?
and the whole story of 8bpp and 16bpp x265. wasn't that an very important part why x264 10 bit was much better thanks to 10 bit bpp instead of 8.
is the 8 bpp x265 even alive?
how does AV1 handle this is this even comparable?
isn't a 8 bit quality pipe absolutely impossible anyway? just the level/RGB conversation pretty much proves it or even the chroma upscale.There is a significant but shrinking number of devices, mainly handheld, that support 8-bit HEVC but not 10-bit.
More common is devices with 10-bit decoders that then feed into 8-bit compositors, GPUs, or display controllers.
Sent from my SM-T837V using Tapatalk
foxyshadis
10th December 2019, 00:52
Do we have any evidence that 8-bit sources encode better in 10-bit than 8-bit in AV1?
If nothing else, 8b RGB to 10b YUV to 8b RGB is substantially higher quality than 8b YUV to 8b RGB (especially near white and black), and 10b YUV can decode to higher depth RGB wherever the panel supports it. It sure beats having to deband afterward, or drop in a shedload of noise (and bits) to hide banding. Those with 6b RGB panels will continue to wail, because hardware companies refuse to incorporate simple tricks like dithering on budget hardware until they're promoted as a requirement, but that's can't be helped.
It's true that AV1 and HEVC fixed AVC's rather staggering difference in quality between 8- and 10-bit encoding, but there's still some utility if the toolchain wants to keep quality at the forefront.
Blue_MiSfit
10th December 2019, 03:07
Premium content is NOT >=10-bit 4:2:2. 10-bit 4:2:0 is all that anyone is distributing to consumers via streaming and disc.
Of course. That's what I said (specifically referencing masters being >= 10 bit)
benwaggoner
10th December 2019, 20:29
Of course. That's what I said (specifically referencing masters being >= 10 bit)Ah, my apologies for the confusion, then.
Sent from my SM-T837V using Tapatalk
huhn
10th December 2019, 22:44
There is a significant but shrinking number of devices, mainly handheld, that support 8-bit HEVC but not 10-bit.
More common is devices with 10-bit decoders that then feed into 8-bit compositors, GPUs, or display controllers.
i don't see anything that diminished the benefit of using more bit deep in encoding here.
If nothing else, 8b RGB to 10b YUV to 8b RGB is substantially higher quality than 8b YUV to 8b RGB (especially near white and black), and 10b YUV can decode to higher depth RGB wherever the panel supports it. It sure beats having to deband afterward, or drop in a shedload of noise (and bits) to hide banding. Those with 6b RGB panels will continue to wail, because hardware companies refuse to incorporate simple tricks like dithering on budget hardware until they're promoted as a requirement, but that's can't be helped.
the fact that AMD supports native 6 bit output with there own dithering which helps a lot of sub optimal screens show the benefit of doing this correctly.
the best panel i have ever tested in term of banding is a 6 bit FRC panel the processing just doesn't add banding on such a "bad" device.
and i totally agree using more bits in encoding to delay the dithering until the end device or at least the presentation of a device is a clear benefit.
just reading about 4:2:2 master is letting me wonder if this is just a bad habit that simple didn't stop instead of 16 RGB or higher even the aged BBB has this as a master and what reason could be there not to do that.
Blue_MiSfit
11th December 2019, 01:15
I always figured the 10 bit 4:2:2 thing most likely goes back to SDI / HD-SDI which was traditionally 10 bit 4:2:2. Capturing the full bandwidth of this signal into a file based format in a way that's perceptually lossless was long since considered "good enough" and widely used mezzanine file formats like ProRes do a great job of this - especially relative to the old days of 50-80 Mbps 1080p MPEG-2 service masters.
Given the context of final distribution always being 8 bit 4:2:0 until recently, perceptually lossless 10 bit 4:2:2 was indeed always good enough.
Now that we're doing HDR, 10 bit is mandatory, and 12 or even 16 bit (ProRes 4444 / 4444 XQ or JPEG 2000) is common for source material when encoding for Dolby Vision. Sure, this gets shaped into a 10 bit IPT file but allegedly more of the input can be recovered after Dolby magic.
Studio archival masters are generally 16 bit full range RGB TIFFs or the OpenEXR equivalent (not sure about the specifics of that), and probably DPX for older stuff. Maybe 16 bit lossless JPEG 2000 if the studio is IMF native.
hajj_3
15th December 2019, 16:29
New rav1e build released with 20% speed improvement: https://github.com/xiph/rav1e/releases/tag/p20191215
hajj_3
19th December 2019, 10:25
Rav1e 0.2.0 released, 40-70% faster than than v0.1.0 depending on the encoding settings.
changelog:
Optional serialization/deserialization of the encoding parameters through the feature serialize
Optional cli advanced commands to use it.
The builds are now using the dwarf debug format for the targets that support it, before it was a mixture of dwarf and stabs due to the nasm defaults.
Added a --benchmark hidden flag for the cli for MacOS and Linux.
documentation is now available on docs.rs.
Changes
Segmentation support is now a tunable SpeedSetting and currently it is default off since it can produce desyncs, this does cause a 3% decrease in quality.
Fixes
#1903 - edge-of-frame miscomputation
#1858 - desync on speed 0 and 1 when certain quantizers are selected
Known issues
#1930 - segmentation encoding may cause desync
Mr_Khyron
21st December 2019, 02:01
SVT-AV1 0.8 is out
https://github.com/OpenVisualCloud/SVT-AV1/releases/tag/v0.8.0
x Preset Optimizations
x Single-core execution memory optimization [-lp 1 -lad 0]
x Rate estimation update enhancements
x On / off flags for feature run-time switching
x Auto-max partitioning algorithm support
x Multi-pass partitioning depth support
x Remove deprecated RC mode 1 and shifter RC mode 2 and mode 3 to mode 1 and mode 2 respectively
x Update cost calculation for CDEF Filtering
x Intra-Inter compound for 10-bit
x Eigth-pel optimization
x AVX512 Optimizations
x AVX2 Optimizations
SmilingWolf
24th December 2019, 18:25
Status report!
"Padoru padoru!" edition
1st edition: https://forum.doom9.org/showthread.php?p=1852449#post1852449
2nd edition: https://forum.doom9.org/showthread.php?p=1857587#post1857587
3rd edition: https://forum.doom9.org/showthread.php?p=1860475#post1860475
4th edition: https://forum.doom9.org/showthread.php?p=1871939#post1871939
5th edition: https://www.reddit.com/r/AV1/comments/cfao4x/codecs_performance_report_5th_and_a_half_edition/
Whatever paragraph I don't repeat here can be assumed to be the same as in the aforementioned post
First of all: graphs!
Click to enlarge
Y axis: chosen metric
X axis: bits per pixel
720p:
https://i.ibb.co/jVQm93Q/hvmaf-720.png (https://ibb.co/jVQm93Q) https://i.ibb.co/T0R7ydf/msssim-720.png (https://ibb.co/T0R7ydf) https://i.ibb.co/ysNYgg4/psnrhvsm-720.png (https://ibb.co/ysNYgg4)
1080p:
https://i.ibb.co/gJ5c6Qs/hvmaf-1080.png (https://ibb.co/gJ5c6Qs) https://i.ibb.co/SQkMnXL/msssim-1080.png (https://ibb.co/SQkMnXL) https://i.ibb.co/2vx59Sq/psnrhvsm-1080.png (https://ibb.co/2vx59Sq)
BD rates for 720p:
Codecs ladder: | x264 relative:
x264 -> rav1e | x264 -> rav1e
RATE (%) DSNR (dB) | RATE (%) DSNR (dB)
MSSSIM -23.2215 1.01158 | MSSSIM -23.2215 1.01158
PSNRHVS -24.4172 1.36951 | PSNRHVS -24.4172 1.36951
HVMAF -15.8889 1.38478 | HVMAF -15.8889 1.38478
----------------------------|----------------------------
rav1e -> svtav1 | x264 -> svtav1
RATE (%) DSNR (dB) | RATE (%) DSNR (dB)
MSSSIM -5.56998 0.20639 | MSSSIM -28.4361 1.23111
PSNRHVS -6.41119 0.290801 | PSNRHVS -30.1248 1.65609
HVMAF -15.6008 1.20241 | HVMAF -29.5127 2.45437
----------------------------|----------------------------
svtav1 -> vp9 | x264 -> vp9
RATE (%) DSNR (dB) | RATE (%) DSNR (dB)
MSSSIM 4.20075 -0.151838 | MSSSIM -25.2711 1.10719
PSNRHVS 4.84265 -0.215506 | PSNRHVS -26.5998 1.48648
HVMAF -1.42341 -0.19208 | HVMAF -29.8597 2.33372
----------------------------|----------------------------
vp9 -> x265 | x264 -> x265
RATE (%) DSNR (dB) | RATE (%) DSNR (dB)
MSSSIM -1.31186 0.0508852 | MSSSIM -26.3349 1.18403
PSNRHVS -5.36837 0.2586 | PSNRHVS -30.5296 1.7606
HVMAF -1.84033 0.341593 | HVMAF -30.9904 2.59114
----------------------------|----------------------------
x265 -> av1 | x264 -> av1
RATE (%) DSNR (dB) | RATE (%) DSNR (dB)
MSSSIM -23.2948 0.961133 | MSSSIM -43.3907 2.06939
PSNRHVS -18.5938 0.914589 | PSNRHVS -43.484 2.58526
HVMAF -19.3808 1.25801 | HVMAF -44.4219 3.59302
BD rates for 1080p:
Codecs ladder: | x264 relative:
x264 -> rav1e | x264 -> rav1e
RATE (%) DSNR (dB) | RATE (%) DSNR (dB)
MSSSIM -32.2826 1.19743 | MSSSIM -32.2826 1.19743
PSNRHVS -30.8004 1.43869 | PSNRHVS -30.8004 1.43869
HVMAF -24.0161 1.68074 | HVMAF -24.0161 1.68074
----------------------------|----------------------------
rav1e -> svtav1 | x264 -> svtav1
RATE (%) DSNR (dB) | RATE (%) DSNR (dB)
MSSSIM -7.09315 0.212708 | MSSSIM -37.5209 1.44372
PSNRHVS -6.91838 0.253513 | PSNRHVS -36.0074 1.70477
HVMAF -14.4798 1.05647 | HVMAF -34.513 2.60239
----------------------------|----------------------------
svtav1 -> vp9 | x264 -> vp9
RATE (%) DSNR (dB) | RATE (%) DSNR (dB)
MSSSIM 1.75731 -0.0570674 | MSSSIM -36.6065 1.39502
PSNRHVS 0.61474 -0.0275689 | PSNRHVS -35.7951 1.68578
HVMAF -5.87037 0.0512821 | HVMAF -38.5344 2.64173
----------------------------|----------------------------
vp9 -> x265 | x264 -> x265
RATE (%) DSNR (dB) | RATE (%) DSNR (dB)
MSSSIM 9.95155 -0.292103 | MSSSIM -30.4864 1.15663
PSNRHVS 4.66281 -0.165867 | PSNRHVS -32.9136 1.56443
HVMAF 5.65708 -0.054123 | HVMAF -34.5119 2.53192
----------------------------|----------------------------
x265 -> av1 | x264 -> av1
RATE (%) DSNR (dB) | RATE (%) DSNR (dB)
MSSSIM -31.9542 1.15077 | MSSSIM -52.2509 2.19044
PSNRHVS -26.5316 1.11388 | PSNRHVS -50.3241 2.54395
HVMAF -22.5904 1.34136 | HVMAF -49.1759 3.58479
Encoders:
x264 158-2984-3759fcb
x265 3.2-7-37648fca915b
libvpx-vp9 1.8.2-23-g50d1a4aa7
rav1e 0.2.0-32-g350ac84
SVT-AV1 0.8.0-5303-0952998d
libaom 1.0.0-60-g3e421b069
Cmdlines:
echo -n "16 19 22 26 29 32" | xargs -d " " -n 1 -P 1 -I {} x264 --preset veryslow --tune ssim --crf {} -o test.x264.crf{}.264 orig.i420.y4m
echo -n "16 20 23 27 30 34" | xargs -d " " -n 1 -P 1 -I {} x265 --preset veryslow --tune ssim --crf {} -o test.x265.crf{}.hevc orig.i420.y4m
echo -n "12 21 30 38 47 56" | xargs -d " " -n 1 -P 6 -I {} vpxenc --codec=vp9 --frame-parallel=0 --tile-columns=0 --auto-alt-ref=6 --good --cpu-used=0 --tune=psnr --passes=2 --threads=1 --end-usage=q --cq-level={} --test-decode=fatal --ivf -o test.vp9.cq{}.ivf orig.i420.y4m
echo -n "030 056 082 108 134 160" | xargs -d " " -n1 -P6 -I {} rav1e -o test.rav1e.cq{}.ivf --quantizer {} -s 5 --tune Psychovisual orig.i420.y4m
echo -n "13 21 29 37 45 53" | xargs -d " " -n1 -P1 -I{} SvtAv1EncApp.exe -i orig.i420.yuv -b test.svtav1.cq{}.ivf -w 1280 -h 720 -q {} -scm 0 -enc-mode 8 -fps-num 24000 -fps-denom 1001 -intra-period 47 -output-stat-file fp_stats{}.stat -enc-mode-2p 3
echo -n "13 21 29 37 45 53" | xargs -d " " -n1 -P1 -I{} SvtAv1EncApp.exe -i orig.i420.yuv -b test.svtav1.cq{}.ivf -w 1280 -h 720 -q {} -scm 0 -enc-mode 3 -fps-num 24000 -fps-denom 1001 -intra-period 47 -input-stat-file fp_stats{}.stat
echo -n "12 21 30 38 47 56" | xargs -d " " -n 1 -P 3 -I {} aomenc --frame-parallel=0 --tile-columns=0 --auto-alt-ref=1 --cpu-used=3 --passes=2 --threads=2 --row-mt=1 --end-usage=q --cq-level={} -o test.av1.cq{}.webm orig.i420.y4m
VMAF: model used: vmaf_b_v0.6.3, pooling: harmonic_mean, bagging score (arithmetic mean of 21 models' scores)
Start-to-end times:
x264: 0 hours, 34 minutes and 58 seconds
x265: 4 hours, 13 minutes and 57 seconds
libvpx-vp9: 3 hours, 17 minutes and 52
rav1e: 9 hours, 8 minutes and 25 seconds
SVT-AV1: 5 hours, 6 minutes and 13 seconds
libaom: 7 hours, 26 minutes and 5 seconds
This concludes this report.
As always, I'm open to any kind of feedback to improve my comparisons and my encodes.
foxyshadis
24th December 2019, 22:24
Good to see that VMAF-wise, SVT-AV1 has reached parity with vpxenc and x265, at least for 8-bit, and is only a little slower than x265. That's a pretty big milestone, and the codebase is still in a lot of flux, so there's probably some headroom.
BTW, SVT has supported y4m for a long time now, they just haven't updated their readme.The encoder guide in general is pretty outdated, but I want to see where the holiday cleanup goes before I send pull requests.
Nintendo Maniac 64
25th December 2019, 05:56
Phoronix benchmarked SVT-AV1 v0.8 on a variety of CPUs (i7, i9, Xeon, Ryzen, Threadripper, Epyc):
https://www.phoronix.com/scan.php?page=news_item&px=Intel-SVT-AV1-0.8-Benchmarks
ts1
27th December 2019, 09:55
SmilingWolf, have you made any tweaks to PSNRHVS? AFAIK by default cweight is set to 1, that's too high, set it to 0.125.
Or just use this (https://bugs.chromium.org/p/aomedia/issues/attachmentText?aid=426755) already tweaked and improved version from libaom (only with this version first video must be original and distorted is second).
VincAlastor
29th December 2019, 10:36
SVT-AV1 0.8 is out
https://github.com/OpenVisualCloud/SVT-AV1/releases/tag/v0.8.0
Hello thank you for the update notification.
I have experimented with the SVT-AV1 encoder. Except that the cmd windows in staxrip freezes after 10.000 frames (but the encoding continues), there is a big problem.
I muxed the finished elementary stream in mkv. If I now play it and jump back or forward for more than 5 min, the player (with current LAVfilters) needs a lot of time until the playback continues.
If I download a youtube AV1 video in mp4 container the player doesn't need too much time to skip.
However, mp4box does not want to mux .opus streams, so I cannot really use the mp4 container. Do you have a solution?
ffmpeg -i -nostdin -f rawvideo -pix_fmt yuv420p - | SvtAv1EncApp.exe -i stdin -fps-num 24000 -fps-denom 1001 -n 548203 -w 1920 -h 816 -enc-mode 6 -q 48 -b
edit:
if i mux youtube AV1 mp4 into mkv, i also can skip fast trough the file. So is there a SVT-AV1 bug?
fg118942
29th December 2019, 16:23
Hello thank you for the update notification.
I have experimented with the SVT-AV1 encoder. Except that the cmd windows in staxrip freezes after 10.000 frames (but the encoding continues), there is a big problem.
I muxed the finished elementary stream in mkv. If I now play it and jump back or forward for more than 5 min, the player (with current LAVfilters) needs a lot of time until the playback continues.
If I download a youtube AV1 video in mp4 container the player doesn't need too much time to skip.
However, mp4box does not want to mux .opus streams, so I cannot really use the mp4 container. Do you have a solution?
ffmpeg -i -nostdin -f rawvideo -pix_fmt yuv420p - | SvtAv1EncApp.exe -i stdin -fps-num 24000 -fps-denom 1001 -n 548203 -w 1920 -h 816 -enc-mode 6 -q 48 -b
edit:
if i mux youtube AV1 mp4 into mkv, i also can skip fast trough the file. So is there a SVT-AV1 bug?
Try adding "-irefresh-type 2" to the command line.
VincAlastor
29th December 2019, 22:53
Try adding "-irefresh-type 2" to the command line.
thank you very much. That parameter was very helpful.
But let's hope we don't need it long, because open GOPs are more efficiently pretty sure.
foxyshadis
31st December 2019, 05:42
thank you very much. That parameter was very helpful.
But let's hope we don't need it long, because open GOPs are more efficiently pretty sure.
That's a decoder bug; open GOPs don't need to decode previous GOPs, but dav1d is still pretty new. Anyway, open GOPs only give you noticeable gains if you have a very short GOP. Of course, SVT-AV1 has disabled scene detection and the default (-2) is at as close to 1s as possible anyway.
Just using a longer GOP will help much more.
Beelzebubu
31st December 2019, 14:58
That's a decoder bug [..] but dav1d is still pretty new
No, it's a player or file bug. The player likely selects keyframes from the index (in Mkv parleance: seekhead), and for whatever reason, the muxer didn't add invisible keyframes (non-IDR in AV1) to the index (i.e. broken file), or the demuxer (player) doesn't recognize them and ignores them.
dav1d simply decodes frames in order presented by the player (along with some reordering etc.), it cannot seek further back by itself.
VincAlastor
1st January 2020, 03:00
No, it's a player or file bug. The player likely selects keyframes from the index (in Mkv parleance: seekhead), and for whatever reason, the muxer didn't add invisible keyframes (non-IDR in AV1) to the index (i.e. broken file), or the demuxer (player) doesn't recognize them and ignores them.
dav1d simply decodes frames in order presented by the player (along with some reordering etc.), it cannot seek further back by itself.
Friends and me experimented last week with SVT-AV1 files - more than one, in many players with different versions of filters and mkvtoolnix. The only thing what solve the problem was disabling open GOPs for now. Please try, hope you can show us a solution to use SVT-AV1 with default open GOPs.
Spyros
3rd January 2020, 15:46
LG announced (http://www.lgnewsroom.com/2020/01/lg-to-unveil-2020-real-8k-tv-lineup-featuring-next-gen-ai-processor-at-ces-2020/) that their new 8K TVs will support AV1.
Not only do LG 8K TVs deliver Real 8K, they are also future-proofed to provide customers peace of mind with multiple ways to enjoy the Real 8K experience. The new models offer the capability to play native 8K content thanks to support of the widest selection of 8K content sources from HDMI and USB digital inputs, including codecs such as HEVC, VP9 and AV1, the latter being backed by major streaming providers including YouTube. LG’s 8K TVs will support 8K content streaming at a rapid 60FPS and are certified to deliver 8K 60P over HDMI.
As far as I know these are the first TVs with AV1 hardware decoding. I wonder if they will also support Opus.
birdie
6th January 2020, 10:01
LG announced (http://www.lgnewsroom.com/2020/01/lg-to-unveil-2020-real-8k-tv-lineup-featuring-next-gen-ai-processor-at-ces-2020/) that their new 8K TVs will support AV1.
As far as I know these are the first TVs with AV1 hardware decoding. I wonder if they will also support Opus.
These TVs will feature quite powerful SoCs and Opus is a very low complexity codec, so I see no reason not to support it.
Besides, YouTube started using it years ago along with VP9 and each device which supports YT must support Opus by default.
Spyros
6th January 2020, 14:06
Samsung will also support AV1 (https://news.samsung.com/global/samsung-electronics-unveils-2020-qled-8k-tv-at-ces):
At the same time, the QLED 8K lineup is among the first in the industry to support the playback of native 8K content. In 2020, consumers will be able to enjoy and stream AV1 codec videos filmed in 8K on QLED 8K TVs. All Samsung TVs in the 2020 8K line will ship with this capability built-in.
These TVs will feature quite powerful SoCs and Opus is a very low complexity codec, so I see no reason not to support it.
Besides, YouTube started using it years ago along with VP9 and each device which supports YT must support Opus by default.
That's true, thanks to Youtube (and all the other streaming services that will follow) manufacturers have reason to support both codecs.
My previous message was based on seeing Samsung support VP9 but only Vorbis (https://developer.samsung.com/tv/develop/specifications/media-specifications/2019-tv-video-specifications), not Opus (as of 2019), but I was wrong. This is only for the .webm container, Opus is already supported in .mp4, .mkv etc. I don't know if LG has similar documentation somewhere.
soresu
6th January 2020, 15:20
These TVs will feature quite powerful SoCs and Opus is a very low complexity codec, so I see no reason not to support it.
Besides, YouTube started using it years ago along with VP9 and each device which supports YT must support Opus by default.
I think he may have meant ASIC support?
As you say it's low complexity so it wouldn't require an ASIC to run it in a wall powered device like a high end 4K TV, but anything that reduces thermal output is welcome, it's annoying hearing a fan coming from a TV.
Atak_Snajpera
8th January 2020, 20:23
Does aomenc really support piping from ffmpeg?
my cmd line (simplified)
ffmpeg.exe -loglevel panic -i 1.avs -strict -1 -f yuv4mpegpipe - | aomenc.exe --cq-level=20 --cpu-used=3 --skip=0 --limit=1343 --kf-min-dist=0 --kf-max-dist=240 --output=1.ivf -
No ETA and crash at the end.
https://i.postimg.cc/44PWn9yw/Capture.png
poisondeathray
8th January 2020, 21:41
Does aomenc really support piping from ffmpeg?
2pass definitely works ok for yuv4mpegpipe or rawvideo pipe - so the pipe works
I haven't done 1pass encoding with aomenc, but that suggests something wrong with the 1pass syntax
EDIT: try adding --end-usage=cq --passes=1 , that completes works here for 1pass
utack
9th January 2020, 10:14
Does aomenc really support piping from ffmpeg?
No ETA and crash at the end.
As of recently I have also had it crash at the very end, and the last few frames of the stream never got encoded
Current git version on linux
I will try to reproduce it and see where it started
Atak_Snajpera
9th January 2020, 14:05
2pass definitely works ok for yuv4mpegpipe or rawvideo pipe - so the pipe works
I haven't done 1pass encoding with aomenc, but that suggests something wrong with the 1pass syntax
EDIT: try adding --end-usage=cq --passes=1 , that completes works here for 1pass
Ok thanks. Those two extra switches solve my issue.
benwaggoner
9th January 2020, 19:50
I think he may have meant ASIC support?
As you say it's low complexity so it wouldn't require an ASIC to run it in a wall powered device like a high end 4K TV, but anything that reduces thermal output is welcome, it's annoying hearing a fan coming from a TV.
Vorbis is quite low complexity. The CPUs in SoCs that can do AV1 decode will have ample power to decode Opus in a small fraction of available MIPS.
A bigger challenge can be if the SW audio decoder is integrated into the DRM system, which is sometimes hinky.
That said, xHE-AAC support is growing rapidly and has or will eclipse Opus's. Since it's just a new feature of the AAC porting kit, it'll get deeper integration in many cases.
I'm really impressed that we already have TVs launching with these specs for AV1. I'd love to know how much the SoC had to get bigger for that AV1 support. The extra transistors required for AV1 support is a really important factor in AV1's market viability.
Blue_MiSfit
10th January 2020, 03:54
Do you usually have to encrypt audio? I usually get a pass on that, specifically since the lower security of the audio path mandates multi key (to avoid stealing the audio key and using it for the video) and that's tricky on a lot of platforms.
I want to use Opus today, but the spec for putting it in fMP4 and using it in DASH are not finished (last I checked)! I even whined about this but sounds like there's not a ton of motivation to finish it, which is a real shame.
benwaggoner
10th January 2020, 23:42
Do you usually have to encrypt audio? I usually get a pass on that, specifically since the lower security of the audio path mandates multi key (to avoid stealing the audio key and using it for the video) and that's tricky on a lot of platforms.
The need to encrypt audio is becoming less common overall. But some platforms have had perf issues mixing encrypted and non encrypted media at the same time. And encryption would be required for high quality streaming audio.
I want to use Opus today, but the spec for putting it in fMP4 and using it in DASH are not finished (last I checked)! I even whined about this but sounds like there's not a ton of motivation to finish it, which is a real shame.
I think xHE-AAC has stolen Opus's thunder, since the licensing and code is free for any AAC licensee. Plus it outperforms Opus some.
vBulletin® v3.8.11, Copyright ©2000-2026, vBulletin Solutions Inc.