FranceBB
11th December 2019, 02:52
Last month, the very first movie I encoded premiered at the cinema.
I was very happy with that, especially 'cause it comes after I asked my boss to work on something that wasn't sport or news a year ago.
Today I wanna make a few points about what I learned from this whole experience. Many of you are well known encoders and engineers, but hopefully my words are gonna be useful for some people approaching the encoding world for the very first time.
1) Shooting
Although there are contrasting opinion about this, I strongly suggest you to always shoot in Log with a camera that has a decent amount of stops. It doesn't matter whether you're gonna go to BT709 SDR or BT2020 SDR or BT2100 HDR HLG or PQ, shooting in Log will almost certainly satisfy your needs. The reason why I say this is that you expect camera operators to be expert and know their shit and many of them are indeed expert, but it's very easy to mess things up and if you don't shoot Log, you have virtually no margin to adjust the scene in post-production. The director may or may not like it and, you know, recording a scene again from scratch because of a mistake made by an actor is understandable, but shooting it twice because of a "technical" mistake/imperfection is not, so...
Anyway, shooting Log also allows you to capture as many info as the sensor of your camera can get, which basically means that blacks are not gonna be crushed and whites are not gonna be clipped out, so you can make the best of it. Besides, shooting Log will probably save you from errors like overshooting and so on (I did received slightly overshoot footages sometimes, but thankfully, since it was Log, it didn't look terribly awful once graded).
2) Audio Recording
There are countless things that can go wrong; many movies are recorded with audio from the very top, which is fine, but if you are making a documentary like in my case I would totally go for the microphone on the chest of each and every person interviewed. It doesn't matter if it looks awful, the quality of the audio is gonna be so much better, it's gonna be crystal clear. If people touch it by mistake or move their head up and down while they talk, don't don't don't say "It's fine, I'll fix it" 'cause in the end you'll never do it and it won't be nearly as good as it could have been if it was recorded properly, so make yourself clear and make them understand that it's really important that they don't touch it, that they don't move their head up and down and that they don't rub their chest. Of course, always listen to it while it's recording live, record it lossless AT LEAST in PCM 48'000Hz 24bit and make sure everything is fine by bringing along a monitor like the ones made by Tektronix.
3) Editing
My preferred software for editing is Avid Media Composer, which is also used by many famous companies like HBO, Marvel, Lucasfilm and so on. It's been developed for years and it has a lot of nice tools to edit a movie. Anyway, it's really important that you DON'T edit the masterfile directly; this is because it would be extremely hard as it's probably going to be a huge file. If your camera supports a lossless mode like .dng, use it. I often call .dng files "the elephant". Although you might laugh at it, the reason why I call it like that is that it's huge and impossible to handle. Most lossless recording modes save the file as a sequence of images, like literally every frame is stored individually, which makes them huge and impossible to handle. If you are working with work-spaces (i.e shared network raid storages), you almost definitely have not enough bandwidth to do anything with such a huge file. This is where the encoder comes in and makes a lossy mezzanine file which is meant to be used for editing. You can use whatever you want as a mezzanine file, just make sure the bitrate doesn't skyrockets so I suggest you to have a constant bitrate one. The quality of the mezzanine file doesn't matter, it doesn't have to be perfect because you're gonna re-link to the original lossless file once you have done your editing, so it's literally just a temporary one. Once you finished your edit, enable the dynamic relink option in AVID, re-link to the original lossless file and export the .aaf file with rendered effects if they need to be rendered. An .aaf file is simply going to be a metadata file which will allow the sequence you edited in the timeline to be used by another editing programme like Davinci Resolve. Of course, your secondary editing programme won't have all the effects of the one you used, so some of them will have to be rendered. If you are going to render effects and you can't afford uncompressed lossless, make sure you choose something with a resonably high quality like DNXHR 4:4:4 12bit. I like it 'cause it has a very high bitrate, it's 4:4:4 and 12bit should be enough to avoid any sort of banding. If you're going to use different mezzanine files in your workflow, please please please make sure that they're lossless but if you have to go lossy, it's mandatory that you use the same codecs or very similar codecs (i.e codecs that use the very same transform and have similar parameters). This is because different codecs with different transforms cut off different kind of frequencies, so what you end up with has the worse of one and the other.
4) Color Grading
My de facto choice for this is Davinci Resolve. It's no doubt the best software to do color grading. It's very easy to make masks in Resolve, so don't be afraid to use them to individually grade objects or different parts of the background of different parts of the scene like the body of a person and so on. Don't worry about motion, it has a built-in analyzer which will convert blocks and macroblocks into frequencies, work in the frequency domain by assigning to them a value (literally a number) and make vector-motion calculations for you to keep track of moving objects within your mask. Another great tool is the Face Refinement which works great for many scenarios and it can help you in many occasions. As to the log I talked about, if you wanna go SDR you can go SDR, if you wanna go HDR and your camera has a decent amount of stops, you can go to HDR, but please if you're going to HDR (either HLG or PQ) and you also have to make the SDR version, please don't be lazy, grade the HDR version first and then re-grade the entire movie in SDR. The reason for this is that if the image you shot in Log had let's say a building and the sky, you're gonna be able to keep both actors walking in front of the building and the blue sky with the clouds. The clouds of a sunny day are gonna be around 1000 nits and you're easily gonna retain all the shadows made by the sun of actors walking in front of a building and the details on them. When you grade in SDR though, all you have is 100 nits, so it's worth making a mask, bring the sky down to 100 nits, grade the actors differently and try not to crush the shadows too much. The thing though is that in order to retain the details of the sky and the clouds, those have to be brought down to 100 nits and the tonality of it in BT709 at 100 nits will look very differently from how it looks in BT2100 HDR (either HLG or PQ). This is because if you record the very same scene in BT709 SDR 100 nits directly, you would have the sky completely clipped out, like completely white, no blue, no clouds, just horrible uniform white. This is because you don't have enough nits to record it, so anything beyond 100 nits will just be considered as pure white. When you record it in log, however, you have as many nits as many stops the sensor you are using has, so the sky could easily peak at 1000 nits with all the information perfectly recorded. So... how do you bring it to 100 nits without clipping it out? Well, with a mask, but (and this is the point I'm trying to make) it will look anything but natural because you're "forcing" something to be where it shouldn't, like you're darkening on purpose. The director may or may not like it, so it's very important to grade both the HDR and the SDR version starting from the original Log mastefile using Davinci Resolve with the .aaf metadata file created by Avid Media Composer.
For instance, this is how a sky that peaks at 1000 nits looks at 100 nits SDR BT709: Link (https://i.imgur.com/bEoaESs.jpg)
Do you see something weird? Of course you do! The tonality and the brightness of the sky is completely different from the one you would see with your naked eye. It would be very very bright and almost unwatchable with your naked eye, but here is definitely watchable and almost relaxing to watch 'cause it has been brought down to 100 nits in order to be displayed on SDR monitors, so this is the compromise you have to make and it's something that it's always done in movies.
For instance, this is what you would get if the very same scene is shot in BT709 SDR 100 nits directly (left) or in Log 1000 (or more) nits and then graded to 100 nits:
https://i.imgur.com/zsARwR4.jpg
I'm not gonna say anything about the best practices of grading and so on, just make sure everything is consistent and that you're respecting the Limited TV Range all the time, even if you're working with 12bit precision and output (256-3670) with a proper waveform monitor. In HDR also keep in mind at how many nits you're going to grade and how it's related to the stops of your camera... and please use HDR the proper way, which is NOT to make over-saturated mind blowing super huge mega colors and contrast, but it's to represent scenes with a high range so that it's as close as possible to the natural light viewed by the human eye.
5) Post-processing with Avisynth and final encode
There are many different frameservers, but Avisynth is the best one and it has so many different plugins that it's so useful that it can and must be used on professional productions. For instance, I used it to restore and do the upscale and frame-rate conversion of each and every archive footages for documentaries. Another time I used Avisynth was to do a very light denoise to the final version exported by Davinci which was a DNxHR 4:2:2 12bit. The reason why I applied denoise is because the final file which is gonna be delivered to users like in the BD has a way lower bitrate than the exported final masterfile, so the noise caught by the camera would turn into random crap averaged out by the codec. At the very beginning I mentioned log footages and I gotta say that although log is great, whenever you shoot Log you get a lot of noise, like tons of it. The reason is that you get exactly everything, literally everything that the sensor can get and part of this is noise which isn't supposed to pass through once is debayered and saved without such a hugely wide color curve. Besides, since you gotta encode different versions, Avisynth can very well handle masterfiles and you can use it as an input to encode with different codecs.
5.1) Delivery to UHD BD
Avisynth can be used as an input to x265 which can be used to encode official UHD blurays. Of course you gotta respect the standard and make sure that everything is correctly set, but once you do that, you're ready to go. I gotta say that the reason why I use it is that x265 is better than other professional encoders like Ateme and provides better quality at the same bitrate.
5.2) Delivery to FULL HD BD
Avisynth can be used as an input to x264 which can be used to encode official FULL HD blurays. x264 is very reliable and has a way better grain retention than many professional encoders and it also scores higher with metrics like ssim and psnr (vmaf wouldn't be of much use with such an high bitrate), so it's my de facto choice.
5.3) Delivery to Cinema
The Cinema is a little bit trickier, 'cause theaters use DCP (Digital Cinema Package) which is basically a Motion JPEG2000 4:4:4 12bit planar in the XYZ color space. The tricky part is indeed to bring it to the XYZ color space correctly. There are many different ways of doing this; for instance, you can make your own LUT according to both mathematical calculations and your taste and use the Cube plugin inside Avisynth, or you can use HDRTools inside Avisynth. Please note that even though your final output would be 12bit, it's very strongly suggested that you do the conversion with 16bit planar precision and then you dither it down to 12bit for output. Also please note that Avisynth doesn't officially support XYZ as it's considered an intermediary color space and not a final/delivery color space, so the stream is gonna be flagged as RGB even though it's XYZ. In other words, don't worry if you see colors messed up in like AVSPmod and so on. Just make sure that whatever you use to decode the uncompressed Avisynth input (like ffmpeg) is aware that you're sending out an XYZ and NOT an RGB (even though it's flagged that way). Once you do that, you can use ffmpeg to encode each and every frame as .tiff and then OpenDCP to encode them as JPEG2000 and then mux them together in an .mxf file. Of course OpenDCP has a series of built-in converters that can also take care of the conversion if you make .tiff in - for instance - YUV. Honestly, it really depends on you, your eyes and your taste.
I was very happy with that, especially 'cause it comes after I asked my boss to work on something that wasn't sport or news a year ago.
Today I wanna make a few points about what I learned from this whole experience. Many of you are well known encoders and engineers, but hopefully my words are gonna be useful for some people approaching the encoding world for the very first time.
1) Shooting
Although there are contrasting opinion about this, I strongly suggest you to always shoot in Log with a camera that has a decent amount of stops. It doesn't matter whether you're gonna go to BT709 SDR or BT2020 SDR or BT2100 HDR HLG or PQ, shooting in Log will almost certainly satisfy your needs. The reason why I say this is that you expect camera operators to be expert and know their shit and many of them are indeed expert, but it's very easy to mess things up and if you don't shoot Log, you have virtually no margin to adjust the scene in post-production. The director may or may not like it and, you know, recording a scene again from scratch because of a mistake made by an actor is understandable, but shooting it twice because of a "technical" mistake/imperfection is not, so...
Anyway, shooting Log also allows you to capture as many info as the sensor of your camera can get, which basically means that blacks are not gonna be crushed and whites are not gonna be clipped out, so you can make the best of it. Besides, shooting Log will probably save you from errors like overshooting and so on (I did received slightly overshoot footages sometimes, but thankfully, since it was Log, it didn't look terribly awful once graded).
2) Audio Recording
There are countless things that can go wrong; many movies are recorded with audio from the very top, which is fine, but if you are making a documentary like in my case I would totally go for the microphone on the chest of each and every person interviewed. It doesn't matter if it looks awful, the quality of the audio is gonna be so much better, it's gonna be crystal clear. If people touch it by mistake or move their head up and down while they talk, don't don't don't say "It's fine, I'll fix it" 'cause in the end you'll never do it and it won't be nearly as good as it could have been if it was recorded properly, so make yourself clear and make them understand that it's really important that they don't touch it, that they don't move their head up and down and that they don't rub their chest. Of course, always listen to it while it's recording live, record it lossless AT LEAST in PCM 48'000Hz 24bit and make sure everything is fine by bringing along a monitor like the ones made by Tektronix.
3) Editing
My preferred software for editing is Avid Media Composer, which is also used by many famous companies like HBO, Marvel, Lucasfilm and so on. It's been developed for years and it has a lot of nice tools to edit a movie. Anyway, it's really important that you DON'T edit the masterfile directly; this is because it would be extremely hard as it's probably going to be a huge file. If your camera supports a lossless mode like .dng, use it. I often call .dng files "the elephant". Although you might laugh at it, the reason why I call it like that is that it's huge and impossible to handle. Most lossless recording modes save the file as a sequence of images, like literally every frame is stored individually, which makes them huge and impossible to handle. If you are working with work-spaces (i.e shared network raid storages), you almost definitely have not enough bandwidth to do anything with such a huge file. This is where the encoder comes in and makes a lossy mezzanine file which is meant to be used for editing. You can use whatever you want as a mezzanine file, just make sure the bitrate doesn't skyrockets so I suggest you to have a constant bitrate one. The quality of the mezzanine file doesn't matter, it doesn't have to be perfect because you're gonna re-link to the original lossless file once you have done your editing, so it's literally just a temporary one. Once you finished your edit, enable the dynamic relink option in AVID, re-link to the original lossless file and export the .aaf file with rendered effects if they need to be rendered. An .aaf file is simply going to be a metadata file which will allow the sequence you edited in the timeline to be used by another editing programme like Davinci Resolve. Of course, your secondary editing programme won't have all the effects of the one you used, so some of them will have to be rendered. If you are going to render effects and you can't afford uncompressed lossless, make sure you choose something with a resonably high quality like DNXHR 4:4:4 12bit. I like it 'cause it has a very high bitrate, it's 4:4:4 and 12bit should be enough to avoid any sort of banding. If you're going to use different mezzanine files in your workflow, please please please make sure that they're lossless but if you have to go lossy, it's mandatory that you use the same codecs or very similar codecs (i.e codecs that use the very same transform and have similar parameters). This is because different codecs with different transforms cut off different kind of frequencies, so what you end up with has the worse of one and the other.
4) Color Grading
My de facto choice for this is Davinci Resolve. It's no doubt the best software to do color grading. It's very easy to make masks in Resolve, so don't be afraid to use them to individually grade objects or different parts of the background of different parts of the scene like the body of a person and so on. Don't worry about motion, it has a built-in analyzer which will convert blocks and macroblocks into frequencies, work in the frequency domain by assigning to them a value (literally a number) and make vector-motion calculations for you to keep track of moving objects within your mask. Another great tool is the Face Refinement which works great for many scenarios and it can help you in many occasions. As to the log I talked about, if you wanna go SDR you can go SDR, if you wanna go HDR and your camera has a decent amount of stops, you can go to HDR, but please if you're going to HDR (either HLG or PQ) and you also have to make the SDR version, please don't be lazy, grade the HDR version first and then re-grade the entire movie in SDR. The reason for this is that if the image you shot in Log had let's say a building and the sky, you're gonna be able to keep both actors walking in front of the building and the blue sky with the clouds. The clouds of a sunny day are gonna be around 1000 nits and you're easily gonna retain all the shadows made by the sun of actors walking in front of a building and the details on them. When you grade in SDR though, all you have is 100 nits, so it's worth making a mask, bring the sky down to 100 nits, grade the actors differently and try not to crush the shadows too much. The thing though is that in order to retain the details of the sky and the clouds, those have to be brought down to 100 nits and the tonality of it in BT709 at 100 nits will look very differently from how it looks in BT2100 HDR (either HLG or PQ). This is because if you record the very same scene in BT709 SDR 100 nits directly, you would have the sky completely clipped out, like completely white, no blue, no clouds, just horrible uniform white. This is because you don't have enough nits to record it, so anything beyond 100 nits will just be considered as pure white. When you record it in log, however, you have as many nits as many stops the sensor you are using has, so the sky could easily peak at 1000 nits with all the information perfectly recorded. So... how do you bring it to 100 nits without clipping it out? Well, with a mask, but (and this is the point I'm trying to make) it will look anything but natural because you're "forcing" something to be where it shouldn't, like you're darkening on purpose. The director may or may not like it, so it's very important to grade both the HDR and the SDR version starting from the original Log mastefile using Davinci Resolve with the .aaf metadata file created by Avid Media Composer.
For instance, this is how a sky that peaks at 1000 nits looks at 100 nits SDR BT709: Link (https://i.imgur.com/bEoaESs.jpg)
Do you see something weird? Of course you do! The tonality and the brightness of the sky is completely different from the one you would see with your naked eye. It would be very very bright and almost unwatchable with your naked eye, but here is definitely watchable and almost relaxing to watch 'cause it has been brought down to 100 nits in order to be displayed on SDR monitors, so this is the compromise you have to make and it's something that it's always done in movies.
For instance, this is what you would get if the very same scene is shot in BT709 SDR 100 nits directly (left) or in Log 1000 (or more) nits and then graded to 100 nits:
https://i.imgur.com/zsARwR4.jpg
I'm not gonna say anything about the best practices of grading and so on, just make sure everything is consistent and that you're respecting the Limited TV Range all the time, even if you're working with 12bit precision and output (256-3670) with a proper waveform monitor. In HDR also keep in mind at how many nits you're going to grade and how it's related to the stops of your camera... and please use HDR the proper way, which is NOT to make over-saturated mind blowing super huge mega colors and contrast, but it's to represent scenes with a high range so that it's as close as possible to the natural light viewed by the human eye.
5) Post-processing with Avisynth and final encode
There are many different frameservers, but Avisynth is the best one and it has so many different plugins that it's so useful that it can and must be used on professional productions. For instance, I used it to restore and do the upscale and frame-rate conversion of each and every archive footages for documentaries. Another time I used Avisynth was to do a very light denoise to the final version exported by Davinci which was a DNxHR 4:2:2 12bit. The reason why I applied denoise is because the final file which is gonna be delivered to users like in the BD has a way lower bitrate than the exported final masterfile, so the noise caught by the camera would turn into random crap averaged out by the codec. At the very beginning I mentioned log footages and I gotta say that although log is great, whenever you shoot Log you get a lot of noise, like tons of it. The reason is that you get exactly everything, literally everything that the sensor can get and part of this is noise which isn't supposed to pass through once is debayered and saved without such a hugely wide color curve. Besides, since you gotta encode different versions, Avisynth can very well handle masterfiles and you can use it as an input to encode with different codecs.
5.1) Delivery to UHD BD
Avisynth can be used as an input to x265 which can be used to encode official UHD blurays. Of course you gotta respect the standard and make sure that everything is correctly set, but once you do that, you're ready to go. I gotta say that the reason why I use it is that x265 is better than other professional encoders like Ateme and provides better quality at the same bitrate.
5.2) Delivery to FULL HD BD
Avisynth can be used as an input to x264 which can be used to encode official FULL HD blurays. x264 is very reliable and has a way better grain retention than many professional encoders and it also scores higher with metrics like ssim and psnr (vmaf wouldn't be of much use with such an high bitrate), so it's my de facto choice.
5.3) Delivery to Cinema
The Cinema is a little bit trickier, 'cause theaters use DCP (Digital Cinema Package) which is basically a Motion JPEG2000 4:4:4 12bit planar in the XYZ color space. The tricky part is indeed to bring it to the XYZ color space correctly. There are many different ways of doing this; for instance, you can make your own LUT according to both mathematical calculations and your taste and use the Cube plugin inside Avisynth, or you can use HDRTools inside Avisynth. Please note that even though your final output would be 12bit, it's very strongly suggested that you do the conversion with 16bit planar precision and then you dither it down to 12bit for output. Also please note that Avisynth doesn't officially support XYZ as it's considered an intermediary color space and not a final/delivery color space, so the stream is gonna be flagged as RGB even though it's XYZ. In other words, don't worry if you see colors messed up in like AVSPmod and so on. Just make sure that whatever you use to decode the uncompressed Avisynth input (like ffmpeg) is aware that you're sending out an XYZ and NOT an RGB (even though it's flagged that way). Once you do that, you can use ffmpeg to encode each and every frame as .tiff and then OpenDCP to encode them as JPEG2000 and then mux them together in an .mxf file. Of course OpenDCP has a series of built-in converters that can also take care of the conversion if you make .tiff in - for instance - YUV. Honestly, it really depends on you, your eyes and your taste.