Log in

View Full Version : madVR - high quality video renderer (GPU assisted)


Pages : 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 [520] 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 559 560 561 562 563 564 565 566 567 568 569 570 571 572 573 574 575 576 577 578 579 580 581 582 583 584 585 586 587 588 589 590 591 592 593 594 595 596 597 598 599 600 601 602 603 604 605 606 607 608 609 610 611 612 613 614 615 616 617 618 619 620 621 622 623 624 625 626 627 628 629 630 631 632 633 634 635 636 637 638 639 640 641 642 643 644 645 646 647 648 649 650 651 652 653 654 655 656 657 658 659 660 661 662 663 664 665 666 667 668 669 670 671 672 673 674 675 676 677 678 679 680 681 682 683 684 685 686 687 688 689 690 691 692 693 694 695 696 697 698 699 700 701 702 703 704 705 706 707 708 709 710 711 712 713 714 715 716 717 718 719 720 721 722 723 724 725 726 727 728 729 730 731 732 733 734 735 736 737 738 739 740 741 742 743 744 745 746 747 748 749 750 751 752 753 754 755 756 757 758 759 760 761 762 763 764 765 766 767 768 769 770 771 772 773 774 775 776 777 778 779 780 781 782 783 784 785 786 787 788 789 790 791 792 793 794 795 796 797 798 799 800 801 802 803 804 805 806 807 808 809 810 811 812 813 814 815 816 817 818 819 820 821 822 823 824 825 826 827 828 829 830 831 832 833 834 835 836 837 838 839 840 841 842 843 844 845 846 847 848 849 850 851 852 853 854 855 856 857 858 859 860 861 862 863 864 865 866 867 868 869 870 871 872 873 874 875 876 877 878 879 880 881 882 883 884 885 886 887 888 889 890 891 892 893 894 895 896 897 898 899 900 901 902 903 904 905 906 907 908 909 910 911 912 913 914 915 916 917 918 919 920 921 922 923 924 925 926 927 928 929 930 931 932 933 934 935 936 937 938 939 940 941 942 943 944 945 946 947 948 949 950 951 952 953 954 955 956 957 958 959 960 961 962 963 964 965 966 967 968 969 970 971 972 973 974 975 976 977 978 979 980 981 982 983 984 985 986 987 988 989 990 991 992 993 994 995 996 997 998 999 1000 1001 1002 1003 1004 1005 1006 1007 1008 1009 1010 1011 1012 1013 1014 1015 1016 1017 1018 1019 1020 1021 1022 1023 1024 1025 1026 1027 1028 1029 1030 1031 1032 1033 1034 1035 1036 1037 1038 1039 1040 1041 1042 1043 1044 1045 1046 1047 1048 1049 1050 1051 1052 1053 1054 1055 1056 1057 1058 1059 1060 1061 1062 1063 1064 1065 1066 1067 1068 1069 1070 1071 1072 1073 1074 1075 1076 1077 1078 1079 1080 1081 1082 1083 1084 1085 1086 1087 1088 1089 1090 1091 1092 1093 1094 1095 1096 1097 1098 1099 1100 1101 1102 1103 1104 1105 1106 1107 1108 1109 1110 1111 1112 1113 1114 1115 1116 1117 1118 1119 1120 1121 1122 1123 1124 1125 1126 1127 1128 1129 1130 1131 1132 1133 1134 1135 1136 1137 1138 1139 1140 1141 1142 1143 1144 1145 1146 1147 1148 1149 1150 1151 1152 1153 1154 1155 1156 1157 1158 1159 1160 1161 1162 1163 1164 1165 1166 1167 1168 1169 1170 1171 1172 1173 1174 1175 1176 1177 1178 1179 1180 1181 1182 1183 1184 1185 1186 1187 1188 1189 1190 1191 1192 1193 1194 1195 1196 1197 1198 1199 1200 1201 1202 1203 1204 1205 1206 1207 1208 1209 1210 1211 1212 1213 1214 1215 1216 1217 1218 1219 1220 1221 1222 1223 1224 1225 1226 1227 1228 1229 1230 1231 1232 1233 1234 1235 1236 1237 1238 1239 1240 1241 1242 1243 1244 1245 1246 1247 1248 1249 1250 1251 1252 1253 1254 1255 1256 1257 1258 1259 1260 1261 1262 1263 1264 1265 1266 1267 1268 1269 1270 1271 1272 1273 1274 1275 1276 1277 1278 1279 1280 1281 1282 1283 1284 1285 1286 1287 1288 1289 1290 1291 1292 1293 1294 1295 1296 1297 1298 1299 1300 1301 1302 1303 1304 1305 1306 1307 1308 1309 1310 1311 1312 1313 1314 1315 1316 1317 1318 1319 1320 1321 1322 1323 1324 1325 1326 1327 1328 1329

nevcairiel
15th April 2014, 12:32
madshi: Do you have any plans for using ivtc with 50/59/60 fps sources? It can be a large performance gain, in my case would allow nnedi3 doubling of 720p. I have some clips of 23, 25, 29 and 59 progressive in 720p59 if they are of any use.

Did you try forcing it to IVTC? It should be able to detect any cadence, even if its 6:4 instead of 3:2 due to being 60 fps.
Just toggle the content type with Ctrl-Alt-Shift-T to "Film", and it may just work.

kasper93
15th April 2014, 13:37
What we really need is auto detection when to use film mode. I'm sure madshi will do it when the time comes :)

Tapatalk 4 @ GT-I9300

DragonQ
15th April 2014, 14:14
It's probably quite high on madshi's priority list. Combined with profiles it'd mean settings don't need to manually changed per video any more. :)

madshi
15th April 2014, 14:43
madshi I have noticed something that may or not be an issue. Its not just happening with your latest windowed mode updates, happened before so let me know if you think its worth creating a bug report for. And it actually doesn't seem to negatively effect me at all but you may like to look into it.

I notice sometimes when I start a movie, sometimes the render queue starts at say 15-16/16 but after a few moment drops to 10-11/16 (always exactly this amount), and even if I leave it there for hours with <50% cpu usage, it never grows beyond 10-11/16. The only way I can get it to fill up again is to seek after which it immediately returns to 15-16/16.

I can only seem to reproduce it when actually starting a movie from scratch (I have madVR set to wait for queues to fill up) and when in non fullscreen mode (both FSE or windowed fullscreen don't seem to show it). After seeking I never see it drop below 15-16/16. All other queues arn't effected. Should I report this on the bug tracker incase it does end up causing frame drops on someones machine at some point?
Things like this are very hard to find the cause for. Maybe by looking at very long logs for a very long time I could find something, or maybe not. At the moment, I'd prefer to not touch this, as it would cost a *lot* of time to investigate, and the chance is high that there isn't anything I can do about it, anyway.

The no cadence breaks for a whole episode with new window mode was definitely an exception. Others have been 5-10 down from 15-20 but only usually happen during the black frames between segments and is not noticeable. New window has reduce frame drops from 5-15 to 0-5. Render times have dropped about 10% which gives enough to disable overlay with ivtc+smoothmotion in my case. Secondary display with a different refresh rate then primary now works fine in new window mode, no need for overlay or fse for me anymore.
Sounds good!

Hello everybody, do you think that we can join BlurayDisc "Mastered in 4k" in their full quality? what settings should I set in lav and madvr?
There's a thread on AVSForum where some users analyzed some "Mastered in 4k" discs (by using AviSynth scripts etc) and found that there isn't really any xvYCC data worth talking about. At this point in time I think it's mostly a marketing gimmick by Sony and nothing else.

HTPCs cannot output untouched YCbCr, there is generally always a RGB step in between, so its doubtful this is ever going to work.
True. However, if the display is calibrated to a bigger than BT.709 gamut (e.g. DCI) and if madVR is configured accordingly, the extended colors might still be transported to the display just fine. I'm not really 100% sure, though. The whole gamut topic will get interesting when we get the first 4K Blu-Rays with hopefully a bigger native gamut. At that point I guess I'll have to investigate how to transport that data to the display properly.

I thought xvYCC used negative values (or are values below 16 considered to be "negative"?) to expand the gamut beyond the normal range.

So:
+100% green
−100% red
−100% blue
(i.e. −100% magenta, being the opposite of green)

Would be equal to 200% green.

Which should be possible to represent in RGB if you are doing the appropriate color management and have a display capable of displaying that range of saturation.

Displays calibrated to Adobe RGB are common with photo editing for example.
The encoding is done in YCbCr, but yes, RGB values after YCbCr -> RGB conversion could become negative, in relation to the BT.709 gamut. After conversion to a bigger gamut they might get back into positive range. But I'm not sure whether all of this works correctly in madVR atm, to be honest. If the colors are just slightly outside the valid range, then there should be no problem because madVR has some headroom due to doing all math in video levels. But if the colors are waaay outside of BT.709, then I'm not sure if madVR isn't maybe clipping them, not sure...

I haven't had much time to test, but I think I found a situation where new windowed fullscreen is worse than old windowed fullscreen. It looks like the new version has slightly lower rendering times, but it starts dropping frames "earlier" (at lower frame times) than the old windowed path.

Try cranking up enough features to have average rendering time close to 1/fps. e.g. on 23.976 fps content, around 40 ms. For me, new path drops frames for the same settings that old path doesn't. (With things turned down more reasonably so we're not skirting the 90%+ render times and the GPU barely keeping up with the source, new path works perfectly.)
Yeah, that's quite possible. That will be especially true if your display refresh rate is higher than the movie framerate.

I tried a whole bunch of configs/reformat/updates/official-beta drivers.. giving up...

No go with NED 2x luma on 1920x1080 to 2560x1600 ...

Tried both my 7970cfx and 7870xt downstairs..

/Cry really loudly

I'm willing to buy a 290x.... Anyone try NED 2xluma @ 19x10 to 25x16.
You could try the interop test builds (see next post) to see if they make a difference for you. But at this point I guess your best bet might be an NVidia 750ti, because it seems that with your mainboard the AMD OpenCL interop cost is too high. Alternatively you could replace your mainboard, but I think going over to green land would be the easier solution.

In 87.7 and 87.9 MadVR does not switch into FSE correctly on the first attempt for me. Alt+Enter to go back to windowed mode, and then another time to go to full screen again works. 87.4 worked okay on the first try. I haven't tried any of the intermediate releases.

I'm using Win 7 x64 with MPC-HC 1.7.3 using the Intel HD4600 graphics in my i7-4770k.
Can you please try to find out which exact madVR build introduced this problem? Also a debug log might help figuring out why the switch fails. Please don't switch back and forth in the debug log. Just let it fail, then stop, I don't need to see it working in the log, I only need to see the fail. Please enable the OSD (Ctrl+J) while creating the debug log, because otherwise the log will not contain all important information.

Aliased?

That's irrelevant when you get 1:1 scaling.. from 1280x720 to 2560x1600(1440)

That is as much information as is in the file.. anything ONTOP of that is an approximation.:p
You're aware of that almost all 1280x720 files were downscaled from 2K or 4K masters? The downscaling is usually done with linear interpolation. Something like Lanczos. The key thing to understand here is that a video file like that does not describe rectangular pixels. You need to think of each pixel more of a circle. If you display these files in 2x resolution with Nearest Neighbor scaling you're throwing away potential image quality. Why? Because what you're doing is this:

4K master -> Blu-Ray downscale -> 720p downscale -> 1440p upscale

All the downscales were done with linear interpolation. Which means that each pixel also contains a small portion of the original neighboring pixels. The best way to watch such movies is to upscale them with a good upscaling algorithm. This will get you nearer to the way the image looked in its original resolution.

If you don't believe me, try this:

(1) Take a sharp and detailed photo.
(2) Downscale it with your favorite image editor to 50%, by using a good downscaling algorithm (e.g. Cubic or Lanczos).
(3) Upscale it again 200% to get back to the original resolution.

Now for step (3) try Nearest Neighbor scaling and compare it to e.g. Lanczos. Check which upscaled image looks nearer to the original photo. This test is very valid for video playback, too. After all you don't just want to see what is in the video file, you want to see an image which is as near to the original film scan as possible, don't you?

720p to 25x16 or 15x14 using NN is akin to LOSSLESS conversion
Lossless to the 720p file, yes. But the 720p file itself has a much lower resolution and quality compared to the original film negative. By using a good upscaling algorithm you would get nearer to the original film negative. Do you want to stay lossless to the 720p downscaled source? Or do you want to get as near to the original film negative as possible?

I'd really like to see some more development put into NNEDI to attempt to improve it, I wonder the chances of it improving in future?
NNEDI3 is what it is. There's no way for *me* to improve it. I could just post-process it (e.g. sharpen it), or alternatively I could try to create a new algorithm from the ground up, but I'm not sure if I could even reach NNEDI3 quality with such a new algorithm. If you want the NNEDI3 algorithm itself to be improved, you'd have to talk to tritical who created NNEDI3 in the first place. But I don't think you can expect big improvements.

Basically it sounds like someone needs to create an external Darby-like neural network upscaler.
Darbee is a sharpener, not an upscaler, and Darbee doesn't use a neural network. If you want sharper images, use a sharpening algorithm after NNEDI3 scaling.

Madshi, how does madVR's NNEDI settings stack up vs the default settings in Avisynth filter? Was wondering if we'd see options available for nsize and qual etc.
I've done some image quality comparisons and found an nsize setting of "8x4" to produce the least amount of artifacts for image doubling. It happens to also be the fastest nsize setting available. So the choice was simple for me. I don't plan to offer nsize or qual options in madVR, because the performance cost would not be worth the quality gain. The best way to improve quality is to increase the neuron count, so that's the only option I'm offering. No plans to change that.

1080p: 16 neurons (Not with NNEDI3 ChromaUpscaling, and it's not used often.)
720p < 24fps: 64 neurons
720p > 24fps: 32 neurons
720p > 30fps: 32 neurons
SD < 24 fps: 128 neurons
SD > 24 fps: 128 neurons
SD > 30 fps: 32 neurons
Nice. Looks faster than my AMD7770, although I haven't compared in detail.

Are you using PCI-E 2.0 or 3.0?
PCIe version is only important for AMD users.

madshi: Do you have any plans for using ivtc with 50/59/60 fps sources? It can be a large performance gain, in my case would allow nnedi3 doubling of 720p. I have some clips of 23, 25, 29 and 59 progressive in 720p59 if they are of any use.
It's on my to do list, but not for soon.

Did you try forcing it to IVTC? It should be able to detect any cadence, even if its 6:4 instead of 3:2 due to being 60 fps.
Just toggle the content type with Ctrl-Alt-Shift-T to "Film", and it may just work.
Hmmmm... You're right, it does seem to work. At least it lists 6:4 and plays just fine in 60Hz. Not sure whether the IVTC decimation timestamp manipulations will work properly, though. I guess at 24Hz it would probably play fine. But playing this at 60Hz with Smooth Motion FRC turned on might fail to achieve smooth motion.

What we really need is auto detection when to use film mode.
Yes, we do need that. Unfortunately it's not that easy to implement properly. Especially if we want to take mixed sources (e.g. film content with video overlay) into account.

@madshi:
Blaire linked to your recent workaround (your changelog for 0.87.9) for the NV driver issue and asked me via PM, if there still is a driver fix needed.

I kinda feared this would happen, since your woraround takes the pressure off of NV to fix an issue no one else (besides you and people that use NV hardware with madVR) seems to care about (that's how I interpret it).

It looks to me that they were in the process of working on the fix, but they re-checked if they have the newest madVR version to test against.

Now, what should I tell him? Some technical details would probably be helpful. Also how you (if memory serves right, a madVR user actually came up with the idea) worked around the bug, so NV knows where and what to search for.
To be honest, madVR doesn't need a fix, anymore. The workaround works fine and doesn't have any negative side effects. That said, it's a clear bug in the NVidia drivers, from what I can see, so they might still want to fix it.

Basically the old madVR builds did this:

for each video frame do
{
clTargetTex = clCreateFromD3D9TextureNV(...);
clEnqueueAcquireD3D9ObjectsNV(clTargetTex);
clSetKernelArg(clTargetTex);
clEnqueueNDRangeKernel(...);
clEnqueueReleaseD3D9ObjectsNV(clTargetTex);
clFinish(...);
clReleaseMemObject(clTargetTex);
}
With this code, older NVidia drivers worked fine, but newer NVidia drivers either do nothing, or write zeroed out data to the target texture.

The latest madVR builds now use the following approach instead, which works around the issue:

clTargetTex = clCreateFromD3D9TextureNV(...);
for each video frame do
{
clEnqueueAcquireD3D9ObjectsNV(clTargetTex);
clSetKernelArg(clTargetTex);
clEnqueueNDRangeKernel(...);
clEnqueueReleaseD3D9ObjectsNV(clTargetTex);
clFinish(...);
}
clReleaseMemObject(clTargetTex);

madshi
15th April 2014, 14:50
Here's a new test build set for AMD users wanting to do NNEDI3:

http://madshi.net/madVRinteropTest.rar

In the rar file are two madVR.ax files which use different methods to try to improve the interop problem. Unfortunately the improvement is probably not as large as I had hoped, but there should be a small improvement at least. Probably one build will work better than the other build. Please try both and let me know which build works better for you. I've intentionally removed the rendering times from the OSD (only for these test builds, of course) because due to the way these 2 test builds work, judging them by looking at the rendering times would be misleading. So please judge these builds by testing which build allows you to use higher/more quality settings.

Looking forward to your feedback!

(FWIW, I've concentrated on NNEDI3 luma doubling, with NNEDI3 chroma upscaling and NNEDI3 chroma doubling disabled. Enabling those might still work, but I've not tested that.)

James Freeman
15th April 2014, 15:11
Is there a problem with NNEDI3 and AMD?
Not long ago it was Nvidia that didn't work at all, now its AMD?

michkrol
15th April 2014, 15:21
Is there a problem with NNEDI3 and AMD?
Not long ago it was Nvidia that didn't work at all, now its AMD?

On AMD it's a performance only problem - it works correctly, just slower than it should, because of the way AMD('s driver) goes around DX->OpenCL interop.

DragonQ
15th April 2014, 15:31
Here's a new test build set for AMD users wanting to do NNEDI3:

http://madshi.net/madVRinteropTest.rar

In the rar file are two madVR.ax files which use different methods to try to improve the interop problem. Unfortunately the improvement is probably not as large as I had hoped, but there should be a small improvement at least. Probably one build will work better than the other build. Please try both and let me know which build works better for you. I've intentionally removed the rendering times from the OSD (only for these test builds, of course) because due to the way these 2 test builds work, judging them by looking at the rendering times would be misleading. So please judge these builds by testing which build allows you to use higher/more quality settings.

Looking forward to your feedback!

(FWIW, I've concentrated on NNEDI3 luma doubling, with NNEDI3 chroma upscaling and NNEDI3 chroma doubling disabled. Enabling those might still work, but I've not tested that.)
Whilst playing a 640x480p/25 file with 16 neurons and Smooth Motion enabled (60 Hz):

v0.87.9: 35-40 dropped frames per second; render queue is 1-2/8; present queue is 0-1/8; GPU load ~95%
Test 1: 1-2 dropped frames per second; render & present queues are 0-4/8 or 1-5/8 typically; GPU load ~80%
Test 2: 0 dropped frames per second; render & present queues are 4-7/8 or 5-8/8 typically; GPU load ~82%

Test 2 seems the best for me. Still can't use 32 neurons though, I get a dropped frame every few seconds and GPU usage rises to 89%.

MS-DOS
15th April 2014, 15:54
Here's a new test build set for AMD users wanting to do NNEDI3:

http://madshi.net/madVRinteropTest.rar

Let's see. On my 5870 (Win 7 x64, 13.12) the interop cost was insane, as I posted here (http://forum.doom9.org/showthread.php?p=1673786#post1673786) (the image is dead, argh).
Tested the new builds on 480 -> 1080 (+J3AR) content in FSE (new path), which gave me about ~8-10 dropped frames per second even with 16 neurons before.

TestBuild1 - Seems to work smoothly up to 64 neurons, 128 starts to give loads of presentations glitches and the playback stutters quite a lot, but it doesn't report any dropped frames, thou. GPU load is stuck at ~63%.
TestBuild2 - Seems smooth up to 128 (!) neurons with no dropped frames or presentation glitches, ~64% GPU load. Setting it to 256 neurons puts 99% load on the GPU and I'm starting to get frame drops.

The improvement overall looks very large to me, TB2 is a beast. Could you implement these two in your OpenCL benchmark? I'd really like to see the raw numbers :D

Great work!

huhn
15th April 2014, 16:24
Hmmmm... You're right, it does seem to work. At least it lists 6:4 and plays just fine in 60Hz. Not sure whether the IVTC decimation timestamp manipulations will work properly, though. I guess at 24Hz it would probably play fine. But playing this at 60Hz with Smooth Motion FRC turned on might fail to achieve smooth motion.

IVTC with something else like 3:2 normally never works fine. madvr doesn't drop the right frame with right detected 4:2:2:2 and playback is unwatchable and this on a 23 hz tv.

@tesbuilds

for me on a r9 270 the build 1 is "faster"
i tested 256 neuron 480p23 to 1080p. with the old build it is impossible with both new builds it works but with test 2 all queue drop but no frame is dropped. with test1 all queue fill up after some time so i think this is working better.


i get 82 % gpu usage test1 and 84% with test2 both drop like crazy with opend gpu-z so they should't be judge with gpu-z

TheLion
15th April 2014, 17:05
Let's see. On my 5870 (Win 7 x64, 13.12) the interop cost was insane, as I posted here (http://forum.doom9.org/showthread.php?p=1673786#post1673786) (the image is dead, argh).
Tested the new builds on 480 -> 1080 (+J3AR) content in FSE (new path), which gave me about ~8-10 dropped frames per second even with 16 neurons before.

TestBuild1 - Seems to work smoothly up to 64 neurons, 128 starts to give loads of presentations glitches and the playback stutters quite a lot, but it doesn't report any dropped frames, thou. GPU load is stuck at ~63%.
TestBuild2 - Seems smooth up to 128 (!) neurons with no dropped frames or presentation glitches, ~64% GPU load. Setting it to 256 neurons puts 99% load on the GPU and I'm starting to get frame drops.

The improvement overall looks very large to me, TB2 is a beast. Could you implement these two in your OpenCL benchmark? I'd really like to see the raw numbers :D

Great work!

This is very exciting news indeed. My 5870 prevented me from using NNEDI3 at all. I will try these test builds as soon as I can - here is hope that at least chroma upsampling for 1080p is now possible, as well as SD->1080p.

tFWo
15th April 2014, 17:14
Same as @huhn with my 270x.

Build1 is slightly better than build2. Slightly lower gpu load and (maybe) faster queue filling.

Both builds allow much higher NNEDI settings than 87.9. :)

720p24->1680x1050@60

87.9 using both chroma upscaling 32neurons and luma doubling 32neurons was just below the treshold for smooth playback (40.5ms)

new builds allow 64 neurons on both settings (around 85% load)

SD@24->1680x1050@60

87.9 128 neurons was usable on both

new builds allow 256 on both or 128 on both + 32quad for luma (also around 85%)

aminfri
15th April 2014, 17:20
About time i reported some stats too:

Using the latest test builds with Hi10 720p to 1080 and these settings:

Jinc 3 AA, Chroma upscaling,
jinc 3 AA, Image upscaling,
Catmull-Rom AA SLL, image downscaling,
Smooth Motion enabled,
Dithering, Error Diffusion 1,

I could easily get 64 Neurons with both test builds, but the usage with the first build (76%) was just a bit lower that the build 2 (78%). Previously i couldn't enable Image doubling without frame drops. So these builds are definitely huge improvements.

On 128 Neurons i was getting frame drops left and right with both builds.

System specs in sig.

seiyafan
15th April 2014, 17:32
Here's a new test build set for AMD users wanting to do NNEDI3:

http://madshi.net/madVRinteropTest.rar



How do I use it? Just paste into MadVR folder?

leeperry
15th April 2014, 17:36
How do I use it? Just paste into MadVR folder?
Backup your existing mVR folder, then copy all the files in there and alternatively rename both builds to madVR.ax

w00t, moar testing :)

I suppose that implementing those changes in the test app woulda been too much work but it didn't work on my box anyway.

Was kinda looking for a reason to avoid going green, let's see how that goes :p

Farfie
15th April 2014, 17:37
On my Win7 x64 HD5850 machine, TB1 is a very clear winner going from 720p -> 1440p. I'm able to use 64 neurons now, which is very close to my GTX680. With TB2, I get about 1 frame drop per "OSD refresh tick," and with the original I get anywhere between 3-5 per. Of course, this is at an overclock of 800mhz for the core (above 725 default), so any AMD user might want to push for this, since this was enough to get 64 neurons for luma doubling at this resolution with TB1 :)

I don't know why my results differ from DragonQ and MS-DOS with TB1 being better than TB2 very clearly. Perhaps it has to do with the resolution size. I speak for nothing though, as madshi will probably know why :)

seiyafan
15th April 2014, 17:48
Great work Madshi! 1080->1440 Before it's dropping 10-15 frames a second, now 0!

Now a question, for movies which of the following provides more visual improvement? debanding or ED?

huhn
15th April 2014, 17:51
Great work Madshi!

Now a question, for movies which of the following provides more visual improvement? debanding or ED?

if needed debanding for sure.

TheLion
15th April 2014, 17:59
Great work, madshi!

On my Win 8.1 64bit i7 system with AMD 5870 (latest beta Catalyst) both test builds show huge improvements for NNEDI3 (doubling as well as chroma upscaling).

Now I can finally use it at all - the limits to the max settings are the same for both builds. testbuild2 seems to die more gracefully when "overloaded": TB1 shows massive amounts of repeated frames in addition to the dropped.

chroma upscaling for 1080p works now up to 32 neurons - it wasn't fast enough before at all.

noee
15th April 2014, 18:11
Win7 x64, HD6570 PCI-E 2.0x16

24p 720x368 P010 (OrderedDith NNEDI luma doubling/32n/SMFRC off) => {1080 playback@59.942Hz}

879: GPU ~95+%, Render(1-4/14) - Present(5-7/10), occasional frame drop
OP1: GPU ~89+%, Render(12-14/14) - Present(7-10/10), no drops
OP2: GPU ~85+%, Render(12-14/14)/Present(8-10/10), no drops

seiyafan
15th April 2014, 18:14
if needed debanding for sure.

what if the video quality is high, like blu-ray? Would it still benefit more from debanding than dithering?

MS-DOS
15th April 2014, 18:21
I hope it's just some kind of a bug with TB1, which causes constant presentation glitches to me when GPU load is above a certain value, and can be fixed. Because, like for most people posted above, to me TB1 has slightly lower GPU cost than TB2.
I tested with SM disabled, ordered dithering, and SC80 chroma upscaling, all Q4P disabled, except subtitles optimization and "don't render frames when fade in/out detected".

James Freeman
15th April 2014, 18:24
what if the video quality is high, like blu-ray? Would it still benefit more from debanding than dithering?

When you are at the edge of the big dilemma of "Visible Quality" vs "Machine Power", I suggest to pick the one which is more visible or beneficial for the picture quality.
In that case, go for Debanding instead of a heavier and almost invisible (imo) dithering algorithm.

Same goes true for NNEDI3 vs Smooth Motion for example.
Judder free playback outweighs slight improvement in scaling aliasing a hundredfold.

baii
15th April 2014, 18:36
Also factor in fan noise when you push the gpu hard. Especially in a laptop set up.

leeperry
15th April 2014, 18:38
No night/day difference between both builds on my 7850/Haswell rig, both run quite a bit faster than 0.87.9. If anything the first picture of a movie shows up faster with the first build, GPU memory and D3D usage are also the lowest. I vote 1 :)

64x nnedi for chroma & luma 29.97 960x540@1080p:

1: http://thumbnails110.imagebam.com/32102/695090321017680.jpg (http://www.imagebam.com/image/695090321017680) 2: http://thumbnails112.imagebam.com/32102/502f8d321017681.jpg (http://www.imagebam.com/image/502f8d321017681) 0.87.9: http://thumbnails111.imagebam.com/32102/4ea789321017682.jpg (http://www.imagebam.com/image/4ea789321017682)

128x nnedi for chroma & luma 25fps 640x480@1080p:

1: http://thumbnails111.imagebam.com/32102/33658a321017715.jpg (http://www.imagebam.com/image/33658a321017715) 2: http://thumbnails109.imagebam.com/32102/883541321017713.jpg (http://www.imagebam.com/image/883541321017713) 0.87.9: http://thumbnails109.imagebam.com/32102/833901321017714.jpg (http://www.imagebam.com/image/833901321017714)

flashmozzg
15th April 2014, 18:57
No night/day difference between both builds on my 7850/Haswell rig, both run quite a bit faster than 0.87.9. If anything the first picture of a movie shows up faster with the first build, GPU memory and D3D usage are also the lowest. I vote 1 :)

Try without HW monitoring tools.

kasper93
15th April 2014, 19:31
Here's a new test build set for AMD users wanting to do NNEDI3:

Good work. Seems to be a lot faster. build2 is better for me. I can do 32 neurons on 720p->1080p while with build1 it drops frames even with 16 neurons. So my vote is for build 2 :) Comparing to stable this is BIG improvement.


Still we should somehow reach AMD and made them fix that ;/


EDIT:

Build2 is memory hungry, 3.6GB of "system commit" was freed after closing player. I needed to close some programs, because I got only 6GB RAM, 3GB pagefile, 1GB of gpu mem which is full, but I used to it already ;p Windows notified me during playback that I run out of memory. But I had already around 5GB used.

iSunrise
15th April 2014, 19:58
...To be honest, madVR doesn't need a fix, anymore. The workaround works fine and doesn't have any negative side effects. That said, it's a clear bug in the NVidia drivers, from what I can see, so they might still want to fix it.

Basically the old madVR builds did this:
...

Thanks. I just forwarded everything to Blaire. It's their decision now.

turbojet
15th April 2014, 20:39
Forcing ivtc with deint=ivtc on film in 59 fps source detects 6:4 cadence but doesn't remove frames and gpu load remains high.

29i works fine with force film mode now, it didn't last I checked months ago. Unfortunately 59i doesn't even when double framerate deinterlacing allowed it still detects 2:2 and plays at 29 fps.

leeperry
15th April 2014, 21:09
Try without HW monitoring tools.
I initially did, reason why I thought hard figures would be more meaningful.

ThurstonX
15th April 2014, 21:25
Finally found some time to do a few quick tests. AMD Radeon R7 200 Series; passively cooled (SAPPHIRE Ultimate 100368USR Radeon R7 250 1GB 128-Bit GDDR5 PCI Express 3.0); Catalyst 14.2; Core i5-3470; 8 GB RAM
Display is an old Sharp Aquos LC-32GA5U running at native 1366x768 via DVI

tl;dr
v0.87.9 couldn't run without dropping frames; Test1 ran with Luma doubling forced at 16 neurons (32 was too much) and Jinc 3 AR; Test2 could only handle Lanczos 3 AR, so I vote for Test1.

Hope this helps, and thanks for the test builds. Definitely a step in the right direction!


I started with v0.87.9 trying to force NNEDI3 to double Luma resolution using 16 neurons. Plenty of dropped frames.

Settings
Chroma upscaling: Bicubic 75 (No AR)
Image doubling: use NNEDI3 to double Luma; always - if upscaling is needed; 16 neurons
Image upscaling: Jinc 3 AR
Image downscaling: Catmull-Rom scale in linear light
No Debanding
Smooth motion: Enable, only if there would be motion judder without it...
Dithering: Ordered; use colored noise; change dither for every frame
Trade quality for performance: first five items checked
Exclusive mode settings at default

With v0.87.9 I got the following:
Queues
Decoder: 13-16/16
Upload: 6-8/8
Deinterlace: 5-8/8
Render: 2-4/8
Present: 0-2/8
Tons of dropped frames

with Test1
Decoder: 14-16/16
Upload: 7-8/8
Deinterlace: 6-8/8
Render: varied from 5-7/8; 6-7/8; 6-8/8
Present: varied from 4-5/8; 4-6/8
1 frame repeat every 3.63 secs
NO dropped frames

Source video (a VHS capture using an old Hauppauge card)

Format : MPEG-PS
File size : 8.90 GiB
Duration : 1h 40mn
Overall bit rate : 12.7 Mbps

Video
ID : 224 (0xE0)
Format : MPEG Video
Format version : Version 2
Format profile : Main@Main
Format settings, BVOP : Yes
Format settings, Matrix : Custom
Format settings, GOP : M=3, N=15
Duration : 1h 40mn
Bit rate : 12.0 Mbps
Width : 720 pixels
Height : 480 pixels
Display aspect ratio : 4:3
Frame rate : 29.970 fps
Standard : NTSC
Color space : YUV
Chroma subsampling : 4:2:0
Bit depth : 8 bits
Scan type : Interlaced
Scan order : Top Field First
Compression mode : Lossy
Bits/(Pixel*Frame) : 1.159
Time code of first frame : 00:00:00:00
Time code source : Group of pictures header
Stream size : 8.46 GiB (95%)

Audio
ID : 192 (0xC0)
Format : MPEG Audio
Format version : Version 1
Format profile : Layer 2
Duration : 1h 40mn
Bit rate mode : Constant
Bit rate : 384 Kbps
Channel(s) : 2 channels
Sampling rate : 48.0 KHz
Compression mode : Lossy
Delay relative to video : -111ms
Stream size : 275 MiB (3%)

~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~

Been playing an NTSC DVD (Peter Gabriel Live in Athens 1987). Same resolution, interlaced, but dropped frames using Jinc 3 AR. OK switching to Lanczos 3 AR. Didn't check any other differences between the two sources.

2nd edit:
I think the difference is that the DVD is really 16:9, while the VHS capture is 4:3. That's my guess, anyway.

Asmodian
15th April 2014, 21:56
what if the video quality is high, like blu-ray? Would it still benefit more from debanding than dithering?

ED is a large performance hit and ordered dither is quite good, if you have banding debanding + OD is better than no debanding + ED. debanding + no dither is not a reasonable option.

leeperry
15th April 2014, 22:00
I believe this was already mentioned but 48/96/192 neurons NNEDI would be pretty cool for when you can't quite run for 64/128x and still got potential cycles unused. I still kinda find NNEDI too sharp for chroma but sometimes I still don't have enough horse power to go for the next level and I can't run 256 neurons luma for SD@1080p, I would welcome the opportunity to try 192 if technically doable :)

It's even more true now that AMD boards have magically earned extra headroom. :thanks:

Fullmetal Encoder
15th April 2014, 22:13
I'd love to test the new build but I keep getting "dxva processing failed" message from madVR. I don't have any of those options selected (for dxva) in the UI and I'm using Radeon 5850 on Windows 7 Pro. I would add that this is with ED 2 selected.

kasper93
15th April 2014, 22:20
@Fullmetal Encoder: Make sure to disable all "enhancements" in gpu driver. Especially "dynamic contrast" there seem to be bug in drivers which breaks DXVA processing, for example deinterlacing will fail to initialize when opening, yet you can re-enable it later.

QBhd
15th April 2014, 22:45
Time for my report on the test builds.

System:
R9 270X factory OC to 1120/1400 (GPU/memory)
PCI-e 2.0 (GA-990FXA-UD7)
Windows 8.1

Target resolution:
1024x768 (rectangular pixels)

Source:
1280x720p24

Settings:
Chroma Upscaling - NNEDI3 x32
Image Upscaling - Jinc 3 AR
Luma Doubling - NNEDI3 x64
Chroma Doubling - NNEDI3 x32
Image Downscaling - Catmull-Rom AR LL
Debanding - med/med

With previous release I could only do Ordered Dithering (Error Diffusion just pushed over the limit of GPU)

Testbuild1 - Still dropped frames with ED
Testbuild2 - NO dropped frames with ED

So my vote goes to Testbuild2... It allows for me to go even further than any build to date

QB

Fullmetal Encoder
16th April 2014, 00:06
Scaling from 720x480 to 1920x1200 with ED2 and NNEDI doubling at 32 neurons on luma using a Radeon 5850 I am getting:

- 52 frame drops/refresh with 87.9
- 40 frame drops/refresh with test 1
- 31 frame drops/refresh with test 2

I don't know why others with the 5850 are getting so much better performance though.

Fullmetal Encoder
16th April 2014, 00:08
@Fullmetal Encoder: Make sure to disable all "enhancements" in gpu driver. Especially "dynamic contrast" there seem to be bug in drivers which breaks DXVA processing, for example deinterlacing will fail to initialize when opening, yet you can re-enable it later.

Thank you very much! I don't know how, but all of those "enhancements" were on in CCC. Although I'm not sure how they got turned on since I turned them off long ago :o

tickled_pink
16th April 2014, 00:50
Win7 x64, HD 7750 PCI-E 2.0x16

720x404@25fps with 64 neurons tested

Test build 1 uses slightly less (55% vs 57%) GPU and less graphics memory (~40 MB or 10%) than test build 2.

0.87.9 used ~60% GPU and similar amount of memory as test build 1.

Neither improved performance enough to allow more neurons but a 10% overall improvement is certainly welcome!

sajara
16th April 2014, 01:23
This test came a bit as a shock because I do remember being unable to use NNEDI3 even with 16 neurons when first release and didn't bothered to try again.

AMD 5730M 650Mhz core /800Mhz GDDR3 mem

H264 clip 720x304 -> 1366x768

87.9 - 16 Neurons ~86.7% / 32 Neurons - slideshow
Test 1 - 16 Neurons ~57.6% / 32 Neurons ~86.5%
Test 2 - 16 Neurons ~60% / 32 Neurons ~89.7%

Queues the same in test 1 and 2.

So again beyond words on the improvement and as much, amazed being able to do 32 neurons.

ryrynz
16th April 2014, 01:49
Do the test builds improve anything on Nvidia hardware at all?

Procrastinating
16th April 2014, 06:40
Considering the previous responses, and the less meaningful low render times, I changed the survey defaults to double luma, and added a default video of tears of steel. Remember that any data is good data, and this survey/spreadsheet will not only help madshi, but those interested in what the optimal media GPU for them might be.

To Madshi in particular, I think it will be interesting to see across the various hardware configurations, how the improvements appear between versions. I will probably try the AMD builds at some point.

Asmodian
16th April 2014, 07:00
Do the test builds improve anything on Nvidia hardware at all?

Yes, but only with SLI on. SLI is much better with the test builds though still not as fast as without SLI.

madVR 87.9:
1280x720p24 -> 2560x1440 @ 72Hz, Bicubic75 AR chroma, Bicubic75 AR image, NNEDI3 128 Luma doubling, No smooth motion, no debanding, Ordered Dither, 3DLUT calibration, Windowed Overlay.

SLI on 41.2ms
GPU0 81% @ 1097 MHz, 19% PCI-E, 7% memory controller
GPU1 18% @ 836 MHz, 16% PCI-E, 0% memory controller

SLI off 29.9ms
GPU0 67% @ 1097 MHz, 5% PCI-E, 7% memory controller
GPU1 00% @ 324 MHz, 0% PCI-E, 0% memory controller

GTX Titans, 3770K @ 4.6 GHz, Z77 chipset, each GPU is on PCI-E 3.0 x8, 32GB DDR3-2133CL9.

madVR interopTest1 & interopTest2 (the two are identical as far as I can tell):

SLI on
GPU0 73% @ 1097 MHz, 11% PCI-E, 7% memory controller
GPU1 09% @ 836 MHz, 7% PCI-E, 0% memory controller

SLI off
GPU0 67% @ 1097 MHz, 5% PCI-E, 7% memory controller
GPU1 00% @ 324 MHz, 0% PCI-E, 0% memory controller

I can also run my "720p24" profile with SLI on which used to drop a lot frames. Jinc3 chroma, Jinc3 Image, NNEDI3 128 luma doubling, ED2, no debanding, no smooth motion.

I did recheck 87.9 immediately after these tests and it does perform as it did before so this isn't an accidental setting or system change. :)

Nvidia Driver 337.50

I believe this was already mentioned but 48/96/192 neurons NNEDI would be pretty cool for when you can't quite run for 64/128x and still got potential cycles unused. I still kinda find NNEDI too sharp for chroma but sometimes I still don't have enough horse power to go for the next level and I can't run 256 neurons luma for SD@1080p, I would welcome the opportunity to try 192 if technically doable :)

It's even more true now that AMD boards have magically earned extra headroom. :thanks:

Sadly I don't think finer grained neuron settings are possible. From the NNEDI3 docs:

nns -

Sets the number of neurons in the predictor neural network. Possible settings are
0, 1, 2, 3, and 4. 0 is fastest. 4 is slowest, but should give the best quality. This
is a quality vs speed option; however, differences are usually small. The difference
in speed will become larger as 'qual' is increased.

0 - 16
1 - 32
2 - 64
3 - 128
4 - 256

Default: 1 (int)

Another impressive update madshi, and I don't even have an AMD GPU. Thanks again!

James Freeman
16th April 2014, 08:34
Another impressive update madshi and I don't even have an AMD GPU. Thanks again!

I'm pretty sure you're wrong, unless madshi is a real living magician.... :)

Asmodian
16th April 2014, 08:37
Huh? did you read my post?

James Freeman
16th April 2014, 08:45
Ohhhhhh.... I see.
There should be a comma there.

Like so:
Another impressive update madshi, and I don't even have an AMD GPU. Thanks again!

Not like so (what I thought):
Another impressive update, madshi and I don't even have an AMD GPU. Thanks again!

:D

Asmodian
16th April 2014, 08:51
OH! haha yes, I never saw that reading. :)

Procrastinating
16th April 2014, 11:54
Alright, after testing the new test builds, I can confirm that, on my HD6770 I go from

Old: Many drops, render times ~46ms on 720p->1080p source, using luma 32 doubling
New interop 1: ~0.2 drops per second
New interop 2: ~ 1 drop per second.

The problem now, is that I'm no longer seeing render times in the debug window (with the new builds).

I can conclude however, that I am at least getting the fastest render times from test build 1, and the improvements are at least enough to prevent noticeable framedrops on a particular source now.

romulous
16th April 2014, 11:57
The problem now, is that I'm no longer seeing render times in the debug window (with the new builds)!

Quoting from that same post in which madshi posted the download link (two lines under the link itself in fact):

I've intentionally removed the rendering times from the OSD (only for these test builds, of course) because due to the way these 2 test builds work, judging them by looking at the rendering times would be misleading. So please judge these builds by testing which build allows you to use higher/more quality settings.

Procrastinating
16th April 2014, 12:06
My bad, but the conclusion stands at least.

That said, I wonder where the difference in results for the two builds lie, with some people reporting better results in either. It doesn't appear to be related to overall architecture, so possibly clocks?