Log in

View Full Version : ANN: SubRip 1.50 b4


Pages : 1 [2] 3 4 5 6 7 8 9 10 11 12 13 14

GrofLuigi
14th May 2005, 00:14
Originally posted by ai4spam
Well, in my case it's a combination of everything: text color and outline color tolerance, filling. Beta 13 introduced some stuff to help with the big ugly areas (most of which should disappear when you click "use outline"), namely drawing lines between the text lines and filling the open areas from there. Also, keep in mind that "large" means that the width or the height is more than 10 times the text line width, so if that is detected incorrectly (max=10), then only areas of more than 10*10=100 in width or height will be filled. The value 10 came from the fact that the widest letter (W) and the tallest letter (E) have some less than 10 lines forming them. Basically, you need to manually lower the text line witdth to the point where text is still detected correctly, but large areas are filled. I'll change it to work better when the height of the text is set in the inter-line options panel.
Well, I have some pretty good ideas what to do, but unfortunately, the first "bug" mentioned prevents me from doing any serious work/testing. Will try the new version and report back.
Originally posted by GrofLuigi
The source is very clean, BTW.
Oops, I meant my source material... :eek:
Originally posted by ai4spam
And you're welcome, it's nice to be appreciated.
And I really mean it.

GL

ai4spam
14th May 2005, 05:29
Well, I have some pretty good ideas what to do, but unfortunately, the first "bug" mentioned prevents me from doing any serious work/testing. Will try the new version and report back.

Please do. Can you specify whether it was in native mode or in MediaPlayer mode? You may also want to try on a different computer, I think it's really weird that the focus rectangle gets reset.

Oops, I meant my source material... :eek:

There I was, thinking I was complimented ;).

ai4spam
14th May 2005, 10:47
SubRip 1.20 Beta 15 is available.

Changes:
- GUI improvements
- support for saving bitmaps for later subtitle removal (sequential bitmaps and a file with rows of the format below)
<index>,<first frame>,<last frame>,<left coord>,<top coord>
Basically, this is all the information needed for subtitle removal. Anyone volunteers to write a VirtualDub plugin ;)?

GrofLuigi
15th May 2005, 04:18
Originally posted by ai4spam
Please do. Can you specify whether it was in native mode or in MediaPlayer mode? You may also want to try on a different computer, I think it's really weird that the focus rectangle gets reset.

It's gone with beta 14, you produce programs faster than I can download them. :D

Ehm... mediaplayer mode or... I just "open hard subbed video files".
As I said, through Avisynth script:
LoadPlugin ("C:\Program Files\DGMpgDec\dgdecode.dll")
MPEG2Source ("D:\VIDEO\DVD\1.d2v")
ConverttoYUY2 ()
Crop (12,76,-8,-76, align=true)
Maybe it was the crop, non-mod-something value (source = PAL DVD). I just took out the black borders and the subtitle nearly touched the bottom.

GL

Kurtnoise
15th May 2005, 10:05
Great stuff guys...I'm trying the hardcode subtitles detection and it works pretty good for the moment. :)


May I suggest to add some new subtitles format for the next releases : USF (http://usf.corecodec.org/), TTXT (http://gpac.sourceforge.net/auth_text.php), and why not the QuickTime TeXML (http://developer.apple.com/documentation/QuickTime/QT6_3/index.html) ?

I can provide some samples if you want...

E-Male
15th May 2005, 12:11
Originally posted by ai4spam
SubRip 1.20 Beta 15 is available.

Changes:
- GUI improvements
- support for saving bitmaps for later subtitle removal (sequential bitmaps and a file with rows of the format below)
<index>,<first frame>,<last frame>,<left coord>,<top coord>
Basically, this is all the information needed for subtitle removal. Anyone volunteers to write a VirtualDub plugin ;)?

i'll put an avisynth plug-in on my (to long) todo-list

ai4spam
15th May 2005, 16:50
Originally posted by E-Male
i'll put an avisynth plug-in on my (to long) todo-list
No need, I'm almost done with it ;). Just stripped down DeLogo.

screw
16th May 2005, 08:49
Few remarks from my side:

1. If I interrupt the OCR with "Pause", and than modify Character matrix, "Continue" is not available (button is grey) after that, but I can restart OCR from beginning. Is that a feature or a bug (in older versions it was possible to "Continue" with OCR after matrix modification)?
P.S. Seems fixed in Beta 15.

2. This SubRip version has the same bug as 1.17 and older versions with OCR of Latvian subs - SubRip crashes, if Latvian subtitles are selected for OCR. It seems that language string "Latvian, Lettish" in SubRip EXE file is the reason (maybe it is too long). If patched to "Latvian" and length indicator set to 7, it works just fine.

Screw

zuggy
16th May 2005, 09:31
Originally posted by screw

2. This SubRip version has the same bug as 1.17 and older versions with OCR of Latvian subs - SubRip crashes, if Latvian subtitles are

My fault...
It was fixed in upcoming (never released) v1.17.2. I will fix it in 1.20 final.

ai4spam
16th May 2005, 20:34
Originally posted by ai4spam
No need, I'm almost done with it ;). Just stripped down DeLogo.
Looks like I spoke too soon: the filter works fine in preview mode, but for the life of me I can't understand why it won't work in the main VirtualDub window (RunPorc is never called?!?). Do you have an example trivial filter code that I can try out? Thanks.

Esc
17th May 2005, 05:26
I tried to run a hardsubbed file. Out of curiosity. Here is exactly what I did.
I had a sucessful run of DVD subtitles recognition. So I had a Character matrix open and some text in Subtitles window. Both were saved but not cleared.
I opened an AVI file. Pressed Play. Waited for the first subtitle to appear. Pressed 'Pause' (same button). Checked 'Use' box. Clicked with the mouse on the subtitle several times in different places. Pressed 'Run'. Waited a bit. Got bored. Pressed 'Stop' (same button). A pop-up window came out with some gibberish image and a prompt to enter a character. Whenever I tried to press 'pause/abort' or just close that window by the upper-right cross button I would get an error message. The header said: SubRip - 100%. The body said: List index out of bounds (5). And the pop-up window would stay open.

ai4spam
17th May 2005, 10:48
Well, I usually reply only to people who actually need the program, not just use it "out of curiosity" :p, but here goes: pressing "Run" won't do you any good (it'll give you garbage/gibberish) if you're not clicking inside the subtitle and changing the settings to get the right image in the first place ;). Basically, it should be white subs on a black background and not much else. For example, if the subtitles are on top of some white object, and that shows up as a big white blob, you should lower the subtitle color tolerance or check the "fill large areas" checkbox. If the subs are too thin or they look like they're "eaten by ants", you should increase the text and/or outline color tolerance. There are other tips in the (now outdated, but still fairly useful) section of the manual/readme.
I'll try to reproduce the bug you reported (adding to a non-empty sub), then fix it.

NN
17th May 2005, 14:39
1. When I use SubRip to convert .idx/.sub files to .srt files I need to supply a name for the .srt file. It would be a lot more conveniant if SubRip would default to the name of the .idx file (with .srt extension).

2. If I forget to clear the subtitle text window before processing the next .idx/.sub file I get the following message:
Subtitles text file isn't empty. Add to the end of file? (OK) (Cancel)
I would like a third option: Clear the subtitle text window before proceeding.

3. I had trouble adding the %-character to the matrix. SubRip first detected the first o and after using 'take with next' it included the /, but I cound't get it to include the second o.

4. The subtitles I converted contained a lot of italic i, l and ! characters which where all detected as /. Is it possible to modify the detection system to recognise that a / within an italic word is probably not a / but an i or an l ?

Thanks, NN

Esc
17th May 2005, 15:32
Maybe I wasn't clear enough. I didn't just 'press any key'. I clicked inside the subtitle. I did it several times to make sure I hit the right spot. Those subtitles were fairly small. And they were not plain white-on-black but some relatively bright color with a darker outline. It was an anime fansub.

ai4spam
17th May 2005, 16:11
@Esc:
Increasing the text and outline color tolerances would deffinitely help you get the whole letters. Use the other features (draw lines between, fill open, fill large) to get rid of the false guesses. Again, adjust the parameters until you see something that makes sense in the preview (white text with red outline).
Also, anime fansubs sometimes use different text and outline colors during the same movie. Nothing can be done at this level, you need to stop processing and click on the text again when the colors change.

ai4spam
17th May 2005, 21:12
@NN:
1) is feasible and easy enough
2) is also feasible, but not really justified, all you have to do is press "Cancel" and then "Clear"
3) is relatively hard, but I'll see what I can do (I had the same problem)
4) you're probably encountering ths problem because you're reusing character matrices from other DVDs, instead of making new ones :p. That's ok in most cases, but you should increase the OCR sensitivity when you get these errors.

Esc
17th May 2005, 21:28
BTW, is it me or ver 1.20 is really so much faster than 1.17?!

Also, I'm curious. Where does it get the names for the subtitles. It called Director's commentary subtitles from my last DVD "English caption for children" or something like that. It's not a real problem. Just looks funny.

ai4spam
17th May 2005, 21:50
@Esc:
It's compiled with a different version of Delphi. Also, Zuggy took out some components that were generating the P4 crash bug.
The names are read from the .ifo/.idx file. If you open it in a text editor, you'll find them there.

NN
17th May 2005, 22:27
ai4spam, thanks for the quick response to my first post !

ps.
2. I know, but it would save me 5 mouse-clicks (and me feeling stupid every time I forget to use Clear).
4. The subtitles with both the / and the italic i, l and ! came from the same movie. Maybe this is something that could be added to the post-OCR correction process (like fixing l vs. I) ?

I forgot one more suggestion:

5. Could SubRip take the delay in an .idx file into account when creating a .srt file ?

Example:

id: en, index: 0
delay: -02:06:50:700
timestamp: 02:07:04:396, filepos: 000000000
timestamp: 02:07:08:200, filepos: 000001800
...

ai4spam
17th May 2005, 22:47
@NN:
2: ok, I'll see what I can do. I've been sick for the last few days, thus unable to work much :(.
4: increase the OCR sensitivity all the way to 1000, that should solve your problem for now. Post-OCR for such cases is possible, but I don't know if I'll be the one to implement it (I got things in my personal life that I need to take care of first).

For the new suggestion: just use the time adjuster for now. We'll put it on the "sensible requests" list ;).

masken
17th May 2005, 23:50
@NN, what you need to do is work a technique for building your character matrix. The solution is simple: simply OCR the first "o" as "%", at the next OCR stop, most likely "/" and part of either "o", just hit enter. Same with the third "o". This will build a correct OCR directly as the other characters are automatically skipped the next time they appear.

There's other examples of this. With the "Extend/Reduce" feature request á la SubResync, the idea was to be able to expand such selections so the whole "%" was to be automatically selected.

ai4spam
18th May 2005, 03:59
Unfortunately, the first "o" can just as easily be a degree sign, so marking it as "%" is not a good idea. One way would be to say "extend right" for the first "o", set the "o/" to "%" and the second "o" to nothing. Then, each time a degree sign is encountered, it will ask you for a character pair again, because it will be followed either by a space or by a comma or other things, but never a "/". All you have to do is type in the appropriate "o ", "o,", "o.", and so on (here "o" is the degree sign). If the degree sign is found last in a line, the "take with next" character will be replaced with what you input, reverting to single character mode.
Some automatic way to do this might be possible (not easy), but care should be taken with character sets that are not Latin-based, like Chinese or Japanese.

ai4spam
19th May 2005, 08:48
Originally posted by NN
5. Could SubRip take the delay in an .idx file into account when creating a .srt file ?

Example:

id: en, index: 0
delay: -02:06:50:700
timestamp: 02:07:04:396, filepos: 000000000
timestamp: 02:07:08:200, filepos: 000001800
...
I've never seen such a file. Could you please mail me an example (lookup the address in the readme)? I need to know if the delay is per stream or per file. Thanks.

masken
21st May 2005, 00:01
@ai4spam, yeah, I thought someone would mention something about the degree character. Thing is though, in most if not all fonts, I've noticed from (long) experience that the "%"-o and the °-character differs enough for subrip to make them out as two different characters ;)

But yes, you're right of course, this is definetly one of the reasons I wated the "extend" feature too :)

ai4spam
21st May 2005, 07:20
Better safe than sorry I always say ;). No worries, Beta 16 is being worked on.

ai4spam
23rd May 2005, 12:39
Finally, Beta 16 is up.
Changelog:
- support for extending more than 1 character (test and provide feedback, please)
- GUI additions and fixes in video mode
- changed saved bitmaps form .bmp to .pgm and fixed the cropping
- some of the feature requests have been implemented - you'll just have to get it to see which ;)

I'll give the VirtualDub filter one more try...

ATTN: Big boo-boo on my part, I had broken DVD subtitles. It's fixed now, please download again. I made a couple of changes elsewhere anyway (improved fill sides, plus added shortcuts for extend buttons).

Esc
24th May 2005, 02:36
Couple of thoughts here.
1. It doesn't show the build number in About window. I know that I am on build 16 because I have installed it right now. But will I remember in 2 weeks if I find a bug? It would be nice if there was the build number in the version.
2. Non-standard letters turn into standard for some reason. My é becomes e. I do not remember such a problem in 1.17.
3. Subtitle Index offset changes its value every time you walk through it. If you go to output format, select SubRip and just keep hitting tab, every time you walk through that field it changes it's value to opposite.

Longinus
24th May 2005, 04:27
Hello..
Thank you for making SubRipAvi. In the past I used AVISubDetector, but I just guessed the options every time, trying to make it work. Your solution is a LOT easier. :D

But I'm having some problems.. The first is that in the latest beta, the "Video File Viewer" window background is "transparent". You can't read some of the text, and it's drawing everything from another window if you put on top of it.. (It didn't happen in the earlier bets)

Another thing is that I'm trying to OCR a timecode (of the frame number), so I have to get every frame. So I set "skip first", "update every" and "Min duration" to 1. But it doesn't work. Sometimes subrip jumps 10 frames, sometimes less, sometimes more.
Am I doing something wrong?

ai4spam
24th May 2005, 05:23
@Esc:
1) Will take care of this.
2) This happens to me with diactitics in languages like Romanian. It's not SubRip's fault, somehow you changed your system settings (that, plus Delphi makes a hidden ANSI-OEM-ANSI conversion in strings, and it can't be taken out AFAIK). Just go to regional settings and set "language of non-unicode programs" to the language of your choice. Are you sure you're using the é from the default font/charset? In my example (Romanian), some letters would change (t, and s,), but not others (i^ and a^). The way I "fixed" it was to assign the correct characters to some of the 20 buttons on the bottom. If you select the Romanian language in the OCR window, you'll notice that the letters I mentioned still don't show up correctly, even when using the EASTEUROPE charset, but they do show up correctly in the final subs. So, try to play with the charset in the general options window. The rule for characters like your é is: if you copy and paste it on a button (with right click), it should show up as you want it on the button, or else it's not the right character.
3) I fixed it, will upload in next beta.

@Longinus:
Thanks for your appreciation.
Your "transparent window" bug is really weird, there is no reason why it should happen. Maybe there's a problem with your video drivers? Please try it on another machine and let me know if it still happens. A screenshot would also be useful. Ah, and it's worth asking: are you sure you can't see the text because of the new feature (fill to the sides of the text with fuchsia color)?
The skip first, update every and min duration were set there for temporal optimization, without them it would be really slow. Beats me why you would try to OCR a timecode...
Anyway, basically I'm processing every frame, when I detect a sub I skip the first frames (to let the sub appear fully), then accumulate for min duration frames, then just compare and reset every now and then.
In your case, if the timecode is the only thing you OCR, then it doesn't change much from frame to frame. Try also lowering the same sub tolerance to a really small number. However, it will still accumulate at least 1 frame, to process every frame I'll need to make some changes (again, done, but will upload in next beta).

Longinus
24th May 2005, 08:38
I tried it in my other computer, and it worked. It probably is a driver bug, this computer is a long time overdue for a full format. But anyway, here is the link to the screenshot image.

http://www.unkind.org/blog/images/subrip_screen.png

About the timecode... Someone gave me an edited movie (15m), that used 3 big ones (20m each). But the edditor somehow LOST the Final Cut projet file. So I'm stuck with the low-qualy final movie, and I have to re-edit, putting the same parts but in high quality. Of course this will be a pain in the ass to do by hand... but I have the timecode, and If I can get it in text format, I can code a simple program to create a avisynth script that will cut the used parts for me.. It will work.. in theory.. :D

ai4spam
24th May 2005, 12:06
Beta 17 is up, with minor bugfixes.

@Longinus: try now, set "skip first" to 0 and see if it processes all frames.

Esc
24th May 2005, 16:25
Thanks for the quick response.
2) My OEM language is russian. The letter I was trying to use is not from there but from french. It's an e with a stroke above it from the default charset. But I thought that should not be a problem since you can select the charset for the subtitles. I saw a field like that in VSfilter.
When I tried to put that character on a button, it came as a question mark.
So what are you saying? I cannot have a character unless it is included in my OEM charset? That's sad.

zuggy
24th May 2005, 17:00
Originally posted by Esc
2) My OEM language is russian. The letter I was trying to use is not from there but from french. It's an e with a stroke above it from

subrip doesn't support unicode (atm) like vsfilter does. That means you can only use fonts of Windows-installed language-support. If you install the french language layout, you'll see stroked e.

Longinus
24th May 2005, 20:34
Originally posted by ai4spam
Beta 17 is up, with minor bugfixes.
@Longinus: try now, set "skip first" to 0 and see if it processes all frames.

YAyy! :D
Now it works. Thank you very much for making it work soo fast (and probably something that only I will use, ever).

Again, thank you. :)

Kaiousama
25th May 2005, 07:43
subrip doesn't support unicode (atm) like vsfilter does

In order to add Unicode support to subrip you need to accomplish 3 tasks:

1) change all strings declarations to widestring.
2) change all string-related function callings to the equivalent widestring-processing function. please refer to this link (http://mh-nexus.de/unicodenotes.htm) for useful informations.
3) change your input/visualization controls to the equivalent TntUnicodeControl (this open-source collection of unicode controls and unicode helper-functions can be downloaded from here (http://www.tntware.com/delphicontrols/unicode/) )

I hope you are interested in supporting unicode in subrip since this application is very useful.

zuggy
25th May 2005, 08:15
Originally posted by Kaiousama
In order to add Unicode support to subrip you need to accomplish 3 tasks:

Yep, I use TntUnicode already in other projects... But that means subrip will run ("only") on nt-based systems properly.

Kaiousama
25th May 2005, 09:03
But that means subrip will run ("only") on nt-based systems properly.

Unicode works properly only on unicode-supporting OS (>=winNT), running SubRip in older OS will limit its use on ANSI strings only.

Is quite phisiological this behaviour, the good thing is that you don't need separate builds for unicode and non-unicode like C/C++ (VSFilter etc..) if you use TntUnicode.

P.S. did you plan to refactor the current spaghetti-programming stle of subrip code in something more Object Oriented-style? (is quite a pain searching for something in its code).

ai4spam
25th May 2005, 09:24
Well, I've started moving to Tnt for unicode support, but it will take a while. It's unlikely that we'll start overhauling the spaghetti code, it's just too much work.

zuggy
25th May 2005, 10:04
Originally posted by ai4spam
Well, I've started moving to Tnt.

You will be lost in code while porting it into tnt!!! ...postocr correction, while-running string correction, lng files, subtitle files, ... - all have to be unicode.

And anyway is there some other tool than vsfilter that support (reading) unicode/utf subtitles?

ai4spam
25th May 2005, 11:01
Well, controls and such aren't that bad. A new char matrix format is also needed (no biggie, working on it). The problem is saving to text, since each language has its own conventions for mapping 2-byte unicode to 1-byte ansi.
I'm thinking of adding a file like "Language.tbl" for each language, with rows of the form:
<ansi index (0..255)>:<unicode 2-byte char>
SubRip would revert to just truncating the second byte if the file for the current language is not present.

zuggy
25th May 2005, 12:44
Originally posted by ai4spam
"Language.tbl"
There are functions for this conversion --> jedi vcl.

Esc
26th May 2005, 14:56
Originally posted by zuggy
Yep, I use TntUnicode already in other projects... But that means subrip will run ("only") on nt-based systems properly.
It already doesn't run on DOS. So... life is moving on. ;)

Originally posted by ai4spam
Well, controls and such aren't that bad. A new char matrix format is also needed (no biggie, working on it). The problem is saving to text, since each language has its own conventions for mapping 2-byte unicode to 1-byte ansi.
I'm thinking of adding a file like "Language.tbl" for each language, with rows of the form:
<ansi index (0..255)>:<unicode 2-byte char>
SubRip would revert to just truncating the second byte if the file for the current language is not present.
I utterly support this decision! And maybe the language for conversion could default to Windows setting but be selectable. Since I use Russian OEM setting for tons of older programs, but my DVD subtitles are always in English (I am living in Texas now) with occasional appearance of non-English characters in foreign words like fiancée or naïve (hopefully you'll get these symbols correct). It's no biggie but I always wanted to keep them, being a total showoff as I am. :)

ai4spam
26th May 2005, 19:10
@Esc: Not to rain on your parade, but you must be doing something wrong: I'm able to use "special" characters and even diacritics for some languages with no problems. Try switching the language in the OCR window and entering special characters by pressing one of the 20 buttons below the text box. This way, the current version of SubRip works without conversions in my case. Make sure your charset is the default one, then your saved files should look ok, in NotePad (with special characters).

Anyway, SubRip 1.30 with UniCode support is some 90% done, I'm ironing out the final details on ANSI-UniCode conversion for legacy ANSI support. This one will be in beta stage for a while, since by making so many changes I might have broken something.

ai4spam
27th May 2005, 07:09
Does anyone have an example of UniCode subtitles that VSfilter reads? For some reason, I can't get VSfilter to work to test my own homebrew format (2-byte UniCode chars with a $FF$FE in front, as Word/NotePad saves them).
Please send them over to my email (same id @gmail), and reply here so I know I need to read it.
Thanks.

zuggy
27th May 2005, 07:50
Originally posted by ai4spam
Does anyone have an example

Here is one, produced by vobsub: http://zuggy.wz.cz/UnicodeSRT.zip

ai4spam
27th May 2005, 08:14
Thanks, it looks like I was right. Now SubRip can read these kind of files, I still need to implement saving. There was so much other stuff that needed to be changed (like corrections, etc) that it looks this one will really be in beta for a while, until all the bugs I may have introduced are found and fixed :(.

ai4spam
27th May 2005, 12:24
Finally, 1.30 Beta 1 is up on the website. It introduces experimental UniCode support. Basically, you need to change the CharSet for "irregular" characters to show up. The CodePage is used for conversions when loading/saving ANSI files. And for other things like displaying the right captions on buttons. Files are saved by default as ANSI. When a file has "irregular" characters, you get a choice to save it as UniCode, or try converting it to ANSI using the current CodePage. The CodePage changes automatically when you select a charset, but you may also change it independently (sometimes the default ANSI codepage 1252 works best). Bear in mind hat only VSfilter UniCode supports UniCode subtitles, so save as ANSI if you have a choice, then check to see if you have the right conversion.
I've only tested .srt input and output, so other formats may be broken. Please test and notify us here.
Also, I think it's time the translations got an update. If anyone wants to update some language file, notify others here, then send the file to zuggy.

PS (for the 20 people who downloaded so far): My bad, I didn't test the clipboard properly. Please download again.

bourtzovlakas
27th May 2005, 14:30
Thanks for the new version....
Is v.1.20 beta 17 being declared 1.20 Final?

Kurtnoise
27th May 2005, 15:18
Is it just me or the server of http://zuggy.wz.cz/ is very slow ? can't download the lastest beta...:(

Why not move SubRip at SourceForge or Berlios or anything else ?

ai4spam
27th May 2005, 19:49
@burtzovlakas: I fixed a couple of other bugs in the meantime. I guess I could go and fix them in 1.20 too, then declare it final. The idea is to have people heavily test 1.30 for bugs I may have introduced, then move to it if it behaves nicely. I'm particularly interested if it still works under Win98 and WinME, and whether the Charset/CodePage combinations I set in by default work for all languages.

@Kurtnoise13: It may very well be, with people downloading it like crazy. SourceForge requires some maintenance effort that neither I nor zuggy have the time for at the moment. We may consider it in the future.