Log in

View Full Version : ANN: SubRip 1.50 b4


Pages : 1 2 3 4 5 6 7 8 [9] 10 11 12 13 14

jesus2099
20th April 2006, 09:31
Hello ai4spam!

In fact I had the error before going automatic! How come I had this error and not you? What must I do to prevent it?

For the 1 second, I understand.. But in the case of a VHS (where image is almost always distorted) one must mostly count on guesses rather than to correct each and every characters. Because the prompt almost display for every character.
For this project I told myself it would be faster to post correct than to correct inline (which is even longer than typing all the subs). A waiting of 0 second would have been very nice in this situation. The rip would have been very very fast instead of (number of characters x 1s) time.
Thanks!

ai4spam
21st April 2006, 00:23
Again, the memory problem is a Delphi thing. I'm not sure if it'd go away when compiled with Delphi 2006, I haven't tried it yet. I need to experiment some. Sometimes it works by saving everything and restarting.

jesus2099
21st April 2006, 09:44
So you think it's not my file being corrupted?
Next time it happens I try to save and restart ; Thank you!
So Don't you think 0 second can sometimes be useful now? (^_^)

ai4spam
23rd April 2006, 07:42
I'll think about making 0 a valid choice ;).

Atomzk
30th April 2006, 18:58
I've been testing SubRip a bit (version 1.50b3). Overall it's doing a great job. Takes a bit of time to "train" it, but it does seem to save a lot of typing.

However, the main problem I'm having is that I just can't get the settings right. Either the "n" and "m" start to break (too thin), or the characters are surrounded by white spots. Using "outline" "fatten" and the tolerance sliders can't fix this. I've settled for having white spots and noise, however it does seem to ruin the OCR after some time because I keep adding spots to the matrix that Subrip has to ignore.

I get the impression the reason for this is that SubRip is insensitive to differences in colour saturation and just looks at the brightness. Subs are often white with a black border, so without any colour. This opposed to the background where bright areas often have some colour in them. At the moment Subrip may see a bright green spot (as in the attachment) as a character. When I would be able to tell subrip to have a low tolerance towards areas with colour in them, I'm sure seperation of the fonts from the background could be further improved.

Just a suggestion of course.. :)

molitar
30th April 2006, 19:09
Atomzk, that is why I think if a brightness and contrast filter at least was built in or the capability to use fdshow so I could do that it would work more efficiently.

ai4spam
30th April 2006, 22:36
@Atomzk: Nice to see people using the hardsubed avi feature ;). You can set the tolerance per channel, so for example in order to avoid something that is bright green being detected, you have to lower the green tolerance.
@molitar: Good idea. No time to implement :(. You can accomplish the same thing by using a custom AviSynth file with all the filters you want.

Atomzk
1st May 2006, 00:33
@Atomzk: Nice to see people using the hardsubed avi feature ;). You can set the tolerance per channel, so for example in order to avoid something that is bright green being detected, you have to lower the green tolerance.
@molitar: Good idea. No time to implement :(. You can accomplish the same thing by using a custom AviSynth file with all the filters you want.
Hi. Well I can adjust the tolerance for green of course in this case, but then I'd have to adjust the colour settings every few frames. That just takes too much time. I'll look into the AviSynth thing.

Isochroma
2nd May 2006, 18:53
I very often get the "overlapping subs" error, and indeed the subs are overlapping. However, I can extract subs correctly using subresync. So why does SubRip make mistakes on the timecodes?

Zerryk
14th May 2006, 01:02
It happens when converting from srt to sub and vice versa.
<i>first italic line</i>
<i>second italic line</i>
is correctly converted to:
{y:i}first italic line|{y:i}second italic lineBut:
<i>two lines spanned
in single italic tag</i>
results in
"two lines spanned|in single italic tag"
which is incorrect - the italic tags are lost and replaced by quotes. The same happens if the <i> tag is not closed at the end of line (as produced by Subtitle Workshop).
Similar problem occurs when doing a sub -> srt conversion. The {Y:i} tag (capital Y) is taken only for the first line.
I think some parse-time convertor which replicates spanning and unclosed tags for each following line would be the best solution...

Suchy
16th May 2006, 12:45
What about joining matrices (*.sum files)?

I have 2 *. sum files with matrices, and I can't join it into one file.

ai4spam
17th May 2006, 02:32
@Zerrik: I'll add it to my (already long) list.
@Suchy: Not possible at the moment. Someone with some Pascal knowledge could try to write something, by looking at the source code. Meanwhile, put them in the same directory, always open the one with more characters, and whenever there's an unknown one, use the search matrix button to make it check the other file as well. If a match is found, it'll be added to the currently open matrix.

Suchy
17th May 2006, 14:32
@ai4spam: thx. I try his way. I'm going to look at source and maybe try do this.

awx
18th May 2006, 10:13
@ai4spam: thx. I try his way. I'm going to look at source and maybe try do this.
Suchy, if you start on a project to join .sum files, another feaure which I would find useful is to compare two .sum files and report how many entries are duplicates.

Suchy
18th May 2006, 22:09
Hmm.. It allows add only non-duplicate entries to second sum file.

Ok. I rememmber that.

P.S.
But I don't know if and when I start this project. It depends of free time (I have a lot other projects).

iElectric
1st June 2006, 11:58
Great app. I would love feature that app would auto load chars at startup

ukendt
4th June 2006, 14:16
Something wrong with the last release, no matter which server I choose the file can not be read ?

ai4spam
4th June 2006, 17:57
It worked for me (easynews server). SourceForge tends to take a while to mirror things. Too bad the release includes source and executable - not everyone wants both.

Great to see people picking up the slack and carrying this forward ;). It's been a while for me, due to things happening in my personal life (I got engaged). I'll try to squeeze in some time to work on this. Unfortunately, this means I'll have to merge the latest changes (v 1.50b3a) with my current version, which has A LOT of changes going towards a v 2.0. I'll keep everyone posted.

Ummm... in case it wasn't obvious, the other news is we got subrip.sourceforge.net. Does anyone want to volunteer some nice website design? I think such a move deserves new "clothes" ;).

ukendt
4th June 2006, 21:58
Could not read file.

Go back. /home/ftp/pub/sourceforge//s/so/sourceforge/subrip/subrip_1.50_beta3a.7z
Jun 04, 2006 13:58

ai4spam
5th June 2006, 04:38
Try the usual way first: go to http://prdownloads.sourceforge.net/subrip/subrip_1.50_beta3a.7z?download and select a mirror.
Only if that fails, try:
http://easynews.dl.sourceforge.net/sourceforge/subrip/subrip_1.50_beta3a.7z

LRN
5th June 2006, 08:08
Unfortunately, this means I'll have to merge the latest changes (v 1.50b3a) with my current version, which has A LOT of changes going towards a v 2.0.
Just send me 2.0 (or 1.60 beta1, anything) code when it's ready.

P.S. I wish i had subtitles (in *.srt) with color tags, so i could implement color support. Do you have some?

ai4spam
5th June 2006, 14:37
Thanks for the offer, will do.
For colors, just enable colors in 1.50b3 (.srt section), save, then replace FFFFFF with radom stuff (hexadecimal) in a few places, and reopen the file, then convert to your new format.

Esc
10th June 2006, 04:10
Excuse me. Which files do I take from the archive if I don't need the source, just the utility?

LRN
10th June 2006, 07:07
http://lrn.no-ip.info/other/subrip_v1.50_beta3a_binaries.7z
Hope i didn't missed anything :)

Esc
10th June 2006, 15:35
Thanks, LRN! That's what I needed.

chipzoller
12th June 2006, 00:49
If this isn't appropriate for this thread either move it or start it as a new one.

I have the complete DVD set of Jeremy Brett as Sherlock Holmes and the episodes of the show produced in the 1980s-1990s that I wish to OCR the subs with SubRip and submit them to subtitle databases. I'm having lots of problems with these with the current and even prior version of subrip. It seems no matter what settings I try it turns out crappy. I wanted to hear what the experts had to say about this. I think the subs when created were done rather poorly, so there may be nothing to do.

I've posted them both (2 movies from the series) here:

http://arches.uga.edu/~czoller/vampyre.rar

and

http://arches.uga.edu/~czoller/blackmailer.rar

thanks for any suggestion.

ai4spam
12th June 2006, 02:03
Try selecting 2 custom colors, it worked ok for me. And, when there is a letter recognized as 2 parts, set the bigger part to the actual letter, and just press Enter when asked to type in the smaller part.
It would work better with 3 colors, but right now SubRip doesn't let you do more than 2.

PS: Make sure the subs are not protected by copyright before posting them anywhere.

chipzoller
12th June 2006, 03:29
Indeed, that works well. Many thanks for improving upon a great program!

chipzoller
13th June 2006, 17:14
What about your suggestion for this set?
http://arches.uga.edu/~czoller/lagoon.rar

When I OCR them, it re-recognizes the same letter MANY times over, even limiting the colors to 2. I'm not sure what else can be done here.

thanks again.

LRN
13th June 2006, 20:06
Try to decrease OCR's sensetivity (from 1000 or 980 to 950 or even lower)

chipzoller
13th June 2006, 23:54
I've also tried that, which only gives mediocre results. Are there no other settings I can use to achieve optimal results, or am I doomed to entering all sub pictures manually?

ai4spam
14th June 2006, 01:19
Again, not possible with only 2 colors allowed. It'll be on my list of things to change.
Using the best guess should save you some typing, and accuracy improves over time.

InuyashaSama
15th June 2006, 23:45
Since the zuggy.wz.cz site is deadly slow, I want to host SubRip 1.50 and the subtitles ripping guide at my website to make them accessible to all users.

Please send latest version binary to webmaster[at]divxland.org, because I can't get it from the server since the download drops to 0 kbps always. My server has plenty of bandwidth and storage.

ai4spam
16th June 2006, 04:21
Thanks for the offer, let's see what Zuggy says (he's the owner, after all). We were going to move to SourceForge, that's why I asked for any volunteers for a redesign. Would you be willing to host a mirror?
I'll be out of town next week, but will send you the archived contents when I get back next Saturday.

InuyashaSama
16th June 2006, 17:14
I can host it but I'll have to apply the standard page layout of DivXLand.org, including the current site ads and menu system, because I only have permission to host DivXLand.org website in that server. I may also translate the main pages to spanish, since all the site is available in both languages.

ai4spam
19th June 2006, 05:10
I guess that would be ok. If Zuggy is ok with it, I'll sent you the files. Probably the deal would be that you may use the text and apply your own formatting, but keep the copyright notices.

chipzoller
20th June 2006, 16:48
I'm having problems converting a new set of subs here (http://arches.uga.edu/~czoller/eng_subs.rar) and while I can only use 2 colors and adjusting the sensitivity to 950, I still have to recognize almost every character about 5 or more times, and with that the accuracy is still poor. Maybe you could try your hand at them and see if you have better results? Thanks.

steelman
7th July 2006, 22:46
Hi!

I want to report bug. When I try to delete character from character matrix I get accessviolation on adress ... error.

ai4spam
8th July 2006, 02:58
Does it always happen, or does it go away after you restart?
If it happens always, send me (fine my mail in the manual) the char matrix file you're using.

chipzoller
8th July 2006, 16:26
ai4spam, can you give me any hints with the abovementioned vobsubs? Perhaps a best-setting scenario by which to recognize them?

steelman
8th July 2006, 18:05
Does it always happen, or does it go away after you restart?
If it happens always, send me (fine my mail in the manual) the char matrix file you're using.


It happens all the time and it will not go away after restart. Maybe I'm blind but i haven't seen you mail in guide on zuggy web. I'm attaching it here. Oh and I'm trying to delete character with index 00036

EDIT:
btw aren't you plannig remove HI optional option for post OCR correction. Something like building another character matrix with this textes (e.g. [MAN]:)

ai4spam
9th July 2006, 07:52
I said the manual, as in the file Credits.txt ;).
Hearing Impaired removal should not be that difficult, especially with the code in place to do whole words formatting (same principle). I'll think about it.

steelman
12th July 2006, 19:03
Have you checked it?

Teebeeke
13th July 2006, 03:56
I tried that hardcoded thing, yesterday, but after 5 hours i still hadnt reached 50 %. I have a hard time getting the right colors for the program to reckognize the subs. I converted a tape to avi, which makes it hard i guess.

ai4spam
14th July 2006, 00:55
@steelman: I can't, the attachments has not been approved yet. Mail it to me (see Credits.txt for the address).
@Teebeeke: It takes a while, what can I say. Try lowering the OCR sensitivity.

SvenBent
20th July 2006, 08:26
Sorry i didnt read thorug 22 page of post.
just spank me for being a bad boy :-)

anyway the "format whoel words" cotains some bugs

E.G

Original: Dennis<i>...</i>
"fixed": <i> Dennis... </i>
Correct would be : Dennis ...

personaly i disable the function (but it wouldnt save it :-( ) and modify the punct.dic


ooh and also a suggestions
you should be able to selecte multiple languages from a list. The list should be a long list with checkboxes and a selection for "correction by langaug" for the post ocr correction.
And then it would go throug alle the files and saving it by the original language name.

That would save the user some time.
maybe even run each languange in its own threads. Rhat way the whole project would not stop for just one letter being uknown, and the speed would scale on a SMP system.

ukendt
20th July 2006, 08:38
@steelman: I can't, the attachments has not been approved yet. Mail it to me (see Credits.txt for the address).
They were approved short after your comment:D
Sorry i didnt read thorug 22 page of post.
just spank me for being a bad boy :-)

Spank:sly: beat:sly: hit:sly: ...

(furthemore in danish...)
Slå, flå, rive:D ....God varm sommer Sven Bent:D :D

ai4spam
21st July 2006, 01:23
Hmm, does it happen when the word is the first in the sentence, or anywhere?
The principle is: it looks for how many characters are italicized and how many are not, and depending on which number is greater, it formats the whole word or not.

SvenBent
26th July 2006, 21:33
Hmm, does it happen when the word is the first in the sentence, or anywhere?
The principle is: it looks for how many characters are italicized and how many are not, and depending on which number is greater, it formats the whole word or not.

I have sofar only found the errors when thers is only one word and then the triple dot

maybe jut make on corret one letter in a word. (or ...)

also i can oresse that this may also make more mistakes ten it fixed
E.G
the original lin :either/or
detectes as: either<i>i</i>or
fixed as: Eitherior

actually i would preffer te non fixes version is its graphically better reprenst what should be there

this is just my 2 cent.

thusfar i have fixed most errors without a wrong fixation by using these lines in the punc.dic

<i>ø</i>
ø
</i>ø<i>
ø
<i>/</i>
/
</i>/<i>
/
<i>...</i>
...
</i>...<i>
...

Works wonders

iElectric
27th July 2006, 01:33
Someone i know gets this error at start:
"Bad Packet Header at LBA 2. Do you want to continue?"
Then LBA just grows up: 2,3,4..