View Full Version : HD subtitle ripping (SUPread)
Pages :
1
2
[
3]
4
5
6
7
8
9
10
11
12
13
14
15
16
manusse
14th March 2007, 00:01
I've started to work on it. However you'll have to wait a bit because I don't have much free time for it. Be patient...
Cheers
Manusse
MichalHabart
14th March 2007, 09:13
There you are here MichalHabart ;)
I was watching into this thread long time but so far there was no working procedure for converting sup into srt :)
Pelican9
14th March 2007, 09:58
I was watching into this thread long time but so far there was no working procedure for converting sup into srt :)
Fortunately, it's not true! :)
I'm working on an OCR routine.
Sagittaire
14th March 2007, 10:42
Fortunately, it's not true! :)
I'm working on an OCR routine.
Well it's perhaps useless. Subripp can make that. Make compatible files with subripp and use subripp for OCR is perhaps a better/simple way ... ???
You make fantastic work with EVODemux and SUPread ... :thanks:
Deckard2019
14th March 2007, 11:29
Well it's perhaps useless. Subripp can make that. Make compatible files with subripp and use subripp for OCR is perhaps a better/simple way ... ???
Read previous pages. We already tried but Subrip works weirdly with image sequence.
Even with a sequence coming from a DVD :eek:
MichalHabart
14th March 2007, 11:56
Fortunately, it's not true! :)
I'm working on an OCR routine.
Cool, keep going :)
derebo
14th March 2007, 14:14
WOW! all this great news about SC and HD-DVD... thanks for your effort manusse and pelican!
greetings,
Pelican9
15th March 2007, 16:53
I'm almost there... :)
Kiriakos
15th March 2007, 18:46
I'm almost there... :)
May the force be with you, master. :thanks:
PS: Cant wait.
Pelican9
15th March 2007, 19:07
May the force be with you, master. :thanks:
PS: Cant wait.
Why don't you download the newest version... :-)))
Kiriakos
15th March 2007, 21:46
Why don't you download the newest version... :-)))
WOW!
:eek: Thanks m8.
Rectal Prolapse
15th March 2007, 22:23
Wow thanks Pelican9!
Is there a way to get the OCR to "learn" characters it can't identify? It would be nice if you can correct its mistakes and if it could remember what you typed in for characters it had problems with.
But - I guess that would be a lot of work. This is a great addition!
Hmmmm, any tips on using Scenarist ACA for creating the subs? Does Scenarist create a subpicture stream usable by today's currently available subtitle display filters, like vsfilter?
Rectal Prolapse
15th March 2007, 22:29
Slight bug report:
It seems the italic form of "rt" is not recognized. When the r and the t characters run together. This might be a problem for other character pairs that are very close apart ("tt", etc.).
Also, sometimes the letter "f", at the beginning of the word, is misidentified as italic when it isn't - and the whole word gets flagged as being italic. Manually running the OCR on it again with "Normal" checked fixes it. :)
The letter "I" (eye) gets confused for lower case "L" all the time when it is the beginning of the word, and is followed by an apostrophe, such as "I'd like to..." being converted to "l'd like to...", with an L.
Pelican9
16th March 2007, 00:52
Changes for v0.3
- Convert bitmaps to text (OCR)
- New options for OCR (skew, cut level)
- 'OCR all' function
- Jump to the next and the previous error
- Enable tags (<i></i>)
Download: SUPread v0.3 (http://pel.hu/down/SUPread.exe)
This is still a beta version, but it can read the characters of the English alphabet and some special chars ([]()$.,'"-)
The 'rt', 'tt', 'ff' problem at the italic font is feasible sometimes with adjusting the level value.
The 'I' and 'l' have the same pixel representation on the bitmap, they recognizable only use of the text context but it differs in the different languages. I'm thinking about it.
The leading f,t,O problem not yet solved. You can use the manual override of the font style.
Rectal Prolapse
16th March 2007, 02:49
I seem to have issues with the small letter n being confused with o, the capital letter E being confused with B, and z being confused with x. Hmmmm.
Is there a server I can upload sample .sup files to you for testing? I guess it may not be necessary if you know about those other bugs. :)
Still, it works well otherwise - it took me about 30 minutes to do the Harry Potter HD-DVD movie!
Rectal Prolapse
16th March 2007, 07:06
I wonder if it would be possible to convert the Scenarist SST file generated by SupRead into .sub/.idx files? Then it would be easy (or easier) to drop it into an MKV for use by DirectVobSub for playback, without the need for OCR.
It would be nice to convert the bitmaps into a usable subtitle picture format - it would save a lot of time. :)
Pelican9
16th March 2007, 08:54
I seem to have issues with the small letter n being confused with o, the capital letter E being confused with B, and z being confused with x. Hmmmm.
Do you use the latest version?
Please share your subtitle file.
JnZ
16th March 2007, 09:29
Hi Pelican, You make very big progress on SUPread from my last checking. OCR seems very promising.
I have one suggestion - when font is italic, put <i> </i> tag into STR file. EDIT: Ups I find this option in settings... :)
:thanks: for your good stuff.
EDIT: Well there is a problem with italic "g","rt","7", which cannot be recognized. Italic "i" reads as "ˇ", italic "2" reads as "g", italic "0" as "8"...but this is minor bugs.
Deckard2019
16th March 2007, 10:08
Sometimes, Font mode switches to "Auto" but OCR result is better in "Normal" :
"tin" becomes "?o" in Auto but is correct in Normal (@Pelican9 : see line 53).
Italic characters are falsely detected as they are normal in the bitmap.
Also, accented characters are not detected. But number 6 is ok now ;)
BTW, an "OCR all" gives a great result. Not so far from 100%.
Do you think it would be a wrong idea to add spelling correction feature ?
DeepBeepMeep
16th March 2007, 10:09
Thanks! Very nice! Now these sup files start to have some use.
Bugs:
- The OCR All options inserts carriage returns in the middle of lines whereas step by step OCR works properly and inserts only CR when there are several lines in a bitmap
- The OCR can't detect the letter "g"
- The OCR confuses quite often Digits (3 and 8 for instance)
- A "i" that was detected is sometime transferred incorrectly as a reversed "!"
- step by step OCR stops a line where there isn't any error
Suggestion
- The "auto" mode of the OCR makes quite often errors assuming a style has changed in the middle of a sentence. In most cases no errors would have been reported if it had assumed the sentence was either ONLY italic or ONLY normal style.
Indeed, it is very rare that a same line contains mixed styles. An option to assume that a style can't change in the middle of words or a line would improved greatly the efficiency of the OCR
Pelican9
16th March 2007, 12:13
I thank everybody for testing.
Some info to understand how the program works.
First step: isolating the characters: the algorythm is very simple, it's trying to isolate the characters with two methods (normal and italic) and the winner is the one which returns the thinner (less width) result. This algorythm fails when non-italic text contains a word starting with f (or some other conditions)
In the image window you can find a Radio group
Normal means all the characters of this subtitle are normal style.
Auto means that the text may contain both styles, the sw recognizes automatically
Italic means all the characters of this subtitle are italic
If the sw finds that all chars of the text are the same style, then it sets the right value, if it finds both styles then it sets the Auto
The wrongly recognized style creates wrongly recognized characters.
DeepBeepMeep
16th March 2007, 12:50
That's interesting... So it seems, it should be then easy to implement quite quickly the suggestion I made in a earlier post:
- for each line of a bitmap, look at the number of chars return of Normal and Italic chars returned by the auto mode and see which types of chars are more numerous, then reprocess automatically the line assuming it is only made of italic or normal chars. You could go even further by applying this processing on each word individually instead of lines
- even simpler, apply each of your three modes (normal, italic, and auto) and keeps the one that has fewer undetected chars than the others
A combination of the these two ideas can give even better results
Pelican9
16th March 2007, 15:52
That is not so simple.
Many words like 'for' gives the same number of chars with the two styles...
Anyway, for me it doesn't ever fail if the style is set properly.
Please share the subtitle with me if there is a mistaken char.
Deckard2019
16th March 2007, 15:58
Please share the subtitle with me if there is a mistaken char.
Did you try mine yesterday ?
Pelican9
16th March 2007, 16:53
Did you try mine yesterday ?
Yes, of course. It works, except the é, á and the other special french chars.
I would like to solve this style problem first.
Rectal Prolapse
16th March 2007, 16:57
Pelican, I just used the one you posted that is slightly different than the other one. It appears that the very first version (before you posted the link) worked a little better - but maybe that was because I turned on the italic tab and it MISSED more obvious errors. I don't know.
Anyways, all the bugs reported by others is true for me.
There doesn't appear to be a way for me to up files easily. I will give Rapidshare a try, even if it sucks. :)
Pelican9
16th March 2007, 17:29
I've tried the OCR with the following movies:
The Chronicles of Riddick
The Last Samurai
Deer Hunter
Serenity
(the English track)
The Riddick works well without any failure except the 'music note' char
The others have some unrecognized char, mostly 'rt', 'tt' where not enough space between the two chars.
RP: Try sendspace.
idamien
16th March 2007, 18:31
Many, many, many thanks for the great tool, Pelican9!
Question: Are you using any kind of dictionary to correct mispellings? e.g. "L" OCR-ed to "I", "aIIey" instead of "alley", etc. If you are not, do you think it is a good idea? Maybe then you could have different dictionary files (.txt files) for different languages which the users themselves could add to. What do you think?
Feedback: I tried doing the OCR on some english subtitles but some things happened which I think could/should be changed:
1. I copied the text from the SRT tab and pasted in a text file with .srt extention. The file would not open under Urusoft´s Subtitle Workshop. The program popped an error saying it wasn´t a valid format.
2. Some subtitle images, when OCR-ed, came out as three, four and I´m guessing even five lines at times. Limiting and distributing text to a maximum of two lines would be a good thing, I guess.
3. The italics tags <i></i> are used once for each line even though all lines are in italics. Isn´t an opening tag at the beginning of the first line coupled with an ending tag at the end of the last line enough for these cases?
Anyway, thank you so much for this app and for EVODemux, Pelican9. Your effort and time, and the great results they are yielding, are greatly appreciated. Many, many thanks again.
Pelican9
16th March 2007, 19:01
I've just made some bugfix and change the method of choosing style.
The .srt export is working now.
The newest version 0.3b3
Rectal Prolapse
16th March 2007, 21:29
I missed the Sendspace suggestion!
I will give you the rapidshare link in PM.
Rectal Prolapse
16th March 2007, 21:39
Mosu in the EVOB demux thread said that a comma should be used instead of periods when specifying the fractional seconds in the SRT file. Maybe that is why Subtitle Workshop doesn't like it?
Also, MKVMerge didn't like that either, but Mosu just linked to a new build of MKVToolNix that accepts periods in the timestamps.
Pelican9
16th March 2007, 22:00
Mosu in the EVOB demux thread said that a comma should be used instead of periods when specifying the fractional seconds in the SRT file. Maybe that is why Subtitle Workshop doesn't like it?
Also, MKVMerge didn't like that either, but Mosu just linked to a new build of MKVToolNix that accepts periods in the timestamps.
I think you missed my post at 19:01...
Not only the comma versus period but the missing leading zero was the problem.
Edit:
Some bug fix in v0.3b5
(and English dictionary)
Some bug fix in v0.3b6
The dictionary solved the 'I' <> 'l' problem
Rectal Prolapse
17th March 2007, 22:11
Thanks Pelican! When I get back to my machine tomorrow I will retest. :)
Pelican9
21st March 2007, 02:36
It's not perfect, but it works for me very well.
Changes for v0.3
- Convert bitmaps to text (OCR)
- New options for OCR (skew, cut level)
- Enable tags (<i></i>)
- Language selection
- 'OCR all' function
- Jump to the next and the previous error
yonta
21st March 2007, 05:35
Hi, Pelican.
I tested SUPread v0.3 on Mummy.
The result is awesome!
only except for the following misreadings.
halfway >> hal??y
of job >> ofjob
of junk >> ofjunk
£100 >> ? 00
netjer >> ne??r
and, I don't think I have to include the old problem of reading 'I' as 'l'.
Pelican9
21st March 2007, 09:37
Thanks.
Please share the sup file!
Did you choose the language?
(It solves the I<>l problem.)
yonta
21st March 2007, 12:52
I did it with the language set to English.
SUPread read all 'Imhotep' instances as 'lmhotep', but this is easily fixable with a find & replace in notepad.
Except that, just a few misreadings on 'I', like;
l, l, I woke him up and I intend to stop him (first two)
Pelican9
21st March 2007, 23:24
Changes for v0.31
- Fixed bug with spaces between cj|fj|tj|yj|fA
- Fixed bug with rt,rw,fw,fy
- New options: Space width values
I need a little help.
What character shall I use for the music (eighth) note?
MichalHabart
22nd March 2007, 08:14
Changes for v0.31
- Fixed bug with spaces between cj|fj|tj|yj|fA
- Fixed bug with rt,rw,fw,fy
- New options: Space width values
I need a little help.
What character shall I use for the music (eighth) note?
Just tried newest version but result is really bad:
1
00:04:16,120 --> 00:04:16,211
D··d·
2
00:04:21,020 --> 00:04:21,122
H····· ··· ··· ·l·idh··
3
00:04:25,058 --> 00:04:25,160
Y····· d····i·d.
4
00:04:29,329 --> 00:04:29,431
··· i· ·b··· M····
5
00:04:36,805 --> 00:04:36,907
l· ·h·· b······
6
00:04:41,677 --> 00:04:41,779
M· d··· b·b· ...
7
00:04:44,582 --> 00:04:44,684
l··· d···i·d ·· b· ·· ·b····i··.
8
00:04:53,757 --> 00:04:53,860
··· ·h· ·h····
9
00:04:56,961 --> 00:04:57,052
·h··
10
00:05:02,100 --> 00:05:02,191
Th· ··· ··· ··ld ·· ·b····
·h· b·······.
What did i do wrong in settings of supread? Movie is Total Recall. If you want, i can upload somewhere sup file for you.
Pelican9
22nd March 2007, 10:33
What options (OCR values) do you use?
If you use the same values, please share your sup file.
http://pel.hu/down/SUPoptions.jpg
MichalHabart
22nd March 2007, 10:53
What options (OCR values) do you use?
If you use the same values, please share your sup file.
http://pel.hu/down/SUPoptions.jpg
I used exactly the same options.
Sup file is here: http://rapidshare.com/files/22214553/L0_mainMovie_L1_mainMovie.8bitRLC.stream.01..sup.html
Pelican9
22nd March 2007, 12:56
Thanks.
The character size is the problem.
I'll fix it tomorrow.
Mtz
22nd March 2007, 13:07
What character shall I use for the music (eighth) note?
Can you use the one from the screenshot. Is at the same position for all Code Pages.
http://img262.imageshack.us/img262/7011/notecharactervw0.jpg (http://imageshack.us)
I'm using it in my settings for subrip.
enjoy,
Mtz
Pelican9
22nd March 2007, 13:48
Thank you.
ai4spam
23rd March 2007, 21:37
Read previous pages. We already tried but Subrip works weirdly with image sequence.
Even with a sequence coming from a DVD :eek:
Well, as per my previous post, SubRip supports image sequences as generated by DVDSupDecode (4 colors, 8 bits per pixel), and the only issue is the size of the characters (might be too large).
Alternatively, use OCRdll like DVDSubEdit.
To see what kind of bitmaps you need to generate, grab a valid SD .sup file, run it through DVDSupDecode, and look at the output.
Hope this helps.
Rectal Prolapse
24th March 2007, 06:01
Would it be possible to resize the bitmaps generated by supread, such that they are compatible with Subrip?
If that will work then maybe we can use a graphics tool with batch processing to rescale the subs to DVD resolution, then use subrip...
ai4spam
25th March 2007, 05:39
It should work. Just use nearest neighbor resizing instead of bilinear or anything fancy, to keep the color count to 4.
Rectal Prolapse
25th March 2007, 07:57
Ok thanks!
dchard
25th March 2007, 17:55
Is there any way to convert an srt file to proper sup, that we can instert(mux) to an HD-DVD?
Thanks!
Dchard
manusse
25th March 2007, 18:10
Maybe with some professional software. But it's on our plans for SubtitleCreator. However we didn't start working on it so it could take a few months before it exists.
I have made good progress in the HD-SUP import. I mean importing a HD-SUP into SubtitleCreator to be able to export it to SD-Sup or VobSub or SRT for further processing. I still have some bugs. However I hope we will be able to release a beta within one month.
Cheers
Manusse
vBulletin® v3.8.11, Copyright ©2000-2026, vBulletin Solutions Inc.