Log in

View Full Version : Probabely Old Idea for OCR based Subtitle Rippers


kilg0r3
26th June 2003, 12:30
O.k. this might be a really old idea, and, I apologize for not taking the time to check that.

One of the worst things about OCR based subtitle ripping is th fact that you have to stand by the whole process because when an unknown element is found the ripping stops until a character the user enters a character and hits o.k.

IMHO, it would be a good thing, modify the programs in a way that they read through the whole stream at once. During this process two things would have to happen.
1. All subtitle images are stored in order to be reprocessed later on.
2. The first column of the matrix will be created. The first encountered character shape gets an index (i.e. gets a slot in the character matrix), the next shape is checked against the alredy stored shape(s). If it differs from the already stored one(s) it gets a new index/slot. If not, it is skipped.

At the end of this process we would have a matrix containig image shapes on the left and empty slots on the right. Now, the user like Omnipage can read the images and enter the values into the table in one sweap.

Another possibility would be to output the character shapes in a specific format. Preferrably all shapes in a row but contained in one large image file. This one could then be fed to programs like Textbridge or Omnipage, which would output a text string. Which could be loaded into the Subrip software filling the right column of the matrix.

Ah, now I feel better. Let's see how long ...:)

ukendt
30th June 2003, 09:20
I moved this thread into development forum.
Sincerely
Ukendt